8,097 Matching Annotations
  1. Last 7 days
    1. Dr. Valaida Wise makes the important distinction that as a theoretical framework, CRT helps us think about and critically analyze systems; as such, it can help teachers think about the right questions to ask regarding potential or actual inequities that may be present in our classrooms. Our concern at the classroom level, therefore, is not the theoretical work of CRT, but that we have a culturally responsive pedagogy (CRP) in place.

      If CRP is just about making lesson relatable to diverse students why do you think people get so defensive when schools announce they are doing CRP?

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1

      Evidence, reproducibility and clarity

      The last several years have seen major advances in our understanding of the basic cell biology of the set of single cell animal relatives, led by the authors and their colleagues. These groups have developed several as models and pioneered remarkable microscopy approaches to examine their cell biology. Here the authors extend this work to the ichthyosporean Sphaeroforma arctica, exploring the fascinating process by which a syncytial life stage cellularizes, and seeking to define the role of microtubules and Golgi trafficking. Their new expansion microscopy delivers impressive images of membrane and the cytoskeleton during this process and has the promise of answering important questions about the roles of microtubules and membrane trafficking. I thus went into my review quite excited. However, as I detail below, while some aspects of the process are carefully quantified, the current manuscript draws multiple broad conclusions that do not seem fully supported by the limited data provided. This substantially reduced my enthusiasm. I'd also note, though I did not take this into account in my evaluation, that many of the images are presented at a size that was difficult to interpret, without my electronically enlarging them, and two of the Figures were mis-labeled.

      The general points raised here are dealt with below. All panels in the main figures are enlarged, zoom-ins are added where they help, and figures that are no longer fitted have been moved to the supplementary figures. Panel labelling has been checked throughout, the two mislabelled figures are corrected, and the labelling is now consistent across the whole figure set. The additional quantification and the associated framework is set out under comment 2.

      1. Interpreting most of their Figures requires understanding the basics of cellularization in this organism. Comparing the diagram in Fig. 1B and the images in 1C left me confused. First, are all the images in 1C and similar images later cross sections? Are nuclei dispersed throughout the cytoplasm at the start or restricted to a region near the cortex. If the former, how do more central nuclei get cellularized? The transition from the unperturbed 20 and 40 minute timepoints left me unclear on the normal process. Panel G may have been helpful in this regard but only shows the treated embryo and no untreated one.

      This comment identified a real shortcoming in how we presented the system, and it prompted the largest single change to the text of the revised manuscript. In writing the original version we took the basic description of cellularization in S. arctica largely for granted, since it has been built up across several previous studies (Dudin et al. 2019, Ondracka et al. 2018, Olivetta et al. 2023 and Shah et al. 2024), and we did not restate it for a reader arriving at the organism for the first time. We have addressed this on three fronts: the imaging planes, the referee's question about nuclei, and the introduction.

      On the images themselves (Figure 1), these panels and the equivalent images throughout the manuscript are equatorial mid-sections. The imaging plane is now explicitly stated in every figure legend, so that single slices, equatorial mid-sections, and maximum-intensity projections can no longer be confused with one another. More usefully, we have added Figure EV1A, which shows an untreated cellularizing coenocyte as both a top view and an equatorial mid-section, at the onset of invagination and 10 minutes later, each at low and high contrast. Figure EV1A makes the two geometries directly comparable, rather than leaving the reader to reconcile the schematic in Figure 1C with a single plane, and it also provides the untreated counterpart to Figure 1G that the referee requested at the end of this comment. All parameters of cellularization measure in the study are now indicated on the sketch in Figure 1C.

      Figure Legend EV1A: (A) Cellularizing control S. arctica coenocytes labelled with FM4-64, shown as a top view and an equatorial mid-section at the onset of invagination (0 min) and 10 min later, each at low and high contrast display. Nuclei are visible as dye-excluded regions (asterisks). Scale bar, 10 µm.

      We would add that an untreated comparison was in fact already available, since Movie EV1 runs a control and an MBC-treated coenocyte side by side through the whole of cellularization at 10-minute intervals. We accept that this was not where the referee would naturally have looked for it, and Figure EV1A now places the comparison in the figures themselves.

      On the referee's central question, we can be unambiguous. Nuclei are not dispersed throughout the cytoplasm From the coenocytic phase through to cellularization, they remain associated with the cortex, and this has been established consistently across the previous work cited above. There are therefore no central nuclei in controls, and the question of how they would be enclosed does not arise. We recognize, however, that the original manuscript never stated this outright, and that a reader had no way of knowing it, and the referee was right to ask. The revised text now says it explicitly and follows it with the geometric consequence.

      Line 70-72: Nuclei divide synchronously and rather than occupying the interior of the coenocyte, associate with the cortex, evenly spacing out along the surface.

      Line 78-82: Recent work has shown that S. arctica cellularization is regulated by the nucleocytoplasmic (N/C) ratio, with cellularization timing tightly coupled to nuclear content relative to cytoplasmic volume. Invaginations initiate at the cell periphery, between adjacent cortical nuclei, and advance inward, carving compartments from the cortex around each nucleus rather than assembling them from the cell interior.

      Finally, the introduction has been substantially expanded so that the process is defined prior to the figures. It now sets out the coenocytic phase, the arrangement of nuclei at the cortex, the direction in which invaginations progress, and the stages through to the end of cellularization, drawing explicitly on the previous characterisations that the original manuscript had assumed the reader would already know. This also addresses the referee's difficulty with the transition between the unperturbed timepoints, which we had described without first establishing what the normal sequence is.

      The authors make a number of conclusions in Figure 1. Some are carefully quantified but others are not. For example, they state "producing uncoordinated ingression with variable rates, diagonal trajectories, and occasional bifurcations, in contrast to the uniform, perpendicular furrows observed in controls". I was not convinced by the single images provided that these were different-for example, spacing in the unperturbed 10 minute time point seems variable and furrows are not "uniformly perpendicular" in the unperturbed 20 minute time point. I also was puzzled by the lack of change in nuclear spacing while furrow spacing was altered. Later on in this section they state "MBC-treated coenocytes displayed significantly decreased and irregular furrow spacing, resulting in some compartments lacking nuclei entirely (Figure 1G & H)", but none of these images visualize nuclei directly-Perhaps the lower level background in H is supposed to indicate this but no parallel wildtype image is shown. Finally, while their TEM images (in the panels that are either I or J due to mislabeling) are lovely, I was not sure what to conclude from them- only the unperturbed images show the embryo surface for orientation and the lefthand unperturbed image is not as straight as they suggest-I certainly don't think these few images support their strong, detailed conclusions here: "MBC treatment resulted in aberrant membrane architecture, with furrows exhibiting bifurcated and convoluted morphology and abnormal fusion at the base of invagination, compared to the smooth, organized structure of control furrows".

      We have considerably expanded both the range and the amount of quantification in the revised manuscript, and every quantity we measure is now sketched out in Figure 1C. These measurements are organised around the four properties of cellularization that the revised manuscript is built on:

      1. Nuclear organisation, measured as inter-nuclear distance and its variability within a coenocyte (Figure EV1G and Figure 1F).
      2. Furrow positioning, measured as furrow spacing, its variability within a coenocyte, and the deviation of the invagination axis from perpendicular to the cortex (Figure 1H, Figure EV1H and Figure EV1I).
      3. Furrow ingression dynamics, measured as ingression rate, its variability, the maximum length furrows reach, its variability, and the duration of cellularization from the onset of invagination to Flip (Figure 1E and Figure EV1B-F).
      4. Developmental outcome, measured as the variability of released cell size (Figure EV1J). The same set of measurements is now applied to every perturbation in the manuscript, so that microtubule depolymerization, centrifugal displacement of nuclei and disruption of membrane trafficking are each characterised against the same quantities rather than described in their own terms. Applied across the whole dataset, these measurements answer the referee's objections directly.

      The descriptors they quote were asserted from single images rather than measured. We have deleted "diagonal trajectories", "occasional bifurcations" and "uniform, perpendicular" from the manuscript, and what remains is carried by measurement.

      Line 113-115: MT loss disrupted furrow ingression dynamics, producing uncoordinated ingression - in contrast to the consistent furrows observed in controls (Fig. 1D, Movie EV1).

      Coordination is now measured within single coenocytes, so that furrows are compared against their own neighbours rather than across the population. This is the comparison the referee implies when looking at a single image and asking whether the furrows within it differ. It is now Figure EV1D.

      Figure Legend EV1D: (C) Variability of furrow ingression rate within a coenocyte, from the rates in (B). DMSO 0.053, MBC 0.111, p = 0.014.

      Perpendicularity is measured in Figure EV1L, on ten furrows per coenocyte at matched ingression depth between conditions. This concedes the referee's specific point: control furrows are not uniformly perpendicular, and deviate by 5.51 degrees on average. We claim only that treated coenocytes deviate roughly twice as far.

      Figure Legend EV1I: (I) Deviation of the invagination axis from perpendicular to the cortex, |angle - 90|, averaged per coenocyte. DMSO 5.51 degrees (n=5), MBC 10.83 degrees (n=4), Ten furrows per coenocyte; p = 0.016. Ingression depth did not differ between conditions.

      On the referee's puzzlement about nuclear spacing, they have identified something we had underplayed. Mean inter-nuclear spacing genuinely does not change, and we do not claim that it does. What changes is its regularity, which roughly doubles. The same pattern recurs for furrow spacing, for the length furrows reach, for ingression rate and for the size of the cells finally released, and it is why the manuscript always argued that microtubules sustain the fidelity of cellularization rather than its execution. However, now it's strengthened by numerical data.

      Line 156-159: While mean inter-nuclear spacing was not significantly different between conditions (6.15 vs 6.22 μm, Fig. EV1G), MBC-treated coenocytes exhibited dramatically increased variability in nuclear positioning (0.143 vs 0.268, Fig. 1F), with nuclei ranging from properly positioned to severely mislocalized.

      On the visualisation of nuclei, we would push back in part. Nuclei are directly visible in FM4-64 as regions from which the dye is excluded, and this is how they were scored throughout; the approach is not introduced here and was used in this same organism in Olivetta and Dudin (2023). We accept that this was never explained and that the display contrast made it difficult to see. The Methods now state the scoring explicitly; Figure EV1A and Figure EV3D show coenocytes at both low and high display contrast, and nuclei are marked with asterisks.

      Figure legend EV1: (A) Cellularizing control S. arctica coenocytes labelled with FM4-64, shown as a top view and an equatorial mid-section at the onset of invagination (0 min) and 10 min later, each at low and high contrast display. Nuclei are visible as dye-excluded regions (asterisks). Scale bar, 10 µm.

      Line 272-274: Throughout these experiments nuclei are resolved in single optical sections as regions from which FM4-64 is excluded and are visible as such at both low and high display contrast (Fig. EV3D).

      The claim about compartment contents has also been moved to data where nuclei are directly labelled rather than inferred. The statement that some compartments lacked nuclei is removed from the furrow spacing sentence, and the observation is now made on U-ExM with a DNA stain, in Figure EV2C and Movie EV5, where it is made on a single section and on the full volume rather than on a projection.

      Figure EV2C: (C) Two MBC-treated coenocytes: a maximum intensity projection of 50 consecutive z-sections (left) and a single z-section of a second coenocyte (right). Tubulin signals persist around a subset of nuclei after MT depolymerization, and individual compartments enclose more than one nucleus. This observation is made on the single section and on the full volume in Movie EV5, not on the projection, since nuclei at different depths overlap in projection. Scale bar, 10 µm. Scale bars are adjusted for expansion factors.

      The mislabelled panels have been corrected throughout, and we thank the referee for catching them.

      TEM images of furrows with visible cell surface marked as yellow asterisks are now included in the main figure. Additional images of furrows (furrow 3 and 4) and zoom-ins of the convoluted membrane at the tip are now included in the supplementary Figure EV1K,L. These are representative images selected from DMSO (n = 25 furrows from 34 tomograms, with cell surface seen in 19), MBC (n = 67 furrows from 72 tomograms, with cell surface visible in 37), total 106 tomograms imaged across different cellularization stages and conditions. This information is now included in the legend of Figure 1J. TEM images of DMSO and MBC-treated full cells are provided in EV1K to provide an overview of furrow consistency in the two conditions. Nuclei are labelled with N to indicate nuclei per compartment.

      The images in Fig. 2B are remarkable and very informative, though as I note above they are presented at such a small size that they require considerable enlargement to appreciate. The surprising accumulation of actin at the invagination front, presumably long before membrane closure begins, was striking, as were the microtubule baskets. However, conclusions drawn again seemed too strong. The authors state "High magnification views further supported that these bundles closely tracked the advancing furrow fronts, with longer MT extensions associated with deeper furrows during later stages (Figure 2D, arrowheads)" (BTW once again this Figure was not labeled in parallel with the text-should be 2C). I did not think the NHS staining provided sufficient resolution of advancing furrow fronts to draw this conclusion. They end this section with some more detailed conclusions, which did not seem to me to be well supported by the single image shown: "Furrows were often misaligned, and compartments frequently enclosed multiple nuclei or, conversely, lacked nuclei entirely. In several cases, nuclei were observed trailing between furrows or located beneath partially formed compartments, suggesting that improper nuclear positioning may interfere with furrow progression and sealing (Figure 2F)." The latter conclusion also seemed to leave me wondering about cause and effect. The final sweeping conclusions in the paragraph on p. 7 top (next time please include page numbers) thus seemed much too broad.

      The descriptive claim has been replaced by a measurement made across the whole dataset, and the causal claim has been removed altogether.

      On presentation, all panels in Figure 2 are enlarged in the revised version. We have also added zoom-ins on a forming compartment in Figure EV2A, and a second late-stage example shown as single channels and merge in Figure EV2B, so that the relationship between the microtubule network and the invaginating membrane can be inspected at a useful magnification rather than inferred from a small panel.

      The sentence the referee quotes has been deleted. Their objection is well founded, since we were reading furrow fronts off the pan-labelling and then drawing a quantitative conclusion about depth from it. In its place, the relationship between microtubule length and furrow depth is now measured directly and reported as a correlation across 68 coenocytes, and we state explicitly what that correlation does not establish.

      Line 209-214: Quantitative analysis of MT networks showed that -MT length scaled with furrow depth (Spearman rho = 0.709, p = 1.3 × 10⁻¹¹, n = 68 coenocytes), consistent with MTs elongating in coordination with plasma membrane invagination (Figs. 2E,F). While this correlation suggests a role for MTs in furrow progression, it remains unclear whether this involves active polymerization at the furrow front or utilization of pre-formed MT tracks.

      On the second part quoted, the speculation that improper nuclear positioning may interfere with furrow progression and sealing has been removed. This addresses the point about cause and effect directly, since we cannot separate the two from these data and we no longer imply that we can. As set out in our response to comment 2, the claim that compartments lacked nuclei is also removed, and what remains of that observation now rests on Figure EV2C and Movie EV5 rather than on a single image, so we do not repeat it here.

      Line 218-221: Furrows were often misaligned, and individual compartments were observed to enclose more than one nucleus (Fig. EV2C, Movie EV5). Nuclei were also seen trailing between furrows or lying beneath partially formed compartments (Movie EV5).

      The panel that should have been cited as Figure 2C is corrected, as part of the labelling pass described under comment 2, and the callouts have been checked so that each panel is cited in the order of its lettering. Line numbers are included in this revised version, as the referee requests.

      Finally, in the closing paragraph of the section, we accept that it drew broader conclusions than the data in that section carried. It has been rewritten so that each claim is tied either to a specific measurement or to a named comparison with another system, and the section now ends on the comparison between our two titratable perturbations across ten measurements in Figure 4F rather than on a general statement about cytoskeletal coordination.

      I thought the use of centrifugation to move nuclei was clever. However, it also is moving many other things-for example it moves whatever organelles are labeled by BODIPY and apparently nuclear associated MTOCs, leading to some caveats and calling into question their claim that it "doesn't disrupt MTs". Once again, broad conclusions were drawn based on an n=1 image: "furrow ingression proceeded with kinetics comparable to controls, confirming that the core machinery for membrane trafficking and actin-driven invagination remained functional (Figure 3B). However, furrow spatial patterning was dramatically altered: in nuclear-depleted cortical regions, furrow initiations were more frequent and closely clustered, but failed to progress deeply. In nuclear-enriched regions, furrows progressed with kinetics comparable to controls but were misaligned, frequently enclosing multiple nuclei per compartment rather than the single nucleus observed in controls" Only one thing was quantified-furrow ingression rate-and this must have been done on the selected set of furrows that progressed, and not, for example, of ones like those at the bottom of the image series presented. None of the other conclusions about spatial patterning were quantified-for example, nuclei are not even visualized in Fig 3C. Finally, they do not even mention the results of the MT perturbation presented in this Figure, and the fact that few differences are apparent calls into question their conclusion that that "nuclei (and their associated MTOCs) serve as spatial landmarks that pattern membrane invagination".

      This comment is well taken on every count, and the centrifugation experiment has been reanalysed accordingly. Every claim in that section is now quantified. The comparisons are paired within single coenocytes (nuclei depleted vs enriched regions) so that each coenocyte serves as its own control. Moreover, the microtubule part of the experiment is analysed and reported rather than left aside. The new quantifications are the following:

      1. Nuclear displacement itself, as the percentage of nuclei in the enriched region per coenocyte before and after centrifugation (Figure EV3B).
      2. The state of the microtubule network after centrifugation, as the length of the longest microtubule of each nuclear aster (Figure EV3C).
      3. Furrow spacing in the nuclei-enriched and nuclei-depleted regions of the same coenocyte and its variability (Figure 3E and Figure EV3F).
      4. Furrow length by region and its ratio (Figure EV3G).
      5. Ingression rate by region and across spin conditions, and the invagination angle by region (Figure 3C, Figure 3D and Figure EV3E). On MTs , we agree that the original claim was not supportable and it has been removed. We now measure the network rather than assert that it is intact. Centrifugation does not leave MTs undisturbed. It relocates them together with the nuclei, and the same sentence reports what is lost from the depleted cortex.

      Line 269-276: Microtubules were relocated with the nuclei as every nucleus retained the associated MTOC and MT network, and the length of the longest microtubule of each network was unchanged by centrifugation (5.97 against 5.08 µm, Fig. EV3C). Throughout these experiments nuclei are resolved in single optical sections as regions from which FM4-64 is excluded and are visible as such at both low and high display contrast (Fig. EV3D). In the nuclei-depleted cortex of the same coenocytes, tubulin was present only as some short fragments, a median of three per coenocyte with a median length of 0.95 µm (Data EV1).

      On the broader point that centrifugation moves more than nuclei, the referee is right and we do not present it as a clean perturbation. The manuscript states that displacement is only partially penetrant (Olivetta et al. 2023), gives the previously reported figures for irregular invagination and lysis under identical conditions, and restricts the analysis to coenocytes with clear nuclear displacement.

      Line 257-263: This approach builds on previous work where we used centrifugation to perturb the spatial relationship between nuclei and the cortex and demonstrate that cellularization in S. arctica is sensitive to local nucleocytoplasmic ratios.26 Centrifugal displacement is partially penetrant, with approximately 40% of centrifuged coenocytes showing irregular plasma membrane invaginations and about 10% undergoing lysis.26 Analyses were therefore restricted to coenocytes showing clear nuclear displacement.

      The observation about selection is also correct, and we have made it explicit rather than leaving it implicit. Measurements were indeed made on furrows that progressed far enough to be traced, and the nuclei-depleted cortex carries many additional small indentations that never sustain ingression. We now state this in the Methods and in the legend of Figure 3C. The failure of those small indentations to progress is part of the phenotype itself.

      Line 308-317: Furrow spatial patterning was dramatically altered: in nuclear-depleted cortical regions, furrow initiations were more frequent and closely clustered but failed to progress deeply (Figs. 3B and EV3D). Among the furrows that did progress, spacing was more variable in the nuclear-free region than in the nuclear-enriched half of the same coenocyte. (Fig. 3E). Since furrow spacing is reliably quantifiable only for furrows that progress, the effective furrow spacing in this nuclear depleted region appeared wider and was abolished upon MBC treatment (Fig. EV3F). Taking into account the high number of furrow initials in the nuclear-depleted region (Figs. 3B and EV3D), and furrow separation in the region enriched with nuclei suggests a minimum furrow exclusion zone around individual nuclei and their MT networks.

      On the visibility of nuclei, and as set out in our response to comment 2, nuclei are resolved in single optical sections as regions from which FM4-64 is excluded. In this figure they are now marked with asterisks * in the time-lapse panel, Figure 3B, and Figure EV3D shows a centrifuged coenocyte at both low and high display contrast, so we do not repeat the general point here.

      The final point is the most important one, and the revised analysis answers it directly. The microtubule part is now reported, and far from showing few differences it reverses the relationship between nuclear position and furrow spacing. In control coenocytes the nuclei-depleted region carries wider gaps than the enriched one, and after MT depolymerization it carries narrower gaps.

      Line 317-322: Furrows in nuclear-depleted regions reached shorter lengths (2.14 against 7.50 µm, Fig. EV3G), forming asymmetrically longer compartments on the nuclear-enriched side. This advantage in nuclear-enriched regions was lost upon MBC treatment confirming the role of nucleus-associated MT networks in maintaining the synchronous invagination and uniform cellular partitioning.

      We have also weakened the conclusion the referee quotes, so that it claims a relationship rather than a mechanism.

      Line 322-326: With centrifugation, though the nucleus-associated MT networks also migrate, the regular spacing of the cytoskeletal network at the cortex was disrupted, and furrows no longer exhibited the uniform spacing characteristic of control coenocytes, suggesting indeed that MT-defined nuclear territories serve as spatial landmarks that define the coordinates of membrane invagination.

      Figure 4 is surprisingly described in a single short paragraph. While some things were quantified, I was not convinced that they could conclude that they observed furrows “mispositioned like those seen upon MT depolymerization”. More broadly, what do we really learn from this?

      This section has been substantially expanded, and the comparison the referee doubted is now made statistically. The section now opens with what is known about membrane supply during cellularization in other systems, including the link between the Golgi and microtubules that makes this perturbation informative in the first place, and it reports a dose series rather than a single condition.

      Line 363-375: While these experiments define how cellularization is spatially patterned, the cellular machinery driving membrane invagination itself remained to be identified. The previous experiments show that furrow invagination proceeds even when furrows are mispositioned. This indicates that the processes of new membrane addition and furrow positioning are controlled independently. To identify the cellular processes driving furrow invagination, we examined the role of membrane trafficking. In early Drosophila embryos, the membrane expands from a reservoir of microvilli localized at the apical cortex.33,34 The first phase of this membrane invagination is supplemented partly by Golgi-derived vesicles. These vesicles are transported in a MT-dependent manner and are stalled with colcemid or colchicine treatment which eventually prevented furrow invagination.18,23,35 It is unclear if such a membrane reservoir is present at the S. arctica cortex where the furrows are first initiated. Brefeldin A treatment in Drosophila embryos inhibits furrow progression in the final stage of cellularization.15,18

      On the dose series, a high dose of Brefeldin A at 3 µg/ml interferes with and in some coenocytes blocks cellularization, which establishes that Golgi-mediated trafficking is required for the process and is shown in Figure EV4A to Figure EV4C. All quantitative measurements are made at 1.5 µg/ml, a dose that leaves ingression intact, and the Methods now state this separation explicitly so that no measurement is read as belonging to the blocking dose.

      On the resemblance to MTs depolymerization, the referee is right that this could not be concluded from the images, and we have therefore tested it. Furrow spacing variability was measured identically in the experiments. They are statistically indistinguishable, and so are their two controls, which is what makes the comparison meaningful.

      Figure 4C: (C) Variability of furrow spacing within a coenocyte, from the same measurements. Methanol 0.212, BfA 0.433, p = 0.0047. Measured the same way in both experiments, the two perturbations are indistinguishable from one another (MBC 0.446 versus BfA 0.433, p = 0.66) and so are the two controls (DMSO 0.213 versus methanol 0.212, p = 0.67).

      On what is learned from the experiment, the answer is that the two perturbations dissociate, and this is now the organising result of the manuscript. Ten measurements are compared between them in Figure 4F, each as a log2 fold change against its own control and grouped by the property it belongs to, and Figure 4G summarises which machinery contributes to which property.

      Line 389-395: Nuclear positioning, by contrast, was not detectably affected as neither inter-nuclear distance nor its variability differed from controls (Figs. EV4E, F). Ingression rate (Fig. 4E), its variability (Fig. EV4G) and the maximum length reached by furrows (Figs. EV4H, EV4I) were likewise comparable to controls. Cellularization nonetheless took longer to complete (from 40 to 50 min, Fig. EV4J), and the cells released at the end were of more variable size (0.140 to 0.266, Fig. EV4K).

      Trafficking is therefore required for cellularization to proceed, since the high dose blocks it, yet at a dose that leaves ingression rate untouched what it contributes is where furrows form and whether cellularization finishes on time. Nuclear organisation is unaffected at that dose, in Figure EV4E and Figure EV4F, which is precisely where the two perturbations part company and is the reason the resemblance in furrow positioning is informative rather than trivial.

      The closing paragraph of this section has been rewritten as described in our response to comment 3, and it now ends on this comparison rather than on a general statement.

      Significance

      As I note in detail in the previous section, Here the authors extend this work to the ichthyosporean Sphaeroforma arctica, exploring the fascinating process by which a syncytial life stage cellularizes, and seeking to define the role of microtubules and Golgi trafficking. Their new expansion microscopy delivers impressive images of membrane and the cytoskeleton during this process and has the promise of answering important questions about the roles of microtubules and membrane trafficking. I thus went into my review quite excited. However, as I detail below, while some aspects of the process are carefully quantified, the current manuscript draws multiple broad conclusions that do not seem fully supported by the limited data provided. This substantially reduced my enthusiasm.

      We are grateful that the referee finds the ExM compelling. The concern that broad conclusions outran the data has driven most of this revision, and every conclusion in the manuscript is now either carried by a measurement made across the dataset or has been removed.

      Reviewer #2

      Evidence, reproducibility and clarity

      Summary

      This interesting paper is a follow-up from Dudin et al.'s seminal 2019 eLife paper describing cellularization in the ichthyosporean Sphaeroforma arctica. In that earlier story, a role for microtubules (MTs) in cellularization was supported by treatment with the MT-depolymerizing drug MBC, which deeply affected nuclear spacing and the regularity of cellularization. However, the difficulty of imaging microtubules at the time had prevented more in-depth functional studies. This technical barrier has now been lifted by ultrastructural expansion microscopy (U-ExM), and this paper thus picks up where the earlier study left off.

      The study combines live imaging, drug treatments, electron microscopy and ultrastructural electron microscopy to support a role for microtubules, nuclei, and membrane trafficking in sustaining the fidelity of S. arctica cellularization. An extensive live imaging dataset and careful image quantifications reinforce and expand the previously published observation that microtubules, while dispensable for cellularization to occur at all, are necessary for it to occur with proper timing and spacing. UEx-M images support the idea that actin and microtubules guide plasma membrane invaginations, and centrifugation experiments support a role for nuclear positioning in ensuring fidelity of cellularization. Finally, Brefeldin A treatment followed by live imaging and U-ExM supports a role for membrane trafficking in furrow positioning and elongation.

      Major comments

      The conclusions are adequately supported by the data, and their limitations are transparently acknowledged: notably, it is not fully clear by what mechanisms nuclei guide cellularization, and whether those mechanisms depend on microtubules or not. Below are a few points where I feel additional data (or better visualization, or additional verbal caveats) could improve the manuscript.

      We are grateful that the referee reads the work as picking up where the 2019 study left off, and that they find the conclusions supported and the limitations transparently stated. Their four points are addressed below.

      1) After 12,000rpm centrifugation, the authors point out that furrows "were misaligned, frequently enclosing multiple nuclei per compartment". This is not obvious in Figure 3C, where nuclei are only visible as empty spaces and only labelled (by asterisks) at a relatively early stage of furrow ingression. This point could be better supported by more explicit images (with nuclei more evident at late stages, even if only as blank spaces), or maybe by co-staining DNA and either membrane or F-actin in centrifuged samples.

      The co-staining the referee suggests is now provided. Figure 3A and Figure EV3A show centrifuged coenocytes by U-ExM with DNA, membrane, actin and tubulin labelled together, so nuclei are seen directly rather than as gaps. Asterisks now mark nuclei throughout the time-lapse in Figure 3B, and Figure EV3D shows a centrifuged coenocyte at both low and high display contrast. As set out in our response to Ref 1-2, nuclei are scored in live imaging as regions from which FM4-64 is excluded.

      2) In centrifuged samples (Figure 3A), some staining is visible in tubulin channel of the nuclei-free part, but does not correspond to discrete, observable microtubules. Could the authors comment on this? Do they think these represent diffuse, perhaps damaged microtubules (nucleated independently of nuclei?), or perhaps free tubulin, or mere background?

      We have measured it rather than interpreted it. These objects are short tubulin-positive fragments, clearly distinct from nuclear asters in both number and length, and we report them separately for that reason. We do not claim to know whether they are remnants of displaced MTs or independently nucleated, and the manuscript says so by describing them rather than assigning them an origin.

      Figure EV3C: In the nuclei-depleted cortex of the same coenocytes, tubulin was present only as short fragments, a median of three per coenocyte and 0.95 µm long (see Data EV1), which are different objects from an aster and are not tested against it.

      3) Similarly: in centrifuged samples Fig. 3C, the nuclei-free part of the cell does seem to cellularize, albeit slower and forming smaller compartments than the nucleated part. This suggests that nuclei (and microtubules?) guide cellularization, but are perhaps not necessary for it. Could the authors comment?

      We agree, and this is now quantified. In the nuclei-depleted region furrows are more widely and more variably spaced, reach shorter lengths and ingress more slowly than in the nuclei-enriched region of the same coenocyte, yet they form and progress. Nuclei therefore guide cellularization without being required for it, which is also what Referee 3 takes from the same experiment.

      Line 308-317: Furrow spatial patterning was dramatically altered: in nuclear-depleted cortical regions, furrow initiations were more frequent and closely clustered but failed to progress deeply (Figs. 3B and EV3D). Among the furrows that did progress, spacing was more variable in the nuclear-free region than in the nuclear-enriched half of the same coenocyte. (Fig. 3E). Since furrow spacing is reliably quantifiable only for furrows that progress, the effective furrow spacing in this nuclear depleted region appeared wider and was abolished upon MBC treatment (Fig. EV3F). Taking into account the high number of furrow initials in the nuclear-depleted region (Figs. 3B and EV3D), and furrow separation in the region enriched with nuclei suggests a minimum furrow exclusion zone around individual nuclei and their MT networks.

      4) The bottom row of Figure 3C presents a series of experiments combining MBC treatment with 12,000rpm centrifugation, but if I am not mistaken these are not discussed at all in the text. This is unfortunate, as it would be interesting to know what the authors want to conclude from these experiments (whose results are however not quantified, perhaps limiting their scope). In my view, one reason to be interested in this dataset is that it could in principle inform epistatic relationships between nuclei positioning and microtubules, that are otherwise left open: do these guide furrow ingression independently of each others (in which case their effects should be additive) or does one act purely through the other (in which case the combined treatment should not be worse than individual treatments)? In any case, I would suggest either commenting on these data explicitly (perhaps with additional quantification guiding interpretations) or removing them.

      The referee is right that these data were not discussed, and they are now quantified and interpreted. As set out in Ref 1-4, combining centrifugation with MT depolymerization does not add to the effect of nuclear displacement but reverses it, so that the region depleted of nuclei carries narrower gaps rather than wider ones. This is the epistatic test the referee proposes, and it argues against two contributions acting independently and in parallel. The nuclear contribution appears to run through the MT network.

      Minor comments

      • Figure 1 contains two distinct panel H's.
      • Typo p. 6: "actomyosing"
      • Page 7: "longer MT extensions associated with deeper furrows during later stages (Figure 2D, arrowheads)" rather seems to refer to Figure 2B or C.
      • Line and page numbers are missing
      • All corrected.

        Significance

      This paper will be of broad interest to readers interested in the evolution of multinucleated cells and cellularization, and the recurrent involvement of actin and microtubules in these processes (see notably recent papers by the Brugués lab). While it could appear incremental in an ichthyosporean-centric perspective (notably compared to the 2019 eLife paper that had anticipated some of the conclusions), its comparative implications gives it an additional scope in my view. Perhaps this is something the authors themselves could emphasize a bit more in their own conclusion. Finally, although most techniques had been established in earlier papers by the same authors, this study further confirms their robustness and the power of S. arctica as an emerging cell biology model - another nice touch.

      We have followed this suggestion. The closing section now sets our results against cellularization in Drosophila and in chytrids. Also, it compares it against recent work showing that partitioning by MT asters is intrinsically unstable, so that losing its control broadens the distribution of compartment sizes rather than shifting it. That is the pattern we observe in our study. We have hinted at this in the discussion, so that the comparative implications the referee points to are stated rather than left to the reader.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      In their manuscript, Araújo et al. described their investigation into the cellularisation process in Sphaeroforma artica during the multicellular stage of its life cycle. The coenocytes are relatively small at 50 µm in diameter, but the authors developed expansion microscopy and used a new actin probe to provide a detailed description of the ingression of the membrane furrow between adjacent nuclei. They compared the formation of the furrow in the presence or absence of microtubules, as well as in conditions involving regularly spaced or clustered nuclei. They concluded that, while microtubules were not essential to the formation and ingression of the furrow, they were required for the fidelity of the process, as stated in the title of the manuscript.

      A better quantification of the ingression pattern would have improved the data. The centrifugation experiment is particularly interesting, as it forces the accumulation of nuclei on one side of the coenocytes. Surprisingly, this only partially perturbs the position of the furrow. This showed, on the one hand, that nuclei contribute to the positioning of the furrow close to them, and, on the other hand, that an additional mechanism exists independently of them. The analysis does not clarify whether the absence of microtubules in these centrifuged states affects the position of the furrow. Figure 3C seems to suggest partial rescue (suggesting that the contribution of nuclei is microtubule-dependent, but that the peripheral membrane has its own partitioning mechanism), but this has not been quantified.

      The quantification of ingression is considerably expanded, and the framework it now detailed in Ref 1 - 2.

      On the centrifuged coenocytes lacking microtubules, this is now quantified. As set out in Ref 1 - 4, the relationship between nuclear position and furrow spacing does not merely weaken when microtubules are removed, it reverses, which supports their inference that the contribution of nuclei is microtubule-dependent. In the same experiment, the local slowing of ingression where nuclei are absent persists whether or not MTs are present, which is consistent with the second half of their reading, that the cortex retains a partitioning mechanism of its own.

      One minor concern is that it remains unclear whether microtubule disruption affects the positioning of the nuclei, as stated in the text but not confirmed by the quantification (Figure 1F).

      Figure 1F now reports the variability of inter-nuclear distance within each coenocyte, which roughly doubles upon microtubule depolymerization, while the mean is unchanged and is shown separately in Figure EV1G. The claim in the text is therefore about regularity rather than about mean spacing, and it is now matched by the panel that supports it. This is set out in full in Ref 1 - 2.

      Reviewer #3 (Significance (Required)):

      The work is primarily descriptive, and the processes determining the position of the furrow and driving its ingression remain unknown. Nevertheless, the work provides beautiful images of a poorly described yet potentially informative system. These are distant relatives of animals with interesting common characteristics, such as the cellularisation process. The conservation of this process may reveal some key fundamental properties of multicellularity. Despite its limited conceptual advances, the manuscript is thus innovative and interesting.

      We thank the referee for their assessment. The revision strengthens what the paper claims. Comparing our two titratable perturbations across ten measurements shows that furrow positioning and furrow ingression dynamics can be disturbed independently of one another, and that neither prevents cellularization from completing (Figure 4F and Figure 4G). What is degraded is fidelity rather than execution of cellularization.

    2. Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.

      Learn more at Review Commons


      Referee #1

      Evidence, reproducibility and clarity

      The last several years have seen major advances in our understanding of the basic cell biology of the set of single cell animal relatives, led by the authors and their colleagues. These groups have developed several as models and pioneered remarkable microscopy approaches to examine their cell biology. Here the authors extend this work to the ichthyosporean Sphaeroforma arctica, exploring the fascinating process by which a syncytial life stage cellularizes, and seeking to define the role of microtubules and Golgi trafficking. Their new expansion microscopy delivers impressive images of membrane and the cytoskeleton during this process and has the promise of answering important questions about the roles of microtubules and membrane trafficking. I thus went into my review quite excited. However, as I detail below, while some aspects of the process are carefully quantified, the current manuscript draws multiple broad conclusions that do not seem fully supported by the limited data provided. This substantially reduced my enthusiasm. I'd also note, though I did not take this into account in my evaluation, that many of the images are presented at a size that was difficult to interpret, without my electronically enlarging them, and two of the Figures were mis-labeled.

      1. Interpreting most of their Figures requires understanding the basics of cellularization in this organism. Comparing the diagram in Fig. 1B and the images in 1C left me confused. First, are all the images in 1C and similar images later cross sections? Are nuclei dispersed throughout the cytoplasm at the start or restricted to a region near the cortex. If the former, how do more central nuclei get cellularized? The transition from the unperturbed 20 and 40 minute timepoints left me unclear on the normal process. Panel G may have been helpful in this regard but only shows the treated embryo and no untreated one.
      2. The authors make a number of conclusions in Figure 1. Some are carefully quantified but others are not. For example, they state "producing uncoordinated ingression with variable rates, diagonal trajectories, and occasional bifurcations, in contrast to the uniform, perpendicular furrows observed in controls". I was not convinced by the single images provided that these were different-for example, spacing in the unperturbed 10 minute time point seems variable and furrows are not "uniformly perpendicular" in the unperturbed 20 minute time point. I also was puzzled by the lack of change in nuclear spacing while furrow spacing was altered. Later on in this section they state "MBC-treated coenocytes displayed significantly decreased and irregular furrow spacing, resulting in some compartments lacking nuclei entirely (Figure 1G & H)", but none of these images visualize nuclei directly-Perhaps the lower level background in H is supposed to indicate this but no parallel wildtype image is shown. Finally, while their TEM images (in the panels that are either I or J due to mislabeling) are lovely, I was not sure what to conclude from them- only the unperturbed images show the embryo surface for orientation and the lefthand unperturbed image is not as straight as they suggest-I certainly don't think these few images support their strong, detailed conclusions here: "MBC treatment resulted in aberrant membrane architecture, with furrows exhibiting bifurcated and convoluted morphology and abnormal fusion at the base of invagination, compared to the smooth, organized structure of control furrows".
      3. The images in Fig. 2B are remarkable and very informative, though as I note above they are presented at such a small size that they require considerable enlargement to appreciate. The surprising accumulation of actin at the invagination front, presumably long before membrane closure begins, was striking, as were the microtubule baskets. However, conclusions drawn again seemed too strong. The authors state "High magnification views further supported that these bundles closely tracked the advancing furrow fronts, with longer MT extensions associated with deeper furrows during later stages (Figure 2D, arrowheads)" (BTW once again this Figure was not labeled in parallel with the text-should be 2C). I did not think the NHS staining provided sufficient resolution of advancing furrow fronts to draw this conclusion. They end this section with some more detailed conclusions, which did not seem to me to be well supported by the single image shown: "Furrows were often misaligned, and compartments frequently enclosed multiple nuclei or, conversely, lacked nuclei entirely. In several cases, nuclei were observed trailing between furrows or located beneath partially formed compartments, suggesting that improper nuclear positioning may interfere with furrow progression and sealing (Figure 2F)." The latter conclusion also seemed to leave me wondering about cause and effect. The final sweeping conclusions in the paragraph on p. 7 top (next time please include page numbers) thus seemed much too broad.
      4. I thought the use of centrifugation to move nuclei was clever. However, it also is moving many other things-for example it moves whatever organelles are labeled by BODIPY and apparently nuclear associated MTOCs, leading to some caveats and calling into question their claim that it "doesn't disrupt MTs". Once again, broad conclusions were drawn based on an n=1 image: "furrow ingression proceeded with kinetics comparable to controls, confirming that the core machinery for membrane trafficking and actin-driven invagination remained functional (Figure 3B). However, furrow spatial patterning was dramatically altered: in nuclear-depleted cortical regions, furrow initiations were more frequent and closely clustered, but failed to progress deeply. In nuclear-enriched regions, furrows progressed with kinetics comparable to controls but were misaligned, frequently enclosing multiple nuclei per compartment rather than the single nucleus observed in controls" Only one thing was quantified-furrow ingression rate-and this must have been done on the selected set of furrows that progressed, and not, for example, of ones like those at the bottom of the image series presented. None of the other conclusions about spatial patterning were quantified-for example, nuclei are not even visualized in Fig 3C. Finally, they do not even mention the results of the MT perturbation presented in this Figure, and the fact that few differences are apparent calls into question their conclusion that that "nuclei (and their associated MTOCs) serve as spatial landmarks that pattern membrane invagination".
        1. Figure 4 is surprisingly described in a single short paragraph. While some things were quantified, I was not convinced that they could conclude that they observed furrows "mispositioned like those seen upon MT depolymerization". More broadly, what do we really learn from this?

      Referees cross-commenting

      Unfortunately I remain convinced that this manuscript has significant issues with the match between the data presented and the claims made--I laid these issues out clearly and stand by them

      Significance

      As I note in detail in the previous section, Here the authors extend this work to the ichthyosporean Sphaeroforma arctica, exploring the fascinating process by which a syncytial life stage cellularizes, and seeking to define the role of microtubules and Golgi trafficking. Their new expansion microscopy delivers impressive images of membrane and the cytoskeleton during this process and has the promise of answering important questions about the roles of microtubules and membrane trafficking. I thus went into my review quite excited. However, as I detail below, while some aspects of the process are carefully quantified, the current manuscript draws multiple broad conclusions that do not seem fully supported by the limited data provided. This substantially reduced my enthusiasm.

    1. Cognitive Process Description Example Cognitive accessibility Some schemas and attitudes are more accessible than others. We may think a lot about our new haircut because it is important to us. Salience Some stimuli, such as those that are unusual, colorful, or moving, grab our attention. We may base our judgments on a single unusual event and ignore hundreds of other events that are more usual. Representativeness heuristic We tend to make judgments according to how well the event matches our expectations. After a coin has come up heads many times in a row, we may erroneously think that the next flip is more likely to be tails. Availability heuristic Things that come to mind easily tend to be seen as more common. We may overestimate the crime statistics in our own area because these crimes are so easy to recall. Anchoring and adjustment Although we try to adjust our judgments away from them, our decisions are overly based on the things that are most highly accessible in memory. We may buy more of a product when it is advertised in bulk than when it is advertised as a single item. Counterfactual thinking We may “replay” events such that they turn out differently—especially when only minor changes in the events leading up to them make a difference. We may feel particularly bad about events that might not have occurred if only a small change might have prevented them. False consensus bias We tend to see other people as similar to us. We are surprised when other people have different political opinions or values. Overconfidence We tend to have more confidence in our skills, abilities, and judgments than is objectively warranted. Eyewitnesses are often extremely confident that their identifications are accurate, even when they are not.

      GUIDE ON ALL THIS CHAPTRS VOCAB WORDS!!

    1. are common in our society and significantly impactthose afflicted with the disease. Similarly, Barryet al. (2009) examined how obesity metaphors,such as “obesity as sinful” (gluttony), affect indi-viduals’ support for diff

      I think this discussion is important to keep in mind as we discuss a patient's condition and consider the expeirence of their illness - not only the physical aspects and treatment plans, but the connotations surrounding their disease and how those may impact their incentive to seek care, trust medical providers, and precieve themselves.

    1. It may not directly mention national health insurance,but it is nevertheless a strong political statement advocating government involvement inhealth care at a time when there was little political consensus on precisely this issue

      This idea of "health for all" is fascinatingly contradictory to the rhetoric of the time and even more interesting, still such a point of contention today. The battle between the human right to health care and a fear of communism/socialism is ongoing as the two are contradictory. In order to have health care for all and to truly care for the poor, health care must not depend on one's income, and insurance ought not depend on one's employment status. I think if we deem something to be a human right, the govement is well positioned to fund that. However, this requires the cooporation of all voting members of a democratic society to work well in spite of the structure of healthcare funding. For example, in the US where healthcare is privatized, the quality of care is unmatched, but prices are astronomical. But in countries with socialized healthcare, the price is affordable yet the quality often suffers.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this paper, Chen et al. identified a role for the circadian photoreceptor CRYPTOCHROME (CRY) in promoting wakefulness under short photoperiods. This research is potentially important as hypersomnolence is often seen in patients suffering from SAD during winter times. The mechanisms underlying these sleep effects are poorly known.

      Strengths:

      The authors clearly demonstrated that mutations in cry lead to elevated sleep under 4:20 Light-Dark (LD) cycles. Furthermore, using RNAi, they identified GABAergic neurons as a primary site of CRY action to promote wakefulness under short photoperiods. They then provide genetic and pharmacological evidence demonstrating that CRY acts on GABAergic transmission to modulate sleep under such conditions.

      Weaknesses:

      The authors then went on to identify the neuronal location of this CRY action on sleep. This is where this reviewer is much more circumspect about the data provided. The authors hypothesize that the l-LNvs which are known to be arousal promoting may be involved in the phenotypes they are observing. To investigate this, they undertook several imaging and genetic experiments.

      While the authors have made improvements in this resubmitted manuscript, there are still multiple concerns about the paper. I think the authors provide enough evidence suggesting that CRY plays a role in sleep under short photoperiod. The data also supports that CRY acts in GABAergic neurons. However, there are still major issues with the quality of the confocal images presented throughout the paper. In many cases it appears that the images are oversaturated with poor resolution, making it hard to understand what is going on. In addition, none of the drivers used in this study are specific to the neurons the authors aim to manipulate. Therefore, the identity of the GABAergic neurons involved in this CRY dependent sleep mechanism remains unclear. Similarly, whether l-LNvs are the target of this GABA mediated sleep regulation under short photoperiod is not fully demonstrated. The data presented suggests that but does not prove it.

      Major concerns:

      (1) While the authors provided sleep parameters like consolidation or waking activity for some experiments. These measurements are still not shown for several experiments (for example Figures 2E, 3, 4, 5, and 6). These data are essential, these metrics must be reported for all sleep experiments.

      These metrics have now been added to Fig.2 S4 and 5, Fig.3 S1-3, Fig.4 S2 and 3, Fig.5 S2 and 4, as well as Fig.6 S1.

      (2) Line 144 "We fed flies with agonists of GABA-A (THIP) and GABA-B receptor (SKF-97541) (Ki and Lim, 2019; Matsuda et al., 1996; Mezler et al., 2001). Both drugs enhance sleep in WT," The proper citation is needed here, Dissel et al., 2015 PMID:25913403. Both THIP and SKF-97541 were used in that paper.

      Thank you for pointing this out. We have modified our manuscript accordingly.

      (3) Figure 2C and 2F: it appears that the control data is the same in both panels. That is not acceptable.

      Thank you for pointing this out. We are now using data from control flies that were monitored in the same experiments as the experimental groups.

      (4) Figure 4A: With the quality of the images, it is impossible to assess whether GABA levels are increased at the l-LNvs soma.

      We apologize for the poor quality. Unfortunately, the GABA immunostaining does not work very well in our hands and thus the background is high. We have now commented on this issue in the fourth paragraph of discussion and have toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript.

      (5) Fig 4 S1A shows colabeling of l-LNvs and Gad1-Gal4 expressing neurons. They are almost 100% overlapping signals. This would indicate that the l-LNvs are GABAergic themselves, or that there is a problem with this experiment.

      Fig 4 S1A demonstrates the expression pattern of SYT-GFP driven by Gad1GAL4, which should label the synaptic terminals of GABAergic neurons. Therefore, the labeling observed at l-LNvs suggest that GABAergic neurons project to l-LNvs. This is further validated by the GRASP and trans-Tango experiments.

      (6) Fig 4 S1B: Again, I can see colabelling of the GFP and PDF staining, suggesting that Gad1-Gal4 expresses in l-LNvs.

      Fig 4 S1B demonstrates anatomical sites where GABAergic neurons project to and form synaptic connections with PDF neurons. Therefore, GFP signals at the l-LNvs suggest that these cells receive synaptic inputs from GABAergic neurons, echoing the results shown in Fig 4 S1A.

      (7) Line 184: "Consistently, knocking down Rdl in the l-LNvs rescues the long sleep phenotype of cry mutants (Figure 4-figure supplement 1D)." This statement is incorrect as the driver used for this experiment, 78G01-GAL4 is not specific to the l-LNvs, so it is possible that the phenotypes observed are not coming from these neurons.

      Thank you for pointing this out. We have modified our manuscript to note this.

      (8) Figure 4G-K: None of these manipulations are specific to the l-LNvs. The authors describe 10H10-GAL4 and 78G01-GAL4 as l-LNvs specific tools, but this is not the case. Why not use the SS00681 Split-GAL4 line described in Liang et al., 2017 PMID: 28552314? It is possible that some of the effects reported in this manuscript are not caused by manipulating the l-LNvs.

      Thank you for pointing this out. We have now modified our manuscript to avoid misleading remarks. We have used SS00681 Split-GAL4 to express TrpA1 but did not observe any substantial effect on sleep duration under short photoperiod. Therefore, we did not use this line for further experiments.

      (9) Similarly for the manipulation of s-LNvs, the authors cannot rule out effect that are coming from other cells as R6-GAL4 is not specific to s-LNvs.

      We have now modified our manuscript to avoid misleading remarks.

      (10) The staining presented in Fig 5 S1 is not very convincing. Difficult to see whether Gad1-GAL4 only expresses in the s-LNvs.

      We have now quantified the GFP signal in the l-LNvs and s-LNVs in Fig.5 S1B and D. As can be seen, the s-LNvs show prominent signal above the background while the l-LNvs do not.

      Reviewer #3 (Public review):

      Summary:

      In humans, short photoperiods are associated with hypersomnolence. The mechanisms underlying these effects is however, unknown. Chen et al. use the fly Drosophila to determine the mechanisms regulating sleep under short photoperiods. They find that mutations in the circadian photoreceptor cryptochrome (cry) increase sleep specifically under short photoperiods (e.g. 4h light: 20 h dark). They go on to show that cry is required in GABAergic neurons and that the effects of the cry mutation on sleep are mediated by alterations in GABA signalling. Further, they suggest that the relevant subset of GABAergic neurons are the well-studied small ventral lateral neurons that they suggest inhibit the arousal promoting large ventral neurons via GABA signaling

      Strengths:

      Genetic analysis to show that cryptochrome (but not other core clock genes) mediates the increase in sleep in short photoperiods, and circuit analysis to localise cry function to GABAergic neurons.

      Weaknesses:

      The authors' have substantially revised their manuscript, and the manuscript is better for the revisions. However, the conclusion that the sLNvs are GABAergic is unfortunately still not well supported by the data. A key sticking point remains the anti GABA immunostaining, and specific driver lines for sLNvs and lLNvs.

      The authors should tone down their conclusions to reflect the fact that their data, as presented, does not support the model that cry acts in sLNvs to modulate GABA signalling onto lLNvs and thus modulate sleep.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript in the Introduction, Results and Discussion.

      Reviewer #4 (Public review):

      Summary:

      Short photoperiod is an important experimental manipulation in neurobiology, endocrinology, and metabolism studies. However, the molecular mechanisms by which short photoperiod gives rise to behavioral phenotypes that are seen in seasonal affective disorders remain unknown. Using the classic circadian model organism Drosophila, this study examines short photoperiod-induced hypersomnolence and identifies the circadian photoreceptor cryptochrome as a regulator of GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions. The discovery has broad implications for understanding how short photoperiod modulates neural inhibition in circadian circuits in regulating sleep.

      Strengths:

      The Drosophila model provided a powerful platform to dissect the molecular mechanisms underlying short photoperiod-induced hypersomnolence. A battery of behavioral, imaging, circuit-manipulation approaches was employed to test the novel hypothesis that the circadian photoreceptor cryptochrome modulates GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions.

      Weaknesses:

      The current model proposed by the authors suggests that the small ventral lateral neurons of the Drosophila clock circuit are GABAergic; however, this remains unclear. At present, the field lacks sufficient data and validated reagents to definitively establish the GABAergic identity of these neuropeptidergic neurons.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript.

      Recommendations for the authors:

      The manuscript has improved after revisions. However, the evidence in support of the claim that the sLNVs secrete GABA onto the lLNvs remains unconvincing. The evidence that loss of cry in GABAergic neurons modulates sleep is solid. However, the authors' claim that the sLNVs are the relevant GABAergic neurons is not sufficiently backed up by the evidence presented. We suggest that the authors tone down their conclusions to reflect this.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript in the Introduction, Results and Discussion.

      Reviewer #3 (Recommendations for the authors):

      Minor points:

      (1) The authors suggest that the effects of cry on sleep and mediated by the Rdl receptor, and use the GABA agonist THIP as support of this argument. However THIP acts on the Lcch3 and Grd receptors, not Rdl

      Thank you for pointing this out. We have modified relevant discussion accordingly.

      (2) In several instances (e.g. line 66, line 148), the authors use 'consistently' in the sense of 'consistent with previous data'. It would be better if they rephrase this.

      This has been fixed.

      Reviewer #4 (Recommendations for the authors):

      It is my pleasure to serve as a reviewer for this revised manuscript. The authors have carefully revised the manuscript in response to the critiques raised by all previous reviewers and have used all the available reagents to conduct additional experiments to assess the GABAergic properties of the small ventral lateral neurons (sLNv). Although it remains unclear in the field whether sLNvs co-transmit GABA, this study raises this possibility within an interesting biological relevant context. I recommend this manuscript for final publication.

      Thank you for your comments.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors characterize the phospholipid scramblase Xkr in Drosophila. They generate null mutants in both S2 cells and flies and find that phosphatidylserine (PS) exposure is reduced during apoptosis; they show reduced engulfment of apoptotic cells, and that the protein is localized partially within the cytoplasm, overlapping with the ER. They go on to identify Xkr binding partners and show that they overlap with plasma membrane-ER contact sites, suggesting that Xkr facilitates PS transfer from the ER to PM. Overall, this reveals a new role for Xkr and identifies new binding partners, which are valuable contributions to the field.

      Strengths:

      (1) The generation of new Xkr reagents in both S2 cells and flies to analyze its function. Tools are used to quantify both PS exposure and efferocytosis, and the effects of Xkr knockout are significant.

      (2) The discovery of new binding partners of Xkr which also affect PS exposure and efferocytosis.

      (3) The authors demonstrate that the binding partners are conserved in mammalian cells.

      Weaknesses:

      (1) Throughout the manuscript (e.g, lines 105, 165, 274 and discussion), the authors describe Xkr as being activated in a caspase-independent manner, and use this as the rationale for identifying binding partners. However, this is never shown in the manuscript or clearly referenced. Interestingly, there is a TEVDA sequence in the fly ortholog at the same location as the caspase cleavage site in C. elegans Ced-8 (Figure S1), suggesting the caspase cleavage site is conserved. This should be further investigated, or the statements regarding caspase independence should be modified. I don't think the N- and C-terminal GFP fusions indicate caspase independence, especially since apoptosis was not induced in Figure 1A, B. If cleavage occurred at the TEVDA site in Figure S1A, it would not lead to a noticeable change on the Western blot, although the size does look a bit smaller in Figure S2B at the 8 h time point.

      We thank the reviewer for pointing out that, as E/DXXD has been considered a conserved caspase-3 cleavage site, TEVDA has also been validated as a caspase-6 cleavage site, which we have missed. We will further confirm this using site-mutated expression vectors in S2 cells.

      (2) The authors examine overlap between tagged Xkr and cellular compartment markers and find substantial overlap with Lamp (and other vesicle markers to a lesser extent) (Figure S2). This is not addressed in the paper and could indicate engulfment of other cells since S2 cells are macrophages. To test this, the staining could be tested on the mixed cells (vesicle-GFP tagged S2 + apoptotic xkr-mcherry). Similarly, calreticulin is an eatme signal that gets translocated to the PM of apoptotic cells. This could affect interpretation of colocalization (Figure 2J), and ideally another ER marker should be used.

      We thank the reviewer for the suggestion. We will attempt to label Xkr-mCherry under apoptosis with other vesicles and change the ER marker to Cnx99A (Calnexin ortholog in Drosophila).

      (3) There are some places where there is over- or incorrect interpretation, and these instances should be corrected.

      We thank the reviewer for their careful reading, and we will correct the mistakes in the revised manuscript.

      Specific examples:

      a) Line 342 "Relative expression analysis by RT-qPCR showed that all three mutants were likely null alleles." This does not make sense since there is still mRNA present. In Figure S7A, the tm9sf4 allele is expressed at 75% of the control. The others show a greater reduction, but this is not proof of a null allele.

      We agree with the reviewer’s opinion. These mutants from the BDSC are not completely deleted but partially deleted; therefore, the RT-qPCR assay may not be very accurate. We will detect the mRNA levels of tm9sf4, dorp9, and sac1 using RT primers from different cDNA regions to make the results more convincing.

      b) Figure S3I - It looks like mCherry-Lact:C2 does get localized to the PM with AcD treatment in the xkr[ko], although the authors conclude "this disrupted PS localization to the PM could not be restored by apoptosis induction". However, the PM localization does look disrupted in the tm9sf4 and sac1 knockdowns.

      We thank you for raising this intriguing hypothesis. Indeed, PM localization of Lact:C2 was reduced in xkr<sup>ko</sup> cells, and the distribution could not be rescued after apoptosis. Unlike xkr<sup>ko</sup>, tm9sf4, and sac1 RNAi-treated cells displayed weak PS disorder, which may be due to the efficiency of knockdown. However, the statistical results indicated that the ratio of PM/Cyto was reduced in tm9sf4 and sac1 RNAi-treated cells.

      c) Figure 3I. The control Lact:C2 staining looks very different from the staining in Figure 2J, with abundant Lact:C2 outside the cell. Given the variability in the staining, were the contact sites quantified? On lines 287-288, it is stated that "fewer ER-PM MCSs were detected in xkrko cells than in WT", but no quantification is provided.

      We thank for the reviewer’s suggestion. We will add the statistical results of Fig. 3I in the revised version.

      d) Line 299-300 - "the interaction between Xkr and dORP9 was enhanced after apoptosis induction". The interaction does not look enhanced in Figure S5F, so this statement should be removed or data supporting the statement should be provided. The interaction between Xkr and dORP2 looks enhanced upon apoptosis induction, but also paradoxically looks even more enhanced when apoptosis is blocked.

      We thank you for raising this intriguing hypothesis. We will delete the relevant statement to eliminate unnecessary misunderstandings.

      e) The data in Figure S6 are highlighted in the abstract. If this is a major conclusion, it would be best to move it to the main text and provide quantification.

      We thank for the reviewer’s suggestion. We will move this to the main text and provide quantification in the revised version.

      f) Lines 392-4. The concluding statement seems overstated given that there was only a modest inhibition of PS exposure in the osbpl5 knockdown (Figure 6A) and no defects in efferocytosis (Figure 6C). The osbpl8 showed a stronger effect on PS exposure but still a very modest effect on efferocytosis.

      We thank for the reviewer’s suggestion. We will weaken the statement in the Results section of Figure 6 and perform osbpl9 knockdown to observe efferocytosis in Raw264.7 cells, as OSBPL9 interacts with Xkr8 strongly.

      Reviewer #2 (Public review):

      In this study, the authors investigate the mechanisms underlying phosphatidylserine (PS) exposure during efferocytosis in Drosophila. They first show that Xkr promotes PS exposure and apoptotic cell clearance in both S2 cells and Drosophila embryos. As Drosophila Xkr lacks the canonical caspase cleavage site found in mammalian XKR proteins, the authors further explore the underlying mechanism by which Xkr regulates PS externalization. Through protein interaction studies, they identify TM9SF4 as an interacting partner of Xkr that regulates PS distribution and show that non-vesicular PS transport contributes to apoptotic PS exposure and efferocytosis. Using protein interaction studies, they further demonstrate that Xkr interacts with the lipid transfer protein dORP9 at ER-PM contact sites to facilitate non-vesicular PS transport to the plasma membrane. Loss of these proteins affects PS externalization and efferocytosis in Drosophila. Finally, using human cells, they demonstrate that human OSBPL8 interacts with XKR8 to regulate apoptotic PS exposure. Overall, the study supports a model in which Xkr promotes efferocytosis by facilitating lipid transport in addition to its role as a phospholipid scramblase.

      Thank you for your comprehensive and generous assessment of our work and for the time and expertise you have devoted to reviewing our manuscript. We will revise the manuscript accordingly and provide a point-by-point response in the revised version.

      Reviewer #3 (Public review):

      Summary:

      The manuscript investigates the function of the Drosophila Xkr protein, a homolog of mammalian Xkr8 that lacks the canonical caspase-cleavage motif. The authors show that apoptotic stimuli increase Xkr protein abundance through a post-transcriptional mechanism and that Xkr promotes phosphatidylserine (PS) exposure during apoptosis. Using immunoprecipitation coupled with mass spectrometry, they identify TM9SF4 as an Xkr-interacting protein and further implicate TM9SF4, Sac1, dORP2, dORP9, and Vap33 in regulating apoptotic PS exposure and efferocytosis. Based on these findings, the authors propose that Xkr regulates PS transport at ER-PM contact sites. Similar observations are also presented in human cells.

      Strengths:

      Overall, this is an interesting study. The authors provide convincing evidence that Drosophila Xkr participates in apoptotic PS exposure and employ multiple complementary approaches to support the involvement of several proteins in this pathway. The identification of TM9SF4 as a potential regulator of Xkr-mediated PS exposure is likely to be of broad interest.

      Weaknesses:

      I am less convinced by the evidence supporting the proposed role of ER-PM contact sites, and several mechanistic conclusions appear to extend beyond the data presented. Addressing the following points would substantially strengthen the manuscript.

      We sincerely thank you for your careful reading and accurate summary of our manuscript. We appreciate the time, effort, and expertise you have dedicated to evaluating our work, and we will try our best to improve our manuscript according to your suggestions.

      Major concerns:

      (1) In Figure 2A and related text, it is unclear whether the mass spectrometry analysis was performed using untreated cells or AcD-treated cells. If the objective was to identify apoptosis-associated Xkr interactors, it would be helpful to clarify the experimental condition and explain whether apoptosis-specific interactors were analyzed separately.

      We thank for the reviewer’s suggestion. We used AcD-treated S2 cells and untreated S2 cells to perform mass spectrometry. To clarify this, we will add a detailed method description in the method section.

      (2) In Figure 2B, 2E, and several other co-IP results, a negative control of Flag tag only is required to exclude experimental errors like insufficient washing, etc.

      We thank for the reviewer’s suggestion. We used anti-HA magnetic beads to perform immunoprecipitation, and single HA-TM9SF4 was used as a negative control.

      (3) In Figure S3B, S3F, and several other BiFC results, an mVC-only negative control would be important to exclude nonspecific fluorescence complementation.

      We thank for the reviewer’s suggestion, we will add the negative control for BiFC results in the revised version.

      (4) In Figure 2G, the quantitative values appear inconsistent with the flow cytometry histograms. The peak shift following Sac1 knockdown appears smaller than that of TM9SF4 knockdown, whereas the quantified values suggest the opposite. Please clarify this apparent discrepancy.

      We sincerely thank you for the careful consideration of our statistical results, which were obtained from 3 repeats. We will choose another flow cytometry histogram of tm9sf4 and sac1 to make the data and images more consistent.

      (5) I find the interpretation in Lines 223-227 difficult to reconcile with the data. Knockdown of both tm9sf4 and sac1 impaired apoptotic PS exposure to a similar extent as xkr knockout. However, while xkr deficiency significantly reduced efferocytosis, sac1 knockdown produced only a modest, statistically insignificant effect. These observations suggest that impaired PS exposure alone may not fully account for the efferocytosis phenotype observed in xkr-deficient cells. These results appear difficult to reconcile with the proposed model, which needs careful discussion.

      We sincerely thank the reviewer for their careful and thoughtful observations. Given the results we have observed, we will add this to the discussion section in the revised version.

      (6) In Lines 274-275, the authors state that 'increased Xkr may accelerate non-vesicular PS transport for efficient apoptotic PS exposure'. However, Xkr protein levels increase only ~8 h after AcD treatment, whereas PS exposure occurs much earlier. Thus, alternative explanations like Xkr relocalization (Figure S5C), rather than increased abundance, may also explain how Xkr mediates PS transport. An Xkr overexpression experiment could be helpful to support this statement.

      We thank for the reviewer’s suggestion. We will overexpress Xkr with or without AcD treatment to observe whether the localization or amount of Lact:C2 changes and to re-evaluate the role of Xkr in PS exposure.

      (7) The interpretation of the MAPPER experiments requires further clarification. In Line 283, the authors refer to "the intracellular proportion of the signal for each protein overlapping with MAPPER." Since MAPPER is designed to label ER-PM contact sites, which are located on the plasma membrane, intracellular MAPPER fluorescence likely represents the ER network rather than bona fide ER-PM contacts. Throughout the manuscript (including Figure S6, etc.), intracellular MAPPER puncta appear to be interpreted as ER-PM contacts, which may not be appropriate. In contrast, the peripheral MAPPER puncta observed along the cell cortex (e.g., Figure S5C after AcD treatment) are more consistent with authentic ER-PM contact sites. It is also not obvious that these cortical MAPPER signals colocalize with Xkr(Figure S5C). Thus, while the data support a role for the ER, they do not yet convincingly demonstrate Xkr clustering at ER-PM contact sites.

      We thank the reviewer for the suggestion, and we believe that the TIRF technique can help us demonstrate the ER-PM signal. Since our college has no TIRF microscope, we will try our best to seek cooperation from other colleges to achieve this experiment.

      (8) In the Xkr knockout cells, all fluorescence signals appear substantially low in intensity. Differences in protein distribution are difficult to interpret when overall probe expression also appears altered. It would be helpful to demonstrate that probe expression levels are comparable between conditions. Furthermore, as noted above, intracellular MAPPER signal may primarily represent ER rather than ER-PM contacts. Finally, despite the reduced signal intensity, the remaining MAPPER and PS signals still appear well colocalized in the knockout cells, similar to the observations in Figure 2J. The interpretation in Lines 285-288 should therefore be reconsidered.

      We sincerely thank the reviewer for this careful and thoughtful observation, and we agree that the interpretation in Line 285-288 is overstated. To explain this, we plan to detect the Lact:C2 and MAPPER signals in S2 and xkr<sup>ko</sup> cells with or without AcD to confirm how Xkr regulates PS via ER-PM under apoptotic conditions.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Some of the authors proposed in a PNAS paper in 2016 the occurrence of the EntnerDoudoroff (ED) pathway in cyanobacteria and plants, on the basis of several lines of biochemical and genetic evidence. However, more recent results indicated that one of the two specific enzymes of the ED pathway (EDD) is missing in Synechocystis PCC 6803. The authors carried out additional experiments, which demonstrated that EDD is missing, and one of the enzymes (ED aldolase) is a promiscuous enzyme which seems to be involved in proline metabolism and is not actually participating in the ED pathway as initially believed. The results described in this paper are strong evidence that this new interpretation is appropriate, and therefore, it corrects the previous proposal, providing an honest description of the reasons why the authors had reached the wrong conclusion about the existence of the ED pathway in cyanobacteria and plants.

      We thank Reviewer 1 for the summary and comments. Based on our finding that EDA is a promiscuous aldolase that, in addition to the cleavage of KDPG to GAP and pyruvate (a reaction of the ED pathway) catalyzes other reactions in vitro, we proposed potential in vivo functions of EDA, including its involvement in proline metabolism. However, these assumptions require further experimental testing. We do not yet have definitive findings regarding the function of the promiscuous aldolase EDA in Synechocystis in vivo, but respective studies are currently underway.

      Strengths:

      Thorough reanalysis of the experimental results obtained in previous studies, which led to the publication of the PNAS paper in 2016.

      New experimental evidence to confirm that enzymes previously considered as participating in the ED actually are not catalyzing the ED biochemical reactions, but are involved in other metabolic pathways. Also, the authors completely discarded the occurrence of the GDH/GK shunt in Synechocystis PCC 6803. Generally speaking, the manuscript is very clearly written, with a precise description of the previous findings, the mistakes which took place in the 2016 paper, and the strategies they have used to address those issues, in order to reach a thoroughly revised vision of the glucose metabolic pathways in Synechocystis PCC 6803. In this regard, the drawings shown in Figures 1 and 7 are very helpful for the reader to follow the story and understand the possible metabolic transformations depending on the working hypothesis.

      Also, I commend the authors for openly describing previous mistakes. In this paper, they reassess past observations in light of more recent findings and to integrate the information in this manuscript. The scientific conclusions are solid and very interesting, and besides, they use the opportunity to offer valuable advice to researchers. This is especially focused on the importance of careful biochemical characterization of enzymes, which should always be carried out when studying proteins which have been identified as a specific enzyme on the basis of sequence homology. In a similar way, they found that an insertional mutant was the cause of the absence of specific metabolites, which had been attributed to particularities of a metabolic pathway in that mutant, when it was actually due to a nucleotide insertion; this could have been easily prevented by confirming the correct generation of the mutant by DNA sequencing.

      We thank the reviewer for this kind comment. We agree that biochemical characterization of enzymes as well as DNA sequencing to check deletion mutants, are important and valuable tools. As outlined in the manuscript and additionally in more detail in a recently submitted article, which is available at bioRxiv (Theune et al. 2026, doi: https://doi.org/10.64898/2026.04.08.717167) and is currently under review at PLOS One, we suggest that genome sequencing of deletion mutants in combination with complemented strains as controls are required to minimize the risk of misinterpretation based on secondary mutations (1). During the early stages of our research on the ED pathway, and later as well when we were already trying to resolve the conflicting results that had accumulated concerning the ED pathway, genome sequencing for Synechocystis mutants was not affordable as a routine procedure (2-4). Therefore, we could not have easily prevented this misconception based on this technique at that time. However, we strongly encourage genome sequencing of deletion mutants (in combination with complemented strains) as routine procedures these days (1).

      Weaknesses:

      The authors propose that EDA might be involved in the PEP-pyruvate-OAA node, or in the proline metabolism, but this requires further experimental work for clarification; what their results indicate clearly is that this enzyme is not actually catalyzing the transformation of KDPG to GAP, which is the second specific enzyme of the ED pathway. But the real physiological function in this cyanobacterium is still unconfirmed.

      As stated above and in the manuscript, we agree that the in vivo role of EDA requires further experimental work which is in progress. However, our results demonstrate that EDA splits KDPG into GAP and pyruvate in vitro, but we assume that this reaction does not play a role in vivo due to the absence of its substrate.

      Another aspect which could be improved is that the recombinant expression of some genes was carried out in E. coli; even if this is a useful and valid research strategy, in studies like this (where there is a strong focus on the physiological function of enzymes in the original organism, Synechocystis PCC 6803), I think it would have been more appropriate to express the 6803 genes in another cyanobacterium easily amenable for genetic transformation and gene expression, which would produce the protein in a physiological environment more similar to another cyanobacterium (compared to E. coli, which is an heterotrophic bacterium). I am not sure this would change any of the obtained results, but it certainly would confer additional robustness to the enzymatic results.

      Synechocystis is easily amenable to genetic manipulation, and we agree that expression and purification of all enzymes from this host would have been ideal. However, the first characterization of Synechocystis EDA was performed with proteins that were purified from Synechocystis and showed activity on KDPG at comparable rates as proteins that were purified from E. coli in this study (2). Moreover, most biochemical characterizations of EDAs from archaea, bacteria and plants were performed after recombinant expression in E. coli and yielded highly active enzyme as in the case of Synechocystis is this study (5-7). Therefore, we currently have no reason to worry that the expression in E. coli might affect the enzymatic activity of EDA. The main reason for utilizing E. coli as an expression strain in this study was to gain higher yields of protein for in-depth analyses.

      Bibliography:

      I think the list of papers used in this manuscript is complete and up to date. However, I do miss recent papers which addressed one aspect that was proposed in the original 2016 PNAS paper: the authors wrote, "We therefore suggest that Prochlorococcus might oxidize glucose via the ED pathway under mixotrophic conditions, as shown for Synechocystis." Recent studies checked this hypothesis and have shown that the ED pathway seems to be also missing in Prochlorococcus and marine Synechococcus, and I think this manuscript is a good place to cite them, since these results are consistent with the findings of this paper.

      We will include a references from Moreno-Cabezuelo et a. 2023 (DOI: 10.1128/spectrum.03275-22) in which the proteomes of three marine Prochlorococcus and three marine Synechococcus strains were investigated upon exposure to glucose (8). Protein levels of EDA were either downregulated or not affected while proteins involved in OPP pathway and CBB cycle were upregulated. The authors of this study conclude that this indicates that the latter processes rather than the ED pathway are involved in photomixotrophy in these strains. However, flux analyses are still missing.

      Reviewer #2 (Public review):

      Summary:

      The study presents novel results on the presence of the Entner-Doudoroff pathway in Synechocystis sp. PCC 6803. In contrast to an earlier study, compelling evidence is given that this strain lacks both an ED pathway and a glucose dehydrogenase/glucokinase bypass but contains a promiscuous aldolase, which also decarboxylates oxaloacetate and cleaves 2-keto-4-hydroxyglutarate (as it occurs in proline degradation). The study concludes with successfully reconciling data from different studies and with lessons learned from the previous misconception.

      Strengths:

      Solid biochemical data are presented to reconcile contradicting data of earlier studies and to serve as a basis for disclosing possible functions of a promiscuous aldolase. Earlier misconceptions and lessons to be learned are well discussed.

      Weaknesses:

      The materials and methods section is rather lengthy, suffering from a lack of conciseness and repetition, and nevertheless misses some specifications.

      We thank Reviewer 2 for the kind summary and comments and will improve the materials and methods part accordingly in a revised version.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Some additional aspects that could be improved:

      (1) L182-184: Are there any known in vitro attempts to determine whether some DHADs can accept 6PG as substrate? If not, did the authors check this possibility in the lab?

      To the best of our knowledge, as mentioned in the manuscript, some DHADs are tested for their activity towards gluconate and some other substrates but not 6PG. It was suggested that since gluconate is smaller it might fit in the catalytic site of DHADs normally occupied by DHIV (7). We discussed this point in the manuscript (Line 177-185).

      In this study we tested the DHAD Slr0452 from Synechocystis and DHAD from Synechococcus with 6PG as substrate but the enzyme did not catalyze 6PG dehydration. The result for Slr0452 is shown in Figure S3.

      (2) L234-241: This paragraph shows how important it is to avoid relying only on sequence alignments to assign functions to proteins (and enzymes in particular). The authors stress this idea elsewhere in the paper, but I think it should receive even more attention in a manuscript like this. Physiological characterization of the protein function is paramount, especially in the case of enzymes. Also, the GDH1 overexpression mutant was done in E. coli; this adds an additional layer of uncertainty, since the protein processing in E. coli might not be entirely identical to that carried out in cyanobacteria, as I mentioned above. This, in turn, could be one of the reasons for not finding the expected enzymatic activity. This comment is also relevant for the results shown in Table 1 (page 12).

      As outlined following lines L234-241 we tested crude cell extracts from Synechocystis WT, a Synechocystis strain overexpressing putative GDH1 (Sll1709) and a Dzwf deletion mutant which can be assumed to upregulate a GDH/GK bypass if present in Synechocystis for GDH activity but did not find any. This strongly indicates that sll1709 does not code for an active GDH in Synechocystis. We thereafter also overexpressed GDH1 in E. coli and again could not detect any GDH activity. In case of GDH1, enzyme activity was therefore tested both in Synechocystis and E. coli and yielded similar results.

      (3) L309-310: "we currently have no explanation for the gluconate that was detected in previous IC-ESI-MSMS measurements in Synechocystis". This is a serious issue, given how important this evidence was for the conclusions of the 2016 PNAS paper and the hypothesis of the ED pathway in cyanobacteria. I would suggest that the authors provide some possible explanations, even if it is based on studies from other teams.

      It would be very speculative and therefore in our eyes not helpful to search for explanations in this case as indeed the measurements were done in another lab. One explanation might be that 6P gluconate got dephosphorylated and yielded gluconate as an artifact. However, as we have no experimental validation for this idea and furthermore cannot test it. We therefore prefer to not further comment on this aspect.

      (4) L347: The presence of an insertion in the sequence of zwf in the ∆gnd mutant is a welcome explanation for the observed results: abolition of 6PG production in this mutant. This is another important message to stress in this manuscript: construction of mutants should always be double checked by DNA sequencing of the relevant genomic regions, to ensure that these kinds of side problems do not appear, leading to confusing results.

      We agree with the reviewer and would get even one step further and rather suggest combining whole-genome sequencing with complemented mutants as it is difficult to know all relevant genomic regions that should be sequenced. We discuss this issue in even more detail in another manuscript that is currently available as online preprint in bioRxiv and is under review (9). This work was now added as a citation in line 566 and the end of the following statement (line 564-566): Routine complementation of deletion mutants and sequencing of selected genes or the entire genome are effective means of identifying secondary mutations that can lead to misleading phenotypes (9).

      (5) L418-419 and Table 2: The results observed for the authors (i.e., that EDA could also catalyze reactions with OAA and KHG, albeit with substantially lower catalytic efficiency than with KDPG), is a matter for concern: if their hypothesis is correct (meaning that this EDA is fundamentally involved in the PEP-pyruvate-OAA node and/or proline metabolism, and not with the ED pathway), then why should it keep in the evolution of these organisms such a strong preference for KPDG, when it is not being used physiologically for the ED pathway)? Furthermore, the Km for KDPG is lower than for OAA or KHG, leading to a very big difference in the Kcat/Km values.

      To solve these questions further respective studies are underway. As EDD is absent from Synechocystis no KDPG should be available in the cells so that catalytic activity on KDPG should be irrelevant in vivo. The in vivo role of Eda requires further clarification.

      Hereafter, I will mention some aspects, following the instructions of eLife, which are related to suggestions for improved experiments/data/analyses, improvements of writing and presentation, and minor corrections to text/figures.

      (1) L81: I would modify the text to "Accordingly, this raises further questions...".

      Thanks for this hint. We modified the text accordingly.

      (2) L189: Add "pages" after "following".

      Thanks for pointing this out. We replaced “following” by “below”.

      (3) Page 7: The whole beginning of the Results section is actually more discussion than description of results, and the first mention of figures appears in L207 of page 8. Given the content of the paper, I think the authors might reconsider using "Results and Discussion" rather than different, specific sections for Results and Discussion. This is one of the papers where I think the combined use of both makes sense and will allow an easier understanding of the message.

      We thank the reviewer for this valuable suggestion and changed the heading to Results and Discussion.

      (4) L181: Add reference regarding the llvD/EDD superfamily.

      We added the following references in lines 172-176 and 185-189:

      (1) Melse, O., Sutiono, S., Haslbeck, M., Schenk, G., Antes, I., & Sieber, V. (2022). Structure Guided Modulation of the Catalytic Properties of [2Fe− 2S]-Dependent Dehydratases. ChemBioChem, 23(10), e202200088.

      (2) Ren, Y., Vettenranta, E., Penttinen, L., Jänis, J., Rouvinen, J., & Hakulinen, N. (2025). The engineered dimer of L-arabinonate dehydratase from Rhizobium leguminosarum bv. trifolii: The role of intersubunit interactions in IlvD/EDD family. Biochemical and Biophysical Research Communications, 757, 151610.

      (3) Ahmed, H., Ettema, T. J., Tjaden, B., Geerling, A. C., Van Der Oost, J., & Siebers, B. (2005). The semi-phosphorylative Entner–Doudoroff pathway in hyperthermophilic archaea: a reevaluation. Biochemical Journal, 390(2), 529-540. ff

      (4) Bräsen, C., Esser, D., Rauch, B., & Siebers, B. (2014). Carbohydrate metabolism in Archaea: current insights into unusual enzymes and pathways and their regulation. Microbiology and Molecular Biology Reviews, 78(1), 89-175.

      (5) Figure 2: The data shown in column plots in Fig 2A, B and C, and 4B, could be presented in tables, which would save space while providing the same information: basically, very little/no activity in some cases vs high levels of activity in others.

      We would like to keep the column plots as we still think that they visualize our data well.

      (6) L255 "Unfortunately, we were not able to overexpress putative GDH2". It would be interesting to give more details about the possible reasons for this fact.

      We tested different growth conditions for recombinant GDH2 expression. The expression culture was incubated at 37 °C for 3 hours as well as overnight at 18 °C for overnight after induction. Both experiments did not resolve the expression problem.

      (7) L409-410: I think this sentence should include a brief part explaining the kind of essay used to test this activity.

      We added now that the LDH-coupled continuous assay was used (see line 403).

      (8) Figure 6E: Please give the specific activity in U/mg, as in Fig 6F, instead of percents.

      100% is given in U/mg units in the figure legend as “control without effector (100 %; specific activity of 4.3 U/mg)”. For easy comparison of effectors, the relative activity (%) is often used. We would therefore prefer to keep the current data presentation.

      (9) L576: Provide the origin of the utilized PCC 6803 strain, given there is a certain level of variability in this strain (glucose tolerance, etc). Also, even if there are some methods which are very widely used, I think the Materials and Methods section should either properly describe them or else cite the source. For instance, BG11 medium is mentioned, but no further information is given.

      We included the information that the glucose-tolerant Synechocystis strain was utilized and added the receipt of and a citation for BG11 medium (10).

      (10) L582 Generation of mutants: This section mentions the Gibson Assembly cloning method, but I miss further information to allow the reader to reproduce the methodology with as many details as possible, or at least cite papers which do so.

      We added a reference in which Gibson Assembly is described (11). Together with the primers listed in Table S3 the generation of mutants is now reproducible.

      (11) L609 Please give information in g, not rpm, for centrifugation. Also, mention the model and brand of the centrifuge and rotors used. Also, immunoblotting is very loosely described. This is also valid for other sections, for instance, L618, L636.

      We now added the following information: Cells were harvested by centrifugation at an RCF (relative centrifugal force) of 3,992 x g in a Beckman Coulter with a JLA-8.1000 Rotor for 20 minutes at 4°C. We now added a reference (12) in which immunoblotting is described in more detail.

      (12) L613: Describe the "small scale purification".

      We now added the information that the small-scale purification was performed using a 50-ml aliquot of the large culture which was treated as described below for the remaining sample.

      (13) L619-620: Describe the composition of the lysis buffer.

      The composition of the lysis buffer is already described as follows: lysis buffer (50 mM NaPO<sub>4</sub> pH=7.0; 250 mM NaCl; 1 tablet complete protease inhibitor EDTA-free

      (Roche) per 50 mL)

      (14) L691: Specify which amounts of auxiliary enzymes in coupled enzymatic assays were used.

      Thanks for pointing this out. We have now integrated the information that 1U of each of the auxiliary enzymes was utilized in the coupled enzymatic assays.

      (15) L716: The authors mention several times using a "double beam spectrophotometer". Please provide the model and brand.

      Model and brand were now added for the double-beam spectrophotometer (Uvikon 810, Kontron, Augsburg, Germany).

      (16) L717 and 718: define "∆absorption".

      In line 715, the information is given that absorption was measured at 340 nm, "∆absorption" is accordingly the ∆absorption at 340 nm. This information was added.

      (17) L723: "Synechocystis" should be in italics.

      Synechocystis is now written italics.

      (18) L749-759: This section should be described in more detail: preparation of protein extracts, SDS, immunoblotting, etc, or cite references of the same team where these methods were properly described.

      In this section the listed methods are already described in detail.

      (19) 798-799: "frozen cell pellets". Please provide numbers to specify the amount of material used.

      Thanks for pointing this out. We now added the information that frozen cell pellets with a wet weight of 2.4 g wet weight were resuspended.

      (20) L871-872: "It was ensured that auxiliary enzymes were not rate-limiting. One unit (1 U) of enzyme activity is defined as 1 µmol substrate consumed or product formed per minute" is repeated several times in the manuscript (L 907-909, L936-938). I would advise using it the first time, and on other occasions, refer to the same conditions as described above.

      We have accordingly circumvented the repetition of 1 U definition from the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Interpunctuation, especially comma placement, should be improved.

      We improved interpunctuation, especially comma placement to the best of our knowledge.

      (2) Line 63: delete the first "which".

      “Which” was deleted.

      (3) Lines 81/82: revise sentence.

      We revised the sentence to: Accordingly, this raises further questions about the presence of the ED pathway in cyanobacteria and plants.

      (4) Line 164: "presumed" instead of "presumes".

      The word was changed accordingly.

      (5) Line 228: by others? especially in reference 1?

      We deleted by others as the reference is given.

      (6) Line 240: "or" instead of "no".

      “no” was replaced by “or”

      (7) Figure 2: The axes are not well visible, and the explanation for the positive control in panel B is missing in the legend.

      Axes from figures 2, 3 and 4 were enlarged. For Figure 2B the following information was added in the legend: As a positive control, 0.05 U glucose dehydrogenase from Pseudomonas sp. was added to Δzwf cultures and to purified putative GDH1 (Sll1709). Axes from figures 2, 3 and 4 were enlarged.

      (8) The investigation on the general absence/presence of the GDH/GK bypass in cyanobacteria may not be necessary for this study.

      We included this data in this manuscript as the mistaken assumption that the GDH/GK bypass exist in Synechocystis lead among other observations to the misinterpretation of an existing ED pathway in Synechocystis. We would therefore prefer to keep these data in the manuscript.

      (9) Line 316: delete "on".

      “on” was deleted.

      (10) Line 352: delete "or".

      “or” was deleted.

      (11) Lines 352/353: ZWF expression level appears to be reduced accordingly. This should be stated.

      We agree that Zwf expression might be lower, however, we are not entirely sure if this is truly valid and would rather test this assumption further as described in the following lines.

      (12) Figure 4: The axes are not well visible.

      Axes from figures 2, 3 and 4 were enlarged.

      (13) Line 363: values are not only normalized to protein content, but also to activity found for the WT.

      We now added: The values are normalized to Zwf enzyme activity found in the WT based on protein content.

      (14) Line 408: no separate subsection required.

      The title for a new subsection was deleted.

      (15) Lines 437-439: These are results descriptions, which should not be placed in the legend, but in the main text, as is partially done.

      We deleted these result descriptions in the legend.

      (16) Lines 441/442: formatting: one or no bracket pair.

      The brackets were corrected.

      (17) Lines 443/444: refer to Table 2 instead of giving the values in the legend to avoid duplication.

      We deleted the values and now refer to Table 2.

      (18) Line 503: delete "identified".

      We deleted the second “identified” in the sentence and changed the wording to: Apart from four identified cyanobacteria that possess potential EDDs. In addition, we also added the names of the four cyanobacteria that were found including the sequence IDs of the putative EDDs.

      (19) Lines 529/530: revise sentence and format.

      We added one sentence and revised the following sentence: In contrast to Synechocystis EDA, EDA from Synechococcus prefers OAA over KDPG. The catalytic efficiency of Synechococcus EDA on oxaloacetate is rather low (OAA 0.437 s<sup>-1</sup> mM<sup>-1</sup>), however, its activity can be enhanced by NADP<sup>+</sup>(13).

      (20) Line 542: revise sentence.

      We revised the sentence to: It remains to be investigated whether this reaction could play a role in vivo, with KDPG potentially acting as a regulatory metabolite at low concentrations.

      (21) Line 577: The glass tubes used for cultivation should be specified.

      We now added the following information: Custom-made glass tubes were placed in a photobioreactor (manufactured by Willi Hilke, Uslar, Germany).

      (22) Lines 584-585: unclear, was the resistance cassette not placed in the gene to be deleted?

      Yes, the resistance cassette was placed in the gene to be deleted and was fused for homologous recombination to two DNA fragments approximately 200 bp directly upstream and downstream of the gene. This information was now added.

      (23) Line 607: cultivation equipment to be specified.

      The following information was now added: For the purification of GDH1 from Synechocystis, a 6 L photoautotrophic culture of the P3-His-GDH1 overexpression strain was grown in a 10 L glass flask at 28°C, illuminated with constant light (50 µmol m<sup>-2</sup> s<sup>-1</sup>) and gassed with filter sterilized ambient air to an OD<sub>750</sub> of about 1.

      (24) Line 670: GTS should be specified.

      Thank you for this hint. This was a typo. GST was meant not GTS. This was now corrected and GST was specified as Glutathione S-Transferase.

      (25) Line 679: delete "gluconate kinase (GK) and".

      The second gluconate kinase (GK) was deleted and sentence was revised to:

      For gluconate kinase (GK) activity measurements in Synechocystis crude cell extracts the GK reaction was enzymatically coupled to 6-phosphogluconate dehydrogenase (GND) reaction, the latter providing NADP<sup>+</sup> reduction, which was monitored photometrically at 340 nm.

      (26) Line 686: again GK activity? Difference unclear. Was the previously described procedure for GND activity determination?

      GK activity measurements in Synechocystis crude cell extracts and GK activity measurements with recombinant enzyme that was expressed in E. coli were done in two different labs with different protocols. Therefore, the first description refers to measurements with Synechocystis while the second measurement refers to measurements with E.coli. This is now specified more clearly.

      (27) Type/supplier of spectrophotometers and centrifuges used should be given.

      Model and brand were now added for the double-beam spectrophotometer (Uvikon 810, Kontron, Augsburg, Germany). As this study was performed in two different labs over a period of 8 years including one lab moving to a new location, it is now impossible to specify all centrifuges that were utilized. Even though we agree that it would be good to provide this information, we now would have difficulties to be specific.

      (28) Consider the referencing of published methods to streamline the materials and methods section.

      We now streamlined the materials and methods section by deleting repetitions as outlined below. However, as protein expression, protein purification and enzymatic tests were performed in different labs, in some cases several protocols are given.

      (29) Line 757: give specifics of anti-rabbit antibody and define PBS-T and PBS-T Cytiva.

      Specifics were added to the text.

      (30) Lines 761ff: It is not given for all genes used how they were derived. All synthesized?

      We now added detailed information for all genes.

      (31) Line 762: codon-optimized for? E. coli?

      The information was added that genes that were expressed in E. coli were codon-optimized for E. coli.

      (32) Lines 782-786: Rationals for experimental strategies do not belong to materials and methods sections, but to results sections.

      The part was deleted in the materials and methods section and transferred to the results section.

      (33) Lines 818/819: repetitive.

      We removed the repetition and refer to the purification method as stated above in the materials and methods section.

      (34) Lines 847-851: True for all EDA-type assays? Kinetic parameters are shown in Table 2 rather than Table 1.

      Yes, true for all EDA-type assays. We changed the Table number to 2.

      (35) Line 863: delete "in".

      “in” was deleted

      (36) Lines 888-893: sounds repetitive.

      The lines were modified accordingly.

      (37) Lines 908/909: repetitive.

      The repetitive comment on the definition of 1U was deleted.

      (38) Lines 928-938: repetition

      The repetition was deleted.

      References

      (1) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in <em> Synechocystis </em> PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (2) X. Chen et al., The Entner–Doudoroff pathway is an overlooked glycolytic route in cyanobacteria and plants. Proceedings of the National Academy of Sciences 113, 5441-5446 (2016).

      (3) D. Schulze et al., GC/MS-based 13C metabolic flux analysis resolves the parallel and cyclic photomixotrophic metabolism of Synechocystis sp. PCC 6803 and selected deletion mutants including the Entner-Doudoroff and phosphoketolase pathways. Microbial Cell Factories 21, 69 (2022).

      (4) A. Makowka et al., Glycolytic Shunts Replenish the Calvin–Benson–Bassham Cycle as Anaplerotic Reactions in Cyanobacteria. Molecular Plant 13, 471-482 (2020).

      (5) V. Zaitsev et al., Insights into the Substrate Specificity of Archaeal Entner–Doudoroff Aldolases: The Structures of Picrophilus torridus 2-Keto-3-deoxygluconate Aldolase and Sulfolobus solfataricus 2-Keto-3-deoxy-6-phosphogluconate Aldolase in Complex with 2-Keto-3-deoxy-6-phosphogluconate. Biochemistry 57, 3797-3806 (2018).

      (6) J. S. Griffiths et al., Cloning, isolation and characterization of the Thermotoga maritima KDPG aldolase. Bioorg Med Chem 10, 545-550 (2002).

      (7) S. E. Evans et al., Plastid ancestors lacked a complete Entner-Doudoroff pathway, limiting plants to glycolysis and the pentose phosphate pathway. Nature Communications 15, 1102 (2024).

      (8) J. Moreno-Cabezuelo, G. Gómez-Baena, J. Díez, J. M. García-Fernández, Integrated Proteomic and Metabolomic Analyses Show Differential Effects of Glucose Availability in Marine Synechococcus and Prochlorococcus. Microbiol Spectr 11, e0327522 (2023).

      (9) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in Synechocystis sp. PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (10) R. Y. Stanier, R. Kunisawa, M. Mandel, G. Cohen-Bazire, Purification and properties of unicellular blue-green algae (order Chroococcales). Bacteriol Rev 35, 171-205 (1971).

      (11) D. G. Gibson et al., Enzymatic assembly of DNA molecules up to several hundred kilobases. Nature Methods 6, 343-345 (2009).

      (12) M. Boehm et al., Comprehensive study on ferredoxin isoforms in the cyanobacterium Synechocystis sp. PCC 6803. bioRxiv 10.64898/2026.04.08.717189, 2026.2004.2008.717189 (2026).

      (13) N. Xie, C. Sharma, K. Rusche, X. Wang, Phosphoketolase and KDPG aldolase metabolisms modulate photosynthetic carbon yield in cyanobacteria. The Plant cell 10.1093/plcell/koae291 (2024).

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      We provide below a point-by-point response to reviewers’ comments, describing changes that were made to the manuscript. We also uploaded a "Full Revision" file that contains a general statement in addition to these responses to reviewers. We sincerely thank both reviewers for their thorough reading and suggestions that we believe contributed to a better concision and clarity of the revised manuscript.

      Reviewer #1

      Major Comments

      1. TDH3 has been a major model in the study of the evolution of gene expression as developed by the authors. While most conclusions derived from this system relate to genetic mechanisms, this study attempted to identify the molecular/cellular mechanisms. Although I greatly appreciate the author's comprehensive efforts, I would suggest that conclusions regarding molecular/cellular mechanisms should be made with greater caution, especially avoiding over-generalized conclusions that may be specific to P_TDH3. Thus, I suggest going over the manuscript again and adjusting some statements if they are over-generalizing, as well as adding an explicit discussion of this limitation.

      As suggested by the reviewer, we went over the manuscript and made sure that we specifically mentioned the TDH3 promoter in conclusions that may be specific to this promoter. In addition, we further addressed this limitation by including new results of an experiment where we measured the effects of yme2, chs1 and msh1 mutations on expression noise of three additional yeast promoters (PFBA1, PACT1, PTEF1). We found that the three mutations had similar effects on expression noise driven by PTDH3, PACT1 and PTEF1, but different or no effect on PFBA1 expression noise. Therefore, some of our conclusions may not be specific to the TDH3 promoter, such as the importance of mitochondrial state in regulating expression noise of nuclear genes. We described these findings in Results, in the new Figure 7 and in associated supplements. We also added a paragraph in the Discussion and we included new co-authors who performed these additional experiments. We believe these new results will increase the impact of the study.

      1. On the surface, it may seem reasonable to ask for changed noise by unchanged mean in order to distinguish independent regulators. However, from a mechanistic perspective, this is quite demanding. In the current prevailing model (transcriptional burst) of noise origination, mutations are not assumed to be noise-specific without affecting the mean. See e.g. PMID: 31320634. If noise and mean are intrinsically coupled, looking for noise-only regulatory mechanisms would imply something very different. This may mean, for example, that we are seeking one single mutation that changes noise and mean through one mechanism, while at the same time reverting mean to its wildtype value through another mechanism. It is likely that this is one of the reasons why causal mutations are difficult to identify.

      This is a very sensible comment. We agree that in the model of noise originating from transcriptional bursts, cis-acting mutations are not expected to alter noise without affecting the mean. However, trans-acting mutations may affect transcriptional bursts of a given gene via different mechanisms that compensate each others at the level of mean expression but not at the level of expression noise. Even though mutations impacting noise without altering mean expression may be less common than mutations affecting both mean and noise, they do exist as we showed here. It is possible that we did not identify mutations in transcription factors known to regulate TDH3 promoter activity because such mutations would affect both mean expression and noise. We added two sentences in the discussion to mention this hypothesis. In addition, while the transcriptional burst model can explain the coupling between intrinsic noise and mean expression, it does not apply to mutations impacting extrinsic noise.

      Regarding the difficulty to identify mutations altering only expression noise, we found a causal mutations in three out of five mutants analyzed. Interestingly, the three successes corresponded to mutations that increase noise, and the two failures to mutations that decrease noise, suggesting that mutations decreasing noise may be particularly difficult to identify. Finally we note that expression noise is a complex genetic trait in natural populations. An interesting case is shown in Fehrmann et al. 2013, where three trans-acting alleles (loci on chr7, 8 & 13) had a strong individual contribution on noise, partially coupled to mean changes. When combined together, their cumulative effect was very strong on both mean and noise, although the wild strain from which these alleles originate showed an elevated noise but no remarkable change in mean as compared to a reference strain. This implies that other (unmapped) loci “buffer” expression mean in this wild strain against the action of the three noise-acting loci, a scenario comparable to the one suggested by the reviewer.

      1. L209-214. The authors state that they selected 254 mutants with the largest noise changes from a library of 1,241 strains (L209-214). However, I cannot match this statement when contrasting figure1-figure supplement 1 (254 mutants) and Figure 1a (1,241 strains). Is this caused by experiments conducted in different labs ? Are there any intuitive ways for the authors to show the strains (and their parameter distribution) that produced consistent results across the two batches of data? E.g. using gray/black dots for un-/repeatable strains. Also, are mean expression levels similarly (un-)repeatable compared to noise ?

      Both fluorescence screens were performed in the same laboratory, using the same instrument and following the same protocol. We clarified in the text and figures how the 254 mutants were selected for the secondary screen. In particular, we colored dots on Figure 1 and on Figure 1 – figure supplement 1 to highlight strains with reproducible or non-reproducible change in expression noise in both assays. The apparent lack of reproducibility between the first screen of 1241 strains and the second screen of 254 mutants is not surprising because the 254 strains were not picked randomly among the 1241 initial strains: they were picked because they showed the largest expression changes (either for mean or noise) in the first screen. Statistically, the most extreme effects on mean expression and expression noise observed among the 1241 strains are expected to be over-estimated on average. This phenomena is similar to the “Winner’s curse” in economy. When measuring a quantitative trait for a large number of samples with a certain degree of uncertainty, values that fall in the tails of the distribution (the most extreme values) are statistically expected to be less accurately estimated than values falling near the mean of all samples. To address the last question of the reviewer, mean expression levels were found to show better repeatability than expression noise, probably because error bars (variation among replicate samples) tended to be much smaller (one order of magnitude) for mean expression than for expression noise.

      1. How were the five strains analyzed chosen? Are they the only strains fulfilling the criteria on L211-214?

      The five strains were picked arbitrarily among nine strains that matched the criteria mentioned in the text. We added a sentence to mention this point. We moved to Supplementary File 5 the section describing how the five strains were chosen to make the main text shorter and easier to read.

      1. L243-247. I didn't understand the logic why m2 is included, please elaborate.

      We modified the text to clarify why we picked mutation m2 as a candidate. The logic is that there was no strong statistical evidence to exclude the mutation (because of lower statistical power relative to other mutations).

      1. The equation for extrinsic noise (L899) seem to be slightly differently from that in Fu and Pachter 2016. The product of mean(RFP) and mean(YFP) is multiplied by 2 here, but not in Fu and Pachter 2016.

      We made a typo in the text and corrected it in the revised version. We verified in our R scripts that we used the correct version from Fu and Pachter (with product of mean(RFP) and mean(YFP) not multiplied by 2), which was the case. We are particularly thankful to the reviewer for the thorough proofreading of the manuscript.

      1. The experimental design to exclude noise from partitioning for yme2 is really nice. It would have been great if we had gotten to the bottom of this. (This is not a question so no response is needed)

      No response requested.

      1. The authors demonstrate that a nonsense mutation in CHS1 increases extrinsic noise via impaired chitin septum reparation in daughter cells. However, glucosamine treatment itself alters cell size (Figure 5-figure supplement 3a), which correlates with noise levels. This raises the question: do the observed changes in extrinsic noise stem from glucosamine-induced changes in cell size or from the impaired chitin repair caused by the CHS1 mutation itself? To disentangle these effects, an alternative approach to modulating chitin synthesis that does not alter cell size should be employed.

      We do not think that glucosamine-induced changes in cell size can explain the effect of glucosamine treatment on extrinsic noise in chs1 mutant. We observed that glucosamine treatment had a stronger impact on cell size in WT cells than in chs1(G1752a) mutant cells (Figure 5 – figure supplement 3a). However, glucosamine treatment had a much stronger impact on extrinsic noise of chs1(G1752a) mutant cells than WT cells (Figure 5g). Therefore, there is no direct correspondence between the effect of glucosamine on cell size and the effect of glucosamine on extrinsic noise. Even though glucosamine drastically reduced cell size in WT cells, it had almost no impact on expression noise in these cells. We added a sentence in the revised Results to clarify this point.

      Our hypothesis is that glucosamine increases chitin synthesis not only during repair of the chitin septum, but more globally at all stages of the cell cycle (as shown by Bulik et al., 2003), which may reduce cell size. The global impact of glucosamine on cell wall chitin levels could rescue defects caused by chs1 mutation on chitin septum repair. Previous studies showed that CHS1 was not involved in global chitin synthesis, but only in the repair of chitin septum in daughter cells.

      1. Why did glucosamine doses not significantly impact cell growth rates during the first phase after addition (Figure5 -figure supplement 4) ? Additionally, I can seem to find the experimental details for glucosamine dosing in the Methods.

      We specified the dose of glucosamine in the revised Methods. We did not observe a significant impact of glucosamine on growth rates during exponential growth either in the first growth phase or in the second growth phase after addition. However, glucosamine increased the duration of the lag phase in the second phase of growth. We do not know exactly why, but we speculate it is because the chitin cell wall becomes thicker after diauxic shift in presence of glucosamine, leading to a delay to resume cell division after cells are exposed to fresh medium with glucose.

      1. Figure 2-figure supplement 2g-l are missing.

      We included the missing panels in the revised figure.

      1. The current manuscript is a bit lengthy (although nicely comprehensive). After deciding the journal, I suggest it would need to be more concised and logically streamlined.

      We agree that the main text is lengthy, with methodological explanations sometime disrupting the main message. For this reason, we included in the revised version a new Supplementary File 5 where we moved these explanations that were important yet not essential for the reader to understand the main conclusions.

      Minor points

      1. P12,L345, "may not only by caused by" should be "may not only be caused by", ,and "YFP an RFP" should be "YFP and RFP" in the same sentence

      We corrected these mistakes.

      1. P19,L565, "sensitivite" is misspelled and should be "sensitive".P21,L630, "mitochondria dysfunction" should be "mitochondrial dysfunction."

      We corrected these mistakes.

      1. Typo in Figure 3's legend "** 0.001 > P {greater than or equal to} 0.001", which should read "0.01 > P {greater than or equal to} 0.001." This error appears again in Figure 4's legend.

      We corrected these mistakes.

      1. P9, L219 "Table 1" should be "Supplementary File 1"?

      We added the number of mutations per strain in Table 1.

      1. Inconsistent tetrad numbers: methods state 22 tetrads (L844) , results mention 21 tetrads ( L262 ) , and figure( figure1 -figure supplement 3) legends indicate 20 tetrads. Please clarify the correct number.

      Thank to the reviewer for mentioning this inconsistency. In fact, we dissected 22 tetrads but only included 21 tetrads that showed the expected segregation of all genetic markers in the fluorescence assay. Finally, we reported fluorescence measurements for 20 tetrads due to a possible contamination for the remaining tetrad. We clarified this in the Methods.

      1. The speculated retrograde response pathway is interesting. Can the authors propose some specific experiments to test that ?

      To test the involvement of the retrograde pathway in PTDH3 intrinsic noise, one could mutate negative or positive regulators of the retrograde signaling in wild-type cells or in chs1 and msh1 mutant cells and quantify the effect on intrinsic noise. We proposed this experiment in the revised discussion. In previous studies, null alleles of rtg1, rtg2 or rtg3 were shown to impair the retrograde response, while specific mutations in rtg2 and deletion of mks1 were shown to activate the retrograde pathway (Garrigos et al., 2024; Jazwinski and Krete, 2012). We expect to observe an elevated intrinsic noise in wild-type cells, but not necessarily in yme2 and msh1 mutant cells, when we activate the retrograde signaling. Conversely, we expect yme2 and msh1 mutations to not alter intrinsic noise anymore when the retrograde pathway activity is impaired by mutation.

      Reviewer #2

      The manuscript by Martin et al. titled 'Trans-acting mutations reveal non-nuclear modulators of both intrinsic and extrinsic gene expression noise in a eukaryote.' identifies genetic mutations in yeast that can has regulate gene expression noise in trans. The manuscript is well written, and the experiments have been performed in replicates. The authors also clearly highlight the experiments where the replicates do not agree in their outcomes.

      However, there are some issues that the authors need to address:

      1. Introduction is too long and needs to be concise

      We have shorten the introduction in the revised version. We have also moved parts of the main text in Supplementary File 5 to be more concise.

      1. Lines 150-153: Do we have enough studies yet for generalizations?

      We do not know other studies/examples that compared the effects of cis-acting and trans-acting mutations on mean expression and expression noise of a target gene. Therefore, we cannot generalize the results obtained for the TDH3 promoter. However, these results show that cis- and trans-acting mutations can significantly differ in their effects on expression noise (but we do not know for how many genes it is the case).

      1. Line 209 - Why 254 strains? Please justify

      We clarified why and how we chose these 254 strains for the secondary screen. We also added colors on Figure 1 to highlight these 254 strains.

      1. Why are the authors choosing median expression and not mean expression (which is usually the norm)?

      We used the median to quantify the average expression among cells as we did in previous studies with the same fluorescent reporter system because median is more robust than mean to rare outliers. However, we found the difference between median and mean fluorescence to be really small. We included a new figure (Figure 2 – figure supplement 4) showing that the effect of yme2, chs1 and msh1 mutations on expression noise were almost identical when using mean and median to calculate the noise. This is because we measured fluorescence from large number of cells for each sample (~5000) and because the distributions of fluorescence levels among cells are always unimodal with very rare outliers (as showed in Figure 3 – figure supplement 1; Figure 5 – figure supplement 1 and Figure 6 – figure supplement 1).

      1. Do EMS mutants have intra-population genetic heterogeneity? This should be discussed in the text.

      We sequenced the genome of the 5 EMS mutants at a coverage of ~100x, but did not find evidence of genetic heterogeneity in these strains: all mutations detected were found at a frequency near 1. In another project, we sequenced the genomes of 288 EMS mutants and detected genetic heterogeneity in 3 strains: mutations were not fixed in these strains, but found at a frequency near 0.75. We therefore expect the number of mutants with genetic heterogeneity from the collection analyzed in figure 1 to be very small. In addition, genetic heterogeneity cannot impact our conclusions because no genetic heterogeneity was detected in the 5 EMS mutants analyzed and because we constructed two independent clones to investigate the effects of each mapped mutation. We added a sentence in the main text mentioning that no genetic heterogeneity was detected in the 5 EMS mutants.

      1. Line 238-240: Shouldn't the change in frequency in low, mid and high- subpopulations be tested relative to the expected distribution from the wild-type strain? This could also alter how mutations are chosen for validation. This should be mentioned in the results section and the text should be rephrased to reflect this point, although it is mentioned in the methods section

      We are not completely sure to understand what statistical test the reviewer has in mind. We could not easily compare the observed frequency of mutant and wild-type alleles in low, mid and high subpopulations to expected frequencies, because these frequencies depend not only on the effect of the mutation on fluorescence among cells but also on the effect of the mutation on growth rate (which is unknown). Our strategy to compare mutation frequency in medium bulk vs low and high bulk was designed to detect mutations changing expression noise independently from their potential effect on mean expression or growth rate.

      1. Line 243: The mutant name YPW2162 suddenly appears in the text - where did this strain come from?

      This is the name of one of five mutants analyzed, as mentioned in Table 2 referenced in the same sentence. We modified the sentence to make it clearer: “For a fourth mutant (YPW2162), ...”

      1. Line 286: 'increase' instead of 'increased' .

      We corrected this error.

      1. Could genomic rearrangements/copy number variation alter expression noise? For example, for m4, m5 and m6 mutants. The authors have genome data of these strains, so this can be checked.

      According to the reviewer’s comment, we performed additional analyses showing that CNVs and rearrangements did not contribute to variation of expression noise in the five EMS mutants included in the mapping experiments. To detect large CNVs and aneuploidies, we computed sequencing depth in 1-kb sliding windows along the genome for each mutant. The profiles were uniform and similar to coverage profiles obtained for the reference strain. To detect rearrangements, we analyzed sequencing data using the GRIDSS module that can detect junctions between non-contiguous parts of the genome from the mapping location of paired-end reads. Using this tool, we only detected 5 rearrangements that were previously known to be present in the genome of all mutant strains relative to the reference genome (deletions at ho and ura3 loci and duplications of TDH3 promoter, CYC1 terminator as well as 41 bases from chromosome I in the PTDH3-YFP transgene inserted at the ho locus). None of these rearrangements can explain variation of expression noise among mutants. Results from GRIDSS analysis are included in Supplementary File 4 and reported in the main text.

      1. Figure 2 - y-axis: What is the measure of expression noise used here? This should be mentioned in the figure captions throughout to avoid confusion.

      We used the same measure of expression noise for all figures, as mentioned in the Methods. We added it in the figure captions as suggested by the reviewer to avoid confusion.

      1. Line 359 - please mention the effect size here and wherever possible throughout the manuscript

      In the sentence mentioned by the reviewer, we used the forward scatter signal (FSC.A) as a relative measure of cell size. A difference of FSC.A between two samples is known to reflect a difference of cell size. However, we cannot estimate the effect size on cell size because the relationship between FSC.A and cell size depends on the instrument and settings, and we have not characterized this relationship empirically. Therefore, we removed “a small effect” from the sentence and we instead only mentioned that the effect on cell size was statistically significant. Indeed, we cannot be sure that the small (and significant) reduction of FSC.A we observed corresponded to a small reduction of cell size. 12. Figure 6 - figure supplement 2 - Please mention the strains represented by grey and orange boxes

      To make it more visible, we moved this information from an inset in panel b to the top of the figure.

      1. One could envisage that there are many more genetic regulators of expression noise which may have not been discovered yet. This point perhaps could be discussed.

      Absolutely. We added a sentence in the discussion to acknowledge that many genetic modulators of noise may still remain unknown.

      1. The figure captions are too long - they should be made concise

      We reduced the length of the longest figure legends. In particular, some of the text from Figure 4 legend was moved to Supplementary File 5.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank you for the time you took to review our work and for your feedback! We have performed the following additional analyses:

      (1) We analyzed the optomotor mismatch response as a function of time spent in the optogenetic closed-loop session.

      (2) We show the correlation between optomotor mismatch response and visuomotor mismatch response for functionally identified PE neurons.

      The main changes to the manuscript are:

      (3) Clarification of our terminology (moved and refined definition of the teaching signals).

      (4) A new figure panel summarizing the functional influence patterns we identified.

      (5) Clarification of our argumentation for the JEPA-inspired proposal.

      All comments are addressed individually in the following.

      Public Reviews:

      Reviewer #1 (Public review):

      Vasilevskaya and Keller test different models of cortical function through the lens of predictive processing, a powerful framework for the brain to learn and predict the statistics of the world via generative internal models. The authors use a clever combination of behavioral perturbations in closedloop and open-loop visuomotor virtual reality assays, a paradigm the Keller lab pioneered and used effectively in the past decade, in conjunction with two-photon imaging of neuronal calcium responses and targeted optogenetic perturbations of activity. They specifically put to test proposed hierarchical vs. non-hierarchical circuit implementations of predictive processing by analyzing the logic of inter-lamina interactions (superficial vs. deep; L2/3 vs. L5/6).

      The authors conclude that both versions of predictive processing architectures they analyze are likely invalid, and instead formulate an alternative novel model of cortical function based on a recently developed machine learning algorithm for self-supervised learning (joint embeddings of predictive architectures, JEPA) and its further refinements. JEPA borrows elements from predictive processing, engaging two encoder networks and training the output of one network to predict the output of the other. In their new model of cortical computations, prediction error neurons in L2/3 compare the deep layers (L5/6) activity, which is taken as a teaching signal, to a local, L2/3 prediction of this latent representation.

      Specifically, the authors build on their previous work and reports from other groups that different sets of L2/3 neurons compute positive prediction errors (fire when sensory stimuli appear unexpectedly with respect to the movements of the animal; e.g., grating onsets in the absence of locomotion) and respectively negative prediction errors (fire when sensory stimuli are absent, while the brain expected them to be present; e.g. mice locomote but visual flow is suddenly halted - visuomotor mismatches). These L2/3 positive and negative prediction error neurons exchange messages with neurons in the deeper cortical layers that, the authors propose, build an internal representation (R) of the sensory stimuli given the animals' movements.

      In the hierarchical model, internal representation neurons (R) are supposed to act as a teaching signal for both types of prediction error neurons; the output of the positive prediction error neurons is assumed to suppress activity of R such that the error between the teaching signal and the prediction is minimized; similarly, in the non-hierarchical version, R serves as a prediction for the prediction error neurons, and in turn it receives excitatory drive from the positive prediction error neurons and negative input from the negative prediction error neurons.

      The authors find that the functional impact of L5 neurons on L2/3 neurons is not compatible with the non-hierarchical architecture they and other groups proposed, but rather in accordance with the hierarchical model. At the same time, the functional impact of L2/3 neurons (positive vs. negative prediction error neurons) on L5 neurons (internal representation) appears not compatible with the hierarchical model, but rather in accordance with the non-hierarchical implementation.

      They further hypothesize that L2/3 prediction error neurons don't use sensory input, but rather the L5 activity as a teaching signal, and test it using perturbations (halts) of optogenetic stimulation of L5 neurons coupled with locomotion (Figure 7).

      All in all, the question is topical, and the new model addresses a decades-long quest to develop a unifying model of cortical function. The findings reported here transform our understanding of cortical computations, opening new, exciting avenues for future investigation. The experimental design and execution are rigorous; the arguments are clearly laid out (in spite of ample potential for confusion given the numerous loops and sign flips). These include a discussion of why the non-hierarchical model proposed by the same group does not hold, as well as potential caveats in interpreting the results and novel testable proposed experiments emerging from the JEPA-like model.

      I have several questions about the interpretations of some of the claims and suggestions for potential additional experiments and analyses.

      We thank the reviewer for their comments. We address them below.

      (1) Some of the pieces of the puzzle remain to be identified and demonstrated: the existence of internal representation neurons in L2/3 and ascertaining that the L5/6 neurons analyzed function indeed as internal representation neurons. The authors find that stimulation of L2/3 positive prediction error neurons enhances activity of L5 neurons...If L5 neurons hold a latent representation that serves as a teaching signal for L2/3 neurons (as the authors posit), wouldn't one expect that the input they receive from the positive prediction neurons be suppressive, such that the error is further minimized?

      Not necessarily - this depends on the model one has for how cortex works. In the hierarchical predictive processing model, PE+ neurons are expected to suppress the local internal representation neurons in L5. Our data, however, are not consistent with this model. This is one of the key arguments we build the idea on that JEPA is a better model for cortex than hierarchical predictive processing. In JEPA we would not expect the L2/3 prediction error to update L5 directly, but instead update the prediction of L5 activity (the source of this signal remains to be identified; see our speculations on the origin of prediction signals in the comments to question 3), and drive plasticity in the local L5 encoder network. That something acts like a teaching signal for the L2/3 comparator does not, by itself, determine how the outcome of the comparison influences the source of the teaching signal.

      (2) Do the authors envision any specific differences between the representations of the two encoder networks posited to exist in L2/3 and L5 in the JEPA-like implementation? Are they synchronous/offset in their temporal representations, or any other features?

      Given theoretical work (Mohammadi et al., 2025), one would expect to find differences in learning rates between the predictor and the encoder networks. Assuming the predictor network is also implemented in L2/3, we would expect to see faster learning rates in L2/3 compared to L5. Implementations inspired by related theoretical works also predict differences in learning rate between the two encoder networks (Grill et al., 2020), again with higher learning rates expected in L2/3 compared to L5. Beyond that, however, we are not aware of any experimentally observable differences one might expect to find. We are hoping the computational community will remedy this soon.

      (3) Where is the prediction coming from onto L2/3 neurons? Is it emerging locally in L2/3 from the putative internal representation neurons, or is it long-range - as work from the authors previously proposed? Or a mix of both?

      We expect the predictions to come from long-range inputs. In classical (hierarchical) JEPA one would expect these to be the lateral communication within L2/3. In a non-hierarchical implementation that is capable of operating on arbitrary graphs (as would be necessary for it to work in cortex), we suspect that the L2/3 network will use both long-range L2/3 and long-range L5 input. In JEPA terminology: the encoder A network uses information from non-local sources of both networks to predict local activity of encoder B network (assuming that information is statistically useful in predicting that activity).

      (4) What is the role of the indiscriminate L4 input that appears to enhance activity of both positive and negative prediction error neurons in L2/3?

      The short answer is, we don’t know. We would have indeed expected to find some asymmetry of influence. There are a few options: A) We might be failing to activate a specific interneuron that mediates feedforward inhibition on this pathway by the artificial stimulation of Scnn1a neurons. B) We know that Scnn1a neurons are only a subset of L4 neurons – there are other populations of L4 neurons that might exhibit the opposing influence. C) The Scnn1a population is more than 1 layer of a JEPA away from the comparator and not part of the predictor that forms the actual representation compared against L5 (e.g. L4 could provide an input to L2/3 internal representation neurons, and thus only indirectly influence L2/3 PE neurons). D) The JEPA analogy is wrong.

      (5) Does Figure 7D change in a meaningful manner if the authors plot the correlation between optomotor mismatch response and visuomotor mismatch response specifically for the negative prediction error neurons in L2/3 (Adamts-2) rather than for all L2/3 cells sampled?

      We might be misunderstanding. If the reviewer means genetically identified negative prediction error neurons (Adamts2), we do not have the data to address this question, as we did not perform any recordings of molecularly defined Adamts2 population in L2/3. If the reviewer means functionally identified negative prediction neurons, this would be the two rightmost data points in Figure 7D (the x-axis in this panel is the visuomotor mismatch response strength we use to functionally identify PE- neurons). These two data points on the right-hand side of the plot correspond exactly to what we classify as negative prediction error neurons throughout the rest of the manuscript (15% of the most responsive neurons to visuomotor mismatch).

      Does the reviewer mean, is there also a positive correlation between optomotor mismatch response and visuomotor mismatch response when looking only at neurons that we identify as PE-? If so, the answer is yes (Author response image 1). Interestingly, while there is a positive correlation, there is also nonuniformity in response patterns. We think that this is expected, given that our visuomotor coupling paradigm captures only a very small subspace of stimuli that prediction error neurons are tuned to. Bulk stimulation of L5, in contrast, might work better to separate all PE neurons. In other words, we speculate that L5 stimulation is a better predictor of the functional role of an L2/3 neuron.

      Author response image 1.

      Optomotor mismatch response as a function of visuomotor mismatch response for neurons that are functionally classified as PE. Red line shows a linear fit estimated with a bootstrap approach.

      (6) Do the optomotor mismatch responses in L2/3 neurons depend on how long the closed-loop coupling of optogenetic stimulation of Tlx3 L5 neurons and locomotion speed has been in place for?

      No, not that we can measure. We performed an analysis in which we split the optogenetic closed-loop session into two equal parts, early and late. We then quantified the average optomotor mismatch response independently for early and late parts of the session (Author response image 2A). Based on this quantification, we find no evidence of a change as a function of time in the closed-loop session. We also performed a sliding window analysis on a shorter timescale. While it did look like the responses may be smaller in the first few minutes, we do not have sufficient data to address this. None of the differences in response size were significant (Author response image 2B).

      Author response image 2.

      Optomotor mismatch response as a function of experience with artificial closed-loop coupling. (A) Mean L2/3 population response to optomotor mismatch in the first half (Early MM) and in the second half (Late MM) of the optogenetic closed-loop session. (B) Mean L2/3 population response to optomotor mismatch as a function of time in the optogenetic closed-loop session. Error bars indicate SEM. Differences between the mean values were estimated by hierarchical bootstrap and are not significant.

      Reviewer #2 (Public review):

      This manuscript reveals the functional connectivity of two different classes of cortical neurons that respond in opposite ways to mismatches between sensory and top-down inputs. These data are very valuable because different theories of information processing in the cortex make different predictions on the patterns of connectivity of these neurons. Therefore, these data strongly constrain possible theories of cortical processing.

      We thank the reviewer for their comments. We address them below.

      General comments:

      (1) The methods of statistical testing are insufficiently described. I did not understand the description in lines 1105-1119. The authors should provide sufficient details so the reader can reproduce their analyses. For example, it may be helpful to provide specific details of the testing procedure for one of the comparisons (e.g. the first comparison in Table S1).

      We assume the reviewer is not familiar with hierarchical bootstrapping in general. If so, the explanation below would summarize the procedure. This is the procedure with particular emphasis on its application to neuroscience data is described in the paper we reference in that part of the methods (Saravanan et al., 2020). Given that the analysis has become relatively standard (and is described in the reference provided), we think it might be an overkill to add the full explanation below to the manuscript. In addition to the general procedure, there are only 2 pieces of information relevant to fully reconstructing the analysis:

      (1) What are the “levels” (mice, recording sites, neurons, trials)?

      (2) What is the number of bootstrap samples used.

      Thus, we think all information is already provided in the methods. We now also explicitly added the levels when describing the nested structure of the data in the manuscript to increase clarity.

      Hierarchical bootstrap analysis:

      Our data are naturally nested: multiple neurons are recorded within a single mouse, and multiple mice are tested within an experimental group. Standard bootstrapping (sampling with replacement from the entire pool of neurons) fails because it assumes all observations are independent. In reality, neurons from the same mouse are more similar to each other than to neurons from a different mouse. The hierarchical bootstrap (or multi-level bootstrap) preserves this nested structure, ensuring your confidence intervals are not artificially narrow due to pseudoreplication.

      To illustrate the problem, assume you have 100 neurons from Mouse A and 10 neurons from Mouse B, a simple bootstrap will be heavily biased toward Mouse A. Furthermore, the simple bootstrap ignores the fact that the true variance in your population comes from two sources:

      (1) Between-mouse variance (differences in surgery, genetics, or behavior).

      (2) Within-mouse variance (differences in tuning or activity between individual cells).

      Hierarchical bootstrap addresses this problem, and is implemented as follows: To estimate the mean response while accounting for different sample sizes per mouse, a two-level resampling scheme is used

      (1) Resample the higher level (Mice)

      First, you account for the variability between animals.

      Suppose you have N mice in total.

      Randomly draw N mice with replacement from your original pool.

      Note: Because this is with replacement, a single mouse’s data might be included multiple times in one bootstrap iteration, while another mouse might be left out entirely.

      (2) Resample the lower level (Neurons)

      For each mouse selected in Step 1, you must now account for the variability within that specific animal.

      Look at the number of neurons actually recorded from that mouse (let’s call it k<sub>i</sub>).

      Randomly draw k<sub>i</sub> neurons with replacement from that mouse’s specific pool of recorded cells.

      This step is crucial: you always resample the same number of neurons that were originally recorded for that specific mouse. This maintains the "weight" or "influence" that animal had in the original dataset.

      (3) Calculate the resampled statistic

      Calculate the mean of all neurons collected in this “bootstrap sample”.

      (4) Iterate

      Repeat Steps 1–3 many times (typically B = 1,000 or 10,000 iterations).

      The distribution of these B bootstrap means represents your sampling distribution and is used to calculate confidence intervals and p-values. 

      (2) The authors should clarify how the problem of multiple comparisons was addressed for comparisons performed in multiple moments of time, where significance is indicated by a black bar (e.g. in Figure 2F).

      There is no family-wise error correction in cases of comparing response time courses implemented in our analysis. If the reviewer has a good suggestion for how to implement family wise error correction, we would be happy to implement it. We are not aware of anything that is better than what we currently do (the time bin-wise comparison). To briefly explain the problem: In most neuroscience papers, response curves are compared by choosing a time window (e.g. 0.5 to 1 s following a trigger) and calculating mean values of the curves in these windows. This hides the problem of multiple comparisons that arises from the fact that the experimenter is free to choose a response window used for analysis. Note, this also creates a strong incentive to “optimize” choice of an analysis window – a part of the analysis that a reader is typically completely blind to. We could of course also choose an analysis window that sounds reasonable and yields significant differences for our analyses. However, to provide a more unbiased image of the data, we have come to do bin-wise comparisons with a fixed p-value (typically 0.05). This is used in all of our papers at the moment. Given that samples from neighboring timepoints are correlated via a combination of actual responses and a subset of noise sources, the samples are not independent. We now also implemented a correction for spurious positive values by requiring at least 2 neighboring bins to have a p-value below 0.05 to be shown. If the samples were independent, this would mean a false positive rate of 0.0025. Given that they are not, this is a lower bound only. Additionally, any family-wise error correction would be a function of the number of time bins we show in the plot. This would mean that our choice of the time window shown in a plot (-1s to +4s, or +5s, etc.) would change the p-value we consider significant. Thus, there is no explicit family-wise error correction, and we have come to the conclusion that the bin-wise comparison with a fixed p-value is the most unbiased representation of the data we can provide. The alternative would be to additionally plot z-scores or the p-values as a function of time, but in our experience these types of plots are even harder to read for readers not used to it.

      (3) It would be helpful to add a figure in the Discussion summarising the functional connectivity suggested by all experiments.

      We now added a panel that summarizes the functional connectivity observed in our experiments to Figure S10. 

      (4) Throughout the manuscript, the authors use the term "teaching signals", but I am unclear what they mean by it: after reading the definition in lines 45-46, I thought that they corresponded to values (as they are compared to sensory signals). Later (428-430), the text suggests that they correspond to error neurons. But then lines 605-607 say it is not an error signal. The authors should define teaching signals very precisely or remove this term.

      The formal definition of the teaching signal was in footnote 1 of the manuscript. We assume the reviewer may have missed this. We now moved this definition into the main text to increase clarity.

      We use the term teaching signal to mean exactly this definition throughout the manuscript. We suspect, a second source of confusion may come from ambiguity in regards to anatomical vs. functional definitions. We have attempted to emphasize that the definition of teaching signal is a functional one, not an anatomical one (as is the case for ‘prediction’ – predictions are functionally defined, not anatomically – hence it makes sense to ask questions of the form “what are the potential sources of predictions” etc.). A teaching signal is a signal that is compared against a prediction (the ‘ground truth’ the prediction is compared and trained against). In different circuit implementations of predictive processing, different inputs function as predictions and teaching signals. In the hierarchical implementation, activity of the internal representation neuron at the lower level serves as a teaching signal for a prediction signal that is formed by the internal representation neuron from the higher level. In the non-hierarchical implementation, external inputs from the thalamus or other cortical areas serve as teaching signals for the respective prediction signals that are formed by internal representation neurons. This terminology is most intuitive when thinking from the perspective of a prediction error neuron, since a prediction error neuron computes the difference between two inputs signals – one of which functions as a teaching input, and the other one as a respective prediction. Hence, lines 428-430 specify that layer 5 input onto prediction error neurons of layer 2/3 serves as a teaching input.

      Reviewer #2 (Public review):

      Vasilevskaya and Keller set out to experimentally distinguish between two variants of predictive processing: a hierarchical and a non-hierarchical variant. The hierarchical variant assumes a hierarchical organization in which internal representation neurons (believed to be a subset of layer 5 excitatory neurons) serve as a source of a teaching signal for local prediction error neurons as well as for the next higher level of the hierarchy, while simultaneously providing prediction signals to the preceding lower level. In contrast, the non-hierarchical variant posits that these layer 5 internal representation neurons provide local predictions to layer 2/3 prediction error neurons.

      The interaction between internal representation neurons and prediction error neurons differs fundamentally between the two variants. In the hierarchical variant, internal representation neurons excite positive prediction error neurons and inhibit negative prediction error neurons, while at the same time being inhibited by positive prediction error neurons and excited by negative prediction error neurons. In the non-hierarchical variant, this pattern of connectivity is reversed.

      This work is very exciting, timely, and carefully executed. The authors functionally, and later molecularly, identify layer 2/3 prediction error neurons in V1 and probe their interactions with genetically defined neuron types in cortical layers 5 and 6 using optogenetics. They demonstrate that the functional influence of putative prediction error neurons in layer 2/3 onto layer 5 is incompatible with the hierarchical variant, whereas the influence of layer 5 onto putative prediction error neurons in layer 2/3 is incompatible with the non-hierarchical variant. They then test an alternative hypothesis, in which layer 2/3 responses resemble prediction errors with respect to perturbations of artificial layer 5 activity patterns. To investigate this, they designed an experiment in which optogenetic activation of L5 IT neurons was closed-loop coupled to the mouse's locomotion speed in the absence of visual feedback, allowing them to probe the causal influence of L5 activity on layer 2/3 responses.

      Finally, the authors hypothesize that their data are more consistent with a joint embedding predictive architecture (JEPA) and outline experimentally testable predictions arising from this framework.

      We thank the reviewer for their comments. We address them below.

      While the work is overall convincing and significantly advances our understanding of the circuit-level implementation of predictive processing, there are a few weaknesses that should be addressed or discussed:

      (1) The authors define putative positive prediction error neurons as the 15% of neurons most responsive to grating onset and putative negative prediction error neurons as the 15% most responsive to visuomotor mismatch. While this selection would be expected to overlap with negative and positive prediction error neurons, the criterion is not sufficiently stringent (independent of the exact percentage chosen). In particular, classification of a neuron as a prediction error neuron should ideally be accompanied by evidence that it does not exhibit a significant increase in activity when the prediction matches the sensory input or teaching signal.

      We understand the reviewer’s intuition. We can indeed use other stimuli to identify prediction error neurons, like the relative suppression of responses in closed-loop running onset vs open-loop running onsets. This was the reason behind including Figure S1, to show that our selection results in expected pattern of running onset responses. We don’t typically use the running onset responses as they have an additional confound we have not fully understood. This is that running onset always tends to result in an increase of calcium activity in all neurons. This could have a variety of reasons: A) Contamination of hemodynamic occlusion signals (blood vessels tend to constrict at running onset, making it appear like an increase in calcium activity) – see Yogesh et al., 2025. B) The virtual coupling in our VR is not good enough to provide a true “closed loop” experience. Humans typically notice lags of larger than 30ms – in our VR it is approximately 100 ms. C). Running onset in head-fixed animals is not accompanied by a vestibular input. D) Predictive processing is wrong. We tend to think it is a combination of the three.

      We can also use combinations of the two criteria to select neurons - if the reviewer has a specific selection criteria in mind (top XXX% MM responsive AND top XXX% closed-loop suppressed, etc.) we are happy to repeat the analysis for that specific set of criteria, but the fundamental problem that we are using a functional response to select these neurons does not go away. We know that our functional selection criteria mean we select a population of neurons that is enriched for prediction error neurons. If the enrichment is too weak, we would expect to find no effects in terms of functional influence. It is hard to explain, however, how a weak enrichment could result in a strong effect on functional influence. Our arguments in more lengthy form, for why the visuomotor mismatch is a good stimulus to identify negative prediction error neurons can be found here: Attinger et al., 2017; Jordan and Keller, 2020; Leinweber et al., 2017; O’Toole et al., 2023; Vasilevskaya et al., 2022; Zmarz and Keller, 2016.

      (2) The authors "speculate that the prediction error responses in layer 2/3 may not be computed with respect to sensory input, but with respect to layer 5 activity as a teaching signal." However, it is unclear how this perspective differs from earlier statements in the manuscript. In the Introduction, the authors note that "these signals, typically referred to as sensory signals, we will refer to as teaching signals," and later describe the hierarchical variant as one "in which internal representation neurons act as a source of the teaching signal." Given this framing, it is difficult to identify what is conceptually novel in the updated view. Is the key distinction that layer 2/3 neurons are now proposed to generate predictions in an internal representation space rather than in sensory input space, as briefly suggested in the Discussion? Or are the authors introducing a distinction between an external (sensory) and an internal (cortical) teaching signal? If so, this distinction should be made explicit. Clarifying this point would considerably strengthen the manuscript.

      There might be a misunderstanding regarding our usage of the term teaching signal. In hierarchical predictive processing the teaching signal is typically referred to as a sensory signal, as e.g. in: “prediction error neurons compare predictions to sensory input”. In non-hierarchical predictive processing, or far away from the sensory input (think prefrontal cortex), or for cross-modal interactions “sensory” input is misleading. Also, in non-predictive-processing type models (like JEPA), sensory input has a different functional role. Thus, we operationally define teaching signal as the signal that is compared against the prediction by prediction error neurons.

      The two primary options we are comparing are:

      (1) Is the teaching signal to the L2/3 comparator a bottom-up input to V1 (as one would expect in predictive processing). 

      (2) Is the teaching signal to the L2/3 comparator L5 input (as one would expect in JEPA).

      Our data argue in favor of option 2. We have rephrased parts of the manuscript to try to make this clearer. 

      (3) The authors propose that "L2/3 neurons predict L5 activity, hence making predictions in the internal representation space rather than the input space," and further suggest that, since both deep and superficial cortical layers receive thalamic input, the cortex may function like a JEPA. This idea appears closely related to the model introduced by Nejad et al. (2025), which effectively implements a JEPA-like architecture: L5 activity serves as a target against which L2/3 predictions are compared in a selfsupervised manner, with both L5 and L2/3 (via L4) receiving thalamic input. It would be helpful for the authors to clarify how their framework differs from that model, and to specify the key conceptual or mechanistic distinctions between the present proposal and the approach described by Nejad et al.

      The two proposals indeed share similarities in assuming that bottom-up input for both L2/3 and L5 arrives from thalamus, and that representations formed in L2/3 are used for predicting the activity of L5. However, there are a few important differences between the JEPA implementation proposal formulated here and the Nejad et al. model.

      (1) There is no proposed mapping of computations in the Nejad et al. model onto different JEPA networks. We assume that the suggested mapping would be L4 and L5 as encoder networks, and L2/3 as a predictor network? In that case, it is different to our proposal, in which L2/3 is part of the encoder network.

      (2) Our proposal contains explicit prediction error neuron cell types within L2/3, while prediction errors in Nejad et al. are encoded in the gradients, and the layer origin of these signals is hypothesized to be L5 (‘the learning-driving error signal originates in L5’). Hence, also the role of L5-L2/3 connection is distinct between the two proposals. In Nejad et al. this connection serves the role of error propagation and update for predictions in L2/3, while in our proposal this connection contains teaching signal (target representations) that are compared to predictions within L2/3. Similarly, the functional role of L2/3-L5 connection is also different, since in Nejad et al, it is supposed to carry predictions of L5 activity, whereas in our proposal we expect it to drive plasticity in L5 encoder.

      (3) The difference outlined above also makes it evident that the two proposals should differ in how deep and superficial layers are expected to influence the activity of one another. Indeed, the proposal in Nejad et al. is based on the cortical column idea, and according to eq. 2 and 3 in the Methods, activity in L5 is a function of activity in L2/3, while activity in L2/3 is not a function of activity in L5. Our proposal is based on idea of layers forming parallel networks, where horizontal communication is the dominant mode of cortico-cortical interactions, and activity in deep layers serve as a teaching signal for L2/3. In our case, we expect the opposite - that activity in L2/3 depends on activity of L5, while activity of L5 is not immediately dependent on activity of L2/3 (only via plasticity route). This led us to propose one of direct tests for our framework – silencing L2/3 in a familiar setting should result in no immediate changes to L5 activity and behavior of the animal.

      (4) The proposal in Nejad et al. relies on input reconstruction or variance maximization within the L5 autoencoder network to avoid collapse. Instead, our proposal has no reconstruction objective.

      (5) Lastly, there is a time delay between inputs to L5 and L2/3 that is proposed in Nejad et al., while this is not something inherent to our proposal.

      We expect that the most useful future models should move beyond JEPA, with the emphasis on nonhierarchical models capable of operating on arbitrary graphs. We think cortex functions according to principles of a JEPA (predictions in latent space), and that the role of cell types and the exact computational organization remain to be constrained.

      REFERENCES

      Attinger, A., Wang, B., Keller, G.B., 2017. Visuomotor Coupling Shapes the Functional Development of Mouse Visual Cortex. Cell 169, 1291-1302.e14. https://doi.org/10.1016/j.cell.2017.05.023

      Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P.H., Buchatskaya, E., Doersch, C., Pires, B.A., Guo, Z.D., Azar, M.G., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M., 2020. Bootstrap your own latent: A new approach to self-supervised Learning. https://doi.org/10.48550/arXiv.2006.07733

      Jordan, R., Keller, G.B., 2020. Opposing Influence of Top-down and Bottom-up Input on Excitatory Layer 2/3 Neurons in Mouse Primary Visual Cortex. Neuron 108, 1194-1206.e5. https://doi.org/10.1016/j.neuron.2020.09.024

      Leinweber, M., Ward, D.R., Sobczak, J.M., Attinger, A., Keller, G.B., 2017. A Sensorimotor Circuit in Mouse Cortex for Visual Flow Predictions. Neuron 95, 1420-1432.e5. https://doi.org/10.1016/j.neuron.2017.08.036

      Mohammadi, A.G., Halvagal, M.S., Zenke, F., 2025. Understanding cortical computation through the lens of joint-embedding predictive architectures. https://doi.org/10.1101/2025.11.25.690220

      O’Toole, S.M., Oyibo, H.K., Keller, G.B., 2023. Molecularly targetable cell types in mouse visual cortex have distinguishable prediction error responses. Neuron 111, 2918-2928.e8. https://doi.org/10.1016/j.neuron.2023.08.015

      Saravanan, V., Berman, G.J., Sober, S.J., 2020. Application of the hierarchical bootstrap to multi-level data in neuroscience. Neurons Behav. Data Anal. Theory 3, https://nbdt.scholasticahq.com/article/13927-application-of-the-hierarchical-bootstrap-tomulti-level-data-in-neuroscience.

      Vasilevskaya, A., Widmer, F.C., Keller, G.B., Jordan, R., 2022. Locomotion-induced gain of visual responses cannot explain visuomotor mismatch responses in layer 2/3 of primary visual cortex. https://doi.org/10.1101/2022.02.11.479795

      Yogesh, B., Heindorf, M., Jordan, R., Keller, G.B., 2025. Quantification of the effect of hemodynamic occlusion in two-photon imaging of mouse cortex. eLife 14, RP104914. https://doi.org/10.7554/eLife.104914

      Zmarz, P., Keller, G.B., 2016. Mismatch Receptive Fields in Mouse Visual Cortex. Neuron 92, 766–772. https://doi.org/10.1016/j.neuron.2016.09.057

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the Reviewing Editor for their careful evaluation of our manuscript and for their constructive and insightful comments. Their suggestions have helped us to improve the clarity, rigor, and presentation of our work. In response to these comments, we have substantially revised the manuscript and performed several additional analyses and experiments, as summarized below.

      Major additions and modifications made during revision

      In response to the reviewers' comments, we have substantially revised the manuscript and performed several additional analyses and experiments:

      New analyses

      - Quantification of MyoF recovery following auxin washout using MyoF-mAID-HA immunofluorescence (Figure 8B, Figure S10D).

      - Quantification of maternal MIC2 fluorescence intensity following 24 h MyoF depletion and subsequent redistribution after auxin washout (150 micronemes per condition; Figure S10A,B).

      - Pearson correlation analysis of ANKER1-Halo and HDEL-GFP localization (Pearson's R = 0.92 ± 0.04; n = 30 parasites).

      - Additional probability-based analysis supporting regulated microneme inheritance.

      - Expanded analysis of microneme redistribution across larger replication stages (Figure S5D).

      New figures

      - Figure S6: Dual-labelling analysis of additional Group 2 organelles (ER, apicoplast, glideosome).

      - Figure S10A, B: MIC2 fluorescence intensity analysis following RB retention and redistribution.

      - Figure S10D: Correlation between MyoF recovery and phenotype rescue.

      - Figure S11: Schematic overview of quantification and analysis workflow.

      Additional experimental efforts

      - Generation of a MIC2-Halo / IMC1-mKATE / Cb-Emerald parasite line to improve visualization of RB-associated trafficking.

      - Multiple attempts to perform higher-temporal-resolution live-cell imaging. However, prolonged acquisition resulted in severe phototoxicity, replication arrest, and parasite death, preventing reliable long-term recordings.

      Textual and methodological revisions

      - Expanded Materials and Methods section with detailed descriptions of fluorescence quantification, colocalization analyses, and statistical procedures.

      - Re-evaluation of statistical analyses using two-tailed tests throughout.

      - Revision of manuscript text to clarify the evidence supporting RB-associated trafficking and to better acknowledge current limitations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work asks the question of how different organelles and structures in the apicomplexan parasite Toxoplasma gondii are recycled and/or segregated to the daughter cells during cell replication. In particular, they consider an unusual cell structure called the residual body that links replicating cells during the intracellular infection stage of this parasite. The residual body has historically been considered a 'dumping ground' for unnecessary relics of the mother cell during division, but this notion is increasingly being revised. Indeed, cell replication in Toxoplasma is often misinterpreted as cell division (cytokinesis), but in fact, the cell replicates its organelles and structures to multiple 10s of copies in seemingly distinctly formed daughter cells, but cytokinesis is delayed for many such cycles and typically only occurs simultaneously with parasite egress from its host cell. The residual body is, in fact, the connection between these pre-cytokinetic replicated daughters, and effectively, this is still a single cell at this stage. The authors have previously shown that an actin network extends through the residual body between these daughter cells, and ER and mitochondria common to all cells are also linked through this structure. This study examining the fates of organelles during cell replication is timely for continuing our understanding of how this fascinating component of the cell participates in these processes. The authors use Halo-tags as their principal tool to track discrete populations of proteins, labelling their organelle locations, and this provides beautiful insight into these processes.

      Strengths:

      Using dyes conjugated to Halo tags, this work elegantly tracks the fates of proteins synthesised by an original 'mother' cell over several replication cycles of pre-cytokinetic 'daughters'. Using this tool, they show that some organelles are made intact just once and that some of these can be subsequently sorted to the daughters (micronemes and rhoptries) while others are dismantled (IMC) and the daughters must make their own. A third set of organelles (largely synthesis, sorting, and metabolic compartments) is divided and inherited, and new daughter-synthesised proteins are added to the preexisting maternal proteins in these structures. A role for actin and myosin is clearly demonstrated for micronemes and rhoptries, and this correlates with their relatively late inheritance into the developing daughters. Overall, this work gives clarity to the behaviours of several cell structures during replication and paves the way to a better understanding of the mechanisms that drive the differences between structures and the universality of these processes in other apicomplexan parasites.

      Weaknesses:

      In addressing the question of residual body participation in sorting of organelles, it would be useful to clearly define this structure and when and where it is delineated from the posterior of a mother cell during the formation of daughter structures. This might seem like a moot point, but it would give clarity to notions of recycling and 'reservoirs'. Mother cells retain their active invasion apparatus until very late in daughter formation, and the need for micronemes and rhoptries to be released from this service late in the process might explain why they are only then trafficked to the cell posterior and then into the daughters. So, is this a distinct 'residual body' body function/reservoir or just a spatial constraint of this sequence of daughter formation? In subsequent cell replications (4, 8, 16... stages), is there a separation between the residual body that links them all and the posterior of each new 'mother cell', and if so, when is this distinction lost? This is important because without a definition, we might be confusing different processes.

      We thank the reviewer for this excellent and thoughtful question. The residual body (RB) emerges at the end of the first replication cycle, where it is delineated by the basal complex and persists as an IVN-associated compartment connecting all daughter parasites through both plasma membrane and cytoplasm. Previous EM and live-cell studies, including ours, have shown that the RB is not a passive remnant but a dynamic structure dependent on F-actin and unconventional myosins, supporting recycling, inter-parasite connectivity, and synchronous growth (Delbac et al., 2001; Muñiz-Hernández et al., 2011; Frénal et al., 2017; Periz et al., 2017).

      In the present study, the MyoF reversibility experiment provides strong support for a model in which RB functions as an active recycling hub. Upon MyoF depletion, maternal microneme and rhoptry proteins accumulate within the RB. Following restoration of MyoF expression, this material is redistributed to daughter organelles. We interpret this reversible phenotype as evidence that the RB represents a distinct and regulated trafficking intermediate rather than simply a by-product of late daughter cell formation.

      We agree with the reviewer that mother cells retain a functional invasion apparatus until very late during daughter formation, and that the delayed release of micronemes and rhoptries likely contributes to their late trafficking toward the cell posterior. However, our data indicate that once released, these organelles transit through a defined RB compartment that actively participates in their recycling rather than merely reflecting positional constraints. This has been previously well illustrated for micronemes, which are trafficked along F-actin filaments within the residual body (Periz et al., 2019).

      At later rounds of replication (4, 8, 16 parasites), previous studies have demonstrated the presence of multiple residual body centres within the same vacuole. However, the precise temporal and structural distinction between the RB linking parasites within the vacuole and the posterior of newly formed mother cells remains insufficiently resolved and is beyond the scope of the present study. Importantly, available ultrastructural and live-cell imaging supports the persistence of shared RB compartments connecting parasites within a vacuole, arguing against a simple conflation of posterior membranes and residual body material.

      While the primary aim of the current work was to investigate the RB's role in organelle recycling, we fully agree that a more precise definition of when and how the RB is formed, remodelled, and ultimately resolved during successive replication cycles will be essential to distinguish recycling from spatial constraints. We have revised the Discussion to better acknowledge this limitation and to avoid overinterpreting the role of the RB in organelle inheritance.

      Are rhoptries/micronemes that originate in one 'mother' able to be sorted to the 'daughters' from a distinct mother in this syncytium? If so, this would make it a sorting centre, but otherwise we could be just capturing the activities at the posterior of any given cell during replication. The authors' further thoughts on this would be very interesting.

      We agree with the reviewer that our current data do not definitively demonstrate whether rhoptries or micronemes originating from one “mother” parasite can be redistributed to daughters derived from another mother within the same syncytial vacuole. Nevertheless, our MyoF chase experiments are consistent with a model in which the RB/IVN functions as an active recycling and sorting hub rather than simply representing posterior trafficking events associated with individual parasites.

      Upon MyoF depletion, maternal micronemes accumulated within the RB. Following restoration of MyoF expression, these accumulated micronemes were subsequently redistributed to daughter parasites. This reversible redistribution is more consistent with an active recycling process than with passive accumulation alone.

      To further support this interpretation, we expanded the analysis presented in Figure S5 by including additional vacuoles and larger replication stages (new panel D). These analyses show that maternal micronemes are redistributed broadly and relatively evenly among daughter parasites. We additionally performed a probability-based analysis demonstrating that the recurrent and homogeneous redistribution patterns observed are highly unlikely to arise from stochastic capture events occurring independently at the posterior end of each parasite during replication. Together, these analyses support the interpretation that microneme redistribution is a regulated process.

      Direct demonstration of recycling between all parasites within a vacuole would require a system allowing simultaneous differential labeling of (i) daughter parasites derived from a specific mother cell and (ii) the maternal organelles originating from that same mother during a subsequent replication cycle. To our knowledge, such an approach is not currently technically feasible. Nevertheless, our live-cell imaging experiments provide additional support for communal redistribution, as microneme material accumulated within the RB was subsequently observed redistributing, albeit unevenly, across multiple tachyzoites within the same vacuole.

      The Group 2 structures are described as those that are divided between daughters and receive newly synthesised proteins that add to the maternal protein of these compartments. While this is a logical conclusion for several that are mentioned, where the maternal protein signal is seen to be depleted with replication (including for the apicoplast, ER, glideosome, and Golgi). Data for the addition of new proteins to these existing structures is actually only presented in direct support of this for the Golgi.

      We thank the reviewer for this important clarification. We initially selected the Golgi as a representative example because its morphology and restricted localization provide the clearest visualization of the dual-labeling dynamics. However, the same experimental approach was applied to all Group 2 organelles analyzed in this study. To address the reviewer's concern more directly, we have now included a new supplementary figure (Figure S6) showing that the same pattern is also observed for the apicoplast, ER, and glideosome.

      We would also like to clarify that the maternal protein signal is not lost during replication but instead becomes progressively diluted as these organelles expand, are partitioned into daughter parasites, and incorporate newly synthesized proteins. The Golgi was originally highlighted because these dynamics are most readily visualized in this compartment; however, the same principle applies to all Group 2 organelles analyzed in this study, as now illustrated in Figure S6.

      Reviewer #2 (Public review):

      Summary:

      Toxoplasma gondii is an obligate intracellular parasite and the causative agent of Toxoplasmosis. Parasite invasion into host cells, intracellular replication, and then egress, which results in the destruction of the infected cell, is central to pathogenicity. This manuscript focuses on understanding how maternal resources (in this case, cellular organelles) are shared between daughter parasites during cell division. Many organelles are single copy, meaning that division and inheritance by the daughters is crucial for successful replication. The major strength of this study was the use of a Halobased pulse chase assay to characterize patterns of organelle inheritance. The results show that both microneme and rhoptries (secretory vesicles) previously thought to be synthesized de novo are inherited by daughter parasites. Thus, this paper adds new insight to our understanding of cell division in this important parasite.

      Strengths:

      This study demonstrated that pulse labeling of proteins can be used to monitor protein synthesis, turnover, and movement. This approach will be of great interest to the field. Using this method, the authors demonstrate three main modes of organelle inheritance.

      (1) Organelles, where there are multiple copies (such as secretory vesicles, micronemes, and rhoptries), are divided between the daughter parasites, with additional contribution of newly formed vesicles. New and old material remain as separate entities in the cell.

      (2) Single-copy organelles, which are expanded to include newly synthesized material prior to division, such as the Golgi and apicoplast.

      (3) Cytoskeletal structures that are synthesized anew during each round of division. These studies provide more refined insight into patterns or organelle inheritance and demonstrate that secretory organelles are not made de novo during each round of division as previously thought. The paper has a logical flow, and overall, the data is presented in a clear and organized fashion.

      Weaknesses:

      (1) Descriptions of methodology and statistical analysis were incomplete.

      We agree with the reviewer that the description of the methodology and statistical analyses required further clarification. To address this, we have added a new supplementary figure (Figure S11) illustrating the experimental workflow, quantification strategy, and analysis pipeline. We have also expanded the Materials and Methods section to provide detailed descriptions of the experimental design, fluorescence quantification procedures, statistical analyses, and the number of biological replicates. These revisions provide a clearer and more comprehensive description of the methodology and data analysis.

      (2) There are inconsistencies between the data in Figures 1 and 5. In Figure 1, a small amount of maternal IMC is visible in stage 2 parasites. Although this is a ~90% reduction, these parasites should be quantified as parasites with material IMC. However, the graph in Figure 5C indicates that no material parasites have GAPM1a, given that graph 5C is a binary measure (present vs. absent), one would expect a non-zero percent of parasites to have maternal material.

      We agree with Reviewer 2 that, based on the raw fluorescence signal, one might expect a non-zero percentage of parasites to retain maternal IMC material after the first replication. The apparent discrepancy between Figures 1 and 5 reflects our thresholding strategy rather than inconsistent data.

      Figure 5C presents a binary analysis (presence versus absence) using a threshold calibrated from stage 1 parasites and applied uniformly across all markers. Under this criterion, the residual GAPM1a signal after the first replication falls below the detection threshold, resulting in 0% positive vacuoles. Although normalization to stage 2 parasites would detect this weak residual signal, such a protein-specific threshold would compromise direct comparison across the dataset.

      To clarify this point, we have updated the Figure 5C legend to explain the analytical approach and the asterisk associated with GAPM1a. The residual maternal IMC signal visible in Figure 1 represents a rare example selected to illustrate the remaining ~10% signal and is consistent with the absence of detectable maternal IMC1 after replication in Figures 2C and 5E.

      (3) The conclusion from Figure 6 was not justified based on the data. I agree with the author's conclusion that the accumulation of micronemes and rhoptries in the residual body was timedependent. In Figure 6A, the signal observed in the residual body at times 6:30, 13, and 14 hours is not observed in subsequent time points. However, the fate of these micronemes and rhoptries is unclear. It cannot be concluded that these vesicles are recycled back to the mother. They could also have been degraded. In fact, the graphs of microneme inheritance in Figure 2B show a decrease in maternal signal from 100% to 80% between stages 1 and 2, indicating that some microneme degradation is taking place.

      We agree with the reviewer that Figure 6 alone does not definitively establish the fate of micronemes and rhoptries accumulating within the residual body (RB), and that both recycling and degradation remain possible interpretations. Our conclusion that maternal micronemes are predominantly recycled is therefore based on the integration of Figure 6 with our MyoF depletion and recovery experiments, additional quantitative analyses, and previous work demonstrating F-actin-dependent microneme trafficking through the RB (Periz et al., 2019).

      Consistent with this model, MyoF depletion results in the accumulation of maternal micronemes within the RB, whereas restoration of MyoF expression following auxin washout leads to their redistribution across multiple tachyzoites within the same vacuole (Figure 8). Furthermore, maternal microneme signal remains detectable even after prolonged MyoF depletion (up to 48 h) and multiple rounds of replication (Figures 7 and 8), arguing against extensive degradation.

      To further address this possibility, we quantified the fluorescence intensity of individual maternal MIC2-positive micronemes retained within the RB after 24 h of MyoF depletion and following redistribution after auxin washout (150 micronemes per condition). No significant difference in fluorescence intensity was observed compared with control maternal micronemes (Figure S10A,B), indicating that maternal microneme signal is preserved during RB retention and redistribution.

      We therefore interpret the decrease in maternal microneme signal observed between stages 1 and 2 in Figure 2B primarily as a consequence of redistribution and dilution rather than degradation, consistent with the stable fluorescence intensity of individual micronemes (Figure 3). Regarding rhoptries, we note that the majority (~90%) are incorporated into daughter parasites before budding is complete, limiting their accumulation within the RB and suggesting that RB-associated trafficking primarily reflects redistribution rather than bulk degradation.

      (4) To convincingly demonstrate that the redistribution of micronemes and rhoptries was due to recovery of MyoF protein levels after auxin washout, a Western blot should be performed to show MyoF protein levels over time. In addition, the decrease in mMIC2 protein levels in the residual body in Figure 8F should be measured and normalized for photobleaching. Both apical and basal signals appear to be reduced over the time course of imaging.

      We agree with the reviewer that demonstrating MyoF recovery following auxin washout is important. Rather than performing a Western blot, we monitored MyoF recovery by immunofluorescence using the HA tag in the MyoF-mAID-HA strain, allowing direct correlation between MyoF reappearance and microneme redistribution at the single-vacuole level. These data are now included in Figure 8B, with the corresponding MyoF presence–phenotype association analysis presented in Figure S10D.

      Regarding photobleaching, we agree that fluorescence loss during time-lapse imaging is an important consideration. However, in this experiment, changes in fluorescence intensity reflect not only photobleaching but also biological redistribution of micronemes and movement of parasites in and out of the imaging plane. In the absence of a stable internal reference fluorophore, applying a standard photobleaching correction could therefore introduce additional inaccuracies. For this reason, we did not quantify fluorescence intensity during the redistribution phase.

      Instead, to assess whether maternal microneme signal is lost during RB retention and redistribution, we quantified the fluorescence intensity of individual maternal MIC2-positive micronemes following 24 h of MyoF depletion and subsequent auxin washout (Figure S10A,B). No significant difference was observed compared with control maternal micronemes, supporting the conclusion that redistribution occurs without substantial loss of the maternal microneme pool.

      Reviewer #3 (Public review):

      Summary:

      Knoerzer-Suckow et al. explore the mechanisms of organelle inheritance during endodyogeny in Toxoplasma gondii using an innovative dual-labeling approach to track the distribution of maternal organelles into daughter parasites. They can clearly distinguish between maternal and daughterderived organelles using their dual-labeling Halo Tag approach. They reveal that different organelles are trafficked to daughter parasites in three broad patterns, which they have binned into groups. Their findings reveal a role for MyoF in the inheritance of micronemes and rhoptries, and notably, they observe that the inner membrane complex (IMC) is not recycled. Instead, the IMC undergoes a pronounced relocalization to the posterior of the maternal cell, where it is likely targeted for degradation.

      Strengths:

      The data surrounding their MyoF knockdown experiments, IMC degradation, and trafficking of MIC2 after auxin washout are compelling. These data add to the knowledge of how organelle inheritance occurs in T. gondii, increasing the field's understanding of endodyogeny.

      Weaknesses:

      (1) The evidence provided to support the claim that microneme and rhoptry inheritance specifically traffics through the residual body does not sufficiently substantiate the claim. The temporal resolution of the imaging is inadequate to precisely trace the path of microneme and rhoptry inheritance. From the data shown in the manuscript, it can be concluded that at least some of the micronemes and rhoptries might be recycled through the residual body, but it is unclear whether many or most of these organelles do so.

      We thank the reviewer for this important comment and refer also to our response to Reviewer 1 above.

      Previous work has demonstrated F-actin-dependent trafficking of micronemes within the residual body (RB) (Periz et al., 2019). Consistent with these findings, our data support a model in which RB-mediated trafficking contributes to maternal microneme recycling. We acknowledge, however, that the temporal resolution of our imaging does not allow continuous tracking of every individual organelle throughout the entire replication process.

      In contrast, our observations indicate that the majority of maternal rhoptry material is incorporated into daughter cells before replication is complete and therefore does not necessarily transit through the RB under normal conditions (Figure 6). Nevertheless, rhoptry inheritance remains dependent on the actin–MyoF trafficking machinery, as MyoF depletion results in the accumulation of rhoptry material within the RB (Figures 6 and 7).

      Taken together, our data support a model in which the RB serves as an important recycling hub for maternal micronemes and can, under conditions of impaired trafficking, also transiently accommodate rhoptry material. However, our imaging resolution does not allow us to conclude that all microneme or rhoptry inheritance obligatorily transits through the RB, and we have revised the manuscript to reflect this limitation more explicitly.

      (2) The absence of specific markers for the residual body brings into question whether microneme inheritance occurs through a discrete residual body or simply via the basal end of the maternal parasite. The authors need a robust way to visualize and define the residual body to claim that micronemes and rhoptries are specifically transported through this structure.

      We agree with the reviewer that the absence of a dedicated residual body (RB) marker remains a limitation and that such a tool would improve the precision of our analyses. To date, no specific RB marker has been identified (see also our response to Reviewer 1). The most reliable proxy currently available is the F-actin chromobody, which labels the dense F-actin network associated with the RB. Using this approach, previous work demonstrated F-actin-dependent trafficking of micronemes within the RB (Periz et al., 2019).

      Building on these findings, our data support a model in which RB-associated trafficking contributes to maternal microneme recycling, whereas rhoptries are more frequently incorporated directly into daughter cells without obvious RB transit. In addition, functional perturbation of the actin–MyoF transport machinery, through MyoF depletion and subsequent recovery, supports the interpretation that the RB represents a discrete actin-associated compartment involved in organelle redistribution. Nevertheless, we acknowledge that our current imaging resolution does not allow us to determine the extent to which all microneme or rhoptry inheritance occurs through the RB, and we have revised the manuscript accordingly.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Comments for revision where either the clarity or accuracy could be improved:

      (1) The methods could do with some further detail with respect to the fluorescence intensity measurement. For example, where the Z-series was taken, and for measurements, were maximum projects taken or single Z-planes? Were all measurements made on unprocessed images, or was any deconvolution, etc, undertaken?

      We updated the Materials and Methods for better clarity and now read as: “Parasites were labeled as described above and allowed to replicate for 24 h on HFF-coated Ibidi live-cell dishes. Approximately 15 fields of view were imaged using Z-stacks spanning 3 μm centred on the vacuoles. For each replication stage (1, 2, 4, and 8 parasites per vacuole), individual tachyzoites were sampled across multiple vacuoles.

      Maximum-intensity projections were generated from non-deconvolved images. Fluorescence intensity (FI) was quantified as the maximum grey value measured within regions of interest (ROIs) drawn on individual tachyzoites from each vacuole stage. ROIs excluded overlapping parasites, neighbouring vacuoles, and regions with atypical signal intensity. For each biological replicate, up to 25 tachyzoites per replication stage were analyzed, and the mean FI value was calculated for each stage. The highest mean FI observed among the stages within a replicate was defined as 100%, and FI values for the other stages were expressed relative to this maximum. Relative FI values were then averaged across three independent biological replicates. Data are presented as mean ± SD.”

      (2) Line 93: Is more JF549 added at each of stages 2, 4, and 8? I assume so to get the progressive increase, but it would help to clarify this here.

      As illustrated in Figure 1a, the second dye is added only once, after 24 h of replication, and is then washed off before imaging. It is not maintained throughout the replication steps. The observed increase in signal is related to the presence of newly synthesized (de novo) proteins generated during replication. The Halo tags of these new proteins are initially free of ligand, as no ligand is present during replication. During the second labelling step, all free Halo tags can bind the dye. The signal intensity at each replication step therefore reflects the amount of de novo material present, which is higher at step 4 than at step 2, because the proteins generated de novo during stage 2 are also present in stage 4. The signal will be determined by both the number of newly produced molecules and their concentration at the same localization.

      (3) Line 103: Check that the Carruthers and Sibley, 1997, ref for Tic20 in the apicoplast is correct. I don't think this could be correct given the date.

      The reference it will be corrected to G van Dooren et al. 2008.

      (4) Figure legends: It would be useful to state what form of microscopy was used in each figure.

      Following the reviewer's advice the legends has been updated.

      (5) Figure 2, S2: How are single organelles tracked, such as rhoptries? I'd assume that a cell will gain new de novo organelles as well, and that this would reduce the signal per cell. Stage 2 has a rhoptry signal in both daughters, so I'd expect the signal to be roughly half for the whole cell, unless the authors can resolve individual rhoptries (which would surprise me with this microscopy). If individual rhoptries were resolved, how was this done, and what was the confidence in this (were there controls?)

      We thank the reviewer for this important point. We do not resolve individual rhoptries with the imaging conditions used in this study. Instead, fluorescence intensity was measured as the maximum grey value within a representative region of interest (ROI) encompassing the apical rhoptry signal while excluding overlapping parasites and regions with atypical fluorescence intensity. The same ROI selection strategy was applied consistently across all replication stages, allowing direct comparison with stage 1 parasites.

      The analyses presented in Figures 3 and S2 show an increasing proportion of tachyzoites lacking detectable maternal rhoptry signal as replication progresses, while the fluorescence intensity of the remaining maternal signal remains relatively stable. Together, these observations are consistent with the redistribution of intact maternal rhoptries rather than a progressive loss of rhoptry fluorescence. 

      (6) Line 150: it is unclear what is meant by 'regulated partitioning'. The more equal inheritance of micronemes versus rhoptries might not indicate a 'regulated partitioning' but just a more uniform distribution, given the larger number of micronemes versus rhoptries.

      We agree with the reviewer that this statement required clarification, and we have revised the text accordingly. Our intention was to emphasize that microneme inheritance appears to be a regulated process, rather than to directly compare it with rhoptry inheritance. The observed differences between these organelles are likely influenced, at least in part, by their different abundances.

      Across successive rounds of replication, daughter parasites consistently inherit comparable amounts of maternal micronemes, even at later replication stages (Figure S4). Given that a single mother parasite contains approximately 30–40 micronemes, whereas successive rounds of endodyogeny can generate up to 32 daughter parasites, a purely stochastic segregation would be unlikely to produce the relatively uniform distribution observed (~1–2 maternal micronemes per tachyzoite). To support this interpretation, we performed an additional probability-based analysis, which indicates that the observed redistribution patterns are unlikely to arise by chance alone. We therefore interpret these findings as supporting the existence of mechanisms that promote balanced microneme inheritance during parasite replication.

      (7) Line 184: How is the maternal signal measured without detecting the internal daughter signal? Is this an average signal for the full parasite, or just for a cross-section of the IMC? And if the latter, how are the different profiles of mother and daughter accounted for? Also, the abbreviation in the brackets doesn't make sense here.

      We thank the reviewer for this important point. We have revised the Materials and Methods section to provide a clearer description of the fluorescence intensity (FI) measurements and added a new supplementary figure (Figure S11) illustrating the analysis workflow.

      Briefly, FI measurements were performed on maximum-intensity projections generated from Z-stack images without deconvolution. A representative region of interest (ROI) was selected, and the maximum grey value was used for quantification. This approach minimizes variability arising from differences in ROI size and provides a robust metric for comparison across replication stages.

      Daughter cell fluorescence was measured using the same approach while excluding overlapping signals from neighboring daughter cells and the maternal IMC. Maternal and daughter signals were distinguished based on their spatial localization and fluorescence labeling. Finally, the abbreviation in brackets has been corrected for clarity.

      (8) Line 191: Subheading a bit unclear. Distinct from other organelles, or are miconeme and rhoptry pathways distinct from each other?

      We agree with the reviewer and have updated the subheading to “Whole-organelle inheritance of micronemes and rhoptries occurs via distinct recycling pathways”

      (9) Line 193: The site of disassembly of the IMC (suggested RB here) might not be the same as the site of degradation. I suggest using 'disassembly' instead here.

      We agree that “disassembly” is an appropriate term to describe the breakdown of the IMC at the residual body (RB). However, we also believe that the RB represents the primary site of IMC degradation, for two reasons. First, if IMC material were not degraded at this site, we would expect to detect Halo-positive signal elsewhere following IMC collapse, which we do not observe. Second, transport of IMC material to an alternative degradation site would be required, but no IMC-positive vesicles are observed, arguing against significant redistribution. Together, these observations support the conclusion that the RB is both the site of disassembly and degradation of maternal IMC.

      (10) Line 215: The conclusion for a difference in timing of microneme and rhoptry segregation is not clearly supported by the data presented. Also, if there are more micronemes than rhoptries, then the frequency of observing a microneme being trafficked through the RB would need to be higher than for rhoptries if the mechanisms were the same. So, a difference in frequency here cannot be used to argue for a different mechanism.

      We agree with the reviewer and the text have been edited to soften our conclusion. 

      (11) Line 234: 'segregation' might be a better term than 'recycling' here because it is actually the sorting into daughter cells that is the important process.

      The text have been edited

      (12) Line 235: I don't think this can be what the authors intend to say. If the maternally-inherited rhoptries are not trafficked through the RB (every time), then how do they get into the daughters? Perhaps this is a case where a clear definition of the RB is required.

      Our observations indicate that maternally inherited rhoptries are frequently incorporated into daughter cells before collapse of the mother cell and establishment of the residual body (RB). Although F-actin is enriched within the RB, an actin network is also present throughout the parasite cytoplasm, where MyoF is likewise localized. We therefore propose that, unlike micronemes, rhoptries do not necessarily transit through the RB during every replication cycle but can be incorporated directly into developing daughter cells while still relying on the same actin–MyoF-dependent trafficking machinery.

      (13) The MyoF Rescue, the experimental plan is not fully described in order to be clear. If the endomembrane architecture was disrupted by MyoF depletion, and this secondary effect caused the segregation phenotype, restoration of MyoF might also simply restore the endomembrane system. So a direct role for MyoF doesn't seem to have been tested in this case.

      We appreciate the reviewer's concern that the segregation phenotype could, in principle, arise indirectly from disruption of endomembrane architecture following MyoF depletion. However, although Golgi morphology is altered in MyoF-depleted parasites, its core functions appear largely preserved. This is supported by the normal biogenesis of de novo micronemes, their correct targeting to the apical pole, their efficient secretion, and the previously reported preservation of parasite invasion. In addition, Golgi markers are not detected in the residual body, where maternally inherited micronemes accumulate, arguing against Golgi-mediated trafficking as the primary cause of the segregation phenotype.

      Taken together, these observations support the interpretation that the segregation defects are more likely to reflect a direct role of MyoF in organelle trafficking and inheritance than a secondary consequence of generalized disruption of endomembrane organization.

      (14) Line 279: Why call it a checkpoint? What is the evidence for its presence here being sensed before a further process is activated, which is what a checkpoint does?

      We agree the reviewer that checkpoint is a misleading term and have been updated to trafficking hub. 

      (15) Line 287 confuses replication of the daughters from cytokinesis, which only happens when each cell loses cytoplasmic connectivity with the other.

      We will clarify this point. In Toxoplasma gondii, cytokinesis represents the final step of daughter cell formation, during which the two fully assembled daughter parasites separate from the mother cell following collapse of the maternal cytoplasm. Historically, the residual body was proposed to arise simply as leftover material from this process. However, multiple studies have now shown that residual body formation is an active and regulated process, dependent on specific cytoskeletal and trafficking factors. Importantly, although cytokinesis marks the physical separation of daughter cells from the mother, parasites within a vacuole remain connected via the residual body and continue to share cytoplasmic and plasma membrane components until egress. 

      (16) Line 298: I don't think there is direct evidence of degradation in the RB. There might be disassembly, but degradation implies proteolysis, which hasn't been tested for.

      We agree with the reviewer that our data do not provide direct biochemical evidence of proteolysis within the residual body (RB) and primarily demonstrate disassembly of the maternal IMC at this site. However, several observations are consistent with local degradation. Following IMC collapse, we do not detect Halo-positive signal elsewhere in the parasite, nor do we observe IMC-positive vesicles or other structures that would suggest transport to a distinct degradation compartment.

      In addition, previous work identified the E3 ubiquitin ligase CSAR1 as a mediator of protein turnover within the RB, supporting the idea that this compartment is associated with degradation-related processes (O'Shaughnessy et al., 2023). While we cannot formally demonstrate proteolysis, these observations support a model in which IMC disassembly is closely coupled to local degradation within the RB.

      (17) The paragraph structure gets a bit confusing at times. See single sentence paragraph, Line 224. Does this sentence justify its own paragraph?

      The text has been edited.

      (18) Make sure Toxoplasma gondii is in italics throughout.

      The text has been edited

      (19) Line 279 cites Figure 10. But there is none.

      The text has been edited

      (20) I advocate introducing a few new acronyms, like DCs. I find that this ultimately reduces the ease with which readers read the work if they don't learn them all quickly.

      We agree that excessive use of acronyms can negatively impact readability. In the present manuscript, all abbreviations used in the text are introduced at their first occurrence in the Introduction, including DCs (line 32), IMC (line 41), PV (line 35), ER (lines 38–39), RB (line 50), and IVN (line 49). We have carefully limited the use of abbreviations to commonly used terms in the field and to those that recur frequently throughout the manuscript, with the aim of balancing clarity and readability. Nevertheless, we are happy to reduce or remove specific abbreviations if the reviewer feels this would further improve clarity.

      Reviewer #2 (Recommendations for the authors):

      (1) Descriptions of methodology and statistical analysis were incomplete as follows:

      (1a) It was unclear how the fluorescence intensity measurements (used to evaluate inheritance vs. new synthesis) were carried out. The y-axis on the graph is labeled average fluorescence intensity (% of max intensity). However, it does not state what was averaged (average fluorescence per vacuole?) and what was max intensity (max pixel intensity in each image or time point with the highest average intensity, relative to the other time points?)

      We agree with the reviewer that the original description of the fluorescence intensity (FI) measurements lacked clarity. We have therefore revised the Materials and Methods section and added a new supplementary figure (Figure S11) illustrating the analysis workflow.

      Briefly, vacuoles were imaged as Z-stacks, and maximum-intensity projections were used for analysis. FI was quantified as the maximum grey value measured within representative regions of interest (ROIs) drawn on individual tachyzoites, rather than as an integrated fluorescence signal across the vacuole. This approach minimizes variability arising from differences in ROI size and allows direct comparison between replication stages.

      For each biological replicate, up to 25 tachyzoites per replication stage were analyzed. The mean FI for each stage was normalized to the highest mean value within that replicate, and data from three independent biological replicates were subsequently averaged.

      (1b) Given the uncertainties with how these measurements were performed, it is difficult to interpret the data. For example, one would expect that the fluorescence intensity of newly synthesized IMC1 in 8-parasite vacuoles would be 4 times higher than that of a 2-parasite vacuole; however, based on the graph in Figure 1B, the measured increase was only 30%.

      We agree that the original description of the fluorescence intensity (FI) measurements required further clarification and have revised the Materials and Methods accordingly. As the reviewer correctly notes, a fourfold increase in FI between 2- and 8-parasite vacuoles would be expected if total IMC fluorescence across the entire vacuole had been measured. However, this was not the parameter quantified.

      Instead, FI was measured as the maximum grey value within representative regions of the daughter IMC, providing a per-cell rather than a whole-vacuole measurement. Using this approach, FI increases between the 2- and 4-parasite stages and then reaches a plateau.

      This behavior is consistent with the biology of IMC biogenesis. Although the total amount of IMC per vacuole increases with parasite number, the amount of IMC protein incorporated into each daughter parasite remains relatively constant. Consequently, once daughter IMCs are fully assembled from de novo-synthesised material, additional rounds of replication increase the total IMC content per vacuole but not the fluorescence intensity measured for individual parasites.

      (1c) T. gondii replicates in an asynchronous manner, so that at the 24-hour time point, a single dish can contain vacuoles containing 2, 4, and 8 parasites. This should be stated explicitly so readers unfamiliar with T. gondii's growth patterns can understand how the experiment was performed.

      We agree with the reviewer and have updated the text line 93. “As Toxoplasma gondii replicates in an asynchronous manner, after 24 of replication, vacuoles containing 1, 2, 4, and 8 parasites can be observed in a single dish.”

      (1d) Colocalization package in Fiji used for ANKER1-Halo with HDEL-GFP and MIC2/RON2 with CbEmeraldFP should be specified.

      We thank the reviewer for this suggestion. Following this recommendation, we performed Pearson correlation analysis for the ANKER1–HDEL-GFP experiment using the Coloc 2 plugins of FiJi. ANKER1Halo and HDEL-GFP showed a strong spatial correlation (Pearson's R = 0.92 ± 0.04, n=30 parasites from three independent biological replicates), supporting localization of ANKER1 to the ER.

      We note, however, that this analysis should be interpreted as evidence for co-distribution within the same organelle rather than direct molecular colocalization, as ANKER1 is a transmembrane protein whereas HDEL-GFP labels the ER lumen.

      For all the rest of our analysis, no automated colocalization package or plugin in Fiji was used for the analyses involving ANKER1-Halo with HDEL-GFP or MIC2/RON2 with Cb-EmeraldFP. Colocalization was assessed manually across all experiments by inspecting both full Z-stacks and maximum-intensity projections to ensure robust spatial overlap.

      For MIC2 and RON2, the presence of signal within the Cb-Emerald–positive filament was scored as either cytoplasmic, on the residual body or absence of colocalisation.

      In total, more than 300 and 500 vacuoles were analyzed for MIC2 and RON2–Cb-Emerald colocalization respectively (stable expression of both markers), and more than 150 vacuoles were analyzed for ANKER1-Halo and HDEL-GFP colocalization (transient expression of HDEL-GFP).

      This information has now been added to the Methods section.

      (1e) Statistical methods should be described on an experiment-by-experiment basis. The authors should justify why a one-tailed t-test was conducted. A two-tailed t-test seems more appropriate.

      We agree that statistical methods should be clearly justified on an experiment-by-experiment basis. All the statistical analysis have been performed using two tails and updated in the figures.

      (1f) In Figures 5E and 5F, using boxes to indicate the exact areas of the cell that were used in the fluorescence intensity measurements, rather than arrows, would make this data easier to interpret.

      The figure has been updated.

      Reviewer #3 (Recommendations for the authors):

      The current time-lapse images and videos do not clearly demonstrate microneme movement from the maternal parasite apical end to the residual body and back to the apical end of daughter parasites. As such, the route by which micronemes enter daughter parasites remains inconclusive. To strengthen their claims, the authors should employ higher temporal resolution imaging to definitively capture the movement of micronemes from the maternal apical region into the daughters. From the current data, it also seems plausible that the micronemes may be trafficked into the daughters through the conoid as well, as there is no evidence provided showing a movement of micronemes away from the apical end of the maternal parasite before being present in the daughter parasites.

      We agree with the reviewer that higher temporal resolution imaging would provide a more definitive view of microneme trafficking. However, long-term live imaging of replicating Toxoplasma gondii requires a compromise between temporal resolution and parasite viability. In our experiments, images were acquired every 15–30 min over periods of up to 16 h, as more frequent acquisition consistently induced phototoxicity and prevented completion of parasite replication.

      Despite this limitation, our imaging reliably tracked maternal micronemes over successive rounds of endodyogeny and consistently showed microneme signal associated with the residual body. These observations are in agreement with previous high-temporal-resolution studies, which demonstrated F-actin-dependent microneme trafficking within the residual body over shorter imaging periods (Periz et al., 2019).

      We cannot formally exclude the possibility that some micronemes are transferred directly to daughter parasites through the apical end. However, together with previous studies showing enrichment of F-actin at the basal region of developing daughter cells rather than at the apical tip (Periz et al., 2017), our observations support a model in which RB-mediated trafficking contributes to maternal microneme inheritance.

      The lack of a clear residual body marker needs to be addressed, as the distinction between the basal end of the maternal cell and a bona fide residual body must be explicitly defined to substantiate the major claim of the study. As it stands, it remains unclear whether micronemes and rhoptries as a whole travel through the residual body to be transported into the daughter parasites.

      We agree with the reviewer that a marker specific to the residual body would strengthen this study. Unfortunately, no such marker has been identified to date. The F-actin chromobody is currently the best available proxy, as previous studies have shown that the F-actin network is enriched within the residual body (Periz et al., 2017; Kellermeier et al., 2024). Moreover, high-resolution live-cell imaging has previously demonstrated F-actin-dependent microneme trafficking within this compartment (Periz et al., 2019). We have revised the manuscript to more clearly acknowledge this limitation.

      The conclusions drawn from the actin colocalization data in Figure 6C are based entirely on fixed samples, despite all experimental tools being compatible with live-cell imaging. Supplementing the fixed imaging with live cell data would increase its biological relevance. Published studies have shown that fixation of the actin chromobody results in the loss of resolution of an appreciable amount of the cytosolic F-actin network, and while the localizations analyzed here are primarily along the periphery, since the quantification and text make claims about the colocalization within the cytosol, this potential loss of cytosolic F-actin becomes an issue as there may be more actin available for analysis that is lost due to fixation within the cytosol of the parasites.

      We thank the reviewer for raising this important point and agree that conventional fixation can compromise preservation of the F-actin network. However, we used the same fixation protocol described by Periz et al. (2019), which allows reliable visualization of the RB-associated F-actin network. We have also corrected the description of the fixation protocol in the Materials and Methods.

      Fixation was necessary to image entire vacuoles with sufficient spatial resolution and signal-to-noise ratio for the volumetric analyses presented in Figure 6C. Although some loss of cytosolic F-actin cannot be excluded, this would be expected to reduce, rather than artificially increase, the detection of organelle–actin associations.

      Importantly, previous live-cell imaging studies demonstrated F-actin-dependent microneme trafficking (Periz et al., 2019), and our observations are consistent with these findings. Moreover, the defects observed following MyoF depletion provide independent functional evidence that the trafficking events described here rely on the actin–MyoF transport machinery.

      The statement of colocalization should be backed up by quantitative coefficients like Pearson's coefficient.

      We thank the reviewer for this helpful suggestion. Following this recommendation, we performed a Pearson correlation analysis of ANKER1-Halo and HDEL-GFP using the Coloc 2 plugin in Fiji. ANKER1-Halo showed a strong spatial correlation with HDEL-GFP (Pearson's R = 0.92 ± 0.04, n = 30 parasites from three independent biological replicates), supporting localization of ANKER1 to the ER. As ANKER1 is a transmembrane protein and HDEL-GFP labels the ER lumen, this analysis should be interpreted as evidence of co-distribution within the same organelle rather than direct molecular colocalization.

      In contrast, we do not consider Pearson's coefficient appropriate for evaluating the association of micronemes or rhoptries with F-actin. These organelles are predominantly concentrated at the apical pole and, when associated with F-actin, are typically positioned along rather than directly overlapping the filaments. Consequently, Pearson's coefficient would underestimate these biologically relevant associations. We therefore relied on morphological and spatial criteria, which we consider more appropriate for assessing organelle–cytoskeleton interactions.

      In addition, from the methods and presented figure images, specifically in Figure 6C for RON2, how the cytosolic and residual body actin is separated is difficult to discern, as there is a clear residual body actin signal overlapping a parasite. The methods for how this was separated and analyzed should be clearer to remove doubts about how this area was measured, as the current description raises concerns about the counting of the residual body actin within the cytosol.

      We agree with the reviewer that the distinction between cytosolic and residual body (RB)-associated F-actin required further clarification. All analyses were performed manually, as described in our response to Reviewer 2 (comment 1d). The RB was identified by the presence of thick, bundled F-actin filaments at the basal pole that formed a continuous structure connecting parasites within the vacuole, whereas cytosolic F-actin was defined as the thinner filamentous network within the parasite body.

      No automated or threshold-based segmentation was used because the marked differences in filament morphology and fluorescence intensity make reliable thresholding difficult and prone to misclassification. Manual annotation based on spatial localization and filament morphology was therefore considered the most appropriate approach. We have clarified these criteria in the Materials and Methods section.

      Line comments:

      (1) 113: round to rounds - "did not obtain maternal organelles after successive rounds of replication...".

      The text has been updated

      (2) 146: grammatical, "As consequence a progressive" -> "As a consequence", or "Consequently".

      The text has been updated

      (3) 164-166: "Autonomous duplication" implies the separation and duplication of the Golgi occurs on its own, i.e., without any outside intervention, when we know from Carmeille et al. 2021 and Figure 7C here that the Golgi becomes fragmented over rounds of division in the absence of MyoF. I think this is primarily a word choice error with "autonomous".

      We agree with the reviewer and the word autonomous has been removed

      (4) 184: The wording suggests that DC's refers to daughter IMC's, when DC has already been given as an abbreviation for daughter cells previously.

      The text has been updated to correct this error

      (5) 189: de novo is not italicized.

      The text has been updated

      (6) 192: The data shown so far do not show that the RB plays a selective role in organelle recycling.

      The text has been edited to fit better our results “Our data suggest that the organelles trafficking through the residual body (RB) have different fate”

      (7) 208: State that it depends on F-actin, but never show that it is dependent on F-actin through actin disruption, such as cytochalasin D treatment or a specific conditional disruption of F-actin.

      We agree with the reviewer that we did not repeat F-actin disruption experiments (e.g., cytochalasin D or jasplakinolide treatments) in this study. These experiments were performed in our previous work, where pharmacological disruption of F-actin was shown to impair microneme trafficking (Periz et al., 2019). We therefore chose not to repeat these assays.

      Instead, the present study provides complementary evidence by demonstrating that depletion of Myosin F (MyoF), a motor that uses F-actin as a transport track (Kellermeier et al., 2024), disrupts the trafficking of both maternal micronemes and rhoptries. Together, our previous F-actin perturbation experiments and the MyoF depletion data presented here support the interpretation that these trafficking events depend on the actin–MyoF transport machinery. 

      (8) 209: This suggests that the chromobody was transiently expressed in the RON2-Halo line, but the methods suggest MIC2-Halo and RON2-Halo were integrated into a parasite line stably expressing Cb-Emerald.

      The text has been edited.

      (9) 229: The section is confusing with the mention of (now maternal). If I understand correctly, the point being made is that the de novo synthesized MIC2 at stage 2 is now the maternal MIC2 for stage 4, but coloring-wise within the figure, the now maternal MIC2 at stage 4 from stage 2 would still be green. The methods suggest these images were all taken simultaneously, and not at specific timepoints of the same vacuole, so the now maternal line remains confusing.

      The text has been revised for clarity and now reads: “Because Toxoplasma gondii replicates asynchronously, vacuoles at different replication stages coexist within the same culture after 24 h. The second labeling step marks all proteins synthesized since the beginning of the experiment, allowing discrimination between proteins present in the original mother parasite and those synthesized during subsequent replication cycles. As daughter parasites form, they inherit material from their mother, such that proteins synthesized during one replication cycle become maternal proteins in the next. Under MyoF depletion, these newly synthesized protein pools accumulate within the residual body instead of being redistributed to daughter parasites during subsequent rounds of replication (Figure 7, stage 4).”

      (10) 238: The sentences here indicate that Golgi inheritance occurs without issue: "golgi inheritance remained unaffected by MyoF depletion". But, it is evident from the images shown that the Golgi is extremely fragmented, with many more Golgi fragments by stage 8 than there are parasites. This could be solved by rewording and including a line along the lines of "in accordance with the results found in Carmeille et al. 2021".

      The reference to this article is already stated later in the text now line 299-303 “This active role is further supported by the dependence of RB-mediated recycling on F-actin and the class XXII myosin MyoF, which we show to be essential for retrieval of maternal MIC2 and RON2 but dispensable for Golgi inheritance although we noticed a fragmentation of the Golgi, which has been described to depend on MyoF (Carmeille et al., 2021).”

      (11) 319: missing a comma, "Many of these, particularly...".

      The text has been edited.

      (12) 437: Images -> imaged.

      The text has been edited.

      (13) 439: a fresh media -> and fresh media.

      The text has been edited.

      (14) 439: Replication -> replicate.

      The text has been edited.

      (15) 440: images -> imaged.

      The text has been edited.

      Other grammatical issues within the methods:

      (1) Figure 5C: Within the y-axis label of Figure 5C, there is an asterisk with no asterisk explanation within the legend.

      We thank the reviewer for pointing out this oversight. The figure legend has been updated to include an explanation of the asterisk, providing a clearer understanding of the results.

      (2) Figure 5E: The addition of an IMC1 label to match the magenta color would be helpful to readers.

      The figure has been updated

      (3) Figure 6D: No statistics showing significance of colocalization.

      We thank the reviewer for this comment. In Figure 6D, we report the frequency of observed colocalization between organelles (MIC2 and RON2) and F-actin filaments, based on manual analysis across >300 vacuoles for MIC2 and >500 vacuoles for RON2. Specifically, we observed cytoplasmic F-actin colocalization in 94% of vacuoles for MIC2 and 79% for RON2, and residual body F-actin colocalization in 85% of vacuoles for MIC2 and 11% for RON2. This analysis is descriptive and is not intended to compare MIC2 versus RON2 quantitatively; rather, it illustrates the general association of these organelles with F-actin filaments. Standard statistical measures such as Pearson correlation are not meaningful in this context because the organelles are punctate and primarily localized along filaments rather than overlapping continuously. Importantly, the percentages reported represent the fraction of vacuoles in which colocalization can be observed, not the percentage of colocalization between Cb-Emerald and MIC2/RON2 within individual vacuoles.

      (4) Figure 7: The legend title of Figure 7 is at the end of the legend of Figure 6.

      The text has been edited.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Here, Pinto and colleagues set out to investigate whether the cow udder is a potential mixing site for the influenza virus. The authors have demonstrated that bovine mammary epithelial cells can be infected with both avian and human influenza A viruses, supporting the idea that the cow udder may be a potential site for reassortment. Furthermore, they demonstrate that the bovine-adapted IAV replicates to similar titers in avian epithelial cells when compared to an AIV precursor virus. Thus, suggesting there is no fitness trade-off, and confirms the potential for spill-back of the cattle B3.13 into poultry, which has already been observed. Overall, I believe the authors achieved their aims. However, there are instances in which the results do not entirely support the conclusions (noted in weaknesses). Given the ongoing questions surrounding highly pathogenic avian influenza A virus in dairy cows, this work provides valuable evidence for the potential of the cow udder as a site of reassortment. These findings highlight the need for surveillance of influenza A virus incursions into livestock species, particularly cows. Some specific strengths and questions regarding weaknesses have been outlined below.

      Strengths:

      (1) The authors use a diverse range of cell types and influenza A virus strains, as well as a wide range of techniques to address the questions at hand.

      (2) The use of cells from multiple bovine breeds for the MAC-T, bMEC and explants suggests the phenomenon is not unique to a single breed.

      (3) The results suggesting there is no fitness trade-off for Cattle Texas in an avian host are interesting, and confirm the potential for spill-back of the cattle B3.13 into poultry, which has been observed.

      Weaknesses:

      I have listed my complete questions/concerns below. However, there are two main weaknesses of the article in its current state. Firstly, there is no apples-to-apples comparison in terms of determining a preference for IAV to infect the cow udder over other organs (Q4). The mammary gland and respiratory tract are represented by epithelial cells, but for other organs, fibroblasts were chosen. I think the fairer comparison would be to compare epithelial cells from different organs to demonstrate a preference for the mammary gland. Secondly, the main premise of the article relies on bMEC and MAC-T (primary and immortalised mammary epithelial cells), facilitating higher viral growth than the cells from other organs. Yet throughout the article, a 10x higher dose of IAV is used in the bMEC cells compared to everything else (Q6). This raises the question of how much of the results are due to a preference for the mammary epithelial cells, and how much is simply due to the increased dose.

      (Q4) When we set out to test if cow mammary gland cells were particularly susceptible to IAV infection compared to other bovine cell types, we used what was available in the Roslin Institute – a mix of primary and continuous cells from various anatomical sites: three epithelial cell types (two mammary, one respiratory tract) two immune cell types and four sets of fibroblasts from various organs. Given the representation of different anatomical sites, cell types and differentiation statuses, we considered this a suitably diverse panel with which to characterise infection dynamics of a broad range of IAVs, before more focussed investigations using the bMEC and explant tissues. Both mammary epithelial cell types grew our library of influenza challenge strains significantly better than the BAT-II respiratory epithelial cells, as well as the two immune cell types and all four fibroblast populations. Of the fibroblast cells, those derived from the brain grew IAV significantly better than the skin and turbinate fibroblasts, while blood-derived macrophages grew virus significantly better than the lymphocytes and non-brain fibroblasts. So there are “apple-to-apple” comparisons as well as apple-to-pear comparisons that give significant differences. We therefore think that our conclusions (in the abstract) that mammary cells are particularly replication competent for IAV, (at the end of the introduction) that “a wide range of cow-derived cells are susceptible” and that (in the results section) that “mammary cells showed the highest susceptibility” are justifiable. However, we agree that testing a wider variety of epithelial cells would be useful and have added text to the Discussion (lines 224-228) to acknowledge this.

      (Q6) We used a higher MOI for bMECs because test experiments with WT PR8 and the Cattle Texas 6:2 reassortant virus showed that MOI 0.01 infections gave more variable results than those run at MOI 0.1, perhaps because of the intrinsic variability of mixed primary cell populations. However, the end-point titres between the two conditions were not significantly different, so we therefore chose to go with the higher MOI. Accordingly, we do not think this choice is a confounding issue. This explanation (line numbers 340-345) and a new Supplementary Figure 11 showing the results of the two MOI tests have been added to the manuscript.

      Reviewer #2 (Public review):

      The authors use a library of influenza A viruses from different strains, classified in lab-adapted, human, avian, and swine according to the animal from which they were isolated. They propose that the cow mammary gland serves as a mixing vessel for influenza A viruses. As a first approach, the authors assess susceptibility to infection across different cell types, including continuous and primary cell lines, bovine mammary cells, and mammary explants. All these cells support polymerase activity. Then, they analyzed changes in the bovine virus's viral fitness relative to an avian precursor. The authors use single-gene replacement to study whether and which RNP segments improve viral transcription. As part of this section, they also test IFN-specific antagonism by NS1 to assess the input of segment 8. Quantitative glycomic analysis was performed on the continuous bovine mammary cell line to demonstrate the presence of both a2,3 and a2,6, which is consistent with their observation that these cells can be co-infected with human and avian IAVs simultaneously. The main question, however, is: what is the glycome in the explants, or directly from tissues?

      We report quantitative glycomics for the primary bovine mammary epithelial cells as well as the continuous line the referee highlights. However, we agree with R2 that a detailed glycomic analysis of primary bovine mammary tissue would allow a better understanding of the actual glycosylation status in vivo. This has been undertaken by the authors and is available as a bioRxiv preprint. This is now cited (ref 25) in the relevant part of the results (line 184-185)

      Overall, the manuscript is clearly written and provides new insights into the behaviour of the cattle isolate, now compared with a representative group of model or precursor HAs of different origins.

      It would be great if a consistent nomenclature for the IAV strains could be used in the study. There is a mix of origin (Texas), animal from which the virus was isolated (mallard), or abbreviations that do not follow guidelines (IAV07). Are the USSR and Udorn not lab-adapted?

      We chose the abbreviated names for a variety of reasons. Partly from common usage (e.g. PR8, Udorn), partly for consistency with other already published papers from the FluTrailMap consortia (e.g. Cattle Texas; Dholakia et al 2026), partly to make diversity obvious in certain figures (e.g. H3N1, H5N2 etc) and partly to avoid confusion between viruses that originate from the same geographic area (e.g. AIV07, AIV09, H5N8-20 etc which are all A/Ck/England/isolate numbers). Overall, we found it more confusing to use the expanded nomenclature. Re AIV07 which the referee criticises for not following naming guidelines – if this is a reference to the EURL nomenclature, AIV07 is the abbreviation for the specific virus A/Chicken/England/053052/2021, our representative virus for EURL genotype EA-2020-C, as we say in the text. This nomenclature has now been added to Table 1, to provide a fuller cross-reference for all the names.

      As to whether USSR and Udorn are lab-adapted – that depends on definitions. There is a continuum of adaptive changes and/or sequence drift starting from the very first growth cycle of an isolate in the laboratory. The viruses we define here as lab adapted are ones that have been deliberately adapted to other host species or which have very long passage histories in multiple laboratory systems resulting in known functionally significant changes; for example, one lineage of PR8 was passaged 77 times in mice, 717 times in cell culture, 30 times in chick embryos, 5 times in ferrets and a further 50 times in chick embryos (https://www.medscape.com/viewarticle/812621_3?form=fpf), rendering it unarguably lab-adapted. We admit that A/USSR/77 and A/Udorn/307/1972 are probably further along this adaptive pathway than more recent isolates such as A/Norway/3433/2018, but are unaware of any specific reason that would put them into our lab-adapted category.

      The experimental setup includes bovine mammary primary and continuous cells, as well as mammary explants. Some of the most significant differences, for example, in viral fitness studies and co-infection experiments, are observed in these explants. Perhaps there could be some additional focus on this observation. The implications in comparison to the results obtained in cultured cells could be described. How will the human and other HA subtype viruses fare in the explants?

      We agree that this is an important and interesting question, and had already tested the strains we used for co-infections: human seasonal pdm09 H1N1 “Norway” and low pathogenic avian influenza “H3N1”, in the mammary explants. Both replicate the avian virus to 20-fold higher titres. We have added this information to the revised manuscript as new Figure panels S6E-H, called out on line 200-201 of the results.

      Reviewer #3 (Public review):

      Summary:

      This excellent manuscript by Pinto, Sharp, and colleagues examines bovine tissue tropism for influenza viruses. They find that bovine flu, as well as other strains, has strong replication in mammary tissue. They also map the genetic changes to influenza that improve replication in bovine cells. Overall, the study is well designed and executed, and the results are very timely.

      Strengths:

      (1) The experiments are well-controlled.

      (2) The figures are well-constructed and easy to follow.

      (3) The Methods and legends are detailed, with sufficient information.

      Weaknesses:

      (1) A comparison to human cells would strengthen the overall impact of the results. Are human mammary cells also uniquely susceptible to influenza? Are bovine mammary cells special in some way?

      This is an interesting question, but we have not tested mammary gland cells from humans (or any other species of mammal). We have however reported elsewhere (Dholakia et al., Nat Commun. 2026 Jan 16;17(1):1603. doi: 10.1038/s41467-026-68306-6.) that Cattle Texas grows well in a variety of human respiratory cells. Here, we are considering the bovine mammary organ as a potential reassortment site for IAVs because of the ongoing viral mastitis epidemic in US dairy cattle; human mammary organs seem unlikely to create a similar opportunity.

      (2) For the virus infection studies with segment 8 swaps, it should at least be noted that some of the phenotypes could be driven by NEP.

      We agree; as Table S1 indicates, NEP has two changes (one shared with NS1) between AIV07 and our B3.13 isolate, so we should not have conflated segment and NS1. We have changed the text to acknowledge this throughout the results (lines 127, 137 and 149) and in the discussion (line 243-244).

      (3) The data demonstrating that bMEC can support co-infection are compelling and important, but would be strengthened with a comparison from a different cell type or species. Do mammary cells uniquely support higher co-infection?

      We have data showing that co-infection also occurs in the continuous MAC-T udder cell line and have now included these data in a revised Figure 4D (described/called out on lines 198-206). We have not tested bovine cells from other organs for co-infection potential as they do not seem to be significant sites of infection in vivo.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) How nasal turbinate and cardiac fibroblasts are acquired/cultures is missing from the methods.

      Apologies for the omissions and thank you, because rectifying this brought to light an error in cell naming. The cells originally called bovine cardiac fibroblasts were in fact a second independent preparation of skin fibroblasts. We have corrected the labelling in Figs 1, 2 and S2. The nasal turbinate cells were bought in from ATCC (code CRL-1390). This information has now been added to the methods (line 310).

      (2) Please specify what cell types make up the 2D enteroids, the mammary explants and the nasal turbinates.

      The composition of the 2D enteroids is described in detail in reference [60]. In precis, they are comprised predominantly of epithelial cells, including Paneth, goblet and enteroendocrine cells, as well as stem cells. This information has been added to the methods (lines 374-376) The mammary explants include duct epithelium, connective and muscle tissue, defined by H&E staining of cut sections (see new Figure S12 and text added to the methods on lines 384-386). We have also added the person who did the histology (Rebecca Ross) as an author and to the credit taxonomy (line 877). The nasal turbinate cells appear to be predominantly fibroblast morphology (Methods line 310-311).

      (3) Epithelial cells are the main target for IAV, so why were fibroblasts chosen as a comparison to mammary epithelial cells? This needs justification in text. Without justification, how much can be attributed to the results being mammary-specific, rather than epithelial-specific? The brain (choroid plexus epithelium), heart (epicardium) and skin all contain epithelial cells.

      We think this query is what the referee calls “Q4’ in the public part of their review. Please see our answer above.

      (4) Figure 1, S1, S2 and S3 seem to suggest that mammary cells are more susceptible to IAV infection than cells from other organs. But Figure 2 demonstrates that when it comes to Cattle Texas and AIV07, most of the cell types show high viral titers. If the question is about whether the cow udder is the primary mixing site, would it not be more relevant to investigate which cell type facilitates the best growth of the potential precursor viruses, similar to Figure 4A?

      Figs 1, S1, 2 and 3 all use “full” viruses whereas Fig 2 uses 6:2 reassortants between PR8 and the HPAIVs for biosafety reasons. WT PR8 replicates well in most of the bovine cells tested (Figs S1-3) so we do not see any contradiction. Figure 2 examines the contributions the internal genes make to replication in bovine cells. Re the question over the udder being a potential mixing site – this is where the virus is replicating in the real world; at least in part because of the transmission mechanism, but also because the mammary gland epithelium is highly susceptible to infection, as we show.

      (5) It is misleading to compare the bMEC infection of MOI 0.1, to all other infections of MOI 0.01. This is a consistent problem throughout the article - Figures 1, S1, 2, 4. Why was the bMEC infection at a 10x greater dose? The main premise of the article relies on bMEC and MAC-T (primary and immortalised mammary epithelial cells), facilitating higher viral growth than the cells from other organs. If we compare the MOI 0.01 experiments alone, then the evidence relies on the immortalised MAC-T cells, compared to primary cell types. In this case, how much can be said about it being mammary specific, rather than immortalised vs primary? I do wonder, for example, how the epithelial nasal turbinates or type II pneumocytes would compare to the primary mammary epithelial cells if they were at the same MOI.

      Please see our answer to this query earlier in the rebuttal.

      (6) Why are the 2D enteroids excluded from Figure 1?

      We had limited supplies of a difficult-to-grow cell model, so we only used them to test the 6:2 viruses (Fig 2).

      (7) The colour scheme for Figure 3 is confusing. In Figure 3A, blue indicates European ancestry, and yellow represents North American ancestry. However, in Figure 3B, these colours now mean something different. To a reader, when there is a colour-coded schematic, it is instinctual to think that this then corresponds to the following panel(s). Since consistently throughout the article, yellow has been used for Cattle Texas, and blue has been used for AIV07, I would suggest choosing different colours to represent European and North American ancestry in Figure 3A

      We’ve changed the figure as the referee suggests and modified the Fig 3 legend accordingly (line 895)

      (8) I am unsure about the conclusions drawn from the results of Figure 3B. In the results, it is framed as trying to determine which segments contributed to the improved activity of Cattle Texas compared to AIV07. In lines 110-111, "... PB2 or PA from AIV07 significantly decreased Cattle Texas minireplicon activity". If PB2 is indeed significant, there is a missing yellow asterisk in Figure 3B.

      Apologies, there was indeed a missing asterisk on the figure; now added.

      Given the significance of PA, why was it not investigated in terms of growth kinetics similar to Figure 3C? Was it overlooked because it doesn't have a North American ancestry? The results of Figure 3B suggest that the 4 amino acid mutation in PA has significantly contributed to changes in polymerase activity.

      The PA changes do indeed matter for minireplicon activity – the key change is K497R, as detailed in our related publication in Nat Comms (citation 17). However, it is less important than changes in PB2, and the PA segment swap by itself has little effect on overall virus replication.

      (9) Similarly, in Figure 3C and lines 116-117, the error bars on the graph are overlapping at 48 hours, suggesting no difference in overall replication. The kinetics are slowed for AIV07 seg1-3, but not for AIV07 seg 1, indicating PB1 does not have an effect. This would then suggest that something in segment 2 or 3 contributes to the slowed kinetics in Figure 3C, which, from the Figure 3B results, is unlikely to be due to PB2. While it was not reassorted, the PA segment is potentially the driver, with its 4 amino acid mutations. I think it is worth performing growth kinetics with and without these 4 amino acid changes in PA.

      We agree that visually on a log<sub>10</sub> scale, the titres of the “WT” 6:2 Cattle Texas and 5:2:1 segment 1 reassortment appear close, but the average titres are 5 and 7-fold different at 24 and 48h respectively, while a 2-way ANOVA with Dunnet’s multiple comparison post-test gives statistical significance at 48h. We have added this information to the figure and its legend (lines 905-906).

      (10) In Figure 3C and E, why did you choose to perform the growth kinetics in the immortalised cell line, when you have access to primary cells? The primary cells would be a more accurate representation of what happens in situ.

      The primary cells were difficult to work with and only available intermittently, so we used what was available at the time.

      (11) In lines 145-147, "thus overall, the reassortment event that replaced segments 1, 2 and 8 alongside drift adaptations in segment 3 may have contributed to the ability of the B3.13 genotype virus to infect cattle". This is not clearly supported by the evidence presented. In terms of segment 1/PB2, the growth kinetics of Figure 3C have overlapping error bars at 48 hours. Where is any evidence presented for the role of segment 2/PB1? There is no change in Figure 3B.

      The referee is correct, calling out seg2 here was an error; we have revised the text (line 144).

      Segment 3 is overlooked in Figure 3 (as highlighted in Q9 and 10), and shows no difference in Figure S4.

      Please see response to Q8; we think segment 3 contributes via PA adaptation, not via PA-X.

      (12) In line 146 "... drift adaptation in segment 3". Make it clear here that you are talking about genetic drift. However, is this likely to be genetic drift? The 4 amino acid mutations are shown to have a significant impact on polymerase activity in Figure 3B, and in Figure 3C, PA potentially contributes to the reduced kinetics. When there are amino acid mutations that correspond to a beneficial phenotypic change, attributing this to drift alone rather than host adaptation is strange.

      Yes, wording clarified (line 145). “Drift” was used to distinguish it from reassortment but we agree this was an incorrect term in the context.

      (13) Figure 4A MAT-C cells: this is ostensibly the same experiment as Figure 3C in terms of the Cattle Texas and AIV07 viruses. If this is the case, how can you explain the difference in kinetics and overall titer? In Figure 3C, Cattle Texas reaches 10^6, and in Figure 4A it reaches almost 10^9. That's almost 3 log difference. Similarly, in Figure 3C Cattle Texas reaches 10^3, but in Figure 4A it reaches 10^6, a 3-log difference. At 24 hrs, they have roughly a 3-log difference between them in Figure 3C, but in Figure 4A this difference is much smaller. As far as I can tell, these are the same viruses, same dose and same cell model. The t0 titer is also vastly different between the two experiments.

      The experiments were done at different times (several months apart, so different cell passage numbers and/or serum batches) and by different people. We have no explanation other than biological variability. However, both groups of experiments include genuine biological replicates done over the course of 2-3 weeks, so in our view represent coherent tests within themselves.

      (14) In Figure 4, why weren't the a-2,6 a-2,3 proportions analysed for the explants and/or used for the co-infection experiment? It showed the greatest difference between the Cattle Texas and precursor viruses in Figure 4A.

      Our data on the proportion of 2,6 and 2,3 SA in bovine udder tissue are now available in a separate preprint (now cited as [25] in our MS). We did not use the explants for co-infection experiments because it would have been technically difficult to read out the outcome by flow cytometry.

      (15) Figure 4D requires a supplementary figure demonstrating the gating strategy, including one of the samples as an example.

      We have compiled a figure of this and added it as new Figure S13 (called out line 556).

      (16) In Figure 4D, why was an MOI of 5 chosen instead of the MOI of 0.01 used throughout the article for MAC-T cell infection? An MOI of 5 (so in a co-infection, a total of 10 virus particles per cell) is completely overloading the cells. At this dose, 10% of the cells were able to be co-infected, but how representative is this of a real co-infection scenario? While it demonstrates it is possible, it potentially remains highly unlikely, similar to the discussion around the swine respiratory tract in lines 235-240.

      We had also performed the co-infections at lower MOI (1) with very similar results – this is now included in Figure 4D. Furthermore, we redid the experiments at a lower MOI of 0.05 and still see co-infection; this now replaces the MOI 5 data in Figure 4D. The text has been revised accordingly (lines 202-205)

      (17) In Figure 4D, why were immortalised cells used when primary bMEC and mammary explants are available? Primary cells would provide more convincing evidence for the potential of the cow udder to be a mixing vessel. Considering that throughout the paper, a 10x higher viral dose is used in the bMEC culture, I wonder if you would need a significantly higher MOI than 5 to produce similar results in a co-infection experiment. The bMEC also has a more even a-2,3 to a-2,6 ratio compared to MAC-T in Figure 4C.

      The bMECs in the original figure are primary cells. In response to other queries, we now include data from the immortalised MAC-T cell line as well.

      Reviewer #2 (Recommendations for the authors):

      Figure 1A, the coloured underline to discriminate continuous and primary cells is lost upon printing... perhaps another way is better?

      We have changed the primary cells to italic text to make the distinction clearer.

      We have made some other minor changes to wording to correct grammatical errors or improve clarity as we went through.

    1. Reviewer #2 (Public review):

      Summary:

      The authors use simulations and empirical data fitting in order to demonstrate that informing a decision model using noisy single-trial estimates of an underlying fixed non-decision time can guide the model to more reliable parameter estimates, especially when the model has collapsing bounds.

      Strengths:

      The paper is well written and motivated, with clear depth of knowledge in the areas of neurophysiology of decision-making, sequential sampling models, and in particular, the phenomenon of collapsing decision bounds.

      Two large-scale simulations are run to test parameter recovery, and two empirical datasets are fit and assessed; the fitting procedures themselves are state-of-the-art, and the study makes use of a very new and well-designed ERP decomposition algorithm that provides single-trial estimates of the duration of diffusion; the results provide inferences about the operation of decision bound collapse - all of this is impressive.

      Weaknesses:

      This is an interesting and promising idea, but a very important issue is not clear: it is an intuitive principle that information from an external empirical source can enhance the reliability of parameter estimates for a given model, but how can the overall BIC improve, unless it is in fact a different model?

      Comment on revised version.

      Thanks to the authors for their responses and inclusion of additional analyses and simulations. Thanks, in particular for clarifying a crucial detail, that the ndt-informed model actually assumes, like the uninformed model, that there is no variability in the non-decision time, and the idea is that the variable single-trial measurements of non-decision time are noisy estimates of an underlying, constant ndt. The revised paper itself has not made this clear - for example, throughout the Intro, there is no statement that the behavioural model assumes a trial-invariant ndt, and line 231 still calls tau the 'mean' non decision time, implying there is a distribution rather than an invariant single value in the behavioural model.

      One implication of the above is that if the lognormal sigma is purely measurement noise that does not relate to actual variation in the underlying decision process generating behaviour, then the HMP latencies should not relate to behaviour, e.g. shorter latencies predicting shorter RT. I assume that even if the authors did find such a relationship, the principle still stands that a model with fixed ndt is more accurately fit when there are single-trial ndt estimates whose mean provides a constraint on that ndt value, than without such measurements. Still, given ndt variability is a core feature of many decision models, the authors could comment on whether the strategy would work in theory for a model with ndt variability (in the behavioural part), where the single trial estimates would then presumably reflect a mix of measurement noise and genuine ndt variability.

      Another more important implication is that since it is in fact the same model being compared with and without the HMP data guiding the fixed ndt estimate, the reason the fit quality improves with HMP-information is not because it is a better model per se (it is the same model) but because without the HMP guidance, the search algorithm somehow gets lost and fails to find the 'optimal' parameter vector. That is, the parameter vector (just the parameters that relate to the behavioural model itself, not the HMP lognormally-distributed noise associated with VEP measurements) identified as optimal in the HMP-informed version of the model exists in the parameter space of the model without HMP information, but it is just not found? I raised this implication before, and it is still not clear whether it applies. I'm sorry to press on it, but it is critical for readers to understand why it is that neural information can improve overall model fit. Again, the enhancement of parameter recovery (like in Nunez 2025) makes sense, but the enhancement of the "model's fit to behavioural data" does not, without pointing to a deficiency in the search algorithm / fitting procedure.

      The authors state in their replies that the onset of bound collapse is set at accumulation onset and imply that setting it instead at stimulus onset could "mathematically resolve the issue" but they don't do it because it is implausible. It is in fact not only plausible but clearly evidenced in empirical data - collapsing bounds are implemented neurally through urgency signals, and these can begin to dynamically build toward threshold well before, let alone at, stimulus onset. There is nothing bizarre about this - we can prepare movements without sensory input, and indeed even if choosing actions based on a sensory discrimination, motor preparation can launch well before the sensory evidence (e.g. Stanford, Salinas et al 2010) and this in effect collapses the bound on cumulative evidence for triggering action before any evidence actually arrives. So, Urgency/bound-collapse does not need to be triggered by a stimulus; it can start in anticipation of the stimulus. It seems critical, therefore, for the authors to clarify this point - does re-defining the onset of the collapse at stimulus onset remove the trade-off and render unnecessary the neurally-informed ndt estimation?

      Related to this, it is still not clear how bias in the estimation of nondecision time would not be a problem. What if, for example, it is the end of the N2 rather than the peak of the N2 that marks accumulation onset, and/or there is an additional fixed motor time that adds to the N2-based marker to make the full nondecision time that applies in the underlying decision process. By definition (and I think this is essentially what the authors' new simulations verify), because of the trade-offs, this bias would simply be absorbed in shifted estimates of theta and lambda describing the bound collapse function. But wasn't the whole point of the exercise to more accurately estimate those parameters? The obvious implication is that the parameter-estimation accuracy of the ERP-informed model is determined by the accuracy with which the proposed ERP marker directly pinpoints the full nondecision time without bias, but this is not at all obvious in the paper as written. Importantly, in the example scenario I describe above where accumulation onsets when N2 ends, there may still be a perfect correlation of N2 peak latency with underlying ndt across trials - they could still be very strongly "linked" statistically, but we can't know what size offset might be involved.

    2. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      This paper proposes a non-decision time (NDT)-informed approach to estimating timevarying decision thresholds in diffusion models of decision making. The manuscript motivates the method well, outlines the identifiability issues it is intended to address, and evaluates it using simulations and two empirical datasets. The aim is clear, the scope is deliberately focused, and the manuscript is well written. The core idea is interesting, technically grounded, and a meaningful contribution to ongoing work on collapsing thresholds.

      Strengths:

      The manuscript is logically structured and easy to follow. The emphasis on parameter recovery is appropriate and appreciated. The finding that the exponential NDT-informed function produces substantially better recovery than the hyperbolic form is useful, given the importance placed on identifiability earlier in the paper. The threshold visualisations are also helpful for interpreting what the models are doing. Overall, the work offers a well-defined, methodologically oriented contribution that will interest researchers working on time-varying thresholds.

      We appreciate the positive and constructive feedback. We have addressed your comments in the revised manuscript, as detailed in our responses to the specific comments below.

      Weaknesses / Areas for Clarification:

      A few points would benefit from clarification, additional analysis, or revised presentation:

      (1) It would help readers to see a concrete demonstration of the trade-off between NDT and collapsing thresholds, to give a sense of the scale of the identifiability problem motivating the work.

      Thank you for this constructive suggestion. We conducted a new simulation study in which we considered the non-decision time as a fixed parameter and estimated the remaining parameters of the collapsing threshold model. In this simulation study, we contaminated the non-decision time with noise at six levels (i.e., 0%, 2%, 4%, 6%, 8%, and 10%). The simulation results showed that increasing the noise level in non-decision time worsens the estimation of the starting threshold and decay rate parameters, indicating a trade-off between non-decision time and collapsing threshold parameters. The results of this simulation are presented in Appendix 1 in the new version of the manuscript.

      (2) Before moving to the empirical datasets, the manuscript really needs a simulation-based model recovery comparison, since all major conclusions of the empirical applications rely on model comparison. One approach might be to simulate from (a) an FT model with across-trial drift variability and (b) one of the CT models, then fit both models to each of the simulated data sets. This would address a longstanding issue: sometimes CT models are preferred even when the estimated collapse in the thresholds is close to zero. A recovery study would confirm that model selection behaves sensibly in the new framework.

      We are grateful for this constructive comment. The revised manuscript includes a model recovery study (see Simulation study 2, in particular Table 1 and the accompanying text). In this simulation, as the reviewer suggested, we generated data from FT-DDM with across-trial variability in drift rate and CT-DDM with hyperbolic and exponential collapsing thresholds. Then, for each generated dataset, we fitted two models and classified the datasets based on goodness-of-fit and estimated decay rate. The results show that incorporating the decay rate, which can be reliably estimated by the NDT-informed modeling framework, significantly improves the precision of model recovery.

      (3) An additional subtle point is that BIC is defined in terms of the maximised log-likelihood of the model for the data being modelled. In the joint model, the parameter estimates maximise the combined likelihood of behavioural and non-decision-time data. This means the behavioural log-likelihood evaluated at the joint MLEs is not the behavioural MLE. If BIC is being computed for the behavioural data only, this breaks the assumptions underlying BIC. The only valid BIC here would be one defined for the joint model using the joint likelihood.

      We thank the reviewer for raising this important methodological point. We agree that, strictly speaking, the behavioral log-likelihood evaluated at the joint maximum likelihood estimates is not guaranteed to be equal to the behavioral maximum likelihood estimate. Therefore, if one were to interpret BIC<sub>Behavior</sub> for joint (NDT-informed) models as a conventional BIC derived from a purely behavioral maximum likelihood fit, this would indeed violate the standard assumptions underlying BIC. Our intention, however, was not to claim that BIC<sub>Behavior</sub> represents a formally valid BIC in the strict information-theoretic sense. Rather, we used it as a diagnostic measure to assess how well the jointly estimated parameters account for the behavioral data relative to the uninformed models. Importantly, in our datasets, the behavioral log-likelihood evaluated at the joint estimates is higher than that obtained from the uninformed models. This suggests that the additional constraint introduced by the non-decision time information helps guide the optimization procedure toward parameter regions that provide a better account of the behavioral data. In other words, uninformed models appear to converge to suboptimal parameter estimates, but the joint modeling framework helps regularize the estimation process and yields better estimates of the optimal parameters. We fully acknowledge that this use of BIC<sub>Behavior</sub> for the joint models departs from standard practice in the cognitive modeling literature. For this reason, we rely primarily on the joint likelihood–based model comparison as the formally valid criterion. The behavioral BIC is reported only to provide additional intuition regarding the goodness of fit on behavioral data and is used exclusively to compare NDT-informed models with their uninformed counterparts under identical evaluation criteria.

      Also, to warn readers about this limitation, we included the following statement in the section where we defined BIC measures:

      “It is worth noting that comparing joint and behavioral models using BIC<sub>Behavior</sub> is uncommon and constitutes a limitation of the model comparison study.”

      (4) Table 1 sets up the Study 1 comparisons, but there’s no row for the FT model. Similarly, Figures 10 and 13 would be more informative if they included FT predictions. This matters because, in Study 1, the FT model appears to fit aggregate accuracy better than the BIC-preferred collapsing model, currently shown only in Appendix 5. Some discussion of why would strengthen the argument.

      In the revised manuscript, we have included the results for NDT-informed FT-DDM in the main text. However, we kept the FT-DDM with drift variability in the appendix, since none of the models in the main text include drift rate variability.

      (5) In Figure 7, the degree of decay underestimation is obscured by using a density plot rather than a scatterplot, consistent with the other panels of the same figure. Presenting it the same way would make the mis-recovery more transparent. The accompanying text may also need clarification: when data are generated from an FT model with across-trial drift variability, the NDT-informed model seems to infer FT boundaries essentially. If that’s correct, the model must be misfitting the simulated data. This is actually a useful result as it suggests across-trial drift variability in FT models is discriminable from collapsing-threshold models. It would be good to make this explicit.

      Regarding the visualization of decay rate estimation in Figure 7, we would like to clarify that, in the data-generating process for the FT model with across-trial drift variability, the decay rate is fixed at zero. That is, unlike the other parameters, there is only a single true value on the x-axis. If we were to present the decay rate recovery using a standard scatter plot (true vs. estimated values), all points would lie vertically above the single true value (zero), resulting in a vertical strip of points. While such a plot would technically be consistent with the other panels, it would not clearly convey the distributional properties of the estimated decay rates, specifically, which values are more likely under model misidentification. For this reason, we chose a density plot to more transparently illustrate the distribution of inferred decay rates when the true generating process is an FT model. We believe this representation more effectively communicates the extent and structure of mis-recovery. To avoid confusion, we have added a clarifying footnote in the revised manuscript explaining why a density plot was used in this specific panel.

      Regarding distinguishing fixed-threshold models from collapsing-threshold models, as suggested in your comment 2, we conducted an additional simulation, and the results showed that when using the NDT-informed diffusion model, we can distinguish between both models with high accuracy. Specifically, we showed that incorporating the estimated decay rate value provided by the NDT-informed modeling approach can significantly improve the model recovery accuracy.

      (6) Given the large recovery advantage of the exponential NDT-informed function over the hyperbolic one, the authors may want to consider whether the results favour adopting the former more generally. Given these findings, I would consider recommending the exponential NDT-informed model for future use.

      Consistent with the reviewer’s argument, we included the following text in the general discussion:

      “Importantly, the exponential collapsing threshold exhibited substantially better parameter recovery and superior model recovery performance, suggesting that this specification may be preferable in future cognitive modeling applications.”

      (7) In Study 2 (Figure 13), all models qualitatively miss an interesting empirical pattern: under speed emphasis, errors are faster than corrects, while under accuracy emphasis, errors become slower. The error RT distribution in the speed condition is especially poorly captured. It would be helpful for the authors to comment, as it suggests that something theoretically relevant is missing from all models tested.

      Thank you for mentioning this point. Because the models considered in the main text do not include across-trial variability in the starting point, the model cannot predict the fast error pattern observed in the speed condition of the second study. In the revised manuscript, we included a note on this point:

      “The NDT-informed models’ predictions are depicted in Figure 13. This figure shows that the FT-DDM overestimates the last RT quantiles for both correct and incorrect responses in both speed and accuracy conditions. However, the qualitative predictions of CT-DDMs align more closely with the empirical data. It is also worth noting that all the considered computational models misfit the incorrect responses in the speed condition. This misfit is to be linked to the presence of fast errors, specifically in the speed condition. Including the starting-point variability parameter in the model enables the model to predict fast errors. However, as the aim here was not merely to fit the data with the best possible model, but to test the NDT-informed modeling framework, we did not include starting-point variability in the model.”

      (8) The threshold visualisations extend to 3 seconds, yet both datasets show decisions mostly finishing by 1.5 seconds. Shortening the x-axis would better reflect the empirical RT distributions and avoid unintentionally overstating the timescale of the empirical decision processes.

      In the new version of the manuscript, we shortened the x-axis in these plots. See Figures 9 and 12 in the new version of the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript should explicitly state how critical the log-normal assumption for NDT is, and whether there are caveats if it doesn’t hold.

      We agree with this suggestion. Therefore, we conducted an additional simulation study in which we assumed a normal distribution for non-decision time measurements and replicated the main results reported in simulation study 1 (see Appendix 8 in the new version of the manuscript). These simulation results reveal that independent of distributional assumptions on non-decision time measurements, constraining non-decision time can improve the estimation of the collapsing threshold.

      (2) On page 14, one paragraph refers to five models and another to six. It is unclear which is correct.

      We are sorry for the confusion. In the new version of the manuscript, we addressed this issue.

      (3) In the Discussion: Hawkins & Heathcote (2021) found that NDT estimates in the TRDM recover well, but the timer-offset parameter does not (and hence is set to a fixed value). It would be interesting to test whether the NDT-informed approach could extend to that parameter. The TRDM can also generate error RT distributions that are faster than correct RT distributions, which is the qualitative pattern that none of the present models capture in the speed condition of Study 2.

      We appreciate the reviewer’s suggestion, as it would be another important demonstration of the framework we built in the paper. Nevertheless, given the number of models already considered in the manuscript and the focus on collapsing threshold diffusion models, we believe that this addition would blur the focus of the present manuscript. Therefore, we have decided not to include TRDM in the manuscript. Nevertheless, we have addressed the reviewer’s concern regarding the non-decision time estimation in the TRDM in the revised manuscript:

      “Notably, this model can also predict faster error responses than the correct response, the pattern that is observed in the speed condition of Study 2. However, the authors reported poor parameter recovery for the onset of the timing process (Hawkins and Heathcote, 2021). Thus, the reliability issue here specifically concerns the estimation of the shift parameter of the timing accumulator. Informing the model with external estimates of non-decision time might, therefore, improve parameter recovery in this model as well.”

      Reviewer #2 (Public review):

      Summary:

      The authors use simulations and empirical data fitting in order to demonstrate that informing a decision model on estimates of single-trial non-decision time can guide the model to more reliable parameter estimates, especially when the model has collapsing bounds.

      Strengths:

      The paper is well written and motivated, with clear depth of knowledge in the areas of neurophysiology of decision-making, sequential sampling models, and, in particular, the phenomenon of collapsing decision bounds.

      Two large-scale simulations are run to test parameter recovery, and two empirical datasets are fit and assessed; the fitting procedures themselves are state-of-the-art, and the study makes use of a very new and well-designed ERP decomposition algorithm that provides single-trial estimates of the duration of diffusion; the results provide inferences about the operation of decision bound collapse - all of this is impressive.

      We appreciate your feedback and comments. Below, we provided a response for each comment.

      Weaknesses:

      (1) This is an interesting and promising idea, but a very important issue is not clear: it is an intuitive principle that information from an external empirical source can enhance the reliability of parameter estimates for a given model, but how can the overall BIC improve, unless it is in fact a different model? Unfortunately, it is not clear whether and how the model structure itself differs between the NDTinformed and non-NDT-informed cases. Ideally, they are the same actual model, but with one getting extra guidance on where to place the tau and/or sigma parameters from external measurements. The absence of sigma (non-decision time variance) estimates for the non-NDT-informed model, however, suggests it is different in structure, not just in its lack of constraints. If they were the same model, whether they do or do not possess non-decision time variability (which is not currently clear), the only possible reason that the NDT-informed model could achieve better BIC is because the non-NDT-informed model gets lost in the fitting procedure and fails to find the global optimum. If they are in fact different models - for example, if the NDT-informed model is endowed with NDT variability, while the non-NDT-informed model is not - then the fit superiority doesn’t necessarily say anything about an NDT-informed reliability boost, but rather just that a model with NDT variability fits better than one without.

      To respond to this comment, we would like to note that the structural difference between NDT-informed and uninformed models is the assumption about non-decision time. In principle, the behavioral parts (i.e., parameters related to the diffusion part) of both NDT-informed and uninformed models are identical. However, during the estimation, the non-decision time in the NDT-informed model is subject to an additional constraint imposed by the neural data. In other words, in the uninformed model, non-decision time is estimated using the behavioral data by maximizing the likelihood of a CT-DDM. However, the NDT-informed model incorporates an additional data type, resulting in a different likelihood function (see Equation (3)). Specifically, in the NDT-informed model, we make an additional assumption regarding the non-decision time: it is set to the mean of the trial-level non-decision time measurements distribution derived from neural data (Equation (3) specifies the joint model structure). This additional assumption constrains the search space of the non-decision time parameter in the NDT-informed model. As in the main text, we assumed that non-decision time measurements obtained from neural data are log-normally distributed, where sigma is the shape parameter of the log-normal distribution. Therefore, sigma can represent the variability of the non-decision time measurements, and it is not the trial-to-trial non-decision time variability parameter. Indeed, the non-decision time parameter is fixed across trials in both models. Also, it should be clear from Equation (3) that sigma only appears in the second term of the joint likelihood, which corresponds to non-decision time measurements and not the CT-DDM term. Therefore, the behavioral parts of the NDT-informed models and Uninformed models are identical, and the NDT-informed models include one additional parameter (i.e., sigma) corresponding to the variability in non-decision time measurements.

      Also, to explain how constraining non-decision time improves the BIC, we would like to clarify that constraining non-decision time using an additional data source constrains the search space for non-decision time, thereby leading to better parameter identification and, consequently, a better fit to behavioral data. We included the following text in the discussion section to make this explicit.

      “This improvement likely reflects more accurate parameter estimation enabled by the additional information. In other words, constraining the non-decision time using neural measurements led the optimizer to estimate the CT-DDM parameter more accurately and, as a result, improve the fit to empirical data.”

      (2) One reason this is unclear is that Footnote 4 says that this study did not allow trial-to-trial variability in nondecision time, but the entire premise of using variable external single-trial estimates of nondecision times (illustrated in Figure 2) assumes there is nondecision time variability and that we have access to its distribution.

      We are sorry for the confusion. To respond to this comment, we would like to highlight that this modelling approach does not include any mechanism for across-trial variability in the behavioral part, as the likelihood of the choice behavior (see Equation (3)) does not include any variability parameter. As mentioned before, sigma represents the variability in non-decision measurements. However, the model does not propagate the across-trial variability of the neural data on the behavior side (unlike the models in Ghaderi-Kangavari et al. 2023). In other words, in this approach, none of the diffusion model’s parameters include across-trial variability, and we have only considered and estimated neural data variability in the NDT-informed model. Also, to improve the manuscript’s coherence, we removed this footnote.

      (3) It is good that there is an Intro section to explain how the tradeoff between NDT and collapsing bound parameters renders them difficult to simultaneously identify, but I think it needs more work to make it clear. First of all, it is not impossible to identify both, in the same way as, say, pre- and postdecisional nondecision time components cannot be resolved from behaviour alone - the intro had already talked about how collapsing bounds impact RT distribution shapes in specific ways, and obviously mean (or invariant) NDT can’t do that - it can only translate the whole distribution earlier/later on the time axis. This is at odds with the phrasing “one CANNOT estimate these three parameters simultaneously.” So it should be first clarified that this tradeoff is not absolute. Second, many readers will wonder if it is simply a matter of characterising the bound collapse time course as beginning at accumulation onset, instead of stimulus offset - does that not sidestep the issue? Third, assuming the above can be explained, and there is a reason to keep the collapse function aligned to stimulus onset, could the tradeoff be illustrated by picking two distinct sets of parameter values for non-decision time, starting threshold, and decay rate, which produce almost identical bound dynamics as a function of RT? It is not going to work for most readers to simply give the formula on line 211 and say ”There is a tradeoff.” Most readers will need more hand-holding.

      We are grateful for this comment. In response to this comment, which was also partly mentioned by the first reviewer, we first highlight that, in the presence of a nonlinear collapsing threshold, the effect of non-decision time is no longer linear, as it forces the threshold to take a specific value at the final stopping point. To illustrate the tradeoff between imprecise non-decision time estimation and collapsing threshold estimation, we conducted a simulation study. The results for the simulation study are presented in Appendix 1. In this study, we contaminated the non-decision time with noise at six levels (i.e., 0%,2%,4%,6%,8%, and 10%). The simulation results showed that increasing the noise level in non-decision time worsens the estimation of the starting threshold and decay rate parameters, indicating a trade-off between non-decision time and collapsing threshold parameters.

      Regarding the second point, we would like to clarify that in this paper, consistent with other works on collapsing threshold, we assumed that the collapsing starts with evidence accumulation and not with stimulus onset, as the decision makers need a short amount of time for perceiving and encoding the stimulus (i.e., perceptual encoding time). Fixing the collapsing onset to the stimulus onset introduces an ad hoc assumption into the model (that the encoding time is zero), which is cognitively implausible. Therefore, although this assumption can mathematically resolve the issue, it is not cognitively plausible. However, one important point we did not consider in the previous version is the two-stage accumulation process models in which the collapsing onset occurs later than the evidence-accumulation onset. We have discussed the estimation of such models as a limitation in the general discussion:

      “The collapsing threshold dynamics considered in this work (e.g., exponential and hyperbolic) impose a monotonically decreasing threshold over time. Although these dynamics are theoretically well motivated (Fudenberg et al., 2018; Frazier and Yu, 2007) and have been employed in several previous studies (e.g., Olschewski et al., 2025; Milosavljevic et al., 2010; Voskuilen et al., 2016), some research has proposed delayed-collapsing threshold models (e.g., Diederich and Oswald, 2016), in which the onset of threshold collapse does not coincide with the onset of evidence accumulation. Estimating such delayed-collapsing models may require more than simply constraining non-decision time, as the onset of threshold collapse must also be identified. A promising approach for addressing this challenge is the HMP method, which may provide additional temporal information about distinct cognitive processing stages. In particular, HMP may allow the onset of threshold collapse to be estimated as a separate cognitive stage. Future research should therefore investigate the estimation of multi-stage evidence accumulation models (e.g., Diederich and Oswald, 2016; Diederich and Colonius, 2021) within the HMP framework.”

      (4) A lognormal distribution is used as line 231 says it “must” produce a right-skew. Why? It is unusual for non-decision time distribution to be asymmetric in diffusion modeling, so this “must” statement must be fully explained and justified. Would I be right in saying that if either fixed or symmetrically distributed nondecision times were assumed, as in the majority of diffusion models, then the non-identifiability problem goes away? If the issue is one faced only by a special class of DDMs with lognormal NDT, this should be stated upfront.

      We would like to clarify that this assumption is about the non-decision time measurements and not about the across-trial variability parameter of non-decision time. Although for computational convenience, non-decision time is often assumed to follow a normal distribution in diffusion modelling literature, some empirical studies have shown that the perceptual encoding time and total approximated non-decision time follow a right-skewed distribution. Moreover, as HMP estimations of non-decision time are usually right-skewed, we formalized the model with a log-normal distribution. However, we would like to note that the identifiability issue in collapsing-threshold diffusion models is not related to the distributional assumption over the non-decision time measurements, as poor parameter recovery of the collapsing threshold was also reported in Evans et al. (2020). To show that the distributional assumption about the non-decision time measurements does not affect the results, we conducted an additional simulation study in which we assumed a normal distribution for non-decision time measurements and showed that constraining non-decision time improves parameter recovery for collapsing threshold parameters. The results for this simulation are presented in Appendix 8 in the new version of the manuscript. We also revised line 231 as follows (see line 236 in the new version):

      “Empirical studies on non-decision time measurement usually have reported a right-skewed distribution for their measurements (e.g., Weindel, 2021; Weindel et al., 2025). For instance, the measured perceptual encoding time and motor execution time reported by Weindel et al. (2025) are right-skewed. Therefore, we assume that the observed non-decision time measurements Z<sub>n</sub> follow an approximate log-normal distribution, which is right-skewed (we will discuss how this distributional assumption can affect the results later). Thus, we model the non-decision time measurements Z<sub>n</sub>, using a log-normal distribution with parameters µ and σ<sub>z</sub>. The available measurements (i.e., observed data) at trial n can be represented as

      follows:”

      (5) In the simulation study methods, is the only difference between NDT-informed and non-informed models that the non-NDT-informed must also estimate tau and sigma, whereas the NDT-informed model “knows” these two parameters and so only has the other three to estimate? And is it the exact same data that the two models are fit to, in each of the simulation runs? Why is sigma missing from the uninformed part of Figure 4? If it is nondecision time variability, shouldn’t the model at least be aware of the existence of sigma and try to estimate it, in order for this to be a meaningful comparison?

      As mentioned in the response to your first comment, the difference between NDT-informed and uninformed models is the access to an additional source of data related to non-decision time, and sigma represents the shape parameter of the non-decision time measurements distribution. Therefore, sigma belongs only to the NDT-informed model, and, as in the uninformed model, there is no additional data source, so the model does not include sigma.

      (6) I am curious to know whether a linear bound collapse suffers from the same identifiability issues with NDT, or was it not considered here because it is so suboptimal next to the hyperbolic/exponential?

      Thank you very much for this comment. The main reason we did not include linear collapsing threshold models was the assessment by Evans et al. (2020), which indicated that, with sufficient trials, these models can be estimated reasonably well. However, to investigate whether constraining non-decision time can also improve the estimation of linear collapsing threshold models, we conducted an additional simulation study, which is reported in Appendix 3 in the new version of the manuscript. The simulation results confirm those reported by Evans et al. (2020) and show that the parameters of the uninformed linear models can be identified using more than 500 trials. Constraining the non-decision time using the NDT-informed diffusion modeling framework still improves parameter estimation in the linear collapsing threshold model and reduces the required number of trials for reliable estimation to 250. Especially, the estimation of the starting threshold improves significantly.

      (7) The approach using HMP rests on the assumption that accumulation onset is marked by the peak of a certain neural event, but even if it is highly predictive of accumulation onset, depending on what it reflects, it could come systematically earlier or later than the actual accumulation onset. Could the authors comment on what implications this might have for the approach?

      Thank you for mentioning this point. We first would like to point out that Weindel et al. (2024) showed that HMP can predict the underlying generative distribution of cognitive states with high precision and without systematic bias. Second, it is worth highlighting the results reported in Appendix 5 (i.e., “Bias in non-decision time measurements”). In this appendix, we discussed how bias in non-decision time measurements (i.e., systematic underestimation or overestimation) affects the estimation of the collapsing threshold. Particularly, see Figure 3 in Appendix 5. To make these results clearer in the main text, we included the following paragraph at the end of the results section in simulation study 1:

      “Additionally, we examined the effect of systematic bias in non-decision time measurement on parameter recovery. Appendix 5 presents the parameter recovery simulation results in the presence of biased non-decision time measurement (i.e., systematically underestimated or overestimated). The results revealed that, even in the presence of biased non-decision time measurement, the actual generating parameters show a high correlation with the estimated parameters. Underestimation in non-decision time leads to overestimation in the starting threshold and decay rate. Conversely, overestimation in non-decision time leads to underestimation in both the starting threshold and non-decision time.”

      (8) Figure 7: for this simulation, it would be helpful to know the degree to which you can get away with not equipping the model to capture drift rate variability, when the degree of that d.r. variability actually produces appreciable slow error rates. The approach here is to sample uniformly from ranges of the parameters, but how many of these produce data that can be reasonably recognised as similar to human behaviour on typical perceptual decision tasks? The authors point out that only 5% of fits estimate an appreciable bound collapse but if there are only 10% of the parameter vectors that produce data in a typical RT range with typical error rates etc, and half of these produce an appreciable downturn in accuracy for slower RT, and all of the latter represent that 5%, then that’s quite a different story. An easy fix would be to plot estimated decay as a scatter plot against the rate of decline of accuracy from the median RT to the slowest RT, to visualise the degree to which slow errors can be absorbed by the no-dr-var model without falsely estimating steep bound collapse. In general, I’m not so sure of the value of this section, since, in principle, there is no getting around the fact that if what is in truth a drift-variability source of slow errors is fit with a model that can only capture it with a collapsing bound, it will estimate a collapsing bound, or just fail to capture those slow errors.

      Thank you for this comment. We would like to first note that the aim of this section is to illustrate that NDT-informed modeling enables us to distinguish between CT-DDM and FT-DDM with drift variability (i.e., the two competing models that are relatively hard to distinguish). Therefore, we changed the name of this section to “Simulation study 2: Model recovery”. In the new version of the manuscript, we included a cross-model fitting simulation and a model recovery simulation in this section. The estimated decay rates in the cross-model fitting study indicate that, when the underlying generating model is FT-DDM with drift variability, parameter estimation using the NDT-informed approach yields precise inference. This is due to the estimated decay rate, which is very close to zero. The model recovery results also confirmed that incorporating the decay rate value into the model inference improves precision.

      Moreover, to address the concern regarding the slow-error pattern in the simulated data, we examined the relationship between the slow–fast accuracy difference and the estimated decay rate. Specifically, we computed the difference between the accuracy of responses with response times below the median (ACC1) and those above the median (ACC2), and plotted this difference against the estimated decay rate. Author response image 1 presents the resulting scatter plot, with color indicating the drift rate. As shown in Author response image 1, incorrect inferences about the decay rate primarily occur at high drift rates. This pattern emerges because high drift rates produce very fast responses, leaving little time for the threshold to meaningfully collapse. Consequently, the behavioral signatures of collapsing-threshold and fixed-threshold models become increasingly similar under high drift conditions.

      Author response image 1.

      Illustration of the difference between the accuracies of responses with response time below the median and those with response time above the median against the estimated decay rate. Colour shows the drift rate value.

      Reviewer #2 (Recommendations for the authors):

      (1) Abstract: improves fit to behaviour in what way? Reliability or absolute quant fit to behavior? I.e., is it just helping constrain it so it finds the global opt?

      (2) Line 72 - Are these “neuroimaging” studies? Perhaps use the broader term “neuroscience.”

      (3) Line 95 - It’s important because it may confuse readers how something dynamic like a collapsing bound could be resolved with, say, fMRI.

      (4) Line 85 – “greater variability”... in what?

      (5) Line 87 -Revise to avoid misconstruing a time on task effect - e.g., “higher error rate for trials with longer RT” is more explicit.

      (6) Line 89-94 - It’s not clear what findings are being referenced here.

      (7) Line 103 - External biases such as priors and relative value?

      (8) Line 175 - Clarify this applies to any ddm, not just ctddm.

      (9) Line 268 - Explain what portion N200 accounts for.

      (10) Line 327 - Is this equation supposed to be for Delta-X, as opposed to X(t+delta-t)? If you want X(t+delta-t) on the LHS, then X(t) must be added to the RHS.

      We are very thankful for such a precise evaluation and constructive comments. We addressed all the comments raised by the reviewer in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The current paper addresses an important issue in evidence accumulation models: many modelers implement flat decision boundaries because the collapsing alternatives are hard to reliably estimate. Here, using simulations, the authors demonstrate that parameter recovery can be drastically improved by providing the model with additional data (specifically, an EEG-informed estimate of nondecision time). Moreover, in two empirical datasets, it is shown that those EEG-informed models provide a better fit to the data. The method seems sound and promising and might inform future work on the debate regarding flat vs collapsing choice boundaries. As an evidence-accumulation enthusiast, I am quite excited about this work, although for a broader audience, the immediate applicability of this approach seems limited because it does require EEG data (i.e., limiting widespread use of the method or e.g., answering questions about individual differences that require a very large N).

      We are very grateful for your positive evaluation and your comments.

      Reviewer #3 (Recommendations for the authors):

      This is a very decent study, very well written and properly executed. Most of my comments below are suggestions for the authors as to how to make the story more compelling.

      (1) I think the authors can do more to explain why the NDT-informed models fit better to empirical data. If the NDT-estimates are equal to the ground truth, then isn’t it possible that such models win simply because they have one parameter less? However, this is not what the authors are claiming, though, on l.567 it says that the fit is better because parameters are better estimated. I think it should be possible to dissociate these two accounts.

      Thank you for this important comment. For clarification, both NDT-informed and uninformed models estimate the non-decision time parameter, and neither treats it as equal to the ground truth. However, the difference between these two models is that, in the NDT-informed model, the non-decision time is subject to an additional constraint; therefore, the number of behavioral parameters (i.e., parameters of the diffusion part) is identical for both NDT-informed and uninformed collapsing threshold models. Although the number of behavioral parameters is identical in the NDT-informed model, one parameter is constrained, thereby limiting the model’s complexity and flexibility compared to uninformed models. Consequently, the improvement in the fit can only be attributed to better parameter estimation in the NDT-informed models, resulting from constraining the non-decision time using neural measurements. To clarify that in the manuscript, we included the following text in the revised version:

      “This improvement likely reflects more accurate parameter estimation enabled by the additional information. In other words, constraining the non-decision time using neural measurements led the optimizer to estimate the CT-DDM parameter more accurately and, as a result, improve the fit to empirical data.”

      (2) Given that so much of the writing focuses on the conclusion that collapsing boundary models ¿ flat models, it is very odd that there is no comparison to a flat boundary model reported in the text. Why is the model in Supplement 5 not just included in the main text (and in Table 1)? This would make it so much easier for the reader.

      To address this comment and also the similar point raised by Reviewer 1, we included the NDT-informed FT-DDM in the results section and compared the FT-DDM with CT-DDMs with respect to BIC<sub>Joint</sub>. However, we retained the other FT-DDM in the appendix because this model includes across-trial variability in drift rate, whereas the models in the main text do not.

      (3) I would have appreciated a bit more background about the importance of the number of trials per participant. Given that the proposed method requires collecting EEG data (which is time and labour-intensive) I wonder to what extent you get similar improvements in parameter recovery by collecting more data per participant (which is usually cheap). Put differently, I would appreciate an additional matrix in Figure 5 for vanilla models.

      The new version of Figure 5 in the revised manuscript now contains the goodness of parameter recovery for the uninformed CT-DDMs for different numbers of trials. As the simulation results suggest, even with 1000 trials, the parameters of the uninformed CT-DDMs are still not reliably identifiable. Therefore, these results suggest that although increasing the number of trials can slightly improve the parameter recovery of the CT-DDMs, it cannot fully resolve the reliability issue in their parameter recovery.

      In addition, it is worth clarifying that, although the paper focuses on extracting non-decision time from the EEG signal using the HMP method, as discussed in the general discussion, the NDT-informed approach can also be employed with purely behavioral methods.

      (4) L116-117: minor detail: I don’t think it’s fair to write that it’s an open question whether or not thresholds collapse. I think it’s fair to say that this is hard to show, and that the conditions under which it appears are unclear; but saying that it’s still unclear whether this occurs at all seems unfair with regard to previous work.

      Thank you for mentioning this point. We revised the mentioned sentence as follows in the new version of the manuscript:

      “This issue is particularly critical given that the conditions under which individuals adjust their decision thresholds during a single trial remain an open question.”

      (5) Figure 5: Minor detail: It would be useful to mention in the figure or caption how many trials underlie these simulations, which is now somewhat buried in the text.

      As Figure 5 shows sensitivity to the number of trials, and the number of trials is explicitly mentioned in the figure, we assume that the reviewer intended Figure 6. In the revised manuscript, we explicitly specify the number of trials for the results reported in Figure 6. We simulated 500 trials for each noise level in this graph.

      (6) Figure 10 and similar figures: I tried to figure out how corrects and incorrects differ, but couldn’t see the difference. Can the authors use something more colorblind friendly (hope I didn’t give up on being anonymous here)?

      We are so sorry for the inconvenience. In the revised manuscript, we used different symbols to distinguish correct and incorrect data.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study investigates the peptide-binding principles of promiscuous chicken MHC molecules. The data from crystallography, mass spectrometry, and modeling are convincing. However, the presentation would benefit from streamlining and clear links between data and conclusions. This paper will be of broad interest to immunologists and those interested in vaccine development.

      Overall, we are delighted and grateful to the eLIFE editors and the two reviewers for the careful and thoughtful assessments and reviews of our paper. We are glad that the strengths of the paper were apparent and appreciated. And of course, every paper has weaknesses, especially for a story as complex as this one.

      We made only minor changes to accommodate the reviewer comments, along with additions for which we only became aware upon this submission of a revised manuscript. In particular, we shortened the title and abstract to fit what is usual for an eLIFE paper, added Key Resources table with accompanying references, changed the numbering of the figures throughout the manuscript to ensure that each page represented a figure (rather than panels of a figure), moved the figure legends from the embedded figures to a list near the end of the manuscript, and split the supplemental spreadsheet into two renamed Data Source files.

      Before answering the comments and questions directly, perhaps a few points would help clarify why the paper is as it is.

      First, the experiments cover over three decades of work, with the first gas phase sequencing results done in 1992. Unlike some of the chicken class I alleles which immediately gave completely clear stringent motifs (B4, B12 and B15 in Wallny et al 2006 PNAS, B19 in Han et al 2023 J Immunol), we harvested nothing but confusion from the B21 class I results (Fig. 1). Initially, we thought that the lack of a clear motif for B21 was due to multiple well-expressed class I molecules but only one dominantly-expressed class I molecule was found (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol) and, to our surprise, bacterially-expressed BF2*21:01 heavy chain and b<sub>2</sub>-microglobulin refolded with two synthetic peptides without sequence in common, and the crystal structures showed that this molecule remodeled the binding site to accommodate two such disparate peptides (Koch et al 2008 Immunity). This was the beginning of our understanding of the spectrum of class I alleles from promiscuous generalists to fastidious specialists, which we have explored in a series of further papers (in particular, Chappell et al 2015 eLIFE, Tresgaskes et al 2016 PNAS, Kaufman 2018 Trends Immunol, Tregaskes and Kaufman 2022 Mol Immunol).

      Second, over these many years, we continued to explore the binding properties of BF2*21:01 in ever more detail, resulting in the current manuscript. We learned only slowly how to probe this unexpected promiscuity, unprecedented in the MHC literature, so that the experiments proceeded with our best understanding at the time, including taking advantage of new approaches as they become available. Each experiment built on the previous set of experiments and each brought us closer to an understanding.

      Third, having amassed a collection of data, we chose eLIFE exactly because it allows us to present the entire story from beginning to end without compromise, not just the highlights with the major points illustrated by a few main figures and with the supporting data in many supplementary figures. We include all the data, because it is all part of the story, and so interested researchers to look at the data from their own perspective. Although mostly we provide bar graphs, we include spreadsheets for the raw data (or close to them) for the final experiments (illustrated by Figs. 10 and 14-22) in the two source data files, so these can be assessed easily by others in the field, perhaps using approaches that we may not feel competent to perform.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Combining in vitro refolding, SEC-based assembly assays, peptide-library screening, MALDI-TOF, LC-MS/MS, structural analysis and immunopeptidomics, this manuscript investigates the peptide-binding principles of the promiscuous chicken MHC-I molecule BF2*21:01.

      Strengths:

      Although the peptide motif of BF2*21:01 is highly complex, this manuscript identified several principles, including a preference for 10-mer peptides, co-variation between P2 and Pc-2, effects of P3 and Pc-3, and a strong cellular preference for Leu at Pc. The results are important for avian MHC biology and poultry vaccine epitope prediction.

      Weaknesses:

      The manuscript is sometimes difficult to follow because the authors present a large amount of peptide-library, structural and immunopeptidomics data. without always clearly explaining how these datasets support the proposed simplifying principles.

      We are delighted and grateful to the reviewer 1 for the careful and thoughtful comments and questions concerning our manuscript. We are glad that the strengths of the paper were apparent and appreciated, and acknowledge the weaknesses that come with such a complex story with experiments performed over decades.

      Major Issues - Points Requiring Clarification or Additional Support:

      (1) Line 282-301, 537-545)

      The immunopeptidomics conclusions are mainly based on one B21 cell line with one biological replicate and at least two technical replicates. Given the complexity of the BF2*21:01 peptide repertoire, this is a major limitation. The authors should either provide additional biological replicates or clearly state this limitation in the Abstract, Results and Discussion.

      This limitation is clearly stated in lines 537-545, as part of a paragraph covering the various ways in which the data presented in this manuscript could be improved. In fact, we have performed immunopeptidomics of several different B21 cell types, with many replicates and found similar data as presented, giving us confidence in our interpretations. However, these other experiments belong in different stories, so it is not appropriate that the data be reported in this manuscript.

      (2) (Lines 290-313)

      The B21 cell preparations contain both BF2 and the lowly expressed BF1 molecule. Some peptides, especially 8-mers or peptides with atypical motifs, may derive from BF1*21:01. The authors should clarify how BF2*21:01-bound peptides were distinguished from possible BF1-derived peptides, or interpret the immunopeptidomics motif more cautiously. The authors should also provide or cite evidence confirming the B21 haplotype identity of the cell line and chicken materials used for immunopeptidomics.

      The concern about the contribution of BF1*21:01 to the immunopeptidomics is clearly stated in the manuscript, both lines 290-313 and as part of the paragraph describing the limitations of the experiments (lines 542-543). In fact, the expression of BF1 molecules has long been known to be less than 10% of BF2 molecules at the RNA level, and much less at the protein level (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol). The proportion of 8mers identified by immunopeptidomics is also low (Fig. 14), and it is not impossible that most 8mers are due to BF1*21:01. We have used assembly assays with peptide libraries, immunopeptidomics and a crystal structure to determine the peptide motif for typical BF1 molecules, of which BF1*21:01 is one and found it may contribute to 8mer peptides but very seldom to longer peptides. This work is unpublished but gives us confidence that the characteristics of BF2*21:01 are not misrepresented by the data in this manuscript.

      The sources of the chicken samples and the cell lines are described in detail under Materials and Methods (lines 577-590), citing relevant publications.

      (3) (Lines 217-221, 243-253)

      The authors acknowledge that MALDI-TOF cannot reliably distinguish peptide combinations with identical or similar masses, nor determine residue positions in some cases. Therefore, MALDI-TOF results should not be overinterpreted as precise evidence for residue preference. The authors should clearly indicate which conclusions are supported by LC-MS/MS.

      As described, the experiments follow each other in temporal sequence, so that we started with single peptides, then peptide libraries that varied in one position, then peptide libraries that varied in two positions first analysed by MALDI-TOF and later by LC-MS/MS. The final experiment (Fig. 10, with the original data in the supplementary spreadsheet) directly compares MALDI-TOF and LC-MS/MS results for six peptide libraries, so that the strength of the evidence for residue preference is clear. Throughout the manuscript, we do our best to not to overstate conclusions based on the data of any particular experiment.

      (4) (Lines 297-301, 316-330)

      The authors suggest that longer peptides may bulge in the middle or extend out of the groove at the C-terminal end. The rationale for the C-terminal extension is not clearly explained. Why is the C-terminal extension considered rather than the N-terminal extension? If the binding register is uncertain, long peptides should be analyzed separately from canonical-length peptides.

      When the first sequence of a chicken class I cDNA was determined, an immediate mystery was why one of the so-called invariant residues that coordinate the N- and C-termini of the bound peptide is not conserved (Kaufman et al 1992 J Immunol). In fact, this residue Tyr at position 86 in HLA-A2 and the equivalent position in all mammalian classical class I molecules is an Arg in the classical class I molecules of all non-mammalian vertebrates and is common with class II molecules (Kaufman et al 1995 Semin Immunol). Similar to class II molecules, this Arg in chicken class I molecules allows the peptide to extend out of the C-terminus, as shown by a crystal structure (Xiao et al 2018 J Immunol). The concern that we might be misidentifying the C-terminal amino acid was the basis for the analysis in Figs. 23 and 24, but in the absence of crystal structures, we are not able to provide a final answer this question. Perhaps relevant is the fact that a chicken class II molecule can bind exactly the same peptide in two conformations, one with a canonical 9mer core and the other with an unexpected 10mer core (Goryanin et al 2026 J Virol).

      By contrast, N-terminal extensions are only found for some class I alleles and thus far depend on the substitution of small amino acid sidechains for W166 (Li et al 2011 J Virol for bovine, Ma et al 2020 J Immunol for Xenopus, Wei et al 2022 J Immunol for ovine). Thus far, no chicken BF2 sequences have this substitution, consonant with the many crystal structures, including those for BF2*21:01 (Koch et al 2008 Immunity, Chappell et al 2015 eLIFE, this manuscript). However, in unpublished data, we find that most BF1 sequences have sequence differences that could allow N-terminal extensions, although we have no crystal structures to support this possibility.

      (5) (Lines 406-439)

      In vitro assembly assays show that several hydrophobic residues can be tolerated at Pc, whereas immunopeptidomics shows a strong Leu preference at this position. The authors should clarify whether this Leu preference reflects intrinsic BF2*21:01 binding specificity, TAP-mediated peptide transport, antigen processing, peptide loading, or a cell-line-specific effect. Additional experimental support, such as TAP transport analysis, would strengthen this conclusion.

      The preference for Leu at the final position of the peptide by immunopeptidomics of the B21 cell line is strong but not absolute and is certainly affected at the least by the length of the peptide (Figs. 23 and 24). Unpublished immunopeptidomics results (mentioned above) show that this is not a cell line-specific result. The evidence from assembly assays of various peptides is that several hydrophobic amino acids are tolerated with sufficient stability of BF2*21:01 that they are detected in the assay (Figs. 3, 5, 9 and 10). Thermostability assays (Fig. 6) show that peptides with these same hydrophobic amino acids are stable to at least body temperature of chickens. These experiments show that such stability is peptide-dependent (that is, whether a particular amino acid is tolerated depends on the stability conferred by the rest of the peptide). Finally, peptide translocation assays using B21 cells have been done (Tregaskes et al 2016 PNAS) and show that peptides with several hydrophobic amino acids can be pumped into the lumen of the endoplasmic reticulum. However, the assays are with single synthetic peptides, so the data are not extensive enough to separate the effects of the final amino acid from the rest of the peptide. Certainly, peptides with amino acids other than Leu at the C-terminus can be translocated. So, it is not yet clear at which point the preference for Leu at the C-terminus of the peptide arises.

      (6) (Lines 172-178, 243-279, 442-457)

      The structural analysis explains some residue combinations, such as Arg at P2 with Glu at Pc-2 or Trp at Pc. However, the structural interpretation is not fully integrated with the large-scale peptide library and immunopeptidomics results. Representative high- and low-frequency combinations should be discussed structurally.

      Six crystal structures show that BF2*21:02 remodels the binding to accommodate a variety of anchor residues (Koch et al 2008 Immunity, Chappel et al 2015 eLIFE). These crystal structures are representative of sequences found by the immunopeptidomics from very frequent (H-E at roughly 15% 8-12mers) to moderately frequent (E-L at roughly 6% 8-12mers) to infrequent (N-F, A-D and E-D at roughly 1.5%, 1.6% and 0.7% 8-12mers) based on Fig. 18. All but one of the structures has Leu at the C-terminus, with the last one having Val which is found but not frequently by immunopeptidomics.

      Similar numbers are found by LC-MS/MS of double-substitution libraries of the two original peptide sequences in Fig. 10 with H-E found frequently (8.1% in P390, 3.8% in P498) and the others infrequently (0.1, 0.9, 1.0, 0.3% in P390, 0, 1.4, 1.0, 0.3% in P498), as calculated from the numbers in the Supplementary data spreadsheet. As discussed in the manuscript, for single-substitution peptide libraries of the two original peptides, Ile/Leu at the C-terminus was very frequent but at the same or slightly less level as Phe, with Met less frequent and Val even less so (Fig. 7).

      In addition, there are two more structures along with models explicitly testing some substitutions (Fig. 5). Attempting more current modelling approaches, we found AlphaFold 3 was unable to correctly predict most of the conformations that are found in the crystal structures of BF2*21:01, so we don’t feel confident in using them to predict unknown structures of this kind.

      (7) The inference of co-variation between P2 and Pc-2, as well as the modulatory effects of P3 and Pc-3, should be better explained. At present, some conclusions appear to be based mainly on residue-frequency patterns, and the logical connection between these observations and the proposed binding principles is not always clear. Statistical analyses, such as mutual information, chi-square tests or permutation tests, and representative structural explanations would strengthen this conclusion.

      We endeavored to do our best to explain the data, our interpretations and our reasoning, so we apologise if we have not managed to be as clear as might be desired. We have included as close to raw data as possible for the LC-MS/MS and MALDI-TOF (Fig. 10) and for the immunopeptidomics (Fig. 14 and 18) in the Supplementary Data spreadsheet, exactly so that competent practitioners can carry out further analyses (including the sophisticated statistical tests mentioned).

      Reviewer #2 (Public review):

      Summary:

      The study presents an in-depth analysis of the peptide repertoire bound by a promiscuous chicken MHC molecule using mass spectrometry, x-ray crystallography and modelling. While the MHC can bind a very diverse set of peptides, the authors have found some new rules that govern peptide binding to this MHC that could help to build a predictive model to study the repertoire of pathogen-derived peptides.

      Strengths:

      The study uses a range of well performed experiment across multiple techniques and provides an in-depth analysis of the peptide repertoire, including peptide sequences, length, preferred residues, stability and MHC presentation.

      Weaknesses:

      The data overall support the analysis and conclusion well. The only caveat is linked to Figure 4, which does not describe the stability of the peptide-MHC complex, but instead shows refold yield, and the two are not always linked.

      We are grateful for the clear understanding of the strengths of the work. With regards to Fig. 4, we agree with the reviewer that there are differences in refold yield but that measure may not be correlated with stability of the peptide-MHC complex. However, we were basing our interpretation of stability on the position and quality of the monomer peak, as illustrated by the trace in Fig. 2, in which a sharp peak at the monomer position represents a stable complex (as seen for the 10 and 11mer peptides) and later peaks represent unstable complexes falling apart during the chromatography (as seen for the 7, 8 and 9mer peptides).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Issues: Editorial and Data Presentation Modifications

      (1) Lines 53-62, 155-170, 303-314: The terms Pc, Pc-2 and Pc-3 should be clearly defined early in the manuscript and figure legends.

      The Abstract introduces the abbreviations as “peptide positions P<sub>2</sub> and P<sub>c-2</sub>” followed by P<sub>C</sub>, which are standard usage and seem clear. The first usage in the text is “anchor residues at three positions, with co-variation between the anchor residues at P<sub>2</sub> and P<sub>c-2</sub>” which again seems clear, particularly in the context of the text and the figure. However, a parenthetical description has been added to read “…anchor residues at three positions, with co-variation between the anchor residues at P<sub>2</sub> and P<sub>c-2</sub> (position 2 and the position two before the C-terminus) …”. Given these usages, it seems unlikely that the reader will fail to understand P<sub>C</sub> and P<sub>C-3</sub>.

      (2) Lines 255-279: The term "peptide backbone" should be defined as the fixed sequence context outside the randomized positions, if this is what the authors mean. Suggest clarifying the meaning of "peptide backbone".

      The phrase reads “double-substitution libraries based on four peptide backbones”, which in context of the figures seems clear. However, a parenthetical description has been added: “The examination of double-substitution libraries analysed by MALDI-TOF was expanded to other backbones (that is, other sequences in which two positions were randomised): …”.

      (3) Several figures are complex. The authors should add brief take-home messages to figure legends.

      Every figure legend starts with a (sometimes quite long) take-home message. It is not clear what more should be added.

      (4) A concise summary table of the proposed binding rules, including preferred peptide length, P2, Pc-2 and Pc preferences, and the effects of P3/Pc-3, would be useful for readers.

      A concise summary table of binding rules would be very helpful, but the rules are complex, both qualitative and quantitative. For example, immunopeptidomics shows that 10mers are preferred, but that fails to capture the quantitation. The co-variation of P<sub>2</sub> and P<sub>c-2</sub>, which is the strongest and best characterized correlation, nevertheless is complex since the occupancy of P<sub>2</sub> by different amino acids (presumably independently of P<sub>c-2</sub>) varies considerably. At our current level of understanding, it is hard to imagine a table that is both concise and precise. With time and more data, perhaps code for quantitative prediction might be constructed (dare one suggest machine learning…), which is the hope of presenting all the available data in one place.

      (5) The Results section contains a large amount of peptide-library, structural and immunopeptidomics data. The authors should improve the logical flow and add clearer transition sentences to explain how each dataset supports the proposed simplifying principles.

      The reviewer is of course correct that any written text can be improved (although each critic may have a different idea about which part should be fixed), but we have done our best with the material and time available. We could respond productively to a more detailed critique.

      (6) Some statements in the Abstract and Discussion should be softened, especially those related to in vivo peptide preferences, BF2*21:01 promiscuity and peptide prediction, given the limited biological replication and uncertainty in peptide assignment.

      The statements in the both the Abstract and Discussion are very general, reflecting what we believe to be careful interpretations based on the data. We could respond productively to concerns about specific claims.

      Reviewer #2 (Recommendations for the authors):

      Overall, the data presented in this study are interesting; however, it is complex, and some results could be merged and simplified, as well as the figures. The data provide an in-depth analysis, using mass spectrometry, of the interplay between the different positions of the peptide and the residues favourable to bind within the antigen-binding cleft.

      (1) From the abstract, the concept of "promiscuous generalists and fastidious specialists" is not explored after or defined within the results.

      The Abstract introduces the concept of promiscuous generalists and fastidious specialists to provide the basis for exploring the peptide-binding specificity of the most promiscuous class I known, BF2*21:01. This overall concept is described in enormous detail in several publications cited in the current manuscript, but it is not particularly germane to the analyses.

      (2) From the abstract "These simplifying principles may eventually allow predictions of pathogen peptides", I'm not sure how "simplifying" the principles are with the data, if anything, it does show a rather complex interplay between the different residues of the peptide that enable the MHC to bind a large and diverse number of peptides.

      Compared to any combination of anchor residues being permitted at equal frequency, there are clear preferences which the experiments identify. Instead of an enormous range of possible peptide lengths, roughly 50% of peptides are 10mers. Of course, the structural reasons behind these results are certainly complex and likely must be understood in detail in order to attempt peptide predictions. 

      (3) Line 48. "less well-expressed". Less than what? Do you mean the level of expression was lower? And if yes, are there values of fold change for comparison?

      The sentence in the Abstract reads “Chicken BF2 alleles … are less well-expressed on the cell surface … while certain human HLA-B alleles … are well-expressed …”. Read as a complete sentence, the meaning is clear. This difference for chicken BF2 alleles has been quantified as reported in several publications cited in the current manuscript, ranging from 3-5 fold for peripheral blood lymphocytes to ten-fold for erythrocytes (Kaufman et al 1995 Immunol Rev, Chappel et al 2015 eLIFE), with similar numbers for a few HLA-B alleles on human peripheral blood lymphocytes and monocytes (Chappell et al 2015 eLIFE).

      (4) Line 149. "with individual peptides" which peptides are we referring to here?

      This introductory sentence to a paragraph outlines the general method used for the experiments in this section of the Results, “refolding in vitro … with individual peptides.” Which “individual peptides” are described in the following paragraphs, with the next section of the Results using “refolding in vitro … with peptide libraries”.

      (5) Line 170. If 9 mers and below are not stable, which is not really quantified or shown with the data on Figure 4, why is refolding material observed for peptides with different lengths of 9 aa and below on Figure 4? A lower yield of refolded material can have a different origin, and there is no association between stability and refold yield. The notion of stability here probably needs to be changed, as it does not apply to the data.

      As described in our response to a similar concern above, we are not basing our interpretation of stability on the quantity of refolded material, but on the position and quality of the monomer peak, as described clearly in the legend to Fig. 4: “The original 10mer (REVDEQLLSV) and 11mer (GHAEEYGAETL) peptides refold with BF2*21:01 to give stable monomers as do 11mer and 10mer derivative peptides, but 9mer, 8mer or 7mer peptides give heavy chain only peaks” and “The peptides 11mer GHAEEYAETL (top panel), 10mer REVDEQLLSV (middle panel), 11mer GHAEAAAAETL and 10mer GHAEAAAETL (bottom panel) gave sharp monomer peaks, while the 9mer GHAEAAETL, 8mer GHAEAETL and 7mer GHAEETL gave a delayed broad peak indicative of unstable binding or heavy chain.” This concept is illustrated by the trace in Fig. 2 (discussed in the text at the beginning of this section of the Results), in which a sharp peak at the monomer position represents a stable complex, and later peaks represent unstable complexes falling apart during the chromatography or free heavy chains.

      (6) Line 175. As the 3BEV structure had a P2-His, is a comparable structure expected?

      This sentence reads “Structures with amino acid substitutions in the 11mer peptide GHAEEYGAETL bind with Asp at P<sub>c-2</sub> and either His or Arg at P<sub>2</sub> (5AD0 and 5ADZ), comparable to the original structure (3BEV) (Fig. 5A).” Minor changes in positions and orientations of individual amino acid sidechains are expected and are clear from the crystal structures presented in Fig. 5A, but are overall comparable to the original peptide in 3BEV, which has a His at P<sub>2</sub> and a Glu at P<sub>c-2</sub>.

      (7) Line 175 "modelling the substituted". How was the modelling done?

      The sentence reads “Modelling the substituted peptide with Arg at P<sub>2</sub> and Glu at P<sub>c-2</sub> shows a steric clash that can explain why this peptide did not refold with BF2*21:01 (Fig. 5A).” The legend to Fig. 5A states that “modelling done as detailed in Materials and Methods”, but apparently that section was omitted. A section has been added now to the Materials and Methods which states “Beginning with known crystal structures, modelling was carried out using PyMol with the protein mutagenesis Wizard, the rotamer toggle, show bumps and show surfaces.”

      (8) Line 176 "shows a steric clash" with what? Figure 5 is too small to see, and there is no label on the residue to follow where the steric clash is coming from.

      As stated in the legend to Fig. 5, “Structures were determined for GHAEEYGAETL (3BEV), GHAEEYGADTL (5AD0) and GRAEEYGADTL (5ACZ), which all refolded successful to give stable monomers, while GRAEEYGAETL did not (see Fig. 3), all models of which showed steric clashes (one depicted, red arrow).” In the model shown, the clash is between R9 of the BF2*21:01 with the Glu at peptide position 9. Parenthetically, this depiction of the key MHC residues for the co-variation has been used repeatedly, starting with Chappell et al 2015 eLIFE. The picture can be zoomed to make it large enough to see.

      (9) Line 247 "many combinations". It is not clear here if combinations are referring to a set of double-substitutions or different peptides?

      Each peptide in the library has a different double-substitution, so the meaning is the same either way. The point of Fig. 10 is to compare the identification of peptides from LC-MS/MS (in which each peptide is identified exactly) with the identification of sets of peptides from MALDI-TOF (in which the order of the double substitution is not clear, as well as the exact amino acid in the cases of I/L and Q/K). As the figure shows, the numbers are generally very similar, but this sentence describes the percentage of cases for which only one method or the other identified a peptide.

      (10) Lines 252-253. The conclusion of the MS data that the LC-MS/NS is more accurate and sensitive than MALDI-TOF is not very surprising. What was the rationale for using both?

      Our examination of the binding properties of BF2*21:01 for peptides and peptide libraries developed over a long time-span, so in the beginning we looked at single peptides with size exclusion chromatography peaks as the measure, then single- and later double-substitution libraries, first by MALDI-TOF and later by LC-MS/MS. Each set of experiments built on the previous work, so that together they tell the story. For example, after we optimized the use of double-substitution libraries, we were worried about the effects of temperature, so we tested that by MALDI-TOF. While repeating the temperature experiment once we optimized the LC-MS/MS approach might have yielded some additional data, there were other questions to answer.

      (11) Line 272-273 "support the idea that P3 is an important position within the peptide (Figures 8-10) despite not contacting the MHC molecule" The structure of 3BEV clearly shows interaction between the P3-Ala and the Tyr156 of the MHC. Residues that are fully or partially buried in the MHC cleft almost all contact the MHC molecule. Maybe I've missed something, but I think this statement is inaccurate.

      We agree with the reviewer that nearly every peptide residue contacts the MHC molecule (but of course some much more than others). The statement is now changed in the text to read “support the idea that P<sub>3</sub> is an important position within the peptide (Figures 8-10) despite not being an anchor residue."

      (12) Line 294. Figure 14 clearly shows that the number of 8 and 9-mer peptides eluted is at the same level as the 11mer and above, so how does the data fit with the statement that 9mer and shorter peptides are not stable with the MHC?

      Fig. 4 shows that 7, 8 and 9mer derivatives of the original 11mer failed to refold to give a single peak of stable monomers, while Fig. 6 shows that the 10mer derivative of the original 11mer was more thermostable than both the 9mer derivative and the original 11mer. The reason why the 9mer yielded so little monomer in Fig. 4 while giving enough to test by thermostability in Fig. 6 is no longer remembered. A key point is that these results are peptide sequence-specific, so it is not impossible to imagine stable binding of an appropriate 8mer (or perhaps even a 7mer). Another uncertainty, mentioned in the Results and Discussion, is that BF1*21:01 molecules bind primarily 8mers, and the contribution of peptides from BF1*21:01 is not known with certainty.

      (13) Line 316. What was the rationale for choosing peptides > 12aa to see if there is some overhang or bulge? As even 11-12mer can exhibit such features.

      We were just looking for any obvious patterns, but we didn’t find any.

      (14) The section starting at line 366 would have benefited from some structure prediction or modelling to illustrate the findings.

      We would have been delighted to model peptides, but we have used AlphaFold3 to compare the models to our crystal structures for seven chicken class I alleles (including BF2*21:01) with one or more peptides. The models sometimes fit the experimental data but they often didn’t, often by a wide margin, and with BF2*21:01 the worst (presumably because the system is not trained on MHC molecules which utilise charge transfer). Therefore, we do not feel confident in using any such modelling approaches except in the most defined situations (such as illustrated in Fig. 5).

      (15) Typo - Alleles should be italic, and in vitro as well.

      Alleles of genes are in italics, but alleles of proteins are not. To write in vitro in italics is customary, which we have corrected.

      (16) Figures

      (a) Some of the figures could be merged together.

      Of course, any presentation can be improved, but which figures we should merge (some already being three pages in length) is not clear. We could productively respond to more detailed suggestions.

      (b) Figure 1. I can't see the different colours mentioned in the figure legend

      Our apologies if the colours are not clear enough, but they are present only as an aid (as elsewhere in the manuscript), with red D and E, blue H, K and R, green N, Q, S and T, and all other amino acids black.

      (c) Figure 12. Is this figure only with 11mer peptides?

      Figure 12 shows the percentage of peptides with particular amino acids at P<sub>2</sub> and P<sub>c-2</sub> for three 11-mer double-substitution peptide libraries, the sequences of which are written on the graphs and described in the figure legends.

      (d) Table 1. The name of the protein should be added to the table.

      OK.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Several aspects of data interpretation, however, require clarification or reanalysis-most notably the single-channel analyses (event counts, Po metrics, and mixed parameters), the statistical treatment, and the interpretation of purported "OFF currents." Additional issues include PKD2L1-TRPP3 nomenclature consistency, kinetic comparison with ASICs, and the physiological relevance of the extreme acidification paradigm. Addressing these points will substantially improve reproducibility and mechanistic depth.

      Overall, this is a scientifically important and technically sophisticated study that advances our understanding of CSF sensing, provided that the analytical and interpretative weaknesses are satisfactorily corrected.

      (1) The authors should re-analyze electrophysiological data, focusing on macroscopic currents rather than statistically unreliable Po calculations. Remove or revise the Po analysis, which currently conflates current amplitude and open probability.

      We agree with the reviewer that the Po analysis has strong limitations, particularly in experiments where the recording times are short, like when extracellular pH is changed either by photolysis (Figure 4D) or puff applications (Figure 3Aa). In order to circumvent that problem and not to rely only on Po estimations, we used alternative methods as well, including the analysis of the current membrane charge that we have used extensively during the manuscript (Figures 3A and 4D, for example) or the analysis of the event latencies (Figure 4G). Nevertheless, single-channel recordings clearly contain information that is not included in the macroscopic current analysis. We intend to stress in the revised version that the elementary current amplitude is conserved by manipulations such as pH changes, leaving the total number of channels (N) and the channel open probability (Po) as possible culprits for the current changes. Since these changes are rapid and reversible, it is likely that N stays constant and that Po changes. In order to address the reviewer’s concern, we propose the following changes/reanalysis: i) to report in each condition the minimum N (maximum observed openings; for example, in Figure 3Aa the minimum N goes from 4 in control conditions to 1 during the puff of the pH 6.4 solution). This method (to estimate N by counting the maximum number of open channels), while imperfect, provides a tentative estimate of Po; ii) following the previous point, we propose to reword the text (and images) and use the expression “apparent Po” instead of “Po”; iii) to report the fraction of time that the channels remain open. Also, we acknowledge that some traces are confusing (Figure 3Aa, top) as they seem to show macroscopic currents. We will modify those figures by plotting the amplitude histograms (as in Figure 1Bb) in order to show unambiguously that current recordings from CSFcNs only show single-channel activities.

      (2) PKD2L1-TRPP3 nomenclature should be clarified and all figure labels, legends, and text should use consistent terminology throughout.

      We agree with the reviewer that the nomenclature concerning polycystin members is confusing. In this manuscript we have followed the nomenclature that has been proposed in a recent, comprehensive review on polycystin channels by Palomero, Larmore and DeCaen (Palomero et al. 2023), where the authors refer to the channels by their gene name. In this review, the authors indicate that PKD2L1 channels correspond to TRPP2 (formerly TRPP3, their table 1). In another, recent review on TRP channels, however, the authors refer to the PKD2L1 channel as TRPP3 (Zhang et al. 2023). In order to avoid any confusion we will remove from the text any reference to the TRPP nomenclature and stick to the PKD2L1 name.

      (3) The authors should reinterpret the so-called OFF currents as pH-dependent recovery or relaxation phenomena, not as distinct current species. Remove the term "OFF response" from the manuscript.

      We concur with the reviewer that the term “OFF response” is not very helpful from the biophysical perspective and conveys the idea that it is another current. We will remove the term “OFF response” or “OFF current” in the revised manuscript and replace it by the term “photolysis-evoked PKD2L1 current”. Also, we will condense two sections (“The proton-induced current is an off-current” and “The off-current is mediated by the activation of PKD2L1 channels”) into a single new section (“The photolysis-induced current is mediated by PKD2L1 channels”), as separating the description of the photolysis-evoked PKD2L1 current compromises its description. Finally, we will rewrite the discussion to better describe this current.

      (4) Evidence for physiological relevance should be provided, including data from milder acidification (pH 6.5-6.8) and, where appropriate, comparisons with ASIC-mediated currents to place PKD2L1 activity in context.

      This is partly addressed in Figure 3. The data there indicates that PKD2L1 channels are very sensitive to pH variations around physiological pH. In order to make this conclusion stronger, we will add to the figure the EC50 values drawn from the fittings. In terms of the ASIC-mediated currents, one of our main conclusions is that ASIC channels are not present in the ApPr, as the effects of proton photolysis in the ApPr and are not blocked by ASIC channels blockers. Our results indicate that PKD2L1 channels are the exclusive pH sensitive channels in the ApPr, while ASIC channels are probably the acid-sensitive channels in the soma, although we have not studied the latter in detail. Following this and the editor’s comments, the subsection “the involvement of ASICs” in the Discussion has been modified.

      (5) Terminology and data presentation should be unified, adopting consistent use of "predominant" (instead of "exclusive") and "sustained" (instead of "tonic"), and all statistical formats and units should be standardized.

      Following the suggestions of the reviewer, an exhaustive work will be performed to unify terminology, data presentation and correct the text following the reviewer’s suggestions.

      (6) The Discussion should be expanded to address potential Ca<sup>2+</sup> -dependent signaling mechanisms downstream of PKD2L1 activation and their possible roles in CSF flow regulation and central chemoreception.

      This is indeed a very interesting and currently unresolved point in the physiology of CSFcNs. Published data indicate that calcium flowing into the cell through PKD2L1 channels is a key regulator of apical process physiology: on the one hand, PKD2L1 channels are calcium permeable and at the same time, they are inhibited by intracellular calcium (DeCaen et al. 2016). Also, ultrastructural data indicate that the ApPr is rich in mitochondria and tubulo-vesicular structures resembling the Golgi apparatus (Bjugn et al. 1988; Bruni et Reddy 1987), which are intracellular organelles that contribute to calcium homeostasis. Altogether, this evidence suggests that intraApPr calcium concentration needs to be finely regulated, both in space and in time, in order for the ApPr to fulfil its physiological roles. Based on what has been published in the literature, we can speculate that these calcium signals can be decoded by several systems: i) calcium can act as the second messenger linking the activation of the multimodal PKD2L1 channels to changes in CSFcNs excitability, which in turn regulate spinal neuronal networks controlling locomotor activity; ii) calcium could initiate neurosecretion of different molecules from the ApPr to the central canal (as has been proposed by the Wyart group in the zebrafish in the context of bacterial infections(Prendergast et al. 2023)); iii) calcium could activate the Hedgehog signaling pathways (as has been shown by (Delling et al. 2013)); iv) calcium could modulate CSF flow directly (by modulating ciliary activity) or indirectly (by modulating ependymal cells ciliary activity through paracrine interactions). Resolving these downstream pathways is essential to fully define the role of CSFcNs as integrators of CSF homeostasis. We will expand this issue in the Discussion of the revised ms.

      Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments relate to their multimodal CSF-sensing function remains unclear.

      The authors confirm that in GATA3:GFP mice, where these cells are labeled, that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show that functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding of how these sensory cells can detect the changes of pH in the CSF so finely.

      Weaknesses:

      The following do not constitute weaknesses; rather, they are minor requests that this reviewer considers would complete this beautiful study.

      (1) It would be nice to quantify further the relation in spontaneous as well as in acidic or basic pH between the effects observed on channel opening and holding current: do they always vary together and in a linear way?

      Following the reviewer’s suggestion, we have performed a Spearman’s rank correlation test, which shows a significant correlation between the changes in the apparent open probability and holding current (paired experiments; ctrl vs pH 6.4 pressure applications; p < 0.05, Spearman r = 0.72 and critical value = 0.67). The Pearson correlation coefficient calculated on the same data set = 0.63 and the critical value is 0.632, which indicates that the correlation is not linear. We will add this analysis to the manuscript.

      (2) Since CSF-cNs also respond to changes in osmolarity (Orts Dell Immagine 2013) & mechanosensory stimulations in a PKD2L1 dependent manner (Sternberg NC 2018), it would be nice to test the same results whether the same results hold true on the role of PKD2L1 in AP for pressure application of changes in osmolarity.

      This is a very important point. As the reviewer mentions, previously published experimental evidence indicates that CSFcNs are also sensitive to osmolarity changes and mechanical stimulation in a PKD2L1-dependent manner. It is therefore reasonable to assume that, as for the pH sensitivity, osmotic and mechanical sensitivity depends on channels segregated to the ApPr. For the mechanosensitivity, the spatial segregation could be tested by “touching” either the ApPr or the soma with a piezo-controlled blunted pipette (see, for example, Hao et al. 2013). However, the sensitivity to osmotic changes is much more difficult to assess, as pressure application does not have enough spatial resolution to discriminate among compartments in such a compact cell such as the CSFcNs. In theory, a highly spatially localized osmotic jump could be reached with photolysis, but a caged compound releasing many osmotic particles simultaneously should be used. In typical photolysis experiments, a localized osmotic jump is produced, but it is very low (in the order of 1 to 2 mOsm).

      In mice, like in fish (Sternberg et al, NC 2018), we can observe throughout the figures that a large fraction of the channel activity occurs with partial and very fast openings of the PKD2L1 channel. I recommend the authors analyse the points below:

      (a) To what extent do these partial openings of the channel contribute to the changes in holding current and resting potential?

      As the reviewer indicates, these partial and very fast openings are a characteristic of PKD2L1 single-channel activity that seems to be present in different species. However, estimating what is the exact contribution of these events to the sustained current would require a detailed model of the channel that it is still lacking. Indeed, the exact mechanism by which CSFcNs show this prominent sustained current is unknown and should definitely being addressed in future works.

      (b) In the trace from the outside out AP, it looks like the partial transient openings are gone. Can the authors verify whether these partial openings are only present in somatic recordings?

      The outside-out recordings from the ApPr also show some partial openings (please look at the upper trace in Figure 4Db). We will specifically mention this important point in the revised version of the ms.

      (3) Previous studies have observed expression of metabotropic Glutamate receptors in CSF-cNs (transcriptome from Prendergast et al CB 2023). The authors only used blockers for ionotropic glutamate receptors in their recordings: could it be that these metabotropic receptors influence the response to uncaging of MNI-Glu when glutamate is co-released with a proton?

      We thank the reviewer for pointing out the presence of metabotropic glutamate receptors in CSFcNs. However, our evidence indicates that there is no contribution of metabotropic receptors when uncaging MNIglutamate because: i) the response obtained when uncaging MNI-gLGG (where there is no glutamate release; Figure 5Ab) and ii) the response obtained when uncaging protons from DPNIGABA (a GABA cage that has similar photochemistry than MNI cages which also release a proton when photolysed; data not shown), are the same. Indeed, in both experiments (MNI-gLGG or DPNI-GABA uncaging) a clear photolysis-evoked PKD2L1 current can be observed.

      (4) In the outside out patch of the AP, PKD2L1 unitary currents appear rare. Could it be that the disruption in the cilium or underlying actin/myosin cytoskeleton drastically alter the open probability of the channel?

      Although we have not quantified it, the reviewer is right that the opening frequency of PKD2L1 channels in the outside-out patches is lower than in the whole-ApPr recordings. We interpreted this difference as a difference in channel number. However, another plausible interpretation is that, as the reviewer suggests, the biophysics of the channels are affected because the protein is taken out from its normal ionic environment and/or loses important interactions with regulatory proteins.

      (5) Could the authors use drugs against ASIC to specify which ASIC channels contribute to the pH response in the soma?

      As described in the manuscript, we did perform experiments with ASIC channel blockers, although we did not attempt to characterize the specific ASIC channel involved in the somatic response. Based on what has been published in the literature, we used both psalmotoxin-1 (which blocks ASIC1 channels) and APETx2 (which blocks ASIC3 channels). The presence of ASIC1 channels in mice CSFcNs has been shown by (Orts-Del’Immagine et al. 2012; Orts-Del’Immagine et al. 2016), while the presence of ASIC3 in the lamprey CSFcNs has been shown by (Jalalvand et al. 2016). When we puff an acidic solution aiming at the soma, we can record an inward current that is blocked by psalmotoxin-1, although there is always a small component remaining (as originally shown by Orts-Del’Immagine in the aforementioned articles); however, we have not attempted to block this small component that remains after psalmotoxin-1 bath application.

      (6) This is out of the scope of this study, but we did observe in fish a very rarely-opening channel in the PKD2L1KO mutant. I wonder if the authors have similar observations in the conditions where PKD2L1 is mainly in the closed state.

      We have never seen such kind of openings in our recordings (when the channel is closed or in the presence of dibucaine).

      Bjugn, R, H K Haugland, et P R Flood. 1988. “Ultrastructure of the mouse spinal cord ependyma.” Journal of Anatomy 160 (octobre): 117‑25.

      Bruni, J. E., et K. Reddy. 1987. “Ependyma of the Central Canal of the Rat Spinal Cord: A Light and Transmission Electron Microscopic Study”. Journal of Anatomy 152 (juin): 55‑70.

      DeCaen, Paul G., Xiaowen Liu, Sunday Abiria, et David E. Clapham. 2016. “Atypical Calcium Regulation of the PKD2-L1 Polycystin Ion Channel”. eLife 5 (juin): e13413. https://doi.org/10.7554/eLife.13413.

      Delling, Markus, Paul G. DeCaen, Julia F. Doerner, Sebastien Febvay, et David E. Clapham. 2013. “Primary cilia are specialized calcium signalling organelles”. Nature 504 (7479): 311‑14. https://doi.org/10.1038/nature12833.

      Hao, Jizhe, Jérôme Ruel, Bertrand Coste, Yann Roudaut, Marcel Crest, et Patrick Delmas. 2013. “Piezo-Electrically Driven Mechanical Stimulation of Sensory Neurons”. In Ion Channels, édité par Nikita Gamper, vol. 998. Methods in Molecular Biology. Humana Press. https://doi.org/10.1007/978-1-62703-351-0_12.

      Jalalvand, Elham, Brita Robertson, Peter Wallén, et Sten Grillner. 2016. “Ciliated Neurons Lining the Central Canal Sense Both Fluid Movement and pH through ASIC3”. Nature Communications 7 (janvier): 10002. https://doi.org/10.1038/ncomms10002.

      Orts-Del’Immagine, Adeline, Riad Seddik, Fabien Tell, et al. 2016. “A Single Polycystic Kidney Disease 2-like 1 Channel Opening Acts as a Spike Generator in Cerebrospinal Fluid Contacting Neurons of Adult Mouse Brainstem”. Neuropharmacology 101 (février): 549‑65. https://doi.org/10.1016/j.neuropharm.2015.07.030.

      Orts-Del’immagine, Adeline, Nicolas Wanaverbecq, Catherine Tardivel, Vanessa Tillement, Michel Dallaporta, et Jérôme Trouslard. 2012. “Properties of Subependymal Cerebrospinal Fluid Contacting Neurones in the Dorsal Vagal Complex of the Mouse Brainstem”. The Journal of Physiology 590 (16): 3719‑41. https://doi.org/10.1113/jphysiol.2012.227959.

      Palomero, Orhi Esarte, Megan Larmore, et Paul G. DeCaen. 2023. “Polycystin Channel Complexes”. Annual Review of Physiology 85 (Volume 85, 2023): 425‑48. https://doi.org/10.1146/annurev-physiol-031522-084334.

      Prendergast, Andrew E., Kin Ki Jim, Hugo Marnas, et al. 2023. “CSF-Contacting Neurons Respond to Streptococcus Pneumoniae and Promote Host Survival during Central Nervous System Infection”. Current Biology 33 (5): 940-956.e10. https://doi.org/10.1016/j.cub.2023.01.039.

      Zhang, Miao, Yueming Ma, Xianglu Ye, Ning Zhang, Lei Pan, et Bing Wang. 2023. “TRP (Transient Receptor Potential) Ion Channel Family: Structures, Biological Functions and Therapeutic Interventions for Diseases”. Signal Transduction and Targeted Therapy 8 (1): 261. https://doi.org/10.1038/s41392-023-01464-x.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers were very impressed with your work and definitely feel it is a scientifically important and technically sophisticated study that advances our understanding of CSF sensing. They, however, request some re-analysis of the data and discussion with minimum new experiments, if any. I think, if feasible, this will improve the quality of the study and would look forward to receiving a revised version.

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 1 - Molecular identity and localization of PKD2L1

      Major

      Nomenclature clarity.

      Please clarify the distinction between PKD2L1 and TRPP3. Several parts of the text and figure labels appear to conflate these names. PKD2L1 corresponds to the TRPP3 subfamily member and should not be interchanged with PKD2/TRPP2. Please confirm and update the nomenclature consistently throughout the manuscript (text, figure labels, and captions).

      We have addressed this issue in response to reviewer 1. There seems to be some confusion in the literature concerning the nomenclature of PKD2L1 channels as in some recent publications the PKD2L1 channels are still named as TRPP3. However, the nomenclature of PKD2L1 channels or TRPP2, was updated in 2016 (Wu, Sweet and Clapham, Pharmacological Reviews, 2010). As indicated in response to reviewer 1, we have removed from the text any reference to the TRPP nomenclature and stuck to the PKD2L1 name.

      Physiological meaning of apical restriction.

      Expand the discussion of why apical-restricted localization matters. Specifically, address how segregation to the ApPr could support directional sensing of CSF flow and/or detection of localized pH gradients.

      Following the reviewers and editor’s comments, we have revised the discussion in order to take this and other comments into account.

      Open-state annotation (O3).

      Please include the O3 state in panels Ba and Bb; the figure clearly shows an additional open level consistent with O3.

      The editor is right in that there is another state that presumably corresponds to O3. Following the editor’s recommendation, we now indicate this 3rd level and add a short sentence explaining this in figure 1 legend.

      Minor

      Indicate the ROI definition and background-subtraction method used for fluorescence quantification (ApPr vs soma).

      Not applicable for this figure.

      Ensure the same intensity scale (lookup table and range) is used across panels to enable direct comparison.

      Done.

      (2) Figure 2 - Electrophysiological characterization of ApPr and somatic recordings Major

      Definition of "PKD2L1-dependent current."

      Define this term precisely at its first appearance. Specify whether it denotes currents inhibited by dibucaine, abolished in PKD2L1-knockout preparations, or both.

      Done

      Statistical power of single-channel analysis.

      The number of observed openings (< 1000 events) is too low to estimate open probability (Po) reliably. Please re-analyze the data using macroscopic current traces rather than Po-based kinetics.

      Confounded Po analysis.

      The current Po analysis mixes current amplitude and Po in the same calculation, conflating independent variables. Re-evaluate or remove this analysis.

      Unknown channel count.

      Because the number of channels in each patch is unknown, Po and "closed probability" values cannot be interpreted meaningfully. Focus instead on the averaged macroscopic current density.

      General analytical validity.

      The single-channel analyses in Figure 2 are not interpretable under these experimental conditions. Closed-time distributions and Po-based metrics (e.g., "Po1," "P2") depend critically on channel number and event sampling. Moreover, the manuscript applies essentially the same Po methodology across conditions (Po1 vs P2), which adds no mechanistic resolution and risks circular interpretation.

      Actionable recommendation:

      Remove Po- and closed-time-based analyses from Figure 2 and from the manuscript as a whole. Reanalyze the data using metrics that remain valid when the channel number is unknown:

      Macroscopic current analysis (leak-subtracted current density, I-V relationships, activation time constants).

      Single-channel conductance only (amplitude histograms and unitary slope conductance), without attempting Po or dwell-time inference.

      Report filtering bandwidth and sampling rate, and restrict statistical treatment to these robust parameters.

      Following the reviewers (see above) and editors’ recommendations, we have reanalyzed the data in order to avoid the analysis based on Po and Pc. We have instead calculated from the recordings other 2 parameters, n<sub>max</sub> (the maximum number of channels that open simultaneously during a 500 ms time window) and the total open time of a single channel during the same 500 ms time window. The main text, figures and corresponding figure legends, and the Materials and Methods section have been changed accordingly. Notably, the Po and Pc analysis were removed from Figures 3, 4 and Supplementary Figure 3, and replaced by the above-mentioned parameters. Also, the fact that the recordings are not long enough to calculate Po is now specifically mentioned in the Materials and Methods section, lines 785 to 790. In addition, the analysis in Figure 3Ce has been redone so that the activity of the channel as a function of pH is now plotted as the normalized apparent Po (relative to the apparent Po value at pH 7.4).

      Minor

      State whether input-resistance values (1.8-4.4 GΩ) were leak-subtracted and series-resistance-compensated.

      As already mentioned in the methodology section (line 693), series resistance was not compensated for during the experiments. We have now added a sentence in the methodology section indicating that in the voltage range that was chosen for the analysis of the input resistance, no voltage-dependent conductance was activated (lines 809 to 812).

      Ensure unit consistency: use Po or normalized Po rather than frequency (Hz) throughout.

      (3) Figure 3 - pH-evoked currents and kinetics

      Major

      Invalid Po analysis (Fig. 3Ca-Ce).

      The Po- and Pc-based single-channel analyses in panels 3Ca-3Ce should be deleted. As noted earlier, the event count is insufficient, and the number of active channels in each patch is unknown. Under these conditions, Po and Pc values have no quantitative meaning and could mislead readers. These panels do not contribute additional mechanistic insight beyond the macroscopic current data and therefore, should be removed. If retained for illustrative purposes, they must be explicitly labeled as representative traces without any statistical quantification.

      As we mentioned above, we removed the Po and Pc analysis from the manuscript.

      Minor

      Present regression equations and r<sup>2</sup> values for the linear fits shown in Fig. 3D.

      Done (lines 947951).

      Confirm that all axes include units and identical scaling between conditions for direct comparison.

      Done.

      (4) Figure 4 - Laser photolysis and local stimulation experiments

      Major

      Laser timing annotation.

      Clearly mark laser-pulse timing (e.g., arrow or shaded region) on all current traces to facilitate interpretation.

      We thank the editor for pointing out the inconsistencies in terms of the laser pulse timing. To indicate the laser pulses, we have now added an arrowhead in cases where a single sweep is shown (for example, Figure 3D), and an arrowhead and a dotted magenta vertical line in cases where multiple sweeps are shown (for example, Figure 3C).

      pH calibration within the laser spot.

      Provide quantitative calibration of pH changes induced by laser photolysis, including information on spot size, local diffusion, and estimated pH recovery kinetics.

      This is an important point and we thank the editor for mentioning it. We have now completed the subsection untitled “photolysis” where we provide information on the lateral and axial dimensions of the photolysis laser spot used in this work (lines 734 to 737). We have also rewritten part of Figure 5A legend to highlight the fact that the experiments presented there (photolysis on top of the ApPr and next to it) are compatible with a high spatial resolution of proton release (lines 1023 to 1026).

      On the other hand, we have attempted to perform pH calibrations in the setup using the pH-sensitive dye pyranine (or HPTS: 8-Hydroxypyrene-1,3,6-trisulfonic acid). HPTS is a very useful tool for pH calibrations in the physiological range: its pKa value is close to 7.2, and it can be used as a ratiometric dye (its fluorescence is pH-independent at 405–410 nm and pH-dependent at 450 nm). Unfortunately, when trying to perform a calibration under the conditions of a real experiment,

      where photolysis occurs in a tiny volume (approximately 1 µm³ in a total bath volume of more than 1 ml), we encountered the following problem, which made it impossible to obtain any useful data: the 405 nm uncaging pulse bleaches the dye, and any useful information is lost. Also, our imaging system is not fast enough to follow the pH change. As it is discussed in the Materials and Methods section, subsection “Estimation of the pH drop induced by photolysis” (line 814), the fast protonation of bicarbonate indicates that the pH change induced by the photolysis recovers in the submillisecond time range.

      Repeated stimulation effects.

      Discuss whether repeated photolysis induces adaptation or desensitization of PKD2L1 currents, and indicate whether current amplitude decreases across successive trials.

      This issue is now specifically mentioned in the Materials and Methods section, lines 739 to 741.

      Invalid interpretation of the "OFF response."

      The interpretation of the so-called "OFF response" in Figure 4C is not supported by the presented data. There is no evidence for a bona fide OFF current, and the literature cited does not demonstrate such a phenomenon for PKD2L1 alone. Rather, previous studies implicate PKD1L3-dependent mechanisms in similar biphasic responses. Please reconsider the cited references and remove claims of an OFF current attributed to PKD2L1.

      Done.

      Actionable recommendations:

      Do not use the term "OFF response" throughout the manuscript. Recast these transients as pH dependent recovery or relaxation of current following cessation of acidification.

      Done. We have performed extensive rewriting and reorganization of the Results and Discussion in order to take into account both the reviewer’s and editor’s comments. Please also take a look at comment #3 of Reviewer 1 and point 11 below.

      Include continuous-illumination controls (sustained local acidification) to test whether a steady state current is maintained. This will clarify whether the post-stimulus transient reflects recovery kinetics rather than a distinct current species.

      We thank the editor for suggesting this experiment. However, continuous laser illumination is a difficult manipulation and does not necessarily lead to an acidification of the illuminated volume. Indeed, with continuous illumination the cage is lost from the illumination spot and needs to be replaced by diffusion from the non-illuminated volume, leading to non-homogeneous concentrations. Also, the chances of inducing photo damage are higher. We thus designed a similar experiment where instead of performing continuous illumination we photolysed with short and high frequency trains in order to produce a long-lasting acidification. The results of these experiments have been added to the manuscript as part of the results section and in Figure 5H. Similarly to what is seen with single illuminations, the photolysis trains induce a current that appears almost exclusively at the end of the train, implying that the current is indeed a PKD2L1-dependent recovery current.

      Align the current time course with measured or estimated local pH (or calibrated proxy) to demonstrate causal coupling and avoid implying a separate conductance.

      We have added the calculated pH change to the inset of Figure 4C as an example.

      Revise the schematic/model figure and textual description accordingly, restricting the framework to phasic vs sustained activation modes without invoking a separate OFF current for PKD2L1.

      Done.

      Minor

      Include scale bars, sample numbers (n), and laser parameters (duration, power) in all panels.

      In order not to make the figure and the panels very heavy in the original version, we tried to limit the number of scale bars. We have now performed some modifications, added the missing scale bars, and changed the figure legend in order to take into account the editor’s comments. We have also corrected a few values that were wrongly reported.

      Standardize p-value formatting (e.g., p = 6 × 10 ⁶) throughout the figure and legend.

      Done.

      (5) Figure 5 - Single-channel recordings

      Major

      Mixed parameters (current amplitude and Po).

      The current analysis improperly mixes single-channel current amplitude and Po within the same figure, conflating distinct parameters. These quantities must be analyzed and presented separately, or the Po data should be removed entirely if not independently supported.

      Insufficient event count.

      Given the very limited number of observed openings, Po-based statistics are not meaningful. Please report only representative single-channel traces and corresponding amplitude histograms without attempting quantitative Po estimation.

      Minor

      Convert frequency (Hz) values to Po for consistency with earlier analyses, or remove frequency metrics altogether if Po analysis is omitted.

      Figure 5 does not include Po or event frequency analysis, so we think there must be a misquotation of the figure. However, the Po issue has already been addressed before and alternative analysis have been proposed.

      (6) Introduction

      The introductory paragraph mentions the "five senses" as a framing concept. However, this statement lacks scientific grounding in the context of CSF-contacting neurons and chemosensory physiology. The traditional "five senses" classification is not an evidence-based neurophysiological framework and may be misleading to readers. I recommend removing or rephrasing this part, focusing instead on molecular and cellular mechanisms of sensory transduction (e.g., chemical, mechanical, and pH sensing) rather than on classical sensory categories.

      Following the editor’s recommendation, we have removed this part.

      The manuscript refers to PKD2L1 using the term TRPP2 in some parts of the introduction. This is incorrect, as PKD2L1 corresponds to TRPP3, not TRPP2. Please correct this nomenclature and ensure consistent use of "PKD2L1 (TRPP3)" throughout the entire manuscript to avoid confusion with the distinct PKD2/TRPP2 protein, which belongs to a different subfamily with separate physiological roles.

      We thank the editor for pointing this out. As we mentioned in the responses to the “public reviews”, the literature is confusing, so we decided to remove from the manuscript any mention to TRPP channels.

      (7) Discussion

      The current Discussion reads largely as a descriptive summary of results and lacks conceptual depth. It does not effectively integrate the biophysical properties of PKD2L1 with its physiological role as a neuronal pH sensor, nor does it develop a broader interpretation relevant to CSF homeostasis or chemoreception.

      Following the reviewers and editor’s recommendations, we have now added a new section in the Discussion untitled “PKD2L1 downstream signaling mechanisms”.

      (8) Insufficient biophysical analysis

      The discussion of channel gating and pH dependence is superficial and does not explore the energetic or structural mechanisms underlying proton sensitivity. The authors should analyze their data in the context of known PKD/TRPP family biophysics-for example, protonation sites, subunit composition, or gating kinetics-and explain how these confer bidirectional (acidic vs alkaline) sensitivity within physiological ranges.

      In this work, we studied the pH sensitivity of PKD2L1 channels in the context of CSFcN sensory physiology. From a pure biophysical perspective, the pH sensitivity of PKD2L1 channels has been studied by multiple groups; however, it is still unknown how the gating of the channel responds to pH changes, although it can be proposed that some polar residues in the protein regulate the state of the pore. Likewise, the mechanism of the “off-response” is also unknown. To the best of our knowledge, there is only one article in which the authors have attempted to relate pH, PKD2L1 channel structure, and function. In this work (Su et al., Nature Communications 2018), the authors compare PKD2L1 channels with another pH-sensitive member of the TRP family, TRPML3, whose structures at pH 7.4 and 4.8 are known (Zhou et al., Nature Structural and Molecular Biology, 2017). We have rewritten some sentences of the Discussion in order to be more specific about the pH dependence of PKD2L1 channels and its proposed mechanisms.

      (9) Weak physiological context

      The manuscript does not adequately address how PKD2L1 functions as a true physiological pH sensor. The discussion should connect channel activity to realistic CSF pH fluctuations (6.8-7.6) and to relevant physiological or pathophysiological conditions (e.g., respiratory acidosis, neurogenic regulation of CSF composition). Without this, the relevance of large, artificial acidification (pH 3-3.5) remains unclear.

      We have added a new section in the Discussion where we speculate on how PKD2L1 channels may be activated in physiological and pathophysiological conditions. However, we would like to insist here that the main goal of the photolysis experiments (which induce short and large acidifications) was to assess the spatial segregation of PKD2L1 channels. We now mention this point specifically and also speculate on the conditions that could eventually give rise to the “recovery” current.

      (10) Over-interpretation of unsupported points

      The paragraph describing voltage propagation from the ApPr to the soma/axon is speculative and unsupported by any data in the manuscript. Please delete this section entirely, including the citation to Orts-Del'immagine et al., unless new electrophysiological evidence is added.

      We think this point (the propagation of signals originating from the ApPr to the soma) is important in the context of our work, so we have decided to make new experiments in order measure directly the degree of coupling between the 2 compartments. To do that we made simultaneous, current-clamp and voltage-clamp recordings from the ApPr and the soma. In these conditions we were able to measure experimentally and for the first time both the coupling coefficient and coupling conductance, which confirm that the propagation of voltage signals from the ApPr to the soma is extremely efficient. These new results are now described in a new subsection and in a new Figure 6.

      (11) Clarify the role of "OFF currents."

      The Discussion repeatedly refers to an "OFF response," but this phenomenon is not experimentally demonstrated for PKD2L1 alone. It likely represents pH-dependent recovery rather than an independent current. All discussion of "OFF currents" should be removed or reformulated accordingly.

      Following the editor and reviewer’s comments, we have deleted the term “off response” and “off currents” from the ms and have replaced them with the term “recovery current”. We have also changed the discussion accordingly.

      (12) Integration with ASICs and compartmental sensing

      While the manuscript briefly mentions ASIC involvement, it does not articulate how PKD2L1- and ASIC-mediated signals might complement each other in different compartments (ApPr vs soma). The authors should discuss the potential division of labor between these sensors and how such compartmentalization enhances pH detection in CSFcNs.

      Following the editor’s comments, we have rewritten the part of the subsection ‘the involvement of ASICs’ in the Discussion.

      (12) Broadened physiological perspective

      The Discussion should close by considering Ca<sup>2+</sup> -dependent downstream pathways activated by ⁺ PKD2L1 and their implications for CSF flow regulation, neurosecretion, and central chemoreception. These translational aspects would substantially improve the impact and readability of the manuscript.

      Done

      Overall, the Discussion must evolve from a descriptive narrative to a mechanistically and physiologically integrative synthesis, highlighting why PKD2L1 is not merely present in the ApPr but is a key molecular transducer linking ionic microenvironment to neuronal excitability.

      As it has been detailed above, we have performed several changes in the Discussion that follow the reviewer’s and editor’s recommendations.

    1. Historians have begun to think about the Enlightenment in a newly global way. .ose creaky wooden ships carried ideas across the boundaries of continents, languages, and religions just as the Internet does now (although they were a lot slower and perhaps even more perilous).

      This comment sticks out to me a lot because I feel like a large amount of people like to think of the internet as the first time that humans had contact in such a broad way, and it is very interesting to see that that may not be as true as we have all assumed. Humans seem to be more interconnected than we believed originally, and lots of ideas come from interacting with one another.

    1. Of maybe 40 people who spoke, only two were women, and I got the impression that they were listened to only because of their personal relationships with a couple of the male leaders. A few of the men confidently flexed their intellectual muscles before the crowd, using sarcasm and non-stop rhetoric to bludgeon other people into accepting their ideas. Ironically, these men were able to play such an elitist role precisely because of the anti-leadership tendencies in the room. Since there was no structured presentation of the issues– “We don’t need a lecture,” the line ran, “we need some participation” – only those who already had a grasp of the information could find a way through the chaos.

      Multiple problems are presented here. The first and more obvious being the fact that women were overlooked and given little opportunity to speak at the meeting. I find this especially frustrating as often women are the most victimized and targeted during times of war. Acts of rape and violence from both foreign imperialist/colonizers and countrymen are not uncommon. While western women may not be the best representatives to voice such concerns, I do think acknowledging the realities of the patriarchy are vital in such discussions. The secondary problem and slightly more egregious in regards to the goal of such a meeting is the unwillingness to teach and educate, and have no clear structure. Such groups do not have the luxury of not educating, having no structure, and letting few speak. To speak to a large meeting, with no baseline of where the majority of the individuals are in regards to understanding such rhetoric and theory is foolish and egotistical.

    1. Our self-concept is also formed through our interactions with others and their reactions to us. The concept of the looking glass self explains that we see ourselves reflected in other people’s reactions to us and then form our self-concept based on how we believe other people see us (Cooley, 1902). This reflective process of building our self-concept is based on what other people have actually said, such as “You’re a good listener,” and other people’s actions, such as coming to you for advice. These thoughts evoke emotional responses that feed into our self-concept. For example, you may think, “I’m glad that people can count on me to listen to their problems.”

      this is a pretty interesting topic to me because to be fully honest, none of us truly know who we are. You might have a really strong understanding of your personal beliefs, the things you like and dislike, maybe strong self awareness of how you act, but at the very root of it all, it’s hard to define exactly who we are. Subsequently, it’s even harder for another person to define who they think you are. Even if you do have strong self awareness, if somehow you could play a recording of the internal thoughts of everyone you’ve ever talked to in relation to what they thought about you, you’d probably be really surprised by what you heard. That’s simply because of the fact that no one’s ideologies are identical and no one has lived exactly the same life, and everyone’s opinions are going to change based on the life they’ve lived. So when you internalized something that someone else thought about you, that thought could’ve been polar opposite to what someone else thought about you, and therefore we are left in an infinite paradox of what is true about ourselves and who exactly we are.

    1. Where does ChatGPT fit into the Framework for Information Literacy?

      Introduction

      The above screenshot (figure 1) was what was generated when we asked ChatGPT, the generative AI system that has been the subject of a thousand hot takes about how it’s disrupting academia-as-we-know-it, to describe itself for an academic librarian audience. Perhaps it’s learning a bit too much from the public relations documents that were a part of the vast amounts of data it was trained on, when it describes itself as “highly relevant,” “invaluable,” and “accurate.” It did not, however, bring up the caveat that greets you when you open up ChatGPT itself: that it “may occasionally generate incorrect information,” that it “may occasionally produce harmful instructions or biased content,” or that it has “limited knowledge of the world and events after 2021.”1 In addition, it doesn’t bring up the reddest of academic red flags—that ChatGPT provides an easy way for students to cheat and plagiarize. The Atlantic has claimed that because of ChatGPT and other AI, “the undergraduate essay [which] has been at the center of humanistic pedagogy for generations

      . . . is about to be disrupted from the ground up.”2 A writer at Times Higher Education has suggested that allowing AI to replace a student’s creative voice means “abandoning our responsibilities as educators.”3 For as many handwringing accounts of how generative AI will destroy academia, there seem to be twice as many researchers, teachers, technologists, and pundits embracing what AI (and specifically ChatGPT) can do for teaching and learning. They suggest using it for overcoming writer’s block, generating outlines, creating summaries, generating prompts for discussion, asking for definitions, or generating flawed examples for critique.4 One compelling argument by Christopher Grobe in the Chronicle of Higher Education suggests that what generative AI can help us with is to “provide new starting points for some of the processes we routinely use to think.”5 We agree with Grobe’s argument that ChatGPT can give us a good starting point from which to work. The text generated by ChatGPT in the screenshot at the start of this article is an overly optimistic and idealized view of itself. We hope that in this article we can add the nuance that it lacks. Academic librarians serve their students and faculty to help them navigate the research process. Therefore, when a new technological tool blazes through higher education, as ChatGPT has over the last few months, it becomes increasingly important that librarians are aware of the tool and its uses so that they can serve their students and faculty. After decades of the ACRL Information Literacy Competency Standards for Higher Education, the ACRL Framework for Information Literacy for Higher Education was established with a much more flexible route for integration into curricula. The Framework provides librarians and disciplinary faculty with a customizable way to provide information literacy instruction that meets the needs of students and enables them to become participants in the information that they are producing (not just consuming). Because of the Framework’s flexible nature, librarians can incorporate new technology, like ChatGPT, more easily into their instruction. We have found that the idea of ChatGPT (and generative AI more broadly) can be connected to many of the knowledge practices and dispositions from the six frames of the ACRL Framework. In some places, the Framework enables us to embrace ChatGPT as an exciting new tool that adds value to information literacy instruction. In other places, the Framework’s discussions of evaluating authority and examining bias shines light on the inherent flaws of ChatGPT. In the next section, we will review each of the frames and discuss how ChatGPT fits into each of those Frames.

    2. Research as Inquiry

      Research as Inquiry

      One of the knowledge practices for the Research as Inquiry frame states that “learners who are developing their information literate abilities deal with complex research by breaking complex questions into simple ones, limiting the scope of investigations.”11 We have found that students will often come to research consultations or library instruction sessions with broad or vague research questions that they often do not know how to simplify or narrow down to research writing sufficiently scoped to a level that they can tackle in five to eight pages. While they might be interested in, let’s say, “the problem of poverty” or “the abortion debate,” they cannot digest the enormous amount of research in multiple academic

      disciplines that have attempted to address these types of questions. This is where we believe ChatGPT can help students (and even seasoned researchers) in generating ways to break complex problems down. ChatGPT can help refine research questions, determine search terms, come up with synonyms and related terms or phrases for searching, help decide which databases to search, generate textual concept maps, and even help generate citations. For example, here was the response it gave when we asked it “What search terms should we use for our hypothetical research question, ‘why aren’t college athletes paid?’”

      ChatGPT can also create textual concept maps to help think through various aspects of a research topic. This can be useful for students who need to narrow or refine their topic. Simply asking the AI to help narrow a research topic can be useful as it will give you a variety of ways to explore a topic. For example, we asked it to help us narrow our search on why college athletes aren’t paid, and it gave us detailed options in an easy-to-read format:

      You can see clearly how ChatGPT can help students push past that inquiry threshold. Often, they don’t even know where to start in their search for information, or how to probe the nuances or facets of a large complex question for a scope they can grasp. Instead of typing their entire research question into Google or a database (as we have all seen students do) and having to sift through a mountain of results, they can type it into ChatGPT and ask “How do I start searching? Where could I go from here? What’s manageable?” As students are growing in their information literacy abilities, ChatGPT can help scaffold their skills enabling them to accomplish this task more confidently in the future. One caveat that always bears repeating: ChatGPT has biases. It is trained on a large dataset of material from the internet. It may not produce underrepresented or less well-researched aspects of a topic. Because of this, it is important for students to explore topics holistically, with ChatGPT as one tool in their toolbelt.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1

      Evidence, reproducibility and clarity

      The manuscript presents a comparative genomic analysis of chemosensory systems across the order Vibrionales. By combining phylogenetic analyses, genomic context, MCP repertoires, and structural comparisons, the authors investigate the diversity and evolution of chemosensory systems within Vibrionales. The study addresses an interesting question and assembles a substantial genomic dataset. The manuscript is generally well illustrated and contains several potentially useful observations regarding the distribution and organization of chemosensory systems across Vibrionales genomes. However, there are a number of conceptual, methodological, and presentation-related issues that should be addressed before the evolutionary conclusions can be fully supported.

      The claim that F8 represents a novel chemosensory system appears to be incorrect Lines 407-424, as well as several other places throughout the manuscript, describe F8 as a "previously uncharacterized lineage" and a "newly identified" system. However, F8 chemosensory systems were previously described by Wuichet and Zhulin and have been part of the established chemosensory system classification framework for more than a decade. Furthermore, F8 systems are already annotated as such in MiST. For example, the genome GCF_003390675.1 (Vibrio anguillarum), which is included in this study, contains CheA, CheR, and CheB proteins assigned to an F8 system in MiST.

      Consequently, describing F8 as a novel, newly identified, or newly designated system appears inappropriate. This issue requires substantial revision throughout the manuscript. The authors should clearly distinguish between the previously established F8 chemosensory class and any novel observations reported here. If the novelty lies in the distribution of F8 systems within Vibrionales, their genomic organization, their evolutionary history, or some other aspect, this should be stated explicitly.

      __Response: __We thank the reviewer for appreciating our work and further commenting on the points to improve this study. As correctly mentioned by you, F8 CSS has been discovered by Wuichet and Zhulin and already annotated in the MiST database. Since this study focuses on experimentally studied CSS clusters within Vibrio organisms, where F6, F7, and F9 have already been studied very well; however, the F8 is not studied experimentally at least in Vibrio cholerae in any literature. We have revised the statement about F8 CSS cluster being novel throughout the manuscript. We agree that MiST database reports different Che proteins, such as CheA, CheR, and CheB, assigned to F8 class, but the information about the entire F8 gene cluster in those Vibrio species has not been provided. Following this, we have identified this entire F8 gene cluster in our study and discussed it as an ‘experimentally unexplored CSS’ in the manuscript.

      The conclusions regarding horizontal gene transfer and vertical inheritance are not sufficiently supported. Lines 429-445 contain evolutionary interpretations that appear internally inconsistent. The manuscript interprets the sporadic distribution of F7 as evidence of vertical inheritance coupled with lineage-specific adaptation, whereas several lines later the similarly patchy distribution of F8 is interpreted as evidence of horizontal gene transfer. Similar distribution patterns should not be used to support contrasting evolutionary scenarios without additional supporting evidence.

      More broadly, patchy phylogenetic distributions alone are generally insufficient evidence for horizontal gene transfer. Alternative explanations, including differential gene loss, genome reduction, incomplete sampling, or rapid sequence divergence, should also be considered. Furthermore, the conclusions regarding vertical inheritance and HGT imply reconstruction of deep evolutionary history across broad bacterial groups. Based on the methods presented, these inferences appear to rely primarily on CheA phylogenies and analyses of homologous sequences. While informative, these analyses may not be sufficient to confidently infer ancestral origins. The authors should either provide additional phylogenetic evidence supporting these conclusions or moderate the language throughout the manuscript. In their current form, the proposed evolutionary scenarios would be more appropriately presented as hypotheses rather than demonstrated conclusions.

      __Response: __We thank the reviewer for this insightful comment. We agree that the evolutionary interpretations presented in the original manuscript need additional supporting evidence. Since our inferences regarding vertical inheritance and horizontal gene transfer were primarily based on phylogenetic distribution patterns and CheA-based phylogenetic analyses. In response, we have carefully revised the relevant sections of the manuscript and moderated the language throughout. The previously presented statements as a conclusion have been reformulated as hypotheses or possible evolutionary scenarios. It must be noted that comparative analysis of protein architecture of F9 CSS, along with phylogeny, revealed a high degree of similarity between Vibrionales F9 protein architecture and their Alphaproteobacterial counterparts, putatively suggesting the HGT event-based acquisition of these CSS within order Vibrionales.

      References are frequently missing, incomplete, or potentially inappropriate. One major concern is the quality and completeness of referencing throughout the manuscript. Multiple statements either lack references altogether or appear to cite sources that do not directly support the associated claims. For example, lines 76-78 cite Ulrich et al. (2005) in support of the statement that two-component systems constitute a dominant signaling paradigm in prokaryotes, whereas the cited article is entitled "One-component systems dominate signal transduction in prokaryotes." The text and/or citation should therefore be reconsidered. Similarly, the statement in lines 100-101 that 17 classes of flagellar chemosensory systems have been designated should be accompanied by an appropriate reference. In addition, numerous ecological, physiological, and evolutionary statements throughout the Introduction and Results sections either lack citations or would benefit from more precise supporting references. I recommend that the authors carefully review all references and ensure that each citation directly supports the corresponding statement.

      __Response: __We thank the reviewer for this comment. We have added the correct reference at the place of Ulrich et al 2005. Along with this we have further checked the references throughout the manuscript and corrected them wherever it is necessary.

      The manuscript would benefit from substantial restructuring and shortening The manuscript is considerably longer than necessary, and several sections appear only loosely connected to the central biological question. In particular, the Introduction contains extensive discussions of Vibrio ecology, virulence, motility, host interactions, and general signal transduction. While these topics are relevant, the overall narrative currently reads more like a broad review article than an introduction to comparative genomics study. I recommend substantially shortening and restructuring the Introduction so that the central biological question and the specific objectives of the study become more apparent to the reader.

      Similarly, the section entitled "Multipartite genome and extensive RNA gene repertoire reflect niche adaptation in Vibrionales" (lines 302-342) contains several observations regarding genome size, tRNA counts, and rRNA copy numbers. However, it remains unclear how these analyses contribute to the primary conclusions regarding chemosensory system evolution. This section should either be shortened substantially and more explicitly connected to the central theme of the manuscript or moved to supplementary material. The manuscript would also benefit from substantial language editing. Numerous grammatical and stylistic issues are present throughout the text, including awkward phrasing, subject-verb agreement errors, and overly long sentences. A thorough language revision would improve readability and help the reader focus on the scientific content.

      __Response: __We thank the reviewer for this comment. We agree that this manuscript is considerably larger and we have shortened several parts in the introduction such as Vibrio ecology, virulence, motility, host interactions and focused more on study objectives. We have also moved the result tilted as “Multipartite genome and extensive RNA gene repertoire reflect niche adaptation in Vibrionales” in the supplementary part. We have carefully revised the entire manuscript to address grammatical errors, improve sentence structure, correct subject–verb agreement issues, and eliminate awkward phrasing. We have also streamlined several lengthy sentences and paragraphs to enhance clarity, readability, and the overall presentation of the scientific content.


      Specific comments:

      Line 20: The phrase "28 Vibrio clades" requires clarification. It is not clear what these clades represent, how they were defined, or whether "clades" is the most appropriate term. Please define these groups more clearly. Perhaps the term "genera" would be more appropriate.

      __Response: __We thank the reviewer for this comment. We agree that the terms clade and genera can be confusing. Jiang et al., 2022 have given this clade classification for Vibrionaceae family members based on phylogenetic analysis of 8 core genes. We have adopted this classification for our study and cited it wherever it is needed. We have explained this in the introduction section.

      Lines 23 and 39: F8 is described as a "novel lineage" and a "previously uncharacterized system." As discussed above, F8 systems have already been described and are annotated in MiST. Please revise these statements.

      __Response: __We thank the reviewer for this comment. We have revised this sentence throughout the manuscript.

      Lines 35-36: The statement could be interpreted as implying that the involvement of chemosensory systems in host colonization is unique to Vibrio. Since chemosensory systems are broadly distributed across bacteria and frequently contribute to host interactions, I suggest rephrasing this sentence to avoid overstatement.

      __Response: __We thank the reviewer for this comment. We rephrased this sentence.

      Lines 41-42: The conclusion that CheA and MCP proteins represent promising drug targets is not directly supported by the analyses presented in this manuscript. The study does not evaluate essentiality, druggability, inhibition, or therapeutic feasibility. I recommend removing this statement.

      __Response: __We thank the reviewer for this comment. We agree that this study does not directly evaluate the essentiality of CheA/MCP proteins as a therapeutic target. Therefore, we removed this part from the manuscript.

      Line 66: The sentence describing responses to environmental cues via "chemotaxis, quorum sensing, and response to nutrients" is awkwardly phrased, since chemotaxis itself often represents a response to nutrients. Please revise for clarity.

      __Response: __We have revised this sentence.

      Lines 69-70: The statement that signal transduction in Vibrio species is "highly precise" and senses chemical gradients "very accurately" require both a reference and a clearer explanation. Relative to what system or organism is this precision being evaluated?

      __Response: __We removed this sentence.

      Lines 76-78: Please reconsider the citation to Ulrich et al. (2005) and revise the associated statement accordingly.

      __Response: __We have corrected this part.

      Lines 100-101: Please provide a reference supporting the statement that 17 classes of flagellar chemosensory systems have been designated.

      __Response: __We have added reference for this sentence.

      Line 106: The manuscript states that our understanding of CSS architecture and function derives primarily from a limited number of model organisms, particularly Escherichia coli. However, later sections highlight the extensive literature on Vibrio cholerae chemotaxis. Since V. cholerae itself is one of the better-characterized organisms in this field, the rationale for emphasizing E. coli alone is unclear.

      __Response: __We agree with your point. We have removed the part of chemosensory system of Escherichia coli and focused on Vibrio cholerae.

      Lines 278-286: This paragraph appears to contain contradictory statements regarding the number of genera included in the analysis. Please clarify how many genera were included, which were excluded, and the criteria used for inclusion.

      __Response: __We have clarified the criteria for inclusion of genera in our study.

      Lines 289-291: Please provide a reference supporting the statement regarding the mutualistic association between Aliivibrio fischeri and squid.

      __Response: __We thank the reviewer for this comment. We have provided the necessary reference for this statement.

      Lines 291 onward: Several ecological and physiological statements in this section require appropriate references.

      __Response: __We have provided the necessary references for these statements. Also, we have shortened some parts of it.

      Lines 302-342: This section would benefit from substantial shortening or a clearer connection to the central theme of chemosensory system evolution.

      __Response: __We have moved this part into the supplementary material, since it is not directly related to the central theme of the manuscript.

      Lines 369-371: This statement is difficult to interpret. MCPs are generally much more variable in abundance than core chemotaxis proteins, and conservation of abundance alone does not demonstrate essentiality. Please clarify and revise this conclusion.

      __Response: __We agree with the reviewer. We have revised this statement in the manuscript. MCP protein shows more abundance than other che proteins.

      Lines 429-445: The contrasting interpretations of F7 and F8 distributions require additional supporting evidence or a more cautious presentation.

      __Response: __We thank the reviewer for this insightful observation. We have rephrased the statement in the manuscript.

      Lines 538-541: This sentence is difficult to follow and would benefit from reformulation. In addition, CheA domain architectures, including those associated with Vibrionales F6, F7, F8, and F9 systems, have recently been described in detail by Berry et al., 2023. The authors should discuss their observations within the context of this work.

      __Response: __We thank the reviewer for this insightful observation. We have reformulated this result part in accordance with the Berry et al., 2023 paper. We have mentioned the possible reasons behind the absence of P2 domain in Che-F6 as well as extra structured insertion between F8 and F9.

      Line 545: The phrase "additional insertion domain" is not well defined. If this region corresponds to a recognized domain, it should be identified explicitly. If it represents an insertion or a poorly structured region, more appropriate terminology should be used.

      __Response: __We thank the reviewer for this comment. We have used the terminology “an insertion” as it is not a recognizable domain.

      Figure 4: It would be helpful to include the identifiers of the proteins used in the structural comparisons.

      __Response: __We thank the reviewer for this comment. We have added the identifiers for the CheA proteins.

      __

      __

      __ __

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      This article has to be completely rewritten because it is now very confused and difficult to read, starting with the title: landscape, architecture, and bipartite organization are difficult terms to employ when describing the "chemosensory system" of bacteria. Systems, in my opinion, does not refer to assemblies, and chemosensing in bacteria refers to anything that involves quorum sensing. Nevertheless, not all bacteria use quorum sensing for flagellar motility, and if these Che-As (and MCPs) are found in "non-chemosensory" systems and are mapped in the abstract without much explanation, this should probably be thoroughly further discussed before debating the organization of the genome and the division of "chemosensory" genes.

      __Response: __We thank the reviewer for these comments regarding terminology and conceptual framing. We agree that bacterial signalling encompasses a wide range of signaling mechanisms beyond those involved in flagellar motility. In our manuscript, use of chemosensory system refers specifically to CheA/CheW/CheY/MCP-based signalling pathways as defined by Wuichet & Zhulin, 2010. We have revised the introduction to clarify the difference between chemosensory systems and other bacterial sensory systems such as quorum sensing. We also agree that not all CheA-containing pathways are necessarily associated with flagellar motility and because of this, we have used chemosensory array and chemostaxis in different contexts. Accordingly, we have moderated several statements throughout the manuscript and mentioned that some chemosensory pathways may regulate other cellular functions. We also agree that here wording “driven by bipartite genome architecture” should be changed to distribution of chemosensory systems across different replicons. Overall, following your suggestions, we have significantly changed the title, introduction, results and discussion.

      There are often few attempts to take into account the "modern" literature to describe the evolution of bacteria and vibrionales and their capacity for quorum sensing, particularly in the introduction and discussion. Such two lengthy, in-depth paragraphs about the biology and diversity of Vibrionales that span more than two pages and only include six references are rather inappropriate for a research article, particularly when the topic is chemosensing and genetics. The description of the Che family, whose abbreviation is never explained, is highly ambiguous and very confusing in comparison to very basic understanding about vibrionales. Despite their significance intracellularly, we typically gain little from reading this section of Ches. Without making a distinction between chemotaxis and quorum sensing, the third paragraph is intended to a lesson about one component systems, two component systems, and chemosensory systems. Which kind of audience do the authors hope to reach? Microbiologists? Examples of these systems that are better characterized can be found throughout the literature.

      Bacterial chemosensory systems are then abruptly introduced. "Chemosensory systems are extremely modular (?) and are typically organized as gene clusters within bacterial genomes". Which sensory genes? Which clusters? What families of bacteria? Which references? When it comes to the classification of chemosensory systems, only Gumerov et al. 2021 is cited. However, what role do other groups play in the study of bacterial Ches and their distribution among the vast diversity of bacteria, such as proteobacteria, actinomycetes, and firmicutes? In order to give a more pertinent Introduction, there are fewer acronyms to employ and undoubtedly more efforts to thoroughly analyze the literature.

      __Response: __We thank the reviewer for this detailed and constructive comment. We agree that the earlier introduction was giving more space to the general biology and diversity of Vibrionales and did not provide sufficient context on the development, diversity, and evolution of bacterial CSS. We have therefore substantially revised and shortened the introductory sections describing Vibrionales biology and diversity, while expanding and restructuring the section on bacterial signal transduction and chemosensory systems. Specifically, we have

      • reduced the number of acronyms and simplified the description of OCSs, TCSs, and CSSs;
      • explicitly distinguished chemosensory signaling from quorum sensing, emphasizing that quorum sensing is primarily a population-dependent signaling process, whereas chemosensory systems detect environmental cues and regulate cellular behaviors such as motility;
      • defined the Che (stands for Chemosensory) proteins and introduced the major components of the chemosensory machinery before discussing their organization; and
      • expanded the discussion of the diversity and evolutionary distribution of bacterial chemosensory systems beyond Vibrionales, incorporating additional literature covering their occurrence and diversification across major bacterial lineages. We have also revised the text describing chemosensory systems as “modular” to clarify that this refers to the combinatorial organization and evolutionary diversification of core signaling proteins, accessory components, and sensory receptors, rather than simply the co-occurrence of che genes within a genomic region. We have further added relevant references to provide a broader and more contemporary overview of bacterial chemosensory system diversity and evolution.

      It learns a little bit more about the chemosensory clusters (whose genes?) in V. cholera, but this information is not really helpful for the study's objective, which at the very end of this incredibly long and badly worded introduction is still ambiguous and essentially unknown. The non-chemosensory/chemotactic motile Vibrionale mutants (line 143) were built by which research team? The authors? Every sentence the authors utilize in the introduction, such as "motility and chemotaxis directly or indirectly contribute to the pathogenicity of bacteria, is somewhat disorganized without a citation (lines 152-155). "The complex (?) interaction between chemotaxis and virulence gene expression in pathogenic Vibrio species" is not better to consider. There are no references included, even while discussing the last three decades of V. cholera research. This is not really appropriate for publication. After a lengthy introduction to the many proteins that mediate chemosensing (quorum sensing) in bacteria, the study is restricted to informatics work, see material and methods, and finally limited on CheAs, which is fairly harmful. For correlation analysis, phylogeny, and structure modeling, the authors use previously available information. Therefore, what distinguishes all of these tables and figures from what is currently understood about bacterial genetics and evolution?

      Response: We thank the reviewer for this important comment. We agree that the original introduction contained excessive discussion of well-characterized chemosensory systems in Vibrio cholerae. We have therefore shortened this section, clarified the study objectives, and added appropriate references to statements concerning motility, chemotaxis, and virulence.

      Regarding the novelty of the study, our objective was not to rediscover individual chemosensory proteins, but to provide a comparative, order-wide analysis of chemosensory system evolution across Vibrionales. Using 116 representative genomes and ~10,000 additional genomes/MAGs, we resolved four distinct CSS lineages (F6-F9), characterized their contrasting genomic distributions and replicon localization, and identified F8 as an experimentally unexplored system. We further provide evolutionary evidence for distinct trajectories of these systems, including probable horizontal acquisition of F9, and reveal a conserved F6 core alongside more dynamic accessory system. We have also revised the introduction to make these objectives, findings, and the significance of the comparative analysis more explicit.


      Reviewer #2 (Significance (Required)):

      The majority of the figures are too little to make any sense. Additionally, there are some unexpected "surprises". For example, the authors' in silico data on Photobacterium toruni, E. coli, and Vibrio qinghaiensis while Introduction led us to anticipate or pick V. cholera as the primary target (see last part of introduction). Bootstrap analysis and a clear display of the clades are necessary for phylogenetic validation. A well-established protein structure (Che-A? Che-B? Other Ches?) is required as an unambiguous reference in order to validate protein structure modelling.

      I would also add that since Vibrionaceae is a family of g-proteobacteria in the order Vibrionales, it is not surprising that they are common traits in the genomes of proteobacteria and vibrionales. However, I'm not sure what the authors mean when they say that patchy and replicon-flexible groups are vertically inherited from g-bacteria. A set of genes (discrete CSS types?) that would be horizontally acquired from alphaproteobacteria are the subject of the same critical point. What is the duration of the convergence of Alphaproteobacteria and Vibrionales? Please refer to Sonnenberg and Haugen (2023) about "bipartite" genome and horizontal transfer.

      __Response: __We thank the reviewer for these suggestions. We have improved the readability of a few figures by increasing font size, clarifying clade annotations, and adding bootstrap support values to the phylogenetic trees. We have also revised the introduction to clarify that the study focuses on chemosensory systems across Vibrionales, rather than V. cholerae alone, and have highlighted Photobacterium toruni and Vibrio qinghaiensis because of their unusual CSS distribution across replicons. In order to have a well-established protein structure from AFDB for structures validation, we have used the pLDDT and pTM score for already available and predicted structures, respectively, and we have added this point in our methodology. Finally, we have revised our interpretation of vertical inheritance and horizontal acquisition, moderating the relevant claims. We have further refined the evolutionary interpretation of F6, F7, and F8 by placing their distribution within the broader Gammaproteobacteria context, while F9 is discussed separately based on its close phylogenetic association with Alphaproteobacteria and the evidence supporting its possible horizontal acquisition.

      __ Reviewer #3 (Evidence, reproducibility and clarity (Required)):__

      This manuscript by Rawool and Sharma is a comprehensive bioinformatics analysis of the chemosensory systems in Vibrionales, their chromosomal locations, likely evolution and acquisition, and differences in CheA architectures. The Abstract and Discussion sections are beautifully written, but the language in the rest of the paper, especially in the Introduction section, is difficult to follow at times (I have many comments below under minor points). An impressive amount of data was acquired and analyzed for this paper, including the identification of F8 systems in Vibrios, and I commend the authors for filling the knowledge gap with such a comprehensive data set. However, the paper overall is extremely long and very detailed, and it took me a very long time to plow through all of it.

      The Introduction section, for example, reads like a literature review from a dissertation, but without figures, and could be streamlined without losing impact. Lines 132-154, include many details on environmental sensing and virulence studies, but it reads like a long list of disconnected sentences with findings from different studies, rather than an integrated summary of current knowledge. The Methods section is also very detailed, and although I am not a bioinformatician, the level of detail often seems excessive. In the Results section, data sentences are often followed by discussion or qualification sentences, which makes the Results section even longer. So, while the science in this paper and its interpretation is sound, the paper needs reworking to make it more palatable for most readers. I think it will be difficult for most readers to remain engaged through the entire manuscript the way it is currently presented.

      Response: We thank the reviewer for their appreciation of our study. Following their suggestions, we have shortened the introduction part as well as the overall manuscript. In the methods part, we have included the detailed parameters used in our analysis, so that result can be reproducible for others. We agree about the lengthy result part, and we have shortened this part as well.


      Minor points:

      1. General - sometimes clades are written in capitals, sometimes not, and sometimes they are in italics and other times not. Was this intentional? Shouldn't the formatting be consistent throughout?

      __Response: __We thank the reviewer for noting down this inconsistency. There should not be variation in the formatting. We have carefully reviewed the entire manuscript and standardized the formatting of clade names throughout the text, figures, figure legends, and supplementary materials to ensure consistency.

      Line 35 - "moving towards nutrients and away from harm" is a very narrow interpretation of chemosensory systems, referring solely to chemotaxis, which only 1 of the 4 chemosensory systems in Vibrio likely supports. The statement here should be more inclusive.

      __Response: __We thank the reviewer for this important clarification. We agree that the original statement focused primarily on chemotaxis and did not adequately reflect the broader functional diversity of bacterial chemosensory systems. The sentence has been revised to emphasize that chemosensory systems mediate the detection of environmental cues and can regulate a variety of cellular behaviors, including but not limited to motility.

      Line 56 and line 702 - "V. cholerae" not "V. cholera" - a common victim of autocorrect

      __Response: __We thank the reviewer for identifying this typographical error. "V. cholera" has been corrected to "V. cholerae" at the indicated locations and throughout the manuscript.

      Line 64 - "marine sea"? Marine = of the sea. So, effectively "sea sea" = redundancy.

      __Response: __We thank the reviewer for noting this. We have corrected this part in the manuscript.

      Lines 67-68 - the English on these lines doesn't make sense to me. Perhaps "...reaching swimming speeds of 40-200 mm/sec, which requires 1-2 orders of magnitude more energy for propulsion than that required for Escherichia coli".

      __Response: __We thank the reviewer for noting this point. Since this sentence is not directly related to the chemosensory system, we have removed this line from the manuscript.

      Line 69 - Why mention V. alginolyticus here as an example of a Na+-driven flagellar motor when it's relevant for other Vibrios as well? Also, no reference is provided.

      __Response: __We thank the reviewer for noting this point. Since this sentence is not directly related to the chemosensory system, we have removed this line from the manuscript.

      Lines 79-95 - While introducing the proteins found in a chemosensory pathway, chemotaxis itself is given as the "pathway", whereas it should be indicated that it is an example pathway. Not all chemosensory pathways have CheYs that interact with FliM. And not all chemoreceptors in Vibrios have periplasmic sensing domains - some are cytoplasmic.

      __Response: __We agree with the reviewer on this point. We have revised the phrasing in the manuscript.

      Lines 91-93 - awkward sentence where the last clause reads like a non-sequitur.

      __Response: __We thank the reviewer for this observation. We have revised and streamlined the relevant paragraph to improve its clarity, organization, and overall flow.

      Lines 110-113 - The "function" of chemosensory systems are determined by their output, whereas the signals recognized determine their specificity. Please correct.

      __Response: __We thank the reviewer for this clarification. We agree that the signals recognized by chemoreceptors determine the specificity of a chemosensory system, whereas its function is defined by the downstream cellular response it regulates. Accordingly, we have revised the text to distinguish between signal specificity and system function.

      Line 111 - Aer is not an MCP as it is not a "methyl-accepting" receptor in E. coli. Change "MCP proteins" to "chemoreceptors" to be accurate.

      Response: We have changed this part.

      Line 115 = 43 MCPs; Line 149 = 45 chemoreceptors; Line 225 = 46 MCPs - there is variation in the total number of MCPS in Vibrios. Perhaps give a number range where appropriate (line 115, V. cholerae in general), and specific numbers where specific strains are mentioned.

      __Response: __We thank reviewer for noticing this variation in MCP gene numbers. We have revised this statement.

      Line 116 - "forms"

      __Response: __We have changed this part.

      Line 119 - Change "the" to "a" and what is meant by "double-layered membrane structure" since it isn't in the membrane? Please use a more accurate description.

      __Response: __We thank reviewer for this comment. We agreed that it is not a double-layered membrane, rather F9 CSS cluster form a double-layered appearance in the cytoplasm. We have rephrased this in the manuscript.

      Line 121 - Replace the comma with a semi-colon before "overall".

      __Response: __We have changed this part.

      Line 124 - You've already told us that F9 is a cytoplasmic array

      __Response: __We have changed this part.

      Line 226 - Change "during" to "via"

      __Response: __We have changed this part.

      1. Line 133 - "which is involved" and "indicating a link"

      __Response: __We have changed this part.

      Lines 137-139 - If V. cholerae shows a chemotaxis response to these chemicals, then it isn't clear why you would say that they sense the environment "through these CSS clusters" - only 1 cluster (F6) is known to be involved in chemotaxis.

      __Response: __We thank reviewer for this comment. We have corrected this statement.

      Line 140 - "for epithelial colonization"

      __Response: __We have changed this part.

      Line 151 - gene names should be in italics

      __Response: __We have changed this part throughout the manuscript.

      Line 153- "contributes"

      __Response: __We have changed this part.

      Line 156 - "are associated"

      __Response: __We have changed this part.

      Line 162 - "Studies" don't perform anything. It is an inappropriate subject. But kudos for using the word "lacuna" so eloquently!

      __Response: __We have changed this part.

      Line 179 - "from which a pie chart”!

      __Response: __We have changed this part.

      Line 187 - "using the ggsignif"

      __Response: __We have changed this part.

      Line 209 - "A total of 154"

      __Response: __We have changed this part.

      Line 219 - "using parameters the same"

      __Response: __We have changed this part.

      Line 225 - "the 46 MCP proteins were aligned"

      __Response: __We have changed this part.

      Line 240 - "was converted"

      __Response: __We have changed this part.

      Figure 1 - the labels under parts C, D and E are too small

      __Response: __We have increased the font size of the labels and revised the figures.

      Line 346 - "encodes"

      __Response: __We have changed this part.

      Line 349 - "with average values"

      __Response: __We have changed this part.

      Line 352 - defined HK and RR on lines 347-348

      __Response: __We have changed this part.

      Line 365 - replace "chemotaxis-associated" with "chemosensory-associated". Chemotaxis is not inclusive.

      __Response: __We have changed this part.

      Line 380 - shouldn't "cheA" be in italics?

      __Response: __We have changed this part.

      Line 281 - again, "chemosensory proteins" not "chemotaxis proteins"

      __Response: __We have changed this part.

      Line 382 - "in trends with" makes no sense

      __Response: __We have changed this part.

      Line 283 - "an average of 33"

      __Response: __We have changed this part.

      Lines 387-388 - I don't understand the comment in brackets - why was the protein count limited to 40?

      __Response: __We thank the reviewer for this comment. We agree that the rationale for the cutoff value was unclear. Therefore, we have revised the analysis and replaced the previously used cutoff of 40 MCP proteins with the average MCP protein count of 34 calculated from the dataset. The corresponding text has been updated in the revised manuscript.

      Lines 398-405 - It isn't clear which species have lost flagella, and did they also loose the F6 system? Did they retain the other systems? Lines 427-429 doesn't make it any clearer and I am interested to know.

      __Response: __We thank reviewer for this comment. The species from three clades namely Marinum, Rumoiensis, and Halioticoli don’t have any CSS proteins present in it and all these organisms are well-known to be non-motile, suggesting that they don’t have flagellar proteins. These species lost F6 as well as all other F classes.

      Another organism, Vibrio qinghaiensis Q67 does not encode F6 system, however it has F7 system.

      Line 423 - "HubP protein functions" makes no sense

      __Response: __We thank the reviewer for this comment. The text has been revised to clarify that HubP is a polar landmark protein that acts as an anchoring factor for the localization of chemotaxis arrays through its interaction with the ParC/ParP complex. We have given proper reference for this.

      Line 470 - why are your reporting 2 F7s and then explaining further on on Line 474 that this is a genome assembly artifact? The information is dislocated and should be streamlined.

      __Response: __We thank reviewer for this comment. As per our analysis, two F7 clusters are present in Vibrio qinghaiensis Q67 species. When we checked further, these clusters reside within nearly identical (~99% sequence identity) terminal regions spanning ~20.3 kb at both ends of the replicon. This is artifact during genome assembly process which involves duplicated terminal sequences. We have corrected the order of writing these details in the text.

      Line 499 - "with the smaller"

      __Response: __We have changed this part.

      Lines 536-538 - The F7 system in Vibrios isn't involved in chemotaxis signaling - please correct your language here.

      __Response: __We have revised the statement for clarification and added the figure number.

      Lines 550-551 - "Structural superpositions analyses similarity...." makes no sense. I don't understand what you are trying to say.

      __Response: __We have rephrased this sentence.

      Lines 612 and 630 report the same thing - sensing of bile and mucin. Are both instances necessary?

      Response: We have removed the redundant sentence.

      Figure 6B - "against" not "againts"

      __Response: __We have changed this text.

      Line 634 - "accounting for 69% of sensory inputs"

      __Response: __We have changed this part.

      Line 659 - Hiremath et al., 2015b isn't an appropriate reference for the sensory repertoire of PAS domains. Please reference a PAS domain review instead, e.g., Stuffle and Watts, 2021. PMID: 33647528; PMCID: PMC8169565.

      __Response: __We thank reviewer for this comment. We agreed that the reference cited was not completely based on PAS domain, but the authors have mentioned about PAS domain in their paper. However, we have also cited paper for PAS domain which you mentioned in the comment.

      Line 664 - "ligand" not "legend"?

      __Response: __We have changed this part.

      Line 740 = "chemosensory" instead of "chemotaxis"

      __Response: __We have changed this part.

      Line 744 - "between the extra domain..."

      __Response: __We have changed this part.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We believe that the manuscript has been substantially strengthened through the revision process. The main changes are summarized below:

      We substantially revised the Introduction and Discussion sections to better position our work relative to previous studies on starvation-dependent thermotaxis plasticity, neuropeptidergic modulation, and AWC function.

      We clarified throughout the manuscript the distinction between negative thermotaxis in innocuous thermal ranges and thermonociceptive responses to noxious heat. We now discuss more explicitly that these behaviors involve at least partly distinct molecular, cellular, and circuit-level mechanisms.

      We performed new experiments in ins-1 mutants. Unlike what was previously reported for thermotaxis plasticity, ins-1 does not appear required for starvation-dependent thermonociceptive plasticity in our paradigm (new Figure 6—figure supplement 1).

      We revised the analysis and terminology used for AWC calcium imaging data. We no longer use the “deterministic/stochastic” terminology and instead describe a starvation-induced shift from predominantly excitatory responses to a mixed distribution of excitatory and inhibitory responses. We also added new quantitative analyses and histogram representations of response distributions, as directly suggested by reviewers, to better illustrate this point.

      We performed new genetic interaction experiments using eat-4; flp-6 double mutants. These analyses revealed that glutamatergic and FLP-6 signaling act largely in parallel to mediate heat-evoked reversals after early food deprivation, while prolonged starvation reveals a hierarchical interaction between these pathways.

      We revised and clarified the mechanistic model figures accordingly, particularly regarding the proposed ASI → AWC signaling pathway and the role of ASI-derived neuropeptides.

      We improved the presentation and statistical rigor throughout the manuscript, including:

      - Replacement of heating power values by corresponding temperature increases,

      - Clarification of the rationale for using the 1-hour off-food condition as reference,

      - Expanded statistical reporting and multiple-comparison procedures,

      - Additional methodological details for calcium imaging, rescue validation, and cell ablation approaches,

      - Clarification of replotted datasets in figure legends.

      We also simplified the manuscript by removing experiments whose interpretation remained ambiguous (notably the nsy-1 and nsy-7 analyses).

      Below, we provide a detailed point-by-point response to all reviewer comments.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Thapliyal and Glauser investigates the neural mechanisms that contribute to the progressive suppression of thermonociceptive behavior that is induced under conditions of starvation. Several previous studies have demonstrated that when starved, C. elegans alters its preferences for a variety of sensory cues, including CO2, temperature, and odors, in order to prioritize food seeking over other behavioral drives. The varied mechanisms that underlie the ability of internal states to alter behavioral responses are not fully understood; however, there is growing evidence for a role of neuropeptidergic signaling as well as the capacity for functionally distinct microcircuits, formed by distinct internal states, to trigger similar behavior outcomes.

      Within the physiological range of C. elegans (~15-25{degree sign}C), starvation triggers a profound reduction in temperature-driven thermotaxis behaviors. This reduction involves the recruitment of the amphid sensory neuron pair AWC. The AWC neurons primarily act to sense appetitive chemosensory cues; however, under starvation conditions begin to display temperature responses that previous studies have linked to the reduction in thermotaxis navigation. Here, Thapliyal and Glauser investigate the impact of starvation on thermonociceptive responses, innate escape behaviors that are triggered by exposure to noxious temperatures above 26{degree sign}C or rapid thermal stimuli below 26{degree sign}C. They compare the strength of thermonociceptive behaviors, specifically heat-triggered reversals, in worms experiencing either early food deprivation (1 hour off food) or prolonged starvation (6 hours off food). Their experiments demonstrate a progressive loss of heattriggered reversals that is mediated by AWC and ASI neurons, as well as both glutamatergic and neuropeptidergic signaling.

      At the level of neural activity, this study reports that the transition from early food deprivation to prolonged starvation reconfigures the temperature-driven activity of AWC neurons from largely deterministic to stochastic. This finding is interesting in light of previous work that reported the opposite transition (from stochastic to deterministic) in temperature-driven AWC responses when comparing well-fed worms to those kept from food for 3 hours. This study also identifies neural and genetic mechanisms that contribute to differences in thermonociceptive responses at +1 versus +6 hours of starvation; confusingly, these mechanisms are partially distinct from those that contribute to differences in negative thermotaxis behaviors in well-fed and +3 hours of starvation worms (Takeishi et al, 2020). A limitation of this manuscript is that these differences are not particularly acknowledged or addressed, other than the hypothesis that independent mechanisms underlie negative thermotaxis versus thermonociceptive stimuli. However, this suggestion is not experimentally verified.

      We thank this reviewer for pointing to the interest of our work. The difference between previous work focusing on negative thermotaxis in the range of innocuous temperatures and our work focusing on thermo-nociceptive response is important and indeed deserves further clarification and a deeper discussion in the manuscript.

      Two major empirical evidence for a distinction between negative thermotaxis (as assessed in previous studies) and thermonociceptive plasticity (as assessed in our paradigm) were already included in the initial article version. First, we reported a decrease in average response in AWC neurons due to a shift in the distribution of response polarities from mostly up-response to a mix of ‘up-response’ and ‘downresponse’ after starvation, while previous results showed an increase in response probability of AWCs after starvation (Takeishi et al, 2020). Second, contrary to starvation-evoked thermotaxis adaptation, ASI neurons are required to orchestrate starvation-evoked plasticity in thermonociception. These observations already indicate differences at the circuit and cellular level. For the revision, we conducted further experiments to address the molecular level. We tested ins-1 mutants (see also specific point 3 by reviewer 2, below) and deepened this aspect in the discussion section of the revised manuscript. Previous study found that INS-1 signaling from the intestine is a major mediator of negative thermotaxis plasticity. In contrast, our new data show that INS-1 peptide does not seem critical in regulating starvation-dependent thermonociceptive plasticity (see new Figure 6-supplement 1). Taken together these three lines of empirical evidence support the notion that negative thermotaxis and thermonociceptive starvation-evoked plasticity involves at least partially distinct mechanisms and it seems therefore inappropriate to qualify this notion as purely hypothetical.

      Modification in the revised manuscript include extended introduction about the known thermotaxis regulation mechanisms (Introduction section), new Figure 6-supplement 1 about ins-1 and accompanying text in the result section as well as extended discussion about these differences (Discussion section)

      Multiple additional aspects of this study make the results difficult to synthesize with existing knowledge, including

      (1) Differences in - and insufficient discussion of - the magnitude and kinetics of thermal stimuli;

      We have included a better description of the stimuli characteristics in the revised methods section. The discussion section was deepened to better emphasize that different types of thermal stimuli have been used in different studies.

      (2) This study's use of "heating power" rather than temperature values when presenting behavioral results;

      Thanks for noting this point, which was indeed an unnecessary complication in the result display of the initial manuscript. We have changed ‘power values’ to corresponding ‘temperature increase’ in the revised figures.

      (3) The use of +1 hours starvation as a baseline instead of well-fed worms. Indeed, this last point reflects a noticeable experimental result that differs from previous studies, namely that at room temperature, the basal movements of well-fed and starved worms are not different. Such a surprising result warrants further quantification of worm mobility in general and could have prompted a set of experiments directly testing previously published thermal conditions to demonstrate that the new effects reported arise specifically from the use of thermonociceptive stimuli, as hypothesized.

      The consideration of on-food and off-food behavioral state is an important point indeed. We found that, at room temp, worms shift from dwelling on food (a state with high spontaneous reversal rate) to global search off-food (a state with low spontaneous reversal, after 1hr starvation) (see Figure 1B). Therefore, unlike the reviewer’s statement, the reported data highlighted key differences in the basal locomotion of worms in fed and 1hr starved conditions. Furthermore, these behavioral states have been characterized very deeply using high-content worm behavioural tracking in our recent publication: Thapliyal et al. 2023 (PMID: 37236963). Our choice of using 1hr as a baseline is primarily driven by the fact that an elevated baseline of spontaneous reversals on food decreased the dynamic range to monitor changes in heat-evoked reversals. 1-hour early food deprivation reduced spontaneous reversals and led to a mild attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. Additionally, we observed a clear progressive decrease in heat-evoked reversals with increased duration of food-deprivation, which we further used to dissect the mechanism underlying this plasticity. The choice of the 1hr food deprivation timepoint as a reference is further justified below.

      Finally, a previous report (Yeon et al, 2021) demonstrated differences in the impact of chronic versus acute neural silencing on starvation-dependent plasticity in the context of negative thermotaxis. We therefore wonder whether similar developmental compensation impacts the neural circuits that contribute to starvation-dependent plasticity in the thermonociceptive responses.

      Indeed, this is an interesting question. Our conclusions are so far based on ablation (with chronic effects). In order to gain insight on this question, future studies could address the impact of chronic vs acute silencing approach in starvation-dependent thermonociceptive responses. We have added an opening on this question in the discussion, as follows:

      “Another open question is whether ASI action takes place during development (prior to starvation), or more acutely with active signaling after starvation.”

      A weakness of this manuscript is that the introduction is insufficiently scholarly in terms of citations and the description of current knowledge surrounding the impact of internal state on sensory behavior, particularly given previous work on the impact of feeding state on thermosensory behavioral plasticity (Takeshi et al 2020, Yeon et al 2021) and chemosensory valence (Banerjee et al 2023, Rengarajan et al 2019, etc).

      To address this weakness, we have revised the introduction section of the manuscript and cited previous relevant research on the impact of internal states on animal behavior, including the papers suggested by the reviewer. We note that 2 out of 4 suggested citations were already present in the initial manuscript (though in the discussion section).

      Similarly, the authors' commanding knowledge of the distinction between thermotaxis navigation (especially negative thermotaxis) and thermonociceptive behaviors could be communicated in more depth and clarity to the readers, in order to contextualize this study's new findings within the previous literature.

      As mentioned above, we have deepened this aspect in the discussion section of the manuscript (with a dedicated paragraph). It is quite clear that starvationinduced plasticity in negative thermotaxis and thermonociceptive behaviors engage distinct mechanisms (at least in part). These differences include distinct alterations in AWC calcium activity, role of ASI neurons and INS-1 neuropeptide.

      Nevertheless, this study represents a solid addition to the growing evidence that C. elegans sensory behaviors are strongly impacted by internal states, and that neuropeptidergic signaling plays a key role in mediating behavioral plasticity. To that end, the authors have provided solid evidence of their claims.

      We thank this reviewer for the efforts in evaluating our manuscript, for the positive assessment of our work, and for highlighting some weaknesses which, we believe, have been addressed through the revision.

      Reviewer #2 (Public review):

      In this work, Thapliyal and Glauser tried to provide a mechanistic understanding by which animals modulate their neural circuit responses to control nociceptive behavior on the basis of the dynamic internal feeding state. It is an important study that adds to the growing body of evidence coming from multiple model systems. They have used elegant genetics, behavioral, and Ca-imaging experiments to demonstrate how the auxiliary thermosensory neuron pair, AWC, and one of the internal state-sensing interneuron pairs, ASI, respond to dynamic internal starvation state to modulate behavioral response to noxious heat. Interestingly, these neuron pairs use distinct molecular mechanisms along with some other unidentified neurons to suppress heat-induced reversal response under short-term and prolonged starvation. The experiments are well performed, supporting most of the claims and providing an important framework for future studies.

      I have some queries that, if answered, will certainly enhance the study.

      (1) The results suggest that ASI is one of the primary drivers for the starvation-evoked behavioral plasticity, which regulates AWC activity under prolonged starvation. It raises many important questions, including: (a) how starvation modulates ASI response to heat?, and (b) under prolonged starvation, whether ASI also promotes other, non-AWC, glutamatergic inhibitory neurons to suppress heat-induced reversal, and how?

      We agree with this reviewer that the mechanisms by which ASI detects and mediates starvation-evoked changes in our model is a very interesting (unsolved) question. However, addressing these questions empirically represents a substantial body of work that would go beyond the scope of the present report. E.g., is temperature-dependent activity in ASI even relevant? At present, we envision that ASI could either work acutely (during heat stimuli) or be modulated over much longer time frames (hours of starvation) as an internal state sensor. Therefore, there will be quite some exploration needed before we figure out the ASI-level regulation more fully (including the critical temporal aspect regarding cell activity, as well as quantitative and qualitative transmission aspects). It will be very interesting in future work to address these questions.

      (2) How does ASI regulate AWC activity? In the proposed model (Figure 8) authors suggested an independent, unknown signal, other than INS-32 and NLP-18, from ASI to regulate AWC activity. However, from the results, the existence of another signal is not very clear.

      Thanks for raising this point, which reveals a weakness in our graphical representation (in Fig. 8) that was not properly conveying our point. Our current work shows INS-32 and NLP-18 to be important in modulating heat-evoked reversals upon starvation. However, at the moment, we don't know if INS-32, NLP-18, both, and/or other neuropeptides from ASI modulate AWC activity patterns. The calcium imaging experiments in single, double and potentially triple mutants would answer these questions but are not within our current reach, given the time needed to carry out these experiments. However, we acknowledge this point and have changed the figure and its legend to state that the arrow connecting ASI to AWC activity pattern could potentially reflect the action of these neuropeptides.

      (3) Previously, Takeishi et. al. showed that ins-1 dynamically modulates AWC-AIAmediated thermotaxis behavior based on the feeding state of the animal. It raises questions whether ins-1 also contributes to noxious heat-induced reversal behavior.

      We thank the reviewer for this question. We have now quantified the phenotype of ins-1 mutant in our paradigm. Our data shows that INS-1 neuropeptide is not critical in mediating starvation-evoked thermonociceptive plasticity, unlike plasticity in thermotaxis behavior (See Figure 6- Supplement 1). Together with the differential activity patterns in AWC and the differential need for ASI neurons, these new data further consolidate the notion that starvation-evoked thermotaxis adaptation and noxious-heat avoidance engage separable molecular, cellular and circuit-level modulatory mechanisms. A specific discussion paragraph was added too.

      (4) Experiments with AWC fate conversion mutants (nsy-1 and nsy-7) were very good ideas; however, the results obtained were confusing. flp-6 mutant data suggest AWCoff would be essential for heat-induced reversal, especially at the low intensity stimulus level. However, the nsy-1 mutant-forming two AWCon neurons showed complete rescue at the low heat level, which is quite opposite. Similarly, although less prominent, eat-4 rescue experiments suggested both nsy-1 and nsy-7 should behave normally at high heat conditions, which was not the result observed.

      We appreciate this comment and the legit attempt to infer what we should expect from a worm with two AWCon or two AWCoff, respectively. From previous studies so far, it's not quite clear if cellular properties of newly formed AWCs in nsy-1 and nsy-7 mutants, including response to sensory cues, formed synapses and their partners, expression of neuromodulator and gap junctions, synaptic output are similar or different. We think further studies are required to first establish if FLP-6 and glutamate signaling (expression, release and action) from altered AWCs in nsy-1 and nsy-7 mutants are the same or different. Therefore, direct comparison between cell fate conversion mutants with flp-6 and glutamate would rely on too many assumptions at this stage. Considering this comment, the limited additional value of the data with nsy1 and nsy-7 mutants (in the absence of additional analyses) and the confusion it could trigger, we have decided to remove these non-essential data of the manuscript.

      Reviewer #3 (Public review):

      Summary:

      Thapliyal and Glauser show that hunger alters how C. elegans responds to noxious thermal stimuli. Using targeted neural ablation, mutant analysis, and live-cell functional imaging, the authors demonstrate that hunger changes the properties of AWC sensory neurons, which sense noxious heat. The authors further show that the effects of hunger on nociception require ASI neurons, which are known to respond to hunger and mediate the effects of food deprivation on behavior. Finally, the study uses mutant analysis to implicate glutamate and specific neuropeptides in thermal nociception and in the modulation of nociceptors by hungerresponsive neurons.

      Strengths:

      The study clearly shows a strong effect of hunger on nociception and documents a striking effect of hunger on the intrinsic properties of AWC sensory neurons, which respond to noxious heat. The study also clearly and compellingly demonstrates that ablation of hunger-responsive ASI neurons blocks the effects of hunger on nociceptive AWCs. These data, which constitute the kernel of the manuscript, are striking and exciting.

      Weaknesses:

      The study has some weaknesses that the authors should address.

      (1) Ablation of AWC neurons alters the basal sensitivity to noxious heat stimuli. This should be clearly noted in the description of the result and warrants some discussion.

      We thank this reviewer for raising this legitimate point. We have clarified this aspect in the results section of the revised manuscript, reading as follows:

      “Removal of AWC nearly abolished heat-evoked reversal behavior across all stimulus intensities and timepoints (Figure 2B and E). While one should keep in mind that potential indirect developmental effects might take place in neuro-ablation lines, this observation suggests that AWC plays an essential role in mediating the thermonociceptive response under both early food deprivation and prolonged starvation. Notably, in AWC-ablated animals, the residual response level was unaffected by starvation, suggesting that AWC might also be required for the expression of starvation-dependent plasticity.”

      The contrast with known function of the best-characterized sensory neurons mediating thermal nociception (AFD and FLP) is discussed as follows:

      “...Therefore, noxious heat-evoked activity in AWC varies widely according to context, which is in line with previous literature [18, 20, 41]. Interestingly, the role of AWC is distinct from that of AFD and FLP neurons, which are canonically linked to thermosensation and nociception [5, 14, 42, 43], but contribute only modestly to heat-evoked behavior in our assay conditions with between 1 and 6 hrs of food deprivation.”

      (2) Throughout the study, it seems that data are replotted in multiple figure panels. The authors should clearly indicate in the figure legends when this occurs. Also, the authors should ensure that statistical tests requiring multiple comparisons are correctly implemented and reflect the number of times experimental data are compared to a single set of control data.

      Thanks for raising this important point. We have clarified this aspect in the revised figure legends of the manuscript, and in the method section. In some instances, we reconducted some analyses to be perfectly rigorous in multiple comparison accounting. This did not lead to significantly different conclusions. The one exception was that the small effect of eat-4 mutation on spontaneous reversal went below significance threshold. We therefore removed this aspect of the result reporting and of the corresponding interpretation scheme, which became slightly simpler (Figure 3). Globally, this makes the story more focused.

      (3) How ASIs modulate AWCs remains unclear. The authors find that loss of INS-6, an insulin-like peptide provided by ASIs, partially recapitulates the effect of ASI ablation. This observation is not further developed, and instead, the authors characterize other secreted factors that seem to mediate sensitization of animals to noxious heat stimuli. While it is interesting that there are multiple opposing inputs into the nociceptor circuit, the essential connection between ASIs and AWCs that underlies the foundational observations in Figures 1 and 2 is not sufficiently characterized.

      Whereas we agree that how ASI modulates AWCs is only partially solved by our study, we should emphasize that our work identified two ASI-expressed neuropeptides that function to decrease reversal response after starvation: INS-32 and NLP-18. We initially set a lower priority on INS-6 because the reversal response level in starved mutants appeared lower than that in nlp-18 and ins-32. It is important to note that ins-32 and nlp-18 are not ‘generally potentiated’ mutants, but display reversal upregulation selectively following starvation, which placed them as strong candidates to selectively mediate ASI regulation. This said, it is also true that these two mutants (and ins-6 too) display reduced responsiveness at the early food deprivation time point. Therefore, none of the neuropeptide mutants was strictly identical to ASI ablated line, suggesting that the peptides might also work via non-ASI cells at the early food deprivation timepoint.

      Following this reviewer’s comment, we have attempted to complement our story with the idea of using a similar approach and rescue ins-6 with its endogenous promoter or ASI-specific promoter. Unfortunately, we failed to obtain rescue effects, and therefore these data (with a negative result) remain inconclusive (as we cannot guarantee that the rescue constructs were functional). We decided to keep these data aside in the revised manuscript. Globally, our point made graphically in Figure 7F remains valid. We have complemented the figure legend to mention that INS-6 could also potentially work from ASI, but it is not depicted as no ASI-specific data are available. In summary, our data suggests that the connection between ASI and AWC(s) might be established by the integrated action of multiple peptides and their receptors. Further calcium imaging experiments in single, double and potentially triple mutant(s) of peptides and receptors would be required go deeper in this question, which could be performed in future work.

      “...Additional neuropeptides (such as INS-6) may also be involved, but in the absence of direct evidence for their origin from ASI, they were not included in this scheme.”

      (4) The assertion that 'starvation reshapes AWC responses from deterministic to stochastic' is not clearly supported by the data. AWC neurons seem capable of showing different responses to thermal stimuli, and the probabilities associated with these responses change after fasting. The different kinds of responses are seen under basal and fasted conditions.

      We thank this reviewer for the comment. There is an activity response shift that is quite solidly described, including with new quantitative analyses of distributions (histograms in new Fig. 4CD and new Fig. 5C-D, accompanied by Kruskal-Wallis tests). Yet, we totally agree that the wording choice was inappropriate. We have furthermore changed our terminology to avoid using the terms “stochastic” or “deterministic” that were indeed a cause of confusion. We now use the terms “stimulus-locked responses” and describe the shift as “shift from mostly excitatory responses to a mix of both excitatory and inhibitory responses”. We have also included detailed methodology for characterization of traces and statistical analysis in the revised method section of the manuscript, together with the new analyses on peak polarity distribution.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers agree that the study is clearly presented and makes good use of behavioral, genetic, and imaging approaches to link starvation state with changes in AWC and ASI function. To strengthen the manuscript and ensure clarity for readers, we ask you to address the following points in revision:

      (1) Positioning and citations.

      Clarify how your findings relate to Takeishi 2020, where the opposite trend in AWC activity was reported, and make a clear distinction between thermonociception and thermotaxis. The introduction should also include additional citations in two specific areas: prior work on AWC and noxious thermal stimuli, and studies demonstrating starvation-dependent behavioral changes via altered neuropeptide release (e.g., Banerjee 2023; Rengarajan 2019).

      We have clarified this aspect with extension of the work cited in the introduction and extensive rewriting of the discussion sections.

      Our data shows that mechanisms underlying starvation dependent changes in thermonociception and thermotaxis show differences at the molecular, cellular and circuit levels. First, we see a decrease in average response in AWC neurons due to shift from mostly excitatory to a mix of excitatory and inhibitory responses in response to noxious heat after starvation, while previous study found an increase in response probability of AWCs after starvation (Takeishi et al, 2020). Second, contrary to thermotaxis behavior ASI neurons are required to orchestrate starvation evoked plasticity in thermonociception. And, finally, previous study found that INS-1 signaling from the intestine regulates thermotaxis behavioral plasticity while INS-1 peptide does not seem critical in regulating starvation-dependent thermonociceptive plasticity (new data in Figure 6 supplement 1).

      We have revised the introduction section of the manuscript and cited previous relevant research on AWC and noxious thermal stimuli and studies demonstrating starvationdependent behavioral changes via altered neuropeptide release including the papers suggested by reviewers. The extended discussion section reads as follows:

      “Starvation regulates thermonociceptive and negative thermotaxis plasticity via at least partly different mechanisms

      Previous studies showed that AWC plays an important role in starvationdependent plasticity in the negative thermotaxis behavior in an innocuous thermal range between 15 and 25°C [26, 33]. Negative thermotaxis involves the detection of thermal changes created by animal movement in spatial thermogradient (0.5°C/cm), the magnitude of the expected thermal changes approximating 0.01°C/s [26]. The starvation impact on negative thermotaxis was shown to (i) involve an up-regulation of AWC cell activity, (ii) rely on INS1 neuropeptide produced in the intestine and (iii) to occur independently of ASI neurons. In contrast, our study used thermo-nociceptive stimuli, with faster raising thermal slopes (~0.5-2°C/s, hence 50-200 times faster than those occurring for thermotaxis) and covering noxious temperatures (up to 28°C). Our results indicate that the regulation of thermo-nociceptive response by starvation (i) is linked to a shift in the distribution of AWC activity response polarities from mostly excitatory to a mix of excitatory and inhibitory response, (ii) relies on ASI and specific neuropeptide produced in ASI, and (iii) works independently of INS-1 neuropeptide. Therefore, our study complements our understanding of the modulation of temperature-dependent behavior in C. elegans with previously undocumented mechanisms at the circuit, cellular and molecular levels.”

      (2) ASI → AWC mechanism.

      Because ASI is central to your conclusions, please expand on how ASI is thought to act on AWC and/or other neurons. If an additional ASI signal is proposed beyond INS-32/NLP-18, mark this as speculative unless further rationale can be provided, and adjust the model figure accordingly.

      We have revised Figure 8 and its legend to clarify what is still hypothetical in the way ASI could affect AWC activity patterns and reversals. Note that the figure was also modified to integrate the conclusions made from epistasis analysis of eat-4 and flp-6.

      (3) AWC response description.

      The data support a shift in response distributions rather than a categorical switch from "deterministic to stochastic." Please adjust the language accordingly and provide a clear description of how traces were classified as "up, variable, or down," ideally with a quantification of the distributional shift.

      We agree that the term “stochastic” can convey different things, and because it was used for something different for AWC in the past, we should have avoided it. We have revised the nomenclature. What we observe can indeed be better described as a shift in the response polarity distribution. The article was revised accordingly. We also included the quantitative analysis and histogram representation, suggested in one of the specific comments, and added detailed methodology on the categorization of traces.

      When ASI is intact, we see a shift from mostly excitatory responses to an ~equal mix of excitatory and inhibitory responses (new Figure 4C-D). This effect is lost when ASI is ablated (new Figure 5C-D).

      (4) Presentation and statistics.

      In figure legends, indicate where datasets are replotted across panels and confirm that multiple-comparison corrections take account of repeated comparisons to the same controls.

      We have included these details in the revised figure legends, and a statement in the method section.

      (5) Methods clarity.

      Provide justification for using 1-h off-food as the baseline, with quantification of baseline mobility/reversal rates. Expand the calcium-imaging methods to describe the processing pipeline (ΔR calculation, baseline period, drift correction), and add a brief rationale if the approach deviates from common normalization procedures. Clarify how cell ablations were performed and verified for specificity, and how cell-specific rescues were confirmed. Please also acknowledge the potential for developmental compensation with chronic ablation.

      The justification of using 1hr off-food as baseline was made more prominent in the revised manuscript.

      Revised result section:

      “Starvation downregulates thermonociceptive responses in C. elegans

      To assess how the feeding state modulates thermonociceptive behavior in C. elegans, we compared responses across different durations of food deprivation (Figure 1A). Synchronized first-day adult animals were stimulated with a series of 4-s infrared pulses of increasing heating power (100, 200, 300, 400 W), causing temperature increase of +2°C, +4°C +6°C and 8°C at the surface of the plate (Figure 1A). Fed animals on food produced robust heat-evoked reversal response to heat, but they also displayed a very elevated baseline of spontaneous reversals (~38%). A 1-hour off-food condition reduced spontaneous reversals (from ~38% to ~10%) and led to an attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. More prolonged food deprivation led to a striking progressive reduction in thermonociceptive responses at every heating level, with responses after 6 hours of starvation approaching baseline spontaneous reversal rates (Figure 1B and C). This suggests a robust inhibition of nociceptive behavior caused by prolonged starvation. To determine whether this attenuation was due to the absence of nutrients or chemosensory cues, we conducted similar starvation experiments in the presence of food odor, with OP50 bacteria present on the petri dish lid (Figure 1D). The reduction in thermonociceptive response persisted, indicating that the effect is driven by the internal starvation state rather than external olfactory input.

      Although fed animals showed high sensitivity to noxious heat, they also displayed an elevated baseline of spontaneous reversals, which limited their utility as a control group by strongly reducing the dynamic range of heat-evoked reversal quantification and by complicating the quantitative comparison with food-deprivation conditions with much-reduced reversal baseline (Figure 1B). In addition, technical limitations in our calcium imaging setup would have prevented the intended follow-up analyses in fed animals. Based on these observations and technical considerations, we focused subsequent analyses, aiming at dissecting the circuit and molecular underpinnings of starvation-dependent plasticity, to the comparison of two off-food conditions with similar spontaneous reversal baseline: the early food deprivation condition (1-hour off-food, with high responsiveness to noxious heat) and the prolonged starvation (6-hour off-food with almost abolished noxious heat responsiveness).”

      In addition, the method section was modified as follows:

      - Calcium imaging details were added regarding ΔR calculation, baseline period, drift correction.

      - We now explicitly refer to the original respective articles describing the neuroablation lines.

      - We clarify that cell-specific transgene expression for rescue was confirmed using SL2::mCherry co-marker

      In the result section, we now explicitly address potential developmental compensation in genetic ablation backgrounds in the result section as follows: “…one should keep in mind that potential indirect developmental effects might take place in neuro-ablation lines”.

      Reviewer #1 (Recommendations for the authors):

      (1) The data availability statement is missing from the reviewed manuscript and should be included.

      Thanks, we have included the data availability statement in the revised manuscript.

      (2) We request additional information on how n's were determined for individual experiments, as well as the inclusion of post-hoc power measurements for all quantification.

      n were determined in agreement with previous studies using similar measures. No a priori power analyses were performed. A posteriori power analyses are not informative beyond the reported effect sizes and p-values (now reported in File S2). We clarified this in the statistical subsection of the method section.

      (3) In many cases, the specific statistical tests used are not clear or justified; more details should be provided, including the non-post-hoc test used. Are all tests one-way ANOVAs? For comparisons across genotype and starvation duration, two-way ANOVAs would likely be more appropriate. Also, the authors switch between Bonferroni post-hoc tests and Holm-Bonferroni post-hoc tests. What determined the use of one versus another?

      We have now clarified the statistical analyses used and provided full details in File S2. We have now more systematically applied two-way ANOVAs across all relevant analyses (with detailed parameters reported in File S2). When particularly relevant (e.g epistasis analysis between eat-4 and flp-6 mutations) the results of the two-way ANOVAs, is also explicitly stated in the result section.

      We also note that all multiple-comparison corrections were performed using the Bonferroni method. Previous mentions of Holm-Bonferroni correction were inaccuracies, and we apologize for this confusion; these mentions have now been corrected throughout the manuscript.

      (4) The use of heating power instead of the temperature experienced by the worms is an unwelcome abstraction. We strongly recommend revisiting that choice.

      We do agree. We have revised the figures to label the axis with temperature increase.

      (5) For calcium imaging, how are the traces categorized into "calcium up", "calcium down", or "no change"? Were those determined blindly - i.e., by individuals unaware of the experimental condition? Did the response direction need to be consistent across different temperatures? Did the change from baseline need to hit a specific threshold, consistent with previous studies in the field (i.e., +/- 3xSD for a minimum amount of time)? We encourage the authors to include these details in their methods section.

      We have complemented the method section to clarify the criteria for the qualitative classification of traces. More importantly, new quantitative peak polarities comparisons were added (see specific points below and above, about histograms).

      (6) For the experiments showing that exposure to food odor does not prevent response reduction, we suggest that feeding worms heat-killed bacteria would be a helpful control for the importance of bacterial nutritional status. In addition, showing that the impact of starvation was reversible with re-feeding would have been a useful experiment in line with standard experimental design in the starvation field.

      Thanks, indeed with our current work we cannot pinpoint the role of additional sensory cues (except food odor) to be mediating starvation-evoked plasticity. Together with refeeding, these are all extremely interesting questions that we aim to answer and potentially link with ASI and AWC activity in our future work.

      (7) For the various AWC rescue experiments, we found it curious that there wasn't an AWCon+off rescue, only each neuron individually.

      Previous studies have identified similar or opposite responses of both AWCs for distinct sensory cues. Though our calcium imaging experiments point to both AWC on and off having similar response patterns to heat, we cannot rule out the possibility that their output (ability of control reversals) is distinct possibly due to recruited neuromodulators. Therefore, in the present work, we examined where these cell types act via the same or distinct combinations of neuromodulators to control reversals.

      Reviewer #2 (Recommendations for the authors):

      Experiments suggested:

      (1) The authors should look into the Ca-dynamics in ASI.

      How does the spontaneous and heat-evoked activity of ASI differ in fed, early food-deprivation and prolong starvation and its link to releases of neuromodulators, modified AWC activity to alter output of thermal nociception are very interesting questions. However, these questions are extremely exploratory (see argumentation above in response to the public review) and addressing them goes beyond the scope of our current manuscript.

      (2) The authors should check AWC activity in ins-32 and nlp-18 mutant animals.

      In this study, we focused on the roles of INS-32 and NLP-18 released from ASI in modulating heat-evoked reversals, as these mutants exhibit relatively strong behavioral phenotypes. However, these effects remain less pronounced than those observed following ASI ablation. In addition, we cannot exclude the contribution of additional signaling molecules, including INS-6 and other neuropeptides.

      A comprehensive analysis of AWC activity in this context would require calcium imaging across multiple genetic backgrounds, including single, double, and potentially higher order peptide and receptor mutants, combined with cell-specific rescue experiments. While we appreciate the suggestion, such an approach would represent a substantial extension of the present work and will be important to pursue in future studies to further elucidate the underlying mechanisms.

      (3) Short-term food deprivation completely eliminated heat heat-induced reversal response to 100W stimulus, while the response to 400W stimulus remained unaffected. This suggests fed, short-term starvation, and prolonged starvation are three distinct states, and authors should also test the response of AWC and ASI ablated animals in the fed conditions.

      We agree that analyzing thermal nociception in fed states, in addition to short-term and prolonged food deprivation states is an important and interesting question, as these 3 conditions likely represent 3 distinct internal states that may recruit different neural pathways.

      Several reasons led us to set the fed condition aside for this study, and we realize we insufficiently explain them in the initial manuscript. There are 2 main reasons.

      (1) It is of paramount importance to consider the ‘baseline’ reversal rate (spontaneous reversals not triggered by heat, but visible in our dataset as the first point in the ‘dose-response’ curve). In Fed animals spontaneous reversal rate is very high (~38%) compared to the 1hr and 6hr food-deprivation conditions (>10%). This has two consequences: first a decreased dynamic range for quantify heat-evoked reversal, and, second, the difficulty in judging quantitative differences in heat-evoked reversals with such major differences in baseline reversals.

      (2) Experimentally, assessing calcium responses in truly fed animals presents technical challenges. With our current setup, animals must be removed from food for at least ~5 minutes prior to recording (followed by ~5 minutes of imaging), which effectively corresponds to a “freshly starved” condition rather than a fully fed state. While previous studies have used serotonin to mimic aspects of the fed state, such manipulations can be difficult to interpret in this context.

      A systematic comparison including fully fed animals, as well as AWC- and ASI-ablated conditions across these states, would be a valuable direction for future work, in particular once the methodological barriers associated with point 2, have been overcome.

      We have clarified these choices in the result section as follows:

      “Starvation downregulates thermonociceptive responses in C. elegans

      To assess how the feeding state modulates thermonociceptive behavior in C. elegans, we compared responses across different durations of food deprivation (Figure 1A). Synchronized first-day adult animals were stimulated with a series of 4-s infrared pulses of increasing heating power (100, 200, 300, 400 W), causing temperature increase of +2°C, +4°C +6°C and 8°C at the surface of the plate (Figure 1A). Fed animals on food produced robust heat-evoked reversal response to heat, but they also displayed a very elevated baseline of spontaneous reversals (~38%). A 1-hour off-food condition reduced spontaneous reversals (from ~38% to ~10%) and led to an attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. More prolonged food deprivation led to a striking progressive reduction in thermonociceptive responses at every heating level, with responses after 6 hours of starvation approaching baseline spontaneous reversal rates (Figure 1B and C). This suggests a robust inhibition of nociceptive behavior caused by prolonged starvation. To determine whether this attenuation was due to the absence of nutrients or chemosensory cues, we conducted similar starvation experiments in the presence of food odor, with OP50 bacteria present on the petri dish lid (Figure 1D). The reduction in thermonociceptive response persisted, indicating that the effect is driven by the internal starvation state rather than external olfactory input.

      Although fed animals showed high sensitivity to noxious heat, they also displayed an elevated baseline of spontaneous reversals, which limited their utility as a control group by strongly reducing the dynamic range of heat-evoked reversal quantification and by complicating the quantitative comparison with food-deprivation conditions with muchreduced reversal baseline (Figure 1B). In addition, technical limitations in our calcium imaging setup would have prevented the intended follow-up analyses in fed animals. Based on these observations and technical considerations, we focused subsequent analyses, aiming at dissecting the circuit and molecular underpinnings of starvationdependent plasticity, to the comparison of two off-food conditions with similar spontaneous reversal baseline: the early food deprivation condition (1-hour off-food, with high responsiveness to noxious heat) and the prolonged starvation (6-hour off-food with almost abolished noxious heat responsiveness).”

      (4) The authors should test the effect of ins-1 in noxious heat-mediated dynamic reversal behavior.

      We thank the reviewer for this valuable suggestion. We have now tested the phenotype of ins-1 mutants in our paradigm. Our data shows that INS-1 neuropeptide is not critical in mediating starvation-evoked thermonociceptive plasticity unlike plasticity in thermotaxis behavior (New Figure 6 Sup1). This molecular aspect adds to our initially presented evidence at the cell activity and circuit levels, that noxiousevoked reversal and thermotaxis behaviors are regulated in a clearly separable manner.

      (5) Whether Glutamate and flp-6 work in parallel or in the same pathway to regulate reversals?

      We thank the reviewer for this question. We tested eat-4; flp-6 double mutants and found:

      (1) After short-term food deprivation (1hr), flp-6 and eat-4 separately contribute to heat-evoked reversal at high & low heat and they act in parallel pathways to explain ~90% of animal responsiveness (new version of Fig 3)

      Corresponding new text:

      “Next, we focused on eat-4 and flp-6 mutants, showing the strongest phenotype. We addressed whether glutamate and FLP-6 signaling act dependently of each other in controlling heat-evoked reversal, by testing eat-4; flp-6 double mutants. The residual response seen in each single mutant (Figure 3 A and B) was almost entirely abolished in the double mutant (Figure 3C). A two-way ANOVA for the highest heat stimuli with eat-4 and flp-6 genotypes as factors (two levels each: mutant or wild type) showed significant main effects of eat-4 (F<sub>(1,67)</sub> =48.70, p<.001, η<sup>2</sup>p=0.421) and flp-6 (F<sub>(1,67)</sub> =58.50, p<.001, η<sup></sup>p=0.466), respectively, but no interaction effects (F<sub>(1,67)</sub> =0.093, p=.761, η<sup>2</sup>p=0.001). The significant cumulative effect of the two mutations indicates that the two signaling pathways act mostly independently of each other to mediate heat-evoked reversals.”

      (2) After prolong starvation (6hr), flp-6 mutation has a dominant impact on plasticity and eat-4 mutation cannot cause loss of plasticity, pointing to a hierarchy in this context (new version of Fig. 6, including revised hierarchy in the model in panel G, and also revised model in Fig.

      8).

      Corresponding revised text:

      “Second, we tested whether starvation-dependent plasticity was preserved in eat-4 and flp6 mutant backgrounds, which we had suggested to represent the main AWC transmitters controlling heat-evoked reversals under the early food deprivation condition (Figure 3). Even if the heat-evoked response upon early food-deprivation was reduced relative to wild type in flp-6 mutants, a significant further decline was seen after prolonged starvation (Figure 6B). These results indicate that starvation-dependent plasticity can operate independently of FLP-6. In contrast, eat-4 mutants displayed markedly elevated heat-evoked responses after prolonged starvation, even exceeding the response level seen in the early food deprivation condition for low heat stimuli (Figure 6C). This potentiated response in eat-4 mutants was entirely dependent of an intact FLP-6 signaling, since reversal responses in eat-4; flp-6 mutants were entirely abolished, like in flp-6 single mutant (Figure 6C-E, a two-way ANOVA indicating a significant interaction effect between the two mutations: F<sub>(1,70)</sub> =0.093, p<.001, η<sup>2</sup>p=0.247). Moreover, the potentiated response in eat-4 single mutant could not be rescued by expressing eat-4 rescue transgene selectively in either AWC<sup>OFF</sup> or AWC<sup>ON</sup> neurons (Figure 6F). Interestingly, AWC<sup>OFF</sup>-specific rescue produced a further potentiation of heat-evoked reversal response to high heat stimuli (Figure 6F, 6 and 8°C thermal increases), aggravating the phenotype of eat-4 mutants. These results are consistent with a model in which glutamatergic signaling regulates heat-evoked reversals in starved animals via two bidirectional drives (Figure 6G). On the one hand, glutamatergic signaling—originating from AWC<sup>OFF</sup>—up-regulates reversals in response to high heat stimuli, thus contributing to prevent starvation-induced thermonociceptive plasticity. On the other hand, glutamatergic signaling—originating from neurons other than AWC— down-regulates reversals over a broad range of heat intensities, thus promoting starvation-induced thermonociceptive plasticity. The latter glutamatergic signaling inhibitory effect seems to be more dominant and to depend on intact FLP-6 signaling.”

      (6) The authors should perform flp-6 and eat-4 mutant/rescue experiments in the nsy-1 and nsy-7 background to clarify the results.

      Our results indicate that glutamate release via EAT-4 from both AWC<sup>ON</sup> and AWC<sup>OFF</sup>, as well as FLP-6 from AWC<sup>OFF</sup>, contributes to heat-evoked reversals regulation. The experiments suggested by the reviewer would, in principle, provide further insight into the interaction between AWC subtype identity and the respective roles of glutamatergic and peptidergic signaling.

      However, as discussed in more details above in the public review, the extent and nature of AWC<sup>ON/OFF</sup> remodeling in nsy-1 and nsy-7 mutant backgrounds remain incompletely understood. This introduces significant uncertainty in interpreting results obtained from combining these mutations with eat-4 and flp-6 manipulations. As a result, such experiments would be difficult to interpret in a definitive manner at this stage.

      We therefore consider this an important direction for future work, once the roles of nsy1 and nsy-7 in AWC subtype specification and function are more clearly established. As our preliminary results with nsy-1 and nsy-7 mutants added more confusion than clarity, we have chosen to set them aside (former Fig. 3-figure supplement 2 has been removed).

      Minor comments:

      (1) What is food odor? The experiment should be clearly mentioned.

      Thanks for spotting this unintended omission. Food odor experiments were performed by adding OP50 bacteria on the inward side of the petri dish lid instead of the NGM surface. We have added this description in the method section of the revised manuscript.

      (2) Panel 3C is coming before 3B. This should be rearranged.

      Thanks for pointing this out. Panel arrangement was entirely reorganized in revise Fig. 3, with the addition of eat-4 x flp-6 genetic interaction analysis.

      (3) In Figure 6, if the panels are arranged horizontally, it would be easier to follow.

      Thanks for pointing this out. Panel arrangement was entirely reorganized in revise Fig. 6, with the addition of eat-4 flp-6 genetic interaction analysis.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors should consider moving measurements of AFD-ablated animals into the main Figure 1. AFD is a well-known thermosensor, and it is worth showing that responses to noxious thermal stimuli persist in animals lacking AFD.

      Thank you for this suggestion. We have moved measurements of AFDablated animals to the main figure (revised Fig. 2).

      (2) AFD ablation does affect responses to noxious heat. The authors could consider ablating/silencing AFDs and AWCs simultaneously to determine whether these two neurontypes account for the behavior.

      We agree that investigating the combinatorial contributions of thermosensory neurons, including AFD and AWC, to thermal nociception is an important and interesting question. In principle, simultaneous ablation or silencing of these neuron types could indeed reveal unexpected interactions.

      In our experimental paradigm, however, we observe only a minimal contribution of AFD neurons to heat-evoked reversals, whereas ablation of AWC nearly abolishes the response. Based on these observations, we chose to focus the present study on AWC, which appears to play a more prominent role in this behavior.

      A more detailed dissection of the potential interactions between AFD and AWC, including combinatorial manipulations, would be a valuable direction for future work. In particular, the possibility that AFD exerts a modulatory influence remains an interesting hypothesis to explore.

      (3) Given that EAT-4/VGLUT and FLP-6 neuropeptides each contribute to nociception, the authors should consider testing an eat-4; flp-6 double mutant to determine whether this combination of neurochemical signals accounts for AWC signaling to downstream circuits.

      We thank the reviewer for this suggestion, which is similar to point 5 of Reviewer 2 (above).

      We tested eat-4; flp-6 double mutants and found:

      (1) After short-term food deprivation (1hr), flp-6 and eat-4 separately contribute to heat evoked reversal at high & low heat and they act in parallel pathways to explain ~90% of animal responsiveness (new version of Fig 3)

      Corresponding new text:

      “Next, we focused on eat-4 and flp-6 mutants, showing the strongest phenotype. We addressed whether glutamate and FLP-6 signaling act dependently of each other in controlling heat-evoked reversal, by testing eat-4;flp-6 double mutants. The residual response seen in each single mutant (Figure 3 A and B) was almost entirely abolished in the double mutant (Figure 3C). A two-way ANOVA for the highest heat stimuli with eat-4 and flp-6 genotypes as factors (two levels each: mutant or wild type) showed significant main effects of eat-4 (F<sub>(1,67)</sub> =48.70, p<.001, η<sup>2</sup>p=0.421) and flp-6 (F<sub>(1,67)</sub> =58.50, p<.001, η<sup>2</sup>p=0.466), respectively, but no interaction effects (F<sub>(1,67)</sub> =0.093, p=.761, η<sup>2</sup>p=0.001). The significant cumulative effect of the two mutations indicates that the two signaling pathways act mostly independently of each other to mediate heat-evoked reversals.”

      (2) After prolong starvation (6hr), flp-6 mutation has a dominant impact on plasticity and eat-4 mutation cannot cause loss of plasticity, pointing to a hierarchy in this context (new version of Fig. 6, including revised hierarchy in the model in panel G, and also revised model in Fig. 8).

      Corresponding revised text:

      “Second, we tested whether starvation-dependent plasticity was preserved in eat-4 and flp6 mutant backgrounds, which we had suggested to represent the main AWC transmitters controlling heat-evoked reversals under the early food deprivation condition (Figure 3). Even if the heat-evoked response upon early food deprivation was reduced relative to wild type in flp-6 mutants, a significant further decline was seen after prolonged starvation (Figure 6B). These results indicate that starvation-dependent plasticity can operate independently of FLP-6. In contrast, eat-4 mutants displayed markedly elevated heat-evoked responses after prolonged starvation, even exceeding the response level seen in the early food deprivation condition for low heat stimuli (Figure 6C). This potentiated response in eat-4 mutants was entirely dependent of an intact FLP-6 signaling, since reversal responses in eat-4; flp-6 mutants were entirely abolished, like in flp-6 single mutant (Figure 6C-E, a two-way ANOVA indicating a significant interaction effect between the two mutations: F<sub>(1,70)</sub> =0.093, p<.001, η<sup>2</sup>p=0.247). Moreover, the potentiated response in eat-4 single mutant could not be rescued by expressing eat-4 rescue transgene selectively in either AWC<sup>OFF</sup> or AWC<sup>ON</sup> neurons (Figure 6F). Interestingly, AWC<sup>OFF</sup>-specific rescue produced a further potentiation of heat-evoked reversal response to high heat stimuli (Figure 6F, 6 and 8°C thermal increases), aggravating the phenotype of eat-4 mutants. These results are consistent with a model in which glutamatergic signaling regulates heat-evoked reversals in starved animals via two bidirectional drives (Figure 6G). On the one hand, glutamatergic signaling—originating from AWC<sup>OFF</sup>—up-regulates reversals in response to high heat stimuli, thus contributing to prevent starvation-induced thermonociceptive plasticity. On the other hand, glutamatergic signaling—originating from neurons other than AWC— down-regulates reversals over a broad range of heat intensities, thus promoting starvation-induced thermonociceptive plasticity. The latter glutamatergic signaling inhibitory effect seems to be more dominant and to depend on intact FLP-6 signaling.”

      (4) The authors should consider representing AWC responses to thermal stimuli as histograms to illustrate how fasting increases the probability of some responses and decreases the probability of others.

      We thank this reviewer for the suggestion. The proposed histograms nicely convey the concept of “shift in response polarity distribution” that we observed (using the new terminology we now use, instead of using the term “stochastic”). To create such histograms, we computed the magnitude of the peaks on a trial-by-trial basis. When ASI is intact, we see a shift from mostly excitatory responses to an ~equal mix of excitatory and inhibitory responses (new Figure 4C-D). This effect is lost when ASI is ablated (new Figure 5C-D).

      (5) It seems important to better understand the ins-6 mutant phenotype and determine whether ASI-to-AWC signaling involves this insulin-like peptide (ILP). The authors should consider using some of the tools available for disrupting ILP signaling to more clearly demonstrate that a specific neurochemical signal mediates modulation of AWCs by ASIs.

      Our work identified two ASI-expressed neuropeptides that function to decrease reversal response after starvation: INS-32 and NLP-18. We initially set a lower priority on INS-6 because the reversal response level in starved mutants appeared lower than that in nlp-18 and ins-32. It is important to note that ins-32 and nlp-18 are not ‘generally potentiated’ mutants, but display reversal up-regulation selectively following starvation, which placed them as strong candidates to selectively mediate ASI regulation. This said, it is also true that these two mutants (and ins-6 too) display reduced responsiveness at the early food deprivation time point. Therefore, none of the neuropeptide mutants was strictly identical to ASI, suggesting that the peptides might also work via non-ASI cells at the early food deprivation timepoint (as follow up data indicated at least for nlp-18).

      Following this reviewer’s comment (and the similar one in the public review), we have attempted to complement our story with the idea of using a similar approach and rescue ins-6 with its endogenous promoter or ASI-specific promoter. Unfortunately, we failed to obtain rescue effects, and therefore these data remain inconclusive (as we cannot guarantee that the rescue constructs were functional). We decided to keep these data aside in the revised manuscript. Globally, our point made graphically in Figure 7F remains valid. We have complemented the figure legend to mention that INS6 could also potentially work from ASI, but it is not depicted as no ASI-specific data are available. In summary, our data suggests that the connection between ASI and AWC(s) might be established by the integrated action of multiple peptides and their receptors. Further calcium imaging experiments in single, double and potentially triple mutant(s) of peptides and receptors would be required to go deeper in this question, which could be performed in future work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the three reviewers for their encouraging and constructive comments. We have addressed them by increasing clarity of the writing and adding details to the methods section that were previously missing or not stated clearly enough. We have included additional experimental data. Specifically, we tested how the osmotic effect of lactulose depends on lactulose dosage and colonization state, added data on cecum sizes of mice with different microbiomes, and tested how lactulose treatment in the active phase affects feeding behavior.

      Public Reviews:

      Reviewer #1 (Public Review):

      Greter et al. provide an interesting and creative use of lactulose as a "microbial metabolism" inducer, combined with tracking of H2 and other fermentation end products. The topic is timely and will likely be of broad interest to researchers studying nutrition, circadian rhythm, and gut microbiota. However, a couple of moderate to major concerns were noted that may impact the interpretation of the current data:

      (1) Much of the data relies on housing gnotobiotic mice in metabolic cages, but I couldn't find any details of methods to assess contamination during multiple days of housing outside of gnotobiotic isolators/cages. Given the complexity of the metabolic cage system used, sterility would likely be incredibly challenging to achieve. More details needed to be included about how potential contamination of the mice was assessed, ideally with 16S rRNA gene sequencing data of the endpoint samples and/or qPCR for total colonization levels relative to the more targeted data shown.

      We thank the reviewer for pointing out that we have not made the experimental setup clear in the text. One of the unique features of our metabolic cage setup is that the mice do not need to be housed outside gnotobiotic isolators, but that the whole system is placed inside an isolator. We have developed and published this system recently (Hoces et al, PLOS Biol 2022), including extensive testing for sterility/gnotobiosis. We have now adapted the main text to increase clarity on this issue (lines 76ff).

      Given that 16S sequencing of germ-free mice will typically produce false-positive reads, we used Blautia pseudococcoides as an indicator strain for contamination. This strain is present in our SPF mouse colony, forms spores that are highly resilient to decontamination measures, and has been the most likely contaminant in our gnotobiotic system. We have checked for presence of this strain in the cecum content of all our animals at the end of each experiment, and only included experiments which had a B. pseudococcoides signal below threshold level. We have now added this information to the methods section (lines 389ff).

      (2) The language could be softened to provide a more nuanced discussion of the results. While lactulose does seem to induce microbial metabolism it also could have direct effects on the host due to its osmotic activity or other off-target effects. Thus, it seems more precise to just refer to lactulose specifically in the figure titles and relevant text.

      We have adapted all figure legends to not contain interpretations, but rather state what was done in the experiments shown in the figures. We have also adapted the text in multiple places to soften the language and avoid overinterpretation of our experimental results.

      Additionally, the degree to which lactulose "disrupts the diurnal rhythm" isn't clear from the data shown, especially given that the markers of circadian rhythm rapidly recover from the perturbation. It is probably more precise to instead state that lactulose transiently induces fermentation during the light phase or something to that effect.

      We tried to make the argument that what we call disruption of the diurnal rhythm is acute, meaning that it is not disrupting the rhythm "chronically" (i.e., for longer), but that it recovers rapidly from this transient disruption. Given the confusion this wording is causing we are introducing this conceptually in the new version of the manuscript (lines 56ff).

      The discussion could also be expanded to address what methods are available or could be developed to build upon the concepts here; for example, the use of genetic inducers of metabolism which may avoid the more complex responses to lactulose.

      We also appreciate the mention of concepts from our study that can be built on in future studies, and we added a paragraph on potential further research. (lines 301ff).

      Despite these concerns, this was still an intriguing and valuable addition to the growing literature on the interface of the microbiome and circadian fields.

      We thank the reviewer for all their encouraging and constructive remarks!

      Reviewer #2 (Public Review):

      Summary:

      The authors aimed to investigate how microbial metabolites, such as hydrogen and short-chain fatty acids (SCFAs), influence feeding behavior and circadian gene expression in mice. Specifically, they sought to understand these effects in different microbial environments, including a reduced community model (EAM), germ-free mice, and SPF mice. The study was designed to explore the broader relationship between the gut microbiome and host circadian rhythms, an area that is not well understood. Through their experiments, the authors hoped to elucidate how microbial metabolism could impact circadian clock genes and feeding patterns, potentially revealing new mechanisms of gut microbiome-host interactions.

      Strengths: 

      The manuscript presents a well-executed investigation into the complex relationship between microbial metabolites and circadian rhythms, with a particular focus on feeding behavior and gene expression in different mouse models. One of the major strengths of the work lies in its innovative use of a reduced community model (EAM) to isolate and examine the effects of specific microbial metabolites, which provides valuable insights into how these metabolites might influence host behavior and circadian regulation. The study also contributes to the broader understanding of the gut microbiome's role in circadian biology, an area that remains poorly understood. The experiments are thoughtfully designed, with a clear rationale that ties together the gut microbiome, metabolic products, and host physiological responses. The authors successfully highlight an intriguing paradox: the significant influence of microbial metabolites in the EAM model versus the lack of effect in germ-free and SPF mice, which adds depth to the ongoing exploration of microbial-host interactions. Despite some methodological concerns, the manuscript offers compelling data and opens up new avenues for research in the field of microbiome and circadian biology.

      We thank the reviewer for their encouraging remarks, specifically on the surprising findings that microbial metabolism seems to affect circadian clock gene expression and behavior differently in EAM and SPF mice.

      Weaknesses:

      The manuscript, while providing valuable insights, has several methodological weaknesses that impact the overall strength of the findings. First, the process for stool collection lacks clarity, raising concerns about potential biases, such as the risk of coprophagia, which could affect the dry-to-wet weight ratio analysis and compromise the validity of these measurements.

      We thank the reviewer for pointing out that our description of the specific methods used for collecting feces were presented in a somewhat confusing manner. In short, dry and wet faecal weights were determined based on fecal pellets that were freshly produced and directly collected from restrained mice. To determine total fecal output over time, we collected all fecal pellets produced in a 5-hour window in a cage, determined their dry weight, and then used the water content determined for fresh faeces to calculate wet weight. Using this method, we cannot account for potential differences in coprophagia between the groups. However, this is not likely to affect the dry-to-wet ratio of faecal output in our results. We have now adapted the section in the methods to increase clarity (lines 440ff), and changed the quantity shown in Figure S2C and E to "water content", which is a more intuitive measure for the same thing.

      Additionally, the use of the term "circadian" in some contexts appears inaccurate, as "diurnal" might be more appropriate, especially given the uncertainty regarding whether the observed microbiome fluctuations are truly circadian.

      Similarly to our answer to reviewer 1 above, we appreciate this remark about imprecise language and have addressed this issue in the text and the figure legends. Indeed, we do not think the fluctuations in microbiota activity are truly circadian, but likely a result of the entrainment through the host's food intake.

      Another significant issue is the unexpected absence of an osmotic effect of lactulose in EAM mice, which contradicts the known properties of lactulose as an osmotic laxative. This finding requires further verification, including the use of a positive control, to ensure it is not artifactual.

      This is a good point. We have used this lactulose dosage specifically to induce microbial metabolism without causing osmotic diarrhoea and went to some lengths do demonstrate this (FigureS2C-E). In response to this comment (and one by reviewer 3 below about transit time), we have now performed additional experiments using higher lactulose dosage (new FigureS3). Our results indicate that the effect of lactulose on water content and transit time depends on microbiota complexity, with no change in fecal water content in 3MM mice even when treated with higher lactulose doses, and a stronger change in SPF mice. We now address this in the main text (lines 127ff).

      The presentation of qRT-PCR data as log2-fold changes, with a mean denominator, could introduce bias by artificially reducing variability, potentially leading to spurious findings or increased risk of Type I error. This approach may explain the unexpected activation of both the positive and negative limbs of the circadian clock.

      While we agree that our description of the qPCR method used for measuring circadian clock gene expression was lacking detail, we do not see how our analysis would lead to an increased risk of Type 1 error.

      Briefly, we use the standard ΔΔCt method to analyze our RT-PCR results. We first normalize gene expression values for each gene of interest to an internal housekeeping gene. Then, we use these normalized values to compare gene expression values in treatment vs control groups (or to the control group at time point 0 in the case of Figure 3C). We then use a log2 transformation to convert the logarithmic RT-PCR values to fold changes. We apologize for the confusing labeling in the figures, where we called the values "log2(fold changes)", which we have now changed to "log2(ΔΔCt) of expression" in Figures 3B, C and S4.

      The simultaneous activation of both limbs of the circadian clock is indeed a surprising result and somewhat complicates interpreting the effect of lactulose treatment on clock gene expression. We take it as evidence that it generally interferes with clock gene expression, while a clearer understanding of the effects would require further research.

      Moreover, the lack of detailed information on the primers and housekeeping genes used in the experiments is concerning, particularly given the importance of using non-circadian housekeeping genes for accurate normalization.

      It seems like the resource table describing these important experimental details was omitted in the original submission. We have now included it in the revised version (Table S1).

      The methods for measuring metabolic hormones, such as GLP-1 and GIP, are also not adequately described. If DPP-IV/protease inhibitor tubes were not used, the data could be unreliable due to the rapid degradation of these hormones by circulating proteases.

      We thank the reviewer for pointing out this omission. We have now added details of how we measured the metabolic hormones to the methods section, including the fact that we have added a DPP-IV inhibitor to the tubes at sampling (lines 469ff).

      Finally, the manuscript does not address the collection of hormone levels during both fasting and fed phases, a critical aspect for interpreting the metabolic impact of microbial metabolites.

      While we agree that it would be interesting to measure hormone levels also in the fed phase, a more thorough examination of hormone levels over the diurnal cycle, as suggested by reviewer 3, would be relevant for a full-scale follow-up. Given our data, we of course cannot exclude that there may be time-point-specific differences and therefore have softened the language around this conclusion to state that hormone levels are not acutely changed after a lactulose intervention at the time-points examined. (lines 318ff).

      These methodological concerns collectively weaken the robustness of the study's results and warrant careful reconsideration and clarification by the authors.

      Because of these weaknesses, the authors have partially achieved their aims by providing novel insights into the relationship between microbial metabolites and host circadian rhythms. The data do suggest that microbial metabolites can significantly influence feeding behavior and circadian gene expression in specific contexts. However, the unexpected absence of an osmotic effect of lactulose, the potential biases introduced by the log2-fold change normalization in qRT-PCR data, and the lack of clarity in critical methodological details weaken the overall conclusions. While the study provides valuable contributions to understanding the gut microbiome's role in circadian biology, the methodological weaknesses prevent a full endorsement of the authors' conclusions. Addressing these issues would be necessary to strengthen the support for their findings and fully achieve the study's aims.

      We thank the reviewer again for their careful and critical reading of our work, and for their constructive input. In the revised version of our manuscript, we address the reviewer's concerns by providing more methodological detail and additional experimental data.

      Despite the methodological concerns raised, this work has the potential to make a significant impact on the field of circadian biology and microbiome research. The study's exploration of the interaction between microbial metabolites and host circadian rhythms in different microbial environments opens new avenues for understanding the complex interplay between the gut microbiome and host physiology. This research contributes to the growing body of evidence that microbial metabolites play a crucial role in regulating host behaviors and physiological processes, including feeding and circadian gene expression.

      We thank the reviewer for their encouraging remarks!

      Reviewer #3 (Public Review):

      Summary:

      In the manuscript by Greter, et al., entitled "Acute targeted induction of gut-microbial metabolism affects host clock genes and nocturnal feeding" the authors are attempting to demonstrate that an acute exposure to a non-nutritive disaccharide (lactulose) promotes microbial metabolism that feeds back onto the host to impact circadian networks. The premise of the study is interesting and the authors have performed several thoughtful experiments to dissect these relationships, providing valuable insights for the field. However, the work presented does not necessarily support some of the conclusions that are drawn. For instance, lactulose is administered during the fasting period to mimic the impact of a feeding bout on the gut microbiota, but it would be important to perform this treatment during the fed state as well to show that the effects on food intake, etc. do not occur.

      We thank the reviewer for this important point. In the revised version, we include an experiment where we administer lactulose during the fed state and do not observe a significant change in food intake. We describe this in the text (lines 189ff) and in the new Figure S5C and D.

      To truly draw the conclusion that the current outcomes are directly connected to and mediated via an impact on the host circadian clock, it would be ideal to perform these studies in a circadian gene knock-out animal (i.e., Cry1 or Cry2 KO mice, or perhaps Bmal-VilCre tissue-specific KO mice). If the effects are lost in these animals, this would more concretely connect the current findings to the circadian clock gene network.

      We agree that these would be interesting experiments to follow up on the question how the observed effects are actuated by host functions. However, they would require a large amount of preparatory work (including rederiving the KO mice to get them germ-free in our gnotobiotic facility), we argue that they are beyond the scope of this study.

      Despite these reservations, the work is promising.

      We thank the reviewer for their encouraging assessment.

      Strengths:

      Attempting to disentangle nutrient acquisition from microbial fermentation and its impact on diurnal dynamics of gut microbes on host circadian rhythms is an important step for providing insights into these host-microbe interactions.

      The authors utilize a novel approach in leveraging lactulose coupled with germ-free animals and metabolic cages fitted with detectors that can measure microbial byproducts of fermentation, particularly hydrogen, in real-time.

      The authors consider several interesting aspects of lactulose delivery, including how it shifts osmotic balance as well as provides calculations that attempt to explain the caloric contribution of fermentation to the animal in the context of reduced food intake. This provides interesting fundamental insights into the role of microbial outputs on host metabolism.

      Thank you!

      Weaknesses:

      While the authors have done a large amount of work to examine the osmotic vs. metabolic influence of lactulose delivery, the authors have not accounted for the enlarged cecum and increased cecal surface area in germ-free mice. The authors could consider an additional control of cecectomy in germ-free mice.

      We thank the reviewer for pointing out the potential effect of the anatomical differences of germ-free and conventionally colonized mice. We agree that when comparing germ-free mice to SPF mice, the enlarged cecum area in germ-free animals could lead to differences in water release or uptake. However, this difference is smaller in gnotobiotic mice colonized with our minimal microbiota, even though their ceca are still slightly smaller than those of germ-free mice (new Figure S2F). While we agree that including control of cecectomy in germ-free mice could be interesting, we do not have the option of doing surgery on germ-free mice given our current experimental setup. We have now added information on cecum weight, a good proxy for cecum size, in the new Figure S2F done.

      The authors have examined GI hormones as one possible mechanism for how food intake is altered by microbial fermentation of lactulose. However, the authors measure PYY and GLP-1 only at a single time point, stating that there are no differences between groups. Given the goal of the studies is to tie these findings back into circadian rhythms, it would be important to show if the diurnal patterns of these GI hormones are altered.

      We fully agree that a deeper investigation of the diurnal fluctuations of hormone levels would be an interesting next step in studying whether perturbations in food intake can disturb these rhythms. Doing this for the whole rhythm would really require a full second study.

      In response to the reviewer's comments, we have changed the statements made around these data to point out just that hormone level fluctuations could not be detected during specific time points after lactulose treatment and therefore do not seem to explain the imminent behavioral changes (lines 318ff).

      Considerations of other factors, such as conjugated vs. deconjugated bile acids, microbial bile salt hydrolase activity, and bile acid resorption, might be an important consideration for how lactulose elicits more influence on ileal circadian clock genes relative to cecum and colon.

      We absolutely agree that investigation of microbial bile acid modification and their metabolism by the host would be an interesting topic for a follow-up study.

      Measurements of GI transit time (both whole gut and regional) would be an important for consideration for how lactulose might be impacting the ileum vs. cecum vs. colon.

      This is also an interesting point. While we did not add an experiment in which we specifically measure transit time to the revised version, we now measure total faecal output in a 5 h time period after PBS or lactulose treatment (Figure S3C). Faecal output is known to be a good proxy for transit time, and we see no significant difference between lactulose treatment (even with a two-fold higher dose than used before, new Figure S3) and PBS treatment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      (1) Line 126 - see the point in the public review, this data argues against disrupting the rhythm.

      See our response above

      (2) Line 156 - the metric used for water content is confusing. Why not just subtract dry weight from wet weight to get water content? The ratio is much harder to think about for me. Perhaps more importantly, this data is very confusing given that colonization seems to impact the activity of lactulose, which complicates the interpretation. Could be an interesting area for future study that you might highlight more in the discussion.

      We thank the reviewer for pointing out that our presentation of water content could be clearer. We have changed the dry/wet ratio we have used in the previous version to the "water fraction" (new Figure S2C, E; new Figure S3A,B), i.e., the per cent of weight of the wet sample that is made up by water. We would argue that this is a measurement that is easier to interpret than the difference suggested by the reviewer, because it is independent of the absolute sample weight.

      We also agree that it is surprising that colonization state changes the osmotic state of the gut environment and have now added text discussing that (lines 127ff).

      (3) Line 188 - the lack of expression changes in the distal gut (cecum/colon) potentially conflicts with the model, warranting additional discussion/qualifications. Is lactulose getting metabolized in the small intestine? Alternatively, does lactulose have a direct effect on the host? The current text seems to imply that lactulose is fermented in the colon, the fermentation products are absorbed, and then they only impact the ileum through circulation, which doesn't seem physiologically possible.

      It is possible that lactulose affects small intestinal tissue directly, but we show that the effect depends on microbial activity. Microbial activity is much larger in the large intestine than in the small intestine, which is why we hypothesize that it acts through systemic signals rather than locally. These systemic signals might be fermentation products impacting the ileum through circulation. We would argue that this is not implausible, given that there are well-documented systemic effects of fermentation products in circulation (e.g., den Besten et al, 2013). It is, however, also possible that, e.g., metabolism of fermentation products in the liver triggers a secondary signal that acts systemically.

      (4) Line 241 - need to weaken this sub-heading. The experimental design shows that fermentation products impact feeding behavior, but this does not necessarily imply that fermentation products are responsible for the lactulose effect.

      Done.

      (5) Line 256 - The lack of an effect in SPF mice is surprising and potentially conflicts with the model proposed. Given this and other confusing results (gene expression site specificity, osmolarity effects, etc) it seems prudent to be more cautious as to the potential mechanisms through which lactulose supplementation impacts host gene expression and feeding behavior, which would likely require a lot more experiments to provide a definitive answer.

      While the lack of an effect in SPF mice is surprising, we would argue that this is rather points towards the need for a better understanding of host-microbiota-diet interactions than a conflict with our interpretation. We do agree with the reviewer that more work is necessary for providing definitive proof of the underlying mechanisms of the observed effects and have adapted the language throughout the text.

      (6) Line 260 - A lot of text is devoted to the potential caloric effects of the fermentation products and the lactulose itself. Was this in response to a prior reviewer? Either way, it seems too speculative to me and I would recommend trimming it down and moving it to the discussion.

      We have rewritten this section to make it more concise and increase readability (lines 224ff).

      (7) Line 305 - Not fasting, just lower caloric intake, as shown by Figure 1c.

      We have changed the text to reflect that.

      (8) Line 371 - Too definitive given the current data. Need to qualify the interpretation here.

      We have adapted the text and qualified the interpretation.

      (9) Figure 2 - Need to revise the title, no data showing that the rhythm is disrupted.

      Done.

      (10) Figure 3b - Clarify what timepoint is shown in the legend.

      Done.

      (11) Figure 3c - label when the treatment started. Consider changing to 2-way ANOVA which is probably more appropriate than t-tests.

      We have added information on treatment start to the figure legend and have changed the statistical analysis to a two-way ANOVA (time, treatment).

      (12) Figure 4c - move to supplement as this negative data isn't sufficient to rule out an impact on the hypothalamus or liver. Even when only considering transcript levels it's possible that the timepoint is just not ideal.

      We agree that this data only represents a snapshot of gene expression at one timepoint after treatment and does not rule out involvement of hypothalamus or liver in this process. We have now moved the previous Figure 4C to the supplementary information (new Figure S6A).

      (13) Figure 5 - defined the "fermentation products" and their concentrations in the legend. Consider moving panels d, and e to the supplement - negative results with unclear interpretation. Clarify in the legend how many calories/g were assumed for the fermentation products and provide a scientific rationale for this decision. Modify the title to be more cautious - as discussed above.

      We thank the reviewer for pointing out that this was not clear. We have now added the formulation of the fermentation products to the figure legend and moved panels D and E to the supplement. We have also combined the former Figures 4AB and 5ABC into the new Figure 4.

      (14) Figure S1 - The patterns in panel a are really intriguing and could be discussed more. For example, what do you think is driving the rapid oscillations in E. rectale?

      We agree that the patterns are potentially interesting, but the fact that the patterns we observe do not replicate well between the two light-dark cycles we monitor suggest a large contribution of experimental noise. We therefore refrain from interpreting too much into this dataset.

      (15) Figure S2 - Could changing the metric for water content be more easily interpretable? Modify the title to better match the data shown.

      We thank the reviewer for pointing that out. We have now changed the metric in Figure S2 (and the new Figure S3) to % water in feces/cecum content, which is easier to interpret.

      (16) Figure S3 - need to specify the multiple testing correction used.

      Done.

      (17) Figure S4 title - replace "inducing microbial metabolism" with "lactulose".

      Done.

      (18) Figure S5 title - modify to weaken the causal claim.

      We have adapted all figure titles to conform to this comment.

      Reviewer #2 (Recommendations For The Authors):

      Greter and colleagues present an insightful manuscript investigating the effects of microbial metabolites, such as hydrogen and SCFAs, on feeding behavior and circadian gene expression at a single time point. Notably, they observed that these metabolites exert a significant influence on behavior in a reduced community model (EAM). However, this effect was not evident in germ-free or SPF mice, highlighting an intriguing paradox. The manuscript is well-written, with experiments that are thoughtfully designed and clearly rationalized. The study contributes valuable data to the poorly understood relationship between the gut microbiome and host circadian rhythms. While I find the manuscript compelling, I believe that providing additional experimental and analytical details would enhance clarity and rigor.

      Major Comments:

      (1) Additional clarity is needed regarding the stool collection process. Were the samples collected as fresh specimens, or was there a possibility of coprophagia before collection? Clarification on this point is important, as it could impact the results, particularly the dry-to-wet weight ratio analysis. Ensuring the collection process did not introduce this bias is crucial for the validity of these measurements.

      Thank you for pointing out that this was not stated clearly. Wherever we assessed water content in faeces, we used fresh samples that were directly collected from a live animal and immediately frozen in a closed container or analyzed. When we assessed total faecal output, we collected the total bedding from a cage, sorted out the faecal pellets, and only measured dry weight. Total wet weight output was then assessed by using the water content of fresh faeces of animals with the same microbiota and the same treatment as a correction factor. We have now made this clear in the methods section (lines 435ff).

      In cases where we performed total output measurements by collecting bedding, there was the possibility for coprophagia. We did not control for this but assumed that coprophagia will have a small effect that is likely similar between groups and should thus not affect the comparison.

      (2) Caution is advised in the use of the term 'circadian,' which is sometimes used when 'diurnal' might be more appropriate. For example, the title of the first results section could be revised to 'Host Feeding [or Diurnal] Rhythms Influence Microbial Metabolic Fluctuations.' Additionally, line 357 should likely use 'diurnal' instead of 'circadian.' It's important to note that 'circadian' implies that cyclical fluctuations would persist without environmental cues (e.g., feeding). Since it's not clear whether most microbiome compositional or functional fluctuations are truly circadian, 'diurnal' is likely the more accurate term.

      We thank the reviewer for pointing out our imprecise use of the term circadian. We have now adapted this throughout the manuscript.

      (3) The lack of an osmotic effect of lactulose in EAM mice is quite surprising, given that lactulose is known to be an osmotic laxative. Was this finding specific to EAM mice, or was a similar lack of osmotic effect observed in SPF mice? A positive control is necessary to verify that this unexpected result is not artifactual. If there is no osmotic laxative effect in SPF mice, an explanation is needed as to why this medication is not functioning as expected in these mice.

      This is a good point. While we initially thought that the lack of an osmotic effect was due to the specific lactulose dose we were using, we also did not observe an osmotic effect (measured by the water content of fresh faecal pellets produced after treatment) when we used twice the amount of lactulose in EAM mice (new Figure S3). However, in SPF mice, we did see an increase in faecal water content after treatment. This intriguing result suggests that the effect of lactulose as an osmotic laxative depends on the presence of a complex microbiota. We now show this data in the new Figure S3A and discuss it in the text (lines 127ff).

      (4) The authors state that 'To account for faulty measurements due to disruptive events and for measurement noise, some datapoints were excluded from the raw datasets.' It would be important for the authors to confirm that this data exclusion was unbiased, meaning it did not disproportionately affect one group over another, and that any exclusions affected groups randomly.

      This is a good point, and we analyzed this for the experiments we show in Figs 4A/S5A, 4B/S5E, 4C, and S5F, and discuss this in the methods part (lines 493ff). The resulting statistics is not fully conclusive: using a Chi-square test to check whether the probability of excluding values differs between experiments, we get significant differences (p=3.9 x 10<sup>-18</sup>). We would, however, argue that this is not surprising, as different experiments sometimes different in the number of times we needed to do maintenance work on the isolators, which could lead to actuation of the scales measuring feed values, and thus faulty measurements.

      We face the same problem within experiments: we found a significant difference in the probability to exclude values between treatment and control in the experiment shown in Figs 4A/S5A (p=1.7 x 10<sup>-6</sup>), but no significant differences in the experiments in Figs 4B/S5E and S5F (p=0.71, and p=0.10, respectively). While it is hard to strictly exclude an influence of our data exclusion strategy, these findings speak against a systematic effect of treatment.

      (5) The authors should provide details on the primers and housekeeping genes used in their experiments. It's crucial that the housekeeping gene is not circadian and is stable at all time points (PMID 17878933).

      In the previous submission, the main resources table was omitted by accident. We have now added it (Table S1), including details on the primers and reagents used for all experimental work.

      (6) The presentation of qRT-PCR data as log2-fold change is confusing. It's unclear what the numerator and denominator represent for this ratio or why such normalization was deemed necessary. Ideally, transcripts should be normalized to a housekeeping gene (as noted in a previous comment), not to a baseline measure of other same genes acquired from other mice. Log2-fold change is typically appropriate when comparing two measures from the same mice; however, in this study, the mice were euthanized, and the denominator is a mean of genes measured from other samples. This approach could introduce bias and might explain why both the positive and negative limbs appear to be activated by the microbial metabolites. It would be more rigorous to present these values as absolute gene expression levels.

      We thank the reviewer for pointing out that this was not described clearly in the previous version of our manuscript. We have now adapted the methods part to explain that all data showing RT-PCR data is normalized to a housekeeping gene. Only after that, we compare the gene expression levels of the treatment group to the control group (or to gene expression of the control group at timepoint 0 in the case of Figure 3C).

      (7) The use of log-ratios, with a mean as the denominator, could artificially reduce the variability in the data, potentially leading to spurious findings or an increased risk of Type I error.

      As pointed out above, we have used internal normalization to a housekeeping gene before comparing the resulting values of the treatment and control groups. This method (commonly known as ΔΔC<sub>t</sub> method) is a standard way of comparing gene expression values obtained by RT-PCR. The use of a log2 transformation is commonly used to convert the logarithmic data resulting from the RT-PCR measurement to a linear fold-change measurement. We do not see that this data analysis strategy should lead to an increased risk for producing false positives. We now explain this better in the methods section of the manuscript (lines 541ff), and have adapted the labels in Figures 3B, C and S4 to avoid confusion.

      (8) It is unusual that both the positive and negative limbs of the circadian clock are overexpressed following lactulose administration. The authors should provide data confirming that these genes are in counter phase to each other at baseline. This clarification would help readers better understand the effects of the experimental interventions. As it stands, this critical part of the results is quite confusing.

      We agree that this result does not allow a clear interpretation of the effect of lactulose on the diurnal rhythm of the host. While we agree that this would be interesting to understand in detail, we are merely taking this as a first indication that actuation of microbial metabolism during the inactive phase of the diurnal rhythm can lead to changes in clock genes. This claim is supported by our data.

      (9) Could the effects of the microbial metabolites be mediated by AMPK, a known nutrient sensor that can influence the post-translational modification of Cry proteins? It would be beneficial for the authors to explore whether these metabolites have a more direct, previously unknown mechanism of affecting the circadian clock, or if their effects are mediated through known signaling pathways such as AMPK.

      We agree that this would be a valuable path to continue investigating the effect of microbial metabolism on clock gene activity.

      (10) The methods section does not specify how the metabolic hormones (e.g., GLP-1, GIP, leptin, ghrelin) were measured in the experiments. It is important for the authors to confirm that DPP-IV/protease inhibitor tubes were used for hormone measurement, as these proteins can be rapidly degraded by circulating proteases. Without the use of appropriate tubes, this data cannot be reliably interpreted. Additionally, it would have been ideal to collect these hormone levels during both the fasting and fed phases, but it appears this was not done. This represents a significant limitation of the study and should be addressed in the discussion.

      We thank the reviewer for pointing out this omission, we have now added a description of our protocol to measure metabolic hormones to the methods section (lines 469ff).

      (11) If the samples were appropriately collected in DPP-IV/protease inhibitor tubes, the authors should consider measuring active GLP-1, as this would likely provide a more accurate assessment of GLP-1 activity.

      We agree that this would be a valuable next step, in addition to testing the effect of changes in microbial metabolism on the time traces of hormone levels.

      Minor Comments:

      (12) Line 350 appears to have an incomplete sentence, as it seems part of the first sentence in the paragraph has been inadvertently deleted. This should be reviewed and corrected for clarity.

      Done.

      Reviewer #3 (Recommendations For The Authors):

      Major comments:

      (1) Could the authors provide a deeper description about what they are referring to in the following statement? "...higher order interactions and microbial metabolism are variable..." it is difficult to interpret as written. Do the authors mean cross-feeding interactions?

      We have changed this sentence to clarify the meaning.

      (2) Could the authors explicitly state their hypothesis in the introduction and provide a brief, but deeper explanation of the intervention prior to the results section?

      This is a good point, we have adapted the text accordingly (lines 53ff).

      (3) Could the authors include a bit more information regarding the diet provided to the mice? If grain-based chow, please provide insights into the fiber source, etc.

      While we agree that it would be interesting to know what part of the mouse diet is available to the microbes, this is hard for the standard mouse chow that we (and most others doing experiments with mice) feed the experimental animals. We have now added more detailed information on the specific type of chow the mice were fed (lines 347f). We would argue that, because the control and treatment groups were always fed the same chow, the effect of the fiber source and other specifics are controlled for, even if we do not know them.

      (4) In figure 1A - cells/g does not seem to be the correct unit - # of copies/g feces perhaps?

      Cells/g is the appropriate unit, but it seems like we have not explained the way we arrive at this unit in sufficient detail. In short, we use a qPCR run on a known standard curve of bacterial counts (known cells/g values) to estimate these numbers from qPCR results from faeces. We have now explained this better in the methods section (lines 385f).

      (5) Figure 1B/C and Figure S1B/C are confusing - the legend states these measurements were taken over two days, however, the plot shows a single 12:12 LD period. Was the data averaged? It might be best to show each day separately (i.e., over a 48-hour period) rather than in one 24-hour plot. Then, the authors could also show the averages in the light period vs. the dark period in a separate, complementary graph.

      We thank the reviewer for pointing this out. The previous figures were indeed averaged over the two days of measurement and projected onto one 24h period for plotting. We have now changed Figures 1B,C and S1B,C to show the full 48h time windows.

      (6) Line 109 - The reviewer concurs that lactulose is a non-nutritive, synthetic disaccharide, however, in theory, lactulose may have a high heat increment, which could cause the animal to undergo metabolic responses to defend core body temperature (which also exhibits diurnal rhythmicity). Have the authors considered core body temperature rhythms, their connection to microbial metabolism, and the core circadian clock gene network in their model?

      This is an interesting thought. We have not measured body temperature in our experiments. As the heat increment from food is typically associated with metabolic activity, and lactulose is not metabolized, we do not expect its heat increment to be high, at least in GF mice. In mice with a microbiota, we agree that metabolic heat will be produced upon lactulose metabolism by the microbes, which could be a contributor the observed effect.

      (7) Figure 2A and corresponding text in lines 121 - 123 - indicate at what time the bar and whisker plots were taken (assuming at ZT8, but please be explicit in the figure and corresponding text). Could the authors also include statistics for these waveforms?

      Done.

      (8) In Figure 2B, the authors state that microbial metabolism had been restored to normal levels by ZT15, however, did this persist into the next light phase? It would be ideal if the authors could present these data in a similar manner to that shown in Figure S2A for SPF mice. Further, what are the statistical considerations here to describe changes in phase, amplitude, periodicity, etc.

      This is a good point. Unfortunately, we do not have H2 measurements for EAM and GF mice over comparable time periods as shown for SPF mice in Figure S2A. However, the food intake data shown in Figure S5BCD is a indicates that the food intake normalized in the second dark phase after treatment.

      (9) In lines 147 - 149 and in Figure S2B figure legend - assuming these measurements are from individual bacteria? Could this be stated clearly in the text or legend? Also, what are the statistical considerations? Were there significant differences in SCFA production between bacteria?

      We thank the reviewer for pointing out that this was unclear. We have now adapted the legend to explain the way these data were collected and added a statistical analysis.

      (10) The authors have done a large amount of work to examine the osmotic vs. metabolic influence of lactulose delivery - however, have the authors accounted for the enlarged cecum and increased cecal surface area in germ-free mice? Would an additional control be cecectomy in germ-free mice to be more in-line w/ SPF animals? Further, could the authors tie in these findings more explicitly and state how they pertain to the overall goal of the study? Is this simply to draw the conclusion that microbial biomass is increased w/ lactulose?

      We wanted to make sure that the effect we see with lactulose is due to microbial metabolism and not due to the induction of osmotic diarrhoea or other host-dependent effects, and we have added a statement to that effect (lines 127ff). We have not corrected for the change in cecum size between GF and EAM mice, as EAM size (and many other gnotobiotic mouse models) also have enlarged ceca relative to conventionally colonized mice, but have added a dataset showing how cecum size of EAM and GF mice compare and discuss this in the text (new Figure S2F, lines 135ff).

      (11) Is Figure 3A necessary?

      It might not be strictly necessary, but it can help with understanding the relation of the genes tested in B and C, and we would therefore like to keep it in.

      (12) Line 155 - 157 - the authors make the statement that dry/wet feces weight ratio is decreased in GF mice, but this does not appear to be statistically significant. Please adjust to state numerical differences were observed or provide statistics.

      We have changed the statement in the text.

      (13) Could Figure S3A be moved to the main Figure 3 as this may provide a more logical flow? qRT-PCR data is expressed as - log2 (fold expression), but relative to what? Could the authors provide further info about the control?

      The former Figure S3A (new S4A) and Figure 3B show the same data in slightly different ways. We therefore opted to keep only one of those illustrations in a main figure. The fold changes are always relative to the average of the PBS control group, which we now state explicitly in the figure legend.

      (14) The authors state that the qRT-PCR data shows that microbial metabolism of lactulose impacts peripheral circadian gene expression, but this conclusion seems simplified. Lactulose treatment only impacted the ileum circadian gene expression. Additional peripheral tissues (liver, adipose tissue, etc.) could be moved to this figure, i.e., move Figure 4C data. Why do the authors think the ileum was most impacted beyond GI hormones as discussed later in the manuscript? Could changes in bile acid deconjugation (i.e., BSH activity?) and/or bile acid resorption by the host in the distal ileum due to lactulose delivery be involved? Or is it simply due to differences in GI transit time (which was not measured in the current study)? Further, lactulose had minimal impact in SPF ileum, and in fact, shifted Cry1 in the opposite direction relative to EAM mice. Could the authors provide more insight into these disparate observations (line 196 - 200)?

      We agree that the statement "lactulose impacts peripheral gene expression" is oversimplified, and we have now adapted the text to avoid the impression that this is our conclusion. No tested tissues other than the ileum showed significant differences in gene expression at the time point tested, which does not rule out that other tissues would react to the treatment at that time point, or the tested tissues would do so at the tested time point. As we don't have a good enough understanding of what mechanism causes the gene expression changes in the ileum, we refrain from speculating in the text, even though we agree that this is an intriguing question.

      (15) Could the authors provide more insight into the statistical approaches used to assess amplitude, peak, nadir, etc. in Figure 3C? Was the co-sinor waveform tested?

      This is a good question. Even though this was a highly work-intensive experiment using many animals, we would argue that the noise level is too high and the coverage of the time analyzed too sparse to infer meaningful statistics on the fluctuations of gene expression over time. In the new version, we have changed our statistical analysis of this dataset to a two-way ANOVA (treatment, time) to better analyze this dataset.

      (16) The food intake decrease and interpretation following treatment (Figure 4A and S4A) is curious - all animals were gavaged and in EAM mice, many animals, regardless of PBS or lactulose are trending down in food intake rate/total intake. It seems to be more of an impact of gavage and not of treatment, which the authors somewhat acknowledge in Lines 228 - 230.

      We agree that gavage is a possible factor in future food intake of experimental animals, which is precisely why we used the PBS gavage as a control. Even though the difference is not large, we see a significant change in food intake when lactulose is given, but not when PBS is given (Figure 4A, S4A).

      (17) Could the authors provide a deeper rationale for their line of thinking for lines 234 - 240? What is the evidence that systemic effects are likely to occur 3 hours after lactulose delivery? Further, as stated in comment 13, could brain and liver data be moved to Figure 3/Figure S3 as an additional example of peripheral tissue clocks?

      We have added an explanation for the rationale we use to justify the 8h time point (it is 3h after the peak of H2 production, as shown in Fig2AB, which happens 5h after lactulose delivery). While we agree that the brain and liver data would also fit into the Figure 3, we have now moved all negative data to the supplementary information (in response to a comment by reviewer 1, and in a general effort to clean up the data in the manuscript).

      (18) The authors measure PYY and GLP-1 at a single time point and state there are no differences, yet, the goal of the studies is to tie this back to circadian networks. Would it be possible to measure these GI hormones over a 24-hour period to show that the diurnal patterns are altered?

      We fully agree that measuring the metabolic hormones over time would be very interesting. It is possible but would represent a major effort using many animals and a large amount of work. We would therefore argue that it is beyond the scope of this revision, but a good starting point for a follow-up study.

      (19) The authors state that the administration of fermentation products acutely altered circadian food intake, but the studies do not support that this change is connected to the circadian network. Suggest softening the interpretation of the findings.

      We have changed the language there to soften the interpretation.

      Minor comments:

      (1) The authors should consider when it is appropriate to refer to rhythms as diurnal vs. circadian, as each has a distinct meaning. Diurnal follows entrainment cues while circadian is endogenously driven (i.e., line 39, line 58).

      We thank the reviewer for pointing out this important difference, we have adapted this in the whole text accordingly.

      (2) Circadian rhythm should be plural throughout the manuscript (circadian rhythms).

      Thank you, we have changed that where we refer to host circadian rhythms generally.

      (3) Lines 54 - 63. Fermentation should be capitalized when used at the beginning of a sentence.

      Done.

      (4) Line 289 - This should be Figure 5D and 5E.

      Done.

      (5) Line 290 - heart should be cardiac.

      Done.

    1. Author response:

      eLife Assessment

      This paper introduces a valuable optical method for simultaneous in vivo multiphoton imaging of the mouse brain combined with DMD-based one-photon patterned photostimulation in different axial planes. The evidence for effective optical separation of excitation and imaging is convincing, although the in vivo data in the olfactory bulb suggest potential confounding factors arising if stimulation not only affects cell bodies but also neuronal processes. The work will be of broad interest to neurobiologists working in circuit and systems neuroscience, as well as to specialists in optical microscopy.

      We thank the reviewers for their constructive feedback. We are planning to address their concerns as detailed below. Specifically, in the revised manuscript, we will streamline and consolidate the text to include:

      (1) An in-depth discussion of the awake recording results in the main text.

      (2) Further discussion of light scattering, photo-stimulation specificity of targeting individual glomeruli and resolution.

      (3) A summary (including also a table) of the operating regime and comparisons with alternative techniques for patterned photo-stimulation and imaging of the ensuing responses. We will highlight the advantages and constraints of the current implementation of ADePT. In particular, here we explored a small set of spatiotemporal parameters to understand the limits of our technique and provide a proof of principle of the strategy. These parameters can be varied further depending on the exact research question.

      (4) Implementation considerations and technical guidelines for calibration and long-term stability of the rig for ADePT (i.e. ‘a how-to guide’).

      (5) A bill of materials and estimated hardware costs.

      Furthermore, we will provide additional controls and rephrase some of the statements in the text as suggested (e.g. replace ‘accessible’ with ‘simple’, etc.). We will correct the unfortunate grammatical errors, improve clarity of text, and update the references accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      In this methods paper, the authors introduce a novel and innovative imaging approach for simultaneous in vivo multiphoton imaging of the mouse brain combined with DMD-based one-photon patterned photo-stimulation in different axial planes. This is a highly exciting technique that enables the axial decoupling of optical imaging of deep neural circuits from surface photo-stimulation of spatially precise (tens of micrometres) brain spots. This method builds on previous developments from the same laboratory, combining DMD-based patterned photo-stimulation with in vivo electrophysiological recordings. To my knowledge, this is the first instance in which patterned photo-stimulation has been combined and axially decoupled from two-photon (2P) imaging.

      Beginning with a thorough characterisation of the optical resolution of the photo-stimulation system, the authors applied this method to the olfactory bulb (OB) network, in which sensory inputs are topographically organised at the surface of the OB and thus ideally suited to demonstrate the relevance of this approach. They first showed that this technique can be used to rapidly reveal connectivity patterns of OB output neurons and to identify sister mitral cells. In addition, they manipulated a specific glomerular inhibitory population and demonstrated that these neurons provide spatially heterogeneous long-range inhibition of OB output neurons, with differential effects on mitral and tufted cells (a result previously observed in a paper from the same lab: Banerjee et al., 2015, Neuron). Altogether, the data demonstrate that this technique is well-suited for high-throughput functional mapping of neural circuit properties. The results are compelling and illustrate both the significant advance represented by this method and its feasibility.

      We thank the Reviewer for their constructive input.

      Despite my initial enthusiasm, there are several concerns in the present study that must be addressed in order to rule out confounding observations and to resolve remaining uncertainties regarding photo-stimulation resolution. These include the following:

      (1) Spatial resolution: Although the authors provide convincing data on spatial resolution in vitro, several observations throughout the paper suggest that the effective photo-stimulation precision may be lower than initially reported. For instance, in Figure 2, the authors observe repeated responses in neighbouring glomeruli (e.g., glomeruli #3 & #5, #4 & #6). To what extent could light scattering along the X/Y/Z-axis above the targeted glomerulus recruit en passage axons, resulting in the inadvertent activation of multiple glomeruli?

      Indeed, we cannot rule out this possibility. Fibers of passage are a potential concern. This is why we systematically sample different light intensities and assess their impact on specificity of dendritic mitral and tufted cell responses within the glomerular layer (same axial-plane optical stimulation and imaging experiments, Fig. 2). We identify a range of intensities that on average result mostly in activation of the targeted glomeruli. Within the range of intensities used for identifying sister cells, >90% of responses were on the diagonal (targeted glomeruli) and ~5% pixels that cleared the signal significance criterion used were in off-target glomeruli, as stated in the text and quantified in Fig. 2f. In the revised manuscript, we will further clarify and expand on these points.

      A further observation concerns the presence of "inhibited" sister mitral cells (Figure 3). The authors claim this is reminiscent of the differential spike-timing reported between sister cells (Dwawale et al., 2010, Nat Neuro). However, observing both excitatory and inhibitory responses following stimulation of glutamatergic inputs is an altogether different matter, particularly given that sister mitral cells are reciprocally connected via gap junctions. This observation requires further clarification and raises serious questions about the effective resolution of the stimulation. Could the inhibited cell simply correspond to a non-sister mitral cell receiving disynaptic feed-forward inhibition?

      This is indeed what we think it is happening (i.e. disynaptic feed-forward inhibition as the Reviewer points out). We observe inhibition in some of the mitral cells in the field of imaging when we stimulate not their parent glomerulus, but other glomeruli in the neighbourhood. As the Reviewer points out, sister cells are connected via gap junctions, but they also receive inhibitory chemical synaptic inputs via their secondary (and primary dendrites) from other (not-their-parent) glomeruli mediated by numerous types of interneurons including the DAT+/GABAergic (a.k.a. superficial short axon cells) and granule cells. Our data is consistent with differential inhibitory input from other glomeruli on sister cells getting input from the same parent glomerulus. As it appears that we failed to present this point clearly in the initial submission, we will further expand along these lines in the revised manuscript. Briefly:

      First, we identify a photo-stimulation regime that results mostly in the activation of a given targeted glomerulus (and not of other glomeruli in the field of stimulation). To this end, we photo-stimulate and image ensuing neuronal responses in the same axial optical plane. We strobe (alternate) between monitoring dendritic mitral and tufted cell (enhanced) GCaMP responses within the targeted glomerulus and other glomeruli in the field of imaging, while varying systematically the light intensity (Figs. 2,3a; Suppl. Fig. 4b, Suppl. Fig, 5b-d;h-j). We use as criterion for specificity a condition when >95% of significantly responding pixels (above a statistically defined signal response threshold) lie within the anatomical boundaries of the targeted glomerulus. For each glomerulus (or pixel within a glomerulus) we compared the average light response across trials with the baseline reference distribution in the absence of light stimulation. If this value crossed the 99th percentile of the baseline distribution, the glomerulus/pixel within glomerulus was classified as responsive to the photo-stimulation.

      Second, using the minimal light intensity regime experimentally identified as ‘specific’ for targeting individual glomeruli in the field of photo-stimulation (< 5% significant activation of off-target pixels), we decouple photo-stimulation in the glomerular layer from monitoring responses of mitral and tufted cells in the deeper layers of the olfactory bulb (100-250 µm axial displacement). This approach enables us to map cohorts of sister (daughter) cells associated with any specific target glomerulus in the field of view (1,2,3…n) by monitoring excitatory (enhanced) responses of mitral and tufted cell bodies. A cohort of sister cells associated with glomerulus x<sub>i</sub> (daughters of glomerulus x<sub>i</sub>) is defined by those cells which show statistically significant excitatory responses (4 SD - standard deviations - above their baseline fluctuations) specifically in response to photo-stimulation of glomerulus x<sub>i</sub>.

      Third, in the process, as we photo-stimulate different glomeruli in the field of stimulation, we also observe at times suppressed (inhibitory) responses in a subset of the mitral and tufted cells (exceeding 3 SD in the negative direction their baseline fluctuations). These suppressed responses occur in response to photo-stimulating not the parent glomerulus of a given cell, but other glomeruli in the field. These experiments revealed that within a cohort of sister cells (daughters of glomerulus x<sub>i</sub>), only a subset of cells are suppressed by activation of glomerulus x<sub>j</sub>, and, in a few example cases, different cells are suppressed by activation of different glomeruli (e.g. x<sub>j</sub> vs. x<sub>k</sub>), presumably through disynaptic feed-forward inhibition (Fig. 3b iii; 3d; Suppl. Figs. 5f,g; l,m). These preliminary observations suggest that sister cells receive differential inhibitory inputs from glomeruli in the neighborhood. In the revised manuscript, we will expand to further clarify these points.

      To verify sister cell identity, the authors could confirm that the predicted sister cells share a similar odour receptive field compared to randomly selected mitral cell pairs. In their previous study employing analogous DMD-based photo-stimulation (Dhawale et al., 2010, Nat. Neurosci.), sister mitral cells did not exhibit such opposite response profiles (firing rate correlation of ∼0.7 between sister cells). Could the authors verify that a comparable activity correlation is also observed among the sister cells identified using ADePT in the present study? In Figure S6, the authors show recordings and stimulation of the same neurons co-expressing GCaMP and ChR2. Applying this experimental design to the mitral/tufted cell population (using a Tbet-Cre mouse transduced in the OB with both GCaMP and Chrimson virus) would constitute a valuable control to clarify the nature of these "inhibited" sister cells.

      We thank the Reviewer for the suggestion. We consider that the experiments shown here are proof-of-principle in nature, highlighting the potential of ADePT for mapping functional neural circuit connectivity. In our opinion, further investigating the logic of similarities and differences in the odor responses of sister mitral cells and the nature of inhibitory glomerular interactions forms the focus of future studies. We also note that firing rate correlations can be notoriously difficult to compare and interpret across experimental regimes (spikes vs. calcium imaging).

      An additional concern relates to the 21 out of 162 mitral cells that were activated by two distinct glomeruli - a finding that is incompatible with the established OB wiring diagram and that further challenges the claimed stimulation resolution.

      Indeed, this reflects some degree of non-specific activation of the targeted glomeruli as discussed above. In the revised manuscript, we will further highlight this issue.

      A critical control experiment is also absent: in a Thy1-GCaMP6 mouse lacking any light-sensitive opsin, do the authors observe any unintended side effects of photo-stimulation?

      In the revised manuscript, we will include an additional control as suggested by the reviewer. Within the range of intensities used, we did not observe significant modulation of GCaMP6s activity in mitral and tufted cells in mice lacking light-sensitive opsins.

      Regarding sister cells (Figure 3), tufted cells are not analysed alongside mitral cells in this dataset, whereas this is elegantly performed in Figure 5 using the DAT+ model. Could the authors also demonstrate how the technique can reveal the complete family portrait of sister mitral and tufted cells?

      We thank the Reviewer for the suggestion. We think that the differences between mitral and tufted cells are indeed very interesting to investigate, but in our opinion form the subject of future studies.

      (2) The authors have explored only a limited set of photo-stimulation parameters, primarily varying light intensity. They should present additional tests, such as varying the spot size (which appears to be arbitrarily fixed at 30-50 µm) and the z plane of stimulation. The level of activation can vary considerably: for example, in Figure 3a(iii), identical stimulations elicit responses of markedly different amplitudes (see glom#3 and #4). In Figure 2, 5 out of 15 glomeruli failed to respond - could the choice of z-plane account for this variability? The stimulation duration (50-150 ms) also appears somewhat arbitrary: can the authors demonstrate that the technique is compatible with finer temporal patterns (e.g., 10 Hz stimulation for 500 ms using 20 ms light pulses)? What are the spatiotemporal and axial scanning limits of this approach, and can two or three glomeruli be targeted simultaneously with temporally patterned stimulation?

      Indeed, here we explored a limited set of spatiotemporal parameters to understand the limits of our technique and provide proof of principle. These can be varied depending on the exact question. In the revised manuscript, we will clearly state what the constraints of the current implementation are and provide context for further optimizations. Briefly, in the current version, individual as well as multiple glomeruli can be photo-stimulated together and 20 ms per pulse regime in trains of pulses is feasible.

      (3) One particularly relevant application of this method would be to guide photo-stimulation based on prior functional measurements - for instance, by generating a photo-stimulation mask specifically targeting odour-responsive glomeruli. In the DAT-Cre × Thy1-GCaMP6 experiment shown in Figure 5e, which glomeruli are activated by a given odour, and how does this odor responsiveness influence the efficiency of DAT+ cell-mediated inhibition?

      We thank the Reviewer for the suggestion. We are thinking along exactly the same lines. In particular, we would like to investigate the relationship between the degree of overlap in odor responses of individual glomeruli and the strength and specificity of their inhibitory interactions mediated by DAT+ interneurons. In the revised manuscript, we will further expand on discussing this venue of study. We feel however that this investigation is beyond the scope of this technical report.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Koh and colleagues describe ADePT (Axially Decoupled Photo-stimulation and Two-photon Readout), a modular approach for combining patterned one-photon optogenetic stimulation with two-photon calcium imaging in independently controlled axial planes. The method relies on a digital micromirror device together with a motorized holographic diffuser to generate spatially confined stimulation patterns while imaging deeper neuronal populations. As proof-of-principle applications, the authors use the system to map excitatory and inhibitory functional connectivity in the mouse olfactory bulb by stimulating superficial glomerular circuits and recording responses from mitral and tufted cells in deeper layers.

      This is a well-executed Tools and Resources manuscript. The technical implementation is described in considerable detail, the optical performance is systematically characterized, and the biological experiments provide convincing demonstrations of the types of circuit questions that can be addressed using the method.

      Strengths:

      The greatest strength of the manuscript is the comprehensive technical characterization of the optical system. The authors carefully benchmark the spatial resolution, axial confinement, registration accuracy, calibration procedure, and practical operating limits of the setup. I found the extensive optical benchmarking particularly helpful, as it gives readers a realistic sense of the operating regime and practical limitations of the approach.

      Another strength is the high level of methodological transparency. The optical design, calibration procedures, stimulation strategies, and analysis pipeline are described in sufficient detail that an experienced laboratory could realistically evaluate whether the system is suitable for its own applications. This level of documentation is particularly appropriate for a Tools and Resources article.

      A further strength is the clear positioning of ADePT relative to existing approaches. The authors are transparent about the trade-off between spatial resolution and implementation complexity: ADePT does not provide single-cell photostimulation, but offers flexible axial separation, a large stimulation field, and cellular-resolution two-photon readout in deeper planes without requiring a full holographic stimulation system. This defines a credible and potentially useful experimental niche.

      The biological applications convincingly demonstrate the utility of ADePT. The experiments identifying sister mitral/tufted cells through selective glomerular stimulation and the mapping of heterogeneous inhibitory influences from DAT-positive interneurons illustrate the types of functional connectivity questions that become experimentally accessible with this approach. Importantly, the authors generally avoid overstating these biological findings and appropriately present them as proof-of-principle demonstrations of the technology.

      We thank the Reviewer for their constructive input.

      Weaknesses:

      The primary limitation is inherent to the method itself rather than the execution of the study. Because ADePT relies on one-photon patterned illumination, photo-stimulation remains restricted to relatively superficial structures and does not achieve single-cell spatial resolution. The authors appropriately acknowledge these constraints and clearly position the method within this operating regime. Consequently, ADePT occupies a useful niche for interrogating spatially organized functional units such as olfactory glomeruli or cortical barrels, rather than applications requiring single-cell precision or deeper tissue penetration.

      We agree. In the revised manuscript, we will further expand on these points, highlighting the limitations and advantages of ADePT compared to other techniques in a table discussing various operating regimes.

      Although the manuscript describes the approach as relatively simple and cost-effective, implementation still requires careful optical alignment, registration, calibration, and optimization. This does not diminish the value of the approach, but terms such as modular or accessible may better reflect the practical implementation than simple. Likewise, a brief bill of materials, approximate add-on cost, and indication of which components are essential versus substitutable would help prospective users assess the accessibility of the system.

      We agree. We will change the text accordingly and provide the additional information as suggested.

      Finally, the manuscript provides an impressive level of technical characterization, but much of the practical guidance for adopting the system is distributed across the Results and Discussion. Bringing together the principal limitations, recommended operating regime, expected calibration workflow, evidence for long-term alignment stability, and the circumstances in which ADePT is preferable to alternative approaches would further strengthen the manuscript as a community resource.

      We agree. We will proceed accordingly in the revised manuscript.

  2. Aug 2026
    1. Reviewer #3 (Public review):

      Summary:

      This manuscript asks whether two forms of explicit strategy use in visuomotor adaptation, i.e., algorithmic mental rotation and retrieval of a cached aiming solution, differentially influence implicit recalibration. The question is relevant because much prior work treats explicit strategy as a unitary process, whereas the algorithmic/retrieval distinction is theoretically meaningful and grounded in cognitive theory. Across three experiments, the authors report that algorithmic strategy conditions initially produced broader fitted implicit generalization functions than retrieval conditions, but that this difference was reduced or eliminated when reach variability and sensory prediction errors were more tightly controlled.

      Strengths:

      The paper is clearly written, theoretically well-motivated, and employs a commendably transparent and progressive experimental logic. The three-experiment structure, in which confounds are systematically identified and addressed, represents a strong model of cumulative experimental design (I will certainly use it in teaching courses on experimental methods):

      Experiment 1 establishes an apparent difference in implicit generalization breadth. Experiment 2 attempts to reduce error spillover from Non-Critical targets by increasing angular separation and using delayed endpoint feedback. Experiment 3 uses an error-clamp design to decouple variable reaching from error feedback. This sequence is appropriate for testing whether the initial difference reflects a strategy-dependent change in implicit recalibration or instead follows from the distribution of movement plans and error exposure. The authors also provide reaction-time and performance data that are broadly consistent with the intended distinction between algorithmic and retrieval-like task performance.

      Weaknesses:

      The evidence does not support the strongest claims made in the manuscript, namely that algorithmic and retrieval strategies generally do not reshape implicit recalibration.

      In general, I am skeptical of the authors' interpretation of null results. Several central conclusions depend on non-significant group differences, especially in Experiment 3. Non-significant tests are repeatedly treated as evidence that groups are equivalent or that confounds are absent (e.g., implicit recalibration magnitude (Algorithmic: 11.43 {plus minus} 6.43{degree sign}; Retrieval: 15.49 {plus minus} 8.99{degree sign}; t(38) = −1.65, p = .11), adaptation level before Exclusion probes (F(1,256) = 3.04, p = .08) and Exclusion RT differences (F(1,266) = 3.15, p = .08), whereas a modest model-dependent breadth effect (bootstrap p = .02) is treated as meaningful (for more on the model-dependent breadth effect, see below).

      Without confidence intervals, equivalence tests, or Bayesian analyses, I think that the authors' interpretations comprise an inferential gap. A failure to find a significant difference is not equivalent to evidence of equivalence, particularly given that the implicit recalibration signal gets progressively attenuated across experiments (Experiment 1: ~16-17{degree sign}; Experiment 2: ~11-15{degree sign}; Experiment 3: ~7-8{degree sign}). With a substantially diminished signal in Experiment 3, the null result could partly reflect reduced statistical sensitivity rather than true equivalence.

      My main technical concern is the analysis of generalization breadth already alluded to. The central claims rely on group-level Gaussian fits to only seven Exclusion probe locations spanning −45{degree sign} to +45{degree sign} around the Critical target. In several cases, the fitted centers and widths are poorly constrained by the sampled range. For example, in Experiment 2 the algorithmic group's fitted center is shifted to approximately 29{degree sign}, meaning that the probe range samples the function asymmetrically relative to its own peak. In Experiment 3, fitted centers are near or outside the sampled range, while estimated widths are very broad. Under these conditions, the width parameter may partly reflect extrapolation or parameter trade-offs between center, amplitude, and width rather than a genuine difference in generalization breadth.

      Lastly, I think that the authors' use of an error-clamp paradigm is, from an experimental point of view, quite elegant. By controlling the sensory prediction error independently of reach direction, they can isolate implicit recalibration from the confounds identified in Experiments 1 and 2. However, I see a fundamental problem or question concerning construct validity here: In Experiments 1 and 2, the algorithmic strategy was operationalized as participants computing a counterrotated aiming direction in response to a visible cursor rotation. This is a naturalistic context where mental rotation is both required and meaningfully connected to task success. In Experiment 3, however, there is no visuomotor rotation to compensate for. The error-clamp renders the cursor feedback task-irrelevant. Instead, participants are instructed via text commands (e.g., "move towards 45{degree sign}") to reach invisible locations, rendering the "algorithmic strategy" in this context essentially an instructed spatial navigation toward arbitrary angular locations, not genuine visuomotor mental rotation driven by an error signal.

      To put it differently, are we sure that the cognitive process engaged by the algorithmic group in Experiment 3 is the same as the algorithmic mental rotation strategy in Experiments 1 and 2? If not, then the null result in Experiment 3 may not speak to the original question about how algorithmic strategies interact with implicit recalibration after all. Instead, it may reflect the absence of a genuine strategy manipulation.

      To their credit, the authors report a compelling RT dissociation that mirrors Experiments 1 and 2: The algorithmic group shows slower RT, which is decreasing over training (0.98s → 0.76s), whereas the retrieval group exhibits faster, stable RT (0.52s → 0.45s). While this pattern is consistent with genuine strategy differences persisting in Experiment 3, it could also reflect the greater spatial precision demands of reaching to invisible targets from text instructions, rather than genuine mental rotation per se. Reaching to an invisible location defined by a verbal angular label is inherently more demanding than reaching to a visible target, regardless of strategy type, and this demand is asymmetrically present in the two groups, since Non-Critical targets are invisible for the algorithmic group but visible for the retrieval group.

      Thus, from my point of view, experiment 3 should not be used as definitive evidence that algorithmic and retrieval strategies during standard visuomotor adaptation cannot differentially influence implicit recalibration.

      Overall, the manuscript addresses a meaningful question and the multi-experiment structure is useful. The evidence is incomplete for the broad claim that implicit recalibration is insensitive to strategy type. The study would make a clearer contribution if the authors narrowed the claims, strengthened the generalization analyses, and treated null effects with appropriate inferential tools.

    2. Author response:

      Reviewer #1 (Public review):

      A previous study from the same team (McDougle & Taylor, 2019) demonstrated that explicit strategies during visuomotor adaptation can be dissociated into retrieval-based and algorithmic strategies. However, whether these distinct forms of explicit processing differentially influence implicit recalibration has remained unresolved, with previous studies providing evidence both for relatively independent explicit and implicit processes and for interactions between them. This study addresses this question through a series of experiments that used Critical and Non-Critical targets to induce distinct strategic modes while maintaining comparable adaptation at the Critical target.

      Experiment 1 replicated previous findings showing broader implicit generalization under algorithmic strategies. However, this broader generalization could be explained by spillover effects arising from adaptation at the Non-Critical targets. Experiment 2 was designed to reduce such spillover effects by increasing the spatial separation between the Critical and Non-Critical targets. Although broader generalization was still observed in the algorithmic condition, this effect was interpreted as reflecting greater variability in reaching behavior at the Critical target. Finally, Experiment 3 introduced additional controls using an error-clamp paradigm, and the difference in generalization width between the two strategies largely disappeared.

      We appreciate the reviewer’s thoughtful and comprehensive summary of our study.  One thing we would like to clarify is that the non-critical targets used in Experiment 1 received the same type of online continuous feedback as the critical target.Therefore the extent of implicit recalibration should be comparable between the critical and non-critical targets in Experiment 1. In contrast, for Experiment 2, we tightened the control for implicit recalibration at the non-critical targets by: 1) delivering delayed endpoint feedback for all non-critical targets while keeping the online continuous feedback for the critical target, which is known to suppress implicit recalibration, and 2) we widened the spatial gap between the critical and non-critical targets, to minimize spillover 2) As a result, the implicit recalibration was diminished at the non-critical targets, contributing to the shrinkage of the generalization curve around the critical target. 

      One other issue that we would like to make clear is that Experiment 1 was not a straight replication of a previous study, at least to our knowledge. We believe that the reviewer is referring to our previous study (McDougle and Taylor 2019), which found broader generalization for algorithmic strategies (Experiment 4). However, in that study cursor feedback was always delayed. As such, the observed broader generalization was most likely due to the strategy itself and not implicit recalibration. 

      Together, these findings led the authors to conclude that implicit recalibration is relatively insensitive to the type of explicit strategy employed and is primarily shaped by the statistics of the movement plans on which learning occurs.

      The experimental design using Critical and Non-Critical targets is particularly interesting and represents a creative approach to manipulating strategy use. Reaction times were generally longer in the algorithmic group, even at the Critical target, suggesting that the manipulation was at least partially successful in biasing participants toward algorithmic versus retrieval-based strategies. The results that the implicit recalibration is independent of the explicit strategy (how you aim) but depends on the aiming point by the explicit strategies (where you aim) are basically reasonable.

      We are glad to know that our primary finding and conclusion appears reasonable. While we acknowledge that the finding doesn’t appear to be particularly exciting at face value, it does speak to larger questions regarding the independence of different learning systems and how just statistical or surface-level differences in training can result in relatively large differences in apparent behavior that could be easily misinterpreted as the result of system interactions.

      I would like the authors to clarify two points.

      First, how reasonable is it to infer the use of distinct explicit strategies primarily from reaction time differences? While longer reaction times in the algorithmic group are consistent with greater computational demands, it remains unclear whether the longer reaction times observed at the Critical target necessarily reflect different strategy implementations at that location. In particular, could the increased cognitive demands associated with the Non-Critical targets in the algorithmic condition have carried over to the Critical target, thereby prolonging reaction times without implying qualitatively different strategies at the Critical target itself?

      This is a fair concern, as RT is an indirect marker of strategy use and, by itself, cannot establish that participants used different strategies at the Critical target. Prior work, however, provided guidance for the design of our experimental manipulations. Algorithmic strategies, in which an aiming solution is computed online, are associated with longer RTs, whereas retrieval of a previously cached stimulus–response association produces substantially shorter RTs (McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Moreover, caching becomes increasingly difficult as the number of target-specific solutions increases, particularly beyond approximately four targets (Velazquez-Vargas and Taylor, 2024; Bejjanki and Taylor 2026). Our manipulation was designed around these findings: participants in the Algorithmic condition learned the 45° rotation across 10 targets, whereas participants in the Retrieval condition repeatedly encountered the 45° rotation only at the Critical target. As expected with this experimental design, RTs at the Critical target were significantly longer in the Algorithmic condition across all three experiments.

      We agree with the reviewer, however, that this RT difference could in principle reflect a more general carryover of cognitive demands from the Non-Critical targets rather than online computation at the Critical target itself. We can address this possibility more directly by asking whether RT at the Critical target exhibits the parametric signature expected of an algorithmic process. A defining feature of mental rotation is that RT scales with the magnitude of the computed aiming solution (Georgopoulos and Massey, 1987; Bhat and Sanes, 1998; McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Although rotation magnitude was fixed in the present experiments, participants’ actual reach angles varied naturally from trial to trial. Indeed, McDougle and Taylor (2019) originally demonstrated this relationship using actual reach angle rather than imposed rotation magnitude. We can therefore test whether trial-by-trial RT covaries with reach angle at the Critical target in the Algorithmic condition but not in the Retrieval condition. Such a relationship would be difficult to explain as a nonspecific carryover of cognitive load and would instead provide direct evidence that preparation time at the Critical target reflects the computation of the aiming solution.

      We also observe a second, independent difference at the Critical target: reach angles are consistently more variable in the Algorithmic condition than in the Retrieval condition across all three experiments (Figure S6). This pattern is consistent with repeated online computation producing variability in the selected aiming solution, whereas retrieval of a cached stimulus–response association produces a more stable response. Importantly, a general carryover account based solely on increased cognitive demands does not readily explain why movements to the Critical target should also be systematically more variable. Nor is this pattern easily explained by a speed–accuracy tradeoff: the Algorithmic group had more preparation time yet nevertheless exhibited greater variability. Consistent with this notion, Velázquez-Vargas and Taylor (2024) found that retrieving cached solutions produced less variable and more precise movements than movements that are not cached in the memory trace. 

      Third, we can conduct additional analysis to compare RT variability between algorithmic and retrieval groups at the critical target location. According to Logan instance theory (Logan, 1988), the retrieval of cached stimulus-response associations produces a stable RT profile, in contrast, trial-by-trial algorithmic computation can result in more variable trial-by-trial RT differences. 

      While we agree that RT differences alone should not be taken as definitive evidence of distinct strategies, our specific experimental design, the longer and more variable RTs at the Critical target, the greater trial-by-trial variability at that same target, and, if confirmed, a parametric relationship between RT and reach angle provide converging evidence that participants in the two conditions relied on different strategy implementations when preparing movements to the Critical target.

      Second, the interpretation of Experiment 3 is not entirely clear to me. The manuscript argues that the algorithmic group continued to exhibit greater reaching variability than the retrieval group. If this variability indeed reflects greater variability in movement plans, one might expect a broader implicit generalization function in the algorithmic group. However, the generalization widths were comparable between groups. Could this result instead suggest that the implicit recalibration process itself generalized more narrowly in the algorithmic group, thereby offsetting the broader distribution of movement plans? More generally, I would appreciate further clarification regarding the relationship between reaching variability, movement-plan variability, and the resulting width of the implicit generalization function.

      We appreciate the reviewer’s thoughtful comment and agree that this is an important distinction. First, we would like to clarify the relationship among reaching variability, movement-plan variability, and the width of the implicit generalization function. Previous work has shown that when implicit recalibration occurs at a particular target location without an explicit aiming strategy, its generalization across the workspace can be described by a Gaussian-shaped function centered near the trained target location (Morehead et al., 2017). When an explicit aiming strategy is involved, however, implicit recalibration is centered closer to the planned aiming location rather than the visual target itself (McDougle et al., 2017). Thus, implicit recalibration is greatest near the direction in which the movement is planned, and a broader spatial distribution of movement plans can, in principle, produce a broader aggregate implicit generalization function. In the present study, we therefore use trial-to-trial variability in endpoint hand angle as a behavioral proxy for variability in planned movement direction, while recognizing that endpoint variability may also contain contributions from execution-related noise.

      This framework motivated the progression from Experiments 1 to 3. In Experiment 1, participants in the algorithmic condition exhibited substantially greater reaching variability, consistent with the idea that they sampled a wider range of movement plans across trials. Because error feedback associated with these different movement plans can induce implicit recalibration around each planned direction, greater variability in strategy use could contribute to the broader implicit generalization observed in the algorithmic group. In Experiment 2, we imposed stricter controls on spillover from noncritical targets, which reduced the overall breadth of generalization; nevertheless, model fits still suggested a modestly broader implicit generalization function in the algorithmic group, consistent with the remaining difference in reaching variability.

      Experiment 3 was designed to further reduce the direct influence of strategic variability on the induction of implicit recalibration by using a modified error-clamp paradigm. Error-clamp feedback has been shown to elicit implicit recalibration independently of task success and the participant’s explicit strategy (Morehead et al., 2017). We therefore used error-clamp feedback as an incidental signal to induce implicit recalibration while participants implemented either algorithmic or retrieval-based strategies. Importantly, however, Experiment 3 did not completely eliminate between-group differences in reaching variability: the algorithmic group continued to show greater variability around the critical 45-degree location than the retrieval group. We agree with the reviewer that, in principle, comparable generalization widths could arise if this broader distribution of movement plans were offset by a narrower local generalization of implicit recalibration in the algorithmic group. 

      However, not all variability is equivalent—or well described by a Gaussian distribution. Depending on the direction of the skew, variability in aiming can produce different effects on the implicit recalibration function, appearing as either broader generalization or greater amplitude. These effects are difficult to appreciate in Figures 3 and 5 for Experiments 2 and 3, respectively. We therefore sought to illustrate this more clearly in Figure 6, which shows how differences in the underlying aim distributions can shape the resulting implicit recalibration function in directionally complex ways. What complicates matters further is that implicit recalibration can asymptote (Morehead et al., 2017; Kim et al., 2018; Wilterson & Taylor, 2021). As a result, plan-based generalization can distort the implicit recalibration function in different ways depending on which side of the aim the error falls. The block-by-block analysis, suggested by reviewer 2, may shed light on this issue because we can get a sense if implicit recalibration has reached asymptote. 

      At a minimum, in the revised manuscript, we work to make clearer that subtle changes in the reach distribution may have a corresponding impact on the shape of implicit recalibration’s generalization function.  

      Reviewer #2 (Public review):

      This study addresses an important question in motor learning: whether algorithmic versus retrieval-based explicit strategies differentially shape implicit recalibration. The progressive experimental logic across three experiments is commendable, and the plan-based generalization account is a plausible and interesting interpretation. However, several methodological concerns limit the strength of the conclusions. I recommend the authors temper their claims accordingly, in the results/discussion section.

      Concerns

      (1) The retrieval group received 5 pre-exposure trials before main training began, which the algorithmic group did not. Faster RTs in the retrieval group could therefore reflect task familiarity from extra practice rather than efficient memory retrieval per se. I might have missed this, but I did not see performance data from these pre-exposure trials. The early training advantage in the retrieval group might be confounded with the 5 pre-exposure trials they received. Unless there is a direct comparison between the pre-exposure trials for the caching group and the first 5 trials of the algorithmic group, the claim that "storing and retrieving a memory from a short-term memory cache confers more rapid performance improvements than executing an algorithmic strategy" seems somewhat unwarranted.

      We appreciate the reviewer raising this potential confound. We agree that the five pre-exposure trials in the Retrieval condition introduce a small difference in initial task familiarity. However, this account makes a straightforward prediction: if the shorter RTs in the Retrieval condition simply reflect five additional trials of general task experience, then the RT difference should disappear once the Algorithmic group has received a comparable amount of practice.

      We can test this directly by comparing the five pre-exposure trials in the Retrieval condition with the first five trials of the Algorithmic condition. We will also compare these pre-exposure trials with a later five-trial window from the Algorithmic condition to determine whether additional practice substantially reduces Algorithmic RTs. Assuming the observed pattern is as expected, RTs in the Algorithmic condition remain substantially longer even after considerably more than five trials of practice. Thus, the group difference cannot be explained simply by the Retrieval group having five additional trials of task familiarity. This persistent RT difference, together with our prior work showing characteristic RT differences between algorithmic computation and retrieval of cached aiming solutions, supports our interpretation that the groups relied on different strategy implementations.

      We are less certain what the reviewer means by the “early training advantage.” If this refers to angular error, we agree that the Retrieval group shows somewhat better performance very early in training, but this difference is not a central focus of the present study and largely disappears by the second or third training block. We will clarify the text so that we do not overinterpret this transient difference.

      If instead the reviewer is referring to RT, then the matched-trial analysis directly addresses the concern. Five additional familiarization trials could plausibly produce a brief initial advantage, but such an effect should dissipate within a small number of subsequent trials. In contrast, the RT difference between the Algorithmic and Retrieval conditions remains robust throughout training. We therefore do not think that general task familiarity provides a sufficient explanation for the observed RT differences.

      The algorithmic group also visited the critical target approximately 40% of trials across 356 trials (about 140 trials?). McDougle & Taylor (2019) showed that 300 trials of practice with 2 targets is enough transition from algorithmic to caching strategies. It seems likely that the number of visits to the critical target here was sufficient for caching to develop in the algorithmic condition. This concern about caching in the algorithmic group has implications for the implicit recalibration measurements. As I understand it, the 7 exclusion blocks were distributed throughout training, and so, implicit recalibration was measured across both early and late practice. If caching emerged in the algorithmic group during late practice, then the generalization functions - averaged across all 7 exclusion blocks - conflate early algorithmic strategy and later caching. The broader generalization function observed in the algorithmic group may therefore be driven primarily by early exclusion blocks, while later exclusion blocks may increasingly resemble the retrieval group as caching develops. This is testable in the data: if generalization breadth in the algorithmic group narrows across the 7 exclusion blocks while remaining stable in the retrieval group, that would be consistent with a strategy transition occurring during training. The authors should either report exclusion block-by-block generalization functions separately for each group, or acknowledge that the averaged generalization functions may obscure a strategy transition in the algorithmic group.

      The reviewer raises an interesting possibility. In McDougle and Taylor (2019), however, the transition from algorithmic computation to retrieval occurred in a condition with only two targets in the task set. With repeated practice, participants needed to retain only two target-specific aiming solutions, making it feasible to replace online computation with retrieval of cached stimulus–response associations. By contrast, the Algorithmic condition in the present study contained 10 target locations. Our prior work suggests that caching becomes increasingly difficult once the number of target-specific solutions exceeds approximately four, at least over the timescale of several hundred trials (Velázquez-Vargas and Taylor, 2024; Bejjanki and Taylor, 2026). Thus, our task was designed to maintain pressure toward an algorithmic strategy throughout training.

      The reviewer nevertheless raises a more specific possibility that is not ruled out simply by the size of the target set: participants might selectively cache the aiming solution for the frequently sampled Critical target while continuing to use an algorithmic strategy at the remaining targets. We think the existing behavioral data argue against such a clear transition.

      First, reaction times at the Critical target in the Algorithmic condition remained substantially longer than those in the Retrieval condition throughout training. If participants had cached the aiming solution for the Critical target, we would expect preparation times at that location to be the same as the Retrieval condition. However, the RTs for the Algorithmic and Retrieval conditions are significantly different in the last block of training. 

      Second, within the Algorithmic condition, reaction times at the Critical target remained similar to those at the Non-Critical targets. Selective caching of the Critical target predicts a different pattern: preparation should become faster at the Critical target than at the surrounding locations, where participants would still need to compute the appropriate aiming solution. However, we do not observe a significant difference between RTs at Critical and Non-critical targets for the Algorithmic conditions at the end of the training block. Taken together, these two observations suggest that the Critical target continued to be treated similarly to the other members of the 10-target set rather than becoming a privileged, cached stimulus–response association.

      Third, participants would have to single out the Critical target as being distinct. All targets had the same visual appearance, the Critical target was never presented on consecutive trials, and participants were not informed that it had a special role in the experiment. However, we acknowledge that its higher sampling frequency, its somewhat greater separation from neighboring targets, and the location of the subsequent exclusion trials could nevertheless have made it more salient. Thus, we cannot rule out selective caching solely from the task structure.

      For this reason, we agree that the reviewer’s proposed analysis provides a useful additional test. If the Critical-target strategy progressively transitioned from algorithmic computation to retrieval, one prediction is that the generalization function in the Algorithmic condition should become narrower across successive exclusion blocks and increasingly resemble that of the Retrieval condition. We will therefore attempt to estimate the width of the generalization functions as a function of the training block between the Algorithmic and Retrieval conditions.

      There is, however, an important limitation to interpreting block-by-block generalization functions in this experiment. Implicit recalibration is both plan-based and temporally labile. Generalization is centered around the planned aiming direction (McDougle et al., 2017), and recent work indicates that implicit adaptation can decay over relatively short intervals (Zhou et al 2017; Hadjiosif et al 2023). Consequently, the first trial of an exclusion block provides the cleanest sample of the current state of implicit recalibration. Across later trials in the block, the measured response can be influenced both by temporal decay and by where the sampled target falls relative to the participant’s current aiming direction.

      To minimize systematic sampling bias, the starting exclusion target was randomized across participants. This means that these effects should average out at the group level, but individual exclusion blocks do not provide equally precise samples of the entire generalization function. A fully balanced estimate of every position within each exclusion block would require substantially more participants than were included in the present experiments. We will therefore present the blockwise analysis while interpreting changes in the estimated breadth cautiously.

      (2) The error-clamp paradigm in Experiment 3 introduces two problems. First, it breaks the relationship between planned movement direction and feedback of movement direction, likely reducing the sense of agency over movement feedback (indeed, typical error clamp study instructions tell participants to ignore the movement feedback). Reduced agency may itself suppress differences between algorithmic and caching conditions. First, if strategy type exerts its influence on implicit recalibration via the explicit plan - as the plan-based generalization account predicts - then severing the link between intended movement and feedback might close off the channel through which strategy could shape the implicit system, regardless of which strategy is used. Second, reduced agency could modify the explicit strategies themselves. For caching, the stimulus-response association might be reinforced by a consistent relationship between intended movement and observed outcome; the clamped feedback may make it more difficult to reinforce the cached response, weakening the stimulus-response association. For the algorithmic strategy, effortful mental rotation may depend on the perception that the computation meaningfully determines the outcome; as participants understand that clamped feedback does not depend on their behavior (although yes, the text-based "Excellent/Good Move feedback) does depend on their behavior, they may engage in somewhat less complete mental rotation. Both possibilities could contribute to convergence between groups in generalization. It is noted that the preserved RT difference between groups in Experiment 3 partially argues against a loss of effort under the algorithmic condition, but it does not rule out weakened formation of stimulation-response associations during caching.

      The reviewer raises an important point. By design, the error-clamp manipulation in Experiment 3 decouples the participant’s planned movement from the visual consequence of that movement. While this gives us precise control over the error driving implicit recalibration, it could reduce agency over the cursor and thereby alter the interaction between explicit strategy and implicit learning in ways that are difficult to rule out completely. In particular, as the reviewer notes, reduced agency could potentially weaken either the influence of the explicit plan on implicit recalibration or the strategies themselves. Because Experiment 3 was intended to test for the absence of a strategy-dependent difference in implicit recalibration, we acknowledge that higher-order interactions of this kind represent an inherent limitation of our study if Experiment 3 is taken in isolation. 

      There are nevertheless several observations that make us think that reduced agency is unlikely to provide the primary explanation for the convergence between groups. First, the progression across Experiments 1–3 is informative. In Experiment 1, where participants retained normal control over the cursor, the broader generalization function in the Algorithmic condition closely mirrored the broader distribution of reach directions. This relationship suggests that the apparent difference in implicit generalization could arise from differences in where participants planned their movements rather than from a direct effect of strategy type on the implicit system. In Experiment 2, we sought to reduce the difference in the distribution of planned movements while preserving normal action–outcome contingencies and, importantly, the generalization functions became correspondingly more similar. Experiment 3 then controlled the error signal itself and again produced similar generalization across strategy conditions. Taken together, this progression favors the interpretation that strategy affects the measured generalization function indirectly, through differences in the distribution of movement plans, rather than directly altering the underlying implicit recalibration process.

      We nevertheless agree that these experiments cannot exclude all possible interactions between explicit and implicit learning systems. Indeed, whether these systems interact directly has been an important and persistent question in the sensorimotor adaptation literature. Several studies have reported evidence consistent with direct interactions (e.g., Albert et al., 2022; Maresch and Donchin 2021; t’Hart and Henriques 2024), whereas our own work has generally pointed toward indirect interactions mediated by factors such as movement planning and the current state of implicit adaptation (Taylor et al 2010; Taylor and Ivry 2011; McDougle et al 2017). Indeed, our recent study was designed specifically to distinguish these possibilities under tighter experimental control (Chen and Taylor, 2026), yet we found that the interaction between explicit and implicit processes is more complex than a simple independent-versus-interacting dichotomy. Going forward, we think it is more cautious to first rule out low-level statistical or distributional differences that could account for apparent effects before invoking higher-order interactions between learning systems.

      Finally, one motivation for the present study was that much of the literature on implicit generalization trains participants at a single target location before measuring generalization across the workspace. Under such conditions, participants have ample opportunity to retrieve a stable target-specific aiming solution. If algorithmic and retrieval strategies fundamentally alter implicit generalization, then many existing estimates of generalization may characterize implicit learning under retrieval-like conditions rather than providing a strategy-independent property of the implicit system. Across the present experiments, we find little evidence for such a fundamental difference once the distribution of movement plans and the experienced error are better controlled. We therefore think the most parsimonious interpretation of the current results is that algorithmic and retrieval strategies primarily influence implicit generalization indirectly through how movements are planned. 

      We plan to revise the manuscript to acknowledge that Experiment 3 cannot completely rule out higher-order effects associated with reduced agency under error-clamp feedback. We will also provide additional validation that participants were implementing distinct strategies, beyond the group-level RT differences, by testing whether RT in the Algorithmic condition scales with the instructed rotation magnitude, and whether RT variability shows group-level difference (Logan, 1988). If present, this relationship would provide stronger evidence that participants continued to engage the intended strategy under the clamp. We agree, however, that confirming distinct strategy use would not by itself rule out the possibility that reduced agency altered how those strategies interacted with implicit recalibration.

      Reviewer #3 (Public review):

      Summary:

      This manuscript asks whether two forms of explicit strategy use in visuomotor adaptation, i.e., algorithmic mental rotation and retrieval of a cached aiming solution, differentially influence implicit recalibration. The question is relevant because much prior work treats explicit strategy as a unitary process, whereas the algorithmic/retrieval distinction is theoretically meaningful and grounded in cognitive theory. Across three experiments, the authors report that algorithmic strategy conditions initially produced broader fitted implicit generalization functions than retrieval conditions, but that this difference was reduced or eliminated when reach variability and sensory prediction errors were more tightly controlled.

      Strengths:

      The paper is clearly written, theoretically well-motivated, and employs a commendably transparent and progressive experimental logic. The three-experiment structure, in which confounds are systematically identified and addressed, represents a strong model of cumulative experimental design (I will certainly use it in teaching courses on experimental methods):

      Experiment 1 establishes an apparent difference in implicit generalization breadth. Experiment 2 attempts to reduce error spillover from Non-Critical targets by increasing angular separation and using delayed endpoint feedback. Experiment 3 uses an error-clamp design to decouple variable reaching from error feedback. This sequence is appropriate for testing whether the initial difference reflects a strategy-dependent change in implicit recalibration or instead follows from the distribution of movement plans and error exposure. The authors also provide reaction-time and performance data that are broadly consistent with the intended distinction between algorithmic and retrieval-like task performance.

      We thank the reviewer for this thoughtful and constructive assessment of the manuscript. We especially appreciate their recognition of the progressive experimental logic across the three experiments and of the broader theoretical motivation for distinguishing algorithmic and retrieval-based strategies. We are also grateful for the reviewer’s comments on the clarity and transparency of the work.

      Weaknesses:

      The evidence does not support the strongest claims made in the manuscript, namely that algorithmic and retrieval strategies generally do not reshape implicit recalibration.

      In general, I am skeptical of the authors' interpretation of null results. Several central conclusions depend on non-significant group differences, especially in Experiment 3. Non-significant tests are repeatedly treated as evidence that groups are equivalent or that confounds are absent (e.g., implicit recalibration magnitude (Algorithmic: 11.43 {plus minus} 6.43{degree sign}; Retrieval: 15.49 {plus minus} 8.99{degree sign}; t(38) = −1.65, p = .11), adaptation level before Exclusion probes (F(1,256) = 3.04, p = .08) and Exclusion RT differences (F(1,266) = 3.15, p = .08), whereas a modest model-dependent breadth effect (bootstrap p = .02) is treated as meaningful (for more on the model-dependent breadth effect, see below).

      Without confidence intervals, equivalence tests, or Bayesian analyses, I think that the authors' interpretations comprise an inferential gap. A failure to find a significant difference is not equivalent to evidence of equivalence, particularly given that the implicit recalibration signal gets progressively attenuated across experiments (Experiment 1: ~16-17{degree sign}; Experiment 2: ~11-15{degree sign}; Experiment 3: ~7-8{degree sign}). With a substantially diminished signal in Experiment 3, the null result could partly reflect reduced statistical sensitivity rather than true equivalence.

      We agree with the reviewer that our original interpretation of several non-significant effects was too strong, especially without providing some form of equivalence test. Our central hypothesis predicts little or no difference between algorithmic and retrieval-based strategies under conditions in which movement plans and error exposure are controlled, and we therefore face the inherent difficulty of drawing conclusions from an expected null effect. As the reviewer notes, a non-significant conventional hypothesis test does not by itself provide evidence that two conditions are equivalent.

      We therefore plan to supplement the existing analyses with quantitative assessments of the strength of evidence for the null/equivalence, using Bayesian factor analyses to confirm whether two conditions are equivalent. These analyses will allow us to distinguish between effects that are sufficiently small to support our theoretical interpretation and effects for which the data are simply inconclusive. We will also revise the manuscript throughout to avoid treating p > .05 as evidence of equivalence in the absence of such supporting analyses.

      We also now appreciate that the magnitude of implicit recalibration decreases progressively across experiments. This reduction could diminish our sensitivity to differences between the Algorithmic and Retrieval conditions and therefore represents an important qualification on the null result. Because the experiments used similar trial structures and were conducted with the same experimental equipment, the source of this reduction is not immediately clear. The block-by-block analysis suggested by Reviewer 2 may provide useful insight into how implicit recalibration evolves over the course of training and whether this attenuation emerges gradually within experiments.

      My main technical concern is the analysis of generalization breadth already alluded to. The central claims rely on group-level Gaussian fits to only seven Exclusion probe locations spanning −45{degree sign} to +45{degree sign} around the Critical target. In several cases, the fitted centers and widths are poorly constrained by the sampled range. For example, in Experiment 2 the algorithmic group's fitted center is shifted to approximately 29{degree sign}, meaning that the probe range samples the function asymmetrically relative to its own peak. In Experiment 3, fitted centers are near or outside the sampled range, while estimated widths are very broad. Under these conditions, the width parameter may partly reflect extrapolation or parameter trade-offs between center, amplitude, and width rather than a genuine difference in generalization breadth.

      Based on prior work characterizing implicit generalization in relative isolation from explicit strategy (Morehead et al., 2017; Poh and Taylor, 2019), we expected a relatively narrow generalization function, with a full width at half maximum of approximately 30°. We therefore expected probes spanning −45° to +45° around the Critical target to capture most of the function. At the same time, prior work on plan-based generalization predicts that the function should shift toward the participant’s aiming direction (Day et al., 2016; McDougle et al., 2017; Chen and Taylor, 2026), which complicates the choice of where to center the probes. Expanding the range and density of probe locations is also not cost-free, because additional exclusion trials increase temporal decay (Hajiosif et al., 2023) and begin to overlap with trained locations.

      We nevertheless agree with the reviewer that, when the fitted center approaches the edge of the sampled range, estimates of Gaussian width can become poorly constrained and may partly reflect parameter trade-offs or extrapolation beyond the observed data. We therefore plan to test whether the group differences persist when the Gaussian fits are constrained so that their centers fall within the sampled range. We will also examine complementary nonparametric measures of generalization breadth, such as the area between the group generalization curves across the sampled probe locations. Convergence across these approaches would provide stronger evidence that the reported differences reflect the observed shape of the generalization functions rather than instability in the Gaussian parameter estimates. If the results are not robust across approaches, we will revise the manuscript to qualify the interpretation of the fitted width estimates accordingly.

      Lastly, I think that the authors' use of an error-clamp paradigm is, from an experimental point of view, quite elegant. By controlling the sensory prediction error independently of reach direction, they can isolate implicit recalibration from the confounds identified in Experiments 1 and 2. However, I see a fundamental problem or question concerning construct validity here: In Experiments 1 and 2, the algorithmic strategy was operationalized as participants computing a counterrotated aiming direction in response to a visible cursor rotation. This is a naturalistic context where mental rotation is both required and meaningfully connected to task success. In Experiment 3, however, there is no visuomotor rotation to compensate for. The error-clamp renders the cursor feedback task-irrelevant. Instead, participants are instructed via text commands (e.g., "move towards 45{degree sign}") to reach invisible locations, rendering the "algorithmic strategy" in this context essentially an instructed spatial navigation toward arbitrary angular locations, not genuine visuomotor mental rotation driven by an error signal.

      To put it differently, are we sure that the cognitive process engaged by the algorithmic group in Experiment 3 is the same as the algorithmic mental rotation strategy in Experiments 1 and 2? If not, then the null result in Experiment 3 may not speak to the original question about how algorithmic strategies interact with implicit recalibration after all. Instead, it may reflect the absence of a genuine strategy manipulation.

      This concern is closely related to that raised by Reviewer 2. We are fairly confident that participants in Experiment 3 were nevertheless engaging in the intended strategy manipulation. Participants in the Algorithmic condition showed substantially longer RTs than those in the Retrieval condition, and they were able to accurately generate the instructed angular reach directions across trials. The two conditions also differed in the variability of both RT and reach direction, consistent with online computation of an aiming solution in the Algorithmic condition and retrieval of a more stable cached response in the Retrieval condition. We can provide an additional validation by examining the relationship between RT and reach angle and comparing RT variability across two conditions. If RT scales parametrically with instructed reach angle in the Algorithmic condition but not in the Retrieval condition, this would provide stronger evidence that participants were engaging an online mental-rotation-like computation rather than simply following arbitrary spatial instructions. Moreover, RT yielded by the retrieval strategy would tend to be less variable than the algorithmic strategy.

      We agree, however, that this does not fully address the reviewer’s broader concern. Experiment 3 necessarily changed the context in which the strategy was implemented. In Experiments 1 and 2, mental rotation was used to counteract a visuomotor perturbation and was therefore directly tied to successful control of the cursor. In Experiment 3, the error clamp removed this instrumental relationship: participants still had to compute and execute different angular reach directions, but those computations no longer determined the visual cursor outcome. In that sense, the algorithmic process in Experiment 3 was less naturally embedded in the task and could reasonably be viewed as a somewhat different instantiation of the task.

      We therefore acknowledge that Experiment 3 cannot establish with certainty that the same higher-order cognitive process was engaged in exactly the same way as in Experiments 1 and 2, nor can it rule out the possibility that this change in task relevance altered how explicit strategy interacted with implicit recalibration. At the same time, when considered together with Experiments 1 and 2, we think the overall pattern remains informative. The apparent strategy-dependent difference in generalization was largest when movement plans and error exposure differed most, became smaller when these factors were better controlled while normal action–outcome contingencies were preserved, and was eliminated when the error signal itself was experimentally controlled. This progression is more consistent with an indirect influence of strategy through differences in movement planning and error exposure than with a robust direct effect of strategy type on implicit recalibration.

      Nonetheless, we agree that Experiment 3 should not be interpreted as a definitive test of whether algorithmic strategy, in its more relevant visuomotor adaptation context, can lead to different interactions with implicit recalibration compared to a retrieval strategy. We plan to revise the manuscript to make this limitation explicit.  

      To their credit, the authors report a compelling RT dissociation that mirrors Experiments 1 and 2: The algorithmic group shows slower RT, which is decreasing over training (0.98s → 0.76s), whereas the retrieval group exhibits faster, stable RT (0.52s → 0.45s). While this pattern is consistent with genuine strategy differences persisting in Experiment 3, it could also reflect the greater spatial precision demands of reaching to invisible targets from text instructions, rather than genuine mental rotation per se. Reaching to an invisible location defined by a verbal angular label is inherently more demanding than reaching to a visible target, regardless of strategy type, and this demand is asymmetrically present in the two groups, since Non-Critical targets are invisible for the algorithmic group but visible for the retrieval group.

      Thus, from my point of view, experiment 3 should not be used as definitive evidence that algorithmic and retrieval strategies during standard visuomotor adaptation cannot differentially influence implicit recalibration.

      We agree that the RT difference in Experiment 3, by itself, cannot rule out the possibility that the Algorithmic condition imposed greater spatial precision demands because participants were reaching to invisible locations specified by angular instructions. We can, however, test for a more diagnostic signature of algorithmic computation by examining whether RT scales parametrically with the instructed reach angle. A general cost associated with reaching to invisible targets could increase overall RT, but it would not necessarily predict the characteristic increase in preparation time with the magnitude of the required angular transformation.

      We will therefore examine the relationship between RT and instructed reach angle in Experiment 3. If RT increases systematically with angular displacement in the Algorithmic condition, this would provide additional evidence that the longer RTs reflect online computation of the instructed aiming direction rather than simply the greater difficulty of reaching invisible targets.

      We can also compare this RT–angle relationship across experiments. If participants are engaging the same underlying algorithmic computation in Experiments 1–3, we would expect the slope relating RT to angular displacement to be similar across experiments, even if the overall intercept differs because of differences in task structure and spatial demands. A comparable slope would therefore provide converging evidence that the same computational process was engaged despite the altered task context in Experiment 3.

      We acknowledge, however, that similarity of the slopes would itself require yet another inference from a null difference and should therefore be interpreted cautiously. As with the generalization functions, we will use a Bayes factor analysis to quantify the strength of evidence for the null.

      Overall, the manuscript addresses a meaningful question and the multi-experiment structure is useful. The evidence is incomplete for the broad claim that implicit recalibration is insensitive to strategy type. The study would make a clearer contribution if the authors narrowed the claims, strengthened the generalization analyses, and treated null effects with appropriate inferential tools.

      Based on the reviewers’ comments and the additional analyses they have suggested, we think we will be able to place our conclusions on a firmer empirical footing while also tightening and narrowing them. In the revised manuscript, we will strengthen the generalization analyses, use more appropriate inferential tools for interpreting null effects, and temper our broader claims about the insensitivity of implicit recalibration to strategy type. We will also more explicitly acknowledge the limitations of the present experiments, especially Experiment 3.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary and Strengths:

      Shin et al deepen our understanding of high-frequency oscillations in the frontal cortex during REM in a manner that sheds important light on the roles of these events. In particular, they reveal that cortical HFOs are modulated by theta oscillations, occur in chains and recruit cortical neuronal activation patterns in a manner that is distinct from other high-frequency events during non-REM or in the hippocampus. They also show that these events occur during increased oscillatory cross-talk between hippocampus and cortex and may protect cortical neurons from downregulation of firing during sleep. Overall, this is important work with several novel observations pointing towards an important role for these events that will become increasingly understood over time.

      I also wanted to comment that 2D is a beautiful illustration of separate and essentially exclusive communication channels used during HF events in NREM vs REM. They almost perfectly complement each other's frequencies.

      Weaknesses:

      I have only one major scientific critique: I believe we need to see quantification of how phasic REM theta waves with versus without HFOs differ. What do REM HFOs add to the "normal" theta oscillation? Without this comparison, it is more difficult to interpret the meaning of these events. Given that HFO chains have IEIs around the time of a theta cycle duration, are the repeating spiking activities stronger during HFO repeats than during adjacent theta waves without HFOs?

      Here, we provide additional analyses to demonstrate that the phasic, theta-modulated PFC activity that we observe during HFOs is specifically tied to the occurrence of HFOs and not a strong phenomenon during non-HFO-associated theta periods. In Figure S5 M (middle and right), we find that aligning PFC multiunit activity to theta periods in phasic REM but temporally distant from HFOs does not elicit the same degree of theta-modulated activity as aligning to HFOs (as in Figure 2A and Figure S5 M, left).

      Additionally, we provide analyses of the theta periods adjacent to HFOs (at different temporal thresholds) and demonstrate that this theta-modulated spiking activity is largely absent (Figure S5; compare to Figure 2A). Unlike Figure S5, this analysis was not restricted to putative phasic REM.

      We have now added Supplementary Figures S5L-N to the revised manuscript.

      What percentage of theta waves contain HFOs, and what is the firing rate during those theta waves with vs without HFOs? Is there differential firing rate modulation? The authors may even consider that all REM-HFO-specific quantifications should be shown as differential from phasic theta cycles without HFOs.

      Although theta oscillations are continuously expressed during REM sleep, HFOs occur only intermittently, such that only a small subset of theta cycles contain HFOs. Across all animals and epochs included for analysis, we found that ~7.4% of theta cycles contain HFOs. We present an epoch-level quantification of this in Author response image 1, where the proportions were calculated across all, tonic, and phasic theta cycles. As expected, a higher proportion of putative phasic theta cycles contain HFOs.

      Regarding differential firing rate modulation, we refer reviewer to normalized MUA plots in the manuscript (Figures 2A and 4F). We would like the emphasize that what we show is PFC multiunit activity that is normalized by the mean population firing rate during REM sleep. Thus, these figures, specifically the chain HFO aligned figure, indicate that there are peaks in activity around baseline level in the background of an overall decrease in population activity relative to baseline (i.e. the troughs surrounding the peaks have lower activity compared to baseline, as in Figure 9C). While this suggests that there is an overall decrease in firing rates during HFOs as compared to baseline theta periods without HFOs, this simply provides a qualitative account of this difference. In Figure S5N, we present firing rate comparisons during HFO chains versus theta periods during putative phasic REM bouts at least 4 theta cycles away from HFOs. We find that HFO chain-associated PFC neuron firing rates are lower compared to non-HFO theta periods, supporting our finding of activity suppression during HFOs.

      Author response image 1.

      Proportion of theta cycles with HFOs. (A) Proportion of theta cycles with HFOs across all, tonic, and phasic cycles (***p = 4.90e-05, rank sum test).

      Lastly, we appreciate the reviewer's suggestion that REM HFO quantifications could be framed relative to phasic theta cycles without HFOs. We agree that such comparisons are informative and ensure that the results we present are specific to periods with detected HFOs. In response, we have added additional analyses of theta periods outside of HFOs in Figure S5. Furthermore, while we found that a larger proportion of chain HFOs occurred during bouts of putative phasic REM compared to isolated events (Figure S5 and Figure S3C), most of the analyses that we performed were on events pooled across putative tonic and phasic, since putative phasic REM is relatively scarce (<10%).

      We also refer the Reviewer to our response to Reviewer 2’s major comment #1 below where we reiterate several of our findings that demonstrate the HFO-specificity of the reported dynamics, as well as our extended response to Reviewer 3’s Public Review comment #1 where we show that the dynamics associated with REM HFOs are absent during HFOs detected during awake behavior (comparable theta state) on the W-Track (Figure S11). We hope that the additional control analyses we present as well as our expanded explanations adequately address the Reviewer’s concerns.

      As a non-scientific comment on the manuscript itself: unfortunately, the paper is difficult to read and understand at times, requiring great effort by the reader. This is to an extent that communication is hindered. The paper is dense with changing methods, often from panel to panel. Unfortunately, the panel quantifications are not explained in the results section in a manner that readers can understand without going to read the methods, often for each individual panel. These measures should be explained in a way that lets readers understand the conclusions of each panel and what gross calculations were used to reach those. Instead, too much jargon is used rather than clear descriptions of the overall calculations being done for each panel.

      We have now split and updated the figures in a more logical progression of ideas:

      Figure 1: Prefrontal cortical HFOs in REM sleep using spectral analyses.

      Figure 2: Characteristic spiking modulation in PFC during REM HFOs.

      Figure 3: HFO and gamma distinction in theta cycles, and PFC-CA1 coherence (including chain/ isolated HFOs in phasic and tonic REM stages).

      Figure 4: Differential modulation of PFC spiking activity during REM PFC HFOs vs. NREM PFC ripples.

      Figure 5: Characteristic PFC population activity during REM HFO chains.

      Figure 6: Comparison of PFC reactivation during REM PFC HFOs vs. NREM PFC ripples.

      Figure 7: Differential engagement and excitability modulation of CA1 neurons by REM HFOs.

      Figure 8: REM theta phase shifting CA1 neurons preferentially engaged by REM HFOs.

      Figure 9: (Model) Network model with ACh reproduces spiking modulation during REM PFC HFOs vs NREM PFC ripples.

      Figure 10: (Model) Model reproduces restricted REM coactivity vs. widespread NREM coactivity.

      We have also rearranged the figures to parallel the main figures and results section.

      The authors mention in the discussion section that they see increased functional connectivity between mPFC and CA1, but most data suggesting this seems to be based on LFP rather than spiking. Functional connectivity is best defined by spiking-spiking relationships. And these authors have spiking data. So I believe either the descriptive language should be pulled back to something like "oscillatory coupling" or more analyses should be dedicated to showing spike-spike coordination across regions.

      We have updated the manuscript accordingly. Specifically, we have modified the text on Page 13, Lines 23-24:

      “These chains are associated with increased measures of oscillatory coupling between PFC and CA1…”

      Reviewer #1 (Recommendations for the authors):

      (1) Please ensure that analytical methods are presented in the same order in the methods section as they are in the results section - panel by panel. That said, the methods section is well-written and presented.

      We apologize for any confusion this may have caused. We have now reorganized the methods section to ensure they presented in the same order as the panels in the main figures.

      (2) Please specify whether the recordings/behaviors occur during the animal's light circadian phase.

      We have added a statement in the methods under the “Behavior” section on Page 34, Lines 9-12 indicating that the experiments took place during the light phase:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Please specify how many tetrodes are in mPFC and how many in CA1?

      We have added this clarification on Page 33, Lines 27-28:

      “Tetrodes were split equally between PFC and CA1 (15, 16, or 32 tetrodes in each region).”

      (4) All mentions of "coherence" should have frequency bands specified. "Theta coherence", for example.

      We have now added this information to all relevant mentions of “coherence”.

      (5) The intuitive logic of the phase slope index (2K) should be briefly explained for maybe half a sentence in the results section. The intuition should be explained better in the methods section devoted to it.

      We have added more clarification on the phase slope index method, both in the legend of Figure 3 on Page 21, Lines 22-23:

      “Phase slope index (PSI), which is a measure of phase lag consistency across different frequencies…” and in the methods section under “Phase slope index” on Page 39, Lines 21-26:

      “In practice, PSI is used to assess the consistency of phase lag relationships between two signals across different frequency bands and is a measure that is weighted by oscillatory coherence. We opted to use PSI to estimate the directional flow of information instead of other methods, such as Granger causality, since it has been demonstrated that PSI is less prone to false positives.”

      (6) "Cofiring" and "coactivity" should be defined clearly as measures - preferably in the results section if possible. They sound similar and are somewhat jargon-y without self-explanatory meaning (or difference from each other). How should readers understand and interpret them?

      We apologize for the confusion regarding these two terms, which are both used throughout the manuscript. Here, we specifically used to term “cofiring” to specify the explicit quantification of coincident activity between pairs of neurons (e.g. Figure 4E, left) as described in the Methods section under “Ripple and HFO co-firing”. There were instances where “coactivity” was used to refer to this quantification, and they have been changed to “cofiring”. We have now added a statement to the manuscript to clarify that “cofiring” is a quantification of coincident activity between neurons on Page 7, Lines 34-35:

      “Overall cortical co-firing, which is a measure of coincident activity between neuron pairs during discrete events”

      Additionally, the term “coactivity” in the manuscript is used when describing neuronal activity in the model or when there are mentions of coincident activity other than the explicit quantification described above.

      (7) The temporal threshold for "cofiring" should be stated in the results to enable interpretation of the results.

      We apologize for the lack of clarity regarding the cofiring metric. Here, the temporal threshold that we are imposing is determined by the ripple/HFO event times (see Methods section “LFP event detection”). If both neurons in a pair emit spikes within the defined window of an event, they are considered to have “cofired” (Cheng and Frank 2008, Singer and Frank 2009, Sosa, Joo et al. 2020).

      We have now clarified this in the results section on Page 7, Lines 34-35:

      “Overall cortical cofiring, which is a measure of coincident activity between neuron pairs during discreet events…”

      We have also added a clarifying statement in the Methods section under “Ripple and HFO cofiring” on Page 40, Lines 24-25:

      “Here, cofiring assesses coincident activity between neuron pairs within the start and end times of events, and thus, no explicit temporal threshold was implemented.”

      (8) Please clarify more systematically in which region the NREM ripples were detected. The natural assumption is the hippocampus, but at times it is mentioned that they are detected in the cortex. Are they cortical in all analyses? Readers could easily get confused about this and misinterpret "ripples". To clarify further, if these are always cortical events, I suggest renaming "ripples" in this text to "cortical NREM HFOs".

      We apologize for the confusion. For all mentions of “ripples” in the text, we are referring to PFC ripples specifically in NREM sleep. When hippocampal ripples are mentioned, we differentiate them by explicitly using “sharp-wave ripples” or “SWRs”. Our decision to use “ripples” for NREM sleep was based on previous studies that investigated NREM high-frequency events in cortex (Khodagholy, Gelinas et al. 2017, Helfrich, Lendner et al. 2019, Vaz, Inati et al. 2019, Aleman-Zapata, Morris et al. 2022, Ghosh, Yang et al. 2022, Shin and Jadhav 2024). Furthermore, we elected to use the “HFO” nomenclature for high-frequency events in REM sleep, since it has been used in previous studies to describe these events (Tort, Scheffer-Teixeira et al. 2013, Bueno-Junior, Ruckstuhl et al. 2023). Thus, to remain consistent with the literature, we decided to use these terms to describe these sleep-state-specific events throughout the manuscript:

      Hippocampal sharp-wave ripples (SWRs) in NREM sleep

      PFC ripples in NREM sleep

      PFC HFOs in REM sleep

      We have now added the following statement on Page 5, Lines 1-3:

      “However, to avoid ambiguity, and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout.”

      (9) The analysis performed for 4B is not explained clearly. It is somewhat better explained in the methods. I believe the reader should understand that each HFO is treated as an event, and spiking participation per unit was measured, and then the similarity of that pattern was assessed between HFOs. I also find the x-axis being quartiles makes understanding this graph particularly difficult.

      Why not label by raw lag and show a correlation plot rather than break down by quartiles? Alternatively, labeling the millisecond values of these quartiles on the x-axis labels may make comprehension much easier.

      We apologize for the lack of clarity regarding this method. We have added points to clarify the analytical procedure used for Figure 4B (Now Figure 5B) on Page 8, Lines 25-28:

      “We represented each HFO as a binary vector of PFC neurons active during the event, and computed the Pearson correlation between the vectors of every pair of consecutive HFOs” As well as in the Figure 5 legend on Page 24, Lines 7-10:

      “Here, the PFC spiking activity during each HFO was binarized across all neurons and the Pearson correlation coefficient was calculated between adjacent events as a measure of pattern similarity. Then the relationship between pattern similarity and IEI was reported.”

      In addition to the quartile labels, we have now added the average inter-event interval (IEI) of each quartile to the x-axis labels of Figure 5B to improve comprehension as well as a statement in the figure legend on Page 24, Line 11:

      “Below each quartile is the average IEI of that quartile in milliseconds.”

      (10) "Rank order analysis" should be again defined in the results in a manner that the reader can follow the point of the analysis and figure. For example, the goal is to assess the spike sequence across HFO events by looking at the regularity of spike timing rank for each neuron in each HFO event. Also, could this correlation be more simply calculated and presented as just a standard deviation around the mean of that cell's rank?

      We apologize for not including an explanation of this analysis in the results section. We have now added more detail about the rank order analysis to the Figure 5 legend on Page 24, Lines 18-21:

      “The mean rank-order correlation from the leave-one-out cross-validation procedure. Each event’s rank was correlated with the averaged rank across all other events. The average across all events compared to a distribution of means generated by jittering (n = 1000) spike times is shown.”

      We also provided a short description of the procedure in the results section on Page 8, Lines 32-36:

      “Furthermore, we examined whether the order in which individual PFC neurons fired during HFOs within a chain was preserved across chains. For each chain, we extracted each cell's first-spike rank order and compared it to a leave-one-out template constructed from the average normalized rank across all other chains.”

      Regarding the Reviewer’s second comment about the correlation, if we presented the result as the standard deviation around the mean of the cell’s rank, the result would be similar to the example rank order in Figure 5C, which we provided as a visualization of the rank-order template procedure we used. However, this would not necessarily demonstrate the consistency of population activity across HFO chains. To demonstrate consistency in activity across HFO chains, we used a leave-one-out cross-validated approach where each chain event was assessed separately. For each event, the neurons firing during that event was ranked and normalized 0-1. Then, the average rank of each neuron across all other events was calculated. Lastly, the correlation between the ranks of the left-out event and the average ranks of the template was taken to assess similarity in sequential activity. We opted to use this method since it has been demonstrated to be effective for evaluating similarities in sequential across events (Stark, Roux et al. 2015, Valero, Viney et al. 2021).

      (11) In Figure 5B and the related results section, I gather that "spatial" relates to place field location rather than anatomical/tetrode location of the neuron? This was not my original understanding and should be stated clearly.

      We apologize for the lack of clarity regarding this result. Yes, the term “spatial” refers to spatial rate map correlation between CA1 neuron pairs, specifically during the behavioral W-Track session prior to the sleep session where cofiring was assessed. We have updated the y-axis label of Figure 5B (Now Figure 7B) to include “rate map” and have updated the results section on Page 9, Lines 32-33:

      “…we observed a higher degree of spatial rate map correlation for high cofiring pairs…”

      (12) Figure 5E and its caption are almost totally unable to be explained since axes aren't explained well in either. It is only by inference from the results text that meaning can be assumed.

      We apologize for the lack of clarity regarding Figure 5E (Now Figure 7E). We have now updated the y-axis labels on both updated figures (Figures 7E and 8B) and have reworded and added more detail to the figure legend on Page 28, Lines 1-5:

      “High REM PFC HFO cofiring CA1 neurons exhibited a greater degree of suppression during NREM PFC ripples. For a description of modulation index, see Methods section Ripple/HFO aligned modulation. Here, since CA1 neurons exhibit a robust decrease in activity in response to NREM PFC ripples, we refer to the modulation as suppressive.”

      (13) Figure 5F is interesting, supporting the concept that HFOs "protect" neurons from downscaling. However, how can a "neuron" be cofiring? Would cofiring not be defined in a pairwise manner, and so each unit of the cofiring measure would be a pair of neurons? This question applies to other panels in this figure. Please clarify this.

      We apologize for the lack of clarity regarding the exact cofiring metric that we use. For Figures 7D-F and 8A, since we wanted to relate cofiring to changes in firing rate and modulation state during NREM PFC ripples, we calculated single cofiring values for each CA1 neurons by averaging across all pairings with PFC neurons. Thus, this metric gives us an estimate of the overall cofiring strength of each CA1 neuron. We have now added a new section in the Methods under “Calculation of a single cofiring metric and separation into populations of high and low cofiring neurons” on Page 44, Lines 34-43:

      “Since we wanted to relate the above cofiring metric to other measures, we needed to obtain a single-value cofiring metric for each neuron. To do this, we averaged the cofiring values across all pairings for a neuron (e.g. 1 CA1 neuron paired with all PFC neurons) and reported it as the cell’s cofiring. Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0 (High cofiring > 0; low cofiring < 0). Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions (High cofiring > mean; low cofiring < mean).”

      We have also clarified this in the Figure 7 legend on Page 27, Lines 25-28:

      “For comparisons between cofiring and other metrics (e.g. firing rate), a single cofiring value was calculated for each neuron by averaging the cofiring metric across all neuron pairings. Additionally, high and low cofiring CA1 neurons were split based on average cofiring values > 0 and < 0, respectively.”

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate high-frequency oscillations (HFOs) in the prefrontal cortex during REM sleep. They identify a specific pattern where these HFOs occur in "chains" that are phase-locked to theta oscillations, primarily during the "phasic" periods of REM. The study contrasts these events with isolated HFOs and NREM ripples, suggesting a unique role for these chains in coordinating activity between the prefrontal cortex and the hippocampus. Most notably, the authors report that a specific subset of hippocampal cells-those that co-fire with the prefrontal cortex during these HFOs-increase their firing rates over the course of sleep, suggesting a potential mechanism for selective memory consolidation.

      Strengths:

      The study addresses an under-explored area of sleep physiology: the fine-grained temporal coordination between the cortex and hippocampus during REM sleep. The identification of HFO "chains" and their association with higher theta power provides an interesting framework for understanding how the brain might organize information transfer outside of NREM sleep. The observation that specific hippocampal populations show differential firing rate changes based on their participation in these HFO events is a striking finding that warrants further investigation.

      Weaknesses:

      The primary weakness of the study lies in the lack of a clear distinction between global brain states and the specific events being analyzed. Because the authors compare HFOs across different sleep stages (NREM, tonic REM, and phasic REM) without sufficient controls, it is difficult to determine if the observed differences are intrinsic to the HFOs themselves or simply a reflection of the different physiological states in which they occur.

      We would first like to note in case it was unclear – as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly spiking activity patterns observed in prefrontal cortical-hippocampal circuits during these event-types that occur in the two sleep stages. As noted in response to Reviewer 1’s Comment #8 and Reviewer 3’s Comment #1, to avoid ambiguity and to remain consistent with existing literature, we use the following terms to describe these sleep-state specific events throughout the manuscript: Hippocampal sharp-wave ripples (SWRs) in NREM sleep, PFC ripples in NREM sleep, and PFC HFOs in REM sleep.

      We refer the Reviewer to our Response to Reviewer 1’s first comment where we provide additional analyses of theta periods outside of HFO times (i.e. baseline REM periods), including Figure S5. We also further address these comments as a response to Reviewer 2’s major comment #1 below.

      Furthermore, the evidence for "structured reactivation" is not yet convincing. The temporal alignment of these reactivation events appears inconsistent, with peaks occurring well before the HFO itself, and the analysis does not sufficiently control for pre-existing cellular assembly strengths.

      We have now addressed this as a response to Reviewer 2’s major comments #3 and #8 below, including Figure 6.

      Additionally, some of the sleep architecture presented appears atypical, such as very short REM bouts and direct NREM-to-REM transitions that bypass standard progression, raising questions about the consistency of the sleep detection across animals.

      We have now addressed this as a response to Reviewer 2’s major comment #2 below, including Figure S1 and Figure S3. We expect that addition of these figures will mitigate concerns about our sleep staging procedures and will clarify certain points raised by the Reviewer.

      Finally, the study does not account for potential confounds like baseline firing rates when interpreting the behavior of "high-cofiring" neurons, which may simply be the most active cells in the population.

      We have now addressed this as a response to Reviewer 2’s major comment #6 below.

      Reviewer #2 (Recommendations for the authors):

      In this study, the authors detect activity periods during REM sleep that feature high-frequency oscillations in the prefrontal cortex. They term these REM HFOs and report that they occur in "chains" phase-locked to theta oscillations. They contrast these with HFOs that appear in isolation and with ripples observed during NREM sleep, either in the PFC or in the hippocampus. The data presented is very interesting in places. The authors show that these HFOs are observed primarily during "phasic REM" periods that have higher theta power. It appears that overall firing is substantially lower surrounding these events. Intriguingly, it appears that CA1 cells that co-fire with the PFC during HFOs increase their firing rates over the course of sleep, whereas other neurons show a firing decrease. This seems to be the most striking finding of this study.

      Major:

      (1) The findings are generally intriguing, but the study makes some choices that are hard to understand. They begin comparing HFOs that occur in chains to HFOs that occur in isolation, even though it appears that chained and isolated HFOs likely occur at different times, some during tonic REM, some during phasic REM, and others during NREM. The lack of control for sleep state makes it difficult to determine if the reported differences are intrinsic to the HFO patterns or merely reflections of the underlying global brain state. Overall, I just didn't quite understand the motivation for comparing isolated and chain HFOs, as it seems natural that chains would occur during periods of greater synchrony.

      We appreciate the Reviewer for raising these points and agree that the properties of ripples and HFOs cannot be interpreted independently of the underlying brain state. As we mentioned in the initial response, we do expect that the generation of these ripples/HFOs in NREM and REM sleep are inextricably linked to global brain state (ex., cholinergic tone, as shown in the model in Figures 9-10), which results in differing patterns of activity across sleep states. We noted in the response to the public review comment #1 above that as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly, spiking activity patterns observed in cortical-hippocampal circuits during these event types that occur in the two sleep stages.

      Sleep state: Our primary goal in comparing isolated and chain HFOs (chain vs. isolated HFOs are only compared in REM periods, not NREM periods) was not to suggest that the differences that we observe are state-independent, but rather to provide evidence that temporal clustering of HFOs underlies distinct PFC dynamics, as well as enhanced CA1 engagement. Rather than attempting to dissociate ripple/HFO occurrence and the differential physiology that underlie NREM and REM sleep states, we view them as complementary—both sleep states are permissive to generation of high-frequency oscillations. Similarly, putative tonic and phasic REM substates (low vs high theta power) are differentially permissive for isolated and chain HFOs (Figure S3C). We add additional detail on tonic vs. phasic REM substages in Supplementary Figure S3, showing that the rate of HFOs and HFO chains is significantly elevated in putative phasic REM. While we would have liked to analyze REM HFOs during tonic and phasic states separately to be able to make more concrete conclusions, the scarcity of phasic REM sleep made it difficult to make accurate comparisons between the two, especially for spiking modulation. We have therefore removed the qualifying term “phasic REM” from the abstract.

      Also, we would like to clarify that while we investigated chains of events (ripples) in NREM sleep as a comparison to REM HFOs, we do not refer to them as HFOs in NREM anywhere in the manuscript. In NREM, while there are PFC ripples that are clustered into chains based on our definition (separation of <200 ms), we do not observe a prominent peak in the IEI distribution that would suggest entrainment by other oscillations (e.g. theta or spindles). This analysis was only to show that associated results are specific to REM HFO chains, and not seen during comparable chains of NREM cortical ripple events.

      Chain vs isolated HFOs: Regarding the Reviewers comment about why we chose to compare isolated and chain HFOs, the motivation is clearly demonstrated by differences in these events in spectral properties (Figure 3H-I) and spiking modulation (Figure 4F-G). Our initial motivation to investigate these chains of events in REM sleep came from our observation that PFC population activity aligned to REM HFOs was theta-modulated (multipeaked, suggesting multiple high-frequency events over a short duration). This led us to hypothesize that there may be chaining of events in REM sleep, which in line with previous studies demonstrating the clustering of events, such as spindles in NREM sleep (Darevsky, Kim et al. 2024). In that study (Darevsky, Kim et al. 2024), the authors showed that reactivation of motor patterns was more persistent during trains of spindle events as compared to isolated spindles. Furthermore, trains/chains of multiple hippocampal SWRs have been shown to underlie the replay of extended experience (Davidson, Kloosterman et al. 2009). Thus, the separation of high-frequency events into isolated and chained events has precedence and may have functional significance. Furthermore, since we see that isolated and chain events have a bias for occurring during putative bouts of tonic and phasic REM sleep, respectively (Figure S3C), characterization of both event types is an important step for understanding potential differences in interregional interactions during REM sleep. Furthermore, a recent appreciation for the role of sleep stage sub-states (Chang, Tang et al. 2025) further emphasizes the importance of investigating these events separately. We expect future studies to further dissect the roles of tonic and phasic REM states in memory and cognition, and our finding provides an account of the existence of different events for future reference.

      HFO chains during periods of greater synchrony: While it may seem natural that chaining would occur during periods of high synchrony (theta synchrony here), we show that our result is HFO-specific, especially for spiking activity modulation (Figures 4F-H and Supplementary Figures S5K-N), which makes it novel and important. Previous studies have focused on theta-gamma cross-frequency phase amplitude coupling, which has been proposed as a mechanism where slower theta oscillations temporally organize faster local population activity, thereby synchronizing neural ensembles within and across brain regions (Belluscio, Mizuseki et al. 2012). Here, we present a comparison of HFOs and gamma events and show that while gamma events can also occur in chains (added to Supplementary Figure S4H), possibly due to increased synchrony as the Reviewer stated above, we do not observe theta-modulated population activity aligned to these events (Supplementary Figure S4J), which is a defining property of HFOs that we propose underlies our results. This indicates that our reported results are unique to periods of synchrony associated with HFOs.

      Control analyses for theta periods: Regarding the comment about lack of control for sleep state, we refer the Reviewer to our response to Reviewer 1’s comments. We have performed additional control analyses comparing PFC activity during REM theta periods adjacent to detected HFOs and demonstrate that phasic PFC population activity is largely absent (Figure S5).

      We are aware that it is difficult to dissociate the generation of ripples and HFOs from the underlying brain state, since they are so tightly linked. However, we do expect that our clarifying points in addition to the REM theta state controls provide strong evidence that the results that we present are intrinsic to HFOs and not general reflections of activity during baseline theta activity in REM. Here, we further reiterate a few of the main results that demonstrate this:

      (1) Phasic spiking modulation of PFC population activity associated with HFOs is not present during baseline theta periods (Supplementary Figure S5K).

      (2) Phasic spiking modulation during HFOs is not linked to extracted gamma events (Supplementary Figure S4J).

      (3) Phasic spiking modulation, as well as activity suppression, is strongest during chains of HFOs (Figure 4F, Supplementary Figure S4J).

      (4) PFC-CA1 theta coherence surrounding HFOs increases relative to baseline (Figure 3E, z-scored relative to baseline coherence).

      (5) Assembly peaks are sequentially organized surrounding HFO chains but not during isolated or shuffled (baseline) HFO times (Figure 6C, Supplementary Figure S7E, and Author response image 2).

      (6) A higher proportion of CA1-CA1 pairs are high cofiring during HFOs as compared to baseline periods (Figure 7B), indicating specific CA1 engagement during HFOs (in addition to coherence).

      Lastly, we show that HFOs detected during periods of active behavior (high theta) on the W-Track are not associated with many of the defining features of HFOs in REM sleep, thus demonstrating the specificity of REM HFOs despite similar background theta activity (Figure S11, in response to Reviewer 3 comment #1).

      (2) It is crucial that the study provides REM-specific sleep examples for each of the data sessions, marking tonic and phasic REM and indicating when isolated and chain HFOs are observed. It remains unclear how interspersed these events are. Do some REM episodes just have isolated HFOs and others chains?

      We thank the Reviewer for raising this point and apologize for the lack of clarity. For additional transparency, we now provide example sleep state plots for each animal (Supplementary Figure S1H-I).

      We also provide hypnograms that show bouts of putative phasic REM on top of REM periods (Supplementary Figure S3). Additionally, as a compact way of demonstrating the validity of our separation of putative tonic and phasic bouts, we provide a plot showing the average velocity and spectrogram surrounding putative phasic REM bouts across all putative phasic REM transitions (Supplementary Figure S3A). Regarding the incidence of HFOs in tonic and phasic REM, we refer the Reviewer to Supplementary Figure S3D, where we show that HFO rate is significantly higher during putative bouts of phasic REM, during which they tend to be organized in chains as compared to putative tonic bouts (Supplementary Figure S3E). In line with this, we additionally report that although both isolated and chain HFOs occur in putative tonic and phasic bouts of REM sleep, the proportion of chain HFOs during phasic REM is significantly higher than that of isolated HFOs (Figure S3), indicating a bias for chains to occur in phasic REM sleep, potentially due to stronger theta input. However, since putative phasic REM accounts for <10% of REM sleep in our dataset, consistent with previous reports using similar methods (Mizuseki, Diba et al. 2011), we were unable to restrict spiking analysis to HFOs in phasic REM. We found that chain HFOs during both putative tonic and phasic REM sleep elicited theta modulated population activity in PFC, thus we pooled HFOs across REM states for analysis.

      While a larger proportion of chain events occur in putative bouts of phasic REM sleep as compared to isolated events (Figure S3), chain and isolated events occur in both putative tonic and phasic substates (Proportion of events in putative tonic REM is (1 - proportion in phasic) shown in Figure S3C, right). Thus, the majority of analyses comparing isolated and chain HFOs, especially for spiking data, were pooled across states. Instead of showing example plots for all 22 sleep epochs (22 out of 36 epochs with >5 s of putative phasic REM), we expect that this analysis will be sufficient to illustrate the distributions of isolated and chain HFOs across putative REM sleep substates. We have also removed the qualifying term “phasic REM” from the abstract.

      To clarify this, we have added a statement to the new “Limitations” section of the manuscript on Page 16, Lines 10-20:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states. (Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult. Finally, although our model predicts that distinct cell-type activity profiles shape REM sleep HFO dynamics, we did not record a suDicient number of interneurons to test these predictions directly. Future studies using appropriate behavioral tasks, longitudinal sleep recordings, and cell-type specific opto-tagging will be able to resolve these limitations and further clarify the roles of high-frequency oscillations in REM sleep.”

      We have now added Supplementary Figure S3.

      This is also important because some of the REM sleeps detected and shown in Figure S1H seem unusual. For example, in Animal 1, there are multiple bursts of REM that seem very irregular. In some other sessions, animals occasionally appear to enter REM sleep with very little preceding NREM, which goes against existing literature. It's not clear which ones of these meet the 30 s duration threshold. Are the chain events occurring in these periods?

      In reference to the hypnograms shown in Supplementary Figure S1H (Now Supplementary Figure S1I), these were generated by concatenating all 9 sleep epochs regardless of whether they passed the inclusion criterion of > 30 s of total REM sleep. Furthermore, only a subset of the epochs shown were included for analysis based on a secondary, manual inspection that is performed to confirm inclusion. Sleep state plots (e.g. Supplementary Figure S1H) were visually inspected to further confirm the transition into REM sleep, ensuring absence of noisy T/D ratio or spurious detection due to noisy signals – epochs where microarousals or persistent subthreshold fluctuations in animal movement induced noisy TD ratio increases, and thus inaccurate REM designation, were excluded. We thus used a total of 36 sleep sessions from a possible 90 sleep sessions.

      We apologize for not specifying what portions of data represented by the hypnograms were included. We have now provided updated hypnograms only illustrating the sleep epochs included for analysis (Figure S1I).

      We have now added Supplementary Figure S1.

      Regarding the Reviewer’s point “Are the chain events occurring in these periods?”: Yes, all of the chain events (and isolated events) come from these updated epochs that are now shown. No events in the excluded epochs were included, since we wanted to only analyze the data that came from curated REM epochs that we were confident in.

      (3) The analysis supporting structured reactivation was not generally convincing. Figure 4E does not provide convincing evidence of this. Indeed, reactivation strength is lowest around the time of the event, and appears highest 0.5s before. The example panel seems rather anecdotal. It's also not clear why REM reactivations should be compared to NREM ones here. I could not follow what was done in Figures 4F-H. Why should the first half and second half of an HFO event be correlated?

      We appreciate the Reviewer for raising these concerns. We agree that this section is somewhat dense and at time hard to follow, so we will clarify with further explanations and analyses (see also response to Comment #8 with new reactivation figures).

      In Figure 6B (originally Figure 4E), our intention was to show that if you simply align assembly activation to REM HFOs, a relatively flat response is observed when averaged across all assemblies, in stark contrast to assembly reactivation during NREM cortical ripples. This could suggest a couple of things: 1) There is no real assembly activation in response to REM HFOs or 2) Assembly activation is organized differentially (compared to NREM) surrounding HFOs. What we hypothesized, and then quantified based on observations, is that assembly activity is sequentially organized around HFOs. If this were true, it would suggest a consistent temporal relationship between REM HFOs and assemblies (for example, assembly 1 tends to be active 115 ms after the onset of chains, assembly 2 is active 230 ms after, etc.). Of course, the sequential pattern that we show in Figure 6A may arise trivially, especially since these types of sequential visualization plots can simply arise from noise. Thus, a cross-validation method must be utilized to ensure the sequences are indeed reflective of an underlying computation.

      In the methods section “Assembly sequence detection surrounding HFOs” we explain the splitripple/HFO procedure that we used to compare assembly sequences across two halves of the data, similar to methods used in hippocampal place cell sequence cross-validation (Plitt and Giocomo 2021, Sosa, Plitt et al. 2025). For every split and assembly alignment, a sequence similar to Figure 6A, right is generated based on the first half of aligned data. Then the second half of aligned data is sorted based on the peak indices of the first half of data and the correlation between the peak reactivation indices across all assemblies for the two datasets is calculated. A high correlation indicates high sequence similarity across the two halves of data (not two halves of an HFO as the Reviewer mentioned), suggesting temporal consistency of assembly reactivation. By utilizing this method, we show that assembly reactivation sequences across randomly chosen halves of data are most similar for chained HFOs (Figure 6C and Supplementary Figures S7C-F). We have now provided an additional control analysis that investigates this sequential assembly activity for time-shifted chain events (Author response image 2). As in Figure 6, assembly activity was aligned to the first event in each time-shifted chain. In addition to the analyses in Supplementary Figure S7, this control further indicates that sequential assembly activity is preferentially restricted to HFO chains.

      Author response image 2.

      Assembly sequences surrounding time-shifted HFO chains. (A) Distributions of r values calculated from the Pearson correlation between peak reactivation bins across all assemblies for two randomly chosen halves of the shifted HFO-aligned data. There was no difference between r values for time-shifted chains and shuffled data, indicating no structured assembly activity during periods outside of real HFOs chains.

      To further quantify this, we calculated two additional metrics: 1) the slope difference between fitted lines for the two halves of data for each split and 2) the absolute peak difference between assemblies in the two halves of data. First, for the slope difference metric, the slope of the best-fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data in Figure 6D indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. Secondly, Figure 6E is a quantification of the peak reactivation displacement between the two halves of data. If there is a high probability of small peak differences, as in the real data, this indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data.

      We apologize for the omission of the description of the slope and peak difference metrics that we used in Figures 6D,E. We have added information in the Methods section under “Assembly sequence detection surrounding HFOs” on Page 43, Lines 23-33:

      “Furthermore, the slope and peak differences were calculated as additional metrics of sequence and temporal reactivation consistency between the two halves of data, respectively. For the slope difference metric, the slope of the best fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. For peak reactivation difference, the temporal displacement of the peak reactivation index between the two halves of data was calculated and compared to shuffle. A high probability of small peak differences indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data. Shuffling of assembly strength was carried out as above.”

      (4) The authors argue that the occurrence of HFOs, rather than theta power, is the reason for lower MUA activity, but the analysis for this (Figure S4I) is quite confusing. The left panel actually seems to indicate that MUA is indeed lower when theta power is high.

      We apologize for the confusion regarding this figure and appreciate the Reviewer’s point that periods of high theta power can also appear to be associated with reduced MUA. We agree that the original presentation may not have clearly separated the contributions of theta and HFOs. What we convey with Supplementary Figure S5K (originally Supplementary Figure S4I) and new Supplementary Figures S5L-M is that the observed theta-modulated PFC spiking response (Figure 2A) cannot be solely explained by baseline theta periods (outside of HFOs) in REM sleep. Since theta oscillations are ubiquitous during REM sleep, an important control is to demonstrate that the theta-modulated population activity is not simply a consequence of ongoing theta activity. Thus, we aligned PFC activity to theta oscillations of varied power and show that the fluctuating, theta-modulated population activity is absent, indicating that HFOs associated with the theta oscillation are driving this phasic response.

      We are not solely arguing that the presence of HFOs is the driver of decreased PFC multiunit activity. In both the data and model, we show that the magnitude of theta power detected in PFC (possibly input to PFC) is inversely related to multiunit activity (Figure 9E). Our interpretation is not that theta is unrelated to MUA, but rather that HFO occurrence provides additional explanatory power beyond theta alone. Accordingly, we show directly in Supplementary Figures S5L-M, in response to Reviewer 1’s comment #1, that theta periods in phasic REM not associated with HFOs do not elicit the MUA activity suppression similar to HFOs.

      The phase-alignment performed is hard to follow and is not being applied to HFO periods. If the study is trying to argue that high-theta periods without HFOs in the same recording sessions show lower MUA, then perhaps some sort of shuffle or jitter would be more suitable. For example, in Figure 1B, it seems there are some high-theta periods that don't have HFOs and appear to have higher MUA.

      We expect that the new analyses where we provide additional baseline theta controls for periods adjacent to HFOs in Supplementary Figures S5L-M now clarifies this point. Briefly, alignment to theta phases at different temporal distances from detected HFOs does not exhibit the same fluctuating PFC activity as HFO alignment.

      Regarding the phase alignment procedure that we used for the control analysis, since REM sleep is characterized almost entirely by ongoing theta activity, the control condition was not a separate brain state but rather theta periods outside of HFOs. We therefore needed a systematic way to select comparable theta cycles and phase bins in order to make a valid comparison with HFO-aligned PFC population activity. Since we demonstrated that there is significant phase amplitude coupling between theta and HFOs, thus a theta phase preference of HFOs, we used that specific phase bin across multiple theta cycles to align PFC activity. This phase bin selection was performed separately for each epoch to account for inter epoch and animal variability in phase preference. We reasoned that this procedure would allow for a valid comparison as compared to random alignment, since activity was aligned to similar phases in the baseline theta vs HFO conditions.

      (5) As far as I could tell, the study does not distinguish between putative excitatory and inhibitory neurons in the PFC, but only in the CA1, even though these play very different roles in the model. What is the rationale for not separating these? How are reactivations to be interpreted among interneurons?

      We apologize for the lack of emphasis on this point, which was originally highlighted in Supplementary Figure S2G (now in Supplementary Figure S2F), and for omitting the explanation as to why we did not separately analyze putative excitatory and inhibitory neurons. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with recording primarily from pyramidal neurons (Supplementary Figure S2F, left). We therefore decided to pool and not separate the populations into putative pyramidal cells and interneurons for the spiking analyses presented. To further validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO-aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript, as we did not have enough interneurons to investigate them separately. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      Regarding the interpretation of reactivation in the context of interneurons, we expect reactivation reflects coordinated ensemble activity with excitatory neurons encoding task-relevant information, and inhibitory interneurons shaping timing and neural synchronization. We however did not record enough distinct interneurons to test the predictions of the model, which is now noted in the Limitations on Page 16.

      Relatedly, we find that PFC assemblies detected from pooled data have task relevant representations (Figures 5F-G).

      (6) Are the high-cofiring CA1 neurons generally higher-firing than the other cells? Could this perhaps explain why they behave differently?

      We apologize that this information was not more evident in the manuscript, as it is an important control. In Supplementary Figure S9A, we show that there was no difference in baseline firing rate between low and high cofiring CA1 neurons.

      We have now explicitly referenced this figure in the main text on Page 9, Lines 42-45:

      “Analysis of low and high cofiring CA1 neurons during REM HFOs showed that high cofiring neurons exhibited elevated activity during chained events as compared to low cofiring neurons (Figure 7D), independent of baseline firing rates (Figure S9A).”

      (7) It appears that the decreased firing around HFO's could be a consequence of the stronger firing modulation around these periods, related to time averaging, rather than suppression per se. How does the firing rate compare to other periods with similar modulation that might not have HFOs?

      We thank the Reviewer for raising this important point. We expect that the new analyses, where we provide additional baseline non-HFO-associated theta periods, and theta periods adjacent to HFOs at different temporal distance as controls in also Supplementary Figure S5L-N now clarifies this point. These figures show that the decreased firing rate is specific to HFO chains (Supplementary Figure S5N). Indeed, if the suppression that we observe is related to time averaging, or another analytical artifact, our claims of suppression during HFOs would not be valid. We present a number of results and provide further explanations to support the accuracy of our characterization of PFC population suppression.

      First, event-aligned multiunit activity was quantified as baseline-normalized population firing relative to detected events. For both NREM and REM, activity was normalized by the mean population firing rate during a baseline period within the same sleep state in which events were detected. Values are therefore expressed as deviations from baseline. This normalization allows comparison of relative changes in firing around events within each state. Importantly, values below baseline reflect reductions relative to the state-matched baseline period.

      Second, we show that smoothing activity with a larger gaussian kernel preserves the dip in population activity, consistent with suppression of activity surrounding HFOs (Figure 9C, note that this is a data figure presented in the context of the model). However, as raised by the Reviewer, this normalized measure does not on its own distinguish sustained suppression from transient deviations introduced by event-locked temporal structure in firing.

      Third, to address this, we employed an alternative method to demonstrate that HFO chains tend to occur during periods of PFC suppression (Figure 4G). Briefly, we detected events in PFC where activity fell below a threshold and calculated the probability of HFOs surrounding these “suppression” events. We refer the Reviewer to the Methods section under “Detection of population suppression” where we explain this procedure in more detail. We found that chain ripples, during which the strongest suppression is observed (Figure 4F), are associated with decreases in PFC activity (Figure 4G and Supplementary Figure S6).

      Fourth, we refer the reviewer to Figure S5N, where we compare the firing rates of PFC neurons during HFO chains and theta periods outside of HFOs during putative phasic REM bouts. The observed reduction in firing during true HFO chains compared to non-HFO periods therefore reflects HFO event-specific activity suppression rather than a common occurrence during baseline REM periods.

      Lastly, we performed a control analysis complimentary to Supplementary Figure S5K where we aligned PFC activity to the preferred theta phase of HFOs and investigated how distance from detected HFOs modulates PFC activity (Figure S5). We found that there was no consistent theta-modulated activity aligned to HFO-adjacent theta phases.

      Regarding the final comment, if the Reviewer meant “modulation” as in the theta modulation or suppression observed when PFC population activity is aligned to HFOs (Figures 2A and 4F), we are not aware of any other REM periods where this strong theta modulation or suppression of PFC population activity is present. To our knowledge, we are the first to demonstrate such a modulation of PFC population activity in REM sleep. The closest comparison that we can make is PFC activity aligned to gamma events that are coordinated with HFOs (suppression of PFC), but this is explained only with association with HFOs, as shown in Supplementary Figure S4J.

      (8) The assembly reactivation measure does not control for pre-existing assemblies. The term "activation strength" would therefore be more appropriate.

      We thank the reviewer for this important methodological point. The concern that ICA-based reactivation strength does not, by itself, distinguish behavior-induced reactivation from pre-existing assembly activity/structure is well-taken, and we have implemented several complementary analyses that directly address these concerns.

      First, the interleaved structure of our recordings (8 run epochs interleaved with 9 sleep epochs) allows us to investigate the within-session pre/post assembly strength differences (i.e. each W-Track run session has a preceding (pre) and following (post) sleep session). An increase in assembly strength from pre to post is a hallmark of behaviorally relevant assembly reactivation (Kudrimoti, Barnes et al. 1999, Peyrache, Khamassi et al. 2009). For each run epoch, the same run-derived templates were projected onto the preceding and following sleep epochs, and the pre and post strengths were compared. The distribution of post-minus-pre reactivation differences across all epoch pairs is significantly skewed toward positive values (Figure 6F), indicating that templates from the run epochs are more strongly expressed in the post-sleep epochs. This asymmetry cannot be explained by pre-existing assembly structure, which would predict similar assembly strengths.

      Second, reactivation strength in post-experience sleep increases across the experiment, with templates from later running epochs producing the strongest reactivation in the following post-sleep (Figure 6G). This increase in reactivation strength over time cannot be explained by preexisting assembly structure, which predicts similar assembly expression strength independent of experience.

      Third, the detected assemblies carry behaviorally meaningful structure. Assembly activation maps computed during running exhibit spatially organized "assembly fields" similar to single-cell place fields (Figure 6H), demonstrating that the detected assemblies represent specific spatial locations or task variables rather than behavior-independent states. Pre-existing co-firing structure unrelated to ongoing experience would not be expected to produce spatially tuned, task-locked assembly activation. Furthermore, this spatial tuning of assemblies was verified by comparison with surrogates, where assembly maps were generated using circularly shuffled activation times (1000 shuffles). Assemblies with p < 0.05 (z > 1.65) were considered to have significant spatial structure (Figure 6I).

      These results establish that what we measure is the selective re-expression of behaviorally relevant assemblies in subsequent sleep epochs, consistent with the use of "reactivation" in the established literature (Peyrache, Khamassi et al. 2009, Lopes-dos-Santos, Ribeiro et al. 2013). We have therefore retained the term "reactivation strength" and have added text to the manuscript noting these new results that justify the use of “reactivation” strength.

      We have now added Figure 6.

      We have also added the procedure for the calculation of spatial information to the Methods section under “Spatial information of assembly fields” on Page 44, Lines 8-23.

      (9) Can the study rule out that the rank-ordering in Fig 4C is related to firing rates? Higher-firing rates tend to fire earlier, and lower-firing cells later.

      We thank the Reviewer for raising this interesting point. Here, we assume that the Reviewer meant the baseline firing rates of the neurons, not the intra-HFO firing rates of the neurons. Indeed, when we look at baseline REM firing rates of these PFC neurons, we do find that neurons with higher firing rates tend to fire earlier than low-firing-rate neurons (Author response image 3). This is also true when PFC rank and firing rate are assessed for isolated REM HFOs and NREM PFC ripples (Author response image 3). Similarly, we also observe this relationship in CA1 during SWRs, during which rank order correlation is typically assessed as a method for replay detection. In line with this, a previous study has shown that CA1 neurons with high excitability at animals’ current location tend to initiate replay events (Karlsson and Frank 2009). Furthermore, high-firing-rate, rigid CA1 neurons are more active during SWRs than low-rate, plastic neurons (Grosmark and Buzsaki 2016), and there are distinct populations of neurons in both hippocampus and PFC that are preferentially active during immobility in sleep epochs (Jarosiewicz, McNaughton et al. 2002, Kay, Sosa et al. 2016, Tang, Shin et al. 2017), potentially biasing replay activity during high-frequency events. Similar dynamics may underlie activity during PFC ripples and HFOs in NREM and REM sleep, respectively. The critical point here is in the leave-one-out cross-validation that we implemented to determine sequence similarity—each left out event’s cell rank was correlated with the averaged rank of the template that was generated from all other events. This analysis provides a basis for our claim that there is preserved sequential PFC activity across HFO chains. We did not observe neuron firing consistency during isolated HFOs or during pseudo-HFO chains (coherently shifted chain HFO times), which indicates that REM HFO chains are unique temporal windows during which PFC activity proceeds in a more structured manner.

      Author response image 3.

      Firing rate difference of low and high rank neurons (A) Comparison of baseline firing rates of PFC and CA1 neurons split by average rank across all PFC REM HFOs, PFC NREM ripples, or CA1 SWRs. Baseline rates were calculated separately for NREM and REM sleep.

      Minor:

      (1) It gets confusing that the authors sometimes (but not always) refer to HFOs during NREM as "ripples" but not if they occur during REM. The terminology is inconsistent. When they refer to HFO chains, it seems they now pool between REM and NREM periods, as well as across phasic and tonic REM periods, which is confusing.

      We apologize for the confusion regarding the terminology. In the revised manuscript, we now use NREM ripples exclusively for NREM events and REM HFOs exclusively for REM events. We have removed mixed labels such as “ripple/HFO” except where a collective term is explicitly defined. We also clarified that HFO chains refer to REM events only and revised the relevant text/figure legends to avoid any implication that chain analyses pool NREM and REM events.

      (2) P7 L8: It might be helpful to emphasize "broader temporal distribution".

      We thank the Reviewer for the suggestion. We have updated the text on Page 8, Line 18:

      “Since we observed a broader temporal distribution of activity…”

      (3) P8 L9: What do they mean by spatial? Do they mean the place-fields of these same neurons during a previous task period?

      We apologize for the confusion. The Reviewer is correct. Here, we calculated the spatial rate map correlation between CA1 neurons as a measure of place field similarity during the W-Track session prior to the sleep epoch being assessed.

      For clarification, we have added “rate map” to the text on Page 9, Line 33:

      “…spatial rate map correlation…”

      We have also updated the y-axis label for Figure 7B for clarity.

      (4) P8 L27: What do they mean by "coordinated SWRs"? As opposed to what?

      Here, we are referring to our previous study where we investigated ripples in NREM sleep and showed that ripples and SWRs in PFC and CA1, respectively could either be independent from or coordinated with events in the other region (Shin and Jadhav 2024). A main result in the study showed that CA1 neurons are strongly suppressed during independent PFC ripples and that there was a relationship between activity suppression and reactivation during coordinated SWRs (CA1 SWR-PFC ripple coordination in NREM). We specifically mentioned “coordinated” since these are SWRs that are also coupled with SOs and spindles as compared to SWRs that are independent from PFC ripples (Shin and Jadhav 2024). Overall, we wanted to frame this result in the context of oscillatory coupling and mechanisms of memory consolidation.

      (5) P37 L28 says "we observed a bimodal distribution" but L31 says "unimodal". Which is it?

      We apologize for the confusion. We observed a bimodal distribution for CA1-PFC cofiring in REM sleep, but a unimodal distribution in NREM sleep. Because of these two observations, we decided to split the CA1 population into high and low cofiring neurons based on two different thresholds:

      (1) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) 0.

      (2) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) the mean of the distribution of averaged cofiring values.

      Using two separate thresholds to split high and low cofiring CA1 neurons demonstrates the robustness of the firing rate change result in Figures 7F and Supplementary Figures S9B-D.

      We added a statement that clarifies that the bimodal distribution was seen in REM sleep only on Page 44, Lines 37-43:

      “Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0. Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions.”

      Reviewer #3 (Public review):

      Summary:

      Shin et al. examine hippocampal-prefrontal interactions during sleep using simultaneous CA1 and prefrontal cortex recordings in rats performing a spatial memory task. They identify high-frequency oscillation (HFO) events in PFC during REM sleep that occur in theta-modulated chains and are associated with increased CA1-PFC coherence and sequential, sparse reactivation of cortical ensembles. This pattern contrasts with the synchronous reactivation observed during NREM cortical ripples. Together with a simple cholinergic network model, the authors propose that REM HFO chains represent a distinct mechanism for hippocampal-cortical coordination that complements NREM ripple-mediated processing during sleep.

      Strengths:

      A major strength of the work is the extensive electrophysiological dataset, which includes simultaneous recordings of large neuronal populations in both hippocampus and prefrontal cortex across behaviour and subsequent sleep. The analyses linking high-frequency events to population dynamics, interregional coherence, and ensemble reactivation are technically sophisticated and provide an incredibly detailed description of REM-associated cortical activity patterns. In particular, the demonstration that REM HFOs occur in chains aligned to theta phase and organise sequential activation of cortical assemblies represents a potentially important advance in understanding the neural structure of REM sleep activity. The integration of experimental data with a computational model further provides a useful framework for interpreting the observed differences between REM and NREM network states in terms of neuromodulatory influences.

      Weaknesses:

      While overall this study provides a highly valuable body of work, there are two primary limitations, which, if overcome, would provide substantially more significance to the overall characterisation of REM HFOs. Specifically:

      (1) Distinction from wake HFOs

      The results largely support the authors' claim that REM HFO chains represent a distinct pattern of neural coordination compared to NREM cortical ripples. The analyses consistently show differences between REM and NREM events in terms of neuronal modulation, ensemble structure, and interregional coupling. However, similar high-frequency events during wake are not examined. Since REM sleep shares several network features with wakefulness, including strong theta oscillations, evaluating whether comparable PFC HFOs occur during wake would provide clarity on whether these events are specific to REM sleep (and its associated functions) or represent a more general theta-associated phenomenon.

      To investigate PFC high-frequency oscillations during running behavior on the W-Track, events were extracted in the same manner as NREM and REM events (Methods). Events during wake were subset by periods where the animals’ velocity was >4 cm/s to provide a comparison of events during periods of high theta. While we were able to detect HFOs during wake that exhibited a similar spectral profile in the high frequency band, we did not observe 1) strong association with gamma or theta oscillations, 2) prominent HFO chaining, 3) HFO aligned theta modulated PFC activity, 4) comparable levels of theta phase amplitude coupling, 5) association with population suppression, or 6) a relationship between peri-event theta power and multiunit activity (Supplementary Figure S11). Many of the defining features of PFC REM HFOs are absent during wake, indicating REM specificity of the results we present.

      (2) Link to memory consolidation

      The manuscript proposes throughout that REM HFO chains may contribute to memory consolidation by coordinating hippocampal-cortical reactivation, but the evidence for this functional role remains indirect. The authors do highlight this as a limitation of the study - the inability to link their findings to learning - but it is not clear why. Further details of the behaviour results should be included. If no learning occurred across the eight behavioural sessions, this should be reported. If learning did occur, but could not be linked to HFO events, this should also be reported.

      To address these concerns, we have now added an explicit “Limitations” section in the main text of the manuscript that includes a statement about learning. We have also added Supplementary Figure S1 in the manuscript, which illustrates the performance of all 10 animals on the W-Track task. Finally, we have also included Figures 6F G in the manuscript, showing that PFC assembly reactivation strength during sleep epochs increases during learning.

      Reviewer #3 (Recommendations for the authors):

      Most of my specific comments were related to further clarification that will help the reader's understanding.

      (1) I'd recommend simplifying terminology. Open to debate, but would it not be simpler and clearer to just say NREM HFO vs REM HFO? Obviously, there is a need to mention how NREM HFOs have previously been referred to as cortical ripples, but I'm not sure it is such a helpful terminology to continue for the field, given, as you state, how different cortical ripples are from hippocampal SWRs. If not, I'd at least provide a clearer explanation early in the manuscript, distinguishing NREM ripples from REM HFOs but collectively still calling them 'cortical events'.

      We appreciate the Reviewer’s suggestion regarding terminology and agree that it would be simpler and clearer to use a single term, ripple or HFO, to describe these events. Initially, we had used a unified term (ripples across both states) but ultimately decided to switch to state-specific terminology due to previous comments we received on the manuscript and to emphasize the distinctions between NREM and REM events. Ultimately, we decided on calling them ripples in NREM and HFOs in REM since there is precedence for both terms in each respective sleep state (Khodagholy, Gelinas et al. 2017, Vaz, Inati et al. 2019, Bueno-Junior, Ruckstuhl et al. 2023, Shin and Jadhav 2024), but we do agree that this distinction can be confusing if not clearly stated. Thus, we have added an additional statement in the manuscript on Page 4, Line 44 to Page 5, Lines 1-3 for clarity:

      “Similar criteria were used to detect cortical high-frequency events in NREM and REM states; however, to avoid ambiguity and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout”

      (2) I assume experiments occurred during the light phase, but it would be good if this could be stated explicitly.

      We have now added text specifying that these experiments took place in the light phase on Page 34, Lines 9-12 of the Methods section under “Behavior”:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Page 2 Line 14: Rephrase to make clearer, e.g. 'that have a shift in...'.

      We have rephrased the sentence for clarification on Page 2, Lines 13-15:

      “REM HFO chains also preferentially engage CA1 neuronal populations that demonstrate a shift in their preferred theta-phase from behavior to REM sleep.”

      (4) Page 4 Line 36 - 'and find coherent shifts in TD', this is self-fulfilling. I would rephrase to something like 'resulting in...'.

      We have rephrased the sentence on Page 4, Lines 36-38:

      “We separated NREM and REM sleep stages based on theta-to-delta (TD) ratio in CA1, which revealed coherent shifts in TD ratio across CA1 and PFC at the onset and offset of REM sleep…”

      (5) Figure 1H left - clarify how many animals or multiunits this is based on.

      We have now updated Figure 1 legend to specify the number of animals and epochs included on Page 19, Line 4:

      “REM HFO aligned multiunit activity (MUA) in PFC (n = 10 animals, 36 epochs)…”

      (6) Figure 1I (right), it would be useful to see the x-axis frequency start from 0, since you are cutting the peak in power.

      We thank the Reviewer for this suggestion. We had initially set the frequency limits to 4 and 12 to specifically illustrate the absence of theta-modulated activity during NREM ripples. However, as suggested by the Reviewer, it is informative to expand the frequency range to ascertain the location of the peak frequency for NREM. Indeed, the peak frequency of NREM ripple-aligned PFC activity tends to be lower than 4 Hz, which is consistent with a single peak of activity that lasts <1 s.

      We have now updated Figure 2C with these new panels.

      (7) Figure 5 E/I - use of ** is confusing, it looks like a significance comparing e.g. quartile 2 to 1, but I think this is the correlation significance. I'd move ** to the top right corner and ideally include rho values.

      We thank the Reviewer for this suggestion. We have now added the r values for the correlation to Figures 7E and 8B to resolve any ambiguity. In addition, we updated Figure 7F, right and Figure 5B to maintain consistency across main figures.

      (8) Figure 7D - Why are stimulated and non-stimulated cells so different at baseline? This is not the case for the NREM results.

      The y-axes in Figures 10C,D (Formerly Figure 7) show the fraction of cells with at least one spike per 10ms bin, either for all pyramidal cells or all interneurons. We report in the figure legend that the stimulated cells represent only 30% of the network for any given stimulus (Page 32, Line 18), so at baseline in both the NREM and REM simulations, there are roughly 3x as many non-stimulated cells with a spike per 10ms bin compared to stimulated cells, as these populations are active at roughly equal firing rates per neuron outside of stimulation periods.

      (9) Methods - there is limited info on spike sorting procedure - can you provide a reference with further details (I couldn't find)? Is this method equally valid for identifying PFC units?

      Matclust is a MATLAB-based spike sorting graphical user interface that allows for manual curation of neuron clusters through the visualization of spike waveform amplitude, peak-to-trough, and principal components. Polygons or boxes are drawn around spike data points, and single unit clusters are resolved through refinement in multiple dimensions. It was developed by Mattias Karlsson and was first used in a publication reporting replay of remote experiences in the hippocampus (Karlsson and Frank 2009). Although there is no formal reference, it can be found at https://bitbucket.org/mkarlsso/matclust/src/master/. Other labs have used the software for clustering neurons from cortical areas (Yu, Liu et al. 2018, Proskurin, Manakov et al. 2023), which demonstrates its robustness across multiple brain areas.

      Additionally, examples of clustered neurons in PFC over the course of the experimental paradigm used here can be found in our previous publication (Shin, Tang et al. 2019). In addition to the aforementioned references, we show that PFC neurons can be accurately clustered and that neurons are stable over time, according to a number of cluster metrics.

      (10) Why were there no further analyses of pyramidal cells and interneurons beyond Figure S2?

      We thank the Reviewer for bringing up this important point, which is similar to Reviewer 2, comment #5 above. We repeat our response here. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with primarily recording from pyramidal neurons (Supplementary Figure S2F, left). All results are similar if we exclude putative interneurons. To validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      (11) Looking at Figure S1B, the second to last main block of NREM sleep shown has a clear peak passing TD threshold, but oddly not classed as REM - I can only assume this is due to the duration limits on your classification?

      Yes, this is due to the REM duration threshold that we implement in our sleep scoring algorithm. We used a minimum REM bout threshold criterion of 10 s for inclusion, as in previous reports (Rothschild, Eban et al. 2017, Zhang, Zhang et al. 2020). Additionally, we have included example sleep plots for all 10 animals in Supplementary Figure S1.

      (12) Include details of how head speed was calculated - just based on the 30fps video?

      Yes, the head speed of the animal was determined by tracking the animals’ position and calculating the speed based on cm/pixel values. We have now added more detail on this in the Methods section under “Surgical implant and electrophysiology” on Page 33, Line 44 to Page 34, Lines 1-2:

      “Additionally, the animals’ speed was calculated based on predetermined cm/pixel values and the position displacement between frames captured at 30 fps.”

      (13) Page 31 line 7 - With the reference you cite, they didn't really show tonic and phasic REM can be segregated based on theta frequency - they just defined it as such. You've done it for some, but I'd ensure all references to phasic REM are defined as putative - mostly missed within the discussion - I'd also add this as a brief limitation (without recording of eye movements or PGO waves).

      We thank the Reviewer for these suggestions. We have now ensured that all mentions of “phasic REM” are qualified with “putative” and have added the requested limitation in the Limitations section on Page 16, Lines 10-15:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states.(Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult.”

      We have also modified the wording in the Methods section under “Theta inter-peak intervals during bouts of high and low theta power” on Page 37, Lines 14-15 to indicate that the referenced article simply used the theta frequency-based method to define tonic and phasic REM – not to explicitly separate the two states:

      “Since previous studies have segregated putative tonic and phasic substages of REM sleep based on CA1 theta frequency…”

      References:

      Abdou, K., M. Nomoto, M. H. Aly, A. Z. Ibrahim, K. Choko, R. Okubo-Suzuki, S. I. Muramatsu and K. Inokuchi (2024). "Prefrontal coding of learned and inferred knowledge during REM and NREM sleep." Nat Commun 15(1): 4566.

      Aleman-Zapata, A., R. G. M. Morris and L. Genzel (2022). "Sleep deprivation and hippocampal ripple disruption after one-session learning eliminate memory expression the next day." Proc Natl Acad Sci U S A 119(44): e2123424119.

      Belluscio, M. A., K. Mizuseki, R. Schmidt, R. Kempter and G. Buzsaki (2012). "Cross-frequency phase coupling between theta and gamma oscillations in the hippocampus." J Neurosci 32(2): 423–435.

      Bueno-Junior, L. S., M. S. Ruckstuhl, M. M. Lim and B. O. Watson (2023). "The temporal structure of REM sleep shows minute-scale fluctuations across brain and body in mice and humans." Proc Natl Acad Sci U S A 120(18): e2213438120.

      Cairney, S. A., S. J. Durrant, R. Power and P. A. Lewis (2015). "Complementary roles of slow-wave sleep and rapid eye movement sleep in emotional memory consolidation." Cereb Cortex 25(6): 1565–1575. Chang, H., W. Tang, A. M. Wulf, T. Nyasulu, M. E. Wolf, A. Fernandez-Ruiz and A. Oliva (2025). "Sleep microstructure organizes memory replay." Nature 637(8048): 1161–1169.

      Cheng, S. and L. M. Frank (2008). "New experiences enhance coordinated neural activity in the hippocampus." Neuron 57(2): 303–313.

      Darevsky, D., J. Kim and K. Ganguly (2024). "Coupling of Slow Oscillations in the Prefrontal and Motor Cortex Predicts Onset of Spindle Trains and Persistent Memory Reactivations." J Neurosci 44(43). Davidson, T. J., F. Kloosterman and M. A. Wilson (2009). "Hippocampal replay of extended experience." Neuron 63(4): 497–507.

      Ellenbogen, J. M., P. T. Hu, J. D. Payne, D. Titone and M. P. Walker (2007). "Human relational memory requires time and sleep." Proc Natl Acad Sci U S A 104(18): 7723–7728.

      Foster, D. J. and M. A. Wilson (2006). "Reverse replay of behavioural sequences in hippocampal place cells during the awake state." Nature 440(7084): 680–683.

      Ghosh, M., F. C. Yang, S. P. Rice, V. Hetrick, A. L. Gonzalez, D. Siu, E. K. W. Brennan, T. T. John, A. M. Ahrens and O. J. Ahmed (2022). "Running speed and REM sleep control two distinct modes of rapid interhemispheric communication." Cell Rep 40(1): 111028.

      Grosmark, A. D. and G. Buzsaki (2016). "Diversity in neural firing dynamics supports both rigid and learned hippocampal sequences." Science 351(6280): 1440–1443.

      Helfrich, R. F., J. D. Lendner, B. A. Mander, H. Guillen, M. PaD, L. Mnatsakanyan, S. Vadera, M. P. Walker, J. J. Lin and R. T. Knight (2019). "Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans." Nat Commun 10(1): 3572.

      Jarosiewicz, B., B. L. McNaughton and W. E. Skaggs (2002). "Hippocampal population activity during the small-amplitude irregular activity state in the rat." J Neurosci 22(4): 1373–1384.

      Ji, D. and M. A. Wilson (2007). "Coordinated memory replay in the visual cortex and hippocampus during sleep." Nat Neurosci 10(1): 100–107.

      Karlsson, M. P. and L. M. Frank (2009). "Awake replay of remote experiences in the hippocampus." Nat Neurosci 12(7): 913–918.

      Kay, K., M. Sosa, J. E. Chung, M. P. Karlsson, M. C. Larkin and L. M. Frank (2016). "A hippocampal network for spatial coding during immobility and sleep." Nature 531(7593): 185–190.

      Khodagholy, D., J. N. Gelinas and G. Buzsaki (2017). "Learning-enhanced coupling between ripple oscillations in association cortices and hippocampus." Science 358(6361): 369–372.

      Kudrimoti, H. S., C. A. Barnes and B. L. McNaughton (1999). "Reactivation of hippocampal cell assemblies: effects of behavioral state, experience, and EEG dynamics." J Neurosci 19(10): 4090–4101. Lee, A. K. and M. A. Wilson (2002). "Memory of sequential experience in the hippocampus during slow wave sleep." Neuron 36(6): 1183–1194.

      Leemburg, S., V. V. Vyazovskiy, U. Olcese, C. L. Bassetti, G. Tononi and C. Cirelli (2010). "Sleep homeostasis in the rat is preserved during chronic sleep restriction." Proc Natl Acad Sci U S A 107(36): 15939–15944.

      Lopes-dos-Santos, V., S. Ribeiro and A. B. Tort (2013). "Detecting cell assemblies in large neuronal populations." J Neurosci Methods 220(2): 149–166.

      Mizuseki, K., K. Diba, E. Pastalkova and G. Buzsaki (2011). "Hippocampal CA1 pyramidal cells form functionally distinct sublayers." Nat Neurosci 14(9): 1174–1181.

      Nitsche, M. A., M. Jakoubkova, N. Thirugnanasambandam, L. Schmalfuss, S. Hullemann, K. Sonka, W. Paulus, C. Trenkwalder and S. Happe (2010). "Contribution of the premotor cortex to consolidation of motor sequence learning in humans during sleep." J Neurophysiol 104(5): 2603–2614.

      Peyrache, A., M. Khamassi, K. Benchenane, S. I. Wiener and F. P. Battaglia (2009). "Replay of rule-learning related neural patterns in the prefrontal cortex during sleep." Nat Neurosci 12(7): 919–926.

      Plitt, M. H. and L. M. Giocomo (2021). "Experience-dependent contextual codes in the hippocampus." Nat Neurosci 24(5): 705–714.

      Proskurin, M., M. Manakov and A. Karpova (2023). "ACC neural ensemble dynamics are structured by strategy prevalence." Elife 12.

      Rothschild, G., E. Eban and L. M. Frank (2017). "A cortical-hippocampal-cortical loop of information processing during memory consolidation." Nat Neurosci 20(2): 251–259.

      Shin, J. D. and S. P. Jadhav (2024). "Prefrontal cortical ripples mediate top-down suppression of hippocampal reactivation during sleep memory consolidation." Curr Biol 34(13): 2801–2811 e2809.

      Shin, J. D., W. Tang and S. P. Jadhav (2019). "Dynamics of Awake Hippocampal-Prefrontal Replay for Spatial Learning and Memory-Guided Decision Making." Neuron 104(6): 1110–1125 e1117.

      Siapas, A. G. and M. A. Wilson (1998). "Coordinated interactions between hippocampal ripples and cortical spindles during slow-wave sleep." Neuron 21(5): 1123–1128.

      Simor, P., G. van der Wijk, L. Nobili and P. Peigneux (2020). "The microstructure of REM sleep: Why phasic and tonic?" Sleep Med Rev 52: 101305.

      Singer, A. C. and L. M. Frank (2009). "Rewarded outcomes enhance reactivation of experience in the hippocampus." Neuron 64(6): 910–921.

      Sosa, M., H. R. Joo and L. M. Frank (2020). "Dorsal and Ventral Hippocampal Sharp-Wave Ripples Activate Distinct Nucleus Accumbens Networks." Neuron 105(4): 725–741 e728.

      Sosa, M., M. H. Plitt and L. M. Giocomo (2025). "A flexible hippocampal population code for experience relative to reward." Nat Neurosci 28(7): 1497–1509.

      Stark, E., L. Roux, R. Eichler and G. Buzsaki (2015). "Local generation of multineuronal spike sequences in the hippocampal CA1 region." Proc Natl Acad Sci U S A 112(33): 10521–10526.

      Tang, W., J. D. Shin, L. M. Frank and S. P. Jadhav (2017). "Hippocampal-Prefrontal Reactivation during Learning Is Stronger in Awake Compared with Sleep States." J Neurosci 37(49): 11789–11805.

      Tort, A. B., R. Scheffer-Teixeira, B. C. Souza, A. Draguhn and J. Brankack (2013). "Theta-associated high-frequency oscillations (110-160Hz) in the hippocampus and neocortex." Prog Neurobiol 100: 1–14.

      Valero, M., T. J. Viney, R. Machold, S. Mederos, I. Zutshi, B. Schuman, Y. Senzai, B. Rudy and G. Buzsaki (2021). "Sleep down state-active ID2/Nkx2.1 interneurons in the neocortex." Nat Neurosci 24(3): 401–411.

      van de Ven, G. M., S. Trouche, C. G. McNamara, K. Allen and D. Dupret (2016). "Hippocampal Offline Reactivation Consolidates Recently Formed Cell Assembly Patterns during Sharp Wave-Ripples." Neuron 92(5): 968–974.

      van der Helm, E. and M. P. Walker (2011). "Sleep and Emotional Memory Processing." Sleep Med Clin 6(1): 31–43.

      Vaz, A. P., S. K. Inati, N. Brunel and K. A. Zaghloul (2019). "Coupled ripple oscillations between the medial temporal lobe and neocortex retrieve human memory." Science 363(6430): 975–978.

      Wilson, M. A. and B. L. McNaughton (1994). "Reactivation of hippocampal ensemble memories during sleep." Science 265(5172): 676–679.

      Yang, S. R., H. Sun, Z. L. Huang, M. H. Yao and W. M. Qu (2012). "Repeated sleep restriction in adolescent rats altered sleep patterns and impaired spatial learning/memory ability." Sleep 35(6): 849–859.

      Yu, J. Y., D. F. Liu, A. Loback, I. Grossrubatscher and L. M. Frank (2018). "Specific hippocampal representations are linked to generalized cortical representations in memory." Nat Commun 9(1): 2209.

      Zhang, L. B., J. Zhang, M. J. Sun, H. Chen, J. Yan, F. L. Luo, Z. X. Yao, Y. M. Wu and B. Hu (2020). "Neuronal Activity in the Cerebellum During the Sleep-Wakefulness Transition in Mice." Neurosci Bull 36(8): 919– 931.

    1. Reviewer #1 (Public review):

      We appreciate the authors have provided answers to many of the points we raised, and the changes made to their manuscript, which we think strengthen the overall evidence presented. However, we find that some important controls are still missing across experiments.

      Major comments:

      (1) Shortcomings in Immunofluorescence experiments:

      a. Antibody cross-reactivity was only tested against CK1ɛ, but should also be tested against CK1α, which is abundant in U2OS cells, and is also known to be involved in cell-cycle regulation.

      b. Fig. 1: Statistical analyses are missing from the analysis.

      c. Fig. 2: No colocalisation analysis shown for figure 2, only some arrowheads pointing to puncta. Appropriate colocalisation statistics are important since for practical reasons, only a few representative images can be shown on the figure.

      d. Fig. 6: Even if the figure is illustrative, it is important to show centrosome staining to visualise CK1ẟ's recruitment to the centrosome in G2/prophase, especially since this information is used to propose the model in figure 7.

      e. For all figures: Please mention the number of independent biological replicates in the figure legends (1, 2, 6). For figure 1, if there are 3 independent biological replicates, the quantification should take all of them into account (as opposed to the data points corresponding to 10 cells), and statistics must be done appropriately, taking those independent replicates into account. Same for the colocalisation analysis in figure 2 once you include it.

      (2) Shortcomings in biochemistry experiments:

      On CalA control, this is not a matter of confirming that CalA treatment works in principle, but rather to confirm that CalA treatment worked in this specific replicate. Aliquots may lose potency (e.g. with freeze-thaw cycles / exposure to light), hence checking for enrichment of phospho-proteins is essential to confirm the treatment was successful in this particular instance. In the worst-case scenario, the company may have sent the wrong compound altogether! A positive and a negative control is the basis for every experiment to make meaningful interpretation. On a separate note, many experiments have control and siRNA or compound treatments on two different gels - this should be rectified as they are meaningless if different exposures have been selected for different immunoblots.

      (4) As the authors mention, the kinase is not fully inactive when tail phosphorylated. Recent research has also suggested that tail-phosphorylated CK1ẟ may show increased catalytic activity for a few select, specific substrates, in the co-occurrence of pT220 (Cullati et al., 2022; Cullati et al., 2024). It is thus tricky to directly infer that phosphorylated CK1ẟ is inhibited, when no positive control for CK1ẟ inhibition was shown in the evidence presented. It would be necessary to either nuance your claim or include a positive control for CK1ẟ inhibition. Please revise statements in the manuscript accordingly.

      (5) It would be important to include statistical analyses for the immunofluorescence data in Fig. 1 and 2.

      (9) The authors mentioned "In the eLife study, we show that inhibition of kinase activity by PF670462 stabilizes CK1δ and that the overexpressed kinase-dead mutant CK1δ-K38R is stable." Unfortunately, the data from biochemical analyses presented in the eLife publication is uninterpretable due to a lack of loading controls.

      (10) While the data presented in Penas et al. strongly suggests a link between CK1ẟ stabilisation and the APC/C-Cdh1 complex, it is the only study to have shown it. Given that (1) science relies on data reproducibility and (2) your proposed model relies heavily on the relationship between CK1ẟ stabilisation and the APC/CCdh1 complex, it would be appropriate to include the investigations mentioned in our original comment.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study identifies a non-canonical essential role for acyl carrier protein in maintaining apicoplast metabolism and blood-stage survival in Plasmodium falciparum. The main conclusions are largely supported by strong genetic and biochemical evidence, although some claims regarding the dispensability of fatty acid synthesis pathways remain incomplete. The work provides novel mechanistic insight into ACP-mediated stabilization of pyruvate kinase II and will be of broad interest to the malaria and apicoplast biology communities.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We kindly ask that the editorial assessment be revised to focus on the major conclusions. Alternatively, we would respectfully suggest revising the second sentence to something akin to “The main conclusions are largely supported by strong genetic and biochemical evidence, while function of fatty acid synthesis pathways in low-lipid conditions will require future studies to fully resolve.”

      Public Reviews:

      Reviewer #1 (Public review):

      This study provides evidence that the apicoplast-locaized isoform of acyl-carrier protein (ACP) has acquired important non-enzymatic functions in the malaria parasite. Previous studies have shown that the apicoplast-located FASII-dependent pathway of fatty acid synthesis is not essential in Plasmodium blood stages. In contrast, genome-wide knockout studies suggested that ACP, a key protein in this pathway, is essential in these stages, indicating that it may have additional non-canonical functions. In this study, the authors confirm that ACP is essential in Pf blood stages (using both apicoplast IPP rescue and conditional knockdown); show that this essential function requires modification with 4-phosphopantetheine and use proximity biotinylation and complementary immunoprecipitation pull-down approaches to provide compelling evidence that ACP binds to and stabilizes the apicoplast-located isoform of pyruvate kinase II. Notably, these interactions appear to differ from those associated with the binding of mitochondrial isoforms of ACP to proteins involved in Fe-S biosynthesis. Loss of ACP was shown to lead to a decrease in PKII levels and apicoplast DNA/RNA synthesis, consistent with loss of NTP synthesis in this organelle. The data are clear and very well described, and the findings represent a significant advance in our understanding of metabolic regulatory mechanisms in apicomplexan apicoplast studies.

      Strengths:

      The study uses a variety of complementary genetic approaches to demonstrate the essentiality of ACP and the enzyme involved in its activation with 4-PP in Pf blood stages, demonstrating that the ascribed non-enzymatic function is mediated by holo-ACP. Similarly, a number of complementary biochemical approaches, including proximity biotinylation, immunoprecipitation, and co-expression of PfACP and PK-II in a heterologous bacterial expression system, are used to confirm the physiological significance of the PfACP and PK-II interaction. The study also reports additional findings, such as the independence of P. faciparum blood stages on exogenous (media) fatty acids, indicating that intracellular stages can salvage all of their requirements from the red blood cell.

      Weaknesses:

      Overall, this is a very strong study. While questions remain around the function of other apicoplast ACP-interacting proteins detected in this study, I don't have any suggestions for significant improvements.

      We thank the reviewer for these positive comments.

      Reviewer #2 (Public review):

      This study focuses on revealing the essential divergent function of the Acyl Carrier protein (ACP) in the deadliest human malaria parasite, Plasmodium falciparum. More precisely, using inducible KO, cellular and biochemical approaches, the authors determined that instead of a canonical role for ACP allowing the de novo synthesis of fatty acids in the apicoplast (essential relict plastid) of the parasite, the enzyme couples with pyruvate kinase II to generate nucleoside triphosphate to maintain parasite survival during blood stages. The study is novel, well-designed, providing interesting new data on Plasmodium and apicomplexa biology. The results convincingly support the major claim of the study. However, it is currently incomplete to support some claims on the essentiality of some apicoplast pathways.

      In this study, Geher et al. focused on deciphering the role of the Acyl Carrier Protein (ACP) present in the relict non-photosynthetic plastid, i.e. the apicoplast of the most lethal human malaria parasite, Plasmodium falciparum. More particularly, they determined an essential function of ACP independent of its usual/typical function as the central protein for the normal function of the apicoplast Type II fatty acid synthesis (FASII) pathway. Rather, the protein seems to associate with the apicoplast Pyruvate Kinase II, together generating an essential nucleoside triphosphate (NTPs) source to fuel the apicoplast and parasite survival instead.

      By generating a TetR-DOZY-based inducible KD line for ACP, they confirmed that the protein is indeed essential to maintain apicoplast integrity and parasite survival during asexual blood stages, as previously predicted and experimentally shown. They showed that ACP requires a biochemical modification, typically activating the protein for its function in the FASII pathway, i.e. binding of the 4-PP group by holoACP synthase. Then, they showed that the other enzymes of the FASII pathway are likely dispensable during the blood stage, as they were able to generate a KO line of the first enzyme of the pathway, FabD (which was predicted to be essential in P. falciparum). Based on a cell culture approach in a controlled culture medium, they further claimed that, unlike current evidence-based hypotheses, the FASII pathway (and thus a potentially FASII-linked ACP) has no role/activity during blood stages. Using a proximity biotinylation approach, they determined that ACP associates with the apicoplast pyruvate Kinase II (PKII), previously shown to generate NTPs in the apicoplast for energy and DNA/RNA maintenance (Xia et al. 2019), and not to fuel the FASII pathway as its main function in blood stages. Finally, they showed that the disruption of ACP induces the reduction of the presence/content in PKII in the parasite, as well as the drastic reduction of the apicoplast DNA and RNA content. Together, they concluded that the main function of ACP is indeed the NTP formation via its association with PKII, rather than its canonical role for the generation of fatty acids in the apicoplast.

      To clarify, we conclude that the essential function of ACP in blood-stage P. falciparum parasites includes a critical stabilizing interaction with pyruvate kinase II. Apicoplast ACP presumably still plays a central biochemical role in FASII pathway function, but that role in FASII is dispensable for blood-stage parasites.

      This study is novel and focuses on a topic of particular interest in malaria biology, but also for most of the apicomplexa-related diseases, and beyond for plastid bearing orgnaisms and this unusual role for ACP. The study is well thought out with proper biochemical approaches that convincingly point to this association of ACP with PKII for NTP synthesis as a major function during P. falciparum blood stages. However, there are currently some important experimental issues/flaws, missing experiments that induced wrong interpretations and thus do not support some important claims of the study, notably for the role of FASII and the interaction between ACP and PKII.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We elaborate on these points and address the reviewer’s critiques below.

      Therefore, at this point, the study is only partial and would require major additions and/or important text edits/revisions before being considered for acceptance.

      We note that the manuscript has already been accepted for publication in accordance with the current eLife publishing model.

      Major points:

      From the graph of P. falciparum growth, we can see that in the lipid-rich condition, where both FabH KO and ACP KO can survive, the addition of mevalonate was essential for the growth of ACP KO. Along with the other evidence (PKII association, DNA levels...), we therefore agree that PfACP is involved in the mevalonate pathway.

      To clarify, our model is that ACP supports IPP synthesis by the apicoplast nonmevalonate/MEP pathway indirectly by stabilizing and thus supporting function by pyruvate kinase II that supplies the pyruvate and NTPs required for IPP synthesis by the MEP pathway.

      The authors claim that the FASII pathway is inactive/not essential in the P. falciparum blood stage. However, the authors have not shown any evidence on whether ACP is or not involved in the FASII pathway during the asexual blood stage.

      To clarify, there is overwhelming data in the prior published literature that we cite (including refs. 13, 14, and 32) to establish that FASII is dispensable for blood-stage Plasmodium growth in vivo in rodent parasites and in vitro culture in human parasites. Prior studies also strongly support a role for apicoplast ACP as the central scaffold for FASII-mediated acyl chain synthesis. However, our and prior studies support the conclusion that essential ACP function in blood-stage parasites is independent of its role in FASII.

      As currently designed, the experiments presented cannot conclude on that point for several reasons. Indeed, it was previously shown that (i) the expression of the protein from the FASII pathway are all present in blood stages and are significantly upregulated in patients that are under under "nutrient starvation" (Daily et al. Nature 2007), (ii) that, growing parasites under similar low lipid conditions in vitro induces an activation/upregulation of FASII, which can be measured by stable isotope precursor labelling and lipidomics (Botté et al. 2013).

      We are aware of these prior studies and cite and discuss the Botté et al. 2013 reference in our manuscript, which provided isotope-labeling evidence to support FASII activity in low-lipid growth conditions for P. falciparum. We note that neither study addresses whether FASII activity is required for growth in low-lipid conditions.

      (iii) that growing the PfFabI KO line under deprived lipid conditions leads to parasite death (Amiar et al. 2020), indicating that the FASII pathway can become critical, if not essential, depending on the host nutritionnal content together correlating patients' data and metabolic adaptation for the same reasons in the related parastie Toxoplasma gondii (Amiar et al. 2020, Krishnan et al. 2020, Liang et al. 2020, Primo et al. 2021, Charital et al. 2024, Dass et al. 2024, Bitew et al. 2025).

      All of the studies cited by the reviewer focus primarily or exclusively on Toxoplasma gondii parasites. We agree with the reviewer that these and other studies provide strong evidence that FASII activity contributes to growth of T. gondii parasites, including roles for apicoplast ACP that appear to differ from what we have unveiled for P. falciparum malaria parasites. We acknowledge and discuss these differences from T. gondii in the final section of the Discussion section and think that exploring these differences will be a fascinating area for future study.

      The Amiar et al. 2020 paper cited by the reviewer is the only study we are aware of that has directly tested the ability of a ∆FASII parasite (in this case, ∆FabI) to grow in low-lipid conditions. We acknowledge that they observed little to no growth of ∆FabI parasites in these conditions. Our growth assays with ∆ACP and ∆FabD parasites indicated a different outcome in which both WT and ∆FASII parasites grew similarly in low-lipid conditions. Our results thus contrast with the prior study. As noted below, the minimal lipid growth conditions explicitly reported in the methods section of the Amiar et al. paper are identical to those used in our study: fatty acid-free BSA, 30 µM palmitic acid, and 45 µM oleic acid (all sourced from Sigma) with daily media changes. Thus, the basis for these differences is unclear and additional follow-up work will be needed to explore and resolve these differences.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

      Here, the authors are expecting to show that FabH (and thus the FASII pathway) is not essential in an experiment that is not designed to be in low lipid conditions but rather in lipid rich conditions: Such high lipid conditions of culture in this study is granted by daily feedings with high fatty acid supplement (30-90 uM palmitic acid and 30-60 uM oleic acid). These fatty acid concentrations were used previously by Mitamura et al. (2005) and Miichi et al.(2007) to replace non-determined supplements such as Serum or Albumax supplement to grant similar growth by a completely controlled culture medium.

      This means the concentrations above do not represent limited fatty acid concentrations, especially not with daily feeding (representing an excess supplied amount of lipids, unlike regular 48h feedings) that allowed the authors to easily reach very high non-physiological parasitaemia of more than 20%!! Amiar et al. previously showed essentiality of FabI in P. falciparum in the limited fatty acid culture at a lower concentration (<30uM 16:0, <45um 18:1), than the Mi-Ichi et al. controlled medium with regular 48 h culture feeding. Therefore, with the current experimental settings, the FAH KO is placed in high lipid conditions, thus preventing any conclusion on its essentiality under low lipid conditions.

      The basis for the reviewer’s statements here is unclear, as this critique and the conditions it describes do not conform to the published conditions reported in the Amiar et al. 2020 paper. The methods section of that study for “Plasmodium falciparum growth assays” explicitly states (page e7):

      “Media was replaced daily, sub-culturing were performed every 48 h when required, and parasitemia monitored by Giemsa-stained blood smears. Growth assays in lipid-depleted media were performed by synchronizing parasites before transferring trophozoites to lipid-depleted media as previously reported (Botte ´ et al., 2013; Shears et al., 2017). Briefly, lipid-rich AlbuMAX II was replaced by complementing culture media with an equivalent amount of fatty acid-free bovine serum albumin (Sigma), 30 µM palmitic acid (C16:0; Sigma) and 45 µM oleic acid (C18:1; Sigma).”

      We used identical culture conditions to those described above: fatty acid-free BSA in place of lipid-rich AlbuMAX, 30 µM palmitic acid, and 45 µM oleic acid. We thus obtained growth results that contrast with the prior study and suggest that additional, future studies will be required to understand and resolve these differences.

      Furthermore, it is too uncertain to conclude that ACP is only essential for the mevalonate pathway.

      Please see our response above that clarifies our model for ACP function in supporting pyruvate kinase II and the many apicoplast pathways that appear to depend on PKII.

      This would be a similar discussion to the Yeh et al. 2011 and the Swift et al., where induced Apicoplast knockout caused parasites to require IPP to survive, but there were always remnant apicoplast vesicles and thus the putative presence of an active FASII in the parasite, where de novo fatty acid synthesis could be maintained.

      It is extremely unlikely that FASII remains active upon apicoplast disruption and loss of the apicoplast genome. The apicoplast-encoded SufB is lost upon apicoplast disruption and can no longer participate in making Fe-S clusters. Without Fe-S synthesis, the apicoplast lipoate synthase (LipA) cannot make lipoate to activate pyruvate dehydrogenase (E2 subunit) and produce the acetyl-CoA needed for FASII activity. There is no experimental evidence that FASII remains active upon apicoplast disruption and loss of the apicoplast genome.

      Amiar et al. (2020) and Krishnan et al. (2020) showed that disruption of FASII and absence of de novo FA synthesis in T. gondii could be compensated by the exogenous supplementation of myristic acid, C14:0.

      As explained above, we acknowledge that FASII contributes to Toxoplasma gondii growth and includes functions that appear to differ from P. falciparum.

      Here, high fatty acid supplementation using commercially available fatty acids may include unexpected fatty acid species such as myristic acid in palmitic acid or oleic acid, since all commercially available fatty acids guarantee only >99% but not 100%. If P. falciparum requires a very, very low amount of myristic acid to survive, the amount of possible contamination, like 1 nM, may be sufficient to maintain their survival. Thus, ACP and FabH might be very important to generate de novo fatty acids within parasites, but this was not shown by the authors.

      As noted above, we used identical culture conditions and commercial sources of defined fatty acids to those reported in the Amiar et al. study. We do not see a basis for the reviewer’s critique that the two studies utilized differing culture conditions. Nevertheless, we agree that future studies are needed to understand and resolve these differences.

      Therefore, the manuscript currently contains incorrect conclusions on the potential essentiality/use of FASII, against current experimental evidence.

      As explained above, we do not see a basis for the reviewer’s critique here or for viewing one study as more or less definitive than the other, as identical culture conditions were used yet contrasting results were obtained for reasons that remain uncertain. Future studies beyond the scope of the present manuscript will be required to fully understand and resolve these differences.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      We either request more solid experimental evidence showing the absence of fatty acid synthesis at low fatty acid conditions by re-doing the growth assay in the lower fatty acid feeding conditions without daily feeding to clarify if the ACP and FabH are essential in the blood stage growth, or not; as well as showing the absence of fatty acid synthesis at low fatty acid conditions using isotope labelled precursor. Without these, the authors cannot conclude on this important point. Alternatively, toning down the text to acknowledge the possibility of FASII being active and critical under certain conditions would be acceptable.

      We note that the Amiar et. al 2020 study cited by the reviewer reported growth assays for WT and ∆FabI NF54 P. falciparum in low-lipid conditions that similarly lacked direct tests of FASII activity by isotope labeling.

      We agree that our results, which utilized distinct ∆ACP and ∆FabD NF54 (PfMev) lines, contrast with the results and conclusions of the Amiar et al. study. We fully agree with the reviewer that future studies, utilizing tandem growth assays and isotope-labeling metabolic flux assays (e.g., mass spectrometry), will be required to fully test and understand the dependence of FASII activity on the lipid content of the growth medium and the functional dependence of P. falciparum growth on FASII activity in low-lipid conditions.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

    1. Author response:

      The following is the authors’ response to the current reviews.

      We again thank the editor and reviewers for their detailed attention to our work. In our previous revisions we endeavored to address the principal concerns raised by reviewers that we were capable of addressing. We recognize that asymptomatic pertussis transmission represents a particularly thorny area of epidemiology and public health, where multiple (and sometimes overlapping) mechanisms have been proffered even as empirical evidence remains thin, particularly in low-resource settings such as sub-Saharan Africa. A key finding of our work is that prospective surveillance in such a low-resource setting revealed abundant evidence of otherwise unobserved asymptomatic incidence. Moreover, as we note in our prior revisions, this finding is supported by recent work in South Africa and elsewhere (Kayina et al., 2015; Moosa et al., 2019, 2025). As such, we believe that further prospective surveillance in similar settings would be highly informative, a point that we have sought to emphasize in our present manuscript (and associated commentary).

      While we broadly agree with many of the concerns raised by the reviewers, we believe that we are unable to significantly strengthen the present work through further revisions. We do, however, wish to respond to several points raised in these reviews. Of note, a reviewer raises the possibility of a "breakthrough strain" without reference to existing literature. We agree that we cannot test this hypothesis, and though it is not incompatible with our own findings, it is, however, not consistent with recent molecular surveillance in South Africa (Moosa et al., 2023). The reviewer also raises the potential of low adult vaccination coupled with recent reintroduction. This hypothesis relies on our investigation looking "at the right place and the right time", and further does not explain how immunologically naive adults would have escaped morbidity. We have adopted what we believe is a more parsimonious interpretation of our results (i.e., that asymptomatic infection represents evidence of previous immune exposure), though we agree that a more thorough exploration of this particular issue is warranted, particularly in light of our persistently colonized mothers. In addition, we have noted similar studies in sub-Saharan Africa that also found widespread evidence of asymptomatic pertussis, which we believe is inconsistent with a “right time, right place” interpretation.

      A reviewer also pointed to the work of Warfel et al. (2014) and Althouse and Scarpino (2015). We are familiar with both of these studies and agree with their broad relevance to the field (our apologies for omitting Warfel et al.). The reviewer states that, "If the mechanism underlying the results in Zambia is that either WP or natural infection does not block transmission (in the absence of a breakthrough strain), that would upend many of the assumptions in pertussis research." Critically, we believe that our prospective field study of human patients in a real-world public health system complements previous research, including animal trials and simulation studies. Simply put, given that our study was unable to establish the prior vaccination or exposure status of participants, we do not claim to have shown evidence for transmission despite wP vaccination or prior infection.

      Regarding Althouse & Scarpino (2015), we believe that, for the majority of readers, the most compelling analysis in their paper was the examination of genome sequences that pointed to substantial asymptomatic transmission in the US. This conclusion emerged from their population model, which required that “births” (representing transmission events) exceeded “deaths” (representing recovery of infectious individuals) in order to be consistent with the sequence data. Unfortunately, this paper does not provide a detailed explanation of their methods and data sources, nor is this work directly reproducible through, for example, an open-access code/data repository. We have explored the availability of US genome sequences over the time period of their study and were able to find only 36 sequences: 2 from the pre-vaccine era, 8 from the wP vaccine era, and 26 from the aP vaccine era. Given this notable imbalance in the number of sequences (and thus sequence diversity) that was biased in favour of the most recent time period, is it then surprising that the “birth rate” in their model had to exceed the “death rate” in order to match the genetic diversity in the data? Based on a careful inspection of this work, we do not consider its conclusions to represent a gold standard against which all subsequent studies should be judged. We also note that genomic surveillance and analysis of pertussis remains sparse relative to other fields, though recent works have added dramatically to the corpus of available sequences (Bridel et al., 2022).

      Finally, we note that our previous revisions addressed several concerns raised in the present reviews. For example, we previously sought to address reviewers' about our presentation of the strength of our evidence. In this regard, we broadly agree with the reviewers, and we now state that our results "suggest that pertussis transmission occurs between minimally symptomatic mothers and their newborn infants." We believe this largely addresses a present reviewer's concern that 'the mother-to-infant transmission pathway should be framed as "highly suggestive" rather than "confirmed"'. We also note that our results examine three different threshold Ct values (survival analysis, Fig 4), a point that we believe partially addresses a reviewer's suggestion to "including a sensitivity analysis using a stricter cut-off" and concerns about "the decision to use a Ct<45 threshold, as this is higher than standard clinical cut-offs". Indeed, we discuss the issue of clinical cut-offs (and their appropriateness) at some length in the section, "Test reliability, disease surveillance, and public health where we state, "We recognize that such weak and potentially ambiguous signals may not be appropriate for clinical diagnosis. However, our results demonstrate that they nonetheless contain valuable information about pathogen presence and infection intensity that can (and should) be leveraged for disease surveillance." We have also included in the present work a detailed discussion of qPCR sensitivity and efficiency that we believe should interest others working in pertussis surveillance.

      We do not view our own research as the "last word" in this rather controversial subject. In this spirit, we have attempted to present our work transparently, state our claims carefully, and underscore future activities that we believe would benefit the pertussis research community going forward. For example, we agree that further attention to shared exposure and functional immune data among low-resource communities could provide valuable insights into epidemiology and ecology of pertussis. However, we also believe the trade-offs of including one set of activities over another should be clearly acknowledged by researchers, clinicians, and public health officials. To simply state that we must measure more fails to account for the very real resource constraints that we all face.

      Althouse, B. M., & Scarpino, S. V. (2015). Asymptomatic transmission and the resurgence of Bordetella pertussis. BMC Medicine, 1–12. https://doi.org/10.1186/s12916-015-0382-8

      Bridel, S., Bouchez, V., Brancotte, B., Hauck, S., Armatys, N., Landier, A., Mühle, E., Guillot, S., Toubiana, J., Maiden, M. C. J., Jolley, K. A., & Brisse, S. (2022). A comprehensive resource for Bordetella genomic epidemiology and biodiversity studies. Nature Communications, 13(1), 3807. https://doi.org/10.1038/s41467-022-31517-8

      Kayina, V., Kyobe, S., Katabazi, F. A., Kigozi, E., Okee, M., Odongkara, B., Babikako, H. M., Whalen, C. C., Joloba, M. L., Musoke, P. M., & others. (2015). Pertussis prevalence and its determinants among children with persistent cough in urban Uganda. PLoS One, 10(4), e0123240.

      Moosa, F., du Plessis, M., Weigand, M. R., Peng, Y., Mogale, D., de Gouveia, L., Nunes, M. C., Madhi, S. A., Zar, H. J., Reubenson, G., & others. (2023). Genomic characterization of Bordetella pertussis in South Africa, 2015–2019. Microbial Genomics, 9(12), 001162.

      Moosa, F., du Plessis, M., Wolter, N., Carrim, M., Cohen, C., von Mollendorf, C., Walaza, S., Tempia, S., Dawood, H., Variava, E., & others. (2019). Challenges and clinical relevance of molecular detection of Bordetella pertussis in South Africa. BMC Infectious Diseases, 19, 1–11.

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018: Findings from the PHIRST study. Journal of Infection, 106550.

      Warfel, J. M., Zimmerman, L. I., & Merkel, T. J. (2014). Acellular pertussis vaccines protect against disease but fail to prevent infection and transmission in a nonhuman primate model. Proceedings of the National Academy of Sciences, 111(2), 787–792. https://doi.org/10.1073/pnas.1314688110


      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study investigates the role of asymptomatic pertussis carriage in transmission between mothers and their infants, in particular. The authors used a longitudinal cohort study that involved 1,315 mother-infant dyads in Lusaka, Zambia, and they utilized qPCR-based detection of IS481 to track Bordetella pertussis transmission over time. Insights from the study suggest that minimally symptomatic or asymptomatic mothers may act as a reservoir for B. pertussis transmission in the infants, thus challenging the traditional surveillance methods that focus on symptomatic cases. Additionally, the study also identified a subgroup of persistently colonized individuals where mothers were majorly asymptomatic despite sustained bacterial presence.

      The authors aimed to improve comprehension of pertussis transmission dynamics in high-burden low-resource settings, and they advocated for enhanced molecular surveillance strategies to capture full pertussis infection, including those that might have gone undetected.

      Strengths:

      The strengths are the use of innovative study design, especially the longitudinal approach and routine sampling, rather than symptom-driven testing that minimizes bias in the study. The methodology was also rigorous and transparent by evaluating the IS481 signal strength to classify pertussis detection and conducting retesting to assess qPCR reliability. There were also important epidemiological insights, and the findings challenge the traditional wisdom by suggesting that pertussis transmission may frequently occur outside of symptomatic cases. The findings also showed its relevance to global health and policy by arguing for the incorporation of molecular tools like qPCR for surveillance of pertussis in low-resource settings.

      Weaknesses:

      These include reliability on qPCR-based detection without additional validation measures like confirmatory culture or serology. There are also potential alternate explanations for transmission patterns observed in the study such as shared environmental exposure or household transmission. Additionally, there is limited generalizability as the study was done in a single urban site in Zambia. There is also a lack of functional immune data.

      Reviewer #2 (Public review):

      Summary:

      In this paper, the authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage. This work represents an important contribution to our understanding of the global burden of pertussis. Also, it highlights the still under-appreciated role of asymptomatic transmission across many infectious diseases (including vaccine-preventable ones).

      Strengths:

      Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage.

      Weaknesses:

      While I am quite enthusiastic about the work, I am concerned that a number of likely relevant confounders were not discussed and that the broader implications of their findings were not well grounded in the existing literature. For example, I could not find information on the vaccination status of the mothers in the study. Given the conclusions about asymptomatic transmission and the durability of immunity, it is important to know the vaccination status of the mothers. Moreover, did the authors have other metadata on the mother/infant dyads, e.g., household size, vaccination status of household members, etc.? Given the potential implications of more widespread asymptomatic transmission associated with pertussis infection, I believe the authors should better couch their results in the context of the broader debate around asymptomatic transmission.

      We appreciate the reviewers' detailed feedback. We provide an overview of our responses here and we address specific recommendations below. In light of reviewers’ comments, we have revised our manuscript in order to improve the clarity of our presentation and to better situate our results within the context of the existing literature. Unfortunately, as the field study has been concluded, many of the reviewers’ recommendations are not possible. These include additional testing (i.e., culture or serology) or sequencing. We have updated the manuscript to more clearly indicate our knowledge regarding maternal vaccine status and immunological immunity of study participants. We have also provided a more comprehensive overview of existing pertussis studies, including genomic surveillance and details regarding sub-Saharan Africa and Zambia in particular. Finally, we have revised the formatting of Figure 4 (survival analysis) to more clearly highlight differences between mothers and infants and to better align with the text, and note that the underlying results are unchanged.

      A particular concern raised in the reviews that we wish to address is the recommendation of culture- or serology-based tests as "confirmatory". We have revised the manuscript in light of this feedback to better reflect our own position on this matter. We believe these recommendations do not adequately account for important trade-offs between testing sensitivity and specificity that are widely recognized in both clinical practice and epidemiology (Enøe et al., 2000; Florkowski, 2008; Swift et al., 2020). When the results of different testing methodology disagree, rarely is one method, a priori, correct. Rather, the disagreement may point to specific test limitations or important biological questions about the study system.

      In the case of pertussis detection, cell culture is recognized for its very low sensitivity, while serological detection is complicated by debate around appropriate threshold levels and time horizons for seroconversion and subsequent decay (Lee et al., 2018; van der Zee et al., 2015). Furthermore, while anti-PT antibodies are a common target of serological detection, these are not reliable correlates of protection (Mills, 2001; Wilk et al., 2019), nor are they reliably generated in response to colonization (de Cellès & Rohani, 2024; Graaf et al., 2020). Overall, the detailed relationship between exposure, carriage, transmissible infection, and the dynamics of anti-PT serology remains poorly characterized (Craig et al., 2020; de Cellès et al., 2025).

      While we agree that these are important questions in epidemiology and public health, we nonetheless wish to highlight that there is no “free lunch": each additional test and protocol comes with additional cost and complexity that should be evaluated based on the specific goals of the intended surveillance. In our case, the repeated sampling of longitudinal surveillance serves as a low-cost "confirmatory" testing regime. We disagree that cell culture would have added value to the present study and would not recommend its addition to future studies (primarily due to low sensitivity). While we agree that before-and-after serology of mothers would have added important context to the present study, we nonetheless expect that significant ambiguity would have surrounded any such results (e.g., Moosa et al. (2025)).

      One area that we strongly agree warrants further attention is the household dynamics in pertussis transmission, particularly in low-resource settings where crowding is common. In our study we were not able to rule out environmental and/or shared transmission events, though our survival analysis did demonstrate a greater impact of mothers on infants than vice versa, results which suggest a causal mechanistic role. In previous studies we detailed the demographics of household size, number of children, and mothers' age (Gill et al., 2021; Gunning et al., 2020), though we have not conducted formal analyses of these important covariates here. We also note that the code and data are freely available, allowing for others to build on our work.

      We believe that future prospective studies are an invaluable tool for directly tracking pertussis disease transmission, including both community and household studies. We have argued here for the value of qPCR-based community surveillance, which could integrate into existing public health activities. Regarding household studies, we note that a key challenge in implementing these studies is selecting an appropriate sampling interval and duration to best capture epidemiological linkages. Our results suggest that qPCR-based real-time population-level surveillance could be used to initiate such a prospective household study during a pertussis outbreak so that a higher sampling frequency (e.g., weekly swabs) could be gainfully employed over a shorter time period.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Enhance Validation of qPCR Findings:

      We address these comments above at greater length. We also note that the term “false positive” is rather ambiguous here, as we lack a clear distinction between carriage and transmissible infection. We note that our manuscript includes a considerable discussion of qPCR validation, including negative controls and sample retesting. We agree that test sensitivity and specificity remains an important question, and have endeavored to clearly indicate these concerns throughout the manuscript.

      (2) Clarify Transmission Dynamics:

      While we agree that this is an important question, we lack the relevant sequence data to test it. Notably, we suspect that any such phylodynamic linkage would require genomic sequencing due to the relatively low genetic diversity observed in pertussis (see, e.g., population-level estimates of time to most common ancestor in Lefranc et al., (2022). However, sequencing pertussis genomes remains resource-intensive, and we expect that the deployment of such sequencing at scale would be cost-prohibitive in low-resource settings.

      We have revised our manuscript to underscore uncertainties around shared/household exposure. We also direct the reviewer’s attention to our survival analysis, where a notable asymmetry exists between mothers and infants. Here, mothers’ prior qPCR signals exhibit a larger impact on their infants than infants on their mothers (Fig 4). If shared exposure was the principal cause of the observed increase in hazard in ego from alter, then we would expect no such asymmetry between mothers and infants.

      (3) Expand Discussion on Public Health Implications:

      As noted above, we have revised our introduction and expanded our discussion to better account for the existing literature. And, while we are hesitant to put forward specific recommendations based on our study, we feel confident in stating that active pertussis surveillance in low-resource settings is A) almost entirely absent B) possible to achieve, and C) necessary to resolve long-standing questions about pertussis epidemiology at regional and national levels. We have endeavored to clarify these points, particularly within the discussion.

      (4) Address the Role of Immunity More Directly:

      While we lack such immune data, we have pointed to recent work from the notable PHIRST study in South Africa, as well as highlighted ambiguities surrounding these data.

      Reviewer #1 (Recommendations for the authors):

      (1) Do we know the vaccination status of the mothers in the study? … If these data are not available, I think that the paper must be re-framed to acknowledge that all the conclusions are statistical in nature, based on publicly available vaccine coverage data from Zambia.

      We do not have information on the immunization status of mothers, though we cite national rates for Zambia across the relevant time period. We have clarified this point in the revised manuscript. We have endeavored to clearly acknowledge that many of our conclusions are statistical in nature and to clearly quantify the strength of evidence.

      We strongly disagree that anomalously low vaccination rates amongst mothers (i.e., relative to national averages) would materially alter the interpretation of our findings.

      Overall, our findings strongly suggest ongoing pertussis transmission in this population. Based on this, we expect that mothers in our study who were not vaccinated would likely have some degree of infection-derived immunity. Indeed, some have argued that the preponderance of mild/asymptomatic infections in mothers is, of itself, evidence of prior immunological exposure (Fine & Clarkson, 1982).

      (2) Do we know anything about rates of pertussis in Zambia, especially in the study site?

      We address this important question in the discussion. In particular, we state that: “As a populous, middle-income and primarily urban country, Zambia offers an evocative example of pertussis surveillance, where no cases have appeared in official WHO reports since 2009”.

      (3) I couldn't find information in the paper related to the severity of infection in the infants. It's mentioned in the section describing results in Figure 5, but I only saw analyses with symptoms (as opposed to severe symptoms). Do you have outcome data from infants testing positive?

      This question was addressed in more detail in our previous work (Gill et al., 2021), which we briefly summarize in the Introduction. We also show the frequency of severe symptoms in Fig 5 (bottom panel), and detail mild versus serious symptoms in our subgroup analysis (Fig 7D).

      (4) Do we know anything about vaccine-resistant strains of pertussis in Zambia?

      We are not aware of any such work. As we noted above (and now address in our Discussion), widespread genomic surveillance and microbiological characterization of pertussis are sorely lacking across Africa.

      (5) While I believe sequencing is beyond the scope of the current study, the authors should comment on the potential utility of sequencing elements of the pertussis genome and use that to demonstrate causality and direction of transmission more strongly.

      We believe that existing literature has addressed the potential of sequence data and phylodynamics to infer transmission, particularly for pathogens with high mutation rates such as RNA viruses. To date, research into the phylodynamics of pertussis has focused exclusively on population-level dynamics (Lefrancq et al., 2022), where estimates of time to most recent common ancestor (TMRCA) are long, indicating low genomic variability at the scale of countries and years. To our knowledge, no work on pertussis has directly inferred transmission chains from sequence data. Given the existing evidence, we expect that any such work would require genome-level sequencing, which would likely be cost-prohibitive in low-resource settings.

      (6) … However, it would be helpful to understand more about how your results fit into the broader story around pertussis resurgence. … if the infant cases were all mild, they might never have been captured in surveillance data sets.

      We believe that a key result of our study is the remarkable mismatch between country-level symptoms-based surveillance and prospective surveillance, which demonstrates that such mild cases have almost certainly not been captured. These findings are mirrored by recent work in South Africa (now addressed in our Discussion, see Moosa et al. (2025)). We believe that prospective surveillance, particularly in under-surveilled regions, is critical to understanding pertussis transmission writ large, which we have attempted to communicate throughout our discussion.

      (7) Relatedly, if there are still high rates of asymptomatic mother-to-infant transmission with whole cell vaccination, then why is there an observed drop in infant pertussis following whole vaccination in most countries?

      In previous work, we demonstrated that some infants in this cohort exhibited asymptomatic infection (Gill et al., 2021). We note that a drop in pertussis incidence amongst infants after the roll-out of the whole-cell vaccine is not contradictory with our findings. We want to clarify that our results, and evidence that mother-to-infant transmission can occur, does not imply that the whole-cell vaccine fails to protect against transmission.

      We have previously used epidemiological evidence to infer the population-level impacts following the roll-out of whole-cell pertussis infant immunization. For example, we observed an increase in the inter-epidemic period that, together with the drop in infant cases, are consistent with a reduction in transmissible infections (Broutin et al., 2010; Rohani et al., 2000).

      References

      Broutin, H., Viboud, C., Grenfell, B. T., Miller, M. A., & Rohani, P. (2010). Impact of vaccination and birth rate on the epidemiology of pertussis: A comparative study in 64 countries. Proceedings of the Royal Society B: Biological Sciences, 277(1698), 3239–3245.  https://doi.org/10.1098/rspb.2010.0994  

      Craig, R., Kunkel, E., Crowcroft, N. S., Fitzpatrick, M. C., Melker, H. de, Althouse, B. M., Merkel, T., Scarpino, S. V., Koelle, K., Friedman, L., Arnold, C., & Bolotin, S. (2020). Asymptomatic Infection and Transmission of Pertussis in Households: A Systematic Review. Clinical Infectious Diseases, 70(1), 152–161. https://doi.org/10.1093/cid/ciz531

      de Cellès, M. D., & Rohani, P. (2024). Pertussis vaccines, epidemiology and evolution. Nature Reviews Microbiology, 1–14. https://doi.org/10.1038/s41579-024-01064-8

      de Cellès, M. D., Wong, A., Dalby, T., & Rohani, P. (2025). Natural immune boosting biases pertussis infection estimates in seroprevalence studies. Nature Communications, 16(1), 8883. 

      Enøe, C., Georgiadis, M. P., & Johnson, W. O. (2000). Estimation of sensitivity and specificity of diagnostic tests and disease prevalence when the true disease state is unknown.  Preventive Veterinary Medicine, 45(1–2), 61–81.

      Fine, P. E. M., & Clarkson, JacquelineA. (1982). The recurrence of whooping cough: Possible implications for assessment of vaccine efficacy. The Lancet, 319(8273), 666–669.  https://doi.org/10.1016/S0140-6736(82)92214-0 

      Florkowski, C. M. (2008). Sensitivity, specificity, receiver-operating characteristic (ROC) curves and likelihood ratios: Communicating the performance of diagnostic tests. The Clinical Biochemist Reviews, 29(Suppl 1), S83.

      Gill, C. J., Gunning, C. E., MacLeod, W. B., Mwananyanda, L., Thea, D. M., Pieciak, R. C., Kwenda, G., Mupila, Z., & Rohani, P. (2021). Asymptomatic Bordetella pertussis infections in a longitudinal cohort of young African infants and their mothers. eLife, 10, e65663. https://doi.org/10.7554/elife.65663

      Graaf, H. de, Ibrahim, M., Hill, A. R., Gbesemete, D., Vaughan, A. T., Gorringe, A., Preston, A.,  Buisman, A. M., Faust, S. N., Kester, K. E., Berbers, G. A. M., Diavatopoulos, D. A., & Read, R. C. (2020). Controlled Human Infection With Bordetella pertussis Induces Asymptomatic, Immunizing Colonization. Clinical Infectious Diseases: An Official Publication of the Infectious Diseases Society of America, 71(2), 403–411.  https://doi.org/10.1093/cid/ciz840

      Gunning, C. E., Mwananyanda, L., MacLeod, W. B., Mwale, M., Thea, D. M., Pieciak, R. C., Rohani, P., & Gill, C. J. (2020). Implementation and adherence of routine pertussis vaccination (DTP) in a low-resource urban birth cohort. BMJ Open, 10(12), e041198.

      Lee, A. D., Cassiday, P. K., Pawloski, L. C., Tatti, K. M., Martin, M. D., Briere, E. C., Tondella, M. L., Martin, S. W., & Group, C. V. S. (2018). Clinical evaluation and validation of laboratory methods for the diagnosis of Bordetella pertussis infection: Culture, polymerase chain reaction (PCR) and anti-pertussis toxin IgG serology (IgG-PT). PLoS One, 13(4), e0195979.

      Lefrancq, N., Bouchez, V., Fernandes, N., Barkoff, A.-M., Bosch, T., Dalby, T., Åkerlund, T.,  Darenberg, J., Fabianova, K., Vestrheim, D. F., Fry, N. K., González-López, J. J.,  Gullsby, K., Habington, A., He, Q., Litt, D., Martini, H., Piérard, D., Stefanelli, P., … Brisse, S. (2022). Global spatial dynamics and vaccine-induced fitness changes of Bordetella pertussis. Science Translational Medicine, 14(642), eabn3253.  https://doi.org/10.1126/scitranslmed.abn3253 

      Mills, K. H. G. (2001). Immunity to Bordetella pertussis. Microbes and Infection, 3(8), 655–677. https://doi.org/10.1016/s1286-4579(01)01421-6 

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018:  Findings from the PHIRST study. Journal of Infection, 106550. 

      Rohani, P., Earn, D. J., & Grenfell, B. T. (2000). Impact of immunisation on pertussis transmission in England and Wales. The Lancet, 355(9200), 285–286. https://doi.org/10.1016/S0140-6736(99)04482-7 

      Swift, A., Heale, R., & Twycross, A. (2020). What are sensitivity and specificity?  Evidence-Based Nursing, 23(1), 2–4.

      van der Zee, A., Schellekens, J. F., & Mooi, F. R. (2015). Laboratory diagnosis of pertussis.  Clinical Microbiology Reviews, 28(4), 1005–1026.

      Wilk, M. M., Allen, A. C., Misiak, A., Borkner, L., & Mills, K. H. G. (2019). The immunology of Bordetella pertussis infection and vaccination. In Pertussis: Epidemiology, Immunology & Evolution. Oxford University Press.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Fisher et al describes the molecular mechanism underlying how G beta gamma subunits engage with the beta 3 isoform of PLC. The paper used a combination of cryo EM, BRET assays, and biochemical assays of PLC beta activity. A key discovery is that G beta gamma is not sufficient to drive membrane binding by itself, and instead promotes G alpha activation. The work is important, but suffers slightly from some ambiguity in the actual interface that is present in their cryo EM model, as crosslinkers could stabilise a transient and non-native complex. This is somewhat abrogated by the careful mutational analysis, which shows that mutation of any of these three sites does somewhat block PLC beta G beta gamma activation. However, there could be some improvement in the presentation of this data, as well as possible mutant selection. Overall, this paper is a nice complement to the Falzone et al paper, showing the membrane-bound complex of PLCB3 on membranes, with this work building on this work, highlighting the importance this will have in our full understanding of PLC beta activation.

      Thank you for the positive feedback.

      Major concerns:

      My biggest concern is the potential that this interface is artefactual based on the crosslinking strategy utilised. Here are thoughts on how this could be better validated, presented in a more convincing way.

      (1) The authors' main claim is that there is a degree of plasticity of G beta gamma binding to the PLC beta 3 isoform, with three possible binding sites. The main complication of this is, of course, the possibility that the crosslinking stabilises a non-native complex, driven by a mutated cysteine.

      Because of this, any other additional details about this interface are going to be critical for the scientific audience to judge if this is accurate.

      What would greatly help Figure 1 is an evolutionary conservation analysis of the novel Gbg interface in PLC, to see how well this is conserved, and compare this to the conservation of the previously annotated sites. Conservation of these sites on both the G beta gamma and PLC side would help justify this as a native complex.

      This will also help orient the reader to the identity of the mutated residues assayed in Figure 3.

      We agree that crosslinking can capture non-physiologically relevant interfaces. However, because we do not observe any crosslinking between Gβγ and a PLCβ3 variant that retains a cysteine in the X–Y linker or between PLCβ3 and any other cysteine in the Gβγ heterodimer, we believe it is site-specific.

      The question about sequence conservation in the Gβγ–PLCb3 interfaces is interesting and we have included this information in Figures S8 and S9.

      (2) The g beta gamma orientation is also different than what I have observed in previous g beta gamma effector structures. Is there any precedent for this as an effector interface? A supplemental figure comparing this structure to other g beta gamma interfaces from other enzymes, for example recent Tesmer structure with PI3K.

      We agree that the orientation of Gβ in the crosslinked structure is different. We include a comparison of this reconstruction to other published Gβγ–effector complexes as Figure S6.

      (3) The mutational analysis in Figure 2D-G seems to give some strange results, and I have some question why certain residues were chosen rather than others. Mutation of the Gbg side will be more complicated, as of course that can affect any of the three surfaces. My main question is that, from the way Figure 2A is oriented, the main salt bridge in their novel interface to me looks like R199-D228, with K183 being in the wrong orientation to E226, and D167 being far from any charged residues. Why did the authors not make the corresponding R199 to D or E mutation?

      Thank you for pointing this out, and we expanded our analysis to include this residue. The R199A and R199E mutations had no defects in basal or Ga<sub>q</sub>-stimulated activities. However, R199A had 3-fold lower activation by Gβγ, while R199E was not activated in this assay. The R199E mutation also had significantly decreased agonist-dependent BRET with Gβγ and decreased PI(4,5)P2 hydrolysis, while retaining robust recruitment to the plasma membrane by Gα<sub>q</sub>. These data is included in the main text and Figures 2-4.

      (4) To help the reader's interpretation of Figure 2A, I would recommend a supplemental figure showing the density for interfacial residues, as that also would increase confidence in the interface.

      Thank for the suggestion. In revised Figure S3, we show the Gβγ–PLCb3 D892-PH<sub>cys</sub> complexes determined in this study at different contour levels.

      Reviewer #2 (Public review):

      In this manuscript, the authors dissect how Gβγ potentiates PLCβ3 signaling in cells. Using engineered crosslinking to stabilize a Gβγ-PLCβ3 complex, single particle cryo-EM, and cell-based functional assays, they identify and map multiple putative Gβγ interaction surfaces on PLCβ3, including a previously unrecognized binding mode. Structure-guided mutagenesis supports the functional relevance of these interactions and suggests that Gβγ potentiation is not primarily mediated by PLCβ3 membrane recruitment, but instead enhances PLCβ3 activity after the lipase is already at the membrane.

      Previous reconstitution work on the membrane surface (Falzone & MacKinnon, 2023) proposed a recruitment/partitioning-centric model in which Gβγ increases PLCβ3 output largely by elevating its membrane surface concentration, whereas Gαq primarily increases catalytic turnover; under those reconstitution conditions, the two inputs can combine approximately multiplicatively. In receptor-driven cellular signaling, however, PLCβ3 is robustly recruited to the plasma membrane upon Gαq activation, which raises the question of whether Gβγ contributes mainly through additional recruitment or through a post-recruitment mechanism once PLCβ3 is already at the membrane.

      This manuscript helps address that gap by using membrane-anchored PLCβ3 and complementary cellular readouts to separate "getting PLCβ3 to the membrane" from "boosting activity once PLCβ3 is already there." Their results argue that, in cells, membrane recruitment is largely dominated by Gαq·GTP, while Gβγ can further potentiate PIP2 hydrolysis after membrane association, consistent with a modulatory role at the membrane rather than primary recruitment.

      Overall, the work provides a structural and mechanistic framework for Gβγ-PLCβ3 cooperation and helps clarify the basis of Gq pathway amplification. The manuscript is generally strong, but some issues need to be addressed.

      Thank you for the positive comments.

      Major comments:

      (1) BMOE/BM(PEG)2 crosslinking may enforce a non-native docking geometry, potentially compromising the physiological relevance and precision of the Gβγ-PLCβ3 interface as described. Although a >50% 1:1 crosslinked complex is formed and remains active, the solution maps show lower local resolution for Gβγ, consistent with a dynamic, potentially heterogeneous, interface. One interface is captured via a single engineered cysteine pair (PLCβ3 E60C-Gβ C271), which could potentially bias the pose. It would be helpful if the authors could provide additional orthogonal support (e.g., alternative crosslinked sites) and bolster the clarification of its uniqueness and relevance.

      We did attempt to isolate other crosslinked complexes. PLCβ3-D892 self-crosslinked under all reaction conditions, while PLCβ3-D892 XY<sub>Cys</sub>, which retains an endogenous cysteine within the X–Y linker (C516), did not result in any crosslinked product when incubated with Gβγ. Only the PLCβ3-D892 E60C crosslinked to Gβγ. With the exception the C68S mutation at the C-terminus of Gg to eliminate its prenylation site, all endogenous cysteines were retained in both Gβ and Gγ. Indeed, Gβ contains two solvent-exposed cysteines in its canonical effector binding surface (C204 and C271), but we did not observe any crosslinker density involving C204. While we cannot exclude the possibility that crosslinking occurred between PLCβ3-D892 E60C and other residues in Gβγ, we were unable to identify any 2D classes corresponding to these alternative conformations. These observations, together with the high efficiency of crosslinking, are consistent with a stable and persistent interaction.

      (2) In the crosslinked structure, the authors report that GβD228 interacts with PLCβ3 R199 and K183. In Figure 2A, R199 appears closer to Gβ D228 than K183, yet only K183 is functionally tested. Testing R199 (e.g., R199E/R199A) would strengthen the structure-guided validation of this interface.

      We agree, and functional analysis of PLCb3 R199E is included in the revised manuscript (see Figures 2-4).

      (3) The mutagenesis strategy appears inconsistent across figures/assays, which makes it difficult to interpret phenotypes and directly link the functional data to the proposed interfaces. For example, in Figure 2E, we see R185L but R215E, while residue L40 is mutated to Gly in the IP accumulation assays but to Glu/Lys (L40E/K) in the BRET assays (Figures 3B/3D/3F). The authors should (i) clearly justify the rationale for each substitution (conservative vs charge-reversal, interface disruption, etc.) and (ii), where possible, test the same mutants across assays (or provide evidence that alternative substitutions yield consistent conclusions).

      Mutagenesis experiments were initially carried out independently in the Lambert and Lyon Labs. As the study progressed, additional mutants were identified and/or designed based on results from both groups. The residues subject to mutagenesis are overall consistent across the different assays, with differences in the identity of the mutation varying in some cases. The L40G mutation is one such example, where given its modest impact on Gβγ-mediated activation in the IP accumulation assay, more impactful changes were made (L40E and L40K) for the BRET and signaling assays. In the revision, we now state that mutations were designed to maximally disrupt the three observed interfaces, such as by changing the size of the side chain and/or introducing charge reversal mutants.

      Reviewer #3 (Public review):

      Summary:

      PLCβ3 is activated by both Gαq and Gβγ subunits. This paper follows previous solutions and cryoEM studies of PLCβ3 / Gβγ, trying to understand the molecular details of activation using cellular BRET assays and cryoEM.

      Strengths:

      The authors find evidence for multiple binding sites on PLCβ3 for Gβγ and suggest that Gβγ is not bone fide activator per se but enhances Gαq activation by positioning the catalytic site towards substrate, although this is not completely convincing. Although these sites may not naturally be operative, the authors might want to develop the potential role of these sites.

      The authors also find that this activation is not through recruitment of the enzyme to the membrane by Gβγ released upon G protein activation, in accord with other PLCβ enzymes, but not for PLCβ3, and again, the authors might want to develop this point further.

      Thank you for the suggestions. We are investigating whether the other PLCb isoforms contain multiple Gβγ binding sites and the relative importance of preactivation by Ga<sub>q</sub> for a manuscript in preparation.

      Weaknesses:

      (1) I'm confused as to why the authors feel that their mechanism is distinct from the two-state enzyme, the synergistic activation proposed by Ross in 2011, using a primarily thermodynamic argument. As written, the authors appear to be very reliant on structural and BRET studies that do not give the details that would disprove this interpretation. The main issue is that the author's mechanism does not fully explain how Gβγ activation occurs for PLCβ2 in reconstituted systems in the absence of Gαq subunits.

      The reconstitution experiments are under extremely artificial conditions, using nM-µM of purified proteins and liposomes that contain up to 30% PI(4,5)P2. Under these conditions, we think the increased activity is due to interfacial activation promoted by Gβγ binding to the lipase once it is associated with the liposome surface. This would be sufficient to account for the dose-dependent increase in both PLCb2 and PLCb3 activity as a function of Gβγ concentration. Given the higher basal activity of PLCβ2 and its decreased sensitivity to activation by Ga<sub>q</sub>, one possible explanation is that this isoform differs in its autoinhibition and/or structure of its proximal CTD that Ha2’ displacement is not a prerequisite for activation. In addition, Gβγ may also be a direct activator of PLCβ2. Further studies, ideally in cell-based systems, are needed to answer these questions.

      (2) In a recent study, McKinnon presents a model showing that Gαq and Gβγ activate PLCβ3 by two distinct pathways and that activation by Gβγ occurs through membrane recruitment. It is not surprising that the authors find that this is not true since the pelleting method used by McKinnon is subject to error. The authors should directly address the limitations of this previous work and the changes in proteoliposomes with sedimentation that alter partition coefficients. Although the inability of Gβγ to drive membrane binding is in accord with the quantitative studies of Scarlata, showing that the affinity of PLCβ3 to Gβγ is fairly weak as compared to the intrinsic membrane partition coefficient.

      We have added some of the limitations of proteoliposome sedimentation experiments to the discussion.

      (3) It was proposed many years ago that in signaling complexes Gαq - Gβγ may not have to fully dissociate when binding PLCβ, but rather shift their relative orientation when binding to PLCβ to allow activation. Is their model consistent with this? Is it possible that PLCβ3 keeps Gβγ from diffusing to enhance the rate of Gq / Gβγ re-association?

      Our crosslinked complex is compatible with simultaneous binding of a Gα<sub>q</sub>-Gβγ heterotrimer to the PLCb3, without disrupting the observed interface. If Gαq were to interact with the Gβγ molecules bound to the PH or EF hands, the interaction would be mediated by the N-terminal helix of Gα<sub>q</sub>. It is possible Gβγ–PLCβ3 interactions may slow heterotrimer reassociation, but this may be complicated by the intrinsic GAP activity of the lipase.

      (4) The authors find that Gβγ binds multiple sites, and it is clear that the PH domain site is the primary one in accord with previous work. Could these weaker sites be an artifact of the elevated concentrations used in cryoEM and BRET assays?

      While more studies have focused on the PH domain as a Gβγ binding site, our data does confirm the EF hands are also functionally relevant. To our knowledge, the role of the EF hands has not been investigated in this capacity until very recently, and so we hesitate to label them primary or secondary. It is possible the EF hands may be a lower-affinity site for Gβγ and the protein concentrations needed in cryo-EM drive complex formation. However, it is also possible the concentration of free Gβγ adjacent to an activated receptor may be high enough to saturate the PH and EF hand binding sites.

      (5) Although their assays infer differences in binding affinities, it would strengthen the paper if the authors could estimate the association energies of these different binding sites. This estimation would also address the concern stated above.

      We appreciate this suggestion and quantifying the affinities of the Gβγ–PLCβ3 interactions is the subject of future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please correct PIP2 to the correct PIP2 (many examples throughout).

      These have been corrected.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) Figure S1B: The lane-condition labels above the third gel appear to be incorrect, as both lanes are marked identically (Gβγ +/+, PLCβ variant +/+, BMOE +/+) despite clearly different banding patterns. Please confirm.

      We have confirmed the markings above the gels are correct.

      (2) In Figure S2, the authors show three fitted models, but the helical density for Gβγ cannot be seen in two of them at the displayed contour level. The authors should provide views of the map at different contour levels (thresholds) to better support the model fitting. Otherwise, it is difficult to assess whether the Gβγ subunit could adopt alternative orientations (i.e., whether it may be rotated) within the density.

      We have included a new figure (Figure S3) that provides images of the maps at different contour levels.

      (3) Page 5: "where PLCβ3 is increased by the overexpression of either Gβγ or Gαq" should be revised to "where PLCβ3 activity is ...".

      This sentence has been corrected.

      (4) Figure 2: Please label residue R215 in Figure 2A/2B (or the relevant structural panel), since R215E is tested in 2E but the position is not shown.

      R215 is now included in Figure 2C.

      (5) Page 19, Figure 2 legend: "Changes ... Figure S3" should be "Changes ... Figure S5".

      We have corrected this figure call.

      Reviewer #3 (Recommendations for the authors):

      The studies seem well carried out, although more details regarding the BRET controls and the significance of the values should be included.

      We have revised the captions to provide more details about the experimental controls and a brief description of significance. Individual p-values are included in the supplemental tables.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      I thank the authors for the revised manuscript and for the detailed responses.

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

      Thank you for your critical review and insightful comments.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid, and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei support the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      (b) Published work links PLK to cell division, FAZ elongation, etc... The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc....

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      - The authors have now addressed this question by assessing what % of KING is phosphorylated at T301 and adding commentary on this point in the revised paper.

      - I would suggest that the model (new figure 8) include a dephosphorylation step, as that is proposed by the authors in the text. Also include in the legend some commentary on the role of phosphorylation, which is the center point of this paper, but not currently mentioned.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function. Thank you.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      - The authors have addressed this question by demonstrating that T301 phosphorylation is reduced upon treatment with a PLK inhibitor, thus supporting that PLK phosphorylated T301 in vivo. It is noted that one might consider an alternate kinase is also able to phosphorylate T301 in absence of PLK activity, as that could explain the relatively low (~27%) reduction in phosphorylation by PLK inhibitor treatment.

      Thank you.

      Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      The authors have addressed prior weaknesses in the manuscript through additional experimentation and rewording of the conclusions.

      Thank you for your critical review of our manuscript and for the very constructive comments and suggestions to improve the manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (There is some redundancy below with my comments in the public review, but I've included here for clarity and further explanation.)

      The authors have addressed my primary concern, as treatment with PLK inhibitor reduces phosphorylation of T301, while also providing some comment on relative impact of PLK-mediated KIN-G phosphorylation.

      It is notable that phosphorylation of T301 was reduced by only ~27%, while phosphorylation of S569 was reduced by ~100% in the presence of PLK inhibitor. The authors note that this might be explained by slower dephosphorylation of T301. In the absence of a phenotype, and with cell doubling continuing unabated in presence of the inhibitor, it is intriguing that more loss is not observed. An alternative explanation is that an alternate kinase might also be able to phosphorylate T301 in the absence of PLK activity, and the authors should consider that possibility.

      We added a sentence in the main text to suggest an alternative explanation.

      The model shown in figure 8 should include a dephosphorylation step, per the authors comments in the text regarding the small fraction of T301 that is phosphorylated and proposal of a phosphorylation/dephosphorylation cycle. The Fig 8 legend needs to have some commentary on the role of phosphorylation, as phosphorylation is the center point of this paper.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function.

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Fig 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Fig 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces motility of KIN-G. These statements are contradictory.

      (3) p. 8, and Fig 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      (4) Fig 7B. Please explain labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      (5) Fig 4, 5, and 7: "% Cells" is reported. Please indicate what number of cells total were examined.

      These minor comments have already been addressed in the previous revision.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Marconcini et al. report results of an ambitious study on the genetic mechanisms that contribute to resistance of Drosophila flies to the toxin octanoic acid (OA). This study was motivated by two observations: first, Drosophila sechellia, a close relative of D. melanogaster, has evolved specialized feeding on fruits of Morinda citrifolia, which contain high concentrations of OA and second, that artificial selection on Drosophila simulans, a sister species of D. melanogaster, can generate higher resistance to OA. Previous studies had performed genetic mapping studies between D. simulans and D. sechellia that implicated certain genomic regions in resistance to OA and, in particular, implicated several Osiris gene paralogs as contributing to resistance, though the molecular mechanisms of resistance remain unclear. In this study, Marconcini et al. performed two major experiments. First, they performed evolution-and-resequence on Drosophila simulans populations exposed to OA for 50 generations and identified candidate regions with excessive shifts in allele frequencies as candidate regions containing OA resistance genes in D. simulans. Second, they performed a CRISPR knock-out screen in a D. melanogaster cell line to identify genes that contribute to OA resistance and susceptibility.

      Evolve-and-resequence yielded many candidate genomic regions with extreme allele frequency shifts, which may be regions containing OA resistance genes, or linked genes, or regions that happen to show a strong shift in all replicate populations by chance. As the authors note, detecting significant shifts in allele frequencies is a challenging problem, and the authors use two measures of allele frequency shifts (the Cochran-Mantel-Haenszel method and Bait-ER) and perform simulations under neutrality to estimate a reasonable significance threshold. I am not entirely convinced by this method of estimating significance levels, because the simulations involve assumptions that may not be met by the real populations. I would think that a permutation test would provide an assumption-free method of estimating significance levels. I have tried to think whether there is something about the design of these experiments that would preclude the use of permutation tests (which are used widely for genome-wide studies, such as QTL), but I can't think of one. Perhaps the authors are aware of a reason permutation tests would be invalid here, and if so, they should state this reason.

      Significance thresholds have been estimated using a variety of approaches in the literature, including false discovery rate (FDR) control, simulations under neutral drift with predefined cut-offs, arbitrary significance thresholds, haplotype-based analyses, permutation tests, amongst other methods. Our choice was guided by a review (doi:10.1186/s13059-019-1770-8), which compared the performance of several of these approaches and found that the relatively simple assumptions underlying the Cochran–Mantel–Haenszel (CMH) test often performed as well as, or better than, more complex alternatives. As with most significance thresholds used in genome-wide analyses, the CMH threshold applied here is ultimately based on a degree of arbitrariness, although we aimed to be as stringent as possible. As a complementary method, we used Bait-ER; here the threshold used followed the recommendations provided in the original publication describing this method (doi:10.1111/jeb.14134).

      One reason that permutation-based approaches may be less widely adopted than theoretical null models is the extensive linkage disequilibrium among SNPs, which results in large blocks of correlated variants that cannot be considered independent observations (and therefore not shuffled). In principle, permutations could be performed at the haplotype-block level; however, defining haplotype blocks itself requires selecting thresholds or criteria that are often user defined or based on significance cut-offs. Although several tools are available for this purpose, our experience has been that the resulting block definitions remain sensitive to these choices and therefore introduce a comparable degree of arbitrariness. In reality, there is not yet a standard method in the field that has emerged as the “best” one.

      There is overlap between regions detected by the two methods, but the methods disagree for many regions. The authors state that a "majority of prominent peaks were found by both methods," but I am unclear on what "prominent" means here. It would be more helpful to be more quantitative about the extent of overlap.

      To quantify the agreement between the two methods, we calculated the overlap in genomic coverage (base pairs) between CMH and Baiter candidate regions. For G25, the two methods shared 1.90 Mb of candidate regions. The overlap encompassed 65.3% of the genomic span identified by CMH and 64.1% of the span identified by Bait-ER. For G50, the methods shared 4.77 Mb of candidate regions. The overlap encompassed 63.6% of the CMH candidate span and 99.9% of the Bait-ER candidate span. We have now replaced the admittedly qualitative statement the reviewer highlighted ("majority of prominent peaks were found by both methods”) with this information.

      The authors hypothesized that the response would be at similar genomic loci in all populations (line 222). It seems at least possible that epistatic interactions would lead to different combinations of alleles evolving in each population. I wonder if it would be possible to test whether there is heterogeneity in the responses across the replicate populations.

      We agree that epistatic interactions could lead to different allelic combinations being favored in different replicate populations. However, both BaitER and the CMH test are designed to detect parallel evolutionary responses across replicates, which was the focus of our study. One approach for testing population-specific responses is the LRT-2 test (doi:10.1534/genetics.118.301824). We did not pursue this analysis because it would likely generate many additional candidate loci, making interpretation more challenging, while providing limited additional insight into the repeatable genomic responses that were the primary focus of this work.

      The evolve-and-resequence method yielded many possible regions contributing to OA resistance in D. simulans, but perhaps too many regions to test directly or even to build sensible hypotheses about the genes involved. Thus, the authors performed a second experiment to try to narrow down the list of possible candidate genes. They performed a CRISPR knockout screen in a D. melanogaster cell line for genes that contribute to resistance or susceptibility to OA. The authors identify several limitations of this experiment, but they nonetheless identified several genes where knockouts contribute to OA susceptibility or resistance. Intersecting top hits with regions that experienced selection identified two "resistance" genes: kraken and Alkbh7. The selection hit at kraken is quite compelling, whereas the evidence at Alkbh7 is less strong because only two SNPs were marginally significant. Further functional assays, including gene knockouts in D. melanogaster and D. sechellia, provide some support for the claim that both of these genes can contribute to resistance to OA in flies.

      Beyond the few issues raised above, I do not have significant questions about methodology or the results. I do think, however, that the authors should be more conservative about the implications and significance of their results. For example, on line 139, the authors claim that this intersection approach provides a "powerful paradigm to investigate ecotoxicology." I am not sure I agree that the identification of two genes that may contribute to OA resistance, after a seemingly heroic selection experiment and CRISPR screen, suggests that this method is all that powerful. It seems that most of the genes that contribute to the selection response remain unidentified.

      We agree with the reviewer and have modified the sentence accordingly. While we believe that integrating the approaches discussed in this paper can provide valuable insights into the genetic basis of ecotoxicological traits, these approaches are not a panacea for traits with highly complex genetic architectures, such as OA resistance. Nevertheless, the identification of two candidate genes with some evidence of contributing to the trait represents a meaningful advance toward understanding its underlying genetic basis.

      Finally, given that one motivation of this project was to identify genes that contribute to evolved resistance to OA, I am surprised that the authors did not generate CRISPR alleles of kraken and Alkbh7 in D. simulans and then use these together with the existing alleles in D. sechellia to perform reciprocal hemizygosity tests to determine if these two genes actually contribute to evolved resistance in D. sechellia. This test is simpler to perform and may be more sensitive than the allelic replacement that the authors propose (lines 446-449).

      While generating null alleles for the candidate genes in D. simulans is beyond the scope of this revision, we note that we did attempt reciprocal hemizygosity tests using the mutants available in D. melanogaster. However, the resulting hybrids were recovered in low numbers, and these animals were rather weak, making them unsuitable for the severe OA exposure conditions employed in our assays. We agree that, where feasible, future studies should incorporate reciprocal hemizygosity tests, as they represent a powerful approach for validating the contribution of candidate genes to the trait of interest. We have revised the final sentence of the Results section accordingly.

      Reviewer #2 (Public review):

      Summary:

      The authors studied the resistance against octanoic acid, a compound of the noni fruit, in D. simulans, using experimental evolution and resistance/susceptibility in D. melanogaster cells. They identified novel candidate genes and performed functional tests.

      Strengths:

      The idea of using experimental evolution of a non-resistant species to develop resistance is interesting, and the idea of narrowing down a large list of candidate loci by CRISPR-based gene knockout in cell culture is innovative. The reviewer also liked the (easy) follow-up experiments to validate the results.

      Weaknesses:

      The reviewer is not convinced of the conceptual idea behind their approach: the intersection of the two approaches implicitly assumes that null alleles (or at least compromised alleles) should be selected during experimental evolution. The reviewer considers this unlikely, and the authors made no attempt to test this implicit hypothesis in their data.

      We respectfully disagree with the reviewer’s interpretation of the conceptual idea behind our approach. Our strategy did not assume that experimental evolution selects for null alleles, but rather for any type of variant that could contribute to the trait being selected for (i.e., increases in OA resistance), pointing to candidate genes contributing to the trait. Like many evolve-and-resequence experiments, this approach identified hundreds of candidate genes. This is why we took an orthogonal, genome-wide CRISPR screening approach, where loss-of-function mutations could lead to increases or decreases in OA tolerance of cultured cells. However, we stress that the naturally selected alleles – which could be gain or loss of function – may have much subtler phenotypic effects than the null alleles used for functional validation.

      Along the same lines, it is not clear how to reconcile an upregulation of candidate genes in resistant flies with the knockout experiments.

      We respectfully disagree that these findings are difficult to reconcile. The observed upregulation of the candidate genes in the selected, resistant D. simulans is consistent with a role in promoting OA resistance, while the knockout experiments (whether in cultured cells or in whole animals) demonstrate that loss of gene function reduces resistance. These observations are complementary: increased expression is associated with enhanced resistance, whereas complete loss of function impairs it. The knockout experiments of kraken and Alkbh7 were intended as functional validation of gene involvement and do not imply that the alleles selected during experimental evolution are loss-of-function alleles (we rather hypothesize that the selected alleles are gain-of-function through some, as yet undetermined, mechanism).

      The experiments to validate the effect of candidate genes did not match the experimental evolution conditions.

      This is correct. As our results suggest that the phenotype is shaped by multiple genes, such that the effect of any individual gene is likely modest compared to their combined contribution. Consequently, detecting and validating the effect of a single gene requires more stringent OA conditions than those needed to observe the overall phenotypic response over the course of several generations. In addition, the shorter-term plate assay was more practical for higher temporal resolution of the analysis of mortality in the presence of OA.

      The statistical analysis suffers from some problems and an insufficient description of the analyses performed.

      Although D. simulans GWAS data are available, the authors did not make an attempt to estimate the effect of selected variants in the candidate genes in the GWAS data set.

      We agree that comparing the experimental evolution and GWAS results (from our previous work, doi:10.1093/g3journal/jkag032) is of considerable interest. We have now expanded the Discussion to explicitly discuss the relationship between the two datasets. Overall, the overlap between the approaches was limited, although two GWAS candidate genes, bez and CG13003, fall within genomic regions exhibiting significant CMH signals at generation 25 (but not generation 50) of the evolve-and-resequence experiment. (We note that these genes could not have been identified in our CRISPR screen, as they are not expressed in S2R+ cells). More broadly, differences between the GWAS and evolve-and-resequence results likely reflect the distinct evolutionary processes captured by each approach: GWAS maps standing phenotypic variation among isofemale lines, whereas experimental evolution tracks allele frequency changes under sustained selection. Understanding why some signals are shared whereas others are not – whether due to effect size, genetic background, epistasis, pleiotropic costs, or the contribution of initially rare variants – remain important open questions.

      The reviewer would have liked to see more connection between the experimental evolution and the GWAS data. As some D. simulans genotypes have similar resistance to D. sechellia, it would have been interesting to test whether this genotype contributed to the observed resistance.

      While D. simulans genotypes displayed a range of OA resistance levels, none approached D. sechellia levels of resistance (see Figure 2e from our previous work, doi:10.1093/g3journal/jkag032) (The reviewer might have conflated the data from our GWAS of D. melanogaster strain, shown in Figure 2b of that paper, where some lines of that species exhibit comparable resistance to D. sechellia under the conditions of that assay). Regardless, our experimental evolution data do not provide sufficient resolution to identify the specific favorable alleles underlying the response to selection. Instead, we detect genomic regions containing many linked variants whose frequencies change under selection. Consequently, a direct comparison between evolved genotypes and GWAS-associated genotypes is currently difficult. We note, however, that expression of the D. sechellia kraken allele in D. melanogaster did not produce a significant effect on resistance, suggesting that even the most promising candidate alleles might not have strong effects in isolation.

      At several places, the authors discuss the challenge of studying a polygenic trait, but at the same time, they claim to have detected and validated candidate genes. It would be helpful if the authors could discuss why they consider that their assays could really detect the contribution of single loci to the polygenic trait. In particular, when GWAS did not detect their candidate genes.

      Our results do not imply that kraken and Alkbh7 are major-effect loci or that they explain a substantial proportion of the phenotypic variation. Rather, our data indicate that these genes make measurable contributions to OA resistance, consistent with the expectation that complex traits are influenced by many loci of individually modest effect. The absence of these genes among the top GWAS candidates does not preclude their involvement, as GWAS and experimental evolution interrogate different aspects of the genetic architecture and differ in their power to detect loci of varying effect sizes and allele frequencies.

      It is not clear to the reviewer why the authors did not pay more attention to the highly significant peaks emerging from the experimental evolution study. Their functional validation would have been biologically more plausible.

      We agree that the significant peaks identified in the evolve-andre sequence experiment represent promising targets for future investigation. However, these peaks typically span large genomic regions containing tens to hundreds of genes, making it difficult to prioritize individual candidates based on the experimental evolution data alone. In this work, we focused our functional validation on genes independently supported by the CRISPR screen, which provided gene-level resolution. We fully acknowledge that additional causal genes are likely to reside within the selected regions and remain to be functionally characterized.

      Impact:

      Given the obvious challenges of functional testing of polygenic traits and the clear limitations of the interpretation of the results, the study will be helpful for future studies aiming to characterize polygenic traits. Unfortunately, the results are just another piece of controversial results regarding resistance against octanoic acid, a trait that is rather easy to evaluate.

      The reviewer appears to imply that a trait being straightforward to phenotype necessarily implies that its genetic basis should also be straightforward to resolve. Many classic complex traits, such as human height, are simple to measure yet have an extraordinarily complex, highly polygenic genetic architecture. We believe that OA resistance represents a similar challenge: while the phenotype is readily assayed, differences in assay conditions, genetic backgrounds, and the contribution of many loci of individually modest effect make its genetic basis difficult to dissect. We have strived to be cautious in our conclusions, in particular the evolutionary interpretations; nevertheless, to our knowledge, this is the first study to provide functional evidence supporting the contribution of specific genes to OA resistance in D. sechellia, combining both loss-of-function phenotypes and expression data. As emphasized by the title of our manuscript, we view the principal contribution of this work as demonstrating how complementary experimental approaches (both of which are fairly novel for study of toxin susceptibility/resistance genetics) can be integrated to prioritize and functionally evaluate candidate genes underlying complex adaptive traits.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers propose to include additional statistical and quantitative measures to strengthen the results. Furthermore, certain parts of the manuscript could be rephrased to make sure that the readers can understand more clearly the implications and significance of the results. To test whether the genes identified in the present manuscript (as contributing to octanoic acid resistance) are also involved in the evolution of the resistance, reciprocal hemizygosity tests using CRISPR alleles of kraken and Alkbh7 in D. simulans would be a plus, but such experiments are not required because the main focus of the paper is on the genetic basis of octanoic acid resistance, and not on evolution.

      We thank the Reviewing Editor for these constructive comments. We have revised the manuscript to clarify the interpretation and significance of our findings and have addressed the reviewers’ comments throughout. Regarding reciprocal hemizygosity tests, we agree that they would provide a valuable means of assessing the evolutionary contribution of candidate genes. However, generating the necessary reagents in D. simulans represents a substantial undertaking beyond the scope of the present study, particularly given the expected modest effects of individual loci underlying this highly polygenic trait. We have nevertheless revised the end of the Results section to mention reciprocal hemizygosity tests as an important direction for future work.

      Reviewer #2 (Recommendations for the authors):

      (1) Provide more details about the selection tests: which sequences were used? Please report p-values. It would also be important to discuss the possibility of false positives caused by a bottleneck in D. sechellia. A genome-wide analysis could help to see if the bottleneck increased the signal of positive selection.

      We are not entirely sure what additional analyses are being suggested. The sequences and methods used for the selection analyses are described in the Methods, and the statistical support for these analyses is reported in the Supplementary Material (MK test p-values and FUBAR posterior probabilities). We are also unclear as to how a genome-wide analysis would address the interpretation of the gene-specific selection analyses presented here. If the reviewer intended a different analysis, we would appreciate further clarification.

      (2) The significance level of the CMH test needs to be determined with the effective population size, not with the census size, as done by the authors. This is important, as it is not clear if the candidate genes remain significant after significance adjustment based on the effective population size.

      We thank the reviewer for this comment. In our analyses, effective population sizes were explicitly incorporated by Bait-ER to model the effects of genetic drift. The CMH significance thresholds, however, were obtained from neutral forward simulations following the recommended workflow for the method, which requires census population sizes rather than effective population sizes as input. To make these simulations as realistic as possible, we therefore used the observed census population sizes at each generation of the experimental evolution. Estimating generation-specific effective population sizes for use in an alternative simulation framework would require substantially more temporal data than are available here (e.g., sequencing many additional time points) and is beyond the scope of the present study.

      (3) Include the allele frequency trajectory across time for the candidate genes.

      We thank the reviewer for this suggestion. However, plotting allele frequency trajectories for the candidate genes is not straightforward because the evolve-and-resequence analysis identified broad linked genomic regions rather than individual causal variants. Each candidate region contains numerous SNPs spanning several kilobases, with each SNP exhibiting its own allele frequency trajectory across the 10 replicate populations. Consequently, there is no single representative trajectory for a given candidate gene, and plotting all SNPs within each region would be difficult to interpret. For kraken, however, we identified seven candidate regulatory SNPs and have plotted their individual allele frequency trajectories, which we include in Author response image 1 to illustrate the diverse patterns observed:

      Author response image 1.

      (4) Estimate the effect of candidate loci in the GWAS data (independent of significance).

      We thank the reviewer for this suggestion. However, we are not entirely sure what analysis is being proposed. In particular, it is unclear which variants the reviewer is referring to, as our candidate genes are associated with multiple linked variants rather than a single causal SNP. Consequently, we are unsure how the effect of a candidate locus should be estimated in the GWAS dataset. Moreover, there is no guarantee that the same variants are represented in both datasets. For example, variants filtered out during the GWAS because of low allele frequency may subsequently have increased in frequency during experimental evolution and therefore contributed to the evolve-and-resequence signals. We would appreciate further clarification of the analysis the reviewer has in mind.

      (5) Discuss the challenge of false positives (see: 10.1016/j.cub.2020.12.023).

      We thank the reviewer for this suggestion. We agree that false positives (as well as false negatives) are an important consideration when studying complex polygenic traits, both through the initial “screening” efforts (e.g., GWAS, experimental evolution) and follow-up functional validation (e.g., RNAi, mutant, overexpression analyses). We believe this issue is already addressed in the manuscript through our discussion of the limitations of the individual approaches, the polygenic nature of OA resistance, and our cautious interpretation of the functional validation results. Our conclusions are limited to identifying candidate genes that contribute to OA resistance, rather than claiming to have identified all of the loci or variants underlying the evolution of this trait. We are therefore not sure what additional discussion the reviewer has in mind and would appreciate further clarification if a specific point from the cited study is intended.

      (6) The figures with the Manhattan plots should be improved to indicate the overlapping genes in the Manhattan plots, rather than in the circle figures below. By using different colors this should be quite easy and clean.

      We thank the reviewer for this suggestion. However, we believe Author response image 2 more appropriately illustrate the overlap between the CMH and BaitER analyses. Although the Manhattan plots display individual SNPs, our candidate loci are defined by broader genomic regions comprising blocks of linked significant SNPs rather than by single variants. Simply highlighting SNPs within overlapping regions would therefore add visual complexity without providing additional biological insight beyond that already captured by the circos plots. As an illustration, we provide here an example of the generation 50 Manhattan plot with overlapping SNPs highlighted in red:

      Author response image 2.

      (7) A more focused discussion of the assumption that functional data from D. melanogaster can explain resistance in D. simulans or D. sechellia. At some places, epistatic interactions are mentioned, but the reviewer feels that the entire screen is based on the idea that similar effects are found across species, hence, this needs to be adequately reflected in the discussion.

      We thank the reviewer for raising this important point. We would like to emphasize that our study was not based on the assumption that functional effects identified in D. melanogaster or D. simulans necessarily explain the evolution of OA resistance in D. sechellia. Rather, our motivation stemmed from the longstanding difficulty of identifying individual genes underlying this highly polygenic trait using mapping approaches alone. We therefore sought to combine orthogonal experimental approaches to identify genes contributing to OA susceptibility (of D. melanogaster and D. simulans) and resistance (D. sechellia). We hoped, but did not assume, that genes supported by multiple independent lines of evidence would be informative for understanding the natural evolution of OA resistance in D. sechellia. Indeed, we acknowledge in the manuscript that the cell-based CRISPR screen, experimental evolution, and the natural evolution of D. sechellia occurred under very different selective contexts and timescales. As reflected in the title of our manuscript, our conclusions are intentionally framed around the identification of novel toxin resistance loci, while remaining cautious about their evolutionary interpretation.

      (8) Discuss that the controls in the RNAi test were quite variable. Could this reflect some problems with the assay?

      We do not believe that the variability among the control lines reflects a problem with the assay. Rather, the different controls (Gal4, UAS-RNAi etc.) represent distinct genetic backgrounds, each of which may exhibit a different baseline level of OA resistance. Indeed, in our recent GWAS of OA resistance (doi:10.1093/g3journal/jkag032), we observed substantial natural variation in OA resistance among D. melanogaster and D. simulans lines. Importantly, each RNAi line was compared with its corresponding genetic background control, so differences among control lines do not affect the interpretation of the individual RNAi experiments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines Müller glia (MG) reprogramming in the uninjured mouse retina through a combination of Notch signaling inhibition and AAV-induced proliferation. Building on their prior work showing that Cyclin D1 overexpression and p27^Kip1^ knockdown (CCA) promotes MG proliferation with very limited neurogenesis, the authors now demonstrate that Rbpj deletion alone induces a modest degree of MG-to-neuron conversion without proliferation, in agreement with recent work in the field. However, combining Rbpj deletion with CCA-mediated proliferation substantially enhances MG dedifferentiation and the generation of retinal neuron-like cells. Through genetic lineage tracing, histological analyses, and single-cell transcriptomics, the authors provide evidence that MG-derived cells acquire molecular features of bipolar (ON, OFF, and rod bipolar) and amacrine neurons. Most MG-derived cells appear to survive long-term (up to 9 months).

      Strengths:

      Overall, the study is carefully designed and executed, and the manuscript is clearly written with well-presented figures. While the work does not significantly expand the repertoire of neuronal types generated from mammalian MG beyond what has been previously reported in the field, it provides a valuable and improved strategy for inducing robust MG proliferation and neurogenesis in the mammalian retina.

      Weaknesses:

      (1) It would be better to include a negative control AAV when evaluating the effect of CCA AAV in the Rbpj KO background. This could help distinguish the specific contribution of the CCA construct from potential effects of intravitreal AAV injection itself, which can induce mild inflammation, known to influence MG reprogramming.

      To address this concern, in the revised manuscript we included the result from Rbpj KO eyes injected with a negative control AAV (AAV<sub>7m8</sub>-GFAP-GFP) (Fig. S14a). MG reprogramming efficiency, quantified as the proportion of tdT<sup>+</sup>Otx2<sup>+</sup> cells among total tdT<sup>+</sup> cells, was then compared between the AAV-GFP–treated and Rbpj KO–only eyes. At 4 months post-injection, the percentage of tdT<sup>+</sup>Otx2<sup>+</sup> cells in the AAV-GFP–treated eyes was comparable to that of Rbpj KO alone (Fig. S14b–c), and substantially lower than in the CCA-treated eyes. Together, these results indicate that the enhanced MG reprogramming observed in the Rbpj KO+CCA group is driven by transgenes expressed rather than by nonspecific effects of AAV or injection.

      (2) The extent of MG transduction by the CCA AAV is not clear. As quantifications are normalized to total MG (GFP^+^ or TdTomato^+^) or retinal length, it would be useful to clarify whether near-complete transduction is assumed, or if additional information on transduction efficiency can be provided.

      In our previous study (Wu, Liao, et al., 2025, eLife), we have demonstrated that high-dose (4E10vg/injection) AAV7m8 effectively transduced the whole retina, with near-complete MG transduction observed in the vicinity of the injection site, as evidenced by virtually all MG expressing GFP in these regions. In the revised manuscript, we clarified the transduction efficiency in Line 108-110 on Page 5 and Line 625-626 on Page 27.

      (3) In Figure S10, the reduced MG proliferation observed in the CCA + Rbpj deletion group could also potentially reflect decreased GFAP promoter activity in dedifferentiated MG following Rbpj deletion. Alternatively, MG-derived cells may be more fragile under these conditions.

      We thank the reviewer for these excellent insights. We agree that a down-regulation of GFAP promoter activity following Rbpj-mediated dedifferentiation is a highly plausible explanation for the moderate reduction in proliferation, as lower promoter activity would diminish AAV transgene expression. We have included this possibility in the data interpretation (Line 176-178, page 8). Regarding the alternative possibility of increased cell fragility, we agree that cell death cannot be ruled out, but occasional apoptotic cells over a long period of time are difficult to capture experimentally.

      (4) In the CCA + Rbpj deletion condition, do MG undergo single or multiple rounds of cell division?

      We have previously demonstrated that MG typically undergo a single round of cell division in wild type mouse retina following CCA treatment (Wu, Liao, et al., 2025, eLife). Given our observation that Rbpj deletion suppresses CCA-induced MG proliferation (Fig. S11), it is unlikely that the addition of Rbpj deletion would trigger multiple or continuous rounds of cell division beyond the single-round baseline established by CCA alone. While we did not re-evaluate cell division kinetics in the current study, we reason that CCA similarly drives MG to undergo a single round of division in the Rbpj KO context.

      (5) What fraction of neuron-like cells (bipolar- and amacrine-like) arises from proliferation versus direct transdifferentiation? Quantification of MG-derived cells expressing neuronal markers (e.g., Otx2, HuC/D), with and without EdU labeling, would help distinguish these mechanisms.

      The percentages of MG-derived cells expressing neuronal markers with and without EdU labeling, were shown in Fig 3d-e and Fig S19d-e. In the Rbpj KO-only group, neuron-like cells arise exclusively through direct transdifferentiation without cell division, as no EdU incorporation was detected in Rbpj-deficient MG. In this group, a small fraction of MG-derived cells expressed the neuronal marker Otx2 or HuC/D (Fig 3e, Fig S19e). In contrast, the Rbpj KO+CCA group achieved a substantially higher neurogenesis rate, with a significant proportion of Otx2+ or HuC/D+ MG-derived cells also being EdU+ (Fig. 3d, Fig. S19d), indicating that they arose through de novo neurogenesis. By subtracting the contribution of direct transdifferentiation observed in the Rbpj KO-only group, we estimate that majority of MG-derived neuron-like cells in the Rbpj KO+CCA group were generated through proliferation-mediated de novo neurogenesis.

      (6) In Figure S18a, the authors state that "while the neuron-like clusters were best classified as BC-like and AC-like based on their distinct marker gene expression, they also exhibited mixed expression of genes associated with other retinal neuronal types, including RGC markers (e.g., Tubb3, Myt1l, Grin1) and photoreceptor markers (e.g., Crx, Prom1, Epha10, Gucy2e, Scg3) (Fig. S18a), suggesting that the regenerated cells exist in a hybrid state" and "MG derived neuron like cells also expressed genes characteristic of RGCs and photoreceptors, indicating enhanced lineage". However, many of these genes are not specific to RGCs or photoreceptors and are instead broadly expressed in retinal neurons or enriched in bipolar/amacrine populations. Therefore, it is unclear whether these cells exhibit hybrid RGC or photoreceptor identity.

      We thank the reviewer for this insightful comment and for pointing out the need for greater precision in our terminology regarding these markers. While individual markers may lack absolute, 100% cell-type exclusivity, genes such as Tubb3 and Gucy2e serve as widely accepted lineage-associated genes that characterize RGC and photoreceptor programs, respectively (Soto et al., 2008; Sato et al., 2018; Sotani et al., 2024). We have revised the manuscript to replace terms "RGC-specific genes" and "photoreceptor-specific genes" with "RGC signature genes" and "photoreceptor signature genes", respectively. Furthermore, these RGC- and photoreceptor-signature genes are co-expressed across the entire Otx2+ MG population rather than being segregated into distinct, specialized subpopulations (Fig. 4d, Fig. S20). This uniform distribution indicates that these cells possess a hybrid transcriptional program that concurrently incorporates elements of both RGC and photoreceptor identities.

      (7) The authors provide a thorough molecular characterization of MG-derived cells through immunostaining and single-cell sequencing. However, their morphological features, synaptic connectivity (e.g., synaptic marker expression), and electrophysiological properties remain largely uncharacterized. While these experiments may be technically challenging, this limitation should be discussed.

      We agree with the reviewer that characterizing the precise morphological features, synaptic connectivity, and electrophysiological properties of MG-derived cells is a crucial step for any neuronal regeneration study, and we acknowledge that this represents an important limitation of our current study.

      As demonstrated by snRNA-seq data, the MG-derived neuron-like cells exhibit an incompletely mature state, characterized by hybrid transcriptomic signatures. By immunostaining, we did not observe any MG-derived cells with photoreceptor outer segment or typical RGC morphology. Therefore, it is highly likely that these cells have not established functional synaptic connectivity or acquired mature electrophysiological properties. Performing functional or circuitry assessments at this stage would be premature.

      We have added a comprehensive discussion regarding this limitation, along with future directions for long-term functional validation, in the revised manuscript (Line 502-515 on Page 22).

      (8) The conclusion that CCA + Rbpj deletion induces neurogenesis without compromising MG supportive functions or retinal homeostasis appears somewhat oversold. This claim is primarily based on gross retinal morphology and ZO-1 staining. Given the extent of MG dedifferentiation and ectopic cell generation in the ONL and INL, it is likely that retinal function is affected. Functional assessments (e.g., ERG) would be required to support this conclusion. The authors should consider tempering this statement.

      To address the concern raised by the reviewer, we performed electroretinography (ERG) to evaluate both scotopic and photopic retinal function in the Rbpj KO+CCA-treated eyes compared to contralateral untreated controls (Supplementary figure S23e-h). In addition, we conducted optomotor response testing to assess whether visual behavior is affected following treatment (Supplementary Figure S23d). The results demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal function.

      (9) Regarding the mechanism by which CCA-induced proliferation enhances MG reprogramming in the Rbpj knockout background, one plausible explanation is that chromatin states (e.g., histone modifications and DNA methylation) are transiently reset during DNA replication and cell division. While this alone may be insufficient to activate neurogenic programs, it could synergize with Rbpj deletion to allow neurogenic transcription factors (such as Ascl1, Otx2, NeuroD1, and NeuroD2) to access previously inaccessible chromatin regions, thereby promoting MG reprogramming.

      We thank the reviewer for the insightful suggestion on the model, which aligns well with our experimental findings. Our snATAC-seq data demonstrate that CCA-induced proliferation broadly increases chromatin accessibility at key neurogenic loci, including Neurod2, Dll1, and Otx2, in active MG compared to resting MG (Figure 6f–h). This chromatin remodeling alone is insufficient to drive neurogenesis, as CCA-only treated MG largely revert to a quiescent glial state. However, when combined with Rbpj deletion, which derepresses downstream neurogenic transcription factors such as Ascl1 and Neurog2 by relieving Notch-mediated transcriptional repression, these newly accessible chromatin regions can be effectively occupied and activated by the available neurogenic factors. The concept that cell division facilitates epigenetic resetting to enhance reprogramming efficiency is well established in the somatic cell reprogramming field, where proliferation rate is directly proportional to reprogramming success by promoting the erasure of lineage-restrictive epigenetic marks and the re-establishment of new transcriptional circuits. In the revised manuscript, we incorporated this mechanistic discussion to provide a more comprehensive interpretation of how proliferation and Notch inhibition converge to promote MG neurogenesis in Line 462-477 on Page 20-21.

      Reviewer #2 (Public review):

      Summary:

      The inability of the mammalian retina to regenerate poses a major clinical challenge. Much has been learned about the regenerative potential of the retina from teleost fish, where Müller glia (MG) are able to proliferate and produce new neurons after injury. However, MG do not retain this potential in the mammalian retina. The authors showed previously that forcing MG to re-enter the cell cycle by downregulating p27 and upregulating cyclin D1 could induce MG to dedifferentiate, but the results were transient, and these cells eventually reverted back to MG and did not form neurons. Here, they expand on this to show that in MG, coupling forced cell cycle re-entry with deletion of Rbpj, which inhibits the transcriptional effects of Notch signaling, induces some MG to proliferate and take on features of multiple cell types, including MG precursor cells, amacrine-like cells, and bipolar-like cells. This work lends valuable insight into the regenerative potential of mammalian MG, particularly when Notch signaling is manipulated.

      Strengths:

      The major claims of the authors are well-supported. They show convincingly - and through multiple methods including immunostaining, single-nucleus RNA sequencing, and in situ hybridization - that coupling notch inhibition with cell cycle reactivation induces the expression of neuronal markers in mammalian MG. The snRNA-seq data are particularly valuable in demonstrating the induction of bipolar-cell subtypes. Edu labeling is effective in demonstrating the induction of proliferation, and the long-term viability of the generated neuron-like cells is intriguing.

      Weaknesses:

      Whether the newly generated neurons are functionally integrated remains unclear, and the effect of the manipulation on the function of the retina was not tested. Imaging data suggests that many of the newly generated neurons persist for months, but often appear mislocalized. It is also not clear if the manipulation of MG affects long-term MG function. Cell death was not evaluated, and although the authors evaluated the long-term effect on tight junctions, this data was not quantified, and further analysis on morphology or function was not done. Control eyes were untreated, not vehicle-injected.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The transgenic line may be Glast-CreERT, not Glast-CreERT2.

      We appreciate the reviewer for bringing this to our attention. The formal allele symbol for this transgenic line is Tg(Slc1a3-cre/ERT)1Nat, while this strain is generically classified as "Cre/ERT2" by the Jackson Laboratory. Some published studies referred to this line as Glast-CreERT and others as Glast-CreERT2. To maintain consistency with the formal allele symbol, we have adopted "Glast-CreERT" throughout the revised manuscript.

      (2) For snATAC data in Figure 6 e,f, and Figure 19b. It is most likely gene activity, not gene expression, since these are snATAC, not snRNA data.

      For this inaccurate terminology, we have corrected all relevant figure labels and associated text in the revised manuscript to clearly state "gene activity" instead of "gene expression."

      (3) Some text in Figure 6 is a bit too small to read.

      We have increased the font size of the text elements in Figure 6 to ensure readability and have also reviewed all other figures for consistency. Revised figures with improved legibility have been included in the updated manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) There are multiple instances where further elaboration of methods or tools in the test would improve readability and comprehension by a broader audience. It would be helpful to (early, often, and clearly) explain precisely which cell types are labeled in your mouse line and how. Someone unfamiliar with the mouse line may struggle to understand what is labeled by the tdT or GFP. Likewise, it would help to consistently define what cell types are labeled by tdT+ vs. Sox9+, tdT+, etc.

      We have added a clear and detailed description of the mouse lines and labeling strategy early in the Results section, specifying which cell types are labeled by tdT and GFP and how the labeling is achieved. We have also ensured that the definitions of cell type identifiers (e.g., tdT<sup>+</sup> for MG-derived cells, Sox9<sup>+</sup>/tdT<sup>+</sup> for MG remaining in a glial state) are consistently stated upon first use and maintained throughout the manuscript to improve readability for a broader audience. In addition, we added headings for the quantification graphs to improve readability in all quantification figures.

      (2) It is unclear what the difference is between Figure 1c and S1c, and these should be quantified as the % of positive cells, as described in the text.

      We have removed Figure S1c and moved Figure 1c to supplementary figure 1. The MG labeled by EdU and Sox9 or Otx2 were quantified as % of the EdU+ MG.

      (3) S2e: Clarify what pixel level means, is this pixel intensity?

      Yes, "pixel level" in Figure S2e refers to pixel intensity. We apologize for the ambiguous wording and replaced "pixel level" with "pixel intensity" in the revised figure legend to ensure clarity.

      (4) Figure 2: In the magnified image of the GFP+, Sox9- cell, the GFP is also very faint. Could these cells be dying? Analysis of the expression profile of these cells (or ruling out apoptosis) would better support a dedifferentiation argument.

      The faint GFP signal observed in GFP<sup>+</sup> Sox9<sup>-</sup> cells is a sign of ongoing dedifferentiation rather than cell death. This is likely due to chromatin remodeling during reprogramming. A similar decrease in reporter signal intensity during MG dedifferentiation has been previously reported by Le et al. 2024, 2025, supporting the interpretation that reduced fluorescence is a characteristic feature of this process. It is possible that a small fraction of GFP<sup>+</sup> Sox9<sup>-</sup> cells may undergo cell death over an extended period, which would be difficult to detect using apoptosis assays. Our long-term survival experiments demonstrate that more than 80% of MG-derived neuron-like cells survive for at least 9 months following treatment (Figure 7), indicating that majority of these cells are viable. The discussion is included in line 112-114 on page 5.

      (5) Figure 3: The Crx labeling appears everywhere except the identified cell. This seems the opposite of the point you are making.

      Crx signal of the MG-derived cell (tdT<sup>+</sup> Crx<sup>+</sup>), which is pointed out by arrowhead, is in a ring-like pattern. This pattern is consistent with the euchromatin region in inverted nucleus of rod. Crx labeling appears in other cells in the image as Crx is highly expressed in native photoreceptors.

      (6) I think it would be nice to address, in the discussion, the apparent disorganization and mislocalization of cells in the long-term images.

      We thank the reviewer for highlighting this critical observation. During retinal development, precise laminar positioning of neurons is guided by a coordinated interplay of cell-intrinsic transcriptional programs and extrinsic cues including cell adhesion molecules, guidance factors, and interactions with neighboring cells. In the adult retina, many of these developmental cues are no longer present or active, which likely contributes to the failure of MG-derived neurons to migrate to their appropriate laminar positions. Interestingly, the vast majority of our divided MG cells remained localized within the outer nuclear layer (ONL). Because the ONL is the physiological location of photoreceptors, this preferential position could serve as an advantageous baseline layout for driving targeted photoreceptor differentiation in future work. To address reviewer’s feedback, we have expanded our discussion section (Line 543-559, page 23-34) to cover the mechanisms underlying this structural disorganization and its downstream implications for functional circuit integration.

      (7) I'm not convinced that ZO1 alone is sufficient to suggest MG function normally or that retinal homeostasis is maintained. I suggest tempering that conclusion in the text.

      For the revision, we have performed additional experiments to address this concern. The optical coherence tomography (OCT) images revealed that retinal layer organization and ONL thickness were comparable among the uninjected eyes, GFP AAV-injected control eyes, and CCA-treated eyes, demonstrating that overall retinal architecture was well-preserved (Fig. S23a–c). Optomotor response testing revealed no significant differences in visual acuity across groups, suggesting that visual function remained intact (Fig. S23d). Furthermore, electroretinography (ERG) demonstrated that scotopic and photopic a- and b-wave amplitudes were unaffected by the treatment, confirming that light responses from photoreceptor and inner retinal neuron were preserved (Fig. S23e–h). Taken together, these findings demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal structure and functional visual circuitry.

    1. Author response:

      The authors thank the reviewers for their thorough and fair assessment of our manuscript. We are currently working to edit the manuscript based on the critiques and guidance offered by the reviewers. This will consist of fixing grammar and typos, expanding the material and methods section to include more information on the behavioral assays, modifying graphs for clarity between visuals and interpretations, and correcting our mistakes in neuroanatomical labeling.

      Public Reviews:

      Reviewer #1 (Public review):

      In their submitted manuscript, Harkinish-Murray and colleagues from the Kozol lab present convincing evidence for a genetically encoded shift in the odor perception of cavefish compared to their surface ancestors. Surface Astyanax, just as zebrafish, are attracted to food odors and are repelled by death odors and the alarm substance Schreckstoff (released from damaged skin by specialized club cells). Based on the experimental evidence in this manuscript, however, their cavefish counterparts are attracted to these odors as well. This would make sense, in an evolutionary framework, as predation is less likely in cave settings and decaying fish are a valuable source of nutrients for their living counterparts.

      Using an F2 hybrid cross scheme between surface fish and cavefish, authors also provide compelling evidence that genetic factors are behind this behavioral shift. Furthermore, they also show that this behavior (i.e., attraction to skin and decay extracts) can be observed in surface fish given long enough food deprivation. This latter observation also makes sense in the light of evolution and is genuinely interesting as it also provides a plausible roadmap to the shift in behavior through Waddingtonian genetic assimilation.

      The manuscript is generally well written and clear, we have identified only few weaknesses, some regarding the presentation of the data.

      (1) For Figure 3, on the x-axis of panels b, e, and h, supposedly we see surface fish vs. different cavefish populations. This is currently missing and makes the figure harder to interpret. Also, two populations (panel e) show a bimodal distribution upon indirect white light exposure, suggesting that some fish still acted as if they were exposed to direct light, while others acted as if they were in darkness (infrared light). We believe this warrants more consideration as it could tell us something about the existing (and relevant) genetic variance within this population. It is also notable that the third cavefish population also showed increased odor indices under indirect white light and infrared light conditions, suggesting that increasing the number of observations could have yielded a statistically significant result.

      We agree with the reviewer that our light testing data suggests complexity in the response to indirect white light within certain cave populations. In addition, an expanded sample size would likely provide clarity on whether individuals fall within two groups, behavior that looks like direct light or infra-red light, that could relate to genetic variation within cavefish populations. We are currently working to reassess the current data and determining the best course of action for continued studies related to light exposure.

      (2) Some extra details about the methods could also be provided to enhance the reproducibility of the experiments.

      We agree with both reviewers that the methodological section on behavior needs to be expanded. We are currently editing our methods section to include more detail on water exchanges, odor preparation, timing, biological replicates, and binning.

      (3) A more serious concern is about the anatomical designation of particular brain regions in Figure 7d and consequently Figure 7f. Whereas we would agree with the positioning of the medial pallium (Dm), we think the region depicting the thalamus is in fact still part of the telencephalon, and the real thalamus should be more posteriorly. On the other hand, we think that the preoptic areas should be under the pallium and not posterior to it (see PMID: 22586363 for corresponding zebrafish anatomy). We would suggest, therefore, that the authors revisit this issue (a minor one, considering the depth of the results presented in the manuscript), and provide a better anatomical annotation - e.g., the identity of particular brain regions could be backed up by Hybridization Chain Reaction experiments for region-specific transcripts. (Disclaimer: we do not consider ourselves experts in adult cavefish neuroanatomy; therefore, we consulted in this case a colleague with much more knowledge on this topic.)

      We agree that our annotation was incorrect or more accurately mislabeled in our write-up of the preprint and submitted manuscript. Therefore, we have now re-assessed the regions using the tissue cleared and light sheet collected zebrafish atlas, Adult Zebrafish Brain Atlas (AZBA; doi: 10.7554/eLife.69988). We are now editing the resubmission in the following manner: our initial labeling of the ventromedial thalamus will be changed to the lateral olfactory tract (nLOT) of the pallium and the preoptic region to the ventromedial thalamus (VM). We will provide a comparable z-slice of the AZBA segmentation file to illustrate the similarity in position. This would also support a known continuous circuit of olfactory integration, with information flowing from the lateral olfactory tract-to the piriform cortex-to the thalamus. We also agree that a more accurate assessment in Astyanax would require HCR in situ hybridization of markers for those specific brain regions or a neurocomputational brain atlas for adult Astyanax populations. Finally, we assert that this small dataset is preliminary at best and only provides regions of shared activity that could explain anything from perception related processes to relay of odor signaling unrelated to perception. Further work with larger sample sizes and additional populations are currently underway for a follow-up study on the neurobiological basis of olfactory processing and perception in adult cavefish.

      (4) It would also be useful to expand the brain imaging data displaying results for similar tests in surface fish, to see if skin and decay extracts trigger different or similar brain activity in those fish.

      We agree with the reviewer that the brain mapping section lacks a sufficient sample size and no control group for comparison (surface fish). However, we found the variation in pERK intensity (notably the putative nLOT) to be informative and decided to include the dataset in the manuscript. We are currently working to fill in these data gaps by sampling all populations and increasing the Pachon cavefish sample size. This will be a follow-up study as mentioned above in the last rebuttal paragraph.

      Further work will surely be able to discern the more precise genetic changes that made the shift in behavior possible. Once these causative variants (or at least linked markers) are determined, it will be quite revealing to see if these variants are indeed already present in the surface population (as hinted by the authors), and also, if besides the Surface x Tinaja F2 hybrids, crosses between other cave populations and surface fish can be performed, we could also see how much evolutionary convergence happened in the parallel evolution of different cave morphs. Were there multiple possible pathways for similar behaviors in different cave populations, or - as in freshwater stickleback populations - do we see broadly the same genetic playbook repeated each time?

      We agree with the reviewer that the hybrid results setup a promising follow up project to map these traits genetically. We are continuing to test odor perception in other hybrid populations and have started Quantitative Trait Locus mapping experiments.

      Another outstanding question, also demonstrated and discussed, albeit briefly, in this paper relates to the behavior-modulating effect of light in cavefish. What is the physiological relevance for a dark-dwelling animal to have this capacity? Is this just the chance result of occasional gene flow from surface populations, or does it have a genuine evolutionary significance?

      Reviewer #2 (Public review):

      Summary:

      The authors tested whether the olfactory cues that drive attraction or avoidance behavior have diverged between surface‑dwelling and cave‑adapted strains of the Mexican cavefish Astyanax mexicanus. They use high‑throughput odor‑discrimination assays between known attractants and repellents by calculating an "odor index" per fish (=the difference in time spent in an odor zone versus a control zone). Further, hybrid crosses to probe heritability, starvation experiments to assess plasticity of odor perception, and whole‑brain pERK detection/mapping to link behavioral changes with known localized neural activity. The results support the hypothesis that the extreme cave environment has selected for an approach response to stimuli that are ancestrally aversive (like alarm or death odors) but in harsh environments can be used as guidance to the rare food sources in this ecosystem.

      Strengths:

      The odor index analysis is convincing, and the experiments for odor attraction/avoidance are robustly performed. The light-to-darkness shift reflected by avoidance to attraction in cavefish towards skin odors is compelling and carefully analyzed. The analysis of odor indices of three cave-dwelling populations in comparison to surface fish highlights a similar regime, yet with differences among the different populations, suggesting population-specific genetic variation.

      Another strength of the paper is exactly this genetic inheritance study by generating F2 hybrids of cave-dwelling and surface-living individuals. The hybrids displayed a continuous range of odor indices for social, alarm, and death odors, indicating that these traits are heritable and likely based on additive genetic markers. Further, the authors uncovered a sexual dimorphism: only female cavefish exhibited approach behavior to social odors, whereas males remained neutral. This result aligns with known differences in olfactory organ morphology between sexes of other species from harsh environments.

      Although limited in number, the neurophysiological correlation using whole‑brain pERK mapping after 10 min of odor exposure is convincing. The data revealed overlapping activation in the thalamus and pre‑optic region for food and decay odors, suggesting that these brain areas mediate the evolved attraction response to previously repellent stimuli.

      Overall, the manuscript presents a concise story: cavefish have evolved attraction to alarm and death odors as a result of shifting from ancestral avoidance-driven to attraction by genetic changes and physiologically similar activation of specific neural circuits. The evidence is robust, with multiple independent experiments (behavioral assays, hybrid genetics, starvation experiments, and brain mapping) that collectively support the conclusions.

      Furthermore, exposure to unpleasant odors can not only be tolerated but can even serve as a trigger for foraging. This plasticity demonstrates that genetic predispositions can be put into practice through active changes in physiology in species or organisms confronted with (drastically) changing environmental conditions.

      Weaknesses:

      I value that the authors are critical of their own data, indicating low numbers in the pERK/brain experiments. Yet this is a weak point as the statistical power is thus limited. However, their reasoning is careful, based on the results and not over-interpreting.

      We agree with the reviewer and direct their attention to the same critique by reviewer 1. We believe this is predominantly preliminary data that was included due to the conspicuous increase in pERK signal from the putative lateral olfactory tract (nLOT) for food and decay exposed cavefish. We are continuing to work on odor stimulated brain mapping and look forward to publishing a comprehensive dataset across wildtype and hybrid populations.

      The layout/design of the ethograms (bout category plots) for both individual and population-wise are not easy to follow. Reworking these display items to convey the information is necessary.

      We agree with the reviewer that the ethograms are challenging to read, especially due to our use of different colors for odor categories. We are currently preparing alternative graphs for displaying ethograms that reduce confusion and make following bout transitions for individual traces and bout probabilities for populations easier on the eyes.

      Taken together, the manuscript uses odor perception and attraction/avoidance behavior studies to show that environmental changes (light-to-darkness) have an immediate impact on smell perception and behavior. Attraction to otherwise repellent odors is used by cavefish to likely adapt to harsh environments with low food sources. The manuscript convincingly demonstrates this plasticity, which is an interesting idea to follow up for other traits spreading among a population. This also underlines that a genome may be fixed and the blueprint for behavioral traits, but extrinsic cues can readily be adapted to change wired behavior even to the extreme as reported here: changing avoidance to attraction.

    1. Author response:

      We thank the editors and reviewers for their thoughtful evaluation and constructive feedback. We are pleased that all three reviewers recognize the importance of mapping Cas9-induced sister chromatid exchanges (SCE) as a previously invisible repair outcome, and that the RDCP analysis is a notable feature of the study.

      We note that since our manuscript was posted, two companion studies in Science have provided direct biochemical evidence for the TRAIP-dependent pathway we discussed:

      (1) Fujisawa & Labib (Science, 2026; DOI: 10.1126/science.aeh2300) showed that TTF2 bridges CDK1-phosphorylated TRAIP to DNA Polymerase epsilon in the replisome, triggering mitotic CMG helicase disassembly, fork cleavage, and repair via SCE. Loss of this pathway reduced replication stress-induced SCE approximately two-fold in mouse ES cells.

      (2) Can et al. (Science, 2026; DOI: 10.1126/science.aeh1834) independently identified the same CDK1-TTF2-TRAIP axis in Xenopus egg extracts and validated it in HCT116 cells, showing that disrupting the TRAIP-TTF2 interaction reduced common fragile site deletions.

      We will incorporate these references in the revised discussion while still framing our RDCP observations as consistent with, rather than definitive proof of, this pathway.

      Below we briefly address the main points raised in the public reviews.

      Reviewer #1:

      We agree that the mechanistic interpretation of the RDCP signature should be presented more cautiously. We will reframe the URR/TRAIP discussion as a model, replacing language such as "direct genetic evidence" with "consistent with." We will add a summary table of RDCP data. We will also expand the description of rescued SCE calls. We will clarify what the DNA repair gene targeting experiment can and cannot answer (delayed protein loss, essential-gene selection) - we think that there is a notable difference at the bulk vs. at the single-cell level depending on the nature of the assay. Fig.1 fonts, labels, and pileup plot descriptions will be improved.

      We agree that the high-SCE subpopulation is particularly interesting and we cannot currently distinguish higher RNP uptake, a permissive cell-cycle state, altered expression, or stochastic variation. This may be better explored by future co-assays with sci-L3-Strand-seq.

      Reviewer #2:

      We agree with the limitations that Cas9-induced DSBs can be dependent on the cell cycle stage and the number of times cuts are made. We will add a brief discussion on this limitation in extrapolating the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements.

      We agree that "a single Cas9 cut" should be revised to "Cas9 targeting of a single genomic locus" to accurately reflect the experimental design. We will clarify the possibility of multiple rounds of cutting at the same sites. We will also clarify "non-local" by modifying Fig.1 - we used this term specifically to include the possibility of inter-sister NHEJ.

      Reviewer #3:

      We will improve the self-contained nature of the manuscript so that readers need not consult the earlier NAR papers, and the companion preprint to understand the key results.

    1. Joint Public Review

      Summary:

      In this study, the authors investigated the developmental and molecular basis of the unusual metamorphic program of the black soldier fly, Hermetia illucens, which differs from the canonical holometabolous life cycle by inserting a distinct, non-feeding prepupal instar between the final larval stage and pupation. Most insects that undergo complete metamorphosis molt to the final instar and then develop into the prepupal stage without molting. H. illucens, however, undergoes a molt before entering a non-feeding prepupal stage. Thus, it is an unusual, novel developmental strategy, and its regulation has remained a mystery. Through an integrated approach combining detailed morphological characterization, developmental gene expression profiling, and RNAi-mediated functional analyses of the core components of the Metamorphic Gene Network (MGN), the authors examine the developmental identity of this prepupal stage and how the temporal deployment of conserved metamorphic regulators has been reorganized to accommodate this atypical developmental program. In particular, they show that the prepupal stage expresses a distinct combination of the key genes known to regulate life history transitions, including unusually high levels of Br-C expression.

      Strengths:

      The study represents a valuable contribution to insect developmental biology. A major strength is the comprehensive characterization of postembryonic development, which establishes a robust developmental framework for H. illucens. This is complemented by detailed expression profiling and RNAi-based functional analyses of the Metamorphic Gene Network (MGN), comprising the temporal specifier factors, Kr-h1, chinmo, Br-C, and E93. The results show that these conserved regulators are deployed in a modified temporal sequence that accommodates the distinctive prepupal stage while largely preserving their canonical developmental functions. Together, the morphological, molecular, and functional data support the conclusion that the prepupal stage of H. illucens is a distinct developmental transition associated with a characteristic configuration of the metamorphic gene network. The results are supported by solid methodology and approaches and will serve as valuable resources for future investigations into insect development, the evolution of metamorphosis, and the diversification of insect life-history strategies.

      Weaknesses:

      While the study successfully establishes the developmental identity of the prepupal stage and its association with a modified temporal deployment of the MGN, some aspects of the proposed regulatory model are less directly supported by the experimental evidence.

      (1) Several regulatory interactions within the MGN remain inferential rather than experimentally demonstrated in H. illucens. In particular, the proposed relationship between juvenile hormone (JH), Kr-h1, and chinmo is based primarily on expression dynamics and RNAi-induced transcriptional changes. Although these observations are consistent with the proposed model, they do not directly demonstrate that JH induces chinmo expression or establish the regulatory relationship between Kr-h1 and chinmo in this species. As a result, the corresponding regulatory interactions presented in the final model should be regarded as plausible hypotheses rather than experimentally validated mechanisms.

      (2) A second limitation concerns the developmental role assigned to Br-C and E93 during the larval-to-prepupal transition. The authors conclude that sustained Br-C expression is a defining molecular feature of the prepupal stage and discuss its potential role in prepupal specification. However, the functional analyses of both Br-C and E93 were initiated only after larvae had already entered the prepupal stage. Consequently, while the RNAi experiments convincingly demonstrate essential roles for Br-C during the prepupal-to-pupal transition and for E93 during adult differentiation, they do not directly address whether either factor is required to trigger the formation of the prepupal stage itself. Therefore, the molecular mechanisms governing the initiation of this distinctive developmental transition remain unresolved. In particular, the proposed lack of repression of E93 by Br-C is only weakly supported, yet may be an essential feature of the prepupal stage of Hermetia illucens.

      (3) Although knockdowns of Kr-h1 and chinmo knockdowns look superficially similar, it would be good to confirm this with higher-magnification views of the cuticles for all three treatments (control, Kr-h1 RNAi, and chinmo RNAi). In other species, Kr-h1 knockdown leads to premature adult cuticle development, whereas chinmo knockdown typically leads to premature appearance of pupal characteristics. Similarly, in Fig. 4A and 4D, higher-magnification images of the cuticle would be helpful.

      (4) (Relating to Line 336 and Figure 7): "This low but persistent prepupal Kr-h1 expression, together with modest chinmo expression from PPD0 to PPD8, may be correlated to a JH-dependent antimetamorphic effect that maintains the prepupal stage." However, we are not aware of a function of JH in extending the prepupal stage. In addition, in most insects, the prepupal stage expresses high Kr-h1 expression; this peak likely prevents the animal from turning into an adult instead of the pupa. We presume the same holds true for H. illucens (although the lower expression of Kr-h1 during that stage is curious). As a result, we suggest that Fig. 7D be revised as it may be difficult to distinguish between pupal formation and prepupal maintenance given the experimental set-up. Fig. 7E may also need to be modified since the development of the pupa may require Kr-h1. It is worth noting that at the prepupal stage, JH and Br-C are co-expressed in many insects. If the authors think that Kr-h1 expression needs to be low at this time, this would imply a novel interaction between Kr-h1 and Br-C, and should be discussed.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public reviews):

      Weaknesses:

      Reliance on self-reports

      A primary limitation of this study, acknowledged by the authors, is its reliance on self-reports of participants’ emotional states. Although considerable effort was made to minimize expectation effects, further research is needed to confirm that the observed behavioral changes reflect genuine alterations in emotional states. Additionally, the generalizability of the findings to long-term remediation strategies remains an open question.

      We agree with this characterisation and have strengthened the corresponding acknowledgment in the Discussion. We would also note that, while self-report measures are inherently subjective, the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ”genuine” emotional state that supersedes self-report currently exists. We have added a sentence to the Discussion making this point explicit: ”While emotional self-reports are inherently subjective, we note that the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ‘genuine’ emotional state that supersedes self-report currently exists.”

      Additionally, we agree that what we have described is limited to a short-term intervention and change. Whether these changes bear on longer-term changes remains to be assessed. Furthermore, the mechanisms or processes that would support such a maintenance are of substantial interest, and will be the focus of future work.

      Statistical analysis and interpretation of the dynamics matrix

      Second, the statistical analysis, particularly the computational approach, sometimes lacks sufficient detail and refinement. While I will not elaborate on specific points here, one notable issue is the interpretation of the intrinsic matrix (A). The model-free analysis reveals correlations between emotions at a given time or within an emotional state across time points. However, it does not provide evidence to support lagged interactions across states that would justify non-diagonal elements in A. The other result concerning the dynamics matrix only highlights a trend in the dominant eigenvalue, which is difficult to interpret in isolation. The absence of a statistically significant group x intervention interaction furthermore makes this finding a little compelling. This weakens the study’s conclusions about the importance of intrinsic dynamics, as claimed in the title.

      We thank the reviewer for raising this important methodological point. We address it in three parts.

      (i) Justification for the full dynamics matrix. In response to this comment, we have added a diagonal-constrained variant of A to the model comparison. The full A model was selected by BIC over the diagonal-constrained version, meaning that cross-emotion lagged interactions contributed to model fit over and above what could be explained by individual emotion autocorrelation alone, thereby justifying the inclusion of off-diagonal elements. We have updated the model comparison results section and figure caption accordingly.

      (ii) Interpretation of the dominant eigenvalue. We agree that the dominant eigenvalue alone is difficult to interpret. Our primary claim regarding intrinsic dynamics rests on model comparison (which identified a change in A as necessary to explain post-intervention data in the distancing group) together with the change in the direction of the dominant eigenvector (tested via Hotelling T<sup>2</sup>, p = 0.019). The eigenvalue magnitude is reported as a complementary, interpretable summary of the stability shift.

      (iii) General analysis of the dynamics matrix. We have added a visualization of the full dynamics matrix A to the Supplementary Materials, reporting per-emotion diagonal elements (persistence) and their variability across participants. This shows that the emotion calm exhibited the highest persistence, consistent with our hypothesis, whereas sadness showed lower-than-expected stability, and other emotions exhibited low mean persistence with substantial individual variability. We have added the following sentence to the Results: ”A general analysis of the dynamics matrix A revealed that while calm exhibited the highest persistence, consistent with our hypothesis (0.59 ± 0.85), sadness did not show the expected stability (0.09 ± 0.58), and other emotions demonstrated low persistence (0.13 to 0.24) alongside substantial individual variability.”

      Terminology (controllability, stability, sensitivity, Gramian)

      Finally, to avoid potential misunderstandings of their work, the authors should be more careful about their use of terms pertaining to the control theory and take the time to properly define them. For example, the ”controllability” of emotional states can either denote that those states are more changeable (control theory definition), or, conversely, more tightly regulated (common interpretation, as used in the abstract). This is true for numerous terms (stability, sensitivity, Gramian, etc.) for which no clear definition nor references are provided. Readers unfamiliar with the framework of control theory will likely be at a loss without more guidance.

      This is an excellent and important point. We have made the following changes throughout the manuscript.

      (i) Terminology table. We have added a new Table 1 in the Methods (Conceptual Definitions) providing a side-by-side mapping of the formal control-theory definition and the psychological interpretation for the three key terms: Controllability, Stability, and Sensitivity.

      (ii) Abstract. The abstract previously used ”controllability” in a way that could be read in the psychological sense. We have replaced the sentence referencing the Controllability Gramian with a clarified version describing the measure as ”continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”

      (iii) Introduction. We have added a paragraph explicitly distinguishing the control-theoretic definition of controllability from the psychological concept of emotion control or emotion regulation, noting that high controllability in the formal sense does not imply tight regulation or reduced variability. Definitions of sensitivity and stability are also provided.

      (iv) Discussion. Occurrences of ”controllability” that could be ambiguous are annotated with brief clarifications of which sense is intended.

      (v) Gramian. We now use the ”controllability matrix” (C) throughout, and its SVD-based analysis is described using ”left singular vectors” rather than ”eigenvectors of the Gramian”.

      Reviewer #2 (Public reviews):

      Online recruitment, selection effects, and remuneration

      Acquiring data online inevitably gives rise to selection and self-selection effects. This needs to be acknowledged clearly. Exacerbating this, participant remuneration seems low at an amount below the minimum or living wage in Western countries (do the authors know where their participants came from?).

      We thank the reviewer for raising this. All participants were recruited from the UK via Prolific Academic and reimbursed at £7.50/h. Remuneration rates were comparable to other experimental settings, in keeping with other online studies, UK living wage recommendations, and ultimately determined according to institutional ethical guidance. We acknowledge that online recruitment via Prolific may introduce self-selection effects and that the sample may not be representative of the general population. We have added a corresponding statement to the Limitations section: ”Online recruitment via Prolific Academic may introduce self-selection effects, and the sample may not be representative of the general population.”

      Intervention ongoing during the second block

      Another concern is that the intervention does not simply take place before the second block begins but is ongoing during the whole of the second block in that it is integrated into the phrasing of the task on each trial. It is therefore somewhat misleading to speak of a period ’after the intervention’, and it would have been interesting to assess the effect of this by including a third group where the phrasing does not change, but the floating leaves intervention takes place.

      This is a valid and important observation. In the distancing group, the trial-by-trial question phrasing during the second block included a reminder, meaning the intervention was reinforced on every trial. We have acknowledged this as a procedural difference and potential confound in the Limitations section, noting that this reminder may have encouraged a form of retrospective reappraisal rather than in-the-moment distancing. The design choice was intentional (the reminder was included to encourage continued application of the strategy, mirroring how such techniques are deployed in practice) but we agree it is an imperfect feature of the current design and have discussed it openly. We have added to Limitations: ”Relatedly, a procedural difference existed in that only the distancing group received a strategy reminder during the second video block. While intended to facilitate real-time regulation, this prompt may have inadvertently encouraged retrospective reappraisal to align with the distancing narrative.”

      Observation noise

      As mentioned in the Limitations section, observation noise was assumed and not estimated. While this is understandable in this case, the effect of this assumption could have been assessed by simulation with varying levels of observation (and process) noise.

      We would like to clarify that both observation noise (Γ) and process noise (Σ) were in fact estimated from the data, constrained to be diagonal. We have extended the parameter recovery analysis (Supplementary Section: Parameter Recovery) in which, for each of 104 subjects and both time periods (N = 208 observations), we generated 100 surrogate trajectories from the fitted model and re-estimated all parameters. Recovery quality is quantified via Pearson correlations between true and recovered parameters (A, C, Σ, and Γ), alignment of dominant eigenvectors and left singular vectors, and bias analysis for dominant eigenvalue and singular value magnitudes. The analysis confirms that A and C matrix parameters and the critical eigenvector-based metrics all exceed our target threshold of r = 0.7. Noise covariances are poorly recovered, consistent with typical Kalman filter behaviour on short time series. We have expanded the Limitations to note that the Gaussian observation noise assumption does not fully capture the bounded nature of 0–100 rating scales, and suggest that future work could use a truncated or censored observation model.

      Reliance on formal model comparison

      Relatedly, the reliance on formal model comparison is unfortunate since the outcome of such comparisons is easily influenced by slight changes to assumptions such as noise levels. An alternative approach would have been to develop a favoured model based on its suitability to address the research question and its ability, established by simulation, to distill relevant changes of behaviour into reliable parameter estimates.

      We appreciate this methodological concern but would argue that formal model comparison is wellsuited to our research question. Our central aim is not simply to fit emotion trajectories, but to determine which components of the dynamical system (intrinsic dynamics (A), input weights (C), or both) are altered by the distancing intervention. This is inherently a model comparison question: without comparing models that do and do not allow each component to change, no principled inference about the mechanism of action is possible. A single favoured model, however well-motivated, would presuppose the answer.

      We also note that the reviewers concern, that outcomes are sensitive to underlying assumptions, applies equally to the favoured-model approach: simulations rely on predefined structures and noise specifications that shape parameter recovery, and a misspecified favoured model risks confounding the parameters intended to capture the intervention effect with artefacts of model structure. By contrast, our approach evaluates a principled, nested set of models of increasing complexity, using BIC to penalise unnecessary parameters, which guards against overfitting. The models are intentionally simple (linear, Gaussian, time-invariant) to limit the degrees of freedom available to absorb noise.

      Critically, the approach is not model comparison alone. We followed established best-practice procedures for computational modelling, including posterior predictive checks (simulated trajectories closely matched observed data; Fig. 4C), and parameter recovery establishing that the key metrics are reliably recovered (see Supplementary Materials F Parameter Recovery). We have also added a Random-Effects Bayesian Model Selection (RFX-BMS) to characterise individual heterogeneity in model preferences. Together, these provide converging evidence for the validity of our inference. We have clarified this reasoning in the revised manuscript.

      Statistical limitations; Bayesian inference

      The statistical analyses clearly show the limitations of classical statistical testing with highly complex models of the kind the authors (commendably) use. Hunting for statistically significant interactions in a multivariate repeated-measures design relying on inputs from time series- derived point estimates is a difficult proposition. While the authors make the best of the bad 3 situation they create by using null-hypothesis significance testing, a more promising approach would have been to estimate parameters using a sampler like Stan or PyMC and then draw conclusions based on posterior predictive simulations.

      We agree that fully Bayesian parameter estimation via Stan or PyMC would be a valuable methodological advance. Implementing this for 104 subjects across two time periods with a 5-dimensional Kalman filter is, however, a substantial undertaking beyond the scope of the current revision. In the interim, the RFX-BMS analysis added in response to Comment 4 above provides a group-level Bayesian perspective on model uncertainty that partially addresses this concern.

      Reviewer #3 (Public reviews):

      Dual meanings of controllability

      An interesting but perhaps at present slightly confusing aspect of their described results relates to the ’controllability’ of emotions, which they define as their susceptibility to external inputs. Readers should note this definition is (as I understand it) quite distinct from, and sometimes even orthogonal to, concepts of emotional control in the emotion literature, which refer to intentional control of emotions (by emotion regulation strategies such as distancing). The authors also use this second meaning in the discussion. Because of the centrality of control/controllability (in both meanings) to this paper, at present it is key for readers to bear these dual meanings in mind for juxtaposed results that distancing ”reduces controllability” while causing ”enhanced emotional control”

      We are grateful for this observation, which echoes Reviewer 1’s concern about terminology. We have addressed this comprehensively; please see our response to Reviewer 1 Comment 3 above.

      Strategy reminder and possible reappraisal

      As above the authors use an active control – a relaxation intervention – which is extremely closely matched with their active intervention (and a major strength). However, there was an additional difference between the groups (as I currently understand it): ”in the group allocated to the distancing intervention, the phrasing of the question about their feelings in the second video block reminded participants about the intervention, stating: ”You observed your emotions and let them pass like the leaves floating by on the stream.” I do wonder if the effects of distancing also have been partially driven by some degree of reappraisal (considered a separate emotion regulation strategy) since this reminder might have evoked retrospective changes in ratings.

      This is a well-founded concern. As noted in our response to Reviewer 2 Comment 2, we have added an explicit acknowledgment of this procedural difference and the potential for retrospective reappraisal to the Limitations section. We note, as the reviewer themselves observe in the Strengths section, that demand effects are unlikely to account for the specific pattern of dynamic changes observed: uniform demand effects would be expected to produce flat reductions across emotions, whereas our findings show emotion-specific changes in eigenmode structure and controllability direction. Nevertheless, a partial contribution of reappraisal cannot be ruled out from the current design.

      Mechanism of distancing effects (eye movement and oculomotor avoidance)

      Not necessarily a weakness, but an unanswered question is exactly how distancing is producing these effects. As the authors point out, there is a possibility that eye-movement avoidance of the more emotionally salient aspects of scenes could be changing participants’ exposure to the emotions somewhat. Not discussed by the authors, but possibly relevant, is the literature on differences between emotion types on oculomotor avoidance, which could have contributed to differential effects on different emotions.

      We thank the reviewer for raising the oculomotor avoidance hypothesis. Research suggests that different emotions elicit distinct patterns of gaze behaviour: disgust is associated with visual avoidance, whereas anxiety and other negative emotions show increased attentional bias following fear conditioning. These emotion-specific oculomotor patterns could have contributed to the differential effects we observe on the input weight matrix C. What would be particularly interesting to examine in future work is whether a distancing intervention induces multiple, emotionally-specific gaze behaviours, or a single undifferentiated avoidance response. We have expanded the Limitations to: ”[...] The literature on emotion-specific oculomotor avoidance suggests that gaze patterns differ across emotion categories, which could contribute to differential effects on the input weight matrix for specific emotions. [...]”

      Recommendations for the Authors:

      Reviewer #1 (Recommendations for the Authors):

      (1) In the procedure description, the authors suggest that some emotions (e.g. disgust) would be more volatile and stimulus driven, while others (eg. sad) would be more stable. Is this hypothesis reflected in the model-based state dynamics, typically in the diagonal elements of A?

      Yes, we added a supplementary figure showing the dynamics matrix A, which confirms that calm exhibited the highest persistence, consistent with our hypothesis. Disgust showed lower persistence, also in line with this expectation. However, sadness did not display the expected stability and instead showed relatively low persistence, contrary to our hypothesis.

      (2) Could the authors elaborate on the emotional space covered by the chosen ratings? If the axes are positive-negative and slow-fast, why 5 and not 4?

      We added a sentence in the Methods clarifying that five emotions were selected based on the specific affective qualities of the video stimuli provided in the validated databaseto and to better capture the high-dimensional, nuanced states elicited by the stimuli rather than to fit a traditional four-axis model.

      (3) The whole methods section crucially lacks references. As an example, the whole derivation of the most important metrics (eigenvalues of the Gramian, energy ellipse, etc) leaves the reader completely on its own.

      References have been added throughout the Methods, including for the controllability matrix, SVDbased analysis, and eigendecomposition.

      (4) Before equation 1, when introducing x and u: a) time appears twice (typo), and b) the 1Tˆ notation is not standard (especially without bold) and unclear until way below when the one-hot encoding is mentioned.

      The typo has been corrected.

      (5) Why use one hot-encoding rather than the original ratings from the video database? Videos must vary if not in spread (as suggested in Figure B.1) at least in intensity. Ignoring this variance surely diminishes the accuracy of the modelling.

      We used one-hot encoding so that the input weight matrix C can directly estimate participantspecific intensity and sensitivity, rather than fixing input magnitudes to database averages. Using database ratings would assume that the emotional intensity of each video clip generalizes perfectly to our sample; any mismatch would be absorbed as error in C, potentially biasing parameter estimates and obscuring individual differences in emotional reactivity, which are central to our analysis.

      (6) Could the authors develop the rationale behind the bias in Equation 1?

      We added a sentence clarifying that the bias term h captures the steady-state baseline of the emotional system, i.e. the mean rating toward which emotions converge in the absence of external inputs.

      (7) While I can understand why the authors included a set of models with a diagonal C matrix, I do not see why they did not do the same with A. While the diagonal elements are necessary to persist emotional states and induce some autocorrelation in the ratings, as observed empirically, the influence of the non-diagonal elements is not justified (and Figure 5G seems to confirm that). This is critical as, in the end, the controllability metrics will highly depend on those non-diagonal elements which remain very obscure throughout the manuscript.

      We now included both diagonal and full variants of A in the model comparison; still the full A was selected by BIC.

      (8) Concerning the model comparison, I am not sure what the authors mean by using the BIC at the group level. Did they just sum them across participants? This approach is known to be highly susceptible to outliers and cannot be relied upon in general. So-called ”random effect analysis” tends to be regarded as the gold standard and can be easily implemented by taking - 0.5 * BIC as an approximation to the model evidence. Such an approach would also allow to properly test for group differences (cf. Rigoux et al. 2014).

      We have added a Random-Effects Bayesian Model Selection analysis as a new Supplementary Section; see Comment 4 of the public review response above.

      (9) The sentence ”proportion of the total amount of predictive power provided by the full set of models contained in the model being assessed” does not make any sense to me. Please rephrase.

      The sentence has been rephrased: ”Cumulative model weights (w<sub>j</sub>) normalize raw BIC scores so they can be interpreted as the relative probability that a specific model is the best one among the set being compared:”

      (10) ”the largest eigenvalue of the dynamics matrix A identifies the most stable combination of emotions” is only true if the eigenvalues are below 1, which is not granted.

      We have added the qualifier that this holds provided the dominant eigenvalue lies within the unit circle (|λ| < 1), indicating a system that converges to a steady state.

      (11) Equation 3 does not define the Gramian but the controllability matrix, a confusion that goes through the manuscript. The Gramian is formally defined as W = P(AkBB′A′k). Luckily, for discrete systems, it can be approximated by CC′ and therefore the singular values of C can be used to approximate the eigenvalues of W, which are the usual metrics used to define the energy ellipse and so on. While the results reported in the manuscript are correct (up the the approximation), the general description is wrong or misleading.

      We thank the reviewer for this important correction; we now consistently refer to C as the ”controllability matrix” throughout, and its SVD-based analysis is described in terms of left singular vectors rather than eigenvectors of the Gramian.

      (12) Figure 2: what is the matrix V? If it’s from the singular value decomposition W = USV, then (if I am not mistaken) the direction of the ellipsoid is defined by U. Again, a reference would help.

      We have corrected the figure and caption to refer to ”left singular vectors” throughout. We retain the variable name V rather than adopting the standard SVD convention of U to avoid confusion with the input vector u, which appears throughout the model equations.

      <(13) Correction for multiple comparisons is mentioned as a way to correct for the number of conducted tests. However, later on, some post-hoc analyses are reported with the mention that the correction is done across emotions (so p/5), while multiple tests are run for each emotion. This is critical when all the pairs across the cells of an ANOVA are tested and no correction seems to be applied, which is inducing a huge risk of false positives.

      Along the same line: the correct way to demonstrate the effect of the intervention is to first do an ANOVA to reveal an interaction between group and time, and then only to do post-hoc tests to pinpoint where the interaction is coming from, and not the other way around as reported in the manuscript. Further, a difference in significance is not equivalent to a significant difference, and showing that a time effect is significant in one group but not in the other does not imply that the intervention differs between groups, only testing the interaction can confirm this.

      We have added reporting of the significant group × time interaction effects (F(5,208) = 2.6, p = 0.026 for mean ratings; F(5,200) = 2.5, p = 0.03 for the most controllable direction) and flagged these in the figure captions. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      (14) The notation DV = b0 + b1IV*b2G is confusing as a full model (interaction + main effects) should have 3 parameters in addition to the intercept. Also, why use different models for testing the main effect and the interaction?

      The regression equation has been corrected to DV = β<sub>0</sub> + β<sub>1</sub>IV + β<sub>2</sub>G + β<sub>3</sub>(IV × G) + ϵ, making the interaction term explicit.

      (15) Figure 4: the control subject in panel C seems to rate close to 0 in all emotions except for ”calm”. How was this subject fitted? Does model selection (at the subject level) correctly identify a change of dynamics in this case? I don’t see how the behaviour after the intervention could be realistically fitted with a 65-parameters dynamical system. What type of checks were operated to ensure the quality of the fit beyond the recovery analysis (see below)?

      The top participant in Figure 4C was from the distancing group and the participant rating close to zero on most emotions except calm shown at the bottom was from the control group. For the distancing participant, model selection correctly identified a change in dynamics and input weights (BIC = 4806) over the same-parameters model (BIC = 4821). For the control participant, model selection similarly favoured a change in dynamics and input weights (BIC = 3723) over the sameparameters model (BIC = 4011), suggesting that the relaxation intervention produced a comparable effect on emotional dynamics to the distancing intervention. Note that these two participants were selected randomly to illustrate the visual quality of model fit (i.e. that simulated trajectories closely resemble the empirical rating curves) and are not intended to be representative examples of group differences.

      Regarding the data-to-parameter ratio: the 65 parameters are estimated from 55 observations per emotion per block, giving a more favourable ratio than it might appear.

      Beyond the visual trajectory overlays in Figure 4C, we have now added R<sup>2</sup> and peak cross-correlation as a quantitative measure of individual fit quality. Mean R<sup>2</sup> across all subjects and emotions was 0.6 and mean temporal correlation was r = 0.74−0.80, confirming that both the timing and magnitude of emotional responses were well reproduced. Notably, for the specific control participant shown in Figure 4C, R<sup>2</sup> values were 0.74, 0.83, 0.81, 0.83, and 0.83 for disgusted, amused, calm, anxious, and sad respectively, confirming that even for this visually striking participant the model fit was adequate across all five emotion dimensions.

      (16) What do the authors mean by ”eigenmodes”? In the following sentence, what does ”This component” refer to?

      Eigenmodes is defined in the Stability section as the independently evolving combinations of emotions obtained by projecting the state vector onto the eigenvectors of A, and ”This component” has now an explicit referent.

      (17) Figure 5: see above the comment about the necessity for testing the interactions, which should also be reported in the figures.

      Interaction effects are now included in the relevant figure caption (Figure 5).

      (18) When looking for the relationships between questionnaires and controllability, looking only at the most controllable direction seems rather inefficient due to the multiple comparisons correction. Why not compute the angle (or other measure of similarity) with an ideal ”calm” unit vector?

      Along the same line, it’s unlikely that the most controllable direction will smoothly rotate as a function of symptoms. More likely, the winning (highest eigenvalue) direction will switch from one to another, creating some discontinuity in the summary statistic used for the correlation with clinical scores. How could one work around this issue?

      This is an interesting suggestion; we have not implemented it in the current revision, but we acknowledge it as a promising analysis for future work.

      (19) More generally, it would be interesting to see if there are some regularities in the dynamics across participants. If this is the case, one could construct a canonical emotional dynamics and project all participants on this eigenspace. Emotional trajectories would then differ only in their controllability (eigenvalues) in this common space, making a comparison across participants more straightforward. Could the authors comment on this?

      This is a valuable suggestion that we have not pursued in the current revision, as constructing a common eigenspace across participants requires additional methodological choices.

      (20) I am a bit puzzled by the hypothesis that questionnaires should mediate the intervention effect. Shouldn’t questionnaires be related to before-intervention controllability only? Similarly, could one use initial controllability to predict the intervention response (irrespective or not of the clinical score)?

      Our hypothesis was that participants with greater difficulties in emotion regulation (high DERS-18) would show a smaller intervention effect, as their emotional system might be less amenable to brief distancing. However, this was not the case. We also note that DERS-18 scores were not significantly related to the overall magnitude of pre-intervention controllability (norm of the controllability matrix), but were related to its direction: participants with higher DERS-18 scores showed a most controllable direction pointing toward disgust and away from amusement and calmness, suggesting that trait-level regulation difficulties are linked to the specific emotional configuration of the system rather than its overall controllability.

      Regarding using initial controllability to predict the intervention response: we agree this is a mechanistically appealing question, but it is unfortunately not straightforward to address here. The intervention effect would be quantified as the change between pre- and post-intervention controllability, and since pre-intervention controllability is a constituent of that change score, any correlation between the two would be partially circular by construction.

      (21) The recovery procedure should be way more detailed. How many surrogates were run, etc.? Which kind of quality checks were used to ensure the recoverability was sufficient at the subject level, especially as parameter recovery seems relatively low for some subjects?

      The parameter recovery section now reports that 100 surrogate trajectories were generated per subject (N> = 104) per time period (before and after intervention: N = 208 observations total), with per-subject recovery quality reported across simulations together with across-subject variability; see Supplementary Materials F Parameter Recovery.

      (22) As the result of the eigen decomposition is the endpoint of the analysis, it would be a nice addition to test the recoverability of those measures (eg. correlation between simulated and inferred eigenvalues).

      Recovery of the dominant eigenvalue/vector and dominant singular value/left singular vector is now reported explicitly, including bias analysis and scatter plots of true versus recovered values. Beyond subject-level recovery, we also assessed whether the observed group differences in emotional dynamics and controllability could be reliably recovered at the group level. See Supplementary Materials F Parameter Recovery.

      (23) Although this comment comes close to last, this is a major concern of mine. I am not convinced that the recoverability procedure is sufficient to prove that the inference is working. The model assumes that the observation noise is Gaussian, which is clearly not the case in the data. By simulating surrogate time series with normally distributed noise, the authors do not account for any saturating effects that could destroy a large part of the behavioural information necessary for a successful inversion (eg Figure 4.C showing that simulated data contains a lot of negative ratings). A workaround would be to bind the surrogate time series to mimic the saturation caused by the rating scale, and then run the model estimation on those capped time series.

      We acknowledge this important limitation: the Gaussian assumption does not capture the bounded 0–100 scale, and we have added this to the Limitations with a suggestion that future work use a truncated or censored observation model.

      (24) Supplementary tables with placeholders (v1, v2) that can vary in meaning depending on the line are extremely hard to decipher.

      The supplementary tables have been restructured to a hierarchical format.

      (25) Table I9 is not referenced in the manuscript.

      A reference to this table has been added in the appropriate Results section.

      (26) It’s a shame that neither data nor analysis code has been made available.

      Fully anonymised data and analysis code are now publicly available on GitHub (https://github.com/huyslab/emotioncon public).

      Reviewer #2 (Recommendations for the Authors):

      (1) Abstract: By some definitions, controllability is binary, present or absent, according to whether the controllability Gramian is positive definite. Mention that you use a continuous definition, otherwise ’quantified’ leads to confusion.

      The abstract now describes the measure as: ”Controllability was assessed using continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”, making the non-binary usage explicit.

      (2) p 5: ’on [not in] the recruiting platform’.

      Corrected.

      (3) p 8: Clarify notation of x<sub>t</sub>t = 1<sup>T</sup> (and same for u). What does this mean?

      The notation has been corrected.

      (4) p 8: Define h in the paragraph following Eq 1, don’t wait until the next section.

      The definition of the bias h (steady-state baseline) has been moved to immediately after Equation 1.

      (5) p 9: Give clear references for your methods here. There are different definitions of controllability etc. than the ones you use.

      References have been added at each key definition.

      (6) p 28: ’20 videos per emotion category were chosen resulting in 50 videos per sequence’ - doesn’t make sense.

      This has been clarified: 20 videos per category across 5 categories yield 100 videos in total. These were split into two matched sequences of 50 videos each. Including 2 videos that were repeated twice resulted in 54 videos per sequence and 108 videos in total.

      (7) Figure B2: 54 videos are listed, and categorized into five categories. How does this relate to the remark right above?

      A clarifying note has been added to the supplementary explaining that each sequence of 54 clips includes repeated videos and is drawn from the pool of 100, with emotion-category sequences matched between blocks.

      (8) p 29: What were the process noise Σ and observation noise Γ assumed in the parameter recovery exercise? What were the consequences of that assumption as assessed by simulation?

      Both Σ and Γ were estimated from the data constrained to be diagonal; the parameter recovery section now explicitly reports their recovery quality.

      (9) p 30: How are results affected by including the excluded participants? The level required to pass attention checks seems arbitrary. How was it chosen?

      We have added a sensitivity analysis including the four excluded outliers, showing results remained qualitatively and statistically similar. With 10 binary attention checks, chance performance is 50%, meaning a participant scoring below 70% is performing only marginally above chance and likely not attending consistently. At the same time, 70% is permissive enough to retain participants who may have missed one or two checks due to momentary distraction.

      (10) Figure F4: Typographically distinguish capital letters referring to panels in the figure from those referring to matrices.

      Panel labels in Figure F4 are formatted in bold to distinguish them from italicised matrix notation.

      (11) Table G1: Showing that differences between groups were non-significant before the intervention but significant after is not enough, you need to show that there was a significant interaction between time point and intervention. [I wrote this after reading the supplementary but before reading the main text. It turns out you know what I’m telling you here. You should mention it more prominently though, including in the abstract and the discussion, because in your chosen null-hypothesis significance testing framework, this is the crucial test of your study. I don’t think there’s any harm at all in being up-front about this - certainly much better than making excuses like the one about randomization at the top of page 11, which I recommend removing].

      This is very right. The significant group × time interaction effects are now reported prominently in the main text Results and figure captions. We also wish to be transparent: the interaction tests were conducted post-hoc rather than as the primary analysis, which is the reverse of the correct order. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      Reviewer #3 (Recommendations for the Authors):

      (1) I would encourage the authors to re-add some basic details regarding their power analyses from the supplement to the main text so the reader can immediately reference the intended effect size, which analysis/analyses were considered primary for the power analysis, etc.

      More information about the power analysis have been added to the Participants section of the main text.

      (2) Similarly, I wondered if there could be a little extra information on how the test-retest reliability was calculated (page 13) - on the first and last views of a video pre-intervention? I wasn’t sure - why do the authors only present ICCs for amusement/joy and disgust/horror? Seems useful to present all (particularly because there may be individual differences in habituation to some emotions).

      We have expanded the test-retest section to clarify that reliability was assessed using duplicate videos shown three times pre-intervention, and now report ICCs with confidence intervals, Cronbach’s α, and habituation/sensitization tests for both disgust and amusement; the selection of these two categories reflects which videos were repeated in the design, they were chosen at random during experimental design.

      (3) Regarding my comments about controllability/emotional control, I think the authors probably have two choices - address this head-on (e.g. with a note describing the relationship/distinction between these two concepts of controllability), or else avoid using it in one of the senses (I would suggest the mathematical sense since overriding the concept of cognitive control seems harder - the authors could use phrases like ’impact of emotional inputs’ instead of ’controllable’). In particular, the abstract could be clearer about the nature of controllability as implemented by the authors - this seems critical for communicability. Because of the high relevance of both of these ’control’ concepts to the paper, if the authors agree with my concern, I would also suggest changes throughout, such as in the results section phrasing: ”In those participants with high DERS-18 scores, the most controllable direction pointed towards disgust ( = 0.26, p = 0.006), and away from amusement ( = 0.26, p = 0.005) and calmness ( = 0.24, p = 0.011; though this did not survive Bonferroni correction).”

      Thank you for this comment. Please see our response to the public review comments above, which we hope address this.

      (4) Regarding the control intervention, which is great, is it possible the follow-up question/reminder affected the results – e.g., is there reason to believe that prospective regulation was the primary difference in the distancing group and not a retrospective effect via this question?

      See our response to Reviewer 2 public review Comment 2 and the corresponding Limitations addition.

      (5) Lastly, purely for interest, the authors could consider elaborating on their brief interpretation as to why difficulties in regulating emotions were specifically linked to the controllability of disgust, amusement, and calmness, but not other emotions (anxiety/sadness) (page 20). I wonder if there is a brief space to discuss the emotional specificity of these results further given the relevance to the wider literature on specific emotion types, e.g. fear vs disgust.

      We have expanded the Discussion to elaborate on the emotional specificity of these findings. Amusement and disgust are strongly influenced by external events, suggesting that stimulus-driven controllability is particularly relevant for these emotions. By contrast, anxiety and sadness are maintained through internally generated processes such as rumination and anticipatory cognition, exhibiting greater emotional inertia over time, and their regulation may therefore be less sensitive to momentary stimulus controllability. This provides a mechanistic account of why controllability effects emerged selectively for disgust, amusement, and calmness, and aligns with growing evidence that emotion regulation is emotion-specific rather than domain-general.

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This important study partially fills the gap in the knowledge of olfaction at the level of the Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) with functional magnetic resonance imaging, electrophysiology, and modeling. The methods used are convincing. Some of the findings confirm ongoing hypotheses, such as the behavioral importance of AON for odor source discrimination. Other results shed light on the dynamics of the connection between the olfactory system and the rest of the brain.

      We sincerely thank the editors and reviewers for the thorough review of our manuscript. We appreciate the insightful comments, which have significantly contributed to improving our work. In this revision, we addressed all the concerns posed by the reviewers, including conducting additional analyses and providing the data generated and codes used in this study.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript combined rat fMRI, optogenetics, and electrophysiology to examine the large-scale functional network of the olfactory system as well as its alteration in an aged rat model.

      Strengths:

      Overall methodology is very solid and the results provided an interesting perspective on large-scale functional network perturbation of the olfactory system.

      Weaknesses:

      The biological relevance and validation of the current results can be improved.

      We thank the reviewer for the comments and suggestions regarding our manuscript. They have been invaluable and instrumental in enhancing our work. Please see R1-1 to R1-8 below for our corresponding responses and revisions.

      Comments:

      (R1-1) Figure 1A, on the top of the figure, ChR2 may be replaced by ChR2-mCherry, as only mCherry is fluorescent. And also, it’s somewhat surprising that in AON and Pir regions (where only axon fibers should be labelled as red), most fluorescence appeared dot-like and looked more similar to cell body instead of typical fiber. The authors may want to double-check this.

      (a) Thank you for pointing out the missing fluorescent marker for ChR2 expression in the label of Figure 1A. We distinguished the transfection of ChR2 in olfactory bulb (OB) neurons from axonal terminals in AON and Pir by identifying the expression patterns in these three regions (see Figure 1). In OB, the cellular nucleus (i.e., labelled in blue by DAPI) was surrounded evenly by ChR2-mCherry expression (red fluorescent, indicated by green arrows in Figure 1), which clearly traced the neuronal cell body shape. Meanwhile, the expression of mCherry in AON and Pir consisted of concentrated red fluorescent dots that were sparsely assembled near the cell nucleus, indicating the presence of synaptic boutons at the axonal terminals. Further, the patterns of expression here did not trace the shape of the neuronal cell body.

      (b) In this revision, we amended the label of Fig. 1A to “ChR2-mCherry”.

      In this revision, changes were made to the confocal images of AON and Pir, based on Figure 1, for clarity and the figure caption was also edited accordingly.

      (R1-2) The authors primarily presented 1 Hz stimulation results. What is the most biologically relevant frequency (e.g., perhaps firing frequency under natural odor stimulation) among all frequencies that were used?

      (a) There are various firing rate and oscillatory frequency of neural activities within the olfactory system (i.e., OB, Pir) in rodents during olfaction. For example in OB, neural oscillations ranging from 1 to 12 Hz (i.e., slow to theta) are driven by sensory stimulation and are closely linked to respiration [1]. Specifically, odor-evoked responsive excitatory bursts of mitral and tufted cells in OB are coupled with the low-frequency respiration rhythm (1-4 Hz) [2,3]. Such coupling has also been documented during light anaesthesia [4]. Meanwhile, higher frequencies such as beta oscillations (15 ~ 30 Hz) have been associated with odor learning and sensitization [5,6], and gamma oscillations (40 ~ 80 Hz) evoked by sensory stimulation are associated with fine olfactory discrimination and odor learning [6-8].

      The piriform cortex (Pir) also exhibits natural oscillatory activity across various frequency bands, including slow, theta, beta, and gamma oscillations [9-11]. As in OB mitral and tufted cells, slow oscillations (< 1.5 Hz) in Pir are also correlated with the respiratory rhythm, indicating the intrinsic respiration-related oscillatory properties within the olfactory system [12]. However, while it has been shown that the primary burst of firing in the anterior Pir is locked to respiration, its coding strategies differed from OB during olfaction [13].

      Despite not being able to entirely mimic the evoked neural activities under natural odor stimulation due to the synchronizations induced by our optogenetic stimulations, our frequencies were chosen to be within the range of firing rates of neurons in the olfactory system under natural circumstances. Our selection of 1 Hz frequency for optogenetic stimulation was based on the balance of experimental simplicity of inducing neural excitation and the biological significance of low-frequency neural oscillations in OB, Pir and AON. The robustness of evoked brain-wide long-range activations is one of our criteria for the selection of 1 Hz stimulation frequency at OB, considering that the goal of this study is to examine long-range olfactory networks. Furthermore, we also examined other frequencies of stimulation covering theta, beta and gamma frequency bands, which provided insights into the response characteristics within the olfactory networks.

      (b) In this revision, we added a brief statement in the Results section that 1 Hz was the most biologically relevant frequency based on discussions in R1-2a above.

      (R1-3) In Figure 2, the statistical thresholding is confusing: in the figure legend, it was stated that “t > 3.1 corresponding to P < 0.001” but later “further corrected for multiple comparisons with thresholdfree cluster enhancement with family-wise error rate (TFCE-FWE) at P < 0.05”? Regardless of the statistical thresholding, such BOLD activation seemed to be widespread (almost whole-brain activation). Does such activation remain specific to the optogenetic stimulation, or something more general (e.g., arousal level change)? Furthermore, how those results (I assume they are group-level results) were obtained was not described very clearly. Is it just a simple average of individual-level results, or (more conventionally) second-level analysis?

      (a) We thank the reviewer for drawing our attention to clarity issues regarding the generation of group-level BOLD activation maps and statistical thresholding methods used.

      In brief, we first generate the BOLD activation map of individual animal through conventional general linear model (GLM) analysis. These individual activation maps then underwent a two-step statistical thresholding method to generate the group-level results. Firstly, uncorrected one-sample t-tests were conducted with a threshold of P < 0.001. Secondly, these thresholded activation maps were further corrected using nonparametric inference with threshold-free cluster enhancement multiple comparison correction of family-wise error rate (TFCE-FWE, P < 0.05) before they were averaged. Hence, the BOLD activation maps presented in our manuscript were the result of group-level analysis instead of individual-level analysis.

      (b) We note the reviewer’s concern about whether the observed widespread BOLD activations were caused by the animal’s general brain state (e.g., arousal levels) rather than the specificity of optogenetic stimulation. First, such widespread activations were only specific to 1 Hz stimulation of OB neurons and OB afferents at AON (Figs. 2A, B and 3A, B) whereas activations were localized to regions in the primary olfactory network with increasing stimulation frequencies. Furthermore, varied neural activity adaptation properties were observed (Fig. 3A, B) following repeated 1 Hz optogenetic excitation of OB neurons and OB afferents at AON indicating that such widespread propagation of evoked neural activity was not driven by the animal’s general brain state, which would be relatively random across animals as we interleaved the presentation of each stimulation frequencies. While we cannot discount that certain characteristics of brain states differ under anaesthetized and awake conditions, we showed that light (1.0% isoflurane) anaesthesia minimally affects the BOLD fMRI activations and/or the propagation of optogenetically-evoked neural activity across thalamo-cortical, hippocampal-cortical and vestibulo-cortical networks [14-20].

      Second, the neural representations in OB have been shown to be brain state-independent to ensure the high fidelity of olfactory inputs to higher olfactory cortices, despite differences in spontaneous baseline activity across states [21-24]. In fact, OB neural activity (i.e., slow to delta oscillations) remains highly coupled to respiration rhythms under low anaesthesia as in awake conditions [4]. Further, low-dose isoflurane (i.e., 1% as in the present study) does not affect cortical gamma oscillations (40 Hz) [25-28]. Although odorant-evoked responses in olfactory cortices (e.g., Pir and olfactory tubercle, Tu) can be modulated by overall brain state, various interactions between such cortical circuits to process odor inputs persist under anaesthesia [29-31].

      (c) In this revision, we clarified the two-step statistical thresholding method that was applied to generate the group-level BOLD activation maps, as discussed in R1-3a above, in the captions of Fig. 2 and the Methods section of the manuscript.

      In this revision, we also included a statement in the Discussion section that the widespread BOLD fMRI activations were specific to the 1 Hz optogenetic stimulation and not the animal’s general brain state, as discussed in R1-3b above.

      (R1-4) In Figure 2, why use AUC to quantify the activation, not the more conventional beta value in the GLM analysis?

      (a) We thank the reviewer for raising this concern. We are aware of the more conventional beta value, b, in GLM analysis and have used them for quantification and comparison of the amplitudes of fMRI activations in our previous rodent fMRI studies [18,32,33]. However, in this study, we chose to utilize area under curve (AUC) as it offers a more comprehensive measure of BOLD signal change over time, including shape, duration, and magnitude, thereby capturing the bulk of neural activities and their dynamics throughout the stimulation period. b primarily represents the peak amplitude of BOLD responses (i.e., the % BOLD signal change) [34] and can be constrained by the assumptions and limitations of the GLM analysis, such as the shape of the canonical hemodynamic response function (HRF). Therefore, AUC provides greater accuracy in capturing different aspects of neural responses across various brain regions, such as transient peaks and/or sustained responses.

      (b) In this revision, we have included the justifications for using AUC over the more conventional b value to quantify BOLD fMRI activations, as in R1-4a above, in the Methods section.

      (R1-5) For Figure 2D, the way that it was quantified can be better described as “relative” activation within one condition, and I don’t know how to interpret the comparison among the relative fraction of activated regions. Perhaps comparison using percentage change (i.e., beta values) is more straightforward.

      (a) We thank the reviewer’s comment and suggestion here. “BOLD activation strength” is the AUCs normalized to the respective sum of BOLD activations of their respective stimulation target. We are of the opinion that “strength” can accurately reflect our purpose for conducting this comparison, which is to distinguish the brain networks that were predominantly recruited by AON- or Pir-driven neural activities compared to OB stimulation. 

      (b) We conducted a comparison using relative fraction for fair comparison considering the differences in absolute BOLD signal amplitudes upon different stimulation targets. As we found the overall amplitudes of activations evoked by Pir stimulation were appreciably weaker, our normalization step ensures that the contribution of Pir is not underrepresented. Without this normalization process, the Pir stimulation seems not to activate most of the downstream targets based on the relatively weak activations. As discussed earlier in R1-4a above, comparison using percentage change/b value would be skewed towards the absolute amplitude of BOLD responses and would be erroneous for subsequent interpretation of brain networks that are primarily recruited by OB vs. AON vs. Pir.

      (b) In this revision, we chose not to include additional statements as we have already described in the Results section the reasons for normalizing AUC to reflect and subsequently compare the BOLD activation strength across various networks that were recruited by optogenetic stimulation of three distinct olfactory regions (i.e., OB, AON and Pir). 

      (R1-6) For Figure 3, it may be more convenient for readers to include the results of 1st activation for direct comparison. The current layout makes it difficult to make direct, visual comparisons among all 3 activations. Again, I think using beta values (instead of AUC) may be more conventional.

      (a) We agree that it will be more convenient for readers if we include the 1st activation maps in Figure 3 for direct comparison. Please refer to R1-4 for the usage of AUC rather than beta values.

      (b) In this revision, we modified Fig. 3 to include activation maps from the 1st fMRI session for easy comparison with the 2nd and 3rd sessions.

      (R1-7) Can the DCM results (at least part of it) be verified using the current electrophysiological data? For example, the long-range inhibitory effective connectivity of AON is rather intriguing. If that can be verified using the electrophysiology data, it would be really great. In the current form, the DCM and electrophysiology results seem to be totally unrelated.

      (a) We thank the reviewer for raising this concern, and it’s a great suggestion to causally link the outcomes from dynamic causal modeling (DCM) analysis and electrophysiology findings.

      (b) In principle, the recorded local field potentials (LFPs) can be treated as the neural dynamic component in DCM analysis under specific conditions. They are as follows:

      (1) LFPs are recorded at the identical brain state as the optogenetic fMRI experiments with identical stimulation paradigms.

      (2) The electrode locations of LFP recordings must match the nodes of the a priori matrix defined for the existing DCM analysis (Fig. 4A).

      However, in the present study, we did not have LFP recordings at the entorhinal cortex (Ent), which will likely influence the modeling of effective connectivity strength as one of the nodes defined is dropped from the analysis. Previous studies have showed that changes in the a priori matrix will result in different connectivity estimations [35-38]. We chose to model the connectivity involving Ent due to its documented interactions with the primary olfactory cortices and hippocampal regions [39,40]. Hence, it would be incomplete without modelling the contributions from Ent.

      (c) In this revision, we decided not to redefine the a priori matrix by removing Ent as a node in the DCM analysis to estimate new effective connectivities. Instead, we acknowledge the absence of Ent recordings.

      (R1-8) In Figure 6, it would be great if the adaptation of BOLD and electrophysiology signals can be correlated at the brain region level. The current figure only demonstrated there is adaptation in the electrophysiology recording, but did not show if such adaptation is related to the BOLD adaptation.

      (a) We thank the reviewer for pointing out that we should associate the electrophysiology and fMRI data to validate the olfactory adaptation. As such, we conducted additional correlation analysis between LFP power and BOLD signal profiles. The method for computing the correlation coefficient is described below:

      (1) For each 1 Hz optogenetic stimulation target (i.e., OB, AON, and Pir), regions of interest (ROIs) that have both LFP recordings and BOLD activations were selected. These ROIs were AON, Pir, ventral caudate putamen (vCPu), visual cortex (V1), and ventral hippocampus (vHP) for OB stimulation; OB, Pir, vCPu, V1, amygdala (Amg), and vHP for AON stimulation; and OB, AON, vCPu, V1, Amg, and vHP for Pir stimulation. Note that we excluded the respective stimulated region as a ROI because of BOLD fMRI signal dropout caused by the implanted optical fibre.

      (2) We first calculate the individual animal AUC differences of 2nd stimulation session vs. 1st session and 3rd session vs. 1st session in LFP power and BOLD signal profiles, respectively.

      Subsequently, the mean of the differences for each ROI across animals was then calculated.

      (3) Correlation coefficients and P values of AUC differences between LFP power and BOLD signal profiles were computed using Spearman nonparametric correlation.

      (b) The relationship between the differences of LFP power and BOLD signal profiles showing neural adaptation was described by the correlation coefficient, r, and the corresponding P value. The r and P values upon the three distinct stimulations were OB: r = 0.77 (P = 0.01), AON: r = 0.46 (P = 0.13), and Pir: r = - 0.06 (P = 0.85), respectively. This indicates that the decrease of LFP power is highly correlated with the decrease of BOLD activations upon OB and AON stimulation compared to those upon Pir stimulation.

      (c) In this revision, we added the description of the correlation analysis conducted, as in R1-8a above, in the Methods section.

      In this revision, we also added several statements in the Results section to indicate that neural adaptation observed is significantly related to the decreased BOLD activations when stimulating OB excitatory neurons or OB afferents at AON, as discussed in R1-8b above. Additionally, we also included the outcomes of the correlation analysis as a new Supplementary Fig. 7.

      Reviewer #2 (Public review):

      Summary:

      Ma and colleagues presented a study on the characterization of brain-wide spatio-temporal impact of olfactory cortical outputs. They take advantage of multi-modal techniques on rats: fMRI, optogenetics, and electrophysiology. In addition, they used cutting-edge analytical techniques and modeling to support and interpret their data. The main findings of the study are:

      (1) The neurons in the Olfactory Bulb (OB) predominantly activate primary olfactory network regions, while stimulation of OB afferents in Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) primarily orthodromically activates hippocampal/striatal and limbic networks, respectively.

      (2) Non-specified adaptation or habituation mechanisms may play a significant role in modulating olfactory outputs over subsequent fMRI sessions.

      (3) Artificially induced aging in rats induces profound modification in the functional interaction between olfactory cortices and multiple brain regions.

      The results on AON are of particular interest because of the lack of functional information on this region, despite its recognized importance in shaping OB output and behavior (odor localization tasks).

      Strengths:

      The manuscript is very accurate. The figures are well-crafted, and clear and provide much information with the most appropriate plots and graphics. The study’s amount and data quality are remarkable, and the experimental size adequately addresses the scientific questions. I particularly appreciated the details in the description of the methods regarding the missing data and the size of the different animal groups. The supplementary data complete the leading figures and provide information at a single animal level.

      We are grateful for the reviewer’s appreciation of our study’s breadth and depth, and the quality and quantity of our data. The recognition of the strengths in our methods and the inclusion of single-animal-level data in the supplementary information is highly encouraging. We are also appreciative for the reviewer’s constructive comments below, which has help us to address specific weaknesses of the study and thereby improving the manuscript text and flow. We have read through the comments carefully and have made the necessary corrections. We hope that our responses to R2-1 to R2-11 below has sufficiently addressed all the reviewer’s concerns.

      Weaknesses:

      (R2-1) One of the main reasons the Piriform Cx is understudied in rodents is because of the proximity to air, which creates artifacts in fMRI images. This issue becomes more critical at ultra-high magnetic fields, but I would expect it also at 7T. One main achievement of this study is, indeed, the acquisition of fMRI data from Piriform, and this point should be highlighted by showing raw functional data from a rat. The best would be if an fMRI data sample for a rat, no matter which stimulation, is shared on a public repository, like Zenodo or similar. I am curious to check the quality of the BOLD data from such an ‘enormous’ field of view, particularly in the OB, with a single-shot sequence. Also, the visual inspection of raw data is essential to appreciate how many 0.5 x 0.5 x 1 mm voxels fit into AON, and others analyzed small brain structures, like the amygdala, etc. Was the amygdala entirely visible in BOLD, or did the air in the ear channel make an artifact partially shadowing it?

      (a) As per the reviewer’s suggestion, we displayed the representative EPI images after preprocessing at a matrix size of 128 x 128 with a pixel size of 0.25 x 0.25 mm in Supplementary Figure 2. Up-sampling was not applied out-of-plane. Note that the atlas-based region-of-interest (ROIs) used to extract BOLD signal profiles (i.e., indicated by colored overlays in Supplementary Figure 2) were drawn based on the up-sampled EPI images (from the acquired 64 x 64 to 128 x 128). Notably, the piriform cortex (Pir) and amygdala (Amg) are visible with no appreciable signal dropout at these regions. To clarify, the Pir in our study represented the anterior Pir.

      Signal dropout is pronounced in the posterior ventral regions of the brain (e.g., Bregma- 4 mm, below the Ent) and as expected in a localized region where the implanted optical fiber region that targeted the AON.

      (b) In this revision, we included as a new Supplementary Figure 8 and added a statement in the Results section to describe the absence of appreciable signal dropouts at OB, Pir and Amg, which could affect subsequent quantitative analyses.

      (R2-2) Surprisingly, the only information missing in the methods is the post-surgery period and the time between two consecutive fMRI sessions. How much time was accorded to rats to recover from the surgeries, and what time interval between two scans? This information is crucial for interpreting the decrease in most BOLD responses in subsequent recordings. The supposed adaptation should fit into the known time frames for odor adaptation. Usually, fast adaptation does not last for days (and it should be measured within a single experiment: is it the case?), while for long-lasting adaptation the stimulus (odor or opto) should be maintained constantly ON. This does not seem to be the case in this study. The hypothesis, alternative to adaptation, of a less efficient light activation, for example, due to gliosis around the fiber tips, should be discarded with more evidence than the preservation of OB > Pir responses or acknowledged in the manuscript.

      (a) We thank the reviewer for drawing our attention to the missing information regarding the timeline of each optogenetic fMRI experiment. We conducted fMRI experiments immediately after the surgical procedure for implanting the optical fiber cannula. Such an implantation procedure typically last for 45mins under 1.2-2.0% isoflurane. The animal is then moved from the surgical table to the magnet to begin the optogenetic fMRI experiment. Intervals between two fMRI scans were within one minute. As the stimulation frequencies (i.e., 1 Hz, 5 Hz, 10 Hz, 20 Hz and 40 Hz) were pseudorandomized, the interval between any two identical stimulation frequency was on average 30 minutes. In this case, we are measuring fast adaptation (on the order of tens of minutes) rather than long-lasting adaptation. Further, as the fMRI experiment was conducted immediately following fiber implantation, the risk of gliosis around the fiber tips is expected to be low.

      (b) The main representation of olfactory adaptation is the attenuation of neural responses upon continuous or repeated stimuli [41,42]. Several fMRI studies in rodents have reported that olfactory adaptation was detected in primary olfactory cortices (i.e., OB, AON, and Pir) upon repeated odor stimulation [43,44]. Specifically, these studies showed robust decreased responses in primary olfactory cortices under repeated odor stimulation with an odor interval of 30 minutes, which were comparable to those observed in our study upon repeated 1 Hz optogenetic stimulation. As such, the pronounced neural adaptation that we demonstrated when stimulating OB excitatory neurons or OB afferents at AON is unlikely to be caused by less efficient light activation.

      (c) In this revision, we added statements in the Methods section, as discussed in R2-2a above, to clarify the timeline of our optogenetic fMRI experiment.

      (R2-3) The D-galactose experiments were conducted only after administering the aging molecule, with no baseline/reference data on the same animals. Then, comparisons were made with healthy rats, but the two groups not only can be discriminated with respect to D-galactose administration but also with age (10 VS 18 weeks). A control group for 18-weeks-old rats with no D-galactose treatment would better compare the D-galactose effect and avoid any potential bias from group comparisons of rats at different ages. Do you confirm that D-galactose was injected into each rat 56 times/day in a row, or am I mistaken?

      (a) We appreciate the reviewer’s comment here and understand the concerns regarding the absence of an age-matched control group with saline administration instead of D-galactose. We acknowledge that the inclusion of a control group would have provided a more direct comparison to evaluate the differences between healthy and aged animals. However, we were unable to include this group in the present study due to logistical constraints.

      To clarify, our experiments were conducted on healthy rats at the median age of 14 weeks to ensure the animals had reached adulthood. The median age of the D-galactose injected animal was 18 weeks. In terms of the lifespan of rats, they reach sexual maturity at approximately 6 weeks of age [45], at which point we conducted the optogenetic viral vector injection. The age of rats at 14 and 18 weeks can both be regarded as teenage adult [45,46]. The primary purpose of the experiments on the aged animal model and corresponding analyses was to provide some potential clues into the dysfunction of olfactory networks at the system level as animals aged. Despite the age difference between the two groups, we believe that the comparison between healthy and aged rats still offers valuable insights into the ageing-related dysfunctions within the olfactory system. 

      (b) Note that D-galactose was injected into each rat 56 times in total over the whole experiment (i.e., one injection per day over the course of 8 weeks).

      (c) In this revision, we acknowledge that the comparisons made were not with age-matched healthy controls in the Discussion section.

      In this revision, we also clarified that D-galactose was administered once daily for 8 weeks (i.e., a total of 56 injections) in the Methods section.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Minor:

      (R2-4) The look of the activation maps will greatly improve if the displayed brain slices are coronal (as in the supplementary figures) with no angle.

      We thank the reviewer for this suggestion. The purpose of such a display was to maintain the consistency of displaying the overlay of BOLD fMRI activation maps on the 3D renders of rat brain anatomical images. In addition, such 3D display can better impress readers that the optogenetically evoked BOLD activations were indeed long-range and brain-wide. 

      As such, in this revision, we decided to maintain the existing displays of BOLD activation maps in the main figures.

      (R2-5) The choice of testing different stimulation frequencies deserves more justification. Why was it performed in the OB, which is naturally activated on the breathing rhythm?

      (a) We thank the reviewer’s comment here to seek further clarification. OB is widely recognized as the first stage of olfactory information processing in the brain. By activating the OB excitatory neurons, we can ensure that our optogenetically-evoked activations are olfactory-related.

      The choice of various frequencies ranging from low (i.e., 1 Hz) to high (e.g., 40 Hz) was made to cover a range of firing rate and oscillatory frequency of neural activities neural oscillations that exist within the rodent olfactory system during olfaction. For example in OB, neural oscillations ranging from 1 to 12 Hz (i.e., slow to theta) are driven by sensory stimulation and are closely linked to respiration [1]. Specifically, odor-evoked responsive excitatory bursts of mitral and tufted cells in OB are coupled with the low-frequency respiration rhythm (1-4 Hz) [2,3]. Such coupling has also been documented during light anaesthesia [4]. Meanwhile, higher frequencies such as beta oscillations (15 ~ 30 Hz) have been associated with odor learning and sensitization [5,6], and gamma oscillations (40 ~ 80 Hz) evoked by sensory stimulation are associated with fine olfactory discrimination and odor learning [6-8].

      (b) In this revision, we added a statement in the Results section to further describe our justifications for testing various stimulation frequencies.

      (R2-6) The absence of natural (odor) stimulation should be acknowledged as a limitation for the findings of this study in the Discussion.

      We thank the reviewer for the suggestion. In this revision, we made the acknowledgement in the Discussion section. 

      (R2-7) Line 46: overstatement, the role of the Piriform cortex at the system level is well known and was not discovered by this study. (see Gordon Shepherd’s book: The Synaptic Organization of the Brain, Figure 10.6)

      As per the reviewer’s suggestion, we decided to use “distinguishes” rather than “uncovers” for an accurate description of our study’s findings, which demonstrated the differences between the piriform cortex and anterior olfactory nucleus in driving the downstream activations of olfactory and non-olfactory targets.

      (R2-8) Line 75: please rephrase for an easier understanding.

      As per the reviewer’s suggestion, we edited this statement to “Further, the need to present a multitude of odor combinations also makes it challenging to efficiently and reliably interrogate long-range olfactory networks and their properties”.

      (R2-9) Line 313: the intermediary region between the Piriform and Entorhinal cortices is the Perirhinal Cortex. No doubt about that.

      As per the reviewer’s comment, we modified the corresponding statement by adding “such as the perirhinal cortex” and cited the appropriate references [47,48].

      (R2-10) Line 391: there is no evidence from this study to exclude that the downstream targets would be less activated because of a decreased OB activation in favor of broad inhibition.

      As per the reviewer’s suggestion, we modified the statement by replacing “rather than” with “in addition to” for a more precise statement.

      (R2-11) Line 585: which was/were the regressor/s used in GLM? Any convolution with common HRFs?

      We regarded the block-designed optogenetic stimulation as a task. The regressors are then obtained by convolving the expected task-evoked neural activity (i.e., boxcar function for block design) with the canonical hemodynamic response function, HRF (SPM12, Wellcome Department of Imaging Neuroscience, University College London, UK).

      In this revision, we added a statement in the Methods section to provide clarification on the regressors used in GLM and the convolution step with canonical HRF.

      References

      (1) Kay, L. M. & Stopfer, M. Information processing in the olfactory systems of insects and vertebrates. Semin Cell Dev Biol 17, 433-442 (2006). https://doi.org/10.1016/j.semcdb.2006.04.012

      (2) Cang, J. & Isaacson, J. S. In vivo whole-cell recording of odor-evoked synaptic transmission in the rat olfactory bulb. J Neurosci 23, 4108-4116 (2003). 

      (3) Margrie, T. W. & Schaefer, A. T. Theta oscillation coupled spike latencies yield computational vigour in a mammalian sensory system. J Physiol 546, 363-374 (2003). https://doi.org/10.1113/jphysiol.2002.031245

      (4) Fontanini, A. & Bower, J. M. Variable coupling between olfactory system activity and respiration in ketamine/xylazine anesthetized rats. Journal of Neurophysiology 93, 3573-3581 (2005). https://doi.org/10.1152/jn.01320.2004

      (5) Kay, L. M. et al. Olfactory oscillations: the what, how and what for. Trends in Neurosciences 32, 207-214 (2009). https://doi.org/10.1016/j.tins.2008.11.008

      (6) Gervais, R., Buonviso, N., Martin, C. & Ravel, N. What do electrophysiological studies tell us about processing at the olfactory bulb level? Journal of physiology, Paris 101, 40-45 (2007). https://doi.org/10.1016/j.jphysparis.2007.10.006

      (7) Beshel, J., Kopell, N. & Kay, L. M. Olfactory bulb gamma oscillations are enhanced with task demands. J Neurosci 27, 8358-8365 (2007). https://doi.org/10.1523/JNEUROSCI.119907.2007

      (8) Martin, C., Beshel, J. & Kay, L. M. An olfacto-hippocampal network is dynamically involved in odor-discrimination learning. J Neurophysiol 98, 2196-2205 (2007). https://doi.org/10.1152/jn.00524.2007

      (9) Kay, L. M. Theta oscillations and sensorimotor performance. Proc Natl Acad Sci U S A 102, 3863-3868 (2005). https://doi.org/10.1073/pnas.0407920102

      (10) Lowry, C. A. & Kay, L. M. Chemical factors determine olfactory system beta oscillations in waking rats. J Neurophysiol 98, 394-404 (2007). https://doi.org/10.1152/jn.00124.2007

      (11) Vanderwolf, C. H. & Zibrowski, E. M. Pyriform cortex beta-waves: odor-specific sensitization following repeated olfactory stimulation. Brain Res 892, 301-308 (2001). https://doi.org/10.1016/s0006-8993(00)03263-7

      (12) Fontanini, A., Spano, P. & Bower, J. M. Ketamine-xylazine-induced slow (< 1.5 Hz) oscillations in the rat piriform (olfactory) cortex are functionally correlated with respiration. J Neurosci 23, 7993-8001 (2003). https://doi.org/10.1523/JNEUROSCI.23-22-07993.2003

      (13) Miura, K., Mainen, Z. F. & Uchida, N. Odor representations in olfactory cortex: distributed rate coding and decorrelated population activity. Neuron 74, 1087-1098 (2012). https://doi.org/10.1016/j.neuron.2012.04.021

      (14) Leong, A. T. et al. Long-range projections coordinate distributed brain-wide neural activity with a specific spatiotemporal profile. Proc Natl Acad Sci U S A 113, E8306-E8315 (2016). https://doi.org/10.1073/pnas.1616361113

      (15) Chan, R. W. et al. Low-frequency hippocampal–cortical activity drives brain-wide resting-state functional MRI connectivity. Proc Natl Acad Sci U S A 114, E6972-E6981 (2017). https://doi.org/10.1073/pnas.1703309114

      (16) Leong, A. T. L. et al. Optogenetic fMRI interrogation of brain-wide central vestibular pathways. Proceedings of the National Academy of Sciences 116, 10122-10129 (2019). https://doi.org/10.1073/pnas.1812453116

      (17) Leong, A. T. L., Wang, X., Wong, E. C., Dong, C. M. & Wu, E. X. Neural activity temporal pattern dictates long-range propagation targets. Neuroimage 235, 118032 (2021). https://doi.org/10.1016/j.neuroimage.2021.118032

      (18) Leong, A. T. L., Wong, E. C., Wang, X. & Wu, E. X. Hippocampus Modulates Vocalizations Responses at Early Auditory Centers. Neuroimage 270, 119943 (2023). https://doi.org/10.1016/j.neuroimage.2023.119943

      (19) Wang, X. et al. Functional MRI reveals brain-wide actions of thalamically-initiated oscillatory activities on associative memory consolidation. Nat Commun 14, 2195 (2023). https://doi.org/10.1038/s41467-023-37682-8

      (20) Xie, L. et al. Brain-wide resting-state fMRI network dynamics elicited by activation of single thalamic input. Nat Commun 16, 11247 (2025). https://doi.org/10.1038/s41467-025-66104-0

      (21) Li, A., Gong, L. & Xu, F. Brain-state-independent neural representation of peripheral stimulation in rat olfactory bulb. Proc Natl Acad Sci U S A 108, 5087-5092 (2011). https://doi.org/10.1073/pnas.1013814108

      (22) Lang, J. et al. Odor representation in the olfactory bulb under different brain states revealed by intrinsic optical signals imaging. Neuroscience 243, 54-63 (2013). https://doi.org/10.1016/j.neuroscience.2013.03.057

      (23) Chery, R., Gurden, H. & Martin, C. Anesthetic regimes modulate the temporal dynamics of local field potential in the mouse olfactory bulb. J Neurophysiol 111, 908-917 (2014). https://doi.org/10.1152/jn.00261.2013

      (24) Wachowiak, M. et al. Optical dissection of odor information processing in vivo using GCaMPs expressed in specified cell types of the olfactory bulb. J Neurosci 33, 5285-5300 (2013). https://doi.org/10.1523/JNEUROSCI.4824-12.2013

      (25) Hudetz, A. G., Vizuete, J. A. & Pillay, S. Differential effects of isoflurane on high-frequency and low-frequency gamma oscillations in the cerebral cortex and hippocampus in freely moving rats. Anesthesiology 114, 588-595 (2011). https://doi.org/10.1097/ALN.0b013e31820ad3f9

      (26) Joliot, M., Ribary, U. & Llinas, R. Human oscillatory brain activity near 40 Hz coexists with cognitive temporal binding. Proc Natl Acad Sci U S A 91, 11748-11751 (1994). https://doi.org/10.1073/pnas.91.24.11748

      (27) Murthy, V. N. & Fetz, E. E. Oscillatory activity in sensorimotor cortex of awake monkeys: synchronization of local field potentials and relation to behavior. J Neurophysiol 76, 3949-3967 (1996). https://doi.org/10.1152/jn.1996.76.6.3949

      (28) Tallon-Baudry, C., Bertrand, O., Delpuech, C. & Pernier, J. Stimulus specificity of phaselocked and non-phase-locked 40 Hz visual responses in human. J Neurosci 16, 4240-4249 (1996). https://doi.org/10.1523/JNEUROSCI.16-13-04240.1996

      (29) Murakami, M., Kashiwadani, H., Kirino, Y. & Mori, K. State-dependent sensory gating in olfactory cortex. Neuron 46, 285-296 (2005). https://doi.org/10.1016/j.neuron.2005.02.025

      (30) Wilson, D. A. & Yan, X. Sleep-like states modulate functional connectivity in the rat olfactory system. J Neurophysiol 104, 3231-3239 (2010). https://doi.org/10.1152/jn.00711.2010

      (31) Schreck, M. R. et al. State-dependent olfactory processing in freely behaving mice. Cell Rep 38, 110450 (2022). https://doi.org/10.1016/j.celrep.2022.110450

      (32) Gao, P. P., Zhang, J. W., Chan, R. W., Leong, A. T. L. & Wu, E. X. BOLD fMRI study of ultrahigh frequency encoding in the inferior colliculus. Neuroimage 114, 427-437 (2015). https://doi.org/10.1016/j.neuroimage.2015.04.007

      (33) Gao, P. P., Zhang, J. W., Fan, S. J., Sanes, D. H. & Wu, E. X. Auditory midbrain processing is differentially modulated by auditory and visual cortices: An auditory fMRI study. Neuroimage 123, 22-32 (2015). https://doi.org/10.1016/j.neuroimage.2015.08.040

      (34) Goddard, E. & Mullen, K. T. fMRI representational similarity analysis reveals graded preferences for chromatic and achromatic stimulus contrast across human visual cortex. Neuroimage 215, 116780 (2020). https://doi.org/10.1016/j.neuroimage.2020.116780

      (35) Friston, K. J., Harrison, L. & Penny, W. Dynamic causal modelling. Neuroimage 19, 12731302 (2003). https://doi.org/10.1016/s1053-8119(03)00202-7

      (36) Friston, K. J., Kahan, J., Biswal, B. & Razi, A. A DCM for resting state fMRI. Neuroimage 94, 396-407 (2014). https://doi.org/10.1016/j.neuroimage.2013.12.009

      (37) Bernal-Casas, D., Lee, H. J., Weitz, A. J. & Lee, J. H. Studying Brain Circuit Function with Dynamic Causal Modeling for Optogenetic fMRI. Neuron 93, 522-532 e525 (2017). https://doi.org/10.1016/j.neuron.2016.12.035

      (38) Zeidman, P. et al. A guide to group effective connectivity analysis, part 1: First level analysis with DCM for fMRI. Neuroimage 200, 174-190 (2019). https://doi.org/10.1016/j.neuroimage.2019.06.031

      (39) Salimi, M. et al. Disrupted connectivity in the olfactory bulb-entorhinal cortex-dorsal hippocampus circuit is associated with recognition memory deficit in Alzheimer's disease model. Sci Rep 12, 4394 (2022). https://doi.org/10.1038/s41598-022-08528-y

      (40) Chen, Y. N., Kostka, J. K., Bitzenhofer, S. H. & Hanganu-Opatz, I. L. Olfactory bulb activity shapes the development of entorhinal-hippocampal coupling and associated cognitive abilities. Curr Biol 33, 4353-4366 e4355 (2023). https://doi.org/10.1016/j.cub.2023.08.072

      (41) Pellegrino, R., Sinding, C., de Wijk, R. A. & Hummel, T. Habituation and adaptation to odors in humans. Physiol Behav 177, 13-19 (2017). https://doi.org/10.1016/j.physbeh.2017.04.006

      (42) Sinding, C. et al. New determinants of olfactory habituation. Sci Rep 7, 41047 (2017). https://doi.org/10.1038/srep41047

      (43) Zhao, F. et al. fMRI study of olfaction in the olfactory bulb and high olfactory structures of rats: Insight into their roles in habituation. Neuroimage 127, 445-455 (2016). https://doi.org/10.1016/j.neuroimage.2015.10.080

      (44) Zhao, F. et al. fMRI study of the role of glutamate NMDA receptor in the olfactory adaptation in rats: Insights into cellular and molecular mechanisms of olfactory adaptation. Neuroimage 149, 348-360 (2017). https://doi.org/10.1016/j.neuroimage.2017.01.068

      (45) Sengupta, P. The Laboratory Rat: Relating Its Age With Human's. Int J Prev Med 4, 624-630 (2013). 

      (46) Quinn, R. Comparing rat’s to human’s age: How old is my rat in people years? Nutrition 21, 775-777 (2005). https://doi.org/10.1016/j.nut.2005.04.002

      (47) Burwell, R. D. & Amaral, D. G. Perirhinal and postrhinal cortices of the rat: interconnectivity and connections with the entorhinal cortex. J Comp Neurol 391, 293-321 (1998).

      (48) Kajiwara, R., Takashima, I., Mimura, Y., Witter, M. P. & Iijima, T. Amygdala input promotes spread of excitatory neural activity from perirhinal cortex to the entorhinal-hippocampal circuit. J Neurophysiol 89, 2176-2184 (2003). https://doi.org/10.1152/jn.01033.2002

    1. Reviewer #2 (Public review):

      Summary:

      Overall, the authors aimed to provide evidence that clarifies two debates within metacognition research concerning subjective confidence reports:

      (1) Does the post-decision confidence report arise from the same process that drives the initial decision, or does a separate, independent process support confidence computation?

      (2) How do we stop accumulating evidence for the post-decision confidence report? Is it based on a self-imposed time limit, or on accumulated evidence crossing a boundary?

      For the investigation, the authors constructed four models (2 × 2 factorial) to compare each combination of processes to account for random-dot motion tasks data with speed/accuracy manipulations. The models are generally embedded in the drift diffusion model framework, retaining basic parameters such as drift rate, boundary separation, starting point, and non-decision time, while adding linearly collapsing boundaries to model the initial choice. For the single vs. distinct process dimension, the difference lies in whether post-decision evidence accumulation is referenced to the endpoint of the initial decision process or restarts from a new, freely estimated starting point. For the time- vs. boundary-based stopping rule dimension, the key difference is that post-decision evidence accumulation stops either at a deadline sampled from a normal distribution or when the accumulated evidence hits a collapsing boundary.

      Based on model comparison, the boundary-based stopping rule clearly outperformed the time-based stopping rule. However, models with the boundary-based stopping rule performed similarly regardless of whether a single or distinct process was used. Here, the authors drew additional insights from EEG recordings during the task, focusing on the centro-parietal positivity (CPP), which has been proposed as a neural correlate of the evidence accumulation process. By simulating evidence accumulation trajectories (with additional assumptions) and comparing the patterns of those trajectories with observed ERP waveforms, the authors argued that the single-process model provided a better match to the CPP findings and was therefore preferred. This was specifically demonstrated by the model's superior ability to match the pre-response CPP amplitude differences conditioned on the post-decision confidence-related variables.

      Strengths:

      (1) The authors translated existing theories into computational models of decision-making and systematically compared different cognitive processes by assessing model fits to the data. This provides strong evidence supporting the idea that post-decision confidence reports could be better explained by boundary crossing rather than a self-imposed deadline to respond.

      (2) Beyond model evidence, an important result is that CPP amplitude predicted confidence before the initial choice was reported, which is a unique prediction of the single-process model. The use of EEG as an independent validation measure provided additional evidence in favour of this model.

      (3) Combining points 1 and 2, this study successfully addressed the two key debates with solid evidence to favour one theory over another.

      (4) Another strength of this study is the data quality. The high number of trials provided a strong foundation for model inference as well as ERP analysis. The experiment also contained a speed-accuracy manipulation to evaluate model performance across diverse situations.

      Weaknesses:

      I have two main concerns around the modelling work and neural analyses, which in my opinion could have limited the interpretation of the findings. My responses here will be lengthier, but this reflects the nature of the modelling work rather than implying stronger criticisms than those suggested by the strengths discussed above.

      (1) There are a few assumptions in the models that lack psychologically meaningful interpretations, and this study placed more effort into model comparison while lacking discussion of the cognitive processes inferred from parameter estimates.

      To start, I think some of the parameterisations were not properly justified. For the boundary models, it is not very clear why the upper and lower boundaries were different and collapsed at different rates for confidence decisions, given that a single boundary parameter and collapse rate were used for the initial decision. This allows more flexible shifts in the model's predictions of confidence ratings without strong justification. Specifically, it is unclear why the boundary-single model has such an implementation while the boundary-distinct model was only equipped with one boundary parameter (a2, compared to a2up and a2down).

      Similarly, the inclusion of metacognitive noise creates another layer of flexibility in the predictions of confidence ratings. In most existing evidence accumulation models with a diffusion process, noise comes from two sources: within-trial noisy evidence accumulation and across-trial variability (e.g., drift rate variability). Beyond these, such models almost always assume that the decision is made deterministically once the evidence reaches a specific boundary. The inclusion of metacognitive noise here sounds more like a noisy decision-to-action mapping.

      I also have similar doubts about allowing the non-decision time parameter for confidence accumulation in the distinct model to be negative. The authors argued that confidence accumulation may begin during initial evidence accumulation. However, this is a flawed implementation, as the non-decision time was simply added to the evidence accumulation time rather than being incorporated within it. Allowing negative non-decision times may achieve similar predictions, but it is ad hoc.

      The inclusion of a collapsing boundary mechanism in the post-decision confidence accumulator helped the model reach more diverse levels of accumulated evidence and ultimately improved predictions of confidence ratings. However, no strong argument is presented for this implementation beyond the observation that the model performs worse without it. The collapsing boundary mechanism has traditionally been interpreted as reflecting a sense of urgency. For the boundary models, I noticed that the collapse rate of the upper boundary differed significantly between speed and accuracy conditions, which is consistent with the urgency interpretation. Overall, I would like to see more discussion of the specific model mechanisms included by the authors, interpreted in light of parameter estimates.

      (2) While the ERP findings provided external evidence and validation of the modelling results, I find the simulation practices not particularly useful and potentially misleading for naive readers. Specifically, the authors attempted to draw a parallel between patterns of simulated evidence accumulation traces and observed CPP waveforms. While the CPP has received support as a correlate of the evidence accumulation process, the DDM is by no means a neural model capable of generating predictions of neural observations. To my understanding, the superior fit of the boundary-single model was primarily due to the fact that pre-response CPP amplitude predicts post-decision confidence ratings. Therefore, as the boundary-distinct model did not connect the two phases of evidence accumulation, it would fail to account for this observation. I think this point could be clearly demonstrated without the need to introduce additional assumptions into the model simulations in order to directly compare averaged trajectories with averaged ERP waveforms. While the authors did not explicitly claim otherwise, this approach creates an illusion that the model can mechanistically account for ERP data. I would like the authors to provide explicit clarification on this point.

      Appraisal:

      Overall, the authors have provided solid evidence in support of their research aims. The findings contribute to longstanding debates with insights from model mechanisms and neural findings that should not be overlooked by future studies on this topic. This study also offers a good starting point for future model development and refinement in broader contexts of confidence reporting, such as paradigms involving simultaneous initial decisions and confidence judgements. The high quality of the behavioural and EEG data will make a valuable contribution to future research.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPES.

      We thank the reviewer for noting these strengths of the work along with the below encouragement to revise the manuscript to better highlight the key findings and their implications.

      Weaknesses:

      The technical proficiency and complexity of the study and analysis also present a clear limitation and challenge for interpretation. As a reader, even those who are quite knowledgeable about the methods, constructs, and questions being addressed will often struggle (as this reviewer did) to keep the large set of findings in mind and gain an understanding of how they all fit together.

      Indeed, it seems like there are many threads running together in the paper, which makes it challenging to find the through-line of the key findings, or to understand how they might relate to some pre-existing hypotheses, rather than merely interesting patterns detected in the data. In the Introduction and Discussion, it seems as if the key question is to understand the pathways by which MPEs impact cognition, but this is a rather broad topic, so it is not clear exactly what the authors are aiming at with this question and study design.

      As an example, authors operationalize frontal theta power as an index of cognitive control demand, and one of the pathways by which MPEs impact cognition. But this point becomes somewhat circular, since it is not clear how or why the Mismatch x Strength interaction in frontal theta reflects that demand. It would have been better to set this pattern up in the Introduction as a theoretically driven hypothesis, since it currently appears more like a post-hoc interpretation. This is mirrored by how the issue is first brought up in the Introduction, where it states somewhat vaguely: "whether MPEs are followed by an increase in frontal theta... warrants closer examination".

      Again, we appreciate the reviewer’s thoughtful feedback on where the manuscript can be clearer, especially given the rich set of results it reports. Following the reviewer’s guidance, we restructured and revised the Introduction to further motivate the hypotheses that (a) MPEs increase both attention/arousal (grounded in studies of Event Segmentation Theory) and cognitive control (given findings on reward prediction errors and other types of prediction errors), and (b) there are greater increases in these processes triggered by strong compared to weak MPEs. To better link these hypotheses to resulting statistical tests, we note that hypothesis (a) was tested in our trial-level regression models in Fig. 1 by examining main effects of Strength, whereas hypothesis (b) was tested in our models in the Mismatch x Strength interactions. On point (b) and potentially circularity, we note that previous work indicates that frontal theta scales with negative reward prediction errors; as such, we hypothesized that stronger MPEs would elicit more frontal theta, as evidenced by a robust Mismatch x Strength interaction during the probe period.

      Later in the results, there are findings relating frontal theta to pupil dilation, posterior alpha suppression and then subsequent memory. It was hard to understand how all the findings might be linked together functionally or conceptually. Are the authors potentially postulating a mediating or mechanistic pathway, in which the MPE leads to increased cognitive control (frontal theta), which then leads to enhanced subsequent memory of those events? If this is the case, then maybe a formal path analysis would be the best way to test or state this hypothesis. It would also be useful to specify more clearly how the pupil components and alpha suppression factor into this mediating path, since it was not clear.

      Relatedly, the authors suggest that internal attention and arousal also play relevant roles in this pathway, but these are also not clear. In some cases, it is stated as if this is a distinct pathway from the cognitive control one, since there is a focus in the results on the independence of frontal theta and posterior alpha, but elsewhere they seem to be treated as two aspects, or distinct steps, within a single pathway. Again, these different threads of the findings were quite challenging for the reader to follow. Pathway analyses, such as with multiple mediation or moderated mediation, could be a useful way to address this question. For example, it seems as if readiness-to-remember is another behavioral outcome (like subsequent memory) that could be used in the search for mediators.

      We thank the reviewer for highlighting these ambiguities in the original submission and for the thoughtful encouragement to leverage mediation models to more formally test the hypothesized relationships between MPE magnitude and constructs of control, attention, and arousal. In the revision, we now more clearly hypothesize that the effects of strong MPE-driven increases in attention and arousal might be explained, in part, by cognitive control (as indexed by frontal theta) upregulating attention and arousal. To more explicitly test this model of the relationships between our measures, as recommended by Reviewer #1, we now include multivariate mediation analyses to assess whether, at a trial-level, changes in posterior alpha and the immediate pupil effect PC3 are explained in part by increases in frontal theta. Because changes in posterior alpha following MPEs and PC3 scores did not predict subsequent memory, mediation analyses addressing the hypothesis that our attention/arousal measures mediate the effect of frontal theta on subsequent memory were not conducted. Following insights from a cross-correlation analysis, as recommended by Reviewer #2, we tested an additional model to examine whether the effects of MPE magnitude on frontal theta were explained in part by changes in posterior alpha. Examination of the posterior distributions of the indirect effects did not favor our hypothesized model, nor the alternative model that attention upregulates control. Altogether, these outcomes suggest that increases in control, attention, and arousal following strong MPEs may be elicited independently. Yet, we also note that the current set of experiments may be underpowered for these mediation analyses; future work can further investigate the directionality between these effects. Finally, with respect to the readiness-to-remember findings, they were removed in the interest of space, as recommended by Reviewer #2.

      At the minimum, it would be quite helpful to have diagrammatic figures that specify the hypothesized and observed relationships between independent variables (Strength, Mismatch), physiological indices (pupil dilation components, frontal theta, posterior alpha) and key outcome measures (accuracy, RT, next-trial retrieval success, subsequent memory), so that the reader can refer back to them as each component of the analyses is conducted.

      To further increase conceptual clarity, we also followed this helpful suggestion, adding diagrammatic figures to illustrate our hypotheses regarding interactions between the effects (Fig. 3a) and to summarize the observed relationships between MPEs, control, attention, and arousal (Fig. 6).

      Minor Points:

      Many figures had x-axes showing a pupil component or EEG power metric broken down by quartile or quintile. Yet nowhere is it ever explained why this graphical (or analytic?) approach is used and what it reflects, or how it is decided which break down to use (quartile/quintile). If the data are analyzed as a correlation, why is a scatterplot not shown instead?

      In the linear mixed effects models, continuous values were used to assess relationships between variables. In the figures, the continuous variables were binned into quartiles or quintiles for ease of visualization. We opted to visualize the data using this approach, rather than with a scatterplot, given the large number of trials. We updated the figure captions to clarify the approach.

      It was surprising that, unlike readiness-to-remember, which was analyzed via logistic regression and odds-ratio, subsequent memory was not analyzed in the same fashion (i.e., as a binary outcome variable predicted by frontal theta), rather than in a reverse chronological one (subsequent memory predicting frontal theta). Historically, it was the case that subsequent memory was analyzed in this manner, but that was before the era in which trial-level linear mixed-effect models were in wide usage, as they are implemented in this study. Thus, the choice seems like a wasted opportunity or a step backwards analytically.

      We thank the reviewer for this encouragement and agree with the point. In the revision, we note that the readiness-to-remember results were removed in the interest of space and clarity, as was recommended by Reviewer #2. With respect to the subsequent memory analyses, they are now analyzed via logistic regression.

      Reviewer #2 (Public review):

      Strengths:

      The study has a clear behavioral paradigm with multiple measures - behavioral, EEG, and pupillometry that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

      We thank the reviewer for noting these strengths of the work along with the below encouragement to revise the manuscript to better highlight the key findings and their implications.

      Weaknesses:

      The methods are rigorous, and the claims are mostly supported by the data, but there are a few weaknesses or places that could be improved:

      (1) The authors conduct PCA analysis to identify different components of the pupillary response to MPE and relate them to behavior. Specifically, the authors identify components PC3 and PC4, which they interpret as related to MPE. However, some parts of the interpretation could be clearer or better justified:

      (a) The authors refer to PC4 as "post-decision cognitive processing". But, given that RT was between .5-.7s, and PC3 peaked after more than 1s, wouldn't it be cautious to interpret PC3 as postdecision as well?

      Thank you for raising this point. Given that pupil is a relatively sluggish response, it is possible that both components reflect post-decision cognitive processing, even if PC3 peaks before PC4. Following the reviewer’s guidance to adopt more cautious language, we replaced “post-decision” with “post-MPE”.

      (b) MPEs overall elicit longer RTs in this study, suggesting that long RT is a behavioral marker of MPE. Nonetheless, the authors argue on p. 12: "Altogether, these findings indicate that when stronger mnemonic predictions (as indexed by shorter RTs) were violated." And, PC3 is correlated with shorter RTs for mismatches, meaning that behaviorally, these trials were more similar to matches. Thus, how do the authors interpret shorter versus longer RTs for MPEs, and what processes do these RT reflect?

      We thank the reviewer for stressing the need for greater clarity regarding the relationships between RT and the constructs of interest. With respect to RT, we interpret the condition-level difference in RTs between mismatch and match trials as a behavioral marker of an MPE. However, when comparing mismatch trials within a given strength condition to each other (i.e., an analysis at the trial-level), shorter RTs may reflect a stronger prediction, greater certainty that the probe is a mismatch, and therefore the experience of a stronger MPE. Note that while larger PC3 scores were associated with shorter RTs for mismatches (Fig. 2b, right), the mismatch RTs in the largest PC3 quartile were still longer than those on match trials in the corresponding quartile (in other words, there was still a condition-level difference in RTs in the largest PC3 quartile, suggesting that the mismatch trials in this bin are behaviorally still likely to be different from match trials in the corresponding bin).

      To clarify these relationships and our interpretation, we modified the referred to text: “(as indexed by shorter strong mismatch RTs).” Moreover, we added text to the Discussion, further delineating our reasoning here and the implications of our findings for understanding the mechanisms giving rise to and triggering by MPEs of varying strengths. This includes adding an explicit summary of the logic and findings that notes that, at the trial-level, shorter RTs may reflect stronger predictions; at the match/mismatch condition-level, longer RTs may reflect the experience of a MPE. For PC3, shorter RTs (trials with a stronger prediction) in the Strong Mismatch condition were associated with a larger pupillary response. The added text notes that “while longer mean RTs for mismatches compared to matches are a behavioral marker of a MPE, within-condition differences in RTs (i.e., between mismatch trials) may reflect more subtle differences in MPE magnitude, with shorter RTs reflecting stronger predictions and thus stronger MPEs. Strong mismatch RTs were used in the mediation models as a proxy measure of MPE magnitude; this estimate may be noisy because RTs in this experiment are likely sensitive to factors independent of mnemonic prediction strength (e.g., preparatory attention (Supplementary Fig. 6) or memory strength of the mismatch probe). This limitation may have additionally reduced sensitivity to detecting indirect effects.”

      (2) The brain to pupil relationship (p. 13-14): If I understand correctly, this was done on a trial-by-trial basis, but the high temporal resolution allows doing the analysis in a time-resolved manner - does brain activity at a certain time point preceding/following the pupil response correlate with the pupil response? It might be that cognitive control influences attention mechanisms or vice versa (because there is some overlap in the response). Although not testing causality, this temporally resolved correlation would be an interesting way to start probing how signals might influence each other.

      Thank you for this suggestion. We now report a cross-correlation analysis (Fig. 3d) that suggests that cognitive control increases precede attention decreases at retrieval, whereas in response to a strong MPE, cognitive control increases follow attention increases. There were no significant clusters for the temporal relationships between frontal theta and pupil, nor for posterior alpha and pupil.

      (3) The relationships the authors find between brain measures and pupil components were largely not specific to mismatches/matches. However, are they specific to this task? I think it would benefit the paper to show that these relationships are potentially specific to making match/mismatch memory decisions, versus, e.g., any stimulus processing. For example, the authors could run the same analyses locked to stimuli in the study phase, anticipating a different pattern, if indeed these findings are specific to the associative memory task.

      Many of the associations between our measures indeed did not show an interaction with Mismatch. Due to jitter in the ISI in the study phase, some of the analyses in the retrieval phase cannot be performed in the exact same way for the study phase. We will leave these questions to be addressed in future research. We agree that an important question for future research is to address whether these responses depend on making match/mismatch decisions and now include consideration of this point in the Discussion: “Finally, the magnitude of observed increases in control, attention, and arousal following strong MPEs may be influenced by the decision-making process engaged when making match/mismatch judgments. Not all MPEs necessitate behavioral responses. Whether similar magnitudes in neurocognitive responses and consequences for learning are observed upon detection of an MPE, but in the absence of a decision remains unclear.”

      (4) During memory retrieval (i.e., before the probe), the authors find that frontal theta, a marker of cognitive control, was associated on a trial-by-trial basis with more posterior alpha (i.e., less alpha suppression, potentially reflecting less attention), and that this association was stronger for weaker predictions. The authors interpreted this as weaker predictions necessitating more cognitive control, and that more cognitive control was recruited specifically in trials where retrieval included less content (memory reinstatement) to attend to. Generally, cognitive control is recruited to facilitate memory retrieval. If so, one possible interpretation is that this correlation reflects cognitive control effort that has failed to produce enough memory reinstatement. The other possibility is that this correlation reflects more specific retrieval of the correct probe, without retrieval of interfering items (i.e., overall less content). I believe that the former explanation predicts that this correlation would be associated with longer RTs (more difficult decisions), while the latter predicts shorter RTs (easier decisions due to successful retrieval), at least for matches.

      Thank you for these insightful comments. Because this analysis is not key to the main questions about MPEs and given both reviewers’ concerns that the manuscript can be overwhelming for the reader given the sheer number of findings reported, we opted to move this point from the main text to the Supplement. However, following the reviewer’s guidance here, we conducted the proposed analyses and tested these alternative accounting by modeling RTs. The results favour the former interpretation:

      “Greater control being associated with less attention could reflect failure in controlled retrieval efforts to reinstate sufficient memory evidence of the probe and thus fewer retrieval products to which attention is allocated. Alternatively, greater cognitive control could increase the likelihood of retrieval success, eliciting selective retrieval of the correct probe and inhibition of interfering items. To address these alternatives, we examined how the association between frontal theta and posterior alpha related to the difficulty of a trial, as assayed by RTs. The former failure-of-control account would predict that a stronger positive frontal theta- posterior alpha association would relate to longer RTs, whereas the greater-retrieval-specificity account would predict that a stronger positive association would relate to shorter RTs. In a model predicting RTs as a function of frontal theta, posterior alpha, Strength, Mismatch, and their interactions, we found a two-way frontal theta × posterior alpha interaction (β=0.016, CI=[0.001, 0.031], p=0.033), such that a stronger positive association between frontal theta and posterior alpha predicted longer RTs. This relationship did not differ as a function of Strength (no frontal theta × posterior alpha × Strength interaction: β=-0.015, CI=[-0.033, 0.004], p=0.124), Mismatch (no frontal theta × posterior alpha × Mismatch interaction: β=-0.016, CI=[-0.050, 0.018], p=0.342), or interact with Strength and Mismatch (no frontal theta × posterior alpha × Strength × Mismatch interaction: β=0.024, CI=[-0.014, 0.062], p=0.218). Together, these outcomes support the idea that during memory retrieval, positive coupling between frontal theta and posterior alpha may reflect failure or inefficiency of cognitive control efforts to rapidly accumulate mnemonic evidence to which to attend in support of a memory decision.”

      (5) In section 3, the authors found a positive relationship between alpha during memory retrieval and PC3 during MPE. If I understood correctly, this means that less attention during retrieval (less suppression) is correlated with a stronger PC3 response. How do the authors interpret this? Maybe along the same lines as in (5), specifically retrieving the correct information (i.e., less retrieved content to attend to) means a stronger prediction, leading to a stronger MPE, and a stronger MPE response, as reflected by PC3?

      We appreciate this comment, as it highlights a need for greater clarity here. The observed relationship was actually negative (Fig. 4f), meaning that more attention during retrieval was associated with a stronger PC3 response. We suspect that the lack of clarity here may be due to the original statement that “there was a positive relationship between posterior alpha suppression during memory retrieval [and PC3 scores]”. To increase clarity, we have modified this statement to “there was a negative relationship between posterior alpha during memory retrieval [and PC3 scores]”. We interpret the negative relationship with posterior alpha (i.e., positive relationship with posterior alpha suppression) to indicate “that greater attentional allocation during memory retrieval, which occurs when memories are stronger and more retrieval products can be reinstated and attended to (Fig. 1e; Fig. 4c), predicts the magnitude of immediate pupil responses to MPEs.”

      (6) The results with subsequent memory are important and address a major gap in the field that largely did not relate neural effects of MPE to subsequent memory. However, one major limitation of the study is that the authors did not test memory for matches. I understand the logic of avoiding testing matches. Because matches were repeated more times in the study, it's not a fair comparison, and could change participants' overall criterion for old/new decisions. However, one possibility would have been to test only the weak prediction; this could have given some specificity to the neural subsequent memory findings.

      We were indeed concerned about the change in decision criterion and did not include match items for this reason. Nonetheless, this is a useful suggestion that future work could include the weak match items for a better comparison of subsequent recognition memory. We now comment on this in the Discussion: “Whether control-associated enhancements in learning are specific to learning from MPEs can be further tested in future work by including a test of subsequent memory for weak match probes as an additional control condition for comparison.”

      (7) The authors nicely characterized the different PC of pupillary MPE response. But, with respect to subsequent memory, they only present pupil size. Unless there is some methodological reason that prevents testing subsequent memory on the PC, I think this will be very informative about the potential mechanisms underlying memory of MPE.

      The pupil PCs were not associated with subsequent memory, though there are some interesting trends in the 48-delay condition which could be explored in future work. These findings are now reported in the Supplement (Supplementary Fig. 16).

      (8) This paper includes many interesting findings, and I am not sure how they all come together into a cohesive mechanistic understanding of MPE response and subsequent memory. I think the paper would benefit from either a conceptual mechanism figure or, in the Discussion, have a summary of a proposed mechanism integrating the findings together.

      We thank the reviewer for stressing this point, which also was raised by Reviewer #1. To better emphasize the novel contributions of the work and to assist the reader’s understanding of the key findings, we: (a) revised the Introduction to more explicitly describe our hypotheses about the relationships among our measures as they relate to MPE responses and subsequent memory; (b) now report mediation models to more directly test these relationships and include diagrams of the hypothesized relationships; and (c) include a schematic figure at the end of the Results to highlight the key mechanistic relationships supported by the data.

      (9) Relatedly, the section "Immediate, strength-sensitive neurocognitive impacts of MPEs" does not link the arguments to specific data points, so it's hard to follow which data specifically the authors are interpreting.

      The discussion in this paragraph rests on the Strength × Mismatch interactions observed in Figures 1d-f, summarized in the first sentence of the section. To increase clarity, we changed the title of this section to “Neurocognitive impacts of MPEs are strength-sensitive”.

      (10) If I understand correctly, the authors did not find improved memory for strong compared to weak MPE. First, I think this behavioral result should be incorporated in the main paper and in the interpretation of the results. Second, given that the neural effects the authors tested either correlated with memory for strong MPE or did not show a relationship with memory, what neural/pupil response could explain memory for weak MPE?

      Thank you for raising these points. The behavioral result and the frontal theta trial-level regression model for subsequent memory are now described in the main results and included in the Discussion. As noted in the Discussion, the trial-level regression model indicates pre-probe and post-MPE frontal theta effects may explain memory for weak (and strong) mismatch probes. While the magnitude of probe-period MPE frontal theta is additionally predictive of memory for strong mismatch probes (as indicated by the Subsequent Memory × Strength interaction in the trial-level regression model in Fig. 5c and the logistic regression in Fig. 5b), frontal theta during this period does not additionally enhance memory for strong mismatch probes above that of weak mismatch probes (Fig. 5a). We added more discussion on why memory for strong and weak mismatch probes in the current experiment did not differ. We also discuss potential directions for future research to further probe mechanisms underlying MPE-driven learning.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It is recommended that the authors determine whether formal path analyses, testing for mediation and moderation, would provide a useful approach from which to better integrate the disparate set of findings and make clear their causal/functional implications.

      At a minimum, adding a diagrammatic figure is recommended to visually depict the key components of the study (independent variables, physiological indices, outcome measures) and how they relate to each other both conceptually (ideally in a theoretically hypothesized manner) and in terms of the observed findings. Such a figure will help the reader keep track of the many types of findings and results threads, and with the goal of better organizing the results into a clearer narrative through-line.

      Thank you for the suggestions to add mediation analyses and diagrammatic figures. To more formally address our hypothesis that increases in attention/arousal following MPEs are explained in part by increases in cognitive control, we tested two mediation models: one with posterior alpha as the measure of attention and the other with pupil PC3 scores. We now diagram this hypothesis in Fig. 3a. Given the outcomes of the cross-correlation analysis suggested by Reviewer #2, we also tested an additional model where posterior alpha might explain the impact of strong MPEs on frontal theta (now diagrammed in Fig. 3e). We did not find credible evidence for an indirect effect in any of the models, suggesting that strong MPE-driven increases in control, attention, and arousal may be elicited independently. Finally, in a newly added summary diagram (Fig. 6), we highlight the observed effects of strong vs. weak MPEs on RTs, control, attention and arousal; differences in control, attention, and arousal for strong vs. weak predictions preceding the MPE; trial-level mismatch-specific or strong-specific effects on subsequent memory; and mismatch-specific interactions between processes during retrieval and responses to MPEs. We hope these revisions address Reviewer #1’s concerns and that these figures help the reader keep track of the key hypothesized relationships and main findings.

      Reviewer #2 (Recommendations for the authors):

      (1) The relationship between event segmentation and prediction errors has been reviewed recently in two papers (Nolden et al., 2024, Neuroscience & Biobehavioral Reviews; Rouhani et al., 2024, JOCN for a potentially relevant computational model). I wonder if insights from these papers can inform the Introduction/Discussion of the current manuscript.

      Thank you for these suggestions. These papers are now incorporated into the Introduction and the Discussion.

      (2) Brod et al. (2022, Psych. Bull. Rev.) have previously reported increased pupil dilations for MPE correlating with subsequent memory, specifically for strong prediction errors. I think it's worth including this paper in the Introduction as the finding is highly relevant. The Brod paper might provide more direct evidence of "MPE-related increases in pupil size" than the evidence the authors provide (p. 3).

      Thank you for this suggestion. This paper is now incorporated into the Introduction and Discussion.

      (3) I'm confused about the temporal analysis: "Trial-level regression analyses were conducted on the frontal theta, posterior alpha, and pupil time series from the associative retrieval test to identify temporal clusters that were sensitive to the factors of Mismatch (i.e., mismatch vs. match probes) and/or Strength (i.e., strong vs. weak associative pairs). For each participant and each time point, a linear regression model testing main effects of Mismatch and Strength, and a Mismatch × Strength interaction was run using R to compute beta weights for each regressor." What regression exactly was run? A separate model for each participant and time point? Across trials, then? Later, the authors mention that beta weights were averaged across participants and t-tests and permutation tests were conducted. However, if the data were averaged, what t-test was conducted? And how was the permutation test conducted? It's also unclear what the authors mean by "the sign of each participant's beta weights" - what sign?

      Thank you for raising this point. A regression was run separately for each participant and each time point, across trials: neurocognitive measure ~ Mismatch + Strength + Mismatch: Strength. Beta weights were averaged only for visualization; t-tests were conducted on each set of beta weights (across participants), separately for each time point. We modified the text of the Methods to describe our procedure more clearly.

      (4) The associative memory and recognition accuracy data are presented as d'. In addition, the authors should provide hits and false alarms to facilitate a better interpretation of the results.

      We now report these outcomes in Table 1 and refer to them in the main text.

      (5) The authors argue regarding the frontal theta that "Qualitatively, the main effect of Strength emerged later than the main effect of Mismatch, suggesting that the increase in frontal theta evoked by the probe was more sustained for weak compared to strong trials (or, as a corollary, that the greater control elicited by strong MPEs enabled more rapid resolution of conflict and ultimate choice selection)." It was unclear to me how that stems from the data.

      This statement has been removed altogether.

      (6) Especially in Figure 1, I think clarity can be improved if the authors would indicate the specific subsection they are referring to in the text, because even within, e.g., 1d, there are different graphs, so mentioning which graph is relevant for which statement would be helpful to the reader.

      Thank you for this suggestion for improving the clarity of the manuscript. Fig. 1d-f includes multiple parts because we wanted to show the raw data (the mean time series on the left) as well as the model coefficients (time series on the right). We are hopeful that, given that each subplot is titled and that the main text refers to both the subplots on the left and on the right, this will be clear to the reader as is. We welcome further guidance if this remains a concern.

      (7) This seems highly speculative to me: "Elevated PC4 scores on weak hits, where no error was made, may instead reflect retrieval practice and thus internally oriented attention. On these trials, when cueelicited retrieval may have been weaker, the probe may have provided additional support for pattern completion of the learning episode for that association, and this engagement in memory retrieval may have elicited pupil dilation (c.f., Strength effect in Figure 1f). Overall, PC4 may therefore reflect an attentional orienting response that, depending on the relative success of memory retrieval and probe identity, may direct attention internally or externally. " (p.13). In my opinion, the authors make a lot of assumptions about underlying processes. I'd consider removing.

      We removed this text.

      (8) In Figure 3c, should the x-axis be PC3 quintile? And (c) is not in the figure caption.

      Thank you for this note. We addressed these issues.

      (9) The carryover effects the authors report are interesting, but they seem detached, and it is unclear how they fit with the additional findings. In a paper that already includes many findings, I'd recommend either integrating better or removing.

      Following this guidance, we removed the readiness-to-remember findings in the interest of space and clarity.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      "Learning is a fundamental source of individuality," by Manna and colleagues, interrogates different sources of variation in individual behavior. The authors place individual flies in a Y-shaped arena, which is a common design in the field, and illuminate the arms of the Y with blue versus green light. They track the color preference of individual animals and also perform operant conditioning, meaning that they teach the fly to avoid a particular color/arm by generating a foot shock when the fly enters that arm. There are a number of things that are impressive about this setup: The authors are able to collect data on thousands of individual flies of many different strain backgrounds, and they demonstrate a strong change in color preference after conditioning. This is nice, because in past papers, visual learning ability has been modest and difficult to study. To put a number on it, in this paper, animals on average don't show a color preference at the start of the assay, spending around 30% of their time in the one arm illuminated green, and the remaining time in the two arms illuminated blue. After conditioning, the average animal spends only 23% of its time in the green arm.

      The authors run 64 animals through the assay for each of 88 wild-type strains (maybe? see Major Point 1 below) and see considerable strain-specific (genetic) variation in the change in time spent in the shocked color after conditioning. Some strains show no learning, while others spend <10% of their time in the shocked color after conditioning. They also, I believe, see that some strains have more variability across individuals, which would suggest that some strains have stronger canalization at the development or circuit function level than others, i.e., some genotypes produce more consistent copies of the individual, others less consistent copies. (Or, some genotypes produce robust circuits, and others produce noisy circuits.)

      Finally, the authors argue statistically that learning itself increases variability in individual performance. This makes a lot of sense to me intuitively. Learning changes the physical/chemical properties of circuits in the brain, and because it evolves over time and interacts with environmental variables, it seems like it should send different animals down different channels. Or, at a conceptual level, if I learn to play the piano and my sister doesn't (because of some genetic difference between us or something stochastic), this learning experience will cause all sorts of other differences in our behavior as time passes. I also think the authors do have enough data to be able to make this finding. However, the presentation of the argument in this portion of the paper is hard for me to understand, and I am not an expert in statistics, so the strength of the result is difficult for me to evaluate.

      Major points

      (1) It's difficult to track through the paper the number of animals tested for different assays. At the beginning, it says N=5632, which works out to 64 flies for each of the 88 DGRP strains. 64 happens to be the number of parallel Y arenas they have. Later in the methods, there's a description of more variation within the set of 64 for each strain, two different parent sets per strain, different sexes, conditioned and unconditioned. And, while the results text focuses on the color learning, the methods discuss additional assays (place learning, multi-day learning).

      Given the numbers, does each run of the 64 mazes include all the tested flies of one strain, or are flies of many strains included in each batch? Do different flies do different assays (color, place, multi-day), or do they all do all the assays? Perhaps there is a table including this information already in the supplement, but I recommend making it much clearer in the main results text and methods. While the dataset is large, if it is split over many conditions and/or if batch and genotype confound each other, this will affect the robustness of the results and how strong the conclusions can be.

      (2) The data presentation in Figure 1 is elegant and easy to follow, but getting into Figure 2 and subsequently, I get lost in the statistics and have trouble understanding what is being measured. My understanding of the big picture is that while genetics and individual randomness contribute a lot to behavior, the evidence for learning as an amplifier of individuality is that variance in behavior among animals of the same strain increases over time in the conditioned group (i.e., the group that is doing the most learning, or a specific kind of learning), but not in the control group. This idea is illustrated in the flattening distributions in the cartoons in Figure 1A. The authors should include graphs of the real data that use the same format as in that cartoon. Instead, the graphs present "residuals," and I don't know what those are. I suspect it's "variation left over after accounting for effects of strain and individual stochasticity." I see the residuals being tracked per strain over time in Figure 2H, but I don't see the change over time in other graphs. I'm looking for something simple like, "variation within the strain at the beginning of learning and at later time points in learning." (But I'm not sure exactly what instantaneous measurement would be the focus in longitudinal analyses of learning behavior.)

      (3) Figure 3 is a cool stab at tracking down the precise mechanism by which a stochastic environment interacts with learning to send individuals along different behavioral routes. But again, like in Figure 2, I don't have the sophisticated understanding of statistics to understand exactly what the graphs are telling me, or how they relate to the underlying measurements. I'm relying on the results text alone to reach a conceptual understanding, and just taking the graphs on trust.

      So, overall, the authors have a very nice body of work here, and with the potential to add a new facet to our understanding of the origins of diversity in animal behavior. In addition to the interpretations they focus on here, this dataset also represents an advance in studying visual associative learning in general, and quite an amazing ability to make longitudinal measurements of many behavioral decisions within the same animals. Improving the data presentation to make it easier to follow for a larger swathe of researchers, especially in figures 2 and 3, will increase its potential impact.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to test the extent to which differences in learning capacity and experience contribute to behavioural variation in a genetically identical population under identical environmental conditions.

      Strengths:

      The authors developed and used a scaled-up version of a simple two-choice behavioural paradigm, allowing them to test thousands of individuals across multiple genotypes. They then deployed clever and powerful statistical analysis methods and provided compelling evidence for a role of variability in learning in the expression of behavioural variation.

      Weaknesses:

      There are no major weaknesses, although some level of longitudinal analysis to strengthen the evidence for a strict definition of individuality would be a welcome extension of a future study. In addition, it would have been very interesting, although understandably beyond the current scope, to delineate a potential source of learning variability in the brain.

      Following our provisional response to the reviewers, we have implemented these additions to the manuscript:

      (1) We have added 7 additional tables (Table 1-6 and table 8) to the supplementary that detail how many individual flies were used in which of the seven separate experiments, how the individuals were distributed across genotypes, replicates and sexes, and how many were filtered out before the final analysis. At the bottom of each table, we added a short description of the type of experiment and a brief explanation of the filtering. The four smaller experiments where we tested the two mutant lines and one wild-type DGRP line were used primarily to test and validate the experimental platform and the behavioural paradigm used for the main experiment. In these four experiments we tested green place learning, blue place learning, green colour learning, blue colour learning in four separate batches of flies. In each of these experiments we used 192 individuals (64 individuals x 3 genotypes x 4 experiments = 768 individuals in total). In the multiday experiment we used 64 flies per genotype per each of the four groups of sequences of learning paradigms, in total 512 individual flies. Here, each of these 512 individuals were retested in different learning paradigms over 4 days (Table 5). The main experiment was the green place learning (Table 6) where we tested all 88 DGRP lines and again the two mutant lines was used to obtain the majority of the main results and conclusions (64 individuals x 90 genotypes = 5760 individuals). Lastly, additional 896 individuals were measured in the experiment using blue place learning paradigm to test the consistency of learning behaviour as opposed to colour bias within genotype (Table 8). In summary, in all experiments, we have always measured behaviour in 64 flies per genotype (full loading of the behavioural platform), and they were distributed almost entirely evenly across replicates, sexes, and conditions (control vs conditioned). No individual was reused across experiments. In most cases, after filtering the data, 60 or fewer individuals were used in final analyses. For the very few deviations from this experimental design (which occurred due to unforeseen events such as dropped/sick vials, flies flying away or accidentally squished during setup, skewed number of males and females etc.) we added a short explanation in the text below the tables. In total, across all reported experiments in this study, we measured behaviour in 7936 individuals.

      (2) We have added a schematic visual representation of classical measurement of individuality (variance of the distribution of behaviour within genotype where genetically identical individuals are raised in the same environment), entropy-based measurement of individuality (residual individuality) and the change in residual individuality, as we use them in this study (Figure 2D). We also provide a list of different DH<sub>resid</sub> measures and what distributions are being compared across the DH<sub>resid</sub> in the same figure. We hope this will serve as a more intuitive explanation of individuality and help readers interpret and follow more easily the results that we report after this figure.

      (3) In the same vein, we added another schematic visual representation to Figure 3 (Figure 3F) where we depict how distributions of individual behaviour may change with every decision and how this change translates to (or can be read out from) the change in residual individuality. We have also renamed the X axis of Figure 3E to “DH<sub>resid</sub> Start”, so that it is clearer what is measured here and matches the explanation in Figure 2D.

      (4) We have added two additional supplementary figures where the reader can inspect in more detail how the distributions of individual behaviour change longitudinally across time for each genotype in control and conditioned (Figure 3 – Figure supplement 2 and  Figure 3 – Figure supplement 3). From these figures one can glance how variance as well as the shapes of the distributions change as the flies learn in the conditioned setting, and how they remain largely the same in the control where flies behave spontaneously. We have added a sentence in the main text to introduce these figures: “We found that the distributions of individual behaviour were broader and their shapes changed substantially over the course of the experiment for the conditioned flies, and not for the control flies Figure 3 - figure supplement 2, Figure 3 - figure supplement 3).”

      (5) As noted in the first provisional response to reviewers, we changed the sentence “In every individual, behaviour is shaped by deterministic, genetic factors and by environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.” to “In every individual, behaviour is shaped by fixed genetic factors and by variable environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.”

      (6) Some sentences were edited in the results so that we can correctly refer to the newly added tables and figures. Context, meaning or interpretation of the results in these sentences was not altered.

      (7) While adding the new table references to the text, we noticed a typo that propagated in the previous version where the number of flies used in the main experiment was stated to be N= 5238, when in fact it should have been N=5239. This is now fixed.

      We once again thank the reviewers for their comments and suggestions – we believe their suggestions helped us improve the presentation and interpretability of our study and we hope the reviewers and readers will agree with this as well.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is solid, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes, although some claims about density dependence, competition, and stress-gradient mechanisms are not fully supported by the experimental design. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management on the Qinghai-Tibetan Plateau.

      Thank you very much for these positive assessments of our work. Below, we provide point-by-point responses to the comments from the two peer reviewers, and we hope these revisions are satisfied by you and the reviewers.

      Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where two herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      The study was not designed to investigate the reasons that pikas clip Stellera chamaejasme. That said, based on other studies and general knowledge of the ecology of these pikas, it is likely that they clip (although do not eat) this plant because its relatively large size hinders predator detection. This species of pika does better where vegetation height is low than where it is higher.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section.)

      We appreciate these positive comments on our work.

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications of the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists.

      Thank you for this suggestion. We have acknowledged this limitation in the Discussion by adding the paragraph below (see the Third point):

      “Despite of these, several questions deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if the documented facilitation of yak nutrition by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition? Third, our study site is located at approximately 3,200 m elevation, relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 272-286 in the Discussion section.

      Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015).

      Citation:

      Harris RB, Wang, WY, Badinqiuying , Smith AT, Bedunah DJ (2015) Herbivory and Competition of Tibetan Steppe Vegetation in Winter Pasture: Effects of Livestock Exclosure and Plateau Pika Reduction. PLoS ONE 10(7): e0132897. doi:10.1371/journal.pone.0132897

      Thank you for this suggestion. We have added more details about the study site, particularly regarding wild ungulates, in the Methods section. Specifically, we have included the sentence of “Wild ungulates, such as Tibetan gazelles (Procapra picticaudata) (Harris et al., 2015), and other small mammals such as rabbits and zokors, occur rarely in the area.”

      See these revisions in Line 333-335 in the Methods section.

      It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats).

      Thank you for the suggestion. We have acknowledged this limitation in the Discussion, by adding a paragraph as: “Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 284-286 in the Discussion section.

      Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

      Thank you for your suggestion. We have acknowledged this limitation in the Discussion, by adding the paragraph as “Second, if the documented facilitation of yaks by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition (Yang et al., 2026)?”

      See these revisions in Line 276-279 in the Discussion section.

      Reviewer #1 (Recommendations for the authors):

      Although no doubt a bit sensitive, it would have been better to reveal a bit more about how pikas were removed.

      We have provided more details about how pikas were removed in the no-pika treatment, by adding “For the no-pika treatment, pikas were trapped once every two weeks using 30 live traps (25 cm high × 25 cm wide × 40 cm long) within each plot and relocated elsewhere in the study site.” in the Methods section. We didn’t recorded how many pikas were removed from the corresponding plots, so no data were available for this point.

      See these revisions in Line 411-413 in the Methods section.

      The authors also missed a few relevant papers worth citing, including Badingqiuying, R. B. Harris, and A. T. Smith. 2018. Summer habitat use of plateau pikas (Ochotona curzoniae) in response to winter livestock grazing in the alpine steppe Qinghai-Tibetan Plateau. Arctic, Antarctic, and Alpine Research 50 (1): e1447190

      We have cited this key paper in Line 75 in the Introduction section.

      Reviewer #2 (Public review):

      Summary:

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via the use of shared resources.

      Strengths:

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good job of demonstrating the potential conservation impacts of their research.

      We appreciate these positive comments on our work.

      Weaknesses:

      What the authors tested is really cool, but their claims go far beyond what they can say based on their experimental design. For example, the authors claim to show that pika impacts on yaks display density-dependent transitions from competition to facilitation. However, their experiment only looked at the presence (at moderate densities) and absence of pikas, and they only tested for facilitation, not competition.

      The paper would also benefit from changes to the framing in the introduction and discussion. For example, the authors pitch the work as a test of the stress-gradient hypothesis. However, there is no abiotic stress gradient in the study, which is an essential component of the SGH. They also pitch the work in terms of density dependence, but there is no significant variation in population densities beyond the presence-absence binary. The paper would be stronger if they focused their framing around the literature on facilitative interactions across mammals of different size classes, especially indirect facilitation via use of shared resources, which is what this paper is really about.

      We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the Stress Gradient Hypothesis (SGH). Thus, we deleted the description on SGH. However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate densities of pikas has the best beneficial effects on yak growth. We added a separate paragraph in discussion to have a clear discussion.

      We have made the major revisions below to address these concerns.

      (1) We have revised the title into “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands ”.

      (2) We have deleted all the statements about facilitation and competition predicted by the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We added a paragraph in discussion (Line 248-259) to have a clear discussion on the humped relation between yak weight gains and pika burrow densities as “Because of the natural variations in pika density in the pika-present treatment, we were able to obtain a hump-shaped relationship between yak weight gains and pika burrow densities in these plots. Compared with the absence of pikas, the facilitative effect reached its maximum at approximately 200 burrows/ha but became competitive at densities exceeding 400 burrows/ha (Figure 3C). This result reveals that pika density modulates the net outcome for yak weight gain, with a facilitation peak at ~200 burrows/ha and a competition onset above 400 burrows/ha. Our findings offer empirical evidence for the non-monotonicity theory, under which the competition-facilitation balance varies with population density: facilitation dominates at low densities, competition at high densities, and these density-dependent shifts may underpin community stability and productivity (Zhang, 2003; Zhang et al., 2015). The theory further holds that the facilitation threshold, not the competition-facilitation transition, is the critical factor governing the stability of interacting species or communities (Zhang et al., 2015).”

      Most importantly, there are inconsistencies in what is visualized in the figures compared to the model results. For example, the results section in several places notes a lack of significant interaction terms in the model but shows interactions in the p-values on the figures.

      In the Results section, there are only two places where we discussed non-significant interactions: Line 175–177 “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Figure 3F, I and Appendix 1—figure 1, table 5,8).” and Line 190–192 “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (Appendix 1—figure 2, table 10).”.

      We have cross-checked both the Results section and the Figures sections mentioned above, and confirmed that they are consistent now.

      The authors also plot smoothed lines rather than their model results and then draw interpretations from those lines that cannot be tested in the models that they used.

      Thank you for the suggestion. Now we have added the Appendix 1—table 3 and Appendix 1—table 7 for the model results of generalized additive models (GAMs) for Figure 2C and Figure 3C that plotted with smoothed lines in Appendix 1.

      There are also missing details that are important for model interpretation, including the distributions used and the sample sizes.

      We have provided the Appendix 1—table 13 to summarize all statistical models used in the study, including the distributions used and the sample sizes in the Appendix 1.

      We have also added a sentence of “A summary of all statistical models used in the study is available in Appendix 1 table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Another major concern with experimental design is in the forage nutrient analyses. The authors picked plants along a grazing trail, then measured nutrient content without standardizing based on plant species, so any differences across treatments could be because of what they happened to grab rather than overall forage quality.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g). We have revised this section as below.

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Figure 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) Introduction

      Line 53 - I wouldn't describe small mammals like rodents as keystone species. They can have strong impacts on ecosystems, but not disproportionate relative to population size, which is a key part of that definition. It may be true when you talk specifically about pikas later on, but not small mammals as a general category.

      We have replaced this term with “consumers” here, see Line 63.

      Lines 58-61 - Good hook

      Thank you for this positive comment.

      Lines 69-74 - I don't think the stress gradient hypothesis is the right pitch for this work. The SGH posits that facilitation increases with abiotic stress, but no abiotic stressors were measured in this study. Population density interacts with abiotic factors in the SGH, but population density in and of itself is not an abiotic stressor. The papers you cite here all look at the interaction between population density and abiotic stressors (e.g., water availability). So these lines set me up to expect a stress gradient in your experiment that didn't exist, then left me confused later on. It would be better to highlight the strength of your work (lines 70-76 pose interesting questions and predictions) rather than trying to make it fit within the SGH.

      Thank you for pointing out this problem. We have deleted the statements about SGH here.

      See these revisions in Line 88-91 in the Introduction section.

      (2) Materials and Methods

      Lines 498-509 - Did you verify beforehand that no plants within the enclosures had been grazed on? How did you know that the consumption or clipping was specifically from those pikas?

      We have clarified here by adding “Before cage installation, we carefully checked the plants within each plot and removed those that had been previously grazed or damaged by herbivores.” in the Methods section. In this case, we made it sure that the consumption or clipping was specifically from those pikas with the cages.

      See the revisions in Line 349-350 in the Methods section.

      Lines 509-512 - I would be careful calling this preference. It's really just a record of what they consumed along paths they were walking, which could be about accessibility and convenience as much as preference.

      We have replaced “diet preferences” with “diet composition” here, see Line 359 in the Methods section.

      Line 514 - Clarify the specific question or hypothesis you're testing with this field survey, beyond just generally testing associations.

      We have modified the sentences here as “In July 2021, we investigated the potential facilitation of pikas on yaks mediated by the poisonous Stellera forbs under unmanipulated field conditions in the study site.”.

      See Line 365-366 in the Methods section.

      Lines 532-534 - The intro for this paper sets it up to be about density-dependent movement from competition to facilitation, but the experiment here is set up to compare pika presence/absence. The framing of the paper needs to be adjusted to better align with this experimental design.

      The same issue as mentioned above. We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the balance between competition and facilitation as predicted by the Stress Gradient Hypothesis (SGH). However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate density of pika has the best benefical effect on yak. We added an separate paragraph in discussion to have a clear discussion about this point in the Discussion section.

      We have made the major revisions below to address this concern.

      (1) We have revised the title as “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands”.

      (2) We have deleted the statements about facilitation and competition and the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We kept the discussion on the humped relation between yak weight gains and pika burrow densities (Figure 3C). We added an separate paragraph in discussion to have a clear discussion about this point. For details, see Line 248-259 in the Discussion section.

      Lines 566-569 - Why did you need to simulate pika clipping when you already had pika presence/absence treatments?

      We conducted Stellera removal treatment by simulating poisonous plant clipping behaviors of pikas because we want to confirm that the shifts in abundance of this dominant poisonous plant species is the key mechanism in driving pika-yak facilitation in our system. If we simply looked the differences in yak weight gain in the pika presence/absence treatments, it should be difficult to secure the underlying mechanism. In addition to the reduction in abundance of the poisonous Stellera, pikas may cause a variety of shifts vegetation properties including plant productivity and diversity, and soil disturbances that can exert direct and indirect effects on yak foraging activities, and thus their weight gains.

      Lines 566-569 - Is there another citation you can give showing that Stellera forbs taller than 20cm are both preferred by pikas and exert greater impacts on plant/animal communities? Those are big assumptions that need to be better supported or explained. If there was a logistical reason that you didn't remove all of the smaller forbs, that needs to be laid out as well.

      We have added one new citation here to support this method here. We have revised this section as “To simulate the clipping behavior of pikas, we clipped only those Stellera forbs exceeding 20 cm in height. This threshold was chosen based on previous observations that pikas preferentially target large forbs of this size (Liu et al., 2009).”

      Liu W, Zhang Y, Wang X, Zhao JZ, Xu QM, Zhou L. 2009. The relationship of the harvesting behavior of plateau pikas with the plant community (In Chinese). Acta Theriologica Sinica 29:40-49.

      Also, we have deleted the sentence of “and can exert significant impacts on the plant community and on yak grazing behaviors (Z.Z., field observations)” mentioned above, because these patterns were observed only by the authors in the field and lack supporting data.

      See these revisions in Line 419-422 in the Methods section.

      Line 587 - sample size per treatment? Was it consistently one species that you measured, and if so, which one? If you collected different species or a mix of species across treatments, then you can't really compare the nutrient values because you haven't accounted for interspecific variation.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g).

      We have revised this section as below:

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Fig. 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Lines 603-613 - Please provide the dependent variables in each model, as well as any interaction terms, in addition to the random effects. Please also state explicitly what you were trying to test with each of these models.

      There are two models that use the tweedie family (forb and sedge bite rate). Indeed, we need to include the tweedie power parameter to help understand the mixture of the three families. We have included the p-value (power) in the Appendix 1—table 13 in the Appendix 1. We chose to use tweedie because a normal gaussian family model fitted the results poorly.

      We have added the Appendix 1—table 13 in the Appendix 1, which provides the summary of all statistical models used in the study, including response variables, model type, distribution family (with Tweedie power parameter where applicable), interaction terms, random effects structure, and sample sizes.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Lines 613-614 - Tweedie is a category of distributions that includes quite a few different options, including Gaussian, Poisson, and Gamma distributions, some of which are normal and some of which are not. So, justifying the use of Tweedie distributions in your model structure doesn't really make sense, and it doesn't really tell me which distribution each model pulled from. Please clarify specifically which distributions you used for each model and why.

      We have now added the reason why we used Tweedie distributions by adding the sentence of “There were two models that used the tweedie family (forb and sedge bite rate). We chose to use tweedie because a normal gaussian family model fitted the results poorly” in Line 468-470 in the Statistical analyses section.

      We have also included the p-value (power) for forb and sedge bite rate in the Appendix 1—table 13 in the Appendix 1.

      Lines 615-618 - Provide citations for R packages described in the text.

      We have provided all the related citations for all R packages described in the text, as listed below.

      glmmTMB: Brooks, M. E., Kristensen, K., van Benthem, K. J., Magnusson, A., Berg, C. W., Nielsen, A., Skaug, H. J., Maechler, M., & Bolker, B. M. (2017). glmmTMB Balances Speed and Flexibility Among Packages for Zero-inflated Generalized Linear Mixed Modeling. The R Journal, 9(2), 378–400. https://doi.org/10.32614/RJ-2017-066

      mgcv: Wood, S.N. (2017). Generalized Additive Models: An Introduction with R (2nd edition). CRC Press. AND Wood, S.N. (2011). Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 73(1), 3–36. https://doi.org/10.1111/j.1467-9868.2010.00749.x

      DHARMa: Hartig, F. (2022). DHARMa: Residual Diagnostics for Hierarchical (Multi-Level / Mixed) Regression Models. R package version 0.4.6. https://CRAN.R-project.org/package=DHARMa

      tidyverse: Wickham, H., Averick, M., Bryan, J., Chang, W., McGowan, L.D., François, R., Grolemund, G., Hayes, A., Henry, L., Hester, J., Kuhn, M., Pedersen, T.L., Miller, E., Bache, S.M., Müller, K., Ooms, J., Robinson, D., Seidel, D.P., Spinu, V., Takahashi, K., Vaughan, D., Wilke, C., Woo, K., & Yutani, H. (2019). Welcome to the tidyverse. Journal of Open Source Software, 4(43), 1686. https://doi.org/10.21105/joss.01686

      See these revisions in Line 472-475 in the Statistical Analyses section.

      (3) Results

      Line 127 - Somewhere in the intro or methods, describe what clipping is and why the pikas do it.

      We have provided more details about the clipping behaviors of pikas by adding “Notably, pikas often clip (although do not eat) the wolf poison S. chamaejasmehas because its relatively large size hinders predator detection (Fan et al., 1998).” in Line 330-332 in the Methods section.

      Line 128 - Change from "preferred" to "consumed greater proportions of"

      See the correction in Line 146 in the Results section.

      Line 136 - You measured yak weight once a month, so give the result in monthly weight gain rather than daily.

      Here we preferred to keep the unit of daily weight gains (converted from monthly ones), as this is the standard presentation for livestock growth performance, also see Fig. 1 in Odadi et al., 2011 Science’s paper.

      W. O. Odadi, M. K. Karachi, S. A. Abdulrazak, T. P. Young, Science 333, 1753–1755 (2011).

      Lines 138-140 - The linear models, as you described them in the methods (lines 603-620), don't test for hump-shaped relationships. Please update the methods to explain how you tested the density relationship and how you got this interpretation.

      The description for Figure 3C here showed estimates from a GAM (family: gaussian), and we did not use a linear model in this figure.

      We have now added a new model summary Appendix 1—table 7 for this Figure 3C in the Appendix 1.

      Line 145 - I don't think you can claim that the total available forage was more nutritious for yaks, as you picked plants that yaks happened to be chewing along a grazing path. You would need to take samples from a consistent set of plant species at random locations to make this claim.

      Sorry for this confusion. We modified “the total available forage” as “the major available forage” here (see Line 172) because we collected the same forage plant species and analyzed their nutrients.

      The same as mentioned above, to clarify the sampling methods, we have also revised this point in the Methods section to provide more detail on how forage samples were collected and their quality were analyzed. See these revisions in Line 439-447 in the Methods section.

      (4) Discussion

      Line 166 - Need to address inconsistency in how you talk about pika density. Here you talk about the impacts of pikas at moderate densities, which I think is a fair claim. Elsewhere, you talk about density-dependence, which I don't think you really measured, given that all your sampling was either in the absence of pikas or within a narrow window of densities that can all be categorized as moderate.

      Done! As mentioned above, we have removed the term of “density-dependence” in the whole manuscript, but keep the term of “moderate density” in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph was removed here), and the References section.

      Line 176-178 - This is a really cool finding.

      Thank you for this positive comment.

      Lines 179-181 - You didn't test competition between plant species, and you didn't measure light, soil moisture, or soil nutrients. So you can suggest competition as a potential mechanism, but you can't say definitively that's what is happening.

      We now have lowered our tone here as “We speculate that these improvements in food availability and nutrition for yaks may be due to the release of grasses and sedges from competition with the forbs for limiting above- and below-ground resources”.

      See the revision in Line 208-211 in the Discussion section.

      Lines 184-186 - You did a good job of it here, suggesting a likely potential mechanism at play without claiming it is for sure happening when it hasn't been measured.

      Thank you for this positive comment!

      Lines 188-199 - Strong paragraph. The impacts of large herbivores on smaller animals are well-studied, but reciprocal impacts are often overlooked.

      Thank you for this positive comment!

      Lines 201-205 - You didn't test the stress gradient hypothesis because there was no abiotic gradient. You also did not take any measurements during outbreaks, so you cannot claim to have compared low-moderate to outbreak pika densities. I think the paper would be much stronger if you removed the stress gradient hypothesis and instead focused more on the literature around facilitation between mammals of different body sizes, as you do in lines 205-209.

      We agreed that our work didn’t specifically design to test the stress gradient hypothesis (SGH) between pikas and yaks, so we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Line 209-215 - Again, you didn't test a competition-facilitation balance because you never tested or demonstrated competition. One of the main strengths of this paper is demonstrating facilitation, so build on that strength rather than referencing things you didn't measure.

      The same as mentioned above, we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Lines 218-222 - Not an accurate description of the relationship between herbivore diet and body size. Larger herbivores typically tolerate lower-quality plants in order to consume sufficient calories, but plenty of them do this via mixed feeding. Grazing in large herbivores and livestock is usually due to specifics of the digestive tract (e.g., hindgut fermentation) rather than specifically about body size.

      We have deleted this description of the relationship between herbivore diet and body size here. Instead, we have modified this statement as “The coexistence of a diverse of herbivore species with different diet selections and size classes can lead to an “compensatory effect” on grass and forb biomass that helps to maintain a balance and diverse plant community” in the Discussion section.

      See these revisions in Line 234-237 in the Discussion section.

      Lines 245-251 - Paragraph addresses an important point. Lines 247-249, though, overstate what you measured. There's no measurement of livestock production or biodiversity in the study.

      We have replaced the term of “livestock production and biodiversity” with “livestock growth performance” here. See Line 288-298 in the Discussion section.

      (5) Figures

      Figure 2C-D - This applies to all figures, but you need to plot the best-fit line generated from your model instead of using geom_smooth, which is what these lines look like. You can do this using functions like predict or ggpredict. Using these smoothed lines implied non-linear relationships that you didn't actually test for.

      We have redrawn Figures 2C and 2D to use model estimates directly and have included the related model summaries as Appendix 1—table 3 and table 4 in the Appendix 1.

      Figure 3B - This figure doesn't show yak weight gain in the presence of pikas. Instead, it shows weight loss when pikas are absent. It's a subtle difference, but very important for interpretation. Yak can maintain weight just fine without pikas as long as Stellera are absent, too. Your results consistently show no Pika x Stellera interactions, but that doesn't match your significance values here. Need to double-check and explain that.

      We have revised the descriptions for Figure 3 and 4 in the Results section, by emphasizing that the absence of pikas REDUCED weight gains of yaks, INCREASED toxic plant abundance, and REDUCED the quantity and quality of palatable grasses and sedges. We have revised these sections as below:

      Abstract (see Line 51-53)

      “Compared to the pika-present treatment, pika removal dramatically increased cover of the poisonous Stellera forbs by two-fold, reducing the abundance and protein content of palatable grasses and sedges, yak foraging efficiency, and yak weight gain by up to 42%.”

      Results (see Line 153-167, Line 169-175, Line 184-192)

      Also, we did find significant Pika x Stellera interactions for yak weight gains, we have provided these details in the Appendix 1—table 5 and table 6 in the Appendix 1.

      (6) Recommendations

      Lines 245-257 - Would recommend combining the last two paragraphs into one.

      We have combined the last two paragraphs, see Line 288-298 in the Discussion section.

      Line 530 - Can you replace large with a more precise measure of area?

      We are unable to provide a more precise measure of area here, so we have deleted the description of “in a large area”, but we have also added the note of “in the study site” by the end of the sentence to better describe the location of the plots.

      See the revision in Line 413 in the Methods section.

      Line 541 - Does this mean +/- 7.8 standard deviations? If so, how big is that range in kg?

      Here should be “115±7.8 kg”, we have done this correction in Line 392 in the Methods section.

      Figure 3 - Would be helpful to use "Stellera/pika present" and "Stellera/pika absent" rather than saying "No Stellera/pika" since you also use the No. abbreviation for numbers a lot in this figure.

      We have revised all the related Figures for this issue in Figure 3, 4, and Appendix 1—figure 1,2.

      Figure 3F and I - Show letters for significance on these two plots as well, even if it is just a row of a's.

      We have redrawn Figure 3F,I to address this issue.

      Figure S1 - Applies to all boxplots. Be consistent about showing significance, even if it is a row of a's indicating no difference between treatments.

      We have re-drawn Appendix 1—figure 1,2 to address this issue.

      Table S1 - Something happened with the line numbers, so they are in the table instead of on the left side.

      We have fixed this problem for Appendix 1—table 1 in the Appendix 1.

      Table S1 - Applies to all tables. Include the type of model that you ran (including distribution if not Gaussian) in the table legend.

      We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Table S9 - This legend has a good description, including the type of model you used and what you were testing. Apply this more detailed legend to the rest of the tables.

      Again! We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides new insight into the regulation of cell organization and division in Trypanosoma brucei through the control of a kinesin motor protein by a polo-like kinase. The authors present solid evidence from rigorous biochemical and imaging analyses showing that phosphorylation modulates kinesin function and cellular organization. However, direct in vivo evidence that PLK phosphorylates kinesin-G is lacking.

      We performed experiments to investigate the effect of PLK inhibition on the phosphorylation of KIN-G in vivo in trypanosome cells by immunoprecipitation and mass spectrometry. The new results showed that treatment of trypanosome cells with GW843682X, a potent TbPLK inhibitor validated previously in procyclic trypanosomes, reduced the phosphorylation levels on Thr301 and Ser569 of KIN-G by ~27% and 100%, respectively. The partial reduction in Thr301 phosphorylation after GW843682X treatment could be attributed to slower dephosphorylation of phosphorylated Thr301 after GW843682X was added to the cell culture. Nonetheless, these new results demonstrated that KIN-G is an in vivo substrate of PLK.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript identifies the orphan kinesin KIN-G as a substrate of Polo-like kinase (TbPLK) in Trypanosoma brucei and demonstrates that phosphorylation of Thr301 inhibits KIN-G microtubule binding and disrupts its cellular function. Using a combination of in vitro kinase assays, phosphosite mapping, microtubule binding and gliding assays, and in vivo complementation with phosphomimetic and phosphodeficient mutants, the authors link TbPLK-mediated regulation of KIN-G to defects in centrin arm integrity, FAZ elongation, Golgi organization, flagellum positioning, and division plane placement. The study provides a mechanistic advance in understanding how TbPLK regulates centrin arm biogenesis and integrates KIN-G into the growing regulatory network controlling hook complex and FAZ assembly. Overall, the work is technically strong, internally consistent, and builds logically on previous studies from this group and others.

      Strengths:

      A major strength of the manuscript is the clear mechanistic link between phosphoryltion of Thr301 and loss of microtubule binding activity. The use of phosphomimetic (T301D) and phosphodeficient (T301A) mutants in an RNAi-rescue framework provides a clean and convincing demonstration of functional relevance in vivo. The integration of biochemical assays with detailed cell biological phenotyping (centrin arm length, FAZ elongation, basal body segregation, and cytokinesis markers) is particularly effective and makes the central conclusion robust. The observed phenotypic cascade from centrin arm defects to FAZ and division plane abnormalities is also well aligned with existing models of trypanosome morphogenesis.

      Weaknesses:

      My (more or less main) concern relates to the interpretation of the Golgi phenotype. The conclusion that phosphorylation of KIN-G "impairs Golgi biogenesis" is currently based on fluorescence microscopy using TbGRASP and Sec13 markers and on quantification of the number and distribution of Golgi/ERES puncta in binucleated cells. While these data convincingly demonstrate altered Golgi/ERES number and spatial organization, they do not distinguish between true defects in Golgi biogenesis or duplication and alternative possibilities such as fragmentation, vesiculation, or mislocalization of Golgi membranes. Given the central role of Golgi-centrin arm organization in the proposed model, ultrastructural analysis (for example, by EM or electron tomography) would greatly strengthen this aspect of the study by providing direct evidence for structural alterations of the Golgi and its association with the centrin arm and ERES. Such data would elevate this part of the manuscript from a descriptive fluorescence phenotype to a true structural cell biological insight. I appreciate that this experiment goes beyond the current dataset, but it would substantially enhance the mechanistic depth of the Golgi-related conclusions and strengthen the causal chain linking centrin arm defects to Golgi abnormalities. However, I have to confess, the inclusion of such data would make this reviewer particularly enthusiastic about the work. If this is not feasible, I would recommend tempering the wording of "Golgi biogenesis" to a more conservative description, such as altered Golgi organization or duplication, and explicitly acknowledging the limitations of fluorescence-based analysis for this conclusion.

      Thanks for these very constructive comments, which are very well taken. We totally agree with this reviewer on these points. Since it is not feasible for us to perform EM or electron tomography, we have revised the manuscript to describe the effect of KIN-G phosphorylation on the Golgi as “altered Golgi duplication” rather than “Golgi biogenesis”. We also explicitly acknowledge the limitations of fluorescence-based analysis of the Golgi for this conclusion and suggest that further characterization with EM or electron tomography would allow one to reveal the potential structural alterations of the Golgi and its association with the centrin arm.

      An additional conceptual point concerns the dual role of TbPLK in centrin arm regulation. TbPLK is known to promote centrin arm biogenesis through phosphorylation of TbCentrin2, yet in this study, TbPLK phosphorylation of KIN-G negatively regulates centrin arm assembly. This dual positive and negative regulatory role is intriguing but could be discussed more explicitly. The manuscript would benefit from a clearer conceptual framework addressing how phosphorylation of KIN-G might serve as a temporal or spatial switch to restrain KIN-G activity at specific stages of centrin arm assembly.

      This is a great point. However, we are not sure whether the previous work on TbPLK phosphorylation of TbCentrin2 could lead to the conclusion that TbPLK promotes centrin arm biogenesis through phosphorylation of TbCentrin2. In the published work (de Graffenried et al., MBoC, 2013), trypanosome cells expressing the phospho-deficient mutant TbCentrin2-S54A only showed minor growth defects, exhibiting growth defects after 5 days (de Graffenried et al., MBoC, 2013). Cells expressing the phosphomimic mutant TbCentrin2S54D, however, showed very strong growth defects (de Graffenried et al., MBoC, 2013). The effects of TbPLK phosphorylation on TbCentrin2 appear to be quite similar to that of TbPLK phosphorylation on KIN-G, although the KIN-G-T301A mutant does not have growth defects (up to 5 days in our experiments). It appears that the primary role of TbPLK in regulating TbCentrin2 and KIN-G is negative regulation. Nonetheless, we have added more discussion on these regulatory roles of TbPLK in the revised manuscript.

      Finally, a schematic model summarizing the proposed regulatory pathway from TbPLK phosphorylation of KIN-G to centrin arm assembly, FAZ elongation, division plane placement, and Golgi organization would aid the reader.

      Thanks for this suggestion. We made a schematic model to summarize the roles of KIN-G and its regulation by TbPLK. This is included in Figure 8.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as a misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding, and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei supports the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do not formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      This is a great point that is very well taken. We also had been puzzled by the observation of no growth defects of T301A mutant. This comment enlightened us. From the new experiments we performed to compare the phosphorylated peptides of KIN-G in cells treated and non-treated with the PLK inhibitor GW843682X, we calculated the percentage of phosphorylated Thr301 in non-GW843682X-treated cells. We found that the percentage of peptides containing the phosphorylated Thr301 is ~14% of the total Thr301-containing peptides (Fig. 1H). This result indicates that T301-P is indeed a small minority of the population in the asynchronous trypanosome cells. It is possible that T301 phosphorylation may occur at a specific cell cycle stage such as early S-phase, during which PLK and KIN-G co-localize at the centrin arm.

      (b) Published work links PLK to cell division, FAZ elongation, etc.. The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc..

      Yes, previous work discovered essential roles of TbPLK in basal body segregation, centrin arm biogenesis, FAZ elongation, and cytokinesis. These functions of TbPLK correlate with TbPLK’s localization to multiple subcellular structures, the basal body, the centrin arm, and the new FAZ tip, and are attributed to the regulation of its substrates at these structures. At the basal body, TbPLK phosphorylates SPBB1, which is required for basal body segregation. Defects in basal body segregation can lead to defective flagellum positioning and FAZ elongation. At the new FAZ tip, TbPLK regulates the cytokinesis regulator CIF1, which is required for cytokinesis. At the centrin arm, TbPLK phosphorylates TbCentrin2 at S54 and KIN-G at T301 (and S569, which was newly identified as an in vivo TbPLK site and has not yet been characterized). However, cells expressing TbCentrin2-S54A have very weak growth defects, and cells expressing KIN-G-T301A have no detectable growth defects. In contrast, cells expressing TbCentrin2-S54D and cells expressing KIN-G-T301D have strong growth defects. Therefore, the essential role of TbPLK in centrin arm biogenesis apparently is not attributed to the phosphorylation of TbCentrin2 and KIN-G. It is possible that phosphorylation of other centrin arm-localized protein(s) by TbPLK may be essential for centrin arm biogenesis, but this possibility remains to be explored.

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      We performed experiments and presented the data in Fig. 1H. We also included commentary in the revised manuscript on the points about the potential role of T301 phosphorylation. Thanks for these great comments that significantly improved the manuscript.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      We treated trypanosome cells with a potent PLK inhibitor GW843682X, which was previously demonstrated to inhibit TbPLK activity in vitro and mimic TbPLK knockdown in vivo in trypanosomes, and immunoprecipitated KIN-G for mass spectrometry. We compared the KIN-G peptides identified by mass spectrometry from trypanosome cells treated with or without GW843682X, and found that two phosphosites (T301 and S569) were reduced by ~27% and 100%, respectively, after GW843682X treatment (Fig. 1G). These results provided evidence to support that TbPLK phosphorylates KIN-G in vivo.

      Reviewer #3 (Public review):

      Summary:

      Here, the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during the early S-phase of the cell cycle. Centrin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescence to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      Some of the broader conclusions are not directly supported by the data. For example, the title states "Polo-like kinase phosphorylation of the orphan kinesin KIN-G negatively regulates centrin arm biogenesis in Trypanosoma brucei," but the data do not directly address the specific role of TbPLK in phosphorylating KIN-G in cells. Moreover, some of the more specific conclusions in the paper, for example, that "phosphorylation of KIN-G" causes various cellular defects, are a bit of an overstatement. The supporting data rely on the expression of a phospho-mimetic mutant of KIN-G. Presumably, phosphorylation in cells is a normal part of KIN-G regulation, and it is not just phosphorylation, but rather hyperphosphorylation that is being mimicked by the mutant. Some rewording of the specific conclusions is warranted, and the broader conclusion would be better supported with additional experimental evidence.

      This is a great point that is very well taken. We performed new experiments to address the in vivo phosphorylation of KIN-G by TbPLK. We treated cells with a potent PLK inhibitor, GW843682X, and then immunoprecipitated KIN-G for mass spectrometry to identify changes in phosphorylation. We found that the phosphorylation levels of T301 and S569 were reduced by ~27% and 100%, respectively, confirming that these two sites are in vivo TbPLK phosphosites.

      We also calculated the ratio of phospho-T301 versus non-phospho-T301 in non-treated cells and found that phospho-T301 accounts for ~14% of the total KIN-G protein. This new result suggests that it is the hyperphosphorylation that causes growth defects. We have revised the manuscript accordingly to reflect this point.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Several statements use rather strong causal language (for example, "thereby impairing Golgi biogenesis, FAZ elongation, and division plane placement"). While the phenotypic correlations are convincing, direct causality is largely inferred from prior literature. Slightly tempering this wording could improve precision. It would also be helpful to state explicitly in figure legends the number of cells analyzed per condition and the number of independent experiments for each quantified phenotype.

      Thanks for these constructive comments, which we agree and appreciate greatly. We have revised the manuscript accordingly to improve precision. The total number of cells analyzed, and the number of independent experiments were included in the figure legends.

      Reviewer #2 (Recommendations for the authors):

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      Thanks for this comment. We have deleted the discussion about TbPLK activities that were published previously.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Figure 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Figure 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces the motility of KIN-G. These statements are contradictory.

      We meant to say that there was a slight but insignificant effect. We agree that such statements are somewhat contradictory and, hence, have been deleted. Thanks.

      (3) p. 8, and Figure 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      We have deleted the wording “ventral side”, as it is not necessary. Thanks.

      (4) Figure 7B. Please explain the labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      Thanks very much. We included two sentences in the revised manuscript to explain this point.

      The sentences read as follows: “The nascent posterior is formed near the mid-portion of the NFD cell through microtubule bundling and cytoskeleton remodeling during late stages of the cell cycle (Wheeler et al., 2013). Consequently, the NFD cell inherits the old, existing cell posterior, whereas the OFD cell inherits the newly formed or nascent cell posterior.”

      (5) Figure 4, 5, and 7: "% Cells" is reported. Please indicate the total number of cells that were examined.

      The total number of cells were included in the figure legends.

      Reviewer #3 (Recommendations for the authors):

      (1) The manuscript should be carefully edited for minor grammatical errors.

      Thanks. We have carefully proofread the manuscript and corrected the grammatical errors.

      (2) A general conclusion is that TbPLK phosphorylation of KIN-G in cells is critical for regulating its motor activity. However, this relies on the expression of phospho-mimetic mutants, which bypass TbPLK. Thus, there really is no direct evidence provided to support the specific role of TbPLK other than the in vitro phosphorylation data. Some additional experiments to assess the specific role of TbPLK in phosphorylating KIN-G in cells would lend support for the general conclusion. Is it possible to deplete or inhibit TbPLK and show that this impacts the phosphorylation of KIN-G in cells?

      This is a great point that is very well taken. We performed a new experiment by inhibiting TbPLK with a potent PLK inhibitor, GW843682X, and then immunoprecipitating KIN-G for mass spectrometry to identify the phosphorylation levels before and after GW843682X treatment. We were able to confirm that T301 and S569 phosphorylation was reduced after treatment. This confirms that TbPLK phosphorylates T301 and S569 of KIN-G in vivo in trypanosome cells.

      Specific:

      (1) Figure 2: For microtubule gliding assays, representative videos should be included as supplementary data. Also, when the data do not show a significant difference between KIN-G and KIN-G + TbPLD-K70R, then the authors should not state there is a slight difference, as this is not supported by the data.

      We have deleted the statement. We have included the representative videos for all the experiments presented in Figures 2 and 3. Thanks.

      (2) Figure 3: As stated above, for microtubule gliding assays, representative videos should be included. Moreover, for KIN-G-T301A, the authors say that the gliding activity was "moderately, but insignificantly, reduced." Again, if the difference is not significant, it cannot be concluded that there is a difference compared with the wild-type protein.

      We have deleted the statement. Thanks.

      (3) Some of the headings in the Results section are not accurate. Regarding the in vivo results, the heading "Phosphorylation of Thr301 in KIN-G by TbPLK causes defective cell proliferation" seems to be an overstatement. It may be more accurate to state that "Expression of a phospho-mimetic mutant of KIN-G causes defective cell proliferation." The same comment applies to the other headings that follow this one. Presumably, there is a population of phosphorylated KIN-G in cells, and phosphorylation/dephosphorylation is a normal part of its regulation. In the text, the authors might more accurately conclude that hyperphosphorylation causes the defects they are seeing.

      This is a great point. We agree and as we responded above, we have revised the manuscript, from the title to the main text, to reflect the point that hyperphosphorylation of T301 by TbPLK causes the defects. Thanks very much for these great comments!

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Response to the points raised by the reviewers.

      Once again, we would like to thank the reviewers for their comments. We have systematically addressed their concerns, as detailed below.

      Reviewer #1

      Evidence, reproducibility and clarity

      This study demonstrates that BICD2, previously known as an adaptor protein for dynein, is involved in regulating centriole engagement during mitosis. First, using different antibodies, it was shown that BICD2 localizes near the mother centriole, as observed by super-resolution microscopy. During G1 and S phases, BICD2 localizes slightly outside the Cep152 ring, while in G2 to mitosis, it localizes near the cartwheel component SAS-6. Moreover, analysis of deletion mutants revealed that BICD2 localizes to the centrosome in a CC domain-dependent manner at the C-terminal end. The localization pattern resembling a ring in the cytoplasm was also observed through the CC3 domain. Next, BICD2 knockout (KO) cells were generated to investigate centriole dynamics. In BICD2 KO cells, the distance between the mother and daughter centrioles was observed to increase from G2 to mitosis compared to controls. Along with this, early centriole disengagement and centriole amplification phenotypes were observed. The increased distance phenotype between centrioles was rescued in BICD2 wild-type (WT) and mutant forms lacking the CC1 domain at the N-terminus, suggesting that this function of BICD2 is independent of dynein. Additionally, BICD2 mutants mimicking phosphorylation at the C-terminus showed reduced centrosome localization and were unable to rescue the phenotypes seen in BICD2 KO cells.

      While the study clearly demonstrates BICD2's contribution to centriole engagement, the underlying mechanisms of how BICD2 is involved in centrosome localization and centriole engagement remain unclear. As it is anticipated that the function of BICD2 is independent of dynein, further exploration of this unknown mechanism would enhance the value of the paper. Below are the concerns that should be addressed, including new experiments.

      Main Points:

      1. __ Fig. 1-3: Regarding the localization of BICD2 to centrioles, during the G1-S phase, its localization appears to overlap with PCM. Experimental investigation should be performed to examine whether BICD2's centrosomal localization is influenced by knockdown of PCM components like PCNT, Cep192, or Cep152.__ We now show that BICD2 localization does not depend on pericentrin (Supplementary Figure S4B). We also show that the two proteins do not colocalize (Supplementary Figure S4A and S4D) and are functionally independent (Figure 5).

      We also show that BICD2 localization does depend on the torus protein CEP152 (Figure 8B). Importantly, our data indicate that BICD2 interacts with the N-terminal region of CEP152, suggesting that this interaction places BICD2 at the outer region of the torus (Figures 8C and 8D). We propose that this provides a mechanistic basis for BICD2 function in maintaining mother-daughter centriole engagement.

      __ Fig. 7: The experiments using BICD2 mutants suggest that the function of BICD2 here is independent of dynein. To further investigate whether BICD2's role in centriole engagement is independent of dynein, experiments should be conducted to examine the effect of dynein knockdown on BICD2 localization to the centrosome and centriole engagement.__

      Using Dynapyrazole-A, a fast-acting and potent dynein inhibitor, we demonstrate that acute inhibition of dynein motor activity affects neither the centrosomal localization of BICD2 during G2 and M phases (Supplementary Figure S3A) nor centriole engagement (Supplementary Figure S3B). Together with experiments using BICD2 mutants deficient in dynein interaction (Figure 7B), these data compellingly demonstrate that the recruitment and function of BICD2 at the centriole is dynein-independent.

      __ Fig. 4: The CC4 domain at the C-terminus of BICD2 is important for its centrosomal localization, but identifying the binder/recruiter responsible for BICD2's centrosome localization would be desirable.__

      We thank the reviewer for prompting us to investigate this further. We are delighted that we now identify the torus protein CEP152 as the BICD2 binder/recruiter at the centriole (Figure 8). As we discuss in the manuscript, our observation that the outward-facing N-terminus of CEP152 interacts with the C-terminal region of BICD2 provides a mechanistic basis for understanding BICD2's role in maintaining mother–daughter centriole engagement.

      __ Fig. 7: Rescue experiments using BICD2 mutants suggest that BICD2's functional domains are critical. Further experiments by creating mutants missing parts of CC2 or CC3 could identify functionally important domains of BICD2 by observing any loss-of-function phenotypes at the centrosome.__

      We fully appreciate the reviewer’s suggestion to examine the roles of the CC2 and CC3 regions. We believe that these domains, and particularly the unstructured loop within CC3, are important for both BICD2 localization, function and regulation at the centrosome. However, given that our current data already establish a clear mechanism for BICD2 centriolar recruitment via CEP152 and the CC4 region, we feel these additional structural studies fall outside the core scope of the present manuscript. We hope the reviewer agrees that the current evidence provides a robust foundation for our conclusions, and we look forward to addressing the roles of CC2 and CC3 in a dedicated future study.

      __ Fig. 6: Regarding the BICD2 KO cell phenotype, is there experimental evidence showing an increase in centriole number during mitosis? For instance, while no abnormality in centriole number may occur during G2, a trend of increase in mitosis should be experimentally demonstrated. Also, how should the slight differences in phenotypes between Ndelta4 and Ndelta5 BICD2 KO cells be interpreted?__

      We thank the reviewer for highlighting this point, but we would like to clarify that we do indeed observe a significant increase in centriole number during both mitosis and G2 phase across multiple cell lines in our BICD2 KO models and RNAi experiments (Figure 4E, RPE-1 KO cells, and 4G, U2OS cells, RNAi) and G2 (Figure 5C, both RPE-1 and U2OS, RNAi). As the main text was not explicit enough on this point, we have revised the manuscript to describe these observations more clearly.

      Regarding the phenotypic differences between BICD2 KO lines, we assign them to the clone-to-clone functional heterogeneity often seen in CRISPR/Cas9-generated cell lines. Importantly both clones show a consistent, statistically significant phenotype (e.g., impaired engagement and increased centriole numbers) compared to wild-type controls, confirming that the overall defect is robust and specific to BICD2 loss. We have added a clarifying note on this in the revised manuscript: “Figure 4F; we assign the differences between BICD2-/- cell lines to standard clone-to-clone phenotypic heterogeneity often seen in CRISPR/Cas9-generated cell lines.

      __ Fig. 8: Regarding the phosphorylation of BICD2 at the C-terminus: The phenotypes of mutants where these two phosphorylation sites are changed to alanine should be experimentally observed. It is expected that the removal of BICD2 from the centrosome during mitosis could be rescued. Additionally, the effect of PLK1 or CDK1 inhibitors on the removal of BICD2 from the centrosome should be investigated.__

      We agree with the reviewer that phosphonull mutants should be added to these experiments. As mentioned above we have decided to remove the preliminary data regarding BICD2 phosphorylation from the manuscript data to present a more comprehensive, dedicated study on BICD2 phosphorylation in the near future. In fact, we have already performed the suggested experiments, including the phosphonull mutants and kinase inhibitor treatments, and would be glad to share these additional results if the reviewers would find them helpful. Interestingly, our experiments show that BICD2 centrosomal amounts are not affected by PLK1 inhibition (using BI 2536); CDK1 inhibition (RO-3306), although not significatively changing the amount of BICD2 at the centrosome, slightly diminishes it. We currently favor a model in which BICD2 is predominantly regulated by CDK1, and we are actively defining the precise molecular mechanism governing this regulation.

      Minor Points:

      __ Fig. 1-3: During G1 and S phases, BICD2 localizes near the mother centriole, and from G2 onward, it colocalizes with SAS-6. How can this be explained?__

      We currently do not have a clear explanation for this transition, as our focus has been in understanding BICD2 recruitment to the centriole (and its role in centriole engagement). We view this as a very interesting question that could be studied together with BICD2 regulation through phosphorylation. Our current hypothesis is that most of BICD2 is removed through phosphorylation in late G2 and M, with a pool remaining at the mother-daughter interface, possibly protected by a yet to be understood mechanism. We have added a sentence in the discussion addressing this (“ A pool of protein could be protected and correspond to the observed remnant of BICD2 at the mother-daughter interface.“). This last pool, as we discuss in the manuscript, could be further phosphorylated at the M/G1 transition or cleaved by separase (although this last point is of course highly speculative).

      __ Fig. 4: The GFP-BICD2 488-820 fragment forms cytoplasmic rings, which is interesting. This domain contains the CC4 domain, so it can localize near the centriole, but why does it not form a perfect ring there? Also, which other centriole/centrosome markers were used for colocalization studies? Does knockdown of PCM1 affect BICD2's centrosomal localization?__

      We show in Figures 6G and 6H that BICD2 488-820 can form a ring around the centriole. Indeed this polypeptide contains the CC4 region, which our results indicate it will guide it to the centriole (through an interaction with CEP152). Once the available CEP152 is occupied with BICD2 we assume that BICD2 488-820 forms oligomers that assemble ring-like structures outside the centriole.

      Other centrosomal markers used are SAS-6.

      We now show that PCM1 knockdown does not affect BICD2's centrosomal localization (Supplementary Figure S4C). Although our results indicate that partial forms of BICD2 such as BICD2 488-820 can colocalize with PCM-1 (Figure 6D), full length endogenous BICD2 (or GST-BICD2) does not seem to colocalize with this protein and thus the centriole satellites (Supplementary Figure S4A). We note this discrepancy in the text: “The presence of BICD2 at the centriolar satellites has been suggested previously (Quarantotti et al, 2019); we ignore the reason why in the conditions used in this study only C-terminal fragments of BICD2 but not the full-length protein”. The relationship between satellites and BICD2 grants further studies. Our data suggests that centriolar localization of the protein may be regulated, possibly by its intramolecular structure, and that regulated binding of BICD2 to a yet to be identified partner may recruit the protein to satellites either for its transport to the centrosome or in order to perform a specific function at the satellites. We have added a sentence to the text to note this: “This suggests that BICD2 satellite localization is regulated (possibly via intramolecular autoinhibition) to mediate BICD2 transport or a distinct satellite-specific function of this protein.”.

      __ Fig. 4A: What are the aggregates observed in the cytoplasm under the GFP-BICD2 + ice condition? Also, does the 1-575 mutant fail to localize to the centrosome upon ice treatment?__

      We currently do not know the nature of the GFP-BICD2 full length aggregates observed upon microtubule depolymerization. We also observe GFP-BICD2 aggregates in cells that express high amounts of the polypeptide, leading us to hypothesize that it may be insoluble and the disappearance of microtubules may liberate it from motor complexes resulting in its aggregation -although of course more work would be needed to clarify this.

      We now show new data (Supplementary Figure S5), showing that the localization of not only GFP-BICD2 1-575 but also the C-terminal fragments 272-820 and 488-820 are not significantly affected by cold-induced microtubule depolymerization. These last results strongly suggest that BICD2 localization at the centriole is microtubule independent and are compatible with our data showing that BICD2 can directly interact with the centriolar protein CEP152.

      __ Can similar phenotypes be observed in other cell types when BICD2 is knocked down? This should be experimentally validated.__

      Our current manuscript now shows that similar phenotypes regarding centriole separation and amplification are observed upon BICD2 depletion in RPE-1 cells (non-transformed, p53-wildtype) and U2OS cells (transformed). These are shown in Figure 4 (RPE-1 knockout, U2OS RNAi knockdown) and Figure 5 (RPE-1 and U2OS RNAi knockdown).

      __ Are there previous studies suggesting that this function of BICD2 is evolutionarily conserved? This should be addressed.__

      To our knowledge there are no previous studies describing BICD2 function at the centrosome, excepting the recent article by Kuang et al., (Kuang W et al. 2025. BICD2 promotes ciliogenesis by facilitating CP110 removal from the mother centriole. EMBO reports 26:5567–5588. DOI: https://doi.org/10.1038/s44319-025-00597-0), that describes a role for BICD2 during ciliogenesis in non-cycling cells. As we mention in our discussion this new role may be related to the distal pool of protein that we observe using ExM, and we don’t think is related to the function of the proximal pool of BICD2 at the torus in cycling cells that we describe in our manuscript.

      Regarding functional conservation, BICD2 orthologs are widely distributed across metazoans (as reflected in OrthoDB, which lists ~5,000 ortholog genes across ~2,500 species). They share a remarkably conserved C-terminal domain that acts as a docking interface mediating subcellular targeting independently of dynein motor activity (i.e. through binding to Rab6, RanBP2 and, as shown here, CEP152). Cross-species analyses show that this C-terminal domain is preserved in most eukaryotic orthologs, including Drosophila melanogaster BICD (UniProt P16568) and Caenorhabditis elegans BICD-1 (UniProt V6CJ04). Interestingly, several predicted orthologous sequences in public databases retain high C-terminal similarity while completely lacking the N-terminal regions containing the CC1 box motif (residues 29–57 in human BICD2) required for dynein interaction (e.g., predicted isoforms in mouse or camels). Thus, dynein-independent scaffolding functions may represent an ancient, foundational role of the BICD protein family, or alternatively (and perhaps most probably, given that basal metazoans like sponges or Cnidaria do show a conserved N-terminus), these truncated forms may have evolved to fulfill distinct cellular roles operating independently of motor-adaptor activity. We have added a passage at the end of the discussion to reflect this.

      Significance

      In this paper, the identification of BICD2 as a novel factor regulating centriole engagement is of significant importance. However, the mechanisms through which BICD2 controls its localization to the centrosome and regulates centriole engagement remain largely undefined. Further exploration of these mechanisms would likely enhance the value of the paper.

      The findings are likely to be of great interest to researchers in the field of cell biology, particularly those focusing on centrosome biology.

      The above feedback comes from a researcher specializing in centrosome studies.

      __ __

      Reviewer #2

      Evidence, reproducibility and clarity

      Montez-Ruiz and colleagues explore the role of a dynein adaptor BICD2 in the engagement of mother and daughter centrioles. Cells need to maintain centriole engagement in interphase to prevent centriole reduplication and in early mitosis to prevent the formation of aberrant mitosis spindles. The authors demonstrate that BICD2 is a centriolar protein that surrounds the mother centriole adjacent to the daughter centriole. It is removed from centrosomes in mitosis, which, in turn, is responsible for centriole disengagement. Further, they suggest that in BICD2 knock-out G2 and early mitotic cells, centrioles disengage prematurely. By conducting rescue experiments, the authors conclude that BICD2 regions CC2, CC3, and CC4, which are dynein-independent, are essential for their function at the centrosome. Finally, they show that the phosphorylation of S817 and S819 of BICD1 controls its centrosome localization.

      Major comments:

      1. __ Based on F1 and SF1, BICD2 is reduced from centrosomes already in early G2. So, it is hard to square how removing a factor that is not present at the centrosomes at the time of disengagement would dysregulate disengagement. The study at this stage does not explain how BICD2 contributes to centriole engagement only in mitosis, while it does not affect centrioles in S.__ We now present new data obtained using expansion microscopy (ExM) that, together with our super-resolution observations, clarifies this point. As shown in the new Figure 3 and Figure EV2 (and supported by Figures EV3 and EV4), although the total amount of BICD2 at centrosomes is significantly reduced from G2 to M, a pool of BICD2 persists at the mother centriole until late mitosis. Importantly, this pool tends to localize close to the daughter centriole. We note this in the text (“BICD2 remained visible in both diplosomes, associated with the SAS-6 foci (Figure 3, Figures EV2-3). Around anaphase, BICD2 was not detectable in some diplosomes, while others retained some protein (again, close to the SAS-6 foci, which at this point were disappearing from the centrioles as the result of the disassembly of the cartwheel).”). Supported by our BICD2 depletion experiments, we propose that this centriolar pool enables BICD2 to contribute to engagement until late mitosis, when the remaining protein at the centrosome is ultimately removed. We highlight this model in the Discussion section (“In mitosis, when the protein progressively disappears from the centrosomes, BICD2 remains functionally relevant -likely via the small pool that persists at the mother–daughter centriole interface.”).

      Based on our new data demonstrating an interaction with CEP152, we propose that BICD2 forms an outer component of the centriolar torus. Centriole engagement is known to be maintained during S phase by the cartwheel (Huang F et al. 2022. Cartwheel disassembly regulated by CDK1-Cyclin B kinase allows human centriole disengagement and licensing. The Journal of Biological Chemistry 298:102658. DOI: https://doi.org/10.1016/j.jbc.2022.102658; Ito KK et al. 2025. Multimodal mechanisms of human centriole engagement and disengagement. The EMBO journal 44:1294–1321. DOI: https://doi.org/10.1038/s44318-024-00350-8), with the torus playing a role in cohesion later in the cell cycle. We note in our manuscript that this is consistent with our observations and supports a model in which BICD2 functions as part of the torus: “During S phase, mother–daughter centriole cohesion is maintained by the cartwheel (Huang et al, 2022; Ito et al, 2025) and, consistently, does not depend on BICD2.

      __ The interpretation that the longitudinal localization of the BICD2 signal coincides with SAS-6 and procentrioles requires further evidence. BICD2 seems largely localized to the other regions around the mother centriole, and in some examples, it does not colocalize with the site of the daughter centriole or SAS-6 (for instance: F2B second row; SF4B, second row; SF5, fourth row; SF6 upper row).__

      We have added an ExM characterization of BICD2 centrosomal localization in the revised manuscript (Figures 2 and 3), that we think further clarify this point, showing that BICD2 longitudinally coincides with the torus and the daughter centriole. This is supported by new additional superresolution images (Figure 2, Figure EV2).

      Note that ExM revealed an additional stable pool of BICD2 at the distal end of the centrioles that is not detected using standard methanol fixation combined with 3D-SIM. As discussed in the text this distal pool may reflect additional centriolar functions of BICD2.

      __ BICD2 is important for centrosome-nucleus tethering during centrosome separation in G2, and its global removal likely affects the dynamics of the spindle assembly. Is G2 and mitotic progression affected in knockouts? Do the knockout cells show issues with chromosome alignment? Such analyses are critically missing from the manuscript.__

      BICD2 knockout cells indeed show a slightly higher mitotic index than their wild type counterparts, and a higher frequency of lagging chromosomes in anaphase and telophase as well (new data, shown in Figure EV5C). We agree with the reviewer that these might result from the role of BICD2 tethering centrosomes to the nuclear envelope to facilitate their separation during the initial steps of spindle formation. We have added a sentence in the text noting this: “As expected from cells with supernumerary centrioles, BICD2-/- cells showed a slightly higher mitotic index and a higher frequency of lagging chromosomes in anaphase and telophase (Figure EV5C), although these mitotic defects might also be partially attributed to the role of BICD2 in centrosome separation (Splinter et al, 2010; Gallisà-Suñé et al, 2023).

      __ In general, SCLT experiments are ambiguous. Centriole disengagement spontaneously occurs during prolonged prometaphase induced by SCLT. Accordingly, F6D shows that many centriole pairs in the control sample are disengaged after 16h of SCLT treatment. Although the distance between centrioles in knockout cells is, on average, larger, without knowing how BICD2 perturbations affect the dynamics of the mitotic spindles and mitosis progression, SCLT experiments do not provide enough insight.__

      After 16h of SCLT treatment, the authors regularly measure centriole distances in mitosis smaller than 500 nm in all samples. This suggests that the used method (which also needs to be described) cannot reliably assess centriole engagement status. Centrioles can be disengaged but adjacent. The authors reference Shukla et al. 2015 to compare the centriole-to-centriole distances here with those from that publication. However, in Shukla 2015, centriole-to-centriole distances increase from S to M. But here, in F6, the control centriole distances in S, G2, and early M are almost identical and less than 500 nm. This discrepancy needs to be addressed.

      We thank the reviewer for these constructive comments. We appreciate the opportunity to further clarify our methodology and experimental rationale.

      Validity and necessity of STLC treatment (Figures 4D and 7)

      We fully agree with the reviewer that prolonged STLC treatment (16 hours) carries inherent limitations and should not serve as the sole experimental system for studying centriole engagement. As the reviewer notes, 16 hours of STLC treatment results in a baseline population of control cells displaying disengaged centrioles. This population likely represents cells that entered mitosis early during the treatment and remained arrested for the longest duration, or cells with inherently less robust engagement machinery.

      However, we would like to highlight two key observations that validate STLC as a useful comparative tool in our study:

      • The significative increase in the number of cells with higher intercentriolar distances that indicate disengagement in particular experimental conditions. We show that depletion of BICD2 consistently leads to a statistically significant increase in mean intercentriolar distances compared to controls under identical STLC conditions, indicating a distinct weakening of centriole engagement in a substantial number of cells (and thus suggesting that BICD2 is part of the engagement mechanism).

      • Validation in unarrested cells: Crucially, this effect is not an artifact of mitotic arrest. Unarrested, normally cycling mitotic cells also display significantly increased intercentriolar distances in the absence of BICD2 (Figure 4C).

      Following the initial characterization in Figure 4D, we restricted the use of STLC exclusively to experiments requiring cell transfection and recombinant protein expression (Figure 7). Human RPE-1 cells offer the key advantage of being an untransformed, p53-wild-type model. However, they also present technical challenges, including lower transfection efficiencies and sensitivity to experimental manipulation. Capturing a statistically robust sample of transfected, unarrested mitotic cells proved technically challenging. STLC treatment provided a necessary tool to enrich for mitotic cells while allowing clear observation of rescue effects.

      We have explicitly clarified this technical rationale in the manuscript text:

      "Although this treatment inherently increased mean intercentriolar distances, it nevertheless enabled clear observation of the effects of BICD2 ablation, while yielding a sufficient number of mitotic cells expressing the recombinant proteins."

      Assessment of centriole engagement

      We agree that centrioles can occasionally be disengaged while still remaining adjacent. To avoid oversimplifying the observed phenotypes, we chose to report raw intercentriolar distances rather than applying an arbitrary binary classification of "engaged" versus "disengaged." Furthermore, we do not rely solely on distance measurements to assess engagement status. We complemented these data by quantifying c-NAP1-positive centrioles in unarrested, cycling mitotic cells (Figure 4E, F). Because c-NAP1 loading marks centriole-to-centrosome conversion (and thus licensing), this functional readout independently confirms that BICD2 loss promotes premature centriole disengagement.

      Intercentriolar distances across the cell cycle and cell-type variation

      Regarding the comparison with Shukla et al. (2015), we note that their study was conducted in HeLa cells, whereas our primary model is RPE-1 (alongside U2OS cells). Variations in centriole engagement dynamics and distance kinetics can likely be attributed to intrinsic differences among these cell types:

      RPE-1 cells: baseline intercentriolar distances in S-phase control RPE-1 cells (0.4–0.5 µm, measured using centrin) match those reported for HeLa cells in S-phase by Shukla et al. However, in RPE-1 control cells, these distances remain relatively constant from S phase through early M phase (Figure 4).

      U2OS cells: U2OS cells exhibit a slight increase from 0.43±0.01 µm in G2 to 0.50±0.01 µm in M (Figure 5), illustrating that slight variations occur between cell lines.

      Other studies similarly report persistent baseline distances around 0.5 µm through early cell cycle stages. For example, Yaguchi et al. (Yaguchi K et al. 2018. Uncoordinated centrosome cycle underlies the instability of non-diploid somatic cells in mammals. The Journal of Cell Biology 217:2463–2483. DOI: https://doi.org/10.1083/jcb.201701151) observed intercentriolar distances close to 0.5 µm in diploid HAP1 cells throughout mitosis and into early G1 phase, with substantial disengagement (>0.8 µm) occurring only well after cytokinesis onset.

      To address this discrepancy, we have added the following sentence to the manuscript text: "Note that in wild-type S-phase RPE-1 cells, intercentriolar distances measured using centrin as a marker were similar to those described in S-phase HeLa cells (Shukla et al., 2015), namely 0.4–0.5 µm; however, in contrast to HeLa cells, these distances remained fairly constant from S to early M phase in RPE-1 cells." We have also updated the Materials and Methods section to provide a precise description of how intercentriolar distances were measured: “Intercentriolar distances were assessed as the distance between centrin foci of the same diplosome in maximum projections of z-stacks

      __ The authors suggest that BICD2's functions at the centrosome are independent of its dynein functions. They show that GFP-BICD2 1-820 DD rescues centriole engagement among several other mutants. However, it is still possible that the expression of the mutants affects some yet uncovered BICD2 function outside of centrosomes. At least, T821A and S823A should be mutated to Ala. From what I gathered, such mutant should remain associated with mitotic centrosomes. The authors should analyze whether mitotic progression remains unperturbed, and centriole engagement status should be analyzed without SCLT treatment in G2, M, and in ensuing G1.__

      We agree with the reviewer that phosphonull mutants should be added to these experiments. In fact, and as mentioned in the responses to Reviewer 1, we have already performed experiments with the phosphonull mutants, observing that they are more retained at centrosomes than the phosphomimetic counterparts. We would be happy to share these results with the reviewers upon request if helpful. Nevertheless, and as mentioned above, we have decided to remove the preliminary data regarding BICD2 phosphorylation from the manuscript data in order to present a separate and more comprehensive study on BICD2 phosphorylation in the near future.

      Significance

      The question explored is relevant to the centrosome field and beyond since the processes leading to premature centriole disengagement and amplification are not fully understood. The study provides some novel insights. However, at the current stage, the study is preliminary. Additional experiments would be needed to strengthen the conclusion that BICD2 directly regulates centriole disengagement.

      My expertise is in centriole and centrosome assembly and the mechanisms that regulate centrosome homeostasis in human cells.

      __ __

      __Reviewer #3 __

      Evidence, reproducibility and clarity (Required):

      Centrosome duplication is tightly control during cell cycle to prevent loss or amplification of centrosome numbers, which are detrimental for cell proliferation. In preparation for centriole duplication in S-phase, mother and daughter centrioles disengaged late mitosis, a process that functions as a licensing factor for duplication. While several mechanism have been proposed to be important for centriole disengagement, differences between systems and organisms exist, suggesting alternative pathways may play a role.

      In this manuscript, Montes-Ruiz and colleagues investigate the role of the dynein adaptor protein BICD2 during centriole disengagement. They found that BICD2 localises to the centrioles, with a peak in S-Phase. Super resolution microscopy suggests that BICD2 localises to the mother centrioles and is mostly absent in mitosis cells after anaphase, when centrioles are disengaging. KO of BICD2 in REP-1 cells does not some t have strong phenotypes, but the authors found that centriole separation is increased, suggesting a role in centriole cohesion. While there is limited mechanist insight about the regulation of BICD2 and its function at the centrosomes, the data presented suggests a role for BICD2 in centriole cohesion that is independent of dynein interaction. There are however several issues with data presentation, image analyses and data interpretation the authors could improve.

      Major comments

      - On page 5, the authors state that figure 1 and supplementary figure1 data strongly suggest that BICD2 associates with mother centrioles and not the PCM. This is not very clear from the images on these figures. In fact, PCM is often associated with mother centriole as well, thus I am not sure they can make these conclusions based on the data presented in these 2 figures. Also, the fact that PCM is more abundant in G2/M, when BICD2 is not, does not mean it does not localize to the PCM. Higher resolution of expansion will be needed.

      The data presented in supplementary figure 3 does not help the conclusion above as it seems form the images that there is co-localization between BICD2 and pericentrin. It is impossible to conclude also that there is co-localization with the satellite marker PCM-1. In fact, they seem to have no overlap from the images provided. Higher resolution of expansion will be needed.

      Following the reviewer’s suggestion we embarked in a full characterization of BICD2 localization using expansion microscopy (ExM). We think that our new data further clarifies this together with new superresolution data.

      Additionally we now have a figure (Supplementary Figure S4) addressing the relation between pericentrin and the localization of BICD2. We show that pericentrin downregulation does not affect centrosomal BICD2 levels. And that both proteins do not colocalize as observed using 3D-SIM.

      Also regarding pericentrin, the revised version of the manuscript now includes a figure that functionally compares the results of its depletion to those of BICD2 (Figure 5).

      We also provide data showing that BICD2 localization does not significatively change upon PCM-1 depletion (Supplementary Figure S4C). As we mention in the text, previous reports have suggested that BICD2 is indeed in the satellites (Quarantotti V et al. 2019. Centriolar satellites are acentriolar assemblies of centrosomal proteins. The EMBO Journal e101082. DOI: https://doi.org/10.15252/embj.2018101082) . But we only observed clear colocalization of PCM-1 with C-terminal fragments of BICD2. Thus, while GFP-BICD2 488–820 strongly colocalizes with satellites, endogenous BICD2 and full-length GFP-BICD2 do not (Figure 6D). We ignore the reason for this, but the data suggests that satellite localization is regulated (possibly via intramolecular autoinhibition) to mediate BICD2 transport or a distinct satellite-specific function of this protein. To address this we have added a sentence to the text that now reads: “The presence of BICD2 at the centriolar satellites has been suggested previously (Quarantotti et al, 2019); we ignore the reason why in the conditions used in this study only C-terminal fragments of BICD2 (but not the full-length protein, see Supplementary Figure S4) colocalize with satellites. This suggests that BICD2 satellite localization is regulated (possibly via intramolecular autoinhibition) to mediate BICD2 transport or a distinct satellite-specific function of this protein.”.

      - In figure 2, to confirm localization to the mother centrioles, could the authors use a mother centriole marker? Such as a distal appendage protein of ninein? CEP152 localizes to both centrioles in the images provided.

      I was surprised that BICD2 localizes to both distal appendages and linker? These are not close to each other. Can the authors comment on this? In supplementary figure 4C orthogonal view it seems like BICD2 is in between distal appendages and linker?

      We believe that the new ExM data (Figures 2 and 3) directly address the reviewer's concerns.

      Regarding the original supplementary figure S4C, indeed in the orthogonal projections of 3D-SIM images the signal corresponding to BICD2 was observed between distal appendages and linker and was quite broadly distributed. We recognize that this could lead to confusion. We have now removed part of this figure (original Figures 4B and 4C) from the manuscript, as we think that the data is made redundant with our new ExM data. Our new data, with a much higher resolution shows that BICD2 localization corresponds to that of the proximal torus (see new Figures 2D and 2E, and Figure 3). Note that in our new ExM images we use a daughter centriole marker (SAS-6) that (in addition to CEP152) we think helps confirm that BICD2 localizes around the mother centriole.

      -The IF data suggests that BICD2 localization to the centrosome is dynamically regulated during cell cycle. Did the authors consider that this protein could be degraded? Is it a matter of recruitment or total protein levels?

      We agree that protein degradation has to be considered when analyzing cell cycle-dependent localization. However, our data suggest that the dynamic behaviour of BICD2 at the centrosomes does not reflect changes in its total protein amount. We have previously shown that total BICD2 levels are not reduced in mitosis, as assessed by western blot (Gallisà-Suñé N et al. 2023. BICD2 phosphorylation regulates dynein function and centrosome separation in G2 and M. Nature Communications 14:2434. DOI: https://doi.org/10.1038/s41467-023-38116-1). To make this clear in the current manuscript, we have additionally added Figure EV1B depicting BICD2 levels in S, G2 and M phase, and the following note to the text : “Total levels of BICD2 remained constant during the different phases of the cell cycle (Figure EV1B and (Gallisà-Suñé et al, 2023))“.

      - The authors propose that the dynamic localization of BICD2 is associated with licensing. However, it is rather surprising that the phenotype of centriole separation they describe is only observe in mitosis when BICD2 in knockdown and not in S-phase when the levels of BICD2 are higher? If the role of BICD2 is to prevent premature centiole disengagement, shouldn't that be observed in S-phase as well? Why only in mitosis when in control cells BICD2 levels are already very low?

      Recent data supports the notion that in S phase centriole cohesion is maintained by the cartwheel (Huang F et al. 2022. Cartwheel disassembly regulated by CDK1-Cyclin B kinase allows human centriole disengagement and licensing. The Journal of Biological Chemistry 298:102658. DOI: https://doi.org/10.1016/j.jbc.2022.102658; Ito et al. 2025. Multimodal mechanisms of human centriole engagement and disengagement. The EMBO journal 44:1294–1321. DOI: https://doi.org/10.1038/s44318-024-00350-8). Our data, including the new results showing that BICD2 interacts with CEP152, suggests that BICD2 is a dynamic part of the mother centriole torus, a structure that does not seem to be implicated in maintaining cohesion in S. We now note this in the manuscript’s text: “During S phase, mother-daughter centriole cohesion is maintained by the cartwheel (Huang et al, 2022; Ito et al, 2025), and, consistently, does not depend on BICD2.”. Of note, after Ito etl al. BICD2 (and the torus) may have a role in late S if the cartwheel is compromised, something that could be tested in future studies by downregulating cartwheel components and BICD2 simultaneously.

      - The images of C-Nap1 localization in figure 6E are not very convincing to illustrate the pint the authors are making in the main text (additional C-Nap1 foci are visible in the KO cells)

      We would like to note that visualizing C-NAP1 in mitosis is technically challenging, as a significant pool of the protein is displaced from the centrioles after phosphorylation in G2. However, the protein has been widely used as a marker of centriole disengagement (e.g. in the seminal Tsou M-FB et al. 2006. Mechanism limiting centrosome duplication to once per cell cycle. Nature 442:947–951. DOI: https://doi.org/10.1038/nature04985). We therefore consider it a valuable tool to support our conclusions regarding centriole engagement. Regarding extra c-NAP1 foci in BICD2 knockout cells, these may reflect additional centrioles that appear in these cells, as a result of abnormal disengagement and early licensing. To have this into account our data quantifies both c-NAP-1 positive centrioles (increased in KO cells, Figure 4E) and number of c-NAP-1 positive centrioles /total centriole number (with an increase in the abnormal >2:4 configuration in BICD2 KO cells, Figure 4F).

      - The authors propose that PLK1 and CDK1 phosphorylation sites regulate the association of BICD2 with the centrioles. Could this be tested with a PLK1 inhibitor?

      As noted above, we have removed the phosphorylation data from the manuscript, as we aim to report these findings in a dedicated upcoming study. Nevertheless, to address the reviewer's query, we now consider BICD2 to be predominantly regulated by CDK1, supported by data using BI 2536 showing that BICD2 centrosomal levels are unaffected by PLK1 inhibition. In contrast, CDK1 inhibition slightly reduces these levels, though this effect does not reach statistical significance under the tested conditions. We would be glad to share these additional results with the reviewers upon request.

      - On page 13, the authors state that their results do not agree with previous literature showing that pericentrin cleavage can result in disengagement. However, it was unclear from this manuscript what is the evidence to demonstrate that this is the case? The data presented in figure 5B for example only demonstrates that pericentrin levels do not change in the absence of BICD2 in what looks like S-phase cells. Did the authors look at pericentrin levels when they observe centriole disengagement in the ko cells? In G2 or early M-phase?

      We recognize that pericentrin is widely considered a crucial factor in centriole engagement, and have added new data in the manuscript studying the relationship between it and BICD2 (the partially new Supplementary Figure S4), and their relative importances for engagement both in G2 and M (the new Figure 5). Our data suggests that both proteins act independently in a partially redundant manner, BICD2 as part of the torus (key for engagement in G2) and pericentrin of the PCM (more important in M).

      We have updated the Discussion to present this and our view on pericentrin importance for engagement more clearly, specially our concerns that its importance may have been overestimated. Specifically we write that “BICD2 depletion reduces its centrosomal levels to a degree that mirrors those naturally observed during late M and early G1 in unperturbed cells. In contrast, experimental depletion of pericentrin reduces its levels far below physiological baselines across any phase of the cell cycle. This severe reduction produces marked centriole separation in mitosis that is likely amplified by spindle-derived forces. Consequently, the individual contribution of pericentrin to regulating physiological centriole cohesion may be somewhat overestimated under standard experimental knockdowns, and this regulation may rely more heavily on torus components, such as BICD2, than previously appreciated.

      Regarding the phases of the cell cycle in which we quantify pericentrin levels in the original Figure 5B (now Figure EV5B), we realize that the figure could lead to confusion as it was, as they were measured in M (when its amount is maximal, as specified in the figure legend) but the figure did not show examples in this cell cycle phase. We have added new examples of mitotic cells to the figure, and modified the figure labels and wording of the figure legend to clarify this.

      Minor comments

      - A more general reference (review) missing in the second paragraph of the introduction that describe the centrosomes.

      We have added a recent general reference when introducing centrioles (Gönczy P. 2025. Critical constituents and assembly principles of centriole biogenesis in human cells. Nature Reviews Molecular Cell Biology 1–18. DOI: https://doi.org/10.1038/s41580-025-00921-5). Later in the paragraph, when centriole duplication is introduced, we now use this reference plus the also recent Fernandes-Mariano C et al. 2025. Centrosome biogenesis and maintenance in homeostasis and disease. Current Opinion in Cell Biology 94:102485. DOI: https://doi.org/10.1016/j.ceb.2025.102485.

      - Some figures are not well organized, difficult to see which panel they correspond to? The authors could consider labelling panels better to make this clear. For example, figure 4 and 6 could benefit from additional panel labels.

      We have added additional panel labels to Figure 4 (now Figure 6) and Figure 6 (now Figure 4), that we have also slightly reorganized with the aim of making it clearer).

      - On page 7, what the authors mean by: "... we ignore the reason why in the conditions used in this study only C-terminal fragments of BICD2 but not the fill-length protein co-localize with these pericentriolar structures"?

      By "pericentriolar structures” we were referring to the centriolar satellites. We realize that that was not clear and updated the wording of the sentence that now reads “we ignore the reason why in the conditions used in this study only C-terminal fragments of BICD2 (but not the full-length protein, see Supplementary Figure S4) colocalize with satellites.”. We subsequently propose a possible a possible explanation for this: "This suggests that BICD2 satellite localization is regulated (possibly via intramolecular autoinhibition) to mediate BICD2 transport or a distinct satellite-specific function of this protein.".

      Reviewer #3 (Significance (Required)):

      In general this work has limited mechanistic insight and BICD2 localization to the centrosomes was known. However, the authors do go into more detail description of the centriole localization of BICD2 . In addition, their established KO cell lines provide some insights into the role of BICD2 in centriole disengagement, which is of interest to the field. But the limited scope of the conclusions does not advance the field significantly as it is.

      this work will interest a specialized audience.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study uses technically compelling long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The authors further propose the provocative model that post-translational mechanisms, rather than the transcriptional-translational processes, may contribute to circadian regulation of neuronal excitability.

      We are pleased that our study is recognized as valuable and that our technically very challenging long-term in vivo recordings and computational modeling are appreciated. We agree that we are proposing a provocative model that opposes the current hypothesis in chronobiology, which suggests that all observed biological circadian rhythms are outputs of a transcriptional-translational feedback loop (TTFL) clock. Instead, we suggest that a cell comprises, in addition to the TTFL clock, other posttranslational feedback loop (PTFL) clocks without the need of daily transcription and daily degradation of its core elements. While the circadian TTFL clock is entrained to the daily light-dark cycle, the circadian PTFL clocks are suggested to be entrained to other daily cues such as to the availability of pheromone, or the daily changing levels of hormones and second messenger levels. Our novel hypothesis proposes that the TTFL and PTFL clocks are coupled and linked, constituting an adaptive, plastic network that can tune and phase-lock to different external and internal Zeitgeber signals. However, we certainly do not claim that the TTFL circadian clock is not at all involved in the circadian control of the ORN’s circadian membrane potential rhythms. We clarified our manuscript accordingly.

      However, the evidence for circadian firing in these neurons […] remains incomplete.

      As requested by the reviewers during the previous round of review, we had provided the results of RAIN analysis (Thaben and Westermark, 2014) of individual animals in our first revision (Fig. 4A), which clearly shows that two-thirds of the population express circadian rhythms in key attributes that we used to characterize the spontaneous spiking activity. In the initial review it was assumed by the reviewers that phase alignment of dispersed rhythms would bias interpretations of rhythmicity. After having established with RAIN that individuals show circadian rhythms, albeit dispersed across the population due to the lack of a zeitgeber in DD conditions, phase-alignment of recordings from DD animals is a valid next step to prepare the data for statistical analysis across the population. Phase-alignment of desynchronized rhythms is a generally accepted and necessary method employed in chronobiology (e.g., for insect ORNs: Gosh et al., 2024). It is proven as prerequisite to find and statistically analyze rhythmicity in complex, desynchronized data.

      As we explained in the previous rebuttal and in our first revision, rhythms in electrical activity of insect ORNs cannot be easily synchronized by the light-dark cycle alone, but appear to require daily cycles of pheromone, as shown in other moth species (Gosh et al., 2024). Please be aware that the animals that we used here have never been exposed to pheromone, as we state in the Methods. Therefore, this lack of pheromone exposure can explain why about one third of our experimental population is not expressing any daily or circadian rhythmicity in spiking attributes. This is an important result of our manuscript, reported for the first time for Manduca sexta, providing evidence for our hypothesis that it is not the LD-entrained circadian TTFL clock that governs electrical activity rhythms in ORNs. As we explained in the first revision, and clarified further here in the second revision, we cannot phase-synchronize our animals with cycles of pheromone application in our experimental paradigm because we are researching circadian rhythms in spontaneous spiking activity and not pheromone responses. We failed to obtain phase-alignment with a single pheromone pulse the night before the experiments started. These data were added as supplementary Figure to Fig. 3 in the first revision. Here, we further revised our manuscript to clarify this important finding.

      Thus, as requested by the reviewers in the initial review, we could successfully confirm our previous results of circadian firing in ORNs and the disruption of these circadian rhythms with Orco antagonist OLC15 with RAIN. In the current review, the reviewers raise no further specific critical points or comments that would doubt our careful rhythm analysis of our long-term recordings. Thus, we conclude that we provided clear evidence for our central, exciting new finding. For the first time we demonstrated an unexpected new task for Orco: Orco controls the circadian firing pattern in the spontaneous activity, and thus, of the ORN’s membrane potential, via its property as leak/pacemaker channel.

      However, the evidence […] for post-translational modification of Orco as the underlying mechanism remains incomplete.

      We agree with the reviewers that there are many more experiments and combined efforts of biochemists, structural biologists, and electrophysiologists required to provide complete evidence for post-translational modification of Orco and to reveal the underlying mechanism of its circadian control. It is beyond the scope of the current manuscript to provide all details of post-translational control of Orco.

      The reviewers asked previously for additional evidence that Orco transcription is not controlled via the TTFL clock. As requested, we provided extended qPCR–based evidence that Orco, in contrast to timeless, is not controlled by the TTFL clock on the transcriptional level (Fig. 6 in Revision 1). Furthermore, we added a new result in Revision 1 to demonstrate cAMP-dependent post-translational modulation of Orco open-time probability (Fig 9 in Revision 1) at a ZT at which antennal cAMP levels are low (Schendzielorz et al., 2015). We already showed in Flecke et al., 2010, that the addition of cAMP at different ZTs increases the spontaneous spiking activity only at specific ZTs. Here, we show that the effect of cAMP depends on Orco. Since Orco’s circadian role is not mediated via TTFL control, it can be concluded that post-translational mechanisms provide daily/circadian temporal control. In this additional Figure we provide statistically significant proof that, in agreement with our model-prediction, Orco’s circadian control of the ORN spontaneous activity could be mediated via the second messenger cAMP. ZT-dependent input for Orco would be provided via daily changes in cAMP levels (Schendzielorz et al., 2015).

      In contrast, the study does provide strong evidence that the application of cyclic nucleotides can modulate Orco-dependent activity at a single time point, and reports that the temporal pattern of Orco transcript abundance is not circadian.

      We appreciate that the reviewer confirms that our revised manuscript with additional experiments now provides strong evidence that cAMP modulates Orco-dependent spontaneous activity of M. sexta ORNs. Since we already published that cAMP levels expresses daily rhythms in M. sexta antennae (Schendzielorz et al., 2015), and in vivo cAMP infusion increases spontaneous activity and sensitizes pheromone detection (Flecke et al., 2010), and our computational model here proves that circadian modulation of open time probability of Orco is sufficient to explain our experimental data, it is sufficient for our conclusions to test just the one specific zeitgeber time when endogenous cAMP levels are low and pharmacological cAMP increase has the strongest impact. To further reveal complete ZT-dependence of cAMP modulation of Orco´s control of spontaneous activity is beyond the scope of the current manuscript and not part of the current research question. The structure of Orco is extraordinarily conserved during evolution, thus, the cited experimental results from other laboratories and other species showing that Orco is a hub for posttranslational modification are very likely generalizable to different insect species. We clarified the manuscript accordingly. In Drosophila, Orco has at least 5 phosphorylation sites for protein kinase C (PKC), is cAMP-dependently sensitized, and has a Ca<sup>2+</sup>/calmodulin binding site that orchestrates the localization of the OR-Orco heteromer to the cilia. However, in fruit fly and other insects, so far, it can only be speculated how circadian control is provided for Orco, since there are no other publications that examine the circadian regulation of Orco in detail. We clarified our manuscript accordingly.

      To summarize, the logical conclusion based on our newly provided data is that the current hierarchical hypothesis in chronobiology based solely on a circadian TTFL clock that controls Orco transcription does not explain our findings in hawkmoth ORNs. Therefore, we suggest a new systemic hypothesis based upon coupled TTFL and PTFL circadian clocks that can also reconcile otherwise inconsistent data published for insect and mammalian circadian clocks (please see reviewed data in: Stengl and Schneider, 2024). We clarified our manuscript in the second revision and added a new Figure 10 to illustrate our novel hypothesis.

      However, the findings are incomplete to exclude a role for transcriptional-translational mechanisms and their associated multi-layered controls in circadian regulation.

      We certainly do not imply excluding a role for the TTFL clock in (indirectly) affecting circadian control of the membrane potential of ORNs. The new qPCR experiments added in the first revision clearly show that the circadian control mediated via Orco is not an output of the TTFL clock via transcriptional control of Orco. Instead, we predict links between a PTFL membrane clock comprising Orco as hub to integrate posttranslational control and the TTFL nuclear clock. We clarified our manuscript accordingly, adding a new Figure 10 to further illustrate and visualize our hypothesis. The predictions of this systemic hypothesis will be challenged in further experiments that, however, are beyond the scope of the current manuscript.

      Joint Public Review:

      This manuscript puts forward the provocative idea that a posttranslational feedback loop regulates daily and ultradian rhythms in neuronal excitability. The authors used in vivo long-term tip recordings of the long trichoid sensilla of male hawkmoths to analyze spontaneous spiking activity indicative of the ORNs' endogenous membrane potential oscillations. This firing pattern was disrupted by pharmacological blockade of the Orco receptor. They then use these recordings together with computational modeling to predict that Orco receptor neuron (ORN) activity is required for circadian, not ultradian, firing patterns. Orco did not show a circadian expression pattern in a qPCR experiment, and its conductance was proposed to be regulated by cyclic nucleotide levels. This evidence led the authors to conclude that a post-translational feedback loop (PTFL) clockwork, associated with the ORN plasma membrane, allows for temporal control of pheromone detection via the generation of multi-scale endogenous membrane potential oscillations. The findings will interest researchers in neurophysiology, circadian rhythms, and sensory biology. However, the manuscript has limited experimental evidence to support its central hypothesis and is undermined by several assumptions that underlie their data analysis and model builds, as well as insufficient biological data including critical controls to validate and/or fully justify the model the authors are proposing.

      We want to remind our reviewers that we used “ORN” as abbreviation for olfactory receptor neuron (= sensory receptor neurons, a.k.a. olfactory sensory neuron (OSN)) and not for Orco receptor neuron, although we focus on the function of Orco. Accordingly, our central finding is that Orco as ion channel is required for the daily/circadian modulation of spontaneous action potential activity generated by the olfactory receptor neurons in the absence of pheromone stimulation.

      We do not understand the specific basis for the conclusions of the reviewers. Therefore, we ask to please specify what experimental evidence is missing to support our central hypothesis that Orco is not directly TTFL- but PTFL clock-controlled, and to name specifically what the “several assumptions” are that undermine our careful data analysis and model builds. Which specific argument in our previous rebuttal was wrong, was not conclusive? Furthermore, please specify your claim that “critical controls are missing”. Which controls are missing for which experiments? We did add a new figure panel in the first revision to demonstrate that neither the addition of DMSO (the OLC15 solvent) nor the repeated attachment of the recording electrode, which was necessary to obtain paired datasets, altered the spontaneous spiking activity (Fig 1B in Revision 1). Furthermore, we expanded the time series of qPCR data and added tim as positive control to Orco (Fig 7), strengthening our argument that Orco expression is not under TTFL control. Dose-response curves of various Orco agonists and antagonists have been published before (see our references in Revision 1) and are therefore not repeated here.

      Our newly added data confirm what our modeling predicted: cAMP increases spontaneous ORN activity dependent on Orco. Previous publications provide evidence for daily rhythms in cAMP concentrations in hawkmoth antennae (Schendzielorz et al., 2012).

      As is true for any other hypothesis, a hypothesis can only be falsified but not validated and needs to be tested by many experiments from many laboratories over a long time until it will be replaced by the next hypothesis that better explains accumulating contradicting evidence. We are very much looking forward to experimental challenges of our provocative new hypothesis by colleagues in the field of olfaction and of chronobiology. We are convinced that our manuscript will greatly stimulate the field, possibly provoking a paradigm switch in chronobiology and in olfactory research.

      Strengths:

      The authors raise several intriguing model-based hypotheses regarding the mechanisms that underlie the generation of olfactory rhythms. The electrophysiological approach and the long-term recording paradigm are elegant and technically impressive. In the revised version, the authors have added additional qPCR data supporting the lack of rhythmic Orco transcript expression and included a new figure suggesting that cAMP can modulate Orco conductance.

      We thank the reviewers for their acknowledgement of our careful work and hope that our further revisions and clarifications help to argue our case.

      Major weaknesses:

      (1) The cAMP experiment was only conducted at one time-point, which is insufficient to support the central claim that "AMP and cGMP may have ZT-dependent effects on Orco conductivity".

      We agree with the reviewers and revised our discussion accordingly to clarify that in this manuscript it is not our central claim that cAMP and cGMP may have ZT-dependent effects on Orco conductivity. Instead, our data show for the first time that Orco controls circadian rhythms of spontaneous activity of ORNs and that the circadian rhythmicity of spontaneous activity is lost when Orco is blocked. Therefore, we provide novel experimental evidence that Orco is a prerequisite to the circadian rhythmicity of spontaneous activity and thus, to circadian rhythms in the membrane potential of ORNs. Furthermore, as requested by the reviewers we provided clear evidence in the first revision that Orco is not controlled at the transcriptional level by the TTFL clock, in contrast to the TTFL clock protein TIMELESS. Thus, it follows logically that Orco is under post-transcriptional control. Since cAMP levels show circadian oscillations and Orco is gated by cAMP (Fig 9 in Revision 1), we used our computational model to show that a cAMP-dependent increase in Orco conductance alone, via daily oscillating concentrations of cAMP, is sufficient to explain our findings. Therefore, we propose here that daily/circadian oscillations of cAMP modify spontaneous spiking activity via Orco on a posttranscriptional level. But it certainly does not provide all evidence for respective mechanisms of how this cAMP modulation of Orco is obtained, since this is beyond the scope of the current manuscript.

      Since we realized that it is difficult for our readers to visualize a circadian PTFL membrane clock we added a new hypothesis-Figure (Figure 10) and considerably focused and clarified especially the discussion of our manuscript. We pointed out that a membrane-associated signalosome that comprises delayed negative feedback mechanisms, and, thus, constitutes an oscillator, a “membrane clock” that generates oscillations. Based upon our data we propose a membrane-associated signalosome constituting a PTFL circadian clock with Orco as central element. This PTFL membrane clock generates superimposed ultradian and circadian rhythms in its outputs: rhythms in the membrane potential, Ca<sup>2+</sup>, and cAMP levels. The PTFL clock comprises positive feedforward elements that upregulate its outputs, resulting in more depolarization, higher Ca<sup>2+</sup>- and higher cAMP levels. Via the clock’s delayed negative feedback mechanisms these outputs are downregulated, again, resulting in hyperpolarization, decreasing Ca<sup>2+</sup>- and cAMP levels. This signalosome comprises the pacemaker channel Orco as a central hub that is controlled via changes in voltage, Ca<sup>2+</sup>, and cAMP levels. Nevertheless, we predict coupling between the multiscale PTFL membrane clock and the TTFL circadian clock in the nucleus to obtain stable circadian rhythms. As likely mechanism of coupling we predict that Ca<sup>2+</sup>- and cAMP-dependent kinases interlink both types of clocks, thereby obtaining robust and at the same time flexible interlinked cellular rhythms.

      We hope to now successfully clarify and to visualize our central hypothesis of our manuscript that Orco is not directly controlled by a TTFL circadian clock but is a central element of a membrane-associated posttranslational feedback loop clock (PTFL) clock that is linked to but not forced by the TTFL clock which is predicted to control intracellular Ca<sup>2+</sup> homeostasis in a circadian rhythm.

      (2) The revised manuscript continues to rely heavily on prior publications or defers key mechanistic questions (or important manipulations) to future studies. In its current form, the evidence presented remains insufficient to support the central claim that a PTFL constitutes the primary underlying circadian clock mechanism. The proposed model is intriguing, but the data provided do not yet directly demonstrate the novel mechanism.

      We do not understand why the reviewers considers it to be problematic that we “continue to rely heavily on prior publications”. Certainly, we built upon previous publications of our lab as well as on manifold experimental data published by other laboratories in the field of insect olfaction. Our ample citations demonstrate that we have an overview both of the current state of literature and relevant previous literature, dating back to the very first experiments that pioneered pheromone transduction in insects. Based upon our extensive knowledge and experimental data collected in insect olfaction and based on very careful, critical, rigorous analysis of our data and data published by others, we were able to come up with a novel interpretation of the current literature about insect olfaction that differs considerably from the current main views. We consider this to be our strength and judge it as good scientific practice and not a flaw of our work. However, since we do not focus on OR-Orco heteromers and their function in pheromone/odor transduction in the cilia in the current manuscript, we considerably shortened this part of the discussion, avoiding pointing out that highly sensitive moth pheromone transduction greatly differs from less sensitive general odor transduction in Drosophila. Furthermore, since here we focus on cAMP, but not on cGMP-dependent modulation of Orco, we also deleted/considerably shortened this part of our discussion.

      We certainly agree with the reviewers that, while the data provided in the current manuscript are a logical basis for developing our novel hypothesis, they are not a direct and sufficient demonstration of proof and we are not able yet to directly demonstrate and explain the novel mechanism predicted. We would like to point out that if we provided this final proof, it would not be any more a novel hypothesis, but only a novel finding.

      We agree with the reviewers that our provocative hypothesis requires rigorous testing by many further experiments, hopefully not only by our laboratory, but hopefully stimulating new experimental challenges by other laboratories employing different species. But certainly, these experiments with proof-of-principle will take many years and are beyond the scope of our current research paper.

      As per eLife’s assessment system we would like to ask the reviewers to provide detailed feedback as to which experiments/results within the scope of this manuscript would complete this work, or how they think this study should be framed in the light of the results that we obtained. Nevertheless, we hope that with our current careful review the reviewers will be more convinced by our arguments and experiments as valid basis for our provocative new hypothesis.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We appreciate that the reviewers provided an overall positive assessment of our manuscript and offered constructive suggestions for improvement. All three reviewers noted that a key strength of our study is the implementation of a gut microbiome model for the characterization of interbacterial antagonism pathways such as the type VI secretion system (T6SS) that approaches natural complexity. They note our work represents a significant advance in microbiome research, and generates resources that will be of use to many researchers in the field. Two of the reviewers point out that the complexity of our model limits the nature of measurements we can make, and suggest we temper the strength of the some of the conclusions we draw. As noted in more detail below, in our revised manuscript, we have used more precise wording to characterize our findings, and we are more explicit about the connection between the measurements we made and what we can conclude about the physiological role of the T6SS in the gut microbiome.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate the physiological role of the Type VI secretion system (T6SS) in a naturally evolved gut microbiome derived from wild mice (the WildR microbiome). Focusing on Bacteroides acidifaciens, the authors use newly developed genetic tools and strain-replacement strategies to test how T6SS-mediated antagonism influences colonization, persistence, and fitness within a complex gut community. They further show that the T6SS resides on an integrative and conjugative element (ICE), is distributed among select community members, and can be horizontally transferred, with context-dependent effects on colonization and persistence. The authors conclude that the T6SS stabilizes strain presence in the gut microbiome while imposing ecological and physiological constraints that shape its value across contexts.

      This study is likely to have a significant impact on the microbiome field by moving experimental tests of T6SS function out of simplified systems and into a naturally coevolved gut community. The WildR system, together with the strain replacement strategy, ICE-seq approach, and genetic toolkit, represents a powerful and reusable platform for future mechanistic studies of microbial antagonism and mobile genetic elements in vivo.

      The datasets, including isolate genomes, metagenomes, and ICE distribution maps, will be a valuable community resource, particularly for researchers interested in strainresolved dynamics, horizontal gene transfer, and ecological context dependence. Even where mechanistic resolution is incomplete, the work provides a strong experimental foundation upon which such questions can be directly addressed.

      Overall, this study occupies a space between system building and mechanistic dissection. The authors demonstrate that the T6SS influences persistence and community structure in vivo, but the physiological basis of these effects remains unresolved. Interpreting the results as evidence of fitness costs or selective advantage, therefore, requires caution, as multiple ecological and host-mediated processes could produce similar abundance trajectories.

      Placing the findings within the broader literature on microbial antagonism, particularly work emphasizing measurable costs, benefits, and tradeoffs, would help readers better contextualize what is directly demonstrated here versus what remains an open question. Viewed in this light, the principal contribution of the study is to show that such questions can now be addressed experimentally in a realistic gut ecosystem.

      We thank the reviewer for this thoughtful summary of our study. We were glad to read they conclude our work will have a significant impact on the microbiome field and that the resources we have developed will be of value to the community.

      Strengths:

      A major strength of this study is that it directly interrogates the physiological role of the T6SS in a naturally evolved gut microbiome, rather than relying on simplified pairwise or in vitro systems. By working within the WildR community, the authors advance beyond descriptive surveys of T6SS prevalence and address function in an ecologically relevant context.

      The authors provide clear genetic evidence that Bacteroides acidifaciens uses a T6SS to antagonize co-resident Bacteroidales, and that loss of T6SS function specifically compromises long-term persistence without affecting initial colonization. This temporal separation is well designed and supports the conclusion that the T6SS contributes to maintenance rather than establishment within the community.

      Another strength is the identification of the T6SS on an integrative and conjugative element (ICE) and the demonstration that this element is distributed among, and exchanged between, community members. The use of ICE-seq to track distribution and transfer provides strong support for horizontal mobility and adds mechanistic depth to the study.

      Finally, the transfer of the T6SS-ICE into Phocaeicola vulgatus and the observation of context-dependent colonization benefits followed by decline is a compelling result that moves the study beyond simple "T6SS is beneficial" narratives and highlights ecological contingency.

      We appreciate this detailed and nuanced characterization of the strengths of our study.

      Weaknesses:

      Despite these strengths, there is a mismatch between the precision of the claims and the precision of the measurements, particularly regarding fitness costs, physiological burden, and the mechanistic role of the T6SS.

      We acknowledge that in some places, our manuscript could benefit from greater precision in the language we use when linking the outcomes we observe in our study to their potential underlying causes. Specific revisions we made to address this concern are described below.

      First, while the authors conclude that the T6SS "stabilizes strain presence" and that its value is constrained by fitness costs, these costs are not directly measured. Persistence, abundance trajectories, and eventual loss are informative outcomes, but they do not uniquely identify fitness tradeoffs. Decline could arise from multiple nonexclusive mechanisms, including community restructuring, host-mediated effects, incompatibilities of the ICE in new hosts, or ecological retaliation, none of which are disentangled here.

      We agree that multiple mechanisms could explain why populations of certain species carrying a T6SS decline over time, and why for others, the T6SS contributes to long-term persistence. Our use of the term “fitness cost” to describe the phenomenon of decline observed for P. vulgatus carrying the T6SS was not meant to imply any particular underlying mechanism, but was rather our attempt to characterize the phenotypic outcome we observed in simplified terms. We note that ecological context is an important determinant of the fitness cost or benefit of any given trait, and our study sheds light on the importance of the presence of the WildR community and the mouse intestinal environment to the fitness contribution of the T6SS to B. acidifaciens and P. vulgatus. Nonetheless, to avoid implying an overly simplistic interpretation of our results, we have modified our language in the manuscript in several places when describing the role of the T6SS in species persistence in mice colonized with the WildR community.

      Second, the manuscript frames the T6SS as having a defined physiological role, yet the data do not resolve which physiological processes are under selection. The experiments demonstrate that T6SS activity affects persistence, but they do not distinguish whether this occurs via direct killing, resource release, niche modification, or higher-order community effects. As a result, "physiological role" remains underspecified and risks being conflated with ecological outcome.

      We acknowledge that our study does not fully resolve the physiological processes under selection that mediate role of the T6SS in maintaining B. acidifaciens populations in WildR-colonized mice. Indeed, several of the outcomes of T6SS activity the reviewer lists, such as target cell killing and nutrient release, are inextricably linked and thus inherently difficult to disentangle. We note that we did attempt to measure higher-order community effects of T6SS activity with metagenomic sequencing, but acknowledge that this approach may not have been sufficiently sensitive to detect small community shifts mediated by a relatively low-abundance species. To address the concern that our current framing implies more of a mechanistic understanding that our study achieves, we have substituted “ecological” for “physiological” where appropriate throughout the manuscript.

      Third, although the authors emphasize context dependence, the study offers limited quantitative insight into what aspects of context matter. Differences between native and recipient hosts, or between early and late colonization phases, are described but not mechanistically interrogated, making it difficult to generalize beyond the specific cases examined.

      We are not entirely clear what the reviewer means by “differences between native and recipient hosts”, but we agree that additional quantitative studies will be needed to address the generalizability of our findings. Future studies are also needed to address the mechanistic basis for the difference in the benefit conferred by the T6SS that we observed between P. vulgatus and B. acidifaciens.

      Fourth is the lack of engagement with recent experimental literature demonstrating functional roles of the T6SS beyond simple interference competition. While the authors focus on persistence and competitive outcomes, they do not adequately situate their findings within recent work demonstrating that T6SS-mediated antagonism can serve additional physiological functions, including resource acquisition and DNA uptake, thereby linking killing to measurable benefits and tradeoffs. The absence of this literature makes it difficult to place the authors' conclusions about physiological role and fitness cost within the current conceptual framework of the field. Without this context, the physiological interpretation of the results remains incomplete, and alternative functional explanations for the observed dynamics are underexplored.

      We thank the reviewer for specifically highlighting the potential pertinence of this literature to our study. Indeed, we did not cite studies indicating a link between T6SS activity and the uptake of DNA and other resources released by targeted cells. As we note above, the release of intracellular contents from target cells is an inevitable consequence of the delivery of lytic effectors. Thus, distinguishing between fitness benefits conferred from the elimination of competitor species and those arising from scavenging the nutrients released during this process is not straightforward. Measuring the benefits deriving from the uptake of certain released molecules, such as DNA, was not immediately feasible in the system employed in this study and instead we focused on the direct lytic consequences of the effectors delivered via the T6SS. We revised the Discussion to include reference to these possible downstream benefits of T6SS activity (Lines 476-479).

      A further limitation concerns the taxonomic scope of the functional analysis. The authors state that the role of the T6SS in the murine environment is functionally investigated using genetically tractable Bacteroides species, citing the lack of genetic tools for Mucispirillum schaedleri. While this is a reasonable, practical choice, it means that a substantial fraction of T6SS-encoding species in the WildR community are not experimentally interrogated. Consequently, conclusions about the role of the T6SS in the murine gut necessarily reflect the subset of taxa that are genetically accessible and may not fully capture community-level or niche-specific functions of T6SS activity. Given that M. schaedleri is represented as a metagenome-assembled genome, its isolation and genetic manipulation would be technically challenging. Nonetheless, explicitly acknowledging this limitation and slightly tempering claims of generality would strengthen the manuscript.

      The reviewer points out that studying the T6SS activity in M. schadleri would potentially expand the generality of our claims. We agree that having an isolate of this species along with genetic tools for its manipulation would allow us to probe the importance of the T6SS in the gut microbiome more broadly. At the suggestion of the reviewer, we have added explicit mention of the potential benefit of studying the T6SS in this organism to the Discussion (lines 538-539), an endeavor that lies outside of the scope of the current study.

      Finally, several interpretations would benefit from more cautious language. In particular, claims invoking fitness costs, selective advantage, or physiological burden should be explicitly framed as inferences from persistence dynamics, rather than as direct measurements, unless supported by additional quantitative fitness or growth assays.

      We agree with the reviewer that invoking fitness costs, selective advantages or physiological burdens should be done cautiously, and have made revisions to our manuscript where we acknowledge that more precise language was needed (line 43, 416, line 417). However, we would also argue invoking fitness costs and benefits when describe strain persistence dynamics in mice has substantial precedent in the literature (Feng et al. 2020, Brown et al. 2021, Park et al. 2022, Segura Munoz et al. 2022), to list a handful of representative examples published by different groups). It is unclear to us what additional in vivo growth measurements could be taken to substantiate our claim that the T6SS provides a fitness benefit to B. acidifaciens during prolonged gut colonization, or that carrying the ICE imposes a fitness cost on P. vulgatus during longterm colonization. Our in vitro experiments evaluating the competitiveness conferred by T6SS activity provide a measure of insight into its fitness benefits, but as our in vivo strain persistence data and the work of many others show, in vitro measurements do not necessarily capture in vivo parameters.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to determine how a contact-dependent bacterial antagonistic system contributes to the ability of specific bacterial strains to persist within a complex, native gut community derived from wild animals. Rather than focusing on simplified or artificial models, the authors aimed to examine this system in a biologically realistic setting that captures the ecological complexity of the gut environment. To achieve this, they combined controlled laboratory experiments with animal colonization studies and sequencing-based tracking approaches that allow individual strains and mobile genetic elements to be followed over time.

      Strengths:

      A major strength of the work is the integration of multiple complementary approaches to address the same biological question. The use of defined but complex communities, together with in vivo experiments, provides a strong ecological context for interpreting the results. The data consistently show that the antagonistic system is not required for initial establishment but plays a critical role in long-term strain persistence. This insight that moves beyond traditional invasion-based views of microbial competition. The observation that transferable genetic elements can confer only temporary advantages, and may impose longer-term costs depending on community context, adds important nuance to current understanding of microbial fitness.

      We thank the reviewer for the positive feedback and are glad they agree our study provides new insight into the role of interbacterial antagonism in natural communities.

      Weaknesses:

      Overall, there is not a lack of evidence, but a deliberate trade-off between ecological realism and mechanistic resolution, which leaves some causal pathways open to interpretation.

      The reviewer makes a good point that the complexity of the experimental system we employ precludes some lines of experimentation that would yield more mechanistic information. As the reviewer notes, we were aware of the tradeoff between mechanistic resolution and ecological realism when selecting our experimental system. Our deliberate choice to favor biological complexity over mechanistic clarity in this study stemmed from our perception that a major gap in understanding of the T6SS and other antagonism pathways lies in defining their ecological function in complex microbial communities.

      Reviewer #3 (Public review):

      Summary:

      Shen et al. investigate the contribution of the type VI secretion system of Bacteroidales in the gut microbiome assembly and targeting of closely related species. They demonstrate that B. acidifaciens relies on T6SS-mediated antagonism to prevent displacement by co-resident Bacteroidales and other members of the microbiome, allowing B. acidifaciens to persist in the gut.

      Strengths:

      Using a gnotobiotic model colonized with a wild-mouse microbiome is a significant strength of this study. This approach allows tracking of microbiome changes over time and directly examining targeting by Bacteroidales carrying T6SS in a more natural setting. The development of ICE-seq for mapping the distribution of the T6SS in the microbiome is remarkable, enabling the study of how this bacterial weapon is transferred between microbiome members without requiring long-read metagenomics methods.

      We thank the reviewer for their enthusiasm toward our study.

      Weaknesses:

      Some conclusions are based on only four mice per condition. The author should consider increasing the sample size.

      We agree that in some experiments it would be beneficial to increase the sample size from four mice. However, the experiments we performed for this study are time and resource-intensive. Additionally, the experiments on which we base our primary conclusions were all independently replicated with similar results. Given these factors, we determined that the extra confidence that might be afforded by increasing our sample size did not merit the delay in publication and investment in resources that would be required.

      Overall, the authors successfully achieved their objectives, and their experimental design and results support their findings. As mentioned in the discussion, it would be important to investigate the role of the T6SS in resilience to disturbances in the microbiome, such as antibiotics, diet, or pathogen invasion. This work represents a step forward in understanding how contact-dependent competition influences the gut microbiome in relevant ecological contexts.

      We agree that investigating the role of the T6SS during perturbations of the microbiome is a key next step for this work and thank the reviewer for highlighting this important future direction.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Beth A. Shen et al. present a comprehensive and carefully executed study investigating the ecological role of the type VI secretion system (T6SS) in maintaining bacterial strains within a native, complex gut microbiome derived from wild mice. By integrating genetic manipulation, metagenomic and sequencing-based tracking approaches, in vitro competition assays, and gnotobiotic colonization experiments, the authors provide compelling evidence that the T6SS functions primarily as a persistence factor rather than a determinant of initial colonization.

      The study is conceptually strong and addresses an important gap in our understanding of how interbacterial antagonistic systems operate in complex, native microbial communities. The manuscript is generally well organized, the data are clearly presented, and the main conclusions are supported by robust experimental evidence. In particular, the demonstration that T6SS-encoding ICEs confer context- and hostdependent fitness effects, including transient benefits and potential long-term costs, adds important nuance to prevailing models of microbial competition.

      That said, several aspects of the study would benefit from clarification and deeper mechanistic discussion. Addressing the points below would further strengthen the rigor and interpretability of the work. Overall, this is a strong and interesting manuscript, requiring some revisions.

      We greatly appreciate the positive summary of our work by the reviewer, which highlights the multi-faceted approach we took to address gaps in our understanding of interbacterial antagonism in the microbiome. It is our hope that the reviewer agrees that our revisions of the manuscript, based on their feedback, clarify our methods and strengthen the interpretability of our work.

      Major comments

      (1) The competition assays in Figure 2B suggest that T6SS-dependent fitness effects are most pronounced among members of the order Bacteroidales. However, these experiments primarily measure population-level competitive outcomes rather than direct T6SS-mediated targeting events. In addition, the limited number of non-Bacteroidales strains included in the assay makes it difficult to conclude that T6SS activity is strictly restricted to closely related taxa.

      The authors should either temper their conclusions regarding target specificity or clarify that these data reflect competitive outcomes rather than direct evidence of targeting. Expanding the discussion to acknowledge these limitations would improve interpretative accuracy.

      With regards to the measurement we employed for assessing T6SS-mediated targeting, we acknowledge that this is, to a degree, an indirect way of determining T6SS targeting. However, there is extensive precedent in the literature for the use of similar assays in assessing targeting by many contact-dependent antagonism systems including the T6SS in Bacteroidales (Russell et al. 2014, Chatzidaki-Livanis et al. 2016, Wexler et al. 2016) and many Proteobacteria (e.g. (Hood et al. 2010)), the T4SS in Xanthomonas citri (Souza et al. 2015), the CDI system in Escherichia coli (Aoki et al. 2005), and the Esx system in Streptococcus intermedius (Whitney et al. 2017). In these studies, targeting was demonstrated by specific depletion of the competitor strain in the presence of a strain encoding an active antagonism system. We acknowledge that the competitive index we report Figure 2B reflects the relative population levels of both species in the assay, and thus does not directly show target species depletion. We opted to use this metric to display the data in the manuscript as a way of efficiently encapsulating and comparing many strain combinations in a single figure, and because the competitive index differences we observed in these derive from differences in target species growth yields (see Author response image 1, indicating growth yields from a representative strain pairing).

      Author response image 1.

      The T6SS of B. acidifaciens targets a WildR-derived P. vulgatus strain. CFUs indicate populations of the indicated strains after co-culture of wild-type or T6SS-inactivated B. acidifaciens with P. vulgatus. Data represent means and standard errors (n=3, *P<0.01, t-test with log ><0.01, t- test with log transformed data)

      We additionally acknowledge that more extensive testing is needed to fully understand the target range of the Bacteroidales T6SS. In our study, we assessed targeting of every WildR species that was readily culturable, which to the best of our knowledge, represents the broadest panel of targets for the Bacteroidales T6SS to be tested to date. We limited our testing to these strains, as the goal of these experiments was to gain insight into which co-residents of the WildR could be targeted by B. acidifaciens. We agree that testing of a broader cross-section of potential targets has merits, but this would require targeted cultivation strategies to obtain these organisms, and lies outside the scope of the current study. We have revised the manuscript to clarify that the target range testing encompassed the diversity of isolates available (p. 10, lines 231-236).

      (2) The bae1 gene encoded in Bacteroides caecimuris F12 contains a frameshift mutation. It would be valuable for the authors to comment on whether such frameshift mutations are a common genomic feature among gut-associated Bacteroides species in murine models. In addition, comparative analysis of human gut metagenomic datasets could reveal whether homologous effector proteins are present in commensal Bacteroides populations, and whether these homologs exhibit similar disruptive mutations.

      More broadly, the manuscript would benefit from a discussion of whether expression of a fully functional bae1 effector might impose a fitness cost on Bacteroidales members, for example, through metabolic burden or altered resource allocation. This is particularly relevant in light of recent studies demonstrating that T6SS effectors can drive physiological trade-offs by modulating metabolic dynamics (PMID: 40592326). Integrating this perspective would strengthen the evolutionary interpretation of effector mutagenesis.

      We agree with the reviewer that the functional and evolutionary significance of the point mutation in bae1 merits further investigation. Following the reviewer's suggestion, we looked in our own datasets and available public datasets from mouse and human microbiomes for evidence of bae1 inactivation. Unfortunately, the gene is present at a low enough frequency that these analyses were inconclusive. In our own metagenomic data from WildR mice, we did not obtain sufficient sequencing depth to assess the frequency at which bae1 is inactivated across genomes. We found a single complete copy of bae1 identical to that of B. acidifaciens in one published mouse microbiome-derived MAG, and detected fragments of the gene in a number of publicly available isolate and MAG genomes, but these were too low of quality to assess whether or not the gene was intact.

      As to whether or not bae1 expression imposes a fitness cost in the producing organism, we think this is unlikely to be significant, given that the impacts of Bae1 will be neutralized by the accompanying immunity protein. We speculate that the point mutation in the B. caecimuris gene is more likely to have arisen through genetic drift than as a result of selection.

      (3) Quantification and tracking of ICE transfer in vivo. In Figure 4D, the authors assess the abundance of resident P. vulgatus populations in germ-free mice co-gavaged with wild-type strains and derivatives carrying either the intact ICE or ICE ΔtssC. Because both ICE variants are capable of horizontal transfer, it is essential to clearly describe how the authors distinguish between (i) the original wild-type strain, (ii) engineered donor strains, and (iii) recipient strains that have newly acquired the ICE or ICE ΔtssC.

      Clarification of the specific molecular or sequencing-based strategies used to discriminate these populations is necessary to ensure accurate interpretation of the colonization dynamics.

      In this experiment, the P. vulgatus strains we introduced which carried the ICE (either the wild-type version or ICE DtssC) also contained an erythromycin resistance cassette (ermG) inserted distal to the ICE insertion site. Populations of the ICE-containing strain were quantified by either qPCR targeting the ermG gene (Fig. 4D, Supplemental Fig. 4D) or by plating on erythromycin-containing media (Fig. 4F). Endogenous P. vulgatus populations were quantified by qPCR targeting the ermG insertion site, which is disrupted in the marked strain. These methodological details have been added to the figure legend for clarity. We acknowledge that transfer of the ICE between introduced and endogenous populations is possible, and would not be detected by these metrics. To assess whether this occurs, we performed ICE-seq analyses on samples collected from mice colonized by the WildR and P. vulgatus ICE at early (7 days) and late (56 days) time points. These analyses revealed that overall, ICE distribution in this experiment was similar to that observed in mice colonized with the WildR alone (Figure 4A and Author response image 2). They additionally provided corroborating evidence that the population of ICE-containing P. vulgatus declined over the course of the experiment. Importantly, the only ICE insertion site we detected in P. vulgatus in these samples was that found in the introduced P. vulgatus strain. Previous studies show that GA1-containing ICE can insert at numerous locations in Bacteroides sp. genomes, a finding supported by our mapping of the ICE insertion sites from in vitro transfer experiments (Supplemental Fig. 4C) (Garcia-Bayona et al. 2021). Thus, our ICE-seq detection of a sole P. vulgatus ICE insertion site indicates that transfer of the element between P. vulgatus populations is likely not occurring in our experiments.

      Author response image 2.

      ICE-seq analysis indicates that introduction of P. vulgatus ICE into WildR-colonized mice has little impact on ICE distribution among endogenous strains. Graphs show frequency of mapped ICE junctions deriving from the indicated species as determined by 5¢ or 3¢ ICE-Seq analysis of DNA extracted from fecal samples collected either 7 or 56 days post-gavage of the WildR and P. vulgatus ICE into germ-free mice.

      (4) The analysis of fitness trade-offs associated with ICE acquisition in P. vulgatus convincingly demonstrates that the benefits of ICE transfer are transient and contextdependent. However, the mechanistic basis of these trade-offs remains underexplored. While the study primarily attributes both benefits and costs to T6SS-mediated antagonism, the ICE likely encodes additional genes that could influence metabolism, regulation, or stress responses.

      We agree with the reviewer that there are many mechanistic questions remaining regarding the benefits and costs associated with ICE acquisition, and acknowledge that we have not investigated the fitness contributions of ICE-encoded genes other than the T6SS. Indeed, as we noted in our discussion of the results from introducing P. vulgatus carrying the ICE into WildR-carrying mice, our data suggest that ICE genes outside the T6SS may be beneficial (lines 463-465). At the reviewer’s suggestion, we have reiterated the importance of considering the fitness contribution of genes beyond the T6SS in determining ICE distribution in the WildR community (line 527).

      Minor comments:

      (1) In lines 319 and 333, the manuscript refers to "Supplemental Figure 3F" and "Supplemental Figure 3G," respectively. However, the provided Supplemental Figure 3 appears to end at panel E. Please clarify or correct these references.

      We have modified the text to reference the correct figure panels.

      (2) Line 1043: The notation for "OD600" should be corrected for consistency and accuracy.

      The notation for OD600 has been updated to be consistent throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      Minor comments:

      (1) Line 144. I would be careful of using "strong correlation, in this sentence. Although it shows a higher correlation than lab mice. Also, the labels in Figure 1A for mouse WildRF7 are confusing and not well explained in the figure legend.

      We modified line 147 (new line in edited manuscript) to say “positive correlation” rather than “strong correlation” to better represent the result. We also revised the legend for Figure 1A to better explain the samples of WildR F7 that were analyzed.

      (2) Line 155. It's unclear which strains were isolated from the WildR community, and the reason for isolating only 15 strains. Also, Supplemental Figure 1 shows 17 isolates, not 15.

      We apologize for the confusion here. We isolated 17 strains, which is the number of distinct strains we were able to readily culture from this community. We obtained genome sequences for 15 of these, and were able to assemble a genome for one more of the strains from metagenomic data.

      (3) Line 236. Is it known what makes B. uniformis resistant to B. acidifaciens carrying a T6SSS? Does it have an orphan immunity protein?

      We do not know why B. uniformis is not targeted by B. acidifaciens under the conditions of our experiments. It does not encode homologs of the immunity genes bai1 or bai2, and does not appear to be intrinsically resistant to targeting by this T6SS given that it is effectively targeted by P. vulgatus carrying the ICE (Fig.4b).

      (4) Line 295. There is a consistent decline in C. acid abundance after 27 days in Figures 3B and 3C. How do you explain this? Is the endogenous B. acid expanding to outcompete C. acid exo since the total C. acid exo remains constant when gavaging 100x B. acid exo?

      We believe that the eventual decline in the introduced population of B. acidifaciens is likely due to a fitness cost imposed by the erm resistance marker we employed. We noted this phenomenon when describing the results depicted in Fig. 3F-H, but neglected to include this explanation earlier. This oversight has been corrected (lines 315-317).

      References

      Aoki, S. K., R. Pamma, A. D. Hernday, J. E. Bickham, B. A. Braaten and D. A. Low (2005). "Contact-dependent inhibition of growth in Escherichia coli." Science 309(5738): 1245–1248.

      Brown, E. M., H. Arellano-Santoyo, E. R. Temple, Z. A. Costliow, M. Pichaud, A. B. Hall, K. Liu, M. A. Durney, X. Gu, D. R. Plichta, C. A. Clish, J. A. Porter, H. Vlamakis and R. J. Xavier (2021). "Gut microbiome ADP-ribosyltransferases are widespread phage-encoded fitness factors." Cell Host Microbe 29(9): 1351-1365 e1311.

      Chatzidaki-Livanis, M., N. Geva-Zatorsky and L. E. Comstock (2016). "Bacteroides fragilis type VI secretion systems use novel effector and immunity proteins to antagonize human gut Bacteroidales species." Proc Natl Acad Sci U S A 113(13): 3627– 3632.

      Feng, L., A. S. Raman, M. C. Hibberd, J. Cheng, N. W. Griffin, Y. Peng, S. A. Leyn, D. A. Rodionov, A. L. Osterman and J. I. Gordon (2020). "Identifying determinants of bacterial fitness in a model of human gut microbial succession." Proc Natl Acad Sci U S A 117(5): 2622-2633.

      Garcia-Bayona, L., M. J. Coyne and L. E. Comstock (2021). "Mobile Type VI secretion system loci of the gut Bacteroidales display extensive intra-ecosystem transfer, multispecies spread and geographical clustering." PLoS Genet 17(4): e1009541.

      Hood, R. D., P. Singh, F. Hsu, T. Guvener, M. A. Carl, R. R. Trinidad, J. M. Silverman, B. B. Ohlson, K. G. Hicks, R. L. Plemel, M. Li, S. Schwarz, W. Y. Wang, A. J. Merz, D. R. Goodlett and J. D. Mougous (2010). "A type VI secretion system of Pseudomonas aeruginosa targets a toxin to bacteria." Cell Host Microbe 7(1): 25–37.

      Park, S. Y., C. Rao, K. Z. Coyte, G. A. Kuziel, Y. Zhang, W. Huang, E. A. Franzosa, J. K. Weng, C. Huttenhower and S. Rakoff-Nahoum (2022). "Strain-level fitness in the gut microbiome is an emergent property of glycans and a single metabolite." Cell 185(3): 513-529 e521.

      Russell, A. B., A. G. Wexler, B. N. Harding, J. C. Whitney, A. J. Bohn, Y. A. Goo, B. Q. Tran, N. A. Barry, H. Zheng, S. B. Peterson, S. Chou, T. Gonen, D. R. Goodlett, A. L. Goodman and J. D. Mougous (2014). "A type VI secretion-related pathway in Bacteroidetes mediates interbacterial antagonism." Cell Host Microbe 16(2): 227–236.

      Segura Munoz, R. R., S. Mantz, I. Martinez, F. Li, R. J. Schmaltz, N. A. Pudlo, K. Urs, E. C. Martens, J. Walter and A. E. Ramer-Tait (2022). "Experimental evaluation of ecological principles to understand and modulate the outcome of bacterial strain competition in gut microbiomes." ISME J 16(6): 1594-1604.

      Souza, D. P., G. U. Oka, C. E. Alvarez-Martinez, A. W. Bisson-Filho, G. Dunger, L. Hobeika, N. S. Cavalcante, M. C. Alegria, L. R. Barbosa, R. K. Salinas, C. R. Guzzo and C. S. Farah (2015). "Bacterial killing via a type IV secretion system." Nat Commun 6: 6453.

      Wexler, A. G., Y. Bao, J. C. Whitney, L. M. Bobay, J. B. Xavier, W. B. Schofield, N. A. Barry, A. B. Russell, B. Q. Tran, Y. A. Goo, D. R. Goodlett, H. Ochman, J. D. Mougous and A. L. Goodman (2016). "Human symbionts inject and neutralize antibacterial toxins to persist in the gut." Proc Natl Acad Sci 113(13): 3639–3644.

      Whitney, J. C., S. B. Peterson, J. Kim, M. Pazos, A. J. Verster, M. C. Radey, H. D. Kulasekara, M. Q. Ching, N. P. Bullen, D. Bryant, Y. A. Goo, M. G. Surette, E. Borenstein, W. Vollmer and J. D. Mougous (2017). "A broadly distributed toxin family mediates contact-dependent antagonism between gram-positive bacteria." Elife 6(Jul 11): e26938.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors ask whether a simple whole-head spectral power analysis of human magnetoencephalography data recorded at rest in a large cohort of adults shows robust effects of age, and their results provide compelling evidence that it does. The relative simplicity of the analysis is a major strength of the paper, and the authors are careful to control for many different confounds - although perhaps highly correlated factors like brain anatomy still pose a slight issue. The paper provides a valuable power analysis framework that should inform researchers across the broader neuroimaging community

      Many thanks to the reviewers and editorial team. This is an insightful and engaging set of reviews with a range of productive suggestions. We’re pleased that the strengths of this approach show through and that this can be a positive contribution to the community.

      We have implemented the large majority of suggestions and believe that the paper is greatly improved with them in place.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a careful, well-powered treatment of age effects in resting-state MEG. Rather than extracting (say) complex connectivity measures, the authors look at the 'simplest possible thing': changes in the overall power spectrum across age.

      Strengths:

      They find significant age-related changes at different frequency bands: broadly, attenuation at low-frequency (alpha) and increased beta. These patterns are identified in a large dataset (CamCAN) and then verified in other public data.

      Weaknesses:

      Some secondary interpretations (what is "unique" to age vs global anatomy) may go beyond what the statistics strictly warrant in the current form, but these can be tightened with (I think, fairly quick) additions already foreshadowed by the authors' own analyses.

      Aims:

      The authors set out to replace piecemeal, band-by-band ageing claims with t-maps, and Cohen's f2 over sensors×frequency ("GLM-Spectrum").

      On CamCAN, six spatio-spectral peaks survive relatively strict statistical controls. The larger effects are in low-frequency and upper-alpha/beta ranges (f2 approx 0.2-0.3), while lower-alpha and gamma reach significance but with small practical impact (f2 < 0.075). A nice finding is that the same qualitative profile appears in three additional independent datasets.

      Two analyses are especially interesting. First, the authors show a difference between absolute and relative spectral magnitude (basically, within-subject normalization). Relative scaling sharpens the spectral specificity of the spatial maps, while absolute magnitude is dominated by a broad spatial mode that correlates positively across frequencies, likely reflecting head-position/field-spread factors. The replication of the main age profile is robust to preprocessing decisions (e.g., SSS movement compensation choices) - the bigger determinant of the effect is whether they apply sensor normalization (relative vs absolute).

      Second, lots of brain-related things might be related to age, and the authors spend some time trying to back out confounds/covariates. This section is handled transparently (in general, I found the writing style very clear throughout) - they examine single covariates (sex, BP, GGMV, etc.) and compare simple vs partial age effects. For example, aging is correlated with reductions in global grey-matter volume (GGMV), but it would be nice to find a measure that is independent of this: controlling for GGMV (via a linear model) reduces age-related effect sizes heterogeneously across space/frequency but does not eliminate them, a nuance the authors treat carefully.

      Thank you for this concise summary of the work. We’re glad that the strengths of the approach come through clearly.

      This is a nice paper, and I have only a few concrete suggestions:

      (1) High-gamma 

      There can be a lot of EMG / eye movement contamination (I know these were RS eyes closed data, but still..) above 30-40 Hz, and these effects are the weakest anyway. Could you add an analysis (e.g., ICA/label-based muscle component removal) and show the gamma band's sensitivity to that step? Or just note this point more clearly?

      Thanks for this suggestion. We agree that there is concern about eye movements for these gamma band analyses. The ICA preprocessing we conducted was relatively thorough, and we were able to remove EOG-related components from the majority of datasets, even though the experimental protocol involved eyes closed at rest. It is possible that some components were missed, but they would require more advanced labelling tools or manual intervention to identify.

      There are alternative automated ICA labelling tools that could be used, but these are either not suitable for MEG (ICLabel; https://labeling.ucsd.edu/tutorial, https://mne.tools/mne-icalabel/dev/api/iclabel.html) or optimised for data from CTF/4D systems (MEGNet; https://doi.org/10.1016/j.neuroimage.2021.118402). It is beyond our capacity to modify one of these tools for CamCAN for the current analyses.

      It was much more straightforward to rerun the analysis without the ICA step to see the impact of reintroducing all ocular artefacts removed from the v1 analysis. These results are shown in Supplemental section XXX and replicated below.

      We have added the following Figure 11 and text to the paper.

      Main Text

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged and low-gamma effects are reduced (see supplemental section A.3). “

      Supplemental materials

      “The age results may be contaminated in some way by residual cardiac or ocular artefacts that are not removed during preprocessing. Though ICA denoising was applied, it is possible that some artefactual components were not identified and removed from the dataset. The results at low frequencies and in the gamma range are most likely to be directly impacted by this contamination.

      To explore the impact this has on our analysis, we reran the core GLM effect of age on the data with no ICA artefact rejection at all, allowing all eye movements and heart rate components to remain in the data (Figure 10).

      This no-ICA analysis has three differences to the original in the publication. The low-frequency decrease with age is stronger and more widespread in the analysis that removes ocular artefacts with ICA. A large negative effect is visible in both analyses, though without ICA several frontal and temporal sensors no longer show significant effects. Similarly, the effect size of the low-frequency effect of age is substantially larger when ICA denoising is applied.

      The low-alpha effect was strongly reduced in the analyses that do not remove artefacts with ICA. A large central-occipital group of sensors shows an effect between 7 and 8.5 Hz in the ICA analysis, but this is reduced to a single sensor at 8 Hz when ICA is not computed.

      At high frequencies, the age effect in frontal sensors is larger with ICA, and the age effect in posterior sensors is larger without ICA, though the position and frequencies of significant effects are largely unchanged. The remaining effects in the high alpha and beta ranges are unchanged by application of ICA.

      Overall, ICA either improves the estimation of age effects (low-frequency, low-alpha, high-gamma) or has negligible effects (high-alpha, beta). Only the posterior low-gamma effect is reduced by ICA. Together, we take this as evidence that our core results are robust to interference by eye movements and that the ICA denoising is working effectively to reduce noise in the analysis.”

      (2) GGMV confound control 

      Controlling for GGMV reduces, but does not eliminate, age effects. I have a few questions about this: a) Could we see the residuals as a function of age? I wonder if there are non-linear effects or something else that the regression is not accounting for. Also, b) GGMV and age are highly colinear - is this an issue? Can regression really split them apart robustly? I think by some cunning orthogonalisation, you can compute the effect of age independent of GGVM. I don't think this is the same as the effect 'adjusted' for GGMV (which is what is shown here if I'm reading it correctly). Finally, of course, GGMV might actually be the thing you want to look at (because it might more accurately reflect clinical issues) - so strong correlations are not really a problem: I think really the focus might even be on using MEG to predict GGMV and controlling for age.

      This is an interesting area with some tricky interpretation. Thanks for the nudge to help us clarify further.

      We have added the following text and Figures 17 & 18 to the paper to clarify these points.

      Main Text Section 2.8

      “It is important to note that the correlation between age and GGMV does impact the interpretation of the GLM results, but does not prevent the model fit. We explore the model validation and diagnostics in detail in Supplemental Section A.6. In brief, the model is able to separate the unique contributions of age and GGMV. However, the correlation between factors leads to an inflation in the standard error of the estimates. A hypothesis test on these estimates is valid. However, the inflated variance reduces our ability to detect statistically significant partial effects.”

      Supplemental Section A.6

      “There is a strong correlation between age and Global Grey Matter Volume (Pearson’s r=-0.75). The shared variance arising from this collinearity adds nuance to the interpretation of the results, which we explore in more detail in this section.

      Firstly, the sum-square residuals for the group-level model fit including both Age and GGMV are shown as a function of frequency in Figure 17. We see that the residuals broadly follow the overall distribution of variance in the data, peaking at low frequencies and in the alpha range. These frequency ranges are where the strongest signal is visible, but also the highest variability between participants. We would expect that the group model would not perform so well in the points of greatest variability. Importantly, though the residuals are relatively high in the alpha, this is still in the context of a very well-performing model with R2 values of around 80%. “

      “Secondly, the correlation between age and GGMV is not inherently problematic for the GLM, but it does add complexity and nuance to the interpretation of the results. Some additional model validation statistics are shown in Figure 18. “

      “The singular value spectrum of the design matrix indicates whether a design is low-rank, the smallest singular value in this case in 0.37 which indicates that there isn’t a rank deficiency which would prevent us from estimating the model. Though we can estimate the model, correlated regressors can reduce its efficiency. The variance inflation factors for the joint AGE-GGMV model are above 1 for both parametric regressors, indicating that the standard errors of their estimates are inflated. The VIF of 2.35 indicates that the standard errors of this joint model are around sqrt(2.35) = 1.533 times greater than they would be in a separate or uncorrelated model. Though there is no hard rule for this, the literature generally suggests that a VIF above 5 (or sometimes 10) indicates severe multicollinearity.”

      Including additional regressors in the model can change the age estimate by ‘partialling’ out the variance that can be attributed to the other variables and by inflating standard errors. In this specific case of GGMV, the partialled estimates are reduced heterogeneously across space and frequency, and the amount of inflation is at a tolerable level. Overall, the regression is able to separate the unique effects of age and GGMV, at the cost of this inflation in the associated standard errors.”

      Reviewer #2 (Public review):

      This paper describes the application of the "GLM-Spectrum" mass univariate approach to examine the effects of age on M/EEG power spectra. Its strengths include promotion of the unbiased approach, suitable for future meta/mega-analyses, and the provision of effect sizes for powering future studies. These are useful contributions to the literature. What is perhaps lacking is a discussion of the limitations of this approach, in comparison to other methods.

      Thank you for this summary and the thoughtful review. The emphasis for this paper is exactly on the points you highlighted, and we’re glad that this has come across well. We agree that a broader comparison to other methods would be a useful addition and have included a series of additional discussion points to address this.

      We will take the opportunity to reclarify that our intention for this method is not to replace other, more complex or targeted approaches, but to establish a more generalisable foundation for their development. We’re fully supportive of other approaches and are working on their application ourselves. On reflection, this was not clear enough in the first submission, and we have added the following text to the introduction to clarify.

      “Reporting of whole-head and full-frequency spectra of effect estimates would make it straightforward to aggregate across studies and eventually enable identification of sub-threshold effects that may be missed in single analyses but are consistent across studies. We argue that this approach provides a generalisable foundation that can support more complex analyses with a frequency component (such as aperiodic slopes, burst detection, and dynamic functional networks) that require more researcher degrees of freedom.”

      An analogy is the mass univariate approach to spatial localisation of effects in fMRI/PET images. This approach is unbiased by prior assumptions about the organisation of the brain, but potentially also less sensitive, by ignoring that prior knowledge. For example, a voxelwise univariate approach is less sensitive to detecting effects in functionally homogeneous brain regions, where SNR can be increased by averaging over voxels.

      In the context of power spectra, the authors' approach deliberately ignores knowledge about the dominant frequency bands/oscillations in human power spectra. This is in contrast to approaches like FOOOF and IRASA, which explicitly parametrise frequency components. I am not saying these methods are better; I just think that the authors should acknowledge that these approaches have advantages over their mass univariate approach (in sensitivity and interpretation; see below). I guess it is a type of bias-sensitivity trade-off: the authors want to avoid bias, but they should acknowledge the corresponding loss of sensitivity, as well as loss of interpretation compared to model-based approaches (i.e, models that parameterise frequency; I don't mean the statistical models for each frequency separately).

      This is an important point, and we are in complete agreement about the importance of giving a balanced description of how this approach fits within the broader literature. We have added the following paragraph to the discussion on limitations to lay this out more clearly.

      “Our approach promotes an exploratory and unbiased approach to quantifying the age effect on neuronal power spectra, which is intended to complement more focused analyses. This has the benefit of reducing researchers’ degrees of freedom and of being broadly generalisable. These come at the cost of a loss in specificity and in sensitivity. Our approach does not specifically quantify features derived from the power spectrum, such as alpha-peak frequency or the aperiodic component of the spectrum. These features are mixed into our full-spectrum estimates but not directly quantified. Thus, they can be challenging to interpret from our approach. Secondly, the mass-univariate approach suffers from a potential loss in sensitivity compared to results that aggregate across spatial or spectral regions that contain consistent results. Where a region or frequency band of interest can be supported from the literature, an approach focusing on a single region has the benefit of reduced noise by averaging estimates from a larger range of observations. Finally, models that consider the whole shape of the spectrum [Donoghue et al. 2021] would also be able to combine information across a range of frequencies rather than depending on a single frequency bin for each estimate. These models have the additional benefit that their parameters are often directly interpretable as features of interest, such as spectral slope or peak frequency.”

      An example of the interpretational loss can be seen in the authors' observation of opposite-signed effects of age around the alpha peak. While the authors acknowledge that this pattern can arise from a reduction in alpha frequency with age, this is an indirect inference, and a direct (and likely much more sensitive) approach would be to parametrise and estimate the peak alpha frequency directly for each participant, as done with FOOOF for example (possibly with group priors, as in Medrano et al, 2025, EJN). The authors emphasise the nonlinear effects of age in Figure 2A, but their approach cannot test this directly (e.g., in terms of plotting effects of age on frequency, magnitude, and width for each participant), so for me, this figure illustrates a weakness of their approach, not a strength.

      We agree that this point might be misleading within its own figure and have moved the result to the supplemental material with a more lightly phrased wording in the main text. Figure 3 on effect sizes has moved to Figure 2, and a new Figure 3 shows the quadratic effect of age as suggested later in the review.

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged, and low-gamma effects are reduced (see supplemental section A.3).“

      This supplemental section contains some additional content relevant to the next comment.

      Then I think the section "Two dissociable and opposite effects in the alpha range" in the Discussion section is confusing, because if there is a single reduction in alpha peak frequency and magnitude with age, then there is only one "effect", not "two dissociable" ones. If the authors do want to claim that there are two dissociable age effects within the alpha range, then they need to do a statistical test, e.g., that the topographies of low and high alpha are significantly different. This then reveals another limitation of the mass univariate approach - that space (channel) is not parametrised either - so one cannot test for significant channel x effect interactions within this framework, as necessary to really claim a dissociation (e.g., in underlying neural generators).

      As above, we agree that this can be misleading and that our intention got muddled in the heading and writing of this subsection. We are not intending to argue that there are definitely two distinct and different effects. We clarify this in the main text of our first version, but not clearly enough:

      “The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak”.

      We choose to present the first paragraph of this section as a discussion of two effects because this is what is already reported in the literature on changes in alpha power with age. Whilst the decrease in alpha frequency with age is well replicated, the decrease in alpha power is less consistently reported, and the literature shows a highly variable picture of the spatial pattern of this effect. We do not claim a strong dissociation based on these results, but both are reported within the literature.

      Our intention is to provide some clarity to this literature by taking a step back and looking at the ‘lay of the land’ in an unbiased way. With this approach, we see different effects of age on magnitude within a canonical alpha range that are completely separated in frequency. With this perspective, it is feasible that different publications taking different regions of interest, frequency band definitions, and processing options could lead to mixed reports of increases and/or decreases in alpha power with age.

      When single papers that take focused but inconsistent approaches report inconsistent results, we would argue that our approach leads to a substantial interpretational gain on the level of collections of papers in the literature.

      We have added the following content to supplemental section A.1 to clarify and Table 3 illustrates the issue. We have retained the paragraph discussing the possibility that these two effects combine to represent a shift in frequency of a single peak.

      Main text section 3.1, replacing the section on ‘Two dissociable effects…’

      “Reconciling conflicting reports of the ageing effect on alpha power.

      The literature exploring how ageing changes alpha power is heterogeneous. Papers that report results in resting alpha power, either from a canonical band or from an individual peak frequency, include reports of a variety of contrasting age effects. This includes positive correlations with age [Rempe et al., 2023, Stier et al., 2023], negative correlations [Thuwal et al., 2021, Lodder and van Putten, 2011, Medrano et al., 2025, Park et al., 2024], both positive and negative effects separated by space [Hoshi and Shigihara, 2020, Pathak et al., 2022], or null results when correcting for individual frequencies and aperiodic slopes [Scally et al., 2018, Merkin et al., 2023] (see supplemental section A.1 more detailed summary). There is broad variability in methodological approaches, which could account for the variety of results. Stier et al. [2023] suggest that analyses in sensor or source space may lead to different effects. Critically for our work, this variability also prevents formal aggregation of results and meta-analyses that could clarify the picture.

      We have proposed that, by taking a step back and tolerating a reduction in sensitivity, we can map out the whole-head whole-frequency structure of the age effect with minimal researcher degrees of freedom and bias. Investigating age effects as a complete spectrum shows that two contrasting age effects on alpha magnitude coexist in close proximity in space and frequency: an increase with a small effect size in central occipital sensors around 7-8.5 Hz and a decrease with a large effect size across a broad set of occipital, temporal and frontal sensors between 9.5-12.5 Hz. Different data samples and different data analysis choices, particularly the selection of regions of interest, source reconstruction, sensor normalisation, or correction of aperiodic components, might emphasise one effect or the other in each analysis. Whilst we have not conclusively explored all possible variants of these analyses, we have provided a framework that would allow future studies to perform formal comparisons and meta-analyses to resolve this bottleneck.”

      “Relationship between effects on alpha power and alpha individual frequency.

      The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak (Seen qualitatively in Figure 1A). This change in alpha peak frequency is highly replicable [Cesnaite et al., 2023, Dustman et al., 1993, Sahoo et al., 2020, Scally et al., 2018, Pathak et al., 2022, Zibrandtsen and Kjaer, 2021] and is a highly predictive spectral marker of ageing [Stier et al., 2024]. Decreases in alpha peak frequency have been linked to a decline in cognitive performance in healthy ageing [Cesnaite et al., 2023, Finley et al., 2024] and MCI [Garc´es et al., 2013, L´opez-Sanz et al., 2016, Puttaert et al., 2021].

      This compelling possibility that the age effect on alpha is a shift in a single peak is complicated by strong evidence for presence of multiple alpha peaks within individuals [Lodder and van Putten, 2011, Chiang et al., 2011, 2008, Klimesch, 1999], with distinct generators and functional relevance [Sokoliuk et al., 2019]. A complex pattern of changes in power, frequency, and spatial distribution likely underlies age-related change in alpha oscillations. Future work will need to explore all three features at the individual level to clearly illuminate the change.”

      Supplemental section A.1

      “Part of our motivation for this method is that variability in the methodological choices in different publications makes it difficult to aggregate varying results across the literature. For example, though a decrease in alpha peak frequency with increasing age is reported highly consistently, the effect of age on alpha power is much more variable. Table 3 shows a representative sample of publications over the last 20 years that report a change in alpha power with age (note that this is intended to be a representative rather than an exhaustive list). Over half of publications (9/15) report a decrease in alpha power, whilst the remaining publications report an increase (2/15), both increases and decreases (1/15), a decrease but only without correcting for aperiodic slope (1/15), a quadratic effect (1/15), and no effect (1/15). Stier et al. [2023] suggest that the choice of analysis space (sensor space or source reconstruction) is likely a source of discrepancies between studies.

      Critically, it is difficult to reconcile these findings with the information reported in the publications. For example, even within the 9 publications that report a decrease, there is little correspondence in the spatial location of the effect. We argue that focused approaches cannot resolve this issue alone, as targeted analyses are more specific to each dataset and less generalisable.

      In the specific case of the mixed literature on change in alpha power with age, our results show that both effects are present and separated in frequency. It is feasible that the different methodological choices and datasets used by each study in our survey means that one or other of these two effects were emphasised. As a result, the literature may not be mixed in scientific terms, but that a rich pattern of results is obscured by methodological variability.”

      While the authors show that normalisation of each person's power spectra by the sum across frequencies helps improve some statistics, they might want to say more about disadvantages of this approach, e.g., loss of sensitivity to any effects (eg of age) that are broadly distributed across majority of frequencies, loss of real SI units (absolute effect sizes) (as well as problems if normalisation were used for techniques like FOOOF, where the 1/f exponent would be affected).

      This is an important point, and we have added the following text to clarify, with one exception. Firstly, the normalisation we applied scales the whole spectrum linearly and would change the 1/f intercept but not the 1/f^alpha exponent.

      Main text section 3.3

      “Both absolute and relative power measures are used throughout the literature, but there is little consensus about their interpretation [Sandre and Troller-Renfree, 2026]. We focus on relative power for most of our results. By normalising each participant’s power spectrum by the sum across frequencies, we found that the results gained specificity in frequency band and reduced concern about wide intersubject differences in overall variance. Though this improved some analyses, relative power has important drawbacks. In particular, it can reduce sensitivity to effects that are broadly distributed across the spectrum and uses arbitrary scaling rather than meaningful physical units. We support calls in the literature to report both relative and absolute power measures [Rempe et al., 2023, Sandre and Troller-Renfree, 2026].”

      The authors should give more information on how artifactual ICs were defined. This may be important for cardiac artefacts, since Schmidt et al (2004, eLife) have pointed out how "standard" ICA thresholds can fail to remove all cardiac effects. This is very important for the effects of age, given that age affects cardiac dynamics (even though the focus of Schmidt et al is the 1/f exponent, could residual cardiac effects cause artifactual age effects in current results, even above ~1Hz?).

      Artefactual components were estimated using standard tools in MNE python. Specifically:

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_ecg

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_eog

      We have clarified the text in the methods to make the overall process clearer.

      We believe that this process has been broadly effective, and we have rejected an average of 2.25 ECG components within each dataset. There remains a strong possibility that residual ECG artefact is present in the data.

      We have added the following text to methods section 4.2

      “Artefactual components relating to eye movements or the heart rate were automatically identified by correlation with the simultaneous EOG and ECG channels. ECG artefacts were identified using cross-trial phase statistics [Dammers et al., 2008] and an automatic threshold based on the sample rate of the data, as implemented in the mne.preprocessing.ICA.find_bads_ecg function in MNE Python. Between 0 and 3 EOG components were rejected in each dataset, with an average of 0.99 (standard deviation: 0.79) across all datasets. EOG artefacts were identified by correlation with the HEOG and VEOG channels, with a threshold set to r = 0.35, as implemented in the mne.preprocessing.ICA.find_bads_eog function in MNE Python. Between 0 and 5 ECG components were rejected in each dataset, with an average of 2.25 (standard deviation: 0.84) across all datasets. The continuous sensor data were then reconstructed without the influence of the components labelled as artefacts.”

      Similar to the response to Reviewer 1, we have not been able to rerun a more advanced ICA algorithm on the data but have repeated the analysis without any ICA to see if including all ocular and cardiac artefacts influences the results. This change does not introduce any new signal components to the results but does attenuate the low-frequency and low-alpha effects. With the assumption that our initial ICA analysis captured the majority of the largest ECG components, we are confident that our core findings are not compromised by cardiac artefacts.

      Please see the response to comments to Reviewer 1 for additional text in the manuscript on this point.

      The authors should clarify the precise maxfilter arguments, and explain what "reference" was used for the "trans" option - e.g., did the authors consider transforming the data to match a sphere at the centre of the helmet, which might not only remove some of the global power differences due to different head positions, but also be best for generalisation of the effect sizes they report to future studies (assuming the centre of the helmet is the most likely location on average)? And on that matter, did head positions actually differ by age at all?

      We have used the maxfilter files as provided by the CamCAN team and an equivalent implementation defined in OSL-ephys for the Oxford and Cambridge MEGUK data. We have clarified the text in section 4.2

      “All MEG data pre-processing was carried out using MNE-Python [Gramfort, 2013] and OSL-ephys [Quinn et al., 2022, van Es et al., 2025] using the OSL batch pre-processing tools. For the CamCAN data, we proceeded with analysis on post-maxfilter processed data provided by the CamCAN team. Briefly, the data were processed using AA [Cusack et al. 2015] with automatic bad channel detection (limited to 7 channels) with the origin set to the centre of a sphere fitted to the individual’s Polhemus head shape points. Maxfilter signal-space separation was performed with the temporal extension enabled (temporal window of 10 second and correlation threshold of r = 0.98). Head position was continuously estimated and compensated for during periods where the HPI coils were on. After maxfilter processing, head positions were translated into a default head position defined as a point relative to each individual’s origin in a head coordinate frame. Data with the full maxfilter processing and with the head position translation or movement compensation were extracted from the CamCAN database. An equivalent pipeline was implemented in OSL-Ephys and applied to the data from the MEG-UK datasets from Oxford and Cambridge. Files from the MEG-UK Nottingham dataset were processed with third-order gradiometry applied.”

      We found that head position in CamCAN does change with age. We have included the following in Supplemental section A.5

      “The head position of participants within the MEG sensor dewar is an important consideration that has the potential to change the signal-to-noise level of each individual data recording. There are significant differences in head position as a function of age in the CamCAN dataset in Y (front-back) direction indicating that older participants are seated further forward in the dewar than younger participants.”

      As a point of reference, this pattern is consistent with the largest change with age estimated using the un-normalised raw power spectra in Figure 5. A 1-95Hz region shows a change with age that is consistent with older participants sitting further forward in the dewar. This effect is absent in the relative power contrasts. 

      We have not fully explored this final point so have not included it in the main paper, but believe it is a useful addition to the discussion of the reviews.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, both reviewers are enthusiastic about your paper, indicating that it provides compelling empirical support for its key claims and that it represents a valuable theoretical advance for the field. They provide a number of comments that you may want to consider prior to finalizing the manuscript for publication. All the best, Redmond O'Connell

      Reviewer #1 (Recommendations for the authors):

      (1) It would be handy to see a single table listing each tested "analysis family" (e.g., sensors×frequency, source parcels×frequency), the multiple controls used, and the permutation count. I kept wanting to see this as I was reading to compare back and forth.

      This is a helpful suggestion, we have included the table in a new supplemental section which is referenced from the main text at the start of the results

      Main text section

      “A summary of all GLM analyses carried out in this work is available in supplemental section A.7.”

      With the following table included in supplemental section A.7

      (2) I loved the power-planning content (section 3.2, the table with peak f2, CIs, contour plot). I think you could somehow make this even more explicit because people will use it a lot - both for this age/MEG domain and more generally as a template for other types of power planning in the field. Perhaps a "How to use this paper to plan N" guide in a paragraph? Power analysis is surely both "important and difficult" - but also not impossible. A flowchart?

      We’re very glad that this section is working well and agree that the practical planning content should have been more constructive! We have added the following text as a guide

      Main text section 3.2

      “We propose the following steps as a practical guide for researchers looking to plan a data sample with a reasonable chance of correctly identifying a particular effect.

      (1) Define research question and identify previous results: Your research question must be defined in advance and well specified. There should be relevant data or literature that can be used to guide your decision.

      (2) Define the smallest effect size of interest and the decision criterion for the sample decision: It is critical to define the parameters of how you will make your sample size decision ahead of time. We recommend considering what the ’smallest effect size of interest’ [Anvari and Lakens, 2021] would be for your question. This is specific to your question and is about more than statistical significance. What effect size would indicate that there is a practically meaningful effect for the future literature to consider?

      (3) Estimate effect sizes to inform your decision: Either by aggregating reported statistics from the literature, or by dedicated processing of previous data, compute an estimate of the effect size. Effect sizes are only estimates, so it is important to compute confidence intervals around your estimate to get a measure of variability.

      (4) Compute power/precision curves assuming these results: Using the estimated effect sizes, compute power curves [Baker et al., 2021] that visualise the relationship between effect size, statistical power, and sample size.

      (6) Select the sample size that meets your pre-defined criteria: Using your definitions from step 2, and when considering the whole power curve, make a decision about what sample size would give you a reasonable chance of replicating the effect of interest.

      (7) Make note of any differences or deviations from this plan during your data collection: There are many practical reasons why your planned sample might not match the data acquired in practice. Such deviations from a plan are ok but should be acknowledged, and any mitigating steps explained [Lakens, 2024].

      (8) Document your process for inclusion in a preregistration or publication: Include details on how a sample size decision was made in your research outputs including preregistrations, preprints, and publications. These details will be useful for future researchers to understand your process and to implement their own.”

      Reviewer #2 (Recommendations for the authors):

      (1) Though they mention in the Discussion, the authors could have noted earlier in Section 2.1 that one does not need to assume a linear effect of age - one could use a polynomial expansion or even local splines within the same GLM framework. Indeed, it seems unlikely a priori that effects of age across the wide range of ages in the CamCAN data are linear for all frequencies.

      This is an important point, and one that was straightforward to implement in our model. Based on feedback on other sections, we have removed the previous Figure 2, moved Figure 3 on effect sizes to the second position, and added a new Figure 3 containing a model of the quadratic age effects. We believe that this is a substantial improvement in the paper.

      Main text section 2.3

      “(2.3) Quadratic effects of age

      The linear effect of age is a convenient and simple regression model. However, it makes a strong assumption that change with age is uniform across the whole age range. A second group-level GLM was computed with an additional regressor to quantify quadratic effects of age, which have been reported in the ageing literature [G´omez et al., 2013, Rempe et al., 2023, Stier et al., 2023]. The spectrum of t-values for the quadratic age predictor (Figure 3) shows significant effects for a U-shaped change with increasing age in the low frequency (1-5 Hz) range in central sensors. Significant inverted-U shaped effects are present in two frequency ranges in the beta band. A low-frequency beta effect is present in occipital and temporal sensors between 16 Hz and 20 Hz, whilst a second high-beta effect is present in central sensors between 24 Hz and 30.5 Hz. The low-frequency and low-beta effects overlap in space and frequency with linear effects, suggesting that the low-frequency change with age has both a linear increase with age and a U-shaped component, whilst the low-beta effect has a decrease with age and an inverted-U shaped component. The high-beta effect does not overlap in frequency with any of the reported linear effects. The effect sizes for quadratic effects range between Cohen’s F 2 values of 0.033 for low frequency to 0.076 for high beta and are generally lower than the effect sizes for linear change.”

      Main text section 2.4

      “(2.4) Sample size planning for effects of age on the neuronal power spectrum

      We use the 95% confidence intervals around the effect sizes to make recommendations for future sample sizes for future samples that plan to replicate these results. The observed power calculations have no bearing on the interpretation of the present results. Instead, they should be used as a general guideline for planning future studies. Table 1 gives a full summary of the peak statistics and future sample size range for the six age effects identified in Figure 1B. These results have implications for future sample planning for resting-state electrophysiology studies of ageing. Using the upper bound of the sample size estimates as a conservative estimate, the linear changes with age that have relatively large effect sizes would have well-powered replications with sample sizes of around 50-60 participants. However, the smaller linear effects and all the quadratic effects would require samples of 200 or more participants to have the same probability of detecting the effect if it is indeed present (Table 1). This indicates that study samples should be planned with the smallest effect of interest in mind and that ageing effects in different frequency bands may not all be well powered within the same sample.”

      Discussion section 3.1

      “Linear increase and inverted-U effects in the beta band. The literature consistently reports an increase in low-beta power with age [Gomez et al., 2013, Heinrichs-Graham and Wilson, 2016, Heinrichs-Graham et al., 2018, Hubner et al., 2018, Koyama et al., 1997, Rempe et al., 2023, Stier et al., 2023, Veldhuizen et al., 1993, Xifra-Porxas et al., 2019]. We observed significant inverted-U-shaped quadratic effects of age in two frequency ranges in the beta band, a posterior low-beta component (centred around 18 Hz) and an anterior high-beta component (centred around 25 Hz). This is broadly consistent with reports of quadratic ageing effects in the beta band in the literature [Rempe et al., 2023], though, to our knowledge, our results are the first to suggest a separation of effects into different parts of the beta range. These spectrum changes may be associated with age-related changes in underlying bursting dynamics [Brady et al., 2020, Power et al., 2023].”

      “Methods section 4.7

      Age plus quadratic age models: To move beyond linear change and explore U and inverted-U shaped effects of age, we fitted a group model with three regressors. One constant term, one z-transformed age regressor, and one z-transformed quadratic age regressor (age − mean(age))<sup>2</sup>.”

      (2) The authors could point out an obvious extension of their approach to mixed-effects models, which could properly model within- and between-participant effects (repeated measures), e.g., for longitudinal effects of ageing, GAMMs, etc, including hierarchical linear models that combine trials and participants in the same model, and potentially model trial-specific/stimulus effects, etc.

      This is an important point; we have added the following paragraph to the discussion section.

      Discussion section 3.5

      “Future extensions

      A clear future extension for this work is to formally incorporate estimates of within-subject variability into a mixed-effect model. These powerful models would enable modelling of both fixed effects and random effects, allowing researchers to account for variation within individuals over time and between individuals. LMMs also provide improved approaches for handling missing observations and unbalanced designs, making them especially useful for longitudinal and hierarchical data. A second expansion of this work could use Generalised Additive Mixed Models (GAMMs) to model non-linear relationships using smooth functions while also accounting for random effects. This may allow for more realistic representations of complex patterns of change across age. Linear Mixed Modelling comes with a substantial increase in researcher degrees of freedom and can be challenging to implement and report accurately [Meteyard and Davies, 2020]. We have used fixed-effects modelling in this work in line with our objectives of maintaining generalisability and minimising researcher degrees of freedom.”

      (3) Do the authors want to comment on why effect sizes in Figure 4Ci are so much higher for the Oxford sample? I know the authors' main point is that smaller samples lead to more variable estimates of the true effect size, but there are also other interpretations for these sample differences, e.g., recruitment differences, scanner differences, etc. Have the authors considered implementing empirically Bayesian approaches (like COMBAT, a type of mixed-effects model) to adjust for site differences in both offset and scaling, which I think should be a fairly simple extension of GLM-spectrum?

      We agree that our explanation of here is somewhat lacking. To be transparent, we have thought long and hard about this difference and cannot identify a clear reason why this difference is so striking. There is no apparent difference in the other covariates, such as head position, age, or gender, which could explain why the Oxford dataset has larger effect sizes in this specific frequency range (the results in the low-frequency and alpha ranges are consistent).

      We have added the following to the main text to highlight dedicated data harmonisation strategies that could be applied in this situation.

      Main text section 2.5

      “This may arise from relatively poor estimates of the population level variability from smaller data samples. It is possible that more systematic differences in the participant sampling, recruitment, and data acquisition process could contribute to between-site differences. At present, our analyses can show that the core effects of ageing are replicable across several datasets, though we have not formally combined these datasets into a single analysis. Formal methods for data harmonisation, such as ComBat [Johnson et al., 2006], could correct for additive and multiplicative differences in data scaling across sites to improve site comparisons.

      (4) There are a number of formatting problems - at least in the PDF provided to reviewers - for some symbols, e.g., "below ¡7 Hz" (I presume "j" was originally "~" or something?). Also, sometimes a figure is cited by a single, hyperlinked number, without the word "Figure".

      (5) Line 402: "influence" = "influenced"

      (6) Line 431: date for Gelman & Loken?

      Thank you for highlighting these, we have done a thorough proofread and fixed a large number of spelling and formatting issues.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1

      Comment: A limitation of this study is that it is mainly descriptive. It lacks a clear indication of the most promising targets that should help to advance our understanding of PMP22 biology in Schwann cells.

      Response: We thank the reviewer for this important comment. We agree that the original version did not adequately prioritize the identified PMP22 interaction candidates. To address this concern, we performed additional cross-cell type enrichment and protein interaction network analyses to identify the most biologically relevant pathways emerging from the dataset. Moreover, we performed new orthogonal validation experiments using NanoBRET and in situ PLA, thereby moving beyond purely descriptive proteomic observations. Changes in the manuscript: • We added a new paragraph to the Results section entitled “Cross-cell type enrichment and network analysis highlight (sphingo-)lipid metabolism as a prominent cluster among PMP22-ALFA PPIs” (p.9, l.23) corresponding to new Figure 5: “Cross-cell type enrichment, PPI network analysis and initial orthogonal validation”. Previous Figure 5B-B’’ was moved to the supplementary material as new supplementary Figure 4. • We introduced a cross-cell type pathway enrichment analysis integrating GO, KEGG and curated pathway databases (new Figure 5B-D). • We added a PPI network analysis highlighting lipid metabolism and sphingolipid biosynthesis pathways as recurring biological themes (new Figure 5C,D). • We added BRET ratios of PMP22-HaloTag and NanoLuc-tagged PPI candidates, demonstrating close proximity in living cells (new Figure 5E). • We added images showing in situ PLA of PMP22-ALFA and endogenous PPI candidates in MDCKII cells (new supplementary Figure 5). • The methodology describing these analyses was added to the Methods section under “PPI and functional annotation enrichment analysis” (p.15, l.31), NanoBRET assay” (p.15, l.48) and “In situ PLA” (p.16, l.8). • We have adapted the part of the previous discussion highlighting enzymes of the de novo sphingolipid synthesis pathway as PMP22 PPI candidates and moved it to the Results section (new sub-heading “PPIs of PMP22 are in line with specialized functions of adhesion and lipid metabolism in epithelial cells”, p.6, l.38), as enzymes of sphingolipid metabolism, in addition to other potentially important lipid-associated proteins (CAV1, PLLP), appear as enriched PPI candidates for the first time in the MDCKII data. • In the Discussion section we additionally highlighted a possible connection between PMP22’s role in regulation of calcium signaling and lipid metabolism through ORMDL3 (p.11, l.1).

      Comment: The Schwann cell experiments are the most important. However, considering the low transfection efficacy (primary SC) or the low differentiation (MSC80) these are very limiting in the present form. A suggestion could be to improve transfection efficiency by nucleofection or by using lentiviral vectors to transduce primary rat Schwann cells.

      Response: We agree that the limited transfection efficacy of primary Schwann cells and the incomplete differentiation status of MSC80 cells represent important limitations of the present study, which we acknowledge in the Results section (p.8, l.9; p.9, l.14) and in the revised Discussion (p.12., l.22). By discussing DNA-editing in more complex cellular systems as a potential approach for future studies, we are also proposing a method that can address these shortcomings all at once (p.12, l.25). While nucleofection or viral transduction may indeed improve transgene delivery, the establishment and validation of these approaches go beyond the scope of the present study, particularly in light of the alternative discussed, which we have in fact already begun to implement as part of a new project. To strengthen confidence in the identified candidates–also in the experiments using Schwann cells–despite the limitations of the present study, we performed new orthogonal validation experiments and expanded the presentation of the top candidates in the volcano plots. Changes in the manuscript: • We added NanoBRET and in situ PLA data described in the previous response (Figure 5E, supplementary Figure 5). • To better illustrate that we have identified multiple myelin proteins in Schwann cells as PMP22 PPI candidates despite the limitations, we have modified the annotations of the hits in the volcano plots so that the 25 top enriched proteins are now displayed in an enlarged inset instead of 15. This revealed several canonical myelin proteins such as MPZ, PLP1, MCAM, and CMTM5 in these plots (Figure 4A, D).

      Reviewer #2

      Comment: This study across different cell types is the most comprehensive to date to document PPIs of PMP22. Disappointingly the PPIs differ enormously between cells.

      Response: We agree with the reviewer that PPI candidates vary depending on cell type and conditions. However, we do not share the view that this is disappointing. Rather, we interpret this finding as a reflection of the different functional roles of PMP22 in the various cell types. The significance of this result is why we state in the abstract: “We confirm known interactors, and uncover distinct, cell type-specific enrichment patterns following functional annotation analysis.” In the revised manuscript, we also discuss as to how our results reflect the multifunctional role of PMP22 (p.11, l.13).

      Comment: The study appears to have been performed well and the results are clearly presented on the whole. This reviewer was expecting to see some sort of illustration in figure 5 (or separately in the Discussion) that distils the principal linked candidates according to biological processes or pathways with a scoring or ranking system to communicate how strong a case we are looking at.

      Response: We thank the reviewer for this excellent suggestion. In response to this criticism, as well as the concerns raised by Reviewer #1 that the study was merely descriptive, we added a new integrative analysis that summarizes the major biological themes emerging from the dataset and highlights (sphingo-)lipid metabolism as an example of particular interest to PMP22 biology (new Figure 5B-D). Changes in manuscript: • Figure 5A summarizing overlap between datasets was adopted after minor calculation errors were corrected. • The original overlap-based presentation was moved to the supplemental material and in the main replaced by a new integrative analysis framework (Figure 5B–D) that prioritizes biological pathways and interaction clusters, as also described in our response to the first comment by Reviewer #1. • Figure 5B presents pathway enrichment analysis integrating GO, KEGG and curated pathway databases across cell types and conditions. Resulting clusters are sorted by significance and P values of representative terms displayed as a heatmap. Additionally, the top 3 enriched PPI candidates per term across all cell types and conditions are indicated along with their highest log2-fold enrichment in the PMP22-ALFA eluate. • Figure 5C presents a PPI network of all identified interaction candidates. • Figure 5D highlights proteins associated with lipid metabolism and sphingolipid biosynthesis, as these analyses identify these pathways as one of the most prominent recurring biological themes across datasets.

      Comment: In the version downloaded from the website the discussion appears as a single paragraph, which affects clarity and masks the limited depth of conclusions. The text needs to be structured better to be clear where discussion of each topic begins and ends, and the repetition from other sections should be removed. ¨ In the Schwann cell line MSC80 we found the term myelin sheath enriched (Fig. 4B), with several known myelin proteins like PLP1, MPZ, MCAM and ANXA1 enriched in the PMP22-ALFA eluates of MSC80 and primary Schwann cells. At the onset of myelination, Schwann cells mount a transcriptional program enabling coordinated synthesis of both myelin proteins and lipids that are required in large quantities for myelin sheath formation (LeBlanc et al, 2005; Pertusa et al, 2007; Fledrich et al, 2018; Kim et al, 2018; Poitelon et al, 2020).¨ The link to sphingolipid synthesis (that follows) could be improved, this reviewer missed entirely the fact that there is a link until the third time of reading.

      Response: We fully agree and have revised the Results and Discussion section. Changes in manuscript: • As already stated in our response to the first comment by Reviewer #1, we adapted the part of the previous Discussion in which we highlighted enzymes of the de novo sphingolipid synthesis pathway as PMP22 PPI candidates and moved it to the Results (p.6, l.38) in order to reflect the central theme of our findings already at this point, which is continued in new Figure 5 (“Cross-cell type enrichment, PPI network analysis and initial orthogonal validation”) and in the newly added paragraph “Cross-cell type enrichment and network analysis highlight (sphingo-)lipid metabolism as a prominent cluster among PMP22-ALFA PPIs” (p.9, l.23). • We have organized the Discussion section thematically into four paragraphs, discussing (1) methodological advances of the ALFA-tag interactomics approach, and how our results can be interpreted with regard to PMP22’s multifunctional role across cell types and subcellular compartments, (2) a novel mechanistic link between PMP22 and lipid metabolism as the central theme of our findings, as well as a possible connection to the described role of PMP22 in calcium signaling, (3) limitations of Co-IP-based interactomics and our study in particular, along with future directions, and (4) summarizing, the significance of our study despite the limitations. • In addition, we have provided the individual parts of the Results section with subheadings so that important findings can be grasped at a glance.

      Comment: Re sphingolipid synthesis, SPTLC1 and 2 are quite far down the list in figure 5 and ORMDLs that are mentioned in the text don´t appear at all. On further reading we come to 2 stronger candidates CERS and KDSR, so the section needs reordering. The fact that SPTLC, CERS and KDSR are not known to interact comes at the end of this section not mixed up in the middle.

      In the same section is interfere the right word? In any case likely seems too strong given the lack of validation. ¨It seems likely that PMP22 can directly interfere with sphingolipid synthesis in the ER.

      Response: We thank the reviewer for this insightful observation. In the revised Results section, we now list the candidates ordered by enrichment (p.7, l.8). ORMDL2, which was previously easy to overlook, is listed along with other candidates in Fig. 5B and D. We also agree that “interfere” is too strong with regard to our results, and now state: “Our results thus indicate that PMP22 associates with multiple ER-resident enzymes of the de novo sphingolipid synthesis, with a possible functional role of contributing to physiological regulation of sphingolipid supply to the myelin sheath and to its dysregulation in CMT1A.” (p11, l.31)

      Comment: The limitations are substantial and largely acknowledged by the authors, above all a complete lack of further investigation of any of the candidates; and with these in mind the study amounts to a list of PPIs of PMP22 that will be of value to the narrow field of those specializing in CMT1A, and might also be given a cursory glance by those with an interest in sphingolipid synthesis. Overall without further validation the work is somewhat preliminary. Generally omic studies are used to build hypotheses, which are then further investigated, whereas this study is limited to proteomic analyses.

      This reviewer is not an expert in CMT or PMP22, but has carried out a number of protoemic studies and has extensive experience of cell and molecular biology. This reviewer is not an expert in CMT or PMP22, but has carried out a number of protoemic studies and has extensive experience of cell biology.

      Response: We thank the reviewer for their critical assessment. Acknowledging the limitations, we argue that our study represents a significant step forward in uncovering the molecular role of the still poorly understood PMP22. We not only reveal a molecular link to (sphingo-)lipid biosynthesis, but also provide starting points for structural and functional investigations in various directions.

      Reviewer #3

      Comment: This is a beautiful proteomics study to identify proteins that interact with peripheral myelin protein 22 (a tetraspan membrane protein) in four different cell types: HEK293T (a generic model mammalian cell line), MDCKII epithelial cells, a Schwann cell line, and primary rat Schwann cells. An impressively-optimized protocol was used based on fusing the recently-developed alpha tag to the C-terminus of PMP22 employed. This work also presents an optimized protocol for solubilizing the PMP22, regardless of which intracellular compartment it is in. While there have been previous proteomic studies of PMP22, this study VERY significantly extends previous results. The manuscript is clearly written and the figures are clear. One suggestion for improvement is that I did not find clear descriptions of the contents of supporting Tables 1 and 2. This is needed and also the columns of those excel tables could be improved to make them easier to grasp.

      Response: We thank the reviewer for the positive assessment of our study, and for their suggestion to improve the supplementary tables. Changes in manuscript: • We revised the legends and layouts of Supplementary Tables 1 and 2. • In Supplementary Table 2, additional rows containing differential hit counts together with explanations of table contents were added to facilitate readability as well as traceability of Figures 2A, 3A, 4A, 5B, Supplementary Figure 4 and the candidate numbers we state in the text.

      Comment:

      **Referee cross-commenting**

      The other two authors have raised concerns and offered suggestions that are excellent and worthy of address. However, my enthusiasm for this work remains high and I think that, even in the present version of this manuscript, this work advances of our understanding of PMP22 biology and pathobiology.

      Response: In response to the concerns raised by the Reviewers #1 and #2, we substantially strengthened the manuscript by: • Adding a new integrative pathway and interaction network analysis (Figure 5A-D). • Performing new NanoBRET validation experiments (Figure 5E). • Performing new in situ PLA validation experiments (Supplementary Figure 5). • Expanding and restructuring the Discussion. • Improving prioritization of biologically relevant PMP22-associated pathways and PPI candidates. • Revising the shown volcano plots to make more PPI candidates visible at a glance (Figures 2B, 3B, 4B) • Revising the supplementary information. We believe that these additions have further strengthened the biological interpretation of the dataset and improved the overall significance of the study.

      Comment: Gene variations that alter the expression levels or sequence of the tetraspan membrane protein PMP22 cause about 2/3 of all cases of Charcot-Marie-Tooth disease (CMT), a peripheral neuropathy that impacts 1:2500 humans, making it a top-10 genetic disorder. There is no treatment or cure for CMT beyond orthotics and physical therapy. The normal function of PMP22 in promoting myelination of peripheral nerve axons by Schwann cells is poorly understood, reflecting understudy. This paper presents results that represent a huge step forward relative to previous proteomics studies of PMP22 both in terms of the approach used and in terms of the novel insights derived. Of particular significance are the results present herein that PMP22 very likely plays a central role in regulating sphingolipid biosynthesis in Schwann cells. This extends previous results that PMP22 is involved in cholesterol homeostasis. Given that Schwann cells expand their membrane areas by a factor of several thousand during myelination and given the lipid-raft like composition of the myelin membranes, recognition that PMP22 is centrally involved in BOTH cholesterol and sphingolipid homeostasis illuminates its native function and suggests possible pathological roles for PMP22 under conditions of CMT. I think this paper will be of interest to the growing CMT research community, as well as to scientists interested the basic biology of the peripheral nervous system. While I would not describe this paper as having any significant weaknesses, as a "pulldown"-based proteomics study, the results are subject to the key limitation that observation that a given protein associates with PMP22 cannot be taken to indicate DIRECT interaction without additional studies. The observed proteomic interaction could be mediated by other proteins. Thus, the results of this work are most important both for supporting results from previous studies and for suggesting relationships of PMP22 to other proteins and their associated pathways that are worthy of future studies.

      Response: We thank the reviewer for placing our study within the context of ongoing research on PMP22 and related neuropathies, as well as within the PNS field as a whole. In the Discussion, we acknowledge the limitation that the approaches used cannot distinguish between direct and indirect interactions, and we suggest alternative techniques to elucidate structural and functional details of the revealed PPIs in future investigations (p.12, l.16).

    1. Reviewer #2 (Public review):

      This study examines how curl in the retinal flow field can be used as a control variable for estimating and controlling the heading of a moving observer. The basic idea (which is not entirely new, see Matthis et al. 2022) is that translation along a path with eccentric gaze (meaning that the subject is not heading toward the point they are looking at) produces a pattern of optic flow on the retina with a rotational component around the point of fixation (which can be captured by the mathematical "curl" operator). The sign and magnitude of retinal curl varies with heading relative to the point of fixation, such that curl can be used as a control variable to steer rightward or leftward to move toward the fixated target. The authors perform behavioral experiments and show that there are biases in perceived heading that seem to be largely governed by retinal curl. They also show that a simple controller model can use curl to steer toward a target, and they provide a neural network model that provides a biologically-plausible implementation of the controller (although there are some questions about that).

      There is a core of interesting work here that I think can be important to the field. However, there is a lack of clarity on several important fronts, including design of the behavioral experiments, presentation of the behavioral data, conceptual framing of what curl can and cannot do, etc. Equally importantly, the manuscript is not written in a manner that will make it accessible to most vision scientists. I consider myself to be pretty knowledgeable about optic flow, and I had to read most of the manuscript 3 or 4 times to be able to understand the bulk of it. And my experience is that most vision scientists do not understand optic flow well, so I fear that most of the readers that the authors should want to reach would struggle to understand the work. As written, this is mainly going to make an impact on a handful of optic flow gurus. Thus, this manuscript is going to need a major overhaul to clarify important issues and make this more accessible.

      Major issues:

      (1) The manuscript contains inconsistent, if not misleading, messaging about what information retinal curl does, and does not, provide regarding heading estimation. In the Abstract, the authors state: "We propose an alternative: the visual system utilizes retinal curl directly to estimate heading, rendering the explicit recovery of the FOE unnecessary." Based on my understanding of the rest of the manuscript, I find this statement to be a misrepresentation for two main reasons:<br /> a. To "directly estimate heading" relative to what? When not qualified, most people interpret "heading" to mean an observer's heading relative to the world (or some allocentric reference frame). But retinal curl only gives information about an observer's heading relative to the point on which their eyes are fixated. Moreover, that point of fixation will change every few hundred milliseconds in natural viewing, so the retinal curl will change with each new fixation even as heading relative to the world remains unchanged. So, I think most readers would grossly misinterpret the claim that retinal curl can be used "directly to estimate heading". Indeed, in the authors' controller model, the initial heading needs to be given and then the controller can work. But from where does the visual system get the initial heading, since it does not come from curl? These issues are left hanging. Thus, while curl can provide a very useful input for steering toward a fixated target, other signals are needed to estimate heading relative to the world. This has to be made much clearer early on, and a conceptual schematic diagram might help. Also, the authors generally do not specify the reference frame of the variables they are talking about, leaving lots of room for misinterpretations. It should be clear each time they are talking about a variable, such as heading, whether it is relative to the fixation target, body, world, etc.<br /> b. It seems to me that retinal curl will depend on other variables, in addition to heading relative to the fixation target. For example, it seems to me that the magnitude of retinal curl will depend on self-motion speed, the depth structure of the scene, the angle of elevation of the fixated target, and perhaps others. This is not discussed at all, and many readers would get the misguided impression that there is a 1:1 mapping from curl to heading (relative to fixation). If I am right that this is not correct, it means that retinal curl can tell the observer whether to steer right or left to move toward the fixated target, but it cannot tell them how much to steer. Indeed, in the authors' controller model, there is a free parameter that calibrates curl to angle. It makes sense that this works to fit trajectory data that are given from a fixed environment, but it is unclear how the brain would use retinal curl to control steering when these other variables are uncertain or changing unpredictably. Moreover, how does the system change the mapping from curl to steering command as the location of fixation changes relative to the current heading? These are issues that need to be brought up in framing the problem and discussed at some length. If the authors can show mathematically that retinal curl is only dependent on heading (relative to fixation) and not any of these other variables, it would be very valuable to show the equations for this relationship.

      (2) The description of the behavioral experiment and presentation of behavioral data leaves a lot to be desired.<br /> a. First, it is stated (line 158) that "Participants continuously reported their perceived direction of self-motion while maintaining fixation on the yellow dot." Again, reference frame is completely unspecified. Participants were reporting their perceived heading relative to what? The fixation target? The world? What exactly were the instructions given to the subjects to perform the task? Based on the description of how perceived paths are computed (line 166-), it seems to be presumed that subjects are reporting their heading relative to the world because those angles are then converted into x and z coordinates in what I presume is a world-centered reference frame. But how do we know that subjects are accurately reporting their heading relative to the world? What if they are biased in their reports by the location of the fixation target relative to the scene, or by some other reference signal? Is it possible for the authors to rule out the possibility that perceptual biases seen in the unaltered curl condition result from observers not fully adopting the assumed reference frame of the task? If this cannot be firmly excluded, it seems to create problems for the rest of the study.<br /> b. I also feel that there is a mismatch between what the behavioral task requires and what the controller model does. Subjects are apparently asked to report their heading relative to the world, but the controller model only controls their heading relative to the point that they are fixating. I understand how this is resolved in the model, but I think this type of distinction is buried and will not be apparent to most readers. Again, the reference frames of what is being measured and controlled need to be specified explicitly in all parts of the paper, and the authors needs to explain how the system would combine curl-based control with some other measures of (at least initial) heading for world-centered heading to be computed. All of the assumptions need to be clearly specified.<br /> c. Second, I found it frustrating that the authors never present raw perceptual data from the observers. Rather, in Figure 2, we see reconstructed trajectories that are perfectly smooth with no indications of noise whatsoever. Since these paths are computed from the perceptual reports, there must be some noise inherent in them. The figures should represent this uncertainty somehow, and it should be explained how these perfectly smooth trajectories are obtained.

      (3) "...the magnitude of retinal curl in the fovea can specify the body trajectory relative to gaze (Matthis et al., 2022)." The main idea put forward by the authors here seems to overlap heavily with this statement that they attribute to Matthis et al. 2022. While I think this paper still adds importantly to the topic, the authors do not discuss how their findings are different from those of Matthis et al. 2022, why they are an important extension, etc. Readers should not have to go read this other paper to have any idea how the present findings are placed in importance relative to the literature.

      (4) The analysis and treatment of eye movements is extremely weak. The authors discarded trials for which gaze deviated from the fixation point by more than 3 degrees (which is a LOT given that the eye speeds are generally in the neighborhood of 0.5 deg/sec), and they provide basic stats on the distribution of positions. But this largely misses the point: it is not small position errors that are likely to matter, but rather velocity errors. Even a small amount of retinal slip of the target while it is being pursued will cause image motion that is going to alter the optic flow field around the fixation target. So, for example, the retinal curl field may no longer be centered on the fixation target. How do we know that some of the perceptual biases are not influenced by image motion resulting from imperfect tracking of the fixation target? This needs to be analyzed and discussed.

      (5) I found the sections of text comparing the separate and joined fits (starting line 287) to be a bit too rosy. The authors show the separate fits in the main text, and it is not very surprising that these fits are good given that the model has 30 parameters, and these data are pretty low dimensional. The authors only show the joined fits in the supplement, and they say that they are almost as good as the separate fits (indeed they are better in a model comparison sense, but this is 30 parameters vs. 2 parameters). However, when I look at the fits of the joined model in the supplement, I don't find them to be very impressive. In particular, the model grossly misses the data for the straight paths for several subjects (e.g., id5, id6, id8, id10). And fitting the straight paths would presumably be easiest. This implies that the joined model is really missing something and that fitting the curved paths interacts strongly with fitting the data for different fixation target locations on the straight path. I think that the authors should discuss the results a bit more soberly and tone down their conclusions here.

      (6) The section of the paper on neural simulations (starting line 387) has a few weaknesses. First, why are only straight paths simulated here? This does not seem to provide a very rigorous test of the model. Second, it is awkward that the simulation results are presented in units of pixels, rather than degrees. Third, the authors seem to downplay the fact that the neural estimates of heading seem to oscillate rather wildly (over a range of hundreds of pixels, whatever that means, see especially Fig. S16). It was far from clear to me how an estimate of heading with these large oscillations is useful. It would seem to require that heading estimates are integrated over substantial lengths of time to be reliable. It was therefore unclear how the model produces such smooth paths from these oscillating estimates.

      Comments on revised version.

      Overall, the authors have done a responsible job of responding to the comments of my previous review, and the manuscript is substantially improved. There are a few points on which I still do not completely agree with the authors, and I think these are important to document for the record:

      (1) Introduction: "Pure visual decomposition should function regardless of 3D depth or whether the rotation stems from an active eccentric fixation." Perhaps in a world of noiseless perfect computation, this might be true. But I generally disagree. When there is more depth structure in an environment, then translation of the observer is generally going to create greater motion parallax. That is a fact that I don't think can be disputed. And greater motion parallax should help to decompose optic flow into components related to translation and rotation (the latter of which is not depth dependent), especially when there is noise in estimating location motion vectors.

      (2) Related to point #9 of my previous review: I had asked why the authors believed that retinal curl was computed in area MSTd. Their response is that previous studies (i.e., Graziano et al. 1994) show selectivity to spiral motion stimuli in MSTd. That is true, but those studies typically placed the spiral stimulus centered on the MSTd receptive field, hence they were not presenting something like retinal curl as defined here. So, I think it is still an open question as to where in the brain retinal curl is encoded, and from which areas it would be possible to decode retinal curl from population responses.

      (3) Related to point #10 of my previous review: I had asked about biological plausibility of the gaze-centered inhibition signal in the model. The authors' response is that parietal neurons show gain fields in which response depends (usually monotonically) on eye position. This is true, but it is not a trivial jump from gain fields in individual neural responses to a gaze-centered inhibition signal, and I think the authors should have been more forthcoming about the lack of an established neural signal that directly signals what they want in their model.

      (4) The authors point out that the perceptual biases they measure take a few seconds to emerge and they attribute this to temporal integration. But in their curl manipulations, they temporally average over a 2.4 second window in computing the curl signals that they use to cancel or over-cancel curl. So, it is not clear whether some of the delay in the behavioral effects might result from their computations.

      (5) Related to point #13 of my previous review: I had asked about empirical evidence for the assumption of a relationship between the heading preferences of MSTd neurons and their receptive field locations. In response, the authors state that such a relationship is built into the Layton and Browning (2014) model. While that is a precedent, citing another model as a response to a question about empirical evidence is not a convincing response. If there is no empirical evidence to support such a relationship, it would have been better for the authors to acknowledge this.<br /> Given the way that the eLife review model works, it is not necessary for the authors to address these comments, but I think they should be included in the public review record.

    2. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editors for their positive assessment of our manuscript, and all the referees for their constructive comments, which have substantially improved this work. We welcome the opportunity to address referee #2's points for the public record, as they highlight key theoretical nuances and valuable future research directions.

      (1) We appreciate the reviewer's point that, in a noisy biological system, the increased motion parallax provided by a rich 3D depth structure naturally aids in separating translation from rotation. We fully agree on this point. Our argument aimed at highlighting a fundamental theoretical distinction. Pure algebraic decomposition algorithms are mathematically capable of solving heading on flat planes. The fact that human perception often shows biases in these zero-depth conditions, unless extra-retinal cues are present, suggests that the visual system does not rely on a generalized, global de-rotation algorithm. Instead, it relies on heuristic, depth-dependent structural signals (like motion parallax and retinal curl). We maintain that while depth certainly reduces noise, its strict necessity points toward an ecologically grounded control strategy rather than a noisy global decomposition process.

      (2) We think the reviewer raises a valid point regarding the exact neural locus of retinal curl encoding. It is true that Graziano et al. (1994) utilized centred spiral stimuli rather than the spatially offset curl geometries defined in our task. We view the spiral tuning of MSTd not as a direct, one-to-one mapping of full-field retinal curl, but rather as the foundational computational building block required to extract such a signal. We fully agree with the reviewer that identifying exactly where and how this population response is decoded into a unified, gaze-relative retinal curl signal remains an exciting and open empirical question for future neurophysiological research.

      (3) We acknowledge the reviewer's call for transparency here. The transition from well-documented multiplicative gain fields (which modulate response amplitude based on eye position) to a direct, localized gaze-centered inhibitory drive is indeed a theoretical abstraction in our model. We utilized this localized inhibition as a functional mechanism to demonstrate how sensory evidence and spatial priors might competitively interact within a standard Mexican-hat recurrent architecture. While gain fields clearly establish that parietal networks integrate gaze position, we agree that the exact local-circuit implementation mapping these gain fields to the specific inhibitory dynamics we modeled has yet to be empirically established.

      (4) The reviewer smartly questions whether the 3-5 second behavioral biases delay emerges from the 2.4-second smoothing window used in our computational flow manipulation. It is important to clarify that this 2.4-second window was used solely to stabilize the computed curl signal against high-frequency gait oscillations. This smoothing was restricted strictly to the modeling phase of the controller and neural model and was not applied to the participants' responses. The gradual build-up of their perceptual bias over 3-5 seconds represents their own intrinsic temporal integration of this trajectory, independent of the smoothing parameters used to smooth the curl used in the fitting of the controller and neural modelling.

      (5) We concede the reviewer's point that citing a computational model (Layton & Browning, 2014) does not constitute direct empirical evidence for a relationship between MSTd heading preferences and their receptive field locations. Our intention was to highlight a successful theoretical framework that elegantly organizes known properties of MSTd into a system capable of bypassing global de-rotation. We readily acknowledge that direct, single-cell empirical validation of this specific topographic relationship is currently lacking in the literature, and we appreciate the reviewer ensuring this distinction is clearly noted for the record.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides an important and biologically plausible account of how human perceptual judgments of heading direction are influenced by a specific pattern of motion in optic flow fields known as retinal curl. By combining psychophysical experiments and neural modeling, the authors demonstrate that what was previously considered an incidental "nuisance" signal actually serves as a functional control signal for estimating heading and steering toward a fixated target. While the evidence for the role of curl signals is convincing and advances our understanding of vision-based navigation, the work's impact would be strengthened by situating these findings among other cues that contribute to heading estimation, and by clarifying both the time course of these computations and their generalizability across different navigational contexts.

      We thank the editors and reviewers for their insightful feedback and positive assessment of our study. In this revised version, we have made substantial modifications to better situate our findings within the broader landscape of cues contributing to heading estimation, while also clarifying the time course of these computations and their generalizability across different navigational contexts. These points are included in new sections in the revised discussion.

      In addition, and following eLife guidelines, we have moved the methods to the end and make sure that the manuscript reads well without needing to go through methods first.

      Next, we address all the concerns raised by the reviewers.

      Reviewer #1 (Public review):

      We appreciate Reviewer #1’s very positive feedback. Incorporating the perspective of ‘incidental’ sensory signals is a valuable suggestion that aligns perfectly with our findings. We agree that this perspective significantly strengthens the impact of our paper.

      In the revised version we have added a last section in the Discussion (Generalizability and Testable Predictions) to comment on the functional utility of 'incidental' signals and incorporated the suggested references. In addition, in the same heading, we briefly elaborate on the predictions and generalizability of the model and possible manipulations that might affect the integration between sensory evidence (curl signal) and straight-ahead prior.

      Reviewer #1 (Recommendations for the authors):

      It would be great if the authors could discuss the implications and predictions of their model.

      First, from a broader perspective, the study forms an important piece in the emerging recognition that incidental sensory signals are not a nuisance to the sensorimotor system, but contain functionally relevant and effectively used visual signals (Rolfs & Schweitzer, 2022). The authors may want to appreciate their contribution to this perspective in the discussion of the impact of their results. Indeed, a similar shift in perspective has been realized in the recognition that saccade-induced motion signals are not entirely suppressed but play a functional role for gaze correction (Schweitzer & Rolfs, 2021).

      Second, the authors could spell out additional predictions of their proposal: What are experimental manipulations that could shift the balance between relying on a straight-ahead prior and sensory estimation of curl? What would happen in extreme cases of such sensory evidence? When would priors become overwhelmingly influential?

      As commented in the public review we have now included these two aspects in the discussion.

      Minor point: After equation 11, the authors may want to specify that, like position, gaze g is also coded as {x,y}, just like position p.

      While the neural model details have been moved to Appendix 2 (including this equation), we added text before Equation 22 (previously Eq. 11) clarifying that gaze is encoded in image coordinates. We omitted point index i because the equation applies to all image points relative to a given gaze g.

      Reviewer #2 (Public review):

      We appreciate the reviewer’s feedback regarding the formalization of our reference frames. We agree that certain definitions were implicitly assumed rather than explicitly stated. We have revised the manuscript to provide all necessary self-contained information, ensuring that the geometry of the task response and the definition of heading are unambiguous. In the last paragraph of the revised introduction, we make clear the response frame of reference which is also included in the caption of fig 1. Also, we have addressed the gap between the task response (in world coordinates) and the functional role of the controller. This is particularly discussed in the discussion (section: reference frames) in which we also provide (and rule out) potential alternatives to our response biases. We also address all the other points raised by the reviewer.

      Major issues:

      (1) The manuscript contains inconsistent, if not misleading, messaging about what information retinal curl does, and does not, provide regarding heading estimation. In the Abstract, the authors state: "We propose an alternative: the visual system utilizes retinal curl directly to estimate heading, rendering the explicit recovery of the FOE unnecessary." Based on my understanding of the rest of the manuscript, I find this statement to be a misrepresentation for two main reasons:

      (a) To "directly estimate heading" relative to what? When not qualified, most people interpret "heading" to mean an observer's heading relative to the world (or some allocentric reference frame). But retinal curl only gives information about an observer's heading relative to the point on which their eyes are fixated. Moreover, that point of fixation will change every few hundred milliseconds in natural viewing, so the retinal curl will change with each new fixation even as heading relative to the world remains unchanged. So I think most readers would grossly misinterpret the claim that retinal curl can be used "directly to estimate heading". Indeed, in the authors' controller model, the initial heading needs to be given, and then the controller can work. But from where does the visual system get the initial heading, since it does not come from curl? These issues are left hanging. Thus, while curl can provide a very useful input for steering toward a fixated target, other signals are needed to estimate heading relative to the world. This has to be made much clearer early on, and a conceptual schematic diagram might help. Also, the authors generally do not specify the reference frame of the variables they are talking about, leaving lots of room for misinterpretations. It should be clear each time they are talking about a variable, such as heading, whether it is relative to the fixation target, body, world, etc.

      In our study, participants were instructed to report their “perceived direction of self-motion” by aligning a rotational encoder (steering wheel) with the direction they felt they were moving within the 3D simulated scene. Consequently, participants reported their instantaneous heading in a world-centered reference frame, from which the 3D trajectories were reconstructed. Since the reviewer had to infer this information, we have clarified this point at the end of the introduction, legend of figure 1, methods (now at the end of the ms.) and discussion to ensure it is immediately evident.

      Participants were informed that the initial heading (i.e. θ<sub>0</sub> in our controller nomenclature) was oriented “straight ahead” relative to their body which was aligned longitudinally with the experimental room. We have modified Figure 1B and revised the Methods section to explicitly clarify this initial alignment and the instructions provided to participants.

      In the revised manuscript, we have clarified that while the participant’s report is world-centered, the retinal curl provides a gaze-relative heading signal. Although this was already mentioned, we emphasize this point. In natural navigation toward a fixated target, a world-centered vector is often unnecessary; an error signal indicating heading relative to fixation is sufficient (as the reviewer also notes). However, the initial alignment of the heading within the 3D scene allows the brain to “calibrate” this internal controller, mapping the retinal curl signal onto the 3D world coordinates required for the task. Ad commented above, a new section in the discussion addresses and hopefully clarifies the relation with the controller.

      The reviewer also asks how we can be certain that participants were reporting in world coordinates rather than an alternative frame, such as “heading relative to the fixation target.” We believe our “Cancelled Curl” (and over-cancelled) conditions provide the most compelling evidence to rule out this alternative. In these conditions, the physical position of the fixation target in the scene remained identical to the unaltered flow condition. If participants were simply reporting heading relative to the fixation target’s spatial location, the observed biases should have persisted regardless of the flow manipulation. Instead, the bias vanished when the curl was removed. This causal evidence proves that the bias is driven by the retinal motion signal (curl) rather than the spatial orientation of the eyes or the target’s position in the scene. Furthermore, the temporal evolution of the response supports a world-centered integration (in agreement with Warren et 2001 Nat Neuro.). For simulated straight paths, the perceived heading remains straight for the first few seconds (consistent with the initial world-centred alignment), with biases only emerging after approximately 3 seconds of integration (a point we elaborate on in our response to Reviewer #3). Had participants been responding based on a simple gaze-relative reference frame from the onset, these biases would have manifested significantly earlier. We have incorporated these points into the revised Discussion to better frame our findings alongside other cues, such as the Focus of Expansion (FOE) and egocentric visual direction that contribute to heading estimation.

      Finally, we have rephrased the abstract sentence for clarity. However, we maintain that the original premise remains valid once the world-centered initial heading is aligned with the gaze-centered reference frame.

      (b) It seems to me that retinal curl will depend on other variables, in addition to heading relative to the fixation target. For example, it seems to me that the magnitude of retinal curl will depend on self-motion speed, the depth structure of the scene, the angle of elevation of the fixated target, and perhaps others. This is not discussed at all, and many readers would get the misguided impression that there is a 1:1 mapping from curl to heading (relative to fixation). If I am right that this is not correct, it means that retinal curl can tell the observer whether to steer right or left to move toward the fixated target, but it cannot tell them how much to steer. Indeed, in the authors' controller model, there is a free parameter that calibrates curl to angle. It makes sense that this works to fit trajectory data that are given from a fixed environment, but it is unclear how the brain would use retinal curl to control steering when these other variables are uncertain or changing unpredictably. Moreover, how does the system change the mapping from curl to steering command as the location of fixation changes relative to the current heading? These are issues that need to be brought up in framing the problem and discussed at some length. If the authors can show mathematically that retinal curl is only dependent on heading (relative to fixation) and not any of these other variables, it would be very valuable to show the equations for this relationship.

      The reviewer notes that we must be clear about the relationship between curl and heading (relative to fixation) and the variables that affect curl. We also thank the reviewer for encouraging to add the equations that show the relation of curl with additional variables. We have now included these equations in appendix 1.

      Beyond the discrepancy between heading (θ) and gaze (ψ), curl is geometrically determined by translational self-motion speed (v), eye height (h), and pitch (α). More specifically, curl = (v.sinψcosα)/h. The derivation is now included in appendix 1. Since h = dsinα, where d is the 3D distance to the fixation point, we could express cos α as a function of distance. Certainly, there is not a 1:1 map from curl signal to heading relative to gaze (e.g. θ-ψ). Participant would need to know v and eye height plus extra-retinal information. Frenz et al (2003, Vis Res.) showed that people can estimate self-motion directly from optic flow, across different simulated eye height and gaze angle; extra-retinal information can, in addition, provide knowledge to ψ and α. It is then plausible that the visual system can use and transform the curl signal from a qualitative directional cue (i.e. steering left or right of fixation) into a quantitative steering command. By combining curl with knowledge of gaze orientation and eye height, the visual system can resolve ambiguities in the flow field and utilize curl as a more precise error signal for locomotor control. These aspects are now included in the new version of the discussion.

      (2B) I also feel that there is a mismatch between what the behavioral task requires and what the controller model does. Subjects are apparently asked to report their heading relative to the world, but the controller model only controls their heading relative to the point that they are fixating. I understand how this is resolved in the model, but I think this type of distinction is buried and will not be apparent to most readers. Again, the reference frames of what is being measured and controlled need to be specified explicitly in all parts of the paper, and the authors need to explain how the system would combine curl-based control with some other measures of (at least initial) heading for world-centered heading to be computed. All of the assumptions need to be clearly specified.

      We thank the reviewer for this point. We have addressed the alignment of the reference frames in our response to Issues 1a and 2a. Once the initial orientation (θ<sub>0</sub>) is established in the world frame, the controller model generates steering adjustments that directly translate into heading predictions within that same world reference frame. By treating the perceptual report as an output of the locomotor controller, we resolve the discrepancy between the steering task and the reported heading.

      (2c) In addition, I found it frustrating that the authors never present raw perceptual data from the observers. Rather, in Figure 2, we see reconstructed trajectories that are perfectly smooth with no indications of noise whatsoever. Since these paths are computed from the perceptual reports, there must be some noise inherent in them. The figures should represent this uncertainty somehow, and it should be explained how these perfectly smooth trajectories are obtained.

      We respectfully disagree with the reviewer’s interpretation regarding data smoothing. The thin lines in Figure 2 represent the mean 3D paths derived directly from the response variable (θ<sub>t</sub>) across trials of identical conditions for each participant (as detailed in the ‘Computation of Perceived Path’ section). No smoothing or filtering has been applied to these plotted trajectories other than computing the mean across trials. We also wish to remind the reviewer that the raw data and analysis code remain publicly accessible for further inspection. Having said that, we include now a supplementary figure showing an example of raw data responses as a function of time. This figure will be a supplemental figure of main Figure 2 (now provisionally included in the Suppl Information).

      Regarding the visual representation: in earlier versions of the manuscript, we included shaded 95% Confidence Intervals (CIs) in Figure 2. However, this addition rendered the plot overly cluttered and obscured the individual trajectories. We therefore chose to present individual participant means (thin lines) alongside group averages (thick lines) to emphasize inter-subject variability. For clarity, the 95% CIs are explicitly displayed in Figure 3, where the data density is more conducive to shaded areas.

      (3) “...the magnitude of retinal curl in the fovea can specify the body trajectory relative to gaze (Matthis et al., 2022)." The main idea put forward by the authors here seems to overlap heavily with this statement that they attribute to Matthis et al. 2022. While I think this paper still adds importantly to the topic, the authors do not discuss how their findings are different from those of Matthis et al. 2022, why they are an important extension, etc. Readers should not have to go read this other paper to have any idea how the present findings are placed in importance relative to the literature.

      We have updated the Discussion to more specifically align our findings with Matthis et al. (2022). We emphasize that our study provides the perceptual validation for their ecological observation that the FOE is often too unstable for reliable use, whereas foveal curl remains a robust signal for path estimation. Our paper provides the causal link, since we manipulate curl in real-time (the ‘cancelled & over cancelled curl’ condition) providing the critical evidence that perceived heading is affected by this signal. The relation with this previous study is made clear in the revised discussion.

      (4) The analysis and treatment of eye movements is extremely weak. The authors discarded trials for which gaze deviated from the fixation point by more than 3 degrees (which is a LOT given that the eye speeds are generally in the neighborhood of 0.5 deg/sec), and they provide basic stats on the distribution of positions. But this largely misses the point: it is not small position errors that are likely to matter, but rather velocity errors. Even a small amount of retinal slip of the target while it is being pursued will cause image motion that is going to alter the optic flow field around the fixation target. So, for example, the retinal curl field may no longer be centered on the fixation target. How do we know that some of the perceptual biases are not influenced by image motion resulting from imperfect tracking of the fixation target? This needs to be analyzed and discussed.

      We thank the reviewer for noting that retinal slip (velocity error) is a more critical metric than positional gaze error. We agree that tracking inaccuracies can introduce translational noise into the flow field. The 3° threshold was established based on the eye tracker’s specifications and the naturalistic setup (1-meter viewing distance without head stabilization). Across all participants, the mean positional error ranged from 1.016° to 1.5° (1 deg is 2.08 cm in our setup). We also calculated retinal slip values, which ranged from 0.12 to 0.27 deg/s (X dimension) and 0.12 to 0.23 deg/s (Y dimension). These values are comparable to natural oculomotor drift (Kowler et al., 1979) and are understandably small given the low velocity of the fixation target. We have added this information about retinal sleep at the beginning of the results section.

      Consequently, it is highly unlikely that retinal slip influenced the results. Furthermore, assuming that tracking error remained consistent across fixation conditions, any present retinal slip cannot explain why the bias followed the retinal curl manipulation as predicted by the controller. We therefore consider retinal slip to be an unlikely confounding factor.

      (5) I found the sections of text comparing the separate and joined fits (starting line 287) to be a bit too rosy. The authors show the separate fits in the main text, and it is not very surprising that these fits are good, given that the model has 30 parameters, and these data are pretty low-dimensional. The authors only show the joined fits in the supplement, and they say that they are almost as good as the separate fits (indeed, they are better in a model comparison sense, but this is 30 parameters vs. 2 parameters). However, when I look at the fits of the joined model in the supplement, I don't find them to be very impressive. In particular, the model grossly misses the data for the straight paths for several subjects (e.g., id5, id6, id8, id10). And fitting the straight paths would presumably be easiest. This implies that the joined model is really missing something and that fitting the curved paths interacts strongly with fitting the data for different fixation target locations on the straight path. I think that the authors should discuss the results a bit more soberly and tone down their conclusions here.

      We thank the reviewer for the opportunity to clarify the logic behind our modeling choices. We acknowledge that the “separate fits” are inherently less informative due to the high number of free parameters relative to the data. Our primary scientific goal was not to achieve perfect descriptive accuracy via 30 parameters, but to test a specific functional hypothesis through the “joint fit.”

      The Logic of the Joint Fit:

      We agree with the reviewer that the joint fit misses some paths in some conditions. Of course, the joint fit reflects a significant compromise. The “Gain” (the weighting of the curl signal) is likely not a static constant but is dynamically tuned based on task demands, confidence in the visual signal, simulated speed, and so on. By using a single Gain parameter, we intentionally ignore this contextual variability to see how much of the behavior can be explained by a “minimalist” controller. In this sense, the 2-parameter joint model is a deliberate attempt to test this limit. By forcing a single Gain parameter to account for all conditions across both straight and curved paths within one flow manipulation (e.g. unaltered flow) we are asking if a single, fixed linear relationship between retinal curl and steering effort/gain can explain the results. We view the joint fit not as a “perfect” model, but as a stronger test of the curl-based control theory. The fact that a 2-parameter model can capture the direction and scale of biases across such a diverse set of conditions (straight/curved paths, five fixation eccentricities) suggests that retinal curl is a robust signal. Upon closer analysis, these discrepancies between the joint model and the data are most pronounced in the over-cancelled condition which is the one when sensory evidence becomes more ecologically inconsistent with the extra-retinal information (gaze direction). While the joint fit successfully demonstrates that a single parameter can capture the general functional role of curl, it fails to account for the complex sensory re-weighting that occurs in ecologically inconsistent conditions (like ‘over-cancelled’ flow). We have updated the manuscript to discuss these limitations in the “fitting the controller” section, framing the model as a parsimonious first-order approximation rather than a complete description of human heading perception based on a minimal set of parameters.

      (6) The section of the paper on neural simulations (starting line 387) has a few weaknesses. First, why are only straight paths simulated here? This does not seem to provide a very rigorous test of the model. Second, it is awkward that the simulation results are presented in units of pixels, rather than degrees. Third, the authors seem to downplay the fact that the neural estimates of heading seem to oscillate rather wildly (over a range of hundreds of pixels, whatever that means, see especially Figure S16). It was far from clear to me how an estimate of heading with these large oscillations is useful. It would seem to require that heading estimates are integrated over substantial lengths of time to be reliable. It was therefore unclear how the model produces such smooth paths from these oscillating estimates.

      We acknowledge that the presentation of the neural model requires more clarity regarding its objectives and its relationship to the behavioral data.

      We first wish to clarify the intended scope of the neural ring-attractor model. Our primary goal was not to provide a comprehensive account of behavioral performance across all conditions (which is the role of the controller model), but rather to demonstrate a biologically plausible mechanism that explains the emergence of the “Opposite-to-Gaze” bias. While the controller demonstrates that the bias follows a specific control law, the neural model shows how such a law can emerge from known primate neurophysiology, specifically, spiral-tuned MSTd neurons, gaze-contingent inhibition, and an egocentric “straight-ahead” prior.

      Why Straight Paths are Sufficient for this Objective. The reviewer asks why only straight paths were simulated. In our study, the straight-path condition with eccentric gaze is the purest test of the bias mechanism. Simulating the straight paths allowed us to isolate the interaction between foveal inhibition and the straight-ahead prior without the confounding variable of path-curvature flow. Given the complexity of the neural network’s parameter space, we focused on these conditions to provide a clear neuro-plausible explanation. We have added text when introducing the model (Neural simulations in the Results section) to make clear why we model straight paths only.

      Units: Pixels vs. Degrees. We acknowledge that the use of “pixels” in the plots of internal neural dynamics may appear awkward. The neural network operates on input stimuli that are defined by the pixel resolution of the videos used in the simulations, we used pixels as the native coordinate system to describe the movement of activity peaks within the network’s internal “map.” We have decided to keep the pixel units in these figures.

      Behavioral Output (Meters): Importantly, the final heading estimates produced by the network are not left in pixels. We use a pinhole camera model to reconstruct the 3D trajectories from the neural activity. These results are expressed in meters, allowing for a direct comparison with the human behavioral data.

      Addressing Wild Oscillations and Smooth Paths. The oscillations observed in the instantaneous heading estimates reflect the stochastic nature of the population peak when tracking high-frequency sensory inputs. In our model, the synaptic time constant (τ) was kept relatively small to ensure a fast, low-latency response to changes in self-motion. While increasing τ would have produced smoother internal dynamics, it would also have introduced delays into the control loop. Instead, we chose to maintain this high sensory responsiveness and applied a temporal moving average later to the network’s decoding to reconstruct the 3D trajectories. This is explicitly stated in the section “Heading Estimation and 3D path reconstruction” in the new appendix 2.

      In addition, the neural activity over time is shown in two ways: the heatmap shows the neuron with preferred heading (one can see more oscillations, specially when the fixation point is closer to the centre (eccentricities -2 and 2), due to larger competition between the sensory evidence and the straight-ahead prior. The other way is the decoded heading. In the ring-attractor model, the decoded heading (φ̂) is not determined by a single neuron but is calculated using a population vector average (equation 19). By summing across the entire population, the decoder effectively integrates sensory evidence from many neurons simultaneously. One can appreciate (see e.g. Fig. 5B) that averaged decoding, leads to a smoother resulting estimate (the white dashed line, whose visibility had been improved in the revised version). Behavioral work by Burr and Santoro (2001) suggests that global motion signals (divergence and rotation in optic flow) are integrated over much longer timescales—roughly 1000ms to 3000ms—compared to local motion units (~200 ms).

      In the previous manuscript, we discussed this aspect in lines 424-426. In the new version, we have added text in the Heading estimation and 3D path reconstruction section (now in appendix 2) stating that we smoothed the decoded signal in agreement with this psychophysical evidence before applying the camera model.

      See also our comment on temporal integration in the responses to reviewer #3

      Reviewer #2 (Recommendations for the authors):

      (7) Line 51: "...a functional role of rotational flow components has been largely neglected in both theoretical and experimental work on heading perception." I feel like this statement is too strong and that the authors try too hard to "sell" their findings by underrepresenting previous work. There are numerous studies (many not cited), both behavioral and electrophysiological, that have examined how heading perception depends on pursuit eye movements, either physical movements or visually simulated ones. These studies directly involve rotational flow components, and several of them have concluded that rotational flow components contribute to estimating heading in the presence of eye movements (just one example is Grigo and Lappe 1999). Because these studies generally involved horizontal pursuit of a target on the horizon, rather than tracking a point in the ground plane (like the authors' work), these studies generally did not involve flow fields with retinal curl around the fixation point. But I consider these older studies just a special case of the more general geometry, and they still involve rotational flow components. Moreover, various previous studies have used stimuli for which there was no FOE present in the visible display (either due to simulated rotation or masking out the FOE), and the authors do not seem to give credit to these works either. In addition, several studies have implicated a role of extraretinal signals in perceiving heading during eye movements, so retinal curl cannot explain everything. Rather than overemphasizing the limitations of previous work, the authors would be better served to explain how their findings extend and generalize from these previous studies.

      We thank the reviewer for pointing out this oversight; it was not our intention to overlook previous work. While our original version cited studies considering rotation-related cue, we have now substantially revised the introduction to include previous work and better acknowledge the informative role of rotation. Our central aim remains to distinguish between models that compensate for rotation to recover a heading vector and our proposal that the visual system exploits retinal curl directly as a primary, functional signal for locomotor control.

      We have now updated the Introduction and Discussion to better situate our work within the context of studies (including Grigo & Lappe, 1999 and some additional ones we have included in the new version) that have investigated the informative role of rotational flow. We now clarify that our study extends these findings by investigating the non-uniform rotational patterns (curl) that emerge during ground-plane fixation, representing a more general and biologically ubiquitous case of locomotor control, while acknowledging previous studies that have also considered the potential role of curl generated by gaze fixation.

      (8) Figure 3: I did not understand why there are two purple and two blue curves in the graphs of the middle column. And the caption does not explain this.

      This a very good observation. This was explained in lines 268-273 (previous version). When the gaze is straight-ahead (same direction as heading), there is no curl. However, we introduced positive or negative curl in the altered conditions. These purple and blue lines refer to these trials and show that the bias re-appears in the expected direction when curl is (unexpectedly) added. We think this adds additional evidence to the curl contributing to heading. Even though this was extensively explained we have added text in the caption of figure 3.

      (9) Line 331: What makes the authors think that retinal curl is computed in area MSTd? They should cite studies to support this idea if it has been shown in physiology.

      Evidence was cited in the introduction (Graziano et al 1994) of sensitivity to spiral motion in addition to neuro-computational models that implement this activity also cited (e.g. work of Leyton et al.)

      (10) The neural network model for computing heading from curl requires a "gaze-centered inhibitory drive" that inhibits activity around where the eyes are looking. This is probably a biologically plausible thing, but is there any evidence to support the idea that this signal exists in the parts of the brain where the authors believe these computations to be happening? They simply posit the existence of this gaze-centered inhibition as though it is common knowledge, but they provide no citations nor discuss any previous evidence for its existence.

      While neurophysiological evidence primarily describes this as gain-field modulation, this process frequently involves localized suppression of neural activity to facilitate coordinate transformations. In parietal areas such as LIP and 7a, eye-position signals do not just enhance responses but can also suppress them, effectively shifting the 'center of gravity' of a population response (Read et al 1997; Born et al 2005, cited in the discussion in the revised section re-evaluating the Focus of Expansion). In the context of our ring-attractor model, this functional modulation is most parsimoniously implemented as a gaze-centered inhibitory drive.

      (11) Lines 482-483: Why should perceptual biases related to retinal curl take seconds to show up?? The curl information itself must be present very quickly, perhaps requiring just a few video frames. So what does this imply about mechanisms? The authors throw out this assertion, but it is left hanging without any further analysis or support.

      The time course reflects the integration requirements of complex motion processing. While local flow is processed rapidly, global patterns like retinal curl require longer temporal windows to reach a stable estimate (Burr et al 2001). In our study, this integration is functionally necessary to filter the higher-frequency 'wobble' induced by gait-cycle oscillations. We now discuss the temporal integration aspects under a new heading in the discussion.

      (12) Line 508: "This suggests that the "bias" observed in our perceived headings may reflect the operation of a control law optimized for action rather than a failure of a perceptual system designed for passive estimation." The authors make this statement to justify why perceptual biases are present with unaltered curl. But I don't fully understand the logic. Are they saying that it is not possible to have a set of computations that can do both things accurately? Is it possible to show this theoretically? Moreover, if it is not possible to rule out other possible sources of the biases, such as those described above (reference frame of judgments, eye movements, etc), then is it necessary to invoke this logic?

      Our logic is that the observed 'bias' is not a representational failure, but a functional byproduct of a control law optimized for active steering. In a closed-loop system, the objective is to null the error signal (retinal curl) to maintain a stable path. When observers are asked to make an open-loop heading report, they likely utilize this same control signal, which manifests as a systematic bias toward the 'null' point of the controller as a result of a sustained fixation in discrepancy with the simulated translation/heading.

      We do not suggest that accurate perception and control are theoretically incompatible; rather, we suggest that perception and action rely in the same underlying information (e.g. work of Brenner & Smeets). While other factors, such as coordinate transformations between reference frames, certainly can contribute to the reporting process, our interpretation provides a parsimonious link between the psychophysical data and the underlying steering mechanism. By framing the bias as a consequence of a 'nulling' strategy, we explain not just the existence of the error, but its specific direction and magnitude relative to the fixated target.

      (13) Line 546: "...MSTd would simultaneously code curvature for trajectory estimation and heading across the neural population, with curvature encoded through the spirality of the most active cell and heading through the visuotopic location of its receptive field center." The latter part of this argument seems to imply a relationship between the heading preferences of MSTd neurons and the locations of their receptive fields. I am not aware of any evidence for such a relationship, so the authors should indicate whether this is based on some experimental data or just a speculation.

      We thank the reviewer for this observation. The proposal that heading is signaled by the visuotopic location of active MSTd populations is a core architectural feature of our model and is supported by several lines of evidence.In the Layton and Browning (2014) framework, MSTd is modeled as a visuotopic map of functional 'hypercolumns'. Each hypercolumn contains neurons tuned to a continuum of spiral patterns, but all neurons in a given hypercolumn share a receptive field center at a specific location in visual space. Consequently, the visuotopic location ($x, y$ coordinates) of the maximally active hypercolumn represents the center of motion (heading), while the spirality (the tuning dimension within that hypercolumn) represents path curvature. We have clarified this in the revised discussion (re-evaluating the FoE) to emphasize that this dual-coding scheme arises from the simultaneous representation of 'where' (population map location) and 'what' (spiral tuning) in MSTd.

      Reviewer #3 (Public review):

      The primary limitation of the paper is that it avoids discussion of some of the inevitable complexities of heading perception. The main issue is what exactly is meant by heading. Different behaviors evolve over different timescales. The geometry of retinal motion defines instantaneous heading, which varies widely through the gait cycle. Time-varying information like this is known to be important in the momentary control of balance. Heading can also be thought of as steering the body toward a distant goal, which evolves over longer timescales. The current manuscript appears to be concerned with heading information integrated over a few seconds and seems to provide evidence that heading is indeed integrated over the gait cycle. The issue of the time scale of the computation is touched on, but it is not related to how it might be used in normal walking or what situations it might apply to. Steering toward a distant goal during walking is not a very difficult problem and may not require evaluation of retinal motion, but control of balance is more challenging and may depend critically on curl. Consequently, the timescale of the computation needs to be considered in order to understand what is meant by heading.

      We thank Reviewer #3 the comments regarding the definition of heading at different time scales, the role of the gait cycle, and the temporal integration of the curl signal. These comments have helped us refine the manuscript’s core arguments.

      We agree that “heading” must be precisely defined within the context of the differing temporal demands of balance and steering. While instantaneous heading provides the high-frequency feedback necessary for momentary postural adjustments and balance, our study is concerned with heading as a gaze-relative signal used for the continuous control of a locomotor trajectory. As such, we have revised the manuscript to specify that the perceived heading measured in our task reflects a signal integrated over the gait cycle to filter out the oscillatory noise induced by head bob and sway (mainly in the Discussion section).

      The reviewer correctly notes that gait-induced head bob and sway produce high-frequency oscillations in the curl signal, yet our behavioral results show smooth, slowly evolving biases. The visual system does not react to “instantaneous” curl, which would lead to jittery, unstable heading estimates. Instead, it integrates flow over a timescale roughly commensurate with a full gait cycle (~500–1000ms). This implies a significant temporal integration process. This temporal integration is consistent with evidence (Burr and Santoro,2001, Vis Res) indicating that optic flow signals (radial and rotational components) are integrated over windows of approximately up to 3 seconds to ensure perceptual stability. Neurally, this likely involves the projection from area MSTd to the Ventral Intraparietal area (VIP), a pathway where fast, eye-centered sensory inputs are transformed into stable, body-centered representations suitable for guiding long-term steering behavior (Chen et al. 2011, JNeurosci.). By grounding our definition of heading in these specific temporal and neural constraints, we tried to clarify how the visual system exploits retinal curl for goal-directed action in natural, dynamic environments and relate our findings to recent studies addressing the role of retinal motion on balance (Powell et al. 2026 Bioarx).

      In our implementation, we explicitly address the high-frequency noise introduced by gait dynamics by smoothing the retinal curl signals computed from the stimulus videos before they are fed into the controller. This temporal filtering allows the fit of the controller’s prediction to the response data while remaining robust to the rapid fluctuations of head bob and sway. In contrast, the neural ring-attractor model would not require an external smoothing step; instead, the integration is an emergent property of the system’s architecture that can be controlled with different parameters, as commented above in a response to Reviewer #2. The dynamics of the synaptic weights and the characteristic “leak” in the population activity naturally implement a leaky integration of sensory evidence, ensuring that the decoded heading reflects a sustained estimate rather than an instantaneous response to visual noise.

      We also agree that we avoided discussing some complexities of the heading perception. In the new version, we also include and integrate the distinction between instant heading and future path in different parts of the ms (introduction) and mainly discussion (temporal integration and steering section) which have been revised substantially.

      Reviewer #3 (Recommendations for the authors):

      There are a number of points that require clarification.

      (1) Head bob and sway were included in the stimulus and need to be addressed in both the analysis of the data and the interpretation. The curl signal in the stimulus varied over time, commensurate with normal gait. However, the results don't reflect the same level of variability that would be produced from curl over the gait cycle. This means that the information must be integrated over some longer timescale. It is not clear from the data analysis what this integration is. Is there an implicit integration with the manipulation of the steering wheel? If subjects indeed appear to be able to use curl to evaluate heading over timescales of seconds, this needs to be explicitly addressed, as it is a novel result. This would require parts of the discussion to be changed/expanded to maintain consistency. For example, line 482 talks about the buildup of biases over time.

      As commented above in the public response, the curl estimated from the optic flow algorithm was smoothed before being input into the controller (path fitting and predictions). The smoothing was only applied to the curl signal, not to the participants responses. Also, as mentioned before, the time course of the bias is consistent with integration times of optic flow reported in the literature. All these aspects are now explicitly included in the new display and conditions section (Flow manipulation conditions).

      (2) Since the experiment included curl variability resulting from the gait cycle, some discussion is needed about the role of retinal motion in the control of balance and posture. There is a large literature about the role of flow in controlling gait and momentary adjustments of the body while walking. Additionally, it should be noted that in the task, head bob and sway from 1 prerecorded individual was shown to all subjects. It is known that gait varies significantly between different individuals, and it should be acknowledged that this may lead to differences at the individual subject level for perceiving heading.

      We have included a point in the discussion addressing the different time scales for different use of optic flow signals (postural control vs locomotion).

      We agree with the reviewer that utilizing a single gait profile for all participants may introduce individual differences in perceived heading, as the simulated head motion might not perfectly match each participant’s unique biological gait signature. However, we prioritized stimulus consistency over idiosyncratic accuracy. By ensuring that every participant viewed the exact same motion profile, we could be certain that the systematic 'opposite-gaze' biases observed across the population were driven by our experimental manipulations of gaze and retinal curl, rather than being confounded by variability in head-motion kinematics. We have added an acknowledgement of this point at the first paragraph of the displays and conditions section in the Methods.

      (3) More information is required about the use of the rotating wheel for the measurement of heading. How easy was it to use? What about time delay, and how does this deal with the bob and sway? Does the wheel impose a de facto integration on the perceptual measurement?

      The rotary encoder provided an intuitive steering-wheel interface that participants found easy to operate. To ensure minimal latency (1–5 ms), the device was interfaced via an Arduino Uno and sampled by a dedicated background Python thread, isolated from the visual rendering loop. We have incorporated these technical details into the Methods (Procedure) section.

      (4) Restructuring the description of the models It remains unclear why the dynamics of the neural network are a necessary inclusion in this paper. It seems interesting, but there is no comparison to actual neural data or other related work. Instead, this appears to be a description of what the network is doing, which is already defined by the equations. This needs to be clarified for its exact interpretation with respect to real neural data, and its importance here for understanding the biases that emerge in heading judgments. The paper would flow better if this section were included as supplementary material or omitted from the paper entirely, as it seems to detract from the other points. If this is a description of why the biases are seen in the controller, then the supplementary material is a good place for it.

      We thank the reviewer for this suggestion. We clarify that the neural model is not intended to simulate specific empirical neural data, but rather to provide a biologically plausible implementation of the controller. This allows us to demonstrate how the observed biases emerge from the dynamics of standard cortical architectures, such as ring attractors. This is now mentioned when introducing the neural simulation results.

      Following the reviewer's suggestion, we have moved the neural model equations to Appendix 2 while retaining the simulation results in the main text (Results). We believe it is essential to present not just the abstract controller, but also its functional implementation, as this provides a mechanistic bridge between retinal signals and locomotor behavior.

      (5) For the modeling approaches, the math would be more appropriate for supplementary materials.

      To ensure a better flow of the paper, we have moved the neural model equations to Appendix 2, while Appendix 1 now details the relationship between measured curl and other variables (speed, yaw, pitch, etc.). We have retained the controller model in the Methods section, consistent with the eLife layout where Methods follows the Discussion.

      (6) How do the models and data analysis deal with the influence of the gait cycle in the input? Do they integrate the information over that timescale? If so, the integration time needs to be specified.

      As commented above, to mitigate gait-cycle fluctuations, we smoothed the computed curl signal before inputting it into the controller and applied a similar smoothing process to the neural model’s readout. Using a LOESS filter, the effective integration window was 2.4 seconds. Close to the integration time reported in Burr et al. 2001. These parameters have now been explicitly specified in the Methods section. For the empirical data, we just utilized trial-averaging.

      (7) What does biologically plausible mean in terms of the neural network model? Especially when control wasn't explicitly a variable in the measured behavior of the subjects.

      By biologically plausible, we mean that our model is constrained by neural architectures documented in the primate brain—specifically ring-attractor dynamics, population coding, and gaze-centered gain-fields. Crucially, the network utilizes recurrent connectivity with a 'Mexican-hat' profile (local excitation combined with lateral inhibition). This is a standard and widely accepted motif in computational neuroscience, representing the consensus on how cortical circuits maintain a stable "bump" of activity to represent spatial variables. Rather than introducing ad-hoc mechanisms, we demonstrate that the observed behavioral biases emerge naturally from these established neural components when they are tasked with maintaining locomotor stability. Since we think this aspect was already emphasized, we haven’t added any additional detail.

      (8) There should be more extensive acknowledgement of the body of literature that has challenged the use of the focus of expansion. That section should also include references to work that has investigated extraretinal signals, as they may also be important.

      We have expanded the Introduction and Discussion to more thoroughly acknowledge research challenging FOE-based models and the critical role of extraretinal signals. These updates, which also align with our response to Reviewer #2, provide a more comprehensive context for our model within the existing body of heading and self-motion literature.

      Minor points:

      (1) Line 69 - In self-generated motion, spiral patterns are almost always centered on the fovea, but many physiological experiments present spirals in the peripheral retina. This is incompatible with the motion generated during self-motion. Therefore, clarify whether the type of spiral motion Graziano investigated was centered on the fovea.

      In the experiments conducted by Graziano et al. (1994), spiral stimuli were centered on the receptive field (RF) of the individual neuron being recorded to accurately characterize its tuning. While the reviewer correctly notes that spiral centers often align with the fovea during active steering (due to fixation on a goal), MSTd neurons possess large RFs that provide a comprehensive 'template' system across the visual field. This population-level representation allows the brain to recover trajectory information even when the focus of motion shifts relative to the fovea—for example, during pursuit eye movements or when fixating on landmarks off the direct path of travel. To keep this part of the text brief, we haven’t add more details concerning this study.

      (2) Line 72 - "Magnitude" instead of "amount".

      This has been changed.

      (3) Line 115 - State explicitly whether the scale of the visual stimulus was matched to the scale of the actual natural images shown in VR.

      We have updated the Methods (Displays and conditions) section to explicitly state that the visual scale was veridical. The virtual camera’s parameters were calibrated such that its field of view (91°) matched the physical dimensions of the projection screen (2.03 m × 1.16 m) at the 1.0 m viewing distance. This ensures that the angular size of the objects and motion gradients in the stimulus were 1:1 with the scale of the simulated natural environment.

      (4) Line 144 - While Farneback is a good dense flow estimation algorithm, it is noisy and may impose biases/variability in the calculation of curl. This should be acknowledged.

      We acknowledge that the Farneback algorithm can introduce variability in curl estimation. To mitigate this, we utilized 10 independent renderings of each experimental trial to compute the flow fields. Although this methodology was reflected in the data previously uploaded to our OSF repository, we have now explicitly added this detail to the manuscript (Flow (curl) manipulation conditions). The computed curl used for the modeling was derived from the aggregate of these different runs, ensuring a robust and stable signal that accounts for potential algorithmic noise.

      (5) Line 206 - "Join fits". Is this a technical term? It sounds awkward. Would "Joint fits" make more sense?

      The referee is right. We have corrected this.

      (6) Line 280 - "Consistent with" (typo).

      This has been corrected.

      (7) Line 286 - Typo in title.

      Also corrected to Fitting the controller.

      (8) Lines 482-493 - There should be more discussion on the time course of integrating the stimulus, and the relationship/generalizability to more natural stimuli.

      This part of the discussion (related to integration time) has been changed considerably to include discussion of postural control in addition to locomotion.

      (9) Lines 516-517 - It is mentioned that retinal flow dynamics override the visual direction cue. This may not be generally true, as the reweighting of the cues might depend on things like task demands or actual stimulus context. In the present experiment, the subject only has access to a large moving textured ground plane, and the body is stationary.

      We agree with the reviewer that cue reweighting is highly context-dependent. However, as noted in the original manuscript (Lines 516-517), we specifically stated that retinal flow dynamics 'can' override the visual direction cue, rather than asserting a universal rule. This phrasing was intentional to acknowledge that while flow is a potent signal—especially in the presence of a large, textured ground plane as used in our paradigm—the relative weighting of these cues remains contingent on the specific sensory and task conditions. We believe this remains a fair and cautious interpretation of our findings.

      (10) Lines 524-256 - It is unclear why Matthis et. al. 2022 is cited for this point. Some of the steering literature, like Wilkie Wann & Allison 2006 or Lappi & Mole 2018 (and some of their other work), would be more relevant for the definition of a control law under these circumstances.

      We agree and this part has been changed substantially.

      (11) Lines 528-532 - Warren et. al. 2001 should be cited in this section because their results were interpreted in terms of focus of expansion, but may result from the curl signal (Powell et. al., 2026. The Role of Retinal Flow in Walking. bioRxiv, 2026-02.).

      We agree with this suggestion and the citation has been added.

      (12) Lines 544-546 - Layton and Browning are focused more on path perception from the implemented spiral tuned cells. Because of this, it wouldn't be appropriate to say it is a shift away from FOE-based heading based on this citation alone. There are more models/psychophysical results that would strengthen this claim.

      We agree with the reviewer that Layton and Browning focus specifically on path perception. Our original intention in citing this work was to emphasize the neurophysiological continuum from radial to circular motion (spiral tuning) rather than discrete expansion/rotation channels. However, we have revised this section to clarify that the reliance on retinal curl represents a mechanism for determining the future path (locomotor trajectory) rather than merely instantaneous heading. This distinction acknowledges that while heading is a momentary vector, the integration of curl signals allows the system to anticipate and control the intended path over time—a framing that better aligns with both the cited literature and our proposed controller model. As commented above this is now extensively discussed in the revised version.

      (13) Figures - There are minor visibility issues for some of the figures. In Figure 2, the thick line is unreadable, and in Figure 1, the x-axis labels are crowded.

      The axis in Fig 1 has been modified to avoid crowdedness. In Fig. 2, the thick (average line) has been modified. We hope they are more visible now.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Poh and colleagues investigate dopamine signaling in the nucleus accumbens (ventromedial striatum) in rats engaged in several forms of Go/No Go tasks, which differed in reward controllability (self-initiated reward seeking or cue-evoked/quasi-pavlovian), and in the specific timing of the action-reward contingencies. They analyze dopamine recordings made with fast scan cyclic voltammetry, and find that dopamine signals vary most consistently to cues that signal a required action (Go cues) vs cues signaling action withholding (No Go cues). Through various analyses, they report that dopamine signals align most clearly with action initiation and with the approach to the reward-delivery location. Collectively, these data support aspects of a variety of frameworks related to accumbens dopamine signaling in movement, action vigor, approach, etc.

      Strengths:

      These studies use several task variants that consolidate a few different components of dopamine signal functions and allow for a broad comparison of many psychological and behavioral aspects. The behavioral analysis is detailed. These results touch on many previous findings, largely showing consistent results with past studies.

      Weaknesses:

      The paper could heavily benefit from some revision to 1) increase clarity of the figures, the methods, and the analysis. 2) The inclusion of many tasks is a strength, but also somewhat overshadows specific points in the data, which could be improved with some revision/reworking. 3) Some conclusions are not fully justified. As shown, support for the conclusion "dopamine reflects action initiation but not controllability or effort" is lacking without more analyses and additional context. 4) Further, the notion that the dopamine signals reported here reflect spatial information could be justified more strongly.

      We thank the reviewer for their detailed evaluation and constructive feedback. We have made substantial revisions to address each concern raised:

      (1) Clarity of figures, methods and analyses

      We have revised the organization of the panels in Figure 1 for clarity.

      We have revised Figure 2 and its caption: we have labelled all comparisons depicted in the figure, and now included a line, “Subject-wise comparisons of dopamine data were made for all alignments”, to clearly show that all statistical tests shown in Figure 2c-e were performed between subjects.

      In the caption of Figure 3, we have now added, “... and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” to clearly show that statistical tests were performed on the trial-wise level for Figure 3c.

      We have now added a table to the Methods section (Table 1), detailing the sample size in each task variant and the number of trials within each No-go classification.

      We have added more information in the Methods section: we now include Videos to show the classified No-go behaviors and other trial types (Go and Free); and provide a schematic of the DLC workflow in the supplementary materials (Supplementary Figure S9).

      (2) Strengthening specific points in the data

      To improve clarity of our main findings, we have revised the layout of the Results section such as including more descriptive headers:

      “Behavioral performance in Go/No-go (“short task”) was unaffected by controllability of reward pursuit”

      “Behavioral performance in Go/No-go/Free (“long task”) was unaffected by controllability of reward pursuit”

      “Motivation to approach the reward magazine was similar between Go and No-go trials”

      “VMS dopamine release encodes reward-related action initiation”

      “VMS dopamine release does not only reflect reward-related action initiation”

      “Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards”

      “Motivational state reflected by No-go behavioral strategy correlates with dopamine signal size during reward approach”

      (3) (4) Conclusions drawn from data

      We have carefully revised our conclusions to accurately reflect our experimental design and analyses, emphasising our core finding that VMS dopamine was consistently increased in Go versus No-go trials, throughout manipulation of the type of trial start (self- and cue-initiated) and effort manipulation (short and long task variants).

      We thank the reviewer for the feedback regarding our VMS dopamine signals reflecting spatial proximity to reward. We performed additional analyses, which we present in Supplementary Figure S10, S11 and S12, to support our interpretation that VMS dopamine encodes spatial proximity to reward.

      We appreciate the reviewer comment relating to the statement, “Dopamine reflects action initiation but not controllability or effort". We have revised the wording of our conclusion to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). This subheading is now changed in the Discussion to “Reward-related dopamine depends on action initiation irrespective of controllability and effort”. The additional analyses that we have performed are shown in Author response image 1entitled: “Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long).

      Additional details on subjects used in each study, analysis details on trialwise vs subjects-wise data, and other context would be helpful for improving the paper.

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures in the original submission of this manuscript.

      To improve the paper, we now include a table in Methods Section 6: Statistical Analysis (Table 1), detailing the sample size in each task variant (with FSCV recordings) and the number of trials within each No-go classification.

      To give more context, we made Author response image 1 to illustrate the number of subjects in each task variant (and their overlap):

      Author response image 1.

      Number of subjects included in each Go/No-go task variant (total n = 21). Values (n) depict overlap of each subject between task variants. One animal was recorded in self-initiated Go/No-go and Cue-initiated Go/No-go/Free (dotted line with arrowheads).

      Reviewer #2 (Public review):

      Here, the authors record dopamine release using fast-scan cyclic voltammetry in the nucleus accumbens/ ventromedial striatum (VMS) while rats perform variants of a Go/No Go task. Two versions are self-paced, in that the rat can initiate a trial by nosepoking at the odor port at any time once the ITI has elapsed, whereas the other two require the rat to wait for a cue-light before responding. Two "long" variants also require either more lever-presses on Go trials, or a longer nosepoke time for No Go trials, and also incorporate "free" trials in which the rat is rewarded for just heading straight to the food tray. The authors find that dopamine levels increase more during the response requirement for Go than No Go trials, indicating a role for invigorating to-be-rewarded actions. Dopamine levels also steadily increased as rats approached the site of reward delivery, and the authors demonstrate quite elegantly that this was not due to orientation to the food tray, or time-to-reward, or action initiation, but instead reflects spatial proximity to the rewarded location. Contrary to previous reports, the authors did not discern any differences in dopamine dynamics depending on whether the trials were cue- or self-paced, and dopamine release did not scale with effort requirements.

      The manuscript is well-written, and the authors use figures to great effect to explain what could otherwise be a hard-to-parse set of data. The authors make good use of the richness of their behavioral data to justify or negate potential conclusions. I have the following comments.

      Re: The lack of relationship between effort to acquire reward in the current study and the magnitude of dopamine release, 1) can the authors unpack this a bit more? 2) Why the difference between the Walton and Bouret studies? Were the shifts in effort requirements comparable across the behavioral tasks? 3) What else could be different between the methodologies?

      We thank the reviewer for the feedback and have responded to each of the three questions below (see points 1-3).

      Firstly, we tried to improve the clarity of our research aims. Our primary comparison throughout the manuscript is between Go versus No-go within each task variant. We ask whether the Go/No-go difference in dopamine signaling persists across different response demands. Thus, testing effort was not central to this main question, but rather a feature of the task that did not affect the Go-No-go dopamine difference.

      (1) Consistent with this aim, we show that VMS dopamine differs between Go and No-go actions persistently across all task variants despite differences in response requirements (action was always accompanied by greater dopamine release compared to action suppression). Our behavioral-training data suggest that the ability to perform short and long tasks differed: rats were first trained to criterion on either a short (∼2 s) or long (∼3 s) Go/No-go variant, with the longer variant requiring substantially more training sessions (short: 18.5 ± 7.6 sessions vs long: 41.1 ± 7.9 sessions; see Author response image 2), indicating behavioral demands were higher for the long-task.

      Author response image 2.

      (2) Regarding the apparent discrepancy with Walton and Bouret (2019), we acknowledge that our original description was imprecise (We wrote: “Previous studies have shown that dopamine signals are influenced by the effort required to obtain rewards”). Our intent was not to suggest a direct contradiction, but rather to emphasize that our findings are consistent with the paper’s broader conclusion that effort encoding by dopamine is limited and highly context-dependent. We have now adjusted the manuscript to better reflect our intent by changing the sentence to, “Previous studies have shown that dopamine signals may be influenced by the effort required to obtain reward but only for particular task conditions (Cousins et al., 1996; Gan et al., 2010; see for reviews, Salamone and Correa, 2024; Walton and Bouret, 2019).”

      For added clarity, these were the main results highlighted in the Walton and Bouret review: Gan et al. (2010) demonstrated that VMS dopamine sensitivity to low-effort costs is prominent early in training (≤ 2 training sessions) and diminishes after extended experience (> 9 sessions). Similarly, Hollon et al. (2014) reported that cue-evoked VMS dopamine primarily tracks reward magnitude with minimal modulation by effort. In line with this literature, our rats were highly trained (≥ 9 sessions until the first recording), and exhibited no difference of average Go minus No-go dopamine between short and long task variants within controllability type (see Author response image 3), supporting the idea that extended training exhibits minimal effort-related modulation of VMS dopamine.

      Author response image 3.

      Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of controllability and effort showed moderate evidence of absence (controllability: BF<sub>10</sub> = 0.324; effort: BF<sub>10</sub> = 0.309). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.101), and the full model including a controllability × effort interaction was the least supported of all models examined (BF<sub>10</sub> = 0.043). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither controllability, effort, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (3) With respect to task comparability and methodological differences, our behavioral paradigm differs in important ways from those highlighted by Walton and Bouret, where effort was often manipulated by training animals to associate cues with different numbers of lever presses within the same session, and typically involved only action initiation. In contrast, our task required both action initiation and action suppression, and changes in response contingencies occurred across separate recording sessions rather than within-session cue-based manipulations. Although these paradigms are not directly comparable, and only had the same dopamine recording technique in common (FSCV), a key takeaway of our results is that regardless of effort differences, VMS dopamine during action initiation is consistently higher than during action suppression.

      I would argue that the cue- vs self-initiated distinction was pretty minor, given that there was a fixed ITI of 5s. How does this task modification compare to those used previously to show that dopamine release corresponds to behavioral controllability? It would help the reader if the authors could spend more time discussing these disparate findings and looking for points of methodological divergence/commonality.

      We agree that clarifying how our manipulation of controllability compares to prior work improves the manuscript, and we have made the necessary adjustments. However, we would first like to correct an incomplete characterization of the task design.

      While the short-task variant used a fixed 5 s inter-trial interval (ITI), the long-task variant employed a variable ITI ranging from 15–25 s. In the long-task variant, the timing of trial onset was less predictable, and we believe this manipulation reduced animals’ ability to precisely estimate when reward pursuit could begin. Under these conditions, whether trials were Self-initiated or Cue-initiated had a substantial impact on animals’ control over the initiation of reward pursuit. That said, we agree that the Self- versus Cue-initiated distinction overall represents a moderate manipulation of controllability compared to those used in studies that focus on controllability.

      A key source of divergence across studies lies in the definition of controllability. We defined controllability as the animals’ ability to choose the time point of beginning the reward pursuit, rather than whether an action was required, and have now added the following sentence in the:

      - Introduction section: “... controllability of reward seeking, defined as the ability to determine when to initiate reward pursuit (Self- vs Cue-initiated trials)...”;

      - Results section: “We defined controllability as the rats’ ability to choose the time point of reward pursuit. In Cue-initiated trials, the time point at which trials could be started was dictated by a cue light, whereas in Self-initiated trials rats were able to choose intrinsically (control) when to attempt a trial start.”;

      - Discussion section

      Importantly, the action requirements for Go, No-go, and Free trials were identical across these trial-start conditions. We found that the degree to which controllability was manipulated in our task was insufficient to modulate the action-specific VMS dopamine signal (Go vs No-go difference), which remained robust across conditions.

      In contrast, controllability has been defined by others as the presence versus absence of an operant action requirement for reward. For example, Goedhoop et al. (2023) directly contrasted operant (lever press required) and Pavlovian (no action required) conditions, removing action execution as a prerequisite for reward. In that context, cues signaling operant control elicited sustained VMS dopamine release, which was interpreted as reflecting anticipation or preparation for executing a learned action. Similarly, Hamid et al. (2021) demonstrated that dopamine “wave” directionality across striatal regions depends on controllability defined by operant versus Pavlovian conditioning.

      Taken together, these comparisons (results from the present study and in the literature) suggest that dopamine sensitivity to controllability may depend on how it is manipulated. We have clarified these methodological distinctions in the revised Introduction, Results and Discussion, and emphasized that more extreme manipulations (such as removing action requirements entirely or increasing uncertainty over trial timing) may be necessary to reveal controllability-dependent changes in VMS dopamine signaling. Alternatively, the apparent discrepancies across studies may primarily reflect differences in the underlying definitions of controllability rather than conflicting results.

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Poh et al. investigated whether dopamine release in the ventral medial striatum integrates information about action selection, controllability of reward pursuit, effort, and reward approach. Rats were implanted with FSCV probes and trained in four Go/No Go task variants:

      (1) trials were self-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (2) trials were cue-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (3) trials were self-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued, and effort was increased,

      (4) trials were cue-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued.

      The authors report that dopamine levels rose during Go trials and slowly rose in No Go trials, but this pattern did not differ across task variants that modified effort and whether trials were cued or initiated. They also report that dopamine levels rose as rats approached the reward location and were greater in rats that bit the noseport while holding during the No Go response.

      Strengths:

      (1) Interesting task and variants within the task paradigm that would allow the authors to isolate specific behavioral metrics.

      (2) The goal of determining precisely what VMS dopamine signals do is highly significant and would be of interest to many researchers.

      Weaknesses:

      (1) This Go/No-Go procedure is different from the traditional tasks, and this leads to several problems with interpreting the results:

      (a) Go/No Go tasks typically require subjects to refrain from doing any action. In this task, a response is still required for the No Go trials (e.g., continue holding the nosepoke). The problem with this modified design is that failure to withhold a response on No Go trials could be because i) rats could not continue holding the response, as holding responses are difficult for rodents, or ii) rats could not suppress the prepotent go response. This makes interpreting the behavior and the dopamine signal in No Go trials very difficult.

      We appreciate the reviewer raising this important methodological consideration. We acknowledge that our Go/No-go task differs from traditional paradigms used in humans and primates (e.g. Raud et al. 2020, 10.1016/j.neuroimage.2020.11658; Eagle, Bari & Robbins 2008, 10.1007/s00213-008-1127-6; Roitman & Loriaux 2013, 10.1152/jn.00350.2013).

      However, our design addresses the specific constraints of studying dynamics in freely moving rodents while maintaining the core feature of Go/No-go tasks: requiring suppression of a prepotent response. Our task accomplishes the primary aim of our study, which is to compare VMS dopamine dynamics during action initiation and action suppression, and below we list the reasons why. Therefore, we do not believe that this difference compromises the validity and interpretation of our results.

      It has been suggested for decades that the two main processes governed by mesolimbic dopamine are reward learning and motivated action, and our study aimed to better understand how VMS dopamine integrates reward-related information and motivated action, rather than studying them in isolation. To do so, we trained rats in a modified Go/No-go task.

      More recent work (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019) demonstrates that VMS dopamine signaling incorporates both action and reward-related information, rather than either of the two alone. Importantly, in freely-moving rodents, examining this relationship requires preventing the approach response that occurs when reward delivery is anticipated (Pavlovian bias, go for rewards). Traditional Go/No-go designs that simply require "doing nothing" would not achieve this control in freely-moving rats, as animals immediately approach the reward magazine as soon as reward is inferred (as seen in our Free trials). Thus, we require a No-go condition, as we and others have defined (Syed et al. 2016), whereby animals have to actively suppress the ‘initiation’ response. Action initiation is defined at the beginning of the Discussion section: “... at two distinct points after trial start: 1) when rats began lever pressing (Go), and 2) when rats walked to the reward magazine, either without action requirement (Free) or after successful trial completion (Go and No-go)”.

      To further strengthen our interpretation that we compare action initiation and suppression, and to facilitate cross-species translation of our results (i.e., rodent to human), we also include Free trials, where reward delivery requires no specific action (which are essentially like “doing nothing” trials in traditional tasks). This addition allowed us to directly compare No-go and Free trials, where animals must actively suppress responding while maintaining task engagement, to a condition where no overt action is required for a reward, respectively. The dramatic difference in VMS dopamine between No-go and Free trials demonstrates that VMS dopamine reflects active action suppression during No-go trials, rather than merely the absence of action requirements. This has now been discussed.

      Finally, we only report correct Go, No-go, and Free trials, which differs from that of human go/no-go studies that focus on the failure of appetitive no-go trials (i.e., inhibiting the pre-potent response). In the present study, the dopamine signals that we interpret are restricted to successful trials only: action initiation (moving the lever press), action suppression (i.e., suppressing the prepotent Go response while maintaining their position in the nose-poke port), or no action (no lever press, not staying in the port). While this design differs from human Go/No-go paradigms, we believe our study of correctly performed Go, No-go, and Free trials are necessary for isolating action-dependent components of dopamine signaling in freely moving rats (action initiation vs action suppression vs action free).

      (b) Most Go/No Go tasks bias or overrepresent Go trials so that the Go response is prepotent, and consequently, successful suppression of the Go response is challenging. 1) I didn't see any information in the manuscript about how often each trial type was presented or 2) how the authors ensured that No Go responses (or lack thereof) were reflecting a suppression of the Go response.

      We appreciate the reviewer's attention to this important methodological consideration. The originally submitted version of the manuscript already addressed both concerns raised.

      Trial type presentation frequencies

      The Methods section describes our trial presentation approach: "On recording days, the trial types were counterbalanced. Within a session, Go left, Go right, and No-go trials were presented with 33% probability each, without replacement. For sessions with Free trials, trials were presented with a 25% chance without replacement."

      This design results in overrepresentation of Go trials overall (66% in the short-task; 50% in the long-task), which establishes the prepotent Go response as intended in standard Go/No-go paradigms. To improve clarity, this detail has now been included in the Methods section.

      Ensuring No-go responses reflect suppression of Go response

      Our paradigm incorporates multiple features that ensure successful No-go performance reflects suppression of the prepotent Go response:

      First, the overrepresentation of Go trials (addressed above) establishes response prepotency. Second, during No-go trials, rats must maintain their snout in the nose-poke port for the duration of the action cue, which creates the requirement to suppress the natural tendency to approach rewards (i.e., Pavlovian bias; Jones et al. 2017, 10.1016/j.bbr.2017.05.044; Guitart-Masip et al. 2014, 10.1007/s00213-013-3313-4; Dayan et al. 2006; 10.1016/j.neunet.2006.03.002). In Go trials, such natural bias does not require suppression as the lever can be approached and pressed during the action-cue period. This conflict between the instrumental No-go requirement and the Pavlovian-instrumental bias toward action makes action suppression particularly challenging (consistent with computational accounts of similar paradigms; Lloyd & Dayan 2023, 10.1371/journal.pcbi.1011569; Jones et al. 2017, Guitart-Masip et al. 2014, Dayan et al. 2006).

      Figure 3 provides behavioral evidence of this challenge: animals frequently left the nose-poke port and developed spontaneous motor strategies (such as biting and digging) to stay in the port, suggesting Pavlovian bias interfering with response suppression for rewards. Importantly, all reported No-go data include only correct trials (i.e., those without lever presses), ensuring that the dopamine signal reflects successful response suppression rather than failed Go attempts.

      (2) The authors observe relatively consistent differences in the DA signal between Go and No Go trials after the action-cue onset. However, the response type was not randomized between trial type, so there is a confound between trial type (Go/No Go) and response (lever/nosepoke). The difference in DA signal may have nothing to do with the cue type, but reflects differences in DA signal elicited by levers vs. nosepokes.

      As stated in the Introduction section and discussed in our rebuttal to point 1a, the focus of our investigation is how VMS dopamine signals differ during action initiation versus action suppression for rewards, as this is a central unanswered question in the dopamine field. More recent work demonstrates that dopamine incorporates not only RPE but also action initiation (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019), and our goal is to further our understanding of action-dependent VMS signals during reward pursuit.

      The reviewer suggests that dopamine differences may reflect differences in lever vs. nosepoke rather than cue type (Go vs No-go). We respectfully suggest this concern reflects a misunderstanding by the reviewer of our experimental question. The cue-action relationship is the experimental manipulation itself. It is not possible to study how dopamine encodes instructed action initiation versus suppression without linking specific cues to specific actions. The suggestion to 'randomize' action type across cue types would eliminate the very phenomenon we are investigating: how dopamine signals differ when cues instruct different action requirements.

      Our experimental design specifically compares reward pursuit with action requirements (Go trials: lever press; No-go trials: sustained hold) to reward pursuit without action requirements (Free trials: direct magazine approach). This design allows us to isolate how action initiation and action suppression influence reward-related dopamine signaling, which can reveal how the timing of action initiation influences RPE-dopamine. And which is the point of the study: to show how actions influence RPE dopamine signaling.

      Supporting this interpretation:

      Firstly, trial types were randomly interleaved, and each auditory cue explicitly instructed a specific behavioral response. Our design directly follows established methods demonstrating that VMS dopamine encodes whether actions are initiated or suppressed following action cues (Syed et al. 2016). That study, like ours, intentionally linked cue identity to a specific action requirement to assess how dopamine reflects instructed behavioral control. Thus, the fact that Go and No-go cues map onto different actions is inherent to the question being addressed, not an unintended confound.

      Second, as discussed in our response to point 1b, the asymmetry between Go and No-go trials is theoretically essential. Go trials align with Pavlovian approach tendencies (action initiation to reward), while No-go trials create conflict with this bias by requiring action suppression despite the cue being associated with a reward. This Pavlovian-instrumental conflict makes suppression particularly challenging (Lloyd & Dayan 2023, PLoS Comput Biol 10.1371/journal.pcbi.1011569) and allows us to examine the role dopamine in overriding prepotent responses.

      Third, the inclusion of Free trials (discussed in point 1a) demonstrates that our findings reflect instructed action control rather than simply motor execution. Free trials require neither lever pressing nor nose poke maintenance, yet show dopamine dynamics distinct from both Go and No-go trials, confirming that dopamine signals encode action requirements beyond motor output per se.

      Finally, we demonstrate that VMS dopamine differs in the same trial type (No-go) and can be classified based on different movement patterns (Biting, Digging, Calm). Importantly, the difference in VMS dopamine only appeared after the action was completed, particularly during reward approach (Figure 3). This data argues against the idea that VMS dopamine is particularly tied to the specific operant manipulanda as suggested by the reviewer, but rather, may reflect an internal motivational state for reward.

      Together, the aim of the present study is not to redefine Go/No-go paradigms for rodents, but to utilize this task structure to investigate action-dependent dopamine signalling for rewards, which cannot be answered without the cue-action mapping that we have used.

      (3) Both Go and No Go trials start with the rat having their nose in the noseport. One cue (Go cue) signals the rat to remove their nose from the noseport and make two lever responses in 5 seconds, whereas the other cue (No Go cue) signals the rat to keep their nose in the noseport for an additional 1.7-1.9 s. The authors state that the time between cue onset and reward delivery was kept the same for all trial types, and Figure 1 suggests this is 2 s, so was reward delivered before rats completed the two lever presses? I would imagine reward was only delivered if rats completed the FR requirement, but again, the descriptions in the text and figures are incongruent.

      The reviewer asks whether reward was delivered before rats completed the two lever presses and notes incongruence between text and figures. We respectfully note that these details were stated in the originally submitted version of the manuscript (see below).

      Reward delivery timing

      The reviewer asks whether reward was delivered before rats completed the two lever presses, which refers to the short-task variant. No - reward was always delivered immediately after the second lever press for all Go trials. This is described in Methods Section 3: Behavioral procedures - Self-initiated task variant. For added clarity, we have now added the term “immediately”: “... food pellet dispensed into the reward-magazine immediately.”

      In the Results section, we report that the average latency to complete two lever presses was 1.8s, which closely matches the 1.7-1.9s nose-poke hold maintenance required for No-go trials in the “short” variant. Thus, the time point of reward delivery was matched between Go and No-go trial types.

      Representation of reward delivery timing in figure and text

      The reviewer's confusion appears to stem from the schematic representation in Figure 1 and the task variant structure. There were two overarching task variants with different trial requirements:

      “Short-task” variants: Go trials required two lever presses (completed on average in 1.8s);

      No-go trials required 1.7-1.9s nosepoke maintenance

      “Long-task” variants: Go trials required a ‘rewarded’ press to occur 2.7-3.2s after cue onset (completed within ~3s); No-go trials required 2.7-3s nosepoke maintenance.

      For simplicity in depicting action-cue onset in Figure 2c, we used grey shading with a speaker icon at approximately 0-2s and 0-3s to represent these two variants. This schematic representation was not intended to indicate the precise reward delivery time, which (as stated in the Methods) occurred only upon successful completion of trial requirements in the short-task variant, or 2s after successful completion of trial requirements in the long-task variant. To improve clarity, we have adjusted the legend of Fig. 2 for more clarity, adding “Shaded gray area depicts approximate duration of action-cue onset for “short” and “long” task variants.”, and included more information under Results: “In the short-task variants, action-cues switched off after trial completion and a reward was delivered immediately” and “n the long-task variants, action-cues switched off after trial completion, or in the case of Free trials after 3s, and reward was delivered 2s later (Figure 1a).”

      (4) The manuscript is difficult to understand because key details are not in the main text or are not mentioned at all. I've outlined several points below:

      (a) The author's description in the manuscript makes it appear as a discrimination task versus a Go/No Go task. I suggest including more details in the main text that clarify what is required at each step in the task. Additionally, providing clarity regarding what task events the voltammetry traces are aligned to would be very useful.

      We respectfully note that the requested details were already present in the originally submitted version of the manuscript (see below). However, we acknowledge that the task design is complex and may benefit from additional clarity in the main text to aid reader comprehension.

      Behavioral task

      The reviewer suggests our task appears more like a discrimination task than a Go/No-go task. We acknowledge that our paradigm differs from traditional Go/No-go tasks used in humans and primates, as discussed in our responses to points 1a and 2. However, we classify this as a Go/No-go task because it shares the defining feature: requiring action initiation (Go) and suppression (No-go). This classification is consistent with established rodent literature examining action initiation versus suppression (Syed et al. 2016).

      Moreover, as discussed in our response to point 1a, we included Free trials specifically to demonstrate that No-go trials require active suppression rather than discrimination alone. The distinct dopamine dynamics across Go, No-go, and Free trials confirm that our task captures action initiation, action suppression, and action-free states, which is an important contrast that we needed to address our research question about action-dependent dopamine signaling.

      The key requirements for each trial type and task variant are described at the beginning of the Results section. A full description of each step required in the task was provided in the Methods Section 3: Behavioral procedures, to avoid repetition in the Results. Specifically:

      Self-initiated Go/No-go task variant (“short”)

      Self-initiated Go/No-go/Free task variant (“long”)

      Cue-initiated Go/No-go and Go/No-go/Free task variant

      To improve clarity, we have added a reference in the Results section to the Methods: Behavioral procedures for additional procedural details.

      Voltammetry trace alignment

      The events to which voltammetry traces are aligned were stated in the legend:

      “... when traces were aligned to action-cue onset…”

      “... aligned to the time when animals departed the nose-poke port…”

      “... we realigned traces to the moment animals arrived at the reward magazine… “

      Figure 2c legend: "c) Traces aligned to action-cue onset"

      Figure 2d-e legend: “d) Traces aligned to nose-poke exit and e) magazine arrival….”

      However, to improve clarity, we have now added “aligned to action-cue onset.. “ to make the trace alignment more immediately apparent when results are first presented, and added: “Dopamine data were aligned to events of interest: action-cue onset, nose-poke exit, and magazine arrival.”

      (b) How many subjects were included in each task variant? The text makes it seem like all rats complete each task variant, but the behavioral data suggest otherwise. Moreover, it appears that some rats did more than one version. Was the order counterbalanced? If not, might this influence the DA signal?

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures, where each task variant description includes the corresponding sample size (“Self-initiated Go/No-go task (‘short’; n =9)”, “Self-initiated Go/No-go/Free task (“long”; n = 5)”, “A total of n = 11 and n = 15 were included in the Cue-initiated Go/No-go and Cue-initiated Go/No-go/Free tasks, respectively.”

      For added clarity, we have also added a table for the separation of No-go trials and the number of subjects that it has come from in the Methods.

      Task variant completion

      Not all animals completed all task variants. As stated in Methods Section 4: Real-time dopamine recordings and analysis, animals had to achieve >60% success rate for each trial type on at least two consecutive training sessions to proceed to recording. Other reasons include electrode degradation before all recordings could be completed (See Author response image 1).

      Training order

      We trained four cohorts of animals. One cohort was trained first in the short-task variant

      (Cue-initiated Go/No-go), and the remaining three cohorts were trained first in Cue-initiated Go/No-go (3s) long-task variant (i.e., without Free trials, behavioral and FSCV data not presented in the manuscript). Training order was not fully counterbalanced due to constraints described above.

      Following additional analyses, our data suggest that training order did not influence our core comparison of Go minus No-go dopamine. To directly address whether training order influenced dopamine signals, we separated animals based on whether they were first trained in the Cue-initiated Go/No-go (2s) or the Cue-initiated Go/No-go/Free (3s). We calculated the average dopamine of each rat during the action-cue period and then calculated the difference between them (Author response image 4). We observed absence of evidence of a difference between the groups.

      Author response image 4.

      Average Go minus No-go dopamine reveals no effect of the initial training variant (2s-first vs 3s-first) or test task version (Go/No-go vs. Go/No-go/Free). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of the initial training variant and test task version both showed moderate evidence of absence (initial training variant: BF<sub>10</sub> = 0.378; test task version: BF<sub>10</sub> = 0.367). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.134), and the full model including an initial training variant × test task version interaction was the least supported of all models examined (BF<sub>10</sub> = 0.069). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither the variant animals were first trained on, the task version administered at test, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (5) There is a major challenge in their design and interpretation of the dopamine signal. Both trial types (Go and No Go) start with the rat having their nose in the noseport. An auditory cue is presented for 2-3 s signaling to the rat to either leave the noseport and make a lever response (Go trial) or to stay in the noseport (No Go trial). The timing of these actions and/or decisions is entirely independent, so it is not clear to me how the authors would ever align these traces to the exact decision point for each trial type. They attempt to do this with the nose-port exit analysis, but exiting the noseport for a Go trial (a rat needs to make 2 lever presses and then get a reward) versus a No Go trial (a rat needs to go retrieve the reward) is very different and not comparable.

      We respectfully disagree with the reviewer’s assertion that our data alignment approach is problematic. Aligning neural activity to specific behavioral epochs that occur at different times and across conditions is a widely used method to investigate the relationship between neural activity and behavior. Just to mention some examples: data collected with fiber photometry (e.g., Tan et al. 2026, doi: 10.1038/s41386-026-02368-4; Hart et al. 2024, doi: 10.1016/j.celrep.2024.113828) and voltammetry (Hamid et al. 2016, doi: 10.1038/nn.4173; Syed et al. 2016, doi:10.1038/nn.4187).

      Alignment method

      We intentionally designed the task so that overall action timing is matched between trial types (as described in our response to point 3), while specific behavioral epochs occur at different times. This allows us to compare dopamine dynamics during comparable behavioral events across Go and No-go trials (e.g., nose-poke exit).

      We align data to three critical behavioral epochs, stated in the Methods Section 4: Real-time dopamine recordings and analysis - FSCV measurement and analysis: action-cue onset, nose-poke exit, and magazine arrival. Each alignment addresses a specific aspect of our research question:

      Action-cue onset alignment captures VMS dopamine dynamics when animals have explicit knowledge of trial type and the required action. This allows us to characterize how dopamine evolves following correct action selection, which is central to our research question about how dopamine differs during successful Go versus No-go action execution, as well as no overt action (Free) in the long-task variant.

      Nose-poke exit alignment captures dopamine dynamics at the moment animals initiate movement. The reviewer suggests that exiting for Go versus No-go trials is "very different and not comparable" because subsequent actions differ (two lever presses vs. direct reward retrieval). However, this is precisely our experimental manipulation: we compare dopamine signals when animals exit the nose-poke port to perform different actions. This comparison is both valid and necessary to address our research question (does dopamine encode action initiation?).

      Magazine arrival alignment captures dopamine dynamics at reward approach. This allows us to differentiate between spatial proximity to reward from other concepts including temporal proximity (how soon is reward) and action requirements.

      We acknowledge that we cannot identify the precise moment of decision formation. In fact, the precise moment of decision formation is irrelevant for our question. However, our aim is to characterize VMS dopamine dynamics during successful action execution for rewards, following cues associated with specific actions.

      (6) The voltammetry analysis did not appear to test the hypotheses the authors outlined in the intro. All comparisons were done within task variants (DA dynamics in Go vs. No Go trials, aligned to different task events), but there were no comparisons across task variants to determine if the DA signal differed in cued vs self-initiated trials.

      Our aim was to investigate whether VMS dopamine signals consistently differed between action initiation and action suppression during reward pursuit. To test this within-variant contrast (Go > No-go), we manipulated how reward pursuit is initiated (self- vs cue-initiated) and the “effort” requirements (short vs long task variants), and our results show that they did not affect the differential between Go and No-go.

      The consistent Go > No-go dopamine that we observed across all task variants, together with the consistent increase during magazine approach, supports our conclusion that VMS dopamine integrates motivated action and reward.

      The reviewer suggests that we should have compared dopamine signals across self- vs cue-initiated task variants. We acknowledge this is an interesting, but entirely different, question and have addressed it in the Discussion section. We note that differences in controllability altered the time course of increased VMS dopamine, presumably by triggering earlier positive RPEs in Cue-initiated tasks as compared to Self-initiated tasks (illumination of the nose-poke light being the earliest predictor of reward). However, since our primary research question relates to the difference in VMS dopamine between action initiation and suppression, our results show that this difference was unaffected in two variants (short and long), strengthening our conclusions about the relationship between action and reward-related dopamine signaling.

      (7) Classification of No Go behaviors was interesting, but was not well integrated with the rest of the paper and was underdeveloped. It also raised more questions for me than answers. For example:

      (a) Was the behavior classification consistent across rats for all No Go trials? If not, did the DA signal change within subjects between biting vs digging vs calm?

      (b) If "biting rats" were not always biting rats on every No Go trial, then is it fair to collapse animals into a single measure (Figure 3C).

      (c) Some of the classification groups only had 2 or fewer rats in them, making any statistical comparison and inference difficult.

      Behavioral classification for each rat was consistent across trials (i.e., 100%, see Author response image 5). Upon reviewing the consistency of classifications within individual animals, we found that “Biting” animals exhibited biting behavior across the majority of their No-go trials. Only one animal (in the Self-initiated Go/No-go "short" variant) showed mixed classifications across trials, occasionally exhibiting digging or calm behavior. For all other animals, the predominant behavioral classification was highly consistent within subjects across sessions. The occasional trials where “biting” animals did not bite were too infrequent to permit meaningful within-animal comparisons. Therefore, we believe collapsing animals by their predominant behavioral phenotype in Figure 3C is appropriate and accurately represents stable individual differences in No-go response strategies.

      As stated in the Methods Section 6: Statistical Analysis - Clustering No-go behaviors and regrouping animals (last-line), we specifically avoided between-subjects statistical comparisons for groups with n≤2, as this would be inappropriate (see Author response image 5), and reported qualitative observations only. These exploratory findings at the individual-trial level suggest behavioral heterogeneity during No-go trials, that others may use for future investigation, but do not form primary conclusions.

      The behavioral classification is integrated with our central findings on VMS dopamine encoding spatial proximity. Our results demonstrate that individual variation in action suppression strategy, in particular Biting behaviors, consistently manipulates the timing of max dopamine release during subsequent reward approach, but not during the action itself (Figure 3b-c). This links our observations of action-dependent dopamine (Figure 2c) with spatial reward approach (Figure 2e). We believe that our findings shed new light into the understanding of how action modulates reward-related dopamine dynamics at the individual level. This has been discussed in Discussion section: Dopamine dynamics are linked to motivated action.

      Author response image 5.

      Rats predominantly stick to a particular strategy to perform No-go trials. Each bar represents an individual animal, and colours represent the % of each classification type.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figures: It would be helpful to have more panel labels on the figures - 1C, for example, labels 4 different dataset panels, similar to a few other cases. This would really help to improve readability, as there are a lot of task types, trial types, behavioral measures, and labels to sift through. The figures overall are very busy, and it is a bit of a challenge to process the tasks and trial comparisons, as well as interpret what the quantification insets mean.

      We thank the reviewer for their suggestion. We have added more panel labels to Figure 1 to improve the ability to follow with the main text. We hope that the reviewer finds it acceptable.

      (2) Figures: A bit more specifically, in Figure 2, the boxplot insets are pretty hard to see, and it's not clear what scale they are on or what data they reflect. Similarly, it's unclear what the horizontal bars reflect in terms of which conditions are being compared. Why are box plots used for some comparisons, and why are some comparisons based on time series bootstrapping, but others are not clear? I would consider broadly reworking this figure and its description for clarity. Figure 3, by comparison, is easier to understand - the quantifications are clearer and labeled.

      We have now labeled the comparisons being depicted by the horizontal bars in Fig 2c. We have also clarified the boxplot analysis in Fig 2d and Fig 2e by adding 'Max dopamine’ labels to the figure, and have made these analysis methods more explicit in the figure legend.

      We used time-series bootstrap analysis to identify when dopamine signals diverged between trial types, which depicts dopamine differences during distinct action requirements. The latency-to-max quantification provides a summary measure to test specific encoding hypotheses (temporal vs. spatial proximity to reward).

      (3) Broadly, more clarity on the FSCV analysis is warranted.

      (a) Targeting: It looks like the dataset contains a mix of medial shell and mostly core accumbens placements. The paper treats VMS as a uniform dopamine region, but it is more standard to separate core and shell (and also other parts of the shell) into subregions. Many of the reported encoding profiles here are known to differ across the accumbens. So, some consideration of this seems appropriate - a minimal signal-behavior comparison for the shell vs core subgroups, for example.

      We thank the reviewer for raising this point. Upon careful re-examination of our histological analysis, we identified an error that occurred when we made the overlay of electrode placements across rostrocaudal planes (to project placements onto a single plane for the sake of simplicity; Fig 2a): We incorrectly assigned some recordings to the nucleus accumbens shell. We corrected this error, which shows that the vast majority of recordings were in the core, with only 2 animals in the shell (black stars). We have adjusted Fig 2a to reflect this, and added an anatomically more complete illustration of electrode placements across rostrocaudal planes as Supplementary Figure S8.

      While we acknowledge reported differences between core and shell dopamine in some contexts, the small number of shell placements precludes meaningful statistical comparison. However, to determine whether average dopamine concentrations differed between core and shell during the action-cue period, we have plotted the values in Author response image 6. The fact that shell data mostly falls centrally into the overall core-data distribution suggests no consistent difference in dopamine release.

      Furthermore, in our experience (and that of colleagues (personal communication)) with appetitive operant tasks, core and shell FSCV dopamine signals do not substantially differ for action-selective encoding and reward approach. Given the sample distribution and our focus on general VMS function in Go/No-go behavior, we believe pooling these regions is appropriate.

      Author response image 6.

      Average dopamine release during the action cue in nucleus accumbens core (circles) and shell (stars) animals showed no distinct separation between regions. Each symbol represents an animal.

      (b) Design: In my understanding of the design, the main distinction between the short and long task variants is a 2-second versus a 3-second required nose poke hold. 3 seconds here is "long" and more "difficult". I'm not sure I agree that a 1-sec distinction really reflects a difference in task difficulty or effort. Can the authors point to a past paper that demonstrates this variation is sufficient to engage a behavioral difference and/or a neural encoding difference? Broadly, some justification of the validity of this manipulation is needed, I think.

      While we lack direct citations for this specific manipulation, we believe that the 2s and 3s hold requirements represent meaningful differences in difficulty based on our extensive rat-behavior experience and behavioral evidence in Author response image 2.

      Importantly, the difficulty of No-go trials does not stem merely from the required time to hold their snouts in the nose-poke port, but from suppressing the motivational/Pavlovian bias to approach reward-associated cues. In No-go trials, subjects must suppress the prepotent tendency to immediately approach reward-related stimuli and instead maintain active suppression of this approach behavior. Even the 2s hold is challenging as animals tend to perform better on Go trials compared to No-go trials (Fig 1b and 1c), demonstrating the inherent difficulty of response suppression even at the shorter duration. The additional 1-second substantially increases this demand, as it represents a 50% increase in hold duration.

      Our training data clearly demonstrate the difficulty in reaching task criterion when increasing the required action (for Go and No-go trials) from 2s to 3s. Across four cohorts of animals trained in Go/No-go task variants, one cohort that was trained first in the 2s task variant, and the remaining three cohorts were trained in the 3s task variant of Go vs No-go. Animals required 41.1 ± 7.9 (n = 27) sessions to learn the 3s hold (approximately 8 weeks), versus

      18.5 ± 7.6 sessions for the 2s hold (approximately 4 weeks; mean ± SEM). This indicates that despite only a 1-second difference in required action performance, animals needed more than double the number of training days to reach criterion, clearly indicating differential effort demands.

      (c) Figure 2 results: the authors state that because there is a greater DA signal to Go vs No Go cues in all the task variants, this means that controllability of reward pursuit and increased task effort do not affect VMS dopamine. But the magnitude of the signals looks different across the task variants - it looks clearly stronger overall in the self-initiated tasks, for example. Given that dopamine signals are not compared across task variants (I think the tasks are all between-subjects?), I don't think the above conclusion is justified.

      We respectfully clarify that our conclusion does not claim controllability and effort have no effect on dopamine magnitude, but rather that these manipulations do not affect the action-selective difference in dopamine (Go > No-go). Our central finding is that the relative difference between Go and No-go remains consistent across all task variants (within-subjects comparison).

      We did not perform across-variant comparisons of absolute dopamine magnitudes because that was not our primary research question. Our focus was to understand whether dopamine differs between action initiation and suppression, and whether this difference can be modulated by controllability or effort.

      We acknowledge the reviewer’s observation that absolute magnitudes appear larger in self-initiated vs cue-initiated task variants. We believe that this likely reflects differences in RPE timing rather than controllability per se: in cue-initiated tasks, the nose-poke light provides an early trial-start signal, distributing RPE temporally across the trial. In self-initiated tasks, trial-initiation and action requirements are temporally integrated. Though understanding how controllability affects absolute dopamine magnitude is an interesting question for future work (e.g., using sophisticated regression-based encoding models), it was beyond the scope of our current investigation, which focuses on action-selective encoding.

      (4) Broadly, I don't think these data, as shown, support the conclusion "dopamine reflects action initiation but not controllability or effort" without more analysis and additional context.

      We have revised the wording of our conclusions throughout the manuscript to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). It is now “Reward-related dopamine depends on action initiation irrespective of controllability and effort”.

      (a) Figure 3 - more description of the classified behaviors would be helpful for interpreting this part of the data. When are the behaviors occurring - during the hold cue? Or is the classification related to what they do immediately after holding? Or something in between> I guess I'm not sure what digging and biting are in the context of a nose poke hold. As described, it's not clear what the signal differences relate to - movement differences? Generally, it's not clear what to make of the behaviors. They seem to emerge spontaneously, but it's not clear whether the specific actions mean anything, so it's a bit difficult to know what to glean from the dopamine is greater during "biting". It's a very different movement pattern, so perhaps this result relates to that, rather than task engagement or motivational drive per se?

      We thank the reviewer for the comment. We have added relevant information in the figure caption and in the Results section to clarify that classified behaviors occurred during the action-cue period (for No-go trials, the hold cue; Figure 3a caption). In addition, we have included Videos to better depict the classified No-go behaviors during the action-cue period.

      We agree with the comment that these classified behaviors, such as biting, seem to emerge spontaneously. Our interpretation of these behaviors is that they may represent the motivational state of each subject. Most importantly, whereas the behavioral differences occurred during the action-cue period (while animals had to suppress actions and stay within the nose-poke port), the difference in VMS dopamine was only observable after this behavior was completed. Thus, the movement pattern per se is likely not relevant to the dopamine release occurring after its completion. This has been discussed in the Discussion section: Dopamine dynamics are linked to motivation action.

      Based on our videos, it appears as though Digging could be perceived as more vigorous (i.e., more general movement in the nose-poke port). However, we did not observe more dopamine during the action-cue period of Digging trials as compared to Biting trials. Furthermore, more vigor during the action-cue period (e.g. Digging trials) did not result in more dopamine during the reward approach period. Together, the data suggest that another process may underlie the large increase in VMS dopamine in Biting trials during reward approach, such as varying attribution of incentive salience.

      (b) In some cases, but not all, dopamine measurement comparisons are done on a total trial basis, and in others, it seems to be subject averages. It's not clear why different approaches are used for different parts of the data. But also, for the trialwise analysis, what statistical steps were taken to incorporate the subject as a random factor in the analysis? If that is not done, then a trial-wise analysis artificially increases the power for the stat (n=trial#).

      We used different analytical approaches depending on sample size and data structure. To compute differences in Go vs No-go dopamine within each animal, as intended by our experimental design, we performed subject-level comparisons (Figure 2).

      For the behavioral classification analysis (No-go, Figure 3), we performed trial-level analyses to increase the statistical power and better characterize this unexpected and interesting phenomenon. We explicitly chose not to perform subject-level group comparisons because

      (1) some groups had only n=2-3 animals, making subject-level statistics underpowered, and (2) behavioral classifications were highly stable within individual animals (see Author response image 5). We acknowledge that formal between-group comparisons (across subjects) are underpowered due to small n, but the stability of within-subject No-go behavioral strategy and qualitatively distinct VMS dopamine profile suggest that these differences may be biologically meaningful and worthy of future investigation in larger samples. We have made these limitations more explicit in the Results.

      This relates to Figure 3, where all trial data are shown next to individual subjects - the subject-wise group comparisons are between 2-5 or so rats, which is quite low. In Figure 2, a subject n of 27 is listed, so it's not clear why this analysis is on such a small set of rats. Generally, it's not clear how many rats/subjects are in each data bit. The methods say only 5 rats are in the long self-initiated task, but 15 in the cue-initiated task. Clarity in all this is needed, including in the figures/captions.

      We have now added detail about the statistical test performed in Figure 3’s caption:

      “After action-cue offset: No-go (trials)’: Individual No-go trials classified by No-go behavior, and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” We have also added in a table in Methods (Table 1), showing the number of trials in each No-go classification, and animals regrouped based on their predominant No-go strategy (see Methods section: Statistical analysis - Clustering No-go behaviors and regrouping animals).

      For the small sample sizes based on the regrouping of animals based on their predominant No-go strategy, we have now added in the caption, “... Rats classified based on their predominant No-go strategy, with no statistical tests performed.”

      (5) I'm also a little confused about the paper's narrative that the dopamine data reflect spatial (but not temporal) proximity to reward - it seems that this conclusion is based on the dopamine signal peaking at magazine entry, but that is different, I think, than a spatial signal per se (space is not manipulated in this study). I think more analysis of the signals during the magazine approach behaviors would be helpful, and possibly comparing rewarded vs unrewarded approaches. The emphasis, including in the title, that a major take-home of the data is that dopamine encodes reward proximity, is not really borne out by the current analyses. Reward expectation is not manipulated independently of the approach action, so it's hard to pin this on "space" vs "reward is soon". This is admittedly a general complexity in characterizing dopamine ramps.

      As the reviewer notes, we acknowledge that 'spatial proximity' and 'reward is soon' are challenging to fully dissociate in appetitive approach paradigms. However, we believe that our data and new additional analyses, which is now included in Results: Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards and Supplementary Figures S10-12, provide compelling evidence that VMS dopamine primarily reflects spatial proximity to the expected reward location, rather than temporal proximity to reward delivery.

      Evidence against full temporal encoding:

      (1) Max VMS dopamine occurred up to 3s before reward delivery in Free trials (Fig 2e, open circles vs triangles). Furthermore, when we calculated the max values of individual trials of realigned traces, maximum dopamine does not consistently coincide with reward delivery across trial types (new Supplementary Figure S11).

      (2) If dopamine encoded temporal proximity from the earliest reward-predictive cue, we would expect consistent accumulation from cue onset (action cue for self-initiated, nose-poke light for cue-initiated). However, realigned trials also did not show a consistent accumulation around these events (new Supplementary Figure S12).

      Evidence for spatial encoding:

      (1) Max dopamine consistently occurred when animals arrived at the magazine across all trial types and task variants, regardless of when reward was actually delivered (Fig. 2e).

      (2) Across individual trials, max dopamine values concentrated around the magazine-panel, with a striking accumulation when animals were in close proximity to it (new Supplementary Figure S10) showing a distance-dependent distribution.

      (3) Assessment of unrewarded magazine approaches during the intertrial interval (ITI) revealed no increase in dopamine release (Fig 2f), indicating that VMS dopamine requires task-relevant reward expectation.

      (6) Examples of the DLC workflow in a supplement would be appropriate. Also, video examples of the 3 kinds of behaviors from the clustering analysis could be useful for understanding what they are/what they mean.

      We have added supplementary figures showing the DeepLabCut workflow (Supplementary Figure S9) and Videos 1-3 demonstrating the three behavioral classifications (biting, digging, calm) during No-go trials.

      (7) Referencing/scholarship: I would suggest broadening the citation pool for the paper to include more older work that has established the notion that dopamine signaling and the accumbens act as a motivation-action interface, as this has been a longstanding notion since at least the 1980s. There is also a sizable literature on dopamine signaling of effort, some of which would be appropriate to cite here.

      We thank the reviewer for the recommendation and have now expanded our citations to include foundational literature to work from the 1980s-90s. These can be found in the Introduction, Results, and Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 353-357- This came as a surprise, as the relevant results are only featured in supplementary information. These should be moved to the main manuscript. As both the biting behavior and faster lever-press completion lead to larger peak dopamine, does this represent response vigor?

      We appreciate the reviewer’s interest in these data. However, we believe that the data that we report in “Supplementary Fig 7: Quartile analysis of Go trials shows coordinated changes in last lever-press timing and VMS dopamine, does not warrant movement to the main manuscript.”

      The purpose of the lever-press timing analysis was to demonstrate that maximum dopamine release coincides with the moment animals arrived at the magazine, rather than with the action period itself. By sorting Go trials based on last lever-press latency, we show temporal coordination between action completion and dopamine timing but critically, dopamine peaks after action completion, not during it.

      This temporal dissociation argues against a 'response vigour' interpretation. If dopamine encoded motor vigour, we would expect the signal to coincide with or precede the vigorous action. Instead, both the lever-press data (Supplementary Figure S7) and the classified No-go trial data (Figure 3, biting behavior) show that changes in max dopamine occur after the actions themselves, during the subsequent approach to reward.

      Together, these findings demonstrate that VMS dopamine reflects spatial approach to the reward location following action completion, not the vigour of the required actions per se. The lever-press analysis serves as supporting evidence for this temporal relationship but does not introduce a novel finding that warrants main figure emphasis. We mentioned this temporal coordination in Results to ensure readers are aware of the converging evidence while maintaining focus on our central findings regarding action initiation versus suppression.

      (2) Line 13- the experimental work cited refers to midbrain dopamine neurons, rather than dopamine release within the VMS. Please correct.

      We thank the reviewer for pointing out this error. We have now corrected the citation to refer to studies measuring striatal dopamine rather than midbrain dopamine neurons.

      (3) Line 166 - Shouldn't this say "consistently delayed for No Go trials"? Looks like peak dopamine occurs later for these trial types.

      This statement refers to traces aligned to nose-poke exit (not action-cue onset). The peak Go dopamine occurred later than other trial types following exit (Green arrows in Figure 2d).

      (4) Lines 391-392 - however however

      Thank you for the comment, we have adjusted the text.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In their important manuscript, Gangadharan, Kober and Rice focus on how Stu2/XMAP215-family microtubule polymerases use their TOG domains to catalytically promote microtubule growth, testing whether their mechanism follows an enzyme-like kinetic model similar to that of actin polymerases. The authors integrate measurements including microtubule polymerization rates and TOG-tubulin binding kinetics to convincingly show that Stu2 follows an enzyme-like model where tight tubulin binding enables efficient polymerization, revealing a shared mechanism with actin polymerases despite their evolutionary divergence. This work will be of general interest to the cell biology and biophysics communities.

      Thank you for the favorable assessment of our manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Gangadharan and colleagues provides significant progress towards a quantitative biochemical mechanism for Stu2 polymerase activity. A key conceptual advance is the novel application of an enzyme-like model, initially developed for the actin polymerase Ena/VASP, to Stu2.

      New refined affinity measurements for a Stu2 TOG domain using Bio-layer interferometry show more than an order of magnitude higher affinity of TOG domains to tubulin compared to previously published reports.

      The findings reinforce the "concentrating reactants" or, more specifically, for TOGdomain proteins, the "tubulin-shuttling antenna" model, compared to the "polarized unfurling" model, a more speculative structural hypothesis.

      The manuscript builds upon a series of previous manuscripts that showcase the profound intellectual engagement with microtubule polymerization mechanisms by TOG-domain proteins from the Rice lab, a thought leader in microtubule polymerization for over a decade.

      Minor remarks:

      (1) A major new experimental finding of this paper is the affinity of TOG domains, which is more than an order of magnitude lower (10 nM) than previous measurements from the same lab (~200 nM). The authors attribute this change to ionic strength differences between buffer conditions, citing the lab's previous work (Ayaz et al., 2014). This argument left me contemplating what the buffer conditions are in both experiments, and I wonder if other readers would feel the same. After going down the rabbit hole, I believe the difference in ionic strength is ~2.3 fold, and at least on the back of my envelope, this works out beautifully with the measured differences in affinities. A short version of this argument may strengthen the manuscript.

      This is a good comment. We should have been clearer about the different buffer conditions. The revised manuscript now explicitly states how the two buffers in question differ in pH and ionic strength. (Page 8, ‘Tubulin binds rapidly …’ section). We tried to perform comparative measurements of TOG:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is stated in the revised manuscript (Page 9, final paragraph before the ‘Unifying measurements …’ section. Along the lines of the reviewer’s ionic strength calculation, and consistent with the increase in affinity we observed with lower ionic strength, we now also state that prior measurements from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with higher ionic strength (100 mM KCl vs 200 mM KCl).

      (2) I am wondering if there may be an alternative explanation to tubulin binding by TOG being the kinetically rate-limiting step for polymerase function:

      TOG + Tubulin ⇌ TOG:Tubulin (fast binding rate, high-affinity binding)

      TOG:Tubulin + MT_end → TOG:MT (tubulin is incorporated into MT, fast transfer rate)

      The binding rate is 3/s, and the transfer rate is 5/s.

      I was wondering if the following step should be considered, which involves a conformational change of tubulin (e.g., straightening) TOG:MT → TOG + MT (ratelimiting straightening and unbinding of TOG from the lattice).

      This is an interesting thought that highlights a gap in the understanding of microtubule dynamics.

      Presumably, the affinity of TOGs for straight tubulin is practically zero for the purpose of this discussion, as there is no lattice binding, which means unbinding is likely very rapid; however, straightening may be the rate-limiting factor here.

      In theory, straightening should also be rapid; however, we lack measurements of how fast or slow this step occurs within the context of a TOG domain, which presumably skews the process towards curved tubulin.

      We agree (based on prior observations) that the affinity of these TOGs for straight tubulin is negligeable in this context. There is much less data about the timescale of tubulin straightening, with or without a TOG domain bound, or even about how tubulin interactions with the microtubule end affect the balance of preferred conformations and/or the rate of conformational change. It’s an extremely interesting topic. Because the straightening process the reviewer envisions is zero-order, the transfer rate in our model could in principle reflect slow straightening (in this view the ‘delivery’ step would need to be very fast, i.e. not rate-limiting). Because there is so little data about this, and because there are not yet methods to study or perturb the timescale of straightening on the microtubule, we prefer not to engage too deeply. We added a sentence to acknowledge this alternative possibility in the revised manuscript (bottom of Page 4 and top of Page 5).

      A hypothetical Stu2, when bound to the microtubule end and with the TOG domain not disengaged from tubulin, would not permit the processivity of that molecule or the binding of a new molecule.

      To emphasize the importance of unbinding, when it is not efficient, as reported for the T238 mutant that results in Stu2 lattice binding (Geyer et al., 2018), the polymerase becomes inefficient.

      The mechanism of polymerase processivity has not been conclusively determined (the Geyer et al. 2018 eLife paper took a step in that direction, though). The model used in this paper is only concerned with how many polymerases are at the microtubule end at steady-state (as opposed to how long a particular polymerase acts before dissociating), so while we appreciate and are interested in these questions, we think it would be better to leave them for future work.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from the Rice lab by Gangadharan et al. investigates the polymerization mechanism of the yeast microtubule polymerase Stu2. The lab has published a number of articles demonstrating the structural basis by which the two TOG domains of Stu2 each bind free tubulin heterodimers, and has developed a tethered polymerization model by which the TOG domains drive polymerization by shuttling those tubulin subunits onto the microtubule plus end. A second model was proposed by Nithianantham et al. (eLife, 2018) based on a closed-to-open transitional state in which Stu2 unfurls and loads two longitudinally associated tubulin heterodimers onto the microtubule plus end. While the second model is not directly tested, the current work aims to further characterize/model the tethered polymerization model using a kinetic framework developed by Breitsprecher et al. for Ena/VASP actin polymerization activity, using a model that is enzymatic (EMBO J., 2011). The general architecture and function of Ena/VASP on actin polymerization versus Stu2 on microtubule polymerization is a reasonable relation and hits upon, as the authors note, potential convergent mechanistic evolution across distinct cytoskeletal networks. The model effectively treats tubulin as the substrate, and the polymerized microtubule plus end as the product. If Stu2 is "enzymatic" in this framework, the model predicts it would behave with Michaelis-Menten kinetics, that there would a Vmax, and polymerase activity would either be "affinity limited" by TOG:tubulin affinity (KD) and/or "kinetically limited" by TOG:tubulin association (Kon) and transfer of tubulin to the microtubule plus end (Kt). The authors find that the Brietsprecher model works well for Stu2 activity, and that Stu2 best aligns with a "kinetically limited" model. The work is interesting and adds to the growing elucidation of the Stu2 microtubule polymerase model. While yeast microtubule polymerases are somewhat distinct in their architecture, there is significant overlap that findings from the manuscript can be utilized to inform the mechanisms of larger, more complex microtubule polymerases such as human ch-TOG.

      Thank you for the nice summary and favorable comments.

      Strengths:

      The manuscript invokes the enzymatic model of Breitsprecher et al. used for Ena/VASP and conducts an elegant series of (mostly established) experiments to determine whether Stu2 microtubule polymerase activity aligns with the model, which they conclude does align, supported by the data/results obtained.

      Weaknesses:

      The authors used biolayer interferometry to measure TOG:tubulin affinity. The affinities obtained were significantly higher than the lab obtained in an earlier publication using analytical ultracentrifugation. While differences in buffer and salt conditions may underlie these differences, additional runs using comparable buffer systems, or the use of a third independent assay to measure affinities, would have added rigor.

      This is a good question that was also raised by reviewer #1. We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is now stated in the revised manuscript (page 9, final paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been independently shown to depend on ionic strength in a way that seems consistent with what we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, final paragraph before the ‘Unifying measurements …’ section).

      The discussion could be expanded to better compare and contrast the results with both existing polymerase models introduced in the introduction, as well as expanded to look at reversible enzymatic activity (microtubule depolymerization at low to zero tubulin concentrations) and microtubule plus versus minus end activity.

      Thank you for the push to be more explicit about the two contrasting models. We made small changes to the introduction (top paragraph on page 3) and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (page 12, penultimate paragraph of the main text).

      The ‘transfer’ reaction is treated as irreversible (analogous to catalysis by an enzyme), so the biochemical model we use for the polymerase cannot account for polymeraseinduced microtubule depolymerization at low to zero tubulin concentration. We added text to state that the model is limited to the growth reaction (page 4, last paragraph) but otherwise prefer to not engage too deeply in questions about the reverse reaction.

      These polymerases are thought to be plus-end specific because of the domain organization of the protein: TOGs bind tubulin such that the N- to C-terminal polarity of the TOG corresponds to the plus- to minus-end polarity of the tubulin, and the basic region used to make a ‘slippery’ connection to the microtubule is located C-terminal to the TOGs. These two factors mean that it is only at the plus-end that TOGs can engage αβ-tubulins with the basic region contacting surfaces ‘deeper’ in the polymer. We added text about these issues, citing prior work, to the legend of Figure 5 (page 11). We chose to not elaborate much since it is not the primary focus of the paper.

      Reviewer #3 (Public review):

      Summary:

      This study by Gangadharan and colleagues seeks to establish a quantitative biochemical model for the microtubule polymerase activity of Stu2. Stu2 is the budding yeast member of the XMAP215 protein family, which is broadly conserved across eukaryotes. XMAP215 proteins play a wide variety of important roles in cells, and these are attributed to effects on microtubule dynamics. Many studies over the last ~20 years have shown that XMA215 proteins selectively associate with microtubule ends, where they increase rates of microtubule assembly and disassembly. More recently, structural biology and biochemical studies by the authors and other groups have shown that the multiple TOG domains on XMAP215 proteins are tubulin-binding domains that selectively bind to curved tubulin, which is present in solution and at microtubule ends, but not to straight tubulin which is present in the walls of the microtubule lattice. This has led to the general model that XMAP215 proteins promote polymerization by delivering soluble tubulin to the growing plus end, and two distinct models have been proposed to explain the mechanism. The 'concentrating reactants' model proposed previously by the authors suggests that TOG domains grab hold of tubulin in solution and concentrate at the microtubule end. The 'polarized unfurling' model proposed by the Al Bassam lab suggests that XMAP215 delivers multiple tubulins to the end, using a step-wise mechanism involving different roles for each TOG domain. The current study seeks to improve our understanding of the mechanism by developing a quantitative model to explain the binding and release of tubulins, the number of Stu2 molecules at the end, and the overall rate of tubulin addition. The authors accomplish this goal using new experimental data. The final model fills in new details of the mechanism. The authors draw a comparison between Stu2 and the actin polymerase, which bears similarity to Ena/VASP, and suggest a convergent strategy for cytoskeletal polymerases.

      Thank you for the good summary and favorable comments.

      Strengths:

      This is a focused and clearly written study that incorporates prior knowledge of XMAP215 and draws inspiration from the actin field. The data are clear and convincing, and the study accomplishes its goal of generating a new, quantitative model for Stu2. The model will be important for microtubule researchers to predict and test key points for altering XMAP215 activity across different organisms and potentially for different tubulin substrates. The comparison to Ena/VASP may also inspire similar comparisons across other microtubule and actin regulators, which could lead to new insights across the cytoskeletal fields.

      Thank you for these comments.

      Weaknesses:

      The study is without major weaknesses, but there are several minor weaknesses worth noting. One is that the final model provides new details regarding the Stu2 mechanism, but does not provide a major new advance in our understanding of how the polymerase works. For example, the discussion does not clearly argue for whether the new results and model rule out either of the prior models. This appears consistent with the 'concentrating reactants' model, but does it clearly rule out the 'polarized unfurling' model?

      Thank you for pointing out what in retrospect was an obvious ‘loose end’ in our discussion. The other referees raised the same point. We made small changes to the introduction and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (top paragraph of page 3 and new penultimate paragraph of the manuscript on page 12).

      A second minor weakness is that the comparison to Ena/VASP is not developed at a deep level based on the final model. I found these ideas exciting and want more critical consideration here, but perhaps it is better suited for a commentary piece to follow.

      We appreciate the enthusiasm and understand where this comment is coming from. Because there has been a fair amount of recent movement in the understanding of TOG domains and what they can do, and because some of the mechanistically interesting parallels entail speculation, we agree with the referee’s suggestion that a future commentary will provide a better venue.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Other minor remarks:

      (1) Figure 2 C is missing the label for what should probably be TOG1*-TOG2.

      Fixed

      (2) Figure 5, lower left, is oddly cropped, showing residuals that are slightly distracting from the beauty of the model.

      Apologies that the figure did not look good in the initial submission. We adjusted it and it looks much better in the revised submission.

      (3) The dynamics assay buffer composition stated in the protein purification Method section is not the same as the PEM buffer used for the dynamics assay. And both are different from BRB80, which, with the chambers, are rinsed. This may very well be accurate, but it raises the question of why not stick to one version, as they are virtually the same.

      Thanks for asking these questions, and sorry for the confusion. First, we should have used different names for the (barely) different buffers. This has now been corrected. Second, the PIPES concentration was not 90 mM, it was 100 mM as in our prior work and this discrepancy failed to get caught in proofreading. Why the other small differences? It’s a good question. The differences reflect an arbitrary decision made at the beginning of the work, there is not a deeper rationale.

      (4) Out of curiosity: Why 90 mM PIPES and not 80?

      Why not 80 mM PIPES? This is just a historical difference. The early measurements of yeast microtubule dynamics (e.g. Gupta … Himes MBoC 2002 and Bode … Himes, EMBO Rep 2003) that partly inspired us to use yeast as a model system used 100 mM PIPES as the working concentration, and we never deviated from that.

      (5) Please state the source of PIPES.

      Sorry for the oversight, we have added the source of PIPES (it is Millipore Sigma P6757).

      Reviewer #2 (Recommendations for the authors):

      (1) Page 4, last paragraph, the authors call out "Fig 1C" which I believe should be "Fig 1D".

      We fixed this, thank for catching this error

      (2) Figure 2C: The authors subtract basal tubulin polymerization (growth rate) from the rates measured in the presence of Stu2 constructs. One assumption in doing this is that Stu2 polymerization activity does not compete for the ability of tubulin (not bound to Stu2) to polymerize on the plus end. I think this is a logical assumption, but it would be beneficial for the authors to state this assumption.

      We said this explicitly in the revised submission (first full paragraph on page 7), borrowing from the reviewer’s phrasing.

      (3) Figure 2C: the label for the last bar is missing - presumably: " Stu2 (TOG1*-TOG2)".

      Fixed.

      (4) Figure 2D: Many of the KM values determined are at the border for points measured, or in one case, beyond the concentration of tubulin sampled. In this regard, the authors should discuss how well the fitted curves correlated with their data. Also, as Vmax and KM are calculated, it would be beneficial if another panel were produced (e.g., Figure 2E) in which the data were presented as a Lineweaver-Burk plot. Doing so, the authors would be able to test their enzymatic model by doping the system with their nonpolymerizable tubulin mutants, which should yield competitive inhibitor behavior but not change Vmax.

      Thanks for pointing this out, we should have been more explicit about this point. We incorporated into the results section an explicit statement about this limitation (first full paragraph on page 7). The suggestion to use blocked mutants and Lineweaver-Burke plots is an interesting one that we hope to pursue in future work using blocked or other mechanism-specific mutants. But we think to do so is complicated enough to be beyond the scope of the present study.

      (5) Figure 3: The authors quantitate the amount of Stu2-GFP fluorescence at microtubule plus ends using line scan analysis of the kymograph. Since the kymographs are processed images, it is more appropriate to integrate intensity from the original frames collected using a circular area. i.e., ID points on the kymograph, and return to the respective position in the corresponding frame to calculate background-subtracted GFP intensity at the plus end.

      This is a fair point. We chose to stick with the kymographbased analysis because the symmetry of the point-spread function and lack of rapid variation in GFP and/or background intensity means that the kymograph analysis is adequate for the intensity-based comparison we were doing.

      (6) Page 7, last line, the authors call out "Fig 2D" which I believe should be "Fig 1D".

      Sorry for the error, we have corrected it.

      (7) Figure 4A: The authors discuss the "sortase epitope," but technically, an epitope is the binding site specifically for an antibody, not to be used in general terms for proteinprotein interaction sites. As such, the authors should describe this as the "sortase recognition sequence" or something similar.

      Thank you for noticing this; we had indeed used ‘sortase recognition sequence’ elsewhere in the paper but we did not catch this instance during proofreading. We have now used that same language in the legend for Fig. 4.

      (8) Page 8, the authors state "KM is approximately equal to Kt/Kon and KM negligeable," but I think they mean "...and KD negligeable".

      Thank you for noticing this typo. We corrected it.

      (9) Page 9, first paragraph last sentence: the readership would be aided by modifying the sentence as follows (adding "Kon" and "Kt"): " ... must be kinetically limited by either the rate of TOG:tubulin binding (Kon) or by the rate of TOG-mediated transfer to tubulin to the microtubule end (kt)."

      Very good suggestion, we implemented it.

      (10) Page 9, second paragraph, the authors call out "Fig 2D", but perhaps they intended to call out "Fig 1D"?

      Sorry for the error, the reviewer is correct and we fixed this.

      (11) Page 9, second paragraph: The authors mention that the transfer rate of tubulin to the plus end via a TOG domain is close to the transfer rate of free tubulin to the growing plus end. Can the authors expand on why they are mentioning this comparison?

      Thanks for the push to be clearer about this. The basic idea is that each ‘delivering’ TOG contributes 50% of the background (uncatalyzed) polymerization rate. So the presence of multiple TOGs (in a single polymerase or from multiple end-resident polymerases) can substantially increase the rate of polymerization. We added brief text to try to make this clearer (first paragraph on page 10).

      (12) Page 9, second paragraph: "TOG-TOG2 polymerases" would be better phrased as "TOG2-TOG2 dimeric polymerases". Noting as well that "2" is missing from the first "TOG".

      This is indeed better phrasing and we have adopted it (also corrected the missing ‘2’) (first paragraph on page 10).

      (13) Page 13, BLI methods: The authors should list the final pH for the PIPES buffer (was it pH 6.9 as in the polymerization assay?).

      Sorry for the oversight, we have added the pH and it was indeed 6.9.

      (14) Page 13, BLI methods: What is "LR1-457"?

      LR1-457 is lab-notebook-speak that did not get purged in editing; it refers to the polymerization-blocked tubulin mutant that also carried a sortase recognition sequence. We replaced ‘LR1-457’ with more evocative phrasing and corrected another typo we found there.

      (15) In Ayaz et al. (eLife, 2014) Stu2 TOG1 and TOG2 affinities for tubulin were measured using AUC, for which the fitted curves appeared to correlate with the data quite well. As the authors note, the values were KD = 70 nM and 160 nM, respectively. This contrasts with the BLI measurement for TOG2-tubulin (~10 nM), which suggests that at least one of the experiments was off the mark - or, as the authors do note, that different buffer and salt condition was used could account for the differences, but that the BLI conditions align with the polymerization conditions (though not exactly) and thus are more appropriate to use. In a supplemental discussion, the authors should run the AUC values through the equation for their model and state what types of differences these values could imply for Stu2 mechanism. If the differences are significant for the Stu2 model derived, the authors should give thought as to whether a third assay should be employed to determine TOG-tubulin affinity. Based on the BLI reagents, it appears the authors would be well-positioned to conduct an assay using SPR. As a potential alternative, the authors could repeat the BLI experiment using the buffer conditions from the Ayaz et al., AUC work (25 mM Tris pH 7.5, 1 mM MgCl2, 1 mM EGTA, 100 mM NaCl, 20 μM GTP) - noting that BSA and Triton X-100 may need to be added as well. If the authors are able to replicate the ~160 nM affinity for TOG2:tubulin, this would be a reasonable way to bootstrap to the conclusion that the BLI is measuring affinity correctly and that the current PIPES-based BLI experiments yielded accurate data.

      We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems. Unexpected challenges and personnel turnover made this slower than anticipated. The reviewer’s suggestions are completely reasonable, but ultimately it was not possible for us to get side-by-side results for TOG:tubulin affinity using the same measurement technique, whether it was biolayer interferometry, analytical ultracentrifugation, or isothermal titration calorimetry. For each technique one or the other buffer caused aggregation, non-specific binding, or some other artifact that prevented analysis. The fact that we were unable to compare the buffer conditions in this way is stated in the revised manuscript (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been shown to depend on ionic strength in a way that is consistent with the changes we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added some text to address the comment about affinity and whether/when the shuttle model would hold (first paragraph on page 11).

      (16) A sentence or two in the discussion, relating how their data aligns (or not) with the Nithiantham model would be beneficial, especially as discussing the two models in the introduction was a central point.

      We completely agree and have now added a paragraph to the discussion to explicitly address the two models and how are or are not supported by the new model and observations (page 12, penultimate paragraph of the main text).

      (17) Discussion: The model in Figure 5 depicts Stu2 engaged with the microtubule, perhaps using its basic linker region (?). The authors could note this in the figure caption for 5A. The authors do not discuss the basis for plus-end polymerization activity versus polymerization activity at both the plus and minus ends. Do the authors propose that this is due to differential Kt values for the two ends and/or differential localization via the basic region to the two ends?

      Thanks for bring this up. The plus-end selectivity of these polymerases is thought to result from the polarity of TOG:tubulin engagement and the positioning of TOG domains relative to the basic region that provides ‘slippery’ binding to the microtubule lattice. We have partially addressed these issues in the legend to Figure 5 (page 11).

      (18) Brouhard (Cell, 2008) demonstrated that XMAP215 can catalyze the depolymerization of GMPCPP microtubules when no free tubulin, or very low levels of free tubulin, are present. This is interesting in that it indicates that the enzymatic activity is reversible. Can the authors comment on how their model would behave in the low-tozero free tubulin concentration regime? Would a different model have to be invoked?

      This is an interesting comment. Because the enzyme-like model treats the transfer step as irreversible, the model cannot account for the kind of ‘depolymerase’ activity Brouhard and others have noted. A more general model that could also encompass the depolymerase activity at low-to-no free tubulin would need to explicitly model the step(s) involved in microtubule association and dissociation. These steps remain a major open question in the field and while it would be quite interesting, trying to address this in a model is beyond the scope of what we can confidently do given the data we have. To be more explicit about this assumption/limitation, we now point this out in the results section where the model is introduced (page 4, last paragraph).

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1D, legend. "...and a transfer rate constant kf that describes how fast...". Should kf be replaced with kt?

      We made this correction, thanks for pointing the problem out

      (2) Figure 2C. The x-axis label under the blue bar is missing. Also, I find the arrows to the left of the bars confusing and unnecessary.

      We fixed the legend problem. We sympathize with the dislike of the arrows but respectfully prefer to keep them in the hopes of avoiding confusion about the fact we are fitting ‘growth rate attributable to Stu2’, not simply growth rate. The figure legend has been expanded to hopefully smooth this over.

      (3) Figure 3. The kymographs are convincing, but it may be helpful for future studies to state here what the polymerization rates are for 0.6 µM and 1.4 µM yeast tubulin. These values are probably different than what one might expect for mammalian tubulin at those concentrations, and the authors could simply determine them from the slopes in the kymographs.

      Good suggestion, we added the growth rates to the legend (as the reviewer expected, they differ from expectations based on mammalian tubulin).

      (4) Results, page 9, line 11: "...yielded a value of 9.6 nM...". Should this be 8.9 nM, which is that value stated in Figure 4C?

      Actually, these different values are correct. We just wanted to point out that whether we used response amplitudes or measured on- and off-rates, we get very similar values for K<sub>D</sub>. We changed wording to hopefully make this clearer: “Calculating the dissociation constant K<sub>D</sub> from the measured rate constants (K<sub>D</sub> = k<sub>off</sub>/k<sub>on</sub>) instead of from the amplitudes yielded a value of 9.6 nM, in good agreement with the amplitude-based determination of 8.9 nM.” (page 9).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The factors that create and maintain diversity in host-associated microbiomes remain poorly understood. A better understanding of these factors will help in the efforts to leverage the adaptive potential of the microbiome to help solve pressing problems in health and agriculture.

      Experimental evolution provides a promising path forward as we can track the causes and consequences in the emergence of novel variants, but experimental evolution remains underutilized in host-microbiome interactions. Here, Gracia-Alvira utilizes a long-term experimental evolution study in Drosophila simulans under hot and cold temperature regimes to identify strain-level variation in an important fly bacterium, Lactiplantibacillus plantarum. They identify three strains of L. plantarum, which are most prevalent in their respective three temperature regimes, suggesting that these are locally adapted bacteria. Then, using a combination of genomics, in vitro, and in vivo, Gracia-Alvira et al attempt to understand the factors that led to the differentiation of the hot and cold L. plantarum and their impacts on the fly host.

      Strengths:

      This is an excellent use of experimental evolution to track the emergence of novelty in the microbiome. The genomic analyses are all solid and appropriate for the data sets. It is especially striking that the comparisons with the other, independent experimental evolution studies in different labs (and across continents between Portugal and South Africa) show a consistent response to temperature. Many have disregarded the microbiome as it is something that is too sensitive to seemingly innocuous variables (particularly in the fly microbiome), such that we cannot find generalities. However, this finding highlights the potential for experimental evolution to uncover these dynamics. The question of how strains emerge and are maintained is timely and is one of the key open questions in host-microbiome evolution currently.

      Weaknesses:

      (1) The framing in the title and throughout the discussion about "subspecies competition" does not match the data that was collected. The subspecies competition requires actually tracking the competitive outcomes between the hot, cold, and unevolved L. plantarum. In the in vivo work, I can see that mixes of the strains were made, but they did not track whether the cold strain outcompeted the hot strain in vivo under cold conditions, for example.

      We thank the reviewer for the honest concern and take this opportunity to defend our claim of "subspecies competition used across the manuscript. As the reviewer states, subspecies competition requires tracking the competitive outcomes between the three clades, and this is what we did by sampling and sequencing across ten years of experimental evolution (Figures 4 and S3). For this reason, we point that the subspecies competition assessment comes from the direct observation of changes in relative abundance across the time series, and not from the follow-up experiments in vivo or in vitro.

      While Figure 4 is suggestive that there is ongoing competition in the hot temperature regime, this is not necessarily shown in the cold, which is dominated by the C clade. It could also be that the bacteria cannot survive in the flies at the different temperatures. The growth curve assays hint that the bacteria can grow, but the plate reader couldn't actually maintain the 18 {degree sign}C temperature (line 455). So all of this evidence is very indirect and insufficient to say that strain competition is driving these patterns.

      We thank the reviewer for the alternative hypothesis that could explain the observed subspecies dynamic. We rule out that dominance of clade C in the cold occurs because the other two clades cannot grow in this regime based on three pieces of evidence:

      (1) In the time series, clades H and U decrease, but never disappear (Figures 4 and S3), even showing some peaks of abundance in specific replicate populations (Figure S3).

      (2) We isolated individuals belonging to clade H in the cold-evolved populations, as shown in figure 2. This is a direct evidence that clade H prevails in the cold-evolved populations, although in low abundance.

      (3) We did grow the three taxa in fly food Petri dishes incubated at both temperature regimes, observing growth in all cases.

      We will include the food growth experiment in the revised manuscript as further supporting evidence for growth in both regimes.

      (2) The in vivo results are interesting in that there appears to be a fitness cost of clade C, but the explanation is underdeveloped. I say under-developed because in Figure 4, the cold L. plantarum remains much higher throughout adaptation to the hot temperature regime than the hot L. plantarum in the cold regime. The hot L. plantarum is low abundance throughout the cold regime. I felt like this observation was not explained, but it seems relevant to understanding the strain dynamics.

      We acknowledge that a strong fitness cost of clade C is observed in axenic D. melanogaster. In the native host, D. simulans, with reduced microbiome, we observed delayed development that could even be an advantage depending on the situation, as pointed out by reviewer 3 in the recommendations.

      Even if we assume that flies colonized with clade C are less fit in the experimental evolution, another caveat is whether the flies can actively select for the L. plantarum clade. Under this assumption, a clade that imposes a fitness cost to the fly (clade C) should be selected against over time because the flies colonized by this clade will have less offspring or develop later than the rest. Alternatively, as the microbiome is shared among all the individuals in the population, the host might not be able to “purge” the pernicious clade, and L. plantarum dynamics might be controlled solely by the relative fitness between clades in the given experimental treatment. We will discuss this hypothesis in the revision as a way to explain the relationship between the abundance of each clade and the effect on the host.

      I will also note that this is not the first time that L. plantarum or other Lactobacillus have been shown to exert fitness costs to Drosophila. Gould, PNAS, 2018, shows that both Lactobacillus plantarum and Lactobacillus brevis in mono-association have lower fitness (measured through Leslie matrix projections using lifespan and fecundity) than axenic flies. Many studies of wild Drosophila fail to find Lactobacillus, or it is low abundance (e.g., Chandler, PLoS Genetics, 2014; Wang, Environmental Microbiology Reports, 2018; Henry & Ayroles, Molecular Ecology, 2022; Gale, AEM, 2025). This might help provide useful context for the in vivo results.

      We thank the reviewer for the references. These observations are compared to our phenotypic results and discussed in the revised version of the manuscript.

      (3) The data in Figure 4 are compelling to focus on the L. plantarum variants. However, I can see from the methods that the competitive mapping included only other strains of Wolbachia.

      We appreciate the thorough reading of the methods by the reviewer. The competitive mapping comprised two steps: first we discarded the reads that mapped to Drosophila, Wolbachia and additional potential contaminants from sequencing facitilies (human, dog...). This step leaves the reads originated from whole the external microbiome of the flies, including L. plantarum. The second competitive mapping step recruits the reads that map any clade of L. plantarum.

      It is not clear how other members of the microbiome changed in response to the temperature regimes. As I note in point #2, given that Lactobacillus is often rare, it is not clear what the rest of the microbiome looks like over the course of adaptation. Indeed, it seems like Mazzucco & Schlotterer, PRSB, 2021 did a broader analysis of the microbiome and found that Acetobacter is by far the most common bacterium (I think this data is also part of the data shown here?). Expanding on why or why not in this context is important and will improve this study, particularly if the focus is on connecting these evolutionary dynamics to ecological competition to explain the emergence of strain diversity.

      We acknowledge that the rest of the Drosophila microbiome is not addressed in this study, as we wanted to focus the storyline around the intraspecific dynamics found in L. plantarum. We consider that a complete characterization of the whole Drosophila microbiome would unnecessarily elongate the paper and thus we treat it as a constant biotic factor.

      We must point out that our dataset is not the one reported by Mazzucco & Schlötterer, which was done in D. melanogaster, rather than D. simulans. Nevertheless, both experiments share the same infrastructure, temperature regimes and fly maintenance.

      We have included a list of taxa that were isolated from the populations, as well as to report L. plantarum prevalence and abundance across the experiment in order to provide context of the microbiome, beyond L. plantarum, to the readership.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gracia-Alvira et al. investigated how environmental temperature affects competition among members of the microbiome, with a focus on intraspecific diversity, using the Drosophila model.

      Notably, the authors identified three clades of Lactiplantibacillus plantarum from a natural population of Drosophila simulans collected in Florida. They tracked the dynamics of these three bacterial clades under two temperature conditions over the course of more than ten years. Using comparative genomics and phylogeny, they showed that these three bacterial clades likely adapted to their host independently in a temperature-specific manner. Further, by combining in vitro culture and in vivo mono-association assays, they demonstrated the functional divergence of these three bacterial clades phenotypically, including their growth dynamics and effects on host fitness. Lastly, they performed pathway analysis and speculated on key genomic variance supporting such functional divergence.

      Strengths:

      The laboratory evolutionary experiment in response to cold or hot environmental temperature is impressive, given its more than ten years of experimental time period. This collection of achieved microbiome samples paired with the fly host data can be a valuable resource for the field.

      Weaknesses:

      The laboratory evolutionary experiment can be limited due to its artificial experimental setup. For example, wild flies rely on a more diverse set of food sources and are constantly exposed to new bacterial inoculations, whereas under laboratory conditions, flies live in a more restricted ecosystem. In addition, environmental temperatures differ among different locations, but they also involve seasonal changes within the same region. This manuscript can be strengthened with further discussions that elaborate on these limitations.

      As the reviewer has correctly noted, our experimental setting is not exempt from limitations. Lab-reared flies are fed with a defined standard diet. Furthermore, although the system is not completely closed to bacterial migration, this is limited as replicate populations are not allowed to mix during the maintenance of the flies. For this reason, we consider our laboratory setting as a compromise between observing wild populations, which undergo all biotic and abiotic stresses but cannot be manipulated, and evolving the bacteria in absence of the host, or in gnobiotic hosts, in which biotic interactions are not fully considered. We will extend on this in the new version of the manuscript.

      Moreover, the extent of host effects involved in these experiments remains ambiguous, because it is unclear whether these Lactiplantibacillus plantarum mostly reside within fly guts or on Drosophila medium. The laboratory evolutionary experiment possibly favored better colonizers on Drosophila medium under either cold or hot temperatures, which subsequently can saturate fly guts. As fully dissociating these variables can be experimentally tedious, the authors may want to comment more on these aspects in the discussion. Or they may want to consider some measurements. For example, measuring the growth rate of these bacteria on Drosophila medium under different temperatures, in addition to the current MRS culture experiments, or measuring the portion of the Lactiplantibacillus on Drosophila medium versus these stably colonizing fly guts.

      The reviewer's point was briefly addressed in the Results chapter: "Phenotypic differences in liquid culture".

      Reviewer #3 (Public review):

      Summary:

      The study presents an analysis of 297 pangenomes derived from 20 populations of Drosophila simulans, at 19 time points for fast-reproducing individuals in a hot environment, or at 10 time points for slow-reproducing individuals in a cold environment, over a period of more than 10 years. The authors select a particular microbial component of the pangenomes and study the dynamics of Lactiplantibacillus plantarum strains in two environments. They discover that the revealed operational taxonomic units could be divided into three phylogenetic clades, which have their own genomic and genetic features, different adaptive capabilities that depend on the environment, and have a distinct impact on the fitness of the host.

      Strengths:

      The authors prove that bacterial microbiome components are sensitive to the environment and could rapidly (years) be fixed in eukaryotic populations. This study establishes a tractable model that potentially enables the study of variability of the physiological influence of distinct strains of an important commensal species, Lactiplantibacillus plantarum, on the Drosophila host. It is clearly shown that this single species consists of several phylogenetically and functionally diverse strains. The authors did not limit their interest to their own model, but rather they have integrated a comparative approach by analysing phylogenetic relationships among 92 described L. plantarum strains.

      Overall, the study is novel and delivers important discoveries of a longitudinal, well replicated experiment, generating a substantial amount of genomic data. It highlights an important dimension of research that environmental selection operates at the subspecies level.

      Weaknesses:

      Even though the authors show only one particular example by conducting their longitudinal experiment, they honestly acknowledge failures important for interpretation of the biological significance of the results (gnotobiotic mono-association experiments was done with D. melanogaster, but not D. simulans) and therefore they state limitations of their conclusions (weaker effects in the non-axenic flies are due to the presence of other taxa or to higher-order interactions with other members of the microbiome). These interactions could significantly affect bacterial growth, metabolism, and physiological influence on the host.

      We agree with the reviewer in that the use gnobiotic animals is a limitation, as by "tuning" the flies' microbiome we are modifying the interactions between members, which can potentially change the phenotypic outcome. Nevertheless, we use it as a complementary approach, rather than the only inference in our study.

      The authors exploit the results of their experiment to speculate about a wide range of evolutionary phenomena, like within-species competition, ecological adaptation and evolution of the host, fitness advantage of bacteria to the host, the benefits of parasitism or mutualism, the domestication of the microbiome, etc. At the end, they conclude that their study "highlights that even subspecies diversity plays a key role in adaptation to environmental temperature". However, the potential mechanisms of such adaptation are barely discussed, so that the focus of the study shifts from the temperature-induced changes in microbial population structures toward metabolism-related adaptations of clade representatives that enable them to diversify their carbon and nitrogen sources. The role of the temperature factor remains elusive.

      We acknowledge that our study does not fully resolve the mechanism by which a different clade ends up dominating each temperature regime. The MRS liquid experiment was an attempt to answer whether differences in optimal growth temperature could explain the temperature-specific abundance of the two clades. Our experiments showed, however, that this was not the case. Beyond this point, it is hard to disentangle the role of the temperature, as it could also act indirectly on the bacteria, for example, through the host or the food.

      A second observation in our time series was that a third clade, U, was unfit in both regimes despite starting the experiment in high abundance. For this reason we also studied what made this clade less fit. Based on our analyses, we propose that the decrease of clade U was driven by the shift to a laboratory diet, shared by all experimental populations.

      In addition to that, the paper has a clearly minimalistic experimental approach to address functional properties of the revealed L. plantarum strains, so that their own fitness, or their relationship with the Drosophila host, is characterised superficially. Therefore, the authors' discourse can be speculative rather than factual (especially when the authors use the expression "likely" to share their guesses in the "Results" section). Nevertheless, these minor drawbacks do not underscore the novelty of the discovered phenotypes and the importance of their further investigation.

      We consider the reviewer's concern and toned down the phrasing when reporting our findings in the revised version of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) One solution to resolve the "competition" issue would be to check that the L. plantarum strains are established at similar or different titers in the in vivo work. Fly phenotypes can be sensitive to microbial load (Keebaugh, iScience, 2018), which might explain some of the counterintuitive in vivo results. In line 227, the authors mention that "bacterial load" is contributing to the magnitude of the effect, but I don't see the data reported anywhere. If this is from the in vitro assays, then the authors need to show that in vitro predicts in vivo L. plantarum abundance.

      Bacterial load inoculated in the in vivo experiment was normalized to OD=0.05 (~5*10<sup>6</sup> CFUs/ml) for the three clades at the beginning of the experiment. Thus, all vials were inoculated with the same titer of L. plantarum. Only the genotype varied between treatments. However, it is possible that, once inoculated, each clade grew to a different titer (as they have different growth rate and carrying capacity).

      Statement in line 227 comes from the differing results in transfers 1 and 2. In transfer 1 we inoculated a fixed load of ~2.5*10<sup>5</sup> CFUs. In transfer 2, however, we inoculated no bacteria to the food, and the flies seeded the vial. Our statement comes from the assumption that bacterial load in transfer 2 has to be lower than in transfer 1 as bacteria seed the vial solely by defecation of the parents.

      Following the reviewer's suggestion, in the revised version of the manuscript we have included a new experiment in which we quantified the bacterial load of each clade in individual flies.

      (2) Tracking the competitive outcomes is tricky, though it could be done with whole genome sequencing. An alternative would be to label the strains with fluorescent proteins (e.g., Obadia Current Biology 2018 has done this in Lactobacillus) and track fluorescence to better understand the results of the "mix" treatment in Figure 6.

      We appreciate the feedback of the reviewer, but consider this rather labour-intensive approach as an interesting option for future follow-up experiments.

      (3) That being said, my main concern with this is the "competition" claim. If the paper were reframed appropriately, this paper could still make an important contribution to the evolution of host-microbiome interactions, but the authors would need to consider what they can and cannot do with this interesting dataset.

      The "competition" claim comes from the changes in relative abundance observed in the time-series data, not from any of the follow-up experiments. Thus, we consider the use of the term "competition" appropriate.

      (4) The text on the figures is very small and hard to read.

      We increased the size of the text in all figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Have you conducted the in vitro culture experiments following the "cold" conditions?

      We have conducted the experiment in "cold" conditions, but with some modifications to the experimental settings, as the plate reader did not have cooling capacity. Instead, we grew a subset of the isolates (four per clade) in glass vials at constant 20 °C, and measured their OD twice a day. We have included the results in the revised version of the manuscript.

      (2) How many technical and biological replicates were measured for the in vitro culture experiments (Figure 5)? Please add this information to the figure legend and method.

      We measured the growth of four isolates from clade U, nine isolates from clade C and sixteen isolates from clade H. Each isolate was grown three times.

      We have included this information, as requested by the reviewer.

      (3) Making the labels in Figures 2, 3, 5, and 6 bigger would be helpful.

      We have increased the font size of all figures.

      Reviewer #3 (Recommendations for the authors):

      (1) Line 268: "Based on our results in experimentally evolved fruit flies, we propose that within-species competition, thus far largely overlooked, could contribute to ecological adaptation and evolution of the host". Overstatement should be avoided, since the evolution of the host was not directly studied here.

      Our results show that reproductive traits of the host differ upon colonization with each clade. Although we don't test the host's evolution, we speculate that flies differing in their offspring number and developmental time might differ in their overall fitness. Finally, we consider the Discussion section as the right place for speculation and development of hypotheses that can be tested in future work.

      (2) Line 258: "These differences do not explain the clade-specific selection, but reflect the different evolutionary histories of the clades". The temperature factor and its possible role in clade selection would be better discussed at least a little bit.

      In this paragraph we described potential metabolic differences between clades using comparative genomics. We did not find enrichment in a function or group of functions that could explain the different dynamics between clades H and C in the temperature regime.

      In the revised version of the manuscript we highlight that we did not find temperature-specific differences from this analysis.

      (3) Line 252: "...This could explain why clade U, which displayed a high growth rate and carrying capacity in liquid culture". The statement could be further developed with a caution. Even if the isolates that belong to the clade U are outcompeted by H or C, it should be noted that the strain U cannot be used as a true reference for fitness, since it could possess its hidden adaptive properties, not being simply "a loser". Such a hypothesis could explain the maintenance of this strain in the wild.

      We agree with the reviewer in that fitness is relative to the selective environment. Clade U is less fit than H and C in our specific experimental conditions, but it could outcompete them in other conditions, such as wild flies or MRS liquid medium. In the revised version of the manuscript we have rephrased this statement to clarify that we specifically refer to clade U's fitness under the new laboratory conditions.

      (4) In a cold environment, association with the clade C induces developmental delay and produces less progeny, which potentially allows the host to survive in case of harsh conditions and potential food limitation. Could the authors speculate and not exclude that this phenotype could be potentially adaptive? It would be curious to check in further studies whether flies associated with C strains are more stress-resistant, for example.

      We thank the reviewer for this alternative hypothesis. In our manuscript we used the Darwinian definition of fitness; reproductive success of an organism in the focal environment. And thus, both higher progeny per female and shorter developmental time would be beneficial in direct competition with other individuals. It is true that delayed developmental time, or less progeny, could be advantageous in specific cases. This could be the case for D. simulans inoculated with clade C. However, we consider that the fecundity levels observed in D. melanogaster upon inoculation with Clade C (average of 0.06 offspring/female*day in the cold) are too low to sustain a population.

      We have included this hypothesis in the Results section.

      (5) It would be highly recommended to add an experiment to complete the story by measuring the quantity of bacteria in the medium and in the flies. This will resolve the hypothesis (Line 785): "Thus, the ability to exploit this ubiquitous source of carbon and nitrogen could be very advantageous in the fly microbiome context, but would not affect the fitness in liquid culture".

      Following the reviewer's recommendation we included two additional experiments. We measured the bacterial load per fly in the native host, D. simulans, inoculated with the three clades. We also compared the clades' growth speed in solid fly food (without host). In the former experiment, we found similar bacterial loads upon inoculation with clades U and clade H. In contrast, in the latter we found delayed growth of clade U relative to H and C in the food. Thus, chitobiose consumption does not seem to provide an advantage in the fly gut to clade H. We attribute the fitness advantage of H and C to their advantage growing on the laboratory fly food, regardless of the host.

      Both experimental results have been included in the revised version of the manuscript, and the comparative genomics paragraph and discussion have been modified in consequence.

      (6) The chapter "Extended clade-specific differences in KEGG metabolic pathways" could be presented in the main text as it contains important results. These results are mentioned in the chapter "Functional divergence on the genomic level", which looks rather humble when it stands alone as it currently does.

      We appreciate the interest of the reviewer in this supplementary chapter. To keep the length of the manuscript digestible for a broad set of readers, we decided to only include in the main text the functional differences that could play a role in adaptation to the new laboratory environment.

      We consider that a full description of the metabolic differences between the three clades has to be published, as it might be relevant for researchers interested in L. plantarum metabolism. However, it does not fully follow the storyline, as the differences reported in the supplementary, such as nitrate respiration or synthesis of molybdenum cofactors, might not be involved in the clade-specific selection observed in the time series.

      (7) Line 773: "Clades C and H encode a shared genetic repertoire related to sugar/riboflavin metabolism that is lacking in clade U". This indeed allows us to hypothesise that the fixation of these clades in fly populations was due to their improved metabolic capabilities. However, the analysis of fitness shows similarity in flies associated with clades H and U, meaning that sugar/riboflavin metabolism in H does not provide an obvious adaptive trait to flies. Moreover, one could say that sugar metabolism in clade C is maladaptive not only for flies, but also for bacteria in liquid cultures. It is recommended to more clearly state the respective limitations of the study.

      Here we have to make a distinction between bacterial fitness and host fitness. The three clades differ in their (bacterial) relative fitness, as evidenced by the time-series dynamics (Figure 4). In the cited statement we hypothesize that a more versatile sugar metabolism repertoire could increase the bacterial fitness of clades H and C (relative to clade U) in the sugar-rich laboratory diet.

      This is independent of the fitness effect that L. plantarum could have in the host. Finally, as it was discussed in the recommendation 3, fitness is specific to the environment. Clade C is the least fit in liquid MRS in hot conditions, but the fittest in cold experimental conditions.

      (8) The authors should better explain why growth in MRS was not performed in a cold temperature regime to further support or refute the hypothesis that capacity and inflection time could partially explain the higher fitness of bacterial strains from clade U.

      We did not perform this experiment in cold conditions due to technical limitations of the plate reader, that does not have cooling capacity. Nevertheless, following the reviewers' suggestion, we have included in the revised version of the manuscript a new MRS growth experiment in cold-like conditions (constant 20 °C).

      (9) When mentioning that L. plantarum can "increase larval fitness of Drosophila melanogaster relative to germ-free flies" (line 196), the authors should specify in which specific conditions this phenotype was observed, and how relevant the mentioned phenotypes are to the current study.

      Following the reviewer's recommendation, we have modified the paragraph in order to clarify the conditions used in other papers and those used in our work. The references cited in this section (PMID: 21907145, 29290388 and 28062579) report that L. plantarum increases the host fitness in protein-poor diets (12 g/l of dried yeast or less), but not in high-protein diet (50 g/l of yeast or higher). Since our experimental diet contains an intermediate amount of protein (24.3 g/l of dried yeast) we were agnostic of whether L. plantarum would benefit the host or not in our conditions. Regarding the phenotypes, we chose two reproductive traits that are affected by changes in the microbiome according to the literature. Developmental time is directly affected by L. plantarum in the aforementioned papers. Offspring number is another fitness component affected by Drosophila microbiome (PMID: 30510004).

      (10) Provide a reference for line 205: "In axenic D. melanogaster none of the L. plantarum clades provided a fitness advantage to the host relative to germ-free controls, contrary to the effects reported in the literature". If the conditions were different from those in the studies referred to, then it would be of no use to compare the fitness advantage (for example, in Reference 24 another type of diet was used).

      Already covered in recommendation 9.

      (11) Please provide more context to this statement (Line 210): "The high content of dried yeast 24.3 g/l in the fly food used in our experiment likely provided already sufficient amounts of essential amino acids, which negated the growth-promoting effects of L. plantarum". It is not clear why amino acids are taken into account, and what the evidence is for the fact that the amount of essential amino acids was sufficient to abolish growth-promoting effects.

      The whole paragraph was modified in order to clarify the relationship between protein input and nutritional fitness benefit of L. plantarum.

      (12) Please provide measurements of bacterial quantity which would support the statement (Line 215): "The fitness reduction was stronger in the first transfer of flies, likely due to a higher bacterial load".

      Upon request of the reviewer, we have estimated the bacterial load per individual fly in D. simulans. Additionally, we have specified the CFUs inoculated in the vials in transfer 1.

      (13) Correct the typo (line 220): "However, the developmental time was significantly extended after inoculation with clade C at cold temperature (Dunn's test, p < 0.05 05 for all significant comparisons)".

      Done.

      (14) Specify more precisely the temperature conditions referred to in line 226: "In summary, we observed that clade C, which is dominant in the cold-evolved populations, decreases host fitness when axenic flies are inoculated". Does it decrease fitness both in hot and cold environments?

      For the axenic flies, we did find a decrease in fitness in both regimes, yes. We specified it in the revised version of the manuscript.

      (15) Please provide evidence for line 227, or otherwise rephrase it: "The magnitude of this effect varies depending on the environmental temperature, the bacterial load, and the presence of other microbial taxa".

      Novel evidence was provided regarding the role of bacterial load on host fitness.

      (16) Correct the following statement, so that it reproduces the results of the original work (reference 19, line 229): "In a low-protein diet, strains that were not isolated from Drosophila enhanced larval growth relative to germ-free individuals, whereas another Drosophila-associated strain did not have any effect".

      This statement was removed from the revised version. This reference was cited in the discussion to state that: " the nutritional symbiosis in L. plantarum is strain-specific".

      (17) Please provide a rationale for using KEGG Orthologs. Why was this database chosen as an appropriate one, even though it is known to be a non-exhaustive metabolomic resource?

      KEGG is a well-known metabolic database that is widely used in comparative genomics (PMID: 40177264) and built in state-of-the-art software for microbial ecology such as Anvi'o (PMID: 33349678). Other similar gene-to-function databases are less focused on metabolic pathways, such as COG or GO, or limited to specific enzymatic activities, like CAZy. Furthermore, the hierarchical organization of KEGG Orthologs in modules and pathways allowed us to map clade-specific orthologs to the broad metabolic context. For these reasons, we considered KEGG to be the best option for this analysis.

      (18) Line 250: "Therefore, we speculate that the ability to exploit this ubiquitous source of carbon and nitrogen in the lab-maintained fruit flies, could be a strong target of selection in the lab environment". This statement concludes the "Results" section but would be more appropriate for the Discussion section, since the authors do not provide any experimental evidence that could support this statement.

      We have modified this chapter, as covered in recommendation 5.

      (19) Line 276: "However, the intraspecific richness of L.plantarum in our flies was three times higher than that estimated in human gut microbiomes". Note that there are other recent studies which show the presence of several OTUs within L.plantarum isolates (for example PMID: 41484402).

      We thank the reviewer for the reference. We comment on it in the revised manuscript.

      (20) Line 287: "Our finding shows that the well-characterized nutritional symbiosis between Drosophila and L. plantarum depends on the bacterial genotype and cannot be generalized to the entire species". Note that such a conclusion has already been previously stated (for example, PMID: 30008290 and 28993620).

      We thank the reviewer for the references. Indeed, these papers show that some L. plantarum strains are beneficial for the host while others are neutral. Furthermore, as commented by Reviewer #1 in the public review, L. plantarum has been shown to reduce the host's fitness relative to axenic flies (Gould, PNAS, 2018).

      Our observations are novel in two ways. (1) The fecundity observed in D. melanogaster, 0.06 offspring/female/day in average, is lethal (in Gould et al. 2018 fecundity never decreased below 1 offspring/female/day). (2) Clade C outcompetes the other clades in the cold, despite being detrimental for the host.

      We have modified the Discussion to account for the previous work.

      (21) Line 343: "In addition, we obtained L. plantarum genomes from two other experimental evolution studies. Two genomes from the South African experiment and seven genomes from the Portugal experiment". Merge two sentences into one.

      Done.

      (22) Line 360: "At sampling, the age of the flies varied between four and eight days for the hot environment and between nine and 16 days for the cold environment". Please comment on the fact that different age of flies (different physiology) is not the reason for bacterial community differences.

      During maintenance, flies are sampled at different ages because the temperature affects their developmental time. We cannot rule out the hypothesis that age difference drives microbiome differences. Temperature could affect clade competition directly (differences in optimal temperature between clades) or indirectly, by affecting either the host (e.g. changes in Drosophila developmental time alters L. plantarum fitness), the surounding microbiome, or the food (e.g. increased metabolic activity in the hot regime changes nutrients profile). We ruled out the direct effect of temperature with growth experiments in liquid MRS medium and solid fly food, but disentangling the indirect effects is not feasible.

      (23) The majority of figures have low-quality labels that are not legible due to the small size of the font. Please improve.

      Done.

      (24) Figure 1 - Correct the legend: There is no "10" label on the picture. Probably by 10, the authors mean "Generation", while by x10 - number of isogenic replicates.

      Done.

      (25) Figure 2 - No numbers at nodes are indicated, whereas it is announced in the legend that they represent bootstrap support values. In addition, it is recommended to show a reference pangenome in the middle panel to clearly refer to the total size of the possible black bar.

      We added high bootstrap support as coloured nodes in figures 2, 3 and S2.

      We do not understand the reference pangenome request. In the middle panel, each black/white bar corresponds to an orthologous gene that can be either present or absent in each of the genomes. These orthologs were sorted based on hierarchical clustering of the their patterns of abundance (columns present in the same set of genomes, together), not by synteny. Thus, a reference pangenome would be simply a black bar.

      (26) Figure 2: It would be advantageous to add a figure that represents the frequency of each strain in each replicate (at the last time point, for instance). It would explain why some “blue” strains appear to be within the “red” cluster. Otherwise, it is confusing to find cold-evolved bacterial strains in hot-evolved fly populations.

      The frequency of each clade in each replicate is shown in figure 4. We think that it would be more confusing to follow the suggestion of the reviewer, as the isolates were sampled at different time points of the experiment. We would not like to call it a confusion that "blue" strains appear in the "red" cluster, but rather the logical consequence of the color code used in figure 2, which corresponds to the temperature regime in which the isolate was sampled (regardless of its clade). Whereas in the following figures colour represents the clade. It is thus possible to find clade H isolates in the cold temperature regime, as this clade is in low frequency but not completely absent in this regime.

      (27) Figure 3 – Add a label for the X-axis.

      We rotated the tree to be able to increase the genome IDs. We have added the label to the Y-axis.

      (28) Figure 4 - Please indicate how the clade relative abundance was assessed.

      Clade relative abundance was inferred by mapping competitively the short reads against the three clades’ reference sequences. It is specified in the legend now.

      (29) Figure 6 - Total number of F1 flies eclosed normalised by day (during which period?). What do T1 and T2 correspond to?

      During the respective number of days that females were allowed to lay eggs: one day in the hot settings and two days in the cold settings in transfer 1. One day and three days, respectively, in transfer two.

      T1 and T2 correspond to the first and second transfers, as described in the Materials and Methods. In first transfer, flies laid eggs in vials pre-inoculated with a set load of L. plantarum. After egg laying, same adults were then transferred to a sterile set of vials and allowed to lay eggs again (second transfer). Bacterial load in these vials was solely seeded by the parents.

      In order to avoid any confusion, in the revised version of the manuscript we have modified figure 6 to show transfer 1 for both Drosophila species, and moved transfer 2 dataset to supplementary figure S5.

      (30) Figure S2 - label the X-axis.

      We guess the reviewer means Y-axis. Done.

      (31) Figure S3 demonstrates the real data and its variability, so it would be better used instead of Figure 4 (which seems to be just a derivative from Figure S3, not a separate dataset and separate type of analysis).

      As the reviewer suggested, we have replaced figure 4 with figure S3.

      (32) Figure S4: Improve plot title: (e.g., C:H:U = 3:3:3).

      Done.

      (33) Figure S6: It is stated that N = 10; however, some datasets do not have 10 points represented. Please specify why. Also, please specify the meaning of "T1/T2".

      For the inoculation experiment in Drosophila simulans, we had nine replicates per treatment, not ten. This has been corrected in the figure and in the Materials and Methods section.

      T1 and T2 correspond to the first and second transfers, already covered in recommendation 29.

      (34) Table S3: provide legend for values (1 - present in all strains, but 0.04 - what does it mean?).

      It means that 4% of the genomes from this clade harbour the specific gene. We have specified it in the legend of the revised table.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Uphoff et al. propose a structural and mechanistic model in which the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, thereby promoting downstream receptor phosphorylation and signaling. Using AlphaFold-based modeling, the authors predict that the CCP20 domain of SVEP1 binds to TIE1, creating a composite surface that facilitates Angiopoietin association and TIE1 activation. The resulting ternary model (SVEP1-TIE1-ANG) offers a structural rationale for how SVEP1 converts TIE1 into a functional, ligand responsive receptor. Additional models and biological assays suggest roles for other domains of SVEP1, such as CCP5-EGF-L7, although these interactions are predicted with low confidence. The authors interpret these findings as the first structural framework for how SVEP1 enables ANG-TIE1 signaling.

      Strengths:

      (1) The central hypothesis - that SVEP1 enables ANG binding to the orphan receptor TIE1 - is biologically compelling and addresses an important question in vascular biology.

      (2) The AlphaFold-predicted ternary complex (SVEP1-TIE1-ANG) is plausible, high-confidence, and structurally consistent with prior functional data (e.g., poly-Ala scanning from Sato-Nishiuchi et al.).

      (3) The authors' model offers a potential explanation for the previously observed role of SVEP1 in enhancing ANG signaling through TIE1, and may represent the first structural insight into TIE1's transition from orphan to ligand-activated receptor.

      (4) The potential clinical implication - that a combinatorial ligand (ANG+SVEP1) can activate TIE1- could have translational relevance for vascular leak and inflammatory disease.

      Weaknesses:

      (1) Lack of structural validation and mechanistic follow-up: Despite the promising AlphaFold model, there are no figures of the predicted interface, no residue-level interactions shown, no ipTM values reported, and no experimental follow-up to test the interface. PAE plots are incorrectly used as confidence justifications, which is not appropriate for complex predictions.

      We have appended the data showing AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity. We also added ipTM scores and confidence plots for the predicted complexes.

      (2) Biophysical validation is missing: No surface plasmon resonance (SPR), ITC, or biochemical assays are included to confirm ternary complex formation or quantify binding kinetics. Given the manuscript's structural focus, this is a major gap. For instance, an SPR experiment where ANG is immobilized, and TIE1 binding is measured {plus minus} SVEP1, would directly test the model. And allow direct comparison to ANG-TIE2.

      We have addressed this question and performed ELISA assays to measure binding affinities between SVEP1 and TIE1 in presence or absence of ANG1 or ANG2, thus confirming that the affinity is increased in the presence of ANG1 or ANG2.

      (3) Missed opportunity for mutagenesis-driven validation: The manuscript does not include any interface-targeted mutations, despite clear opportunities. For example, mutating T2595 in SVEP1 (to R) or mutating the TIE1-specific residues (residues PL 202-203 to LF) could strongly test the model and potentially reveal dominant-negative behaviors. E.g. A T2595 mutant should block ANG binding but not TIE1 binding.

      We have depicted figures of the interfaces including P202-L203 and included the TIE1 P202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study, for the following reason: Modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2, thus reducing its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both TIE1 and ANG2, the data is not included in the manuscript.

      (4) Overinterpretation of weak models: The additional AlphaFold model involving the CCP5-EGFL7 domains binding TIE1 has extremely low confidence (ipTM < 0.15) when reexamined by this reader and should not be emphasized. There is no biophysical evidence or binding data (SPR) to support this interaction, and its inclusion detracts from the much stronger CCP20 model.

      We agree with this point made by both reviewers and have removed the data on CCP5-EGFL7 from the manuscript.

      (5) Language around modeling is overstated and potentially misleading: Terms like "unequivocal," "high-affinity," or "affirms strong binding" in reference to AlphaFold predictions are inappropriate. These are hypotheses -not confirmations - and must be tested at the biochemical level. This should be clarified throughout the manuscript to ensure non-experts do not misinterpret modeling confidence as binding affinity.

      We agree with the reviewer, and have adjusted the wording.

      (6) Negative stain EM data is not informative due to low resolution and lack of defined interfaces; unless replaced by higher-resolution Cryo-EM, this should be omitted. Better would be co-gel filtration, AUC, or SEC-MALLs with ANG-SVEP1-TIE1.

      We have now appended the data by adding gold-labelled TIE/ANG proteins, thus enhancing clarity.

      (7) Disjointed narrative: The manuscript presents a compelling mechanism involving CCP20-driven ANG binding to TIE1, but then becomes fragmented by introducing the low-confidence CCP5-EGFL7 model and speculative higher-order polymerization models that are not experimentally supported.

      We agree with this point and have have centered the manuscript around CCP20. We removed data concerning CCP5-EGFL7 as suggested by both reviewers.

      Reviewer #2 (Public review):

      Uphoff and colleagues present the results of a study focused on characterizing the binding of SVEP1 to TIE1 along with Angiopoietin-2. Starting with computational prediction of SVEP1 binding to TIE1, the authors identify the region of SVEP1 that serves as a high-affinity ligand for TIE1. Advanced studies identify a weak secondary binding site within SVEP1 that appears to be sufficient but not necessary for its interaction with TIE1 based on in vivo rescue experiments. The most novel contribution of the manuscript seems to be the identification of angiopoietin-1 and -2 as co-factors that seem to enhance the binding of SVEP1 with TIE1 and impact downstream AKT signaling. They propose a complex in which SVEP1 binds to TIE1 and ANG2.

      Although the first set of results is essentially confirmatory, the identification of ANG-2 as a "cofactor" enhancing the binding of SVEP1 to TIE1 and associated downstream signaling (i.e., Figures 3 and 4) is novel and is of interest. However, the manuscript and its conclusions would greatly benefit from some clarifying details and additional experiments to ensure rigor and support specific claims.

      We have addressed the reviewers concerns and significantly appended the manuscript. Most importantly, we provide structural validation of AlphaFold models reporting interfaces, residue-level interactions and ipTM values. We have included new biophysical validation of binding kinetics of SVEP1 and TIE1 in the presence or absence of ANG1 or ANG2. Furthermore, we have removed the data on CCP5-EGFL7 from the manuscript in order to retain focus on the CCP20 domain.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The AlphaFold modeling for the CCP20-based interactions is strong (as determined by this reader rerunning and getting ipTM values and visually inspecting the interactions “because this is not in the manuscript”). As presented, the manuscript stops at the hypothesis-generation stage. Validation is needed to fulfill the paper's title and claims. The structure-function link is not demonstrated, despite an obvious and achievable experimental path (mutagenesis, SPR, kinetics).

      Additional Context and Suggestions

      (1) Show and label AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity.

      We have appended the data in the new supplementary figures 1.1, 1.2, 2.1 and 2.2

      (2) Provide ipTM scores and confidence plots for each predicted complex.

      We have added the values and plots in the new supplementary figures.

      (3) Perform SPR assays with ANG-coated surfaces and measure binding of TIE1 {plus minus} SVEP1. Compare to TIE2 binding for context.

      We performed the proposed experiment using an ELISA assay to measure binding affinities between SVEP-1 and TIE1 in presence or absence of ANG2 and included these data in the manuscript in figure 2.

      (4) Test interface mutants: e.g., T2595R in SVEP1 (should impair ANG recruitment but not TIE1 binding), or PL→LF muta on in TIE1 (should disrupt SVEP1 binding).

      We have depicted figures of the interfaces including P202-L203 and included the TIEP202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study. Our modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2 thus reduce its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both, TIE1 and ANG2, the data is not included in the manuscript.

      (5) Clarify in the Introduction that SVEP1 is a large, multidomain ECM protein to help readers contextualize the domain names early on. Do not use terms like CCP before defining them.

      We added: “Svep1 encodes a 3571 amino acid long extracellular matrix protein containing different domains such as Willebrand factor type A domain (vWF), ephrin-receptor like domains, complex control protein (CCP) domains, and Hyalin repeats at the N-terminus. The C-terminus mainly consists of CCP and EGF domains. Svep1 is expressed in mesenchymal cells, but not in endothelial cells, and functions non-cell-autonomously (Karpanen et al. 2017; Morooka et al. 2017)”.

      (6) Replace or remove negative-stain EM unless higher-resolution cryo-EM data are available.

      We would like to retain the EM data, but have now replaced the negative-stain EM by adding gold-labelled TIE/ANG Proteins to verify that the proteins we show are the ones we expect. The reason we would like to retain the data is that TIE1 has been considered for so many years as an orphan receptor, and thus we consider it appropriate to demonstrate SVEP1/TIE1 binding using multiple methods.

      (7) Reframe claims around modeling to avoid overstatement. For example: "The model suggests a plausible mechanism consistent with prior biochemical data" is more appropriate than "unequivocal".

      We agree with the reviewer, and have adjusted the wording.

      (8) Consider narrowing the focus: the CCP20-TIE1-ANG model is a strong story on its own. The CCP5EGFL7 model and polymerization hypotheses are not essential and may dilute the impact.

      Since this point was made by more than one reviewer, we have removed the data on CCP5EGFL7 from the manuscript.

      (9) Properly define TIE1: Tyrosine kinase with Ig and EGF domains.

      We have corrected the full protein name for TIE1 and added: “TIE1 and Tie2 exhibit a high degree of homology with a globular head domain consisting of three immunoglobulin-like (Ig) domains and three epidermal growth factor-like (EGF) modules and a short stalk formed by three fibronectin type III repeats, while the N-terminal two Ig domains of Tie2 harbor the angiopoietin binding site (Macdonald et al. 2006).” We also added: “D1 and D2 refer to the two N-terminal Ig domains, D3 refers to the three EGF domains and D4 to the third Ig domain of TIE1 or TIE2.”

      (10) Use RTK, not tyrosine kinase receptors TKR.

      Tyrosine Kinase receptor was replaced by receptor tyrosine kinases (RTKs)

      (11) PDBs (.cif and .json files) of the models must be supplied for the readers so they don't need to rerun the AlphaFold jobs.

      We are providing all PDBs with the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) In several locations, the authors state that alphafold detects an "unequivocal" and/or "high affinity" interaction. Experts on computational structure prediction can weigh in, but I am not sure it is accurate to say that alphafold predicts affinity. Quantitative estimates of the prediction confidence or other parameters of the alphafold output are not provided.

      Thank you for the comment, we agree with the reviewer and have adjusted the wording. We have also appended the data showing interfaces, ipTM scores and confidence plots for the predicted complexes.

      (2) In Figure 1C, what concentrations of TIE1 and TIE2 are being used in the SPR experiments shown?

      We have added the concentrations of the proteins to the methods section

      (3) In Figure 1C, what is the affinity constant (KD) of the interaction between SVEP1 and TIE1 and SVEP1 and TIE2?

      We have added the values of affinity constants to the manuscript text.

      (4) In Figure 1C, the authors immobilize a "70kD" fragment of human SVEP1 to determine the interaction between SVEP1 and TIE1. How does the affinity they measure between the SVEP1 fragment and TIE1 compare to the affinity between immobilized full-length SVEP1 with TIE1?

      The largest molecule we used in any assay is not full-length SVEP1, but consisted of a C-terminal SVEP1 protein spanning from the first EGF domain to the C-terminus (approximately 295 kDa) as previous studies have shown that SVEP1 is proteolytically cleaved N-terminal to the first EGF domain, generating a protein of this size. In our ELISA assays, this molecule has a higher affinity to TIE1 than the 70 kD fragment. We have included data for the 70 kDa as well as for the 295 kDa fragment in the manuscript. ELISA assays have used larger SVEP1 fragments (as indicated in the figure legends), the SPR assay was performed with the 70 kDa fragment.

      (5) How does the alphafold structure prediction for the "70kD" fragment of hSVEP1 compare to the prediction of the same 70kD fragment from the full-length protein prediction?

      All modellings using different SVEP1 fragments including the 70 kD and full-length version identify the putative binding site at position CCP20. While AlphaFold3 can predict the correct domain folds in both the 70kDa fragment and full-length SVEP1, the orientation of these domains is highly variable due to flexible linkers between each domain module. This flexibility effects the output confidence metrics, thereby hampering our interpretation of the models. Therefore, we conducted the structural prediction of the complexes with smaller fragments and not the full-length SVEP1.

      (6) What are the amino acids for the 70kD fragment?

      The relevant amino acids are 2261-2890. The accession number and amino acids of each protein have been listed in the key resources table.

      (7) In Figure 1D, what is being depicted by the red stars? This is not explained in the text or figure legend.

      We have replaced figure 1D.

      (8) By itself, Figure 1D is not terribly informative and in my opinion does not support the statement that the authors "were able to directly visulalize the attachment of SVEP1 and TIE1." As a minimum, the authors should repeat the same set of images with SVEP1 and TIE2, but other approaches, such as labeling, could be performed.

      We replaced figure 1D with new data and gold-coated protein enhancing clarity. We think it beneficial to demonstrate SVEP1/TIE1 binding using multiple methods as TIE1 has been considered as an orphan receptor for so many years. We have not performed these experiments with TIE2 proteins as we were not able to show binding of TIE2 to SVEP1 with other assays.

      (9) What is being stained in Figure 1D? Full-length SVEP1/TIE1? Or fragments of these proteins?

      We replaced figure 1D by a new assay with labeled proteins using the 150 kDa version of SVEP1 and the ectodomain of TIE1 as well as ANG1 or ANG2 (new figure 1D and new supplementary figure 2.3). TIE1/ANG proteins were gold-labelled. The protein fragments used in this assay are described in detail in the methods section.

      (10) The authors discover CCP6-EGFL7 as a low-affinity binding region of SVEP1 for TIE1. Is this region in physical proximity to CCP20 (the high-affinity binding region for TIE1) based on alphafold prediction? How would the authors think this is binding TIE1?

      We have removed this data set (see comment to reviewer 1’s request).

      (11) What is the affinity constant (KD) for CCP6-EGFL7 with TIE1?

      We have removed this data set (see comment to reviewer’s 1 request).

      (12) The authors claim that ANG1/ANG2 increase affinity between SVEP1 and TIE1 based on immunoblotting. Immunoblots are semi-quantitative at best. If the claim is higher affinity, I think the authors should measure this by SPR and determine the KD between immobilized SVEP1 with TIE1 in the absence and presence of ANG1 and/or ANG2.

      We conducted ELISA assays (figure 2) showing that the affinity is increased in the presence of ANG1 or ANG2 and agree with this reviewer that this strengthens the data.

      (13) In Figure 3, can the authors explain why ANG1/2 does not pull down with SVEP1/TIE1?

      We noticed that upon transfection of TIE1 into HEK cells, ANG1/2 is almost undetectable anymore in the total lysate. Thus, we believe that after the pull down we are below the detection limit.

      (14) In Figure 3, what is "TL"? I assume total lysate, but this is not specified.

      Thank you, we now specify TL as total lysate.

      (15) In Figure 3 "TL" panel (again, I assume this is total lysate), why are the ANG1/2 immunoblots so weak when co-transfected with TIE1?

      We consider it likely that in the presence of TIE1, ANG1/2 proteins are internalized and digested. Another reason for low signals could be that upon transfection of two plasmids, the amount of ANG1/2 protein is reduced as the cell has limited capacity for transcription and translation.

      (16) In Figure 3B, why is the SVEP1 fragment now 150kD when 70kD fragment was previously used? What domains are contained in this 150kD fragment?

      We now better define the domains of the 150kD SVEP1 fragment. The 150 kD fragment was the one produced first and available in high amounts in our laboratory and thus used for functional assays. The 70 kDa fragment together with ANG2 also induces phosphorylation of AKT, but it was not used in as many conditions/replicates as the amounts we had available were lower.

      (17) In Figure 4A, signaling with SVEP1 by itself and ANG2 by itself should be shown to support the claims being made.

      We added the lines for SVEP1 and ANG2, and also the quantification. SVEP1 itself already affects the phosphorylation of TIE1, most likely because hdLECs produce ANG2 by themselves. We show this with the ANG2 blocking antibody for pAKT.

      (18) For pAKT, what are the concentations of proteins being used and the times of incubation?

      This information is provided in the Materials and Methods section. We added the concentration of the anti-ANG2 antibody, which was missing.

      (19) It seems that p-AKT and AKT are being blotted on different gels. If this is correct, loading controls need to be shown for p-AKT blot.

      We added HSC70 as a loading control for both blots.

      (20) It is interesting that anti-ANG2 antibody inhibits SVEP1-induced p-AKT signaling. As the authors may know, SVEP1 has been identified as a receptor for PEAR1, which also leads to downstream p-AKT signaling, which seems to be independent of ANG2. Do the LECs being used here express PEAR1? If these cells express PEAR1, how do the authors think ANG2 silencing will eliminate SVEP1-associated p-AKT signaling?

      hdLECs express PEAR1. However, we show that p-AKT signaling is attenuated after siRNA KO of TIE1. Thus, the downstream signaling is dependent on TIE1 (Figure3).

      (21) Again, experts on computational structure prediction can weigh in, but I am not sure how to interpret the prediction of the 2:2:2 stoichiometry for the theoretical SVEP1/TIE1/ANG1-2 complex. Are there quantitative estimates of the confidence that can be provided? Did the authors attempt to model this complex with different stoichiometries? It is difficult to know how relevant this model is without any experimental results supporting this result.

      Since 1:1:1 is the smallest possible triple complex, it is our starting point. We can model a 2:2:2 version, but anything larger than this AlphaFold will not run. Furthermore, we now provide quantitative estimates of confidence with the pLDDT, PAE, pTM, ipTM scores for all models including the 2:2:2 complexes.

      Minor comments:

      (1) The authors could consider including a reference to alphafold on line 102.

      We have added a reference for AlphaFold2 and 3

      (2) The authors should refer to surface plasmon resonance (SPR) assays by this term as opposed to using the brand name Biacore.

      We agree with the reviewer and have changed the term Biacore to SPR.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the reviewers for their time and for their valuable inputs throughout the review process. We wish to clarify, one final time, the primary scope and empirical grounding of our work for prospective readers.

      Our study was designed to evaluate whether microsaccades track (in a correlative manner) covert visual-spatial attentional shifting, maintenance, or both. We did so within a single dedicated paradigm, across a large sample (N = 48 human participants). Despite remaining criticisms concerning per-participant event counts and microsaccade classification criteria, the key observation remains a striking dissociation (of the link between microsaccades and covert attention) during the initial shifting and the subsequent maintenance of visual-spatial attention. Moreover, we note how the robust effect observed during shifting (but not maintaining) attention, mitigates residual concerns regarding microsaccade sparsity or signal-to-noise ratio.

      We thus remain confident in the empirical foundation of our work and we invite readers to examine the full paper, supplementary materials, and open-access data to evaluate these findings independently.


      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the editors for their careful evaluation of our article and for their valuable input. Building on these suggestions, we were able to further corroborate our main conclusions, make our article more comprehensive, and thereby substantially strengthen the manuscript.

      We have one additional point of our own: we noticed that in our original submission, we had smoothed the data more than intended. Having caught this, we have now reduced the smoothing employed by 2.5 times compared to the original amount of smoothing (the exact smoothing values have also been added to the methods section). Importantly, however, while this has affected how the results look, this has not affected any of our original conclusions.

      General summary

      We would like to first respond to the major points brought forward by both the editorial summary and the public reviews. As we understand, the two main points that were raised regard: (1) the novelty and theoretical importance of our work and (2) the (in)completeness of our results. We start by providing our response to both of these main points below.

      Novelty and theoretical relevance of the work

      Regarding the novelty of our work, we believe the reviews and, by extension, the editorial summary underappreciated the main theoretical value of the question we addressed. Our work set out to investigate whether microsaccades track covert attentional shifting, attentional maintenance, or both. We fully recognise that there are ample prior studies that investigated and reported a link between microsaccades and covert attention, but also underscore how other studies report seemingly contradicting evidence by reporting that there is no such link. One such example is a recent paper by Willett & Mayo in PNAS (2023). Prompted by the recent hypothesis that this seemingly conflicting evidence may be due to prior work investigating attention ‘in different stages’ (van Ede, PNAS, 2023), we set out to address precisely this using a dedicated task that we designed for this purpose. As acknowledged by the summary and public reviews, this helps to reconcile seemingly opposing views in the literature. In our view, such reconciliation has substantial theoretical value.

      While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’). To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. 

      Finally, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Completeness of results

      Regarding the completeness of our results, the editorial summary points to “the absence of independent measures, single-trial analyses, and neutral-condition controls needed to substantiate the central claims”. While the raised points are valuable, they pertain to issues that are tangential to our primary question and stem from unfortunate misunderstandings of key analytical choices, as we now better clarify. We consider our results complete and comprehensive with regards to the main question our studies set out to answer.

      First, regarding the portrayed “need” for independent measures to define the ‘shift window’ of interest, we clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research. 

      Second, regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the single-trial level in our discussion, and prompted by the valuable feedback, we have expanded on this important contextualisation as part of our revision.

      Finally, regarding the portrayed “need” for a neural-attention control condition, we agree that inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ also at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that challenges our main conclusions. Having clarified this, we embraced this valuable reflection and revised the article to ensure that we do not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation, but refer to ‘the presence of an attentional modulation’ instead.

      Taken together, the explicit design and analysis choices that we made align with the theoretical aims of our study, and the central question we set out to address. The raised points are valuable and we are grateful to have been able to leverage them to improve our article, but we hope to have also clarified how they do not render our findings “incomplete” (as currently portrayed) with regards to the key goal of our article.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining the relationship between microsaccades and covert attention. This question has been widely investigated, with numerous studies showing that during sustained fixation, when subjects covertly attend to a peripheral stimulus, microsaccades tend to be biased toward the attended location. Here, the authors ask whether this microsaccade bias reflects a shift of covert attention or the maintenance of covert attention. They conclude that the bias is primarily driven by attention shifts, a finding that also helps reconcile the seemingly conflicting results of prior research, where the bias was questioned in paradigms that largely involved attention maintenance rather than shifting.

      Strengths:

      The paradigm and conclusions appear sound and supported by the results. A large sample size was used.

      We thank the reviewer for this clear and supportive summary of our work.

      Weaknesses:

      Weaknesses are mostly related to how the authors enforced fixation in the task, and clarifications are needed regarding some methodological details. A more direct comparison of the effect in the two experimental conditions is missing.

      We thank the reviewer for raising these valuable points. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”). Regarding the fixation points, we will address them in our point-by-point replies below.

      Reviewer #2 (Public review):

      Summary:

      This study aims to test the hypothesis that microsaccades are linked to the shifting of spatial attention, rather than the maintenance of attention at the cued location. In two experiments, participants were required to judge an orientation change at either a validly cued location (80% of the time) or an invalidly cued location (20% of the time). This change was presented at varying intervals (ranging from 500 to 3,200 ms) after cue onset. Accuracy and reaction times both showed attentional benefits at the valid versus invalid location across the different cue-target intervals. In contrast, microsaccade biases were time-dependent. The authors report a directional bias primarily observed around 400 ms after the cue, with later intervals (particularly in Experiment 2) exhibiting no biases in microsaccade direction towards the cued location. The authors argue that this finding supports their initial hypothesis that microsaccade biases reflect shifts in attention, but that maintaining attention at the cued location after an attention shift is not correlated with microsaccade direction.

      Strengths:

      The results are straightforward given the chosen experimental design. The manuscript is clearly written, and the presentation of the study and its visualisations are both of a high standard.

      We thank the reviewer for this clear summary of our work.

      Weaknesses:

      The major weakness of this paper is its incremental contribution to a widely studied phenomenon. The link between attention and microsaccades has been the subject of extensive research over the past two decades. This study merely provides a limited overview of the key insights gained from these papers and discussions. In fact, it attempts to summarise previous work by stating that many experiments found a link, while others did not, and provides only a relatively small number of references. To make a significant contribution, I believe the authors should evaluate the field more thoroughly, rather than merely scratching the surface.

      We thank the reviewer for this valuable reflection. For an elaborate response to the perceived novelty, please see our general summary reply above. In addition, we have added a more thorough evaluation of the field to the introduction (page 2, find relevant paragraph below). We hope that this will provide more context for the manuscript and strengthen its contribution to the field.

      Revised paragraph from introduction:

      “This link between microsaccades and covert visual-spatial attention has been demonstrated repeatedly. Early studies linked the direction of microsaccades to the deployment of covert attention [14, 15] and these findings were later replicated and extended. For example, it has been demonstrated in both humans [14–25] and non-human primates [26, 27]; during both externally directed perceptual attention [14, 15, 17, 20, 21, 24–27] and internally directed attention within visual working memory [16, 18, 19, 22, 23]; and in both perception and action tasks following directional cues [28]. Several studies have further linked the directional microsaccade bias to task performance [18, 21, 25, 28, 29]. For example, following spontaneous microsaccades, perception of visual targets presented in the same direction is better [25], and visual discrimination benefits may start already prior to microsaccade execution [21]. Recent evidence further suggests that microsaccades may even play a causal role in shaping the perception of peripheral stimuli [30].”

      The authors then present a potential solution to the conflicting past findings, arguing that attention should be considered a dynamic process that can be broken down into an attention shift and a sustained attention phase. Although the authors present this as a novel concept, I cannot think of anyone in the field who considers spatial attention to be a static entity. Nevertheless, I was curious to see how the authors would attempt to determine the precise timing of the attention shift and manipulate the different stages individually. However, the authors only varied the interval between the onset of the attention cue and the test stimulus, failing to further pinpoint their dynamic attention concept.

      The current version of the experiment, therefore, takes a correlational approach, similar to initial studies by Engbert and Kliegl (2003) and Hafed and Clark (2002). Meanwhile, we have learned a great deal about the link between microsaccades and attention. Below, I will list just a few of these findings to demonstrate how much we already know. It is important to note that, while the present study cites some of these papers, it does not provide a clear overview of how the current study goes beyond previous research.

      (1) Yuval-Greenberg and colleagues (2014) presented stimuli contingent on online-detected microsaccades. A postcue indicated the target for a visual task, and the target could be congruent or incongruent with the microsaccade direction. The authors showed higher visual accuracy in congruent trials. The authors cited that paper, but it is still important to emphasize how this study already tried to go beyond purely correlational links on a single trial level.

      (2) The Desimone lab (Lower et al., 2018) showed that firing rates in monkey V4 and IT were increased when a microsaccade was generated in the direction of the attended target.

      (3) However, attention can modulate responses in the superior colliculus even in the absence of microsaccades (Yu et al., 2022)

      (4) Similarly, Poletti, Rucci & Carrasco (2017) observed attentional modulations in the absence of microsaccades, or comparable attention effects irrespective of whether a microsaccade occurred or not (Roberts & Carrasco, 2019).

      Thus, in light of these insights, I believe the current study only adds incrementally to our understanding of the link between microsaccades and spatial attention.

      We thank the reviewer for this insightful comment, and for pointing out several important studies on this topic. While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’).

      To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. Moreover, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly and appropriately address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the singletrial level in our discussion. Prompted by the reviewer’s valuable feedback, we have expanded on this important contextualisation in our discussion section (page 8: “Therefore, even if microsaccades may more reliably track shifting than maintaining attention, as our current findings show, our findings should not be taken as evidence that microsaccades reliably track attentional shifts at the single-trial level.”). We also incorporated the valuable reference suggestions in our revised manuscript, including in the revised paragraph in our introduction where we provide a more extensive overview of the prior literature, as shown in response to the preceding comment and in the discussion where we discuss the relationship between microsaccades and attention on a single-trial level.

      In general, it is important to have an independent measure of the dynamics of an attention shift. I think a shift of 200-600 ms is quite long, and defining this interval is rather arbitrary. Why consider such a long delay as the shift? Rather than taking a data-driven approach to defining an interval for an attention shift, it would be more convincing to derive an interval of interest based on past research or an independent measure.

      We thank the reviewer for their question. We wish to clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research.

      The present analyses report microsaccade statistics across all trials, but do not directly link single-trial microsaccades to accuracy. Similarly, reaction times and accuracy were analyzed only with respect to valid vs. invalid trials. Here, it would be important to link the findings between microsaccades and performance on a single-trial level. For instance, can the authors report reaction times and accuracy also separately for trials with vs. without microsaccades, and for trials with congruent vs. incongruent microsaccades?

      We thank the reviewer for their sincere interest in our findings and for the great suggestion of an additional analysis. We have now investigated whether trials with a congruent, incongruent or no microsaccade in the shift window (where congruent or incongruent was determined as based on the first microsaccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to stress that our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we have decided not to include these analyses.

      The study would benefit greatly from including a neutral condition to substantiate claims of attentional benefits and costs. It is highly probable that invalid trials would also demonstrate costs in terms of reaction times and accuracy. It would be interesting to observe whether directional biases in microsaccades are also evident when compared to a neutral condition.

      We thank the reviewer for this valuable reflection. We agree that the inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial-numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that in any way challenges our main conclusions.

      Prompted by this reflection, we have ensured to not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation throughout the article, but refer to this only as ‘the presence of an attentional modulation’ instead (such changes were made on pages 3 and 9, and we kept this phrasing consistent in our additions on pages 5 and 22).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The results resemble recent findings by Brandolani et al. (2025), who also showed that the microsaccade bias-in a similar task and using a comparable analysis approach-was restricted to a narrow time window, primarily around the time of the attention shift. The authors should reference this work, discuss similarities and differences, and, given their larger sample size, consider whether they observe a similar correlation between response times and microsaccade rate.

      We thank the reviewer for pointing out this useful reference to us. We have included this article in our introduction, when sketching the current state of the field (page 2) and also mentioned this related article in the discussion (page 8).

      In addition, prompted by this comment, we have investigated whether we observe a similar correlation between response times and microsaccade rate in valid trials. We have investigated this for both possible timeframes: (1) the ‘shift’ timeframe and (2) the ‘maintain’ timeframe. For both experiments, there was no consistent correlation between response times and overall microsaccade rate, as shown in Author response image 1.

      Author response image 1.

      Relationship between the reaction time and the overall saccade rate. This figure shows the relationship between the average reaction time and the average overall saccade rate for both the ‘shift’ period (from 200 to 600 ms after cue onset) and the ‘maintain’ period (from 600 to 1400 ms after cue onset). Each dot represents one participant. Throughout the entire figure, the following significance levels were used: *: p< 0.05, **: p < 0.01, ***: p < 0.001, ****: p < 0.0001.

      (2) I could not find information on the average number of trials per condition and the average number of microsaccades per subject per condition. Ideally, these numbers should be reported (e.g., in a supplementary table). Since the analysis is based on microsaccade direction, knowing how many microsaccades each subject contributed per condition is critical. Microsaccade rates vary substantially across individuals, and subjects with very few events may add noise to the analysis, as proportions of toward/away microsaccades become unreliable.

      We thank the reviewer for pointing out that this relevant information was missing. We have now included these numbers in a supplementary table as suggested (page 17).

      (3) Relatedly, it was unclear whether the time-course analyses were based on collapsing all microsaccade events across subjects or on subject-level averages. In the Methods, the authors state that "the permutation distribution of the largest cluster size was acquired by randomly permuting the trial-average data at the group level 10,000 times," but this is ambiguous. Please clarify.

      We thank the reviewer for pointing out this ambiguity. We have changed the methods section to reflect more clearly that we first obtain time-courses of the microsaccade rate per participant, and subject these time courses to second level statistics using cluster-based permutation analysis (page 13). We have also reworded the sentence you quoted to remove any ambiguity (page 13): “A permutation distribution of the largest cluster size was acquired by randomly permuting the condition labels of each participant’s trial-averaged time course data (i.e. randomly flipping the sign of the difference in rate of toward vs. away saccades) 10,000 times and identifying the size of the largest clusters observed in these randomised data after each permutation.”

      (4) The authors analyze only downward microsaccades, but the cutoff definition is not specified. Presumably, this includes all directions between 180{degree sign} and 360{degree sign}, which may also include nearly horizontal events. This should be clearly stated. In addition, the rate of upward microsaccades should still be shown, divided into up-left and up-right quadrants to parallel the toward/away analysis. This would provide informative context on whether upward microsaccade rates change systematically over time.

      We thank the reviewer for pointing out that this information was missing, and for suggesting this valuable additional analysis. In the methods section, we now explicitly state the angular cutoffs used for the main analysis (page 12). Additionally, we have added a supplementary figure that shows the time course of upwards microsaccades over time (page 21), please see Supplementary Figure S4.

      (5) If the dataset contains enough microsaccadic events per subject, it would be useful to test more conservative angular cutoffs for defining "toward" versus "away."

      We thank the reviewer for this insightful suggestion. We have included an additional analysis, where only microsaccades were included with a direction within a 45° angle around the exacttoward and exact-away directions. This replicated our main finding. The results from this analysis are now included in the supplementary materials (page 24), please see Supplementary Figure S8.

      (6) Figure 2C: It is unclear what the "Center" and "Border" lines represent. The figure is also potentially confusing because it shows microsaccade amplitudes rather than landing positions. Small amplitudes may still bring gaze close to the target; this distinction should be clarified.

      We thank the reviewer for pointing this out. We have changed the “centre” and “border” labels to include more information (pages 6 and 19). We have also added an in-text clarification of the distinction between saccade amplitude and landing position (page 5: “Note that Figure 2C does not show saccade landing positions. While it is theoretically possible for multiple small unidirectional saccades to lead to a larger change in gaze position, a complementary analysis shows that fixation was maintained during the period of peak microsaccade rate in both experiments (see Supplementary Figure S6).”). In addition, in response to the related comment below, we have now also added heatmaps of gaze showing that gaze overall remained close to fixation in our tasks.

      (7) From the Methods, it appears that in Experiment 1, there was no automatic criterion for discarding trials in which gaze deviated from fixation. In Experiment 2, trials were terminated if gaze left a 2{degree sign} window, but given that the target was only 5{degree sign} from fixation, this seems a relatively loose criterion. It would be important to show the distribution of gaze positions during the task to assess whether fixation control was adequate.

      We thank the reviewer for this great suggestion. We have now added a figure to the supplementary materials (page 23) that shows the probability density of gaze position throughout the ‘shift’ period, for left cued trials and right cued trials separately, please see Supplementary Figure S6. We hope that this will further show that even in Experiment 1, fixational control was successful. We also show the difference between left cued and right cued trials, which again shows a gaze bias towards the cued item.

      (8) Did the authors examine whether there was a response time benefit (e.g., RT in congruent microsaccade trials minus RT in incongruent microsaccade trials, as in Brandolani et al., 2025) or an accuracy benefit when microsaccades were directed toward the target?

      We thank the reviewer for their interest in our findings and for the great suggestion of an additional analysis. As we discussed also in response to the general summary from reviewer #2 above, we have now investigated whether trials with a congruent, incongruent or no saccade in the shift window (where congruent or incongruent was determined as based on the first saccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to again stress how our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we decided not to include these analyses. However, please note that we did now include the outcomes of another analysis that more directly targeted the relation between the spatial modulations in microsaccades and task performance, as we turn to below. 

      (9) Was there a relationship between the size of the attentional effect and the magnitude of the microsaccade bias?

      We thank the reviewer also for this insightful question. We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (10) The criteria for minimum microsaccade amplitude and duration are not specified. This should be clarified. I recommend excluding events smaller than ~5 arcmin, as these are likely noise-especially since eye tracking was monocular. Monocular "microsaccades" can be spurious, but this can be determined only with binocular tracking. It is also unclear whether subjects used chin/head rests. A main-sequence plot in the supplementary material would be helpful.

      We thank the reviewer for pointing this out. We have included a main-sequence plot in the supplementary materials (page 23). The main-sequence plot can also be found in Supplementary Figure S7 and suggests that our saccade-detection algorithm worked well with detected saccades following the main sequence. We have also stated more clearly in the methods section that subjects used a chinrest (page 11).

      (11) Please specify the asterisk convention in figure captions (i.e., what * vs. ** vs. *** correspond to in terms of p-values).

      We thank the reviewer for pointing out that these significance levels were not mentioned in every figure caption, so we have added this information to every figure caption where they were missing (page 4, 6, 7 and 19).

      (12) The fact that stricter fixation criteria reduced the size of the effect suggests the possibility that gaze drift toward the target might have conferred an eccentricity advantage in this discrimination task. A direct comparison of the effect in the two experiments would be valuable. The authors should comment on this. It would be informative to plot the average gaze position around the time of peak microsaccade rate in both experiments. Reanalyzing the data post hoc with a stricter trial-selection criterion (e.g., excluding trials where gaze deviated more than 1{degree sign} from fixation) could also be very valuable, as it would systematically test how fixation control influences the observed microsaccade-attention relationship. This would be informative for the community studying this topic.

      We thank the reviewer for these valuable reflections. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

      Regarding the average gaze position around the time of peak microsaccade rate: in response to reviewer #1, under point (7), we have included Supplementary Figure S6 that shows the probability distribution of gaze position throughout the ‘shift’ period (the same figure is found on page 23 in the article), which shows that fixational control was successful in both experiments. This period is also the period of peak microsaccade rate in both experiments.

      We wholeheartedly agree that systematically investigating the effect of fixational control is important for the field as a whole, and this is also precisely why we set out to perform the same experiment in two different experimental settings with regards to fixational control, and why we decided to include the results from both experiment variants side-by-side in our article.

      Reviewer #2 (Recommendations for the authors):

      In addition to my general concerns in the public review, I have the following recommendations.

      (1) Did the authors distinguish between the initial and subsequent microsaccades during their analysis? Is it possible to produce multiple microsaccades when shifting attention, or do the authors only consider the first microsaccade to be linked to an attention shift?

      We thank the reviewer for pointing out this ambiguity. We have now stated more clearly in the methods section that we consider all microsaccades for our analyses (page 12: “Crucially, we did not restrict our analyses to initial saccades; rather, all detected saccades were included. This allowed us to examine oculomotor behaviour during later trial phases, where initial saccades are unlikely to occur.”). We also believe this methodological choice is important, as otherwise it would be conceivable that no microsaccade bias can be found during the ‘sustain’ period, simply because no ‘first’ microsaccades occur anymore.

      (2) Two microsaccades had to be separated by at least 100 ms. This is an unusually long delay.

      Could the authors please specify how many microsaccades were discarded using this criterion?

      This inter-saccade-interval is quite large on purpose, as we want to minimise the probability of counting the same microsaccade twice. We have re-analysed the data with a minimum ISI of 50 ms, and this led to an increase of found saccades of a, respectively, 5.1% and 1.9% increase for Experiments 1 and 2. However, two participants in Experiment 1 led to a much higher increase in saccades than all other participants (these participants had z-scores of 3.9 and 2.4 for the number of additionally found saccades with an ISI of 50 ms; all other z-scores for Experiment 1 were between -0.5 and 0.5). When those two participants were removed, in Experiment 1 only 1.6% more saccades were found.

      (3) If I understand correctly, the authors did not use staircase procedures to eliminate differences in task difficulty between participants. Could the authors demonstrate how task difficulty relates to the link between microsaccades and performance? For example, is the time course of the microsaccade direction bias correlated with performance?

      We thank the reviewer for this suggestion (that overlaps with a comment of Reviewer 1). We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (4) The authors reported using equiluminant stimuli. Could the authors please specify the exact luminance?

      We thank the reviewer for pointing out this missing information. We have now included this information in the methods section (page 11: “, with a luminance of 88.5 cd/m<sup>2</sup>.”). We have also included the luminance of the background (page 11: “luminance: 29.0 cd/m<sup>2</sup>”).

      (5) Could the authors please provide a full polar plot showing all microsaccade directions, and colour-code those included in the analysis?

      We thank the reviewer for this great suggestion on how to present our results even more clearly and comprehensively. We have included a supplementary figure showing the full polar histograms (with 20 radial bins), for all three timeframes of interest: the whole trial, the ‘shift’ period and the ‘maintain’ period (page 20). As requested, the saccades included in the main analyses are colour-coded. See Supplementary Figure S3

      (6) Can the authors please directly compare the main effects between experiment 1 and experiment 2 (Figure 2B)?

      We thank the reviewer for this great suggestion (that was also made by reviewer 1). We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

    1. Reviewer #2 (Public review):

      This manuscript describes the fascinating phenomenon of growth hormone (GH)-independent growth occurring in the mother during pregnancy. This growth was most pronounced in dwarf mice that are lacking the receptor for growth hormone-releasing hormone (GHRH) and therefore showing isolated GH deficiency. However, the pregnancy-induced growth could also be observed in wild-type mice, suggesting that it is a normal part of the maternal adaptation to pregnancy. The study falls short of identifying the mechanism(s) driving this pregnancy-induced growth response, but it certainly reveals a novel insight into maternal physiology. The authors have completed a range of experiments in mice to prove that, as well as being GH independent, the pregnancy-induced growth also did not require GH signaling in the liver (i.e. not another pregnancy-specific ligand operating through the GHR to promote IGF). They also provided complementary data from a population of humans with untreated isolated GH deficiency that are broadly consistent with the hypothesis. While it is important to consider the significant species differences between rodents and humans, both in terms of growth physiology and also in terms of evolution of placental somato-mammotrophic hormones, this unique population are a valuable resource and adds credence to the study. Overall, I find this a compelling research story, but disappointingly unfinished. There are some areas where additional information could improve the ability to interpret the data, and some additional concepts that could be considered in the discussion. There are also areas where additional experiments might provide important insights. However, I think that such suggestions can be considered as appropriate for future research, rather than delaying consideration of the current manuscript.

      Main comments:

      (1) Data in Figure 1 are remarkable - not so much the growth in pregnancy in the wildtype mice, because while elevated GH is well known in pregnancy, but growth in the dwarf mice is indicative of GH-independent growth. From these data, it seems that there is good evidence that growth in pregnancy is an adaptive function. However, it is possible that growth is achieved in dwarf mice and that in wildtype mice may have been mediated through different mechanisms. The dwarf mice showed an increase in liver and plasma IGF1, suggestive of an additional ligand driving IGF in pregnancy. One could hypothesize that such an effect could be mediated by an additional pregnancy-specific ligand activating the GH receptor. In humans, placental growth hormone could be such a ligand, but as far as we know, there is no placental GH in mice. In contrast, the wildtype animals showed suppression of liver and circulating IGF1, and low levels of pSTAT5 in the liver during pregnancy. These data (in Figure 5) are very surprising. Given the high circulating GH in pregnancy, as well as high placental lactogen (which would be expected to activate STAT5 in the liver through the Prlr), the low levels of pSTAT5 are unexpected and would seem to indicate some sort of acquired insensitivity to GH. Is this entirely driven by down-regulation of STAT5b protein, or could there be activation of other, negative regulators of STAT signalling, such as SOCS? What is causing such a profound suppression of STAT5? Regardless of the mechanism, this suggests that pregnancy-induced growth in wildtype mice is independent of circulating IGF1 (potentially a different mechanism or in addition to that seen in IGHD mice).

      The data shown in Figure 6 are a major strength of the study, showing that the pregnancy-induced changes are not specific to one particular transgenic model, but still occur in a variety of models affecting GH through different approaches. Given the pregnancy-specific nature of the changes, however, it seems an oversight not to have evaluated the role of placental lactogens. Prlr is highly expressed in the liver, but the function of this hormone in the liver is not well established. Could the extremely high levels of PL be mediating this growth response? Given the low expression of STAT5 in the liver and the fact that plasma IGF1 is not markedly elevated, it seems more likely that this growth response may be mediated by locally produced IGF1 in target tissues.

      I think these possibilities could be addressed by an expanded discussion of species variation in placental hormones, to highlight that humans have expansion of the GH locus, but rodents have expansion of the prolactin axis (see Soares, M. J. The prolactin and growth hormone families: pregnancy-specific hormones/cytokines at the maternal-fetal interface. Reprod Biol Endocrinol 2, 51, 2004). Importantly, placental GH and chorionic somatomammotropins (CSM) in humans are all variants of the GH gene, but CSM have preferential activity at Prlr. This seems to be a fundamental species difference in pregnancy biology, but has been interpreted as an example of convergent evolution, with conservation of prolactin and GH-like functions at the maternal-fetal interface, mediated by different mechanisms, likely contributing to the metabolic adaptations of the mother (see Newbern D, Freemark M. Placental hormones and the control of maternal metabolism and fetal growth. Curr Opin Endocrinol Diabetes Obes. 2011; 18: 409-416). While the preceding function has focused on explaining the evolution of placental lactogens (either prolactin or GH variants), the present data suggest that there are also mechanisms to maintain growth in pregnancy, independent of GH (even in the absence of a placental GH).

      (2) The human data are very interesting, and my initial impression was that it seemed unlikely to be the same phenomenon. Was there any real evidence for "growth" in pregnancy? Pubertal maturation of long bone growth might be expected to prevent further growth in adulthood. However, these issues were appropriately discussed, and it seems well justified to evaluate this unique population of women with IGHD who underwent pregnancy. It would be very interesting to know if these women experienced elevated IGF1 during pregnancy, indicative of placental GH contributing to growth. Mechanistically, this might be more like the dwarf mouse situation of IGHD, that the situation in wildtype mice (associated with liver insensitivity to GH and low IGF1).

      (3) It would be useful to include investigations that isolate the effects of pregnancy and the placental hormones. Such studies could include evaluating growth in pseudopregnant mice with IGHD (pregnancy-like changes in hormones but lacking the placental contribution) and in IGHD animals that experience pregnancy but not lactation (pups removed at birth). I accept that this might be too large an additional study to add for the present manuscript.

      (4) It is an important and translationally relevant observation that pregnancy increased the risk of long-term weight gain, and that after the first pregnancy, the pregnancy-induced growth response was more directed to promoting fat deposition. Does this provide any mechanistic insight? Could a metabolic adaptation result in growth?

    1. Author Response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study reveals distinct representations of task-related information in the dendrites and somata of cortical neurons during sensorimotor learning and behavioral adaptation. The evidence is compelling, combining simultaneous imaging of dendritic and somatic activity during behavior to demonstrate compartment-specific encoding of sensory cues, motor actions, and corrective signals. The work will be of broad interest to neuroscientists studying dendritic computation, motor learning, and the cellular mechanisms underlying adaptive behavior.

      Thank you for this excellent summary. We recommend one change: removing the word “simultaneous”. It could perhaps be replaced with “concurrent” or simply omitted. Tuft dendrites and somata were imaged on alternating days, and most readers will probably interpret “simultaneous” as implying a faster, interleaved sampling rate.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Scheib et al. identify distinct calcium dynamics in the somata and tuft dendrites of layer 5 pyramidal cells in mice performing a licking task. Animals are trained to lick water ports on the left or right following an acoustic cue, and can adjust their targeting when the ports are displaced. For tongue premotor cortical neurons projecting to the ventromedial thalamus, calcium transients in tuft dendrites are tightly locked to the direction-instructive cue, while somatic calcium signals are more broadly dispersed and more frequently synchronized with tongue motion and port contact. Finally, when the targets are shifted, tufts exhibit a sparse but large corrective signal on an improperly-targeted first lick, and the changes in population activity in the tufts and somata differ after adaptation to the new port locations.

      Strengths:

      (1.1) In my opinion, this is a very strong manuscript which reports several novel and significant observations, contains high-quality data and (for the most part) reasonable analyses, and is clear and well-written. Most prior studies of cortical sensorimotor processing have measured the output of neurons using extracellular recording - an approach which obscures potentially important signaling differences between neuronal compartments. This study leverages cutting-edge imaging techniques in mice to document large, time-dependent differences between calcium signals at cortical somata and tuft dendrites. This phenomenon could have major implications at the cellular level for synaptic plasticity, and at the systems and behavioral levels for motor adaptation. As described below, I have only one major technical concern (which should be addressable with additional analysis), along with several relatively minor suggestions for improving the manuscript.

      We thank the reviewer for their insightful summary of the significance of the differences that we identified in the task-related activity of tuft dendrites and somata.

      Weaknesses:

      (1.2) At a conceptual level, the authors may wish to elaborate a bit on what sensorimotor computation they think the circuit is implementing, and how their results help explain this implementation. Several possibilities are raised: tuft activation could "prime" the pyramidal cells in advance of movement initiation (line 319ff), or could track errors to engage plasticity (line 351ff) and solve the credit assignment problem (line 362ff). It might be helpful to make one of these proposals more concrete with a computational model, but this is not strictly necessary.

      We thank the reviewer for this feedback. We absolutely agree that detailed computational models of each proposed computation will be very valuable and constitute an important follow-up to this work. We hope to collaborate with theorists to take that next step. Each possible computation noted by the reviewer reflects distinct differences that we observed in the task-related activity of tuft dendrites and somata. They are not mutually exclusive hypotheses to explain the same phenomenon. As such, we think they are best addressed independently in future modeling work. By making the data and a concise description of the main findings available immediately, we hope to allow computational experts in each of these areas to take advantage of the results of this study without delay.

      (1.3) My only major technical concern relates to the analyses in Figures 4F-H, 5G-I, and 6H-K (c.f. equations 2-5). Typically, one identifies population-level factors by projecting neural activity onto fixed dimensions of interest; this makes it possible to see how activity evolves over time along interpretable coordinates. Here, however, the coding directions are redefined at each time point, so the "choice" activity at time t is actually a different signal from the "choice" activity at t+1. This procedure is a bit like comparing the activity of one neuron at one time point with the activity of a different neuron at a later time point. It also makes the physiological interpretation more complicated: if the dimensions are fixed, one can see how a downstream neuron could "read out" the signal by computing a weighted sum of the activity of upstream neurons, but it is harder to see how this could happen if the weights are always rotating.

      We thank the reviewer for raising this point. We agree that our use of projections along coding directions (CDs) defined at each time point is a less conventional use of coding directions, although nearly identical calculations have been previously used to assess population-level selectivity and code stability in this task (Chen et al., 2017; Yang et al., 2022). As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task-dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task-variable. Finally, because the primary goal of the study was to identify any differences between tuft dendrite and somatic encoding, we think that calculating the population selectivity at each timepoint gives readers a less biased view of the selectivity of the two compartments, whereas calculating a CD over a single arbitrary time window could conflate differences in dynamics with differences in selectivity.

      We agree that calculating the CD at each timepoint makes it hard to see where the code is stable and where it is rotating, and thus how a downstream neuron might “read out” the signal. To provide this information, we have added new panels to the supplement showing the correlation of CDs across time (Figure 4 - figure supplement 1B,D). We also now provide this information for CR-CA in Figure 5—figure supplement 2A (the plots previously presented in 2A were the correlations of CR with CA, rather than CR-CA; an error that has been fixed). The following changes were also made to the Results section to clarify this issue:

      “From the linear model, we calculated coding directions (CDs) at each timepoint that maximally separated Stimulus, Choice, and Outcome activity (Figure 4F; Figure 4—figure supplement 1A) and estimated the direction and selectivity along each dimension across time (Figure 4—figure supplement 1B-E; see Methods). Allowing CDs to rotate in time (see Figure 4—figure supplement 1B,D), although unconventional, ensured that comparisons of population selectivity across the two compartments were not biased by the selection of an arbitrary CD time window.”

      We also identified a mistake in the description of CD orthogonalization in the Methods, which has been corrected as follows:

      “For each timepoint, each selectivity CD was then orthogonalized with respect to the other two selectivity CDs by a QR decomposition in which that selectivity CD was last in the order.”

      (1.4) A few comments on the behavioral task and results. After the port shift, the error rate is quite high, and doesn't diminish much between the early and late epochs (approximately 42% and 38% error rate, respectively; Figure 1I). That is, mice do not seem to fully master the task. Clearly, animals do alter their aim, but even this does not seem to change much between early and late periods (Figure 1J). I recommend that the authors show the behavioral data at a finer level of granularity (e.g., by plotting the change in exit trajectory on all individual trials across sessions, with a loess fit) to allow an assessment of the adaptation rate and when adaptation saturates. It would also be more conventional to refer to the behavioral changes as "motor adaptation," instead of "skill learning." (The latter would be appropriate if the port offset were randomized across trials, and animals received two separate cues for direction and offset, but I suspect this task would be too difficult for mice to learn.)

      We agree with the reviewer that by the end of the late period, performance on the right side (Figure 1I) has still not returned to pre-shift levels. This may reflect mice not fully mastering the task, as the reviewer suggests, or it may reflect that after the shift, the right port is substantially more difficult to reach than the left port. Unfortunately, because of high animal-to-animal and lick-to-lick variability, plotting the post-shift lick angle at a finer level of granularity is not statistically informative.

      With regard to the nature of the learning in our task, we selected “skill learning” as the best description of the motor learning task based on distinctions between adaptation and the learning of motor skills by Krakauer et al., 2019 and Heald et al., 2021. Conceptually, the difference is whether an existing motor controller memory is simply updated with new parameters, or whether the motor context has changed sufficiently that a distinct motor controller memory (which can still use parts of previous memories) is formed. In our task, after the port shift the left port forms an obstacle to reaching the right port. This obstacle was simply not present before the shift. Before the shift, ports were approximately equidistant from the mouth and easily avoided given the port separation and tongue width. Thus, avoiding an obstacle would presumably not be part of the initial motor controller memory and a distinct memory would need to be constructed.

      We agree with the reviewer, however, that given that we do not have fine-timescale dynamics of behavioral changes in response to the shift, and did not conduct other experiments (such as returning the ports to their original location) that would typically be conducted to identify “adaptation-like” or “skill learning-like” dynamics, we cannot empirically distinguish between the two. We now clarify in “Study limitations” that we call the studied behavior “skill learning” based on the nature of the task, but that our behavioral analysis cannot distinguish between adaptation and skill learning:

      “We refer to the behavioral paradigm as motor “skill learning” strictly based on the nature of the task. After the shift, mice must avoid a new obstacle close to the mouth (i.e., the left port), which we assume requires the formation of a distinct motor controller memory and therefore would be considered skill learning (Krakauer et al., 2019). However, we did not confirm that the mice exhibited specific behavioral characteristics of skill learning and it is possible that other kinds of motor learning (e.g., motor adaptation) were dominant.”

      (1.5) This is perhaps a semantic point, but it might not be entirely accurate to refer to the activity evoked by the directional cue as "sensory." Typically, a "sensory" response should encode some feature of a stimulus - in this case, the frequency of a tone. Here, it seems likely that the cue-aligned activity reflects the instructed lick direction, rather than the auditory information per se. (Presumably, these premotor neurons do not have well-behaved auditory tuning curves.) By comparison, in macaques performing center-out reach tasks, activity in dorsal premotor cortex rapidly ramps up following a visual cue instructing the direction of an upcoming reach, but one usually wouldn't refer to this activity as "visual" or "sensory" (though this is sometimes done). I suggest the authors either use "Instruction" or similar (e.g., in Figure 4F), or clarify in the text whether they think the activity is a genuine auditory response or something else.

      We understand how this could cause confusion. “Sensory” was meant to denote the nature of the differences in external events between the trial types used to calculate selectivity, not to imply that the activity was necessarily selective for detailed features of the cues outside the context of the task. Previous work in ALM cortex has labeled this selectivity direction as “stimulus” (Yang et al., 2022; Chen et al., 2024) to better emphasize that it is simply defined by the external cue. Where appropriate, we have revised the article to use “stimulus” or “instructional cues” in place of “sensory” for clarity and to better conform with convention.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to compare functional encoding in the tuft dendrites and somata of a specific cortical cell type during motor planning and learning.

      Strengths:

      (2.1) The investigation of a specific projection type (L5 ET) is a strength that aids reproducibility and interpretation. The elegant approach to increasing the depth of field of dendritic imaging is another strength. The data analyses are largely clear in their methods, scope, and interpretation. The writing is extremely clear and appropriately referenced, with an excellent Introduction, in particular.

      We thank the reviewer for their appreciation of the study design, imaging methods, and scholarship of the article.

      Weaknesses:

      (2.2) It is not obvious whether the selected labeling strategy avoids labeling Layer 6 CT neurons, which would contaminate dendritic recordings. The images provided suggest enrichment in L5, but a discussion of this important potential caveat is warranted, especially since within-cell comparisons of apical dendrites to somata were not performed.

      We thank the reviewer for emphasizing the need to discuss this potential issue. For the following reasons, it is likely that the vast majority of dendrites we imaged in layer 1 originated from layer 5 ET neurons. First, as the reviewer notes, the provided images suggest enrichment in layer 5. This enrichment likely reflects the fact that most L6 CT neurons in motor and premotor cortex send denser projections to other thalamic nuclei than to VM thalamus (Winnebust et al., 2019, Cell), where we targeted our retrograde-Cre injections. Second, L6 CT neurons are predominantly untufted (Ledergerber and Larkum, 2010, J. Neurosci.), including in motor and premotor cortex (Peng et al., 2021, Nature; Ichikawa, 2025, Front. Neuroanat.). A recently identified subclass of L6 CT neurons in secondary motor cortex has dense projections to VM thalamus, but this class also appears to extend minimal dendrites into L1 (Li et al., 2024, bioRxiv). Nonetheless, we did not label post-hoc tissue collected from imaged mice with markers of precise laminar boundaries, and thus cannot definitively rule out the possibility that dendrites from a subclass of L6 CT neurons with tuft dendrites were also imaged. We have added the following paragraph to the “Study limitations” section to make readers aware of these issues:

      “L5 ET neurons in premotor cortex elaborate extensive tuft dendrites in L1, whereas Layer 6 (L6) corticothalamic (CT) neurons are predominantly untufted (Jiang et al., 2020; Peng et al., 2021). Thus, although we cannot rule out the possibility that dendrites from a subclass of L6 CT neurons were also sampled, it is likely that the vast majority of dendrites we recorded in L1 originated from L5 ET neurons.”

      (2.3) The application of DeepInterpolation to dendritic data appears to be novel, and little detail or vetting is provided. The reader is left guessing: Was the model retrained or fine-tuned on dendritic data? How does the denoising affect the resulting segmentation and activity traces? Is denoising necessary for this workflow?

      We thank the reviewer for requesting this useful additional information.

      In all cases, the model was retrained for each dendritic or somatic imaging session. Denoising improved segmentation consistency, as measured by comparing segmentations of individual sessions from the same animal. This is now specified in the Methods as follows:

      “The DeepInterpolation model was trained on each imaging session prior to denoising of that session. Denoising prior to NMF-based segmentation resulted in more robust and consistent dendrite segmentation than NMF-based segmentation without prior denoising (0.79 +/- 0.01 ⍴ vs. 0.46 +/- 0.01 ⍴; mean of the max Spearman correlation of components across sessions; random subsample of N = 3 mice, 15 sessions, 400 components).”

      With regard to how denoising impacts activity traces, examples were shown in Figure 2I, K. To provide more quantitative information to the reader, we calculated estimates of the power and reliability of the spectral content of dendrite activity traces extracted with or without denoising. These data are now shown in the new panel, Figure 2 - figure supplement 2G. The power spectral density of the denoised activity and the estimated reliable power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C). Some frequencies beyond this point have been suppressed beyond what would be expected due to photon shot noise (as estimated by the replicate coherence-weighted PSD, or “recoverable” PSD). Further characterization of the precise nature of the suppressed high-frequency information – which could be suppressed artifacts (e.g., fast brain motion) or lost signal detail (i.e., GCaMP8m rise kinetics) – is beyond the scope of this paper.

      Details of the PSD calculations have been added to the Methods, and the following statement has been added to the Results: “Power spectral density of the denoised traces and the coherence-weighted power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G; Methods), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C).”

      (2.4) The activity patterns of the recorded cells appear to lack the characteristic ramping during the delay epoch previously reported in both calcium imaging and electrophysiology studies. Given that a major contribution to the significance of the work is to constrain models of ALM function, a discussion of how the data aligns with previous measurements in the same circuit would improve the work.

      Preparatory selectivity and ramping activity can be seen in Figure 3H, Figure 6I, and Figure 5 – figure supplement 1B. We note that in ALM cortex, the ramping mode explains a minority of the total variance (~17%, Yang et al., 2022), but it can appear particularly prominent in projections along certain fixed CDs.

      (2.5) It would be very informative to compare differences in signals between dendrites and somata of the same cells. Consistently tracing dendrites to their respective somata would assuage worries of potential contamination from dendrites of deeper cells and enable more direct comparisons of signal transformations between dendrites and somata. It would be good to understand the relationship between dendritic calcium signals and backpropagating action potentials in this task. The authors detect less frequent calcium events in tufts versus somata; is this due to selective backpropagation of action potentials? The dynamics of this process were recently investigated by Adam Cohen's group in vivo and in vitro, and measurements in the present settings could be compared to such work.

      We agree with the reviewer that being able to compare differences in signals between the dendrites and somata of the same cells would be very valuable. However, reliable tracing of tuft dendrites to somata from in vivo 2P anatomical imaging requires extremely sparse labeling, such that very few neurons are recorded per animal (Kerlin et al., 2019, eLife; Otor et al., 2022, Science). As stated in the “Study limitations” section of the Discussion, we suspected (correctly) that some task-related selectivity (i.e., selectivity for corrective action) would be sparsely represented in the dendrites, and thus adopted a labeling and image processing strategy that allowed us to record from many dendrites per animal. This strategy necessarily comes at the expense of generating a labeling density that precludes reliable tracing of tuft dendrites to their respective somata based on 2P morphology alone. As discussed in our response to reviewer comment 2.2 and a new paragraph of “Study limitations,” substantial contamination of the dendrite recordings by dendrites of L6 CT neurons is highly unlikely. Future studies could use simultaneous functional imaging across large volumes combined with activity-based segmentation or post-hoc high-resolution imaging of tissue sections registered to in vivo 2P imaging to accomplish both high-throughput dendritic imaging and reliable tracing.

      We thank the reviewer for pointing out that we could discuss selective backpropagation as a potential mechanism more explicitly. Our results are consistent with previous studies of L5 tufts in vivo (Francioni et al., 2019, eLife), including in ALM cortex (Maristany de las Casas et al., 2026, Science), that reported that rates of multi-branch calcium transients in the tuft dendrites of L5 neurons are lower than somatic spike rates. As discussed in “Study limitations,” there is not a clear approach in our data to determine the precise nature of the events underlying the calcium transients we measured in the tuft dendrites. Selective backpropagation of action potentials is certainly one possibility and we agree that recent research from Dr. Adam Cohen’s group should be discussed. We have added the following to the Discussion:

      “Based on previous calcium imaging of L5 tufts in ALM cortex of mice engaged in similar tasks (Kerlin et al., 2019; Maristany De Las Casas et al., 2026), we suspect that most of the activity we measured was coincident with global tuft or hemi-tree events, as well as somatic spiking. Recent in vivo voltage imaging in the hippocampus has also indicated that most spikes in distal dendrites start as bAPs that have been selectively amplified (Wu et al., 2026; Lee et al., 2026).”

      (2.6) The Coding Direction analyses presented in this work, while consistent with previous literature on population codes in ALM, are at odds with the nature of the measurements here. The changes in representation that occur between the dendrites and soma of an individual cell are probably best thought of in terms of the dynamics of signals themselves within individual neurons, rather than in the information encoded across a population.

      We thank the reviewer for giving us the opportunity to clarify this issue. As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task variable in the population activity. Thus, the analyses are not at odds with the nature of the measurements in the study.

      Nevertheless, it is true that by recalculating the CD at each timepoint, our selectivity projections do not provide the same information as conventional projections along a fixed CD, which can indicate where the selectivity code is stable and where it is changing. To provide this information we have added new panels to the supplement showing the correlation of selectivity CDs across time (Figure 4 - figure supplement 1B, D).

      (2.7) This work is largely observational, describing signals that might reflect computational transformations and/or instruct plasticity, but those possibilities have not yet been deeply investigated. The manuscript does a good job of laying out these as future directions.

      We agree with the reviewer. As noted by the reviewer in comment (2.1), we combined a number of approaches in an innovative manner to explore how tuft dendrite activity differs from somatic activity at the population level during motor learning. These measurements provide the necessary foundation for future mechanistic studies and we think it is appropriate to share them at this stage of investigation and in the format of this article.

      Reviewer #3 (Public review):

      Summary:

      This article by Scheib et al. investigates how layer 5 extratelencephalic (ET) neurons in the frontal cortex encode sensorimotor information during motor learning, focusing on differences between their apical tuft dendrites and somas. The authors alternated recordings among these ET neuronal compartments in the mouse anterior lateral motor cortex (ALM) during a cued directional licking task with a target port shift. They found that while tuft dendrites predominantly encode sensory cues, with a subset selectively active during corrective actions, somatic activity was more strongly associated with action timing. Additionally, learning induced divergent plasticity: tuft dendrites increased their selectivity but decreased response gain, maintaining stable net selectivity, whereas somas showed increased net selectivity early in learning. Together, these findings reveal distinct sensorimotor representations and learning-related plasticity in dendritic and somatic compartments, providing insight into how compartment-specific activity in the frontal cortex may contribute to motor skill acquisition.

      Strengths:

      The authors developed an innovative imaging approach and a comprehensive data analysis pipeline to address a knowledge gap in the literature. By alternating imaging of dendritic tufts and somas in the same animals, they compare compartment-specific activity during motor learning and identify distinct encoding of task variables and learning-related plasticity across these compartments. Interestingly, a subset of dendritic tufts shows activity associated with corrective actions. The findings are discussed in the context of current theories of dendritic computation, credit assignment, and motor learning, providing a useful foundation for future mechanistic studies.

      We thank the reviewer for highlighting interesting findings in the paper and their assessment that it provides a “useful foundation for future mechanistic studies”.

      Weaknesses:

      No major weaknesses were identified.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A very minor suggestion: it would be useful to mention the model organism in the abstract or title.

      (1.6) Thank you for catching this. We have added the model organism to the abstract as follows:

      “Using longitudinal two-photon calcium imaging, we investigated sensorimotor encoding in the apical tuft dendrites and somata of L5 extratelencephalic (ET) neurons in the frontal cortex of mice during learning of a discrete change to a cued dexterous action.”

      Reviewer #3 (Recommendations for the authors):

      Major:

      (3.1) Lines 197-199: It is unclear why the authors conclude that somas have stronger representations of choice and task outcome. In Figure 4G, there is no significant difference between dendrites and somas for Choice or Outcome coding selectivity. The differences in the Sensory/Choice and Sensory/Outcome ratios shown in Figure 4H,I are likely explained by stronger Sensory selectivity in dendrites (Fig 4g), rather than by stronger Choice or Outcome encoding in somas.

      We agree with the reviewer’s interpretation of the data. The statement at 197 - 199 was meant to reflect relative selectivity, but it was imprecise. We have replaced that sentence with the following, more precise sentence:

      “Somatic activity also encoded these features, but the representation of the stimulus was weaker – and the representation of action timing was stronger – than in the tuft dendrites.”

      (3.2) Figure 4B: The authors realign FL-associated IRFs to GO-cue timing using the mean FL latency for each trial type and animal. Because FL timing is jittered across trials and may differ between CL and CR trials, this could smear the realigned traces and complicate the interpretation of contact-associated activity. The authors should consider using trial-by-trial FL timing for realignment or quantify the impact of FL-timing variability on the resulting traces.

      We aligned average GO- and contact-IRFs in Figure 4B so that comparisons of their magnitudes could be drawn from the same time window.

      With regard to jitter across trials, we think the reviewer may have misinterpreted how the mean IRFs in Figure 4B are calculated. The contact-IRF, by its nature, is calculated once per animal and trial type with respect to FL timing and shifted once based on mean FL latency. There is no smearing due to trial-to-trial FL timing.

      With regard to systematic differences in FL timing across animals and CR vs. CL, the reviewer is correct that this could – in theory – smear the realigned mean contact-IRF shown in Figure 4B. However, differences in mean FL latency across animals and trial-types are small compared with the long-timescale contact-IRFs. Thus, the non-realigned (i.e., always FL-aligned) mean contact-IRF looks nearly identical to Figure 4B just globally offset in time, as shown in Author response image 1:

      Author response image 1.

      Since this is nearly identical to data already presented in Figure 4B, we do not think it is necessary to include it in the revised article. However, we have added the following to the Methods:

      “Population averages of contact-IRFs that were not shifted prior to averaging were nearly identical (excluding the overall temporal shift; data not shown), indicating that pooling of mean IRFs across animals and trial types produces minimal smearing of the final population IRF.”

      (3.3) Figure 5A: Are CA trials specific to motor learning, or do they reflect a corrective lick toward the alternative port after an unrewarded lick? An analysis of the second lick on left-error trials or pre-shift right-error trials could help distinguish whether correction licking reflects a general decision change after failed reward, or a motor-command correction specific to post-shift motor learning. The authors should also report the prevalence of CA versus AP trials and clarify whether these trial types are behaviorally distinct.

      We thank the reviewer for highlighting the need to emphasize that CA trials reflect a distinct behavior related to reaching the displaced port.

      By definition, CA trials started as Motor Error trials and thus reflected a corrective lick toward the same port after an unrewarded lick. Almost all first contact licks on Motor Error trials were well outside the distribution of correct left licks both pre- and post-shift (Figure 1 - figure supplement 1B,D), consistent with the interpretation of this first lick as directed toward the right port. Thus, we see no evidence suggesting that CA trials involve a decision change. CA trials are exceedingly rare pre-shift, because Motor Error trials are rare pre-shift (Figure 1I, only ~5% of all right trials).

      With regard to other error types before the shift, most expert-trained mice did not immediately sample the other port with a second lick after an unrewarded lick. They usually either stopped licking immediately or licked the unrewarded port multiple times before switching ports. When port switches occurred pre-shift, timing was highly variable across mice and trials. Even on rewarded trials, some mice would “check” the unrewarded port after consuming the reward, as can be seen in Figure 3I, J. All of these behaviors are clearly distinct from the stereotyped second lick that occurred on CA trials after the shift. We agree that the prevalence of CA and AP trials, as well as the prevalence of immediate port alternation, should be reported, and we have added that information to the article as follows:

      “On Correction Attempted (CA) trials, the first lick made contact with the incorrect port, and the mouse chose to direct a second lick toward the correct port (Figure 5A; prevalence: 54% of motor error trials). We interpreted these licks as a corrective action, because the tongue exit angle shifted further toward the correct target (Figure 5B). Abandoned Port (AP) trials were the same as CA trials, except the mouse either did not make a second attempt or the second lick was directed toward the incorrect port (Figure 5A; prevalence: 46% of motor error trials).”

      (3.4) The classification of pre-shift errors into motor and decision errors is not clear. If error-trial exit angles follow a unimodal distribution (Figure 1- Figure Supplement 1C), then the distinction between motor and decision errors may not be behaviorally well separated. The authors should explain how these categories are validated and whether conclusions depending on this classification are robust to alternative definitions.

      We do not conclude that motor errors and decision errors are distinguishable pre-shift. Pre-shift licks were classified into motor error and decision error categories only to demonstrate that the boundary we established for classifying post-shift licks classifies extremely few (~5%, Figure 1I) pre-shift licks as motor errors. No conclusions were drawn from comparisons between pre-shift licks classified as decision errors and those classified as motor errors. The categorization is defined by the distribution of exit angles pre-shift and validated by the bimodal distribution of exit angles on error trials post-shift. To improve clarity regarding our classification of pre-shift errors, we have added the following to the Results:

      “Exit angles after the shift exhibited a bimodal distribution across error trials (Figure 1G,H; Figure 1—figure supplement 1C,D), supporting this distinction in error type. The frequency of licks classified as motor errors on right-cued trials increased significantly after the shift (median pre-shift 0.06, median post-shift 0.42, p < 0.001; Figure 1I; Figure 1—figure supplement 1C,D), reflecting the new challenge of avoiding the left lickport. In contrast to after the shift, exit angles on error trials before the shift were unimodal (Figure 1—figure supplement 1C). These errors were classified based on the fixed CB in order to demonstrate that very few pre-shift licks qualify as motor errors (Figure 1I), and not to suggest that tongue trajectories before the shift are behaviorally well-separated.”

      (3.5) Figure 1- Figure Supplement 1D, post-shift decision errors: Are these truly decision errors? The lick angles appear similar to those observed before the shift, suggesting that these trials may reflect execution of a "default" or "uncertain" lick trajectory rather than an incorrect choice under the new contingency.

      The post-shift exit angles on right-cued decision error trials (Figure1 - figure supplement 1D, grey) are similar to the lick angles on correct left-cued trials pre-shift (Figure 1 - figure supplement 1A, red) and clearly different from the correct right-cued trials pre-shift (Figure 1 figure supplement 1A, blue). Thus, to the extent that the animal’s intention can be measured from lick trajectory, it was targeting the incorrect (left) port. It is also true that it may still target the previous location of the left port (a “default” left trajectory), but because the decision error makes precise targeting irrelevant to the task outcome (it is easy to reach the left port after the shift), we do not designate it as a joint decision error and motor error. As to whether the deliberative process leading to this action is somehow cognitively distinct from other behaviors typically labeled as decision errors or incorrect choices, we cannot say.

      Minor:

      (3.6) Vocabulary consistency: soma vs somata.

      When data are shown for, or derived from, multiple somata, we use “somata”. When data are shown for an individual soma (such as in a panel with data from a single example soma), we use “soma.” We could not find any use of “somas,” which would indeed be inconsistent.

      (3.7) Figure 1C: I am not sure why the lick trajectories do not depict the tongue exiting the mouse. What time window is shown? Why does it look like the trajectories are shifted to the left?

      We thank the reviewer for identifying this issue. The definition of the location labeled “mouth” was accidentally omitted. The lick trajectories in Figure 1C do depict the tongue tip once it became visible to the cameras. Jaw opening and shifting partly determined the location where the tongue became visible in the videography. These movements varied from mouse to mouse and trial to trial, so exit angle was measured from the approximate midpoint between the temporomandibular joints, which is the grey point in 1C. We have fixed the captions and Methods to precisely define this location. With regard to the appearance of a slight leftward shift in the trajectories, this reflects how the tongue exits the mouth and how the tongue tip curves downward as the tongue approaches the port.

      (3.8) Figure 1- Figure supplement 1: it could ease the comparisons to report population statistics, such as median, from panel A to panel B and D, population statistics from B to D.

      Thank you. We have added these statistics to the Figure 1 - figure supplement 1 caption.

      (3.9) Choice boundary (CB) should be defined in line 100, not 110.

      Thank you. We have fixed this.

      (3.10) Line 109: claim not supported by referenced figure (Figure 1 - Figure Supplement 1). Lick angle histogram to the right port, pre-shift does not overlap substantially with lick angle to the left port, post-shift.

      We thank the reviewer for the opportunity to clarify this. We agree that Figure 1 - figure supplement 1 is not sufficient to support the claim. First, we want to make clear that Figure 1 - figure supplement 1 does not contradict the claim. The new location of the left port can obstruct the tongue during right-cued licks, regardless of the distributions of left licks pre- or post-shift. Second, to confirm that the new location of the left port would obstruct a substantial fraction of pre-shift right-cued lick trajectories, we measured the minimum distance between tongue trajectories and the post-shift location of the left port. Of pre-shift right-cued exit trajectories, 30 +/- 5% came within 1.25 mm – half of the combined tongue width (1.5 mm) and port width (1 mm) – of the port center.

      To make this claim more precise, we have changed the statement as follows:

      “Thus, on right-cued trials, mice continuing to follow the pre-shift motor plan would be biased to more frequently contact the new left port location (Figure 1E,F; 30 +/- 5% of pre-shift trajectories came within a tongue-width of the new location) and receive punishment (i.e., timeout).

      (3.11) Line 113: claim not supported by referenced figure. Figure 1G does not display error trials.

      We have changed the line to refer to “both correct and error trials”, such that reference to Figure 1G is also appropriate.

      (3.12) Figure 2 - Figure Supplementary 3 & method: how is noise estimated?

      Thank you. The following has been added to the Methods:

      “For Figure 2 - figure supplement 3, noise was estimated as the square-root of the geometric mean of the Welch power spectrum in a high-frequency band (0.25–0.5 times the frame rate; Giovannucci et al., 2019).”

      (3.13) Figure 2D: Was imaging during the shift epoch always performed in dendrites? If so, could the imaging schedule bias comparisons between dendritic and somatic activity during learning, especially given that mice show behavioral learning between early and late post-shift sessions (Figure 1J)?

      No, imaging during the shift was not always performed in the dendrites. The following has been added to the Methods to make clear that the post-shift data reflect dendritic and somatic imaging conducted on the day of the shift with roughly similar frequency:

      “For Figure 5 and Figure 6, which make comparisons between dendritic and somatic activity during the post-shift period, 67% of animals providing somatic data (4 of 6 mice) underwent somatic imaging on the day of the shift and 80% of mice providing dendritic data (8 of 10 mice) underwent dendritic imaging on the day of the shift.”

      (3.14) Lines 163-164, "we observed that the onset of tuft activity was consistently time-locked to the GO cue (vertical green line; Figure 3B). This was in contrast to somatic activity, which had more variable timing (Figure 3E)." The authors cite panels B and E in support of this point, but these appear to be example ROIs. It would be helpful to clarify how representative these examples are, since the corresponding population summaries in panels G and H do not make the effect immediately apparent.

      These examples are representative, as supported by the population summary of activity time-locked to the GO-cue versus port contact in Figure 4B.

      (3.15) Figure 4B: It could be useful to add the lick traces here as well. To allow the reader to have an idea of contact timing with respect to the Go cue and compare the sustain response with the licking pattern.

      We understand how this could be helpful. However, since these exact traces are already present in Figure 3I,J, we think that adding them to Figure 4 is unnecessary and would add complexity to an already very busy figure.

      (3.16) Figure 5E: Why are the imaging sessions labeled 0 and +1 rather than 0 and +2? Are the dendritic and somatic imaging not alternated?

      Yes, imaging was not alternated for 3 of the 22 mice. We have clarified this in the Methods, as follows:

      “Somatic and dendritic imaging sessions alternated every other day (19 of 22 mice), except for 3 mice in which only one compartment was imaged daily (dendrite-only: 2 mice, soma-only: 1 mouse). The exceptions were due to brain curvature or the angle of the coverslip with respect to the brain, such that only one compartment could be imaged and the other compartment was underneath skull regrowth or dural thickening that made high-quality imaging impossible.”

      (3.17) Figure 6C, legend: I suppose the authors meant "remapping", not "Post-shit SI distribution" for the description of the right column.

      Thank you. We have fixed this label.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents a useful finding on the effects of arginine vasopressin (AVP) on islet cells in pancreatic tissue slices, using technically sophisticated spatio-temporal calcium recordings to confirm that AVP influences α and β cells differently depending on glucose concentrations. While the study's methods - particularly the calcium imaging techniques and peptide ligand design targeting V1b receptors - are strong, the reviewers were concerned about several aspects of the experimental design. However, the results on βcell responses are incomplete and insufficient to support the manuscript's claims, especially due to the high variability of islet responses and lack of mechanistic and functional (hormone release) data. There are also concerns about the possibility of off-target effects and incomplete receptor specificity, noting that the study would have been significantly strengthened by inclusion of signaling pathway interrogation, hormone output assays, genetic validation (e.g., β cell-specific deletion of V1br), and receptor localization, although the work will still be of interest to researchers studying islet physiology in the context of health and diabetes.

      We sincerely thank the reviewers and editors for their thorough evaluation of our manuscript and their recognition of its technical strengths, including the advanced spatio-temporal calcium imaging and the rational design of selective V1b receptor ligands. We appreciate their acknowledgement of the study’s relevance for understanding AVP effects in a physiologically intact islet context and their positive assessment of our methodological rigour and innovation. The reviewers’ constructive feedback has helped us clarify the boundaries and intent of our study, which focuses on the glucose- and context-dependent modulation of α- and β-cell activity, rather than exhaustive molecular dissection.

      While the reviewers rightly emphasize the importance of receptor specificity and downstream signaling validation, we respectfully suggest that some of their concerns may reflect a lingering bias toward reductionist frameworks. Our interpretation is rooted in the emerging understanding that β-cell behaviour is largely defined by dynamic intercellular interactions within the islet collective, rather than by static gene expression or receptor localization alone (Jin et al., 2025; Korošak et al., 2021; Rutter et al., 2024). Recent studies have demonstrated that roles such as “leader” or “hub” β cells are transient and emergent, governed more by timing, environment, and local network structure than by fixed molecular identity (Postic et al., 2023; Gosak et al., 2018).

      This has profound implications for how we interpret cell responsiveness to agents like AVP: what appears as biological variability may in fact reflect context-sensitive transitions within a non-linear, self-organizing system (Stožer et al., 2021). Hence, we chose to focus on functional collective dynamics using intact pancreatic slices, rather than isolated cell models which fail to preserve the essential network architecture of islets. Although the addition of genetic models or isolated receptor measurements would strengthen receptor-specific conclusions, we argue that such approaches alone cannot resolve the physiological complexity of a system where function arises from cell–cell communication and spatiotemporal context.

      Indeed, the lack of direct correlation between receptor transcript abundance and functional outcomes has been noted in prior studies, reinforcing the view that function cannot be strictly predicted by molecular presence (Rutter et al., 2024). As articulated in our manuscript, the islet behaves as a sensory collective (Fancher & Mugler, 2017), where emergent patterns— not static cell identity—determine behaviour. This perspective aligns with broader shifts in biology away from strict genetic determinism toward causal emergence and collective agency (Ball, 2023; Levin, 2021).

      We therefore believe our study contributes not only new pharmacological insights but also a conceptual reframing of how AVP responses should be interpreted in a complex organ like the pancreas. We have added new data addressing reviewer suggestions—such as glucagon secretion assays, clarifications on the role of forskolin, and an analysis of event timing—that further support our conclusions. We also expanded the discussion on how islet variability is functionally meaningful, not just noise, and explained why β-cell responses to AVP must be interpreted within this probabilistic framework.

      We agree that future work should include receptor-specific knockouts and more direct signaling pathway assays, but these would need to be designed with careful consideration of the islet’s dynamic topology and the emergent nature of β-cell roles. In this light, we see our study not as the final word, but as a necessary systems-level foundation for more targeted interventions. We thank the reviewers again for their careful critiques and hope that our response clarifies both the rationale and scope of our work. Our revisions aim to enhance the paper’s clarity while maintaining its commitment to an integrative, physiology-rooted approach.

      We thank the reviewers and editors for their thoughtful and constructive assessment of our work. We are especially grateful for their recognition of the study’s technical strengths, including the use of spatio-temporal calcium imaging in intact pancreatic tissue and the strategic development of receptor-selective peptide ligands. We also appreciate their acknowledgement that our study contributes to the understanding of glucose-dependent AVP effects in islet physiology. The reviewers’ concerns regarding variability, receptor specificity, and functional validation helped us further clarify the scope and context of our study.

      We respectfully submit that some reservations stem from a reductionist framing that may not fully account for the collective behaviour of islets. As we and others have shown, β-cell function arises from emergent, self-organizing network dynamics, not just from static gene expression or receptor abundance (Jin et al., 2025; Korošak et al., 2021; Postic et al., 2023). In this view, pharmacological heterogeneity across islets is not simply noise or experimental inconsistency, but a signature of dynamic attractor states within the islet network (Stožer et al., 2021). Because an islet functions as a coupled system, most response variability originates from its emergent collective behavior, which eclipses variability in receptor expression or metabolic state.

      For this reason, even single-islet receptor quantification or ATP measurements would provide limited explanatory power: it is the state of the network—not absolute receptor levels—that determines whether a perturbation elicits activation or inhibition. As we illustrate in our graphical abstract, a single islet tested repeatedly under identical glucose conditions can yield divergent responses, simply because it occupies different dynamic states. These findings are in line with systems biology and network science approaches, which have revealed that cell function, especially in the β-cell collective, cannot be fully understood through reductionist parameters alone (Gosak et al., 2018; Ball, 2023).

      We have included glucagon secretion assays and new analyses to address key reviewer suggestions. Still, we chose not to pursue extensive knockouts or cAMP imaging, as these would require a different experimental scope and could risk disrupting the very dynamics we aim to understand. Likewise, while direct measurements of V1bR or IP3R expression would add molecular detail, they are not definitive without network context. The bell-shaped AVP dose-response curve and its explanation through IP3R inactivation are supported by prior studies; we invoke this mechanism not speculatively, but because it provides the most parsimonious explanation for the glucose-dependent shift in β-cell responsiveness.

      We also clarify that our study does not aim to resolve every mechanistic detail, but rather to offer a systems-level insight into how AVP modulates islet dynamics across varying glucose and cAMP contexts. The implications extend beyond AVP pharmacology, suggesting that perturbations to β-cell function must be understood within a probabilistic, state-dependent framework (Fancher & Mugler, 2017). This resonates with emerging concepts in cell physiology that emphasize causal emergence and local agency over static molecular determinism (Levin, 2021; Rutter et al., 2024).

      In summary, we see our work as part of a necessary shift in perspective—from linear receptor-function models to context-sensitive dynamic systems. We are grateful for the opportunity to revise our manuscript in response to insightful feedback and hope our clarifications and new data will strengthen its impact for the islet research community.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors confirmed earlier findings that AVP influences α and β cells differently, depending on glucose concentrations. At substimulatory glucose levels, AVP combined with forskolin - an activator of cAMP -did not significantly stimulate β cells, although it did activate α cells. Once glucose was raised to stimulatory levels, β cells became active, and α cell activity declined, indicating glucose's suppressive effect on α cells and permissive effect on β cells. Under physiological glucose levels (8-9 mM), forskolin enhanced β-cell calcium oscillations, and AVP further modulated this activity. However, AVP's effect on β cells was variable across islets and did not significantly alter AUC measurements (a combined indicator of oscillation frequency and duration). In α cells, forskolin and AVP led to increased activity even at high glucose levels, suggesting that α cells remain responsive despite expected suppression by insulin and glucose.

      Experiments with physiological concentrations of epinephrine suggest that AVP does not operate via Gs-coupled V2 receptors in β cells, as AVP could not counteract epinephrine's inhibitory effects. Instead, epinephrine reduced β cell activity while increasing α cell activity through different G-proteincoupled mechanisms. These results emphasize that AVP can potentiate αcell activation and has a nuanced, context-dependent effect on β cells.

      The most robust activation of both α and β cells by AVP occurred within its physiological osmo-regulatory range (~10-100 pM), confirming that AVP exerts bell-shaped concentration-dependent effects on β cells. At low concentrations, AVP increased β cell calcium oscillation frequency and reduced "halfwidths"; high concentrations eventually suppressed β cell activity, mimicking the muscarinic signaling. In α cells, higher AVP concentrations were required for peak activation, which was not blunted by receptor inactivation within physiological ranges.

      Attempting to further dissect the role of specific AVP receptors, the authors designed and tested peptide ligands selective for V1b receptors. These included a selective V1b agonist; a V1b agonist with antagonist properties at V1a and oxytocin receptors; and a selective V1a antagonist. In pancreatic slices, these peptides seem to replicate AVP's effects on Ca<sup>2+</sup> signaling, although responses were highly variable, with some islets showing increased activity and others no change or suppression. The variability was partly attributed to islet-specific baseline activity, and the authors conclude that AVP and V1b receptor agonists can modulate β cell activity in a statedependent manner, stimulating insulin secretion in quiescent cells and inhibiting it in already active cells.

      We applaud the reviewer to capture the essence of work in their introduction.

      Strengths:

      Overall, the study is technically advanced and provides useful pharmacological tools. However, the conclusions are limited by a lack of direct mechanistic and functional data. Addressing these gaps through a combination of signaling pathway interrogation, functional hormone output, genetic validation, and receptor localization would strengthen the conclusions and reduce the current (interpretive) ambiguity.

      Thank you!

      Weaknesses:

      (1) The study is entirely based on pharmacological tools. Without genetic models, off-target effects or incomplete specificity of the peptides cannot be fully ruled out.

      We partially agree with this comment and acknowledge that genetic models would provide a valuable complementary approach to address possible off-target effects or incomplete peptide specificity. However, genetic models also have important limitations, particularly when the aim is to resolve subtle, population-level physiological differences in beta cell activity. We therefore used pharmacological tools at different concentrations to test whether the observed effects were concentration-dependent and consistent with the expected receptor-mediated actions. An advantage of the pancreatic slice preparation is that it preserves much of the native tissue environment and allows pharmacological manipulation within concentration ranges closer to in vivo efficacy, thereby reducing the likelihood of nonspecific effects. To compensate for the lack of genetic models, we now emphasize the collective activity analysis as an additional strength of the study and have clarified this limitation in the revised manuscript.

      (2) Despite multiple claims about β cell activation or inhibition, the functional output - insulin secretion - is weakly assessed, and only in limited conditions. This aspect makes it very hard to correlate calcium dynamics with physiological outcomes.

      We agree that the functional output needed stronger support and have therefore expanded the hormone secretion experiments. While the effects of AVP and its analogues were tested during a stable plateau phase in the Ca<sup>2+</sup> imaging experiments, this phase provides only a narrow dynamic range for insulin release measurements in mouse slices. We therefore added a sequence of stimulations on the same slices, using 8 mM glucose and 500 nM forskolin, with glucose lowered to a non-stimulatory range between different AVP concentrations. These new experiments better define how AVP-dependent changes in Ca<sup>2+</sup> dynamics translate into insulin secretion under conditions with a broader secretory dynamic range. The new insulin and glucagon secretion data have now been added to the manuscript as Figure 5, and the text has been revised accordingly.

      (3) Insulin and glucagon secretion assays should be provided; the authors should measure hormone release in parallel with Ca2+ imaging, using perifusion assays, especially during AVP ramp and peptide ligand applications.

      We added insulin and glucagon secretion assays for AVP ramp to Figure 5.

      Additionally, there is no standardization of the metabolic state of islets. The authors should consider measuring islet NAD(P)H autofluorescence or mitochondrial potential (e.g., using TMRE) to control for metabolic variability that may affect responsiveness.

      We agree that standardization of the metabolic state of the islets would further strengthen the interpretation of the responsiveness data. We attempted to address this experimentally, but the results were inconclusive and therefore not included in the manuscript. Based on our previous unpublished observations, NAD(P)H levels appear to be significantly higher and less variable in islets within tissue slices than in isolated islets, suggesting that the slice preparation may better preserve the native metabolic state. However, we acknowledge that this remains an important limitation and we now indicate that in the manuscript. Additional experiments will be required to establish a robust and standardized approach, for example by combining NAD(P)H autofluorescence and/or mitochondrial potential measurements with Ca<sup>2+</sup> imaging.

      (4) There is a high degree of variability in response to AVP and V1b agonists across islets (activation, no effect, inhibition). Surprisingly, the authors do not fully explore the cause of this heterogeneity (whether it is due to receptor expression differences, metabolic state, experimental variability, or other conditions).

      This is a well-taken point and has indeed been one of the major bottlenecks in interpreting the results of this study. We agree that the variability in responses to AVP and V1b agonists may reflect several factors, including receptor expression, metabolic state, experimental conditions, and differences in the functional state of individual islets. However, our data also suggest that the beta cell population within an islet should be considered as a dynamic, non-linear system, in which even small differences in initial conditions or collective state can result in qualitatively different outcomes, including activation, no apparent effect, or inhibition. In this framework, the response to AVP is not determined by receptor expression alone, but by the current physiological context of the islet network. This is also why we believe that pharmacological tests are most informative when interpreted within a defined functional state rather than as isolated receptor-specific readouts. As indicated in the graphical abstract, apparently similar islets may occupy different dynamic states and therefore respond differently to the same Gq/PLC/IP3R stimulus. We have now expanded the discussion to make this interpretation more explicit and to acknowledge that receptor expression, metabolic variability, and experimental factors remain possible contributors that will require further targeted studies.

      The following text has been added to expand the discussion:

      “The heterogeneous responses observed across different islets, where some showed increased activity while others showed no detectable change or inhibition, could intuitively be attributed to variability in V1b receptor expression or signaling capacity among β cells. Such an explanation would be consistent with differences in receptor density, coupling efficiency to Gq proteins, or downstream signaling components such as PLC or IP<sub>3</sub> receptors. However, our data suggest that receptor-level variability alone is unlikely to fully explain the observed response spectrum, and that the current functional state of the islet collective must also be considered. The islet behaves as a non-linear dynamic system in which the same molecular perturbation can produce different functional outcomes depending on the current state of the β-cell collective. In such systems, cells or cell populations do not occupy a single deterministic activity state, but rather move within a landscape of possible states, with perturbations shifting the probability distribution of transitions between them. This concept is well established in dynamical systems approaches to biological cell-state transitions, where attractor landscapes, noise, and signaling inputs determine the probability of moving between alternative functional states rather than enforcing a single fixed output.

      In this framework, AVP and V1b receptor-selective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that βcell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone. Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (5) There is no validation of V1b receptor expression at the protein or mRNA level in α or β cells using in situ hybridization, immunohistochemistry, or spatial transcriptomics.

      We agree with the reviewer that spatial validation of V1b receptor expression is important for interpreting the cellular targets of AVP signaling in the islet. We have therefore added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      (6) AVP effects are described in terms of permissive or antagonistic effects on cAMP (especially in relation to epinephrine), but direct measurements of cAMP in α and β cells are not shown, weakening these conclusions. The authors should use Epac-based cAMP FRET sensors in α and β cells to monitor the interaction between AVP, forskolin, and epinephrine more conclusively.

      We agree that direct measurements of cAMP dynamics in alpha and beta cells would provide a more conclusive assessment of the interaction between AVP, forskolin, and epinephrine signaling. We attempted to address this experimentally; however, within the time domain of the Ca<sup>2+</sup> oscillations analyzed here, the temporal resolution and robustness of currently available cAMP readouts were not sufficient to resolve these interactions reliably. Even at slower time scales, cAMP sensor signals can be difficult to interpret quantitatively and may be overinterpreted if not tightly linked to the functional readout. We have therefore moderated the wording of the manuscript and now describe the proposed permissive or antagonistic interaction between AVP/V1b and cAMP-dependent signaling as an interpretation supported by the pharmacological Ca<sup>2+</sup> response patterns, rather than as a directly demonstrated cAMP mechanism. We now explicitly acknowledge in the limitations that most experiments were performed under cAMP-permissive conditions, which increases sensitivity for detecting AVP-dependent modulation but complicates the separation of direct beta cell effects from intra-islet interactions. Future studies using optimized cell-type-specific Epac-based sensors will be required to resolve this interaction.

      (7) Single-islet transcriptomics or proteomics (also to clarify variability) should be provided to analyze receptor expression variability across islets to correlate with response phenotypes (activation vs inhibition). Alternatively, the authors could perform calcium imaging with simultaneous insulin granule tracking or ATP levels to assess islet functional states.

      We agree that single-islet transcriptomics, proteomics, or simultaneous metabolic readouts could provide useful complementary information, particularly for describing molecular variability across islets. However, we do not think that differences in receptor expression or ATP levels alone are sufficient to explain the diversity of response phenotypes observed here. Our interpretation is that the beta cell population behaves as a collective dynamic system, in which the same input can lead to different outcomes depending on the current state of the network and its local physiological context. In such a system, AVP/V1b signaling does not necessarily impose a single deterministic response, but changes the probability distribution of accessible states, including activation, inhibition, or no detectable response. Theoretically and partially confirmed by the preliminary data, even the same islet exposed repeatedly under apparently identical conditions could be expected to display different responses if it occupies a different position within this dynamic state space at the time of stimulation. This concept is summarized in the graphical abstract and is central to our interpretation of the pharmacological data. We have therefore clarified in the Discussion that receptor expression, ATP levels, and other molecular parameters may modulate the response landscape, but are unlikely to fully define the observed functional phenotype without considering the collective dynamics of the islet.

      Added to Discussion section: “In this framework, AVP and V1b receptorselective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that β cell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone (63). Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (8) While the study implies AVP acts through V1b receptors on β cells, the signaling downstream (e.g., PLC activation, IP3R isoforms involved) is simply inferred but not directly shown.

      We agree that downstream signaling was not directly resolved at the level of PLC activation or specific IP3R isoforms. However, we did not infer Gq/PLC/IP3R involvement solely from AVP pharmacology, but used ACh as an independent Gq-coupled receptor reference stimulus in the same pancreatic slice preparation. With ACh concentration ramps, we could reproduce both activation and inactivation patterns observed with AVP/V1b stimulation, supporting the interpretation that these responses arise from modulation of the Gq-dependent Ca<sup>2+</sup> signaling axis.

      In addition, in prelilminary expriments we could observe that inhibition of Gq activity with YM254890, as well as interference with IP3R-dependent signaling using Xestospongin C, diminished the response, although not completely. This incomplete suppression is important, because it suggests that beta cell Ca<sup>2+</sup> homeostasis and collective islet activity are not controlled by a single linear pathway, but by partially redundant and context-dependent mechanisms. We have therefore revised the manuscript to state more cautiously that our data support the involvement of Gq/PLC/IP3R-dependent signaling, while acknowledging that direct measurements of PLC activity and IP3R isoform-specific contributions remain outside the scope of the present study.

      (9) The interpretation that IP3R inactivation (mentioned in the title!) underlies the bell-shaped AVP effect is just hypothetical, without direct measurements. Assays in β (and/or α)-cell-specific V1b KO mice and IP3R KO mice must be provided to support these speculations.

      We agree that the involvement of IP3R-dependent signaling should be stated with appropriate caution. However, the concept of IP3R inactivation as a mechanism contributing to bell-shaped Gq-dependent Ca<sup>2+</sup> responses is not purely hypothetical, since IP3R inactivation has been directly demonstrated in previous studies and provides a parsimonious explanation for the shift from activation to suppression at higher AVP concentrations. In the present study, this interpretation is further supported by the glucose dependence of the AVP concentration-response relationship, where different stimulatory glucose conditions shift the apparent efficacy peak.

      We also agree that cell-specific V1b receptor and IP3R knockout experiments would be valuable future approaches. In this respect, we have obtained preliminary results from a small sample of IP3R triple-knockout mice, which cannot yet be fully included because they are part of an ongoing collaboration. In these experiments, supraphysiological AVP concentrations did not produce the IP3R-like beta-cell response pattern observed in controls, namely reduced halfwidth and increased frequency, whereas alpha cell stimulation was preserved similarly to WT slices.

      At the same time, we believe that definitive knockout experiments must be carefully designed, because the beta cell population behaves as a dynamic collective system in which the response to AVP depends on the current functional state of the islet, glucose context, and intercellular coupling. We therefore now present IP3R inactivation as a strongly supported mechanistic interpretation rather than as a directly proven mechanism in this study, and we explicitly acknowledge that cell-specific V1b and IP3R genetic models will be required to fully resolve this pathway.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Drs. Kercmar, Murko, and Bombek make a series of observations related to the role of AVP in pancreatic islets. They use the pancreatic slice preparation that their group is well known for. The observations on the slide physiology are technically impressive. However, I am not convinced by the conclusions of this manuscript for a number of reasons. At the core of my concern is perhaps that this manuscript appears to be motivated to resolve 'controversies' surrounding the actions of AVP on insulin and glucagon secretion. This manuscript adds more observations, but these do not move the field forward in improving or solidifying our mechanistic understanding of AVP actions on islets. A major claim in this manuscript is the beta cell expression of the V1b Receptor for AVP, but the evidence presented in this paper falls short of supporting this claim.

      Observations on the activation of calcium in alpha cells via V1b receptor align with prior observations of this effect.

      I have focused my main concerns below. I hope the authors will consider these suggestions carefully - please be assured that they were made with the intent to support the authors and increase the impact of this work.

      We thank the reviewer for their detailed input and support to increase the impact of our work and our understanding of important cellular processes overall. We have considered their suggestions carefully to further expand the strenghts of our approach and analysis.

      Strengths:

      The main strength of this paper is the technical sophistication of the approach and the analysis and representation of the calcium traces from alpha and beta cells.

      Thank you!

      Weaknesses:

      (1) The introduction is long and summarizes a substantive body of literature on AVP actions on insulin secretion in vivo. There are a number of possible explanations for these observations that do not directly target islet cells. If the goal is to resolve the mechanistic basis of AVP action on alpha and beta cells, the more limited number of papers that describe direct islet effects is more helpful. There are excellent data that indicate that the actions of AVP are mediated via V1bR on alpha cells and that V1bR is a) not expressed by beta cells and b) does not activate beta cell calcium at all at 10 nM - which is the same concentration used in this paper (Figure 4G) for peak alpha cell Ca2+ activation (see https://doi.org/10.1016/j.cmet.2017.03.017; cited as ref 30 in the current manuscript).

      We thank the reviewer for this important comment and agree that the literature on AVP actions in vivo is complex, with several possible sites of action outside the islet. We have therefore revised the Introduction to make the rationale more focused and to better separate systemic effects of AVP from studies addressing direct actions on pancreatic islet cells. At the same time, we chose not to restrict the Introduction only to the alpha cell V1bR literature, because one of the aims of the manuscript is precisely to address why AVP effects on insulin secretion have remained difficult to interpret across experimental contexts.

      Our results fully confirm a central aspect of the study cited by the reviewer, namely that V1bR activation robustly stimulates alpha cell Ca<sup>2+</sup> activity under non-stimulatory glucose conditions, and that 10 nM AVP does not produce a uniform activation of beta cell Ca<sup>2+</sup> activity. In fact, in a substantial fraction of beta cell populations, 10 nM AVP failed to activate oscillations, consistent with the view that alpha cells are the more sensitive and more direct cellular target of AVP/V1bR signaling. However, we do not think that the available transcriptomic evidence is sufficient to categorically exclude V1bR expression or functional relevance in beta cells. Re-analysis of the published dataset, together with more recent datasets and our newly added RNAscope data, supports a higher relative expression of V1bR transcripts in alpha than in beta cells, but does not justify treating beta (or non-alpha) cell expression as absent.

      We have therefore revised the manuscript to avoid overstating beta cell V1bR expression as a major isolated claim. Instead, we now present the data as evidence that AVP/V1bR signaling acts most prominently through alpha cells, while beta cell responses emerge in a concentration-, glucose-, and statedependent manner within the intact islet. This interpretation is consistent with the reviewer’s concern that 10 nM AVP preferentially activates alpha cells, but it also accommodates our observation that beta cell collective activity can be modulated under defined pharmacological and metabolic conditions. We believe that this is an important distinction, because the absence of a uniform beta cell Ca<sup>2+</sup> activation at one AVP concentration does not exclude beta cell modulation by AVP/V1bR signaling within the intact islet network. The Introduction and Discussion have been revised accordingly to clarify that our study does not simply challenge the alpha cell V1bR model, but expands it by examining how AVP-dependent alpha cell activation, possibly lower beta-cell receptor expression, and collective beta cell dynamics interact in the native pancreatic slice preparation.

      (2) We know from bulk RNAseq data on purified alpha, beta, and delta cells from both the Huising and Gribble groups that there is no expression of V2a. I will point you to the data from the Huising lab website published almost a decade ago (http://dx.doi.org/10.1016/j.molmet.2016.04.007) - which is publicly available and can be used to generate figures (https://huisinglab.com/dataghrelin-ucsc/index.html). They indicate the absence of expression of not only AVP2 receptors anywhere in the islet, but also the lack of expression of V1bra, V1brb, and Oxtr in beta cells. Instead of the detailed list of expression of these 4 receptors elsewhere in the body, it would be more directly relevant to set up their pancreatic slice experiments to summarize the known expression in pancreatic islets that is publicly available. It would also have helped ground the efforts that involved the generation of the V1aR agonist and V2R antagonist, which confirm these known AVP/OXT receptor expression patterns.

      We thank the reviewer for pointing us more directly to the publicly available islet expression datasets. We agree that the expression of AVP/OXT receptors in purified alpha, beta, and delta cells provides an important reference frame for interpreting our pharmacological data, and we have revised the manuscript to summarize these islet-specific datasets more directly rather than emphasizing receptor expression in other organs. These data support the absence or very low expression of V2 receptors in islet endocrine cells and confirm that V1b receptor expression is substantially enriched in alpha cells compared with beta cells.

      At the same time, as outlined in our response above, we do not think that the currently available transcriptomic datasets are sufficient to categorically exclude low-level V1bR transcript expression or functional relevance in beta cells within the intact islet. For this reason, we added independent RNAscope validation to assess V1bR transcripts in the pancreatic slice preparation. These data confirm stronger V1bR expression in glucagon-positive alpha cells, while also showing a broader expression pattern within the islet and pancreas.

      We have also revised the rationale for the pharmacological experiments using V1aR- and V2R-directed tools. We now present these experiments not as evidence for unexpected receptor expression, but as functional controls that are consistent with the known AVP/OXT receptor expression patterns in pancreatic islets. This better aligns the manuscript with the existing transcriptomic literature while preserving the main physiological question of the study: how AVP/V1bR-dependent signaling reshapes alpha-cell activity and beta-cell collective dynamics in intact pancreatic tissue.

      (3) Importantly, the lack of V1br from beta cells does not invalidate observations that AVP affects calcium in beta cells, but it does indicate that these effects are mediated a) indirectly, downstream of alpha cell V1br or b) via an unknown off-target mechanism (less likely). The different peak efficacies in Figure 4G would also suggest that they are not mediated by the same receptor.

      We agree with the reviewer that the absence or very low abundance of V1bR transcripts in beta cells in published transcriptomic datasets would not invalidate the observation that AVP modulates beta-cell Ca<sup>2+</sup> activity. It does, however, raise the important question of whether this modulation is mediated indirectly through alpha-cell V1bR activation, through V1bR expression in beta cells that is difficult to resolve transcriptomically, or through another mechanism. To address this more directly, we have now added RNAscope data, which confirm relatively stronger V1bR transcript enrichment in glucagon-positive alpha cells, but also show a broader V1bR transcript signal within the islet and pancreas. Thus, while our data support alpha cells as the dominant V1bR-positive endocrine population, they do not support a strict absence of V1bR-associated signaling capacity in the beta cell compartment.

      We also agree that different peak efficacies in alpha and beta cells could be interpreted as evidence for distinct receptors or indirect mechanisms. However, we favor a different interpretation: the apparent efficacy of AVP depends strongly on the physiological state in which the cells are tested. This is particularly evident in beta cells, where the AVP efficacy peak shifts with glucose concentration, suggesting that the beta-cell response is shaped by the metabolic and Ca<sup>2+</sup>-handling context rather than by receptor occupancy alone. In this framework, the same V1bR/Gq-dependent input can generate different downstream Ca<sup>2+</sup> outcomes in alpha and beta cells because the two cell types operate in different dynamic regimes.

      We have therefore revised the manuscript to acknowledge this dilemma more explicitly. We now state that beta cell effects of AVP could include indirect alpha cell-dependent components, but given the magnitude and statedependence of the beta cell Ca<sup>2+</sup> response it is unlikely to be driven by alpha cell activation. Instead, our preferred interpretation is that AVP/V1bR signaling acts within the intact islet as a context-dependent perturbation of the collective beta cell Ca<sup>2+</sup> system, with IP3R-dependent mechanisms being modulated by glucose-dependent changes in beta cell excitability and intracellular Ca<sup>2+</sup> handling.

      (4) The rationale for the use of forskolin across almost all traces is unclear. It is motivated by a desire to 'study the AVP dependence of both alpha and beta cells at the same time'. As best as I can determine, the design choice to conduct all studies under sustained forskolin stimulation is related to the permissive actions of AVP on hormone secretion in response to cAMPgenerating stimuli. The permissive actions by AVP that are cited are on hormone secretion, which in many cell types requires activation of both calcium and cAMP signaling. Whether the activation of V1br and subsequent calcium response is permitted by cAMP is unclear. I believe the argument the authors are making here is that the activation of beta cell calcium by AVP is permitted by forskolin. i.e., the cAMP stimulated by it in beta cells. However, the design does not account for the elevation of cAMP in alpha cells and subsequent release of glucagon, particularly upon co-stimulation with AVP, which permits glucagon release by activating a calcium response in alpha cells. This glucagon could then activate beta cells. If resolving the mechanism of action is the goal, often less is more. The activation of Gaq-mediated calcium is not cAMP dependent (although the downstream hormone secretion clearly often is). As was shown, AVP does not activate calcium in beta cells in the absence of cAMP. The experiments in Figures 1, 2, and 4 should have been completed in the absence of cAMP first.

      We agree with the reviewer that the use of forskolin needs to be explained more clearly, and we have revised the manuscript accordingly. Our rationale was based on the established permissive role of cAMP in AVP-dependent endocrine responses, but we acknowledge that this does not necessarily imply that the upstream V1bR/Gq-mediated Ca<sup>2+</sup> response itself is cAMP-dependent. The reviewer is also correct that forskolin elevates cAMP broadly and therefore may affect both alpha and beta cells, including the possibility that AVP-enhanced alpha cell activation and glucagon release secondarily influence beta cell activity.

      In fact, our initial experiments were performed without forskolin and revealed an important difficulty: stimulatory glucose alone can increase cAMP levels to a variable extent, as also supported by our previous work on epinephrine signaling, thereby shifting the apparent peak efficacy of AVP stimulation. Thus, forskolin was originally used to reduce this variability and create a more defined cAMP-permissive background in which alpha and beta cell responses could be compared in the same slice. However, we agree that this design works against isolation of beta cell-autonomous AVP effects.

      Within a scope of another study we have done an independent series of more focused experiments using GLP-1 receptor stimulation, which preferentially increases cAMP signaling in beta cells compared with the broad cAMP elevation produced by forskolin. We have clarified that the modulation of the AVP-dependent pathway by GLP-1 and related ligands at largely supports beta cell-autonomous AVP effects. It is part of ongoing work and will be reported independently, because a full mechanistic dissection of cAMP–AVP interactions goes far beyond the scope of the present study.

      (5) It is unexpected that epinephrine in Figure 2 does not activate the alpha cell calcium? A recent paper from the same group (Sluga et al) shows robust calcium activation in alpha cells in a similar prep by 1 nM epinephrine, which is similar to the dose used here.

      We thank the reviewer for pointing this out, but we would like to clarify that epinephrine did significantly activate alpha-cell Ca<sup>2+</sup> activity in our experiments, as shown in Fig. 3F. This result is consistent with our previous study by Sluga et al., where low nanomolar epinephrine robustly activated alpha cell Ca<sup>2+</sup> signals in the pancreatic slice preparation. The main point of the present comparison was therefore not that epinephrine is inactive in alpha cells, but that AVP produces a substantially stronger and reproducible alpha cell Ca<sup>2+</sup> response under comparable experimental conditions. This is also consistent with the data of van der Meulen et al., supporting the view that AVP/V1bR signaling is a particularly potent activator of alpha cell activity. We have revised the text to make this comparison clearer and to avoid the impression that epinephrine failed to activate alpha cells in our preparation.

      (6) Figure 8 suggests a pharmacological activation of beta cell V1bR in the low pM range. How do the authors reconcile this comparison with the apparent absence of an effect of AVP stimulation at low pM to low nM doses in beta cells (Figure 4A)? I note that there are changes over time with sustained beta cell stimulation with 8 mM glucose, but these changes are relatively subtle, gradual, and quite likely represent the progression of calcium behaviors that would have occurred under sustained glucose, irrespective of these very low AVP concentrations. I will note that the Kd of the V1bR for AVP is around 1 nM, with tracer displacement starting around 100 pM according to the data in figure 5B, which is hard to reconcile with changes in beta cell calcium by AVP doses that start 10-100-fold lower than this dose at 1 and 10 pM (Figure 8).

      We agree that the interpretation of low-pM AVP effects requires caution, particularly when compared with reported V1bR binding affinities. The apparent discrepancy between Fig. 4A and Fig. 8 most likely reflects differences in experimental design, stimulation context, and readout sensitivity. In Fig. 4A, we assessed acute AVP effects under conditions in which beta cell Ca<sup>2+</sup> responses are relatively threshold-dependent and where low AVP concentrations produced little or no activation. In contrast, Fig. 8 analyzes prolonged beta cell population dynamics during sustained stimulation with 8 mM glucose, a physiological stimulatory context in which even weak modulatory inputs may become detectable at the level of collective Ca<sup>2+</sup> activity.

      Importantly, the strongest and statistically significant effect was observed at 100 pM AVP, while lower pM concentrations showed only a trend. We therefore do not interpret the low-pM range as evidence for robust direct pharmacological activation of beta cell V1bR. Rather, these data suggest that AVP may exert permissive or modulatory effects within an already active beta cell network, where glucose-dependent excitability, receptor-effector coupling, and Ca<sup>2+</sup> amplification mechanisms can enhance the apparent efficacy of weak inputs. This interpretation is consistent with the known permissive role of AVP in endocrine responses, where AVP may not act as a primary activator alone but can increase the efficacy of other physiological stimuli.

      We have also clarified that sustained 8 mM glucose alone does not account for these effects, since Suppl. Fig. 1 shows no comparable time-dependent progression of Ca<sup>2+</sup> behavior under sustained glucose stimulation alone. Thus, we now present the low-concentration AVP effects as subtle, contextdependent modulation within the physiological stimulatory range, rather than as evidence for direct beta cell activation at concentrations below the expected receptor affinity range.

      Reviewer #3 (Public review):

      Summary:

      This work aims to better understand the role of arginine vasopressin (AVP) in the control of islet hormone secretion. This builds on previous literature in this area reporting on the actions of AVP to stimulate islet hormones. The gap in literature being addressed by these studies is primarily focused on the glucose-dependency of AVP on both insulin and glucagon secretion. A secondary objective is to explore the role of individual receptors with the use of newly generated peptides and existing tools. The methods include the use of Ca2+ imaging in pancreas slices from mice, with additional outcomes including insulin secretion in some areas. The conclusions presented are that AVP acts through V1b receptors in both alpha- and beta-cells, that this activity occurs in the high cAMP environment, and is glucose dependent.

      Strengths:

      The area of research is emerging with plenty of room for new contributions. The concept of AVP stimulating islet hormone secretion is important and deserving of further insight. The use of pancreas tissue to image primary cells makes the experiments physiologically relevant. The advancement of novel tools in this area should be helpful to other groups investigating the actions of AVP.

      We would like to thank the reviewer for recognizing the potential of our emerging area of research.

      Weaknesses:

      The conclusions are only modestly supported by the data and lack experimental depth and rigor. The rationale for only conducting studies at high cAMP conditions is not entirely clear and limits the conclusions that can be made. The use of Ca2+ is helpful, but it is a surrogate for hormone secretion. Additional measurements of hormone secretion are needed to enhance the robustness of these conclusions. Consideration of paracrine effects between alpha- and beta-cells is only superficially made and is likely essential in the context of the experimental design. For instance, there is clear literature that alpha-cells secrete several factors that work in paracrine interactions on beta-cells and autocrine actions back on alpha-cells. Conducting these studies in a high cAMP context only completely overlooks these interactions, skewing the interpretations made by the investigators. Finally, the clarity of the experiments and results could be significantly enhanced.

      We thank the reviewer for this balanced assessment and for emphasizing several issues that are central to the interpretation of our study. We agree and now explicitly state in the Limitations, that Ca<sup>2+</sup> oscillations are a surrogate readout for hormone secretion and that currently used stimulation protocols are not optimized to directly quantify the relationship between Ca<sup>2+</sup> dynamics and secretory output. To address this limitation, we have now expanded the functional part of the study by adding complete insulin and glucagon secretion measurements during AVP concentration ramps. These new data provide a stronger functional framework for interpreting the Ca<sup>2+</sup> imaging results, while also clarifying that Ca<sup>2+</sup> activity and secretion cannot be assumed to correlate linearly under all stimulation protocols.

      We have also revised the rationale for the high-cAMP experimental condition. The original aim was to reduce variability arising from glucose-dependent endogenous cAMP signaling and to study alpha and beta cell responses in a common permissive background. However, we agree that broad forskolin stimulation complicates the interpretation of cell-autonomous versus paracrine mechanisms. Independent experiments within a scope of another study demonstrate that using GLP-1 co-stimulation, which provides a more beta-cell-oriented cAMP-permissive condition and supports the interpretation that AVP can modulate beta-cell collective activity in a manner that is not solely secondary to alpha-cell activation.

      We fully agree that paracrine interactions within the islet are physiologically important and must be considered, particularly in intact pancreatic slices. Nevertheless, the rapid onset of the AVP effects observed in beta cell Ca<sup>2+</sup> activity argues against a mechanism mediated predominantly by slower indirect paracrine loops. High AVP concentrations, as shown in Fig. 5, significantly shorten the intervals between Ca<sup>2+</sup> events in both alpha and beta cells, but that the activity of the two cell populations remains largely noncoordinated. This temporal dissociation does not exclude paracrine modulation altogether, but it argues against a simple alpha-cell-driven explanation for the beta-cell response.

      We have revised the manuscript to state these points more clearly and to moderate conclusions where the data support modulation rather than definitive cell-autonomous receptor action. We believe that the added secretion experiments and alpha/beta event-timing analysis substantially strengthen the physiological interpretation of the study, and we thank the reviewer for raising these issues.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The paragraph discussing the benefits of slice physiology over islets is not reflective of how most - if not all of your colleagues who do islet experiments conduct these. Many labs have reported for years high-quality GSIS experiments, synchronous calcium responses, and a plethora of studies detailing the mechanism of hormone and neurotransmitter actions using islet models, and have done so well. Slice physiology is a unique and helpful model that can have advantages over other models. This particular reviewer uses both models in their lab and each has benefits and - inevitably - drawbacks. Many of the possible drawbacks cited for islet studies apply equally to slices, including the possibility of altered gene expression, lack of innervation, and circulation. Added drawbacks are the exposure to higher levels of pancreatic enzymes from the slice, which require co-culture with enzyme inhibitors.

      We agree with the reviewer and have revised these limitations accordingly. Our intention was not to imply that isolated islet preparations are generally inferior, since they have provided a highly productive and rigorous experimental platform for GSIS, synchronized Ca<sup>2+</sup> dynamics, and mechanistic studies of hormonal and neurotransmitter regulation with standardized protocols with all their positive and negative sides. We modified the presentation of pancreatic slices as a complementary model with specific advantages, particularly preservation of local tissue architecture, while also acknowledging their limitations. The revised text therefore avoids a comparative hierarchy between slices and isolated islets and instead emphasizes that both models have distinct strengths and drawbacks depending on the experimental question.

      (2) If you want to demonstrate direct actions on beta cells, deconstructing the islet would be a better way to go. Less complicated, not more. Dissociated beta cells, instead of slices, were used just to prove or disprove the hypothesis of direct beta cell effects of AVP.

      We agree that dissociated beta cells can be a useful reductionist model to test whether AVP is capable of acting directly on individual beta cells. However, this approach would also remove the collective beta-cell activity that is central to the physiological question addressed in the present study. Since our data indicate that AVP effects emerge within the intact islet as rapid and extensive changes in coordinated Ca<sup>2+</sup> dynamics, dissociation would not necessarily provide a more informative model for understanding these responses. The fast onset and magnitude of the beta cell response argue against a predominantly indirect non-autonomous mechanism, and this interpretation is further supported by the GLP-1 co-stimulation experiments, which are more consistent with beta cell-autonomous modulation. We have therefore clarified in the revised manuscript that dissociated-cell experiments would be valuable for a narrowly defined receptor-cell autonomy question, but would not resolve the collective islet dynamics that are the focus of this work.

      (3) If you want to sustain the claim of beta cell expression of V1br, you would have to demonstrate this far more directly by staining (if appropriate antibodies exist), by beta cell-specific deletion of V1br, or by highly selective, well-validated pharmacology. This should include a demonstration of Gaqdependence in isolated beta cells.

      We have added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      Minor

      (1) O'Carroll et al. should be cited in the context of islet permissive actions of AVP/cAMP. PMID: 18434353, although that paper offers no evidence that the AVP-dependent potentiation of insulin release is mediated directly by beta cells. It does confirm dependence on PKC.

      We agree and have added O’Carroll et al. in the revised manuscript in the context of AVP/cAMP-dependent permissive actions on islet hormone secretion. We therefore use it as support for the broader concept of AVP-dependent amplification of secretion in a permissive signaling context, rather than as direct evidence for beta-cell-autonomous V1bR signaling.

      (2) Figure 2 E-H, glucose concentration mislabeled.

      We thank the reviewer for pointing this out. The glucose concentration label has been clarified: panels E–H show pooled data from separate experiments performed at 8 mM glucose, whereas panel D shows a representative experiment performed at 9 mM glucose.

      (3) The insulin secretion in 4E is difficult to interpret without a low-glucose control. If this is hard to do in a slice preparation, a separate static islet secretion experiment would help here. The possibility that the inhibition of insulin secretion traces back to the activation of delta cells by AVP could be considered - I struggle to come up with a plausible mechanistic explanation why AVP (which activates calcium in alpha and in beta cells in the presence of 8 mM G plus forskolin according to your data) would inhibit insulin secretion.

      We agree that the original insulin secretion experiment was difficult to interpret without a clearer low-glucose reference condition. To address this, we have added new insulin release experiments in which glucose was lowered to a non-stimulatory range between AVP concentrations, followed by sequential stimulation with 8 mM glucose and 500 nM forskolin in the same slices. These new data provide a broader dynamic range for assessing insulin secretion and allow a more direct comparison between AVP-dependent Ca<sup>2+</sup> modulation and secretory output.

      We also agree that AVP-dependent inhibition of insulin secretion requires careful interpretation. One possible explanation is not simply activation of delta cells, but a failure of the beta cell collective to maintain coordinated activity at very high AVP concentrations. In the Ca<sup>2+</sup> imaging data, high AVP concentrations increase activity in many beta cells, but numerous cells within the islet fail to keep pace with the collective oscillatory rhythm, leading to fragmented and less synchronized population activity. Thus, despite increased frequency of Ca<sup>2+</sup> oscillations in the islet, the integrated beta cell output may become less efficient for insulin secretion. We have added this interpretation to the revised manuscript and now discuss delta cell activation as a possible contributing mechanism, but not as the primary explanation supported by our current data.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the referees for noting the substantive revisions and for the praise of the work. While we each have somewhat different weightings of likelihood, we feel the appraisals are fair and reasonable.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors addressed an important biological question, namely the role of glutamine metabolism in humoral responses, and they obtained solid conclusions. The strength of this study is that the authors used state-of-the-art transgenic mouse models together with in vitro analysis, thereby providing significant insights into the question posed. The following would strengthen the manuscript: i) adding more in-depth functionality/physiological relevance in the discussion part, and ii) regarding the experiments, the inclusion of more appropriate controls and a clearer and more accurate description of the methods.

      We are grateful for the decision of the Editors to select this submission for in-depth peer review and to the Reviewing Editor and referees for the thoughtful and constructive comments.

      We mostly agree with the specific comments and evaluation of strengths of what the work adds as well as with indications of limitations and caveats that apply to the breadth of conclusions. We have edited the text to be more clear and provide more details about certain aspects of the Methods and Legends. In addition, although we try to avoid Discussion sections that are unduly long or have flights of fancy, we will add to the Discussion as well as edit it for directness about potential relevance, basic explorations of mechanisms, and functionality.

      The revised manuscript also contains new data, some of it dealing with comments of the referees, other additions representing work done while the manuscript was under review. While we would be inclined to do more, the sad practical problem is one of limits placed by both the absence of any grant funds and the institution's terminations (RIFs) of the two experimenters in the lab.

      While we believe the original data interpretable as presented originally, up to a point it nonetheless is good to enhance scope or have even better data and add refinements about some of the technical issues. Ultimately, the question becomes "when is enough enough?"

      In the detailed point-by-point response below, we outline changes prompted by the reviewers. We also comment on a few points more expansively that would be suitable for the paper itself, and offer some skepticism or disagreement, (longer and more detailed explanations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Cho et al. present a comprehensive and multidimensional analysis of glutamine metabolism in the regulation of B cell differentiation and function during immune responses. They further demonstrate how glutamine metabolism interacts with glucose uptake and utilization to modulate key intracellular processes. The manuscript is clearly written, and the experimental approaches are informative and well-executed. The authors provide a detailed mechanistic understanding through the use of both in vivo and in vitro models. The conclusions are well supported by the data, and the findings are novel and impactful. I have only a few, mostly minor, concerns related to data presentation and the rationale for certain experimental choices.

      Detailed Comments:

      (1) In Figure 1b, it is unclear whether total B cells or follicular B cells were used in the assay. Additionally, the in vitro class-switch recombination and plasma cell differentiation experiments were conducted without BCR stimulation, which makes the system appear overly artificial and limits physiological relevance. Although the effects of glutamine concentration on the measured parameters are evident, the results cannot be confidently interpreted as true plasma cell generation or IgG1 class switching under these conditions. The authors should moderate these claims or provide stronger justification for the chosen differentiation strategy. Incorporating a parallel assay with anti-BCR stimulation would improve the rigor and interpretability of these findings.

      We edited the manuscript to be clear that total splenic B cells were used in this set-up figure and the rest of the paper. In addition, we performed new experiments to improve this "set-up figure (Fig. 1)" and moved the older data using alternative experimental conditions to a supplemental figure, Figure 1 - supplement 1. We also used new conditions that included styles of stimulating proliferation and differentiation - to foster an increased sense of generality. The findings in no way change the supported conclusions of the work. Specifically, we used mitogenic stimulation with anti-IgM <sup>+</sup> anti-CD40, all with BAFF, IL-4, and IL-5 in addition to the anti-CD40 stimulation of the original manuscript, bearing in mind excellent work from Aiba et al, Immunity 2006; 24: 259-268, and similar papers. In addition, we added a panel with representative flow cytometric profiles. These new data are presented in Figure 4 - supplement 1 (panels ae).

      To be transparent and add to a more open public discussion (using the virtues of this forum), the senior author and colleagues would caution about whether any in vitro conditions exist that warrant complete confidence. That is the reason for proceeding to immunization experiments in vivo. That is not said to cast doubt on our own in vitro data - there are some experiments (such as those of Fig. 1a-c and associated Fig 1 - supplement 1) that only can be done in vitro or are better done that way (e.g., because of rapid uptake of early apoptotic B cells in vivo).

      For instance: Well-respected papers use the CD40LB and NB21.2D9 systems to activate B cells and generate plasma cells. Those appear to be BCR-independent and yet continue in common use. [We found that these cellular systems (CD40LB; NB21.2D9) cannot be used in experiments with a.a. deprivation or the inhibitors due to effects on the engineered stroma-like cells.] In considering BCR engagement, Reth has published salient points about signaling and concentrations of the Ab, the upshot being that this means of activating mitogenesis and plasma cell differentiation (when the B cells are costimulated via CD40 or TLR (4 or 7/8) is also artificial. Moreover, although Aiba et al, Immunity 2006; 24: 259-268 is a laudable exception, one rarely finds papers using BAFF despite the strong evidence it is an essential part of the equation of B cell regulation in vivo and a cytokine that modulates BCR signaling - in the cultures.

      (2) In Figure 1c, the DMK alone condition is not presented. This hinders readers' ability to properly asses the glutaminolysis dependency of the cells for the measured readouts. Also, CD138<sup>+</sup> in developing PCs goes hand in hand with decreased B220 expression. A representative FACS plot showing the gating strategy for the in vitro PCs should be added as a supplementary figure. Similarly, division number (going all the way to #7) may be tricky to gate and interpret. A representative FACS plot showing the separation of B cells according to their division numbers and a subsequent gating of CD138 or IgG1 in these gates would be ideal for demonstrating the authors' ability to distinguish these populations effectively.

      In the revised manuscript, we have added new experimental data (Figure 1).

      We agree that exact placement of divisions and deconvolution by FlowJow is more fraught than might be thought from presentations in many or most papers. We include the data shown to the right as representative FACS plot(s) with old and new data that illustrate the gating on CTV fluorescence. With the representative examples pasted in here and presented in Fig 1 - supplement 1f, g of the revised manuscript, we will aver that using divisions 0-6, and ≥7 was and is entirely reasonable.

      Ditto for DMK with normal glutamine. However, in the spirit of eLife transparency lacking in many other journals, this comparison is more fraught than the referee comment would make things seem. The concentration tolerated by cells is highly dependent on the medium and glutamine concentration, and perhaps on rates of glutaminolysis (due to its generation of ammonia). In practice, DMK becomes more toxic to B cells unless glutamine is low or glutaminolysis is restricted. Thus, the concentration of DMK that is tolerated and used in Fig. 1b, c can become toxic to the B cells when using the higher levels of glutamine in typical culture media (2 mM or more) - at which point the "normal conditions <sup>+</sup> DMK" "control" involves the surviving cells in conditions with far greater cell death and less population expansion than the "low glutamine <sup>+</sup> DMK". condition.

      (3) A brief explanation should be provided for the exclusive use of IgG1 as the readout in classswitching assays, given that naïve B cells are capable of switching to multiple isotypes. Clarifying why IgG1 was preferentially selected would aid in the interpretation of the results.

      On lines ~112-3 and ~182-5, we edited the text in light of the referee's suggestion that we focus the presentation of serologic data on IgG1 in the immunization experiments. We also rearranged figures and panels to be more explicit and harmonize. That said, and [Brief explanation - IgG1 provides the strongest signal and hence better signal/noise both in vitro and with the alum-based immunizations that are avatars for the adjuvant used in the majority of protein-based vaccines for humans. Perhaps for this reason, the majority of papers on molecular mechanisms seem only to analyze IgG1. Nonetheless, since molecular regulation can differ according to isotype, and the more pro-inflammatory mouse IgG2c is more pertinent to some forms of anti-pathogen immunity and some auto-immune disease models, we believe it valuable to retain these data in supplements to the related Figures.]

      (4) The immunization experiments presented in Figures 1 and 2 are well designed, and the data are comprehensively presented. However, to prevent potential misinterpretation, it should be clarified that the observed differences between NP and OVA immunizations cannot be attributed solely to the chemical nature of the antigens - hapten versus protein. A more significant distinction lies in the route of administration (intraperitoneal vs. intranasal) and the resulting anatomical compartment of the immune response (systemic vs. lung-restricted). This context should be explicitly stated to avoid overinterpretation of the comparative findings.

      We appreciate the positive assessment, and agree with the referee that it is possible the conditions of immune challenge or re-exposure may contribute to the observed differences. We edited the text of the revised manuscript accordingly [lines ~152-153; ~159-160]. Certainly, the difference in how the anti-ova response is elicited compared to the anti-NP response in the same mice or with a bit different an immunization regimen might be another factor - or the major factor - explaining why glutaminolysis was important after ovalbumin inhalations (used because emergence of anti-ova Ab / ASCs is suppressed by the NP hapten after NP-ova immunization) but not needed for the anti-NP response unless Slc2a1 or Mpc2 also was inactivated. Thank you prompting addition of this important caveat!

      Nevertheless, it seems fair to note that in Figures 1 and 2, the ASCs and Ab are being analyzed for NP and ova in the same mice, albeit with the NP-specific components not being driven by the inhalations of ovalbumin. With that in mind, when one compares the IgG1 anti-NP ASC and Ab to those for IgG1 anti-ovalbumin (ASC in bone marrow; Ab), the ovalbumin-specific response was reduced whereas the anti-NP response was not. [lines ~171-172]

      (5) NP immunization is known to be an inducer of an IgG1-dominant Th2-type immune response in mice. IgG2c is not a major player unless a nanoparticle delivery system is used. However, the authors arbitrarily included IgG2c in their assays in Figures 2 and 3. This may be confusing for the readers. The authors should either justify the IgG2c-mediated analyses or remove them from the main figures. (It can be added as supplemental information with proper justification).

      We rearranged the Figure panels to move IgM and IgG2c data to Supplemental Figures (Figure 3 - supplements 1, 2, 4, 5 in the eLife system).

      For purposes of public discourse, we note first that in contrast to the premise about weak IgG2c responses, the data [previously, Figure 3(c, g); now in the supplements] show substantial levels of NP-specific IgG2c. The referee is quite right that the class switching and in vitro ASC generation were done with IL-4 / IgG1-promoting conditions.

      To assist readers, the revised manuscript takes note of the important role of IgG2c (mouse - IgG1 in humans) in controlling or clearing various pathogens as well as in autoimmunity [lines ~182-5]. Moreover, we continue to think that these measurements add substantial value both from the standpoint of providing a better sense of generality to the loss-of-function effects, and in considering potential ways of translating the findings to B cell-dependent autoimmune conditions such as systemic lupus erythematosus.

      [As a scientific aside, we speculate that a greater or lesser IgG2c anti-NP response may arise due to different preparations of NP-carrier obtained from the vendor (Biosearch) having different amounts of TLR (e.g., TLR4) ligand. In any case, the points of presenting the IgG2c (and IgM) data were to push against the limiting boundaries of convention (which risks perpetuating a narrow view of potential outcomes) and make the breadth of results more apparent to readers.

      (6) Similarly, in affinity maturation analyses, including IgM is somewhat uncommon. I do not see any point in showing high affinity (NP2/NP20) IgMs (Figure 3d), since that data probably does not mean much.

      As noted in the reply immediately preceding this one, we appreciate this suggestion from the reviewer and moved the IgM and IgG2c to supplemental status.

      Nonetheless, in collegial discourse we disagree a bit with the referee in light of our data as well as of work that (to our minds) leads one to question why inclusion of affinity maturation of IgM is so uncommon - as the referee accurately notes. Of course a defect in the capacity to class-switch is highly deleterious in patients but that is not the same as concluding that recall IgM or its affinity is of little consequence.

      In some of the pioneering work back in the 1980's, Bothwell showed that NP- carrier immunization generated hybridomas producing IgM Ab with extensive SHM (~11% of the 18 lineages; ~ 1/3 of the IgM hybridomas) [PMID: 8487778], IgM B cells appear to move into GC, and there is at least a reasonable published basis for the view that there are GC-derived IgM (unswitched) memory B cells (MBC) that would be more likely, upon recall activation, to differentiate into ASCs. [As an example, albeit with the Jenkins lab anti-rPE response, Taylor, Pape, and Jenkins generated quantitative estimates of the numbers of Ag-specific IgM<sup>+</sup> vs switched MBC that were GC-derived (or not). [PMID: 22370719]. While they emphasized that ~90% of IgM<sup>+</sup> MBC appeared to be GC-independent, their data also indicated that ~1/2 of all GC-derived MBC were IgM<sup>+</sup> rather than switched (their Fig. 8, B vs C; also 8E, which includes alum-PE). And while we immensely respect the referee, we are perhaps less confident that IgM or high-affinity Ag-specific IgM doesn't mean that much, if only because of evidence that localized Ab compete for Ag and may thus influence selective processes [PMCID: PMC2747358; PMID: 15953185; PMID: 23420879; PMID: 27270306].

      (7) Following on my comment for the PC generation in Figure 1 (see above), in Figure 4, a strategy that relies solely on CD40L stimulation is performed. This is highly artificial for the PC generation and needs to be justified, or more physiologically relevant PC generation strategies involving anti-BCR, CD40L, and various cytokines should be shown.

      In line with our response to point (1), we tested BCR-stimulated B cells (anti-CD40 plus anti-IgM with BAFF, IL-4, and IL-5, parallel to the analyses with anti-CD40 but no BCR engagement). These results align with and reinforce the utility of the data with anti-CD40 as the sole mitogen.

      (8) The effects of CB839 and UK5099 on cell viability are not shown. Including viability data under these treatment conditions would be a valuable addition to the supplementary materials, as it would help readers more accurately interpret the functional outcomes observed in the study.

      We added presentation of data that provide cues as to relative viability / cxmsurvival under the experimental conditions used.

      [FSC X SSC as well as 7AAD or Ghost dye panels; we also generated new data that in[ further experiments scoring annexin V staining (see Fig 4 - supplement 1d, e, and Fig 5 - supplement 1e, f)].

      (9) It is not clear how the RNA seq analysis in Figure 4h was generated. The experimental strategy and the setup need to be better explained.

      Including text added at lines ~291-293 and ~582-585, the revised manuscript provides more information in the Results, Methods and Legend for Fig 4j-l. We agree entirely with the concern and apologize that in this and a few other instances we inadvertently sacrificed sufficiency of detail on the altar of attempting brevity.

      [As a synopsis: In three temporally and biologically independent experiments, cultures were harvested 3.5 days after splenic B cells were purified and cultured as in the experiments of Fig. 4a-e. Total cellular RNA was prepared from the twelve samples (three replicates for each of four conditions - DMSO vehicle control, CB839, UK5099, and CB839 <sup>+</sup> UK5099), then analyzed by RNA-seq. RNA-seq data were initially processed using the pipeline described in the Methods. For panels g & h of Fig 4, DESeq2 was used to quantify and compare read counts in the three CB839 <sup>+</sup> UK5099 samples relative to the three independent vehicle controls and identify all genes for which variances yielded P<0.05. In Fig 4g, all such genes for which the difference was 'statistically significant' (i.e., P<0.05) were entered into the indicated Immgen tool and thereby mapped to the B lineage subsets shown in the figure panels (i.e., g, h). In (g), these are displayed using one format, whereas (h) uses the 'heatmap' tool in MyGeneSet.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigate the functional requirements for glutamine and glutaminolysis in antibody responses. The authors first demonstrate that the concentrations of glutamine in lymph nodes are substantially lower than in plasma, and that at these levels, glutamine is limiting for plasma cell differentiation in vitro. The authors go on to use genetic mouse models in which B cells are deficient in glutaminase 1 (Gls), the glucose transporter Slc2a1, and/or mitochondrial pyruvate carrier 2 (Mpc2) to test the importance of these pathways in vivo.

      Interestingly, deficiency of Gls alone showed clear antibody defects when ovalbumin was used as the immunogen, but not the hapten NP. For the latter response, defects in antibody titers and affinity were observed only when both Gls and either Mpc2 or Slc2a1 were deleted. These latter findings form the basis of the synthetic auxotrophy conclusion. The authors go on to test these conclusions further using in vitro differentiations, Seahorse assays, pharmacological inhibitors, and targeted quantification of specific metabolites and amino acids. Finally, the authors document reduced STAT3 and STAT1 phosphorylation in response to IL-21 and interferon (both type 1 and 2), respectively, when both glutaminolysis and mitochondrial pyruvate metabolism are prevented.

      Strengths:

      (1) The main strength of the manuscript is the overall breadth of experiments performed. Orthogonal experiments are performed using genetic models, pharmacological inhibitors, in vitro assays, and in vivo experiments to support the claims. Multiple antigens are used as test immunogens--this is particularly important given the differing results.

      (2) B cell metabolism is an area of interest but understudied relative to other cell types in the immune system.

      (3) The importance of metabolic flexibility and caution when interpreting negative results is made clear from this study.

      Weaknesses:

      (1) All of the in vivo studies were done in the context of boosters at 3 weeks and recall responses 1 week later. This makes specific results difficult to interpret. Primary responses, including germinal centers, are still ongoing at 3 weeks after the initial immunization. Thus, untangling what proportion of the defects are due to problems in the primary vs. memory response is difficult.

      We performed new experiments and added the data on differences prior to a boost [see below; new Fig 3d, e; etc].

      (2) Along these lines, the defects shown in Figure 3h-i may not be due to the authors' interpretation that Gls and Mpc2 are required for efficient plasma cell differentiation from memory B cells. This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. The more likely interpretation is that ongoing primary germinal centers are negatively impacted by Gls and Mpc2 deficiency, and this, in turn, leads to reduced affinities of serum antibodies.

      We have edited the wording of the conclusion to add a possibility we consider unlikely and downplay a conclusion that MBCs bearing switched BCRs are affected once reactivated. [see lines ~221-230] We also have added citations pertaining to the topic, including work from the Victora lab which seems to put the point succinctly: "Recall GCs in mice consist almost entirely of naïve B cells, whereas recall antibodies derive overwhelmingly from memory B cells." [emphasis added] [PMID: 38838672; new ref #83]. While unclear as to the reasoning - as one looks at the data - and skeptical as to the accuracy of the referee's point (2), it suggests that the matter is open to reasonable doubt. In line with the point and the edits, we also have added citation of a bioRxiv preprint from the Victora lab, which touches on the concept of what one could call boost-induced reinvigoration of a pre-existing GC [new ref #82].

      Beyond the textual changes, we performed a new series of experiments to investigate partially, and present the results in Fig 3d, e as well as Fig 3 - supplement 1d, e. Unfortunately, time before lab closure was an enemy both for the period between primary and recall immunizations in performance and multiple replication of work to extend that presented in Figure 3, panels g & h, and the related Supplemental Data (Fig 3 - supplements 4d, 5a-g). Unfortunately, it was not possible to do a longer-term memory experiment with recall immunization out at 8 weeks.

      The intriguing concerns and questions of points 1 & 2 provide a springboard for consideration of generalizations and simplifications. Germinal center durability is not at all monolithic, and instead is quite variable**. It is true that in the literature (especially with the substantially different approach of transferring BCR-transgenic / knock-in versions of an NP-biased BCR) there may be meaningful pools of IgG1 and IgG2c GC B cells. The premise (cognitive bias, perhaps?) in our interpretation is that in our previous work we measured few if any GC B cells - NP-APC-binding or otherwise - above the background (non-immunized controls) three weeks after immunization with NP-ovalbumin in alum. While recognizing that the immunogen can matter, we note for the readers and referee that Fig. 1 of the Taylor, Pape, & Jenkins paper considered above [PMID: 22370719] reported 10-fold more Ag-specific MBCs than GC B cells at day 29 post-immunization (the point at which the boost/recall challenge was performed in our Figure 3g, h. [That work did not use NP-carrier in alum to immunize, or measure the anti-NP response.]

      Viewing Fig. 3i from that perspective, the surmise of the comment is that a major contribution to the differences in both all-affinity and high-affinity anti-NP IgG1 (whose production requires differentiation into plasma cells) derived from the immunization at 4 wk stimulating persistent GC B cells as opposed to memory B cells.

      The issue and question also relate to rates of output of plasma cells or rises in the serum concentrations of class-switched Ab. To this point, our prior experiences agree with the long-published data of the Kurosaki lab in Figure 3c of the Aiba et al paper noted above (Immunity, 2006) (and other such time courses). Readers can note that the IgG1 anti-NP response (alum adjuvant, as in our work) hits its plateau at 2 wk, and did not increase further from 2 to 3 wk. The most likely interpretation is that GC are on the decline and Ab production has reached its plateau by the time of the 2nd immunization in Fig. 3h.

      Assuming we understand the comment and line of reasoning correctly, we also lean towards disagreeing with the statement " This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. Our evidence shows that both low-affinity as well as high-affinity anti-NP Ab (IgG1) were reduced due to combined gene-inactivation after the peak primary response (Fig. 3h; also, see the new data in Fig 3 and Fig 3 - supplement 1). Recent papers show that affinity maturation is attributable to greater proliferation of plasmablasts with high-affinity BCR. Accordingly, the findings with loss of GLS and MPC function are quite consistent with the interpretation that much of the response after the second immunization draws on MBC differentiation into plasmablasts and then plasma cells, where the proliferative advantage of high-affinity cells is blunted by the impaired metabolism. Notwithstanding these issues, the revised manuscript includes the alternative, if less likely, interpretation proposed by the review [lines ~221-230].

      **In some contexts, of course, especially certain viral infections or vaccination with lipid nanoparticles carrying modified mRNA, germinal centres are far more persistent; also, in humans even the seasonal flu vaccine

      (3) The gating strategies for germinal centers and memory B cells in Supplemental Figure 2 are problematic, especially given that these data are used to claim only modest and/or statistically insignificant differences in these populations when Gls and Mpc2 are ablated. Neither strategy shows distinct flow cytometric populations, and it does not seem that the quantification focuses on antigen-specific cells.

      The revised manuscript improves these aspects of the presentation, using old and new data. See Fig 3 - supplement 3a, c; Fig 3 - supplement 4a. We note for readers that many other papers in the best journals show plots in which the separation of, say, GC-Tfh from overall Tfh is based on cut-off within what essentially is a continuous spectrum of emission as adjusted or compensated by the cytometer (spectral or conventional).

      The revised manuscript presents results from new experiments that deal with the subset of GC B cells whose BCRs bind NP-APC with enough affinity to retain a positive signal after washing. These new data are presented in Fig 3 - supplement 3c & 3e. In practice, the new findings suggest that the metabolic requirement applied more to the NP-binding B cells than the overall GC B cell population.

      (4) Along these lines, the conclusions in Figure 6a-d may need to be tempered if the analysis was done on polyclonal, rather than antigen-specific cells. Alum induces a heavily type 2-biased response and is not known to induce much of an interferon signature. The authors' observations might be explained by the inclusion of other ongoing GCs unrelated to the immunization.

      We apologize for ambiguity or insufficient clarity and, as noted above, have edited the text to be more clear that the in vitro experiments do not represent GC B cells and that the RNA-seq data were from experiments that did not involve alum and were not an Ag (SRBC)-specific subset.

      New text in the Results, an expanded Legend, and tweaking the Methods make it more readily clear that the RNA-seq data (and hence the GSEA) involved immunizations with SRBC (not the alum / NP system. That said, we note that the hapten-carrier experiments in which the immunogen was adjuvantized with alum actually generated a robust IgG2c (type 1-driven) response along with the type 2-enhanced IgG1 response, in line with what has been reported by others with alum-adjuvanted vaccination.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors investigate how glutaminolysis (GLS) and mitochondrial pyruvate import (MPC2) jointly shape B cell fate and the humoral immune response. Using inducible knockout systems and metabolic inhibitors, they uncover a "synthetic auxotrophy": When GLS activity/glutaminolysis is lost together with either GLUT1-mediated glucose uptake or MPC2, B cells fail to upregulate mitochondrial respiration, IL 21/STAT3 and IFN/STAT1 signaling is impaired, and the plasma cell output and antigen-specific antibody titers drop significantly. This work thus demonstrates the promotion of plasma cell differentiation and cytokine signaling through parallel activation of two metabolic pathways. The dataset is technically comprehensive and conceptually novel, but some aspects leave the in vivo and translational significance uncertain.

      Strengths:

      (1) Conceptual novelty: the study goes beyond single-enzyme deletions to reveal conditional metabolic vulnerabilities and fate-deciding mechanisms in B cells.

      (2) Mechanistic depth: the study uncovers a novel "metabolic bottleneck" that impairs mitochondrial respiration and elevates ROS, and directly ties these changes to cytokinereceptor signaling. This is both mechanistically compelling and potentially clinically relevant.

      (3) Breadth of models and methods: inducible genetics, pharmacology, metabolomics, seahorse assay, ELISpot/ELISA, RNA-seq, two immunization models.

      (4) Potential clinical angle: the synergy of CB839 with UK5099 and/or hydroxychloroquine hints at a druggable pathway targeting autoantibody-driven diseases.

      We agree and thank the referee for the positive comments and this succinct summary of what we view as contributions of the paper.

      Weaknesses:

      (1) Physiological relevance of "synthetic auxotrophy"

      The manuscript demonstrates that GLS loss is only crippling when glucose influx or mitochondrial pyruvate import is concurrently reduced, which the authors name "synthetic auxotrophy". I think it would help readers to clarify the terminology more and add a concise definition of "synthetic auxotrophy" versus "synthetic lethality" early in the manuscript and justify its relevance for B cells.

      We edited the Abstract, Introduction, and Discussion to try to do better on this score. Conscious of how expansive the prose and data are even in the original submission, we appear to have taken some shortcuts that we will try to rectify or at least mitigate. Thank you for highlighting this need to improve on key concepts !!

      Specifically, the revised text expands a bit on the notion that synthetic auxotrophy represents effects on differentiation that go beyond additional mechanisms of reducing division efficiency and a modest impact on selective death. [see the 10th - 11th lines in Abstract and lines ~84-85, Introduction] Even though decreased population expansion is observed and new evidence supports a model in which the altered metabolism contributes to enhanced death in vivo, at equal division numbers the frequency of CD138<sup>+</sup> progeny is lower once glutaminolysis and mitochondrial pyruvate are reduced by either genetic or pharmacological means.

      This comment of the review raises interesting semantic questions about what represents "physiological relevance". The fundamental point is to explore a basic science question - what, if any, are limits to metabolic flexibility? In principle, shouldn't B cells be able to use fatty acid metabolism to generate enough ATP and provide the backbones for biosynthesis during growth? Put a different way, the point is that a basic curiosity to understand why decreasing glucose influx did not have an even more profound effect than what was observed, combined with curiosity as to why glutaminolysis was dispensable in relatively standard vaccine-like models of immunize/boost, provided a springboard to identification of new vulnerabilities. The manuscript shows one physiological limitation (and hence vulnerability). Be that as it may, the revised text of the Discussion section more clearly addresses this issue (lines ~531-549 at the end of the Discussion).

      While the overall findings, especially the subset specificity and the clinical implications, are generally interesting, the "synthetic auxotrophy" condition feels a little engineered.

      CAR-T cells are 'a little engineered' (or more than a little) and yet they do seem to have had an impact on understanding the centrality of B cells in various autoimmune conditions as well as in the direction of cancer therapy research. So it is a matter of balancing this perspective of the referee against the strengths they highlight in points 1, 2, and 4. In editing the revision, we try to expand and be more explicit about this in the Discussion of the revised manuscript.

      In brief, even were the money not all gone, we would not believe that expanding the heft of this already rather large manuscript and set of data would be appropriate. As matters stand, a basic new insight about metabolic flexibility and its limits leads to evidence of a way to reduce generation of Ab and a novel impairment of STAT transcription factor induction by several cytokine receptors. The vulnerability that could be tested in later work on B cell-dependent autoimmunity includes the capacity to test a compound that already has been to or through FDA phase II in patients together with an FDA-approved standard-of-care agent.

      Therefore, the findings strongly raise the question of the likelihood of such a "double hit" in vivo and whether there are conditions, disease states, or drug regimens that would realistically generate such a "bottleneck".

      Hence, the authors should document or at least discuss whether GC or inflamed niches naturally show simultaneous downregulation/lack of glutamine and/or pyruvate. The authors should also aim to provide evidence that infections (e.g., influenza), hypoxia, treatments (e.g., rapamycin), or inflammatory diseases like lupus co-limit these pathways.

      Again, we appreciate some 'licensing' to be more expansive and explicit, and will try to balance editing in such points against undue tedium or tendentiously speculative length in the Discussion. In particular, we will note that a clear, simple implication of the work is to highlight an imperative to test CB839 in lupus patients already on hydroxychloroquine as standard-of-care, and to suggest development of UK5099 (already tested many times in mouse models of cancer) to complement glutaminase inhibition.

      As backdrop, we note that the failure to advance imaging mass spectrometry to the capacity to quantify relative or absolute (via nano-DESI) concentrations of nutrients in localized interstitia is a critical gap in the entire field. Techniques that sample the interstitial fluid of tumour masses or in our case LN as a work-around have yielded evidence that there can be meaningful limitations of glucose and glutamine, but it needs to be acknowledged that such findings may be very model-specific and, as can be the case with cutting-edge science, are not without controversy. That said, yes, we had found that hypoxia reduced glutamine uptake but given the norms of focused, tidy packages only reported on leucine in an earlier paper [PMID27501247; PMCID5161594].

      Beyond all that, another impetus to and inspiration for these experiments stems from quite data that we generated in a model of short-term protein-restricted diet (loosely akin to kwashiorkor in humans), based on an excellent publication showing that such a regimen quickly led to lower circulating glutamine and mTORC1 activity (**). In brief, we found that a low-protein diet did, in our experiments, preferentially lower glutamine but - importantly - led to reduced Ab responses (which would match what we have modeled here). The findings were not a well-enough connected evidentiary component to include in the "story" but I'll append slides with the relevant data to this Response to Reviews for the referee's perusal (and anyone else who reads this online discourse).

      It would hence also be beneficial to test the CB839 + UK5099/HCQ combinations in a short, proof-of-concept treatment in vivo, e.g., shortly before and after the booster immunization or in an autoimmune model. Likewise, it may also be insightful to discuss potential effects of existing treatments (especially CB839, HCQ) on human memory B cell or PC pools.

      We certainly agree that the suggestions offered in this comment are important next steps and the right approach to test if the findings reported here translate toward the treatment of autoimmune diseases that involve B cells, interferons, and pathophysiology mediated by auto-Ab. As practical points, performance and replication of such studies would take more time than the year allotted for return of a revised manuscript to eLife and in any case neither funds nor a lab remain to do these important studies.

      Concrete evidence for our concurrence was embodied in a grant application to NIH that was essential for keeping a lab and doing any such studies. [We note, as a suggestion to others, that an essential component of such studies would be to test the effects of these compounds on B cells from patients and mice with autoimmunity]. Perhaps unfortunately for SLE patients, the review panelists did not agree about the importance of such studies. However, it can be hoped that the patent-holder of CB839 (and perhaps other companies developing glutaminase inhibitors) will see this peer-reviewed preprint and the public dialogue, and recognize how positive results might open a valuable contribution to mitigation of diseases such as SLE.

      (2) Cell survival versus differentiation phenotype

      Claims that the phenotypes (e.g., reduced PC numbers) are "independent of death" and are not merely the result of artificial cell stress would benefit from Annexin-V/active-caspase 3 analyses of GC B cells and plasmablasts. Please also show viability curves for inhibitor-treated cells.

      This comment leads us to see that the wording on this point may have been overly terse in the interests of brevity, and thereby open to some odd misunderstanding. The CD138<sup>+</sup> events are scored among VIABLE CELLS, so a decrease in the %CD138<sup>+</sup> at similar division number represents an effect independent from (or beyond) survival and division-counting. Accordingly, we expanded the text of the Abstract and elsewhere in the manuscript, to be more clear. In addition, we added data from new experiments addressing death in vitro and among GC-phenotype B cells in vivo. To clarify in this public context, it is not that an increase in death (along with the reported decrease in cell cycling) can be or is excluded. The point is that beyond any such increase, and taking into account division number (since there is evidence that PC differentiation and output numbers involve a 'division-counting' mechanism), the frequencies of CD138<sup>+</sup> cells and of ASCs among the viable cells are lower, as is the level of Prdm1-encoded mRNA even before the big increase in CD138<sup>+</sup> cells in the population.

      (3) Subset specificity of the metabolic phenotype

      Could the metabolic differences, mitochondrial ROS, and membrane-potential changes shown for activated pan-B cells (Figure 5) also be demonstrated ex vivo for KO mouse-derived GC B cells and plasma cells? This would also be insightful to investigate following NP-immunization (e.g., NP+ GC B cells 10 days after NP-OVA immunization).

      We performed a series of new experiments to have enough biologically independent replications for meaningful and statistical analyses. The new results, added in as Fig 5 - supplement 1, showed that the combined pathway interruption by loss-of-function increased ROS, mtROS, and death (annexin V / 7AAD) upon analyzing GCphenotype B cells immediately upon harvest. The findings align well with the data in Fig 5 (cultured B cells).

      (4) Memory B cell gating strategy

      I am not fully convinced that the memory-B-cell gate in Supplementary Figure 2d is appropriate. The legend implies the population is defined simply as CD19+GL7-CD38+ (or CD19+CD38++?), with no further restriction to NP-binding cells. Such a gate could also capture naïve or recently activated B cells. From the descriptions in the figure and the figure legend, it is hard to verify that the events plotted truly represent memory B cells. Please clarify the full gating hierarchy and, ideally, restrict the MBC gate to NP+CD19+GL7-CD38+ B cells (or add additional markers such as CD80 and CD273). Generally, the manuscript would benefit from a more transparent presentation of gating strategies.

      In considering the referee's viewpoint, we further expanded the supplemental data displays to include more of the gating and analytic schemes, which we believe should mitigate one concern noted here. In addition, we now include flow data from the non-immunized control mice that had been analyzed concurrently in the experiments.

      Third and finally, we performed new experiments and analyses in which the focus was the frequencies of memory-phenotype (IgD<sup>neg</sup> GL7<sup>neg</sup> CD38<sup>+</sup> / CD38<sup>hi</sup> aka CD38<sup>+</sup><sup>+</sup>) NPbinding B cells after immunization. While this time, as opposed to previously, the NP-APC staining met our standard for interpretability, the gist of the findings was that the two independent repeat experiments yielded a split decision and a degree of variability. With time being up due to the funds running out, we have elected to delete the issue and the data panel in question.

      That said, it bears noting that in the previous figure panel, the labeling indicated that the gating included the important criterion that cells be IgD<sup>neg</sup>, which excludes the vast majority of naive B cells but measures memory-phenotype B cells independent from consideration of whether or not they were NP-binding.

      [In principle marginal zone (MZ) B cells might fall within this gate. However, the MZ B population is unlikely to explain the differences shown.

      (5) Deletion efficiency - [The] mRNA data show residual GLS/MPC2 transcripts (Supplementary Figure 8). Please quantify deletion efficiency in GC B cells and plasmablasts.

      Even were there resources to do this, the degree of reduction in target mRNA (Gls; Mpc2) renders this question superfluous. To the best of our understanding, the proteins (for which there might be some phenotypic lag) are translated from RNA. Might there be a small subpopulation of B cells (or their PC progeny) with only one, or even neither, allele converted from fl to D? Yes, but they would be a minor subset in light of the magnitude of mRNA reduction, in contrast to our published observations with Slc2a1. As to plasmablasts and plasma cells, the pre-existing populations make such an analysis misleading, while the scarcity of such cells recoverable with antigen capture techniques is so low as to make both RNA and genomic DNA analyses questionable. We also refer readers to the supplemental figure that presents the results of experiments testing the issue one might infer from the question about extents of deletion in PC (i.e., how much counter-selection might have occurred by the PC stage).

    1. Author response:

      eLife Assessment

      The manuscript by von Velsen et al. offers valuable structural insights into the mitogen-activated protein kinase (MAPK) pathway by providing cryo-EM structures of stabilized MEK1-ERK2 kinase-substrate complexes in inactive, active, and nucleotide-free states, complemented by HDX-MS, SAXS, ITC, crystallography, and molecular dynamics. The work provides solid evidence for the overall architecture of the complex and identifies interaction sites that help explain MAPK pathway specificity. However, some mechanistic conclusions are not yet fully supported, particularly the designation of one state as an active phosphoryl-transfer configuration, the claim that substrate binding releases the MEK1 catalytic machinery, the proposed link to processive phosphorylation, and the extrapolation to disease-associated mutations.

      We would like to counter the final statement. We were very careful in our description of state 2, while we describe it as ‘active’ we clearly explain that the resolution of the reconstruction is not sufficient to define all the classical indicators of a kinase active state; however, the map is consistent with the active conformation, the complex is active in vitro, and the MEK1 variant used is the well known DD mutant that is constitutively active. While the A-loop is not observed, this is in agreement with many crystal structures of other DD mutants. We therefore decided to define this state as ‘active’ as the A-loop of ERK2 approaches the active site, the alpha-C helix has moved in and the A-loop of MEK1 no longer occludes the active site - to clarify the state we refer to the classically active confirmation as ‘fully active’. Our supporting data also show that the complex is highly dynamic during turnover, meaning we have captured MEK1 in a number of conformations on the landscape of an active state – we feel that rather than a limitation, this is an important observation in MAP2K studies. Finally, the determination of an 80 kDa complex by cryoEM to resolutions well below 4 Å is a huge technical achievement allowing the first snapshots of the MEK1-ERK2 complex to be visualised.

      Regarding the A-helix release – our observation is that the helix becomes less folded on binding of substrate. There are many studies, which we cite, that show that destabilising this helix leads to release of the catalytic machinery, see Mansour et al, 1996, Biochemistry, 35, 15529-15536 and Jindal et al 2017 J. Biol. Chem. 292, 18814-18820 for initial studies. Our observation shows that this is linked to substrate binding – a very relevant new insight that demonstrates the importance of this helix, in addition to many previous studies, but links unfolding to substrate recognition for the first time.

      For the mechanism of processive phosphorylation – it has been well established that both processive and distributive mechanisms exist. While the way that a distributive mechanism could work is obvious (complete dissociation of the two proteins), it has not been clear how a MAP2K can remain bound to its substrate and exchange nucleotides. While caution should be employed in interpreting our state 3 structure, it clearly shows what nucleotide exchange when bound to substrate can look like and that this low nucleotide affinity state is linked to disorder in the P-loop, the A-helix and substrate binding via the KIM. We would love to perform experiments that could demonstrate this but cannot at present think of an appropriate method – the reviewers did not suggest a route either.

      Finally, for the cancer-causing mutations – there are many studies demonstrating that the mutations lead to a destabilisation of the A-helix. Our study links this to substrate recognition. While this is inference, it seems justified to describe a link between substrate recognition, A-helix unfolding and disease mutations given the large body of literature describing these events.

      We are currently performing a series of in-cell activity assays that should strengthen our claims regarding the A-helix and other observations in the structure - the histidine interactions in particular.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes three conformers derived from a complex between ERK2-T185V, a variant of MEK1-DD with the KIM sequence replaced by the KIM from the p38 activator, GRA24, ADP, and AlF4-. The goal was to try to capture the complex in its active state. The results show contacts between the kinases between their N-lobe and their C-lobes that resemble MKK6-p38 complexes previously reported by the authors. Two MEK1-ERK2 conformers (States 1,3) are deemed inactive based on the lack of access of ERK-Y187 to the MEK1 active site, and the absence of ADP bound to MEK1 in State 3, while one conformer (State 2) is deemed active, but not fully active due to disorder in MEK1 activation loop (A-loop) and an essential salt bridge between strand beta3 and helix aC. HDX-MS and SAXS solution measurements and all-atom MD simulations are used to model the mutant complex and variants with WT ERK2. The study concludes that substrate recognition involves low-energy contacts with MEK, allowing substantial protein flexibility within the complex in a manner that may accommodate processive phosphorylation of ERK2.

      Strengths:

      The strengths of the work are that the findings provide important structural insights for MEK-ERK signaling and protein phosphorylation in general. These are valuable given that atomic resolution structures of kinase-substrate complexes are still limited in number. The authors succeeded in showing key contacts between subunits and conformational variations within the complex.

      Weaknesses:

      Weaknesses were that some of the conclusions about activity state, dynamics, and effects of ligand binding were less convincing. For example, that State 2 truly represents an active configuration seemed ambiguous, given the absence of Mg2+ and AlF4- in the cryoEM structure and disorder in the activation loop and the K97-E114 salt bridge. Conclusions by SAXS that ADP-AlF4 binding increases active site compaction while increasing local flexibility were not rigorously supported by HDX data, given that the latter were performed without ligand. Sections of the narrative and figures throughout were often confusing, and many assertions were made without clear explanation. Data shown in the supplementary materials were not always described in the Results, even those important for the conclusions. Figure legends and text lacked clear descriptions of specific complexes analyzed. Substantial changes are recommended to improve the readability and clarity of the work.

      We thank reviewer #1 for in-depth comments and analysis of our manuscript. However, there is a misunderstanding regarding HDX-MS and SAXS data. First, the HDX-MS data were performed on the ADP.AlF<sub>4</sub><sup>-</sup> inhibited complex - this was not made sufficiently clear in the text, and we will amend this. Secondly, we are not trying to support local flexibility observed in the SAXS data with the HDX data. The HDX data support the interactions observed in the cryoEM structure and demonstrate flexibility in the proline-rich loop, the ERK A-loop and unfolding of the MEK1 A-helix. The SAXS data demonstrate that when the transition state complex is formed, the complex is more compact but flexibility within the complex increases - as observed in the dimensionless Kratky plot, supporting our observations in the cryoEM maps. Therefore, the HDX data and SAXS data are separate observations. We thank reviewer #1 for all the comments and will rewrite the manuscript in order to increase clarity as suggested.

      Reviewer #2 (Public review):

      Summary:

      The authors used Cryo-EM to obtain a complex between MEK1 and ERK2. They used the same method as previously used by the same authors to form a stable complex between MKK6 and p38, an extra-strong KIM replacing the wild-type KIM in MEK1. Three conformers were resolved, with the highest resolution of 3.0 Å. The multiple conformers indicate more flexibility in MEK2 than in ERK2. These data suggest that nucleotide exchange is possible while maintaining MEK1-ERK2 interactions. SAXS and HDX data reinforce the idea of flexibility. They point to interactions between the two N-terminal domains between histidines at the N-terminus of helix C and between the G helices that are maintained in each of the 3 conformers, and sequence and structure suggest these histidines may be a source of specificity in MEK1-ERK2 versus MKK6-p38 interactions. A 2.2 Å structure of a complex between ERK1 (88% identical to ERK2) and the docking peptide used was also presented. Molecular dynamics simulations suggest that the MEK1-ERK2 complex can assume a fully active configuration of MEK1.

      Strengths:

      This is the first structure of a MEK1-ERK2 complex. The structural data are valuable additions to our understanding of MAP2K-MAPK interactions. The discussion points offered in the results section are palatable. These include the origins of specificity and the idea of flexibility in the MAP2K in support of a processive mechanism for the dual phosphorylation activity of MAP2Ks.

      Weaknesses:

      (1) This reviewer considers that the abstract is overstated. Specifically, this paper does not reveal the molecular details of phosphoryl transfer, nor does it demonstrate that substrate binding releases the catalytic machinery.

      (2) The discussion is in some places not supported by evidence and in others has superfluous text. Examples follow:

      - "Once the αG-helix is docked, and the C-lobe histidine triad is in place, the N-lobe interactions must then be fulfilled." The data in this paper does not suggest an order of events.

      - "If the substrate MAPK is incorrect, the N-lobe interaction will not be stabilised, preventing alignment of the MAPK A-loop with the MAP2K active site." This statement could be described as obvious.

      (3) Much of the discussion is embedded in the results, such that it is difficult to separate new facts offered by the paper from speculation.

      We thank reviewer #2 for comments and thorough analysis of our manuscript. We agree that perhaps the abstract should be toned down in terms of claims of an active conformation even though we feel that the combination of the first structure of the MEK1-ERK2 complex combined with MD simulation studies clearly demonstrate how phosphoryl transfer will occur. We are now also performing in-cell assays to support our theory of A-helix regulation. For point 2 we based the order of events on data from Juyoux et al 2023 Science, 381, 1217-1225, where in long-timescale MD simulations and experimentally validated adaptive Markov state model simulations the KIM interaction was the last to dissociate after the alpha-G helix interaction. In our MD simulations of the MEK1-ERK2 complex, the interactions formed by the N-lobe were weaker than those formed by the alpha-G. Indeed, dissociation of the N-lobe was observed in various independent simulations, whereas dissociation of the alpha-G was observed only once.

      Assuming that, as in the MKK6-p38a complex, the association proceeds along the reverse of the dominant dissociation pathway, the simulations suggest that the KIM interaction forms first, followed by the alpha-G and finally the N-lobes. While alternative association pathways are possible, this interpretation is consistent with the MD and in line with the main association and dissociation pathway observed for the MKK6-p38a complex. This is additionally supported by the observation that there is no catalytic activity if the KIM is removed, demonstrating this as the first essential recruitment event. We will expand this section to include our arguments.

      For the second example, we feel this is rather unfair. The statement that if the His-His interaction is absent, catalysis will be prevented is only obvious if one knows about the His-His interaction - this is the first structure showing pathway-specific interactions in the variable loop regions of a MAPK. If it is obvious, why has no one described these residues as important before?

      We have taken on board the comments on the style of the manuscript and will make significant changes as suggested by both reviewers.

    1. AbstractThe transformer architecture in deep learning has revolutionized protein sequence analysis. Recent advancements in protein language models have paved the way for significant progress across various domains, including protein function and structure prediction, multiple sequence alignments and mutation effect prediction. A protein language model is commonly trained on individual proteins, ignoring the interdependencies between sequences within a genome. However, biological understanding reveals that protein–protein interactions span entire genomic regions, underscoring the limitations of focusing solely on individual proteins. To address these limitations, we propose a novel approach that extends the context size of transformer models across the entire viral genome. By training on large genomic fragments, our method captures long-range interprotein interactions and encodes protein sequences with integrated information from distant proteins within the same genome, offering substantial benefits in various tasks. Viruses, with their densely packed genomes, minimal intergenic regions, and protein annotation challenges, are ideal candidates for genome-wide learning. We introduce a long-context protein language model, trained on entire viral genomes, leveraging a sparse attention mechanism based on protein–protein interactions. Our semi-supervised approach supports long sequences of up to 61,000 amino acids (aa). Our evaluations demonstrate that the resulting embeddings significantly surpass those generated by single-protein models and outperform alternative large-context architectures that rely on static masking or non-transformer frameworks.

      This work has been peer reviewed in GigaScience (see https://doi.org/10.1093/gigascience/giag081), which carries out single-anonymized peer review. These reviews are published under a CC-BY 4.0 license and were as follows:

      Reviewer 2:

      The authors trained a new long-context protein language model for viral protein analysis. In their work, they leverage a biologically informed sparse attention mechanism which also incorporated the inter-protein relationship as sparsity priors. They then evaluated and showed that their model has an improved perplexity and embedding quality. Overall, I think the biological question is important and the method part is relatively convincing. Although I also have few concerns that may further help improve the manuscript.

      Major comments:

      1. Viral species, genome sizes and embedded proteins varied dramatically, excepting the deduplication, the authors should discuss clearer how they get a clean and high-quality data for the training. What is the model performance for different viral group?

      2. From the main text, I don't know how the authors selected the 83 viral genomes from the NCBI database step by step. Also, what viruses were chosen? They may picked the ones with PolA, RNR, and HEL proteins, but whether it is enough for the downstream evaluation should be proved. Whether it could be used for unknown pairs as a viral protein language model.

      3. Additional biological validations may be needed, such as known viral protein motifs or domains.

    2. AbstractThe transformer architecture in deep learning has revolutionized protein sequence analysis. Recent advancements in protein language models have paved the way for significant progress across various domains, including protein function and structure prediction, multiple sequence alignments and mutation effect prediction. A protein language model is commonly trained on individual proteins, ignoring the interdependencies between sequences within a genome. However, biological understanding reveals that protein–protein interactions span entire genomic regions, underscoring the limitations of focusing solely on individual proteins. To address these limitations, we propose a novel approach that extends the context size of transformer models across the entire viral genome. By training on large genomic fragments, our method captures long-range interprotein interactions and encodes protein sequences with integrated information from distant proteins within the same genome, offering substantial benefits in various tasks. Viruses, with their densely packed genomes, minimal intergenic regions, and protein annotation challenges, are ideal candidates for genome-wide learning. We introduce a long-context protein language model, trained on entire viral genomes, leveraging a sparse attention mechanism based on protein–protein interactions. Our semi-supervised approach supports long sequences of up to 61,000 amino acids (aa). Our evaluations demonstrate that the resulting embeddings significantly surpass those generated by single-protein models and outperform alternative large-context architectures that rely on static masking or non-transformer frameworks.

      This work has been peer reviewed in GigaScience (see https://doi.org/10.1093/gigascience/giag081), which carries out single-anonymized peer review. These reviews are published under a CC-BY 4.0 license and were as follows:

      Reviewer 1:

      The paper by Dejean, et. al. describes a novel approach, and set of models, to extend protein language models to a much larger context window (60,000 amino acids) using a combination of approaches to extending the context window including biologically-informed attention mechanisms. The approach is interesting, and potentially very useful. The authors apply it to viral sequences, and show in a variety of ways, improvements over existing models. The manuscript could be improved in several ways detailed below: 1. The authors of the evo paper (which is referenced in the current paper), a DNA language model with a large context, thought it necessary to specifically exclude eukaryotic and thus potential human pathogen viral sequences from their training because of the potential safety issues around their generative model. It is A) not clear if any similar filtering was done in the current study (it is not mentioned, so I assume not), B) not clear if this is as necessary as it may be with the evo model - that is, do you consider your models to be usefully generative? At the very least some discussion of this important consideration should be mentioned in the paper. 2. The different models that are developed and used in the paper are overall confusing. This is a significant issue with the paper since it makes it very difficult to figure out what results map to which models, how the pieces fit together, and the significance of the advances. My feeling is that reducing the number of different models and forms of different models being referred to in the main text (see point 3, below). a. If I understand (and I may not), ESM is used in some form (fine tuned, I think) to serve as the first stage that predicts protein interactions, which are then fed into a second stage which (may be) the two models labelled LV-3C and LV-5B. ESM is used in two forms (maybe more?) - which seems to be the original form and a fine-tuned form which might have an extended context window and also might be fine-tuned on the same input sequence? b. The first part of the results section, which talks about predicting protein interactions seems to describe this fine-tuning. But it's not clear. There are two positional encoding and 3 sparse encoding strategies tested (which is fine) but it makes it really hard to figure out what is then used, and why. The motivation for this section, and the choices for the models needs to be made clear. Also, there seem to be variants of the fine-tuning that are used throughout? This makes things more confusing. c. The difference between the two main models, LV-3C and LV-5B are not clear. 3C is trained on genomic segments using 'sparse attention informed by protein-protein interactions' (are these from the first stage?) and 5B is trained on shorter segments - but using transfer learning from ESM-2 and using 'biologically-informed sparse attention'. How is 'biologically-informed' different than 'protein-protein interactions'? Much more clarity around the differences and motivation for each of these models is needed to be able to interpret the following results. d. ESM 'flavors' are used to compare - as baselines - for the other models, but since it's hard to follow what these ESM flavors are exactly it makes interpretation difficult. Also, the same ESM model(s) is used for the 'first stage' to predict interactions as to compare against the final 'second stage' models? This needs to be made clear. Related, I liked the use of the PolA, RNR, and HEL to optimize the number of interactions needed. But then this same complex is used as an example in the final model. Which is confusing, and also, it's not entirely clear that it's not circular (I feel it's not, but this really needs some explanation and clarity) - that is, 'we used PolA, RNR, and HEL to optimize our stage 1, and we find that our stage 2 is really good at predicting these interactions' e. PST is introduced as another model not trained here, but included as a comparator for only one part of the results - Fig 9, evaluation of protein embeddings. It's not clear why it needs to be included in these results, but not for others? f. Methods 3.4 selecting an optimal fine tuning strategy: for what model or portion of the model? 3. The paper would greatly benefit from some revision and streamlining to make the entire process, the models used, and the methods more clear. I would suggest trying to move some of the results to a supplemental section: the discussion of interaction inference (section 4.1) or 'stage 1' and the different things tried there to be confusing - whereas I might be very interested in those results and what was being done if they weren't confusing (moving those to the supplement is one way of accomplishing this, then describing 'we optimized stage 1 (see supplement) and decided on using XXX for the final model because it showed XXX.' 4. Each in the results must have a clear motivation sentence or section right at the start: what is the hypothesis that's being examined here? What are you doing to test this hypothesis (give enough of an idea of the approach/methods to give the reader a reminder)? Then at the end, What do the results tell you and how does it suggest the next section? 5. The Related Work section is well-written and was very useful to give me good background information. However, it reads more like a mini-review paper than a focused examination of the approaches which are used in the current paper - and, importantly, it's not clear that all the methods described are important to understand for the paper. That is, which parts are needed for the reader to understand the approaches tested and/or used in the final work. A short section at the start of this that briefly teases that the final model includes elements from these different approaches would be helpful so the reader knows why they should be paying attention. 6. Perplexity and silhouette score are used extensively in the results section, but never described in the methods section. 7. The final paragraph of the introduction seems out of place. This should be moved to the discussion maybe? It's an odd way to end the introduction, and feels like it's more of a caveat that can be discussed after, than the most important thing we should be taking from the paper. 8. Methods: the description of Splits is not clear. It seems that genomes, and collections of proteins from the same genome were kept together in a single split so as not to contaminate the evaluation, which makes sense, but just needs to be stated in a more careful and clear way. 9. Methods: protein non-redundancy - the use of an identity filter to separate similar proteins is A) welcome, it's an important thing to do, but… B) 90% seems overly generous - that is, it makes the task pretty easy since 89% sequence identity is still very similar. I would welcome a mention of this in the discussion, and some more supplemental results that show a few of the downstream task comparisons using the more stringent 50% filter. 10. Methods: metadata - I can't tell if the first paragraph only is the 'metadata' part (that seems reasonable) but the rest of the section is not about metadata and probably should be titled something on its own. 11. Methods: the section titled 'Ablation' doesn't immediately strike me as being about ablation at all? For this (and some of the other methods sections too) a short lead in of what the method describes would be good 'We needed an approach to limit the number of interactions used for the context so we …'. Also, leading with 'We varied k…' - what is k (it's there in the equation, but the reader won't know that)? 12. Figure listing is out of order it seems? Maybe because methods refer to results section - not sure. 13. Table 1 - what is S2? Also really not clear what the other labels refer to exactly either 14. Figure 5 is fairly easy to digest visually (removing the adjacent proteins improves the ranking of known interactions) but would very much benefit from a quantitative measure of how different these distributions are (likely a p-value from some test) 15. Section 4.2. The motivation for using MMseqs2 is not clear. I gather it's to provide a baseline 'trusted' clustering to compare to? Also this is not described in the methods. 16. I would suggest moving the discussion of analysis of variance and distribution of the embeddings to the first downstream analysis reported. It is useful, but really doesn't say anything about model quality or performance by itself. If you put it first (in section 4.4) it will lead in to the other measures, which are stronger (in my opinion). 17. Comparison of the embeddings to STRING is a nice addition, but not much time is spent on describing how this is done. 146 interactions (I'm assuming these are TPs) are considered. What are used as negatives? What does it mean to have an F1 score of 0.74 (which seems - pretty good) and an AUC of 0.66 (which seems marginal)? The final sentence in this section is not supported (no correlation between STRING confidence scores and attention scores was shown). 18. Figure 12 needs a figure legend. 19. Section 4.3 the authors state that 'The aim of clustering-based evaluation is to identify consistency between embeddings in latent space' - but it's not clear that this is how clustering is being used or evaluated following this? - the performance of the different models on a classification task don't directly evaluate how consistent the embeddings are (you could have very different embeddings that lead to similar performance, e.g.)

    1. AbstractThe performance of long-read mapping is critical yet highly sensitive to parameter choices. We present CycSim, a context-aware simulator that models sequence-context-dependent errors from empirical sequencing data, coupled with a Bayesian optimization framework for systematic parameter tuning. CycSim more accurately reproduces real error profiles than existing simulators, enabling reliable simulation-based optimization. The framework identified parameter configurations that achieved 2.78-fold faster mapping for data from the newly developed Cyclone platform, and consistently improved both mapping efficiency (8.14-32.65% faster) and structural variant calling accuracy (0.75-1.70% higher F1) across ONT, HiFi, and Cyclone datasets, providing a robust and generalizable foundation for analysis-goal-driven parameter refinement.

      This work has been peer reviewed in GigaScience (see https://doi.org/10.1093/gigascience/giag079), which carries out single-anonymized peer review. These reviews are published under a CC-BY 4.0 license and were as follows:

      Reviewer 2:

      In this work, Hu et al describe CycSim, a context aware simulator for diverse long read chemistries. Using this simulator, the authors aimed to optimize the mapping parameters to improve mapping speed and accuracy of variant calls.

      This is important, given the increase in use of long read sequencing, particularly in the wake of upcoming technologies such as Cyclone. However, I have significant concerns about how the results are analyzed and work is presented. Despite it being a short article, I had to do a lot of back and forth reading due to lack of clarity in presentation. It would also help to have line numbers to point out specific places in the manuscript. Further I have some questions which the authors should address with a revision.

      CycSim is a "dual-stage framework" - Does that mean the users are expected to train with every new sample? Or the trained model can now simulate reads for any sample? How do we know the training is not over engineered for the HG002 sample, particularly the 4 chromosome subset?

      The algorithm characterizes many aspects of the read including strand, orientation etc. Are they used for training in any way?

      Are the kmer models built for each "kind" of genomic region (for example LCRs/TRs)?

      Simulation/validation is done for the same set of chromosomes as those that were used for training. How do the various parameters benchmarked in Fig 1 fare for other chromosomes which the model has not seen?

      In LCRs, CycSim outperforms other tools. This is a crucial point. But what is shown is mapping identity - as far as I know, the mapping identity in these regions should be lower. Not sure why it's higher, unless I'm misunderstanding how this is calculated. Also, it would be important to show other features of the reads such as kmer profiles and substitutions, specifically in various genomic regions, rather than mapping identity % alone.

      Regarding the parameter tuning/optimization - I have significant problems with how the data is presented. First of all, radar charts, while they look fancy, are less effective in depicting small changes, which is what the authors are trying to show here. Simple bar charts would have been way more clear. Also, this feels like an independent goal and section, and the connection with their simulation strategy is not clear.

      More importantly, why wasn't this done with reads simulated with other tools? For all we know, the optimization may have worked with reads simulated by BadRead or PBSim as well.

      The 2.78-fold increase in speed for Cyclone data is compared to map-ont preset, which by definition is not optimized for Cyclone data. While this is an important exercise and result, not sure how this is a direct benefit of CycSim. Could the mapping speed be not optimized by trying Optuna on raw Cyclone data directly?

      The authors claim that the framework is robust in optimizing parameters "across diverse sequencing platforms". But in their own words, the gains for PacBio and ONT were <0.1%.

      The truvari refine parameters had a flag -p 0.0, indicating only position (that too up to 1kb distance) was taken into account when calculating the overlap, irrespective of sequence similarity. Most people in practice keep 0.5-0.7, often with reciprocal overlap. Curious to see how the results change with such parameters.

      It is not clear whether the SNP and SV optimized parameters are the same or not. If yes, this should be clarified. If not, authors should include results on what happens to SNP accuracy when using SV specific parameters. While I agree that there is merit in using specific alignment parameters for specific tasks, in practice, users might just use the same BAM file for multiple types of variants - hence having this information can help the user in deciding which parameters they want to use.

      The gains in F1 scores are marginal. Authors should discuss if and why such marginal improvements are important.

      Other comments:

      What happens when you simulate high coverage? Would it recapitulate actual high coverage data or would there be repeated data due to limits of a kmer model?

      Was the simulation done multiple times? This should be clarified. If not done, it should be - to see how reproducible the results are.

      Thanks for providing the optimized parameters for each platform - can there be a comment on why the authors think these parameters outperform default parameters?

      Minor:

      Several typos - such as "framwork", "charaterize", and missing commas etc.

      I could not initially find cutesv results, till I stumbled upon them in the tables. Please tag the table numbers at the appropriate place in the text.

    1. AbstractBackground Live-cell fluorescence microscopy enables the study of dynamic cellular processes. However, fluorescence microscopy can damage cells and disrupt these dynamic processes through photobleaching and phototoxicity. Reducing light exposure mitigates the effects of photobleaching and phototoxicity but results in low signal-to-noise ratio (SNR) images. Deep learning provides a solution for restoring these low-SNR images. However, these deep learning methods require large, representative datasets for training, testing, and benchmarking, as well as substantial GPU memory, particularly for denoising large images.Results We present a new fluorescence microscopy dataset designed to expand the range of imaging conditions and specimens currently available for evaluating denoising methods. The dataset contains 324 paired high/low-SNR images ranging from four to 282 megapixels across 12 sub-datasets that vary in specimen, objective used, staining type, excitation wavelength, and exposure time. The dataset also includes spinning disk confocal microscopy examples and extreme-noise cases. We evaluated three state-of-the-art deep learning denoising models on the dataset: a supervised transformer-based model, a supervised CNN model, and an unsupervised single image model. We also developed an image stitching method that enables large images to be processed in smaller crops and reconstructed.Conclusions Our dataset provides a diverse benchmark for evaluating deep learning denoising methods, and our stitching method provides a solution to GPU memory constraints encountered when processing large images. Among the evaluated deep learning models, the supervised transformer-based model had the highest denoising performance but required the longest training time.

      This work has been peer reviewed in GigaScience(see https://doi.org/10.1093/gigascience/giag071), which carries out single-anonymized peer review. These reviews are published under a CC-BY 4.0 license and were as follows:

      Reviewer 2:

      In the manuscript, the authors present a dataset of 324 paired high- and low-signal-to-noise ratio fluorescence microscopy images designed to improve the training and benchmarking of deep learning denoising models across various biological specimens and imaging conditions. The authors also developed an image stitching method to address GPU memory constraints when processing large images and demonstrated that supervised transformer-based models achieve the highest denoising performance among state-of-the-art methods. Overall I found the exposition quite good but some parts could result a bit confusing. So while I think the work is useful, I have some comments.

      Here are my points:

      I think the way Table 1 is organized is not very clear (at least to me!). Because the table refers to images with different sizes I was quite a bit lost. If for technique A, Sample B, there is an image of 26 MP (which I guess stands for megapixels), is this image then divided in 100 non-overlapping 512x512 images? So of these 100 images 90 are considered for training? How does this really work? In the table, instead of the MP indication, I would put how many paired images are considered for training/testing/validation and their typical size (e.g. technique A, Sample B has 100 images for training, 20 for testing, and 10 for validation. Also I would mention that these images have all size 512x512 or whatever the size was. I actually found a hint of this in the methods section toward the end of the paper. But I would just insert explicit numbers in the table so the reader knows immediately what is happening from the beginning.

      The term "high-resolution images" is used ambiguously ("We introduce a novel dataset of 324 high-resolution images"). Does this refer to high pixel counts (large field of view), high spatial resolution (sampling frequency/Nyquist), or the optical resolution of the objectives used? A clearer definition is required.

      Given the difference in image sizes across the dataset, some samples appear to contribute disproportionately to the training set (but this point could be due to a misreading on my part of how the table in column 1 and colum 2 is built and how it should be interpreted). Could the authors discuss how this imbalance affects the test results? I would expect under-represented features to show lower performance, and this should be reflected in the evaluation metrics.

      The abstract mentions "spinning disk confocal" as an "also included" category. But before that there is no mention of any other modality. It is critical to define all modalities (e.g., widefield vs. confocal) in the introduction/background, as the difference in Z-resolution and PSF (Point Spread Function) means models trained on one may not generalize to the other.

      The authors include excitation wavelengths but omit emission wavelengths. SNR and image quality are in some way dependent on the emission filters and camera quantum efficiency at specific wavelengths; therefore, an emission column should be added to the technical tables.

      Regarding the layout of Figure 3 I think that for better visual comparison, the figures should be rearranged. Since Restormer is identified as the best-performing model, it should be placed immediately adjacent to the "Ground Truth" (high-SNR) image to allow the reader to easily assess its fidelity.

      I'm not sure I got this right but in the abstract the authors mention "12 sub-datasets" but the table contains 15 entries.

      I think the sentence " how different stains may impact denoising accuracy" should be rephrased. I can have stains that mark the same structures but with different fluorophores or mechanisms of attachment, and the features will be the same. It is more the target of the stain (which represents the feature content of the image) that would affect how the trained model can be more or less effective on the new target.

      In conclusion the subject of the paper is interesting in my opinion and I think that overall the authors did a good job in providing very convincing results and in presenting potential applications of their method. The inclusion of a stitching method for high-megapixel images is also a practical contribution for microscopy applications and quite useful.

    1. AbstractBackground The rapid advancement in single-cell, spatial omics, imaging, and genomic technologies requires robust analytical and visualisation platforms capable of managing complex biological data. Tools such as Multi-Dimensional Viewer (MDV) offer comprehensive interfaces for data exploration, but still require manual configuration and computational expertise to generate visualisation outputs, limiting accessibility for many users.Results We present ChatMDV, a natural language interface integrated with MDV that allows users to generate high-quality interactive visualisations through natural language commands. ChatMDV employs a retrieval-augmented generation (RAG) pipeline combined with large language models (LLMs) to translate user queries into reproducible Python code and interactive output. This approach enables exploratory and targeted analysis in diverse biological domains. We demonstrate ChatMDV’s capabilities using three datasets of increasing complexity: the Peripheral Blood Mononuclear Cells 3K (PBMC3K) dataset, the lung cancer atlas dataset hosted at the Human Cell Atlas and the longitudinal TAURUS study single-cell RNA-sequencing (scRNA-seq) dataset.Conclusions By bridging the gap between natural language processing and bioinformatics visualisation, ChatMDV reduces technical barriers, enhances reproducibility, and supports more inclusive scientific inquiry. Its modular design and adherence to FAIR (Findability, Accessibility, Interoperability, and Reuse) principles make it a scalable and adaptable framework for accelerating biological data analysis.

      This work has been peer reviewed in GigaScience (see https://doi.org/10.1093/gigascience/giag073), which carries out single-anonymized peer review. These reviews are published under a CC-BY 4.0 license and were as follows:

      Reviewer 2:

      This manuscript presents ChatMDV, a natural language-driven bioinformatics visualization platform that integrates large language models (LLMs) with the MDV. Through an agent + RAG + code-generation architecture, ChatMDV translates user prompts into executable, reproducible Python analysis code and interactive MDV visualizations, and the authors provide a systematic evaluation on multiple scRNA-seq datasets at scale, including PBMC3K, the Human Cell Atlas lung atlas, and the TAURUS longitudinal study dataset. Overall, the work is technically solid, with strong system engineering and a high completeness. Below are a few remarks that I hope the authors can clarify in a revision.

      1. The reported ">95% success rate" uses an overly generous definition and likely overestimates usability. A central claim is that ChatMDV achieves a success rate exceeding 95% in visualizing the evaluated datasets. However, the manuscript defines success as any rating except 1, i.e. effectively rating 2 - "View present but error-filled". However, the rating rubric explicitly states that rating 2 includes cases where "charts were present but did not address the question or contained significant errors." Counting these cases as "success" is too loose in my opinion, especially when success is interpreted by readers as semantic correctness or user-ready performance. Therefore, I think it would be great to distinguish execution success (rating over 2) and semantic success (rating over 4 and 5). Without this, the abstract ">95% success" risks being misleading.

      2. Evaluation prompts may be too well-formed. The prompt set is carefully designed and scored with a bespoke complexity scheme, and the manuscript notes that prompts were developed with expert input, literature review, and assistance from ChatGPT. While this is reasonable, it raises a concern that the final benchmark may be closer to "experienced user requests" than what many wet-lab or novice computational users would actually type, for example ambiguous phrasing, missing chart specifications, typos, etc. Therefore, adding a small additional benchmark of "inexperienced biologist prompts" would substantially strengthen claims about democratizing analysis.

      3. Handling biological hallucinations is described at an engineering level as the manuscript acknowledges key failure modes and describes retries/error handling. However, these are mostly presented as engineering notes rather than as actionable, interpretable examples for end-users and reviewers. I think it would be useful to provide one or two concrete failure case analyses so that users can understand what ChatMDV can and cannot reliably do and how to proceed when it fails.

      4. The word "democratizing" sounds quite strong as no one would describe the current bioinformatics practice as "autocratic". A more technical but accurate phrasing, like "reducing technical barriers" would be appreciated.

    1. Author response:

      The following is the authors’ response to the current reviews.

      Reviewer #1 (Recommendations for the authors):

      (1) Interpretation of Syllable-Tracking in the RND Condition:

      The finding of greater syllable-tracking in the LL group compared to the HL group in the RND condition warrants cautious interpretation. Currently, there is no direct statistical evidence demonstrating greater PLV at 4 Hz in the Structured versus Random conditions for either group; readers must infer this solely from numeric differences in Figure S5 B and D. Therefore, while the interpretation on Page 14 (Lines 443-446) "successful segmentation may enhance syllable tracking via top-down predictions of the next syllable" is an interesting speculation, it feels somewhat far-reaching. Additionally, the authors should discuss whether this upregulated syllable tracking in the structured condition (which is specific to the HL group) represents an adaptive or maladaptive response.

      The reviewer correctly highlights the lack of direct  comparison between conditions (RND versus STR). We tempered our claims in the cited paragraph and insisted on the speculative nature of this part of the discussion. We also clarified that, to us, it may represent an adaptive compensatory strategy:

      Page 14, line 441: “Interestingly, our supplementary analyses (Supplementary Material Figure S4-5) suggest that syllable entrainment may be differentially affected in HL versus LL infants, depending on the statistical structure of the input stream (RND versus STR). However, as our experiment was not explicitly designed to test stream effects, these results should be interpreted with caution. Future studies could explore how successful segmentation may enhance syllable tracking via top-down predictions of the next syllable in both LL and HL infants. If confirmed, such a mechanism may improve alignment to syllable onsets, potentially constituting a compensatory process allowed by preserved segmentation abilities.”

      (2) Preservation of Statistical Learning in HL Infants:

      The text added on Pages 17-18 (Lines 562-566) regarding a "heightened dependence on bottom-up mechanisms (in autism)" does not appear to be supported by the data or by theories of implicit statistical learning. Because greater syllable-level entrainment was observed in the LL group than the HL group across both the random and structured conditions, the data actually point toward impaired bottom-up processes. Furthermore, implicit statistical learning typically involves an interplay of both bottom-up and top-down mechanisms; the implicit nature of a task does not guarantee a strictly bottom-up process. Consequently, this interpretation is not entirely convincing.

      We agree with the reviewer that the concepts of “top-down” and “bottom-up” were not fully appropriate to support our point in the cited paragraph. We should have used the concepts of implicit versus explicit learning instead, in line with previous literature suggesting increased reliance on preserved implicit learning in autism to compensate for altered explicit processes. The paragraph was slightly modified.

      Page 18, line 564: “According to these studies, autistic impairments in explicit attentional processes, such as social orienting - which are critical for bootstrapping language acquisition (70) - may result in a heightened dependence on implicit mechanisms, including statistical learning. As previously discussed, preserved word segmentation abilities may further compensate for alterations in lower-level implicit processes, such as syllable tracking.”

      Reviewer #2 (Recommendations for the authors):

      Potential typo on line 199 - I think an apostrophe is needed here.<br /> Potential typo on line 255 - do you mean Central electrodes?

      We addressed the typos spotted by reviewer.

      Line 199: variables’

      Line 255: Centro-frontal electrodes


      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports a prospective longitudinal study examining whether infants with high likelihood (HL) for autism differ from low-likelihood (LL) infants in two levels of word learning: brain-to-speech cortical entrainment and implicit word segmentation. The authors report reduced syllable tracking and post-learning word recognition in the HL group relative to the LL group. Importantly, both the syllable-tracking entrainment measure and the word recognition ERP measure are positively associated with verbal outcomes at 18-20 months, as indexed by the Mullen Verbal Developmental Quotient. Overall, I found this to be a thoughtfully designed and carefully executed study that tackles a difficult and important set of questions. With some clarifications and modest additional analyses or discussion on the points below, the manuscript has strong potential to make a substantial contribution to the literature on early language development and autism.

      Strengths:

      This is an important study that addresses a central question in developmental cognitive neuroscience: what mechanisms underlie variability in language learning, and what are the early neural correlates of these individual differences? While language development has a relatively well-defined sensitive period in typical development, the mechanisms of variability - particularly in the context of neurodevelopmental conditions - remain poorly understood, in part because longitudinal work in very young infants and toddlers is rare. The present study makes a valuable contribution by directly targeting this gap and by grounding the work in a strong theoretical tradition on statistical learning as a foundational mechanism for early language acquisition.

      I especially appreciate the authors' meticulous approach to data quality and their clear, transparent description of the methods. The choice of partial least squares correlation (PLS-c) is well motivated, given the multidimensional nature of the data and collinearity among variables, and the manuscript does a commendable job explaining this technique to readers who may be less familiar with it.

      The results reveal interesting developmental changes in syllable tracking and word segmentation from birth to 2 years in both HL and LL infants. Simply mapping these trajectories in both groups is highly valuable. Moreover, the associations between neural indices of brain-to-speech entrainment and word segmentation with later verbal outcomes in the LL group support a critical role for speech perception and statistical learning in early language development, with clear implications for understanding autism. Overall, this is a rich dataset with substantial potential to inform theory.

      Weaknesses:

      (1) Clarifying longitudinal vs. concurrent associations

      Because the current analytical approach incorporates all time points, including the final visit, it is challenging to determine to what extent the brain-language associations are driven by longitudinal relationships vs. concurrent correlations at the last time point. This does not undermine the main findings, but clarifying this issue could significantly enhance the impact of the individual-differences results. If feasible, the authors might consider (a) showing that a model excluding the final visit still predicts verbal outcomes at the last visit in a similar way, or (b) more explicitly acknowledging in the discussion that the observed associations may be partly or largely driven by concurrent correlations. Either approach would help readers interpret the strength and nature of the longitudinal claims.

      We thank the reviewer for this insightful comment. We agree that distinguishing between longitudinal predictive power and concurrent correlations at the final visit is crucial for clarifying the nature of these brain-language associations. Following the reviewer’s suggestion (a), we re-ran the two critical Partial Least Squares Correlation (PLS-c) analyses by excluding all EEG and behavioral data from the final 18–21 month visit (n = 54 recordings kept) to test whether earlier trajectories still predict the final verbal outcome.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2C–D): The PLS-c restricted to the 3- to 15-month visits still identified a single significant component (p=.001, r=.56, 64.0% explained covariance, Figure 2 -figure supplement 3). Bootstrap ratios (BSR) were: contrast (low vs. high autism likelihood) 4.1; mean age −1.3; contrast*mean-age −2.1; delta-age 10.5; contrast*delta-age 2.0; age<sup>2</sup> −6.0; contrast*age<sup>2</sup> −5.8; and notably verbal outcome 7.1; contrast*verbal-outcome −6.3.

      The latent component and its spatial electrode configuration remain highly consistent with the original analysis (Figure 2C–D). This confirms that excluding the final visit preserves the model’s predictive validity: lower syllable entrainment correlates with poorer verbal outcomes at 18–21 months, particularly in the high-likelihood group.

      (2) Late evoked response to novel words (original analysis on Figure 6): The PLS-c analysis on the ERP late time window (1500–3000 ms), excluding the final visit, also revealed one significant component (p=.002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (part-word versus word) 14.5; mean age -5.2; contrast*mean-age 6.2; delta-age -0.8; contrast*delta-age 3.5; age<sup>2</sup> -0.8; contrast*age<sup>2</sup> -12.1; verbal-outcome 20.1; contrast*verbal-outcome -7.2. The latent component closely mirrored the original analysis (Figure 6), with frontal electrodes contributing negatively and posterior electrodes positively. Minor divergences in age-related parameter contributions were observed, likely due to the absence of 18-21 month timepoints, which previously contributed to the convex/concave shapes of the group age trajectories in figure 6B (left panel).

      Crucially, both models (with and without the final visit) positively predicted verbal outcomes (Figure 6B, right panel, and Figure 6 -figure supplement 1B). However, excluding the final visit reversed the direction of the group*verbal-outcome interaction (from 3 to -7.2): This indicates that after ruling out cross-sectional correlations at 18–21 months, the early predictive value of the late ERP to word novelty is more prominently observed in high-likelihood infants, suggesting that the original result was influenced by concurrent cross-sectional correlations at the final visit. This aligns with the syllable entrainment findings (Figure 2 -figure supplement 3), as both 4 Hz neural tracking and late ERP responses to novelty predominantly predict verbal outcomes in infants at high likelihood for autism.

      We reported these supplementary analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 3 and Figure 6 -figure supplement 1. In general, most of figures that were present in Supplementary materials were moved as figure supplements to enhance readability.

      Page 8 lines 234-240 (pages and lines refer to the reviewed uploaded manuscript): “To rule out the possibility that the association between syllable entrainment and verbal outcome was driven by concurrent measures taken at 18–21 months, we re-ran the PLS-c analysis excluding EEG data from the final visit (n = 54 recordings kept). The resulting latent component remain significant (p = .001) and showed contributions from behavioral and EEG variables that were highly similar to those observed in the previous analysis, with a verbal outcome BSR of 7.1 and a group’verbal-outcome interaction BSR of −6.3 (Figure 2 -figure supplement 3).”

      Page 12 lines 387-394: “As we did for neural entrainment to syllables, we conducted a new analysis on late ERP to word novelty, excluding EEG data from the final visit. This PLS-c yielded one significant latent component (p = .002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1) with globally similar EEG parameter contributions and age trajectory modelling. Verbal outcome still significantly contributed to the latent component (BSR=20.1), with a negative verbal outcome*group interaction (BSR=-7.2). These results suggest that, after ruling out cross-sectional correlations at 18–21 months, the late ERP to word novelty predominantly predicts verbal outcomes in high-likelihood infants for autism.”

      Page 17 lines 547-548: “As with syllable entrainment, the late ERP to novel words primarily predicted verbal outcomes in high-likelihood (HL) infants.”

      Page 18 lines 588-590: “Likewise, the absence of a late ERP orientation response in HL participants may represent an early neural signature of altered attention to novelty that can be used both as a non-invasive predictor of language development and as a potential target for early intervention.’

      (2) Incorporating sleep status into longitudinal models

      Sleep status changes systematically across developmental stages in this cohort. Given that some of the papers cited to justify the paradigm also note limitations in speech entrainment and word segmentation during sleep or in patients with impaired consciousness, it would be helpful to account for sleep more directly. Including sleep status as a factor or covariate in the longitudinal models, or at least elaborating more fully on its potential role and limitations, would further strengthen the conclusions and reassure readers that these effects are not primarily driven by differences in sleep-wake state.

      The reviewer is highlighting here a limitation of our study design that comprised sleeping status that varied from one timepoint to another among participants. To rule out any confounding effect of wake status (coded as a binary variable: sleeping or awake during recording) on analyses comparing groups, a linear mixed-effect model with repeated measures was fitted finding no significant difference between high- and low-likelihood participants (p=.769, reported at page 20, lines 646-647). However, as rightly suggested by the reviewer, this doesn’t prevent from a sleep bias on age trajectories, especially given that sleeping status significantly decreases with age in our sample.

      Including sleep status as a covariate in our analyses, as suggested by the reviewer, would be difficult to implement in our PLS-c methods, since a categorical behavioral parameter that varies within participants is not possible in the models provided by myPLS toolbox.

      As an alternative option, we re-ran all analyses that explored the condition effect on the whole sample within the sleeping participants only (n=25 recordings) to confirm that the same age-trajectories of EEG parameters were highlighted. However, negative results should be interpreted with caution since the sample is small for such a multivariate approach, resulting in modest statistical power.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2A–B): The PLS-c identified one significant component (p <.001, r = .78, 85.1% explained covariance, Figure 2 -figure supplement 2 and Figure 3 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (4hz vs. adjacent frequencies) 30.3; mean age -2.9; contrast*mean-age -2.4; delta-age 3.8; contrast*delta-age 3.4; age<sup>2</sup> -1.1; contrast* −2.5. The spatial distribution of contributing electrodes globally matched that shown in Figure 2A. The high contrast BSR (30.3) confirms robust syllable entrainment in sleeping infants. Critically, the contrast*age<sup>2</sup> parameter contributed negatively to the latent component (BSR = −2.5), confirming that the convex age trajectory of syllabic entrainment (Figure 2B) is also present in the sleeping subsample.

      (2) Word entrainment (1.3 Hz) (original analysis on Figure 3A–B): The PLS-c identified one significant component (p <.001, r = .63, 37.3% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (1.3hz vs. adjacent frequencies) 24.9; mean age -7.5; contrast*mean-age -3.2; delta-age 2.9; contrast*delta-age 0.0; age<sup>2</sup> 1.5; contrast* 1.1. The spatial distribution of significant electrodes partially overlaps with the ones in the original analysis, primarily showing fronto-central positive contribution to the latent component. The high contrast BSR confirms a robust word entrainment in sleeping participants, in line with previous studies (e.g., Flò et al, Sci Rep, 2022). However, the lack of a significant contrast* age<sup>2</sup> suggests that the U-shape age trajectory illustrated on Figure 3 might be modulated by wakefulness or due to a lack of power in the present analysis. A non-significant trend towards a U-shape pattern with a 12-month nadir is visible in sleeping participants, but additional data from sleeping 18-21 months sleeping infants would be required to confirm or refute this trend.

      (3) Early evoked response to novel words (original analysis on Figure 4): The PLS-c analysis on the ERP early time window (0–1000 ms) in sleeping participants revealed no significant component. The absence of early response to word novelty in sleeping participant might account for the lack of response observed in the whole sample, illustrated on Figure 4. To test this hypothesis, we conducted the same PLS-c in awake participants (n=58 recordings), which also yielded no significant latent component. This suggests that the lack of a measurable early response to word novelty observed in the whole sample is consistent across both sleeping and awake infants, and not driven by any of the two subsamples.

      (4) Late evoked response to novel words (original analysis on Figure 5): The PLS-c analysis on the ERP late time window in sleeping participants revealed no significant component. This suggests that sleeping participants might present a reduced or even absent late response to novel words. Given this identified effect of sleep on late ERP response, we reran the PLS-c on the late ERP window using group as contrast (original analysis on figure 6), excluding the sleeping participants to avoid any confounds. This PLS-c revealed one significant component (p = .006, r = .68, 30.5% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (low versus high likelihood) 7.6; mean age -5.9; contrast*mean-age -3.8; delta-age 2.5; contrast*delta-age -3.5; age<sup>2</sup> -2.5; contrast*age<sup>2</sup> -4.9; verbal-outcome 8.5; contrast*verbal-outcome -1.0. Behavioral parameters contribute to this latent component with similar magnitude and polarity as in the original analysis. Electrode contributions are also highly consistent, with frontal negative and posterior positive contributions. This confirms that sleeping participants, despite their potentially reduced late response, did not significantly bias the results presented in Figure 6.

      Author response image 1.

      Late evoked response potential (ERP) to word novelty in awake participants. A. Design and brain saliences derived from the significant latent component. Brain topographies of bootstrap ratios (BSR) are displayed at 250ms intervals. Black dots indicate BSR > 2.3. B. Participants’ brain scores for part-word and word conditions, as a function of age (left panel) and verbal DQ (right panel). For details on brain scores, see Figure 6 -figure supplement 1. Linear fitting is used for illustrative purposes only. HL: high likelihood for autism; LL: low likelihood for autism.

      We reported these analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 2A and Figure 3 -figure supplement 1.

      Page 7, lines 218-224: “Because some infants were asleep during the recording session, particularly at younger ages, we performed a supplementary control analysis restricted to this sleeping subsample (n = 25 recordings, Figure 2 -figure supplement 2). This PLS-c also identified a significant latent component (p < .001, r = .78, 85.1% explained covariance), with a significant contrast effect (BSR = 30.3) and a significant negative contrast*age<sup>2</sup> interaction (BSR = −2.5). These findings confirm that the convex age trajectory observed in the main analysis remains present and observable even in sleeping infants.”

      Page 8, lines 252-259: “We further investigated word entrainment in sleeping participants (n=25), which yielded one significant latent component (p<.001, r=.63, 37.3% explained covariance, Figure 3 -figure supplement 1). Centro-frontal electrode contributed to this component, with a high contrast BSR (24.9), confirming a similar word entrainment pattern in the sleeping subsample. The contrast*age<sup>2</sup> was also positive but not significant (1.1), suggesting a trend toward a U-shape age trajectory with a 12-month nadir in sleeping infants. Additional 18-21 month recording would be required to confirm this trend.”

      Page 11, lines 347-349: “The same PLS-c, conducted separately in sleeping (n=25) and awake subsamples (n = 58), yielded no significant latent component, indicating a consistent absence of early response to word novelty in both sleeping and awake infants.”

      Page 11-12 lines 368-370: “The same PLS-c in the sleeping subsample yielded no significant latent component, suggesting that sleep may reduce or even abolish the late response to word novelty.”

      Page 12 lines 382-385: “Given that no late response was detected in sleeping participants, we re-ran the PLS-c analysis using group as a contrast in the awake subsample (n=58). This yielded one significant latent component (p=.006, r=.68, 30.5% explained covariance), with behavioral and electrode contributions highly overlapping with those in Figure 6.”

      Page 18 lines 590-592: “This potential biomarker might nevertheless be modulated by participants’ sleep status, warranting careful consideration of vigilance state in future studies.”

      (3) Use of PLS-c and potential group × condition interactions

      I am relatively new to PLS-c. One question that arose is whether PLS-c could be extended to handle a two-way interaction between group and condition contrasts (STR vs. RND). If so, some of the more complex supplementary models testing developmental trajectories within each group (Page 8, Lines 258-265) might be more directly captured within a single, unified framework. Even a brief comment in the methods or discussion about the feasibility (or limitations) of modeling such interactions within PLS-c would be informative for readers and could streamline the analytic narrative.

      The reviewer raises a valid concern regarding the capacity of PLS-c to accommodate multi-way interactions among categorical and continuous variables. While PLS-c has no inherent theoretical constraints on the number of predictor terms (they can even exceed the sample size in number), practical limitations arise from model stability and interpretability when the ratio of predictors to sample size becomes excessive. As noted by Geladi and Kowalski (1986), exceeding ~10% of the sample size with predictors increases noise sensitivity and overfitting.

      In our study, the PLS-c analyses already reach this ~10% limit, with a maximum of nine predictors for a sample size of n=83. Attempting to integrate both group and condition as contrasts — along with necessary age parameters to account for developmental trajectories — would result in 12 predictors (or 15 if verbal outcome is included). Specifically, the model would require behavioral terms for Group, Condition, Group*Condition, Mean-age, Group*Mean-age, Condition*Mean-age, Delta-age, Group*Delta-age, Condition*Delta-age, Age<sup>2</sup>, Group*Age<sup>2</sup>, Condition*Age<sup>2</sup>, Verbal-outcome, Group*Verbal-outcome, and Condition*Verbal-outcome.

      Although a unified multivariate model capturing the complex dynamics at play in our sample is theoretically appealing, the substantial risk of overfitting precludes its feasibility. Therefore, we opted to use only one categorical predictor per PLS-c analysis to maintain model parsimony and reliability. However, a larger sample could overcome this limitation, allowing a stable and unified model of longitudinal EEG data that simultaneously captures age trajectories, group, clinical outcome, and condition.

      Reference:

      Geladi, P., & Kowalski, B. (1986). Partial least-squares regression: A tutorial. Analytica Chimica Acta, 185, 1–17. https://doi.org/10.1016/S0003-2670(00)82582-3

      We added the following comment in the method section:

      Page 24, lines 773-777: “We limited the number of behavioral variables to nine to mitigate noise sensitivity and overfitting risks associated with exceeding the 10% sample size threshold (Geladi & Kowalski, 1986). This limitation precluded the implementation of a single PLS-c model incorporating group, condition (STR vs. RND), age, and their interactions.”

      (4) STR-only analyses and the role of RND

      Page 8, Lines 241-245: This analysis is conducted only within the STR condition. The lack of group difference observed here appears consistent with the lack of group difference in word-level entrainment (Page 9, Lines 292-294), suggesting that HL and LL groups may not differ in statistical learning per se, but rather in syllabic-level entrainment. As a useful sanity check and potential extension, it might be informative to explore whether syllable-level entrainment in the RND condition differs between groups to a similar extent as in Figure 2C-D. In other work (e.g., adults vs. children; Moreau et al., 2022), group differences can be more pronounced for syllable-level than for word-level entrainment. Figure S6 seems to hint that a similar pattern may exist here. If feasible, including or briefly reporting such an analysis could help clarify the asymmetry between the two learning measures and further support the interpretation of syllabic-level differences.

      The reviewer points to the interesting pattern highlighted in supplementary figure S6, suggesting that group differences in syllabic entrainment might be modulated by the structure of the stream (STR versus RND). Such modulatory effect of stream structure on entrainment to syllables has been suggested by many studies, like Moreau et al (2022), as pointed by the reviewer, and seems at play in our sample, as illustrated on supplementary figure S5 (decline in the 4hz PLV that exceeds the size of confidence intervals, ~90 s after STR onset).

      Following the reviewer’s suggestion, we ran a PLS-c testing group effect on 4hz PLVs in each stream:

      (1) in the RND stream: the analysis yields one significant component (p<.001, r=.49, 52.7% explained covariance, Author response image 2A-B). Bootstrap ratios (BSR) are: contrast (low versus high likelihood) 6.9; mean age -1.0; contrast*mean-age 0.1; delta-age 11.8; contrast*delta-age 0.1; age<sup>2</sup> -4.9; contrast*age<sup>2</sup> 4.6; verbal-outcome 9.7; contrast*verbal-outcome -0.6. Interestingly, the model still highlights a strong link between syllable tracking and group, suggesting that RND also discriminate between HL and LL. However, RND syllable tracking doesn’t appear to be linked to group x verbal-outcome as we observed in Figure 2C-D.

      (2) In the STR stream, we obtained one significant latent component (p=.002, r=.51, 57.8% explained covariance, Author response image 2C-D). Bootstrap ratios (BSR) are: contrast 2.2; mean age -1.6; contrast*mean-age -1.1; delta-age 6.1; contrast*delta-age -0.7; age<sup>2</sup> -4.4; contrast*age<sup>2</sup> 0.5; verbal-outcome 8.2; contrast*verbal-outcome -7.3. Here, the strong association between syllable tracking and group x verbal-outcome is similar to the model presented in Figure 2C-D.

      Taken together, these results suggest that the apparent STR/RND dissociation illustrated in Figure S6 might primarily reflect a Group*Verbal-outcome divergence, with syllable tracking in the STR stream being related to verbal outcome mainly in high likelihood for autism.

      Author response image 2.

      Syllable entrainment within RND (A-B) and STR (C-D).

      These results were reported in the revised manuscript in the Result section (Time course of the entrainment along experiment subheader), implying a slight reframing of the result presentation of supplementary analysis S6. Author response image 2 was added in supplementary material as Figure S5.

      Page 10, lines 307-319: “The group, age and verbal outcome parameters were mainly correlated (BSR>2.3) with the neural entrainment occurring~90 seconds after the onset of the STR stream, coinciding with the time participants began tracking word boundaries (Supplementary material, S3). This result suggests that the group differences in syllable entrainment, as shown in Figure 2C-D, as their associations with verbal outcome, are modulated by the structure of the stream (STR versus RND). We ran one additional PLS-c for each stream separately, using group as contrast. In both streams, the PLS-c yielded a significant LC (p<.001 for RND and p=.002 for STR), with a positive group effect (BSR>2.3) in both LC (Supplementary material, S5). Most strikingly, the group*verbal outcome parameter reached significance exclusively within the STR latent component (BSR:-7.3). These results suggest that while syllable tracking is generally decreased in HL infants across both streams, its association with verbal outcome is prominently driven by the stream containing words (STR).”

      Page 14, lines 443-446: “This temporal overlap suggests that successful segmentation may enhance syllable tracking via top-down predictions of the next syllable, improving alignment to syllable onsets in LL infants as well as in HL with better verbal outcome.’

      (5) Multi-speaker input and voice perception (Page 15, Lines 475-483)

      The multi-speaker nature of the speech input is an interesting and ecologically relevant feature of the design, but it does add interpretive complexity. The literature on voice perception in autism is still mixed: for example, Boucher et al. (2000) reported no differences in voice recognition and discrimination between children with autism and language-matched non-autistic peers, whereas behavioral work in autistic adults suggests atypical voice perception (e.g., Schelinski et al., 2016; Lin et al., 2015). I found the current interpretation in this paragraph somewhat difficult to follow, partly because the data do not directly test how HL and LL infants integrate or suppress voice information. I think the authors could strengthen this section by slightly softening and clarifying the claims.

      We acknowledge the reviewer’s concern regarding the potential ambiguity in the cited paragraph. To address this, we have revised the text to explicitly clarify the aims of our study and its design. Furthermore, we now emphasize the speculative and post-hoc nature of the hypotheses and interpretations presented, thereby ensuring transparency regarding the limitations of our findings.

      Page 16 lines 520-530), as follows: “HL infants, on the other hand, did not show this transient disruption. In this group, word entrainment remained stable over time. To account for this unexpected finding, we followed up on the post-hoc hypothesis proposed above: a reduced sensitivity to social and vocal cues observed in HL infants may have spared segmentation abilities by limiting the interference introduced by speaker variability. If this post-hoc hypothesis holds true, LL and HL infants would differ not in their intrinsic ability to learn statistical regularities per se, but rather in how they integrate or suppress competing cues (such as speaker changes) during the segmentation process. It is important to note, however, that the present study was not designed to isolate and evaluate the specific impact of speaker changes on word segmentation. Consequently, this interpretation remains speculative, and additional research is required to further address this question.”

      (6) Asymmetry between EEG learning measures

      Page 16, Lines 502-507 touches on the asymmetry between the two EEG learning measures but leaves some questions for the reader. The presence of word recognition ERPs in the LL group suggests that a failure to suppress voice information during learning did not prevent successful word learning. At the same time, there is an interesting complementary pattern in the HL group, who show LL-like word-level entrainment but does not exhibit robust word recognition. Explicitly discussing this asymmetry - why HL infants might show relatively preserved word-level entrainment yet reduced word recognition ERPs, whereas LL infants show both - would enrich the theoretical contribution of the manuscript.

      We concur with the reviewer’s observation that our findings imply a theoretically significant double dissociation between HL and LL groups, specifically concerning the asymmetries between word-level neural entrainment and word recognition mechanisms. We believe this point was partly addressed in the subsequent paragraph, where we stated that “in contrast” to LL, HL infants “showed no clear ERP difference between novel and familiar triplets”, while “both groups showed similar word neural entrainment during learning”. We further explored potential explanations for this apparent dissociation, such as a possible deficit in novelty orientation that may be specific to HL infants and unrelated to statistical learning itself. We cited Liu et al (2023) as a reference showing the dissociation between mechanisms underlying implicit versus explicit traces of statistical learning. We acknowledge that we can discuss more in depth the potential preservation of statistical learning in HL infants. We have incorporated the following discussion in the reviewed manuscript, supported by relevant references:

      Pages 17-18, lines 562-566: “Interestingly, this dissociation between spared implicit versus impaired explicit statistical learning in autism has been previously discussed in the literature (Zwart et al, 2018, Kissine, 2021). According to these studies, autistic impairments in top-down attentional processes, such as social orienting — which are critical for bootstrapping language acquisition (Kuhl, 2007) — may result in a heightened dependence on bottom-up mechanisms, including implicit statistical learning.”

      References:

      Zwart, F.S., Vissers, C.T.W.M., Kessels, R.P.C. and Maes, J.H.R. (2018), Implicit learning seems to come naturally for children with autism, but not for children with specific language impairment: Evidence from behavioral and ERP data. Autism Research, 11: 1050-1061. https://doi.org/10.1002/aur.1954

      Kissine, M. (2021). Autism, constructionism, and nativism. Language 97(3), e139-e160. https://dx.doi.org/10.1353/lan.2021.0055.

      Kuhl, P.K. (2007), Is speech learning ‘gated’ by the social brain?. Developmental Science, 10: 110-120. https://doi.org/10.1111/j.1467-7687.2007.00572.x

      References:

      (1) Moreau, C. N., Joanisse, M. F., Mulgrew, J., & Batterink, L. J. (2022). No statistical learning advantage in children over adults: Evidence from behaviour and neural entrainment. Developmental Cognitive Neuroscience, 57, 101154. https://doi.org/10.1016/j.dcn.2022.101154

      (2) Boucher, J., Lewis, V., & Collis, G. M. (2000). Voice processing abilities in children with autism, children with specific language impairments, and young typically developing children. Journal of Child Psychology and Psychiatry, 41(7), 847-857. https://doi.org/10.1111/1469-7610.00672

      (3) Schelinski, S., Borowiak, K., & von Kriegstein, K. (2016). Temporal voice areas exist in autism spectrum disorder but are dysfunctional for voice identity recognition. Social Cognitive and Affective Neuroscience, 11(11), 1812-1822. https://doi.org/10.1093/scan/nsw089

      (4) Lin, I.-F., Yamada, T., Komine, Y., Kato, N., Kato, M., & Kashino, M. (2015). Vocal identity recognition in autism spectrum disorder. PLOS ONE, 10(6), e0129451.https://doi.org/10.1371/journal.pone.0129451

      Reviewer #2 (Public review):

      Summary:

      This article looks at differences in how the brain entrains to, or tracks, the rhythmic presentation of syllables and words in speech in infants at increased likelihood versus low likelihood for autism. The authors first sought to characterize how brain responses are modulated by learning the statistical probability of a given syllable following the one before it over the first two years of life. They then sought to identify at which stages of word learning infants with increased likelihood of autism showed difficulties, and whether those difficulties worsened over time. Finally, they sought to indicate whether infants' statistical learning and word learning abilities could predict later verbal skills. The authors found similar developmental trajectories of neural entrainment to syllables in infants at high and low likelihood for autism, but infants at high likelihood for autism had overall weaker syllable-level entrainment. Infants at high versus low likelihood for autism showed different developmental trajectories for word entrainment. Lower syllable entrainment in high-likelihood infants corresponded with poorer verbal outcomes, but word entrainment was not associated with verbal outcomes. Event-related potential responses to words and part words were positively associated with verbal outcomes, however, but only in low-likelihood infants.

      Strengths:

      Overall, the article provides rigorous statistical analysis of longitudinal EEG data to provide strong support for the claims that neural entrainment to syllable and word features of speech may be a useful marker for language development difficulties, particularly in infants at increased likelihood for neurodevelopmental disorders. The EEG data collection and preprocessing procedures are well within standards in the field. Readers should take care to note that authors indexed neural entrainment to speech using phase-locking values instead of spectral power.

      Weaknesses:

      While the statistical analyses are rigorous, a few of the components of the models are not clearly defined, and some corrections and thresholds for significance warrant further justification. Further, a few stimuli and participant details that could influence results are not specified. It is not clear whether all participants came from majority French-speaking families; differences in the amount of French language exposure (compared to other languages that may be spoken by a participant's family) could influence results. The standardized volume of the stimuli is also not included. As a result, readers should be encouraged to interpret that neural entrainment to speech features is likely a useful mechanism to explain differences in language development, while taking this interpretation with some caution.

      We thank the reviewer for these remarks.

      Regarding the amount of French exposure: while all participants were raised in primarily French-speaking environments (i.e., French as the dominant language at home and daycare), the parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. We did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism. The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      Regarding the volume of stimuli, they were played at 50cm distance with an intensity of 75dB. Both considerations have been included in the new version of the manuscript. In general, we moved most of the figures present in Supplementary material to figure supplements to improve readability.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Comments:

      Figure 6: The figure caption is not complete (there is no description for the right half of panel B).

      We thank the reviewer for this observation, figure 6 caption has been completed.

      Reviewer #2 (Recommendations for the authors):

      Broadly speaking, I would recommend reducing the number of abbreviations in this article, and I would recommend that the authors take care as to where these abbreviations are being introduced. Many of the abbreviated terms are defined in the Materials and Methods section, which is presented after the abbreviations are used in the main results.

      We acknowledge that our extensive use of abbreviations compromises the readability of the manuscript. Consequently, we have removed the following abbreviations:

      - SL (replaced by statistical learning)

      - LC (replaced by latent component)

      - ASD (replaced by autism)

      - TP (replaced by transition probability)

      - MEG (replaced by magneto-encephalogram)

      - MSEL (replaced by Mullen Scale of Early Learnings)

      The remaining abbreviations are:

      HL (high likelihood for autism), LL (low likelihood for autism), EEG (electroencephalogram), PLS-c (partial least square correlation), ERP (event-related potential), RND (random), STR (structured), BSR (bootstrap ratio), AIC (Akaike Information Criterion), PLV (phase locking value), DQ (developmental quotient), APSI (Autism Parent Screen for Infants).

      Moreover, we carefully reviewed how abbreviations were introduced and identified that PLS-c, STR and RND were not defined prior to the Method section. This oversight has been corrected in the reviewed manuscript.

      I would also recommend that the authors be careful with the structuring of the Introduction, particularly with their research questions and hypotheses. The article initially makes clear that the research questions are focused on the developmental trajectory of statistical learning, the levels of word learning that may differentiate high-likelihood versus low-likelihood infants, and the stability of those differences, and associations between statistical learning and various levels of word learning with verbal outcomes. The use of acoustic variability across syllables, while a valuable methodological tool, is somewhat presented as an additional research question, but not clearly stated or tested as such.

      We acknowledge that the introduction (particularly the paragraph from lines 173 to 184) may have implied that speaker variability across syllables was one of our primary research aims. We clarify here that speaker variability was introduced as a mean to increase task difficulty, particularly for high-likelihood (HL) participants, with the aim of amplifying the effect sizes in our analyses.

      To address this, we have removed the theoretical discussion on speaker variability in autism and typical development (lines 173–184) and explicitly stated that speaker variability was not a research question in this study. Crucially, our experimental design did not include a control condition without speaker variability, and thus we could not test its specific effects on statistical learning across age trajectories and groups.

      Page 6, lines 173-176 (pages and lines refer to the reviewed uploaded manuscript): “It is worth noting, however, that our study was not designed to isolate or quantify the specific impact of speaker variability on statistical learning, as the experimental design did not include a baseline control condition omitting this acoustic variation.”

      The authors do a nice job in the Materials & Methods explaining PLS-c and defining the latent components and bootstrapped ratios that will be shared in the Results. An additional brief iteration defining these statistical elements is needed at the beginning of the Results section.

      We thank the reviewer for their appreciation of our Method section. We agree that an additional iteration in the result section would improve readability. We added the following paragraph at the very beginning of the Result section, briefly defining PLS-c and its main statistical output (latent components and bootstrap ratios):

      Pages 6-7, lines 193-202: “Briefly, PLS-c is a data-driven multivariate modelling approach designed to identify significant patterns of electrode clusters (from a brain data matrix containing electrophysiological measures, here PLV) and their associations with “behavioral” variables (from a behavioral design matrix, here age-related parameters). Patterns of brain x behavior associations are called latent components, and their statistical significance is evaluated using permutation testing (n=1000, Bonferroni correction for number of components tested, alpha=.006). Brain and behavioral variables respective contributions to any significant latent component are tested with bootstrapping (500 random samples and replacement), with bootstrap ratios (BSR) greater than 2.3 indicating a stable contribution (for details, see the Materials and Methods section).”

      (1) Page 18 Line 576. The authors need to clarify whether participants were required to be in primarily French-speaking environments and whether there was a minimum amount of French language exposure that participants were required to have if they were exposed to additional languages besides French in their everyday life.

      The reviewer raises a valid concern regarding participants’ language exposure. In this study, all participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare. The parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. However, we did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism.

      The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      To include these considerations, Limitations and Material and methods sections were modified as follows:

      Page 19, lines 604-607: “Second, although all participants were primarily exposed to French, we did not quantify additional language exposure, precluding any analysis of its potential moderator effects on statistical learning in our groups and age-trajectories. However, prior work has reported no effect of bilingualism on auditory triplet segmentation in children (Yim & Rudoy, 2013).”

      Page 20, lines 630-631: “All participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare.”

      References:

      Weiss DJ, Schwob N, Lebkuecher AL. Bilingualism and statistical learning: Lessons from studies using artificial languages. Bilingualism: Language and Cognition. 2020;23(1):92-97. doi:10.1017/S1366728919000579

      Yim D, Rudoy J. Implicit statistical learning and language skills in bilingual children. J Speech Lang Hear Res. 2013 Feb;56(1):310-22. doi: 10.1044/1092-4388(2012/11-0243). Epub 2012 Aug 15. PMID: 22896046.

      (2) Page 18 Line 588. Further, the authors should clarify whether the 7 infants in the HL group, due to early parental concerns were defined by the 18-21-month APSI scores or by parental report prior to study enrollment.

      These 7 infants were recruited based on early parental concerns prior to intake. The APSI score at 18-21 months is only reported to provide an illustration of the amount of early autistic signs that were present in these 7 infants, and to provide an estimation of their probability to develop autism later on based on Sacrey et al., 2018 longitudinal study on the APSI predictive value. We agree with the reviewer that our phrasing suggests that the APSI was used as an inclusion criterion. We rephrased the page 20 lines 642-646 as follows:

      “The 7 other HL infants presented with early parental concerns for autism, based on parental report prior to enrollment. Their Autism Parent Screen for Infants (APSI) total score at their 18-21 months visit was 15.6±6.4, [8-22] range – a score greater than 8 reflecting a 63% positive predictive value for autism in HL populations.”

      (3) Page 20 Line 641. The authors should specify the volume of the stimuli.

      The volume of stimuli was reported in the main text (page 22, lines 695-696) as follows:

      “Stimuli were played on a Bose® Companion 2 Series III at a 50cm distance with an intensity of 75dB.”

      (4) I'd prefer Figure 1 to be reorganized slightly - at present, the placement of the arrows explaining the analysis steps is not intuitive.

      We addressed the reviewer’s comments (4) and (5) together as they both refer to Figure 1B.

      (5) Page 23 lines 718-719. I think it would be helpful to explicitly define each of the interaction variables included in the behavioral design matrix. Further, this matrix should be labeled consistently in both Figure 1B and in the main text.

      We refined figure 1B and its corresponding main text (in Methods section) for clarity. The arrows are now simpler and more parsimonious, labels (e.g., participant i, visit n, behavior design matrix and its parameters) are now standardized between the figure and the main text, and the interaction terms at lines 718-719 are explicitly defined.

      (6) Page 23 lines 726-731: It would be helpful to know whether applying a Bonferroni correction in addition to completing permutation testing is standard when evaluating latent components derived from PLS-c. The authors should also cite justification for a bootstrap ratio cutoff of 2.3 for defining stability.

      In PLS-c analyses, multiple comparisons correction across latent components and bootstrap ratio (BSR) thresholding at 2.3 are commonly adopted practices.

      - Correction for multiple comparisons in PLS-c: PLS-c performs singular decomposition of the data into latent components equal in number to the variables included in the behavior design matrix (7-9 in our study, depending on the inclusion of Verbal outcome as an input variable). Each latent component’s statistical significance is assessed through permutation testing, generating a null distribution for its singular value (Krishnan et al., 2011). Given the multiple tests (one permutation test per latent component), Type I error inflation must be addressed. Recent PLS-c studies commonly applied Bonferroni correction (default procedure in the myPLS toolbox, used by Zoeller et al., 2017, and Delavari et al, 2021), though FDR correction has also been used (Lombardo et al, 2018).

      - Stability threshold for bootstrap and replacement: Within each latent component, saliences’ stability (brain/behavior parameter contributions to each latent component) are evaluated using bootstrapping (Krishnan et al., 2011). The bootstrap ratio (BSR) of each parameter, calculated as the saliency divided by its bootstrap-derived standard error, functions analogously to a z-score under normality assumptions. The BSR can then be used to assess the stability of the saliency (i.e., how stable is its contribution to the latent component). BSR thresholds in the literature typically range from 1.96 to 3.0. Krishnan et al (2011) state that when BSR are “larger than 2 the corresponding saliences are considered significantly stable”. Delavari et al (2021) and our study used a 2.3 thresholding, corresponding to a 99.0% bootstrap confidence interval not crossing the zero line – roughly equivalent to a two-tailed p<.001. Lombardo et al (2018) used a looser threshold of 1.96, corresponding to a 95% confidence interval not crossing the zero line (~two-tailed p<.05), while Zöller et al (2017) used a more stringent 3.0 thresholding (~p<.001, or 99.9% confidence interval not crossing the zero line).

      Thus, our application of Bonferroni correction for multiple comparisons and our 2.3 BSR threshold aligns with established conventions.

      We added following lines in the manuscript:

      Page 25 lines 784-785: “Bonferroni correction was applied to account for multiple comparisons across the 9 tested latent components in the PLS-c, yielding an adjusted alpha of .006 (Zoeller et al, 2017; Delavari et al, 2021).”

      Page 25 lines 789-792: “BSR are analogous to Z-scores and can be used to assess the stability of the saliency. We considered BSR > 2.3 as stable, corresponding to a 99.0% bootstrap confidence interval not crossing zero – roughly equivalent to a two-tailed p<.001 (Delavari et al., 2021; Krishnan et al., 2011).”

      References:

      Delavari F, Sandini C, Zöller D, Mancini V, Bortolin K, Schneider M, Van De Ville D, Eliez S. Dysmaturation Observed as Altered Hippocampal Functional Connectivity at Rest Is Associated With the Emergence of Positive Psychotic Symptoms in Patients With 22q11 Deletion Syndrome. Biol Psychiatry. 2021 Jul 1;90(1):58-68. doi: 10.1016/j.biopsych.2020.12.033. Epub 2021 Jan 18. PMID: 33771350.

      Lombardo, M.V., Pramparo, T., Gazestani, V. et al. Large-scale associations between the leukocyte transcriptome and BOLD responses to speech differ in autism early language outcome subtypes. Nat Neurosci 21, 1680–1688 (2018). https://doi.org/10.1038/s41593-018-0281-3

      Daniela Zöller, Marie Schaer, Elisa Scariati, Maria Carmela Padula, Stephan Eliez, Dimitri Van De Ville. Disentangling resting-state BOLD variability and PCC functional connectivity in 22q11.2 deletion syndrome. NeuroImage, Volume 149, 2017, Pages 85-97, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2017.01.064

      Anjali Krishnan, Lynne J. Williams, Anthony Randal McIntosh, Hervé Abdi, Partial Least Squares (PLS) methods for neuroimaging: A tutorial and review, NeuroImage, Volume 56, Issue 2, 2011, Pages 455-475, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2010.07.034

      (7) I have a few minor grammar/formatting recommendations for the authors as well:

      (a) Should the Geneva Autism Cohort be capitalized? At present, it is not.

      We agree with the reviewer’s suggestion, and we capitalized the Geneva Autism Cohort in the main text (page 18, line 571)

      (b) Page 24, line 750. Do the authors mean that the data was re-referenced to average?

      The preprocessed data is not average-referenced (see section Data pre-processing). Therefore, both for neural entrainment computation and ERPs, the data were average-referenced.

      (c) It would be nice to have a figure of the actual ERP for each condition and age group.

      We agree that PLS-c can be difficult to interpret without the raw actual ERPs on which it was modelled. We direct the reviewer to supplementary figure S6 at page 59, which displays the raw ERPs for each condition (part-word, word, and their subtraction) per age group. Supplementary figures S7-8 at pages 60-61 further illustrate topographical ERPs for each group (high and low likelihood for autism). We deemed these figures too extensive for the main text. Instead, the most relevant ERP topographies are presented in Figures 4-6 to facilitate PLS-c interpretation.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      We sincerely thank the reviewers for the time dedicated to providing us with feedback on our work. In response to their helpful remarks, we have generated new data and made several changes to the manuscript. We hope the reviewers agree that, together with the rebuttal points below, this improved version addresses their concerns.

      Reviewer #1

      Evidence, reproducibility and clarity

      In this report the authors have provided evidence for the involvement of transposable elements in the regulation of gene expression in murine trophoblast cells and placenta. They utilized data that they generated and also published data in their analyses. They concluded that the involvement of transposable elements in the regulation of genes in differentiated trophoblast cells was less than utilized in trophoblast cells in the stem state. They provide evidence for the utilization of intracisternal A particle elements in the modulation of gene transcription of the mouse placenta, which represents a more recent evolutionary adaptation. Overall the report is descriptive presenting correlations with limited testing of specific hypotheses. There is also the impression that the manuscript consists of the merging of two projects, which have not been fully developed. Some concerns with the experimental design and interpretation of the results are provided below.

      1. Some concerns with the model systems used in the analysis. First of all, there are methods for inducing the differentiation of mouse trophoblast stem cells, which usually involves the removal of factors that promote trophoblast stem cell proliferation. The authors do not describe their method for inducing trophoblast stem cell differentiation nor did they show evidence that they directly investigated differentiated trophoblast stem cells.

      We have added information in the Methods section to clarify that differentiation was performed by culturing cells in TS base medium (no conditioned media, FGF or heparin) for 4 days. We also provide RT-qPCR data confirming TSC differentiation (Figure S1A).

      1. There is a published report presenting data from single cell analysis of mouse trophoblast stem cells in the stem and differentiated states that was not acknowledged or used in the authors' analyses. Please see: Angelova et al. 2025 Nature Communications (PMID:39747179).

      We appreciate the reviewer’s suggestion, but the purpose of our single-cell analysis was to assess the expression of IAP elements and their associated chimeric transcripts in vivo. This has more significance than single-cell data from in vitro differentiated cells. We also found that at least some of the chimeric transcripts seen in vivo are not detected in vitro.

      1. Much of the analysis, compared mouse trophoblast stem cells and murine placentas. Interpretation of single nucleus sequencing data from mouse placentas can provide information regarding the behavior of trophoblast cells; however, bulk sequencing of placentas is limited. The placenta contains trophoblast cell and non-trophoblast cell components. More specifically the placenta contains fetal endothelial, immune, and mesenchymal cells and depending upon dissections and the gestational stage of dissections variable amounts of uterine decidua and yolk sac-derived tissues. The authors need to be clear in the comparisons that they are making. More specifically, the authors need to effectively communicate the cell types from the placenta contributing to the results they are describing. Attributing analyses of the placenta to trophoblast cells is problematic.

      We agree with this point. In our original submission we had included a cell type deconvolution analysis to infer the composition of our bulk placental tissue (Figure S1B of the revised submission). This clarifies the heterogeneity of the tissue and shows that most cells are trophoblast. We also used a genetic model to isolate trophoblast from placentas to address this point. Cell type deconvolution confirms that >90% of cells are trophoblast.

      1. How were the newly derived mouse trophoblast stem cells characterized? Do they behave like authentic mouse trophoblast stem cells? Were the newly derived trophoblast stem cells capable undergoing differentiation? What parameters were measured?

      We now include data on trophoblast stem cell and differentiation markers, comparing our newly derived line with the well-established GFP-TSC line (Figure S1A). While not included in this submission, the cells also presented with the expected morphologies when cultured under stem or differentiation conditions.

      1. Were analyses with the newly derived mouse trophoblast stem cells performed in the stem or differentiated states?

      Our initial analyses were only from cells cultured in stem conditions. However, we now also include an analysis of TE regulatory activity in differentiation conditions (updated Figures 1B, 1D, S1C).

      1. TSC derivation and culture section. The authors appear to be initially describing the generation of mouse embryonic fibroblast conditioned medium not trophoblast stem cell conditioned medium as stated. Some clarification will be helpful. As stated above, the authors do not provide any information on the characterization and validation of the newly derived mouse trophoblast stem cells, which is problematic.

      We appreciate the confusion with the nomenclature. To clarify we have changed the start of that section to: “Conditioned medium for the culture of TSCs (TS-CM) was prepared by…”. This conditioned medium is generated using MEFs and is then used to culture TSCs.

      1. Discussion. The authors state that there are fundamental differences between the mouse and human placenta regarding the co-option of transposable element subfamilies. Human trophoblast stem cells represent a highly tractable model and could be compared with mouse trophoblast stem cells to further explore this observation.

      We previously published a paper focused on TE co-option in human trophoblast (PMID: 37012406), and we made a brief comparison to those data in the current manuscript (Figure S1G). Interestingly, in contrast to mouse, many of the TEs with regulatory activity in human TSCs remain active in term placenta.

      Significance

      Efforts to understand roles for transposable elements in the regulation of trophoblast cell gene expression and placental evolution are very important. We recognize significant differences in placentation across various species but do not have a good understanding how this important developmental process evolved.

      Assessment: The authors have a potentially interesting story. However, it appears that they have merged two incomplete research efforts: i) effects of trophoblast cell differentiation on utilization transposable elements to regulate gene expression; ii) IAP involvement in regulating murine placental transcription.

      Whilst we appreciate this viewpoint, our investigation of IAPs as regulators of gene expression was triggered from the analysis in the first part of the manuscript and thus follows logically in our view. Moreover, the overarching theme remains consistent: the effects of TEs (whether IAPs or others) on gene expression/transcription.

      Advance: The scientific advance is somewhat fragmented. There is a reinforcement of our existing understanding of the involvement of transposable elements in trophoblast and placental gene regulation but other new insights are limited or not well developed.

      Both the TE and placental scientific communities largely assume that the placenta is a privileged organ for co-option of TEs as regulatory elements. Here we demonstrate that TE co-option in the placenta can be quite limited and species-specific. Additionally, the roles of IAPs as gene regulators in the placenta had not been previously described. We believe these two novel observations constitute significant advancements in the field.

      Audience: Evolutionary biologists and reproductive and developmental biologists.

      Reviewer #2

      Evidence, reproducibility and clarity

      Summary: This manuscript describes the characterization of transposable elements (TEs) in mouse trophoblast stem cells and in the mature mouse placenta. The authors find that overall, the trophoblast stem or progenitor state of TSCs harbours a greater abundance of active TE elements, while their activity levels decline as trophoblast differentiates. Instead, the dominant repetitive element that is active in the mature placenta are intracisternal A particle (IAP)-derived elements. Indeed, the authors show that these provide the initiation sites for differential isoforms of some 27 chimeric transcripts that are specific to differentiated trophoblast cell types. The authors attempt to epigenetically silence these IAPs in TSCs using CRISPRi methodolgy, and find reduced expression of 4 IAP-driven transcripts and many presumably secondary transcriptional changes. Finally, they also compare IAP activity in the placentas of different mouse species or sub-species, and conclude that IAP insertion sites close to genes can affect their expression in the placenta, with potential consequences for development and evolution.

      Major comments: This is a well-conducted study that brings significant novelty, albeit to a more specialized audience.

      There are several aspects that need clarification, addition and some experimental work: 1. Figure 1A shows carefully separated cell types, in particular extraembryonic mesoderm, that have also been assessed by the various cut&tag and ATAC-seq methods, but are not mentioned in the remainder of the manuscript. This should be added. I.e., is the same activity pattern of IAPs evident in the ExMes cells, or do they follow a more somatic pattern?

      Apologies if additional analyses of extraembryonic mesoderm were not obvious, but we did analyse IAP expression in these cells and show in Figure 2B that it is much lower when compared to trophoblast. We also used the comparison between trophoblast and extraembryonic mesoderm in Figures 2A, 3B, S2A-D and S4B.

      1. Page 4, top: The mention of a "custom pipeline" for cut&tag analysis is vague, and the modifications and what they stand for is hardly mentioned. These details need to be elaborated, so to be more accessible to a wider audience.

      We have tried to clarify the overall strategy of the analysis: “Using CUT&Tag and ATAC-seq data, we aimed to identify TE subfamilies that bear classic hallmarks of active promoters (open chromatin, H3K4me3, H3K27ac) and/or enhancers (open chromatin, H3K27ac, H3K4me1). We used a custom pipeline that selects TE subfamilies bearing more elements overlapping CUT&Tag/ATAC-seq peaks than expected by chance.”.

      1. For differentiated TSCs, only ATAC-seq data were analysed. How do they relate to the various cut&tag profiles, and do they result in a robust detection of putative active repeat elements at a detection limit similar to the chromatin marks? I would think that it might be prudent to include the same cut&tag for differentiated TSCs as well, so to be directly comparable to the other data. This is important to establish whether TE elements are really less active in differentiating trophoblast, or whether this feature is intrinsic to the placenta and not to pure trophoblast cells in culture, in which case it may be influenced by tissue context.

      We are thankful for this important suggestion. We have now carried out CUT&Tag on differentiated cells and include the findings in the revised Figures 1B, 1D and S1C. Consistent with our observations using ATAC-seq data, we find that TE regulatory activity is diminished upon in vitro differentiation.

      1. Are the IAP-initiated chimeric transcripts including new coding regions? If so, a Western Blot analysis of a few of the 27 candidates should be performed to prove this. Suv39h2 is a particularly interesting candidate where such protein analysis would be very informative.

      We performed a search for ORFs in IAP-driven transcripts and identified a putative protein isoform of SUV39H2 that includes a portion of the IAP and that is larger than the canonical form by 28 kDa. However, by Western blot we see no major size shift in the main band when comparing placenta (where the IAP isoform predominates) with TSCs (where only the canonical form is expressed). We now include this in a new Supplementary Figure S5. To note is that in our hands the main SUV39H2 band runs at a lower molecular weight than expected (54 kDa), which could be due to buffer/gel conditions and/or expression of a shorter isoform (ENSMUSG00000026646, 46 kDa). But we are reassured that the antibody used has been validated in multiple human KO lines, as well as in at least one mouse knockdown model (PMID: 32698678).

      1. A WB analysis should for sure be performed on the M. musculus and M. pahari placentas. The IHC staining is not interpretable as to whether or not SUV39H2 levels are reduced in M pahari.

      We appreciate the reviewer’s point, but the main hypothesis to be tested here was whether there was an obvious difference in the spatial distribution of SUV39H2, which we did not find. Any more subtle differences would be cell-type specific and would require complex cell sorting approaches before attempting a western blot. This would not affect our conclusion that, despite differences between species at the transcriptional level, this does not lead to an overt redistribution of SUV39H2 protein expression.

      1. Could the authors please also provide more global proof of the CRISPRi success. The display of two candidate gene tracks is not very telling.

      In the original submission we had included a subfamily-level analysis of IAP expression in the CRISPRi experiment (Figure S6A of the revised version). This shows a mild downregulation of IAP expression overall. Whilst an element-based analysis would be preferable due to potential caveats with subfamily-level analyses, very few TSC-expressed elements are sufficiently mappable to ensure a robust analysis, which is why we only showed two highly expressed loci where the effects of CRISPRi can be evaluated. Importantly, we observe effects on gene expression that, whilst mild, are non-random and support a role for IAPs in regulating the expression of nearby genes (Figure 4D).

      Significance

      General assessment: Collectively, this is a carefully conducted study that needs to be bolstered by some few additional experiments, as suggested above. The discovery of changing patterns of repetitive element activity in differentiating trophoblast cells is important and intriguing, as it has direct impact on the evolutionary divergence of gene expression and, as a consequence, cell type differentiation, through the insertion of IAP and L1 elements close to placenta-expressed genes. This will be a major contributor and even driver of the barrier to inter-species hybridization that the placenta represents.

      Advance: Currently, the main TE elements known to drive placenta-specific gene expression are retrovirally derived LTR elements. Here, however, the authors show that the relevance of these elements diminishes in the mature placenta, and instead is taken over by a different class, the IAP elements. This is important, as many these elements retain the capacity for retrotransposition, and thus actively contribute to ongoing evolutionary divergence of placental gene expression patterns that ultimately may drive speciation.

      Audience: The manuscript is not particularly easy to follow, even for the informed reader, and it appeals to a relatively specialized audience in the field of genome regulation coupled to evolutionary aspects of repetitive element insertion/transposition. The authors should be encouraged to spell out some aspects of their thought process throughout the study in some more detail, so not to "lose" the reader.

      We have made multiple changes throughout the manuscript that we hope improve readability.

      __Reviewer #3 __

      __Evidence, reproducibility and clarity __ The cis-regulatory roles of TEs in human/mouse TSCs have been extensively studied, yet in vivo studies on their roles in placenta tissue is largely absent. In this manuscript, Amante and colleagues compared the regulatory landscape of TEs across the trophoblast cell lines and placenta samples in human and mouse, and after revealing the shared and species-specific patterns (including some that are surprising), they further investigated the regulatory function of the murine-specific IAP retrotransposons in house mouse and other mouse strains. Specifically, it presents several findings regarding the shared and diverged function of TEs across: 1) in vivo vs. in vitro placental models, 2) human vs. mouse, 3) and different mouse strains. The writing is of good quality, the results are well visualized and interpreted, the conclusions are reasonable, and the novelty is high. It significantly extended previous studies from the same group as well as many other researchers. I think this manuscript should fit publication after a minor revision. Below I have a few comments:

      1. In Fig. 1B, it seems the differences of TE enrichment between the same groups of samples (e.g., B6 TSC vs. GFP TSC) is also remarkable. Is this expectable? I am curious if such difference is robust, or it is just due to the TE sub-families with too few copies, whose enrichment can be influenced by just a couple of overlapping counts. The authors may double-check if possible.

      This is an interesting hypothesis, but the main subfamilies that are H3K27ac-enriched in TSCs are quite abundant (e.g., 683 copies of RLTR13D5, 260 copies of RLTR13B3). We believe these are cell line-specific differences, possibly partly driven by genetics, since they were derived from different mouse strains. Nonetheless, there is good agreement between the two lines with respect to the TE subfamilies that are enriched.

      1. The authors demonstrate that the association of TEs to cis-regulatory elements is much weaker in the placenta of mouse relative to human, and in mouse the activation of TEs is indeed similar to most other tissues. And based on this observation, they propose that "Co-option of TEs as regulatory elements within the mature placenta may therefore not be as promiscuous across species as commonly thought" (page 4 paragraph 1). While this finding is quite interesting, how it is related to the popular hypothesis that "maternal-fetal conflict leads to the strong TE activation in placenta"? I am curious if the authors have any idea on this point.

      It is indeed a fascinating topic. We would dispute that the conflict hypothesis leads to TE activation in the placenta, but rather that it creates selective pressures that drive their co-option. But this still requires for TEs to be available for co-option. What we suggest here is that TE co-option opportunities can be tightly constrained by transcriptional silencing mechanisms, even in the placenta. We added the following text to that section of the discussion: “Whilst maternal-fetal conflicts may create selective pressures for TE co-option in the placenta, epigenetic mechanisms can still act as gatekeepers and dictate the frequency of co-option events in this organ.”

      1. For the highly active IAP subfamilies identified in mouse placenta, have the authors tried to identity the enriched motifs, which may be helpful for uncovering transcription factors responsible for their activation?

      This is an interesting question, given the specific expression of IAP elements in the spongiotrophoblast. We now performed transcription factor motif analysis on subfamilies that are highly expressed in the placenta (IALTR1/2). We then filtered this list for motifs that are absent/mutated in lowly expressed IAP subfamilies (IAPLTR3/4) and whose associated transcription factor is highly expressed in spongiotrophoblast. In the revised manuscript we highlight our top candidate, MITF, which is a spongiotrophoblast-specific marker (Figure S3C).

      1. In Fig. 3B, the IAPEY_LTR-adjacent Zfp229 gene is demonstrated, yet this gene is not mentioned at all in the main text. The authors may consider providing more details for this gene.

      Unfortunately, nearly nothing is currently known about this zinc finger protein gene, but we did not feel that should prevent us from using it as a strong example of placenta-specific usage of an IAP-derived promoter. Future work on this gene may indeed be triggered by highlighting this observation.

      1. A few errors for the citations should be corrected. For example, the journal names are missed for ref56 and ref58 at page 22.

      We have reviewed all our references and added missing information

      1. A few typos should be corrected. For example, at page 11 line 2, "of" is missed between "presence this IAP-driven.

      We have corrected this typo and made additional changes to the manuscript to improve readability.

      Significance

      Overall, this is an interesting and technically-sound study with substantial novelty, which significantly extends previous knowledge on TE function in placenta which largely relies on in vitro models.I believe this study will be attractive to the fields about TE function and placenta evolution.

    1. The reasoning goes that if there is always a high level of background risk to humanity, then we should expect to go extinct soon anyway, which means the importance of avoiding any one particular risk is not as valuable as it may seem. For more details see the full report here.

      This seems rather intuitive to me, but it's asking a slightly different question than what the original phrasing might seem to imply.

      I think the initial intuition that more risk means more value of reducing risk, comes from the natural idea that effort spent reducing a particular risk will reduce that risk proportionally. So, spending effort on reducing risks from car crashes, malaria in Africa, or heart disease, all else equal, we yield more value than spending comparable effort on reducing the risks of bear attacks. I guess this is the "importance" part of the ITN paradigm.

      But of course, the benefit of reducing the risk of car crashes is lower if we are facing other impending doom. Let's say we see an asteroid coming toward the Earth, or the threat of incoming nuclear war is high.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      This study elucidates the molecular linkage between the mobilization of damaged rDNA from the nucleolus to its periphery and the subsequent repair process by HDR. The authors demonstrate that the nucleolar adaptor protein Treacle mediates rDNA mobilization, and the MDC1-RNF8-RNF168 pathway coordinates the recruitment of the BRCA1-PALB2-BRCA2 complex and RAD51 loading. This stepwise regulation appears to prevent aberrant recombination events between rDNA repeats. This work provides compelling evidence for the recruitment of the Treacle-TOPBP1-NBS1 complex to rDNA DSBs and demonstrates the critical role of MDC1 in the rDNA damage response. There are some issues with the over-interpretation of results as described subsequently. Some aspects could be strengthened, for example, a potential role of the RAP80-Abraxas axis, the origin of the repair synthesis (HDR vs. NHEJ), and a direct comparison of the RNF8 and RNF168 recruitment in the absence or presence of MDC1.

      We thank the reviewer for the positive assessment of our work and for the constructive suggestions. We agree that certain aspects of the manuscript required clarification, in particular the potential contribution of the RAP80–ABRAXAS pathway, the interpretation of the repair synthesis assay, and the role of RNF168 recruitment. We have addressed these points experimentally where feasible and have revised the manuscript accordingly to avoid overinterpretation.

      Reviewer #1 (Recommendations for the authors):

      Major comments

      (1) In Figures 4C, 4D, and S4B-D, BRCA1 and RAD51, recruitment to nucleolar caps is partially reduced upon RNF168 depletion. Despite this, the authors broadly conclude that recruitment mainly depends on the MDC1-RNF8-RNF168 pathway. Since the RAP80-Abraxas pathway may also contribute, as briefly mentioned in the Discussion, siRNA knockdown of Abraxas would help clarify the relative roles of these two pathways.

      We thank the reviewer for this important suggestion. To directly address the potential contribution of the RAP80–ABRAXAS pathway, we depleted RAP80 by siRNA and analysed BRCA1 and RAD51 recruitment to nucleolar caps following I-PpoI-induced rDNA damage.

      Strikingly, RAP80 depletion strongly impaired the formation of both BRCA1 and RAD51 nucleolar caps in two independent cell lines (U2OS and RPE1) (new Figure 5). These findings demonstrate that the RAP80–ABRAXAS pathway plays a critical role in BRCA1 recruitment at nucleolar caps.

      Together with our observation that RNF168 depletion only partially reduces BRCA1 and RAD51 recruitment, these results indicate that both RNF168-dependent and RAP80–ABRAXAS-dependent pathways contribute to HDR factor recruitment downstream of RNF8-mediated chromatin ubiquitylation.

      We have revised the model Figure (Figure 9) and the Results and Discussion sections accordingly to reflect this dual-pathway model and to avoid overemphasising the contribution of RNF168.

      (2) In Figure 7C, the EdU-γH2AX PLA assay detects DNA synthesis at rDNA breaks, but it remains unclear whether this signal reflects HDR- or NHEJ-mediated repair. Since MDC1 functions upstream of the DSB repair pathway choice, the observed reduction in PLA signal upon MDC1 depletion does not necessarily reflect impaired HDR alone. Synchronizing cells in G2 or using cell cycle markers would help clarify the repair context and strengthen the interpretation.

      We thank the reviewer for raising this important point. We agree that the EdU–gH2AX PLA assay does not exclusively report on HDR-mediated DNA synthesis and may also capture other forms of repair-associated DNA synthesis.

      In the revised manuscript, we have therefore tempered our interpretation and now describe this assay more cautiously as a readout of DNA synthesis at sites of rDNA damage, rather than as a direct measure of HDR activity.

      Importantly, our conclusion that MDC1 promotes HDR factor recruitment at nucleolar caps is based primarily on the reduced accumulation of BRCA1, PALB2, and RAD51, which are well-established markers of HDR. The PLA assay is now presented as supportive evidence for ongoing DNA synthesis at these sites rather than as a definitive indicator of HDR.

      We agree that further experiments, such as cell cycle synchronization or the use of phase-specific markers, would help to more precisely define the repair context, and we have included this point in the Discussion.

      (3) The authors propose that MDC1 is essential for RNF8-RNF168 recruitment, specifically at nucleolar rDNA breaks. A side-by-side comparison of RNF8 or RNF168 localization in the presence and absence of MDC1, with IR-treated conditions, would provide important validation of this model. Including representative images in Figure S2C would further support the claim.

      We agree with the reviewer that a direct analysis of RNF8 and RNF168 recruitment in the presence and absence of MDC1 would provide valuable mechanistic insight. We therefore attempted to address this experimentally.

      However, despite testing multiple antibodies, we were unable to obtain specific and reproducible signals for RNF8 and RNF168 at nucleolar caps, precluding a reliable analysis of its recruitment under these conditions.

      Given this technical constraint, we have revised the manuscript to avoid overinterpretation regarding direct RNF168 recruitment and instead focus on functional readouts of downstream ubiquitylation-dependent signalling, such as BRCA1 and RAD51 accumulation.

      We note that the requirement for MDC1 in BRCA1 and RAD51 recruitment at nucleolar caps is consistent with a role of MDC1 upstream of RNF8-dependent chromatin ubiquitylation, in line with its established function at IR-induced DSBs.

      Minor comments:

      (1) The legend for Figure 8 should more clearly explain the proposed mechanism and include concise titles or descriptions for each sub-panel.

      We agree with the reviewer that the model should be described in the Figure legend. We have thus updated the model to accommodate the new data and wrote a legend that concisely explains the proposed model. We do not think that titles for each sub-panel are required. Instead, we separately referred to the sub-panels in the legend.

      (2) Typos:

      (a) Page 8: PRE1 MDC1, as "RPE1 MDC1;

      (b) S3 Figure legend: Dhermacon";

      (c) Page 29: "80.103"-please clarify or correct.

      We thank the reviewer for pointing out these errors. These have been corrected in the revised manuscript

      Reviewer #2 (Public review):

      Summary:

      DNA double-strand breaks (DSB) in repeated DNA pose a challenge for repair by homologous recombination (HR) due to the potential of generating chromosomal aberrations, especially involving repeats on different chromosomes. This conceptual caveat led to a long-held notion that HR is not active in repeated DNA, which was disproven in groundbreaking work by Chiolo showing in Drosophila that DSBs in pericentromeric repeats are mobilized to the nuclear periphery for repair by HR. A similar mechanism operates in mouse cells, as shown by the Gautier laboratory, but the mobilization goes to the nucleolar periphery, called nucleolar caps. In this manuscript, the authors reexamine the role of MDC1 in the mobilization of DSBs in rDNA in human cells. Previous work has shown that MDC1 is replaced by Treacle, the gene associated with Treacher Collins syndrome 1, in its role as the main adaptor of the DNA damage response, and these results are confirmed here. The novelty of this contribution lies in the discovery that MDC1 is required downstream in the recruitment of BRCA1 and RAD51 to nucleolar DSBs that were mobilized to the nucleolar cap. Using multiple MCD knockout models and DSBs induced by the nuclease PpoI, which cleaves at nuclear sites as well as in the 28S rDNA, convincingly documents this role of MDC1 and shows that it acts upstream of the RNF8-RNF168 ubiquitylation axis. Using a proxy assay of co-localization of EdU incorporation at DSBs (gammaH2AX), evidence is provided that MDC1 is required for HR in rDNA. MDC1 was not required for RAD51 recruitment to IR-induced foci, but it is unclear whether this is related to the different DSB chemistry (enzymatic versus IR) or to the localization of the DSB (rDNA versus unique sequence genome).

      Strengths:

      (1) The manuscript is well-written, and the experimental evidence is nicely presented.

      (2) Multiple MDC1 knockout models are used to validate the results.

      (3) Convincing back-complementation data clarify the relationship between MDC1 and RNF8.

      Weaknesses:

      (1) The recruitment of BRCA2 was not directly demonstrated. This caveat could be recognized, as IF for BRCA2 is challenging.

      (2) PpoI also induces DSBs in the non-rDNA genome. These DSBs would be an ideal control to establish nucleolar specificity of the events described and clarify whether the difference between IR and PpoI is the chemical structure of the DSB or the location of the DSB.

      We thank the reviewer for the positive and insightful evaluation of our work. We appreciate the recognition of the conceptual advance and the robustness of our experimental approaches. We have carefully considered the reviewer’s suggestions and have revised the manuscript to clarify interpretation where appropriate, particularly regarding BRCA2 recruitment and the specificity of I-PpoI-induced DNA damage. Where possible, we have also added new analyses to strengthen the conclusions.

      Reviewer #2 (Recommendations for the authors):

      (1) The claim that the BRCA1-PALB2-BRCA2 is recruited (abstract, end of results section, discussion page 15) should be qualified as BRCA2 recruitment was not directly demonstrated.

      We thank the reviewer for this important point. We agree that BRCA2 recruitment was not directly demonstrated in our study, as reliable immunofluorescence detection of BRCA2 remains technically challenging.

      We have therefore revised the manuscript throughout (Abstract, Results, and Discussion) to avoid overstatement and now refer more precisely to the recruitment of BRCA1, PALB2, and RAD51, rather than implying direct recruitment of a BRCA1–PALB2–BRCA2 complex.

      We note that BRCA2 function is supported indirectly by the observed RAD51 loading, which depends on BRCA2 activity. However, we have clarified this point to ensure that our conclusions remain fully supported by the presented data.

      (2) The temporal sequence established in Figure 1, 1hr BRCA1 and 2 hrs PALB2, argues against recruitment of a stable BRCA1-PALB2-(BRCA2) complex. This should be acknowledged.

      We thank the reviewer for this insightful observation. We agree that the temporal separation between BRCA1 accumulation (1 h) and PALB2/RAD51 recruitment (2 h) argues against the recruitment of a pre-assembled, stable BRCA1–PALB2–BRCA2 complex.

      We have revised the manuscript to reflect this interpretation and now describe the recruitment of HDR factors as a sequential process rather than as the assembly of a pre-formed complex. This is consistent with current models in which BRCA1 promotes subsequent PALB2 and BRCA2 recruitment, ultimately leading to RAD51 loading.

      (3) The model predicts that MDC1-KO cells are proficient for transcriptional repression after nucleolar DSB induction. Has this been tested?

      We did not specifically test this in the current work, but previous results published by our group revealed that siRNA-mediated depletion of MDC1 in human cells had a minimal effect on rDNA transcriptional inhibition after DNA damage (Larsen et al., 2024).

      (4) The nuclear PpoI DSBs could be analyzed as a specificity control, and clarify whether the difference between IRIF and PpoI DSBs relates to the DSB chemistry or location.

      We thank the reviewer for this important point. We agree that I-PpoI induces DNA breaks both within rDNA repeats and at additional genomic loci.

      To address this, we have now analysed the formation of gH2AX-positive nucleolar caps and non-nucleolar gH2AX foci over time following I-PpoI expression (new Figure 1–figure supplement 2). We find that nucleolar caps form rapidly and are prominent at early time points, whereas gH2AX foci accumulate more gradually.

      These results indicate that nucleolar caps and non-nucleolar DNA damage responses can be distinguished both spatially and temporally, and support the use of nucleolar caps as a specific readout for rDNA damage in our study.

      In addition, we note that RAD51 recruitment to IR-induced foci is not affected by MDC1 loss, suggesting that the requirement for MDC1 in RAD51 loading is specific to nucleolar rDNA breaks rather than reflecting differences in DSB chemistry alone. We have clarified this point in the Discussion.

      Additional points:

      (5) Page 4 top: Shieldin.

      Corrected.

      (6) The general reader will be interested to learn about the connection of the Treacle function with Treacher Collins syndrome. Maybe a paragraph could be added to discuss this?

      We thank the reviewer for this suggestion. We agree that the relationship between Treacle and Treacher Collins syndrome may be of interest to a broad readership. Since the developmental pathology of Treacher Collins syndrome is currently thought to arise primarily from impaired ribosome biogenesis and nucleolar dysfunction rather than defective nucleolar DNA damage signalling, we felt that an extensive discussion would be beyond the scope of the present study. We have, however, added a brief statement introducing Treacle as the product of the TCOF1 gene mutated in Treacher Collins syndrome and noting that whether its DNA damage response function contributes to disease pathology remains an open question.

      (7) Figure 7: A short explanation could be added as to why hypoxia conditions were chosen for the p53-deficient cell lines.

      We thank the reviewer for pointing this out. We have added a brief explanation in the figure legend to clarify that hypoxia conditions were used to stabilise replication stress and enhance detection of DNA repair intermediates in p53-deficient cells.

      (8) A short statement on whether the repair of nuclear DBS is affected by Treacle could be added.

      We thank the reviewer for this interesting point. While our study focuses on nucleolar DNA damage, we did not observe evidence that Treacle is required for the repair of non-nucleolar DSBs. We have added a brief statement in the Discussion to clarify that Treacle appears to function specifically in the nucleolar DNA damage response.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely appreciate you and the reviewers for investing time and effort in evaluating our manuscript. After carefully reading the comments and suggestions, we found they are insightful, constructive, and critical for improving the quality of our work. Based on these valuable recommendations, we have substantially revised the manuscript as summarized below.

      Abstract: Inappropriate or ambiguous statements have been revised to improve clarity.

      Introduction: (1) The study purpose and hypotheses have been re-organized in a clearer and more concise way. (2) A mechanistic rationale for sequence-specific degradation has been provided and the use of PMA treatment has been explained. (3) The terminologies related to extracellular DNA and 16S rRNA gene amplicons have been clarified.

      Materials and Methods: 1) More detailed description of the microcosm experiment has been added. 2) The design and rationale of GAPDH F-tagged primers and the use of fusion primers for Illumina library preparation have been clarified. 3) More details about PCR amplification, DNA purification, and pooling strategies have been added. 4) We have corrected and standardized primer naming throughout the manuscript; 5) More details about bioinformatic workflow have been added. 6) We have defined statistical parameters and multiple testing corrections. 7) All abbreviations have been defined and standardized.

      Results: 1) The terminology for PMA-treated DNA has been revised and it has been clarified interpretation as “PMA-treated prokaryotic community” rather than “living community”. 2) The figures and legends have been updated for clarity, and the explicit explanation of “ASV I” and “ASV II” in pairwise comparisons have been added. 3) the figures (e.g., Figs. 2–5, S2–S8) have been reorganized to better reflect results; 4) Inappropriate statements or misleading interpretations have been removed.

      Discussion: A detailed section on technical limitations have been added. The limiatons added mainly include: 1) PCR amplification bias and recommendations for spike-in standards or multi-primer approaches; 2) differential DNA extraction efficiency due to variable cell lysis; and 3) limitations of using 16S rRNA amplicons as proxies for natural extracellular DNA and the limitations of PMA treatment efficiency in soil matrices;

      eLife Assessment

      This valuable study introduces an innovative experimental design to address a crucial and timely issue in microbial ecology: the potential bias in soil microbial community analyses caused by extracellular DNA degradation. While the evidence showing variable degradation rates of extracellular DNA is convincing, additional conceptual, methodological, and statistical clarifications could reinforce the claims and the study's contribution to the field. This research will appeal to microbial ecologists and researchers interested in using molecular techniques to evaluate microbial community structure.

      We sincerely appreciate the editors for the careful assessment of our work and for recognizing the value of addressing extracellular DNA degradation in soil microbial community analyses. We also greatly appreciate the reviewers’ constructive feedbacks concerning the need for additional conceptual, methodological, and statistical clarifications. We agree that further refinement in these areas will strengthen our claims and enhance the study’s contribution to the field. Based on these insightful suggestions, we have carefully revised the manuscript to provide clearer conceptual framework, more detailed methodological descriptions, and more rigorous statistical analyses. We believe these revisions have substantially improved the clarity and robustness of our work. More details about the revisions have been provided in the following responses.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates the degradation dynamics of extracellular DNA in soils and its impact on estimates of microbial abundance and diversity. By combining a broad geographic sampling design with a primer-labeling strategy, qPCR quantification, amplicon sequencing, and PMA treatment, the authors aim to disentangle total versus intracellular DNA signals and explore sequence-specific degradation patterns. The topic is relevant, particularly given the increasing awareness of relic DNA as a confounding factor in microbial ecology. The experimental design is ambitious and potentially impactful. However, several conceptual inconsistencies, methodological ambiguities, and statistical limitations currently weaken the robustness of the conclusions. These issues need to be addressed.

      We sincerely appreciate the reviewer for the constructive assessment of our work. We also appreciate the reviewer’s critical insights regarding the conceptual inconsistencies, methodological ambiguities, and statistical limitations that currently weaken the robustness of the conclusions. We agree with the reviewer that addressing these issues is essential to strengthen our work. Based on these valuable comments, we have carefully revised the manuscript to clarify the conceptual framework. Additionally, we have provided more detailed methodological descriptions, and enhance the statistical rigor of our analyses. We believe these revisions have substantially improved the clarity, consistency, and overall robustness of our conclusions.

      Strengths:

      The manuscript addresses a timely and important question in microbial ecology, particularly given the growing recognition that relic DNA can bias interpretations of community composition derived from amplicon sequencing. The study is ambitious in scope, incorporating a broad geographic sampling design across multiple soil types, which enhances the generalizability of the findings. The use of a controlled microcosm experiment combined with a primer-labeling strategy to track extracellular DNA dynamics is conceptually innovative and provides a structured framework to investigate degradation processes.

      In addition, the integration of multiple approaches, including qPCR for absolute quantification, high-throughput sequencing for community profiling, and PMA treatment to differentiate extracellular from intracellular DNA, represents a comprehensive attempt to disentangle complex sources of bias in soil microbiome analyses. The effort to link degradation dynamics with environmental variables and to explore sequence-level patterns further demonstrates the authors' intent to move beyond descriptive analyses toward a mechanistic understanding.

      We sincerely thank the reviewer for the positive and encouraging comments of our work.

      Weaknesses:

      Several conceptual and methodological issues currently limit confidence in the study's conclusions. Key terms such as "sequence-specific degradation" are not clearly defined or supported by a mechanistic or structural hypothesis, making it difficult to interpret the biological meaning of the results. In addition, the bioinformatic workflow presents inconsistencies, particularly the use of ASVs followed by clustering at 97% similarity, which undermines the resolution required to support sequence-level inferences. Statistical analyses are also insufficiently described, including unclear definitions of "T values," a lack of detail on pairing structure, and no indication of multiple testing correction.

      Furthermore, important methodological details are missing or unclear, including primer design (e.g., GAPDH tag vs ACTF), Illumina library preparation (e.g., adapter and indexing strategy), and validation of PMA treatment efficiency. The interpretation of PMA-treated samples as representing "living communities" is likely overstated, given the known limitations of the method in soil systems. Finally, typographical errors, inconsistent terminology, and unclear phrasing throughout the manuscript reduce readability and further complicate interpretation.

      We sincerely appreciate the reviewer’s thorough and critical evaluation of the manuscript’s weaknesses. The issues raised regarding conceptual clarity, bioinformatic consistency, statistical rigor, methodological transparency, and the interpretation of PMA treatment have been fully acknowledged. We also recognized that typographical errors, inconsistent terminologies, and unclear phrasing largely reduced readability. In response to these valuable comments, the manuscript has been carefully revised as follows. (1) The clearer definition of the term “sequence-specific degradation” has been provided. (2) The bioinformatic workflow was streamlined to ensure consistency. (3) The descriptions of statistical analyses were substantially expanded, including explicit definitions of “t values,” detailed clarification of the pairing structure, and the application of appropriate multiple testing corrections. (4) Missing details regarding primer design, Illumina library preparation, and PMA treatment validation have been added to the Method section. (5) Interpretations of PMA-treated samples have been revised to more accurately reflect methodological limitations in soil systems. (6) The manuscript has been thoroughly proofread to correct typographical errors, standardize terminology, and enhance overall clarity. These revisions are believed to substantially address the concerns raised and significantly strengthen the manuscript. A point-by-point response to the specific comments is provided below.

      Reviewer #2 (Public review):

      Summary:

      This manuscript describes the results of an interesting study examining the rate of degradation of extracellular DNA in soil ecosystems using a clever experimental approach. 16S ribosomal RNA genes were amplified from soil samples, and then purified PCR amplicons, containing a 5' linker sequence on the forward primer, were introduced to soils and monitored over time using real-time quantitative PCR and NGS amplicon sequencing. The study was able to measure rates of overall extracellular DNA degradation, but also sequence-specific degradation rates. I like the idea and execution of the study, and the results are interesting. The manuscript needs some help to improve the overall readability. Please see general and editorial comments below.

      We sincerely thank the reviewer for the positive and encouraging assessment of our study. We have carefully revised the manuscript to enhance clarity, streamline the presentation, and refine the language throughout. We believe these improvements have made the manuscript more readable and easier to follow. We are also grateful for the general and editorial comments provided, which have been addressed as outlined below.

      Strengths:

      Innovative experimental design that is well deployed across a large number of soil types, revealing interesting variability in extracellular DNA degradation.

      We sincerely thank the reviewer for the positive and encouraging assessment of our work.

      Weaknesses:

      (1) The manuscript needs another review to improve the readability of the document.

      We thank the reviewer for this helpful suggestion. We fully agree that improving readability is essential for effectively communicating our findings. Based on the comment, we have carefully revised the manuscript to enhance clarity and readability. We have streamlined sentence structures, standardized terminology, corrected typographical errors, and improved the logical organization of the text. We believe these revisions have substantially improved the overall readability of the manuscript.

      (2) The authors have used 16S genes to look at sequence-specific degradation. But 16S rRNA genes are actually pretty well conserved, and there isn't as much genetic variation across this gene among organisms as there is for other genes. It might be more relevant to look at metagenomic DNA degradation from high AT, high GC organisms, etc. This would be more generalizable than 16S genes.

      We thank the reviewer for this insightful comment. We agree with the reviewer that 16S rRNA genes are relatively conserved compared to functional genes or whole metagenomic DNA, and that studying degradation of more variable sequences (e.g., high‑AT, high‑GC regions, or metagenomic DNA) would provide greater generalizability. However, we would like to clarify the rationale for using 16S rRNA gene amplicons in the present study. First, the 16S rRNA gene remains the most widely used phylogenetic marker in soil microbial ecology (Knight et al., 2018). Demonstrating sequence‑specific degradation with this well‑established marker directly informs a large body of existing research that relies on 16S RNA gene‑based community analyses. Second, despite its conserved nature, the targeted fragment in this study is belong to the highly varied region (V4) of 16S rRNA gene. Accordingly, we indeed observed significant sequence‑specific variation in degradation rates among different 16S rRNA gene amplicon sequence variants (ASVs) (Fig. 2c, 3a). This indicates that even within a conserved marker gene, sequence‑dependent degradation biases exist and can affect diversity estimates. Third, our study was designed as a proof‑of‑concept to establish a methodological framework for quantifying both overall and sequence‑specific degradation rates. Using a single, well‑characterized marker allowed us to develop and validate the primer‑labeling and qPCR/sequencing workflow without the additional complexity of metagenomic DNA (e.g., variable fragment lengths and complex mineral associations). In the revised manuscript, we have added the following sentence to the Discussion section to address the concerns from the reviewer.

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015)”

      (3) Consideration of differential cell lysis during soil DNA extraction needs to be considered as well.

      We thank the reviewer for raising this important technical consideration. We agree that differential cell lysis during soil DNA extraction is a well‑recognized source of bias in microbial community analysis. Different microbial taxa (e.g., Gram‑positive vs. Gram‑negative bacteria, spores, or fungi) vary in their cell wall structure and susceptibility to lysis, which can lead to under‑representation of certain groups and over‑representation of others in the extracted DNA. This bias affects both total DNA extracts and PMA‑treated fractions, potentially influencing our estimates of the relative contributions of intact‑cell derived DNA versus extracellular DNA. However, currently, eliminating these biases are still challenging, and thus we have added the following sentence to the Discussion to address this concern.

      L305-311

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015).”

      (4) It is not clear why the authors didn't put GAPDH linkers on the reverse primer as well. This would have given an easier amplicon to amplify (no degeneracies at all).

      The decision to place the GAPDH linker only on the forward primer (515F) was intentional to balance the need for tracking exogenous extracellular DNA with amplification efficiency, sequencing quality, and cost-effectiveness. Adding linkers to both primers would increase the total amplicon length, potentially reducing amplification efficiency, especially in complex soil samples with degraded or low-quality DNA. More importantly, the reverse primer used in our study is a degenerate primer designed to target the 16S rRNA gene across diverse bacterial taxa, and extending it with an additional GAPDH linker could introduce further complexity, decrease amplification efficiency, and increase primer-dimer formation. Additionally, single-end labeling allows the usage of standard 16S rRNA reverse primers with existing barcodes, whereas dual-end labeling would require synthesis of new barcode-labeled primers, increasing both cost and time. Our preliminary experiments confirmed that single-end labeling produced reproducible amplification curves (~85% efficiency) and high-quality sequencing reads, which were sufficient for quantifying degradation rates. We have added a clarification in the Methods section to explain this rationale.

      L365-371

      “The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Inconsistency between ASV inference and 97% sequence recruitment

      The bioinformatic pipeline presents a major conceptual inconsistency. ASVs are inferred using UNOISE3, but reads are subsequently mapped to ASVs at a 97% similarity threshold, effectively reintroducing OTU-level clustering. Given that the manuscript's central claim is sequence-specific degradation, this step undermines the single-nucleotide resolution that ASVs provide and may obscure biologically meaningful differences. The authors should either reanalyze the data using a consistent ASV framework (exact matching), or explicitly treat the analysis as OTU-like and moderate claims of sequence specificity.

      We appreciate the reviewer's critical evaluation of our bioinformatic pipeline. We also apologized for our unclear statements in our original manuscript. We understand the concern that mapping reads to ASVs at 97% similarity might appear to reintroduce OTU-level clustering. Actually, we used the default pipeline provided by the authors of USEARCH with otutab command to generate the ASV table.

      Following the logic and recommendations of the USEARCH/UNOISE3 developer (Robert Edgar), this approach is a standard procedure for robust noise management rather than a conceptual inconsistency. First, in our pipeline, ASVs (ZOTUs) are first inferred using the UNOISE3 algorithm, which effectively identifies “true” biological sequences at single-nucleotide resolution. Secondly, according to the USEARCH manual, while the ASVs themselves represent exact biological sequences, the raw reads inevitably contain stochastic sequencing errors. Using “exact matching” for recruitment would discard a significant portion of the data that originates from a specific ASV but carries minor random errors. Mapping at 97% identity is the recommended method to recruit these noisy reads back to their correct biological origin (the ASV centroid). Meanwhile, during the recruitment process, reads are not randomly assigned to any ASV within the 97% identity radius. Instead, the algorithm follows a “highest similarity first” principle. For instance, if a specific read exhibits 98% similarity to ASV1 and 97% similarity to ASV2, it is strictly assigned to ASV1. A read is only discarded if its highest similarity to any ASV falls below the 97% threshold. Unlike traditional OTU clustering (where sequences are clustered together based on similarity from the start), our approach maintains the ASV as a fixed biological reference. The quantification of degradation rates is performed on these high-resolution centroids. Thus, our claims regarding sequence-specific degradation remain valid, as the underlying biological variation is defined by the ASVs. To avoid any possible confusion, we have rewritten the relevant paragraph in the Methods section (subsection 4.6) as follows.

      L446-457

      “ASVs were generated using the UNOISE3 non‑clustering denoising algorithm (Edgar, 2016), which infers 100% exact sequence variants by distinguishing biological sequences from PCR/sequencing errors. ASVs with total sequence counts fewer than 9 across all samples were removed to reduce noise. To quantify the abundance of each ASV, an ASV table was generated by mapping the quality‑filtered raw reads back to the ASV set using the otutab command. A 97% similarity threshold was applied for this recruitment to accommodate stochastic sequencing noise while maintaining biological resolution. Crucially, the mapping followed a best-hit priority rule, where each read was assigned to the ASV with the highest per cent identity within the 97% radius. This approach ensures that reads derived from the same biological template are accurately counted toward their respective ASV, preventing the underestimation of abundances that would occur with exact matching while strictly preserving the single-nucleotide resolution of the ASV framework.”

      (2) Undefined "ASV I" and "ASV II" groups

      The manuscript refers to "ASV I" and "ASV II" groups in pairwise comparisons of degradation rates (e.g., Fig. 3), but these groups are not defined anywhere in the text.

      It is unclear whether these represent: predefined biological categories, arbitrary pairwise ASV comparisons, or groupings based on taxonomy, abundance, or degradation rate.

      In addition, the pairing structure underlying these comparisons is not described. While a paired t-test is mentioned, it is unclear how ASVs were paired (e.g., within sites, across samples, or across time points).

      The current terminology ("groups") is potentially misleading and suggests biological structure where none may exist. The authors should explicitly define these terms, clarify the pairing scheme, and revise terminology if these are simply pairwise comparisons.

      We thank the reviewer for this keen observation. We completely agree that the terms "ASV I" and "ASV II" were poorly defined and potentially misleading.

      We would like to clarify that "ASV I" and "ASV II" were not intended to represent predefined biological categories (such as groupings based on taxonomy, abundance, or degradation rates). Instead, they were merely used as a labeling convention to indicate the directionality of pairwise comparisons within the heatmap matrix. Specifically, "ASV I" referred to the ASVs represented in the rows, while "ASV II" referred to those in the columns. In the original Fig. 3, blue indicated that the degradation rate of the row ASV was significantly lower than that of the column ASV, and red indicated the opposite. To avoid any confusion, we have removed the “ASV I/II” terminology throughout the manuscript and figures, replacing them with “Row ASVs” and “Column ASVs”. To address this issue, we have revised the Figure 3 legend to include a more explicit explanation:

      L820-824

      “In the heatmap, each cell represents a pairwise comparison between two ASVs. Blue indicates that the degradation rate of the ASVs listed in the row (row ASVs) is significantly lower than that of the ASVs listed in the column (column ASVs); red indicates that the row ASVs has a significantly higher degradation rate than the column ASV. A positive t value indicates that the row ASVs degrades significantly faster than the column ASVs; a negative t value indicates the opposite.”

      We thank the reviewer for raising the important issue regarding the definition of “t values” in our statistical analysis. We apologize for the lack of clarity in the original manuscript. To clarify, the T values presented in Figure 3a represent the test statistics (t-values) from paired t-tests comparing the degradation rate constants of two ASVs across the 30 study sites. The T-value was obtained from a paired t-test between two ASVs across the same samples. The t-value indicates the magnitude and direction of the difference between the two ASVs’ degradation rates relative to the variability across sites. A positive t-value (colored red in the heatmap) indicates the row ASVs degrades significantly faster than the column ASVs; a negative t-value (colored blue) indicates the opposite.

      L544-548

      “As for the analysis, we performed paired t‑tests across all the study sites. Thus, the degradation rates were essentially compared within each site, with both values originating from a same soil sample under identical incubation conditions. A positive t value indicates that the first ASV has a significantly higher degradation rate than the second one, and a negative t value indicates the opposite. The p values were adjusted for multiple comparisons using the FDR method.”

      (3) Lack of definition and justification of "T values"

      The manuscript reports "T values" for comparisons between ASVs but does not clearly define how these values are calculated. Although a paired t-test is mentioned, it remains unclear how the pairing was constructed, whether assumptions (normality, independence) were evaluated, and whether corrections for multiple comparisons were applied. Given the large number of ASVs, failure to control for multiple testing could inflate false positives. More broadly, the use of a simple paired t-test may not be appropriate given the hierarchical and compositional structure of the data.

      We sincerely thank the reviewer for pointing out the need to clarify the definition and justification of the t-values presented in our manuscript. Each t-value represents the test statistic from a paired t-test comparing the degradation rates of two ASVs across the same set of samples. The paired t-test assumes that the differences between paired observations are approximately normally distributed and that the pairs are independent across columns. We have evaluated the normality of differences using standard diagnostic plots and verified that the assumption is reasonably satisfied given the sample size. We performed a correction for multiple comparisons using the False Discovery Rate (FDR) procedure to control for potential false positives. We have revised the Methods section to clearly define t-values.

      (4) Conceptual validity of "sequence-specific degradation"

      The manuscript repeatedly refers to "sequence-specific degradation" of extracellular DNA; however, this concept is not clearly defined nor supported by a biological or structural hypothesis. It is unclear what "sequence-specific" refers to (e.g., nucleotide composition, GC content, secondary structure, taxonomic identity), whether differences are expected in conserved versus variable regions of the 16S rRNA gene, or what mechanistic basis would explain differential degradation among sequences. Given that the analysis is based on short 16S V4 amplicons, and no structural or biochemical framework is provided, it is difficult to interpret whether the observed differences truly reflect intrinsic sequence-dependent degradation or are instead driven by methodological or statistical artifacts (e.g., abundance effects, amplification bias).

      I believe the authors should explicitly define what is meant by "sequence-specific degradation," provide a biologically grounded hypothesis (e.g., structural accessibility, GC content, stem-loop stability), and align their interpretation with the resolution and limitations of the data.

      We thank the reviewer for this critical conceptual comment. We apologize that “sequence‑specific degradation” was not clearly defined and lacked a biological or structural hypothesis. To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation. This adjustment ensures that the hypotheses are directly grounded in the theoretical framework (e.g., GC content, thermodynamic stability, and secondary structures) presented in the paragraph.

      We now define “sequence‑specific degradation” as statistically significant differences in first‑order degradation rate constants among distinct ASVs, mainly arising from intrinsic DNA properties (base composition, secondary structure, and restriction sites) or differential mineral adsorption.

      L99-104

      “Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      We also expanded the mechanistic discussion to include GC content and secondary structure.

      L230-235

      “We also examined whether GC content could explain the observed sequence‑specific patterns, but no significant correlation was found (Fig. S4), suggesting that simple base composition is not the primary driver in this study. However, this does not exclude the possibility that higher‑order structural features (e.g., hairpin loops) or sequence‑specific nuclease recognition motifs contribute to differential degradation (Wang et al., 2007). This should be tested in future studies using synthetic DNA constructs with controlled structural elements.”

      We acknowledge that inferring sequence‑specific degradation from combined relative abundance and qPCR data is subject to potential methodological artifacts, including compositional effects, PCR amplification bias, and abundance‑dependent detection limits. However, we have taken several stringent steps to minimize these concerns. Specifically, we restricted our analysis to ASVs that were present in more than 90% of the study sites and for which the degradation curve fits yielded R<sup>2</sup> > 0.5, ensuring that only robustly detected and reliably modeled sequences were retained. Because our analysis tracks the ratio of each ASV across a time series, any sequence-specific PCR amplification bias remains constant for that particular sequence. By focusing on the rate of change rather than absolute read counts, such systematic biases are mathematically canceled out during the calculation of degradation kinetics.

      (5) Conceptual ambiguity in "GAPDH F-labeled 16S rRNA genes"

      The manuscript repeatedly refers to "GAPDH F-labeled 16S rRNA genes," which is confusing and may be misinterpreted as targeting GAPDH rather than 16S. It should be clearly stated that GAPDH refers to glyceraldehyde-3-phosphate dehydrogenase, and a GAPDH-derived sequence is used as a synthetic tag appended to a 16S primer. Additionally, the divergence of this tag from microbial sequences should be justified to ensure specificity. There is also an inconsistency in primer naming (e.g., "GAPDH F" vs "ACTF" in the figures), which should be corrected.

      We sincerely thank the reviewer for this important comment. We agree that the phrase “GAPDH F‑labeled 16S rRNA genes” could be confusing, as it may be misinterpreted as targeting the GAPDH gene rather than the 16S rRNA gene. We have revised the manuscript to avoid this ambiguity and to provide clear justification for the use of the GAPDH tag. GAPDH (glyceraldehyde‑3‑phosphate dehydrogenase) is a human housekeeping gene. Its forward primer sequence (GAPDH F: 5′‑CAT TGG CAA TGA GCG GTT C‑3′) was used as a synthetic tag appended to the 16S primer because (i) no homologous sequences exist in soil DNA (confirmed by PCR), and (ii) its melting temperature is compatible with the reverse primer. This tag allows specific tracking of exogenous DNA without interference from native soil sequences.

      Throughout the manuscript, ambiguous phrases such as “GAPDH F‑labeled 16S rRNA genes” have been replaced with more precise terms, “GAPDH F‑tagged 16S rRNA gene amplicon fragments” clarifying that the tag is an appendage and not the amplification target.

      We have checked the entire manuscript and confirm that “ACTF” does not appear anywhere. To avoid confusion, the primer is now consistently referred to as “GAPDH F” in all figures, legends, and text.

      L360-371

      “GAPDH is a primer for a human housekeeping gene and it has no homologous sequences in soils. Subsequently, GAPDH was selected as the label primer based on two criteria. First, this primer was selected to avoid interference from the original soil sequences (Huang et al., 2014; Yang et al., 2021; Arvizu-Hernandez et al., 2025), and no detectable PCR amplification was observed for the primer set GAPDH F-806R across all the soil DNA samples included in this study. Second, the melting temperature (Tm) value of GAPDH F approximately matched that of 806R. The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      (6) Limitations of using PCR amplicons as proxies for extracellular DNA

      The study uses PCR-generated amplicons to simulate extracellular DNA. While useful for controlled comparisons, these fragments may not reflect the physicochemical diversity of natural extracellular DNA (e.g., adsorption to minerals, fragment size variability, protection within aggregates). This limitation should be explicitly acknowledged, and conclusions should be framed accordingly.

      We appreciate the reviewer’s constructive feedback. We fully acknowledge that using PCR-generated amplicons to simulate extracellular DNA (eDNA) has inherent limitations in capturing the full physicochemical diversity of naturally occurring eDNA in soils. Specifically, we agree that PCR fragments may not replicate features such as highly variable fragment size distributions, associations with complex cellular components (e.g., vesicles or protein complexes), or long-term physical sequestration within soil micro-aggregates. Despite of these limitations, the use of uniform primer-tagged PCR amplicons was a deliberate choice to enable precise tracking of exogenous DNA degradation kinetics while eliminating background interference from endogenous soil eDNA. This design is a prerequisite for the high-resolution kinetic modeling of sequence-specific decay. Furthermore, in our bioinformatic pipeline, the 97% mapping threshold was specifically applied to minimize the influence of stochastic sequencing and PCR errors on abundance quantification. In the revised manuscript, these potential limitations have been addressed.

      L295-299

      “First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020).”

      (7) Interpretation of sequence-specific degradation

      Sequence-specific degradation rates are inferred from combining relative abundance data with total qPCR estimates. This approach is sensitive to compositional effects, amplification biases, and abundance-dependent detection limits. It remains unclear whether observed differences reflect true sequence-specific degradation or methodological artifacts. This limitation should be discussed more explicitly.

      We thank the reviewer for highlighting this critical methodological point. In our study, sequence-specific degradation rates were estimated by combining ASV-relative abundances with total qPCR-derived 16S rRNA gene copy numbers. We acknowledge that this approach may be influenced by compositional effects, PCR amplification biases, and abundance-dependent detection limits. However, the degradation rate constant (k) in our study, represents the rate of change for a specific sequence over time. Since PCR amplification biases are generally sequence-specific and consistent across samples processed under identical conditions, these systematic errors are mathematically canceled out when calculating the relative change (slope) for the same ASV across a time series. Second, all qPCR measurements were performed with three technical triplicates with standard curves to ensure quantitative reliability. Third, relative abundances were converted to absolute abundances using total qPCR estimates, allowing cross-taxa comparisons that reduce compositional bias. This approach is widely recognized in microbial ecology as a robust method. To address this concern, we have added some explanations in the revised manuscript.

      L84-86

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequence variants (ASVs) under identical soil and incubation conditions.”

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      L311-314

      “While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017)..”

      (8) Overinterpretation of PMA-treated samples as "living communities"

      The manuscript interprets PMA-treated DNA as representing intracellular or "living" microbial communities. While PMA is useful, this interpretation should be treated with caution in soils. PMA efficiency can be affected by soil matrix complexity, DNA adsorption to particles, incomplete light penetration, and permeability of compromised cells. Importantly, no validation of PMA efficiency is presented.

      We thank the reviewer for this important caution. We agree that interpreting PMA‑treated DNA as representing “living” or “intracellular” communities is an overstatement in soil systems. In the revised manuscript, we no longer describe PMA-treated DNA as a direct proxy for the “living community,” but instead refer to it as the “PMA-treated prokaryotic community”.

      Although we did not directly validate PMA efficiency in this study, we used a standardized PMA protocol that has been widely applied in microbial ecology, and our goal was to obtain a comparative estimate of the influence of extracellular DNA on community analysis across soils under a consistent methodological framework. Based on previous studies (Carini et al., 2016; Du et al., 2025), which found that in similar soil types, PMA treatment can significantly reduce the interference of extracellular DNA and alter the community structure, this indirectly proves the effectiveness of this technique.

      Nevertheless, we agree that future studies should include explicit validation controls, such as live/dead cell mixtures, heat-killed controls, or soil-specific PMA efficiency tests, to better quantify method performance across diverse soil matrices. We have added a dedicated paragraph in the "Methodological Considerations and Limitations" section to discuss how soil-specific properties (e.g., turbidity, adsorption capacity) might lead to incomplete exclusion of extracellular DNA, thereby advising a more cautious interpretation of the "viable" community data.

      L496-500

      “To inhibit amplification of eDNA, soils were incubated with propidium monoazide (PMA), as described previously (Carini et al., 2016). Upon photoactivation, eDNA can form covalent bonds through cross-linking, leading to the inhibition of its PCR amplification. In contrast, microbes with intact cell membranes exclude PMA, and their DNA is not cross-linked with PMA, and remains amenable to PCR amplification.”

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      Minor Comments

      (1) Line 79: Provide examples of how extracellular DNA contributes to nutrient cycling (e.g., P, N sources) and signal transduction (e.g., horizontal gene transfer).

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have added specific examples to clarify how extracellular DNA contributes to nutrient cycling and signal transduction. Specifically, we now note that extracellular DNA can serve as a source of phosphorus and nitrogen following enzymatic degradation, thereby contributing to soil nutrient turnover. We also clarify that extracellular DNA plays an important role in horizontal gene transfer, acting as a genetic reservoir that can be taken up by competent microorganisms and thereby facilitating the spread of functional traits such as antibiotic resistance. These examples have been added to improve the clarity and biological context of this statement.

      L48-52

      “EDNA serves as a critical vector for horizontal gene transfer (HGT), facilitating the uptake of genetic material by competent microorganisms and promoting the spread of functional traits such as antibiotic resistance (Liu et al., 2024). In addition, eDNA participates in soil biogeochemical cycling because its enzymatic degradation releases bioavailable nutrients, particularly phosphorus and nitrogen, which can be reused by soil microorganisms (Ye et al., 2022).”

      (2) Line 79: Replace "for an extended period of time" with a more precise or referenced timescale.

      We agree with the reviewer. We have replaced the vague phrase with a precise timescale. Extracellular DNA can persist in soils for months to years.

      (3) Line 94: Clarify what is meant by "high-level structure" (e.g., secondary structure, environmental association).

      We thank the reviewer for pointing out this ambiguity. In the original manuscript, the phrase “high-level structure” was not sufficiently precise. In the revised version, we have clarified that this refers primarily to higher-order structural properties of DNA molecules, such as secondary structure, local conformational features, and sequence-dependent interactions with minerals or organic matter in soil. These characteristics may influence the accessibility of extracellular DNA to nucleases and thus affect degradation rates. We have revised the text accordingly to improve clarity and precision.

      L84-104

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequences (ASVs) under identical soil and incubation conditions. The potential variations in sequence-specific eDNA degradation rates can be attributed to several factors. First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C: N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021). Third, the persistence of soil DNA is often associated with its adsorption and protection by minerals and humus in soils (Cai et al., 2006b; Vuillemin et al., 2017; McKinney and Dungan, 2020). Thus, sequence-dependent differences in the physicochemical behavior of DNA molecules, including their affinity for soil minerals and organic matter, may also contribute to variation in degradation rates among sequences (Levy-Booth et al., 2007; Morrissey et al., 2015). Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (4) Line 111: The hypothesis is not clearly linked to the rationale. If sequence-specific degradation is expected, clarify whether it relates to conserved vs variable regions or structural features (e.g., stems vs loops).

      We thank the reviewer for this helpful comment. We agree that the original manuscript did not clearly link the hypothesis regarding sequence-specific degradation to its mechanistic rationale. In the revised manuscript, we have clarified that the expectation of sequence-specific degradation is not simply based on conserved vs variable regions of the 16S rRNA gene, but rather on the potential for sequence differences to influence intrinsic physicochemical properties, including base composition, local conformational features, potential secondary structures, and motif-dependent nuclease susceptibility. These factors may alter DNA accessibility to extracellular nucleases, providing a mechanistic basis for sequence-specific degradation. This clarification is now reflected in the Introduction and linked to the formal hypothesis statement.

      To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses (H1–H3) immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation.

      L84-104

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequences (ASVs) under identical soil and incubation conditions. The potential variations in sequence-specific eDNA degradation rates can be attributed to several factors. First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C:N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021). Third, the persistence of soil DNA is often associated with its adsorption and protection by minerals and humus in soils (Cai et al., 2006b; Vuillemin et al., 2017; McKinney and Dungan, 2020). Thus, sequence-dependent differences in the physicochemical behavior of DNA molecules, including their affinity for soil minerals and organic matter, may also contribute to variation in degradation rates among sequences (Levy-Booth et al., 2007; Morrissey et al., 2015). Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (5) Lines 310-311: Clearly indicate which portion of the primers corresponds to the modified (GAPDH-derived) sequence. Provide full annotated primer sequences.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we now clearly indicate which portion of the forward primer corresponds to the GAPDH-derived synthetic tag and which portion corresponds to the 16S rRNA gene primer sequence. We have also provided the full annotated primer sequences in the Methods section to avoid ambiguity.

      Specifically, the modified forward primer is now described as:

      GAPDH-F-515F: 5′-CAT TGG CAA TGA GCG GTT C-GTG CCA GCM GCC GCG GTA A-3′,

      where CAT TGG CAA TGA GCG GTT C is the GAPDH-derived synthetic tag and GTG CCA GCM GCC GCG GTA A is the 16S rRNA gene forward primer sequence (515F).

      The reverse primer is:

      806R: 5′-GGA CTA CHV GGG TWT CTA AT-3′.

      L354-359

      “Briefly, exogenous eDNA was prepared by PCR amplification using a modified forward primer consisting of a GAPDH F tag fused to the 16S rRNA gene primer 515F, together with the reverse primer 806R. The full primer sequences were as follows: GAPDH-F-515F: 5'-CAT TGG CAA TGA GCG GTT C-GTG CCA GCM GCC GCG GTA A-3', in which CAT TGG CAA TGA GCG GTT C represents the GAPDH F tag and GTG CCA GCM GCC GCG GTA A represents the 16S rRNA gene forward primer sequence (515F); and 806R: 5'-GGA CTA CHV GGG TWT CTA AT-3'.”

      (6) Lines 310-311: Explicitly define GAPDH and justify its use as a synthetic tag.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we now explicitly define GAPDH as glyceraldehyde-3-phosphate dehydrogenase, a human housekeeping gene. Specifically, the GAPDH-derived sequence was selected for two reasons. First, it is highly divergent from known soil microbial 16S rRNA gene sequences and did not produce detectable amplification when tested with soil DNA using the GAPDH tagged 806R primer pair, indicating that it would not interfere with endogenous soil DNA signals. Second, its melting temperature was compatible with that of the reverse primer, which allowed stable amplification of the tagged 16S amplicons under our PCR conditions.

      L360-371

      “GAPDH is a primer for a human housekeeping gene and it has no homologous sequences in soils. Subsequently, GAPDH was selected as the label primer based on two criteria. First, this primer was selected to avoid interference from the original soil sequences (Huang et al., 2014; Yang et al., 2021; Arvizu-Hernandez et al., 2025), and no detectable PCR amplification was observed for the primer set GAPDH F-806R across all the soil DNA samples included in this study. Second, the melting temperature (Tm) value of GAPDH F approximately matched that of 806R. The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing”

      (7) Line 346: Start a new paragraph to clearly separate this as a distinct experiment.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have started a new paragraph.

      (8) Line 346: Specify the number of samples analyzed for consistency.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have now explicitly specified the number of samples.

      A total of 120 samples were analyzed in this moisture gradient experiment: 2 ecosystems (Kaiyuan and Dashanbao) × 5 moisture levels (10%, 25%, 50%, 75%, and 100% of water holding capacity) × 6 incubation time points (0, 1, 3, 6, 12, and 24 days) × 2 replicates. Just two technical replicates were performed for this validation experiment, as the primary aim was to assess the trend of moisture effects rather than statistical inference across replicates.

      L408-410

      “This complementary experiment included two sites, five moisture levels, six incubation time points, and two replicates per treatment combination, resulting in a total of 120 soil samples.”

      (9) Lines 351-352: Replace "harvested" with "collected."

      We have replaced “harvested” with “collected” as suggested.

      (10) Line 367: Clarify how Illumina adapters and indices were added (e.g., two-step PCR, fusion primers).

      We thank the reviewer for this helpful suggestion. We have revised the Methods section to clarify how Illumina adapters and indices were incorporated. We used a pooled amplicon library preparation strategy. Individual samples were first amplified with primers containing sample-specific barcode sequences. The barcoded amplicons from multiple samples were then pooled and used for library preparation with the ALFA-SEQ DNA Library Prep Kit. Universal Illumina-compatible adapters were first ligated to the pooled amplicons. After bead-based purification, an indexing PCR was performed using the index primer mix, which introduced the complete P5/P7 sequences and a library-level Illumina index into the library molecules. Thus, sample demultiplexing was based on the sample-specific barcodes introduced during amplicon PCR, whereas the Illumina index was used to identify the pooled sequencing library. We have clarified this procedure in the revised manuscript.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).

      (11) Provide more detail on chimera removal, filtering thresholds, and normalization choices.

      We thank the reviewer for this helpful suggestion. The raw paired-end reads were first merged, and primer sequences were removed using the search_pcr2 script in USEARCH. Reads with more than two primer mismatches were discarded. Quality filtering was then performed using fastq_filter, and sequences with quality scores below 20 were removed. Redundant reads were collapsed using fastx_uniques. Amplicon sequence variants (ASVs) were generated using the UNOISE3 denoising algorithm, which also performs built-in chimaera filtering during ASV inference. In addition, ASVs with total sequence counts fewer than 9 were excluded to reduce the influence of low-frequency noise.

      L443-460

      “Briefly, paired-end reads were merged using USEARCH, and primer sequences (GAPDH-F-515F and 806R) were removed using the search_pcr2 script. Reads with more than two primer mismatches were discarded. Quality filtering was performed using the fastq_filter script, and sequences with quality scores below 20 were removed. Redundant sequences were dereplicated using the fastx_uniques script. ASVs were generated using the UNOISE3 non‑clustering denoising algorithm (Edgar, 2016), which infers 100% exact sequence variants by distinguishing biological sequences from PCR/sequencing errors. ASVs with total sequence counts fewer than 9 across all samples were removed to reduce noise. To quantify the abundance of each ASV, an ASV table was generated by mapping the quality‑filtered raw reads back to the ASV set using the otutab command. A 97% similarity threshold was applied for this recruitment to accommodate stochastic sequencing noise while maintaining biological resolution. Crucially, the mapping followed a best-hit priority rule, where each read was assigned to the ASV with the highest per cent identity within the 97% radius. This approach ensures that reads derived from the same biological template are accurately counted toward their respective ASV, preventing the underestimation of abundances that would occur with exact matching while strictly preserving the single-nucleotide resolution of the ASV framework. Taxonomic annotation of the ASVs was performed in QIIME2 with the Silva v138 database. A total of 89322 prokaryotic ASVs were obtained. To standardize sequencing depth across samples, the read number of each sample was rarefied to 53251 using the rarefy function in the vegan package in R.”

      (12) Line 412: Rephrase to refer to 16S amplicon addition rather than 16S rRNA genes (along the whole text), as only the V4 region is analyzed.

      We thank the reviewer for this helpful suggestion. we have rephrased references to “16S rRNA genes” to “16S rRNA gene amplicon fragments”

      (13) Ensure consistent primer naming throughout (e.g., GAPDH F vs ACTF).

      We have checked the entire manuscript and confirm that only “GAPDH F” is used as the label primer.

      (14) Finally, the manuscript would benefit from careful language editing. Several typographical errors, grammatical inconsistencies, and unclear phrases are present throughout. Examples include:

      Misspellings such as "diffrence" (e.g., figure legends) and inconsistent capitalization. Inconsistent terminology (e.g., "genes," "amplicons," and "fragments" used interchangeably without clarification). Redundant or awkward phrasing (e.g., repeated use of "extracellular 16S rRNA genes"). Occasional subject-verb agreement issues and missing articles.

      We sincerely apologize for the language issues. The manuscript has now undergone a thorough language editing process by a native English‑speaking colleague.

      Recommendation

      Major revision: The manuscript addresses an important problem and presents a promising approach. However, key issues related to conceptual clarity, bioinformatic consistency, statistical rigor, and interpretation of PMA-based results must be resolved. With substantial revision and clarification, the study has the potential to make a meaningful contribution to the field.

      We sincerely thank the reviewer for the thorough, constructive, and critical evaluation of our manuscript. We greatly appreciate the recognition that our study addresses an important problem and presents a promising approach. We also acknowledge the key issues raised regarding conceptual clarity, bioinformatic consistency, statistical rigor, and interpretation of PMA‑based results. We have taken these comments very seriously and have substantially revised the manuscript accordingly, more details about the revisions are described in the following point-by-point responses.

      Reviewer #2 (Recommendations for the authors):

      Editorial comments:

      (1) Title: I recommend removing "across China" from the title. In many ways, the study has nothing to do specifically with China, and you limit the broad applicability of the study. The same work could have been done with soils from Africa, for example. Also, it might be ok to remove 16S rRNA as well. The 16S rRNA genes are a proxy for rates of extracellular DNA degradation, but the study isn't exactly about 16S either.

      We thank the reviewer for this thoughtful suggestion regarding the title. We have revised the title to “The overall and sequence-specific degradation of soil extracellular DNA fragments: rates and influential factors.”

      (3) L44-45: "...such as real-time PCR, high-throughput amplicon sequencing, and metagenomic analysis...".

      We thank the reviewer for this suggestion. We have revised the order according to the suggestions of the reviewer.

      L44-45

      “The investigation of soil microbial abundance and diversity heavily relies on DNA-based technologies, such as real-time PCR, high-throughput amplicon sequencing, and metagenomic analysis.”

      (4) L48: remove "they".

      We agree with the reviewer and have removed the extraneous “they”.

      (5) L51: "noise factor"; "...persistence can lead to...".

      We have revised the sentence as suggested.

      (6) L53: remove theoretical.

      We have removed “theoretical”.

      (7) L58: remove "the".

      We have removed "the".

      (8) L86: Is restriction digestion of DNA a likely extracellular process in soil?

      We thank the reviewer for this thoughtful comment. We agree that the original wording may have overstated the likelihood of classical restriction digestion as a dominant extracellular process in soils. Our intention was not to suggest that intracellular restriction enzyme systems operate directly in the soil matrix in the same manner as they do within living cells. Rather, we aimed to indicate more generally that sequence-dependent nuclease susceptibility could contribute to differential degradation among extracellular DNA fragments.

      L87-95

      “First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C: N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021).”

      (9) L94-99: The authors might also consider the different nitrogen content of different bases; this might also affect sequence-specific selection of DNA for degradation.

      We thank the reviewer for this insightful suggestion. We agree that differences in the elemental composition of DNA bases, including nitrogen content, may provide an additional mechanistic explanation for sequence-dependent degradation. In the revised manuscript, we have incorporated this point into the Introduction.

      L92-95

      “Differences in base composition also alter the elemental stoichiometry (e.g., C:N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021).”

      (10) L110-112: These are not really written in hypothesis form. Also, what about a hypothesis about degradation rates and soil type/temperature/moisture?

      We thank the reviewer for this constructive critique. We have rewritten the hypotheses. To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation. This adjustment ensures that the hypotheses are directly grounded in the theoretical framework.

      L99-104

      “Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (11) L116: "GAPDH F-labeled 16S rRNA gene amplicon fragments....".

      We thank the reviewer for this helpful suggestion. we have rephrased references to “16S rRNA genes” to “16S rRNA gene amplicon fragments”

      (12) L117: "rapidly".

      We agree with the reviewer and have revised.

      (13) L118-120: "After a 48-day incubation period, 0.2 to 3.1% of the initial spike GADPH F-labeled 16S rRNA gene amplicon fragments ...".

      We agree with the reviewer and have revised.

      (14) L125: Spell out SEM in first usage.

      We thank the reviewer for this suggestion. In the revised manuscript, we have spelled out SEM as Structural equation modeling.

      (15) L128: I don't like the idea of putting this Figure in supplemental materials.

      We thank the reviewer for this suggestion. We have moved Figure S2 (moisture gradient microcosm experiment) to the main text as Figure 1f.

      (16) L154: The term "intracellular prokaryotic abundance" is not the right term. This makes one think of an intracellular parasite. I think you want something like: "Approximately 40% of sequences in total soil DNA extraction NGS amplicon libraries were derived from intact cells, while the remaining represented extracellular DNA. Conversely, greater than 80% of observed richness was derived from intact cells." (Please check that I stated this correctly.) I would also suggest some statistics or ranges here.

      We thank the reviewer for this important terminological clarification. We agree that the term “intracellular prokaryotic abundance” is misleading, as it could imply intracellular parasites. In the revised manuscript, we have replaced this with a clearer description and We have also added the across‑site ranges to provide statistical context.

      L163-166

      “The PMA treatment revealed that intact cells accounted for approximately 40% (range: 9–73%) of the total 16S rRNA gene copies. In contrast, over 80% (range: 27–97%) of the observed ASV richness was associated with sequences originating from intact cells (Fig. 4a and b).”

      (17) L168: "...a significant NEGATIVE correlation was observed...".

      We agree with the reviewer and have revised.

      (18) L169: "However, no significant relationship was observed...".

      We agree with the reviewer and have revised.

      (19) L194-195: What about pH and temperature?

      We thank the reviewer for this comment. We agree that pH and temperature are important environmental factors that can influence microbial DNA degradation and community composition. However, our results (Fig. 1c) indicate that soil moisture is the most dominant factor affecting extracellular DNA degradation. Therefore, in the revised manuscript, we have focused the explanation primarily on soil moisture, while acknowledging that pH and temperature may also be important influencing factors.

      L208-211

      “Third, environmental factors, including soil moisture, pH, and temperature, can predominantly govern enzymatic reaction rates (He et al., 2024; Shah et al., 2024). Indeed, strong positive correlations were observed between moisture content and eDNA degradation rates in both the survey and microcosm experiments (Fig. 1d-f).”

      (20) L199: "findings".

      We have revised as suggested.

      (21) L227-229: This sounds more like results.

      We thank the reviewer for this comment. We agree that the original first sentence in L227–229 reads more like results. Our intention was to introduce the discussion by linking extracellular DNA to potential impacts on prokaryotic community analysis, rather than to present specific findings at this point. We have reorganized this section as follows.

      L246-248

      “Accordingly, we further explored how DNA may influence prokaryotic community analyses using PMA treatment, and significant disparities were observed between the profiles of the total and PMA-treated soil prokaryotic communities (Fig. 4).”

      (22) L230: Need to also consider differential cell lysis during DNA extraction.

      We thank the reviewer for this important comment. We agree that differential cell lysis during DNA extraction could influence the observed community profiles, as microbial taxa differ in cell wall composition and resistance to mechanical or chemical lysis. In the revised manuscript, we explicitly acknowledge this limitation in the relevant section. We also clarify that a standardized DNA extraction protocol (DNeasy PowerSoil kit) was used to efficiently lyse a broad range of microbial taxa, but some taxon-specific lysis bias may remain. Future studies could combine multiple lysis methods or spike-in controls to quantify and correct for potential extraction bias.

      L262-265

      “However, as DNA extraction efficiency may differ between intact cells and eDNA, the actual differences between total and living prokaryotic abundance could be smaller than those observed in this study. Similarly, the overestimated prokaryotic richness may arise from historically accumulated microbial taxonomic information stored in eDNA pools (Deshpande and Fahrenfeld, 2023; Wang et al., 2024).”

      L309-311

      “Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015).”

      (23) L232: Need to also consider that PMA treatment is not perfect and can be affected by substrate, the ability of light to access DNA for crosslinking, etc.

      We thank the reviewer for this important reminder. We agree that PMA treatment is not perfect and that its efficiency can be affected by soil matrix properties (e.g., organic matter, clay minerals) and the ability of light to penetrate the sample for DNA crosslinking. In the revised manuscript, we have explicitly acknowledged these limitations in the discussion.

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (24) L240: Can extracellular DNA have an ecological role?

      We thank the reviewer for this thoughtful question. Yes, extracellular DNA (eDNA) does have important ecological roles beyond being a potential bias in molecular analyses. In the revised manuscript, we have added statements to highlight that extracellular DNA can serve as a nutrient source (e.g., nitrogen and phosphorus) for microbes and may also contribute to horizontal gene transfer. This emphasizes that extracellular DNA may actively influence microbial community structure and function, in addition to its role in potentially inflating observed abundance and richness.

      L267-275

      “We observed a significant correlation between eDNA degradation rates and the overall structure of the prokaryotic community, but this relationship was absent in PMA-treated communities (Fig. 5b). This discrepancy highlights the divergent ecological roles of extracellular and intracellular DNA. Analyses of the total community integrate intracellular DNA from metabolically active cells with eDNA which primarily originates from historical microbial residues (Lennon et al., 2018). EDNA incorporates signals that likely reflect the legacy effects of past environmental conditions (Wang et al., 2021). In contrast, the PMA-treated community reflects transient microbial activity driven by current selective pressures. Additionally, eDNA can serve as a nutrient source and facilitate horizontal gene transfer, which may further shape its interactions with contemporary microbial communities (Levy-Booth et al., 2007).”

      (25) L256: Why would microorganisms selectively degrade one DNA sequence vs another? This seems to be likely to be stochastic in terms of which sequences are taken up by microorganisms. However, different DNA sequences might hydrolyze differently or be otherwise damaged, and that could lead to differential degradation of a viable amplicon. It might be interesting to incorporate long pieces of DNA with different internal primer sites and use quantitative PCR to determine how sequences are degrading.

      We thank the reviewer for this important mechanistic insight. We agree that the observed correlation between degradation rate and sequence abundance does not necessarily imply active microbial preference. It could equally reflect stochastic encounter rates or intrinsic chemical differences (e.g., AT‑rich regions hydrolyzing faster). We have revised the corresponding paragraph in the Discussion.

      L279-292

      “This finding suggests that abundant eDNA degrades at a faster rate compared to rare eDNA. As mentioned earlier, this could be explained by several mechanisms. First, as soil eDNA is subject to enzymatic degradation and microbial recycling, abundant DNA sequences may be more likely to be encountered and degraded by extracellular nucleases simply due to their higher copy numbers (Levy-Booth et al., 2007; Nagler et al., 2018). Similarly, if microbes preferentially take up DNA as a nutrient source, they may degrade abundant sequences more frequently as a stochastic consequence of higher encounter rates (Finkel and Kolter, 2001). However, we also found that the relationships between the sequence-specific degradation rates and the effect sizes of extracellular 16S rRNA gene amplicon fragments varied across the study sites (Fig. S1g). The sequence-specific effect sizes of extracellular 16S rRNA gene amplicon fragments are mainly determined by both their production and degradation rates (Pietramellara et al., 2009; Sirois and Buckley, 2019). These inconsistent correlations emphasize the critical role played by the production rates of extracellular 16S rRNA genes in influencing the analysis of prokaryotic communities. Therefore, future studies should systematically determine both the production and degradation rates of eDNA.”

      (26) L282-283: This belongs in the discussion.

      We agree with the reviewer and have revised accordingly.

      (27) L289: "as well as measurements of total organic carbon".

      We agree with the reviewer and have revised accordingly.

      (28) L338: Any water content for these soils?

      We thank the reviewer for this comment. The water contents of soils from all study sites are reported in Supplementary Table 2.

      (29) L349-350: You mean that you measured the total soil extracted DNA and then added 1% as labeled 16S?

      Yes, for each soil sample, we extracted total soil DNA and quantified its concentration (ng DNA per gram of soil). We then added exogenous GAPDH‑tagged 16S amplicon fragments at an amount equal to 1% of this total DNA concentration. This concentration was chosen to mimic a realistic pulse of extracellular DNA input without overwhelming the endogenous DNA pool. We apologize for any confusion caused by the imprecise wording in the original manuscript.

      L392-398

      “The microcosm experiment was conducted using 30 g of soil for each sample. After pre-incubation at 20℃ for one week, each soil was thoroughly mixed with the GAPDH F‑tagged 16S rRNA gene amplicon fragments and incubated further at 20℃ (Fig. S8). The amount of exogenous GAPDH F‑tagged 16S rRNA gene amplicon fragments added to each soil sample was equivalent to 1% of the total DNA concentration naturally present in that soil, as determined fluorometrically prior to the experiment. This concentration was chosen to approximate natural eDNA fluxes resulting from microbial lysis, ensuring experimental relevance to in situ conditions (Table S2).”

      (30) L354: Remember that soil recovery from intact cells is going to be lower than for extracellular DNA. So, you are probably overestimating the contribution of extracellular DNA to the total DNA in the system.

      We thank the reviewer for this comment. We agree that DNA recovery from intact cells is generally lower than from extracellular DNA due to differential cell lysis efficiencies. Consequently, the contribution of extracellular DNA to total soil DNA may be somewhat overestimated in our study. We have clarified this limitation in the revised manuscript.

      L262-265

      “However, as DNA extraction efficiency may differ between intact cells and eDNA, the actual differences between total and living prokaryotic abundance could be smaller than those observed in this study. Similarly, the overestimated prokaryotic richness may arise from historically accumulated microbial taxonomic information stored in eDNA pools.”

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (31) L362: Amplification efficiency is pretty low. I think you would have been better served with GAPDH on both ends, and that would have given you a much higher efficiency qPCR.

      We thank the reviewer for this comment. The actual qPCR amplification efficiency in our assay was approximately 85%, which, although slightly below the ideal range, was still acceptable and produced reproducible amplification curves and reliable quantification for degradation-rate calculations.

      We acknowledge that the amplification efficiency in our qPCR experiments using a GAPDH F-labeled 16S primer on one end was suboptimal. The current design used a single GAPDH tag at the forward primer to avoid potential amplification bias or primer-dimer formation that could arise from extending the degenerate reverse primer. In addition, dual-end labeling would have required synthesis of new barcode-labeled tagged primers, increasing both cost and experimental complexity. Thanks again for the constructive comments, which provided us with the direction for future experiment optimization.

      L365-371

      “The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      (32) L367: Not enough detail on how barcoded libraries were made. UDIs?

      We thank the reviewer for this helpful comment. We have now clarified the library preparation and indexing strategy in the revised Methods section. This amplicon diversity sequencing used a pooled-library strategy. Individual samples were first distinguished by sample-specific inline barcodes introduced during the amplicon PCR step. After amplification, barcoded PCR products from multiple samples were pooled and subjected to library construction using the ALFA-SEQ DNA Library Prep Kit. Universal Illumina-compatible adapters were ligated to the pooled amplicons, followed by an indexing PCR that introduced the complete P5/P7 sequences and a library-level Illumina index. Thus, the Illumina index was used to identify the pooled sequencing library, whereas sample demultiplexing was performed according to the sample-specific inline barcodes. We have revised the Methods section to make this procedure explicit.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).”

      (33) L368: Why was the # of cycles increased?

      Thank you for your question. In the original manuscript (L368), we stated that the number of PCR cycles was increased to 35. This was mainly because the exogenously added GAPDH F‑labeled 16S rRNA genes had a relatively low initial abundance in the soil and gradually degraded during the microcosm incubation, with their copy numbers becoming particularly low at the last time points (see Fig. 1a). To ensure sufficient PCR product for high‑throughput sequencing from samples at all time points (especially those with low abundance at later stages), we appropriately increased the cycle number to 35.

      L429-431

      “The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing.”

      (34) L372: Were sequencing adapters ligated onto the pool?

      We thank the reviewer for this question. Yes, in this amplicon diversity sequencing workflow, sequencing adapters were ligated onto the pooled amplicon products. Briefly, individual samples were first amplified with sample-specific barcode sequences, allowing each sample to be distinguished after sequencing. The barcoded PCR products from multiple samples were then pooled for library construction. Universal Illumina-compatible adapters were ligated to this pooled amplicon library using the ALFA-SEQ DNA Library Prep Kit. After adapter ligation and purification, an indexing PCR was performed to introduce the complete P5/P7 sequences and a library-level Illumina index. We have clarified this pooled-library construction workflow in the revised Methods section.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).”

      (35) L378: "Amplicon sequence variants".

      We agree with the reviewer and have revised accordingly.

      (36) L380: Why were ASVs with fewer than 9 reads removed?

      We thank the reviewer for this question. The threshold of removing ASVs with fewer than 9 total reads across all samples was applied to reduce noise from sequencing errors and PCR artifacts. Our justification is supported by both the default parameters of the UNOISE3 algorithm and common practice in amplicon sequencing analysis.

      The USEARCH manual specifies that the -minsize parameter in the unoise3 command defaults to 8. This means that unique sequences occurring fewer than 8 times are discarded by the algorithm during ASV inference, as they are unlikely to represent true biological variants. Our threshold of 9 is slightly more conservative than the default (9 > 8), ensuring that only ASVs with a minimal level of abundance are retained. This choice is directly aligned with the algorithm’s intrinsic noise‑filtering logic.

      (37) L402: Please don't forget to discuss that PCR bias can contribute to uncertainty in the abundance of each taxon.

      Thank you for this important reminder. We agree that PCR bias (e.g., primer‑template mismatches, GC content differences, and variable amplification efficiency) can contribute to uncertainty in the abundance estimates of each taxon. Following your suggestion, we have now added a paragraph in the Discussion section to address this issue. We state that sequence‑specific degradation rates and PCR bias may jointly affect the accuracy of taxon abundance estimates, and future studies should incorporate internal standards or multiplex PCR strategies to correct for such biases. Thank you for your careful review.

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      (38) L414: Suggest: "To inhibit amplification of extracellular DNA, soils were incubated with propidium monoazide (PMA), as described previously (REF). Briefly, soil (X grams) was mixed with PMA in a total volume of Y (ml).

      We thank the reviewer for this suggestion. We have revised the Methods section to provide a clearer description of PMA treatment, specifying the soil amount (0.50 g) and the total volume (0.5 mL).

      L496-497

      “To inhibit amplification of eDNA, soils were incubated with PMA, as described previously (Carini et al., 2016).”

      L505-506

      “In this study, 0.50 g of soil was mixed with PMA in a total volume of 0.5 mL (40 µM PMA in phosphate‑buffered saline, PBS), while the control soil samples were mixed with PBS without PMA.”

      (39) L416: In contrast, microbes with intact cell membranes exclude PMA, and their DNA is not cross-linked with PMA, and remains amenable to PCR amplification.

      We agree and have revised.

      (40) L418-420: wording/sentence is strange and needs work.

      Thank you for pointing this out. We have reviewed the sentence at L418‑420 and agree that the wording is awkward. Moreover, the content only listed the advantages of the PMA method without acknowledging its limitations, making the statement less balanced. Therefore, in the revised manuscript, we have deleted this sentence. The limitations of the PMA method have been addressed in the Discussion section.

      L502-503

      “Currently, PMA treatment is a widely used to suppress PCR amplification of eDNA (Xue et al., 2023; Canini et al., 2024).”

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (41) L421-422: PMA treatment is a widely used method for inhibiting the enzymatic processing of extracellular DNA (Xue, Canini).

      We agree and have revised.

      (42) L425: include volume of PBA.

      We thank the reviewer for this comment. We have revised the Methods section to include the volume of PMA used

      L505-506

      “In this study, 0.50 g of soil was mixed with PMA in a total volume of 0.5 mL (40 µM PMA in phosphate‑buffered saline, PBS).”

      (43) L429-430: Don't use the word precipitates- use "pellets".

      We agree and have revised.

      (44) L433: "The abundance of 16S rRNA genes was determined using quantitative PCR employing a LightCycler...".

      We agree and have revised.

      (45) L445-: Section 4.9 - needs citations for PERMANOVA, NMDS, SEM, etc.

      Thank you for your suggestion. We have added the necessary citations for PERMANOVA, NMDS, SEM, and other methods in Section 4.9.

      L531-539

      Prokaryotic community structure differences among the study sites and incubation time points were examined through non-metric multidimensional scaling analysis (NMDS), permutation multivariate analysis of variance (PERMANOVA), and Permutational Analysis of Multivariate Dispersion (PERMDISP) (Kruskal, 1964; Anderson, 2001). Random forest modeling was conducted to assess the importance of environmental and soil variables in predicting the overall degradation rates of extracellular 16S rRNA gene amplicon fragments. Structural equation modeling (SEM) was employed to further evaluate the direct and indirect effects of soil moisture, soil pH, MAP, and prokaryotic abundance on the overall degradation rates of extracellular 16S rRNA gene amplicon fragments (Grace, 2006).

      (46) L698: A few comments. It would be nice to know how many different 16S sequences were tracked for differential degradation and shown in the figure.

      We thank the reviewer for this helpful comment. We would like to clarify that Fig. 1A does not track the degradation of individual 16S rRNA gene amplicon sequences, but instead shows the overall degradation dynamics of the total added exogenous DNA pool. The data points are derived from total 16S gene copy numbers measured via qPCR at each incubation time point. Consequently, this quantification inherently includes all sequences present within the added pool. The multiple lines visualized in the figure represent the collective degradation trajectories of the entire DNA pool across different study sites

      To address sequence-level changes, we further analyzed the richness and composition of the GAPDH F-tagged 16S rRNA gene amplicon fragments, which are presented in Fig. 2A and related analyses.

      (47) L699: Better to use "16S rRNA gene amplicon fragment abundance" as the term.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have replaced the original wording with “16S rRNA gene amplicon fragment abundance” where appropriate.

      (48) Y-axis for Figures 1A and 2A should be GAPDH-labeled, not ACTB-labeled.

      We apologize for this mistake. We have corrected this error in the revised manuscript.

      (49) For Figure 1b: Why not use box plots and ANOVA for different soil types?

      Thank you for your valuable suggestion. In the original Figure 1b, we used a bar plot to display the degradation rate constants across the 30 study sites. This choice was intended to emphasize the continuous variation among sites and their gradient relationships with environmental factors (e.g., soil moisture, MAP), which were then used in random forest and structural equation modeling. The bar plot better illustrates the spatial continuum of degradation rates rather than treating ecosystem types as discrete categories.

      Nevertheless, we fully agree that a boxplot grouped by ecosystem type (grassland, forest, cropland, desert) would help readers quickly grasp the overall differences among land‑use types. In the revised manuscript, we have added a boxplot grouped by ecosystem type and performed one‑way ANOVA followed by Tukey HSD post‑hoc tests (Fig. 1.). The results show that degradation rate constants differ significantly among ecosystem types (P < 0.05).

      L124-128

      “The degradation rate constants of the spiked extracellular 16S rRNA gene amplicon fragments displayed considerable variability among the study sites, ranging from 0.05 to 0.16 day<sup>-1</sup> (Fig. 1b). Furthermore, we found that degradation rate constants differed significantly among ecosystem types (Fig. 1c, P < 0.05). Specifically, cropland and forest soils exhibited significantly higher degradation rates than grassland soils (P < 0.05).”

      (50) For Figure 2: Where are PERMANOVA and PERMDISP values for the figure?

      We thank the reviewer for this comment. In the revised manuscript, we have added the PERMANOVA and PERMDISP values corresponding to Figure 2 in the figure legend and Results section (Fig. 2).

      (51) I found Figure 2b to be hard to see. The 48-day circles are almost invisible. Difficult to know what the authors are trying to show here, since there is so much variability associated with soil type.

      We thank the reviewer for this comment. Figure 2b is intended to illustrate the temporal changes in microbial community structure during the incubation. The different colored circles represent samples at different time points (1, 3, 6, 12, 24, and 48 days), showing how communities shift over time. We apologize that in the original Figure 2b, the 48‑day samples were nearly invisible and that the high variability among soil types obscured the intended message. In the revised manuscript, we have added a black border around every data point, which greatly enhances the visibility of the 48‑day samples (and all time points). We now use distinct shapes to represent different ecosystem types (grassland, forest, cropland, desert) in the NMDS ordination, and added PERMANOVA results in both the Results section and the figure legend (Fig. R2b).

      (52) Figure 4A: Y-axis need a label like "16S rRNA gene abundance".

      We agree and have revised.

      (53) Figure 4B: I'd like to see a Shannon index too, not just richness.

      Thank you for your suggestion. We agree that the Shannon index, which integrates both richness and evenness, provides a valuable complement to richness alone. In the revised manuscript, we added an analysis of the Shannon index to compare α‑diversity between total DNA (PMA‑untreated) and intact cell DNA (PMA‑treated) samples (Fig.4).

      (54) Figure 4D: Would be good to have lines linking the intact cell vs total abundance. Also, what about a box plot of Bray-Curtis (or similar) dissimilarity between intact cell and total microbial analysis across the dataset?

      Thank you for your suggestions. Regarding the addition of connecting lines in Figure 4D, after careful consideration we decided not to add them for the following reason: the total and PMA-treated communities from the same site are already coded with the same color (different colors for different sites), which effectively indicates the pairing. Adding lines would greatly reduce readability due to dense overlapping lines, especially given the number of sites. Therefore, we kept the original color‑based pairing design.

      To address your second suggestion, we have added a bar plot showing the distribution of Bray‑Curtis dissimilarities between total (PMA‑untreated) and intact cell (PMA‑treated) communities across all study samples (Fig. R3d).

      (55) Figure 5B: What do correlations with p > 0.05 show? I would remove these from the image.

      We thank the reviewer for this suggestion. We agree that correlations with p > 0.05 do not represent statistically significant relationships and may cause confusion. In the revised manuscript, we have removed these non-significant correlations from Figure 5B.

      (56) Figure 6: "Incubations of 0, 3, 6, 12, 24, and 48 days".

      We agree with the reviewer and have revised as suggested.

      References

      Albertsen, M., Karst, S.M., Ziegler, A.S., Kirkegaard, R.H., Nielsen, P.H., 2015. Back to basics–the influence of DNA extraction and primer choice on phylogenetic analysis of activated sludge communities. PLoS One 10, e0132783.

      Anderson, M.J., 2001. A new method for non‐parametric multivariate analysis of variance. Austral Ecology 26, 32-46.

      Arvizu-Hernandez, E., Ocadiz-Delgado, R., Gariglio, P., 2025. E7HPV16 Oncogene and 17beta-Estradiol Stress Promote Oncogenic microRNA Expression Patterns, Cell Proliferation and Cervical Intraepithelial Neoplasia 1. Cell Biochemistry and Function 43, e70065.

      Buitrago, D., Labrador, M., Arcon, J.P., Lema, R., Flores, O., Esteve-Codina, A., Blanc, J., Villegas, N., Bellido, D., Gut, M., 2021. Impact of DNA methylation on 3D genome structure. Nature Communications 12, 3243.

      Cai, P., Huang, Q., Zhang, X., Chen, H., 2006a. Adsorption of DNA on clay minerals and various colloidal particles from an Alfisol. Soil Biology and Biochemistry 38, 471-476.

      Cai, P., Huang, Q.Y., Zhang, X.W., 2006b. Interactions of DNA with clay minerals and soil colloidal particles and protection against degradation by DNase. Environmental Science & Technology 40, 2971-2976.

      Carini, P., Marsden, P.J., Leff, J.W., Morgan, E.E., Strickland, M.S., Fierer, N., 2016. Relic DNA is abundant in soil and obscures estimates of soil microbial diversity. Nature Microbiology 2, 1-6.

      Deshpande, A.S., Fahrenfeld, N.L., 2023. Influence of DNA from non-viable sources on the riverine water and biofilm microbiome, resistome, mobilome, and resistance gene host assignments. Journal of Hazardous materials 446, 130743.

      Du, Y., Wang, Z., Liu, K., Chai, G., Chi, Y., Li, T., Duan, Y., Xia, T., Liu, D., Che, R., 2025. The performance of different methods in characterizing soil live prokaryotic diversity and abundance is highly variable. iMetaOmics, e70011.

      Emerson, J.B., Adams, R.I., Román, C.M.B., Brooks, B., Coil, D.A., Dahlhausen, K., Ganz, H.H., Hartmann, E.M., Hsu, T., Justice, N.B., 2017. Schrödinger’s microbes: tools for distinguishing the living from the dead in microbial ecosystems. Microbiome 5, 86.

      Finkel, S.E., Kolter, R., 2001. DNA as a nutrient: novel role for bacterial competence gene homologs. Journal of Bacteriology 183, 6288-6293.

      Frostegård, Å., Courtois, S., Ramisse, V., Clerc, S., Bernillon, D., Le Gall, F., Jeannin, P., Nesme, X., Simonet, P., 1999. Quantification of bias related to the extraction of DNA directly from soils. Applied and Environmental Microbiology 65, 5409-5420.

      Grace, J.B., 2006. Structural equation modeling and natural systems. Cambridge University Press.

      He, P., Li, L.-J., Dai, S.-S., Guo, X.-L., Nie, M., Yang, X., Kuzyakov, Y., 2024. Straw addition and low soil moisture decreased temperature sensitivity and activation energy of soil organic matter. Geoderma 442, 116802.

      Heise, J., Nega, M., Alawi, M., Wagner, D., 2016. Propidium monoazide treatment to distinguish between live and dead methanogens in pure cultures and environmental samples. Journal of Microbiological Methods 121, 11-23.

      Huang, C., Xie, D.C., Cui, J.J., Li, Q., Gao, Y., Xie, K.P., 2014. FOXM1c Promotes Pancreatic Cancer Epithelial-to-Mesenchymal Transition and Metastasis via Upregulation of Expression of the Urokinase Plasminogen Activator System. Clinical Cancer Research 20, 1477-1488.

      Knight, R., Vrbanac, A., Taylor, B.C., Aksenov, A., Callewaert, C., Debelius, J., Gonzalez, A., Kosciolek, T., McCall, L.-I., McDonald, D., 2018. Best practices for analysing microbiomes. Nature Reviews Microbiology 16, 410-422.

      Kruskal, J.B., 1964. Nonmetric multidimensional scaling: a numerical method. Psychometrika 29, 115-129.

      Lennon, J.T., Muscarella, M.E., Placella, S.A., Lehmkuhl, B.K., 2018. How, when, and where relic DNA affects microbial diversity. mbio 9, e00637-00618.

      Levy-Booth, D.J., Campbell, R.G., Gulden, R.H., Hart, M.M., Powell, J.R., Klironomos, J.N., Pauls, K.P., Swanton, C.J., Trevors, J.T., Dunfield, K.E., 2007. Cycling of extracellular DNA in the soil environment. Soil Biology and Biochemistry 39, 2977-2991.

      Liu, Q.H., Yuan, L., Li, Z.H., Leung, K.M.Y., Sheng, G.P., 2024. Natural organic matter enhances natural transformation of extracellular antibiotic resistance genes in sunlit water. Environmental Science & Technology 58, 17990-17998.

      Marrone, A., Ballantyne, J., 2008. Sequence Specificity of BAL 31 Nuclease for ssDNA Revealed by Synthetic Oligomer Substrates Containing Homopolymeric Guanine Tracts. PLoS One 3, e3595.

      McKinney, C.W., Dungan, R.S., 2020. Influence of environmental conditions on extracellular and intracellular antibiotic resistance genes in manure-amended soil: A microcosm study. Soil Science Society of America Journal 84, 747-759.

      Morrissey, E.M., McHugh, T.A., Preteska, L., Hayer, M., Dijkstra, P., Hungate, B.A., Schwartz, E., 2015. Dynamics of extracellular DNA decomposition and bacterial community composition in soil. Soil Biology and Biochemistry 86, 42-49.

      Nagler, M., Insam, H., Pietramellara, G., Ascher-Jenull, J., 2018. Extracellular DNA in natural environments: features, relevance and applications. Applied Microbiology and Biotechnology 102, 6343-6356.

      Nocker, A., Sossa-Fernandez, P., Burr, M.D., Camper, A.K., 2007. Use of propidium monoazide for live/dead distinction in microbial ecology. Applied and Environmental Microbiology 73, 5111-5117.

      Pietramellara, G., Ascher, J., Borgogni, F., Ceccherini, M., Guerri, G., Nannipieri, P., 2009. Extracellular DNA in soil and sediment: fate and ecological relevance. Biology and Fertility of Soils 45, 219-235.

      Shah, A., Huang, J., Han, T., Khan, M.N., Tadesse, K.A., Daba, N.A., Khan, S., Ullah, S., Sardar, M.F., Fahad, S., 2024. Impact of soil moisture regimes on greenhouse gas emissions, soil microbial biomass, and enzymatic activity in long-term fertilized paddy soil. Environmental Sciences Europe 36, 120.

      Sirois, S.H., Buckley, D.H., 2019. Factors governing extracellular DNA degradation dynamics in soil. Environmental Microbiology Reports 11, 173-184.

      Vuillemin, A., Horn, F., Alawi, M., Henny, C., Wagner, D., Crowe, S.A., Kallmeyer, J., 2017. Preservation and significance of extracellular DNA in ferruginous sediments from Lake Towuti, Indonesia. Frontiers in Microbiology 8, 1440.

      Wang, X., Ganzert, L., Bartholomaus, A., Amen, R., Yang, S., Guzman, C.M., Matus, F., Albornoz, M.F., Aburto, F., Oses-Pedraza, R., Friedl, T., Wagner, D., 2024. The effects of climate and soil depth on living and dead bacterial communities along a longitudinal gradient in Chile. The Science of the total environment 945, 173846.

      Wang, Y.-T., Yang, W.-J., Li, C.-L., Doudeva, L.G., Yuan, H.S., 2007. Structural basis for sequence-dependent DNA cleavage by nonspecific endonucleases. Nucleic Acids Research 35, 584-594.

      Wang, Y., Yan, Y., Thompson, K.N., Bae, S., Accorsi, E.K., Zhang, Y., Shen, J., Vlamakis, H., Hartmann, E.M., Huttenhower, C., 2021. Whole microbial community viability is not quantitatively reflected by propidium monoazide sequencing approach. Microbiome 9, 1-13.

      Wolpe, J., Guertin, M., 2022. Regional and Single Nucleotide Correction of Sequence Bias in Chromatin Accessibility Data. The FASEB Journal 36.

      Yang, A., Liu, X., Liu, P., Feng, Y.Z., Liu, H.B., Gao, S., Huo, L.M., Han, X.Y., Wang, J.R., Kong, W., 2021. LncRNA UCA1 promotes development of gastric cancer via the miR-145/MYO6 axis. Cellular & Molecular Biology Letters 26, 33.

      Ye, M., Zhang, Z., Sun, M., Shi, Y., 2022. Dynamics, gene transfer, and ecological function of intracellular and extracellular DNA in environmental microbiome. iMeta 1, e34.

    1. Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. In particular, two types of learned associations are studied. One association involves associations between stimulus features and control states (SC), and the other involves stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates.

      The authors conclude that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing that SC and SR correlates are (1) not entirely overlapping in cross-decoding, (2) simultaneously observed on average over trials, (3) independently correlate with RT, and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Response:

      In their last response to the reviewers, the authors write:

      "... constructing a theoretically unbiased decoder requires perfectly counter-balanced training data (i.e., for every training trial of class A that is X trials away from the test data, there must be a training trial of all other classes that is exactly X trials away from the test data). As we were unable to achieve such a perfect design, we chose not to run an additional experiment."

      This isn't really an issue about whether this design is "perfectly" orthogonal. It's an issue regarding a clear confound among the decoded classes for SC/SR decoders. To be clear: of the 8 classes in the SC decoder, 4 are overwhelmingly presented in the first half (PHASE 2) of the session, whereas the other 4 are overwhelmingly presented in the second half (PHASE 3). The same is true for the SR decoder. So, session-half correlated noise could readily contribute to distinguishing among these classes. And counterbalancing this across subjects won't help because decoders lose sign.

      To me, the conducted control analyses don't really make strong contact with this issue. The split-half cross-validation is a nice idea but, as the authors acknowledge, it's also subject to slower cross-session noise, as is the original analysis. This sort of noise is not exactly exotic in EEG. Caps/hair/electrodes shift, gel dries and impedance changes, posture / muscle tension / skin conductance changes, fatigue may wax and wane (e.g., linked to increasing alpha), etc. And the newest analysis didn't really seem to engage with this issue either, as it only assessed minimum distances between classes, on the order of 5 +- 2 SD trials. This seems to assume that the dominant potential sources of noise will be scale-free, such that the strength of the relation at short time scales would generalize to longer ones. I'm not sure why that's expected here.

      Here are some suggestions for alternative control analyses that I think would be more targeted to this issue:

      (1) Explicitly train a decoder to separate the three levels of PHASE from each other. Successful decoding would provide positive evidence for the presence of structured noise at this timescale.

      (2) Specify an RDM for the PHASE variable and regress this component separately from each time-point/trial of the SC and SR decoders. This is a post-hoc band-aid, but it is in the spirit of correcting for a known confound.

      (3) In the spirit of the authors' distance analysis, but without assuming that the noise is scale-free: perform a time-series RSA like that in Alink et al. (2015; https://doi.org/10.1101/032391), Fig. 1 and 3. This would allow one, e.g., to estimate the structure & timescales of the noise processes across the session.

      Other readers may, like me, be puzzled by the selection of this particular experimental design to test this question of SC and SR coding, given the temporal confound among SC/SR classes, and given that there would seem to be many possible designs that are less confounded. For example, why not use a design where ISPC was swapped/shuffled several more times within each subject, so that PHASE is more orthogonal to long-timescale noise? Isn't ISPC learning fast enough to support learning phases shorter than 700 trials? Such readers would likely appreciate a frank discussion of this dilemma, and a motivation for the choice of the present design, within the manuscript.

      Pre-stimulus coding:

      To explain the apparent pre-stimulus coding of several task variables, the newest version of the manuscript proposes that subjects were proactively coding these variables via predictive mechanisms. This is an interesting account of item-specific control. It is also surprising, given that item-specific control mechanisms are typically conceptualized as reactive or stimulus-driven phenomena. But I think support for a proactive control account was incomplete. The mechanistic logic was not presented, and no hypotheses under this account were developed or tested. So I would suggest pinning down some hypotheses here and actually putting this account to the test.

      Outliers & t-values: thank you for checking this!

      Random slopes were omitted due to convergence failure, but this can inflate false positive inferences (e.g., Barr et al. 2013), and doesn't really motivate a minimal model. I'd suggest trying a slightly reduced model (e.g., drop correlations via `slope || subject`) using buildMer automated selection, or switching to brms.

    2. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

      We thank the editors and the reviewers for their positive assessment and constructive feedback of our work, which led us to think more deeply about the conceptual and methodological aspects of this project and further strengthen the manuscript. Based on the comments, we included more control analyses and revised the manuscript accordingly. Please see below our responses to each comment raised in the reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. Two types of learned associations are characterized, one being associations between stimulus features and control states (SC), and the other being stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates, which are used to test properties of their dynamics.

      The conclusion is reached that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing SC and SR correlates are: (1) not entirely overlapping in cross-decoding; (2) simultaneously observed on average over trials in overlapping time bins; (3) independently correlate with RT; and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Thank you for acknowledging the strength in EEG decoding and design. We have addressed all your concerns raised below point by point.

      Weaknesses:

      I still have my concern from the first round that the decoders are overfit to temporally structured noise. As I wrote before, the SC and SR classes are highly confounded with phase (chunk of session). I do not see how the control analyses conducted in the revision adequately deal with this issue.

      In the figures, there are several hints that these decoders are biased. Unfortunately, the figures are also constructed in such a way that hides or diminishes the salience of the clues of bias. This bias and lack of transparency discourage trust in the methods and results.

      I have two main suggestions:

      (1) Run a new experiment with a design that properly supports this question.

      I don't make this suggestion lightly, and I understand that it may not be feasible to implement given constraints; but I feel that this suggestion is warranted. The desired inferences rely on successful identification of SC and SR representations. Solidly identifying SC and SR representations necessitates an experimental design wherein these variables are sufficiently orthogonalized, within-subject, from temporally structured noise. The experimental design reported in this paper unfortunately does not meet this bar, in my opinion (and the opinion of a colleague I solicited).

      An adequate design would have enough phases to properly support "cross-phase" cross-validation. Deconfounding temporal noise is a basic requirement for decoding analyses of EEG and fMRI data (see e.g., leave-one-run-out CV that is effectively necessary in fMRI; in my experience, EEG is not much different, when the decoded classes are blocked in time, as here). In a journal with a typical acceptance-based review process, this would be grounds for rejection.

      Please note that this issue of decoder bias would seem to weaken the rest of the downstream analyses that are based on the decoded values. For instance, if the decoders are biased, in the within-trial correlation analysis, how can we be sure that co-fluctuations along certain dimensions within their projected values are driven by signal or noise? A similar issue clouds the LMM decoding-RT correlations.

      We appreciate the reviewer’s concern with the potential confound of temporally structured noise (TSN) in the EEG data. As we understand it, TSN refers to a process that the noise structure drifts over time. It follows that noise structure should be more similar for temporally closer trials and that the TSN’s bias on decoding accuracy is stronger for test trials that are closer to the training data. In the previous round of revision, we conducted a control analysis that reduced the influence of TSN by maximizing the temporal distance between training and test data (the distance between the centers of the training and test data of the same SC/SR manipulation is about 400 trials given the experimental design) and showed comparable decoding accuracy with the main results. As the reviewer finds this analysis unconvincing, we reason that the reviewer believes that the TSN has a long-term effect, such that it remains relatively stable over time and can be picked up by trials temporally distant from the training data. With this assumption and the assumption that this effect may not be linear, constructing a theoretically unbiased decoder requires perfectly counter-balanced training data (i.e., for every training trial of class A that is X trials away from the test data, there must be a training trial of all other classes that is exactly X trials away from the test data). As we were unable to achieve such a perfect design, we chose not to run an additional experiment. Instead, we focused on testing whether and how much TSN systematically biased the reported decoding accuracy.

      Please note that the existence of TSN in the EEG data is not sufficient to rule that the decoding results are biased. As TSN is stronger for trials closer to each other, the idea that auto-correlation biases decoding results would predict a distance effect, such that if a test trial is closer to a training trial of the same trial type, the higher similarity in TSN between the training and test data would more strongly inflate the decoding accuracy of the test trial, resulting in a negative correlation between distance between a test trial and its closest training trial of the same type and the test trial’s decoding accuracy. To test this predicted negative correlation, in each fold and each repetition of the cross-validation reported in the SC-SC and SR-SR decoders in Fig. 4, we calculated the distance (mean=5.84 trials, SD=2.05, 5th percentile =2.87, 95th percentile=9.45, one trial = 2.4-2.6s) between each test trial and its closest training trial of the same trial type. This distance was used as the predictor to predict decoding accuracy in a linear regression. Note that even if the relation between distance and decoding accuracy is non-linear, the linear relation will be negative because the relation is monotonic (similarity in noise structure decreases monotonically with temporal distance between trials). Similarly, because the effect is monotonic, if a long-range effect exists, it should also exist in short-range and be picked up by the distance range in this analysis. The regression coefficient is averaged across cross-validation folds and repetitions for each subject to match how the decoding accuracy was reported in the main text. Finally, the averaged regression coefficient was tested against 0 using a one-sample t-test. This analysis was conducted at each time point (from -250ms to 1500ms) separately. As shown in the figure below, no time point exhibited the negative correlation as predicted by the auto-correlation account. An alternative explanation is that this result indicates that TSN remains stable over time. If this is the case, TSN will be shared by all trials and will be unable to bias decoding results. Together with the control analysis introduced previously, this new control analysis supports the notion that the decoding results are not inflated by TSN in the EEG data. We included all the control analyses in the revised manuscript (page 13-14). Please note that this analysis is specific for the present dataset and we strongly agree with the reviewer that TSN is a key confounding factor in EEG analysis in general and should be carefully addressed.

      Lastly, we understand the concern with the early onset of above-chance decoding accuracy. Here, we provide an explanation: because of the blocked design (i.e., participants performed hundreds of trials with the same SC/SR associations), it is possible the participants learned the associations and used them to guide proactive cognitive control. As proactive cognitive control is anticipatory and sustained (Braver, 2012; Khan et al., 2025), it may be able to be decoded early on a trial, or even before trial onset. In the revised manuscript, we discussed this account along with the TSN issue as a limitation of the current project and directions for future research (page 24).

      (2) Increase transparency in the reporting of results throughout main text.

      Please do not truncate stimulus-aligned timecourses at time=0. Displaying the baseline period is very useful to identify bias, that is, to verify that stimulus-dependent conditions cannot be decoded pre-stimulus. Bias is most expected to be revealed in the baseline interval when the data are NOT baseline-corrected, which is why I previously asked to see the results omitting baseline correction. (But also note that if the decoders are biased, baseline-correcting would not remove this bias; instead, it would spread it across the rest of the epoch, while the baseline interval would, on average, be centered at zero.)

      Please use a more standard p-value correction threshold, rather than Bonferroni-corrected p<0.001. This threshold is unusually conservative for this type of study. And yet, despite this conservativeness, stimulus-evoked information can be decoded from nearly every time bin, including at t=0. This does not encourage trust in the accuracy of these p-values. Instead, I suggest using permutation-based cluster correction, with corrected p<0.05. This is much more standard and would therefore allow for better comparison to many other studies.

      I don't think these things should be done as control analyses, tucked away in the supplemental materials, but instead should be done as a part of the figures in the main text -- including decoding, RSA, cross-trial correlations, and RT correlations.

      Thank you for your suggestions. we have added the baseline period from 200 to 0 ms prior to the stimulus onset in all the stimulus-locked analyses and tested the significance with cluster-based permutation test (cluster-forming threshold p < 0.001, cluster-level p < 0.05, (Collins & Frank, 2018)) in all the analyses including decoding, RSA, cross-trial correlations and RT correlations. The results showed similar patterns, and they are all reported in the main text (please see all the figures and page 30-32 in the main text).

      Other issues:

      Regarding the analysis of the within-trial correlation of RSA betas, and "Cai 2019" bias:<br /> The correction that authors perform in the revision -- estimating the correlation within the baseline time interval and subtracting this estimate from subsequent timepoints -- assumes that the "Cai 2019" bias is stationary. This is a fairly strong assumption, however, as this bias depends not only on the design matrix, but also on the structure of the noise (see the Cai paper), which can be non-stationary. No data were provided in support of stationarity. It seems safer and potentially more realistic to assume non-stationarity.

      This analysis was included in the supplemental material. However, given that the correlation analysis presented in the Results is subject to the "Cai 2019" bias, it would seem to be more appropriate to replace that analysis, rather than supplement it.

      Regardless, this seems to be a moot issue, given that the underlying decoders seem to be overfit to temporally structured noise (see point above regarding weakening of downstream analyses based on decoder bias).

      Thank you for this important point. We now replaced the previous control analysis with a new one that does not assume stationary noise structure (page 19 in the revised manuscript). In Cai et al (2019), the source of confound is the covariance between observations. Specifically, as the observations in fMRI data are the BOLD signal at different time points, TSN can introduce covariance between nearby observations, which further biases the observed correlation between experimental conditions/trial types. In our case, the observations are decoding accuracy for different trial types. Thus, bias in the correlation may come from covariance between trial types. In this study, potential covariance between trial types includes the constrain that the decoding accuracy of all trial types adds up to 1 for a given trial (although we transformed the accuracy into logits prior to RSA, so the constrain may not hold), and the blocked design (as discussed above). Thus, to establish a baseline level of correlation between SC and SR representation strength, we took a similar shuffling approach as in Cai et al (2019) and randomly shuffled the trial types within each block. The reason to shuffle within each block is to preserve the covariance structure in the blocked design. We then repeated the same analysis using the shuffled data. The results of 10 shuffled analysis were averaged to form a baseline. Please note that (1) this control analysis was performed separately at each time point, hence removing the assumption of stationary noise structure, (2) this analysis also included as noise any covariance introduced by the proactive cognitive control guided by the learned SC and SR associations (see response to comment 1), thus it is more stringent than intended and (3) this control analysis started from decoding and was intended to provide a baseline for all downstream analysis. As shown in figures 2A, 3A, 7C and 8C, the reviewer was correct that the bias was not stationary, as the baseline of correlation coefficient varies over time. Additionally, the SC-SR representation strength correlation remained significantly above baseline between ~100 and ~ 450 ms following stimulus onset and between -180 and + 50 ms relative to response, suggesting that the noise structure (even when including potential proactive cognitive control) cannot fully explain the observed the SC-SR representation strength correlation. Considering the fact that this control analysis treated proactive control as a source of confound, this result does not necessarily contradict the absence of distance effect reported above.

      Outliers and t-values:

      More outliers with beta coefficients could be because the original SD estimates from the t-values are influenced more by extreme values. When you use a threshold on the median absolute deviation instead of mean +/-SD, do you still get more outliers with beta coefficients vs t-values?

      Thank you for your suggestion. We calculated the proportion of outliers with a threshold of median absolute deviation (defined as values beyond median ± 5 median absolute deviation) for each subject. The outliers remained less frequent for t-values than for beta coefficients (t-values: mean = 1.08%, SD = 0.12%; beta-values: mean = 4.45%, SD = 0.28%). Based on these results and to maintain consistent with previous studies employing the methods (Cellier et al., 2022; Kikumoto & Mayr, 2020; Kikumoto et al., 2022a; Kikumoto et al., 2022b; Rangel et al., 2023), we still decided to stay with t-values.

      Random slopes:

      Were random slopes (by subject) for all within-subject variables included in the LMMs? If not, please include them, and report this in the Methods.

      Thank you for your suggestion. The model failed to converge with random slopes of all variables. Thus, we chose not to add random slopes in the LMM. But we have added the random effects structure in the methods (see page 34).

      Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a sold design and the analyses are appropriate for the research questions.

      Strength:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      We thank you for acknowledging our work on design and brain-behavior analysis. We have addressed all your concerns raised below point by point.

      Weakness:

      (1a) The distinction between Phase 2 and Phase 1&3 behavioral results, specifically the opposite effect of MC/MI in congruent trials raises some concerns with regard to the effectiveness of the ISPC manipulation. Why do RTs and error rates under MC congruent condition in Phase 2 seem to be worse than MI congruent?

      Thank you for raising these issues. In Phase 1, one color set (red and blue) was assigned to the MC condition, whereas another color set (yellow and green) was assigned to the MI condition. In Phase 2, these assignments were flipped, and they were flipped back again in Phase 3. Thus, the MC condition consisted of red and blue in Phases 1 and 3 but yellow and green in Phase 2, whereas the MI condition consisted of yellow and green in Phases 1 and 3 but red and blue in Phase 2 (Fig. 1b in the manuscript). This manipulation leads to seemingly opposite patterns between Phases 1 & 3 and Phase 2.

      However, when considering specific colors, the pattern is consistent across phases. In Phase 2, RTs and error rates for yellow and green (MC congruent) were worse than those for red and blue (MI congruent), which mirrors the pattern observed in Phases 1 and 3, where RTs and error rates for yellow and green (MI congruent) were worse than those for red and blue (MC congruent)

      We interpreted the results in Phase 2 as reflecting a typical ISPC effect, which is defined as a smaller conflict effect in the MI condition (MI incongruent – MI congruent) compared with the MC condition (MC incongruent – MC congruent). To our knowledge, the ISPC paradigm does not impose a specific prediction regarding the relative difference between MC-congruent and MI-congruent conditions.

      (1b) Could there be other factors at play here, e.g. order effect?

      We agree that order effect could play a role, such that memory from Phase 1 may influence the pattern in phase 2. For example, in phase 1, yellow and green were assigned to the MI condition, and participants therefore have associated these colors with a high control state (SC) and incongruent responses (SR). These prior associations could interfere with the newly learned mappings in Phase 2, where yellow and green were reassigned to the MC congruent condition (i.e., low control state and congruent responses). As a result, memory from Phase 1 may have weakened the expected MC in phase 2. A similar effect could also apply to the MI condition. Consequently, the same condition does not show parallel performance between phase 1 and phase 2, which may lead to different patterns in the difference between MC congruent and MI congruent conditions in phase 2.

      (1c) How does this potentially affect the neural analyses where trials from different phases were combined?

      Thank you for the question. As we mentioned above, the order effect could slow down the newly learned associations. However, we still found the ISPC effect in each phase, suggesting that all kinds of both SC and SR associations were formed and could be applied to the decoding and the following analyses cross phases. Relatedly, there might be confounded with temporal structured noise (TSN) when the neural analyses on decoding were combined the trials from different phases. However, we have performed the control decoding analyses and distance effect tests and confirmed that our decoding results were not driven by TSN (Please see comment #1 of R1).

      (1d) the manuscript does not mention whether there is counterbalancing for the color groups across participants, so far as I can tell.

      Thank you for the reminder. We have balanced the color groups by randomly dividing the participants into two groups and assigning different color sets to each group. The related interpretations have been included in task overview of the revised manuscript (page 6), which reads:

      “The color groups were counterbalanced across participants by red and blue as the color set of MC in one group while as the color set of MI in another group in the phase 1.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I commend the authors for addressing and clarifying my previous questions. One new comment regarding the newly added Figure 9: the response-locked behavioral correlation is much weaker compared to the stimulus-locked one, even never reaching the significance level. I think this difference should be discussed instead of simply glossing over it.

      Thank you for your suggestion. We discussed the difference in the discussion of revised manuscript (page 25), which reads:

      “Note that we found the negative prediction of the strength of SC and SR to RTs did not reach statistical significance with response-locked analysis as stimulus-locked analysis. It is possible that SC and SR representations have occurred before the stage of response processing, which is usually aligned with stimulus onset (Jiang et al., 2020a; Kang & Yu-Chin, 2024; Khan et al., 2025)”

      References

      Braver, T. S. (2012). The variable nature of cognitive control: a dual mechanisms framework. Trends Cogn Sci, 16(2), 106-113. doi:10.1016/j.tics.2011.12.010

      Cellier, D., Petersen, I. T., & Hwang, K. (2022). Dynamics of Hierarchical Task Representations. J Neurosci, 42(38), 7276-7284. doi:10.1523/JNEUROSCI.0233-22.2022

      Collins, A. G., & Frank, M. J. (2018). Within- and across-trial dynamics of human EEG reveal cooperative interplay between reinforcement learning and working memory. Proceedings of the National Academy of Sciences, 115(10), 2502-2507. doi:10.1073/pnas.1720963115

      Khan, A. U., Hoy, C. W., Anderson, K. L., Piai, V., King-Stephens, D., Laxer, K. D., . . . Bentley, J. N. (2025). Neural dynamics of proactive and reactive cognitive control in medial and lateral prefrontal cortex. iScience, 28(9), 113375. doi:10.1016/j.isci.2025.113375

      Kikumoto, A., & Mayr, U. (2020). Conjunctive representations that integrate stimuli, responses, and rules are critical for action selection. Proc Natl Acad Sci 117(19), 10603-10608. doi:10.1073/pnas.1922166117

      Kikumoto, A., Mayr, U., & Badre, D. (2022a). The role of conjunctive representations in prioritizing and selecting planned actions. Elife, 11. doi:10.7554/eLife.80153

      Kikumoto, A., Sameshima, T., & Mayr, U. (2022b). The Role of Conjunctive Representations in Stopping Actions. Psychol Sci, 33(2), 325-338. doi:10.1177/09567976211034505

      Rangel, B. O., Hazeltine, E., & Wessel, J. R. (2023). Lingering Neural Representations of Past Task Features Adversely Affect Future Behavior. J Neurosci, 43(2), 282-292. doi:10.1523/JNEUROSCI.0464-22.2022

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study of Drosophila mating behaviors has offered a powerful entry point for understanding how complex innate behaviors are instantiated in the brain. The effectiveness of this behavioral model stems from how readily quantifiable many components of the courtship ritual are, facilitating the fine-scale correlations between the behaviors and the circuits that underpin their implementation. Detailed quantification, however, can be both time-consuming and error-prone, particularly when scored manually. Song et al. have sought to address this challenge by developing DrosoMating, software that facilitates the automated and high-throughput quantification of 6 common metrics of courtship and mating behaviors. Compared to a human observer, DrosoMating matches courtship scoring with high fidelity. Further, the authors demonstrate that the software effectively detects previously described variations in courtship resulting from genetic background or social conditioning. Finally, they validate its utility in assaying the consequences of neural manipulations by silencing Kenyon cells involved in memory formation in the context of courtship conditioning.

      Strengths:

      (1) The authors demonstrate that for three key courtship/mating metrics, DrosoMating performs virtually indistinguishably from a human observer, with differences consistently within 10 seconds and no statistically significant differences detected. This demonstrates the software's usefulness as a tool for reducing bias and scoring time for analyses involving these metrics.

      (2) The authors validate the tool across multiple genetic backgrounds and experimental manipulations to confirm its ability to detect known influences on male mating behavior.

      (3) The authors present a simple, modular chamber design that is integrated with DrosoMating and allows for high-throughput experimentation, capable of simultaneously analyzing up to 144 fly pairs across all chambers.

      Weaknesses:

      (1) DrosoMating appears to be an effective tool for the high-throughput quantification of key courtship and mating metrics, but a number of similar tools for automated analysis already exist. FlyTracker (CalTech), for instance, is a widely used software that offers a similar machine vision approach to quantifying a variety of courtship metrics. It would be valuable to understand how DrosoMating compares to such approaches and what specific advantages it might offer in terms of accuracy, ease of use, and sensitivity to experimental conditions.

      (2) The courtship behaviors of Drosophila males represent a series of complex behaviors that unfold dynamically in response to female signals (Coen et al., 2014; Ning et al., 2022; Roemschied et al., 2023). While metrics like courtship latency, courtship index, and copulation duration are useful summary statistics, they compress the complexity of actions that occur throughout the mating ritual. The manuscript would be strengthened by a discussion of the potential for DrosoMating to capture more of the moment-to-moment behaviors that constitute courtship. Even without modifying the software, it would be useful to see how the data can be used in combination with machine learning classifiers like JAABA to better segment the behavioral composition of courtship and mating across genotypes and experimental manipulations. Such integration could substantially expand the utility of this tool for the broader Drosophila neuroscience community.

      (3) While testing the software's capacity to function across strains is useful, it does not address the "universality" of this method. Cross-species studies of mating behavior diversity are becoming increasingly common, and it would be beneficial to know if this tool can maintain its accuracy in Drosophila species with a greater range of morphological and behavioral variation. Demonstrating the software's performance across species would strengthen claims about its broader applicability.

      Reviewer #2 (Public review):

      This paper introduces "DrosoMating," an integrated hardware and software solution for automating the analysis of male Drosophila courtship. The authors aim to provide a low-cost, accessible alternative to expensive ethological rigs by utilizing a custom acrylic chamber and smartphone-based recording. The system focuses on quantifying key temporal metrics-Courtship Index (CI), Copulation Latency (CL), and Mating Duration (MD)-and is applied to behavioral paradigms involving memory mutants (orb2, rut).

      The development of open-source behavioral tools is a significant contribution to neuroethology, and the authors successfully demonstrate a system that simplifies the setup for large-scale screens. A major strength of the work is the specific focus on automating Copulation Latency and Mating Duration, metrics that are often labor-intensive to score manually.

      However, there are several limitations in the current analysis and validation that affect the strength of the conclusions:

      First, the statistical rigor requires substantial improvement. The analysis of multi-group experiments (e.g., comparing four distinct strains or factorial designs with genotype and training) currently relies on multiple independent Student's t-tests. This approach is statistically invalid for these experimental designs as it inflates the family-wise Type I error rate. To support the claims of strain-specific differences or learning deficits, the data must be analyzed using Analysis of Variance (ANOVA) to properly account for multiple comparisons and to explicitly test for interaction effects between genotype and training conditions.

      Second, the biological validation using $w^{1118}$ and $y^1$ mutants entails a potential confound. The authors attribute the low Courtship Index in these strains to courtship-specific deficits. However, both strains are known to exhibit general locomotor sluggishness (due to visual or pigmentation/behavioral defects). Since "following" behavior is likely a component of the Courtship Index, a reduction in this metric could reflect a general motor deficit rather than a specific lack of reproductive motivation. Without controlling for general locomotion, the interpretation of these behavioral phenotypes remains ambiguous.

      Third, the benchmarking of the system is currently limited to comparisons against manual scoring. Given that the field has largely adopted sophisticated open-source tracking tools (e.g., Ctrax, FlyTracker, JAABA), the utility of DrosoMating would be better contextualized by comparing its performance - in terms of accuracy, speed, or identity maintenance - against these existing automated standards, rather than solely against human observation.

      Finally, the visual presentation of the data hinders the assessment of the system's temporal precision. While the system is designed to capture time-resolved metrics, the results are presented primarily as aggregate bar plots. The absence of behavioral ethograms or raster plots makes it difficult to verify the software's ability to accurately detect specific transitions, such as the exact onset of copulation.

      We sincerely thank the reviewers for their constructive and detailed feedback, which has substantially improved the clarity and rigor of our work. Below is a summary of the major revisions.

      (1) Comparison with existing tools. We added a new main figure (Figure 5) and Table 1 systematically benchmarking DrosoMating against Ctrax and FlyTracker on identical low-quality, high-throughput mating videos. Both conventional tools failed under our recording conditions: Ctrax showed severe segmentation instability and fragmented trajectories, while FlyTracker frequently crashed during feature computation. These results demonstrate that DrosoMating's state-detection approach bypasses the pose-tracking limitations that impair established pipelines. We also revised the Discussion to clarify that DrosoMating is specialized for robust extraction of mating timing metrics from low-quality videos, not a replacement for general-purpose tracking tools.

      (2) Statistical analysis. Following Reviewer #2's recommendation, we re-analyzed Figure 3 using one-way ANOVA with Tukey's HSD for strain comparisons, and Figure 4 using two-way ANOVA with Sidak's post-hoc test for learning assays, including the critical Genotype by Training interaction term.

      (3) New supplementary data. We added Figure S4 showing basal locomotor velocity of single-housed males across all four strains to decouple motor defects from courtship deficits in w1118 and y1 mutants. We also added Figure S3 with behavioral ethograms for individual flies to visually demonstrate the system's temporal resolution and detection accuracy.

      (4) Additional improvements. We standardized statistical reporting across all figure legends, explicitly defined the segmentation threshold parameter s in the Methods, consistently used "Mating Duration (MD)" throughout, standardized MB247-GAL4 labeling, clarified that occluded frames are retained for mating-state detection via merged-contour analysis, and corrected minor errors including duplicate references and unnecessary quotation marks.

      We believe these revisions substantially strengthen the manuscript and clearly position DrosoMating within the existing ecosystem of behavioral analysis tools.

      We remain grateful for the valuable feedback from both reviewers and the editorial team, and we hope the revised version meets the standards for publication in eLife.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It's difficult to assess the utility of this tool in relation to the variety of alternative methods available for the automated scoring of courtship behavior, some of which offer more granular behavioral data than what DrosoMating has been presented to produce. A direct comparison to other approaches (e.g., FlyTracker, DeepLabCut, SLEAP) would be highly valuable for understanding what use cases DrosoMating is best suited for. Ideally, this analysis would include (i) a comparison of the accuracy of scoring courtship metrics and constituent behaviors, (ii) an assessment of the sensitivity of the various software to different experimental conditions, such as lighting and alternative chamber designs, and (iii) testing across Drosophila species, which could help highlight the particular strengths of DrosoMating over more established methods. At a minimum, a table comparing key features, requirements, and capabilities would help position DrosoMating within the existing ecosystem of tools.

      We sincerely appreciate the reviewer’s insightful comments on benchmarking DrosoMating against existing tools for automated courtship behavior analysis. We fully agree that direct comparisons are essential to clarify the unique advantages and ideal use cases of DrosoMating. We have extensively revised the manuscript by adding comparative experiments, a new figure(Figure 5) and comparative table (Table1), and expanded descriptions, as detailed below.

      (1) Direct comparison with conventional tracking tools

      We systematically tested Ctrax and FlyTracker on the same low-quality, high-throughput mating videos used for DrosoMating. The corresponding results are presented in Figure 5.

      Ctrax showed severe segmentation instability, including over-segmentation, under-segmentation, and fragmented trajectories, and failed to stably detect two flies.

      FlyTracker was able to generate background models and showed partially improved segmentation after threshold adjustment, but frequently crashed during feature computation and could not complete the end-to-end analysis pipeline.

      These results confirm that tracking-based tools are vulnerable to low contrast, chamber artifacts, and prolonged male–female overlap, whereas DrosoMating bypasses pose tracking and directly detects mating states.

      The revised manuscript text is as follows (line 265-342):

      "DrosoMating is more compatible with low-quality mating videos than conventional tracking-based pipelines.

      To evaluate whether conventional fly-tracking pipelines could be used as alternative tools for extracting mating-duration metrics, we tested Ctrax and FlyTracker on representative low-quality single-chamber mating videos recorded under our standard high-throughput conditions (Fig. 5). Ctrax is designed to estimate the position and orientation of multiple walking flies while maintaining individual identities over time (Branson et al., 2009) (https://ctrax.sourceforge.net/), whereas FlyTracker aims to track fly pose, including position, orientation, body size, wing and leg positions, and to generate trajectory- and feature-based outputs for downstream behavior analysis (Eyjolfsdottir et al., 2014). Because these programs were not readily compatible with our full high-throughput behavioral recording setup, we first cropped the original videos and tested single-chamber videos containing one male and one female fly.

      Using Ctrax, we observed that target detection was highly variable even within single-chamber videos. In representative frames, Ctrax could occasionally identify the two flies correctly (Fig. 5A). However, imperfect segmentation was frequently observed. In some frames, parts of the fly body were detected as additional targets, resulting in over-segmentation and an apparent increase in the number of detected flies (Fig. 5B, C). Conversely, when the male and female were close to each other or physically overlapped, the two animals were sometimes detected as a single target (Fig. 5D). These examples indicate that Ctrax detection was sensitive to the low contrast and overlapping fly bodies present in our mating videos. We then adjusted the Ctrax detection threshold to improve segmentation quality. Although threshold optimization improved detection in some frames, abnormal detections remained evident across randomly sampled frames (Fig. 5E). For example, some frames still showed incorrect target numbers, and a severe segmentation failure was observed in the lower-right example, where the detected objects did not correspond to two clearly separable flies. Thus, even after parameter optimization, Ctrax did not consistently maintain the expected two-target detection state in single-chamber mating videos.

      The instability of Ctrax detection was also reflected in the trajectory output. In the first 500 frames, the generated trajectories were fragmented into multiple colored track segments rather than two continuous trajectories corresponding to the male and female (Fig. 5F). In addition, some trajectories extended outside the chamber boundary, indicating tracking errors and identity instability. Consistently, frame-by-frame quantification of detected target number showed that the detected object count did not remain stable at the expected value of two flies per chamber (Fig. 5G). Because Ctrax failed to maintain stable two-fly detection and continuous trajectories under these video conditions, its output could not be reliably used for downstream extraction of copulation latency or mating duration.

      We next tested FlyTracker on the same type of cropped single-chamber mating videos. FlyTracker was developed to track multiple flies by estimating body position, orientation, size, wing and leg positions, and by maintaining fly identities across video frames; it also outputs per-frame features such as velocity, facing angle, and wing-angle-related measurements for downstream behavioral analysis (Eyjolfsdottir et al., 2014). In our videos, FlyTracker was able to generate a background model, indicating that the program could recognize the overall imaging field and chamber background (Fig. 5H). However, during calibration and segmentation, the default threshold setting produced inconsistent detection results (Fig. 5I). Although some frames were segmented relatively well under the default threshold (Fig. 5J), these successful examples were not representative of the overall tracking process, and the program frequently terminated with runtime errors during tracking or downstream feature computation. To improve detection stability, we lowered the segmentation threshold. Under this adjusted setting, FlyTracker produced more complete fly masks in representative frames (Fig. 5K), and the diagnostic output showed that two flies were detected in many sampled frames (Fig. 5L). Nevertheless, detection remained unstable in some sampled frames, including frames in which zero or one fly was detected despite the expected two flies per chamber (Fig. 5L). The full pipeline ultimately failed during feature computation, producing a runtime error before complete tracking and feature outputs could be generated (Fig. 5M). Therefore, even after threshold adjustment, FlyTracker could not provide a stable end-to-end workflow for extracting mating-duration metrics from these low-quality mating videos.

      This failure mode is relevant because FlyTracker depends on stable segmentation, identity maintenance, and per-frame feature extraction. In our assay videos, low contrast, chamber-edge artifacts, and prolonged male–female overlap during copulation interfered with these requirements. As a result, FlyTracker could occasionally identify the flies in individual frames, but it did not reliably complete the full analysis pipeline required for downstream behavioral quantification. This limitation is especially important for workflows such as JAABA, which use manually labeled examples to train behavior classifiers but still depend on upstream tracking-derived features. Thus, for our low-quality high-throughput mating recordings, FlyTracker-based analysis was substantially less practical than DrosoMating, which directly outputs mating-related timing metrics without requiring continuous high-fidelity two-fly pose tracking."

      (2) Evaluation of accuracy, robustness, and cross-species potential

      We addressed the three key points requested by the reviewer:

      (i) Accuracy: DrosoMating achieved 98–99% agreement with manual scoring. Under our experimental conditions, neither Ctrax nor FlyTracker completed an end-to-end workflow capable of reliably extracting copulation latency and mating duration. Ctrax produced fragmented trajectories with unstable target counts, and FlyTracker terminated with runtime errors during feature computation.

      (ii) Robustness: DrosoMating is highly robust to low-contrast lighting and common behavioral chamber setups, whereas conventional tools require high-quality videos and fail under fly occlusion.

      (iii) Cross-species testing: We did not perform cross-species validation in this revision. DrosoMating relies on mating state detection rather than species-specific morphology, suggesting potential transferability, although this remains to be experimentally validated. We have therefore revised the text to avoid claiming demonstrated cross-species performance and now describe this as a potential future application.

      The revised manuscript text is as follows (line 572-576):

      "Cross-species testing was not performed in this study. Although DrosoMating's state-detection approach is morphology-agnostic and therefore theoretically applicable across Drosophila species, we have revised the text to avoid claiming demonstrated cross-species performance. Formal validation across diverse species remains a promising future direction."

      We added Table1 to compare key features across DrosoMating, Ctrax, and FlyTracker, including low-quality video compatibility, multi-chamber support, robustness to fly overlap, direct output of mating timing metrics, accuracy.

      (3) Clarification of ideal use cases

      We revised the Discussion to emphasize that DrosoMating is not intended to replace general-purpose pose or tracking tools (e.g., FlyTracker, Ctrax) that provide fine-grained behavioral data.

      Instead, it offers a specialized, robust, and high-throughput workflow for scenarios requiring efficient extraction of mating timing metrics from low-quality videos, where conventional pipelines often fail.

      These revisions greatly improve the clarity and positioning of DrosoMating. We thank the reviewer for this valuable suggestion.

      The revised manuscript text is as follows (line 578-620):

      “Comparative Evaluation with Existing Courtship Analysis Tools

      A major advantage of DrosoMating is its compatibility with low-quality, high-throughput mating videos. Conventional tracking-based tools such as Ctrax and FlyTracker are powerful for trajectory- and pose-based behavioral analysis, but they generally require stable object segmentation, identity maintenance, and reliable feature extraction across frames. Ctrax was designed to estimate the positions and orientations of multiple walking flies while maintaining their identities, whereas FlyTracker aims to track detailed fly pose and generate per-frame behavioral features. Under our recording conditions, these requirements were difficult to satisfy because the videos were low contrast and the male and female frequently overlapped during copulation.

      In our tests, both tools showed limited compatibility with these videos. Even after cropping to single-chamber videos and adjusting detection parameters, Ctrax produced unstable target numbers, fragmented trajectories, and tracking errors. FlyTracker could generate a background model and occasionally segment flies successfully, but detection remained unstable and the full pipeline failed during feature computation. These issues are particularly relevant for copulation latency and mating duration analysis: during copulation, the male and female remain physically coupled for a long period, which makes identity-based tracking difficult. For these timing metrics, it is more important to robustly detect the onset and offset of the mating state than to reconstruct detailed individual trajectories. DrosoMating was purpose-built to address these specific challenges. It operates reliably on lower-quality video streams, requires no complex pre-processing or manual ROI definition, and is optimized for high-throughput multi-chamber analysis. While it does not offer the same level of pose or kinematic detail as other tools, it provides a unique solution for laboratories seeking a simple, fast, and robust pipeline to quantify core reproductive timing metrics—copulation latency (CL), courtship index (CI), and mating duration (MD)—without the overhead of more complex systems (Table 1).

      Although we directly tested only Ctrax and FlyTracker, this limitation may also affect workflows that depend on upstream tracking-derived features. For example, JAABA uses tracking-derived features to train behavior classifiers, and DANCE, a recent Drosophila aggression and courtship pipeline, uses JAABA-based classifiers and lists FlyTracker and JAABA as required software. Thus, our conclusion is not that these tools are generally unsuitable for Drosophila behavior analysis, but that DrosoMating provides a more practical workflow for low-quality, high-throughput videos focused specifically on mating timing.

      While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays.”

      We have also added this clarification to the Table 1 legend to avoid overgeneralizing the limitations of Ctrax and FlyTracker (line 786-790):

      "DrosoMating was designed to extract mating-related timing metrics from high-throughput videos without requiring continuous two-fly identity tracking. Ctrax and FlyTracker were tested on cropped single-chamber videos from the same recording setup. Their limitations described here refer specifically to these low-quality mating videos and should not be interpreted as general limitations of the tools."

      (2) The coarse nature of the metrics captured by DrosoMating may limit the usefulness of this tool for many researchers. Consider integrating DrosoMating with one or more behavioral classifiers (e.g., JAABA) and validating its performance to increase its utility across a wider range of uses. Demonstrating the feasibility of this would substantially increase the tool's appeal to those who need access to more discrete behavioral elements.

      We thank the reviewer for this constructive suggestion. We agree that DrosoMating is currently optimized for rapid, high-throughput quantification of core mating timing metrics (CL, CI, and MD) rather than discrete behavioral classification (e.g., wing extension, circling, or aggressive postures). This reflects a deliberate design trade-off: by prioritizing robust state detection over detailed pose tracking, DrosoMating achieves reliable performance on low-quality videos where conventional identity-based pipelines fail.

      We appreciate the reviewer’s vision for expanding the tool’s utility. While DrosoMating does not presently generate the per-frame kinematic features (position, orientation, wing angles, etc.) required as input for JAABA classifiers, its modular architecture and underlying video-processing framework provide a foundation for future integration. In the revised Discussion, we have clarified that extending DrosoMating to export trajectory-derived features compatible with downstream classifiers such as JAABA represents a promising future direction—one that would broaden its applicability to laboratories requiring granular behavioral elements, without necessitating a complete overhaul of the existing pipeline.

      We hope this clarification addresses the reviewer’s concern and accurately reflects both the current capabilities and future potential of the tool.

      The revised manuscript text is as follows (line 614-620):

      "While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays."

      (3) Please clarify how DrosoMating handles fly identification during mating? I would think mounted flies would be occluded, and if these frames are excluded from analysis, it would be expected to skew mating duration scores. It would be helpful if the authors could discuss these details.

      Occluded frames are excluded from CI calculation because individual courtship actions cannot be reliably assigned during prolonged overlap. However, these frames are not discarded from mating-duration analysis. Instead, prolonged merged contours are used as evidence for copulation-state detection, and the MD timer continues until physical separation:

      (1) When two flies are separate, the system detects two contours; when they overlap during mounting, it detects one merged contour. Our code processes both cases, so occluded frames are retained, not skipped.

      (2) To distinguish a single fly from two overlapping flies, we analyze the shape of the merged contour. Two overlapping flies produce a characteristically different aspect ratio (elongated shape) compared to one fly. This geometric cue is fed into the state classifier to label the frame as "mating."

      (3) Because these frames are classified as mating rather than excluded, the mating duration timer runs continuously through the occlusion period. Thus, mating duration is not artificially shortened.

      The high agreement between DrosoMating and manual scoring for mating duration (within 10 s, no significant difference) confirms that this approach does not introduce bias.

      We revised the manuscript as follow (line 191-193):

      "Occluded frames are excluded from CI calculation but retained for mating-state detection via merged-contour analysis, ensuring continuous MD measurement through copulation."

      (4) Please indicate the statistical tests used in Figures 3, 4, and S1. What methods were used to address multiple hypothesis testing?

      We thank the reviewer for this important comment. For comparisons between two groups, two-sided Student’s t-tests were used. For comparisons among three or more groups, one-way ANOVA followed by post-hoc tests were applied. For two-factor experimental designs, two-way ANOVA was used. We have clearly stated these statistical tests in the Statistical Analysis section and have added this information to the figure legends for Figures 3, 4, and S1 in the revised manuscript.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (5) Line 43: "High resolution video tracking enables...".

      We thank the reviewer for noting this imprecise expression. We have revised the statement to accurately describe that our method uses image-based video analysis to identify courtship and copulation events and quantify their temporal parameters. The description has been corrected in the revised manuscript line 44-46:

      "Our image-based video analysis enables precise identification of courtship and copulation events, as well as quantification of their timing and duration under controlled conditions. "

      (6) Line 51: remove quotation mark.

      We thank the reviewer for the careful correction. The unnecessary quotation mark at Line 51 has been removed in the revised manuscript.

      (7) Line 107: Chen et al, 2024 is listed twice in the references.

      We thank the reviewer for the careful correction. The duplicate reference of Chen et al., 2024 has been removed from the reference list in the revised manuscript.

      (8) Line 242: remove quotation mark.

      We thank the reviewer for the careful correction. The unnecessary quotation mark at Line 242 has been removed in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Please re-analyze the data in Figures 3 and 4 using ANOVA followed by appropriate post-hoc tests (e.g., Tukey's HSD). Specifically, use a One-way ANOVA for strain comparisons in Figure 3 and a Two-way ANOVA for the learning assays in Figure 4. The interaction term (Genotype $\times$ Training) is critical for demonstrating specific learning deficits. Update the "Statistical Analysis" section and figure legends accordingly.

      We appreciate the reviewer’s recommendation. We have re-analyzed the data in Figures 3 and 4 using the suggested ANOVA approaches: one-way ANOVA with Tukey’s HSD for strain comparisons (Figure 3), and two-way ANOVA (including the Genotype × Training interaction term) with Sidak's post-hoc test for learning assays (Figure 4). The updated statistical methods are now described in the Statistical Analysis section and figure legends.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (2) To decouple motor defects from courtship deficits in $w^{1118}$ and $y^1$ mutants, please use your tracking data to calculate and present a "General Locomotion" metric (e.g., average velocity or total distance traveled in the absence of a female).

      We thank the reviewer for this suggestion. We have now added Figure S4 showing basal locomotor velocity of single-housed males across all four strains. As expected, w^1118 and y^1 mutants move more slowly than wild-type controls.

      These data reveal that both motor and sensory factors contribute to the observed courtship phenotypes. Reduced basal locomotion likely limits the males' ability to approach and follow females. However, this generalized hypoactivity is compounded by strain-specific sensory deficits: w<sup>1118</sup> males suffer visual impairment that compromises female detection (Krstic et al., 2013), while y<sup>1</sup> males display altered cuticular hydrocarbons that disrupt pheromonal communication (Drapeau et al., 2006). These sensory defects impair courtship initiation and female recognition independent of locomotor capacity. Thus, the reduced CI, CL, and MD in these mutants reflect the combined effects of slower movement and courtship-specific sensory impairments, rather than motor defects alone.

      We have revised the manuscript to incorporate the velocity data and clarify this interpretation (line 219-226):

      "Notably, reduced basal locomotor activity in w<sup>1118</sup> and y<sup>1</sup> mutants has been well documented in previous studies, independent of courtship behavior (Drapeau et al., 2006; Krstic et al., 2013). Consistent with these reports, our tracking data show that single-housed w<sup>1118</sup> and y<sup>1</sup> males exhibit lower average velocity than Canton-S and Oregon-R controls (Fig. S4A). These general locomotor differences are insufficient to fully explain the observed courtship and mating timing phenotypes, indicating that additional courtship‑related processes contribute to the observed behavioral differences."

      (3) Please expand the discussion or provide a small comparative dataset contrasting DrosoMating with established tools like JAABA. Explain the specific advantages of your pipeline (e.g., cost, simplicity, focus on CL/MD) to justify its adoption over these alternatives.

      We sincerely appreciate the reviewer’s insightful comments on benchmarking DrosoMating against existing tools for automated courtship behavior analysis. We fully agree that direct comparisons are essential to clarify the unique advantages and ideal use cases of DrosoMating. We have extensively revised the manuscript by adding comparative experiments, a new figure (Figure 5) and comparative table (Table 1), and expanded descriptions, as detailed below.

      (1) Direct comparison with conventional tracking tools

      We systematically tested Ctrax and FlyTracker on the same low-quality, high-throughput mating videos used for DrosoMating. The corresponding results are presented in Figure 5.

      Ctrax showed severe segmentation instability, including over-segmentation, under-segmentation, and fragmented trajectories, and failed to stably detect two flies.

      FlyTracker was able to generate background models and showed partially improved segmentation after threshold adjustment, but frequently crashed during feature computation and could not complete the end-to-end analysis pipeline.

      These results confirm that tracking-based tools are vulnerable to low contrast, chamber artifacts, and prolonged male–female overlap, whereas DrosoMating bypasses pose tracking and directly detects mating states.

      The revised manuscript text is as follows (line 265-342):

      " DrosoMating is more compatible with low-quality mating videos than conventional tracking-based pipelines

      To evaluate whether conventional fly-tracking pipelines could be used as alternative tools for extracting mating-duration metrics, we tested Ctrax and FlyTracker on representative low-quality single-chamber mating videos recorded under our standard high-throughput conditions (Fig. 5). Ctrax is designed to estimate the position and orientation of multiple walking flies while maintaining individual identities over time (Branson et al., 2009) (https://ctrax.sourceforge.net/), whereas FlyTracker aims to track fly pose, including position, orientation, body size, wing and leg positions, and to generate trajectory- and feature-based outputs for downstream behavior analysis (Eyjolfsdottir et al., 2014). Because these programs were not readily compatible with our full high-throughput behavioral recording setup, we first cropped the original videos and tested single-chamber videos containing one male and one female fly.

      Using Ctrax, we observed that target detection was highly variable even within single-chamber videos. In representative frames, Ctrax could occasionally identify the two flies correctly (Fig. 5A). However, imperfect segmentation was frequently observed. In some frames, parts of the fly body were detected as additional targets, resulting in over-segmentation and an apparent increase in the number of detected flies (Fig. 5B, C). Conversely, when the male and female were close to each other or physically overlapped, the two animals were sometimes detected as a single target (Fig. 5D). These examples indicate that Ctrax detection was sensitive to the low contrast and overlapping fly bodies present in our mating videos. We then adjusted the Ctrax detection threshold to improve segmentation quality. Although threshold optimization improved detection in some frames, abnormal detections remained evident across randomly sampled frames (Fig. 5E). For example, some frames still showed incorrect target numbers, and a severe segmentation failure was observed in the lower-right example, where the detected objects did not correspond to two clearly separable flies. Thus, even after parameter optimization, Ctrax did not consistently maintain the expected two-target detection state in single-chamber mating videos.

      The instability of Ctrax detection was also reflected in the trajectory output. In the first 500 frames, the generated trajectories were fragmented into multiple colored track segments rather than two continuous trajectories corresponding to the male and female (Fig. 5F). In addition, some trajectories extended outside the chamber boundary, indicating tracking errors and identity instability. Consistently, frame-by-frame quantification of detected target number showed that the detected object count did not remain stable at the expected value of two flies per chamber (Fig. 5G). Because Ctrax failed to maintain stable two-fly detection and continuous trajectories under these video conditions, its output could not be reliably used for downstream extraction of copulation latency or mating duration.

      We next tested FlyTracker on the same type of cropped single-chamber mating videos. FlyTracker was developed to track multiple flies by estimating body position, orientation, size, wing and leg positions, and by maintaining fly identities across video frames; it also outputs per-frame features such as velocity, facing angle, and wing-angle-related measurements for downstream behavioral analysis (Eyjolfsdottir et al., 2014). In our videos, FlyTracker was able to generate a background model, indicating that the program could recognize the overall imaging field and chamber background (Fig. 5H). However, during calibration and segmentation, the default threshold setting produced inconsistent detection results (Fig. 5I). Although some frames were segmented relatively well under the default threshold (Fig. 5J), these successful examples were not representative of the overall tracking process, and the program frequently terminated with runtime errors during tracking or downstream feature computation. To improve detection stability, we lowered the segmentation threshold. Under this adjusted setting, FlyTracker produced more complete fly masks in representative frames (Fig. 5K), and the diagnostic output showed that two flies were detected in many sampled frames (Fig. 5L). Nevertheless, detection remained unstable in some sampled frames, including frames in which zero or one fly was detected despite the expected two flies per chamber (Fig. 5L). The full pipeline ultimately failed during feature computation, producing a runtime error before complete tracking and feature outputs could be generated (Fig. 5M). Therefore, even after threshold adjustment, FlyTracker could not provide a stable end-to-end workflow for extracting mating-duration metrics from these low-quality mating videos.

      This failure mode is relevant because FlyTracker depends on stable segmentation, identity maintenance, and per-frame feature extraction. In our assay videos, low contrast, chamber-edge artifacts, and prolonged male–female overlap during copulation interfered with these requirements. As a result, FlyTracker could occasionally identify the flies in individual frames, but it did not reliably complete the full analysis pipeline required for downstream behavioral quantification. This limitation is especially important for workflows such as JAABA, which use manually labeled examples to train behavior classifiers but still depend on upstream tracking-derived features. Thus, for our low-quality high-throughput mating recordings, FlyTracker-based analysis was substantially less practical than DrosoMating, which directly outputs mating-related timing metrics without requiring continuous high-fidelity two-fly pose tracking."

      (2) Clarification of ideal use cases

      We revised the Discussion to emphasize that DrosoMating is not intended to replace general-purpose pose or tracking tools (e.g., FlyTracker, Ctrax) that provide fine-grained behavioral data.

      Instead, it offers a specialized, robust, and high-throughput workflow for scenarios requiring efficient extraction of mating timing metrics from low-quality videos, where conventional pipelines often fail.

      These revisions greatly improve the clarity and positioning of DrosoMating. We thank the reviewer for this valuable suggestion.

      The revised manuscript text is as follows (line 578-620):

      “Comparative Evaluation with Existing Courtship Analysis Tools

      A major advantage of DrosoMating is its compatibility with low-quality, high-throughput mating videos. Conventional tracking-based tools such as Ctrax and FlyTracker are powerful for trajectory- and pose-based behavioral analysis, but they generally require stable object segmentation, identity maintenance, and reliable feature extraction across frames. Ctrax was designed to estimate the positions and orientations of multiple walking flies while maintaining their identities, whereas FlyTracker aims to track detailed fly pose and generate per-frame behavioural features. Under our recording conditions, these requirements were difficult to satisfy because the videos were low contrast and the male and female frequently overlapped during copulation.

      In our tests, both tools showed limited compatibility with these videos. Even after cropping to single-chamber videos and adjusting detection parameters, Ctrax produced unstable target numbers, fragmented trajectories, and tracking errors. FlyTracker could generate a background model and occasionally segment flies successfully, but detection remained unstable and the full pipeline failed during feature computation. These issues are particularly relevant for copulation latency and mating duration analysis: during copulation, the male and female remain physically coupled for a long period, which makes identity-based tracking difficult. For these timing metrics, it is more important to robustly detect the onset and offset of the mating state than to reconstruct detailed individual trajectories. DrosoMating was purpose-built to address these specific challenges. It operates reliably on lower-quality video streams, requires no complex pre-processing or manual ROI definition, and is optimized for high-throughput multi-chamber analysis. While it does not offer the same level of pose or kinematic detail as other tools, it provides a unique solution for laboratories seeking a simple, fast, and robust pipeline to quantify core reproductive timing metrics—copulation latency (CL), courtship index (CI), and mating duration (MD)—without the overhead of more complex systems (Table 1).

      Although we directly tested only Ctrax and FlyTracker, this limitation may also affect workflows that depend on upstream tracking-derived features. For example, JAABA uses tracking-derived features to train behavior classifiers, and DANCE, a recent Drosophila aggression and courtship pipeline, uses JAABA-based classifiers and lists FlyTracker and JAABA as required software. Thus, our conclusion is not that these tools are generally unsuitable for Drosophila behavior analysis, but that DrosoMating provides a more practical workflow for low-quality, high-throughput videos focused specifically on mating timing.

      While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays.”

      We have added this clarification to the Table 1 legend to avoid overgeneralizing the limitations of Ctrax and FlyTracker (line 786-790):

      "DrosoMating was designed to extract mating-related timing metrics from high-throughput videos without requiring continuous two-fly identity tracking. Ctrax and FlyTracker were tested on cropped single-chamber videos from the same recording setup. Their limitations described here refer specifically to these low-quality mating videos and should not be interpreted as general limitations of the tools."

      (4) Complement the aggregate bar plots with behavioral ethograms or raster plots for representative individual flies. Color-code these plots for specific states (resting, following, courting, copulating) to visually demonstrate the system's temporal resolution and detection accuracy.

      We thank the reviewer for this valuable suggestion. To address this comment, we have added a new supplementary figure (Figure S3) that presents behavioral ethograms for individual flies, directly complementing the aggregate bar plots in the main text. The figure displays the full temporal progression of courtship and copulation behaviors for all wells that exhibited successful mating. As requested, behaviors are color-coded (orange: courting, red: copulating) to clearly delineate different states. These ethograms visually demonstrate the system’s ability to resolve behavioral transitions with high temporal precision, confirming the accuracy of our automated detection of courtship initiation, duration, and copulation events. This addition provides critical individual-level validation that supports the aggregate statistical results presented in the main text.

      (5) Standardize the use of estimation statistics (DBMs). If used in Figure 3, they should also be applied to Figure 4, with appropriate statistical comparisons between groups.

      We thank the reviewer for this suggestion. We have standardized the statistical reporting format across all figure legends, with each legend explicitly stating the test used, the post hoc method (where applicable), and significance thresholds.

      Our approach is as follows:

      Fig. 2 and Fig.3D-I (two-group comparison, manual vs. automated scoring): DBM + Student's t-test.

      Fig. 3A-C and Fig. S1B-C, E-F (multi-group comparison, 4 strains or 2 rearing conditions): One-way ANOVA + Tukey's HSD.

      Fig. 4B-D and F (two-factor design, genotype × training): Two-way ANOVA + Sidak's post hoc.

      We have also added a sentence to the Methods clarifying that estimation statistics (DBM) are used for single two-group contrasts, while ANOVA-based approaches are used for multi-group or multi-factor designs. The statistical methods are now consistently documented in both the Methods section and the corresponding figure legends. We hope this clarification addresses the reviewer's concern.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (6) Fix inconsistent labeling (e.g., "MB-247" vs. "MB247") and redundant axis labels.

      Thank you for pointing this out. We have now standardized all labels to "MB247-GAL4" throughout the text and figures.

      We have also removed redundant axis labels from multi-panel figures. And redundant DBM and metric labels have been streamlined in Figures 2 and 4. All statistical tests and metric definitions are now fully described in the corresponding figure legends rather than being repeated on each sub-panel.

      (7) Significantly increase the size of data panels to make individual data points and error bars legible.

      Thank you for this suggestion. We have standardized the figure formatting and simplified the panels by removing redundant axis labels and consolidating descriptive details into the figure legends. We believe the current panel sizes, combined with these clarifications, provide sufficient legibility for both data points and error bars in the final high-resolution PDF.

      Minor Corrections:

      (1) Abstract: Rephrase "lack of certain timing-related behavioral repertoires" (Line 108) to "lack of precise quantification for temporal parameters of post-copulatory behavior."

      Thank you for the suggested rephrasing. We have updated the sentence accordingly.

      (2) Define the physical parameter "s" (e.g., is it a pixel threshold?) to ensure reproducibility.

      Thank you for this important suggestion. We have now explicitly defined the parameter s in the Methods section. Briefly, s is the grayscale intensity threshold (0–255, 8-bit) used for binary segmentation of flies from the background. It is automatically calculated as the maximum grayscale value of three user-selected reference flies plus an offset of 28, and can be manually adjusted to accommodate varying illumination conditions.

      The revised manuscript text is as follows (line 499-503):

      “Note on the s value: The parameter s represents the grayscale intensity threshold (range: 0–255 for 8-bit images) used for binary segmentation of flies from the background. It is automatically calculated as the maximum grayscale value at the three reference fly positions plus an offset of 28, and can be manually adjusted to accommodate varying illumination conditions.”

      (3) Figure 1 Legend: Clarify the "clockwise selection" of pillars and their relation to well numbering.

      Thank you for this suggestion. We have revised the Figure 1 legend to clarify that:

      (1) The four pillars are selected in clockwise order starting from the top-left corner to define the chamber corners for perspective transformation (homography), which corrects for camera angle and standardizes the field of view.

      (2) Well numbering is independent of pillar selection order. After automated perspective correction, wells are numbered sequentially in a left-to-right, top-to-bottom order (1–36) based on the standardized chamber layout.

      The updated Figure 1 legend now reads:

      "Columns (pillars) should be selected in clockwise order starting from the top-left corner to define the chamber boundaries for perspective transformation (Fig. 1A, lower). Well numbering (1–36) follows a left-to-right, top-to-bottom sequence after perspective correction and is independent of pillar selection order."

      (4) Add the citation for Eastwood and Burnet (1977) regarding Courtship Latency.

      Thank you for pointing this out. We have added the citation Eastwood and Burnet (1977) to the Introduction where Courtship Latency is first defined

      (5) Select one term ("Mating Duration" or "Copulation Duration") and use it consistently throughout the text and figures.

      Thank you for this suggestion. We have now standardized the terminology throughout the manuscript and figures. "Mating Duration" (MD) is used consistently in all instances where "Copulation Duration" previously appeared. The abbreviation MD has been retained for consistency with existing figure labels

      References

      Branson, K., Robie, A.A., Bender, J., Perona, P., & Dickinson, M.H. (2009). High-throughput ethomics in large groups of Drosophila. Nature Methods, 6(6), 451–458.

      Claridge-Chang, A., & Assam, P.N. (2016). Estimation statistics should replace significance testing. Nature Methods, 13(2), 108–109.

      Drapeau, M.D., Cyran, S.A., Viering, M.M., Geyer, P.K., & Long, A.D. (2006). A cis-regulatory sequence within the yellow locus of Drosophila melanogaster required for normal male mating success. Genetics, 172(2), 1009–1030.

      Eastwood, L., & Burnet, B. (1977). Courtship latency in male Drosophila melanogaster. Behavior Genetics, 7(3), 359–372.

      Eyjolfsdottir, E., Branson, S., Burgos-Artizzu, X.P., Hoopfer, E.D., Schor, J., Anderson, D.J., & Perona, P. (2014). Detecting social actions of fruit flies. Lecture Notes in Computer Science, 8692, 772–787.

      Gil-Martí, B., Barredo, C.G., Pina-Flores, S., Poza-Rodriguez, A., Treves, G., Rodriguez-Navas, C., Camacho, L., Pérez-Serna, A., Jimenez, I., Brazales, L., Fernandez, J., & Martin, F.A. (2023). A simplified courtship conditioning protocol to test learning and memory in Drosophila. STAR Protocols, 4(1), 101572.

      Greenspan, R.J., & Ferveur, J.F. (2000). Courtship in Drosophila. Annual Review of Genetics, 34, 205–232.

      Hall, J.C. (1994). The mating of a fly. Science, 264(5163), 1702–1714.

      Kabra, M., Robie, A.A., Rivera-Alba, M., Branson, S., & Branson, K. (2013). JAABA: interactive machine learning for automatic annotation of animal behavior. Nature Methods, 10(1), 64–67.

      Kitamoto, T. (2001). Conditional modification of behavior in Drosophila by targeted expression of a temperature-sensitive shibire allele in defined neurons. Journal of Neurobiology, 47(2), 81–92.

      Krstic, D., Boll, W., & Noll, M. (2013). Influence of the White locus on the courtship behavior of Drosophila males. PLoS ONE, 8(9), e77904.

      Levin, L.R., Han, P.L., Hwang, P.M., Feinstein, P.G., Davis, R.L., & Reed, R.R. (1992). The Drosophila learning and memory gene rutabaga encodes a Ca2+/calmodulin-responsive adenylyl cyclase. Cell, 68(3), 479–489.

      Pavlou, H.J., & Goodwin, S.F. (2013). Courtship behavior in Drosophila melanogaster: towards a ‘courtship connectome’. Current Opinion in Neurobiology, 23(1), 76–83.

      Yapici, N., Kim, Y.J., Ribeiro, C., & Dickson, B.J. (2008). A receptor that mediates the post-mating switch in Drosophila reproductive behaviour. Nature, 451(7176), 33–37.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Garcia-Alcala, Kratz and Cluzel investigate to what extent our understanding of bacterial physiology in bulk experiments can be applied to single-cell observations. They find that intrinsic noise may be powerful enough to even inverse the trends found in the bulk. The authors hypothesize that the asymmetric distribution of ribosomes to daughter cells during cell division plays the dominant role in the intrinsic noise and is able to generate the observed phenomenon. They do not show it directly, but the data and its agreement with the model are sufficient to support this claim.

      Strengths:

      The experimental part is convincing: the positive correlation between the elongation rate and promoter activity of unnecessary protein is clear, as well as the negative correlation between the mean values while changing the promoter strength. This was demonstrated in both rich and poor media. The causality between the growth rate and the promoter activity was shown using the negative lag time of the cross-correlation function. A simple, reasonable model accounts well for the data. This paper demonstrates an interesting phenomenon and provides a plausible theory for it, advancing our understanding of bacterial physiology on the single-cell level.

      Weaknesses:

      (1) Mean-reversion timescales were assumed to be longer than the simulation time and much longer than the cell cycle time. It is not clear whether the results are robust in case mean reversion timescales become of the order of the cell-cycle or smaller. If not, is there an argument for such practically infinite reversion timescales?

      Due to an error in the simulation code, the incorrect mean-reversion timescales were reported in the manuscript and have now been corrected. Instead of 1000 h and 100 h for 𝜏<sub>R</sub> and 𝜏<sub>U</sub>, respectively, they are 6.25 h and 0.0625 h, which are both much shorter than the total simulation time (75 h). The error had no effect on simulation results and was purely a timescale conversion mistake. We appreciate the referee’s careful review of the manuscript which allowed us to catch this error.

      Given the correct values, 𝜏<sub>U</sub> is well within the cell-cycle time, as is expected since we assume the source of noise in unnecessary protein expression is from stochastic gene expression. In contrast, 𝜏<sub>U</sub> is still significantly longer than the cell-cycle time and needs to be in order to generate the experimentally observed behavior (i.e., the increase in growth rate with an increase in unnecessary protein expression at the single-cell level). This is consistent with our biological hypothesis that some cells inherit ribosomal surpluses from their mothers which enable bursts in protein production. 𝜏<sub>U</sub> less than the cell-cycle time would correspond to a case where ribosomal composition quickly decays to the population average, meaning that daughter cells would never have time to capitalize on the benefit of receiving a ribosome surplus. Furthermore, 𝜏<sub>U</sub> being longer than the cell-cycle is biologically justifiable as proteins such as ribosomes are passed from mother to daughter at division, thus allowing for memory to persist over longer timescales than a single generation.

      (2) It is not easy to understand the simulation part unless one reads Ref [14]. k(t) is assumed Equation (1) from Reference [14]? Is it crucial that the ribosome noise appears only at the division? The ribosome noise strength σ<sub>R</sub> =0.06 - is it lower or higher than the naively expected binomial division? Also, a more intuitive explanation of the Simpson paradox would help the reader.

      𝜅(𝑡) is indeed Eq. (1) from Ref. [14]. To make the computational results clearer, the methods section has been updated to include the full set of equations used to perform the simulations, and the code used to produce the computational figures is now on GitHub. The relative standard deviation expected by modeling ribosome distribution at division by a binomial distribution with equal probability of being inherited by either daughter cell is in the range of 1-3% (0.01-0.03), assuming N~10<sup>3</sup>-10<sup>4</sup> ribosomes. Thus, 0.06 is reasonable as it is in the same order of magnitude as what is predicted by binomial division. Furthermore, if clustering is present as already demonstrated in [39, 40, 43], we would expect the relative standard deviation to increase as clustering reduces the effective number of proteins which are distributed between daughters.

      (3) It would be useful for the reader to see the raw data and not only the filtered one to appreciate the measurement noise level.

      We have included the direct calculations of cell size and promoter activity in the time-lapse plots in Fig. 1. In addition, we included Fig. S2, which shows typical time traces of Class-2 activity and elongation rate, displaying both the direct measurements and the corresponding smoothed traces for the flagellar reporter strains.

      (4) Negative lag time of the cross-correlation function is visible, but consider adding a statistical test for it.

      We have included a statistical analysis of the lags of maximum correlation of elongation rate and activity for the flagellar reporter strains in Fig. S10. For each strain, we now show an overlay of the cross-correlation functions for all lineages and the distribution of the lag corresponding to the maximum correlation. In addition, we have added a section in Materials and Methods describing in detail how the cross-correlation and the strain-averaged lag were computed.

      (5) Can you make similar cross-correlation plots using the model? Can you infer by using it, whether the data agrees better with the assumption that ribosomal noise appears only at division or continuous fluctuations during the cell cycle?

      The model in its current form is unable to capture the observed cross-correlation (the correlation is sharply peaked at zero). This is because the model coarse-grains transcription and translation into one process of protein production and thus lacks any delay or memory mechanisms which would make a significant positive or negative correlation.

      Both continuous fluctuations and noise from division could in principle contribute to our observation of Simpson’s paradox. We are unable to use the model in its present form to dissect the contributions from both mechanisms. However, our model simulations clearly show that noise at division by itself is sufficient to explain the observed effect with parameter values that are biologically plausible. As Chao et al., [40] has previously shown experimentally that a significant source of ribosomal noise comes from unequal distribution at division, which is the hypothesis we retained for the model.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Garcia-Alcala et al. reports an interesting paradox: the cost of gene expression slows the population-average growth rate, whereas at the single-cell level, expression levels from these genes positively correlate with the growth rate. The effect is observed in the expression of flagellar genes and a gene under a synthetic promoter in E. coli. The findings are explained by the inheritance of growth factors, including ribosomes, during asymmetric division.

      Strengths:

      (1) The manuscript adds strength to an emerging body of literature showing that the population-level bacterial growth laws do not match correlations based on single-cell data. The evidence presented here is more striking than in previous works (such as Pavlou et al., Nat. Commun. 2025), as the trends in population-level data and single-cell data are reversed.

      (2) A relatively simple model correctly explains the trends in the data.

      Weaknesses:

      (1) It is not clear whether flagellar proteins are expressed proportionally to the reporter signal. Furthermore, it is questionable if E. coli bacteria in the mother machine channels are flagellated. If they are, they could potentially swim out of the channels, which is not the case when they do not carry the MotA E98K mutation. The authors should provide some evidence that E. coli expresses the actual filament proteins in the channels.

      We agree that it is important to demonstrate that our reporter reflects the production of functional flagellar structures under our experimental conditions. To this end, we first tested the swimming capabilities of our strains before introducing the MotA E98K mutation, using a standard soft‑agar motility assay. For strains in which only the Class‑1 promoter was modified (Pro2, Pro4, Pro5), after 12 h of inoculation, the diameters of the swarming rings followed the expected order based on promoter strength: Pro2 showed the smallest ring, followed by Pro4, WT, and Pro5. In contrast, control strains carrying Pro4 together with either MotA E98K (non‑rotating motors) or ΔfliC (no flagellin filament) did not form rings, consistent with their inability to swim. We also included an MG1655 strain carrying an IS5 insertion upstream of the Class‑1 promoter, which is known to enhance flagellar expression [56]; this strain showed a larger ring, as expected. A qualitative summary of ring diameters for all strains is provided in Author response table 1. We have included the figure and section “Experimental validation of functional flagella expression in reporter strains” on Supplementary Material.

      Author response table 1.

      We also tested whether cells assemble functional flagella on the mother‑machine. We compared MG1655 WT and MG1655 carrying the MotA E98K mutation under identical microfluidic growth conditions. Many WT cells left the channels during the experiment (∼35% of lineages over ~20 h after the onset of exponential growth inside the device), consistent with active swimming, whereas the non‑motile MotA E98K strain did not leave the channels. Because the only difference between these two strains is the MotA E98K mutation, which disables motor rotation but not flagellar assembly, this result indicates that (i) cells do express functional filaments and motors in the mother‑machine environment, and (ii) the strain used for our main experiments is immobilized by MotA E98K.

      Both experiments are included in the Supplementary Information section “Experimental validation of functional flagella expression in reporter strains” and Fig. S3.

      (2) It is unclear what fraction of the total proteome mVenus represents in different measurements. Some quantification is needed (for example, using the Coomassie staining). Using f_U as high as 14.4% in simulations is questionable.

      We agree that we do not currently know the exact fraction of the proteome occupied by mVenus in our experiments, as we only quantified fluorescence and did not perform Coomassie staining or proteomics. The primary goal of our simulations is not to reproduce exact experimental conditions, but to illustrate how unequal ribosome partitioning affects daughter cells across a range of protein synthesis burdens. Consequently, our use of values up to F<sub>U</sub> = 14.4% in the simulations was intended as an exploratory upper range rather than as a direct estimate of the experimental condition.

      To put our experimental burden in context, we compare our growth-rate reduction to a well– characterized high-burden case in the literature. In the study by T. Hwa’s group [6], overexpression of β–galactosidase such that it constituted approximately 27% of the total proteome led to a 67% reduction in growth rate for E. coli growing in a medium similar to ours (differing only in the carbon source: glucose in their case, glycerol in ours). In our system, overexpression of mVenus from a plasmid result in a substantially smaller growth–rate reduction of 9% relative to the non–expressing control, and this is observed on the poorer carbon source (glycerol), under which burden effects are typically less pronounced [6].

      Although we cannot convert our fluorescence measurements into an exact proteome fraction, the much smaller growth defect compared to the 27% β‑galactosidase case strongly suggests that mVenus does not approach such an extreme fraction of the proteome under our conditions. Under the simplifying assumption that the qualitative relationship between unnecessary‑protein fraction and growth‑rate reduction is similar in the two systems, our data are therefore consistent with a modest fraction of unnecessary protein and make it unlikely that mVenus reaches very high fractions such as 27%. This gives us confidence that exploring F<sub>U</sub> values up to 14.4% in the simulations represents a conservative upper range relative to our experimental burden, rather than an underestimate.

      (3) The data from the MC4100 strain does not directly match the trends of MG1655. The justification for filtering out the low-frequency components of MC4100 is not particularly convincing. It appears unlikely that ribosomes or other growth factors partition significantly differently in the MC4100 strain than in the MG1655 strain. Further discussion and a plot similar to Figure 1 (Left) for this strain are warranted.

      We thank the reviewer for this helpful suggestion. We agree that the filtering analysis from the initial manuscript was not clear enough to be used as robust supplementary information, and we therefore removed it. Instead, we carried out with the ‘unfiltered’ MC4100 strain, the same single-cell analyses as with MG1655, including the binning analysis of instantaneous elongation rate versus Class-2 promoter activity, and a cross-correlation analysis between such measurements.

      Both analyses did not reveal a significant correlation between growth and flagellar promoter activity in MC4100. We now present these results in the revised manuscript (Fig. S15), where we explicitly show the lack of association in a plot directly comparable to Fig. 1.

      While ribosomes partition mechanism is likely to be the same between these two strains, MC4100 is known to exhibit very long oscillations of growth rate over 10 generations, which are absent in MG1655.

      We now present MC4100 as an explicit counterexample to highlight that the positive correlation between short timescale growth fluctuations and flagellar expression observed in MG1655 is not universal across all E. coli strains, especially in strains like MC4100 whose growth rate fluctuations are dominated by long timescales, much longer than the division time. In MC4100 these slow modes are largely decoupled from flagellar gene regulation. We have also revised the text to clarify this point.

      (4) The model needs to be described in more detail. A closed set of equations that have been simulated must be presented, along with all values of the model parameters and their sources. The authors should consider depositing their code on GitHub or another publicly accessible repository.

      The methods section has now been updated to include the full set of equations used to perform the simulations along with all parameter values and their sources. Additionally, the code used to produce the computational figures is now on GitHub. For a full derivation and biological justification of each model component, we still refer readers to ref [14] where this model was first published.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The first paragraph of Results still belongs to the Introduction Section, I feel.

      Thank you for the recommendation, we agreed with the change.

      (2) Figure 1, left, is after some filter (Savitzky-Golay)? It might be useful to see the raw points.

      We have included the direct calculations of cell size and promoter activity in the time-lapse plots in Fig. 1. In addition, we included Fig. S2, which shows typical time traces of Class-2 activity and elongation rate, displaying both the raw (direct) measurements and the corresponding smoothed traces for the flagellar reporter strains.

      Reviewer #2 (Recommendations for the authors):

      (1) Is there a statistically significant positive slope in single-cell data of the elongation rate as a function of Class-2 activity? There should be an analysis of statistical significance for the slopes in Figure 1 (Right) and Figure 4A, including both population-average and single-cell data.

      Thank you for your recommendation. We have now added statistical analyses of the correlations between activity and elongation rate at both the single-cell and population levels. The results for the flagellar strains are presented in Fig. S5, S6, and S7. Also, the corresponding analysis for constitutive Venus expression is shown in Fig. S16.

      (2) How does FlhC-YFP data compare with the CFP signal in the first set of measurements? What would Figure 1, Left look like for this signal?

      At the single‑cell level, the relationship between Class‑1 promoter activity and elongation rate within each strain is still positive, but clearly weaker than for Class‑2, as shown in Fig. S8. When we bin the data, we can observe that the binned averages don’t display a clear positive trend, even though the Pearson correlation for each flagellar reporter strain is still positive. We think this is because Class‑1 controls a much smaller part of the proteome (it only encodes the two subunits of FlhDC) while Class‑2 promoters drive many structural and export proteins. So, changes in Class‑2 activity more directly reflect shifts in global translational capacity and are more tightly linked to growth, whereas Class‑1 activity adds only a small translational load.

      (3) The information from the ER-Activity cross-correlation functions is interesting but has not been interpreted or compared with the model. Which signal precedes the other? What can explain the observed lag time on the order of Tdiv? Why do some cross-correlations show negative values in Fig. S9 while others are positive (as they should)? Can the model explain the experimentally observed cross-correlation function?

      In our cross‑correlation analysis, a negative lag at the maximum means that fluctuations in elongation rate (ER) precede fluctuations in promoter activity (A) by that lag. We now describe this explicitly in Materials and Methods and quantify lag distributions for all reporter strains in Fig. S10.

      In Fig. 13 (previously Fig. 9), we mainly observe positive correlations between ER and PA with negative or near zero peak lags, i.e., ER tends to lead PA by a lag by about a division time. The reason governing the negative lags is not immediately clear, but it is consistent with a resource driven mechanism: cells that inherit more growth factors (e.g. ribosomes) at division, use them to prioritize housekeeping processes and thus increase growth rate first, and only subsequently use these extra resources to increase flagellar promoter activity.

      As for Class–1 activity, the correlation with growth is much weaker as it was already observed in Kim et al. Averaging the cross–correlations across all lineages yields a modest peak (mean correlation 0.10, SD 0.12; see Author response image 1) at a small positive lag of 0.5 h (while average division time is T<sub>div</sub> ~ 1.7h). This indicates a weak but real positive correlation between Class 1 activity and ER at short lags. However, the lags of the individual maxima are widely distributed (SD ≈ 9 h; mean −1.8 h, median −0.28 h, mode ≈ 0), with only ~53% of the cells showing a negative time lag. Thus, delays are roughly symmetrically spread around zero with only a slight negative bias.

      Author response image 1.

      Left: cross−correlation between elongation rate and Class−1 activity for each lineage (N = 96, colored lines), and their average as a function of lag (black line). Right: distribution of the lags at which each lineage’s cross−correlation attains its maximum. The dashed line indicates zero lag, and the full line marks the lag of the peak of the mean cross correlation.

      An in-depth analysis about the sign and magnitude of the lag would require measurements from a broader set of promoters. The lag is likely to depend on what genes the promoter controls (e.g. stress response, housekeeping, or large structural modules), on its strength and regulation, and on growth conditions.

      The model in its current form is unable to capture the observed cross-correlation (the correlation is sharply peaked at zero). This is because the model coarse-grains transcription and translation into one single process of protein production and thus lacks any delay or memory mechanism that could generate a phase shift.

      Nonetheless, this memory-free formulation shows that stochastic redistribution of growth factors at division is by itself sufficient to generate the Simpson’s paradox behavior. Capturing the experimentally observed lag time would require including explicit transcription/translation delays and additional regulatory dynamics between housekeeping and flagellar genes whose information we do not have.

      (4) Figure 3, Left - it is not clear what this plot shows. Red and blue are scattered over the whole plot. How have daughter 1 and daughter 2 been assigned? Perhaps choosing one of the daughters with a higher growth rate and then plotting the data could reveal some trends.

      We agree that the original left panel of Fig. 3 was difficult to follow. In the original version, the daughter labels were assigned as “Daughter–1” for the cell at the closed end of the channel and “Daughter–2” for the cell closest to the open side. After discussing this with Dr Camilla Ulla Rang (whom we now acknowledge in the manuscript), we relabeled the daughters as “new–pole daughter” and “old–pole daughter,” following the convention used in studies of aging and ribosome distribution in E. coli.

      We now show the class–2 promoter activity comparison between new pole vs old pole, and in a separate panel, we plot the mean ratios of elongation rate, and class–1/class–2 promoter activities between the new–pole and old–pole daughters. These ratios show that the new–pole daughter tends to have both greater growth rate and flagellar gene activity than the old–pole daughter. Importantly, this new plot is in line with Rang and colleagues who demonstrated that new–pole daughters have greater ribosome density and faster growth rates than their old–pole sisters. Together, these results further support our hypothesis that excesses of ribosomes inherited at division underlies the observed growth boosts.

      (5) Flagellar activity -> activity of flagellar gene synthesis (presumably no flagellar activity in these cells).

      We corrected the terms used to reference the flagellar gene activity.

      (6) Page 7: "The strain MC4100, known to exhibit slow, long period oscillations in growth ..." - some reference is- needed here.

      We placed the reference some lines after, as such paper also includes the information of MG1655 short-term oscillations. Now the reference is [44]: Tanouchi, Y., et al., A noisy linear map underlies oscillations in cell size and gene expression in bacteria. Nature, 2015

      (7) Page 7: Figure S11B - is Figure S11C perhaps meant?

      The correct panel indeed was Fig. S11C, and we have now corrected and updated the figure label and text accordingly.

      (8) Page 11: Savitsky-Golay filter.

      We have corrected it, thank you.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reply to the reviewers:

      We thank the three reviewers for their thoughtful and constructive comments, which have helped us to better explain our arguments and strengthen the manuscript. In this revision, we have clarified several points raised in the initial review and added new evidence that we believe further supports our conclusions.

      Point-by-point responses:

      Reviewer #1:

      _Additional in vitro experiment using artificial R-loop structures as substrate for lambda-exonuclease could prove the efficacy of the nuclease to digest through R-loop structures._

      We agree this experiment would be the cleanest way to demonstrate λ-exo's behavior on an R-loop substrate under defined biochemical conditions, and we would have liked to include it. However, this work originated in a laboratory that has since been closed due retirement, and further wet-lab experimentation is not feasible for this revision.

      Nonetheless, the manuscript already provides strong convergent evidence for RNA:DNA hybrid-mediated obstruction, short of this direct reconstitution. First, in vivo colocalization analysis shows that tSNS-seq (but not iSNS-seq) signal is enriched at S1-DRIP-seq and S9.6 ChIP-seq sites in an RNase H-sensitive manner, directly implicating RNA:DNA hybrids at the genomic loci where tSNS-seq peaks form. The reciprocal is also true, RNA:DNA hybrid signal is enriched at tSNS-seq sites, but not iSNS-seq sites (nor at known origins), in an RNase H-sensitive matter. Second, this obstruction signature is not a yeast-specific artifact of our analysis. The same asymmetric, tRNA/snoRNA-anchored peak shape is present in independently generated tSNS-seq datasets from Drosophila and C. elegans. In fact, C. elegans shows this artifact is in the non-replicating control, and is eliminated in the RNase control. Furthermore, a recent publication using (strand-specific) tSNS-seq in Trypanosoma bruceireports that large majority (90%) of its peaks map to R-loops (Stanojcic et al. 2026), although they report it as mapping true replication origins, not as a demonstration of the artifact. Finally, we added an analysis comparing R-loops mapped by RIAN-seq in mouse (Li et al. 2025) and two independently generated mouse tSNS-seq datasets (Cayrou et al. 2015 and Pratto et al.2021) that show strong correlation. Taken together, this is not one experiment pointing to hybrid-mediated blockage, but a repeated, multiple-species in vivo-validated pattern observed in tSNS-seq by multiple independent laboratories.

      What our data cannot do is isolate the enzyme's behavior on a defined R-loop substrate in vitro, separate from all other cellular context. That is a distinct and narrower question from whether hybrid-mediated obstruction explains the peaks we observe. The latter is what the convergent evidence above addresses. Future work can nail down the exact enzymology and conditions needed to replicate this in vitro.

      In Fig. 1C, there seems to be a minor increase of lambda-exo activity on R10D30 substrate in the Tris-HCI buffer compared to Glycine-KOH buffer, maybe a quantification using the loading control (D20) would help.

      As suggested by the reviewer, we have quantified band intensities for the RNA-DNA chimera (R10D30), the G4 oligo (D10G427D23), and the all-DNA control (D25) shown in Fig. 1C, normalized to the D20 loading control at each timepoint (new Supplementary Figure S5). As this gel-based densitometry is semi-quantitative at best, the values below should be interpreted as indicative rather than precise measurements; nonetheless, they confirm the reviewer's observation: the RNA-DNA chimera was digested to a somewhat greater extent in Tris-HCl (35% signal remaining at 16 h) than in Glycine-KOH (65% remaining). We have revised the Results text to report these values explicitly and to clarify our claim. The relevant passage in the Results now reads: "However, the entire oligo (D10G427D23) was digested in Tris-HCl buffer very efficiently, with barely visible bands after as early as 1 h of λ-exo digestion (1% remaining at 16 h, comparable to the D25 control). The RNA-DNA chimera (R10D30) was digested slowly in both buffers over the 16 h time course, though to a somewhat greater extent in Tris-HCl (35% remaining) than in Glycine-KOH (65% remaining). This shows that RNA-primers offer partial, but not full, protection from λ-exo digestion in either buffer, distinct from the buffer-dependent protection conferred by G-rich DNA in glycine-KOH buffer. Thus, conditions for better digestion through G4-containing DNA in Tris-HCl buffer may be more suitable for mapping origins with SNS-seq."

      We further note that this observation does not bear on the final iSNS-seq protocol, since iSNS-seq relies on 5'-OH rather than 5' RNA to protect nascent strands from λ-exo digestion. The buffer-dependent difference in RNA-primer protection described here is therefore relevant only to the mechanistic characterization of λ-exo behavior in Fig. 1C, and does not affect the enrichment logic or performance of iSNS-seq itself.

      In Fig. 2D, lower panel, the digestion time of RecJf at different units were not indicated.

      We thank the reviewer for catching this point. We have revised the Figure 2D legend to remove this ambiguity, clarifying that the top panel shows a RecJf timecourse, while the bottom panel shows a titration of RecJf units (0–300 U) at a fixed 16 h digestion time.

      In Fig. 3B, since SNS-seq detect fired origins, an overlap of iSNS rep 1 and iSNS rep 2 peaks with confirmed origins from OriDB that are also early and efficient origins may yield a higher percentage of overlap.

      We thank the reviewer for this suggestion. Rather than stratifying discrete peak-overlap percentages, which are sensitive to peak-calling thresholds and tend to understate genuine concordance, we assessed this using continuous signal enrichment at OriDB-confirmed origins, that were stratified into quartiles by annotated origin efficiency (Hawkins et al., 2013) (see modified Fig. 3C). Both iSNS-seq replicates showed a clear trend of increasing median signal from the least to the most efficient origin quartile. This is consistent with the expected behavior of SNS-seq, which detects origins in proportion to their firing frequency in the asynchronous population, and we now state this explicitly in the manuscript.

      In Fig. 3D, in addition to showing the tSNS and iSNS data on ChrIV with indicated known origins, showing at the same time the corresponding FORK-seq, OK-seq, and ORC ChIP signal will give a more comprehensive view of the performance of different origin mapping methods.

      We agree with the reviewer that including complementary origin-mapping datasets alongside tSNS-seq and iSNS-seq would provide a more comprehensive view of method performance. We have added FORK-seq initiation site midpoints, ORC ChIP-seq signal, and OK-seq origin efficiency metric (OEM) signal to Fig. 3D, alongside the tSNS-seq and iSNS-seq tracks and annotated known origins on Chr IV. In addition, we have added close-up views of the tSNS-seq and iSNS-seq tracks around three selected origins in Fig. 3D, which highlight that the prominent tSNS-seq peaks do not align with known origins, in contrast to iSNS-seq.

      In Fig. 5F, in supporting the model, a comparison of tSNS-seq peak with known budding yeast DRIP-seq datasets would help to see if the tSNS method indeed enrich for RNA:DNA hybrid signals.

      We agree with the reviewer that a direct comparison to known R-loop maps would substantially strengthen the proposed model. We have added this analysis, described briefly below, in a new Results subsection and Figure 6.

      We obtained raw, publicly deposited budding yeast S1-DRIP-seq and S9.6 ChIP-seq data, including RNase H-treated and RNase H-deficient (RNase HΔ) conditions, and called peaks from these datasets ourselves using our own pipeline, so that all datasets were processed identically. A new heatmap colocalization figure (Fig. 6A) shows that tSNS-seq (but not iSNS-seq) signal is clearly elevated at these R-loop peaks, most strongly in the RNase HΔ condition and abolished in RNase-treated controls, consistent with genuine RNA:DNA hybrid dependence. The reciprocal analysis confirms this specificity: R-loop signal is concentrated at tSNS-seq peaks, but shows no enrichment at iSNS-seq, OK-seq, or FORK-seq peaks, or at confirmed origins. R-loop signal is also enriched at tRNA, snoRNA, and snRNA genes, mirroring the tSNS-seq enrichment pattern at these same loci. Together, these results directly demonstrate that tSNS-seq peaks colocalize with bona fide R-loops rather than replication origins, supporting the model in Fig. 5F.

      Finally, to test whether this observation generalizes beyond yeast, we correlated published mouse RIAN-seq signal — an antibody-free, nuclease-based R-loop mapping method developed independently of SNS-seq — with two independent mouse tSNS-seq-type datasets (Cayrou et al. 2015; Pratto et al. 2021). tSNS-seq enrichment correlated significantly with RIAN-seq R-loop signal across replicates and genomic scales, indicating that the same hybrid-mediated obstruction mechanism likely operates in a mammalian system as well. Moreover, analyses of tSNS-seq signal in Drosophila and C. elegans confirms that the enrichment bias at tRNA sites also exists in those systems. The C. elegans data shows it is not dependent on DNA replication, and is dependent on RNA, just like in yeast.

      _Reviewer #2:_

      The authors' sugestion that tSNS miscalling of origins may be due to the presence of RNA:DNA hybrids is largely persuasive in theoretical terms but is, surpisingly, untested. The first issue that is unexamined in whether or not RNA:DNA hybrids would survive the two rounds of 95 degree denaturation that are central to both froms of SNS; can the auhtors comment or provide evidence?

      Second, it is important that the authors test their proposal (Fig.5) of RNA:DNA hybrids being the cause of at least some non-origin tSNS signal by mapping the correspondance of their SNS-seq data with the locations of RNA:DNA hybrids, since there are several forms of such mapping available for S. cervisiae (eg using S9.6 antiserum DRIP-seq, or RNaseH1 ChIP-seq).

      We thank the reviewer for raising these two important points, both of which push us to be more precise about what our model does and does not claim.

      On denaturation and hybrid survival: We agree this is a critical point to clarify, and we think it reflects a distinction we had not made explicit enough in the original submission. We are not proposing that in vivo-formed R-loops survive the 95°C denaturation steps intact. That would indeed be difficult to reconcile with the protocol. Rather, our data suggest that the highly abundant RNA species themselves (tRNAs, snoRNAs) survive denaturation as free single-stranded RNA and re-anneal with their complementary genomic DNA strand at, or before, the λ-exo digestion step, which is carried out at 37°C over an extended overnight incubation. This reannealed RNA:DNA hybrid is what we propose obstructs λ-exo in tSNS-seq. Indeed, the preservation of this RNA throughout the protocol until the λ-exo step is a built-in feature of the traditional protocol, since the RNA is only hydrolyzed after the λ-exo digestion is complete. We have added the following sentence to the Discussion to make this explicit: “Importantly, we do not propose that in vivo R-loops survive the denaturation steps of the iSNS-seq protocol. Rather, we propose that RNA:DNA hybrids reform in vitro after denaturation at genomic loci that are prone to R-loop formation. The likelihood of hybrid reformation would be expected to increase with local RNA abundance. These reformed hybrids would then selectively block λ-exo digestion, producing the characteristic asymmetric enrichments observed in tSNS-seq.

      On testing correspondence with mapped RNA:DNA hybrids: We agree, and we have added a new Results subsection and Figure 6 to directly test this. We generated a new heatmap colocalization figure (Fig. 6A) comparing tSNS-seq and iSNS-seq signal against published budding yeast S1-DRIP-seq and S9.6 ChIP-seq maps, including RNase-treated controls. tSNS-seq signal is clearly elevated over R-loop sites identified by these orthogonal methods, but not peaks called from RNase-treated controls; iSNS-seq signal is flat over both R-loop and control peak sites. We also performed the reciprocal analysis and see that R-loop signal is highly concentrated at tSNS-seq peaks, but not iSNS-seq peaks (Fig. 6B). Importantly, this RNA:DNA hybrid enrichment signal at tSNS-seq peaks is abolished in the RNase H-treated controls and enhanced in RNase deficient mutants that accumulate R-loops, both consistent with genuine RNA:DNA hybrid dependence. Finally, we demonstrate that, like tSNS-seq, the R-loop signal is also highly enriched at tRNA, snoRNA, and snRNA genes (Fig 6C). Analyses of tSNS-seq signal in Drosophila and C. elegans shows this SNS enrichment bias at tRNA sites also exists in those systems. Importantly, the C. elegans data shows it is not dependent on DNA replication, and is dependent on RNA, and is thereby consistent with the RNA:DNA enrichment artifact we discovered in yeast.

      As a further, independent test of generalizability, we correlated tSNS-seq signal with RIAN-seq (which maps R-loops genome-wide) using two published mouse datasets: the original tSNS-seq data from Cayrou et al. 2015 and the strand-specific tSNS-seq data from Pratto et al. 2021. We found a significant positive correlation with RIAN-seq R-loop signal in both, indicating the same mechanism operates in a mammalian system (Figure 6D, E). We also note in the Discussion that a recent report (preprint at the time of this review, Stanojcic et al. 2026) similarly found that 90% of tSNS-seq peaks in Trypanosoma bruceioverlap with mapped R-loops, providing independent, multiple-species support for this interpretation.

      Fig.S7. Can the authors comments on the poor reproducibility between SNS-seq replicates? Though there isa four-fold increase in concordance between iSNS experments compared with tSNS, 80% of potetial origins are missed; can this be explained? In this regard, is the statment 'Both of the iSNS-seq samples showed signal enrichment around most of the confirmed origins' correct: while Fig.3C gives this impression, Fig.3B appears to diagree.____

      We agree the original Fig. 3B (Euler diagram) understated the reproducibility of iSNS-seq and, on reflection, we do not think it was the appropriate metric to lead with. Euler/peak-set overlap is a binary, threshold-dependent measure. Peaks that fall just above the calling threshold in one replicate and just below it in the other are counted as fully discordant even when the underlying signal is well correlated. We have moved the Euler diagram to the supplement (Fig. S9) and replaced Fig. 3B with a cross-replicate F1 curve, and cross-replicate signal heatmaps. Together these analyses show that concordance between replicates is graded and substantially above what the fixed-threshold Euler comparison implied. We have revised the text accordingly, and noted that the aggregate enrichment is not fully captured by peak-overlap statistics for the reasons above.

      The same failure mode of overlap analysis applies to overlap of iSNS (or tSNS) peaks with known origins. While the SNS signal may be fully correlated with origin positions and even origin efficiency, parameter choices in peak calling and using binary overlap statistics can mask the underlying concordance of the data with known replication origins. Moreover, the list of known replication origins, even confirmed origins, is not a list of origins most likely to be active. SNS methods can only detect active origins. Therefore, we assessed concordance of the SNS methods with origin efficiency using signal enrichment at OriDB-confirmed origins that were stratified into quartiles by annotated origin efficiency (Hawkins et al., 2013) (see modified Fig. 3C). Both iSNS-seq replicates showed a clear trend of increasing enrichment signal from the least to the most efficient origin quartile. This is consistent with the expected behavior of SNS-seq, which detect origins in proportion to their firing frequency in the asynchronous population, and we now state this explicitly in the manuscript.

      Given the very nice data in Figs.1 and 2 showing the confounding effect of G4s on tSNS, can the authors comment on why G4s show little overlap with either SNS-seq mapping in Fig.4, and why they saw no increase in tSNS-seq or iSNS-seq signal over G4 motifs in Supplementary Figure S9B? Might this indicate that the in vitro work does not translate well to in vivo mapping?

      We agree this warranted comment and have expanded the Discussion accordingly. We were ourselves somewhat surprised not to see a higher genome-wide G4 signal, though not entirely so. This is not evidence that the in vitro work fails to translate in vivo: the same λ-exo bias has been shown to manifest genome-wide in human cells (Foulk et al. 2015). Rather, we suggest the limited yeast signal likely reflects yeast-specific properties, which has a genome with more uniform local GC content than human (new Supplementary Figure S22), far fewer G4 motifs overall (Wu et al. 2021), and possibly less stable G4 folding in vivo (Tran et al. 2011). Moreover, the G4s in the human genome are non-randomly clustered in the GC-rich isochores, which compounds the G4 bias further with the known GC bias of Lambda exonuclease. These points are now made in the revised manuscript.

      The known G4 problem was the basis for searching for better buffer conditions in this study. While we found conditions that eliminate the G4 bias in vitro, the yeast genome did not provide ample opportunity to test this. However, the yeast genome ultimately allowed us to discover a possibly more dominant systematic bias, which is the correlation with tRNA and other high copy number RNA species that likely form RNA:DNA hybrids in vitro. Moreover, the lack of G4s allows a clean separation of G4s and R-loops, features that are highly correlated in the human genome.

      A requirement to validate the interesting suggestion that RNA:DNA hybrids are a cause of tSNS-seq miscalling of origins; this should be a combination of in vitro tests and colocalisation of tSNS-seq signal and RNA:DNA signal in vivo using available datasets (eg DRIP-seq).

      We agree that a direct in vitro biochemical demonstration that λ-exo digestion is specifically obstructed by an RNA:DNA hybrid substrate would further strengthen this model. However, as the laboratory where the wet-lab work for this study was performed has since been closed due to retirement, additional experiments of this kind are not feasible for this revision. We have instead addressed this request through a combination of new in vivo colocalization analyses, and an additional line of convergent evidence from the literature, which together we believe make a compelling case for the model.

      First, we have now added direct in vivo colocalization evidence: a new heatmap figure (Fig. 6A) comparing tSNS-seq and iSNS-seq signal to published budding yeast S1-DRIP-seq and S9.6 ChIP-seq maps, with RNase H-treated controls. tSNS-seq signal is clearly elevated over R-loop sites identified by these orthogonal methods, but not peaks called from RNase-treated controls; iSNS-seq signal is flat over both R-loop and control peak sites. Importantly, this RNA:DNA hybrid enrichment signal at tSNS-seq peaks is abolished in the RNase H-treated controls and enhanced in RNase deficient mutants that accumulate R-loops, both consistent with genuine RNA:DNA hybrid dependence. Finally, we demonstrate that, like tSNS-seq, the R-loop signal is also highly enriched at tRNA, snoRNA, and snRNA genes (Fig. 6C). Analyses of tSNS-seq signal in Drosophila and C. elegans shows this SNS enrichment bias at tRNA sites also exists in those systems. Importantly, the C. elegans data shows it is not dependent on DNA replication, and is dependent on RNA, and is thereby consistent with the RNA:DNA enrichment artifact we discovered in yeast.

      Second, we further demonstate that R-loop and tSNS-seq correlation extends to mouse as well. A novel R-loop mapping method RIAN-seq (Li et al. 2025) uses λ-exo, together with nuclease P1 and T5 exonuclease, to selectively degrade single-stranded RNA, single-stranded DNA, and double-stranded DNA from digested genomic DNA, while RNA:DNA hybrids resist this treatment and are selectively recovered. That λ-exo digestion is used as one of the core enzymatic steps to enrich for RNA:DNA hybrid-containing DNA is itself a genome-wide demonstration that λ-exo activity is impeded by RNA:DNA hybrid structures. Building on this, we correlated tSNS-seq signal against RIAN-seq signal using two independent published mouse tSNS-seq datasets (Cayrou et al. 2015 and Pratto et al. 2021) and found significant positive correlation in both, directly linking λ-exo obstruction by RNA:DNA hybrids to tSNS-seq peak formation in an independent system (Fig. 6D, E).

      Finally, a recent study (Stanojcic et al. 2026) using tSNS-seq on trypanosomes noted the majority (90%) of their peaks were associated with R-loop structures. Thus, we can defensibly conclude that this correlation is seen in yeast, C. elegans, Drosophila, mouse, and trypanosome datasets.

      An explanation and/or comment on the low reproducibility and low signal-to-noise ratio of iSNS-seq; specifically, although it improves on tSNS, does it provide accurate origin prediction?

      We do not believe reproducibility is in fact poor; the discrete peak-overlap metric we originally used was overly conservative (although still significantly above random). On signal-to-noise: we have added FRiP-vs-cumulative-ranked-peaks curve (Fig. 3E) showing that both iSNS-seq and tSNS-seq achieves higher than random raw FRiP. While tSNS-seq achieved higher FRiP, the iSNS-seq shows greater specific enrichment (SN/PPV against OriDB) at real origins. This is because the majority of tSNS-seq reads are in tRNA-associated non-origin peaks. In other words, tSNS-seq's apparently favorable S:N is driven in large part by non-origin enrichment (e.g., R-loop-forming loci), consistent with our RNA:DNA hybrid-mediated bias hypothesis. In contrast, iSNS-seq's lower overall S:N reflects the removal of much of that strong non-origin signal, leaving a smaller but more origin-specific signal. Whereas the non-origin biases are strongly present in a sample, nascent strands from a replication origin are very rare in the sample in comparison; present only in the fraction of S-phase cells in an asynchronous population where the nascent origin-proximal DNA is Comment on the lack of in vivo evidence, from available mapping data, for the confounding effect of G4s seen in the in vitro experiments.

      In vivo evidence for this λ-exo bias does exist in human cells (Foulk et al. 2015). Its absence in our yeast mapping data is better explained by yeast-specific genomic features than by the effect being artifactual, or in vitro-only. We have revised explicit comment on this in the Discussion:

      Given the clear G4-mediated λ-exo obstruction we observed in traditional conditions in vitro, as well as previously demonstrated genome-wide blockage at G4s in human cells and preference of λ-exo for AT-rich over GC-rich DNA (Foulk et al 2015), we were somewhat surprised not to observe a higher proportion of G4-overlapping fragments in tSNS-seq compared to iSNS-seq. However, this is not entirely unexpected for several reasons. First, the budding yeast genome is markedly more uniform than the human genome: local GC content (100 bp bins) spans a 10th–90th percentile range of 30–46% in yeast versus 26–55% in human (Supplementary Figure S22). In fact, the yeast genome has roughly half the variation of GC content found in the human genome as measured by standard deviation and median absolute deviation (MAD) (Supplementary Figure S18). Budding yeast also has substantially fewer G4 motifs than humans, both in absolute number and as a proportion of the genome (Wu et al. 2021), limiting the opportunity for a genome-wide G4 signal to emerge regardless of any per-motif blocking effect. Finally, we cannot rule out that budding yeast G4 motifs simply form less stable quadruplexes, as has been shown for yeast telomeric G4s specifically (Tran et al. 2011).”.

      Reviewer #3:

      While the method developed (iSNS-seq) appears to improve aspects of the previously used method (tSNS-seq), poor signal-to-noise ratios and reproducibility are major concerns. The signal-to-noise ratio of iSNS-seq, as shown in Figure 3D, appears low. The authors should discuss what practical consequences this has for the applicability of the method, especially in species with less well-defined/efficient origins, and how it affects the reliability of peak calling, and overlap with known origins. Consistently, the overlap of called peaks between biological replicates of iSNS-seq is relatively low (Fig3B). This raises a concern for the usability of the method, the interpretation of the genome-wide data and the quality metrics provided. The authors should discuss these issues.

      We have added a Discussion paragraph addressing this directly. Because the intrinsic signal-to-noise ceiling of SNS-seq-type methods scales with the size and firing efficiency/synchrony of the origin population (as detailed in our existing calculation based on Cadoret et al. (2008)), we now state explicitly that applying iSNS-seq to organisms or cell populations with less efficient, less synchronized, or less well-defined origins than budding yeast should be expected to yield correspondingly lower signal-to-noise and reduced peak-calling reliability, with a likely higher false-negative rate for genuine origins. We now acknowledge the need for further optimization to improve the underlying SNS:background ratio before extending the method to other organisms.

      On reproducibility: the original Fig. 3B (Euler diagram) understated the reproducibility of iSNS-seq and, on reflection, we do not think it was the appropriate metric to lead with. Euler/peak-set overlap is a binary, threshold-dependent measure. Peaks that fall just above the calling threshold in one replicate and just below it in the other are counted as fully discordant even when the underlying signal is well correlated. We have moved the Euler diagram to the supplement (Fig. S9) and replaced Fig. 3B with a cross-replicate F1 curve, and cross-replicate signal heatmaps. Together these analyses show that concordance between replicates is graded and substantially above what the fixed-threshold Euler comparison implied.

      We have replaced the Euler-diagram comparison with a graded, cross-replicate F1 analysis (new Fig. 3B) and supplementary correlation/Jaccard metrics, which show reproducibility is substantially above random for both methods and is in fact tighter for iSNS-seq than tSNS-seq at the level of peak-height correlation.

      Related to the above, please explain/discuss the rationale behind retaining all peaks, even the ones observed in only one of the biological replicates, when comparing called peaks to confirmed ORIs. In Figure 3B, clarify what are the percentages depicted.

      We agree with the reviewer that this point deserves fuller explanation, and we address the two parts in turn.

      Rationale for retaining single-replicate peaks. Origin firing is stochastic across an asynchronous population, and any single replicate is a depth-limited sample of the true origin population rather than an exhaustive one. Requiring a peak to be called in both replicates before comparing it to OriDB would systematically discard true origins that happen to be under-sampled in one replicate, artificially inflating apparent PPV at the direct cost of sensitivity. Since our goal in this analysis was to characterize each method's raw concordance with known origins, we compared each replicate's full peak set to OriDB independently.

      What the percentages in the diagram represent. We have moved this Euler diagram comparison from the main text to Supplementary Figure S9. We have clarified the legend of Supplementary Figure S9 (formerly Figure 3B) to read: "Percentages represent the fraction of the total combined peaks across all three sets being compared (i.e., the union of iSNS-seq rep 1, iSNS-seq rep 2, and confirmed ORIs for panel A; tSNS-seq rep 1, tSNS-seq rep 2, and confirmed ORIs for panel B) that fall into each region of the diagram."

      Why this analysis was moved to the supplement. Binary overlap statistics of this kind are informative but limited: they collapse a continuous relationship (signal strength versus origin identity) into a single threshold-dependent yes/no call, and are sensitive to peak-calling parameters. Moreover, the list of OriDB-confirmed origins is not itself a list of the origins most likely to be active in a given asynchronous population: most methods can only ever detect origins that actually fired, so a "true" origin that failed to overlap a peak may simply reflect biological non-firing rather than a false negative of the method.

      For these reasons, we chose to complement the overlap-based comparison with an analysis that queries origin correspondence directly against the experimental signal rather than against a binarized peak call. Specifically, we stratified OriDB-confirmed origins into quartiles by annotated origin efficiency (Hawkins et al., 2013) and examined SNS signal enrichment across quartiles (modified Fig. 3C). Both iSNS-seq replicates show a clear, monotonic increase in enrichment signal from the lowest to the highest efficiency quartile. This is the expected behavior for a method that detects origins in proportion to their firing frequency in an asynchronous population. We now state this rationale explicitly in the manuscript text: “To assess the reproducibility of iSNS-seq and tSNS-seq more rigorously than a fixed-threshold peak-overlap comparison allows […] Because this approach evaluates concordance at every possible rank cutoff rather than a single arbitrary threshold, it is not subject to the same sensitivity to near-threshold peaks that limits discrete overlap statistics such as Euler diagrams (Supplementary Figure S9).

      The title and abstract do not accurately convey the limitations of the improved method and must be rephrased.

      We thank the reviewer for this comment and have revised the abstract accordingly. Specifically, we have (1) clarified that our benchmarking was conducted in S. cerevisiae precisely because it offers a well-defined set of confirmed origins for direct comparison, (2) made explicit that our enrichment claims are relative to traditional SNS-seq rather than absolute, and (3) added language framing this work as a proof-of-concept benchmark, with extension to metazoan systems as a next step rather than a claim already established here. We believe these changes address the concern while accurately reflecting what our data show.

      We have retained the title, as we believe "enhances DNA replication origin detection and reduces non-origin biases" is a comparative claim relative to traditional SNS-seq, which is directly supported by our benchmarking data, and does not imply resolution of broader field-wide inconsistencies in origin mapping.

      One of the manuscript's central mechanistic claims is that residual cellular RNA reanneals to genomic DNA, creating RNA:DNA hybrids that obstruct λ-exonuclease digestion. While the presented data from cells are consistent with this model, the manuscript lacks a direct biochemical demonstration of hybrid-mediated obstruction in a controlled system. The manuscript would benefit from in vitro studies corroborating this conclusion, using the in vitro system established.

      We agree this would be a valuable addition, and in principle the in vitro system established in this study (Fig. 1, 2) is well suited to such a test. However, the laboratory where this wet-lab work was performed has since been closed due to retirement, and we are unable to carry out new biochemical experiments of this kind for this revision.

      We believe the manuscript already provides convergent evidence for RNA:DNA hybrid-mediated obstruction, short of this direct reconstitution. First, our in vivo colocalization analysis shows that tSNS-seq (but not iSNS-seq) signal is enriched at S1-DRIP-seq and S9.6 ChIP-seq sites in an RNase H-sensitive manner, directly implicating RNA:DNA hybrids at the genomic loci where tSNS-seq peaks form. Second, this obstruction signature is not a yeast-specific artifact of our analysis. The same asymmetric, tRNA/snoRNA-anchored peak shape is present in independently generated tSNS-seq datasets from Drosophilaand C. elegans. Furthermore, a recent publication using stranded tSNS-seq in Trypanosoma brucei reports that a large majority (90%) of its peaks map to R-loops (Stanojcic et al. 2026). Finally, we added an analysis comparing R-loops mapped by RIAN-seq in mouse (Li et al. 2025) and two independently generated mouse tSNS-seq datasets (Cayrou et al. 2015 and Pratto et al. 2021) that show strong correlation. Taken together, this is not one experiment pointing to hybrid-mediated blockage, but a repeated, cross-species in vivo-validated pattern observed in tSNS-seq by multiple independent laboratories. The distinct contribution of our yeast experiments is that a well-defined set of confirmed origins allowed us, for the first time, to directly distinguish genuine replication signal from this R-loop-associated background.

      What our data cannot do is isolate the enzyme's behavior on a defined R-loop substrate in vitro, separate from all other cellular context. That is a distinct and narrower question from whether hybrid-mediated obstruction explains the peaks we observe. The latter is what the convergent evidence above addresses. Future work can nail down the exact enzymology and conditions needed to replicate this in vitro.

      Currently, the RNA-dependence of the tRNA/snoRNA-associated obstruction signal is inferred from indirect evidence (UMAP clustering, feature overlap, peak shape) plus a re-analysis of an RNase-treated control from a previously published C. elegans dataset, not from a control generated within the authors' own yeast system, where the paper's primary genome-wide claims are made. Including such controls would be required to provide direct evidence for this central mechanistic claim.

      We appreciate the opportunity to clarify this, as we believe we already have exactly this control within our own yeast dataset. iSNS-seq itself functions as this RNA-dependence control for tSNS-seq: unlike tSNS-seq, the iSNS-seq protocol includes an RNA hydrolysis step (NaOH) performed before λ-exo digestion, removing the RNA that would otherwise be available to reanneal and form obstructing RNA:DNA hybrids. Both tSNS-seq and iSNS-seq are generated from the same size-selected input material from the same yeast cultures, differing specifically in whether RNA is retained (tSNS-seq) or hydrolyzed (iSNS-seq) prior to the λ-exo step (see Fig 3A, which is now improved to clarify this point). The result is a matched, within-system, RNA-present versus RNA-removed comparison, generated entirely in our own hands in budding yeast: tSNS-seq (RNA intact) shows strong signal and peak enrichment at tRNA/snoRNA loci, while iSNS-seq (RNA hydrolyzed) shows essentially none (Fig. 5). This is, in effect, an RNase-equivalent control built directly into our primary experimental design, rather than one requiring a separate treatment arm.

      We have clarified this point explicitly in the Discussion by noting that the tSNS-seq vs. iSNS-seq comparison itself constitutes the within-system RNA-dependence control, complementing the new yeast S1-DRIP-seq/S9.6 ChIP-seq RNase H colocalization analysis (Fig. 6A-C), which independently corroborates this at the level of orthogonal R-loop mapping methods.

      A substantial portion of the in vitro biochemical work (Figure 1B-C, E, Figure 2B-C with BKGmimic) is dedicated to establishing that G4 motifs cause significant λ-exo obstruction in glycine-KOH buffer, and that the Tris-HCl buffer switch resolves this. However, the genome-wide finding that G4 motif overlap is nearly identical between tSNS-seq and iSNS-seq (4-5% in both), with no significant differences in signal over G4 motifs in either dataset, requires clarification/discussion.

      We have clarified this in the expanded Discussion. We note it was somewhat surprising to us as well not to observe more obstruction in tSNS-seq samples. However, our data is consistent with prior human work showing the same λ-exo bias (Foulk et al. 2015), combined with the yeast genome's comparative GC uniformity (new Supplementary Figure S22) and its markedly lower G4 motif density (Wu et al. 2021), and possibly less stable yeast G4 folding (Tran et al. 2011). We believe these factors sufficiently explain why a robust in vitro effect does not produce a differential genome-wide signal in this organism.

      The documented G4 bias was a motivation to find new buffer conditions in the first place, and our further study of this problem highlighted something important for the field even if it doesn’t matter for the yeast genome: RNA primers are less effective at protecting downstream DNA than G4 motifs. Moreover, that observation led us to inventing a way where nascent strands were more strongly protected than G4s by using the 5’-OH after purposeful RNA hydrolysis. Finally, RNA hydrolysis led to the discovery that there is a major RNA-dependent bias in tSNS-seq. Thus, the paper documents the journey and key insights that led to the iSNS-seq protocol, and major findings in this study regarding the tSNS-seq protocol.

      None of the gel-based panels in Figure 1 (B, C, E) or Figure 2D state whether the result shown is from a single experiment or is representative of multiple independent experiments. Please state explicitly, for each gel-based panel, whether it represents a single experiment or is representative of repeated independent experiments (and if the latter, how many).

      We agree with the reviewer that this information was missing and should be stated explicitly. The gel panels in Figures 1B, 1C, 1E, and 2D arise from an extensive experimental search (for buffer conditions and additives) that would increase digestion through G4 motifs while preserving the RNA-DNA chimera, during which these conditions were tested many times over the course of the study. Across this large number of experiments testing different buffers and additives, the same core observations consistently emerged: (1) -exo digests through the RNA-DNA chimera, (2) -exo has difficulty digesting through G4 motifs in Glycine-KOH buffer, but not in Tris-HCl, (3) a 5'-OH end protects DNA from -exo digestion much more effectively than an RNA primer, and (4) RecJf does not digest through a 5' RNA end (or through G4 motifs), but can reduce background DNA levels. The panels shown are representative of these consistently observed outcomes, which are also consistent with the overall findings presented throughout the manuscript. We have added a statement to this effect in the Methods and to each relevant figure legend.

      In the Introduction section, the authors state: "...traditional SNS-seq enriches non-origin DNA that appears to originate from RNA:DNA hybrids due to contaminating RNA, while failing to enrich known replication origins above random expectation." Replace the term "contaminating RNA" with the term "residual cellular RNA", accurately conveying that this is endogenous RNA persisting through the protocol.

      We thank the reviewer raising this point. We corrected the sentence accordingly, which now reads: “[…]traditional SNS-seq enriches non-origin DNA that appears to originate from RNA:DNA hybrids due to residual cellular RNA, while failing to enrich known replication origins above random expectation.”

      The sentence describing that "the RNA-DNA chimera was not digested when 5' phosphorylation by PNK was performed prior to RNA degradation, rather than after the removal of the RNA primer by NaOH" introduces the PNK/NaOH order-of-operations logic before this concept has been explained to the reader, the significance of treatment order is not established until two sections later. A minimal fix would be adding a forward-referencing clause or relocating this result to a later paragraph.

      We agree that the result appeared to be out of place. We have relocated a slightly modified version of this sentence to a later paragraph, just before explaining the Figure 1C experimental logic. The relocated part now reads: “When characterizing λ-exo properties in vitro, we found that the RNA-DNA chimera was not digested when 5' phosphorylation by PNK was performed prior to RNA degradation, rather than after the removal of the RNA primer by NaOH (Supplementary Figure S4D); also confirming that a 5' phosphate is required for λ-exo activity. This result prompted us to further assess how hydrolyzing RNA primers to obtain 5'-OH DNA protects short nascent strand (SNS) DNA from λ-exo digestion compared to 5'-phosphorylated DNA, and to devise an experiment that varied the order of RNA hydrolysis (NaOH) and phosphorylation (T4 polynucleotide kinase, PNK) (Figure 1D).

      In Figure 2D. (top panel), please state the amount of λ-exo used to perform this experiment in the corresponding figure legend.

      We believe there may be a small mix-up here: Figure 2D shows RecJf digestion throughout (top and bottom panels) and λ-exo was not used in either. However, we agree with the underlying point: the amount of RecJf used in the top panel (timecourse) was not stated in the legend. We have corrected this omission and, as also noted in a related minor comment from Reviewer 1, revised the Figure 2D legend overall for clarity, now stating explicitly that the top panel shows a timecourse of RecJf digestion (with hours indicated) and the bottom panel shows a titration of RecJf units (0–300 U) over a fixed 16 h digestion.

      In Figure 2D. where sonicated, heat-denatured yeast genomic DNA is used as a substrate, please explain how this DNA is prepared and if residual cellular RNA reannealing to complementary genomic sequences would/would not be expected in this prep. Discuss with respect to the RecJf experiments shown.

      The sonicated genomic DNA substrate in Figure 2D (bottom panel) was RNase-treated prior to RecJf digestion, which removes cellular RNA and precludes RNA reannealing to complementary genomic sequences in this specific preparation. We acknowledge that RNase treatment means this experiment does not fully replicate the RNA content of genomic DNA in an actual iSNS-seq reaction, where residual cellular RNA is present. However, we note that in the iSNS-seq workflow, RecJf digestion is used solely as an additional bulk genomic DNA clean-up step prior to λ-exonuclease digestion. Any genomic DNA fragments that resist RecJf digestion due to reannealed residual RNA would still be subject to the downstream PNK, NaOH, and λ-exonuclease steps that carry out the primary enrichment for origin-proximal SNS molecules. We have modifed the Results section to clarify this point: “[…] RecJf significantly reduced the levels of yeast genomic DNA that had been sonicated to the average SNS size range (0.5--2.0 kb), RNase-treated, and denatured by heat (Figure 2D bottom panel). Note that in the iSNS-seq workflow, RecJf digestion serves as an additional bulk gDNA clean-up step upstream of λ-exo digestion. Any genomic DNA species that escape RecJf digestion, including those potentially protected by reannealed residual RNA, remain subject to the subsequent PNK, NaOH (that hydrolyzes residual RNA), and λ-exo steps that establish specificity for SNS molecules.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

      We thank the editors and the reviewers for their positive assessment and constructive feedback on our work. We also thank the editor for communicating with reviewer #1 regarding our thoughts on their comments. We hence revised the manuscript based the new feedback from reviewer #1. Please see below our responses to each comment raised in the reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. In particular, two types of learned associations are studied. One association involves associations between stimulus features and control states (SC), and the other involves stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates.

      The authors conclude that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing that SC and SR correlates are (1) not entirely overlapping in cross-decoding, (2) simultaneously observed on average over trials, (3) independently correlate with RT, and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Thank you for acknowledging the strength in EEG decoding and design. We have addressed all your concerns raised below point by point.

      In my view, the ability to address this issue with additional analyses is relatively limited. I cannot think of a solid way to escape this issue in the present design. Adding a nuisance regressor to their RSA regression, which I suggested in my previous response, may reduce the bias, but the efficacy of this would be limited to controlling only a certain kind of phase-dependent confound (a 'main effect' component of study phase, i.e., one that is constant across other conditions; see next message for discussion).

      Rather than new analyses, I think a more straightforward revision might be to modify the conclusions advanced in the paper, so that they are more solidly supported by the design and evidence. In my opinion, this design is ill-posed to solidly identify SC and SR representations. As a result, I think that any framing that alleviates pressure on this design to yield 'solid' evidence for identification of SC and SR representations, and instead emphasizes stronger areas of this work, would be constructive.

      For example, one potential framing is to advance the idea of SC and SR representations, and discuss an idealized design that could identify them by orthogonalizing stimulus features from ISPC, which are genuinely novel and useful ideas. The current design could then be presented as an opportunistic or initial case study of testing this question, while acknowledging its limitation upfront. The goal here would be to frame the study in a way that allows for the results to be presented with an appropriate grain of salt, while also illustrating the authors thoughtfulness and creativity in devising analyses to test for latent associative representations. In this case the 'solid' label would reference the authors' reasoning and analyses rather than design and conclusions.

      I only intend this example as an illustration; there may be several ways of framing this paper so that it is on more 'solid' ground, and I don't want to dictate how exactly authors should write their paper.

      Nonetheless I think this issue is important, not only to avoid faulty inference, but also to avoid establishing counterproductive precedents in this field. For example, if students read a paper whose conclusions are labeled "solid" but that nevertheless has critical flaws in its design, then those students may be misled in their own work. But if the flaws were discussed transparently and critically, and the strength of the conclusions were de-emphasized relative to other aspects of the paper, students may not only be inspired by the ideas developed in the paper, but also come away with knowledge about the issues of experimental design.

      Discussion of the weaknesses in the conclusions and an additional potential analysis:

      The key goal of this study is to identify SC/SR representations, which requires decoupling stimulus features from item-specific proportion congruency (ISPC), but the study-phase confound contaminates this decoupling. I think this impacts both their cross-phase decoding and RSA analyses.

      In their results (lines 139-144):

      "This standard ISPC manipulation can test whether neural representations of controlled and non-controlled information are wrapped on the same trial by combing with the following EEG analysis (See Methods). However, it potentially mixes color identity with SC, word identity with SR, and the ISPC between SC and SR. To deconfound these factors when estimating SC and SR association representations on each trial, we modified this paradigm by flipping the ISPC contingencies across different phases of the task."

      SC/SR representations are higher-order conjunctive classes, formed by the interaction of lower order variables (Color, Word, and ISPC). This means that successfully identifying these representations relies on demonstrating that each class can be reliably individuated from every other class in a manner that cannot be explained by (1) representation of shared lower-order features, such as stimulus color or response, and (2) trivial nuisance factors such as study phase. However, in this design, the trivial factor of study phase is strongly confounded with the ISPC contingency flip.

      Regarding RSA: If study phase indeed leads to trivial separability, it seems that similarity among conditions within the same phase would be inflated because the study phase does not appear in the RSA model (Figure 12). In which case, the SC/SR coefficients would be inflated, as the SC and SR models are entirely within-phase.

      Nevertheless, an additional control analysis may be possible here. Looking at this regressor set, it seems possible to me to fit a model where the predominant phase is entered as an additional covariate. (If I am reading this correctly, phase would appear as a 2x2 block-diagonal matrix). I suggested this in my last review, but in their recent letter, authors refused. I do not understand why, as I think this regressor set should be identifiable, but perhaps I am wrong here.

      That being said, I do not think adding this nuisance covariate would fully solve the issue. The lower-order regressors (Color, Word, ISPC) are defined as the EEG responses shared across phases of the study. To interpret the higher-order SR/SC coefficients, the lower-order components must be fully partialled out. But the putative impact of study phase contaminates this interpretation, as study phase could trivially decrease the similarity of lower-order terms (e.g., decreasing Color similarity between phase 2 vs 3). In which case, the partialling would be expected to be incomplete.

      This incomplete partialling is the RSA analogue of the issue in interpreting cross-phase decoding discussed above. As there, so too here I do not see a solid way around it in the present design. This is why in my previous review I referred to the addition of a phase covariate in the RSA regression as a "band-aid": it controls for some problems (main effect of phase), but not all (interactions of phase and lower-order terms).

      To summarize, I think that RSA would offer an additional opportunity to control for this potential confound, albeit in a limited sense (study phase effects that are consistent across conditions). But in a more general and rigorous sense, to me, the SC and SR terms in the RSA regression also seem susceptible to the same weakness as the cross-phase decoding analysis.

      We thank the reviewer for taking the additional time and effort to provide the new comments. Following the reviewer’s suggestion, we revised the language regarding the conclusions of this project (page 2,5,27-28) and explicitly discussed the weaknesses of the design on decoding and outline a possible solution for future studies in the Discussion section:

      “A limitation of the current design is that in theory temporally structured noise (e.g., autocorrelation in EEG data) may bias the decoding accuracy due to the blocked design. Although the present data provided no evidence that the decoding results in this study were biased by temporally structured noise, future studies should aim to develop experimental designs that eliminate this potential confound at the source. One potential solution would be to introduce additional phases flipping ISPC manipulations. At the same time, enough trials must be included in each phase to ensure the strength of the ISPC effect within each phase. A careful balance between session number and length will be helpful to optimize the duration of such a design.”

      As eLife also publishes review report, below we also summarize the three control analyses we ran and our reasoning of how they (would) address the issue of temporally structured noise for interested readers to assess:

      We acknowledge the theoretical issue of temporally structured noise (TSN) in our design when classes were decoded across different phases. However, the key question for the current data is whether there is empirical evidence that the decoding results were actually driven by TSN. To clarify this issue, we summarize several lines of evidence suggesting that the decoding results were not attributable to TSN:

      (1) Split-half cross-validation. We split the EEG data from Phase 2 and the combined Phases 1 and 3 into chronological first and second halves. Phases 1 and 3 were combined because they shared the same MC and MI assignments. This resulted in four possible combinations, each consisting of eight classes drawn from different phases: combination 1 included the first half of Phase 2 and the first half of Phase 3; combination 2 included the first half of Phase 2 and the second half of Phase 3; combination 3 included the second half of Phase 2 and the first half of Phase 3; and combination 4 included the second half of Phase 2 and the second half of Phase 3. We trained the decoders on one combination and tested them on another and then averaged the decoding results across all possible training-test assignments. The similar decoding patterns observed across these analyses (Fig. 6a,b) further confirmed that the decoding results were not driven by TSN.

      This analysis is conceptually similar to the “cross-phase” decoding analysis suggested by the reviewer in the first round of review. We also performed an additional distance-based control analysis (see below) to further test whether the decoding results could be explained by TSN, without imposing the constraint used in the split-half cross-validation that trials from the two phases had to fall within a 400-trial window.

      (2) Distance analysis. We predicted that if a test trial is closer to a training trial of the same trial type, the higher similarity in TSN between the training and test data would more strongly inflate the decoding accuracy of the test trial, resulting in a negative correlation between distance between a test trial and its closest training trial of the same type and the test trial’s decoding accuracy. However, we did not observe such a negative pattern (Fig. 6c). Note that this distance was defined with respect to trials of the same type, rather than absolute chronological time.

      (3) Shuffled analysis. If the decoding results were primarily driven by TSN, either at a short-term or long-term timescale, then shuffling the condition labels within each mini block should preserve the TSN structure present in the real data. In that case, the decoding results from shuffled data should not differ from those observed from real data. However, we found the significant difference between real data and shuffled data as shown in Author response images.

      Author response image 1.

      Shuffling analyses with stimulus-locked data support separable SC and SR subspace. (a) Group average decoding accuracy of all 16 experimental conditions as a function of time after stimulus onset. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (b) Group average t values of representational strength for each factor over time. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (c) SC and SR association results from Fig. 1b.

      Author response image 2.

      Shuffling analyses with response-locked data support separable SC and SR subspace. (a) Group average decoding accuracy of all 16 experimental conditions as a function of time after stimulus onset. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (b) Group average t values of representational strength for each factor over time. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (c) SC and SR association results from Fig. 2b.

      We thank the reviewer for the suggestion on RSA with phase. There are some concerns for this analysis:

      First, we think that the suggested analysis may be difficult to interpret. Because the SC and SR conditions differ across phases. Regressing out phase in RSA could also remove SC and SR information.

      Second, based on the reviewer’s comment, we understand that the suggested analysis may still not provide a clear falsifiable criterion for determining whether the results could be driven by the theoretical TSN issue inherent in the design.

      Third, the three control analyses we have performed examine this issue from different perspectives and collectively provide no evidence that the results were driven by TSN.

      Other readers may, like me, be puzzled by the selection of this particular experimental design to test this question of SC and SR coding, given the temporal confound among SC/SR classes, and given that there would seem to be many possible designs that are less confounded. For example, why not use a design where ISPC was swapped/shuffled several more times within each subject, so that PHASE is more orthogonal to long-timescale noise? Isn't ISPC learning fast enough to support learning phases shorter than 700 trials? Such readers would likely appreciate a frank discussion of this dilemma, and a motivation for the choice of the present design, within the manuscript.

      Thank you for your suggestion regarding the design. It is possible that the (re-)learning of ISPC can be fast. That said, enough trials are required to obtain a robust ISPC effect for each phase after the ISPC flips. Given that the EEG scanning (not including capping) in current design was about 1.5 hours, it is impractical to have both more sessions for a more orthogonal design and long sessions for robust within-session ISPC effects. We chose to maximize the latter because flipped behavioral ISPC effect in each session is the basis for the following EEG analysis. We have included the reviewer’s suggestion as a potential design solution for future studies in the Discussion section mentioned above on page 27.

      Pre-stimulus coding:

      To explain the apparent pre-stimulus coding of several task variables, the newest version of the manuscript proposes that subjects were proactively coding these variables via predictive mechanisms. This is an interesting account of item-specific control. It is also surprising, given that item-specific control mechanisms are typically conceptualized as reactive or stimulus-driven phenomena. But I think support for a proactive control account was incomplete. The mechanistic logic was not presented, and no hypotheses under this account were developed or tested. So I would suggest pinning down some hypotheses here and actually putting this account to the test.

      Thank you for raising this important point. Although ISPC effects are considered reactive, in our design the long sessions may create a temporal context for the participants to differentiate the current control demand linked to each color. The maintenance of such contextual information needs to span across trials, leading to pre-stimulus coding that proactively guides the control demand for each color. This claim is not central to this manuscript, which investigates whether SC and SR representations simultaneously guide behavior. Additionally, we do not think the current design is well-equipped to test this hypothesis because the pre-stimulus onset is the only supporting evidence. In the revised manuscript, we discussed this as a future research direction and proposed a design that aims at better isolating proactive control signal on page 25.

      Random slopes were omitted due to convergence failure, but this can inflate false positive inferences (e.g., Barr et al. 2013), and doesn't really motivate a minimal model. I'd suggest trying a slightly reduced model (e.g., drop correlations via `slope || subject`) using buildMer automated selection, or switching to brms.

      We indeed tried both the full model of random effects (i.e., considering covariance between all slopes and intercept) and a reduced model without any covariance (i.e., listing each random slope separately without intercept in lme4). However, neither model converged at all time points.

      Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a solid design and the analyses are appropriate for the research questions.

      Strengths:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      We thank you for acknowledging our work on design and brain-behavior analysis. We have addressed your concerns raised below.

      Weaknesses:

      I still have some doubts on the effectiveness of the experimental manipulation on Phase 2: although the ISPC effect is still present, it is much weaker in comparison, suggesting the participants did not learn the contingency statistics in Phase 2 as well as they did in the other phases, due to either the lingering effect of the previous phase or an inherent bias towards one color pairs. Perhaps by separately plotting the earlier and later blocks of Phase 2 any difference can be revealed if it exists. This behavioral difference could result in unequal levels of SC/SR representation across phases, which may raise problems when data were combined for analyses that assume the neural effects are equivalent.

      Thank you for your concern about this important issue. We agree with the reviewer that the true SR/SC levels may not be equivalent between Phase 2 and Phase 1/3. Nevertheless, because the manipulation of ISPC is binary, the decoders were trained to test whether the neural signals represent the two levels of ISPC (i.e., a higher vs. a lower level) differ systematically. The decoding analysis does not require that the neural effects of SC and SR must be numerically equivalent between phases (i.e., it is not necessary that the two levels are equidistant from the center point of SC/SR. Indeed, the decoding analysis only requires that the two levels are different). For example, if the ISPC level ranges from -1 to 1 and the EEG signals can reliably decode ISPC levels of -0.5 and 0.7 (i.e., two unequal levels), it can still be treated as supporting evidence that ISPC levels are encoded in the EEG signals. The same logic applies to the representational subspace analysis. As to the RSA, as can be seen in Fig. 12, the regressors are also binary, encoding whether two experimental conditions share the same SC/SR level without assuming equivalent neural effects. Thus, we argue that the reported decoding and RSA can still test the encoding of SR and SC. We discussed this issue on page 24.

      Following the reviewer’s comment, we plotted the ISPC effects in first and second half of Phase 2 separately (the figure below). We also tested whether the ISPC effects differ qualitatively between the two halves using a 3-way ANOVAs separately on RT and Error rate. The results showed that the time (the first half vs. the second half) × Congruency × ISPC interaction was not significant for either RT data (F<sub>(1,39)</sub> = 3.40, p > 0.05) or error rate (F<sub>(1,39)</sub> = 2.49, p > 0.05), suggesting that ISPC effect did not systematically change over time in Phase 2 See Supplementary Figure 9.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In the revised manuscript, we have expanded the real-data analyses, clarified the relationship between CMP and prior modulated Poisson models, and added discussion of model limitations and future extensions. In summary, the major changes include:

      (1) We revised the Introduction, Results, Methods, and Goris-model appendix to clarify the relationship between CMP and prior modulated Poisson models. In particular, we now emphasize that the key distinction is CMP’s continuous-time stochastic gain process.

      (2) We moved the simulation-based recoverability analysis from Appendix 3 into the main Results section (Figure 3 in the revised manuscript), making the validation of the inference procedure more visible to readers.

      (3) We added new analyses of the inferred gain process. Specifically, we now show the cross-trial gain mean and cross-trial gain variance in Figure 4A to assess whether gain captures stimulus-locked structure, and we added an analysis of pre- versus post-stimulus cross-trial gain variability in Figure 4C to test for gain-variability quenching during stimulus presentation.

      (4) We clarified the definitions and implementation of the Baseline Poisson, Poisson-GP, and Goris-style comparison models, including the role of the smoothness prior on the stimulus drive.

      (5) We expanded the Discussion to describe future extensions to population recordings, including a GPFA-inspired extension with low-dimensional shared gain activity across neurons.

      (6) We added a Discussion paragraph clarifying that the current CMP model captures Poisson and super-Poisson variability, but not sub-Poisson variability, and outlined possible extensions using spike-history terms, renewal-process likelihoods, or alternative count distributions.

      eLife Assessment

      This work of fundamental significance introduces a novel statistical model of spiking activity that incorporates continuous−time gain modulation. The authors provide exceptional evidence that the model outperforms earlier approaches and alternative candidates in capturing spiking responses across multiple visual areas in the macaque. Beyond its methodological contribution, the study offers new insights into how stimulus−driven variability and internally generated gain fluctuations evolve over time and between brain areas. The framework is likely to find broad application beyond the datasets examined here.

      We sincerely thank the Senior Editor, Reviewing Editor, and both reviewers for their careful evaluation and constructive feedback. We are encouraged by the positive assessment of the work and by the recognition of its methodological and conceptual contributions. We especially appreciate the acknowledgement that the continuous-time formulation provides a useful framework for modeling gain modulation in spiking activity, improves upon earlier approaches in capturing responses across multiple visual areas, and offers new insights into how stimulus-driven variability and internally generated gain fluctuations evolve over time and across brain regions.

      In the revised manuscript, we have addressed the reviewers’ comments by clarifying the relationship between CMP and prior modulated Poisson models, strengthening the presentation of the simulation-based recoverability analysis, adding new validation analyses of the inferred gain process, and expanding the Discussion of model scope, limitations, and future directions. In particular, we now more clearly distinguish the continuous-time gain process in CMP from Goris-style models with constant or piecewise-constant gain, move the simulation recoverability analysis into the main Results, examine trial-averaged inferred gain and gain-variability quenching, clarify the definitions of the baseline and comparison models, and discuss extensions to population recordings and sub-Poisson variability.

      We believe these revisions improve the clarity, rigour, and scope of the manuscript. Below, we address each reviewer comment in turn and describe the corresponding changes made in the revised manuscript.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, Rupasinghe and co−authors introduce a new statistical model for spiking neurons. Building on earlier work, they propose to model spikes as arising from a Poisson process whereby the firing rate is the product of stimulus drive and astimulus−independent gain signal. The critical innovation of this work is that the gain signal is modeled in continuous time. Earlier explorations of this statistical construction treated the gain−signal as constant within a trial. This innovation is elegant and important. It makes the model richer, more plausible, and more broadly applicable. The authors show that the model parameters are recoverable from realistic amounts of data and then apply the framework to previously studied datasets. They show that the new model outperforms earlier models and alternative candidates in capturing spiking data across four visual areas of the macaque monkey. Analysis of the model parameters replicates some earlier findings and uncovers several new insights. The model and fitting methods can be broadly applied to partition different types of signals and noise from spiking data and are likely to be widely adopted in the systems neuroscience community.

      Strengths:

      (1) Through clever use of advanced statistical techniques, the authors manage to infer critical information from single−trial single−cell data.

      (2) The question of which aspect of a spike train is signal and which is noise is omnipresent in neuroscience. By improving our ability to characterize the distinct factors that shape spiking activity, this work makes a fundamental contribution to the literature.

      We sincerely thank the reviewer for the thoughtful and detailed evaluation of our manuscript. We are pleased that the continuous-time formulation and its methodological contributions were viewed as elegant, important, and broadly applicable. We also appreciate the reviewer’s recognition that the framework provides a useful way to separate stimulus-driven and modulatory components of neural variability from single-trial, single-cell data. The reviewer’s comments helped us improve the precision of our framing, clarify the relationship between CMP and prior modulated Poisson models, and strengthen the validation of the inferred gain process. Below, we respond to each point in turn and describe the revisions made in the manuscript.

      Weaknesses:

      Overall, I find the work impressive and important. I have a couple of questions and suggestions.

      (1) The work is entirely focused on single−cell data. While this is a great starting point, expanding the approach to spiking activity in neural populations is an importantfuture goal.

      We thank the reviewer for this important suggestion. We agree that extending the CMP framework to population recordings is a natural and important direction for future work. In the present study, we focus on single-neuron responses to establish the continuous-time model, validate the inference, and characterize how stimulus-driven activity and stochastic gain fluctuations can be separated at the level of individual cells. However, the same modeling principles could be extended to simultaneously recorded neural populations by introducing shared latent structure across neurons. For example, one natural direction would be to combine CMP with ideas from Gaussian Process Factor Analysis [Keeley et al., 2020], using low-dimensional shared gain activity to capture population-wide fluctuations, while retaining neuron-specific stimulus-driven components. Such an extension would allow the model to capture correlated variability and shared modulatory dynamics across neural ensembles. In the revised manuscript, we have expanded the Discussion to describe this possible future extension to population recordings.

      To address this comment, we expanded the Discussion (Page 14: lines 473-478) to describe a possible GPFA-inspired extension of CMP to population recordings.

      (2) Line 49−53: These statements seem incorrect to me. The modulated Poisson model , as introduced in Goris et al (2014), is a process model that can perfectly be used to generate spike trains (within a trial, spiking emerges from a Poisson process, which canbe homogeneous or inhomogeneous). Moreover, the model contains a parameter thatrepresents the duration of the counting window (delta t). The dependency of over− dispersion on the size of the time bins for real neurons is shown in Figure 1b (inset plot) of that paper (and shown to resemble the model prediction). This time− dependency was further explored by the same authors in Goris et al (2018 − Journal ofVision) and also in Henaff et al (2020 − Nature Communications). I suggest that the authors rephrase this argument (here and at some later points in the paper). They could just say that the Goris model makes the simplistic and implausible assumption that, within a given trial, gain does not fluctuate. This is clearly an important limitation and the key difference with the continuous model introduced here.

      We sincerely thank the reviewer for identifying this lack of clarity in our original description. We agree that our original description was not sufficiently precise. The modulated Poisson model introduced by Goris et al. (2014) is indeed a generative process model and can be used to generate spike trains, with spiking arising from a Poisson process that may be homogeneous or inhomogeneous within a trial. We apologize for implying otherwise.

      Our intended point was that, in the original formulation, the modulatory gain is represented as a scalar random variable associated with a counting window or trial, and therefore does not explicitly model gain as a continuously time-varying process within a trial. Thus, the key limitation addressed by CMP is not the use of a Poisson process, but the assumption that gain is constant or piecewise constant over the relevant interval.

      In the revised manuscript, we have rephrased the Introduction to clarify this distinction. We now describe the Goris model more accurately as a modulated Poisson framework in which gain is constant over the counting window, and we emphasize that CMP extends this framework by replacing this assumption with a continuous-time stochastic gain process. We have also added discussion of related time-dependent analyses and extensions [Goris et al., 2018, H´enaff et al., 2020], as thoughtfully suggested by the reviewer.

      In addition, we revised the Results and Methods to clarify how the Goris-style baselines were implemented in our comparisons. Specifically, all Goris-style results reported in the main model comparisons use versions with a smoothness prior on the stimulus drive, where the stimulus-dependent firing rates are set to the smooth firing-rate estimates obtained from the Poisson-GP model. This ensures that the comparisons focus on different assumptions about the temporal structure of the gain process, rather than differences in stimulus-drive estimation. We also clarified the comparison to Goris-style variants without this smoothness prior, in which the stimulus-drive parameters are estimated directly under the corresponding Goris-style likelihood (Figure 5 - figure supplement 2). These results show that the smoothness prior on the stimulus drive substantially improves model performance. Finally, we revised the Figure 1 caption and the Goris-model appendix to make these distinctions explicit.

      To address this comment, we revised the Introduction (Pages 2-3: Lines 49-77), Results (Page 9: Lines 263-266, 273-277, Page 11: Lines 319-326), Methods (Page 21), Figure 1 caption, and Goris-model appendix to clarify that CMP extends the Goris framework by modeling gain as a continuously time-varying process within trials.

      (3) Line 54−55: I think the first part of the claim is a bit misleading. There is nothing in the Goris model that would inherently limit it to homogeneous Poisson processes, as seems to be implied by this description. The model is built on the assumption thatspike generation within a trial arises from a Poisson process. This may very well be an inhomogeneous Poisson process (i.e., a stimulus−dependent time−varying firing rate). Homogeneous and inhomogeneous Poisson processes both give rise to Poisson distributed spike counts (and thus a mixture of Poisson distributions across trials in the Goris model). I suggest the authors clarify this description a bit. Note that the two model variants illustrated in Figure 1b and c were also explored in Henaff et al (2020 − Nature Communications).

      We thank the reviewer for this helpful clarification. We agree that the Goris model is not limited to homogeneous Poisson spiking and can incorporate a stimulus-dependent, time-varying firing rate within trials. We did not intend to imply otherwise, and we have revised the relevant text to avoid this misunderstanding.

      Our intended point was that, in formulating continuous-time extensions of the modulated Poisson framework, we explicitly model the time-varying stimulus drive using a smoothness prior, as in the CMP framework, and then consider different assumptions about the temporal structure of the gain process, including constant gain and independently resampled gain across time bins. This highlights the distinction between piecewise-constant gain assumptions and the fully continuous gain process introduced in CMP.

      In the revised manuscript, we have clarified this distinction in the Introduction, Results, and Methods. We now state that the Goris-style variants use stimulus-dependent, time-varying Poisson firing rates, and that the main difference between these variants and CMP lies in the temporal structure assumed for the gain process. We have also acknowledged related variants explored in Goris et al. [2018] and H´enaff et al. [2020], and clarified that our continuous-time formulations of the Goris model differs by imposing a smoothness prior on the stimulus drive. This allows us to estimate a regularized time-varying stimulus component while comparing different assumptions about gain dynamics, ensuring that the comparison focuses on the temporal structure of the gain process rather than differences in stimulus-drive estimation. We also highlight in Figure 5 - figure supplement 2 that even for the Goris-style models, versions that use a smoothness prior on the stimulus drive outperform versions that do not, which are closer to the original modulated Poisson formulation.

      To address this comment, we revised the Introduction (Pages 2-3: Lines 49-77), Results (Page 9: Lines 263-266, 273-277, Page 11: Lines 319-326), Methods (Page 21) to clarify that the Goris-style variants allow stimulus-dependent time-varying firing rates and differ from CMP primarily in their assumptions about gain dynamics. We also added citations to related time-dependent extensions of the modulated Poisson framework.

      (4) The extension to the continuous case is very elegant!

      We thank the reviewer for the positive comment and are pleased that the continuous-time formulation was viewed as elegant.

      (5) I find the result shown in Appendix 3 critically important. The recoverability of the model for realistic amounts of data is foundational for the rest of the paper. I wouldconsider including this analysis in the main results section. Not all readers may check Appendix 3, but they should know about this result.

      We thank the reviewer for emphasizing the importance of this result. We agree that demonstrating parameter recoverability is foundational to the paper and should be visible to readers in the main Results section. In the revised manuscript, we have moved the simulation-based validation from Appendix 3 into the main Results. This section now describes the synthetic CMP dataset, the inference procedure used to estimate the latent stimulus-drive and gain processes, and the comparison between true and inferred GP hyperparameters. These results show that the proposed inference framework can accurately recover the ground-truth stimulus drives, gain processes, and hyperparameters from realistic amounts of simulated data.

      To address this comment, we moved the simulation-based recoverability analysis from Appendix 3 into the main Results section (Page 6: Lines 200-211 and Figure 3).

      (6) Figure 3: I am wondering whether the inferred gain is capturing some response fluctuations that originate from the cell’s phase−selectivity. Could the authors compute the trial−averaged inferred gain (ideally, aligned to stimulus−phase at the start of the trial if this experimental parameter varied across repeats)? If they have successfully partitioned the response variance, the trial−averaged gain should have no systematic temporal structure. If it has a sinusoidal modulation, it may partially capture stimulus−drive. This could be an interesting test to run on all model fits to further validate that the partitioning into a signal and noise component succeeded as intended.

      We thank the reviewer for this insightful suggestion. We agree that verifying that the inferred gain does not capture stimulus-driven structure is an important validation of the model. In the revised manuscript, we have added the trial-averaged inferred gain to Figure 4A for the example neuron. This analysis shows that the trial-averaged inferred gain is relatively flat and neither resembles the inferred stimulus drive nor exhibits clear stimulus-locked temporal structure. This suggests that trial-specific gain fluctuations largely average out across repeats, consistent with the interpretation that the gain process captures random trial-to-trial variability rather than stimulus-driven activity.

      We also note that a direct comparison of this inferred gain trace across methods is not possible for the Goris-style baselines, because these models do not infer a continuous trial-specific gain process. Instead, they marginalize over scalar or time-bin-independent gain variables when computing likelihoods and Fano factor curves. Thus, the trial-averaged gain diagnostic is specific to the CMP model, where the posterior over the continuous-time gain process is explicitly inferred.

      To address this comment, we added the trial-averaged inferred gain to Figure 4A and clarified that it does not show a clear stimulus-locked temporal structure (Page 7: Lines 229-236).

      (7) One common observation that is currently not explored is the quenching of neuronal response variability following stimulus onset (Churchland et al 2010 − NatureNeuroscience), which was suggested to reflect a quenching of gain variability in Goris et al (2024 − Nature Reviews Neuroscience). Building on the previous suggestion, the authors could compute the temporal evolution of cross−trial gain variability from the inferred gain traces. Do they recognize a reduction in gain variability following stimulus onset? If so, it would be worthwhile to show this.

      We sincerely thank the reviewer for this valuable suggestion. We agree that examining whether gain variability decreases following stimulus onset provides an important test of the inferred gain process. In the revised manuscript, we have added an analysis of the temporal evolution of cross-trial gain variability before and after stimulus onset.

      First, in Figure 4A, we now show the cross-trial variance of the inferred gain for the example neuron. This trace shows larger gain variability during the stimulus-off period and a reduction following stimulus onset, suggesting that the inferred gain captures a stimulus-related quenching of trial-to-trial variability. To quantify this effect across the population, we also added a pre- versus post-stimulus comparison in Figure 4C. Following the approach of Churchland et al. [2010], we compared gain variability in two matched 400-ms windows: a pre-stimulus window ending at stimulus onset and a stimulus-period window beginning 100 ms after stimulus onset. For each neuron and stimulus condition, we computed the cross-trial variance of the inferred gain at each time bin, averaged this quantity within each window, and then compared the pre- and post-stimulus values across neuron-stimulus pairs.

      This analysis revealed a significant reduction in inferred gain variability following stimulus onset (one-sided paired Wilcoxon signed-rank test, p≤ 10<sup>−15</sup>), consistent with gain variability quenching [Churchland et al., 2010, Goris et al., 2024]. We now report this result in the main text and illustrate it in Figure 4A and Figure 4C. This provides additional evidence that the inferred CMP gain captures meaningful trial-to-trial variability and its temporal modulation around stimulus presentation.

      To address this comment, we added the cross-trial gain variance trace to Figure 4A and a population-level pre- versus post-stimulus gain-variability quenching analysis (Page 7 and 8: Lines 239-248) in Figure 4C.

      (8) Line 543−565: I want to make sure I understand the Baseline Poisson model and Poisson−GP correctly. For the baseline model, I had imagined that the authors would simply use the stimulus−conditioned PSTH as an estimate of the time−dependent firing rate, coupled with an inhomogeneous Poisson process assumption. But they additionally assume a Gamma prior on the firing rate to compensate for the sparsenessof the data (sometimes only 5 repeats per condition). The Poisson−GP includesexactly the same model components, but now the time−dependent firing rate is modeled by a Gaussian process. Doing this massively improves the goodness−of−fit (Fig 4A). Do I understand this correctly?

      We thank the reviewer for this careful reading. Yes, this understanding is broadly correct, and we have revised the manuscript to clarify the relationships among the Baseline Poisson, Poisson-GP, and Goris-style models. The Baseline Poisson model estimates a stimulus- and time-dependent firing rate independently for each stimulus condition and time bin, using a Gamma-Poisson formulation to regularize the estimate when the number of repeats is limited. The Poisson-GP model uses the same conditionally Poisson observation model, but replaces these independent time-bin-wise rate estimates with a smooth stimulus-specific Gaussian process model for the log firing rate.

      We have also clarified how the Goris-style models were implemented. All Goris-style results reported in the main model comparisons use versions with a GP prior on the stimulus drive. In these versions, the stimulus-dependent firing rates are set to the smooth firing-rate estimates obtained from the PoissonGP model, and the gain parameters are then fit under either the independent-gain or constant-gain assumptions. We used these GP-smoothed versions as stronger baselines. In Figure 5, Figure Supplement 2, we additionally compare these models to Goris-style variants without the GP prior on the stimulus drive, in which the stimulus-drive parameters are estimated directly under the corresponding Goris-style likelihood. This comparison shows that adding a GP smoothness prior to the stimulus drive substantially improves held-out model fit. Together, these analyses clarify that the GP-smoothed stimulus drive improves the Goris-style baselines, while the continuous-time gain process in CMP provides an additional improvement by capturing temporally structured trial-to-trial variability.

      To address this comment, we clarified the definitions of the Baseline Poisson, Poisson-GP, and Goris-style models (Pages 8-9: Lines 255-259, 263-266, 273-277), and revised the text (Page 11: Lines 319-326) describing Figure 4 - figure Supplement 2 to make explicit how this existing comparison isolates the effect of the GP prior on the stimulus drive.

      Reviewer #2 (Public Review):

      Summary:

      Neurons have varied responses to external stimuli that cannot be explained by naive Poisson models. Previous work has quantified and partitioned higher−than−Poisson variability in the brain into different components. The authors improve on these methods to infer how both the stimulus drive and internal gain dynamics impact neuronal variability continuously in time. The clean and well−reasoned model is rigorously developed and then applied to neural data across the visual hierarchy. This lends new insights into how variability is partitioned, agreeing with and extending previous work on how that variability changes from early visual areas (LGN, V1) through to higher, motion−sensitive areas (area MT). Another key contribution is that this partitioning can be fully addressed as a continuous−time process, which allows for the dissection of how the timescale of fluctuations in these two components changesacross the brain’s processing arc.

      Strengths:

      (1) The model is cleanly derived and thoroughly documented, including usable code shared in a GitHub repo. This makes the method immediately portable to other neural systems.

      (2) This is a clear and well−presented piece of work. The figures and writing are clear and understandable, and all pieces of the derivations are included in the main text and supplementary information.

      (3) Comparisons to other models, particularly the one from Goris et al., 2014 shows how this Continuous Modulated Poisson (CMP) model outperforms previous work.

      (4) New insights about how variability partitioning changes across the visual stream from LGN to MT are revealed, including how the gain fluctuates on longer timescales in higher visual areas. Another key result about the anticorrelation between the variance in stimulus drive and gain fluctuations comports with theories about how neurons maintain efficient, reliable encoding.

      (5) In addition to the results reported here, this work will serve as an excellent tutorial for students and postdocs first delving into the sources of variability in the brain.

      We sincerely thank the reviewer for the thoughtful and positive assessment of our work. We are pleased that the model development, empirical analyses, and presentation were viewed as clear, rigorous, and useful for the broader neuroscience community. We also appreciate the reviewer’s recognition that the continuous-time formulation meaningfully extends prior variability-partitioning approaches by allowing stimulus drive and internal gain dynamics to be characterized across temporal scales. The reviewer’s comments helped us further clarify the positioning of the work, expand the Discussion of model scope and limitations, and better articulate future extensions. Below, we address the specific suggestions raised by the reviewer and describe the revisions made in the manuscript.

      Weaknesses:

      The work is somewhat incremental, building on previous studies of the partitioning of variability in the brain, but it provides important new extensions, as noted above.

      Regarding the comment on incremental contribution, we agree that our framework builds directly on previous variability-partitioning approaches, especially the modulated Poisson framework of Goris et al. However, the main goal of this work is to move this class of models from a count-based formulation to a continuous-time spike-train framework. This extension is important because it allows us to model gain as a temporally structured latent process, characterize how variability depends on the timescale over which spikes are counted, and infer the temporal covariance structure of stimulus-independent fluctuations. In addition, the CMP framework provides analytic expressions for the Fano factor as a function of bin size, introduces the EPL covariance function for slowly decaying gain dynamics, and enables direct comparisons of gain amplitude and timescale across visual areas. In the revised manuscript, we have clarified this positioning and emphasized how CMP extends prior variability-partitioning models while preserving their interpretability.

      To address this comment, we revised the Introduction (Pages 3-4: Lines 108-111 and Lines 118-121) and Discussion (Page 13: Lines 425-429) to clarify better how CMP builds on prior variability-partitioning models while extending them to continuous-time spike-train data.

      The only major gap I would suggest addressing in the Discussion is the observation of sub−Poisson variability in the brain. It seems clear that this model can extend to sub− Poisson variability and its partitioning and perhaps even show how that varies in real time, with an animal’s attentional state. That is, of course, beyond the scope of the current work, but could be mentioned in the Discussion.

      We thank the reviewer for this suggestion. We agree that sub-Poisson variability is an important phenomenon observed in neural data. Because the CMP model uses a conditionally Poisson observation model with stochastic gain modulation, it naturally captures Poisson and super-Poisson variability but does not generate sub-Poisson spike count statistics in its current form. In the revised manuscript, we have clarified this limitation in the Discussion and outlined possible extensions that could address sub-Poisson variability, including spike-history terms, renewal-process likelihoods, and alternative count distributions [Truccolo et al., 2005, Paninski et al., 2007, Aghamohammadi et al., 2024]. We also note that such extensions could allow future models to examine how sub-Poisson and super-Poisson components vary with behavioral state, attention, or arousal.

      To address this comment, we added a Discussion paragraph describing the current model’s limitation for sub-Poisson variability and possible extensions to capture it (Page 14: Lines 460-471).

      References

      Stephen Keeley, Mikio Aoi, Yiyi Yu, Spencer Smith, and Jonathan W Pillow. Identifying signal and noise structure in neural population activity with gaussian process factor models. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 13795–13805. Curran Associates, Inc., 2020.

      Robbe L. T. Goris, Corey M. Ziemba, J. Anthony Movshon, and Eero P. Simoncelli. Slow gain fluctuations limit benefits of temporal integration in visual cortex. Journal of Vision, 18(8):8–8, 08 2018. ISSN 1534-7362. doi: 10.1167/18.8.8. URL https://doi.org/10.1167/18.8.8.

      Olivier J H´enaff, Zoe M Boundy-Singer, Kristof Meding, Corey M Ziemba, and Robbe L T Goris. Representation of visual uncertainty through neural gain variability. Nat. Commun., 11(1):2513, May 2020.

      Mark M Churchland, Byron M Yu, John P Cunningham, Leo P Sugrue, Marlene R Cohen, Greg S Corrado, William T Newsome, Andrew M Clark, Paymon Hosseini, Benjamin B Scott, David C Bradley, Matthew A Smith, Adam Kohn, J Anthony Movshon, Katherine M Armstrong, Tirin Moore, Steve W Chang, Lawrence H Snyder, Stephen G Lisberger, Nicholas J Priebe, Ian M Finn, David Ferster, Stephen I Ryu, Gopal Santhanam, Maneesh Sahani, and Krishna V Shenoy. Stimulus onset quenches neural variability: a widespread cortical phenomenon. Nat. Neurosci., 13(3):369–378, March 2010.

      Robbe L T Goris, Ruben Coen-Cagli, Kenneth D Miller, Nicholas J Priebe, and M´at´e Lengyel. Response sub-additivity and variability quenching in visual cortex. Nat. Rev. Neurosci., 25(4):237–252, April 2024.

      Wilson Truccolo, Uri T. Eden, Matthew R. Fellows, John P. Donoghue, and Emery N. Brown. A point process framework for relating neural spiking activity to spiking history, neural ensemble, and extrinsic covariate effects. Journal of Neurophysiology, 93(2):1074–1089, 2005. doi: 10.1152/jn.00697.2004. URL https://doi.org/10.1152/jn.00697.2004. PMID: 15356183.

      Liam Paninski, Jonathan Pillow, and Jeremy Lewi. Statistical models for neural encoding, decoding, and optimal stimulus design. In Paul Cisek, Trevor Drew, and John F. Kalaska, editors, Computational Neuroscience: Theoretical Insights into Brain Function, volume 165 of Progress in Brain Research, pages 493–507. Elsevier, 2007. doi: https://doi.org/10.1016/S0079-6123(06)65031-0. URL https://www.sciencedirect.com/science/article/pii/S0079612306650310.

      Cina Aghamohammadi, Chandramouli Chandrasekaran, and Tatiana A. Engel. A doubly stochastic renewal framework for partitioning spiking variability. bioRxiv, 2024.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting and well-written manuscript in which the authors set out to answer a simple, old question with a modern toolkit: where in crab evolution did sideways walking arise, how often has it been lost or regained, and is it plausibly linked to the ecological and taxonomic success of true crabs. To do this, they record locomotion from 50 live species, convert each species' movements into a quantitative index that compares forward versus sideways bouts, and then map the resulting states onto a recent crab phylogeny to infer the most likely evolutionary history of locomotor direction.

      We thank the reviewer for this positive summary of the study and for recognizing the value of our comparative behavioral dataset and phylogenetic approach.

      Strengths:

      The strongest part of the study is the dataset itself. Comparable behavioral measurements across dozens of crab species are rare. The authors have done the field and husbandry work needed to make this possible. The overall pattern they recover, that most true crabs are strongly biased toward sideways movement (while a smaller set of lineages move predominantly forward), is interesting and likely to be useful to others. The phylogenetic mapping is also a reasonable way to address the "how many times" question (although this is peripheral to my expertise). The manuscript makes a convincing case that sideways locomotion is not simply a trivial byproduct of a crab-like body plan.

      We appreciate the reviewer’s recognition of the dataset and the overall value of the study. We have revised the manuscript to make the conclusions more robust and better aligned with the strength of the evidence.

      (1) Where I am less convinced is in how strongly the authors describe the discreteness of the behavioral categories and the absence of intermediates. The manuscript states that the Forward-Sideways Index shows a clear separation between two locomotor types with little evidence for intermediates, and it cites a statistical test rejecting a single peak in the distribution. However, the histogram in Figure 3 appears structured within each labeled category, with subclusters inside both the forward and sideways groups rather than a single tight peak per group. This matters because the index is built by first placing each movement bout into "forward" versus "sideways" bins using a fixed angle boundary and then collapsing the result into a single ratio. That approach is simple and transparent enough, but it can also hide mixed strategies. For example, a species that produces substantial amounts of both forward and sideways walking can still end up with a strongly positive or negative index, and therefore be classified as a pure "type," even though the underlying behavior is mixed. In that context, rejecting a single peak in the across-species distribution does not, by itself, justify the stronger claim that intermediates are rare or absent.

      Related to this, a key methodological choice is the use of 60 degrees as the cutoff between forward and sideways bouts. This boundary may be reasonable as a convention, but the paper does not explain why it is the right place to draw the line, and there is a plausible biological concern that a fixed angular cutoff does not mean the same thing across taxa.

      Crabs vary in body shape and in how the legs are arranged around the body. In my own comparative work, for example, some species show an elliptical stance pattern elongated along the preferred direction of travel, while others show a more circular leg arrangement, and the latter can express more mixed forward and sideways behavior. When limb arrangement and body geometry differ across species, the same measured angle can correspond to different underlying mechanics and different functional "degree of sidewaysness." The practical implication is that the reported binary separation may partly reflect the imposed classification rule, rather than a sharp biological divide.

      We thank the reviewer for this important point. We agree that the across-species distribution of FSI values alone does not justify a strong statement that intermediate or mixed locomotor tendencies are absent. We also agree that reducing continuous bout-angle distributions to a single index could potentially obscure mixed directional strategies. We have therefore revised the manuscript to avoid implying a strict absence of intermediates and have added an additional analysis of the underlying continuous angle distributions (Abstract, lines 27-29; Results, lines 191-208; Table S2).

      Specifically, we fitted one- and two-component mixture models to the continuous bout-angle distributions of each taxon and examined the supported number of components, peak locations, and mixture weights (Results, lines 196-204; Table S2). This analysis showed that 14 taxa were best described by a one-component model, whereas 36 taxa were best described by a two-component model. Importantly, among the 36 taxa best described by a two-component model, 33 had a dominant component explaining at least 70% of the distribution, whereas only three taxa showed relatively balanced two-component distributions. Thus, although some taxa do show mixed directional tendencies, most taxa are dominated by a primary directional component rather than showing an even mixture of forward and sideways locomotion.

      We also clarified the rationale for the 60° threshold used in the FSI calculation (Methods, lines 125-130). This threshold was not intended to represent a taxon-specific biological boundary between forward and sideways locomotion. Rather, it was used to divide the 360° space into three equal directional sectors: forward, sideways, and backward. This equal partitioning provides a consistent reference under a null expectation of uniformly distributed movement directions.

      To further assess whether our classification depended on the original FSI-based classification, we performed an additional data-driven check based on continuous bout-angle distributions (Results, lines 204-208; Fig. S2). We extracted the dominant peak location from each taxon’s continuous bout-angle distribution (Table S2) and estimated a boundary from the distribution of these dominant peak locations using a Gaussian mixture model. This yielded a data-informed cutoff of approximately 49.4°. The resulting peak-based classification was identical to the original FSI-based classification, with 15 forward-moving and 35 sideways-moving taxa. Thus, no taxon changed category under this independent classification approach.

      (2) Another limitation that affects interpretation is the decision to use one individual per species. I understand the logistics, and for some questions, a single representative individual can be a reasonable first pass. But it is not strong support for negative claims about intermediates, especially in a group where individuals can change substantially with growth and allometry. Crabs can grow dramatically, often with pronounced allometric shifts in limb proportions that can alter the center of mass location. Size alone can alter the kinematics and choice of locomotor behaviors in crustaceans. In species where appendage proportions change with size, or where certain legs become disproportionately large (or calcified), it is plausible that locomotor direction and the distribution of movement angles shift across ontogeny. That makes it hard to treat a single individual as a complete description of a species-level strategy, particularly for species that fall closer to the boundary between categories.

      We thank the reviewer for raising this important limitation. We agree that using one representative individual per species cannot capture the full range of within-species variation, including ontogenetic, size-dependent, or allometric changes in locomotor behavior. We also agree that this limitation is particularly relevant to strong claims about the absence of intermediates.

      As described in our response to Comment #1, we have therefore toned down statements implying a strict absence of intermediates and added analyses of the underlying continuous angle distributions (Abstract, lines 27-29; Results, lines 191-208; Table S2). These additional analyses showed that some taxa do exhibit mixed directional tendencies, although most taxa were dominated by a primary directional component.

      We have also revised the manuscript to clarify the scope of our conclusions. Specifically, we now state that our single-individual sampling design does not capture possible ontogenetic, size-dependent, or allometric variation within species (Methods, lines 105-107). We also clarify that our conclusions are intended to identify broad interspecific patterns in the predominant direction of locomotion across major brachyuran lineages, rather than to describe the full range of locomotor variation within each species (Methods, lines 107-108). Thus, we no longer treat a single individual as providing a complete description of species-level behavioral variation, but instead use it as a standardized representative observation for broad comparative and phylogenetic analyses.

      In sum, this is a valuable and useful behavioral comparative study with a dataset that many in the field will appreciate. The main conclusions about the likely evolutionary placement of sideways walking are plausible, but several of the stronger claims about discrete locomotor types, the absence of intermediates, and the relationship to diversification would be more convincing if the analysis were less dependent on a fixed angular cutoff and on single individuals per species, or if the manuscript framed those points more cautiously so the conclusions track the strength of the evidence.

      We thank the reviewer for this constructive summary and for recognizing the value of our behavioral comparative dataset. We have addressed these concerns in detail in our responses above and revised the manuscript to make the main claims better aligned with the strength of the evidence.

      Reviewer #2 (Public review):

      Summary:

      The current work investigates the evolution of sideward locomotion in Brachyura in light of a single evolutionary origin. To this end, the authors first analysed the mode of locomotion in 50 crab species and observed mutually exclusive presence of sideways vs. forward movement. The phylogenetic analysis confirmed that there is indeed a single evolutionary origin for sideways movement, which was sometimes followed by several reversions to forward locomotion. This way, authors demonstrate how locomotor movement modes shape evolutionary diversification in animals by showing that species richness is much higher in side-ways-moving crabs than in the nearest groups. This is an interesting work that integrates behavioural analysis and phylogenetic relations, capitalising largely on crabs. I have a few suggestions and questions.

      We thank the reviewer for the positive assessment of the study and for recognizing the value of integrating behavioral analysis with phylogenetic relationships. We address the specific suggestions and questions below.

      (1) Firstly, I think the paper spends too much time on a straightforward analysis of the mode of locomotion.

      We agree that the final classification of taxa into predominantly forward- and sideways-moving groups is conceptually simple. However, because our study compares locomotor behavior across a broad range of crab taxa and then uses these behavioral data for phylogenetic reconstruction, we considered it important to describe the behavioral quantification in a transparent and reproducible way. The purpose of this section is therefore not to make a simple endpoint unnecessarily complex, but to show how discrete locomotor states were derived from raw trajectory data using a standardized procedure. For this reason, we retained the current analytical description.

      (2) I was also wondering whether the phylogenetic analysis could be simply achieved by maximising an objective function in which the modes of movement are inversely coded for two putative groups, with all values calculated at all possible nodes.

      The proposed objective-function approach may be useful for identifying a node that best separates two predefined locomotor groups. However, in the present study, we aimed not only to locate a possible boundary between forward- and sideways-moving lineages, but also to reconstruct the evolutionary history of locomotor transitions under an explicit phylogenetic model.

      For this reason, we used standard ancestral state reconstruction and stochastic character mapping rather than maximizing an ad hoc objective function across possible nodes. This approach allowed us to compare alternative transition-rate models (ER and ARD), estimate uncertainty in ancestral states at internal nodes, and quantify the posterior distribution of gains and reversals. We therefore retained the current phylogenetic framework, as it provides a model-based and more informative reconstruction of locomotor evolution across true crabs.

      (3) Unfortunately, I find that the authors did not sufficiently discuss differences in the ecological niches of species with forward vs. sideways locomotion modes (including challenges of locomotion and substrate).

      Likewise, what are the anatomic correlates of forward vs. sideways locomotion? For instance, how are the advantages assumed for sideways movement associated with a flattened body? Is it possible that the mode of motion is secondary to flattened/narrow body structure, which basically limits the distance between legs and thus makes the forward movement difficult - under this logic, the mode of movement would be a secondary phenomenon to body shape traits. How can one differentiate between this alternative and the one that puts the mode of movement in the centre of the story? On a related note, how do different modes of movement relate to the ability to fit into tight spaces - how does it relate to differences in leg joints?

      Is it possible that the sideways movement maximises the scanned visual field per unit time/displacement, which may be beneficial for mostly forward-moving predators?

      We thank the reviewer for this helpful comment. We agree that the previous version did not sufficiently address the possible relationship between locomotor mode and body shape, especially the alternative explanation that sideways locomotion may be secondary to carapace flattening. In response, we added a new morphological analysis using two carapace shape indices: relative carapace length (CL/CW) and relative carapace depth (CD/CS) (Methods, lines 178–185). In the revised Results, we report that relative carapace length differed significantly between forward- and sideways-moving taxa (phylogenetically informed ANOVA: F = 26.90, p < 0.001), whereas relative carapace depth did not differ significantly between the two groups (F = 1.18, p = 0.403) (Results, lines 209–214; Fig. S3). We also added this interpretation to the Discussion, noting that locomotor mode is associated with some aspects of carapace shape but is not explained by simple carapace flattening alone (Discussion, lines 318–325).

      We also revised the Discussion to address the reviewer’s suggestions about possible functional advantages of sideways locomotion beyond rapid bidirectional escape. Specifically, we now mention that other possible advantages may include movement through confined spaces and visual-field sampling during locomotion (Discussion, lines 310-318).

      Finally, we retained the existing discussion of ecological specializations in forward-moving lineages, including coordinated collective movement in soldier crabs, decoration and concealment in majoid crabs, and life inside confined host spaces in pea crabs. This discussion supports the broader point that the adaptive value of sideways locomotion may depend on ecological context.

      (4) It is really difficult to decipher the information contained in the nodes (circles) in the printed black-and-white version of the manuscript.

      We have changed the color scheme and strengthened the outlines of the node pie charts so that the ancestral-state probabilities can be more easily distinguished (Fig. 5). We also applied the same revised color scheme to Figure 3 and Figure S4 for consistency across the manuscript.

      (5) Briefly, although I find the study interesting, the presented complexity may not be necessary given the endpoints; it can be achieved much more simply. Furthermore, the degree to which the conceptual analysis of different modes of locomotion was exercised was limited. The general approach may serve as a good model for the evolutionary analysis of other traits. The demonstration of traceability of the relations in question is a major contribution of the work.

      We thank the reviewer for this constructive summary and for recognizing the broader value of our approach. We have addressed the methodological and conceptual points raised here in our responses to the specific comments above.

      Strengths:

      The research question and the novel combination of different data types.

      We thank the reviewer for highlighting the research question and the novel combination of different data types as strengths of the study. We have revised the manuscript to further strengthen this integrative framework.

      Weaknesses:

      The complexity of the methods used, along with a limited discussion of the potential dynamics that may underlie the evolution of the sideways movement mode.

      We have addressed these concerns in our responses to the specific comments above, particularly by clarifying the rationale for the behavioral quantification and expanding the discussion of morphology, ecological context, and functional hypotheses.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Unbiased analysis of angle distribution. The authors already extract continuous bout angles prior to binning. I recommend using these distributions directly to assess modality at the species level (e.g., unimodal vs bimodal, peak locations, mixture weights) before collapsing behavior into the Forward-Sideways Index. Even a simple circular density estimate or mixture model would clarify whether species classified as "forward" or "sideways" are behaviorally pure or mixed, and would provide a quantitative basis for claims about intermediacy.

      We added mixture-model analyses of the continuous bout-angle distributions, including modality, peak locations, and mixture weights (Results, lines 196-204; Table S2).

      (2) Justify or stress-test the 60° cutoff. The manuscript should either provide a clear biological or data-driven justification for using 60° as the boundary between forward and sideways bouts, or demonstrate that the main conclusions are robust to reasonable alternative cutoffs (e.g., 45°, 75°). A brief sensitivity analysis in the supplement would be sufficient and would greatly strengthen confidence in the classification. Alternatively (my preference) would be to let the data inform the cutoff.

      We clarified the rationale for the 60° sector definition used to calculate FSI (Methods, lines 125-130; Fig. 2) and added an independent data-driven boundary analysis based on dominant peak locations, which yielded the same forward/sideways classification (Results, lines 204-208; Fig. S2).

      (3) Sampling justification. I recommend explicitly acknowledging that sampling a single individual per species limits the ability to detect ontogenetic, size-dependent, or allometric variation in locomotor strategy. If feasible, adding even limited replication across size classes or individuals for a small subset of taxa (particularly those near the classification boundary) would substantially strengthen the conclusions; otherwise, the manuscript should more clearly delimit which claims do and do not rely on the assumption of within-species invariance.

      We clarified that our conclusions concern broad interspecific patterns of predominant locomotor direction, rather than the full range of within-species variation (Methods, lines 102-108).

      (4) I would suggest toning down or reframing statements about "no intermediates". If additional analyses are not added, I recommend revising statements that imply a strict absence of intermediates to language that reflects what is directly shown (e.g., bimodality in an index derived from binned data). This would better align the claims with the current evidence.

      We revised the manuscript to avoid implying a strict absence of intermediates and now acknowledge that some taxa show mixed directional tendencies (Abstract, lines 27-29; Results, lines 191-208).

      (5) Framing and claims about diversification. The discussion of sideways locomotion as a key innovation would benefit from clearer separation between observed correlations and causal inference. If trait-dependent diversification analyses are not added, I suggest consistently framing this section as a hypothesis supported by comparative patterns rather than a demonstrated mechanism.

      We revised the Discussion to more clearly frame sideways locomotion as a possible key innovation associated with diversification, rather than as a demonstrated causal mechanism (Abstract, lines 31-34; Discussion, lines 294-309).

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors perform an analysis of the relationship between the size of an LMM and the predictive performance of an ECoG encoding model made using the representations from that LMM. They find a logarithmic relationship between model size and prediction performance, consistent with previous findings in fMRI. They additionally observe that as the model size increases, the location of the "peak" encoding performance typically moves further back into the model in terms of percent layer depth, an interesting result worthy of further analysis into these representations.

      Strengths:

      The evidence is quite convincing, consistent across model families, and complementary to other work in this field. This sort of analysis for ECoG is needed and supports the decade-long enduring trend of the "virtuous cycle" between neuroscience and AI research, where more powerful AI models have consistently yielded more effective predictions of responses in the brain. The lag analysis showing that optimal lags do not change with model size is a nice result using the higher temporal resolution of ECoG compared to other methods like fMRI.

      We thank the reviewer for their thoughtful assessment! We agree that the “virtuous cycle” between neuroscience and AI research has been, and will continue to be, a driving force in advancing our understanding of brain function through more powerful predictive models. We are especially pleased that the reviewer appreciated the lag analysis, as we view this as a valuable complement to the existing fMRI work.

      Weaknesses:

      I would have liked to have seen the data scaling trends explored a bit too, as this is somewhat analogous to the main scaling results. While better performance with more data might be unsurprising, showing good data scaling would be a strong and useful justification for additional data collection in the field, especially given the extremely limited amount of existing language ECoG data. I realize that the data here is somewhat limited (only 30 minutes per subject), but authors could still in principle train models on subsets of this data.

      We thank the reviewer for their valuable suggestion. For the revised manuscript, we performed a new analysis where we trained encoding models using subsets of the data (randomly sampling contiguous chunks of 50%, 25%, and 10% of all words in each of the training folds) and tested these models on all words in the test fold. As expected, we found that encoding performance increases as the training dataset size increases, suggesting that model performance scales with data quantity even within the constraints of our relatively small dataset. This result reinforces the importance of collecting dense ECoG data. We have added the following text to our Results section: “We also built encoding models using subsets of the data and found that encoding performance increases as the volume of training data increases (Fig. S6)” and included the results as a supplementary figure 6 in the revised manuscript.

      Separately, it would be nice to have better justification of some of these trends, in particular the peak layerwise encoding performance trend and the overall upside-down U-trend of encoding performance across layers more generally. There is clearly something very fundamental going on here, about the nature of abstraction patterns in LLMs and in the brain, and this result points to that. I don't see the lack of justification here as a critical issue, but the paper would certainly be better with some theoretical explanation for why this might be the case.

      We thank the reviewer for this insightful comment. The general inverted U-shaped trend of encoding performance across layers has been a frequently observed phenomenon in studies comparing LLM representations to brain activity (Goldstein, Ham, et al., 2025; Schrimpf et al., 2021). A potential explanation is the existence of a “two-phase abstraction process” within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the initial layers, models begin by processing relatively low-level input features. As layers get deeper, representations become increasingly abstract and richly contextualized in semantic features relevant for understanding language. These intermediate layers often show the highest correlation with brain activity in language areas, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. These observations suggest that it is primarily the abstractive, contextual features developed in the intermediate layers of LLMs that drive their alignment with brain activity. As models become more potent at prediction, their most predictive layers (for the LLM’s natural language task) and their most generalizable layers (for brain activity) can diverge.

      A key finding in our study is that the initial processing phase does not scale and take up more layers as models scale up in size and layers. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Consequently, the prediction phase may begin relatively earlier in these larger models, and the later layers could develop highly specialized representations that are increasingly divergent from the more general linguistic processing captured in brain activity. For example, these layers may specialize in capturing very specific patterns of language (thus lowering their perplexity) that do not actually occur often or at all in our naturalistic dataset. It is also possible that the later layers of larger models are overall underutilized and do not contribute as much to linguistic processing and next-word prediction (Csordás et al., 2025).

      We have added the following text to our Discussion section:

      “The inverted U-shaped trend of encoding performance commonly found in previous research is likely due to a "two-phase abstraction process" within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the early and intermediate layers of the model, a composition phase occurs, where low-level input features become increasingly abstract and contextualized. The intermediate layers of the model show the highest correlation with brain activity, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers of the model, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. Our results indicate that the initial composition phase does not take up more layers as models scale up in size. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Thus, as LLMs increase in size, the later layers of the model may contain representations that are increasingly divergent from the more general linguistic processing captured in brain activity. It is also possible that the later layers of larger models are overall underutilized and may not significantly contribute to benchmark performances during inference (Csordás et al., 2025; Fan et al., 2024; Gromov et al., 2024).”

      Lastly, I would have wanted to see a similar analysis here done for audio encoding models using Whisper or WavLM as this is the modality where you might see real differences between ECoG and other slower scanning approaches. Again, I do not see this omission as a fundamental issue, but it does seem like the sort of analysis for which the higher temporal resolution of ECoG might grant some deeper insight.

      We appreciate this suggestion. In a separate project, we focused on multimodal audio-to-speech-to-language large language models (LLMs), building encoding models using Whisper embeddings (from both the encoder and decoder stacks) to predict electrocorticographic (ECoG) signals during naturalistic conversations (Goldstein, Wang, et al., 2025). The higher temporal resolution of ECoG enables us to trace the temporal flow of information from the superior temporal gyrus (STG) and somatomotor areas (SM) to the inferior frontal gyrus (IFG) during speech comprehension. Conversely, during speech production, encoding in IFG peaked significantly earlier than in the STG and SM. We agree that scaling encoding models using multimodal approaches and our ECoG conversation datasets could yield valuable insights, and we look forward to exploring this in future work. However, we feel that the added complexity of multimodal encoding models falls beyond the scope of this paper.

      We have modified the following text to our Discussion section:

      “Since we exclusively employ textual LLMs, which lack inherent temporal information due to their discrete token-based nature, future studies utilizing multimodal LLMs integrating continuous audio or video streams, like Whisper or WavLM may better unravel the relationship between model size and temporal dynamic representations in LLMs (Goldstein, Wang, et al., 2025; Millet et al., 2023; Vaidya et al., 2022).”

      Reviewer #2 (Public review):

      Summary:

      This paper investigates whether large language models (LLMs) of increasing size more accurately align with brain activity during naturalistic language comprehension. The authors extracted word embeddings from LLMs for each word in a 30-minute story and regressed them against electrocorticography (ECoG) activity time-locked to each word as participants listened to the story. The findings reveal that larger LLMs more effectively predict ECoG activity, reflecting the scaling laws observed in other natural language processing tasks.

      Strengths:

      (1) The study compared model activity with ECoG recordings, which offer much better temporal resolution than other neuroimaging methods, allowing for the examination of model encoding performance across various lags relative to word onset.

      (2) The range of LLMs tested is comprehensive, spanning from 82 million to 70 billion parameters. This serves as a valuable reference for researchers selecting LLMs for brain encoding and decoding studies.

      (3) The regression methods used are well-established in prior research, and the results demonstrate a convincing scaling law for the brain encoding ability of LLMs. The consistency of these results after PCA dimensionality reduction further supports the claim.

      We thank the reviewer for their thoughtful and positive feedback.

      Weaknesses:

      (1) Some claims of the paper are less convincing. The authors suggested that "scaling could be a property that the human brain, similar to LLMs, can utilize to enhance performance", however, many other animals have brains with more neurons than the human brain, making it unlikely that simple scaling alone leads to better language performance.

      We thank the reviewer for this insightful comment. We agree that simply having more neurons does not automatically confer more complex or human-like cognitive or linguistic capabilities. This suggestion deserves a more nuanced treatment than we had included in the original manuscript.

      Research in comparative neuroscience has argued that human cognitive abilities emerge from scaling up the primate brain (Herculano-Houzel, 2012). However, the critical aspect is not merely the number of neurons, but how these neurons contribute to computational power within a specific evolutionary and cultural context. The uniqueness of human cognition has been argued to result from a global adaptation for increased information processing capacity (Cantlon & Piantadosi, 2024). Moreover, the language network in humans is likely grounded in the evolution of particular structural networks in the primate brain (Friederici & Becker, 2025). This suggests that the way brain regions are connected and the expansion of certain pathways are critical, not just the overall scale. Furthermore, the specialized structure must be tuned by its learning environment and training data. For example, both humans and LLMs learn from language data generated by other humans, which reflects world knowledge that has accumulated over many generations.

      We have modified the following text in the Introduction:

      “Research in comparative neuroscience has suggested that uniquely human cognitive abilities emerge from scaling up the primate brain (Herculano-Houzel, 2012).”

      We also added a caveat to the Discussion on this point:

      “As in the human brain, while scaling alone may yield emergent cognitive abilities (Cantlon & Piantadosi, 2024; Herculano-Houzel, 2012), specialized architectural features likely also play a critical role (Friederici & Becker, 2025).”

      Additionally, the authors claim that their results show 'larger models better predict the structure of natural language.' However, it remains unclear to what extent the embeddings of LLMs capture the "structure" of language better than the lexical semantics of language.

      We appreciate the reviewer's point about how well LLM embeddings capture the "structure" of language versus just lexical semantics. It's true that distinguishing these aspects is complex. From our perspective, a model's ability to predict/produce natural language entails that the model captures various levels of linguistic structure, including morphology, syntax, semantics, and contextual dependencies. We use "structure" inclusively in this sense. A model cannot achieve high predictive accuracy without representing, to some extent, all of these structural elements (Linzen & Baroni, 2021; Manning et al., 2020; Pavlick, 2022). There is a very active field of research into understanding exactly how these models represent these different structures of language (Ameisen et al., 2025; Chemla et al., 2024; Elhage et al., 2021, 2022; Hewitt & Manning, 2019). Our results confirm the core trend that larger models tend to better reproduce the various structures of language (i.e., yield lower perplexity; Fig. 2A).

      In previous work, we have shown that LLM embeddings better predict neural activity during natural language processing than lexical embeddings (e.g., GloVe) that do not contain other elements of linguistic structure (Goldstein et al., 2022; Kumar et al., 2024; Zada et al., 2024). In response to the following comment, we also compare LLMs to simpler models capturing specific speech and language features (see next comment). To clarify our intended use of the word “structure”, we’ve added a brief explanation in the Methods section:

      “In this study, we use the term “structure” to refer to a variety of linguistic patterns (e.g., morphology, syntax, semantics, context) that LLMs encode in order to better predict natural language.”

      (2) The study lacks control LLMs with randomly initialized weights and control regressors, such as word frequency and phonetic features of speech, making it unclear what the baseline is for the model-brain correlation.

      We’ve added several supplementary analyses to the revised manuscript to address these concerns. To establish a baseline, we extracted embeddings from each layer of the SMALL model with randomly initialized weights and constructed encoding models. The encoding performance is significantly higher for pretrained SMALL than for untrained SMALL for every layer (Fig. S4). For the untrained model, the performance is the highest for the 0th layer and decreases in subsequent layers. This is because at the 0th layer, every instance of the same word receives an identical, albeit random, embedding (See Supplementary Figure 4).

      We also compared the encoding performance of LLMs with more classical speech/language features and static GloVe embeddings (Goldstein, Wang, et al., 2025; Kumar et al., 2024). First, we extracted features capturing lower-level speech features. Using the stimulus transcript as input, we created one-hot vectors for phonetic and articulatory features. Phoneme classes (39 total classes) were obtained from the Carnegie Mellon Pronouncing Dictionary (The CMU Pronouncing Dictionary, n.d.). We further classified the phonemes based on their place of articulation (9 classes), manner of articulation (9 classes), and voiced or voiceless status (3 classes), according to the general American English consonants of the International Phonetic Alphabet. Given that each word consists of multiple phonemes, we averaged the one-hot vectors for all phonetic and articulatory features for each word.

      Second, we extracted linguistic features using spaCy (Honnibal et al., 2020), including part of speech (17 classes), tag (50 classes), function or content word (3 classes), dependency (45 classes), whether the word is an alpha character (binary), and whether the word is a stop word (binary). We also extracted prefix (30 classes) and suffix (44 classes) information using the Cambridge Dictionary. We constructed one-hot vectors for each multi-class feature and one-dimensional vectors for each binary feature.

      Third, for each word, we obtained word frequency from the Google Web Trillion Word Corpus (Brants & Franz, 2006) and from our own dataset.

      Fourth, we generated static word embeddings of dimension 50 using GloVe (Pennington et al., 2014).

      We then built encoding models in the same way as the contextual embeddings for each of the three categories of speech features, all speech features concatenated, and the GloVe embeddings. To control for the different dimensions of the embeddings, we also standardized all embeddings to the same size (50 dimensions) using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression. For both ridge and OLS encoding, our contextual embeddings from LLMs showed significantly better performance than the classic speech features and GloVe embeddings.

      We have added the following text to our manuscript and updated our Figures S4, S5, Table S1, and the methods section:

      “To establish a general baseline for encoding performance, we built encoding models using embeddings from the SMALL model with randomly initialized weights. The trained SMALL model exhibits significantly higher encoding performance across all layers compared to the untrained SMALL model (Fig. S4). We also assessed the encoding performance of contextual embeddings from LLMs against classic speech features and static GloVe embeddings (Table S1). The SMALL and XL embeddings achieved markedly higher encoding correlations than the speech features and GloVe embeddings (Fig. S5).”

      (3) The finding that peak encoding performance tends to occur in relatively earlier layers in larger models is somewhat surprising and requires further explanation. Since more layers mean more parameters, if the later layers diverge from language processing in the brain, it raises the question of what aspects of the larger models make them more brain-like.

      We thank the reviewer for this insightful comment; this point was also highlighted by Reviewer 1. We agree that this result is somewhat surprising, and we aim to provide a more detailed explanation in the revised manuscript. The general inverted U-shaped trend of encoding performance across layers has been a frequently observed phenomenon in studies comparing LLM representations to brain activity (Goldstein, Ham, et al., 2025; Schrimpf et al., 2021). A potential explanation is the existence of a “two-phase abstraction process” within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the initial layers, models begin by processing relatively low-level input features. As layers get deeper, representations become increasingly abstract and richly contextualized in semantic features relevant for understanding language. These intermediate layers often show the highest correlation with brain activity in language areas, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. These observations suggest that it is primarily the abstractive, contextual features developed in the intermediate layers of LLMs that drive their alignment with brain activity. As models become more potent at prediction, their most predictive layers (for the LLM’s natural language task) and their most generalizable layers (for brain activity) can diverge.

      A key finding in our study is that the initial processing phase does not scale and take up more layers as models scale up in size and layers. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models.

      Consequently, the prediction phase may begin relatively earlier in these larger models, and the later layers could develop highly specialized representations that are increasingly divergent from the more general linguistic processing captured in brain activity. For example, these layers may specialize in capturing very specific patterns of language (thus lowering their perplexity) that do not actually occur often or at all in our naturalistic dataset. It is also possible that the later layers of larger models are overall underutilized and do not contribute as much to linguistic processing and next-word prediction (Csordás et al., 2025).

      We have added the following text to our Discussion section:

      “The inverted U-shaped trend of encoding performance commonly found in previous research is likely due to a "two-phase abstraction process" within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the early and intermediate layers of the model, a composition phase occurs, where low-level input features become increasingly abstract and contextualized. The intermediate layers of the model show the highest correlation with brain activity, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers of the model, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. Our results indicate that the initial composition phase does not take up more layers as models scale up in size. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Thus, as LLMs increase in size, the later layers of the model may contain representations that are increasingly divergent from the more general linguistic processing captured in brain activity. It is also possible that the later layers of larger models are overall underutilized and may not significantly contribute to benchmark performances during inference (Csordás et al., 2025; Fan et al., 2024; Gromov et al., 2024).”

      Reviewer #3 (Public review):

      This manuscript studies the connection between neural activity collected through electrocorticography and hidden vector representations from autoregressive language models, with the specific aim of studying the influence of language model size on this connection. Neural activity was measured from subjects who listened to a segment from a podcast, and the representations from language models were calculated using the written transcription as the input text. The ability of vector representations to predict neural activity was evaluated using 10-fold cross-validation with ridge regression models.

      The main results are that (as well summarized in section headings):

      (1) Larger models predict neural activity better.

      (2) The ability of language model representations to predict neural activity differs across electrodes and brain regions.

      (3) The layer that best predicts neural activity differs according to model size, with the "SMALL" model showing a correspondence between layer number and the language processing hierarchy.

      (4) There seems to be a similar relationship between the time lag and the ability of language model representations to predict neural activity across models.

      Strengths:

      (1) The experimental and modeling protocols generally seem solid, which yielded results that answer the authors' primary research question.

      (2) Electrocorticography data is especially hard to collect, so these results make a nice addition to recent functional magnetic resonance imaging studies.

      We thank the reviewer for their thoughtful and positive feedback.

      Weaknesses:

      (1) The interpretation of some results seems unjustified, although this may just be a presentational issue.

      (a) Figure 2B: The authors interpret the results as "a plateau in the maximal encoding performance," when some readers might interpret this rather as a decline after 13 billion parameters. Can this be further supported by a significance test like that shown in Figure 4B?

      We agree that this could be a subjective interpretation, so we conducted an additional analysis. We performed paired two-sided t-tests between best layer encoding performances averaged across electrodes (df = 159 electrodes), comparing all models with larger models. We found that after 13 billion parameters, only the encoding performance for OPT-66B, the largest model in the OPT family, is significantly worse than the encoding performance of some other smaller models, supporting the claim that the maximal encoding performance declines after 13 billion parameters. However, we did not find conclusive statistical evidence of a decline in encoding performance for other model families.

      We have added the statistical results as Supplementary Figure 1.

      We have also modified the following text in the manuscript:

      “We also observed a plateau in the maximal encoding performance, occurring around 7 billion parameters (Fig. 2B), with a decline in performance for the OPT-66B model (Fig. S1).”

      (b) Figure S1A: It looks like the drop in PCA max correlation is larger for larger models, which may suggest to some readers that the same trend observed for ridge max correlation may not hold, contra the authors' claim that all results replicate. Why not include a similar figure as Figure 2B as part of Figure S1?

      PCA is an unsupervised dimensionality reduction technique and may discard model features with small eigenvalues that nonetheless contribute to encoding performance. Ridge regression, a supervised method, can capitalize on these features. We suspect that this is why there appears to be a larger drop in model performance for larger models with PCA than with ridge regression. We replicated the logarithmic relationship between model size and encoding performance using PCA and ordinary least-squares (OLS) regression encoding models. We have updated Supplementary Figure 2.

      (2) Discussion of what might be driving the main result about the influence of model size appears to be missing (cf. the authors aim to provide an explanation of what seems to drive the influence of the layer location in Paragraph 3 of the Discussion section). What explanations have been proposed in the previous functional magnetic resonance imaging studies? Do those explanations also hold in the context of this study?

      We suspect that the increased expressivity of larger models - that is, their improved sensitivity to nuanced structure in natural language - yields improved alignment to brain activity (given large enough samples of brain activity) (Antonello et al., 2023). This effect persists even when dimensionality is tightly controlled in our PCA-based analysis, indicating that the improved alignment with the brain is not a modeling artifact of dimensionality alone, but results from the structural representations learned by these larger models.

      We have added the following text to our Discussion section:

      “We suspect that the improved alignment with brain activity in larger models is driven by their increased expressivity and sensitivity to nuanced linguistic structure present in large-scale naturalistic datasets (Antonello et al., 2023).”

      (3) The GloVe-based selection of language-sensitive electrodes (at least to me) isn't explained/motivated clearly enough (I think a more detailed explanation should be included in the Materials and Methods section). If the electrodes are selected based on GloVe embeddings, then isn't the main experiment just showing that representations from larger language models track more closely with GloVe embeddings? What justifies this methodology?

      We selected electrodes based on previously established methods (Goldstein et al., 2022). Our use of GloVe embeddings for electrode selection does not imply that larger language model representations are simply more closely aligned with GloVe embeddings. On the contrary, contextual embeddings from LLMs, which incorporate the word’s previous context, consistently outperform static embeddings like GloVe or word2vec (Fig. S3). Selecting electrodes using LLM embeddings would likely result in a slightly different, potentially larger set of electrodes (Goldstein et al., 2022), but would be more circular (Kriegeskorte et al., 2009). The GloVe-based electrode selection represents a more conservative approach by identifying words encoding linguistic content without biasing the selection directly toward any LLMs.

      We have added the following text to our Method section:

      “We used GloVe embeddings for electrode selection to avoid biasing our main results toward a particular LLM.”

      (4) (Minor weakness) The main experiments are largely replications of previous functional magnetic resonance imaging studies, with the exception of the one lag-based analysis. Is there anything else that the electrocorticography data can reveal that functional magnetic resonance imaging data can't?

      We thank the reviewer for this thoughtful question. While we agree that a key contribution of our work corroborates previous fMRI findings, we would argue that using ECoG is not merely a replication but a crucial validation and extension of that work. It is important to validate these effects across distinct measurement modalities. In our work, we further observed a novel trend where the peak encoding performance tends to occur in relatively earlier layers for larger models. This is supported by recent studies suggesting that later layers of large LLMs may not significantly contribute to benchmark performance (Csordás et al., 2025). While scaling has been an effective method to improve LLM performance, including in encoding models, future research should explore the potential underutilization of the later layers as models scale.

      Furthermore, ECoG data offers temporal resolution on the order of milliseconds, far superior to fMRI’s. Although we did not observe a relationship between model size and temporal lags in this study, future work should investigate the temporal dynamics of encoding that are accessible with ECoG (Goldstein, Ham, et al., 2025; Goldstein, Wang, et al., 2025).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Thank you to the authors for the fun and personally useful read.

      I see in Supplementary Figure 1 the authors show a comparison of the performance between OLS vs. Ridge regression. Is the OLS model the only one that is working over PC features, or are both models using PC features? The current text is a bit unclear. My current understanding is that the comparison is between (OLS + PCA) and (Ridge with no PCA), but I am not sure.

      The OLS model is the only one that works over PC features, following previous methods (Goldstein et al., 2022).

      We have added the following text to our Results and Methods section for clarity:

      “To control for the different embedding dimensionality across models, we standardized all embeddings to the same size using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression, replicating the logarithmic relationship but with significantly lower encoding performance overall (Fig. S2). The PC features are used by the OLS models only.”

      Clarification in the text would be appropriate. If this is the correct understanding, the authors should note in the main text that the ridge approach is more effective than the PCA approach, which is still the dominant approach to building linear encoding models in the field for some unjustifiable reason.

      We thank the reviewer for pointing out the confusion. We have updated Supplementary Figure 2.

      How were the alpha values for ridge regression determined? Do you use the same ridge parameter for all electrodes or fit a different parameter for each electrode? This is not mentioned anywhere.

      The alpha values are determined by cross-validation using the “RidgeCV” method from the “himalaya” package (Dupré la Tour et al., 2022). Specifically, we perform a grid search over cross-validation folds in the training data to find the best-performing alpha. The alpha parameter is specific to each ridge regression model, meaning each fold, lag, and electrode has a different alpha parameter.

      We have added the following text to our manuscript:

      “For each ridge regression model (for each fold, lag, and electrode), the alpha parameter is determined by cross-validation using the “RidgeCV” method from the “himalaya” package (Dupré la Tour et al., 2022).”

      It's not entirely clear to me how the authors handle tokens that do not terminate in words (such as the "there" + "'s" example in the text). My current reading of the text is that authors essentially ignore these half-word embeddings, doing one forward pass per word, rather than per token, but the current text is somewhat ambiguous.

      If a word is tokenized into several tokens, like “there” and “‘s”, we average the token embeddings to get a word embedding.

      We have added the following text to our Method section:

      “To facilitate a fair comparison of the encoding effect across different models, we aligned all tokens in the story across all models. We averaged the token embeddings if a word is split into multiple tokens, resulting in one embedding per word for each model.”

      The authors describe the scaling relationship they find as a "log-linear" relationship. I believe this is a misnomer derived from the original paper describing this relationship in fMRI as log-linear (Antonello et al.) The correct term is simply "logarithmic", and for what it's worth, the authors of the original fMRI work have made this correction as well.

      Thank you! We have made this correction.

      Is the data publicly available? If not, there should be some basic justification as to why (consent reasons, etc.).

      We have recently made the data publicly available (Zada et al., 2025). We have also provided tutorials for preprocessing the data and training encoding models: https://hassonlab.github.io/podcast-ecog-tutorials. For this specific project, the analysis code is available at https://github.com/hassonlab/247-pickling/tree/scaling-paper-0 and https://github.com/hassonlab/247-encoding/tree/scaling-paper-1.

      The authors assert that ECoG has "superior spatiotemporal resolution". While this is unquestionably true for temporal resolution, the story is a bit more complicated for spatial resolution, where ECoG has far less cortical coverage than fMRI. Perhaps this sentence should be revised.

      Thank you for pointing out the typo! We have changed it to “superior temporal resolution”.

      Minor Points:

      The bolded title of Figure 3 probably shouldn't be bolded, as this is just actually the title of Figure 3A.

      Fixed.

      Figure 4d is has a typo: "Best Encoidng Layer".

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      The authors could consider adding control regressors such as word rate, word frequency, phonetic features, and syntactic features like node counts, as well as control LLMs of comparable size to serve as baselines. The authors could also include correlation analyses of the embeddings from different layers of the same LLM to further illustrate how distinct the layers are within the models.

      We have added untrained LLM embeddings as a baseline and included a comparison of encoding models between LLM contextual embeddings and classical speech features. We have also performed some preliminary correlation analyses of embeddings. In some models, we found evidence of the “two-phase abstraction process” (Cheng & Antonello, 2024). However, the result is inconclusive across different LLM families. Since each LLM layer accesses and modifies the residual stream (Elhage et al., 2021), the embeddings across layers are inherently correlated. Future work could instead explore the isolated transformations within each layer to illustrate the distinct information across layers (Kumar et al., 2024).

      The analysis codes and data should be made available.

      We have recently made the data publicly available (Zada et al., 2025). We have also provided tutorials for preprocessing the data and training encoding models: https://hassonlab.github.io/podcast-ecog-tutorials. For this specific project, the analysis code is available at https://github.com/hassonlab/247-pickling/tree/scaling-paper-0 and https://github.com/hassonlab/247-encoding/tree/scaling-paper-1.

      Reviewer #3 (Recommendations for the authors):

      Most of my concrete recommendations are in the public review. Below are some additional minor ones:

      (1) Introduction: "Remarkably, these models learn from much the same shared space as humans: from real-world language generated by humans."

      I think this is an extremely strong claim due to e.g. the different nature of child-directed speech vs. written text corpora, the lack of multimodality and grounding in language models, etc. I might suggest re-wording this sentence or removing it entirely.

      We thank the reviewer for their suggestion! We have removed the sentence from the manuscript.

      (2) Introduction: "EleutherAI, n.d." reference for GPT-Neo

      GPT-NeoX-20B has an associated paper, which the authors might cite instead: https://aclanthology.org/2022.bigscience-1.9

      Thank you! We have added the reference for GPT-NeoX-20B (Black et al., 2022).

      (3) Figure 4D: Encoidng -> Encoding

      Fixed.

      (4) Materials and Methods, Contextual embeddings: "except for GPT-Neox-20b, which assigns additional tokens to whitespace characters."

      What do the authors mean by "additional tokens to whitespace characters?" The tokenizer for GPT-NeoX-20B works in much the same way as that of GPT-Neo, just with a different vocabulary set.

      >>> t1 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-125M")

      >>> t2 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")

      >>> t1.convert_ids_to_tokens(t1("The quick brown fox jumps over the lazy dog.").input_ids) ['The', 'Ġquick', 'Ġbrown', 'Ġfox', 'Ġjumps', 'Ġover', 'Ġthe', 'Ġlazy', 'Ġdog', '.']

      >>> t2.convert_ids_to_tokens(t2("The quick brown fox jumps over the lazy dog.").input_ids)

      ['The', 'Ġquick', 'Ġbrown', 'Ġfox', 'Ġjumps', 'Ġover', 'Ġthe', 'Ġlazy', 'Ġdog', '.']

      If the authors are referring to Ġ as the "additional token to whitespace characters," then these are in all other tokenizers as well (not only that for GPT-Neo, but also those for GPT-2 and OPT).

      We agree that “additional tokens to whitespace characters” is an oversimplification. The GPT-Neo model family, which includes the 125M, 1.3B, and 2.7B models, utilizes the same Byte Pair Encoding (BPE) tokenizer as GPT-2. This common tokenizer has a vocabulary size of 50,257 tokens, providing compatibility and seamless integration across the models.

      The GPT-NeoX-20B model introduces a modified tokenizer to address limitations observed in the GPT-2 tokenizer (Black et al., 2022). As detailed in Section 3.2, this new tokenizer incorporates a few key improvements:

      (1) New BPE tokenizer: A more general-purpose BPE tokenizer was trained using the Pile dataset.

      (2) Space Delimitation: Unlike the GPT-2 tokenizer, which treats tokenization at the start of a string as a non-space-delimited token, the GPT-NeoX-20B tokenizer applies consistent space delimitation regardless. This change resolves inconsistencies related to the presence of prefix spaces in the tokenization input.

      (3) Whitespace Handling: The tokenizer includes tokens for repeated space characters (up to 24 consecutive spaces), enhancing efficiency in tokenizing text with substantial whitespace, such as program source code or LaTeX documents.

      These modifications result in the GPT-NeoX-20B tokenizer representing the Pile validation set with approximately 10% fewer tokens than the GPT-2 tokenizer. This efficiency gain is particularly beneficial for processing texts with extensive whitespace.

      In our analysis, we extracted embeddings by setting `add_prefix_space = True` to all tokenizers, so space delimitation does not result in tokenizer differences. We highlight here examples of the other two tokenizer differences using the Huggingface `AutoTokenizer`:

      >>> t1 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-125M")

      >>> t2 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")

      >>> t1.convert_ids_to_tokens(t1("The Downing Street.").input_ids) ['The', 'ĠDowning', 'ĠStreet']

      >>> t2.convert_ids_to_tokens(t2("The Downing Street.").input_ids)

      ['The', 'ĠDown', 'ing', 'ĠStreet']

      >>> t1.convert_ids_to_tokens(t1("Hello !").input_ids)

      ['Hello', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ!']

      >>> t2.convert_ids_to_tokens(t2("Hello !").input_ids)

      ['Hello', ' ', '!']

      More examples showing the differences between the GPT-2 tokenizer and the GPT-NeoX-20B tokenizer can be found in Appendix F: Tokenizer Analysis (Black et al., 2022).

      We have added the following text to our manuscript for simplicity:

      “All models within the same model family adhere to the same tokenizer convention, except for GPT-Neox-20B, which utilizes a different tokenizer (Black et al., 2022).”

      References

      Ameisen, E., Lindsey, J., Pearce, A., Gurnee, W., Turner, N. L., Chen, B., Citro, C., Abrahams, D.,  Carter, S., Hosmer, B., Marcus, J., Sklar, M., Templeton, A., Bricken, T., McDougall, C.,  Cunningham, H., Henighan, T., Jermyn, A., Jones, A., … Batson, J. (2025). Circuit Tracing:  Revealing Computational Graphs in Language Models. Transformer Circuits Thread. https://transformer-circuits.pub/2025/attribution-graphs/methods.html

      Antonello, R., & Huth, A. (2024). Predictive coding or just feature discovery? An alternative account of why language models fit brain data. Neurobiology of Language (Cambridge, Mass.), 5(1), 64–79.

      Antonello, R., Vaidya, A., & Huth, A. G. (2023). Scaling laws for language encoding models in fMRI. NeurIPS 2023. https://doi.org/10.48550/ARXIV.2305.11863

      Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell,  K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., & Weinbach, S. (2022). GPT-NeoX-20B: An Open-Source Autoregressive Language Model.  Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models. Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models, virtual+Dublin. https://doi.org/10.18653/v1/2022.bigscience-1.9

      Brants, T., & Franz, A. (2006). Web 1T 5-gram Version 1 [Dataset]. Linguistic Data Consortium. https://doi.org/10.35111/CQPA-A498

      Cantlon, J. F., & Piantadosi, S. T. (2024). Uniquely human intelligence arose from expanded information capacity. Nature Reviews Psychology, 3(4), 275–293.

      Chemla, E., D’Ascoli, S., Diego-Simón, P., King, J.-R., & Lakretz, Y. (2024). A Polar coordinate system represents syntax in large language models. Advances in Neural Information Processing Systems 37, 105375–105396.

      Cheng, E., & Antonello, R. J. (2024). Evidence from fMRI supports a two-phase abstraction process in language models. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2409.05771

      Csordás, R., Manning, C. D., & Potts, C. (2025). Do language models use their depth efficiently?  In arXiv [cs.LG]. https://doi.org/10.48550/ARXIV.2505.13898

      Dupré la Tour, T., Eickenberg, M., Nunez-Elizalde, A. O., & Gallant, J. L. (2022). Feature-space selection with banded ridge regression. In bioRxiv. https://doi.org/10.1101/2022.05.05.490831

      Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., & Olah, C. (2022). Toy Models of Superposition. Transformer Circuits Thread.

      Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A.,  Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A.,  Kernion, J., Lovitt, L., Ndousse, K., … Olah, C. (2021). A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread.

      Fan, S., Jiang, X., Li, X., Meng, X., Han, P., Shang, S., Sun, A., Wang, Y., & Wang, Z. (2024). Not all Layers of LLMs are Necessary during Inference. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2403.02181

      Friederici, A. D., & Becker, Y. (2025). The core language network separated from other networks during primate evolution. Nature Reviews. Neuroscience, 26(2), 131–132.

      Goldstein, A., Ham, E., Schain, M., Nastase, S. A., Aubrey, B., Zada, Z., Grinstein-Dabush, A.,  Gazula, H., Feder, A., Doyle, W., Devore, S., Dugan, P., Friedman, D., Brenner, M., Hassidim, A., Matias, Y., Devinsky, O., Siegelman, N., Flinker, A., … Hasson, U. (2025). Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models. Nature Communications, 16(1), 10529.

      Goldstein, A., Wang, H., Niekerken, L., Schain, M., Zada, Z., Aubrey, B., Sheffer, T., Nastase, S. A., Gazula, H., Singh, A., Rao, A., Choe, G., Kim, C., Doyle, W., Friedman, D., Devore, S., Dugan, P., Hassidim, A., Brenner, M., … Hasson, U. (2025). A unified acoustic-to-speech-to-language embedding space captures the neural basis of natural language processing in everyday conversations. Nature Human Behaviour. https://doi.org/10.1038/s41562-025-02105-9

      Goldstein, A., Zada, Z., Buchnik, E., Schain, M., Price, A., Aubrey, B., Nastase, S. A., Feder, A.,  Emanuel, D., Cohen, A., Jansen, A., Gazula, H., Choe, G., Rao, A., Kim, C., Casto, C., Fanda, L., Doyle, W., Friedman, D., … Hasson, U. (2022). Shared computational principles for language processing in humans and deep language models. Nature Neuroscience, 25(3), 369–380.

      Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P., & Roberts, D. A. (2024). The Unreasonable Ineffectiveness of the Deeper Layers. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2403.17887

      Herculano-Houzel, S. (2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences of the United States of America, 109 Suppl 1(supplement_1), 10661–10668.

      Hewitt, J., & Manning, C. D. (2019). A Structural Probe for Finding Syntax in Word Representations. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North (pp. 4129–4138). Association for Computational Linguistics.

      Honnibal, M., Montani, I., Van Landeghem, S., & Boyd, A. (2020). spaCy: Industrial-strength Natural Language Processing in Python.

      Kriegeskorte, N., Simmons, W. K., Bellgowan, P. S. F., & Baker, C. I. (2009). Circular analysis in systems neuroscience: the dangers of double dipping. Nature Neuroscience, 12(5),  535–540.

      Kumar, S., Sumers, T. R., Yamakoshi, T., Goldstein, A., Hasson, U., Norman, K. A., Griffiths, T. L., Hawkins, R. D., & Nastase, S. A. (2024). Shared functional specialization in transformer-based language models and the human brain. Nature Communications, 15(1), 5523.

      Linzen, T., & Baroni, M. (2021). Syntactic Structure from Deep Learning. Annual Review of Linguistics, 7(1), 195–212.

      Manning, C. D., Clark, K., Hewitt, J., Khandelwal, U., & Levy, O. (2020). Emergent linguistic structure in artificial neural networks trained by self-supervision. Proceedings of the National Academy of Sciences of the United States of America, 117(48), 30046–30054.

      Millet, J., Caucheteux, C., Orhan, P., Boubenec, Y., Gramfort, A., Dunbar, E., Pallier, C., & King, J.-R. (2023). Toward a realistic model of speech processing in the brain with self-supervised learning. NeurIPS 2022. https://doi.org/10.48550/ARXIV.2206.01685

      Pavlick, E. (2022). Semantic structure in deep learning. Annual Review of Linguistics, 8(1),  447–471.

      Pennington, J., Socher, R., & Manning, C. (2014). Glove: Global vectors for word representation.  Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar. https://doi.org/10.3115/v1/d14-1162

      Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences of the United States of America, 118(45), e2105646118.

      The CMU Pronouncing Dictionary. (n.d.). Retrieved May 27, 2025, from http://www.speech.cs.cmu.edu/cgi-bin/cmudict

      Vaidya, A. R., Jain, S., & Huth, A. G. (2022). Self-supervised models of audio effectively explain human cortical responses to speech. ICML 2022. https://doi.org/10.48550/ARXIV.2205.14252

      Zada, Z., Goldstein, A., Michelmann, S., Simony, E., Price, A., Hasenfratz, L., Barham, E., Zadbood,  A., Doyle, W., Friedman, D., Dugan, P., Melloni, L., Devore, S., Flinker, A., Devinsky, O., Nastase, S. A., & Hasson, U. (2024). A shared model-based linguistic space for transmitting our thoughts from brain to brain in natural conversations. Neuron, S0896627324004604. Zada, Z., Nastase, S. A., Aubrey, B., Jalon, I., Michelmann, S., Wang, H., Hasenfratz, L., Doyle, W.,  Friedman, D., Dugan, P., Melloni, L., Devore, S., Flinker, A., Devinsky, O., Goldstein, A., & Hasson, U. (2025). The “Podcast” ECoG dataset for modeling neural activity during natural language comprehension. Scientific Data, 12(1), 1135.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors report the results of a tDCS brain stimulation study (verum vs sham stimulation of left DLPFC; between-subjects) in 46 participants, using an intense stimulation protocol over 2 weeks, combined with an experience-sampling approach, plus follow-up measures after 6 months.

      Strengths:

      The authors are studying a relevant and interesting research question using an intriguing design, following participants quite intensely over time and even at a follow-up time point. The use of an experience-sampling approach is another strength of the work.

      Comments on revisions:

      Overall, I think the authors made many improvements to their manuscript. There are, however, still a number of concerns that first need to be addressed, since it is still not currently possible to fully evaluate the analyses, results, and conclusions presented in the paper. I list these points below:

      (1) The authors still use causal language where they must not use causal language. This is true for many places in the manuscript; I am highlighting here just a few places, but the authors nevertheless have to go carefully through the whole manuscript to change these instances.

      We sincerely thank the reviewer for this critical and well-taken point. We fully agree that our design (manipulating DLPFC excitability while measuring procrastination, task value, and aversiveness) does not directly measure or manipulate self-control, nor does it rule out alternative neurocognitive mechanisms. Accordingly, we have conducted a comprehensive, line-by-line revision of the entire manuscript to systematically replace causal claims with cautious, hypothesis-consistent language. In response, we have replaced all the wordings that may imply causal inferences, such as impair, cause, boost, by association-consistent phrasing. Furthermore, as you clearly raised below, we explicitly reframed self-control as a hypothesized theoretical construct rather than an empirically verified mediator throughout the whole re-revised manuscript. Please see specific revisions below:

      Abstract Section (Page 2, Line 53-59)

      “... a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement. In conclusion, these findings are consistent with the hypothesis that enhancing DLPFC function may reduce procrastination by selectively amplifying the valuation of future rewards, not by simply reducing negative feelings about the task.”

      Introduction Section (Page 3, Line 83-86)

      “... Even worse, chronic procrastination has been consistently associated with poor general health conditions, such as immune system disruption, gastrointestinal disturbance, hypertension and cardiovascular disease (Sirois, 2015; Sirois, 2016).”

      Introduction Section (Page 4, Line 143-145)

      “Consistent with this framework, the left dorsolateral prefrontal cortex (DLPFC)—a region frequently implicated in value-based decision-making and top-down regulation—has been associated with procrastination. ...”

      Introduction Section (Page 4, Line 135-139)

      “... Also, given the integrative nature of prefrontal regulatory functions, we hypothesize a third pathway whereby both decreased task aversiveness and increased task-outcome value may jointly contribute to reduced procrastination, potentially reflecting coordinated downstream effects on valuation and affective processing.”

      Introduction Section (Page 5, Line 183-186)

      “... Thus, this study aims to clarify the brain-behavior association between DLPFC neuromodulation and procrastination, and to test whether observed changes in task valuation and aversiveness are consistent with theoretical models of top-down regulation.”

      Results Section (Page 11, Line 534-536)

      “... Thus, these findings are consistent with the view that neuromodulation of the left DLPFC is associated with reduced task aversiveness and increased task-outcome value.”

      Results Section (Page 11, Line 579-584)

      “... In summary, these findings identified a statistical pathway consistent with the theoretical model: neuromodulation of the left DLPFC was associated with increased task-outcome value, which in turn was associated with reduced procrastination.”

      Discussion Section (Page 12, Line 617-620)

      “On balance, our findings provide evidence consistent with the hypothesis that neuromodulation of the left DLPFC is associated with reduced procrastination, primarily through increasing task-outcome value rather than merely reducing task aversiveness. ...”

      Discussion Section (Page 13, Line 664-669)

      “... Building on this foundation, among several theoretical interpretations and cognitive pathways, our study showed the one plausible neurocognitive mechanism of procrastination: the cortical excitability of the DLPFC produced by active neuromodulation may engage prefrontal regulatory networks to increase task outcome value, which in turn is associated with reduced procrastination behavior, statistically supporting the theoretical accounts of temporal decision model (TDM, Zhang et al., 2019).”

      Discussion Section (Page 14, Line 755-762)

      “... Moreover, this study did not collect data for assessing participants' self-control at either baseline or post-neuromodulation. Accordingly, we explicitly note that self-control was not directly measured or manipulated in this study; the observed associations between DLPFC neuromodulation, task-outcome value, and procrastination are consistent with theoretical models positing a role for top-down regulatory processes, but do not constitute direct evidence that self-control mechanisms were engaged. This limitation precludes definitive conclusions about the unique contribution of self-control-related pathways versus alternative neurocognitive mechanisms.”

      Some examples:

      (a) In response to my comment (1) in the previous round, where the authors adjusted their text, the authors still use causal language in their last sentence "... procrastination behavior has been observed to impair general health..." Unless the cited study truly allowed causal conclusions, the causal language should be removed here as well.

      Thank you for pointing out this inappropriate phrasing. As you kindly suggested, we have reworded it as “Even worse, chronic procrastination has been consistently associated with poor general health conditions, such as immune system disruption, gastrointestinal disturbance, hypertension and cardiovascular disease” (Introduction Section, Page 3, Line 83-86).

      (b) The authors still make (causal) claims about the involvement of self-control in their observed results. To reiterate from the previous round of revisions: The authors cannot make any strong claims about the role of self-control processes because they do not directly measure self-control nor do they directly manipulate self-control or have a design that would rule out alternative mechanisms other than self-control. Therefore, their claims about self-control have to be toned down. It is laudable that the authors have added a statement towards the end of their discussion about not being able to make strong conclusions about the role of self-control. But the authors need to use similar careful wording not just at the end of the discussion but throughout the manuscript.

      We appreciate you reiterating this concern. In the re-revised manuscript, we have thoroughly removed or rewritten all the statements implying causal inferences, and have substantially toned-down claims for the roles of self-control in procrastination reduction from this neuromodulation. Please see instances for what we have replied to the Comment #1.

      (i) In the abstract, the authors use the formulation "...conceptualized roles of self-control on procrastination..." -- this wording is still too strong, suggesting that you actually studied self-control.

      Thank you for providing this specific instance. This inappropriate sentence has been removed.

      (ii) In the introduction (page 4, lines162-169), the way the authors formulate these sentences suggests that they directly measured self-control. Again, the authors need to make it explicit that they are not directly measuring self-control but its hypothesized down-stream consequences on valuations/behavior.

      Many thanks. This statement has been removed, and we have reworded it as “... Thus, this study aims to clarify the brain-behavior association between DLPFC neuromodulation and procrastination, and to test whether observed changes in task valuation and aversiveness are consistent with theoretical models of top-down regulation.” (Introduction Section, Page 5, Line 183-186).

      (iii) In the discussion, for example, on page 11, lines 555 and following, the authors write: "One major contribution this study has made is to disentangle the neurocognitive mechanism of procrastination by demonstrating that self-control could increase task-outcome value so as to reduce procrastination."

      As you kindly instructed, we have rewritten this statement as “One contribution of this study is to provide empirical evidence partially consistent with the temporal decision model (TDM), showing that increased task-outcome value—rather than decreased task aversiveness—was statistically associated with reduced procrastination following DLPFC neuromodulation.” (Introduction Section, Page 12, Line 625-628), which no longer implies any conclusions for the role of self-control in this study.

      Again, please be aware that you are NOT demonstrating that self-control does anything, since you only measure procrastination rates, outcome values, and task aversiveness. It is possible that mechanisms other than self-control might be relevant for this. Perhaps neuromodulation directly increases outcome values, without involvement of self-control processes. You simply cannot know that and therefore you cannot make those claims in the form that you are making them. You can write that the observed results are consistent with the idea that neuromodulation might have had an effect on self-control and this in turn might have affected outcome values. But you also need to make it explicit that, to substantiate these claims, you would need more direct evidence that indeed self-control was involved. These more careful formulations would not at all reduce the value of your work, but indeed they would rather demonstrate your carefulness in interpreting the results you obtained.

      We sincerely thank the reviewer for this exceptionally clear and constructive guidance. We fully agree that our study design does not measure or manipulate self-control, and therefore we cannot demonstrate that self-control processes are causally involved in the observed effects. As you correctly note, it is entirely possible that neuromodulation directly modulates outcome valuation or engages alternative neurocognitive pathways (e.g., attentional allocation, feedback learning, or affective processing) without invoking self-control mechanisms.

      In direct response, as we replied above, we have completely rewritten the whole revised manuscript to remove any assertions that we “identified” a role of self-control. The revised text now explicitly states as follow: (1) our findings merely are consistent with the theoretical hypothesis that DLPFC neuromodulation might engage prefrontal self-regulatory functions, which in turn influence outcome valuation; (2) we explicitly acknowledge that substantiating this specific pathway would require more direct evidence. Rather than single sentence, we have applied this careful, hypothesis-consistent framing systematically across the Abstract, Introduction, Results, and Discussion. As you suggested, these revisions more accurately reflect the interpretative boundaries of our data and demonstrate our commitment to rigorous, transparent scientific reporting. Please see specific cases for this revision above.

      (2) I am still puzzled by the power analysis. In the text, you write that a sample size of 18 participants (i.e., 9 per group) would be sufficient to achieve 80% power. I still feel this seems far too optimistic and hard to believe, but that is not my point here. While in the text, you write that you need 18 participants, the G*power output seems to suggest a sample size of 34, not 18. Why this contradiction? Or is it not contradictory? If it is not, then please explain it more fully.

      We appreciate you pointing out this critical typo. In the last round of revision, we mean that 18 participants per group are required to achieve at least 80% statistical power, rather than a total sample size, as shown by the GPower software. We are sorry for this critical typo to confuse you. As you correctly pointed out, the GPower indicated that the minimum sample size to reach 80% power is 34 (i.e., 17 per group). Thus, we selected 36 (i.e., 18 per group) participants as minimum sample size in case of potential drop-out. We have thoroughly corrected this typo, and double-checked no such numeric issues:

      Methods Section (Page 5, Line 234-237)

      “... statistical power was predetermined by G*Power at a relatively medium effect size (1-β err prob = 0.80, f = 0.25), indicating the total sample size at 34 (17 per group) to reach acceptable power. To account for potential attrition, we determined to recruit 36 participants, at least.”.

      (3) I have several comments about the mixed-effects analysis.

      First of all, I want to thank the authors for adding more details, things have become much clearer now. However, I still have a few questions and comments related to these analyses:

      (a) The variable Emotions was within-subjects, as far as I understood. Accordingly, Emotions should most likely be modelled with random slopes varying over participants (in addition to being modelled as a fixed effect).

      We thank you raising this reasonable concern on the mixed-effect linear modeling. Yes, the Emotions reflect daily baseline affect, which is modeled as covariates of no interests to adjust for daily emotional fluctuation (if any). In this vein, this baseline emotion score is included for each participant across all the sessions. Therefore, it should be modeled with random slopes as you assumed indeed.

      In response, we have remodeled this mixed-effects analysis by including the daily baseline emotion as random slopes varying over participants. After centering the variables, we estimated the revised models as “Procrastination Rate ~ Group * Treatment day + Age + Gender + SES + Emotions + (1 + Treatment day + Emotions || SubjectID)” and “Task execution willingness ~ Group * Treatment day + Age + Gender + SES + Emotions + (1 + Treatment day + Emotions || SubjectID)”. Notably, as you correctly assumed, fitting this complicated random-effect structure is likely to result in convergence failure, given the limited sample size in the present study. Therefore, we hypothesized the independence among random effects for model simplification. Consistent with this assumption, model comparisons indicated that the simplified models fit better than original ones (Procrastination Rate model, ∆AIC = -1.0, ∆BIC = -11.8, LRT, χ<sup>2</sup>(3) = 5.04, p = .17; Task-execution willingness model, ∆AIC = -5.5, ∆BIC = -16.4, LRT, χ<sup>2</sup>(3) = 0.51, p = .91).

      Taken together, as you kindly suggested, we have rebuilt the mixed-effect models by adding daily baseline emotion as random slopes varying over participants, and have demonstrated the consistent findings with the original one:

      Methods Section (Page 8-9, Line 412-419)

      “... Given the risks of convergence failure with the two correlated random-effects structure (i.e., treatment days and self-reported emotions), we hypothesized that the random effects are independent, leading to model simplification. Consistent with this assumption, model comparisons favored the simplified independent structure over the full correlated structure for both outcomes. For the actual procrastination model, the simplified model showed lower AIC (∆ = -1.0) and BIC (∆ = -11.8), with a non-significant likelihood ratio test (χ<sup>2</sup> (3) = 5.04, p = .17). For the task-execution willingness model, the simplified model was also preferred (∆AIC = -5.5, ∆BIC = -16.4; LRT: χ<sup>2</sup> (3) = 0.51, p = .91).”

      Results Section (Page 9-10, Line 469-489)

      “For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001; Fig. 3A). In the post-hoc simple effect analysis, it demonstrated a significantly increased task-execution willingness (i.e., decreased procrastination willingness) after neuromodulation in the active neuromodulation group (NM-before: 35.65 ± 30.21, NM-after: 80.43 ± 19.92, Mean Diff = 41.79, SE = 7.58, DF = 103.4, t.ratio = 5.51, p < .0001, Tukey correction), but no such effects were identified in the sham control group (SC-before: 37.57 ± 26.46, SC-after: 47.35 ± 30.49, Mean Diff = 2.58, SE = 7.56, DF = 96.8, t.ratio = 0.34, p = .73, Tukey correction) (Fig. 3B-C). A linear uptrend for task-execution willingness was further observed across multiple sessions in the active NM group, indicating gradually increasing neuromodulation effects (Fig. 3D; p < .01, Mann-Kendall test). For actual procrastination behavior, changes to actual procrastination rates across all the sessions have been detailed in the Fig. 3E. Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group (NM-before: 56.74 ± 39.10, NM-after: 0.00 ± 0.00, Mean Diff = 44.40, SE = 9.36, DF = 110.0, t.ratio = 4.74, p < .0001, Tukey correction), but no such prominent changes found in the sham control group (SC-before: 46.47 ± 40.76, SC-after: 33.35 ± 37.82, Mean Diff = 7.53, SE = 9.28, DF = 102.0, t.ratio = 0.81, p = .42, Tukey correction) (Fig. 3F-G).”

      (b) The analyses still cannot fully be evaluated as I cannot access the scripts and data. The authors mention that the scripts and data should be available via a link they provide (https://doi.org/10.57760/sciencedb.35140). However, when I try to access these materials via this link, no page opens; it seems the link is dead?

      Thank you very much for bringing this case to us. We checked this link and found it to be still active.

      To ensure accessibility for your evaluation, we have uploaded scripts and data into this online submission system. Please do let us know if you are still unable to access them. We are glad to send them to you by other available pathways. This link is a private access to you, and the repository would be openly available for other users upon the final publication.

      (c) What are the results and conclusions if you do not include the covariates of no interest? I.e., please re-run your main models without age, gender, SES, Emotions.

      Thank you for raising this question. As you clearly instructed, we have rerun main models without all those covariates. As shown in the table below, the results for the key predictors of interest (Group, Treatment day, and their interaction) remained largely unchanged in terms of effect size, direction, and statistical significance:

      Author response table 1.

      Comparison to statistics derived from model with covariates (i.e., age, gender, SES, Emotions) and without covariates

      (d) The authors mention that they use GLMMs, which would suggest generalized mixed-effects models, but they do not describe what family/distribution they used. Since they mention lmerTest and seem to report F-tests, my guess is that they used Gaussian models. However, both their DVs (procrastination rates and their ratings) are bounded variables and at least procrastination rates hit the lower boundary. That can mean that their analyses suffer from inflated Type 1 and/or Type 2 rates. Therefore, please repeat the analyses with an appropriate generalized mixed-effects model (perhaps a beta regression type of model?).

      We are very grateful to you for raising this crucial statistical point. As you correctly pointed out, we used the Gaussian distribution in estimating this model. We are sorry to confuse you due to the absence of reporting family/distribution we used. In the original manuscript, we meant “general” linear mixed-effect model, rather than “generalized” one. As you clearly and correctly raised, procrastination rates and willingness are technically bounded, and that procrastination rates frequently reached the lower boundary (0%) in the present study, which are in high risks to be inflated for Type 1 and/or Type 2 error.

      Thus, as you kindly suggested, a beta family distribution with logit function is used to reanalyze those main effects of interest. Results are tabulated in Author response table 2.

      Author response table 2.

      These convergent results confirm that the critical main effect (i.e., Group and Treatment day) and their interaction remain statistically significant across distributional specifications, and that our primary conclusions are not artifacts of the Gaussian assumption. Taken them together, as you kindly suggested, we have repeated the analyses with beta regression family distribution, which replicated our main findings, potentially supporting their statistical robustness.

      Following your suggestion, we have added those results derived from such sensitivity analyses into the revised manuscript:

      Methods Section (Page 9, Line 430--436)

      “... To examine whether our findings were sensitive to the distributional assumptions of the dependent variables, we re-analyzed the main models using an alternative distributional specification. Given that both procrastination rates (ranging from 0% to 100%) and task-execution willingness (measured on a 0-100 visual analog scale) are bounded continuous outcomes, and that procrastination rates frequently reached the lower boundary (0%) in the present study, a Beta regression model with a logit link function was employed for a sensitivity analysis.”

      Results Section (Page 10, Line 506-512)

      “... Furthermore, as a sensitivity analysis, we reran the main LMMs using Beta regression distribution with a logit link function, which is appropriate for the both bounded outcomes mentioned above (i.e., procrastination rate and procrastination willingness). The main effects (i.e., Group and Treatment day) and their interaction remained significant for both procrastination rate and willingness (see SI Results and Tab. S5), confirming that our findings are robust to alternative distributional assumptions.”

      (e) When reporting the results of the mixed-effects models, the authors report the regression coefficient, standard error, DFs and p value, but not the actual test statistic. Please add the information about the test statistic and report all degrees of freedom (in case of F tests that would be the degrees of freedom of the test and the residual degrees of freedom).

      We truly thank you for this nuanced reminder. As you suggested, we have added actual test statistics, including t-values and all degrees of freedom (DF). Please see specific instances below:

      Results Section (Page 9-10, Line 469-489)

      “For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001; Fig. 3A). In the post-hoc simple effect analysis, it demonstrated a significantly increased task-execution willingness (i.e., decreased procrastination willingness) after neuromodulation in the active neuromodulation group (NM-before: 35.65 ± 30.21, NM-after: 80.43 ± 19.92, Mean Diff = 41.79, SE = 7.58, DF = 103.4, t.ratio = 5.51, p < .0001, Tukey correction), but no such effects were identified in the sham control group (SC-before: 37.57 ± 26.46, SC-after: 47.35 ± 30.49, Mean Diff = 2.58, SE = 7.56, DF = 96.8, t.ratio = 0.34, p = .73, Tukey correction) (Fig. 3B-C). A linear uptrend for task-execution willingness was further observed across multiple sessions in the active NM group, indicating gradually increasing neuromodulation effects (Fig. 3D; p < .01, Mann-Kendall test). For actual procrastination behavior, changes to actual procrastination rates across all the sessions have been detailed in the Fig. 3E. Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group (NM-before: 56.74 ± 39.10, NM-after: 0.00 ± 0.00, Mean Diff = 44.40, SE = 9.36, DF = 110.0, t.ratio = 4.74, p < .0001, Tukey correction), but no such prominent changes found in the sham control group (SC-before: 46.47 ± 40.76, SC-after: 33.35 ± 37.82, Mean Diff = 7.53, SE = 9.28, DF = 102.0, t.ratio = 0.81, p = .42, Tukey correction) (Fig. 3F-G).”

      (f) Thank you for adding the analysis where you remove the last two sessions. But currently you present them in the manuscript without explaining/motivating why you do this. Please add this motivation, as otherwise it will be puzzling for the reader why you conduct these analyses.

      Thank you for this very practical requirement to clarify the motivation of reanalyzing main models from removing the last two sessions. Please see specific explanation as follow:

      Results Section (Page, Line 499-506)

      “... To systematically test whether such effects are biased by extreme data points or patterns, we reran the main LMMs by iteratively removing data from the last two sessions, which showed extraordinarily high effectiveness from neuromodulation (e.g., all the participants in the neuromodulation group had no actual procrastination behavior in session #6 and #7). Results showed the significant group*neuromodulation sessions interaction effects across all those nested models (removing session #6, #7 or both, all p < .05; see SI Results and Tab. S3-4), potentially indicating a statistical robustness from the data pattern.”

      (4) Mediation analysis

      In your manuscript, you present some mediation analyses. Please be aware that such mediation analyses cannot establish causality and they suffer from extremely high Type 1 error rates (see, e.g., https://datacolada.org/103). My suggestion would be to completely remove all mediation analyses. However, if you want to keep them, then you need to be extremely careful in how you present the results. You need to explicitly mention that you cannot derive any causal conclusions from them and that simulation studies have shown that such mediation analyses suffer from extremely high Type 1 errors.

      We sincerely thank you for this exceptionally important methodological guidance. We fully agree that mediation analyses, especially those based on observational measures rather than experimentally manipulated mediators, cannot establish causal pathways and are susceptible to inflated Type 1 error rates, as rigorously demonstrated in recent simulation studies (https://datacolada.org/103).

      As you kindly suggested, please allow us to retain those mediation analyses upon explicitly highlighting that this mediation statistical model cannot generate any causal conclusions. In response, we have systematically replaced all instances of “causal mediation” with “statistical mediation” or “exploratory mediation analysis”, and removed causal verbs (e.g., “depends on”, "drives", "explains") in favor of association-consistent phrasing (e.g., “is statistically mediated”, “aligns with the hypothesis that”) throughout the abstract, introduction, methods, results and discussion sections. Furthermore, in the Discussion section, we explicitly reiterated the limitations of extending this mediation associations to causal conclusions. Please see specific modifications underneath:

      Abstract Section (Page 2, Line 52-56)

      “... While the intervention is significantly associated with both decreased task aversiveness and increased perceived task outcome value, a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement.”

      Methods Section (Page 9, Line 448-453)

      “... As these mediation analyses are based on observational measures rather than experimentally manipulated mediators, they do not establish causal pathways. Simulation studies have shown that such analyses can suffer from inflated Type 1 error rates. Results should therefore be interpreted as hypothesis-generating and statistically consistent with the proposed theoretical model, rather than as confirmatory evidence of causal mechanisms.”

      Methods Section (Page 9, Line 438-440)

      “... the Quasi-Bayesian mediation analysis was used to model the association between the effects of tDCS, task aversiveness/outcome and decreased procrastination.”

      Results Section (Page 11, Line 568-573)

      “As an exploratory analysis, results indicated that increased task outcome value was associated with changes in the task-execution willingness (δ = 21.73, p < .01; ζ = 11.25, p = .07, ρ = 32.99, p < .01, simulation = 1,000; see Fig. 5A) and real-world procrastination (δ = 30.75, p < .01; ζ = 3.05, p = .52, ρ = 33.81, p < .01, simulation = 1,000; see Fig. 5B), in the context of ms-tDCS neuromodulation. ...”

      Results Section (Page 11, Line 582-584)

      “... Nevertheless, all mediation findings are now explicitly labeled as “exploratory” and framed as quantifying statistical associations consistent with the TDM pathway, not causal mediation.”

      Discussion Section (Page 14, Line 747-755)

      “Notably, we explicitly acknowledge that exploratory Quasi-Bayesian mediation analyses, based on observational measures rather than experimentally manipulated mediators, cannot establish causal pathways and are susceptible to inflated Type 1 error rates as demonstrated in recent simulation studies. These findings should be interpreted strictly as hypothesis-generating and statistically consistent with the proposed theoretical model, rather than as confirmatory evidence of causal mechanisms. Substantiating the precise neurocognitive pathway will require future studies employing stronger causal designs, such as experimental manipulation of task valuation or longitudinal cross-lagged modeling. ...”

      As an example (but the mediation results are mentioned in several places, for example, also in the abstract): On page 10, lines 501-503: What you can causally conclude is that neuromodulation affects your measured variables (outcome values, procrastination rates, task aversiveness), but you cannot conclude that the effect of neuromodulation on procrastination rates causally operates via outcome values. Thus, please adjust the formulation accordingly. The same applies to the mediation section that follows right afterwards (page 10, lines 505-522).

      Thank you for offering those specific instances. As we replied above, they have been revised accordingly:

      Results Section (Page 11, Line 557-559)

      “... Collectively, these findings provide statistical evidence consistent with the hypothesis that the outcome-value pathway may contribute to procrastination reduction.”

      Results Section (Page 11, Line 563-582)

      “Increased task outcome value is specifically associated with reduced procrastination in the context of neuromodulation

      To explore the potential neurocognitive pathways of procrastination, the Quasi-Bayesian mediation analysis was undertaken, with increased task outcome value as a statistically mediated variable. As an exploratory analysis, results indicated that increased task outcome value was associated with changes in the task-execution willingness (δ = 21.73, p < .01; ζ = 11.25, p = .07, ρ = 32.99, p < .01, simulation = 1,000; see Fig. 5A) and real-world procrastination (δ = 30.75, p < .01; ζ = 3.05, p = .52, ρ = 33.81, p < .01, simulation = 1,000; see Fig. 5B), in the context of ms-tDCS neuromodulation. To ensure the statistical robustness and specificity of these findings, the sensitivity analysis was implemented by changing sampling parameters and outcome variables. By doing so, those findings were validated statistically robust, as shown by replicated observations across bootstrapping sampling subsets (see SI Results and Tab. S6-7). Moreover, the results of the control analysis further validated the specificity of these findings by showing a null statistically mediated effect of this model to predict one’s task aversiveness (see SI Results and Tab. S8). In summary, these findings identified a statistical pathway consistent with the theoretical model: neuromodulation of the left DLPFC was associated with increased task-outcome value, which in turn was associated with reduced procrastination.”

      (5) In the introduction, the authors introduce several theoretical procrastination frameworks (TMT, mood repair, TDM). Do the results of the current paper help to decide which framework might be the most appropriate, at least for the authors data set? It might be of interest to address this explicitly.

      We do thank the reviewer for this insightful theoretical question. We agree that explicitly positioning our findings within the broader theoretical landscape strengthens the conceptual contribution of our work. Upon careful consideration, we believe that our results provide the strongest empirical support for the TDM over alternative frameworks (TMT, mood repair). TDM uniquely posits procrastination as contingent on the dynamic trade-off between task aversiveness and task-outcome value. In the present study, neuromodulation was identified to be associated with both pathways but only increased outcome value statistically predicted reduced procrastination. This aligns precisely with TDM’s hypothesis that value-based processes may dominate aversiveness-avoidance processes in driving behavioral change. Neither TMT (which emphasizes temporal discounting of utility per se) nor the mood repair perspective (which prioritizes short-term affect regulation) explicitly predicts this dissociable pattern. As you kindly suggested, we have extended the Discussion section to explicitly contextualize our findings into this theoretical landscape (Discussion Section, Page 13, Line 669-687).

      (6) The language is sometimes hard to understand and seems in quite some places grammatically incorrect. Thus, I think the paper would profit very much from thorough English proofreading.

      We sincerely thank you for this practical suggestion. We fully agree that the original manuscript contained grammatical inaccuracies and awkward phrasing that could hinder readability. In response, we have engaged a professional academic editing service to thoroughly proofread and polish the entire manuscript. All sentences have been revised for grammatical correctness, syntactic clarity, and academic tone, while strictly preserving the original scientific meaning and technical terminology. We believe this language improvements have substantially enhanced the readability and overall quality of the paper.

      Reviewer #2 (Public review):

      Summary:

      Chen and colleagues conducted a cross-sectional longitudinal study, administering high-definition transcranial direct stimulation (HD-tDCS) targeting the left DLPFC to examine the effect of HD-tDCS on real-world procrastination behavior. They find that seven sessions of active neuromodulation to the left DLPFC elicited greater modulation of procrastination measures (e.g., task-execution willingness, procrastination rates, task aversiveness, outcome value) relative to sham. They show that HD-tDCS reduces task aversiveness and increases task-execution willingness on real-world tasks as quantified by intensive experience sampling methods, providing causal evidence for the role of DLPFC in modulating contextual features to delaying or completing one's goals.

      Strengths:

      • This is a well-designed protocol with rigorous administration of high-definition transcranial direct current stimulation across multiple sessions. The intensive experience sampling approach which probes and assesses self-relevant task goals is innovative and aims to address an important question regarding the specific role of DLPFC in modulating specific features of chronic procrastination behavior (e.g., task-execution willingness, task aversiveness).

      • The quantification of task aversiveness through AUC metrics is a clever approach to account for the temporal dynamics of task aversiveness, which is notoriously difficult to quantify.

      Weaknesses:

      • While the findings that neurostimulation reduces procrastination behavior is compelling, there remain several alternative interpretations for these effects. For example, it could be that the task-execution willingness isn't increased per se, but rather that the goal completion becomes more valuable as participants learn from feedback or become more aware of their successful attainment of or failure to complete task goals. It is unclear whether the effects could be driven by improved working memory or attention to the reported tasks (and this limitation is addressed by the authors). In short, it is also difficult to examine the temporal dynamics of how these goals are selected across time.

      We sincerely thank you for raising these thoughtful and methodologically important points. We fully agree that the observed reductions in procrastination could reflect multiple neurocognitive pathways beyond the value-based mechanism emphasized in our primary analysis.

      In response, we have thoroughly removed claims on the “unique mechanism” of value-based pathways to procrastination reduction, and fully substituted language implying exclusive mediation by “value amplification” with more cautious phrasing (e.g., “statistically consistent with a value-based pathway”; “one plausible mechanism among several processes”). In the revised manuscript, we reiterated that the pattern of results, showing increased outcome value predicting reduced procrastination while decreased aversiveness did not, aligned with the Temporal Decision Model, yet does not rule out concurrent contributions from attention, learning, or executive processes. To clearly bring this interpretative boundary of our primary findings for audiences, we have explicitly warranted such cautions in the Discussion Section. Please see specific modifications underneath:

      Discussion Section (Page 13, Line 664-669)

      “... Building on this foundation, among several theoretical interpretations and cognitive pathways, our study showed the one plausible neurocognitive mechanism of procrastination: the cortical excitability of the DLPFC produced by active neuromodulation may engage prefrontal regulatory networks to increase task outcome value, which in turn is associated with reduced procrastination behavior, statistically supporting the theoretical accounts of temporal decision model (TDM, Zhang et al., 2019). ”

      Discussion Section (Page 13, Line 676-687)

      “... Despite statistically supporting the TDM, we acknowledge that alternative neurocognitive mechanisms could contribute to the observed reductions in procrastination. For instance, repeated exposure to the experience-sampling protocol may have enhanced participants’ awareness of task progress or facilitated feedback-based learning, thereby increasing the subjective value of goal completion independent of DLPFC neuromodulation. Similarly, improvements in working memory for task maintenance, attentional allocation to reported goals, or strategic shifts in goal selection across sessions could plausibly mediate the intervention effects. While our double-blind, sham-controlled design and inclusion of daily emotional covariates help mitigate some non-specific confounds, the present study did not incorporate direct measures of these alternative processes. Consequently, we cannot definitively isolate the value-based pathway posited by the TDM from concurrent contributions of attention, learning, or executive functions.”

      • It is unclear whether the current evidence support long-retention of this neurostimulation intervention. The study includes one 6-month timepoint after the study to examine the long-term retention of the neural stimulation effect. Future studies that evaluate the long-term effects across multiple time points would strengthen the evidence for the robustness of this intervention.

      We genuinely appreciate you for this insightful and methodologically reasonable point. We fully agree that a single 6-month follow-up assessment, while valuable, provides only preliminary evidence for long-term retention, and that multiple follow-up timepoints would substantially strengthen claims about the durability of neuromodulation effects. To carefully address this point, we have rephrased the “long-term retention” as “long-term after-effects” throughout the whole revised manuscript, and overall toned down the claims on the retention effects. Moreover, this limitation has been explicitly elaborated in the Discussion section:

      Abstract Section (Page 2, Line 49-50)

      “... we assessed the effect of anodal HD-tDCS on real-world procrastination behavior at offline after-effect (2-day interval) and long-term after-effect (6-month follow-up).”

      Results Section (Page 12, Line 606-608)

      “... Therefore, beyond short-term effects, the benefits of ms-tDCS neuromodulation on reducing procrastination were still detectable at a 6-month follow-up, providing preliminary evidence consistent with long-term after-effects.”

      Discussion Section (Page 14, Line 721-725)

      “... Thus, the detectable effects at 6 months are consistent with the hypothesis that repeated neuromodulation may induce neuroplastic changes in the DLPFC that support sustained behavioral change. However, we explicitly note that a single follow-up timepoint cannot establish the stability or trajectory of these effects; future studies with multiple longitudinal assessments are required to substantiate claims about long-term retention.”

      Discussion Section (Page 15, Line 771-775)

      “... Finally, while our 6-month follow-up provides preliminary evidence for sustained effects, the use of a single follow-up timepoint limits our ability to characterize the temporal trajectory of retention. Future studies incorporating multiple follow-up assessments (e.g., 1-month, 3-month, 6-month, 12-month) would strengthen evidence for the robustness and durability of this intervention.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please see my detailed comments above (6 points; several of them with subpoints a, b, c, etc).

      Thank you for listing those specific and helpful recommendations above. Please see our detailed response posed above, point-by-point.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study uses an elegant visual-anagram approach to test whether perceived animacy structures visual working memory and attention while controlling for many low-level image properties. The evidence is solid, with converging results across seven preregistered experiments, but the central claim that animacy itself is represented independently of visual features should be tempered, as residual mid-level configural cues, ensemble or category structure, and broader semantic differences may also contribute to the effects. The work will be of interest to researchers studying high-level visual representation, attention, and working memory.

      We thank the Editors and Reviewers for this careful and informed assessment. We appreciate that every Reviewer found our approach to be elegant, our findings to be solid, and our question to be of broad interest. We respond to each Reviewer’s specific comments in more detail below; but we thought to summarize some of the highlights - especially the specific comments that come up in this Assessment - here.

      (1) The Reviewers make the insightful point that, even if our stimuli effectively control for many low-level features, there may be other high-level features that explain performance in our experiments (R2: “Although the anagram paradigm effectively controls low-level visual features […] these stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories.”). We are happy to embrace this possibility. If our results were explained by high-level visual representation of the natural vs. manmade distinction, rather than the animate vs. inanimate distinction, this would still be an appeal to a (not altogether unrelated) high-level property being represented independently from its lower-level features, which was the primary motivation for our study. We framed our work specifically around animacy given the persistent debates regarding perceived animacy, as well as the fact that our stimuli do quite saliently vary along that dimension; but we are certainly open to other nearby high-level categories being at play. We also think this is an empirical question that could be tested in future work. For example, objects like rocks and lakes are natural but inanimate. If they behave more like dogs than like boots in our paradigms, then Reviewer #2 may be right that naturalness was the relevant property all along; but if they behave more like boots than like dogs, then perhaps it really was animacy doing the work. We now discuss this explicitly in our paper, and we appreciate the opportunity to not only clarify our claims but also spur discussion for future work.

      (2) Multiple Reviewers raise the question of whether semantic factors that go beyond the images themselves may be driving our effects. Reviewer #3 raises a particularly interesting question along these lines: “if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results?” We have now taken this question quite literally and run this experiment exactly as described. Of course, much research already explores cognitive processing of animate/inanimate words, finding (for example) stronger memory for animate objects than inanimate ones (e.g., Nairne et al., 2013; Nairne et al., 2017). However, such tasks do not invoke effects of visual processing, whereas the question at issue here is specifically whether the visual system prioritizes animacy independent of its lower-level features. To this end, we conducted a new, pre-registered experiment (now Experiment 8) where participants search for animate/inanimate words on some trials, and animate/inanimate pictures on others. Given the nature of visual search tasks, we should expect to find no search advantage for words (as their meanings are not processed in vision per se) — and we should also expect to replicate (once again) our search advantage for pictures. This is exactly what we found. In other words, linguistic stimuli alone failed to produce the effect, while anagrams did produce the effect. We believe this rules out the strongest form of the semantic labeling account.

      (3) Finally, Reviewers #2 and #4 raise some concerns regarding residual mid-level features such as configural shape and ensemble statistics, which lie somewhere between animacy itself and more basic properties like contrast or spatial frequency. In our paper, we now clarify each of these concerns in greater detail. In short: We think that our stimuli and experiments indeed control for these residual cues. For example, rotating an image preserves its configural shape; and, as we argue below, the specific ensemble statistics argument fails to get off the ground without appeal to animacy itself.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Evidence for visual representation of animacy.

      Strengths:

      This is a very cool paper that casts light on a persistent problem in the psychology and philosophy of visual representation: is there high-level perception? Every vision scientist agrees that low-level features such as shape, color, texture, motion and spatial frequency are represented in visual perception, but there is a great deal of controversy about the representation of high-level properties such as causation, faces, agency and animacy. Animacy is especially problematic because there are large differences in line curvature between stimuli that represent animate and inanimate items.

      This article uses a novel approach-visual "anagrams" that are exactly the same image, except one is rotated 90 degrees relative to the other. They found persistent differences in visual processing between animate and inanimate stimuli. (Of course, the stimuli aren't animate-they represent animate items). For example, there were processing differences between changes between animate and inanimate items (rabbit to boot) that were not present in rabbit to dog. They also showed such differences in two kinds of visual search tasks.

      Of course, there are feature differences that exploit orientation. A classic example is the difference between a square and a diamond that is produced from the square by rotating it 45 degrees.

      They addressed an aspect of this challenge having to do with some features using silhouettes. There was no search advantage for silhouetted stimuli.

      Weaknesses:

      I thought this was an excellent submission. I have two suggestions for revision:

      We are glad to hear this Reviewer recognizes the broad challenge we are tackling in this work (separating high-level from low-level visual features) and found our submission to be “excellent”.

      (1) I thought that experiment 7 should have been described in more detail, with the upshot explained better. What exactly do the authors take it to show?

      Sorry for the lack of clarity here. We think the Reviewer actually gets this right earlier in their review; many feature differences exploit orientation, and our silhouettes control (Experiment 7) shows that those differences alone fail to explain our effects. For example, one might worry that our search effects merely reflect oddities in the aspect ratio or center of mass of the images. Converting the anagrams into silhouettes preserves these features. Thus, the fact that we found no search advantage with silhouettes suggests that these features on their own fail to produce the relevant effects; put the other way around, the effects we observed earlier must go beyond those features. We now discuss this in greater depth in our paper.

      (2) There should be a candid discussion of what the loose ends are and how they might be addressed. It would be good to have some examples like the square/diamond case with some indication of what would address such challenges.

      We agree with this, though we are somewhat limited by the space constraints of the Short Report format. A primary loose end we see is the possibility that high-level properties other than animacy explain our results (as raised by other Reviewers). We have added some discussion of this possibility to the paper.

      We would like to thank this Reviewer for their thoughtful feedback.

      Reviewer #2 (Public review):

      Summary:

      The authors present a creative approach using visual anagrams matched on low-level image statistics to isolate animacy from low-level visual features and report consistent effects of animacy on visual working memory and attention. While this is a thoughtful design and is well executed across seven pre-registered experiments, it remains unclear whether the reported effect is truly driven by animacy, as opposed to broader differences in ensemble statistics or semantic structure across the "mixed animacy" versus "uniform animacy" conditions. As such, the interpretation of a "pure" animacy effect may be overstated.

      Strengths:

      (1) An important methodological advance in controlling low-level confounds that have historically complicated the study of animacy.

      (2) The converging effects across multiple experiments, together with the pre-registered design, strengthen the reliability of the reported findings.

      We are glad to hear this Reviewer found our work to be “creative” and believes it offers an “important methodological advance”.

      Weaknesses:

      (1) Specificity of the animacy effect vs. category-level ensemble structure

      The central claim is that animacy itself drives the observed effects. However, the key manipulation ("mixed animacy" versus "uniform animacy") also introduces differences in category-level ensemble structure. For example, in Experiments 1-2, cross-category change detection (e.g., dog to chair) may be easier not because of animacy per se, but because of a change in overall ensemble statistics (Brady & Alvarez, 2011, 2015). In addition, since each display contains five objects (two in one category and three in the other category), cross-category changes may also alter category balance in a way that further facilitates detection. In contrast, within-category changes preserve both ensemble structure and category composition, making them more difficult to detect.

      Brady, T. F., & Alvarez, G. A. (2011). Hierarchical encoding in visual working memory: Ensemble statistics bias memory for individual items. Psychological Science.

      Brady, T. F., & Alvarez, G. A. (2015). Contextual effects in visual working memory reveal hierarchically structured memory representations. Journal of Vision.

      We appreciate the opportunity to clarify our claims and the support for them. Our claim is indeed that animacy (or a closely related high-level property; see below) drives our effects, over and above its lower-level correlates — i.e., that the explanation for differences in change detection or search across conditions will invoke a high-level property of the images. As we understand the Reviewer’s concern(s), they either (a) are already addressed by our novel methodology, or (b) would still fall perfectly in line with our claim as stated above.

      Consider the Reviewer’s concern that cross-category change detection “may be easier not because of animacy per se, but because of a change in overall ensemble statistics”. Which ensemble statistics change across categories in our stimulus set? Take as an example the case depicted in our figure, where a rabbit changes into either a dog (within-category) or a boot (cross-category). The dog and the boot are the very same image, just rotated; thus, they have the same luminance, curvature, area, spatial frequency, and so on. So if the change from rabbit to dog changes the array’s ensemble statistics with respect to any of those properties, it does so in the very same way as the change from rabbit to boot — and yet detection is still better for rabbit → boot than for rabbit → dog. Indeed, for nearly any ensemble statistic, the difference between the rabbit-display and the dog-display will be identical to the difference between the rabbit-display and the boot-display. To engage with the specific cases discussed in the two cited papers (Brady & Alvarez, 2011, 2015): The dog and the boot are the same size (because they are the same image), so average size is identical (just as average luminance, curvature, area, and spatial frequency are identical). And the very few properties left over (e.g., aspect-ratio) are addressed by later experiments.

      To be clear: We are not saying that there are no differences in ensemble statistics between the rabbit-display and the dog-display; across those displays, we replace one image with a different image, so there are likely all kinds of corresponding differences in ensemble statistics. The key question is whether that change in ensemble statistics differs across trial types in ways that might explain our effect - i.e., whether there is any difference between the rabbit-display and dog-display that is not also present between the rabbit-display and the boot-display. We don’t see how the answer could be yes, at least with respect to the statistics typically considered. A similar logic applies to the search tasks, with the silhouette control (Experiment 7) providing especially strong evidence that certain ensemble statistics or lower-level features cannot explain our effect.

      Now, it’s possible the Reviewer is referring to properties other than the low-/mid-level properties we mention above. Perhaps, for example, many animate stimuli on a display at one time have a striking collective appearance (all these animals are looking at me!) that lots of inanimate stimuli do not (this might be related to the Reviewer’s concern about “category balance”). But as we see it, this explanation just invokes animacy all over again, and so is the sort of explanation we would embrace.

      We now say more about this concern in the paper to be as clear as possible about our claims.

      (2) Limited stimulus set and potential learning effects

      The relatively small stimulus set (six anagram pairs) and repeated exposure raise the possibility of learning or familiarity effects. Does performance change over time? e.g., are there meaningful differences between early and late trials (e.g., first 10% vs. last 10%)? If such differences are present, they could suggest the development of task-specific strategies or increased efficiency with repeated exposure, rather than stable effects driven by the experimental manipulation itself.

      This is an interesting question, and we recognize this analysis absent from our initial submission. To be fair, stimulus sets of this size are not unusual in change-detection and search tasks, which often involve red, green, and blue squares repeated over the course of several hundred trials. Still, we certainly take the Reviewer’s point here and also embrace their analytical approach to addressing it. We’ve now run the “familiarity effects” analyses the Reviewer suggests (as well as some they did not suggest). The top-level headline is that learning or familiarity effects cannot explain our results, and if anything most of these analyses not only fail to support this alternative account but actively point against it. Below are more details.

      First, we worry that the Reviewer’s concern about “the development of task-specific strategies … rather than stable effects driven by the experimental manipulation itself” isn’t actually addressed by the suggested analysis of comparing the last 10% of trials to the first 10%. One reason for this is simply that it’s possible that both mechanisms are at play - i.e., that there is a baseline difference even without any familiarity that is then enhanced by some learning mechanism. (There are other issues as well: For example, one might imagine that participants get quite good at the task during the middle 80% of trials, but then get fatigued at the end. If this were true, then comparing the first 10% to the last 10% of trials could make it seem like there is no learning or familiarity, even if there were such effects. And on top of all this there is just the issue of statistical power, since far fewer trials go into these analyses than into our primary, pre-registered analyses). Nevertheless, we ran the Reviewer’s proposed analyses (using the first and last 10 trials of each type, which offers the best chance to find the pattern the Reviewer is concerned about). If anything, this analysis points in the opposite direction to the Reviewer’s prediction: 4/6 experiments (Experiments 1, 3, 4, and 5) revealed numerically weaker effects at the end of the task than the start, while only 2/6 experiments (Experiments 2 and 6) revealed numerically stronger effects at the end of the task than the start. Moreover, most of these results were non-significant, with only one marginal result (Experiment 5, p< = 0.08) and one significant result (Experiment 2, p = 0.01), and this is before any correction for multiple comparisons, which would make all of these results non-significant. So even though our account could easily accommodate learning effects, it’s not clear that they even exist here in any consistent or reliable way.

      Second, however, we think a more informative way to answer the Reviewer’s question is to ask not about learning over the course of the experiment but rather whether the key effects arise very early in the task. If they do, then any learning effects arising later couldn’t fully account for our results. Now, again, these tests are underpowered and only exploratory (to do this analysis properly, we would want to run entirely new experiments designed for this purpose), but we in fact did find evidence that our key effects arise early. In 5/6 experiments (Experiments 1, 3, 4, 5, and 6), the key effect was significantly (or in one case marginally) present even at the beginning of the experiment (Experiment 1, p = 0.07; Experiment 3, p = 0.01; Experiment 4, p < 0.001; Experiment 5, p < 0.001; Experiment 6, p < 0.01), and most of these results would survive correction for multiple comparisons. (In only one experiment, Experiment 2, was there a numerical disadvantage, but it was not significant; p = 0.34.) So even though our experiments were not designed or powered for this purpose, they do seem to suggest that the effects arise even without much familiarity at all.

      All told, we think these analyses suggest quite strongly that learning alone fails to explain our key effects. There is no evidence that the effects in general are stronger at the end of the experiment than the beginning (if anything it is the opposite); and there is evidence that most of the effects we investigated can be detected even very early in the experimental sessions. We have added discussion of these new analyses to our manuscript.

      (3) Role of semantics

      Although the anagram paradigm effectively controls low-level visual features, it still relies on high-level semantics (e.g., "dog" vs. "boot"). These stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories. From a semantic standpoint, it remains unclear whether the observed effects can be uniquely attributed to animacy or whether they reflect broader conceptual distinctions.

      We agree with the Reviewer here. While we feel comfortable interpreting our effects in terms of a high-level property like animacy as opposed to a lower-level property like curvature, it remains possible that the observed effects reflect some other, closely related high-level distinction (like natural vs. manmade). Our primary concern was to tease apart high-level properties from low-level features, which the Reviewer’s question does not threaten — if attention and memory are sensitive to the natural/artificial distinction, that’s interesting too, and a near neighbor of our actual claim. Still, we agree that this could be addressed, and we even see it as an empirical question testable in future work. Perhaps the most relevant departures between animate/inanimate and natural/manmade include objects like clouds, plants, and rocks — objects that are natural but not “animate” in the sense often used in this literature. If something like our paradigm revealed that rocks behave more like dogs than like boots, that would suggest that naturalness, rather than animacy, was driving the effects; but if rocks behave more like boots than like dogs, that would point to animacy even more strongly. We remain open-minded about this possibility, but it would of course require multiple new experiments with a brand new stimulus set and so goes beyond the present contribution. In any case, we have added a discussion of this issue to the paper and have adjusted our claims accordingly.

      Reviewer #3 (Public review):

      Summary:

      This study makes clever use of generative AI to create stimuli that are pixel-for-pixel identical but which have radically different meanings depending on their orientation, to investigate the perception of animacy while retaining control over low-level image features (so-called 'anagram' stimuli).

      The authors present seven elegantly designed experiments in a commendably compact format.

      Experiments 1 and 2 involved a working memory paradigm in which participants had to spot which of five objects in an array changed after a pause. Importantly, the changed object was an anagram stimulus that in one orientation matched the animacy/inanimacy of the changed object, and in the other orientation was the opposite (e.g., a rabbit is replaced by either a dog or a boot, where the dog and boot stimuli are actually identical, just rotated by 90 degrees). They found a difference in accuracy depending on whether the animacy of the objects matched.

      Experiments 3 and 4 used a visual search task in which the participants had to localize the target, and the distractors were anagrams that either matched the target in terms of animacy or did not. There was a significant cost in terms of response time when the animacy of the target was the same as that of the distractors. Experiments 5 and 6 also used a similar visual search design, except that the task was to determine if the target was present or absent from the display, and the distractors again either matched or differed from the target in terms of animacy. Again, the authors found slower responses when the distractor arrays matched the animacy of the target than when they differed.

      An obvious potential concern about the studies is addressed by Experiment 7. It is unclear if the observed effects are related to the specific orientations of the target and distractor stimuli selected in each condition. For example, it could be that all the animate versions of the anagrams involved tall and skinny shapes, while all the inanimate versions involved wide and short objects, due to the 90-degree rotational difference between the two versions of the stimuli. To control for this, the authors repeated the visual search experiment but with convex-hull silhouettes of each of the stimuli. In other words, all targets and distractors from each trial were replaced by a black splotch with approximately the same overall outline (envelope) as the corresponding stimulus. Importantly, in contrast to the anagram stimuli, the silhouettes had had no meaningful semantic interpretation, and their animacy did not change depending on their orientation.

      Strengths:

      The main strength is the elegant use of stimuli that control almost perfectly for low-level image features.

      Thank you for this kind feedback. This summary perfectly captures both our empirical contribution and the claims we are making.

      Weaknesses:

      My only real concern about the study is whether the findings truly provide evidence for a high-level visual representation of animacy independent of the low-level stimulus characteristics, or whether, instead, the effects are essentially semantic priming, which is independent of visual processing per se. For example, if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results? Words can also access semantic representations of the animacy of objects, and also don't suffer from low-level visual confounds. It would be helpful to add a discussion of this possibility to the article.

      Wow, we love this question! And so we’ve now conducted exactly the experiment the Reviewer suggests here. In a new pre-registered study (Experiment 8), we presented participants with a present/absent search task (as in Experiments 5–7). One half of trials consisted of the anagram stimuli (such that we could, once again, replicate the mixed-animacy search advantage); but the other half of trials consisted of the words describing the anagrams (e.g., “dog”, “boot”, “sheep”, “car”, etc.). The experiment worked beautifully: We found no effect with the words, but replicated the search advantage with the pictures — and also found a significant difference between the effects elicited by the two stimulus types.

      We agree with the Reviewer that this now rules out the possibility that semantic representations alone explain these visual effects. Thank you! 

      Reviewer #4 (Public review):

      In this article, the authors investigate whether perceived animacy influences visual processing independently of lower-level visual features by using "visual anagrams." Across seven experiments, they test whether animacy, isolated from many lower-level visual properties, structures visual working memory and guides visual attention. The central claim is that the visual system may represent animacy itself, rather than animacy emerging solely from associations among low-level visual properties.

      I find this investigation compelling. The experiments described provide strong control over several lower-level visual features, including curvature, texture, and related image properties. However, the visual anagrams are not pixelwise-identical across orientations. Because the images are rotated, the retinal configuration of pixels and the spatial organization of some low- to mid-level shape features also change. As a result, the configural arrangement of mid-level visual features may still contribute to perceived animacy.

      We are glad to hear the Reviewer finds our investigation “compelling”.

      I encourage the authors to discuss how independent perceived animacy is in this context from the contribution of mid-level visual features, such as configural shape cues that are diagnostic of animacy. This distinction would help sharpen the interpretation of the results and more precisely define the level of visual representation isolated by the visual-anagram approach.

      This is a helpful point, and it also echoes a sentiment expressed by Reviewer #2. While configural shape is diagnostic of animacy writ large, it can’t account for our observed effects here because rotating an image does not vary its configural shape. We now mention this in our work, and we agree that it helps sharpen the interpretation of our studies.

      Additionally, previous studies have argued that low- and mid-level curvilinear features may contribute to animate/inanimate categorization, and may in some cases be sufficient to support such distinctions (e.g., PMID: 33798259; PMID: 28654965). I encourage the authors to clarify how these previous findings on curvilinearity and rectilinearity fit with the overarching claim of the current study, namely that the visual system may represent animacy itself rather than animacy emerging solely from associations among lower-level visual properties.

      Yes, many studies from exactly that corner of the field actually motivated the present work, which is why we cited them in our submission. In a way, we are approaching this issue from the other side of the equation. Whereas the papers the Reviewer points to (along with many others) ask whether mid-level features (such as curvilinearity and rectilinearity) are sufficient to support perceived animacy, we ask whether these and other features are necessary to support perceived animacy. Prior work is relatively split on this issue, leaving the question wide open. We take our work to show that differences in curvature are not necessary for differences in perceived animacy, because our anagrams have identical curvature yet differ in animacy — and the visual system capitalizes on that difference. Put the other way around, representation of animacy can and does go beyond representation of its low- and mid-level correlates. Thank you!

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This study approaches an important topic providing insight into the neuronal circuitry that interconnects memory consolidation and sleep. The data were collected and analysed using a solid methodology, contributing new findings for neurobiologists working on how memories are stored and the roles of sleep. However, the data is incomplete to support the proposed role of the PAM-DPM circuits as the link between sleep state and long-term memory consolidation.

      We sincerely appreciate the editor and reviewers’ thoughtful and constructive comments on our study. Your insightful feedback has not only affirmed the significance of our work on the interplay between memory consolidation and sleep, but also provided valuable inputs for improving the clarity, rigour, and impact of our study.

      We have carefully addressed all the comments raised by the reviewers and revised the manuscript accordingly. We have also streamlined the paper with the goal of making it more accessible to readers. We feel this revised version strengthens our conclusion that the PAM-DPM circuits as the link between sleep and memory consolidation.

      The main improvements in terms of data addition are three complementary sets of circuit-specific experiments:

      (1) To better characterize the dynamics of the PAM-DPM circuit following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. This experiment addresses the activity of the microcircuit in a much more relevant time frame than the CRTC data in the previous version of the paper which looked only at the first hour after training. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, both PAM-α1 and DPM neurons displayed similar neural activity over the 3-hour recording period. The first hour was characterized by a synchronous decrease in activity for both neuron types, with hours 2 and 3 achieving a stable baseline. Since the decrease in the first hour is also seen in the trained condition, we think it is likely a reflection of the animals becoming acclimated to the recording tubes.

      Notably, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit in the LTM consolidation time window. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 8C-F). These findings not only support the hypothesized role of the inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics.

      (2) To further support the functional connectivity of the PAM-DPM microcircuit, we conducted in vivo experiments to complement the dissected brain prep P2X2 data. Optogenetic activation of PAM neurons in intact flies via the red light-gated cation channel CsChrimson (Klapoetke NC et al., 2014) resulted in a significant reduction in GCaMP signals within DPM neurons (revised Figure 2B). These findings strongly confirm that PAM neurons exert direct inhibitory control over DPM neurons in the intact brain.

      (3) Further, we investigated how dopamine signaling to the DPM inhibits its activity, and issue which has not been investigated previously. We conducted a series of experiments:

      Firstly, we verified which dopamine receptors (Dop1R1, Dop1R2, DopEcR, and Dop2R) express on the DPM neurons via double-labeling with gene-embedded GAL4 lines. We found that DPM neurons have expression of both Dop1R1 and Dop1R2 (revised Figure 10A).

      Secondly, to clarify which receptors on DPM neurons respond to dopamine and how they signal, in addition to EPAC experiments in the first submission, we recorded neural activity changes when we knocked down Dop1R1 and Dop1R2 in DPM neurons. DPM neurons exhibited a significantly reduced GCaMP level with DA application, regardless of whether Dop1R1 or Dop1R2 was intact or knocked down knockdown in comparison to the no-DA control condition (revised Supplemental Figure 3C-E). These data suggest that either residual Dop1R1 and Dop1R2 remaining in the RNAi condition is sufficient or that the two receptors may coordinate to mediate the inhibition of neural activity.

      Finally, we investigated the behavioral contributions of Dop1R1 and Dop1R2 in DPM neurons to sleep and memory processes (revised Figure 10C-H). Dop1R1 knockdown resulted in a marked reduction in daytime sleep and a significant impairment of 24 h memory expression. In contrast, Dop1R2 knockdown selectively compromised 24 h memory without affecting sleep.

      When integrated with our EPAC assay findings from the initial submission, which demonstrated, that Dop1R1 is the primary receptor mediating dopamine-induced cAMP elevation, these new data collectively delineate a more complex mechanistic framework: dopamine signaling in DPM neurons coordinates the dual regulation of sleep and memory predominantly via Dop1R1. Meanwhile, Dop1R2 are engaged in the selective modulation of memory.

      All newly generated experimental datasets, comprehensive statistical analyses, and their corresponding figure panels (revised Figures 2B, 10, 11 and Supplemental Figure 3) have been fully incorporated into the revised manuscript.

      In addition to adding the experiments described above, we have reorganized and streamlined the paper. First, the CRTC data have been replaced by the Tric-luc data. The CRTC data were taken in the first hour after training and do not shed light on the bulk of the consolidation window. Since the behavioral and sleep effects we see with manipulation of the PAM/DPM microcircuit all occur with a time delay, examining later times in consolidation is more relevant. Additionally, the first hour post-training is quite complex since there are sensory changes and STM processes overlaid on the processes we want to study. Second, we have moved the data in Figure 8 to supplemental (revised Supplemental Figure 2) since they are basically a control for the experiments in Figure 7 validating known requirements for appetitive LTM.

      We have also substantially expanded the Discussion section to contextualize the PAM-DPM circuit within the broader framework of well-characterized memory-regulatory pathways, such as the intrinsic circuits of the mushroom body, and to explicitly delineate the hierarchical interplay between sleep-dependent synaptic plasticity and LTM consolidation.

      We contend that these complementary experimental assays and targeted revisions markedly strengthen the causal evidence underscoring the role of the PAM-DPM circuit as a pivotal regulatory node bridging sleep states and LTM consolidation. We are confident that these revisions essentially address the concerns raised by the reviewers.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behavior, imaging, and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      From a Drosophila sleep researcher's perspective, the investigation follows a clear and logical strategy to collect a huge dataset of sleep, appetitive memory, and live imaging. The authors clearly identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation-dependent manner. The authors also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the authors applied a new way of sleep statistics to demonstrate hour-by-hour changes between treatment and genotypes. Importantly, the authors demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP during the memory consolidation period after training.

      Weaknesses:

      Two investigatory gaps relate to the misalignment between circuital activity and behaviors, due to the nature of large circuital functional analysis like this. Firstly, the central observation of the study indicates that PAM alpha1 activation causes DPM inhibition which disrupts sleep and memory consolidation. Therefore one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found that the endogenous CRTC::GFP reported neuronal activity for PAMalpha1 and DPM are both increased after memory training (Figure 9). This can be due to the difficult functional demarcation among the 14 PAMalpha1 projections. Secondly, the authors acknowledged the contradicting finding that memory defect is detected in PAMalpha1 inactivation (Figure 7C), yet suggested a tight link between sleep and memory consolidation; it is clear loss of PAM subset activity can disrupt memory consolidation without affecting sleep (cf Figure 7C and 7I).

      Thank you for your insightful analysis and the relevant possibilities you've raised. We agree that given that memory consolidation and sleep are time-dependent processes, the 1-hour window employed to capture neural activity changes via the CRTC::GFP reporter may not fully reflect the overall dynamics of neural activity in this microcircuit. To better characterize the dynamics of the PAM-DPM circuit in the consolidation window following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, PAM-α1 neurons displayed stable neural activity over the entire 3-hour recording period. However, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 8C-F). These findings not only support to the hypothesized role of inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics.

      Regarding the second question, the core finding underlying the link between sleep and memory elucidated in the present study lies in the whole PAM-α1-DPM microcircuit rather than the specific DANs alone. MB299B and MB043B, the two split-GAL4 drivers employed to target PAM-α1 neurons, were originally characterized previously (Aso et al., 2014). However, these drivers also exhibit non-specific labeling of additional cells, and we can not rule out the possibility that such off-target labeling may have masked the subtype-specific necessity in sleep or memory processes.

      Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain poorly understood. The authors present novel findings implicating two small neuronal groups with inhibitory connections, PAM-a1 to DPM, in sleep regulation and LTM consolidation. However, whether the PAM-a1 to DPM microcircuit promotes LTM consolidation through sleep regulation requires further investigation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-a1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly at night. Interestingly, disruption of PAM-a1 and DPM neurons impairs sleep and appetitive memory consolidation only under starvation conditions, and pharmacological induction of sleep during the night rescues the LTM defects. These findings suggest that PAM-a1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important findings that advance our understanding of the link between sleep and memory consolidation.

      Weaknesses

      Some claims lack sufficient evidence or clarity:

      (1) All sleep experiments are conducted under the "training" (temperature-change) condition. While genotypic controls are helpful, additional no-training controls are required to confirm that the observed differences are due to training rather than unknown genotype-related factors. The fact that experimental genotypes exhibit significantly altered sleep even before "training" (e.g., Figs. 7H, J, K, 8A, B, D) highlights the necessity of these controls.

      Thank you for raising this important question. We have re-examined the sleep profiles recorded over two acclimation days and one day of baseline sleep, which preceded the implementation of the “training” paradigm (temperature manipulation) and thus served as a valid no-training control. As shown in Author response images 1-4, subtle yet discernible genotype-dependent differences were indeed observed under baseline conditions. However, when animals were subjected to starvation, the experimental manipulations (activation or inactivation of the target cells) elicited marked, statistically significant alterations in sleep patterns that cannot be accounted for by the baseline genotype differences. Collectively, these data confirm that the observed sleep phenotypes are attributable to the “training” intervention, rather than to confounding, pre-existing genotype-related factors.

      Author response image 1.

      Baseline and manipulation day sleep profiles following PAM activation and PAM/DPM inactivation under starvation conditions.

      Author response image 2.

      Baseline and manipulation day sleep profiles following PAM activation and PAM/DPM inactivation under non-starvation conditions.

      Author response image 3.

      Baseline and manipulation day sleep profiles following PAM- α1 activation and inactivation under starvation conditions.

      Author response image 4.

      Baseline and manipulation day sleep profiles following PAM- α1 activation and inactivation under non-starvation conditions.

      (2) Previous studies on disrupted memory due to sleep reduction have primarily examined conditions with severe sleep deprivation. In contrast, this report claims that relatively small decreases in total sleep accompanied by sleep fragmentation are responsible for impaired memory consolidation. It remains unclear whether sleep fragmentation at this level is truly critical for memory consolidation. The authors should cause sleep loss and fragmentation of similar magnitude through other means and determine whether it can impair LTM.

      We appreciate the reviewer’s insightful suggestion. While alternative assays for inducing sleep loss or sleep fragmentation are indeed available, this line of investigation lies beyond the core scope of the present study. We will certainly take this valuable suggestion into consideration for the future studies.

      (3) The authors employed a neural activity reporter to show that starvation increases the basal activity of PAM-a1 but not DPM neurons in untrained flies (Figures 9C-E). They observed small increases in the activity of both neuron groups immediately after training but not one hour later. Given the inhibitory connection from PAM-a1 to DPM, it is unclear why both neuron groups show increased activity after training. Additionally, as the authors acknowledge, it is puzzling how the inactivation of PAM-a1 produces similar effects on sleep and memory as DPM inhibition and PAM-a1 activation. Further experiments are needed to clarify these findings, such as manipulating PAM-a1 activity during the one-hour post-training period and evaluating the effect on DPM activity. Including data from training under fed conditions would provide a more comprehensive understanding of state-dependent neural activity. Even if certain experiments are not feasible, these issues warrant further discussion. It is also important to clarify that the term "synchronized" does not imply single-spike-level synchrony.

      Thank you for raising these critical questions. To deepen our understanding of these issues, we have conducted additional experiments and have incorporated them into the revised manuscript. Below are our specific responses to each of your points:

      (1) Regarding the contradiction between "PAM-α1 inhibition of DPM" and a transient increase in the activity of both neurons immediately after training:

      PAM/PAM-α1 neurons are well-documented to respond to reward signals (Liu et al., 2012, Ichinose et al., 2015), while DPM neurons have been shown to respond to both olfactory stimuli and electric shocks, and to form delayed olfactory memory traces (Yu et al., 2005). Thus, the concurrent increase in the activity of PAM-α1 and DPM neurons immediately following training is likely a response to the olfactory and/or sucrose stimuli in the assay. Given that memory consolidation and sleep are time-dependent processes, the 1-hour window employed to capture neural activity changes via the CRTC::GFP reporter likely does not fully reflect the overall dynamics of neural activity in this microcircuit. Additionally, this time window overlaps with the period in which the animals are adapting to the new tubes and is likely contaminated with other sensory information.

      To better characterize the dynamics of the PAM-DPM circuit following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, PAM-α1 neurons displayed stable neural activity over the entire 3-hour recording period. Notably, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 9F-I). These findings not only support to the hypothesized role of inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics post-training. We have replaced the CRTC data with this more relevant data set.

      (2) Regarding the state-dependent neural activity:

      We agree that investigating state-dependent neural activity would be an interesting extension of our study. However, this falls beyond the scope of the current study and will be considered in future research. Our primary findings, including sleep disruptions and the associated memory impairments, were specifically observed under starvation conditions, which align with the appetitive memory paradigm employed here. Delving into neural activity changes under non-starvation state would not yield direct evidence to support the core conclusions of the present work, as the study’s focus is on the starvation-dependent interplay between sleep, neural circuitry, and appetitive memory consolidation.

      (3) Regarding the terminology of “synchronization”:

      We believe that the use of the term “synchronization” in our study is appropriate. In the context of neural circuitry, synchronization refers to the process by which distinct neurons or neural populations achieve temporal alignment of their activity, a phenomenon that supports neural communication and information integration. In the present work, this specifically describes how PAM-α1 and DPM neurons exhibit phase-related temporal coordination of their activity to regulate the interplay between sleep and memory consolidation.

      (4) The authors considered that PAM-a1 and DPM might function in parallel, independent pathways for sleep and LTM. They rejected this possibility based on the lack of additive effects when both neuronal groups were simultaneously inactivated. However, they found that MB299B-labelled neurons exert stronger memory effects than MB043B-labelled neurons, while MB043B neurons have stronger sleep effects. If sleep is a primary driver of memory consolidation, a stronger correlation between memory and sleep effects would be expected. This observation merits further discussion.

      We appreciate the reviewer’s constructive suggestions. We have performed additional experiments to explore a well-characterized memory-related PAM-α1 recurrent loop in sleep regulation. The new data, along with further discussion, have been incorporated into the revised manuscript.

      The two split-GAL4 drivers (MB299B and MB043B) used to target PAM-α1 neurons were originally characterized previously (Aso et al., 2014). However, these drivers exhibit non-specific labeling of additional neuronal populations, a technical limitation that may have masked the subtype-specific functional requirements of PAM-α1 in sleep and memory processes.

      In addition, we assessed sleep and LTM following the thermoactivation of DPM neurons (revised Supplemental Figure 1), and no significant changes were observed in either phenotype.

      PAM-α1 has previously been demonstrated to drive appetitive LTM formation and consolidation via a recurrent loop with MBON-α1 (Ichinose et al., 2015). To investigate whether MBON-α1 also participates in sleep regulation, we activated or inactivated MBON-α1 neurons under both starvation and non-starvation conditions. Our results revealed that inhibition of MBON-α1 under both starvation and non-starvation conditions resulted in a significant reduction in sleep and a reduced arousal threshold (revised Figure 11B, D), suggesting that MBON-α1 participates in regulating sleep in a state-independent manner. However, no significant changes were observed upon activation of MBON-α1 neurons (revised Figure 11A, C). Combined with our observation that inhibition of MBON-α1 during the memory consolidation phase also impaired 24 h LTM, these new data indicate that MBON-α1-mediated sleep is necessary for effective memory consolidation. Notably, while activation of MBON-α1 during consolidation phase similarly impaired LTM, this manipulation did not alter the sleep profile, suggesting a dissociation between MBON-α1’s mechanistic roles in sleep regulation and LTM processing.

      Taken together (see Author response table 1 and the new schematic diagram of revised Figure 12), these findings reveal a dedicated hierarchical, modular regulatory network that mediates sleep-LTM coupling via an activity-dependent mechanism. Within this network, activation of PAM-α1 acts as an upstream modulator to inhibit the activity of DPM, a downstream integrative hub that coordinates the execution of sleep and memory processes via recruiting different signaling cascades mediated by distinct dopamine receptors. MBON-α1, which is likely inhibited by PAM-α1, serves as parallel pathway to suppress sleep and impair LTM. Conversely, inactivation of PAM-α1 relieves its inhibitory control over MBON-α1, leading to MBON-α1 activation; MBON-α1 then functions as a signal amplifier that further exacerbates the reduced activity of PAM-α1, ultimately resulting in LTM impairment. Inactivation of PAM-α1, together with non-PAM-α1 neurons labeled by MB043B, contributes to the regulation of sleep. Sleep and memory are highly intertwined within this circuit, where distinct neuronal populations exhibit specialized yet interdependent functional roles, with overlapping and divergent regulatory contributions to sleep and LTM. The inherent complexity of this regulatory network thus merits further dedicated investigation in future studies.

      Author response table 1.

      (5) Given prior knowledge that PAM neurons are heterogeneous and that the R58E02 driver is broadly expressed, data in Figures 1-5 concerning PAM are outdated. The use of more restricted PAM-a1 drivers from the outset would make the manuscript easier to read and interpret.

      We sincerely appreciate the reviewer’s point of view regarding the selection of PAM drivers. While we acknowledge the well-characterized heterogeneity of PAM neurons and the broad expression profile of the R58E02 driver, and fully agree that employing subtype-restricted drivers enhances the precision of functional interpretation, this set of experiments serves as an essential foundational step and logical basis for subsequent subtype-specific investigations and thus merits retention in the manuscript. As detailed above, the more specific drivers also have some drawbacks in terms of additional expression, making the broad driver critical for setting the stage.

      (6) Some figures lack relevant data, certain experiments are missing necessary controls, and anomalies are present in some data sets.

      We sincerely appreciate the reviewer’s detailed suggestions, and we have revised the manuscript comprehensively in accordance with them.

      Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence that dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strengths:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuits interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of dopamine's role in cognitive functions.

      Weaknesses:

      While the study is well-designed and presents compelling findings, some aspects require further clarification. The interpretation of dopamine receptor signaling remains incomplete, particularly regarding inhibitory pathways. The role of DPM in memory consolidation is not entirely conclusive, as different genetic approaches yield variable results. Additionally, some inconsistencies in neuronal activity patterns and experimental variability, especially regarding sleep patterns or pharmacological rescue, should be addressed to strengthen the mechanistic framework.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopamine circuits interact to regulate memory consolidation. While the findings are compelling, addressing the points above-particularly receptor signaling and the specific role of DPM and its activity patterns within the microcircuit would further solidify the study's conclusions.

      We sincerely appreciate the reviewer’s constructive feedback and useful suggestions, which have been instrumental in enhancing the rigour and completeness of our study.

      To address these points, we have performed a series of additional experiments that we believe strengthen the mechanistic framework of our work. The key new findings are summarized below:

      (1) Regarding the dopamine receptor signaling

      To define the dopamine receptor (DAR) signaling mechanisms underlying DPM neuron activity and its regulatory roles in sleep and memory, we first characterized DAR expression profile of the DPM neurons. Using double-labeling assay, we detected robust expression of Dop1R1 and Dop1R2 in DPM neurons, whereas no detectable colocalization was observed for DopEcR and Dop2R (revised Figure 10A). Accordingly, we refined our FRET-based EPAC data by removing the DopEcR knockdown group, and now present cAMP changes in DPM neurons following Dop1R1 and Dop1R2 knockdown, in direct comparison with the intact receptor control group (revised Figure 10B). These data conform that Gαs-coupled Dop1R1 is the primary receptor mediating DA-dependent cAMP elevation in DPM neurons.

      To further identify the DARs responsible for transducing DA-induced inhibitory effect on DPM neural activity, we quantified GCaMP levels in DPM neurons with targeted knockdown of individual DARs. Knockdown of either Dop1R1 or Dop1R2 failed to abolish DA-induced Ca<sup>2+</sup> decrease; only Dop1R2 knockdown exhibited a trend toward attenuating this Ca<sup>2+</sup> decrease (revised Supplemental Figure 3C-E), suggesting that the two receptors cooperate to modulate DPM neural activity.

      Finally, to dissect the specific contributions of DARs in DPM neurons to sleep and/or memory regulation, we performed sleep monitoring and memory assays in animals with DPM-specific knockdown of distinct DARs (revised Figure 10C-H). Knockdown Dop1R1 in DPM neurons resulted in statistically significant sleep reduction, decreased arousal threshold, and impaired 24 h LTM memory (revised Figure 10C-E). In contrast, knockdown Dop1R2 in DPM neuron selectively impaired 24 h LTM memory with no effect on sleep (revised Figure 10F-H). Collectively, these findings demonstrate that coupling sleep and LTM requires Dop1R1 in DPM neurons through the modulation of both cAMP signaling and neuronal activity, while Dop1R2 specifically mediates LTM regulation, likely through modulating DPM neural activity alone.

      (2) We have additionally characterized the role of MBON-α1 in sleep, which has been previously shown as a PAM-α1-related recurrent feedback loop in the regulation of memory formation and consolidation (Ichinose et al., 2015).

      To investigate whether MBON-α1 also participates in sleep regulation, we activated or inactivated MBON-α1 neurons under both starvation and non-starvation conditions (revised Figure 11A-D). Our results revealed that inhibition of MBON-α1 under both starvation and non-starvation conditions resulted in a significant reduction in sleep and a reduced arousal threshold (revised Figure 11B, D), suggesting that MBON-α1 participates in regulating sleep in a state-independent manner. However, no significant changes were observed upon activation of MBON-α1 neurons (revised Figure 11A, C). Moreover, inhibition of MBON-α1 during the memory consolidation phase significantly impaired 24 h LTM (revised Figure 11E-F). These results indicate that MBON-α1mediated sleep is necessary for effective memory consolidation. Notably, while activation of MBON-α1 during consolidation phase similarly impaired LTM, this manipulation did not alter the sleep profile, suggesting a dissociation between MBON-α1’s mechanistic roles in sleep regulation and LTM processing.

      Taken together (see Author response table 1 and the new schematic diagram of revised Figure 12), these findings reveal a dedicated hierarchical, modular regulatory network that mediates sleep-LTM coupling via an activity-dependent mechanism. Within this network, activation of PAM-α1 acts as an upstream modulator to inhibit the activity of DPM, a downstream integrative hub that coordinates the execution of sleep and memory processes via recruiting different signaling cascades mediated by distinct dopamine receptors. MBON-α1, which is likely inhibited by PAM-α1, serves as parallel pathway to suppress sleep and impair LTM. Conversely, inactivation of PAM-α1 relieves its inhibitory control over MBON-α1, leading to MBON-α1 activation; MBON-α1 then functions as a signal amplifier that further exacerbates the reduced activity of PAM-α1, ultimately resulting in LTM impairment. Inactivation of PAM-α1, together with non-PAM-α1 neurons labeled by MB043B, contributes to the regulation of sleep. Sleep and memory are highly intertwined within this circuit, where distinct neuronal populations exhibit specialized yet interdependent functional roles, with overlapping and divergent regulatory contributions to sleep and LTM. The inherent complexity of this regulatory network thus merits further dedicated investigation in future studies (See Author response table 1).

      We have modified the schematic diagram in the revised manuscript to illustrate the mechanistic framework (revised Figure 12).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here I listed details for potential clarification or further investigation related to the weaknesses:

      (1) Line 145-147: I suspected the authors used previously verified RNAi lines, but it would be informative to include a citation or their own validation for the effectiveness of these RNAi lines.

      We sincerely appreciate the reviewer’s suggestion. As our double-labeling assays confirmed that only Dop1R1 and Dop1R2 are colocalized with DPM neurons (see our responses to the public review from Reviewer #2 and #3), we refined the revised data to focus exclusively on these two DARs. Corresponding revisions have been made to the Materials and Methods, Results and Discussion sections. Additionally, we conducted qPCR analysis to verify the knockdown efficiency of these DARs, providing further support for our findings that Dop1R1 and Dop1R2 are functionally required in DPM neurons for the regulation of sleep and memory (revised Supplemental Figure 3).

      (2) Line 169: Moving from describing Figure 3A/B to Figure 3C, it is not immediately clear from 3C-H, the authors follow the training paradigm of 3B?

      To enhance clarity, we have added the referenced figure citations in the “Memory assay” section: “For all 24 h sucrose-odour memory, a single training session of sucrose paired with an odour for 2 min was employed (Figure 3A).”

      (3) Line 179: before the PAM inactivation data are shown in Figure 7, the authors seem to be getting ahead of themselves by stating "suggesting that activity of PAM neurons is necessary for the consolidation window or that heterogeneity in the subsets of PAM neurons masks any phenotype." when the Figure 3 data collectively indicate that "suppression" of PAM is necessary.

      This statement is based on our observation that inactivation of the majority of PAM neurons labeled by R58E02 results in sleep disruption but leaves memory intact, and we stand by this conclusion.

      (4) Line 186-188: The labelling of DP1 is not entirely aligned between the figures and the text for a reader to follow which time period is described, as DP1 is embedded within the dark phase in the figures.

      These experiments spanned two consecutive days. LP1 and DP1 denote the light and dark periods on the first day, respectively, whereas LP2 designates the light period on the second day. As only one full dark phase was monitored across the experimental interval, we characterized the relevant phenotype using the general terms dark phase or nighttime, rather than specifying DP1. We thank the reviewer for this thoughtful observation; nonetheless, we consider the original description correct and unambiguous, and thus appropriate for inclusion in the manuscript.

      (5) Line 321: The statistics for Figure 9 CRTC::GFP measurement is crucial for interpretation but the referring and labelling for this on Figure 9 is poor: it is not apparent which comparisons are indicated. There is inconsistency between Table 1 Figure 9D and Table 3 Figure 9D entries: no significant between train and untrain indicated in Table 1 but it is described as significant in the text and Table 3?

      Thank you for this observation. As described above, we have removed these data from the paper and replaced them with Tric-Luc data that capture the consolidation window more completely.

      (6) Line 423: The starvation-mediated sleep suppression is not clear in this manuscript, can the author comment on this? The response to this may also alter the summary concept cartoon.

      This is an important point. To directly address the reviewer’s question regarding starvation-mediated sleep suppression, we have generated a representative response figure comparing sleep duration under starvation versus non-starvation conditions (Author response image 5). This figure clearly demonstrates that sleep is suppressed under starvation, providing straightforward evidence to address this concern.

      Author response image 5.

      Examples of starvation-mediated sleep suppression.

      However, the key focus of our study is that changes in neuronal activity disrupt sleep under starvation conditions but not under non-starvation conditions. To emphasize this critical distinction, we have incorporated additional discussion focused specifically on this point.

      “It is well established that starvation induces sleep suppression (MacFadyen, 1973; Thimgan et al., 2010; Melnattur and Shaw, 2019; Keene et al., 2010; He et al., 2020; Yangkyun et al., 2022), and our results are consistent with these previous findings: all genotypes exhibited less sleep under starvation than under fed conditions (i.e. Figures 4A-B, 5A-B and 11). Under normal appetitive memory training, starvation-induced sleep loss does not necessarily impair memory processing (Thimgan et al., 2010; Chouhan et al., 2021). PAM-α1 neuronal activity is higher in starved, trained flies than in fed or untrained flies (data not shown), suggesting that these neurons act as a critical node for integrating internal motivational and arousal states, as well as conveying positive valence for the normal appetitive memory process, independently of starvation-induced sleep loss. While DPM neurons are less sensitive to starvation, they still exhibit training-induced elevated activity (data not shown), indicating coherent responsiveness to upstream signaling. In the present study, we found that under fed conditions, sleep remained intact even when excessive changes in neural activity occurred within the PAM(-α1)-DPM circuit; in contrast, under starvation conditions, significant sleep reduction and fragmentation were observed. These observations indicate that starvation may trigger a transition from a physiologically normal brain state to an unstable, abnormally active state, which consequently elicits negative behavioral outputs.”

      (7) Line 1121: The data points for Figure 9 D-E are surprisingly low considering there are 14 PAMalpha1 labelled, the data presented here indicated potentially only 1-2 neurons were counted per fly brain. Can this contribute to the large variation and the contradiction of PAM's memory-suppressing role?

      We sincerely appreciate the reviewer’s critical comments regarding the sample size of labeled PAM-α1 neurons in Figure 9D–E. We have revisited our raw data, incorporated additional brain samples, and reanalyzed the dataset. For this updated analysis, we included all clearly distinguished neurons, excluded overlapping ones, and calculated a single NLI per brain for statistical analysis. The key conclusions remain consistent with those in the original submission, confirming the robustness of the observed phenotype.

      Memory consolidation is a time-dependent process. To further elucidate the link between neural activity and behavioral outputs, we performed additional experiments with an extended recording period. A detailed response to this point is provided in the response to public review, and we therefore do not reiterate the details here.

      (8) Line 345: the effect size and data spread of THIP restored memory is different from the controls in Figure 10, perhaps warranting a more conservative interpretation of the role of sleep in memory consolidation.

      We appreciate this critical comment. We fully agree that the role of sleep in memory consolidation requires cautious interpretation, a point we have integrated into the revised manuscript.

      Drug treatment in Drosophila, particularly for group-based assays, can introduce substantial variability at both the individual and group levels. To account for this, we employed a statistically valid sample size for our analyses to ensure robust conclusions. While minor quantitative discrepancies exist in the data, this technical consideration does not significantly alter the core conclusions of the study.

      Reviewer #2 (Recommendations for the authors):

      (1) As mentioned in the public review, all data using the broad PAM-DAN driver should be removed. Concerns regarding the experiments involving the broad driver are not included here.

      A detailed response to this point is provided in the response to public review, and we therefore do not reiterate the details here.

      (2) In GCaMP experiments (Figure 9B), the ΔF/F traces for the AHL and AHL+ATP conditions start diverging before the addition of ATP. The quantification shows they are not significantly different in the first 30 seconds, but the fact that in two separate experiments (2A and 9B), they diverge in the same direction makes me wonder whether the AHL condition is different from the +ATP condition even before the ATP treatment. Also, the traces should include standard errors.

      We observed the same diverging trend in the first 30-second baseline as the reviewer. We reviewed the raw data for each sample and found that this divergence is likely attributable a small number of outliers. Given the absence of a statistically significant difference, this divergence does not affect our conclusions.

      We have also added standard errors to the revised figures.

      (3) Figure 9B. The authors need to show data for a control genotype. +>P2X2; VT064246-LexA > GCaMP6f that does not include MB299B-Gal4 is crucial to demonstrate that expression of P2X2 in PAM-α1 is responsible for the inhibitor effect, as LexA-P2X2 may be leaky.

      One of the UAS-P2X2 lines was found to exhibit leaky expression, so we instead used a non-leaky UAS-P2X2 line for all related experiments. To address the reviewer’s comments and further validate our findings, we have added complementary experiments with a control genotype. In addition, we also added a control to confirm the non-leaky expression of LexA-P2X2 under the driver of R58E02-LexA. As shown in revised Figures 8B, application of ATP in the absence of MB299B-GAL4 failed to induce a significant inhibitory effect. These data strongly and convincingly support our conclusion.

      (4) Figure 9B. Some of the individual data show values lower than -100% ΔF/F0. By definition, ΔF/F cannot be less than -100%, as this would require negative fluorescence, which is physically impossible. The calculation of fluorescence changes using ΔF/F should be carefully reconsidered.

      We thank the reviewer pointing out this potential confusion. We used a standard method of calculating the change in fluorescence over time using △F/F = (Fn-F<sub>0</sub>) / F<sub>0</sub>×100% as we previously described (Liu et al., 2019). Changes of greater than +100% of △F/F would not be unusual, since the reported value is a ratio to the initial level of fluorescence, not a subtraction of the baseline value from the signal (which obviously could not go below 100%). We have included a sentence in the results explaining this (page 7): “As previously described, we used the percent change in fluorescence over time as a ratio to the initial level, △F/F = (Fn-F0)/F0×100% for quantification (Liu et al., 2019).” And we have carefully reviewed our raw and processed data and confirmed that our analysis was correct.

      (5) Figure 2B. The number of UAS transgenes should be controlled, as Gal4 could be diluted with 3 UAS constructs in experimental conditions compared to only 1 UAS construct in controls. Are Dop1R2 and DopEcR significantly different from wt? Why do they present an average ΔF/F in 2A and a maximum in 2B?

      We appreciate the reviewer’s careful observations and valuable comments.

      As the reviewer noted, the EPAC imaging experiment utilizes three UAS transgenes, which enable Gal4 enhancement via Dicer, targeted manipulation of dopamine receptor expression levels, and neural activity monitoring in DPM neurons. All other imaging experiments in the study employ only one or two UAS transgenes. Given the robustness of the observed phenotypes, the potential dilution effect is not a major concern. Knockdown of Dop1R2 and DopEcR showed no significant differences relative to the WT control group; the maximum values presented in Fig. 2B are included solely to illustrate statistical significance. While the EPAC (CFP/YPF) signal reflects an obvious cAMP elevation, no differences were detected in the averaged signal across groups.

      Notably, in the revised manuscript, our double-labeling assays confirmed that only Dop1R1 and Dop1R2 are colocalized with DPM neurons (see our responses to the public review from Reviewer #2). Accordingly, we have refined our data analysis to focus exclusively on these two DARs.

      (6) Figures 7H, J. Why is almost every MB299B>TrpA1 fly sleeping at ZT0?

      To align the starvation protocol for sleep analysis with that used in the memory assay, MB299B>TrpA1 flies and their genetic controls were transferred to fresh sleep tubes containing starvation food during the ZT0–1 time window. This transfer resulted in no detectable locomotor activity during this period, a pattern indicative of sleep in all flies.

      (7) The number of episodes and P(wake) should be presented for all sleep data.

      We have added these two parameters as new panels to all relevant sleep figures. The corresponding statistical analyses have also been included in the supplemental tables.

      Reviewer #3 (Recommendations for the authors):

      The study's findings provide compelling insights into the neural circuits connecting sleep and memory and the role of dopamine in general. While the anatomic dissection of the microcircuit and its overall involvement in sleep and memory is convincing and of high interest to the field and beyond, some statements of the study need further clarification, particularly the interpretation of receptor signaling and the role of DPM.

      Major Points

      (1) Figure 2: cAMP Imaging and Dopamine Receptor Involvement

      The authors present calcium and cAMP imaging to support the inhibitory connection between PAM and DPM neurons. While using both sensors is a robust approach, I am not entirely convinced that cAMP imaging is the ideal approach for identifying the dopamine receptors involved. To my knowledge, only Dop1R1 is classically linked to Gs-mediated cAMP signaling. Dop1R2 is typically coupled to Gq (PLC and DAG), while DopEcR is non-canonical and can engage both pathways. Additionally, these receptors are classically excitatory, yet the authors did not analyze Dop2R, the primary inhibitory dopamine receptor - which would represent the most relevant candidate for an inhibitory PAM-DPM connection.

      We have addressed this point in our response to the public comments, so will not reiterate here.

      While dopamine receptor functions can vary by neuronal context, I would appreciate clarification on the following points:

      (a) Why was Dop2R not tested? Was it omitted or found to have no effect?

      We sincerely appreciate the reviewer’s critical questions. This point has been addressed in our response to the public comments. Briefly, Dop2R is not colocalized with DPM neurons; instead, only Dop1R1 and Dop1R2 are detected in DPM neurons, which is why we focused exclusively on these two receptors in the revised manuscript.

      (b) Why was cAMP imaging chosen for receptor identification? Was calcium imaging performed, and if so, what were the results?

      This is an excellent point, and we sincerely appreciate the reviewer’s valuable input, which has helped to strengthen the logical framework of our analysis on receptor-mediated neural activity. These dopamine receptors are well-characterized as members of the Gas-coupled protein receptor family, and cAMP signaling serves as a reliable readout of their functional activity. To strengthen the logic flow of our analysis on the target inhibitory circuit, we have made the following key revisions to the manuscript: 1) defined the expression profile of dopamine receptors in DPM neurons; 2) refined our cAMP imaging data analyses based on specific receptor subtypes; and 3) assessed DPM neural activity via calcium imaging under conditions of targeted receptor knockdown. For further details, please refer to our response to the public comments.

      (c) Since the data suggest multiple receptor involvements and complex interactions, I encourage a more detailed discussion of the working hypothesis, particularly regarding the unexpected finding that classically excitatory receptors contribute to an inhibitory connection.

      We appreciate the suggestion to elaborate on our working model. Accordingly, we have revised the schematic diagram and refined the manuscript to clearly illustrate the underlying mechanistic framework. For further details, please refer to our response to the public comments.

      (2) Figure 3: DPM Involvement in Memory Consolidation

      The authors show that PAM activation and DPM inhibition during consolidation impair appetitive LTM. However, the role of DPM is critical. While the c316-GAL4 driver yields strong effects, VT064246 inhibition shows only slight significance, requiring more than twice the sample size of other experiments. Given that c316-GAL4 is not DPM-specific and also labels MB Kenyon cells, I suggest using MB-GAL80 to restrict expression - or commenting on the possibility that other neurons like MB-KCs could directly participate in the phenotype. This is particularly relevant since VT064246 efficiently modulates sleep, indicating that it is generally effective in altering behavior. These issues weaken the claim that DPM plays a crucial role in linking sleep and memory, and should be addressed. Minor comment on this Figure: In the Figure legend, the driver and "n" are not mentioned for 3C, while this is the case for all other panels. Moreover, the DPM schematic only depicts the MB, making it somewhat confusing. DPM innervates the entire MB, still, it would be helpful to shade the DPM projections more distinctly within the MB for clarity.

      We thank the reviewer for the suggestion to improve the precision of our figures.

      Regarding the expression specificity concern, in all experiments using c316-GAL4, we had eyeless-GAL80 and MB-GAL80 co-expressed to restrict GAL4-driven expression to DPMs. While complete suppression of expression of MB-KCs was not achievable, we largely eliminated the potential confounding effects from majority of these cells. VT064246-GAL4 is known to exhibit weak expression (Jenett et al., 2011; Haynes et al., 2015), but high relative specificity. Importantly, the overall conclusion derived from experiments using c316-GAL4 with GAL80s and VT064246-GAL4 are consistent, which strongly supports the role of DPM neurons in mediating the link between sleep and memory.

      As suggested, we have added sample sizes for all panels and refined the depiction of DPM projections in revised Figure 3C.

      Minor Comments

      (1) Introduction:

      The authors introduce dopamine's role in forgetting but focus on aversive rather than appetitive memories. To avoid confusion, this distinction should be mentioned explicitly (likewise in the discussion). Regarding references: Zhang et al. (line 95) do not discuss DPM or APL. Donlea et al. (line 97) do not cover dopamine - I think Pimentel et al. (2016) would be a more appropriate citation.

      This is a good point. We have removed Zhang et al. (2013) and replaced Donlea et al with Pimentel et al. 2016 as suggested.

      (2) Figure 9: DPM Activation During Consolidation:

      The authors show that PAM neurons are activated by starvation and further enhanced by appetitive training. Surprisingly, DPM neurons also increase activity post-training, despite the proposed inhibitory connection between PAM and DPM. The authors state that "PAM-α1-DPM microcircuit exhibits synchronized neural activity changes during the consolidation window" (line 326), yet they do not address this apparent contradiction. If I have not overlooked key information, this should be clarified/addressed e.g. in the discussion.

      This is an excellent point. We have addressed this in our response to point (3) from Reviewer #2 in the public comments, so we will not reiterate it here.

      (3) Figure 10D/E: THIP Rescue of LTM Deficits:

      Some inconsistencies in the THIP rescue experiments need clarification:

      (a) In Figure 10D, MB299B activation with THIP appears not to significantly restore memory relative to zero, nor to differ from untreated conditions in Figures 7A or 10E.

      (b) In Figure 10E, MB299B activation +/- THIP shows a much clearer effect.

      (c) Are Figures 7A, 10E, and 10D independent experiments, or were they conducted together?

      (d) Should the left bar in 10D and the right bar in 10E be identical? If not, I do not fully understand the discrepancy and suggest discussing the variation.

      Upon revisiting the raw datasets and conducting a one-sample t-test to analyze the group differences, the experimental group in Figure 7A showed no significant difference from the theoretical mean (set at zero). This group also did not differ from the two genetic controls, indicating that the restored memory was comparable to control levels. In Figure 10E, the group with MB299B activation plus THIP treatment exhibited a significant difference from the theoretical mean (one-sample t-test) and from the non-THIP control group, confirming a significant restoration of memory function. Owing to our laboratory relocation, the starvation duration at the new facility was adjusted based on a recalibrated starvation curve; the higher overall 24 h memory index in Figure 10E is likely attributable to a relatively longer starvation period. However, this experimental parameter variation does not alter the study’s overall conclusions.

      (4) Sleep Phenotypes and Starvation Effects:

      Sleep scores are shown under starvation/fed conditions but not under baseline conditions (without inhibition/activation). Could the authors indicate whether they observe basal starvation-induced sleep changes? The authors frequently state that PAM-DPM effects on sleep are context-dependent, yet mild but significant changes occur under fed conditions. I suggest rewording to clarify that the effect is enhanced in a context-dependent manner rather than strictly context-dependent.

      Starvation-induced sleep reduction is a well-characterised phenotype. Our study focused on the key question of whether altered neuronal activity modulates sleep under innate starvation conditions. Accordingly, all comparisons were made between the experimental and control groups under both starvation and fed conditions. We appreciate the reviewer’s suggestion to improve clarity and have revised the text as suggested.

      (5) Starvation Duration in Methods:

      The authors use 30-46h of starvation, which is longer than the ~20h typically used in appetitive memory studies. Could the authors explain why such extended starvation times were necessary?

      Determining starvation levels via survival curves is a well-established and relatively objective method, one that has been widely adopted in prior studies. For the memory test, we standardized the total starvation duration for each genotype to the time point at which mortality reached 20%. Owing to inherent differences in to starvation resistance across distinct genotypes, the final starvation durations ranged from 20 hours to 46 hours.

      (6) Variability in PAM-α1 Sleep Effects:

      (a) The extent and timing of sleep effects differ across PAM-α1 drivers (e.g. night vs. light-period effects). Could MBON co-targeting by these drivers contribute to the variability?

      We have supplemented additional experiments to investigate the effects of MBON-α1 neurons on 24h memory and sleep. For detailed findings, please refer to our response to your public comments.

      (b) Even within the same driver, results differ (e.g., Figure 7H vs. 10A). A general comment on these differences would be important, e.g. regarding the relevance of day and night sleep for memory consolidation.

      We sincerely appreciate the reviewer’s incisive observation regarding these details. The discrepancy stems from the timing of neuronal activity inhibition, during which a laboratory relocation led to adjustments in starvation duration for memory experiments, which in turn indirectly altered sleep patterns.

      (c) Technical note: Similar y-axis scales for sleep plots (Figures 10A and B) would make comparison easier.

      We have unified the y-axis scales to the same range.

      (7) Discussion, Line 376:

      The phrase "sleep deprivation is important for memory consolidation" is misleading, as it could imply that deprivation aids memory formation. Please clarify.

      We appreciate the reviewer’s suggestion. We have revised the text to: “These results demonstrate that preserving unperturbed sleep during the critical memory consolidation window is essential for stabilizing appetitive long-term memory.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study builds upon a major theoretical account of value-based choice, the 'attentional drift diffusion model' (aDDM), and examines whether and how this might be implemented in the human brain using functional magnetic resonance imaging (fMRI). The aDDM states that the process of internal evidence accumulation across time should be weighted by the decision maker's gaze, with more weight being assigned to the currently fixated item. The present study aims to test whether there are (a) regions of the brain where signals related to the currently presented value are affected by the participant's gaze; (b) regions of the brain where previously accumulated information is weighted by gaze.

      To examine this, the authors developed a novel paradigm that allowed them to dissociate currently and previously presented evidence, at a timescale amenable to measuring neural responses with fMRI. They asked participants to choose between bundles or 'lotteries' of food times, which they revealed sequentially and slowly to the participant across time. This allowed modelling of the haemodynamic response to each new observation in the lottery, separately for previously accumulated and currently presented evidence.

      Using this approach, they find that regions of the brain supporting valuation (vmPFC and ventral striatum) have responses reflecting gaze-weighted valuation of the currently presented item, where as regions previously associated with evidence accumulation (preSMA and IPS) have responses reflected gaze-weighted modulation of previously accumulated evidence.

      A major strength of the current paper is the design of the task, nicely allowing the researchers to examine evidence accumulation across time despite using a technique with poor temporal resolution. The dissociation between currently presented and previously accumulated evidence in different brain regions in GLM1 (before gazeweighting), as presented in Figure 5, is already compelling. The result that regions such as preSMA response positively to |AV| (absolute difference in accumulated value) is particularly interesting, as it would seem that the 'decision conflict' account of this region's activity might predict the exact opposite result. Additionally, the behaviour has been well modelled at the end of the paper when examining temporal weighting functions across the multiple samples.

      In response to reviewer comments, the authors have explicitly tested for the effects of gaze-weighting over and above any main effect of value, and convincingly shown that these effects are both present in the main regions of interest - namely |SV| and gazeweighted |SV| in the vmPFC, alongside |AV| and |AV_gaze| in the pre-SMA. This provides clear evidence in support of the notion of gaze-weighting of value signals in these regions.

      We thank the reviewer for their comments.

      Reviewer #2 (Public review):

      Summary:

      In this paper the authors seek to disentangle brain areas that encode the subjective value of individual stimuli/items (input regions) from those that accumulate those values into decision variables (integrators) for value-based choice. The authors used a novel task in which stimulus presentation was slowed down to ensure that such a dissociation was possible using fMRI despite its relatively low temporal resolution. In addition, the authors leveraged the fact that gaze increases item value, providing a means of distinguishing brain regions that encode decision variables from those that encode other quantities such as conflict or time-on-task. The authors adopt a region-of-interest approach based on an extensive previous literature and found that the ventral striatum and vmPFC correlated with the item values and not their accumulation whereas the preSMA, IPS and dlPFC correlated more strongly with their accumulation. Further analysis revealed that the pre-SMA was the only one of the three integrator regions to also exhibit gaze modulation.

      The study uses a highly innovative design and addresses an important and timely topic. The manuscript is well-written and engaging, while the data analysis appears highly rigorous.

      Weaknesses:

      With 23 subjects the study has relatively low statistical power for fMRI although the within-subjects design and relatively high trial count reduces these concerns.

      We thank the reviewer for their comments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Something seems to have gone slightly wrong in (I think) the labelling of new figure 7 (the correlation matrix between the different regressors). There are five variables on both the x- and y-axes of the figure, and they are the same five variables - meaning the diagonal of the matrix would be all equal to 1 (being the correlation of a regressor with itself - e.g. |AVgaze| with |AVgaze|). In the figure legend, six variables are mentioned, including lagged |deltaAVgaze| - but this doesn't appear on the x or y-axes. I suspect that the authors may need to check that this matrix has been calculated correctly, and isn't being mislabelled?

      We thank the reviewer for noticing this issue. We have now corrected Figure 7.

    1. Reviewer #2 (Public review):

      Summary:

      A relatively new measure of flexible brain state engagement (SEV - State Engagement Variability) is used here. It simply measures time-to-time variation in brain activity in terms of how it matches pre-specified motifs of activity. This metric seems to be predictive of behavioural data measuring cognitive control abilities. This was found to be the case in two independent datasets with different (though related) behavioural measures.

      Strengths:

      Use of multiple datasets is a clear strength. The use of both replication and out-of-sample model prediction is another.

      Weaknesses:

      (1) It is not clear to me how specific the SEV metric is for telling us about brain state engagement flexibility. Resting state fluctuations have been described as quasi-periodic changes that can be mapped onto "states", but the fluctuations could easily be a reflection of vascular flow, which may indirectly correlate with cognition.

      (2) If SEV is calculated using other state descriptors (e.g. a random parcellation of the brain into 4 networks) - would the result still hold? Or are the motifs important (this would rule out, to some extent, the vascular argument from (1) above)?

      (3) Figure 1 confused me a little. Why not show all the combinations (patient v full sample), inhibition vs shift, and main vs validation? Instead, a subset of 4 was selected?

      (4) The inhibition/patient/main correlation seems to be driven by 4 patients with particularly high inhibition measures?

      (5) Why is SEV negative in some cases (e.g., Figure 1) if it's a std measure? Has it been demeaned or orthogonalised wrt another variable?

      (6) The external analysis is great, but why should the model predict a relationship between SEV and inhibition if the claim is that it is only true for patients? Why would it only be true for patients in the first place?

      (7) I can't get my head around the results shown in Figure 3. How can one have both positive and negative correlations being significant or meaningful in the same pairs of networks? I think this set of results could benefit from more explanation.

      (8) I struggled with Figure 4 analysis. What is the SEV network? How do we know that it is specific enough to the SEV concept? Looking at co-fluctuations with the cognitive network, are we not simply looking at the old anti-correlation between the default mode and the rest of the brain (I note that the correlations in the y-axes of Figure 4 are negative)?

    1. Reviewer #1 (Public review):

      The authors address a difficult and well-known problem in systems/computational neuroscience: how to estimate the magnitude of "information-limiting" noise. Existing approaches (direct Fisher-information estimation, decoding + Cramer-Rao, and large-N extrapolation) are data-hungry and unstable, which has left the field with conflicting empirical estimates across systems.

      The central proposal - "split-trial analysis" - is simple and appealing. The recorded population is randomly partitioned into two non-overlapping halves; a decoder (continuous case) or classifier (binary case) is trained on each half using the same trials; and the covariance of the two halves' decoding errors is used to estimate the variance of the information-limiting noise. There is a clean mathematical derivation to support this conclusion (although there are a couple of mathematical errors in the methods section that should be fixed to avoid confusion on the part of the reader).

      They benchmark the method in simulation against three prior methods (Moreno-Bote et al. 2014; Rumyantsev et al. 2020; Kafashan et al. 2021) and report substantially better sample efficiency, lower bias, and greater robustness. They then apply the method to three datasets: (1) mouse head-direction cells (Ajabi et al.), (2) mouse V1 (Stringer et al.), and (3) macaque PFC during a saccade task (Bartolo et al.).

      This is a strong and timely contribution. The core idea is elegant, and the method appears to be more practical than existing alternatives in the finite-data regime that real experiments occupy. The three applications are well chosen, and each yields a non-trivial, biologically interpretable result. I am strongly supportive of the potential of this paper.

      That said, the paper makes several strong empirical claims - most notably that prior V1 estimates were substantial overestimates, and that PFC information-limiting noise is temporally redundant - and the central estimator rests on an independence assumption whose finite-N validity is only partially characterized. Before these claims can be considered well supported, I would like the authors to address the following:

      Major Points:

      (1) The method relies on a key independence assumption that may not always be satisfied in the regime of finite neurons and trials. The author's main idea is to decompose the residuals of two decoders as follows:

      X1 = delta + phi1<br /> X2 = delta + phi2

      The covariance is equal to the scale of information limiting noise, Var[delta], plus three terms:

      Cov[X1, X2] = Var[delta] + Cov[delta, phi1] + Cov[delta, phi2] + Cov[phi1, phi2].

      We can define phi1 as the part of X1 that is orthogonal to delta and likewise define phi2 as the part of X2 that is orthogonal to delta; thus, the cross terms evaluate to zero, and we are left with:

      Cov[X1, X2] = Var[delta] + Cov[phi1, phi2]

      Now the authors introduce an assumption that Cov[phi1, phi2] = 0. This leaves us with Cov[X1, X2] = Var[delta], but the question is: when is it justified to assume that Cov[phi1, phi2] = 0? For example, it is possible that

      phi1 = c(N) * z + e1<br /> phi2 = c(N) * z + e2

      where z is another shared noise dimension that is not information limiting and e1 and e2 are truly independent. Here, c(N) is a constant that goes to zero as the number of neurons used to train the decoder, N, goes to infinity. Thus, in the limit of having very large neural populations at hand for the analysis, the author's assumption of Cov[phi1, phi2] = 0 can be justified. If the authors agree with this analysis, it would be nice to (a) flesh it out and include it in the methods / supplementary notes, and (b) to analyze in simulation how good this approximation is in finite N regimes. I suspect that the assumption works in finite N regimes if noise is low-dimensional, but that if there are many additional dimensions of correlation (i.e. many z's above), you will need a very large number of neurons before Cov[phi1, phi2] approaches zero.

      Along these lines, another worthwhile analysis would be to report outcomes when the neural populations are sub-sampled further. Intuitively, it should fail once you subsample to only a handful of neurons, e.g. 3, but I'm curious where the breaking point is and whether the decline is graceful.

      (2) In point 1, I raised the question of how the method behaves with a finite number of neurons. Another worry is that there is a finite number of trials. In particular, if you train two decoders on the same trials, I would worry that non-information-limiting fluctuations in those trials would induce correlations in the decoders that then would show up as correlations on the held-out test set. A more conservative approach would be to split trials into three disjoint subsets: a training set for decoder A, a training set for decoder B, and a common test set used to compute Cov[X1, X2].

      As a concrete example, suppose that on the particular trials used for training, the animal happened to be more aroused when theta = 1 and less aroused when theta = 0, and that arousal added a fluctuation on top of the neural response. This arousal-related signal is not information-limiting - it would average away given enough trials - but because both decoders are fit to these same trials, each one adjusts its weights to partially discount the same spurious high-arousal/low-arousal trend. Their weights are now distorted in a correlated way, so when both are applied to the shared test set, their errors covary, and the method reads this shared-training artifact as information-limiting noise.

      I think this dynamic should be acknowledged in the text and clarified in more detail. Ideally, simulations could be done to estimate how many trials are needed to average out this sort of confound, and similar to the suggestion in point 1 above, I would be interested in seeing what happens when the authors sub-sample trials before running their analysis. Together with point 1, the feedback is that I'd like to see more about "how many neurons and how many trials" are needed in order to trust your results. Similarly, are there diagnostics or resampling methods (e.g. bootstrapping) that could be helpful for a practitioner to know if they have enough neurons/trials?

      (3) Unless I've fundamentally misunderstood something, there is an error on page 17 in the methods. There we find sigma2 = Var[delta] = ... = Cov[phi1, phi2], but I believe this is meant to be Cov[X1, X2]. Indeed, the method assumes that Cov[phi1, phi2] = 0, as discussed in point 1.

      Additionally, on page 4, the authors introduce the main quantity as Cov[\hat{theta}_1, \hat{theta}_2] instead of Cov[X1, X2]. However, if theta is changing from trial to trial, then these two quantities are not technically equal to each other, so it would be more accurate to write down the conditioning on theta. That is, assuming conditionally unbiased decoders, Cov[X1, X2] = Cov[\hat{theta}_1, \hat{theta}_2 | theta] for a fixed theta.

      More generally, I found it hard to wrap my head around the underlying math on my first read through the paper. The polarization identity, 1/4 * (Var(X1 + X2) - Var(X1 - X2)), seems like a very roundabout way to derive the method. This identity is very helpful for the deconvolution extension, but I would have thought that a simpler and more straightforward derivation would have just used the expansion, Cov[X1, X2] = Var[delta] + Cov[delta, phi1] + Cov[delta, phi2] + Cov[phi1, phi2], as I did in point 1. I suggest the authors revise the mathematical presentation for clarity.

      Minor Points

      (1) A very nice feature of the authors' method is that they make no parametric assumption on the distribution of noise. This is in contrast to Kanitscheider et al. [12]'s finite-sample bias correction using the inverse-Wishart distribution of $\hat\Sigma^{-1}$, which is derived under an assumption of multivariate Gaussianity. I think it is worth adding a sentence to highlight this feature of the model.

      (2) Statistical inference claims (across sessions and population sizes) are supported by reported s.d.'s but no formal tests or confidence-interval-based comparisons. Given that several claims are comparative (split-trial < naive; V1 < prior reports; PFC stable over windows), please add appropriate uncertainty quantification (e.g., bootstrap CIs over sessions) and, where a difference is claimed, a test or effect size.

  3. Jul 2026
    1. Comment on “Ecological constraints to mirror life”

      Deepa Agashe1, Damon J. Binder2, Vaughn S. Cooper3, Kevin M. Esvelt4, Richard E. Lenski5,6, David A. Relman7,8,9

      Authors are listed in alphabetical order. Affiliations: 1 National Centre for Biological Sciences, Tata Institute of Fundamental Research, Bengaluru, India; 2 Coefficient Giving, San Francisco, California, USA; 3 Department of Microbiology and Molecular Genetics, University of Pittsburgh, Pittsburgh, Pennsylvania, USA; 4 Media Laboratory, Massachusetts Institute of Technology, Cambridge, Massachusetts, USA; 5 Department of Microbiology, Genetics, and Immunology, Michigan State University, East Lansing, Michigan, USA; 6 Program in Ecology, Evolution, and Behavior, Michigan State University, East Lansing, Michigan, USA; 7 Department of Medicine, Stanford University School of Medicine, Stanford, California, USA; 8 Department of Microbiology and Immunology, Stanford University School of Medicine, Stanford, California, USA; 9 Infectious Diseases Section, Veterans Affairs Palo Alto Health Care System, Palo Alto, California, USA

      We welcome mathematical modeling of mirror bacterial invasion dynamics, and are glad to see some of our earlier comments taken into account in Version 2 of this preprint. Nonetheless, we continue to have significant concerns about certain points in this preprint. In particular, many of the models presented show that evasion of chirality-dependent sources of mortality could facilitate the invasion of mirror bacteria in diverse environments, consistent with prior work on the subject. Such sources of mortality are ubiquitous in densely populated marine and terrestrial ecosystems, as well as within multicellular hosts. Despite their own models demonstrating invasion is possible once such mortality is included, the text of the preprint repeatedly concludes that ecological dynamics “strongly limit” the ability of mirror bacteria to invade the global environment.

      For a mirror bacterium to invade an ecosystem, it must be able to reproduce faster than it dies. As outlined in the 2024 Science commentary “Confronting risks of mirror life”, mirror bacteria would be largely or wholly resistant to predation and chirality-dependent microbial antagonism, which are major sources of bacterial mortality in many environments. They are similarly expected to evade most immune recognition in multicellular hosts, and thus the downstream responses that are a primary cause of pathogen mortality. The key question is whether the advantage from reduced mortality outweighs the disadvantage from reduced nutrient access. The answer is likely to depend on the specific environment and mirror bacterium. Chapter 8 of the Technical Report on Mirror Bacteria (https://doi.org/10.25740/cv716pj4036), which we co-authored and which accompanies the Science commentary, considers this question in detail, and finds that invasion appears plausible in many settings with bacterial predators, including most biodiverse and human-relevant ecosystems.

      Many of the models highlighted in the main text of the preprint do not address this key issue. They represent mortality as chirality-independent (e.g., δ in the closed system, D in the chemostat), and in these cases the models indicate that mirror bacteria cannot invade. This result follows standard resource-based competition theory (Tilman 1982), in which, in the absence of predation, an invader unable to access a sufficient amount of limiting resource is excluded at equilibrium. But in the real world, the mortality rates of bacteria are not chirality-independent, and so these models do not capture the conditions underpinning the substantial concerns about mirror life.

      When the preprint’s models do include chirality-dependent mortality, for example via a predator targeting the native population (Figure 4), mirror bacterial invasion is shown to be possible, even likely. This result, again, follows standard ecological theory: predation on a dominant competitor can allow invasion by a less competitive but predation-resistant species (Levin, Stewart & Chao 1977; Tilman 1982; Thingstad 2000). The preprint downplays this critical result, however, as “... context-dependent: it applies primarily when the natural ecosystem is already degraded, or when the predator exerts unusually strong top-down control over the natural population despite the availability of resources that could otherwise support growth.” But in resource-rich environments like surface waters, biologically active soils, biofilms, and host tissues, microbial mortality is often dominated by phage lysis, protist grazing, microbial antagonism, immune clearance, and other chirality-dependent processes. (Carlson et al. 2022 is one relevant reference for the surface ocean; many more are provided in Chapter 8 of our Technical Report). Top-down control in these ecosystems is not an aberration, and the evasion of chirality-dependent mortality could allow mirror bacteria to invade a wide range of environments.

      The text of the preprint repeatedly neglects this essential point. For example, the abstract concludes that nutrient limitations and competitive exclusion “constrain [mirror bacterial] growth and persistence across a broad range of ecological conditions”, and the discussion states that “intrinsic nonlinearities associated with resource incompatibility and ecological competition function as an effective form of distributed containment”. Additionally, the Table I caption states that invasion “is highly unlikely under realistic conditions” and that “all models consistently indicate that mirror life faces strong ecological constraints”. Neither the abstract, introduction, results, nor discussion make it clear that this containment is a general result only when mortality is chirality-independent, even though many human-relevant or species-rich real-world environments are dominated by chirality-dependent mortality.

      In fact, the preprint’s updated Supplementary Material (SM) presents additional models showing that invasion is plausible in realistic cases. Part I of the SM models a mirror autotroph (e.g., a mirror Prochlorococcus or Synechococcus), and concludes that invasion is possible “provided [the mirror autotroph’s] reduction in mortality from escaping predators and phages outweighs any catalytic handicap” – which it likely would, as explained in Chapter 8 of our Technical Report. Part II of the SM extends the closed-ecosystem model to explicitly incorporate chirality-dependent mortality from microbial warfare or antibiotics, and it again shows that invasion is predicted for a wide range of parameters (SM Figure 1). Unfortunately, the main text of the preprint neglects to discuss these important results, providing only a one-sentence note that the SM contains two other relevant case studies.

      There are other important considerations that further weaken the nutrient-limitation hypothesis as a potential ecological containment for mirror bacteria, which are also not discussed in the preprint. For example, mirror heterotrophs could be engineered to catabolize natural-chirality sugars (e.g., via incorporation of the Paracoccus laeviglucosivorans pathway; Shimizu et al. 2012), whether for benign reasons like facilitating laboratory studies or possibly for nefarious ends. In any case, such engineering would substantially improve the growth rate and competitiveness of mirror bacteria, pushing the authors' model results deeper into the invasion regime (Figure 4). Mixotrophic and autotrophic mirror bacteria would enjoy still greater advantages. Further, even environments that cannot be stably colonized by mirror bacteria could still harbor significant populations through repeated re-introduction, for example from animal hosts. Both of these scenarios are discussed in Chapter 8 of the Technical Report, and they would expand the conditions under which invasion succeeds in this preprint’s own framework.

      The paper also draws on two other arguments that we think are less than compelling. First, the absence of a "shadow biosphere" is cited as evidence that alternative biochemical systems like mirror life could not persist within the extant biosphere. This is a weak inference, as the absence is more plausibly explained by there being no evolutionary pathway to mirror life from the present biosphere on relevant timescales. Second, the preprint cites evidence that biodiversity can act as a “firewall” to invaders. While we agree that biodiversity can affect the likelihood that an ecosystem is invaded, it is important to note that biological invasion still occurs frequently in the real world, including in biodiversity-rich ecosystems like those in the tropics (Chong et al. 2021). Biodiversity may well raise the bar for invasion in some cases, but its effects can demonstrably be outweighed by the advantages discussed earlier.

      Mathematical models can be useful in clarifying the conditions under which mirror bacterial invasion is possible, and the models presented in the preprint are a valuable contribution. However, it is important to interpret and present the results of these models as comprehensively and accurately as possible. We hope that the authors will consider further revising their article to clarify and emphasize how invasion risk depends crucially on the different types of microbial mortality (chirality-dependent and chirality-independent); to highlight that chirality-dependent mortality occurs across real-world environments; and to more accurately reflect what their models predict.

      References:

      Adamala, K. P., Agashe, D., Belkaid, Y., Bittencourt, D. M. D. C., Cai, Y., Chang, M. W., et al. (2024). Confronting risks of mirror life. Science, 386(6728), 1351-1353.

      Adamala, K. P., Agashe, D., Binder, D. J., Cai, Y., Cooper, V., Duncombe, R., Esvelt, K., et al. (2024). Technical report on mirror bacteria: Feasibility and risks. https://doi.org/10.25740/cv716pj4036

      Carlson, M. C., Ribalet, F., Maidanik, I., Durham, B. P., Hulata, Y., Ferrón, S., ... & Lindell, D. (2022). Viruses affect picocyanobacterial abundance and biogeography in the North Pacific Ocean. Nature microbiology, 7(4), 570-580.

      Chong, K. Y., Corlett, R. T., Nuñez, M. A., Chiu, J. H., Courchamp, F., Dawson, W., et al. (2021). Are terrestrial biological invasions different in the tropics? Annual Review of Ecology, Evolution, and Systematics, 52(1), 291-314.

      Levin, B. R., Stewart, F. M., & Chao, L. (1977). Resource-limited growth, competition, and predation: A model and experimental studies with bacteria and bacteriophage. The American Naturalist, 111(977), 3–24.

      Shimizu, T., Takaya, N., & Nakamura, A. (2012). An L-glucose catabolic pathway in Paracoccus species 43P. Journal of Biological Chemistry, 287(48), 40448-40456.

      Thingstad, T. F. (2000). Elements of a theory for the mechanisms controlling abundance, diversity, and biogeochemical role of lytic bacterial viruses in aquatic systems. Limnology and Oceanography, 45(6), 1320–1328.

      Tilman, D. (1982). Resource competition and community structure. Monographs in Population Biology, 17. Princeton University Press.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Response to the points raised by the reviewers.

      Once again, we would like to thank the reviewers for their comments. We have systematically addressed their concerns, as detailed below.

      Reviewer #1

      Evidence, reproducibility and clarity

      This study demonstrates that BICD2, previously known as an adaptor protein for dynein, is involved in regulating centriole engagement during mitosis. First, using different antibodies, it was shown that BICD2 localizes near the mother centriole, as observed by super-resolution microscopy. During G1 and S phases, BICD2 localizes slightly outside the Cep152 ring, while in G2 to mitosis, it localizes near the cartwheel component SAS-6. Moreover, analysis of deletion mutants revealed that BICD2 localizes to the centrosome in a CC domain-dependent manner at the C-terminal end. The localization pattern resembling a ring in the cytoplasm was also observed through the CC3 domain. Next, BICD2 knockout (KO) cells were generated to investigate centriole dynamics. In BICD2 KO cells, the distance between the mother and daughter centrioles was observed to increase from G2 to mitosis compared to controls. Along with this, early centriole disengagement and centriole amplification phenotypes were observed. The increased distance phenotype between centrioles was rescued in BICD2 wild-type (WT) and mutant forms lacking the CC1 domain at the N-terminus, suggesting that this function of BICD2 is independent of dynein. Additionally, BICD2 mutants mimicking phosphorylation at the C-terminus showed reduced centrosome localization and were unable to rescue the phenotypes seen in BICD2 KO cells.

      While the study clearly demonstrates BICD2's contribution to centriole engagement, the underlying mechanisms of how BICD2 is involved in centrosome localization and centriole engagement remain unclear. As it is anticipated that the function of BICD2 is independent of dynein, further exploration of this unknown mechanism would enhance the value of the paper. Below are the concerns that should be addressed, including new experiments.

      Main Points:

      1. __ Fig. 1-3: Regarding the localization of BICD2 to centrioles, during the G1-S phase, its localization appears to overlap with PCM. Experimental investigation should be performed to examine whether BICD2's centrosomal localization is influenced by knockdown of PCM components like PCNT, Cep192, or Cep152.__ We now show that BICD2 localization does not depend on pericentrin (Supplementary Figure S4B). We also show that the two proteins do not colocalize (Supplementary Figure S4A and S4D) and are functionally independent (Figure 5).

      We also show that BICD2 localization does depend on the torus protein CEP152 (Figure 8B). Importantly, our data indicate that BICD2 interacts with the N-terminal region of CEP152, suggesting that this interaction places BICD2 at the outer region of the torus (Figures 8C and 8D). We propose that this provides a mechanistic basis for BICD2 function in maintaining mother-daughter centriole engagement.

      __ Fig. 7: The experiments using BICD2 mutants suggest that the function of BICD2 here is independent of dynein. To further investigate whether BICD2's role in centriole engagement is independent of dynein, experiments should be conducted to examine the effect of dynein knockdown on BICD2 localization to the centrosome and centriole engagement.__

      Using Dynapyrazole-A, a fast-acting and potent dynein inhibitor, we demonstrate that acute inhibition of dynein motor activity affects neither the centrosomal localization of BICD2 during G2 and M phases (Supplementary Figure S3A) nor centriole engagement (Supplementary Figure S3B). Together with experiments using BICD2 mutants deficient in dynein interaction (Figure 7B), these data compellingly demonstrate that the recruitment and function of BICD2 at the centriole is dynein-independent.

      __ Fig. 4: The CC4 domain at the C-terminus of BICD2 is important for its centrosomal localization, but identifying the binder/recruiter responsible for BICD2's centrosome localization would be desirable.__

      We thank the reviewer for prompting us to investigate this further. We are delighted that we now identify the torus protein CEP152 as the BICD2 binder/recruiter at the centriole (Figure 8). As we discuss in the manuscript, our observation that the outward-facing N-terminus of CEP152 interacts with the C-terminal region of BICD2 provides a mechanistic basis for understanding BICD2's role in maintaining mother–daughter centriole engagement.

      __ Fig. 7: Rescue experiments using BICD2 mutants suggest that BICD2's functional domains are critical. Further experiments by creating mutants missing parts of CC2 or CC3 could identify functionally important domains of BICD2 by observing any loss-of-function phenotypes at the centrosome.__

      We fully appreciate the reviewer’s suggestion to examine the roles of the CC2 and CC3 regions. We believe that these domains, and particularly the unstructured loop within CC3, are important for both BICD2 localization, function and regulation at the centrosome. However, given that our current data already establish a clear mechanism for BICD2 centriolar recruitment via CEP152 and the CC4 region, we feel these additional structural studies fall outside the core scope of the present manuscript. We hope the reviewer agrees that the current evidence provides a robust foundation for our conclusions, and we look forward to addressing the roles of CC2 and CC3 in a dedicated future study.

      __ Fig. 6: Regarding the BICD2 KO cell phenotype, is there experimental evidence showing an increase in centriole number during mitosis? For instance, while no abnormality in centriole number may occur during G2, a trend of increase in mitosis should be experimentally demonstrated. Also, how should the slight differences in phenotypes between Ndelta4 and Ndelta5 BICD2 KO cells be interpreted?__

      We thank the reviewer for highlighting this point, but we would like to clarify that we do indeed observe a significant increase in centriole number during both mitosis and G2 phase across multiple cell lines in our BICD2 KO models and RNAi experiments (Figure 4E, RPE-1 KO cells, and 4G, U2OS cells, RNAi) and G2 (Figure 5C, both RPE-1 and U2OS, RNAi). As the main text was not explicit enough on this point, we have revised the manuscript to describe these observations more clearly.

      Regarding the phenotypic differences between BICD2 KO lines, we assign them to the clone-to-clone functional heterogeneity often seen in CRISPR/Cas9-generated cell lines. Importantly both clones show a consistent, statistically significant phenotype (e.g., impaired engagement and increased centriole numbers) compared to wild-type controls, confirming that the overall defect is robust and specific to BICD2 loss. We have added a clarifying note on this in the revised manuscript: “Figure 4F; we assign the differences between BICD2-/- cell lines to standard clone-to-clone phenotypic heterogeneity often seen in CRISPR/Cas9-generated cell lines.

      __ Fig. 8: Regarding the phosphorylation of BICD2 at the C-terminus: The phenotypes of mutants where these two phosphorylation sites are changed to alanine should be experimentally observed. It is expected that the removal of BICD2 from the centrosome during mitosis could be rescued. Additionally, the effect of PLK1 or CDK1 inhibitors on the removal of BICD2 from the centrosome should be investigated.__

      We agree with the reviewer that phosphonull mutants should be added to these experiments. As mentioned above we have decided to remove the preliminary data regarding BICD2 phosphorylation from the manuscript data to present a more comprehensive, dedicated study on BICD2 phosphorylation in the near future. In fact, we have already performed the suggested experiments, including the phosphonull mutants and kinase inhibitor treatments, and would be glad to share these additional results if the reviewers would find them helpful. Interestingly, our experiments show that BICD2 centrosomal amounts are not affected by PLK1 inhibition (using BI 2536); CDK1 inhibition (RO-3306), although not significatively changing the amount of BICD2 at the centrosome, slightly diminishes it. We currently favor a model in which BICD2 is predominantly regulated by CDK1, and we are actively defining the precise molecular mechanism governing this regulation.

      Minor Points:

      __ Fig. 1-3: During G1 and S phases, BICD2 localizes near the mother centriole, and from G2 onward, it colocalizes with SAS-6. How can this be explained?__

      We currently do not have a clear explanation for this transition, as our focus has been in understanding BICD2 recruitment to the centriole (and its role in centriole engagement). We view this as a very interesting question that could be studied together with BICD2 regulation through phosphorylation. Our current hypothesis is that most of BICD2 is removed through phosphorylation in late G2 and M, with a pool remaining at the mother-daughter interface, possibly protected by a yet to be understood mechanism. We have added a sentence in the discussion addressing this (“ A pool of protein could be protected and correspond to the observed remnant of BICD2 at the mother-daughter interface.“). This last pool, as we discuss in the manuscript, could be further phosphorylated at the M/G1 transition or cleaved by separase (although this last point is of course highly speculative).

      __ Fig. 4: The GFP-BICD2 488-820 fragment forms cytoplasmic rings, which is interesting. This domain contains the CC4 domain, so it can localize near the centriole, but why does it not form a perfect ring there? Also, which other centriole/centrosome markers were used for colocalization studies? Does knockdown of PCM1 affect BICD2's centrosomal localization?__

      We show in Figures 6G and 6H that BICD2 488-820 can form a ring around the centriole. Indeed this polypeptide contains the CC4 region, which our results indicate it will guide it to the centriole (through an interaction with CEP152). Once the available CEP152 is occupied with BICD2 we assume that BICD2 488-820 forms oligomers that assemble ring-like structures outside the centriole.

      Other centrosomal markers used are SAS-6.

      We now show that PCM1 knockdown does not affect BICD2's centrosomal localization (Supplementary Figure S4C). Although our results indicate that partial forms of BICD2 such as BICD2 488-820 can colocalize with PCM-1 (Figure 6D), full length endogenous BICD2 (or GST-BICD2) does not seem to colocalize with this protein and thus the centriole satellites (Supplementary Figure S4A). We note this discrepancy in the text: “The presence of BICD2 at the centriolar satellites has been suggested previously (Quarantotti et al, 2019); we ignore the reason why in the conditions used in this study only C-terminal fragments of BICD2 but not the full-length protein”. The relationship between satellites and BICD2 grants further studies. Our data suggests that centriolar localization of the protein may be regulated, possibly by its intramolecular structure, and that regulated binding of BICD2 to a yet to be identified partner may recruit the protein to satellites either for its transport to the centrosome or in order to perform a specific function at the satellites. We have added a sentence to the text to note this: “This suggests that BICD2 satellite localization is regulated (possibly via intramolecular autoinhibition) to mediate BICD2 transport or a distinct satellite-specific function of this protein.”.

      __ Fig. 4A: What are the aggregates observed in the cytoplasm under the GFP-BICD2 + ice condition? Also, does the 1-575 mutant fail to localize to the centrosome upon ice treatment?__

      We currently do not know the nature of the GFP-BICD2 full length aggregates observed upon microtubule depolymerization. We also observe GFP-BICD2 aggregates in cells that express high amounts of the polypeptide, leading us to hypothesize that it may be insoluble and the disappearance of microtubules may liberate it from motor complexes resulting in its aggregation -although of course more work would be needed to clarify this.

      We now show new data (Supplementary Figure S5), showing that the localization of not only GFP-BICD2 1-575 but also the C-terminal fragments 272-820 and 488-820 are not significantly affected by cold-induced microtubule depolymerization. These last results strongly suggest that BICD2 localization at the centriole is microtubule independent and are compatible with our data showing that BICD2 can directly interact with the centriolar protein CEP152.

      __ Can similar phenotypes be observed in other cell types when BICD2 is knocked down? This should be experimentally validated.__

      Our current manuscript now shows that similar phenotypes regarding centriole separation and amplification are observed upon BICD2 depletion in RPE-1 cells (non-transformed, p53-wildtype) and U2OS cells (transformed). These are shown in Figure 4 (RPE-1 knockout, U2OS RNAi knockdown) and Figure 5 (RPE-1 and U2OS RNAi knockdown).

      __ Are there previous studies suggesting that this function of BICD2 is evolutionarily conserved? This should be addressed.__

      To our knowledge there are no previous studies describing BICD2 function at the centrosome, excepting the recent article by Kuang et al., (Kuang W et al. 2025. BICD2 promotes ciliogenesis by facilitating CP110 removal from the mother centriole. EMBO reports 26:5567–5588. DOI: https://doi.org/10.1038/s44319-025-00597-0), that describes a role for BICD2 during ciliogenesis in non-cycling cells. As we mention in our discussion this new role may be related to the distal pool of protein that we observe using ExM, and we don’t think is related to the function of the proximal pool of BICD2 at the torus in cycling cells that we describe in our manuscript.

      Regarding functional conservation, BICD2 orthologs are widely distributed across metazoans (as reflected in OrthoDB, which lists ~5,000 ortholog genes across ~2,500 species). They share a remarkably conserved C-terminal domain that acts as a docking interface mediating subcellular targeting independently of dynein motor activity (i.e. through binding to Rab6, RanBP2 and, as shown here, CEP152). Cross-species analyses show that this C-terminal domain is preserved in most eukaryotic orthologs, including Drosophila melanogaster BICD (UniProt P16568) and Caenorhabditis elegans BICD-1 (UniProt V6CJ04). Interestingly, several predicted orthologous sequences in public databases retain high C-terminal similarity while completely lacking the N-terminal regions containing the CC1 box motif (residues 29–57 in human BICD2) required for dynein interaction (e.g., predicted isoforms in mouse or camels). Thus, dynein-independent scaffolding functions may represent an ancient, foundational role of the BICD protein family, or alternatively (and perhaps most probably, given that basal metazoans like sponges or Cnidaria do show a conserved N-terminus), these truncated forms may have evolved to fulfill distinct cellular roles operating independently of motor-adaptor activity. We have added a passage at the end of the discussion to reflect this.

      Significance

      In this paper, the identification of BICD2 as a novel factor regulating centriole engagement is of significant importance. However, the mechanisms through which BICD2 controls its localization to the centrosome and regulates centriole engagement remain largely undefined. Further exploration of these mechanisms would likely enhance the value of the paper.

      The findings are likely to be of great interest to researchers in the field of cell biology, particularly those focusing on centrosome biology.

      The above feedback comes from a researcher specializing in centrosome studies.

      __ __

      Reviewer #2

      Evidence, reproducibility and clarity

      Montez-Ruiz and colleagues explore the role of a dynein adaptor BICD2 in the engagement of mother and daughter centrioles. Cells need to maintain centriole engagement in interphase to prevent centriole reduplication and in early mitosis to prevent the formation of aberrant mitosis spindles. The authors demonstrate that BICD2 is a centriolar protein that surrounds the mother centriole adjacent to the daughter centriole. It is removed from centrosomes in mitosis, which, in turn, is responsible for centriole disengagement. Further, they suggest that in BICD2 knock-out G2 and early mitotic cells, centrioles disengage prematurely. By conducting rescue experiments, the authors conclude that BICD2 regions CC2, CC3, and CC4, which are dynein-independent, are essential for their function at the centrosome. Finally, they show that the phosphorylation of S817 and S819 of BICD1 controls its centrosome localization.

      Major comments:

      1. __ Based on F1 and SF1, BICD2 is reduced from centrosomes already in early G2. So, it is hard to square how removing a factor that is not present at the centrosomes at the time of disengagement would dysregulate disengagement. The study at this stage does not explain how BICD2 contributes to centriole engagement only in mitosis, while it does not affect centrioles in S.__ We now present new data obtained using expansion microscopy (ExM) that, together with our super-resolution observations, clarifies this point. As shown in the new Figure 3 and Figure EV2 (and supported by Figures EV3 and EV4), although the total amount of BICD2 at centrosomes is significantly reduced from G2 to M, a pool of BICD2 persists at the mother centriole until late mitosis. Importantly, this pool tends to localize close to the daughter centriole. We note this in the text (“BICD2 remained visible in both diplosomes, associated with the SAS-6 foci (Figure 3, Figures EV2-3). Around anaphase, BICD2 was not detectable in some diplosomes, while others retained some protein (again, close to the SAS-6 foci, which at this point were disappearing from the centrioles as the result of the disassembly of the cartwheel).”). Supported by our BICD2 depletion experiments, we propose that this centriolar pool enables BICD2 to contribute to engagement until late mitosis, when the remaining protein at the centrosome is ultimately removed. We highlight this model in the Discussion section (“In mitosis, when the protein progressively disappears from the centrosomes, BICD2 remains functionally relevant -likely via the small pool that persists at the mother–daughter centriole interface.”).

      Based on our new data demonstrating an interaction with CEP152, we propose that BICD2 forms an outer component of the centriolar torus. Centriole engagement is known to be maintained during S phase by the cartwheel (Huang F et al. 2022. Cartwheel disassembly regulated by CDK1-Cyclin B kinase allows human centriole disengagement and licensing. The Journal of Biological Chemistry 298:102658. DOI: https://doi.org/10.1016/j.jbc.2022.102658; Ito KK et al. 2025. Multimodal mechanisms of human centriole engagement and disengagement. The EMBO journal 44:1294–1321. DOI: https://doi.org/10.1038/s44318-024-00350-8), with the torus playing a role in cohesion later in the cell cycle. We note in our manuscript that this is consistent with our observations and supports a model in which BICD2 functions as part of the torus: “During S phase, mother–daughter centriole cohesion is maintained by the cartwheel (Huang et al, 2022; Ito et al, 2025) and, consistently, does not depend on BICD2.

      __ The interpretation that the longitudinal localization of the BICD2 signal coincides with SAS-6 and procentrioles requires further evidence. BICD2 seems largely localized to the other regions around the mother centriole, and in some examples, it does not colocalize with the site of the daughter centriole or SAS-6 (for instance: F2B second row; SF4B, second row; SF5, fourth row; SF6 upper row).__

      We have added an ExM characterization of BICD2 centrosomal localization in the revised manuscript (Figures 2 and 3), that we think further clarify this point, showing that BICD2 longitudinally coincides with the torus and the daughter centriole. This is supported by new additional superresolution images (Figure 2, Figure EV2).

      Note that ExM revealed an additional stable pool of BICD2 at the distal end of the centrioles that is not detected using standard methanol fixation combined with 3D-SIM. As discussed in the text this distal pool may reflect additional centriolar functions of BICD2.

      __ BICD2 is important for centrosome-nucleus tethering during centrosome separation in G2, and its global removal likely affects the dynamics of the spindle assembly. Is G2 and mitotic progression affected in knockouts? Do the knockout cells show issues with chromosome alignment? Such analyses are critically missing from the manuscript.__

      BICD2 knockout cells indeed show a slightly higher mitotic index than their wild type counterparts, and a higher frequency of lagging chromosomes in anaphase and telophase as well (new data, shown in Figure EV5C). We agree with the reviewer that these might result from the role of BICD2 tethering centrosomes to the nuclear envelope to facilitate their separation during the initial steps of spindle formation. We have added a sentence in the text noting this: “As expected from cells with supernumerary centrioles, BICD2-/- cells showed a slightly higher mitotic index and a higher frequency of lagging chromosomes in anaphase and telophase (Figure EV5C), although these mitotic defects might also be partially attributed to the role of BICD2 in centrosome separation (Splinter et al, 2010; Gallisà-Suñé et al, 2023).

      __ In general, SCLT experiments are ambiguous. Centriole disengagement spontaneously occurs during prolonged prometaphase induced by SCLT. Accordingly, F6D shows that many centriole pairs in the control sample are disengaged after 16h of SCLT treatment. Although the distance between centrioles in knockout cells is, on average, larger, without knowing how BICD2 perturbations affect the dynamics of the mitotic spindles and mitosis progression, SCLT experiments do not provide enough insight.__

      After 16h of SCLT treatment, the authors regularly measure centriole distances in mitosis smaller than 500 nm in all samples. This suggests that the used method (which also needs to be described) cannot reliably assess centriole engagement status. Centrioles can be disengaged but adjacent. The authors reference Shukla et al. 2015 to compare the centriole-to-centriole distances here with those from that publication. However, in Shukla 2015, centriole-to-centriole distances increase from S to M. But here, in F6, the control centriole distances in S, G2, and early M are almost identical and less than 500 nm. This discrepancy needs to be addressed.

      We thank the reviewer for these constructive comments. We appreciate the opportunity to further clarify our methodology and experimental rationale.

      Validity and necessity of STLC treatment (Figures 4D and 7)

      We fully agree with the reviewer that prolonged STLC treatment (16 hours) carries inherent limitations and should not serve as the sole experimental system for studying centriole engagement. As the reviewer notes, 16 hours of STLC treatment results in a baseline population of control cells displaying disengaged centrioles. This population likely represents cells that entered mitosis early during the treatment and remained arrested for the longest duration, or cells with inherently less robust engagement machinery.

      However, we would like to highlight two key observations that validate STLC as a useful comparative tool in our study:

      • The significative increase in the number of cells with higher intercentriolar distances that indicate disengagement in particular experimental conditions. We show that depletion of BICD2 consistently leads to a statistically significant increase in mean intercentriolar distances compared to controls under identical STLC conditions, indicating a distinct weakening of centriole engagement in a substantial number of cells (and thus suggesting that BICD2 is part of the engagement mechanism).

      • Validation in unarrested cells: Crucially, this effect is not an artifact of mitotic arrest. Unarrested, normally cycling mitotic cells also display significantly increased intercentriolar distances in the absence of BICD2 (Figure 4C).

      Following the initial characterization in Figure 4D, we restricted the use of STLC exclusively to experiments requiring cell transfection and recombinant protein expression (Figure 7). Human RPE-1 cells offer the key advantage of being an untransformed, p53-wild-type model. However, they also present technical challenges, including lower transfection efficiencies and sensitivity to experimental manipulation. Capturing a statistically robust sample of transfected, unarrested mitotic cells proved technically challenging. STLC treatment provided a necessary tool to enrich for mitotic cells while allowing clear observation of rescue effects.

      We have explicitly clarified this technical rationale in the manuscript text:

      "Although this treatment inherently increased mean intercentriolar distances, it nevertheless enabled clear observation of the effects of BICD2 ablation, while yielding a sufficient number of mitotic cells expressing the recombinant proteins."

      Assessment of centriole engagement

      We agree that centrioles can occasionally be disengaged while still remaining adjacent. To avoid oversimplifying the observed phenotypes, we chose to report raw intercentriolar distances rather than applying an arbitrary binary classification of "engaged" versus "disengaged." Furthermore, we do not rely solely on distance measurements to assess engagement status. We complemented these data by quantifying c-NAP1-positive centrioles in unarrested, cycling mitotic cells (Figure 4E, F). Because c-NAP1 loading marks centriole-to-centrosome conversion (and thus licensing), this functional readout independently confirms that BICD2 loss promotes premature centriole disengagement.

      Intercentriolar distances across the cell cycle and cell-type variation

      Regarding the comparison with Shukla et al. (2015), we note that their study was conducted in HeLa cells, whereas our primary model is RPE-1 (alongside U2OS cells). Variations in centriole engagement dynamics and distance kinetics can likely be attributed to intrinsic differences among these cell types:

      RPE-1 cells: baseline intercentriolar distances in S-phase control RPE-1 cells (0.4–0.5 µm, measured using centrin) match those reported for HeLa cells in S-phase by Shukla et al. However, in RPE-1 control cells, these distances remain relatively constant from S phase through early M phase (Figure 4).

      U2OS cells: U2OS cells exhibit a slight increase from 0.43±0.01 µm in G2 to 0.50±0.01 µm in M (Figure 5), illustrating that slight variations occur between cell lines.

      Other studies similarly report persistent baseline distances around 0.5 µm through early cell cycle stages. For example, Yaguchi et al. (Yaguchi K et al. 2018. Uncoordinated centrosome cycle underlies the instability of non-diploid somatic cells in mammals. The Journal of Cell Biology 217:2463–2483. DOI: https://doi.org/10.1083/jcb.201701151) observed intercentriolar distances close to 0.5 µm in diploid HAP1 cells throughout mitosis and into early G1 phase, with substantial disengagement (>0.8 µm) occurring only well after cytokinesis onset.

      To address this discrepancy, we have added the following sentence to the manuscript text: "Note that in wild-type S-phase RPE-1 cells, intercentriolar distances measured using centrin as a marker were similar to those described in S-phase HeLa cells (Shukla et al., 2015), namely 0.4–0.5 µm; however, in contrast to HeLa cells, these distances remained fairly constant from S to early M phase in RPE-1 cells." We have also updated the Materials and Methods section to provide a precise description of how intercentriolar distances were measured: “Intercentriolar distances were assessed as the distance between centrin foci of the same diplosome in maximum projections of z-stacks

      __ The authors suggest that BICD2's functions at the centrosome are independent of its dynein functions. They show that GFP-BICD2 1-820 DD rescues centriole engagement among several other mutants. However, it is still possible that the expression of the mutants affects some yet uncovered BICD2 function outside of centrosomes. At least, T821A and S823A should be mutated to Ala. From what I gathered, such mutant should remain associated with mitotic centrosomes. The authors should analyze whether mitotic progression remains unperturbed, and centriole engagement status should be analyzed without SCLT treatment in G2, M, and in ensuing G1.__

      We agree with the reviewer that phosphonull mutants should be added to these experiments. In fact, and as mentioned in the responses to Reviewer 1, we have already performed experiments with the phosphonull mutants, observing that they are more retained at centrosomes than the phosphomimetic counterparts. We would be happy to share these results with the reviewers upon request if helpful. Nevertheless, and as mentioned above, we have decided to remove the preliminary data regarding BICD2 phosphorylation from the manuscript data in order to present a separate and more comprehensive study on BICD2 phosphorylation in the near future.

      Significance

      The question explored is relevant to the centrosome field and beyond since the processes leading to premature centriole disengagement and amplification are not fully understood. The study provides some novel insights. However, at the current stage, the study is preliminary. Additional experiments would be needed to strengthen the conclusion that BICD2 directly regulates centriole disengagement.

      My expertise is in centriole and centrosome assembly and the mechanisms that regulate centrosome homeostasis in human cells.

      __ __

      __Reviewer #3 __

      Evidence, reproducibility and clarity (Required):

      Centrosome duplication is tightly control during cell cycle to prevent loss or amplification of centrosome numbers, which are detrimental for cell proliferation. In preparation for centriole duplication in S-phase, mother and daughter centrioles disengaged late mitosis, a process that functions as a licensing factor for duplication. While several mechanism have been proposed to be important for centriole disengagement, differences between systems and organisms exist, suggesting alternative pathways may play a role.

      In this manuscript, Montes-Ruiz and colleagues investigate the role of the dynein adaptor protein BICD2 during centriole disengagement. They found that BICD2 localises to the centrioles, with a peak in S-Phase. Super resolution microscopy suggests that BICD2 localises to the mother centrioles and is mostly absent in mitosis cells after anaphase, when centrioles are disengaging. KO of BICD2 in REP-1 cells does not some t have strong phenotypes, but the authors found that centriole separation is increased, suggesting a role in centriole cohesion. While there is limited mechanist insight about the regulation of BICD2 and its function at the centrosomes, the data presented suggests a role for BICD2 in centriole cohesion that is independent of dynein interaction. There are however several issues with data presentation, image analyses and data interpretation the authors could improve.

      Major comments

      - On page 5, the authors state that figure 1 and supplementary figure1 data strongly suggest that BICD2 associates with mother centrioles and not the PCM. This is not very clear from the images on these figures. In fact, PCM is often associated with mother centriole as well, thus I am not sure they can make these conclusions based on the data presented in these 2 figures. Also, the fact that PCM is more abundant in G2/M, when BICD2 is not, does not mean it does not localize to the PCM. Higher resolution of expansion will be needed.

      The data presented in supplementary figure 3 does not help the conclusion above as it seems form the images that there is co-localization between BICD2 and pericentrin. It is impossible to conclude also that there is co-localization with the satellite marker PCM-1. In fact, they seem to have no overlap from the images provided. Higher resolution of expansion will be needed.

      Following the reviewer’s suggestion we embarked in a full characterization of BICD2 localization using expansion microscopy (ExM). We think that our new data further clarifies this together with new superresolution data.

      Additionally we now have a figure (Supplementary Figure S4) addressing the relation between pericentrin and the localization of BICD2. We show that pericentrin downregulation does not affect centrosomal BICD2 levels. And that both proteins do not colocalize as observed using 3D-SIM.

      Also regarding pericentrin, the revised version of the manuscript now includes a figure that functionally compares the results of its depletion to those of BICD2 (Figure 5).

      We also provide data showing that BICD2 localization does not significatively change upon PCM-1 depletion (Supplementary Figure S4C). As we mention in the text, previous reports have suggested that BICD2 is indeed in the satellites (Quarantotti V et al. 2019. Centriolar satellites are acentriolar assemblies of centrosomal proteins. The EMBO Journal e101082. DOI: https://doi.org/10.15252/embj.2018101082) . But we only observed clear colocalization of PCM-1 with C-terminal fragments of BICD2. Thus, while GFP-BICD2 488–820 strongly colocalizes with satellites, endogenous BICD2 and full-length GFP-BICD2 do not (Figure 6D). We ignore the reason for this, but the data suggests that satellite localization is regulated (possibly via intramolecular autoinhibition) to mediate BICD2 transport or a distinct satellite-specific function of this protein. To address this we have added a sentence to the text that now reads: “The presence of BICD2 at the centriolar satellites has been suggested previously (Quarantotti et al, 2019); we ignore the reason why in the conditions used in this study only C-terminal fragments of BICD2 (but not the full-length protein, see Supplementary Figure S4) colocalize with satellites. This suggests that BICD2 satellite localization is regulated (possibly via intramolecular autoinhibition) to mediate BICD2 transport or a distinct satellite-specific function of this protein.”.

      - In figure 2, to confirm localization to the mother centrioles, could the authors use a mother centriole marker? Such as a distal appendage protein of ninein? CEP152 localizes to both centrioles in the images provided.

      I was surprised that BICD2 localizes to both distal appendages and linker? These are not close to each other. Can the authors comment on this? In supplementary figure 4C orthogonal view it seems like BICD2 is in between distal appendages and linker?

      We believe that the new ExM data (Figures 2 and 3) directly address the reviewer's concerns.

      Regarding the original supplementary figure S4C, indeed in the orthogonal projections of 3D-SIM images the signal corresponding to BICD2 was observed between distal appendages and linker and was quite broadly distributed. We recognize that this could lead to confusion. We have now removed part of this figure (original Figures 4B and 4C) from the manuscript, as we think that the data is made redundant with our new ExM data. Our new data, with a much higher resolution shows that BICD2 localization corresponds to that of the proximal torus (see new Figures 2D and 2E, and Figure 3). Note that in our new ExM images we use a daughter centriole marker (SAS-6) that (in addition to CEP152) we think helps confirm that BICD2 localizes around the mother centriole.

      -The IF data suggests that BICD2 localization to the centrosome is dynamically regulated during cell cycle. Did the authors consider that this protein could be degraded? Is it a matter of recruitment or total protein levels?

      We agree that protein degradation has to be considered when analyzing cell cycle-dependent localization. However, our data suggest that the dynamic behaviour of BICD2 at the centrosomes does not reflect changes in its total protein amount. We have previously shown that total BICD2 levels are not reduced in mitosis, as assessed by western blot (Gallisà-Suñé N et al. 2023. BICD2 phosphorylation regulates dynein function and centrosome separation in G2 and M. Nature Communications 14:2434. DOI: https://doi.org/10.1038/s41467-023-38116-1). To make this clear in the current manuscript, we have additionally added Figure EV1B depicting BICD2 levels in S, G2 and M phase, and the following note to the text : “Total levels of BICD2 remained constant during the different phases of the cell cycle (Figure EV1B and (Gallisà-Suñé et al, 2023))“.

      - The authors propose that the dynamic localization of BICD2 is associated with licensing. However, it is rather surprising that the phenotype of centriole separation they describe is only observe in mitosis when BICD2 in knockdown and not in S-phase when the levels of BICD2 are higher? If the role of BICD2 is to prevent premature centiole disengagement, shouldn't that be observed in S-phase as well? Why only in mitosis when in control cells BICD2 levels are already very low?

      Recent data supports the notion that in S phase centriole cohesion is maintained by the cartwheel (Huang F et al. 2022. Cartwheel disassembly regulated by CDK1-Cyclin B kinase allows human centriole disengagement and licensing. The Journal of Biological Chemistry 298:102658. DOI: https://doi.org/10.1016/j.jbc.2022.102658; Ito et al. 2025. Multimodal mechanisms of human centriole engagement and disengagement. The EMBO journal 44:1294–1321. DOI: https://doi.org/10.1038/s44318-024-00350-8). Our data, including the new results showing that BICD2 interacts with CEP152, suggests that BICD2 is a dynamic part of the mother centriole torus, a structure that does not seem to be implicated in maintaining cohesion in S. We now note this in the manuscript’s text: “During S phase, mother-daughter centriole cohesion is maintained by the cartwheel (Huang et al, 2022; Ito et al, 2025), and, consistently, does not depend on BICD2.”. Of note, after Ito etl al. BICD2 (and the torus) may have a role in late S if the cartwheel is compromised, something that could be tested in future studies by downregulating cartwheel components and BICD2 simultaneously.

      - The images of C-Nap1 localization in figure 6E are not very convincing to illustrate the pint the authors are making in the main text (additional C-Nap1 foci are visible in the KO cells)

      We would like to note that visualizing C-NAP1 in mitosis is technically challenging, as a significant pool of the protein is displaced from the centrioles after phosphorylation in G2. However, the protein has been widely used as a marker of centriole disengagement (e.g. in the seminal Tsou M-FB et al. 2006. Mechanism limiting centrosome duplication to once per cell cycle. Nature 442:947–951. DOI: https://doi.org/10.1038/nature04985). We therefore consider it a valuable tool to support our conclusions regarding centriole engagement. Regarding extra c-NAP1 foci in BICD2 knockout cells, these may reflect additional centrioles that appear in these cells, as a result of abnormal disengagement and early licensing. To have this into account our data quantifies both c-NAP-1 positive centrioles (increased in KO cells, Figure 4E) and number of c-NAP-1 positive centrioles /total centriole number (with an increase in the abnormal >2:4 configuration in BICD2 KO cells, Figure 4F).

      - The authors propose that PLK1 and CDK1 phosphorylation sites regulate the association of BICD2 with the centrioles. Could this be tested with a PLK1 inhibitor?

      As noted above, we have removed the phosphorylation data from the manuscript, as we aim to report these findings in a dedicated upcoming study. Nevertheless, to address the reviewer's query, we now consider BICD2 to be predominantly regulated by CDK1, supported by data using BI 2536 showing that BICD2 centrosomal levels are unaffected by PLK1 inhibition. In contrast, CDK1 inhibition slightly reduces these levels, though this effect does not reach statistical significance under the tested conditions. We would be glad to share these additional results with the reviewers upon request.

      - On page 13, the authors state that their results do not agree with previous literature showing that pericentrin cleavage can result in disengagement. However, it was unclear from this manuscript what is the evidence to demonstrate that this is the case? The data presented in figure 5B for example only demonstrates that pericentrin levels do not change in the absence of BICD2 in what looks like S-phase cells. Did the authors look at pericentrin levels when they observe centriole disengagement in the ko cells? In G2 or early M-phase?

      We recognize that pericentrin is widely considered a crucial factor in centriole engagement, and have added new data in the manuscript studying the relationship between it and BICD2 (the partially new Supplementary Figure S4), and their relative importances for engagement both in G2 and M (the new Figure 5). Our data suggests that both proteins act independently in a partially redundant manner, BICD2 as part of the torus (key for engagement in G2) and pericentrin of the PCM (more important in M).

      We have updated the Discussion to present this and our view on pericentrin importance for engagement more clearly, specially our concerns that its importance may have been overestimated. Specifically we write that “BICD2 depletion reduces its centrosomal levels to a degree that mirrors those naturally observed during late M and early G1 in unperturbed cells. In contrast, experimental depletion of pericentrin reduces its levels far below physiological baselines across any phase of the cell cycle. This severe reduction produces marked centriole separation in mitosis that is likely amplified by spindle-derived forces. Consequently, the individual contribution of pericentrin to regulating physiological centriole cohesion may be somewhat overestimated under standard experimental knockdowns, and this regulation may rely more heavily on torus components, such as BICD2, than previously appreciated.

      Regarding the phases of the cell cycle in which we quantify pericentrin levels in the original Figure 5B (now Figure EV5B), we realize that the figure could lead to confusion as it was, as they were measured in M (when its amount is maximal, as specified in the figure legend) but the figure did not show examples in this cell cycle phase. We have added new examples of mitotic cells to the figure, and modified the figure labels and wording of the figure legend to clarify this.

      Minor comments

      - A more general reference (review) missing in the second paragraph of the introduction that describe the centrosomes.

      We have added a recent general reference when introducing centrioles (Gönczy P. 2025. Critical constituents and assembly principles of centriole biogenesis in human cells. Nature Reviews Molecular Cell Biology 1–18. DOI: https://doi.org/10.1038/s41580-025-00921-5). Later in the paragraph, when centriole duplication is introduced, we now use this reference plus the also recent Fernandes-Mariano C et al. 2025. Centrosome biogenesis and maintenance in homeostasis and disease. Current Opinion in Cell Biology 94:102485. DOI: https://doi.org/10.1016/j.ceb.2025.102485.

      - Some figures are not well organized, difficult to see which panel they correspond to? The authors could consider labelling panels better to make this clear. For example, figure 4 and 6 could benefit from additional panel labels.

      We have added additional panel labels to Figure 4 (now Figure 6) and Figure 6 (now Figure 4), that we have also slightly reorganized with the aim of making it clearer).

      - On page 7, what the authors mean by: "... we ignore the reason why in the conditions used in this study only C-terminal fragments of BICD2 but not the fill-length protein co-localize with these pericentriolar structures"?

      By "pericentriolar structures” we were referring to the centriolar satellites. We realize that that was not clear and updated the wording of the sentence that now reads “we ignore the reason why in the conditions used in this study only C-terminal fragments of BICD2 (but not the full-length protein, see Supplementary Figure S4) colocalize with satellites.”. We subsequently propose a possible a possible explanation for this: "This suggests that BICD2 satellite localization is regulated (possibly via intramolecular autoinhibition) to mediate BICD2 transport or a distinct satellite-specific function of this protein.".

      Reviewer #3 (Significance (Required)):

      In general this work has limited mechanistic insight and BICD2 localization to the centrosomes was known. However, the authors do go into more detail description of the centriole localization of BICD2 . In addition, their established KO cell lines provide some insights into the role of BICD2 in centriole disengagement, which is of interest to the field. But the limited scope of the conclusions does not advance the field significantly as it is.

      this work will interest a specialized audience.

    1. document-to-knowledge graph system

      That was the round trip from Google Docs creted using TrailMarks Mrk-in intentional notation back in the late wikiNizer Days ~ 2014

      using MindGraph that handled Proporisional Trail very well, but it was clear that what is required as a sweet spot was more like what Vannevar Bush's wrote about in "As We May Think" the new professions of TrailBlazers would be able to make sense of associative trails as they were to organie their learnings in form taat was ready to share, and perhaps co=laborate.

      That is what IndyWiki Brings to any web page on the Web

    1. Author response:

      The following is the authors’ response to the original reviews.

      In the revised manuscript, we have clarified several points that were raised by the reviewers. First, we now state more explicitly that the presence of intact env-containing Ty3/gypsy retrotransposons does not by itself demonstrate their mechanism of transmission, tissue specificity, infectivity, or current activity. We have therefore revised the wording throughout the manuscript to distinguish intact element structure and multicopy genomic expansion from experimentally demonstrated activity.

      Second, we performed targeted host-taxonomy concordance analyses on selected clades of the POL RT tree. These analyses do not exclude local horizontal transfer, particularly between closely related hosts, but they show that horizontal transfer alone is insufficient to explain the broader host-taxonomic structure observed across the dataset.

      Third, we incorporated representative viral and retroelement-associated fusogen proteins into our F-type ENV phylogenetic analysis and HSV/gB-type ENV structural comparison. These additions place the ENV proteins associated with Ty3/gypsy elements in a broader evolutionary context and strengthen the conclusion that these ENV associations are deeply diverged rather than recent derivatives of a single sampled viral lineage.

      Fourth, we added two each of entirely new Supplementary figures (S5 and S7) and Tables (S2 and S3) and substantially modified now Supplementary figure S8. Other figures have also been modified only to increase readability. The four tables from the original manuscript have not been modified although their numbering has changed.

      We believe that the revised manuscript is substantially improved in clarity, terminology and interpretive precision, while retaining the central conclusion that the association between env-like genes and Ty3/gypsy retrotransposons is ancient in metazoan evolution. Sincerely,

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript provides a comprehensive systematic analysis of envelope-containing Ty3/gypsy retrotransposons (errantiviruses) across metazoan genomes, including both invertebrates and ancient animal lineages. Using iterative tBLASTn mining of over 1,900 genomes, the authors catalog 1,512 intact retrotransposons with uninterrupted gag, pol, and env open reading frames. They show that these elements are widespread present in most metazoan phyla, including cnidarians, ctenophores, and tunicates-with active proliferation indicated by their multicopy status. Phylogenetic analyses distinguish "ancient" and "insect" errantivirus clades, while structural characterization (including AlphaFold2 modeling) reveals two major env types: paramyxovirus F-like and herpesvirus gB-like proteins. Although bot envelope types were identified in previous analyses two decades ago, the evolutionary provenance of these envelope genes was almost rudimentary and anecdotal (I can say this because I authored one of these studies). The results in the present study support an ancient origin for env acquisition in metazoan Ty3/gypsy elements, with subsequent vertical inheritance and limited recombination between env and pol domains. The paper also proposes an expanded definition of 'errantivirus' for env-carrying Ty3/gypsy elements outside Drosophila.

      Strengths:

      (1) Comprehensive Genomic Survey:

      The breadth of the genome search across non-model metazoan phyla yields an impressive dataset covering evolutionary breadth, with clear documentation of search iterations and validation criteria for intact elements.

      (2) Robust Phylogenetic Inference:

      The use of maximum likelihood trees on both pol and env domains, with thorough congruence analysis, convincingly separates ancient from lineage-specific elements and demonstrates co-evolution of env and pol within clades.

      (3) Structural Insights:

      AlphaFold2-based predictions provide high-confidence structural evidence that both env types have retained fusion-competent architectures, supporting the hypothesis of preserved functional potential.

      (4) Novelty and Scope:

      The study challenges previous assumptions of insect-centric or recent env acquisition and makes a compelling case for a Pre-Cambrian origin, significantly advancing our understanding of animal retroelement diversity and evolution. THIS IS A MAJOR ADVANCE.

      (5) Data Transparency:

      I appreciate that all data, code, and predicted structures are made openly available, facilitating reproducibility and future comparative analyses.

      Major Weaknesses

      (1) Functional Evidence Gaps:

      The work rests largely on sequence and structure prediction. No direct expression or experimental validation of envelope gene function or infectivity outside Drosophila is attempted, which would be valuable to corroborate the inferred roles of these glycoproteins in non-insect lineages. At least for some of these species, there are RNA-seq datasets that could be leveraged.

      We added a sentence in the discussion, subsection “The survival mechanism of errantiviruses in the genome”, citing our recent work now published (PMID: 41922845), explaining that the defence mechanism against errantiviruses appears to be conserved in insects beyond Drosophila, indirectly suggesting that their biology dependent on the presence of env—may be more universal.

      (2) Horizontal Transfer vs. Loss Hypotheses:

      The discussion argues primarily for vertical inheritance, but the somewhat sporadic phylogenetic distributions and long-branch effects suggest that loss and possibly rare horizontal events may contribute more than acknowledged. Explicit quantitative tests for horizontal transfer, or reconciliation analyses, would strengthen this conclusion. It's also worth pointing out that, unlike retrotransposons that can be found in genomes, any potential related viral envelopes must, by definition, have a spottier distribution due to sampling. I don't think this challenges any of the conclusions, but it must be acknowledged as something that could affect the strength of this conclusion

      We have added a targeted host-taxonomy concordance analysis for two well-sampled POL extended RT/connection subclades: an Annelida-associated clade from tree position A6 and a Lepidoptera-associated clade from tree position I1 (new Fig S5). Rather than attempting to infer exact numbers of duplication, loss and horizontal transfer events, which is difficult across highly expanded and unevenly sampled transposon families, we tested whether host-taxonomic labels were more clustered on the observed POL extended RT/connection topology than expected by chance. In the Annelida clade, highly supported small subclades showed strong host-family and host species concordance under host-label permutation tests. The Lepidoptera clade showed a more mixed pattern, but still contained several highly supported subclades enriched for related host groups at the superfamily or broader taxonomic level. These results do not exclude rare horizontal transfer, particularly between closely related hosts, but support the conclusion that the observed POL extended RT/connection trees retain significant host-taxonomic structure and are not consistent with frequent broad horizontal transfer between distantly related animal groups. We have added a paragraph in the Results section “Multiple intact elements of env-carrying Ty3/gypsy retrotransposons are found widespread across metazoan species” describing these observations, and also revised the Discussion to more explicitly acknowledge the possibilities of lineage-specific loss and the horizontal transfer.

      (3) Limited Taxon Sampling for Certain Phyla:

      Despite the impressive breadth, some ancient lineages (e.g., Porifera, Echinodermata) are negative, but the manuscript does not fully explore whether this reflects real biological absence, assembly quality, or insufficient sampling. A more systematic treatment of negative findings would clarify claims of ubiquity. However, I also believe this falls beyond the scope of this study.

      In the revised manuscript, we have added a targeted analysis of two representative genomes from each phylum. Although we did not detect full-length GAG-POL-ENV elements in these genomes, we recovered multiple full-length, multicopy GAG-POL Ty3/gypsy elements from all four genomes, many of which were flanked by predicted LTR sequences and associated with putative tRNA primer-binding sites. This suggests that the apparent absence of env-carrying elements in these representative Porifera and Echinodermata genomes is unlikely to be due simply to poor assembly quality or a general inability to recover intact Ty3/gypsy-like retrotransposons. We have added these data as the new Supplementary table S3 and revised the Results section “Multiple intact elements of env-carrying Ty3/gypsy retrotransposons are found widespread across metazoan species” to clarify that absence in these phyla may reflect true biological absence, lineage-specific loss, or incomplete taxon sampling.

      (4) Mechanistic Ambiguity:

      The proposed model that env-containing elements exploit ovarian somatic niches is plausible but extrapolated from Drosophila data; for most taxa, actual tissue specificity, lifecycle, or host interaction mechanisms remain speculative and, to me, a bit unreasonable.

      We stressed in the Discussion section “The survival mechanism of errantiviruses in the genome” that the mere presence of env gene does not imply the mechanism of transmission of retrotransposons.

      Minor Weaknesses:

      (1) Terminology and Nomenclature:

      The paper introduces and then generalizes the term "errantivirus" to non-insect elements. While this is logical, it may confuse readers familiar with the established, Drosophila-centric definition if not more explicitly clarified throughout. I also worry about changes being made without any input from the ICTV nomenclature committee, which just went through a thorough reclassification. Nevertheless, change is expected, and calling them all errantiviruses is entirely reasonable.

      We have revised the Results section and discussion where we introduced the term "errantivirus" to clarify that we use "errantivirus" operationally to refer to env-containing Ty3/gypsy retrotransposons identified in this study, rather than as a formal taxonomic proposal. We also now state explicitly that bona fide infectivity and amplification through the Drosophila-like ovarian somatic-cell route have not been experimentally established for most non-Drosophila elements. Our use of the term is therefore intended to distinguish env-containing Ty3/gypsy elements from related non-env containing Ty3/gypsy retrotransposons, while acknowledging that their biology outside Drosophila remains to be determined.

      (2) Figures and Supplementary Data Navigation:

      Some key phylogenies and domain alignments are found only in supplementary figures, occasionally hindering readability for non-expert audiences. Selected main-text inclusion of representative trees would benefit accessibility.

      We agree that clearer navigation between the main text and supplementary figures would improve readability. Although we considered moving selected supplementary phylogenies and alignments into the main figures, the main figures are already data-dense and are intended to provide representative summaries across many host groups and ENV types. We therefore retained the detailed trees and alignments as supplementary figures, where they can be shown at readable scale, but revised the manuscript to improve navigation. Specifically, we added signposting sentences in the Results where supplementary figures are mentioned, expanded the relevant figure legends, and clarified how each supplementary tree or alignment supports the corresponding main-text conclusion.

      (3) ORF Integrity Thresholds:

      The cutoff choices for defining "intact" elements (e.g., numbers/placement of stop codons, length ranges) are reasonable but only lightly justified. More rationale or sensitivity analysis would improve confidence in the inclusion criteria. For example, how did changing these criteria change the number of intact elements?

      We agree with the reviewer that the rationale for the ORF integrity thresholds should be stated more clearly. We have revised the Methods section "Identification of intact genomic copies of Ty3/gypsy errantiviruses" to clarify that the initial length, gap and stop-codon thresholds were deliberately permissive screening criteria, designed to avoid excluding divergent or non-canonical elements at the discovery stage. These initial filters were not used alone to define the final “intact” set. Candidate elements were subsequently subjected to multiple additional curation steps, including confirmation of Ty3/gypsy POL identity, recovery of full-length RT and Integrase domains within continuous ORFs, HHpred-based domain annotation of GAG, POL and ENV, and removal of elements with large domain truncations. Thus, the final set of intact elements is substantially more refined than would be implied by the initial stop codon or length thresholds alone.

      A full sensitivity analysis varying each threshold across the entire iterative discovery and manual-curation pipeline would be difficult to interpret, because changing early permissive filters would alter the candidate pool that then undergoes downstream structural and phylogenetic validation. Instead, we have clarified in the Methods that the early thresholds were intended as inclusive prefilters, whereas final inclusion required intact domain architecture and phylogenetic/domain support.

      (4) Minor Typos/Formatting:

      The paper contains sporadic typographical errors and formatting glitches (e.g., misaligned figure labels, unrendered symbols) that should be addressed.

      We now fixed these issues in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The authors first surveyed metazoan genomes to identify homologs of Drosophila errantiviruses and classified them into two groups, "insect" and "ancient" elements, supporting the hypothesis of an early evolutionary origin for these retrotransposons. They subsequently identified two distinct types of envelope proteins, one resembling the glycoprotein F of paramyxoviruses and the other akin to the glycoprotein B of herpesviruses. Despite differences in their primary amino acid sequences, these proteins display notable structural similarity in their predicted domain architectures. The congruence between the phylogenies of the envelope and pol genes further supports the ancient origin of the envelope genes, challenging earlier hypotheses that proposed recent recombination events with baculoviruses. Additional analysis of the Pol "bridge region" corroborated the divergence among these elements, consistent with a pattern of limited cross-species recombination. Finally, by comparing these elements with non-envelope-containing Gypsy retrotransposons, the authors concluded that errantiviruses originated from multiple elements independently.

      Strengths:

      The conclusions of this study are based on a comprehensive collection of errantiviruses identified across a wide range of metazoan genomes. These findings are further supported by multiple lines of evidence, including phylogenetic congruence and the diverse evolutionary origins of envelope genes. AlphaFold2-assisted protein domain structure analyses also provided key insights into the characterization of these elements. Together, these results present a compelling case that errantiviruses arose independently through multiple evolutionary events, extending well beyond previous hypotheses.

      Weaknesses:

      It would be beneficial to emphasize in the Abstract the potential impact of this work by more clearly articulating the current knowledge gap in the field. While the second paragraph of the Introduction briefly touches on this point, highlighting the broader significance in the Abstract would better capture readers' interest. Additionally, some methodological choices would benefit from clearer justification and explanation. For instance, in Figure 6, the selection of the bridge region/RNase H domain is not explicitly explained, leaving the rationale for its choice unclear. As a minor point, some figure labels and texts are too small and difficult to read, and improving their legibility would enhance overall clarity.

      We have revised the Abstract to more clearly state the knowledge gap addressed by this study: although env-containing Ty3/gypsy elements were known from Drosophila and sporadically reported in other animals, whether their association with env-like fusogen genes reflected recent, lineage-specific acquisitions or a much deeper evolutionary relationship remained unclear. We now highlight this broader significance in the Abstract and frame our results as evidence that env-containing Ty3/gypsy elements represent deeply diverged, genome-resident retroelements rather than a recent insect-specific phenomenon.

      We have also revised the Results, Methods and Figure 6 legend to explain why the

      RNase H-containing bridge region was analysed. Specifically, we now distinguish the Pol extended RT/connection region used for phylogenetic analysis from the RNase H-containing bridge region analysed structurally in Figure 6. We define the bridge region as the canonical RNase H domain together with the C-terminal region between RNase H and Integrase, and explain that this region was selected because RNase H-related and adjacent RNase H-like domains vary among LTR retroelement lineages. The bridge region architecture therefore provides an independent structural feature for comparing the “insect errantivirus” and “ancient errantivirus” groups.

      Finally, we have revised the figures and figure legends to improve readability. In particular, we enlarged labels where possible, clarified figure annotations, corrected cross-references between main and supplementary figures, and added signposting sentences in the Results so that readers can more easily connect the main conclusions to the supporting supplementary trees and alignments.

      Reviewer #3 (Public review):

      Summary and Significance:

      In this work, Cary and Hayashi address the important question of when, in evolution, certain mobile genetic elements (Ty3/gypsy-like non-LTR retrotransposons) associated with certain membrane fusion proteins (viral glycoprotein F or B-like proteins), which could allow these mobile genetic elements to be transferred between individual cells of a given host. It is debated in the literature whether the acquisition of membrane fusion proteins by non-LTR retrotransposons is a rather recent phenomenon that separately occurred in the ancestors of certain host species or whether the association with membrane fusion proteins is a much more ancient one, pre-dating the Cambrian explosion. Obviously, this question also touches upon the origin of the retroviruses, which can spread between individuals of a given host but seem restricted to vertebrates. Based on convincing data, Cary and Hayashi argue that an ancient association of non-LTR retrotransposons with membrane fusion proteins is most probable.

      Strengths:

      The authors take the smart approach to systematically retrieve apparently complete, intact, and recently functional Ty3/gypsy-like non-LTR retrotransposons that, next to their characteristic gag and pol genes, additionally carry sequences that are homologous to viral glycoprotein F (env-F) or viral glycoprotein B (env-B). They then construct and compare phylogenetic trees of the host species and individual encoded proteins and protein domains, where 3D-structure calculations and other features explain and corroborate the clustering within the phylogenetic trees. Congruence of phylogenetic trees and correlation of structural features is then taken as evidence for an infrequent recombination and a long-term co-evolution of the reverse transcriptase (encoded by the pol gene) and its respective putative membrane fusion gene (encoded by env-F or env-B). Importantly, the env-F and env-B containing retrotransposons do not form a monophyletic group among the Ty3/gypsy-like non-LTR retrotransposons, but are scattered throughout, supporting the idea of an originally ancient association followed by a random loss of env-F/env-B in individual branches of the tree (and rather rare re-associations via more recent recombinations).

      Overall, this is valuable, stimulating, and important work of general and fundamental interest, but still also somewhat incompletely explored, imprecisely explained, and insufficiently put into context for a more general audience.

      Weaknesses:

      Some points that might be considered and clarified:

      (1) Imprecise explanations, terms, and definitions:

      It might help to add a 'definitions box' or similar to precisely explain how the authors decided to use certain terms in this manuscript, and then use these terms consistently and with precision.

      (a) In particular, these are terms such as 'vertebrate retrovirus' vs 'retrovirus' vs 'endogenized retrovirus' vs 'endogenous retrovirus' vs 'non-LTR retrotransposon' and 'Ty3/gypsi-like retrotransposon' vs 'Ty3/gypsy retrotransposon' vs 'errantivirus'.

      We agree with the reviewer. We inserted a paragraph at the end of the first Results section, explaining how we define endogenous retroviruses (ERVs), Ty3/gypsy retrotransposons and errantiviruses.

      (b) The comment also applies to the term 'env' used for both 'env-F' and 'env-B', where often it remains unclear which of the two protein types the authors refer to. This is confusing, particularly in the methods, where the search for the respective homologs is described.

      We revised the manuscript and now used F-type env/ENV and HSV/gB-type env/ENV throughout the text. We also modified the method section where we explained the tBlastn search to clarify which ENV proteins were used initially for the search and how we classified them in later analyses.

      (c) Other examples are the use of the entire pol gene vs. pol-RT for the definition of the Ty3/gypsy clade and for the generation of phylogenetic trees (Methods and Figure S1), and the names for various portions of pol that appear without prior definition or explanation (e.g., 'pro' in Figure 1A, 'bridge' in Figure S1C, 'the chromodomain' in the text and Figure 7).

      We revised the manuscript and explained ‘pro’, ‘bridge’ and ‘the chromodomain’ in the Results section or figure legends when they first appear. Please refer to other sections of the response for pol-RT definition.

      (d) It is unclear from the main text which portions of pol were chosen to define pol-RT and why. The methods name the 'palm-and-fingers', 'thumb', and 'connections' domains to define RT. In the main text, the 'connection' domain is called 'tether' and is instead defined as part of the 'bridge' region following RT, which is not part of RT.

      We agree that our previous terminology around Pol domains was imprecise and could confuse readers. We have revised the manuscript to distinguish the region used for phylogenetic analysis from the region analysed structurally in Figure 6. The phylogenetic analysis used an extended RT/connection region, comprising the RT polymerase core together with the downstream connection subdomain. This connection subdomain is treated as part of retroviral RT in structural studies, but corresponds to a partial RNase H-like fold and has been interpreted evolutionarily as a degenerated RNase H-like tether domain. It is therefore broader than the RT polymerase core alone, but it is not the complete canonical RNase H domain.

      We now define the Figure 6 “bridge region” separately as the region spanning the canonical RNase H domain and the C-terminal region between RNase H and Integrase. Figure 6 shows that the invertebrate errantiviruses analysed retain an intact canonical RNase H domain immediately downstream of the extended RT/connection region, but differ in the additional downstream RNase H-like or mini-domain structures before Integrase. We have revised the Results, Methods and figure legends accordingly. We also acknowledge that a phylogeny based strictly on the RT polymerase core alone could differ in some local branch relationships, but the major conclusions are supported independently by the Integrase tree, host-taxonomic structure, ENV-type distribution, Pol bridge-region architecture and ENV structural features.

      (2) Insufficient broader context:

      (a) The introduction does not state what defines Ty3/gypsy non-LTR retrotransposons as compared to their closest relatives (Ty1/copia retrotransposons, BEL/pao retrotransposons, vertebrate retroviruses). This makes it difficult to judge the significance and generality of the findings.

      (b) The various known compositions of Ty3/gypsi-like retrotransposons are not mentioned and explained in the introduction (open reading frames, (poly-)proteins and protein domains, and their variable arrangement, enzymatic activities, and putative functions), and the distribution of Ty3/gypsi-like retrotransposons among eukaryotes remains unclear. The introduction does not mention that Ty3/gypsi-like retrotransposons apparently are absent from vertebrates, and Figure 7 is not very clear about whether or not it includes sequences from plants ('Chromoviridae').

      We agree that the Introduction needed more context on Ty3/gypsy retrotransposons. We have revised it to briefly state that LTR retrotransposons include several major lineages, including Ty1/copia, BEL/Pao, Ty3/gypsy and retrovirus-related elements, and that Ty3/gypsy elements are classified primarily by POL similarity and domain organisation. We also now explain that Ty3/gypsy retrotransposons typically encode GAG and POL proteins, with POL providing the enzymatic activities required for reverse transcription and integration, while noting that ORF arrangement and accessory domains can vary between lineages.

      Please note that we stated that our screen did not identify intact env-containing Ty3/gypsy elements in vertebrate genomes that were homologous to the invertebrate errantiviruses analysed here. This is not to say that non-env-containing Ty3/gypsy elements are also absent in vertebrate genomes. Finally, we revised the Figure 7 legend to make clear that the comparison includes representative non-env-containing Ty3/gypsy elements from animals, fungi and plants, including chromovirus or chromovirus-related elements.

      (c) The known association of Ty3/gypsi-like retrotransposons from different metazoan phyla with putative membrane fusion proteins (env-like) genes is mentioned in the introduction, but literature information, whether such associations also occur in the context of other retrotransposons (e.g., Ty1/ copia or BEL/pao), is not provided. The abstract is somewhat misleading in this respect. Finally, the different known types of env-like genes are not mentioned and explained as part of the introduction ('env-f', 'envB', 'retroviral env', others?)

      We expanded the introduction to introduce literature information of known env-associated retroelements, including Ty1/copia and BEL/pao and explained which ENV types are known to be associated to these elements.

      (d) Some key references and reviews might be added:

      - Pelisson, A. et al. (1994) https://www.embopress.org/doi/abs/10.1002/j.1460-2075.1994.tb06760.x (next to Song et al. (1994), for the identification of env in Ty3/gypsy)

      - Boeke, J.D. et al. (1999) In Virus Taxonomy: ICTV VIIth report. (ed. F.A. Murphy),. Springer-Verlag, New York. (cited by Malik et al. (2000) - for the definition and first use of the term 'errantivirus')

      - Eickbush, T.H. and Jamburuthugoda, V.K. (2008) https://doi.org/10.1016/j.virusres.2007.12.010 (on the classification of retrotransposons and their env-like genes)

      - Hayward, A. (2017) https://doi.org/10.1016/j.coviro.2017.06.006 (on scenarios of env acquisition)

      Thank you. We included these references in the introduction.

      (3) Incomplete analysis:

      (a) Mobile genetic elements are sometimes difficult to assemble correctly from shortread sequencing data. Did the authors confirm some of their newly identified elements by e.g., PCR analysis or re-identification in long-read sequencing data?

      Most newly identified elements are found in contigs/chromosomes that are longer than 100kb. The information of the contig/chromosome size, in which the representative copy of the identified elements are found, can be found in the column “CONTIG_SIZE” in supplementary table S1.

      (b) The authors mention somewhat on the side that there are Ty3/gypsy elements with a different arrangement (gag-env-pol instead of gag-pol-env). Why was this important feature apparently not used and correlated in the analysis? How does it map on the RT phylogenetic tree? Which type of env is found with either arrangement? Is there evidence for a loss of env also in the case of gag-env-pol elements?

      We agree that the non-canonical GAG-ENV-POL arrangement is an important feature that was insufficiently integrated into the analysis. We have revised the Results and figure annotations to make this clearer. Specifically, we now indicate GAG-ENV-POL elements in the POL tree in Fig S4 and in the HSV/gB-type ENV alignment/architecture figure S8. These elements are found in Nematoda, Bryozoa and Platyhelminthes and all carry HSV/gB-type ENV. They do not form a single monophyletic group in the Pol tree, but instead occur in distinct host-associated clades. They also show different HSV/gBtype cysteine-bridge architectures. Thus, the GAG-ENV-POL arrangement is unlikely to represent a single recent rearrangement event shared by all such elements; rather, it appears to be associated with several deeply diverged HSV/gB-type errantivirus lineages.

      We have not inferred specific env-loss events for GAG-ENV-POL elements, because doing so would require a separate analysis of related non-env-containing elements.

      (c) Sankey plots are insufficiently explained. How would inconsistencies between trees (recombinations) show up here? Why is there no Sankey plot for the analysis of env-B in Figure 5?

      We agree that the Sankey plot was insufficiently explained. We have revised the Figure 4 legend to clarify that the Sankey plot was used as a qualitative visual summary of global congruence between the Pol extended RT/connection phylogeny and the F-type ENV ectodomain phylogeny. We now state that ribbon crossing alone should not be interpreted as recombination, because tree drawings can be rotated without changing topology and the relative order of clades in the two displayed trees may differ. Instead, the relevant signal is whether Pol-defined clades map mostly to corresponding F-type ENV-defined clades. Strong discordance, potentially reflecting recombination, env exchange or poor phylogenetic resolution, would be expected to appear as extensive splitting or many-to-many connections between Pol and ENV clades.

      We did not include an equivalent Sankey plot for HSV/gB-type ENV in Figure 5 because we did not construct a global HSV/gB-type ENV phylogeny comparable to the F-type ENV ectodomain tree. Instead, HSV/gB-type ENV proteins were analysed by predicted structural organisation and cysteine-bridge architecture, which are shown in Figure 5 and Supplementary Figure S8.

      (d) Why are there no trees generated for env-F and env-B like proteins, including closely related homologous sequences that do NOT come from Ty3/gypsy retrotransposons (e.g., from the eukaryotic hosts, from other types of retrotransposons (Ty1/copia or BEL/pao), from viruses such as Herpesvirus and Baculovirus)? It would be informative whether the sequences from Ty3/gypsy cluster together in this case.

      We agree that comparison with homologous fusogens outside Ty3/gypsy retrotransposons is informative. We have therefore added an expanded F-type ENV ectodomain phylogeny that includes representative viral and retroelement-associated F-like proteins, including baculovirus F proteins, paramyxovirus and pneumovirus F proteins, and the BEL/Pao-associated Drosophila Roo F-like protein. In this expanded tree, the added viral sequences formed family-level clades within the broader F-type ENV diversity. Errantivirus F-type ENV proteins did not cluster as a shallow Ty3/gypsyspecific group or as a recent derivative of a single sampled viral family; instead, they spanned a level of diversity comparable to that separating major viral F-protein groups.

      For HSV/gB-type ENV proteins, we did not generate an equivalent global phylogeny because the primary sequences and domain organisations of viral class III fusogens and errantivirus HSV/gB-type ENV proteins were too divergent for reliable full ecto domain multiple-sequence alignment. Instead, we added a structural comparison with representative viral and retroelement-associated class III fusogens, including herpesvirus gB, rhabdovirus G, orthomyxovirus GP75/GP64-like proteins, baculovirus GP64 and BEL/Pao-associated gB-like proteins. This analysis showed that viral class III fusogens often retained family-specific cysteine-bridge architectures despite low primary-sequence identity. Errantivirus HSV/gB-type ENV groups showed a comparable pattern, retaining lineage-specific cysteine-bridge architectures despite extensive sequence divergence. We have revised the Results, Methods and supplementary figure legends to clarify these analyses and to distinguish the phylogenetic analysis of F-type ENV from the structural comparison of HSV/gB-type ENV.

      (e) Did the authors identify any other env-like ORFs (apart from env-F and env-B) among Ty3/gypsy retrotransposons? Did they identify other, non-env-like ORFs that might help in the analysis? It is not quite clear from the methods if the searches for env-F and envB - containing Ty3/gypsy elements were done separately and consecutively or somehow combined (the authors generally use 'env', and it is not clear which type of protein this refers to).

      We agree that this was not sufficiently clear. We have revised the Methods to clarify that the search was designed to identify Ty3/gypsy elements carrying ORFs structurally resembling known envelope/fusogen proteins. In the iterative tBLASTn searches, bait sequences representing both F-type ENV and HSV/gB-type ENV were included together in each round, rather than being searched as two entirely separate pipelines. Candidate elements were then annotated and classified by ORF structure, HHpred/domain similarity and structural prediction.

      Among intact Ty3/gypsy candidates recovered by this strategy, we identified two recurrent classes of env-like ORFs: F-type env and HSV/gB-type env. We did not identify an additional recurrent class of env-like ORF among the intact Ty3/gypsy elements analysed here. We also did not identify other recurrent non-env accessory ORFs that were informative for the phylogenetic analyses beyond the GAG, POL and ENV features described in the manuscript.

      (f) Why was the gag protein apparently not used to support the analysis? Are there different, unrelated types of gag among non-LTR retrotransposons? Does gag follow or break the pattern of co-evolution between RT and env-F/env-B?

      We agree that the role of GAG in the analysis should be clarified. GAG ORFs were used during element annotation to identify intact GAG-POL-ENV or GAG-ENV-POL retrotransposon architectures, but we did not use GAG as a major phylogenetic marker because GAG proteins are less conserved and less reliably alignable across deeply diverged Ty3/gypsy elements than the enzymatic POL domains. Our central question was the acquisition and long-term retention of env-like ORFs by POL-defined Ty3/gypsy retrotransposons. We have revised the Methods to clarify that GAG was used for structural annotation and intactness assessment, whereas phylogenetic analyses were based on the Pol extended RT/connection region and Integrase domain.

      (g) Data availability. The link given in the paper does not seem to work (https://github.com/RippeiHayashi/errantiviruses_2025/tree/main). It would be useful for the community to have the sequences of the newly identified Ty3/gypsy retrotransposons listed readily available (not just genome coordinates as in table S1), together with the respective annotations of ORFs and features.

      The GitHub repository that contains suggested data is made public. Please check the link again.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Additional Analyses That Could Strengthen Claims (but I concede might be well beyond the scope of this study):

      (1) Reconciliation and Gene Tree-Species Tree Analysis: Implementing explicit genetree-species-tree reconciliation (e.g., using Notung or ALE) could formally test the frequency of horizontal transfers vs. vertical inheritance in errantivirus evolution.

      We performed targeted host-taxonomy concordance analyses in representative well-sampled clades to address this question, as described in the revised manuscript.

      (2) Functional Validation in Non-Insect Hosts: RNA-seq or proteomics from representative non-insect hosts could reveal whether env genes are expressed, and-if possible-experimental assays for envelope function would move beyond computational inference. This is the one thing that might be done in a revision.

      We agree that direct functional validation in non-Drosophila species would substantially strengthen the conclusions. Such experiments are beyond the scope of the present revision. However, we now cite our recent work (PMID: 41922845) showing conserved piRNA-mediated defence against errantiviruses across insect orders, providing indirect evidence that these elements remain biologically active outside Drosophila.

      (3) LTR Age Dating: Estimating insertion ages using LTR divergence across clades would contextualize the timing of expansion events and help test hypotheses about ancient vs. recent proliferation.

      We agree that LTR divergence-based insertion dating would be informative for estimating the timing of recent expansion events. However, our main evolutionary conclusions concern the deeper history of env acquisition and long-term retention across metazoan lineages, rather than the precise insertion age of individual genomic copies. Because the dataset includes elements from highly divergent genomes with variable assembly quality and many multicopy families, systematic LTR dating across all clades would require additional curation and is beyond the scope of the present revision.

      (4) Comparative Host Defense Analysis: Surveying host antiviral or transposon defense systems (e.g., piRNA, APOBEC) in lineages rich in errantiviruses could test for signatures of recurrent molecular arms races.

      We agree that comparative analysis of host defence pathways, including piRNA and antiviral systems, is an important future direction. However, a systematic survey of host-defence evolution across all errantivirus-rich lineages is beyond the scope of the present revision.

      Reviewer #2 (Recommendations for the authors):

      Apart from my comments in the Public Review, I have a few additional minor recommendations:

      (1) The authors frequently use the term "active" to describe complete retrotransposons. However, in transposon biology, "active" implies recent or ongoing transpositional activity and should therefore be used with caution. Terms such as "complete" or "fulllength" would be more appropriate in this context.

      We agree that intact ORF structure should not be over-interpreted as evidence of recent mobilisation. We have therefore revised the manuscript to distinguish element intactness from evidence of recent expansion. Specifically, we now use “intact” or “full-length” to describe element structure. We also clarify that the presence of multiple highly similar copies in the same genome, defined as >98% nucleotide identity across >98% of the three ORFs, is evidence consistent with recent or ongoing genomic expansion, but not definitive proof of current transposition. These changes have been made in the Abstract, Results, Methods and Discussion.

      (2) In the third paragraph of the Results, the authors conclude that many identified errantiviruses were mobilized recently. However, since the search strategy specifically targets complete and uninterrupted elements, this may introduce a bias toward younger elements, making the conclusion somewhat circular.

      We agree and have revised the wording to avoid over-interpreting completeness as evidence of recent mobilisation. We now distinguish between intact/full-length element structure, which was part of our search strategy, and independent evidence for recent or ongoing mobilisation, such as the presence of multiple highly similar copies. We have therefore softened statements that previously implied that intact ORFs alone demonstrate recent activity.

      (3) The first paragraph of the Introduction requires appropriate references.

      We have now added references to enhance the readability of the first paragraph of the introduction.

      (4) It would benefit readers if the retrotranspositional process were briefly explained, including a description of what the tRNA primer binding site (PBS) is and its role in retrotransposition.

      In the same part of the introduction, we now describe the role of the tRNA primer binding site in retrotransposition.

      Reviewer #3 (Recommendations for the authors):

      Suggestions and requests for clarification currently are part of the public review, as they also point the reader to critical open questions if the authors decide not to amend their version of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this important study, the authors have performed a zebrafish drug screen to identify suppressors of atherogenic lipoproteins. They utilize a well-established LipoGlo assay to find molecules that modulate these lipoproteins, identifying 49 potential hits. They perform some validation experiments, including studies linking enoxolone to its likely inhibitory effect on a specific transcription factor, HNF4alpha. Overall, the results are convincing and robust, and will open up new areas of exploration for those investigators interested in in vivo lipid biology.

      We appreciate the manuscript assessment and provide new data (see below) that significantly increases the “strength of the evidence”.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The authors performed a whole-organism chemical screen with over 3000 agents. Such screens are challenging, and the authors used strict criteria for determining hits. The conclusions of this study are well supported by the presented data.

      Weaknesses:

      There are areas within the study and writing that can be improved and extended, specifically within the gene expression studies.

      We appreciate the Reviewer recognizing the strength of the data and the challenging nature of performing a whole animal small molecule screen. With regard the “Weaknesses”, as the Reviewer suggested we improved and clarified the text and “extended” the study by adding analysis and discussion of six additional hits (see Figure 3 and Sup. Fig.1) and modified the results section with the addition of over a page of text describing the additional phenotyping results:

      “Secondary characterization of selected validated B-lp lowering hits.

      To further evaluate the biological relevance and potential mechanisms of selected validated hits, we performed secondary analyses assessing total B-lp levels, larval morphology, and lipoprotein size distribution. Our lab previously defined the method to measure total B-lp levels from whole animals by using the homogenate of a single zebrafish larva [35]. We confirmed that treatment of animals with 4 µM pomiferin significantly reduced total B-lp levels measured from homogenates collected from whole animals after treatment (p = 8.6x10<sup>-4</sup>; Figure 3A). However, pomiferin treatment (4 µM) produced animals with reduced body length and lethality at higher doses, suggesting developmental toxicity may confound interpretation of its B-lp-lowering effect.

      Treatment of animals with riboflavin tetrabutyrate (Figure 3B) and calcipotriene (Figure 3C) reduced total B-lp levels (p < 2x10<sup>-16</sup>) but did not affect larval morphology. A key feature of B-lps is their size, often a proxy for the total amount of lipid in the particle [35]. Particle size can impact the particle's lifetime (e.g. in metabolically healthy humans, small particles are cleared rapidly by the liver) [54–56]. Thus, we also assessed whether these compounds alter B-lp size distribution. Animals were treated for 48 h with vehicle, 5 µM lomitapide, or a drug of interest, and whole-animal homogenates were prepared and subjected to native polyacrylamide gel electrophoresis followed by chemiluminescent imaging. B-lps were classified into four classes based on gel migration: zero mobility (ZM), very low-density lipoproteins (VLDL), intermediate-density lipoproteins (IDL), and low-density lipoproteins (LDL) as previously described [35]. Lomitapide treatment effectively reduces VLDL particles and increases LDL particles [35] (Figure 3E), whereas riboflavin and calcipotriene did not affect lipoprotein classes. Thus, riboflavin tetrabutyrate and calcipotriene reduce total B-lp levels without overt developmental toxicity or changes in lipoprotein subclass distribution, suggesting they may act through mechanisms that decrease overall particle abundance rather than altering lipoprotein turnover or catabolism.

      Although doxycycline treatment lowered B-lp levels in whole fixed animals in the primary screen and validation studies, we did not observe a reduction in total B-lps in whole-animal homogenates (Figure 3D). However, we detected a slight increase in VLDL levels (p < 2x10<sup>-16</sup>; Figure 3E), suggesting that doxycycline may alter lipoprotein composition or distribution rather than total particle abundance.

      Alternatively, two structurally related compounds, thiethylperazine and prochlorperazine, at 4 µM significantly reduced (p < 1.4x10<sup>-10</sup> and p < 1.2x10<sup>-6</sup> respectively), B-lp levels measured from whole-animal homogenates (Figure 3F and 3H). Furthermore, both 8 µM thiethylperazine and 8 µM prochlorperazine increased relative LDL (p = 1.2x10<sup>-3</sup> and p = 2.3x10<sup>-3</sup>, respectively) and decreased relative VLDL levels (p = 8.4x10<sup>-4</sup> and p = 9.1x10<sup>-4</sup>, respectively; Figure 3G and 3I) suggesting a shift toward smaller lipoprotein particles and a potential alteration in lipid processing or clearance pathways. Together, these results highlight the diversity of mechanisms among validated hits, ranging from compounds that reduce total B-lp abundance without affecting B-lp class composition to those that shift B-lp class distribution, while also underscoring the importance of secondary assays to distinguish true B-lp modulators from those that likely produce a B-lp effect through generalized toxicity.

      Enoxolone significantly reduces B-lps in the larval zebrafish.

      Hit compounds were prioritized for follow-up studies based on reproducible dose-dependent responses, minimal toxicity as indicated by normal morphology over development, lack of direct NanoLuciferase inhibition, and the presence of literature suggesting potential links to lipid metabolism. One compound meeting these criteria was enoxolone, also known as 18β-Glycyrrhetinic acid, (Figure 2 Drug 20, Supplemental Table 1, Supplemental Figure 1T, Supplemental Figure 2L).”

      The Discussion now has the following additional text:

      “Further validation of these hits demonstrated a wide range of potential mechanisms of lipoprotein regulation. We identified hits that affected larval development, some hits that reduced total B-lp levels, and several structurally related compounds that directly reduced B-lp particle size (Figure 3).”

      Reviewer #2 (Public review):

      Strengths:

      The study was methodical and robust, using a published and well-validated zebrafish LipoGlo model. The authors validated the hits from the screen independently and considered the possibility that some drugs may have been detected as false positive results due to effects on the enzymatic activity of NanoLuciferase; only one hit, verteporfin, was shown to be a false positive. Using LipoGlo-Electrophoresis, the authors are able to obtain extra insights into the ApoB-lipoprotein size/subclass distribution. They showed that while enoxolone treatment reduces total B-lps, there are no overt changes in B-lp size distribution compared to vehicle-treated animals, other than a slight increase in the zero mobility (ZM) fraction, which contains very large particles and/or tissue aggregates. In contrast, the positive control, lomitapide, does show a change in B-lp size distribution compared to vehicle-treated animals - an increase in frequency of LDLs (low-density lipoprotein), but a decrease in VLDLs (very low density lipoprotein). This study also assesses the LipoGlo-Electrophoresis profile of HNF4⍺ inhibitors. Work in the zebrafish larvae means that the effect on overall development and an entire vertebrate organism can also be assessed. Finally, the authors applied a thorough statistical measure to define a hit, using the Strictly Standardized Mean Difference (SSMD) method.

      We appreciate that the Reviewer valued the rigour and robustness of our approach.

      Weaknesses:

      While the screen was thorough and well-validated, the authors missed a chance to provide a lot of extra significance to a wide range of readership. While the hits were thoroughly validated and displayed, the authors could have also presented the LipoGlo-Electrophoresis for all validated hits or at least a number of them. This would hugely increase the insights into these compounds. Also, the authors chose to validate and follow up a mechanism for Enoxolone, yet this hit was already known to modulate lipid metabolism through HNF4⍺, therefore, hugely limiting the impact of the paper. So what the authors have shown that is novel is only subtly added to this - consistent in vertebrate models, RNA sequencing of pathways, further validation of the HNF4⍺ pathway, and a profile of resulting B-lp size distribution. It seemed an easy way out to pick such a candidate, and they could have followed up by validating more thoroughly a completely novel drug. Also, the authors' prior paper showing the methodology also depicted complementary EM and LipoGlo-microscopy approaches. The microscopy especially, would have been an easy complementary add-on to the screen to really give extra insights into B-lp metabolism in a whole organism for all candidates. This felt like a missed opportunity.

      Here we agree and added Fig. 3 describing the phenotyping of 6 additional compounds including some LipoGlo-Electrophoresis analyses as suggested by the Reviewer. The text of the Results section was modified as described for Reviewer 1 (see above).

      Reviewer #3 (Public review):

      Strengths:

      The study uses a well-validated in vivo stain (LipoGlo) for measuring lipoproteins in the context of a developing whole organism with a quantitative read-out on a high-throughput platform, allowing for screening of thousands of compounds altering the complex metabolic/physiologic functions necessary for lipoprotein production.

      The use of genetic mutant HNF4alpha to assign the mechanism of action to the prime candidate compound studied (enoxolone) is a powerful approach for this challenging aspect of chemical genetics studies.

      We appreciate that the Reviewer understands how challenging it can be to assign a mechanism to any small molecule and thereby recognizes the power of the zebrafish model combined with our unique lipoprotein phenotyping tools.

      Weaknesses:

      As shown in Figure 5A, the HNF4alpha mutant homozygous -/- already lowers lipoproteins. Is it just that the mutant level is already at a minimum in this homozygous mutant (and thus enoxolone cannot induce even lower lipoprotein levels), or is it true that the enoxolone molecule is primarily acting through this TF (i.e. HNF4alpha homozygous mutant is truly epistatic to enoxolone function) as favored in the text.

      While it is definitely interesting to study enoxolone effects during whole embryo development, the link to HNF4alpha had previously been described in the literature, as pointed out by the authors. The generalizability of the approach to identify truly novel pathways remains to be fully realized, but sharing this available screen data to date will invite further inquiry and be very valuable to the community.

      Here too we agree that a link between enoxolone was proposed in the literature. However, we added quite a lot of additional insight regarding the transcriptional targets shared by HNF4alpha and enoxolone. The goal of identifying the mechanism(s) of action of other novel small molecule hits from the screen is important and that work is ongoing.

      Figure 5 - The same allele of HNF4alpha loss of function/hypomorph (rdu14) is used in both 5A and 5B, but labeled differently in each subpanel. This is explained in the figure legend, but could be updated to use the same nomenclature in both panels to clarify the Figure presentation.

      We thank the Reviewer for catching this and have modified the Figure (now Fig. 6) so the subpanels are labeled identically to avoid any confusion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors describe statistical methods to improve the calling of hits. In the results section, they discuss the use of fold change and of strictly standardized mean difference as criteria to account for the variability that occurs within in vivo screens. The combination of both criteria resulted in a significant reduction in the total number of hits, from 487 (16%) to 49 (1.6%), which is a large decrease in total hits. The authors should comment on whether some of the 487 compounds were randomly tested individually to confirm that they were true negatives and compared to the true hits of 49. For example, did the authors independently test calcipotriene and diphenylboric acid to confirm that these are truly negative?

      Initially, we did not directly retest any of the 487 compounds outside of the top 49 hits with the intention to compare their effect size to the top compound. While we do suspect that there is a chance that some of these compounds may still have a significant effect on lipoprotein levels, we prioritized compounds with the strongest overall biological effect size and largest statistically significant effect. We agree with the Reviewer that it would be interesting to continue to test more of these compounds and examine whether the lipoprotein reduction phenotype correlated with the primary screen effect, this would be a large experimental undertaking, though a randomized subset could be tested. Validation studies of Calcipotriene (Supplemental Figure 2E) showed a small, but significant, lipoprotein reduction at only the 0.25 µM dose across 3 independent experiments. Because the effect is small and not a dose-response, it is not a highly prioritized hit. We did not validate diphenlylboric acid in this study but will in the future.

      (2) The transcriptomics profile studies are not well described. The authors do not provide a detailed list of differential genes for each of the conditions in their dataset. This should be included.

      Supplemental Table 2 includes the fold change and p-values of all genes measured in differential expression analysis. We agree with the Reviewer and added new supplementary tables (Supplemental Table 5 and 7) that contain the differentially expressed genes at each treatment duration.

      (3) The Gene Ontology analysis results are not accompanied by a list of genes that match the GO terms. The authors should include these results. This part was confusing as it was not clear if the cholesterol pathway affected was due to an abundance of genes that were up- or down-regulated.

      We agree with Reviewer 1 and added Supplemental Table 6 that contains each gene ontology term with the associated DE genes that contribute to the significance of that term as well as explicit directionality information.

      (4) Figure 6A shows heat maps of differential gene expression, but there is no key provided in the figure or legend. Are they color-coded for fold change or log2 fold change?

      We agree that Figure 6A can be approved and added a key to the legend as suggested by the Reviewer.

      (5) The overlap of genes that are changed in hnf4a mutants and with enoxolone is provided as percentages, but the actual genes are not listed. Do these genes represent cholesterol biosynthesis pathways? There are other bioinformatic tools that the authors could apply to their dataset for further analyses. For example, enrichr (https://maayanlab.cloud/Enrichr/) is one tool to query for GO Biological processes, cell/tissue type, and even overlap of genes with genetic models and other disease states. An extension of these bioinformatic studies will be useful to determine if other pathways are relevant.

      We agree with the Reviewer and added a table of the genes that overlap between drug treatments and hnf4a mutants (Supplemental Table 7). We also appreciated the Reviewer’s suggestion to deploy Enrichr, which we have done and modified the Results section to now read:

      “Of the 439 differentially expressed genes from 12, 16, and 24 hours post-treatment, 34 differentially expressed genes are shared between all three treatment durations and are associated with gene ontology terms related to carbohydrate metabolism and signaling pathways (Figure 7D). We expanded this analysis using the bioinformatic tool Enrichr [68–70], which largely recapitulated gene ontology results described above. However, Enrichr analysis revealed significant overlap between several late enoxolone-responsive gene sets and transcriptional signatures associated with prochlorperazine, another compound identified in our screen (Figure 2, Figure 3, Supplemental Figure 1AO, Supplemental Figure 2Z). Enrichment of prochlorperazine-associated signatures was observed at 12 (prochlorperazine MCF7 up, 4/58 genes [INSIG1;IRF7;ISG15;ATF3], adjusted p-value = 0.006), 16 (prochlorperazine MCF7 up, 6/58 genes [INSIG1;DDIT4;IRF7;PMAIP1;ISG15;ATF3], adjusted p-value = 0.00008; prochlorperazine PC3 up, 3/29 genes [INSIG1;DDIT4;ATF3], adjusted p-value = 0.006), and 24 hours post-treatment (prochlorperazine PC3 up, 6/29 genes [DUSP5;INSIG1;DDIT3;TRIB3;SQSTM1;ATF3], adjusted p-value = 0.0001; prochlorperazine MCF7 up, 7/58 genes [DDIT3;INSIG1;IRF7;PMAIP1;ISG15;SAT1;ATF3], adjusted p-value = 0.0005). These results suggest that enoxolone and prochlorperazine may perturb overlapping molecular pathways, an observation that warrants further investigation. Ultimately, these data demonstrate distinct early and late responses to enoxolone treatment, and the early response modulates key lipid metabolism pathways.”

      (6) Figures 1E and 1F would benefit if the exact spots/points where enoxolone and other hits mentioned in the text were labelled.

      Great idea, we modified Figure 1F as suggested.

      (7) Figure 3 and Figure 4 graphs should state Fold change on the Y-axis title.

      We are thankful the Reviewer noticed this typo and we have fixed both Figures.

      Reviewer #2 (Recommendations for the authors):

      (1) To boost the impact for more readers, the authors should include the LipoGloElectrophoresis and LipoGlo-Microscopy results from a few more of the validated hits, especially ones that are completely novel (unlike Enoxolone, which already had a known role in lipid metabolism). Results on enoxolone are useful as a validation of the assay, mostly with some minor additional insights.

      We agreed with Reviewer 2 and added more validation testing of for a few hits (see response to Reviewer 1, Strengths and Reviewer 2, Weaknesses). We did not perform these experiments on each drug for technical reasons, mainly because these experiments are low-throughput (especially the Microscopy). Nonetheless, we performed many additional experiments to add phenotyping data for 6 new drugs that included multiple LipoGlo-Electrophoresis panels to an entirely new Figure.

      (2) The authors should include raw data from the screen from all drugs tested in the supplementary and then for which SSMD was calculated for, providing in an excel sheet or similar the values and how these were calculated, i.e. the 487 unique drugs that lower B-lp levels with an SSMD cutoff of < -1.0.

      This information was provided in the supplemental file as separate .csv files with the associated R script, which can be run locally and contains the SSMD functions. Considering the Reviewer comments, we ensured this information in provided in Supplemental Tables 1 and 2.

      (3) Page 3, lines 23-24: What does the 2 to 4 fold chance mean? Perhaps rewrite: Genetic mutations in Lipoprotein(a) increase the chance of heart attack or stroke 2-4 fold greater than without the mutation.

      We agree and the sentence now reads: Patients with genetic mutations in the Lipoprotein(a) encoding gene have a 2 to 4 fold increased risk of sudden heart attack or stroke.

      (4) Figure 1, for C and D, label some of the most significant hits and definitely show where exonolone lies.

      We agree see response to Reviewer 1 Pt6

      (5) Page 6, lines 1-9: I'm a bit confused why this is here if you do not present the data.

      We thought this was relevant information to share for researchers that running drug screens with positive controls and defining hit cutoffs. In light of the Reviewer’s comment, we removed the last sentence from this paragraph.

      (6) Page 6, line 8: This needs better justification of why you are validating enoxolone rather than other hits; otherwise, it could seem like cherry picking. Especially as enoxolone is known to affect lipid metabolism. Otherwise present more details of a couple of validated candidates.

      We agree. As the Reviewer requested, we validated more compounds (described above) and modified the text of the results to elaborate on our justification for selecting enoxolone for further study. The text of the results now reads: Hit compounds were prioritized for follow-up studies based on reproducible dose-dependent responses, minimal toxicity as indicated by normal morphology over development, lack of direct NanoLuciferase inhibition, and prior reports the presence of literature suggesting potential links to lipid metabolism. One compound meeting these criteria was enoxolone, also known as 18β-Glycyrrhetinic acid, (Figure 2 Drug 20, Supplemental Table 1, Supplemental Figure 1T, Supplemental Figure 2L).

      (7) Supplementary Figures 1 and 2: The resolution is too low, and the reader cannot even see the charts or the text.

      We agree and now have uploaded higher resolution images

      (8) Page 7, lines 31-35: Needs a higher resolution and magnified image to merit this 'offhand' statement. Also cite reference [59] here.

      We agree with the Reviewer and added magnified insets of the heart. As far as the suggestion of adding Ref 59, we do not see the connection to that paper (A point mutation decouples the lipid transfer activities of microsomal triglyceride transfer protein PLOS Genetics 16:e1008941)

      (9) Figure 3E: In addition to the proportions graph would be useful to also have an absolute amount of lipoprotein in each class graph.

      While there may be changes in total luminescence values from lane to lane in these gels, we have not fully validated the absolute quantitation of a full lane. We typically use plate-based whole-animal assays to determine total absolute lipoprotein levels and calculate the proportion of the whole lane for each lipoprotein class, as described in our prior publication detailing the assay. We do expect that there is some additional variation incorporated into the native PAGE assay due to sample freeze/thaw, dilution, and loading.

      (1) Page 10 lines 27-28: "Continual statin use for more than 1 year reduced circulating Blps and all-cause mortality by ~30% in individuals with high B-lp levels." This sentence doesn't seem right, intimates continual statin use causes death - I don't think that's right.

      We thank the Reviewer for catching this and have corrected the sentence. It now reads: “Continual statin use for more than 1 year in individuals with high B-lp levels reduced circulating Blps and lowered all-cause mortality by ~30% [12,13].”

      (11) Page 12, line 25: "canlikely" is a typo, should be can likely.

      We fixed that sentence and now reads: “Further, the drug screening paradigm we developed using the LipoGlo system is highly scalable and can be deployed to screen large novel drug libraries to identify many additional B-lp-lowering compounds.”

      (12) Figure 1 legend: "An ordered plot of each SSMD score measured from 5 μM lomitapide treated animals from each 96-well plate (n = 1381) relative to respective vehicle treatment." This comes across as though it's 5uM Iomitapide/vehicle. But it's the SSMD score of each drug compared to Iomitapide and relative to the respective vehicle (I think) - make it clearer.

      We agree and clarified the legend so it now reads: “…(D) An ordered plot of each SSMD score measured from positive control (5 µM lomitapide) treated animals from each 96-well plate (n = 1381) relative to respective vehicle treatment. The solid black line at y = 0 represents the divide in increased and decreased SSMD score, the solid blue line at y = -1.41 represents the curve's inflection point, and the dashed black line at y = -1 represents the SSMD (open circles) cutoff used to define a hit. “

      (13) Figure 1E: What is the x axis?

      Each data point on the x-axis represents each drug at every dose tested, we will clarify the test. The legend now reads: “(E) A plot of SSMD scores measured from each drug at each dose tested, each open circle represents the SSMD score of an individual drug at an individual dose.” In addition, “Compound (each dose tested)” was added to the x-axis of the figure panel.

      (14) Figure 2: Would you not have space to put the drug names in the figures? Where, for example, is enoxolone?

      We agree and have updated the figure accordingly.

      (15) Figure 3A-C: label enoxolone on the x axis.

      We agree and added this text to what is now Figure 4.

      (16) Figure 3D: Looks like delayed development with enoxolone, if left to grow, would the embryos develop normally?

      We did not examine if animals treated from 3-5 dpf develop normally beyond 5 dpf.

      (17) Figure E. Are stars all compared to vehicle control? Perhaps useful to have lines to indicate what are the significantly different relationships.

      Comparison in these experiments are always to the vehicle (negative control) and clarified the legend considering the Reviewer’s comment we modified the legend to now read,”… * <0.05 as compared to vehicle.”

      (18) Figure 3E: As well as the proportion of total lipoprotein, it would also be beneficial to see absolute lipoprotein levels.

      See above response to Reviewer 2 Pt9.

      (19) Figure 4: Does the overall health or size of the animal correlate with the luminescence score?

      While we do know that, in untreated animals, lipoprotein levels vary with age (and, thus, size), we have not examined this more granularly than in 24-hour time points after treatment. Further, we have not examined this in the context of a drug treatment.

      (20) Figure 4D: Again the absolute in each fraction would be meaningful, also the 5078 looks brighter?

      See above response to Reviewer 2 Pt9.

      (21) Figure 4A, 5A: Would the traces (line plots) not be useful here to see the overall dynamics over time?

      We considered presenting Figure 5A this way but decided to keep the plots as is because we wanted to show the individual points which make the line blots very difficult to read. Further, our analysis evaluates individual animals at each time point as it is not possible to follow the same animal over time.

      Reviewer #3 (Recommendations for the authors):

      Figure 4: Consider using standard scientific notation for the very small p values in some of the figure legends.

      We agree and adjusted the p-value notation as suggested.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments following re-submission:

      Overall, I think the authors have done a satisfactory job of addressing most of the points I raised.

      There’s one final issue which I think still needs better discussion.

      I think reviewer 2 articulated better than I have the point I was concerned about: the relationship between JNDs and metamers as depicted in the schematics and indeed in the whole conceptualization.

      I think the issue here is that there seems to be a conflating of two concepts- ’subthreshold’ and ’metamer’-and I’m not convinced it is entirely unproblematic. It’s true that two stimuli that cannot be discriminated from one another due to the physical differences being too small to detect reliably by the visual system are a form of metamer in the strict definition ’physically different, but perceptually the same’.

      However, I don’t think this is the scientifically substantial notion of metamer that enabled insights into trichromacy. That form of metamerism is due to the principle of univariance in feature encoding, and involves conditions in which physically very different stimuli are mapped to one and the same point in sensory encoding space whether or not there is any noise in the system. When I say ’physically very different’ I mean different by a large enough amount that they would be far above threshold, potentially orders of magnitude larger than a JND if the system’s noise properties were identical but the system used a different sensory basis set to measure them. This seems to be a very different kind of ’physically different, but perceptually the same’.

      We are in full agreement with this. Typically, the notion of metamers is about deterministic information loss, which can be modeled as a projection from a high-dimensional physical space to a lower-dimensional perceptual space. This is the topic of the paper, and is analogous to the work on color matching in the 19th century.

      In contrast, sensitivity to small differences within the perceptual space due to internal noise is usually addressed by other methods, such as signal detection theory, a topic which is not the focus of this paper. It is analogous to the work on discriminability of colors such as MacAdam ellipses (MacAdam, 1942). It is instructive to look at progress in the color field. The color matching experiment and the question of metamerism is quite well worked out, whereas the question of how to quantify discriminability within that space has been an ongoing topic of investigation for over a century.

      Here, we aim to develop and test a model of metamers, analogous to the color matching experiments, but we do not attempt to develop a model of discriminability. Nevertheless, while the two types of information loss are conceptually distinct, they are both present in the nervous system of the observer, and both are reflected in the performance vs. scaling plots in our paper. We have added clarifications about this point in the introduction on page 2, starting on line 45, and the discussion, starting on page 16, line 397.

      Finally, regarding physical differences between stimuli: the differences between target images and synthesized metamers are quite large (high mean squared error), as shown in Appendix 5. In no condition did we present subjects with stimulus pairs that were physically similar.

      I do think the notion of metamerism can obviously be very usefully extended beyond photoreceptors and photon absorptions. In the interesting case of texture metamers, what I think is meant is that stimuli would be discriminable if scrutinised in the fovea, but because they have the same statistics they are treated as equivalent.

      The notion of “texture metamers” is perhaps a reference to the work by Freeman and Simoncelli (2011), whose stimuli are similar to ours: when synthesized using a model with sufficiently small scaling, they are indiscriminable, and therefore metamers. The reviewer is of course correct that the stimulus pairs are not metamers when the observers move their eyes due to differences in spatial encoding as a function of eccentricity. That is, they are only metameric under a specific set of viewing conditions, and they are not metameric when those conditions are violated. The same is true for color metamers, as the spectral sensitivity of the cones also differ with eccentricity (Stockman and Sharpe, 2000).

      I think the discussion of this could still be clearly articulated in the manuscript. It would benefit from a more thorough discussion of the difference between metamerism and subthreshold, especially in the context of the Voronoi diagrams at the beginning.

      We agree that a more thorough discussion of the diagrams could help clarify the issues to the reader. We have modified the caption of figure 1 with the goal of clarifying interpretation of the diagrams, and see also our discussion earlier in this note about noise and discriminability.

      It needs to be made clear to the reader why it is that two stimuli that are physically similar (e.g., just spanning one of the edges in the diagram) can be discriminable, while at the same time, two stimuli that are very different (e.g., at opposite ends of a cell) can’t.

      Do the cells include BOTH those sets of stimuli that cannot be discriminated just because of internal noise AND those that can’t be discriminated because they are projected to literally the same point in the sensory encoding space? What are the strengths and limits of models that involve the strict binarization of sensory representations, and how can they be integrated with models dealing with continuous differences? These seem like important background concepts that ought to be included in either the introduction of discussion sections. In this context it might also be helpful to refer to the notion of ’visual equivalence’ as described by:

      This is an important point and we appreciate the reviewer raising it. In brief, as one traverses a region in one of the Voronoi diagrams, the images are changing physically but subject to the constraint that they all project to the same single point in the reduced perceptual space. When one crosses from one region to another, the images now project to a different point in the perceptual space. Whether or not that the two locations in the perceptual space are distant enough to be distinguishable given the internal noise is a question pertaining to the topic of JNDs in the perceptual space, rather than the mapping from physical space to the perceptual space. We do not address that question in detail in this paper, though we do now reference it in the caption of figure 1, as well as in the new sections in the introduction and discussion mentioned earlier in this response.

      We do note that the perceptual space is not discrete: the model outputs are real-valued. The apparent discretization is a limitation of the simplified 2-D schematics.

      Ramanarayanan, G., Ferwerda, J., Walter, B., & Bala, K. (2007). Visual equivalence: towards a new standard for image fidelity.ACM Transactions on Graphics (TOG), 26(3), 76-es.

      Other than that, I congratulate the authors on a very interesting study, and look forward to reading the final version.

      Reviewer #2 (Public review):

      Summary:

      The authors have improved clarity overall and have spoken to most of the issues raised by the reviewers. There are still two outstanding problems however, where issues raised during the review were inappropriately dismissed in the manuscript. These should be explicitly addressed as limitations to the results presented (no eye tracking), and early pilot experiments that informed the experiments as presented (pink noise) rather than brushed off as ’unnecessary’ and ’would be uninformative’.

      Eye tracking:

      It is generally accepted that experiments testing stimuli presented at specific locations in peripheral vision require eye tracking to ensure that the stimulus is presented as expected, in particular, in the correct location. As I stated in the previous round of review, while a stimulus presentation time of 200ms does help eliminate some saccades, it does not eliminate the possibility that subjects were not fixating well during stimulus onset. I am also unclear what the authors mean by ’trained observer’ in this context, though the authors state that an author subject in a different portion of the paper is an ’expert observer’. Does this mean the ’trained observers’ are non-expert recruited subjects?

      Given the conditions tested differ from previous work (Freeman & Simoncelli, 2011) ‘these differences are a main contribution of the paper!’ which DID include eye tracking in a subset of subjects, it is entirely possible to get similar results to this work in the context of non eye-tracking controlled stimulus presentation. The reasons now in the manuscript are not reasons that make eye tracking ’considered unnecessary’.

      I appreciate that the authors now state the lack of eye tracking explicitly, but believe the paper needs to at least state that this is a limitation of the results reported, and eyetracking being ’considered unnecessary’ is unreasonable, nor a norm in this subfield.

      By “trained” observers, we mean people who were recruited from the community of vision science labs at NYU and who have participated in many visual psychophysics experiments. All of the participants are “trained” in this sense, and are thus used to maintaining fixation while performing peripheral tasks. One of these participants, an author, was also an expert in the specific content area of the paper. By “expert”, we mean high familiarity with the stimulus types and models employed in the paper.

      We have now further clarified this in the text in the subsection of the methods on Observers, on page 22.

      We also discuss the issue at greater length in the methods subsection Apparatus, on page 26. We removed the word "unnecessary" and make it clear that while we don’t think our results are undermined, the lack of eye tracking is nonetheless a limitation.

      N=1: The authors now state clearly the limitations of a single subject in the manuscript, and state the expertise level of this subject.

      Large number of trials: The authors now address this and include an enumeration of the large number of trials.

      Simple Models / Physiology comparison: I support the choice to reduce claims regarding tight connections to physiology, and appreciate the explanation of the luminance model.

      Previous Work: I appreciate the author’s changes to the introduction, both in discussing previous work and citation fixes.

      Blurred White, Pink Noise: While the authors now address pink noise, the explanation for such stimuli being expected to be uninformative is confusing to me. The manuscript now first states that pink noise is a natural choice, then claims it would be uninformative, while also stating in the rebuttal (not the manuscript) that they tried it and it indeed reduced the artifacts they note. The logic of the experiments indeed relies on finding the smallest critical scaling value, which is measured by subjects determining if a synthesis is similar or different to a target or second synth. A synthesis free from artifacts would surely affect the subjects responses and the smallest critical scaling measured.

      The statement that the authors experimented with pink noise early on and found this able to address the artifacts should be stated in the manuscript itself, not just in the rebuttal, and the blanket statement that this experiment would be ’uninformative’ is incorrect. Surely this early pilot the authors mention in the rebuttal was informative to designing the experiments that appear in the final paper, and would be an informative experiment to include.

      First, we did render some test stimuli with pink noise seeds, but we did not collect psychophysical data, hence there are no results we could add. Visual inspection of these stimuli was indeed clarifying in the following sense. The pink noise stimuli had fewer high-frequency “artifacts”. If our goal was to synthesize stimuli that are indistinguishable from the original stimulus, as one might do to save compute power when in a device that for foveated rendering, then starting with pink noise would be better than starting with white noise. Our purpose was just the opposite. For our experiments, the artifacts were just what we wanted: the more artifacts, the better. The reason is that a strongest test of a metamer model is whether two stimuli that are as physically different from one another as possible, are nonetheless indistinguishable when their model representations are the same. Stimuli synthesized from pink noise seeds are harder to discriminate from the target stimulus, not easier. Thus using them in an experiment would result in a larger estimate of critical scaling. Since our explicit goal was to estimate the smallest critical scaling window, these stimuli would not bring us closer to our goal. As the reviewer points out, these metamers were “informative” in the sense that they informed our experimental design, but they are “uninformative” (relative to white noise seeds) for estimating the critical scaling.

      We have updated our description in the discussion starting on page 19, line 449, and included a new appendix to demonstrate this point (appendix 2 on page 35).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Typo: p. 19, l. 439: ’Why does asymptotic performance, but not critical scaling, depends on image content?’": remove ’s’ from ’depends’.

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      Recommendations: State that the lack of eye tracking to control stimulus presentation is a limitation of the results presented.

      Remove the claim that pink noise or filtered white noise seeds would be uninformative, and mention the fact that the authors in fact experimented with pink noise seeds in an early version of the experiments (which was surely informative to the experimental setup as presented here).

      Addressed as described above.

      References

      Freeman J, Simoncelli EP. Metamers of the ventral stream. Nature Neuroscience. 2011 aug; 14(9):1195–1201. doi: 10.1038/nn.2889.

      MacAdam DL. Visual Sensitivities To Color Differences in Daylight*. Journal of the Optical Society of America. 1942 may; 32(5):247. http://dx.doi.org/10.1364/josa.32.000247, doi: 10.1364/josa.32.000247.

      Stockman A, Sharpe LT. The Spectral Sensitivities of the Middle- and Long-Wavelength-Sensitive Cones Derived From Measurements in Observers of Known Genotype. Vision Research. 2000 jun; 40(13):1711–1737. doi: 10.1016/s0042-6989(00)00021-3.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      We thank Reviewer 1 for this thoughtful summary of our work.

      Strengths:

      Overall, this is a scientifically solid paper. The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Weaknesses:

      My main concern is that the study's aim is a bit unclear. Modern benchmarking studies comparing physics-based docking with deep learning-based co-folding approaches (e.g., AF3, Boltz-2, Chai-1, and others) are increasingly expected to go beyond aggregate performance metrics.

      Indeed, we have gone into several examples of failures and successes for each of these methods. As we are not developing these methods ourselves, we also think this dataset will be a valuable contribution for improving them further.

      In addition to rigorous dataset construction, transparent methodology, and appropriate statistical evaluation, high-impact benchmarks typically provide actionable guidance on when each method class is most appropriate, reflecting their distinct inductive biases and practical constraints. Failure-mode analyses that link performance differences to protein flexibility, ligand chemistry, or binding-site characteristics are particularly valuable, as they move comparisons beyond "scoreboard" assessments toward mechanistic understanding.

      Right now, we do not observe meaningful trends that separate the failure modes for any individual method. This is covered in Supplementary Figures 6 and 7.

      While full biological validation is not expected, qualitative interpretation grounded in physical and biological principles strengthens conclusions. Providing reproducible workflows or reference pipelines is not mandatory, but it is increasingly viewed as a best practice because it facilitates adoption and helps contextualize results for practitioners.

      We note that our code is available (https://github.com/jongbin99/Cofolding/) and all structural data will be publicly accessible in the PDB alongside publication (we only held it back only for “blinding” during peer review to avoid contamination with any new deep learning methods).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Kim et al. evaluates the performance of three modern AI-based methods in predicting complex structures and binding affinities between proteins and chemical compounds. An honest 'prospective' evaluation is achieved by studying benchmark structures and chemical compounds that did not exist in the PDB at the time the AI structure prediction models (AlphaFold3, Chai-1, Boltz-2) were trained.

      Strengths:

      (1) The study addresses an important question in modern computational biology and drug discovery, and establishes the strengths and limitations of the three tools in solving various computational chemistry tasks, including compound pose prediction, active-inactive discrimination, and potency ranking.

      (2) The conclusions are based on examination of four separate targets and respective compound datasets, where for one of the targets, the authors also obtained numerous X-ray structures to serve as experimental answers for the binding pose prediction task.

      (3) The study reports relationships between structure prediction confidence, predicted energies (DOCK3.7), and affinity predictions (Boltz-2) with the geometric accuracy of compound pose prediction as well as the experimentally measured potency.

      (4) One of the key findings is the limited ability of co-folding methods to predict conformational rearrangements, which does not correlate with their ability to predict binding poses of the compounds inducing these rearrangements.

      (5) The findings could serve as useful guidelines for computational chemists in selecting appropriate software and scoring schemes for each task.

      We appreciate Reviewer 2’s summary of the novelty of the dataset and analysis.

      Weaknesses:

      While I consider this a solid study, several aspects would need to be addressed to make it really strong:

      (1) DOCK3.7 docking and scoring experiments were performed using one experimental structure of Mac1, selected from dozens of structures based on a criterion that is not sufficiently well justified. For sigma2 receptor, dopamine D4 receptor, and AmpC β-lactamase, it is not clear which structures or models were selected for docking at all. It is well known that geometry predictions, scoring, and active-inactive ROC AUCs are all strongly influenced by the selected structure. It would be important to attempt Mac1 docking using all available experimental Mac1 structures, or at least against representative structures in various conformations; it would also be quite insightful to compare results to docking of the same compound sets to AF3, Boltz-2 and Chai-1 predicted structures of Mac1. Same goes for the docking studies of sigma2, D4, and AmpC β-lactamase.

      In any program, a decision has to be made as to which template will be used for docking, we justified the choice in the methods:

      “We used this structure because the inhibitor (Z5014193706) was the most potent molecule with a structure determined around the same time as the ligands in this dataset were tested.”

      We stand by this as a reasonable assumption. Similarly, for sigma2, D4, and AmpC β-lactamase, the template was chosen in the respective papers:

      a) The σ2 receptor bound to cholesterol (PDB ID: 7MFI) was used in the docking calculations.

      - This structure was determined in the paper, the first structure of sigma2 and therefore a worthy template

      b) The D4 receptor campaign used PDB 5WIU

      - This was one of two D4 structures available and chosen because it was not bound to sodium

      c) For AmpC, the campaign used the structure in the Protein Data Bank (PDB) 1L2S

      - This maximizes comparisons to other docking studies that used the same receptor template.

      The major goal of this study is to compare different methods under reasonable (but perhaps as the reviewer points out, not optimal) conditions, not to optimize docking score.

      (2) For binding affinity predictions, as a control, authors should consider compound co-folding with an unrelated protein, or even with a pseudo-peptide that consists of a few random single amino acids - this would provide an honest baseline for such predictions.

      This suggestion would be valuable for understanding the performance for these methods from the perspective of ligand specificity (a valuable, but separate, goal). Surely this will generate some number or some prediction - but what would this baseline mean and how would it be relevant for drug discovery? Therefore, we do not think this suggestion is relevant for the issues being investigated in this manuscript.

      (3) ROC curves Figure 3 and elsewhere should be shown, and AUCs quantified/reported on a log or square-root scaled x-axis, to emphasize early enrichment, which is the area of practical significance for these predictions. For example, Figure 3A currently suggests that the pose prediction performance of AF3 exceeds that of Boltz-2 whereas the early enrichment is clearly better for Boltz-2.

      We agree with this, and added a semi-logAUC plot for Figure 3A. For Figure 5, we also generated a semi-logAUC plot to see early ligand enrichment clearly, added as Supplementary Figure 11. We added the text:

      “Considering its early enrichment performance, Boltz-2 Ligand ipTM was the strongest predictor of pose accuracy based on normalized logAUC (20.5% above random, Fig. 3a). In contrast, although Boltz-2 pIC50 showed poor overall discrimination, it overestimated its ability to enrich true positive poses at low false positive rates, despite having a weak early enrichment behavior”

      (4) 'Trained set' in figures and text should probably be 'training set'? Or otherwise explain this new term the first time it is introduced.

      Thank you for pointing out this for clarification. ‘Training set’ is the correct word, and we made changes appropriately across all figures and texts.

      (5) Figure 1 illustrates a projection onto the first two principal components of a space that apparently had only one (scalar) metric for each compound pair (% maximum common substructure or Tanimoto coefficient); the authors need to better explain the principle behind this analysis and visualization.

      This suggestion is valuable, since we often use PCA to reduce dimensionality for more complex features. For clarification, we actually have a full pairwise similarity matrix for all tested Mac1 compounds based on each of Tc and MCS%. PCA for each MCS% and Tc is a representation of each pairwise similarity matrix. We also made a change in Figure 1 caption to make this point clearer:

      “projection of compounds represented by their full pairwise similarity vectors (by ECFP-4 Tc and MCS%)”

      Reviewer #3 (Public review):

      Summary:

      This study's core conclusions are well-supported by data. It is shown that co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false-positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      (1) Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods.

      (2) Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment.

      (3) The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      We thank Reviewer 3 for pointing out the unprecedented and comprehensive nature of our study

      Weaknesses:

      (1) Limited generalization to diverse protein families (e.g., no ion channels/transporters).

      We agree - we have not explored the entire proteome and these are important target classes that will surely be investigated by future studies. We focused on targets here where we had large number of X-ray crystal structures (Mac1) and affinity/inhibition measurements from docking (the other three targets).

      (2) Ambiguity in the mechanism underlying co-folding's failure to predict rare conformational changes.

      Again, we agree. We are not the developers of these methods. We observe that these methods do not predict conformational changes with high fidelity and this weakness is an area that co-folding methods will surely prioritize in the future.

      (3) Virtual screen comparison is unbalanced (docking-prioritized hit lists bias results).

      We acknowledge this in the results: “An important caveat is that the hit-lists were composed of molecules prioritized by docking in the first place, giving it an advantage on these particular sets.” and discussion: “Finally, comparing co-folding to docking based on hit-lists themselves selected by docking is arguably unfair to co-folding. Counter-balancing this is the inclusion, in each of the three hit lists, of molecules that had mediocre and poor docking scores intentionally selected to test the correlation between docking score and hit-rate. Here too, the correlation between co-folding score and likelihood to bind, what we sometimes call a “dock-response-curve” was no better than docking’s, often worse (SFig.11).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are suggestions for revisions:

      (1) The writing is at times obtuse and hard to follow.

      This happens sometimes when multiple authors are writing together. We apologize and are happy to respond to specific areas that can be streamlined to be easier to follow.

      (2) In the Results section, "A set of 557 previously unreported Mac1 ligand complexes", the authors have compared the ligand poses across different metrics such as Tc - a standard, highly effective method in chemo-informatics and MCS (maximum common substructures); these are standard metrics for quantifying the structural similarity between pairs of small molecules. This part of the analysis checks whether this is memorization; it is critical to compare the two metrics, but it is not sufficient to draw a conclusion.

      Thank you for pointing out about the structural similarity of molecules co-folded to those present in the training set (resolved as Mac1 complexes and deposited in PDB before training dates). We have conducted an analysis where we do a pairwise similarity comparison for all ligands present in the PDB (regardless of the target), by both Tc and MCS, and overlay the cluster of ligands we tested (Mac1, AmpC, sigma2, D4). This should show where our tested benchmark datasets lie in the chemical space covered in the entire PDB. Each cluster (around 500 to 1300 compounds per target system) is overlaid on the cluster of all ligands deposited in PDB (over 50,000 compounds), and each cluster was relatively diverse by both Tc and MCS.

      (3) In the "Co folding can accurately reproduce poses of ligands dissimilar to those trained." Subsection under Results, the authors' conclusions are hard to follow; they state that the co-folding models often mispredict or miss the alternative conformation, but they also predict poses that are distinct from the training set. What does that imply?

      Our interpretation is actually a somewhat unsettling one: co-folding gets the ligand pose right even when it gets the protein wrong, and even when the ligand is novel. This suggests the models may be anchoring on conserved pharmacophoric interactions (like the adenosine-mimicking purine scaffold) rather than truly modeling the physics of the full complex. We added to the results section:

      This result suggests that co-folding reliably recapitulates dominant ligand-binding interactions even in the absence of accurate protein conformational modeling, providing further support to the idea that they are learning specific interaction patterns rather than a deeper physics-based representation (Masters et al. 2025).

      (4) The Discussion section connects the results and conclusions, but it can be challenging to grasp the study's overall message.

      We think the final paragraph hits on three major points:

      - Co-folding accurately predicts ligand poses for known binders, but fails to capture conformational changes

      - Co-folding does not reliably distinguish true binders from false positives in virtual screening hit lists

      - Docking and co-folding are complementary rather than competing tools

      (5) The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment. The value of the paper would be further enhanced by explaining how it differs from seemingly similar results reported in other studies, including the one cited in this manuscript (see https://www.biorxiv.org/content/10.64898/2025.12.04.692352v1).

      The Mac1 results are completely unique. However, the docking datasets are exactly the same as those analyzed in the Menon et al manuscript. We don’t think our results differs from conclusions of the Menon et al manuscript as we wrote: These observations are supported by a fascinating study on some of the same ligand sets as investigated here, using AlphaFold3, reaching similar conclusions (Menon et al. 2025).

      Reviewer #3 (Recommendations for the authors):

      (1) Expand target diversity to include ion channels, transporters, etc., beyond enzymes and GPCRs.

      (2) Investigate the cause of co-folding's failure in predicting rare conformational changes (e.g., adjust sampling, MSA inputs, or add experimental constraints).

      (3) Mitigate docking bias in virtual screens (e.g., re-analyze unbiased compound libraries).

      We addressed these three points in the public review above

      (4) Test Boltz-2's affinity predictions without linear calibration and compare with FEP.

      The data without linear calibration are included in the manuscript. Comparing such a large number of compounds with FEP is currently beyond our capabilities.

      (5) Conduct proof-of-concept to test co-folding-docking integration for better hit rates.

      We think this is well beyond the scope of this manuscript - but look forward to testing this idea in the future.

      We also got one community review that we respond to below:

      Summary

      This manuscript evaluates the performance of co-folding models when tasked with 1) the recapitulation of a large number of experimentally determined co-crystal structures of Mac1 with a series of Mac1 ligands and 2) the rescoring of hits to identify false positives originally derived from a set of large docking-based virtual screens. The evaluation leverages a dataset of crystal structures and affinity data from high-throughput crystallographic and biophysical screens, respectively. These data uniquely enable this report to focus on the ability of co-folding models to handle ligands, resulting in an analysis that is particularly timely given the wide adoption of co-folding models and the relative scarcity of such ligand-focused benchmarks among existing evaluations, which have primarily focused on protein structure prediction or binder design.

      Thank you for this thoughtful summary of our work

      Feedback

      The experiments and analyses in the manuscript are well thought-out and do not have any significant issues. There are a few high-level points that may improve the clarity and completeness of the results. Importantly, none of the suggested additional experiments will affect the conclusions of the paper, but rather help provide additional context for the results:

      The first section presents an exciting opportunity to frame the Mac1 ligands against ligands in the PDB more broadly. It would be informative to assess whether chemotypes that are easier or harder to predict accurately and confidently are over- or under-represented in the PDB as a whole. Note that this is not a recommendation that new scaffold similarity metrics be incorporated into the analysis, but rather that analyses similar to those already performed in the manuscript are performed using all ligands in the PDB. For example, PCA-based analyses similar to those in Fig. 1c could be used to examine Mac1 ligands in the context of all PDB ligands enabling questions such as whether similarity to a nearest PDB neighbor, cluster size in a Tc/MCS PCA space, or other frequency-based measures show any relationship with prediction vs. crystal structure RMSD. Such analyses could provide additional insight into how effectively models leverage ligand information present in the PDB overall, as opposed to biases arising specifically from scaffolds represented in Mac1 structures in the PDB, which are already well covered in the manuscript. The conclusion that Tc/MCS do not correlate with the ligand RMSDs for the ligands already associated with the Mac1 is well supported, and presumably suggests that a correlation would not exist against the backdrop of the PDB, but it would be interesting to see the data using analyses similar to those already done in the manuscript nonetheless.

      We are adding new figures in SFig.1 that consider how different clusters of ligands tested for our co-folding analysis are distributed across the chemical space in PDB. This is done by making a similarity comparison between every ligand in PDB and those tested in our analysis by Tc and MCS%, then plotting in PCA space for each metric. We are excited to see that each dataset covers a wide scope in PCA space, but at the same time, there are unexplored areas in the chemical space of PDB by co-folding.

      Similarly, even though the four proteins used in this manuscript are not themselves the primary focus of the analysis, it would be valuable to perform a high-level assessment of the precedent for each protein in the PDB (beyond the count of liganded structures in Table S6), either in protein sequence space (e.g., MSAs) or structural space (e.g., FoldSeek). An analysis like this would provide important context about whether any of the proteins in the study have close homologs with liganded structures in the PDB, or are generally overrepresented in the PDB. The fact that the AUC for L-pLDDT for AmpC is higher than σ2 and D4, for example, is notable given the relative abundance of liganded AmpC structures in the PDB (this raises potentially interesting questions related to where DOCK3.7 and AF3 actually place the ligands, given the orthosteric β-lactam binding pocket in AmpC, although this is outside of the scope of this manuscript).

      High-level assessment of the precedent for each protein in the PDB will definitely help to understand if proteins we used have close homologs with liganded structures in the PDB. Our Supplementary Table 6 covers the extent to which these liganded structures were available by cutoff dates for AF3, Chai-1 and Boltz-2. AmpC had more homologs than sigma2 and D4, and this may explain a better AUC for AF3 L-pLDDT specifically for this target.

      A discussion of the affinity probability results (`affinity_probability_binary`) from Boltz-2 is likely warranted in the second section in addition to the pIC50s that are already reported (`affinity_pred_value`). The former seems like it would be more applicable for section 2 of the manuscript, but both warrant inclusion—they should both be calculated by default when the affinity pipeline in Boltz-2 is turned on, so it wouldn't involve any more inference.

      As boltz-2 affinity module outputs both affinity probability binary output and affinity predicted value, we kept track of both metrics. So we tried re-ranking hit lists using both metrics. Where boltz-2 performed better (Sigma2, D4), binary probability values were more representative as a metric to differentiate true actives from non-binders. This was more clear in semi-logarithmic ROC plots. However, in AmpC, both Boltz-2 scoring metrics performed similarly. Such inconsistency in trend made it difficult to draw conclusions.

      Minor points

      A more detailed description of the experimental methods used to generate the ground-truth data in the introduction (even though these have been explained in prior works) would help orient the reader early on, and ground the benchmarking aspect of the story. In general, the abstract and introduction would benefit from a more cohesive through-line to tie the two complementary but orthogonal sections of the paper together.

      We will include a more thorough description alongside the PDB depositions. As for the two sections, we have tried to tie them together from the perspective of drug discovery workflows…

      The cutoffs in the "Co-folding can accurately reproduce..." section shift between 2.5 Å (from the ligand center of mass) and 2.0 Å. Is there a reason for this? Along similar lines, mentioning cutoffs for true positives/negatives when introducing the ROC analyses later on in the Mac1 section seems unnecessary since no cutoff should be necessary here.

      We used 2.5A distance to COM to just get at “broadly the correct binding site” for fast filtering and 2.0A RMSD because that is the broadly accepted standard in the field for “relatively correct binding pose”.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This study describes motor cortical activity patterns during food handling in mice, investigating whether the hand/s used is reflected in distinct neural activity. The experiments focus on forelimb M1 and M2 (fM1, fM2) and an oral-manual region LOM. The main findings are that fM1 and fM2 have largely similar relationships with forelimb control, and LOM neurons are more broadly tuned. These conclusions are reached using a variety of analyses spanning straightforward firing rate analyses, selectivity metrics, PCA, and GLM decoding methods to assess tuning generalizability. The study's significance is strengthened by including analyses of bimanual control, and in this sphere, there are descriptive data and analyses that aficionados of cortical control of dexterous behaviors will find useful. The use of unimanual control is useful as a point of comparison here, but less novel overall. There are a number of places where the descriptions of what is being analyzed, what is being concluded, and data reporting should be strengthened and clarified. Additionally, the study could be greatly improved by consolidating figures and the analyses shown, since many are redundant. Many of the analyses need clearer reporting of means and effect sizes in the text, rather than just statistical outcomes. Overall, at this juncture, the study presents analyses of a unique dataset that may seed future investigations of mechanisms of bimanual coordination.

      Strengths:

      There are relatively few studies that compare neural activity across bimanual and unimanual control. This study uses a naturalistic food handling task to explore neural relationships to forelimb kinematics under these conditions. The uniqueness of the task and analysis target is a strength of the study.

      The authors remain fairly conservative and make few strong claims in the study, which may be warranted given the diversity of tuning profiles they observed.

      Weaknesses:

      There are a number of statistical tests that were accompanied by too little information to critically evaluate. Means and effect sizes needed to be better reported; some details of analyses were difficult to parse, making the strength of the conclusions difficult to evaluate.

      We will improve these aspects of the presentation and reporting of statistical analyses in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Barrett et al. examine how neural activity in the mouse motor cortex varies when a movement is performed with the ipsilateral or contralateral forelimb. First, they train animals to grasp and manipulate a pellet of food with either the left forepaw, the right forepaw, or both. Next, they measure activity in the primary and secondary forelimb motor areas (fl-M1 and fl-M2) and in the classical tongue-jaw area (tj-M1 / LOM). While responses in the forelimb areas are diverse, with some neurons preferring ipsilateral or bilateral movements, a plurality of cells prefer the contralateral limb. In LOM, by contrast, little limb selectivity is observed. At the neural population level, structure is preserved across conditions in LOM, but not in the forelimb areas. Finally, paw position can be decoded from activity in all three areas, and the LOM decoder generalized across limbs.

      Strengths:

      While previous studies in macaques have compared motor cortical activity during movement (and perturbation) of the contralateral and ipsilateral arms, no analogous work has been undertaken in rodents. This paper closes this knowledge gap by showing, for the first time, moderate-to-strong lateralization in the forelimb motor cortical areas of mice transporting grasped food pellets to the mouth, and a relative absence of lateralization in the classical tongue-jaw area. On the whole, I think this is a solid paper that reports novel observations of interest to the motor systems community.

      Weaknesses:

      The central question posed is whether cortical activity depends on the effector(s) used (ipsi forelimb, contra forelimb, or both). The corresponding hypotheses (Figure 1) are somewhat coarse-grained and are not mutually exclusive. One might expect to see condition-independent, limb-selective, and uni-/bimanual-selective signals in motor cortex (though their magnitudes could differ substantially), and to find these signals intermingled at the level of single neurons. The authors may wish to consider setting up a more focused question. For example, can bimanual responses be explained as a sum of the unimanual responses from the left and right limbs?

      In the revised manuscript, we will clarify that the possibilities illustrated in figure 1 are not intended as mutually exclusive. We indeed find all of these signals intermingled. In terms of a more focused question amenable to hypothesis testing, this can be expressed as: for each dimension (laterality vs manuality) are the activity patterns in each area closer to those predicted by invariance or dependence, as compared to the other areas? The various statistical analyses in the paper all essentially boil down to testing this question. Broadly, the answer is yes: we see activity closer to the invariant prediction LOM, and activity closer to the dependent prediction in fl-M1 and fl-M2. Testing whether bimanual activity can be explained as a simple linear sum of the left and right unimanual might provide additional insight into this question, and this analysis will be presented in the revised manuscript.

      In the area usually identified as tongue-jaw motor cortex (here referred to as LOM), unit and population activity look quite similar for ipsilateral, contralateral, and bilateral forelimb reaches. The most parsimonious explanation is that the activity is related mostly to mouth and tongue movements, rather than limb movements. Systematic mapping studies with microstimulation in the rat (Neafsy et al., Brain Res. Rev. 1986) and optogenetic stimulation in the mouse (Mayrhofer et al., Neuron 2019) tend to support the idea that tjM1/LOM is specialized for control of the tongue and mouth. Thus, I'm not entirely convinced that it "encodes ingestion-related forelimb parameters necessary for oromanual coordination." The authors could say more about this issue: what specific limb-related parameters do they think are encoded, why would these parameters be effector-independent, what evidence for this encoding is presented here, and how can limb- and mouth-related components be distinguished? The problem could potentially be addressed experimentally, as well, by delivering food pellets directly to the mouth while preventing manipulation with the paws, but this experiment isn't strictly necessary.

      The issue of orofacial movement confounds is an important one that we made a point of addressing in the discussion. There are three main points that we believe cast doubt on this as the most likely explanation for the effector-invariant representation in LOM.

      First, while we do not disagree that LOM has an important role in tongue and jaw control, there is plenty of evidence from mapping and behavioural studies (which we cite in the introduction) that it also plays a role in forelimb motor control as well.

      Second, while we cannot observe all orofacial movements, we have previously shown that the jaw is less active when the hands and LOM are most active (Barrett et al., 2024). Conversely, LOM firing is much lower during chewing, when the tongue and jaw are very active.

      Finally, the correlation between LOM firing and forelimb movements is not merely a coarse-grained one on the timescale of active manipulation vs passive holding phases. LOM firing closely tracks the position of the forelimb(s) on fast timescales and with near-zero lag, giving better decoding than from fl-M1 or fl-M2, as we have shown here and previously (Barrett et al., 2022). If we assume that LOM only encodes orofacial movements, then this result implies that orofacial movements correlate with forelimb position better than fl-M1 or fl-M2 firing correlates with forelimb position.

      The revised manuscript will include an expanded discussion to clarify these and related points.

      Because the corticospinal tract is strongly lateralized, cortical activity presumably has a smaller effect on ipsilateral than contralateral motor output. Somatosensory feedback should also be relatively lateralized for the forelimb areas. The authors could say a bit more about this issue and how it relates to their data and conclusions in the Discussion.

      We will discuss this in the revised manuscript.

      An important limitation of the behavioral task is that it involves only a single stereotyped movement for each limb, instead of multiple directions, speeds, or loads. This issue and its consequences for the analyses (especially those in Figures 7-10) and conclusions could be discussed.

      This limitation applies to the analyses relating to the transport-to-mouth movement (Figures 3-7). The population correlation structure and decoding analyses (Figures 8-10) consider activity throughout the full duration of food handling, which involves a much greater variety of movements (Barrett et al., 2020). Indeed, this was a major motivation for including these analyses. The revised manuscript will clarify this point.

      Reviewer #3 (Public review):

      Summary:

      Barrett et al. compare the responses of different parts of the mouse primary and secondary motor cortex in the context of a task where the animals manipulate and eat food using either or both hands. They find that roughly half the activity is conserved when reaching with one hand vs. the other hand, or with both. Similarity of activity was somewhat higher in the "lateral oral and manual" (LOM) part of the motor cortex, consistent with notions of a more generalized oromanual function there.

      Strengths:

      This work aims at addressing two worthwhile questions in a mouse model of motor control: (1) what specializations do we have for controlling feeding movements, and (2) how are the arms and hands coordinated with one another? The authors develop a simple but innovative apparatus to block either hand during food handling, track the behavior at high temporal fidelity, and record a sizable neural dataset. The analyses come from numerous angles to take good advantage of the data, and succeed in showing multiple lines of evidence for greater invariance in LOM than in the forelimb parts of M1 and M2.

      Weaknesses:

      There are several limitations of the current study. Most importantly, the behavior presents an inherent challenge: there is only one type of movement for each of the three conditions (contra hand, ipsi hand, and bimanual).

      See our response to reviewer #2 above regarding the variety of movements. We agree that this a limitation for the unit-level and PCA analyses, but one that is alleviated by the population correlation and decoding analyses, which relate to complex ongoing movements.

      This is entirely reasonable from the perspective that this is the ethological behavior when feeding, but it limits what analyses are possible. In particular, it precludes disentangling the neural relationship with many correlated aspects of behavior, and limits identifying population-level features of the neural activity meaningfully. This means that there are a number of alternative possible sources of the neuron-level area differences found here, and the population-level features may not be reliable.

      Important behavioural confounds include orofacial movements, non-specific movement initiation signals, and arousal. Orofacial movements we have discussed above in our response to Reviewer 2. Movement initiation signals would likely be transient and well-timed to movement onset, hence this may explain some of the effector-independent activity in fl-M1 and fl-M2 (consistent with e.g. (Kaufman et al., 2016)). However, we do not believe this to be the case in LOM as its activity is delayed and sustained relative to movement initiation. Regarding arousal, the mouse is actively engaged in consuming the food even when the hands are stationary, so there is no a priori reason to believe that arousal varies rapidly during the behaviour. Consistent with this, measurements of noradrenergic activity in the locus coeruleus during consumption suggest that arousal varies on slow timescales, on the order of seconds to tens of seconds (Sciolino et al., 2022). Such slow variation in arousal would not explain the rapid but condition-invariant changes in firing in any of the cortical areas studied here. The updated manuscript will include more detailed discussion of these points.

      Second, the behavior tracking was used at a relatively coarse level, and thus the relationships to various behavioral variables were left less distinguishable than they might have been.

      Behavior tracking was performed with kilohertz temporal resolution and submillimeter spatial resolution.

      Finally, there may be an issue with the coordinates of what is being called forelimb M1 here, which may include some hindlimb M1.

      Our recording coordinates are based on the territory of corticospinal neurons retrogradely labeled from C6 spinal cord, medial to any layer 4 labelling, as reported in our previous study (Yamawaki et al., 2021). Thus we are confident in calling this area forelimb M1.

      References:

      Barrett, J. M., Martin, M. E., Gao, M., Druzinsky, R. E., Miri, A., & Shepherd, G. M. G. (2024). Hand-jaw coordination as mice handle food is organized around intrinsic structure-function relationships. The Journal of Neuroscience, 44(42), e0856242024.

      Barrett, J. M., Martin, M. E., & Shepherd, G. M. G. (2022). Manipulation-specific cortical activity as mice handle food. Current Biology, 32(22), 4842-4853.e6.

      Barrett, J. M., Tapies, M. G. R., & Shepherd, G. M. G. (2020). Manual dexterity of mice during food-handling involves the thumb and a set of fast basic movements. PLOS ONE, 15(1), e0226774.

      Kaufman, M. T., Seely, J. S., Sussillo, D., Ryu, S. I., Shenoy, K. V., & Churchland, M. M. (2016). The Largest Response Component in the Motor Cortex Reflects Movement Timing but Not Movement Type. eNeuro, 3(4).

      Sciolino, N. R., Hsiang, M., Mazzone, C. M., Wilson, L. R., Plummer, N. W., Amin, J., Smith, K. G., McGee, C. A., Fry, S. A., Yang, C. X., Powell, J. M., Bruchas, M. R., Kravitz, A. V., Cushman, J. D., Krashes, M. J., Cui, G., & Jensen, P. (2022). Natural locus coeruleus dynamics during feeding. Science Advances, 8(33), eabn9134.

      Yamawaki, N., Raineri Tapies, M. G., Stults, A., Smith, G. A., & Shepherd, G. M. (2021). Circuit organization of the excitatory sensorimotor loop through hand/forelimb S1 and M1. eLife, 10, e66836.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study is methodologically solid and introduces a compelling regulatory model. However, several mechanistic aspects and interpretations require clarification or additional experimental support to strengthen the conclusions.

      Strengths:

      (1) The manuscript presents a compelling structural and biochemical analysis of human glutamine synthetase, offering novel insights into product-induced filamentation.

      (2) The combination of cryo-EM, mutational analysis, and molecular dynamics provides a multifaceted view of filament assembly and enzyme regulation.

      (3) The contrast between human and E. coli GS filamentation mechanisms highlights a potentially unique mode of metabolic feedback in higher organisms.

      Weaknesses:

      (1) The mechanism underlying spontaneous di-decamer formation in the absence of glutamine is insufficiently explored and lacks quantitative biophysical validation.

      (2) Claims of decamer-only behavior in mutants rely solely on negative-stain EM and are not supported by orthogonal solution-based methods.

      We thank the reviewer for the summary and noting of the strengths. We agree that the evolutionary divergence of metabolic feedback in GS homologs is a fruitful avenue for future studies. With regard to the weaknesses, the di-decamer in the absence of glutamine only forms under high (higher than physiological) concentrations of enzyme. Our primary evidence for the mutant behavior was the lack of crosslinking (Figure 1E), with supplementary support from the negative stain. In the revised version we will soften the language to say “reduced” rather than “did not support” filament formation.

      Reviewer #2 (Public review):

      The authors set out to resolve the high-resolution structure of a glutamine synthetase (GS) decamer using cryo-EM, investigate glutamine binding at the decamer interface, and validate structural observations through biochemical assays of ATP hydrolysis linked to enzyme activity. Their work sits at the intersection of structural and functional biology, aiming to bridge atomic-level details with biological mechanisms - a goal with clear relevance to researchers studying enzyme catalysis and metabolic regulation.

      Strengths and weaknesses of methods and results:

      A key strength of the study lies in its use of cryo-EM, a technique well-suited for resolving large, dynamic macromolecular complexes like the GS decamer. The reported resolutions (down to 2.15 Å) initially suggest the potential for detailed structural insights, such as side-chain interactions and ligand density. However, several methodological limitations significantly undermine the reliability of the results:

      (1) Cryo-EM data processing: The absence of critical details about B-factor sharpening - a standard step to enhance map interpretability - is a major concern. For high-resolution maps (<3 Å), sharpening is typically applied to resolve side-chain features, yet the submitted maps (e.g., those in Figures 1D, 2D, and supplementary figures) appear unprocessed, with density quality inconsistent with the claimed resolutions. This makes it difficult to evaluate whether observed features (e.g., glutamine binding) are genuine or artifacts of unsharpened data.

      (2) Modeling and density consistency: The structural models, particularly for glutamine binding at the decamer interface, do not align with the reported resolution. The maps shown in Figure 2D and Supplementary Figure S7 lack sufficient density to confidently place glutamine or even surrounding residues, conflicting with claims of 2.15 Å resolution. Additionally, fitting a non-symmetric ligand (glutamine) into a symmetry-refined map requires justification, as symmetry constraints may distort ligand placement.

      (3) Biochemical assay controls: While the enzyme activity assays aim to link structure to function, they lack essential controls (e.g., blank reactions without GS or substrates, substrate omission tests) to confirm that ATP hydrolysis is GS-dependent. The use of TCEP, a reducing agent, is also not paired with experiments to rule out unintended effects on the PK/LDH system, further limiting confidence in activity measurements.

      Achievement of aims and support for conclusions:

      The study falls short of convincingly achieving its goals. The claimed high-resolution structural details (e.g., side-chain densities, ligand binding) are not supported by the provided maps, which lack sharpening and show inconsistencies in density quality. Similarly, the biochemical data do not robustly validate the structural claims due to missing controls. As a result, the evidence is insufficient to confirm glutamine binding at the decamer interface or the functional relevance of the observed structural features.

      Likely impact and utility:

      If these methodological gaps are addressed, the work could make a meaningful contribution to the field. A well-resolved GS decamer structure would advance understanding of enzyme assembly and ligand recognition, while validated biochemical assays would strengthen the link between structure and function. Improved data processing and clearer reporting of validation steps would also make the structural data more reliable for the community, providing a resource for future studies on GS or related enzymes.

      We disagree with the reviewer’s overall assessment.

      With regard to sharpening and resolution: we examined sharpened maps and in a revised version will present additional supplementary figures showing these maps side by side. We note that the resolutions reported are global and that the most interesting features are, of course, in the periphery and subject to conformational and compositional heterogeneity. We will include supplementary figures of core side chain densities that are more like what are expected by the reviewer in the revision. With regard to modeling: the apo filament and turnover filament datasets were handled nearly identically. The additional density is therefore likely not artefactual to the symmetry operator - however, the lower resolution in this region noted by the reviewer is worthy of further exploration. The maps are public and we think this is the most plausible interpretation of the density, which we based primarily on the biochemical data and will include more speculation in the version.

      With regard to the biochemical controls: we point the reviewer to Figure S1, which shows that omission of ammonia or glutamate in the wild-type (tagless) system removes any coupling of the reactions. We will perform the additional controls to publication quality in the revised version along with the TCEP control. We note that the reducing agent is present across all experiments, ruling out an effect on any specific result. The inclusion of TCEP is also very standard in other published uses of the Coupled ATPase assay (e.g. PMID: 31778111 and PMID: 32483380 by our first author)

      Additional context:

      Cryo-EM has transformed structural biology by enabling high-resolution analysis of large complexes, but its success hinges on rigorous data processing and validation steps that are critical to ensuring reproducibility. The challenges highlighted here are not unique to this study; they reflect broader issues in the field where incomplete reporting of methods can obscure the reliability of results. By addressing these points, the authors would not only strengthen their current work but also set a positive example for transparent and rigorous structural biology research.

      All the data is public and the reviewer or anyone is free to reinterpret the maps and models - and we encourage that rather than just an interpretation of our static figures. In addition, we will upload the raw micrograph data for the apo filament and turnover filament datasets to EMPIAR prior to submitting the revision.

      Reviewer #3 (Public review):

      In this manuscript, the authors propose a product-dependent negative-feedback mechanism of human glutamine synthetase, whereby the product glutamine facilitates filament formation, leading to reduced catalytic specificity for ammonia. Using time-resolved cryo-EM, the authors demonstrate filament formation under product-rich conditions. Multiple high-quality structures, including decameric and di-decameric assemblies, were resolved under different biochemical states and combined with MD simulations, revealing that the conformational space of the active site loop is critical for the GS catalysis. The study also includes extensive steady-state kinetic assays, supporting the view that glutamine regulates GS assembly and its catalytic activity. Overall, this is a detailed and comprehensive study. However, I would advise that a few points be addressed and clarified.

      (1) In Figure 2D and Supplementary Figure 7, the extra density observed between the two decamers does not appear to have the defining features of a glutamine. A less defined density may be expected given the nature of the complex, but even though mutagenesis assays were performed to support this assignment, none of these results constitutes direct and conclusive evidence for glutamine binding at this site. I would thus suggest showing the density maps at multiple contour thresholds to allow readers to also better evaluate the various small molecules under turnover conditions that cannot be well fitted based on this density map, helping to provide a more balanced interpretation of the results.

      (2) On the same point regarding the density for the enzyme under turnover conditions, more details should be provided about the symmetry expansion and classification performed, and also show the approximate ratio of reconstructions that include this density. Did you try symmetry expansion followed by focused classification, especially on the interface region?

      (3) The interface between the two decamers of the model needs to be double-checked and reassigned, especially for the residues surrounding the fitted glutamine. For example, the side chain of the Lys residue shown in the attached figure is most likely modeled incorrectly.

      We thank the reviewer for the feedback. As noted above, we will include supplemental figures that show maps at multiple thresholds and sharpening schemes. We noted in the manuscript and above that our interpretation here is based on integrating biochemical evidence alongside the density and will make that even more clear in the revised manuscript. The filaments +/- the putative glutamine density were processed nearly identically, but we will attempt various schemes of focused classification/symmetry expansion in the revision as well. However, we point out that there is extensive averaging there that makes modeling a bit trickier than expected given the global resolution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Limitation to Di-decamer Formation:

      Could the authors clarify why hGS, when visualized by cryo-EM, predominantly forms di-decamers rather than extended filaments? Since glutamine bridges the two decameric rings, one would expect this to promote further polymerization. It remains unclear why longer filaments are not observed under the conditions used. Moreover, data in Figure 1B-1D are not mixed with Gln, and the mechanism of the apo-form filament formation is not clearly discussed. If K52 and C53 are in charge of filamentation, it is supposed to form a long filament, not a stack of 2 decamers. The kinetics of wild type and mutations, including K52A and C53A, are different. This confused me as K52 and C53 don't participate in the reactions of glutamine synthesis. Does data reduce catalytic efficiency with K52A or C53A mutation suggest that di-decameric GS exhibits a greater catalytic turnover rate than the pure decameric GS?

      We thank the reviewer for pointing this out and it was indeed filament length that was a point of curiosity during the study. Prior to preprint, we repeated the freezing conditions under identical turnover conditions and with high protein concentration as reported and indeed found much longer filaments - pointing to the capacity of the system to form larger complexes. However, we only captured a couple of ‘screening’ images and did not collect a full second dataset. Therefore, we hypothesize that filament length may be stochastic based on the subtle differences of individual grid vitrification based on the assumption that decamers within a filament can freely and quickly exchange. However, as shown in the time-resolved cryoEM experiment, the fraction of particles that are characterized as participating in a filament form (length-agnostic metric), does not appear to be subject to individual grid vitrification conditions but rather by experimental conditions (concentration of reaction-derived glutamine).

      The filament formation in apo state was not further explored because the concentrations required to achieve filament formation in this case were supraphysiological.

      We have edited the discussion to emphasize this point more clearly:

      “While the enzyme concentrations to achieve robust filamentation in the absence of glutamine are much higher than observed in cells, the protein concentrations used in our time-resolved cryoEM experiments where filamentation is correlated with accumulation of glutamine are within the range of intracellular GS concentrations in S. cerevisiae (Engel et al. 2025) and human cell lines (Wiśniewski et al. 2014).”

      Regarding the ability of apo-GS to form filaments - we identified that these residues were important based on their structural location at the decamer: decamer interface (Figure 1D; away from the active site as pointed out) and because their individual mutation to alanine attenuated the ability to form higher order filaments (Figure 1E). Therefore, these mutants were crucial controls in the steady-state kinetic experiments reported (Figure 2E). Here, we used exogenous glutamine to seed/stabilize filaments because we identified glutamine as serving this function and, crucially, because we did not observe glutamine occupancy in the active site under turnover conditions (which would suggest an orthosteric feedback inhibition mechanism; Supplementary Figure 10). Under these conditions, the wild-type, filament-competent protein displayed a ~3-fold K<sub>M, ammonia</sub> increase with glutamine addition compared to no glutamine, which when considered in the context of the greater E-305 flap conformational heterogeneity speaks to a model of allosteric feedback inhibition. Importantly, K52A and C53A show no difference in K<sub>M, ammonia</sub> between the glutamine and no glutamine condition, suggesting that attenuation of filament formation at this interface via mutation, eliminates the kinetic deficit. Therefore, K52A and C53A are not product-inhibited in the same manner as wild-type GS.

      We have clarified the discussion to emphasize this result:

      “Importantly, point mutations of the interfacial residues do not show a K<sub>M, ammonia</sub> defect in the presence of glutamine, indicative of the importance of the filament form for product feedback.”

      (2) Origin of Di-decamer Formation in the Absence of Glutamine:

      While the manuscript demonstrates glutamine-stabilized filamentation, the spontaneous formation of di-decamers under apo conditions is not mechanistically explained. The observation of 10-mer, 20-mer, and 40-mer species in Figure 1B should be validated against molecular weight standards or through SEC-MALS. The inference of higher-order oligomers based solely on migration is insufficient. Additional characterization (e.g., SEC-MALS, AUC, or mass photometry) would clarify whether these assemblies are biologically relevant or incidental.

      We thank the reviewer for pointing out the low precision of preparatory size exclusion chromatography assignments of GS molecular weight filament depicted in Figure 1. We have included calibration standards and assignment in Supplementary Figure 1 and updated Figure 1 to include the ambiguity of these assignments in panel B. It is important to note that GS has historically been underestimated in size via these methods and was originally assigned as an octamer for which there was previous consensus (PMID: 10708854). The low concentration requirements of mass photometry preclude its use for this purpose and we are not in a position to do SEC-MALS or AUC for this. Hopefully, the negative stain, crosslinking, and cryo-EM results are sufficient to indicate that we have correlated signals with the correct species!

      (3) Validation of Interface Mutants as Decamer-only Species:

      K52A and C53A mutants are used to disrupt di-decamer formation and are shown by negative-stain EM to exist as decamers. While supportive, this is qualitative. The inclusion of quantitative biophysical data (e.g., SEC-MALS or mass photometry) would more convincingly demonstrate that these mutants do not transiently assemble into higher-order oligomers. Furthermore, molecular measurements describing the spatial relationship of interface residues - such as the distance between K52 and E55 or between C53 residues of opposing decamers - would aid interpretation. The use of the term "adjacent" (line 530) is vague and should be made more precise.

      We thank the reviewer for the thoughtful comments. Beyond negative-stain EM we also performed a biochemical validation of the filament interface through bi-functional crosslinking based on the premise that the new filament interface, as defined by the apo-filament structure, presented new/unique pairs of nucleophilic amino acid R-groups in close proximity. We used Bis-sulfosuccinimidyl glutarate (BSG) or bis-maleimoethane (BMOE) to covalently link adjacent primary amines and sulfhydryls respectively (Figure 1E). This experiment defines two key principles of GS filament formation in the absence of glutamine:

      (1) It is concentration dependent. In the wild-type case there is a protein-dependent increase on crosslinking efficiency for both crosslinkers.

      (2) It is dependent on C53 and K52. Mutation of C53 or K52 significantly attenuated crosslinking efficiency.

      To make sure that these results are more prominent, we have now included a table of these intersubunit distances between epsilon amine groups of lysines and gamma sulfhydroxyl of cysteines groups based on the apo-filament structure and labeled this as either participating in the filament interface or not. Furthermore, in line with multiple reviewers comments, we have updated Figure 1D to include a sharpened representation of the map that shows strong side chain density for the amino acid side chains to further support these reported side chain distance measurements.

      We thank the reviewer for pointing out the low precision of the SEC chromatogram interpretation of Figure 1B and the figure has been amended to show filaments of variable length, instead of defined length. We also included a calibration curve to Supplemental Figure 1B and estimated molecular weights. While these estimates of size are lower than ground truth it is important to note that GS has historically displayed smaller than predicted molecular weights via size exclusion chromatography and analytical ultracentrifugation where initial characterization papers defined the oligomeric state as an octamer rather than decamer (PMID: 10708854). These studies and the present indicate potential adherence to resin and/or other factors about the shape of GS that lead to longer retention. Lastly, these SEC procedures were performed as a preparative step rather than for analytical purposes, so resolution was not the ultimate goal.

      (4) Terminology: "Scarless" hGS:

      The term "scarless human glutamine synthetase" is unconventional and potentially confusing. If it refers to the wild-type sequence lacking N- or C-terminal tags or mutations, I recommend using the term "native hGS" for clarity.

      We usually reserve “native” for proteins isolated from the original species and not recombinantly expressed (as here). So we will leave the term scarless in the document.

      (5) Helical Parameters of Filament Assembly:

      The manuscript states a ~26{degree sign} rotation (clockwise or counterclockwise?) between decamers in the filament, yet does not describe how this value was derived. Given that hGS filaments form helices, this parameter could be assessed via helical reconstruction. Is it possible the actual helical twist is ~30{degree sign}, implying 12 stacked decamers per full turn? Please elaborate on how the rotational angle was determined.

      Rotation was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was measured. Helical reconstruction was not pursued in this work owing to the typically short filaments observed in micrographs and the relative ease by which a 20-mer species could be selected via traditional 2D and 3D classification/reconstruction methods.

      We have added to the methods the following to better illustrate this measurement:

      “Decamer: Decamer rotation across the filament interface was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was depicted in Figure 1D.”

      (6) Time-Resolved Cryo-EM and Filament Growth:

      The use of time-resolved cryo-EM is innovative; however, the accessible timescales are relatively short. I suggest complementing this approach with techniques such as dynamic light scattering (DLS) or mass photometry, which allow extended real-time monitoring of filament assembly over longer durations (e.g., hours). These methods can also provide higher temporal resolution and particle size distributions.

      We thank the reviewer for this suggestion and agree that understanding the kinetics of filament formation is critical. While DLS and mass photometry are excellent for monitoring assembly over hours, our data indicates that GS filament formation occurs on a much faster timescale.

      As shown in Figure 2C, when we added ATP and Glutamine directly to GS and vitrified the sample after only 5 minutes, the majority of particles had already formed filaments, indicating that the interaction had reached saturation. This contrasts with our time-resolved experiment, where the kinetics of filament formation were likely rate-limited by the enzymatic generation of glutamine rather than the assembly process itself.

      Consequently, we anticipate that filament assembly occurs on the order of seconds or less—a timescale we interpret as a necessary prerequisite for a rapid and effective cellular feedback mechanism. Therefore, we believe the current cryo-EM data accurately captures the biologically relevant window of assembly.

      (7) Cryo-EM Symmetry Imposition and Loop Flexibility:

      The use of D5 symmetry in cryo-EM reconstructions may obscure conformational heterogeneity in flexible elements, such as the E305 loop. Since the authors used MD simulations to characterize loop dynamics, it would strengthen the study to also perform symmetry expansion followed by non-uniform refinement and alignment-free 3D classification of individual subunits. This could provide experimental validation of the proposed conformational variability.

      We agree with the reviewer that symmetry enforcement can mask conformational heterogeneity, particularly for flexible elements like the E305 loop (the E-flap). To address this, we followed the reviewer’s suggestion and performed symmetry expansion on both the turnover decamer and filament consensus maps. This was followed by focused, alignment-free 3D classification on the asymmetric unit containing the E-flap.

      Our analysis revealed a clear distinction: while 4 out of 12 turnover decamer classes showed partial density for the E-flap (class 3, 7, 8, and 12)—consistent with the flexibility observed in our MD simulations—none of the turnover filament classes demonstrated similar density. To ensure a direct comparison, we utilized a C5-expanded turnover decamer map to maintain an identical asymmetric unit to the D5-expanded turnover filament map. We note here that the turnover decamer consensus volume is different from the deposited map for which no symmetry was applied.

      Despite the different initial symmetries (D5 for filaments vs. C5 for decamers), we utilized a C5-expanded decamer map to maintain an identical asymmetric unit. For transparency, we have uploaded this C5-refined consensus map and all resulting 3D classification maps to Zenodo. We agree that the text is now strengthened given this result and we have added the following to the main text:

      “The differential loop density between turnover-decamer and turnover-filament species was further supported by 3D classification of symmetry-expanded particles, which recovered partial E305-loop density in 4/12 turnover-decamer classes (C5; Supplemental Figure 18) compared to 0/12 classes for the turnover-filament (D5; Supplemental Figure 19).”

      (8) Crosslinking Gel Analysis (Figure 1E):

      The SDS-PAGE gels shown in Figure 1E have molecular weight ladders cropped. For proper interpretation, please include full ladders with size markers and labels. In addition, clarify whether the crosslinked samples were denatured in reducing buffer. Crosslinking efficiency and specificity using BMOE or BSG require verification under reducing conditions (e.g., DTT, β-mercaptoethanol, or TCEP) to confirm covalent linkage between decamers.

      We have added in Figure 1E molecular weight markers estimates to aid in gel interpretation and have included the uncropped gels in Supplementary Figure 1D that contain the full MW ladder. The methods were clarified to indicate that reducing reagent was used in both the crosslinking reaction and all SDS-PAGE samples.

      “Protein samples were diluted to concentrations noted in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.1 mM TCEP) and, reacted with crosslinker to a final concentration of 0.5 mM for 10 mins at room temperature followed by quench in 5X SDS-PAGE sample buffer (225 mM Tris pH 6.8, 50% glycerol, 0.05 % SDS, 0.2 mg/mL bromophenol blue, 1M DTT) supplemented with 100 mM of either NH<sub>4</sub>Cl (to quench BSG reactions only) or DTT (to quench BMOE). Protein concentrations were normalized after quench prior to SDS-PAGE analysis.”

      (9) Missing Reference for NADH-Coupled Assay:

      Line 678-679 refers to an NADH-coupled assay described "previously" without citing a source. Please provide a proper reference to ensure reproducibility.

      The original paper describing the implementation of a coupled-assay to measure ADP production from glutamine synthetase was written by Bennett Shapiro and Eric Stadtman in 1970 and has been included. We will note that the conditions of this assay have been much improved since this time with better buffers, commercially available reagents of combined lactate dehydrogenase and pyruvate kinase, and modern plate readers. We added the following reference:

      “Shapiro, B.M. and Stadtman, E.R., 1970. [130] Glutamine synthetase (Escherichia coli). In Methods in enzymology (Vol. 17, pp. 910-922). Academic Press.”

      (10) Unclear Description of the NADH Assay:

      The stability of NADH is influenced by pH and light exposure. Please specify the pH range used in the assay and whether precautions (e.g., light shielding) were taken. NADH autoxidation at high pH or degradation at low pH could impact assay reliability and should be addressed in the Methods section.

      For clarity and transparency the following text was added to the Methods section.

      “Stocks of ATP, NADH, and phosphoenolpyruvate were made in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.5 mM TCEP) and the pH was adjusted until it reached 7.5 on ice prior to aliquoting, flash freezing, and storage at -80°C in the dark. NADH was only exposed to light upon thawing and assay set-up and no appreciable change in absorbance of control experiments were noted.”

      (11) Ligand Density in Figure 2 and Supplementary Figure 7:

      The density attributed to glutamine, ADP, and phosphate appears broader than expected. Please include cross-correlation (CC) values, estimated occupancies, and Q-factors for ligand fitting. Varying the contour level to assess density consistency would clarify whether the observed volume represents multiple conformations, partial occupancy, or overfitting. A similar concern applies to the cysteine sidechain density.

      We have updated Supplemental Figures to include:

      (1) Globally refined map in comparison to locally refined map where both are sharpened per previous feedback.

      (2) Ligand placement now also include Q-scores and CC values

      We did not include multiple contour levels because these are included in the resolution representative Supplemental Figure and because alternative contours do not influence Q-scores. From this analysis it is apparent that phosphate and ADP are both worse fits to the density.

      Moreover, we have now included Supplementary Table 2 that includes all ligand validation statistics for the reader to evaluate the range of B-factor, CC values, and Q-scores for all ligands in all models.

      (12) Missing Ligand B-factors in Supplementary Table 1:

      The ligand refinement statistics in Supplementary Table 1 are incomplete. Please include B-factors and occupancy values for all ligands.

      We have updated the PDB depositions to include B-factors in .cif files that are now available. We have also included Supplementary Table 2 in the manuscript detailing the ligand statistics for all models including cross-correlation, Qscore, and Bfactor.

      (13) Style and Formatting Issues: format consistently throughout.

      (a) Kinetic Parameters: Please follow the IUPAC and IUBMB-recommended formatting:

      kcat should be italic with subscript.

      KM should be italic K with upright M.

      Use lowercase s<sup>-1</sup>, not uppercase S<sup>-1</sup>.

      Refer to:

      IUBMB enzyme nomenclature guidelines https://iubmb.org/wp-content/uploads/2021/01/Current_IUBMB_recommendations_on_enzyme_nome nclature.pdf

      IUPAC Green Book https://publications.iupac.org/books/gbook/green_book_2ed.pdf

      We have corrected the abbreviations according to the reviewers recommendations.

      (b) Inconsistent Terminology and Typography:

      cryo-EM vs. cryoEM are used inconsistently - standardize throughout.

      FSC 0.143 appears with and without subscript formatting-please unify.

      Line 123: "X-ray" should be capitalized.

      Line 266: CryoEM should cryoEM, lowercase "c"

      Line 571: "100 μg ml<sup>-1</sup>"-use superscript minus; ensure consistency with "mg ml<sup>-1</sup>".

      Lines 605, 606, 626, 628: MgCl<sub>2</sub>-ensure the <sub>2</sub> is subscripted throughout.

      Temperature units (lines 607, 630, 640): Write as "4 {degree sign}C" instead of "4C".

      Microliters (lines 660, 681, 697): Replace "uL" with "μL".

      Line 797: Use superscripts: K<sup>+</sup>, Cl<sup>-</sup>.

      Line 862: CO<sub>2</sub> should appear with subscript.

      We have made all terminology and typography consistent throughout based on these suggestions.

      Reviewer #2 (Recommendations for the authors):

      To strengthen the manuscript and address the methodological and interpretational gaps identified, we recommend the following revisions and additions:

      (1) Data processing and cryo-EM map quality

      (a) B-factor sharpening: Reprocess all cryo-EM maps using standardized B-factor sharpening workflows (e.g., the autoSharpen tool in cryoSPARC or similar methods) to enhance side-chain and ligand density visibility.

      We have updated main and supplementary figures to include sharpened maps. All maps were sharpened using the Autosharpen feature of Phenix, specifically, by half-maps. We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      (b) Document the specific parameters used (e.g., B-factor values, solvent content estimates) in the Methods section to improve transparency.

      See above regarding the additions made to the methods section.

      (c) Map replacement and reanalysis: Replace all figures and supplementary panels displaying raw (unsharpened) maps (e.g., Figures 1D, 2D, 3A/B, 4A, 5B, and Supplementary Figs. 2D, S7B, S10) with the newly sharpened versions. Reanalyze density features (e.g., glutamine binding sites, ATP triphosphate groups) using these revised maps and update results to reflect any changes in interpretation.

      As requested, we have updated figures with sharpened maps and found our original analyses to hold. In particular, we have included multiple metrics of ligand model scoring in Supplementary Figure 7B including Q-score and CC. Additionally, we have included all ligand model statistics in Supplementary Table 2.

      (2) Structural modeling and validation

      (a) Ligand fitting justification: Provide high-resolution ({less than or equal to}3 Å) density slices or side-chain density close-ups (e.g., for phenylalanine rings or glutamine-binding regions) to validate claims of atomic-level detail. For non-symmetric ligands (e.g., glutamine) fitted into symmetry-refined maps, explicitly describe how symmetry constraints were adjusted or applied during fitting (e.g., local symmetry refinement, manual adjustment of ligand orientation) and include validation metrics (e.g., cross-correlation scores, density fit plots) to support the placement.

      We have supplied 5 new supplementary figures to demonstrate the resolution of our sharpened cryo-EM maps (most notably Supplementary Figures 3, 4, 8, 12, 22 and panels in others) . Of particular note is the sharpened map features of R298A decamer under turnover conditions which demonstrates multiple instances of a ring density for aromatic residues.

      See discussion above regarding the placement of glutamine in the interface density and updated handling of symmetry during refinement.

      (b) Ligand density supplements: Include supplementary figures showing representative ligand-density fits (e.g., ATP, glutamine) with clear side-chain or functional group annotations, as is standard in structural biology publications.

      In our revision we have included the following updated figures and figure panels demonstrating ligand density into sharpened maps:

      Turnover Filament Glutamine Ligand: Figure 2D-E (updated representation) and Supplementary Figure 7 (new and updated representations).

      Turnover Filament ATP and Mg(II): Supplementary Figure 10 (updated representation)

      Turnover Decamer ADP and Mg(II): Supplementary Figure 5 (new figure panel)

      Turnover R298A ADP and Mg(II): Supplementary Figure 14 (new figure panel)

      (3) Biochemical assay rigor

      (a) Control experiments: Perform and report the following controls to strengthen enzyme activity claims:

      - A blank control (reaction mixture without GS, ammonia, or glutamate) to quantify background ATP hydrolysis.

      - Substrate omission controls (reactions lacking ammonia or glutamate) to confirm that ATP hydrolysis depends on both substrates and GS catalysis.

      - A TCEP effect control (compare ATP hydrolysis rates with and without TCEP) to rule out reducing agent interference with the PK/LDH coupled assay.

      We have provided blank, substrate omission, and TCEP controls in Supplementary Figure 1. These results demonstrate negligible ATP hydrolysis without complete substrate inclusion and do not indicate any impact from TCEP inclusion.

      (b) Direct activity validation: Consider supplementing the coupled assay with a more direct measure of GS activity (e.g., quantifying inorganic phosphate release via malachite green assay) to cross-validate results.

      On the merits of the PK/LDH coupled assay being used for >55 years to measure steady-state activity of glutamine synthetases and that it is a robust assay as supported by the additional control experiments presented above in Supplementary Figure 1 we have elected not to pursue tedious cross-validation with a non-continuous assay and believe our interpretation of the enzyme kinetic results hold.

      (4) Writing and presentation clarity

      (a) Methods detail: Expand the Methods section to explicitly describe:

      - Cryo-EM data processing steps, including B-factor sharpening parameters, map reconstruction workflows, and any post-processing (e.g., filtering, masking).

      - Criteria used to validate ligand fitting (e.g., density threshold values, manual vs. automated docking).

      See above the revisions made in response to critique from review #1 which we will briefly summarize here:

      We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      Focused masks are represented in Figure 2E, Supplementary Figure 18, and Supplementary Figure 19. The details around focused mask utilization are included in the revised figure captions and the following was included in the Methods.

      “Focused masks were generated in ChimeraX (v.1.7 and above). Focused refinement and 3D classification (3 Å filter resolution, PCA initialization) were performed in cryoSPARC. Strategy of class picking and refinement are noted in Supplementary Figures 18 and 19.”

      Map reconstruction workflows are present in the relevant Supplementary Figures. No post-processing steps beyond map sharpening in Phenix were carried out. In general, human GS represents a straightforward protein to reconstruction via cryo-EM.

      Ligand identification criteria was described throughout the results section. Supplementary Figure 7 was revised to show sharpened density for either globally refined or locally refined maps fit with all three products of the glutamine synthetase reaction (ADP, Pi, and glutamine) individually showing the best CC and Qscore for glutamine. Beyond Supplementary Figure 7 we also combined both biochemical experiments and cryoEM to make this ligand assignment supported by:

      (1) Time-resolved cryo-EM experiments that show increasing filament particles over reaction time (Figure 2B and Supplementary Figures 9 and 10)

      (2) Glutamine+ATP cryo-EM screening (Figure 2C) showing long filaments

      (3) Supplementary Table 2 showing reasonable ligand statistics for glutamine

      To clarify this in the Methods sections we include the following statement:

      “Ligands were placed with ISOLDE (v1.7) and those with >0.5 Qscore and supporting biochemical and/or literature precedent were built.”

      (b) Results framing: In the Results, clearly distinguish between observations supported by sharpened maps and preliminary/unvalidated features. Avoid over interpreting density in unprocessed maps (e.g., referring to "glutamine binding" in Figure 2D without noting current density limitations).

      We have updated our discussion of Figure 2D (and now also Figure 2E) to include discussion of only sharpened maps and noted current density limitations to the interpretation.

      (5) Data and material availability

      (a) Ensure all supporting data are publicly accessible:

      - Upload raw cryo-EM movies, particle stacks, and processed maps to the Electron Microscopy Data Bank (EMDB) with appropriate accession codes.

      - Deposit final atomic models in the Protein Data Bank (PDB) and reference these accession codes in the manuscript.

      We deposited maps and models with accession codes in advance of review. The PDB and EMDB codes are available in Supplementary Table 1.

      Furthermore, for the focused maps and focused classifications that were generated during the review, and for the benefit of not cluttering the PDB/EMDB, we have included these more specific analyses in Zenodo: 10.5281/zenodo.20298855.

      - Provide detailed protocols for biochemical assays (e.g., TCEP handling, enzyme purification) in the Methods or as supplementary information to enable reproducibility.

      See updates above to reviewer #1

      (b) By implementing these revisions, the manuscript will better align with eLife's standards for methodological rigor, transparency, and reproducibility, allowing readers to confidently evaluate the study's contributions to structural and functional biology.

      We agree!

      Reviewer #3 (Recommendations for the authors):

      (1) In line 252, it would be helpful to show negative-stain EM images for each SEC peak, further probing whether any peaks correspond to partially aggregated, as this could affect the measured Kcat and Km.

      We aren’t in a position to do this experiment. We routinely check for aggregation by noting Absorbance at 340nm for non-specific scattering indicative of aggregation and observed no evidence of aggregation in our fractions.

      (2) In Supplementary Figure 6, many of the classes in the "Selected Filament Classes" inset appear to be averages of closely spaced particles, which may bias the calculation and should be excluded. In the "Selected Decamer Classes", I would suggest removing the top-view particle classes, as these particles not only have significantly different ice penetration rates, but are also more difficult to distinguish in 2D classification.

      We agree and have provided an additional, more strenuous cutoff, analysis of the tr-cryo-EM data wherein only classes that show clearly aligned decamers are included and all top views are omitted (Additional Supplementary Figure 6). We are happy to say that even with the more strenuous cutoffs that our main conclusions that filaments increase with forward reaction time holds.

      (3) In line 445, "Figure 5A" should be corrected to "Figure 5B".

      We thank the reviewer for pointing this out and have made the correction.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors performed seqFISH in 26 gastruloids and performed a variety of computational analyses on these novel spatial data sets. Whilst the data is valuable and the computational concepts useful (exposure index, L-metric, ...), the article falls short on novelty and is written using a very clunky language, often with contradictory conclusions.

      We thank the reviewer for their comments about the value of our data and computational concepts. We agree with the reviewer’s critical comments and have endeavored to address them all. We believe the resulting manuscript is greatly clarified and improved.

      Major issues:

      (1) The authors did well in explaining and detailing the provenance of data and the individual experiments performed. However, their 26 gastruloid data still constitute a very limited sampling from their total organoids: one experiment pooled 4 plates at an 80-94% success rate; 6 different aggregation experiments were done, making a total of 1843 gastruloids, sampled 26 (~1-2%). A simple IF stain of 2-3 markers in a bigger sample could have given a more accurate picture of specific domains of interest and their proximity. Regardless, more information should be given about the existing samples: variation across experimental batches, differences between 300-cell vs 100-cell gastruloids that were used.

      This omission was an oversight on our part and we thank the reviewer for catching it. We did the following to address this point:

      (1) Added date labels to Figure S1.2d (now S1.1a) so that the proportion correct for each separate experiment is clear:

      (2) We added the raw images of the gastruloids used in the study taken before fixation.

      (3) We segmented these images and quantified metrics of the masks to address differences across samples in morphology. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation. Elongation was measured as 1-(the ratio of the width and the height of the segmented gastruloid area); the code can be found here:

      https://github.com/arjunrajlaboratory/ImageAnalysisProject/blob/1b2f2119f77083c27f58f8c36b14 c48d97ea706c/workers/properties/blobs/blob_metrics_worker/entrypoint.py

      Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300).

      The literature also supports our assumption that combining experiments with different starting numbers of cells would not dramatically affect the results (Bennabi et al. 2024). In this paper they show that only 35 genes were differentially expressed between gastruloids formed from 100 cells and those formed from 300 cells (compared to 319 for those formed from 1200 cells and 336 for those formed from 50 cells, both compared to 300 cells). The same paper also demonstrates that the positioning of gene expression (as measured by IF staining) for several representative genes (Bra and Foxc1) does not significantly differ between gastruloids formed from 100 and 300 cells when normalized for overall size and AP axis length (as we’ve done in this paper as well).

      We have updated the text and revised Figure S1.1 to reflect these changes.

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. The experiment was performed 3 times on different days, so to ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].

      To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (see also Revised Figure S1.1a-c).

      (2) Language in the manuscript should be revised. Overall the manuscript is very long, descriptive and written "impressions and beliefs" are often not adequately justified and indeed can be contradictory, e.g. in Section 1: the title states "cell types' locations ...are consistent", a few sentences down we find "there was substantial variation" and "within range of what would be considered a 'morphologically normal' gastruloid". "quite consistent", "compelling patterning", "we don't believe"... these types of expressions are best avoided and replaced with data or used and bolstered with quantitative numbers such as percentages when a given cutoff is used. Another example: "location of each cell type relative to gastruloid morphology was quite consistent the posterior region ... mainly consisted in NMPs." Given T expression in the posterior, this result phrased as such appears quite inflated, in fact, looking at cell types in Figures S1, 2a/b/c, this reviewer would state they are all but consistent and indeed it takes sophisticated analyses to find a pattern (of sorts) beyond the coarse domains expected!

      We thank the reviewer for their careful reading of our paper and appreciate that the work would be strengthened by increasing the degree to which quantitative measures are used to justify the statement we make. We have made the following changes to the manuscript to address this criticism:

      (1) We more clearly delineate where we are making qualitative descriptions and have removed summary language (like ‘consistent’, ‘normal’, ‘variable’ etc.) from these sections. For example, the section the reviewer refers to originally read:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we characterized the organization of each by mapping where each cell type was found relative to other types and overall morphology. The approximate location of each cell type relative to gastruloid morphology was quite consistent: the posterior region, although highly variable in size (Figure S1.2a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]....”

      And now reads:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we first qualitatively examined where each cell type was found relative to other types and overall morphology. The posterior region, although variable in size (Figure S1.3a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]...”

      (2) We follow this qualitative description with a quantitative analysis of cell type proportion where we clearly state which variable aspects are statistically significant:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d).”

      (3) We added a summary paragraph at the conclusion of the results from the first two figures which clearly states which aspects of gastruloid organization we find to be variable and which are consistent, with statistical testing:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (3) Figure 6 is one of the most valuable parts of the work, as the authors use the battery of analyses developed to investigate the variable and not-so-robust endothelial clusters in gastruloids. However, this investigation is still very preliminary, and it should be further linked with known biology. It is still unclear what the unique organization of this cell type is (circularity isn't convincing) and whether any signalling cues of adjacent cells could explain it. Is there any evidence that more mature endodermal cell types are generated (like the suggested "liver") to give rise to endothelial cells? It would certainly be interesting to perform IF for this cell type together with mesodermal and endodermal markers to validate seqFISH predictions on a bigger sample.

      We appreciate the reviewer pointing out that the comparisons between different endothelial cell types was interesting, and agree that the clustering methods were insufficiently justified and that a more explicit consideration of the signaling context of the gastruloid could strengthen our findings.

      We have re-evaluated how we calculate differentially expressed genes. We restricted our analysis to only consider genes that are expressed at > 2 counts/cell in at least 50% of the subsets considered. The results are in shown in the revised Figure 6.

      We find that, as the reviewer suggested, some signaling genes are significantly differentially expressed. Specifically, Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in notch signaling in the posterior of the gastruloid. Tek, on the other hand, is more expressed in the somite-associated endothelial cells, and Tek has been annotated to be involved in retinoic acid signaling. These findings align with the reviewer’s observation that signaling from adjacent cells could explain or relate to differentially expressed genes.

      We also did a more thorough review of the literature, and found several papers that reported unique subsets of endothelial precursors, albeit in related systems. In [Rossi 2022] and [Rossi 2021] the authors find a population of endoderm-associated endothelial cells in gastruloids grown with a different protocol that involves Matrigel embedding, treatment with factors that promote blood development, and growth for 168 hours. In [Veenlveit 2020] the authors find a unique somite-associated population of endothelial cells in Trunk-Like Structures, which are similar to gastruloids but model later in development and have more physical organization with discrete somites. To address the reviewer’s request that we further link with known biology we have added the following to the text:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which show more tissue-like organization than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      Finally, we tested several methods of clustering and calculating circularity, and determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and any conclusions drawn from the text.

      (4) Figures 1c and 6b need statistical significance assessments.

      We thank the reviewer for pointing out that without significance testing these plots are difficult to interpret. We have removed plot 6b (see response above about removing the circularity assessments). For plot 1c we appreciate that it is difficult to interpret which cell types vary more than others in their occurrence without significance testing. To address this we did two things: we first calculated the coefficient of variation for the proportion of each cell type across samples:

      Author response image 1.

      To calculate significance, we first considered that since these values are proportions, they must sum to 1 and changes in one cell type will affect at least one other cell type within the same sample. To properly account for this when applying statistical tests, we calculated the CLR-transformed proportion and tested all pairs of cell types for significant variation. The results are shown in the Author response image 2:

      Author response image 2.

      We added a plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We also address said variation in the text:

      “We sought to quantify variability in cell type composition between gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centred log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.3a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.4).

      The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.4d).”

      (5) The article should include an analysis of Hox colinearity expression in these gastruloids as a validation of the system.

      We thank the reviewer for pointing out the importance of these genes in validating the gastruloid system and agree that assessing their expression specifically would help readers assess data quality.

      We analyzed the center of mass of expression along the AP axis for the Hox genes included in our panel (Hoxb6, Hoxc10, Hoxd1, Hoxb9, Hoxc8, Hoxc6, Hoxaas3, and Hoxb1). We highlighted these genes in Figure S1.3: their expression along the AP axis is consistent with previously reported expression in the tomoseq dataset from [van den Brink 2020]. A summary of the correlation coefficients for each individual gastruloid for all genes (blue) and the Hox genes (orange) is shown in the Figure S1.2. The Hox genes have similar correlation coefficients overall, although their variation is higher. This is likely due to differences in gastruloid pseudo-age; in future experiments we plan to include more Hox genes and use their expression to further classify gastruloids (see updated Figure S1.2).

      We have updated the text with these new results:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an ambitious and technically challenging spatial-transcriptomic atlas of 26 gastruloids using seqFISH. The authors introduce quantitative metrics (mixing score, exposure index, L-metric / scL-metric, spatial L-metric, triplets) to characterize spatial organization at multiple scales. The dataset is valuable, and several analyses are original, particularly the rank-based L-metric family for mutual exclusivity.

      Strengths:

      The authors generate one of the most detailed spatial transcriptomic datasets of gastruloids to date. They propose creative computational metrics (L-metric/scL-metric) to quantify mutual exclusivity of gene expression without predefined thresholds, and they explore organizational principles from single-cell topology to cluster-level structure. Many observations align well with known gastruloid biology, such as posterior robustness and anterior variability. The writing is generally clear, and the figures are rich.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      Several central claims rely on metrics whose computation and justification are insufficiently explained, making it difficult to assess how robust or interpretable the results are. Many choices in the analysis appear arbitrary or are insufficiently motivated (normalization schemes, choice of parameters such as the number of neighbors, the distance cutoffs, hierarchical clustering setup, and so on). The interpretations of spatial consistency, gene-program inference, and endothelial heterogeneity are plausible but might be stronger than the evidence currently supports.

      The manuscript would benefit from stronger benchmarking, quantification of uncertainty, and explicit controls for known artifacts in spatial transcriptomics (e.g., spillover, 2D slicing, cell type assignment entropy). The biological insights are promising, but since several depend on methodological assumptions that have not yet been demonstrated to be stable, they would benefit from clearer methodological explanation.

      We thank the reviewer for spending time to give constructive and actionable comments, and we believe the manuscript is greatly strengthened and more consistent and clear as a result of the changes suggested.

      The work is rich and could become a reference dataset. Then, clarifying and validating the quantitative methods will considerably strengthen the impact and reliability of the conclusions.

      Reviewer #3 (Public review):

      Summary:

      Triandafillou and colleagues report a single-cell resolved spatial atlas of gene expression of 26 gastruloids. While previous work had analyzed either single-cell gene expression or spatially coarse-grained patterns of gene expression (van den Brink et al, 2020), the authors here use multiplexed sequential RNA FISH (seqFISH) to create the first gastruloid atlas, which is simultaneously spatially and cellularly resolved. This atlas adds to a growing list of resources cataloging gastruloid development (see also Suppinger et al 2023).

      To analyze this dataset, the authors also describe a novel analytical framework. Their analysis centers around the 'L-metric', which measures the degree to which pairs of genes are either coexpressed or mutually exclusive. While this metric is similar to calculating correlations in gene expressions, it has important differences (including that it can, in principle, be asymmetric; although the authors symmetrize much of their analysis). In addition to the gene-centric L-metric analysis, the authors also analyze cells in their dataset according to the cell type entropy (an information-theoretical measure of confidence in cell type assignment) and the 'exposure index' (a measure of the similarity of nearest cellular neighbors).

      Using this framework, the authors focus their analysis on two major features of development. The first is the differentiation of the bipotent neuromesodermal progenitor (NMP) cells in the posterior of the gastruloid into either presomitic mesoderm (PSM) or spinal cord SC lineages. They use L-metric analysis to compare overlap in marker genes used to separate NMP, PSM, and SC fates. They highlight that L-metric analysis can recover spatial patterns of gene expression (without explicit spatial information) and discern subtle features of marker genes beyond simple binning of cell types (e.g., that Epha5 expression in anterior NMPs may predict future SC differentiation).

      The second is the formation of endothelial (spatial) clusters within the gastruloid. The authors highlight two subtypes of endothelial clusters: (1) smaller clusters within the somitic anterior region, and (2) larger clusters associated with endoderm. While the authors discern some subtle differences in gene expression between these two clusters, their different spatial patterns suggest a potential physiological difference that would not be captured in traditional droplet microfluidic-based scRNAseq pipelines.

      Overall, this manuscript is a sophisticated and technically sound study that will provide a valuable beachhead for future studies of developmental patterning in gastruloids and organoids.

      Strengths:

      The major strengths of this study are the overall technical sophistication of the data set and analysis, as well as its potential generalizability to other developmental systems (both in vitro and in vivo). The data are extensively analyzed and reasonably interpreted, and this atlas makes good use of the variability in gastruloid development to extract the statistical structure of developmental processes. The L-metric offers a parameter-free tool to analyze transcriptomic datasets that could overcome the pitfalls of other approaches.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      The major limitations of this study are the depth and novelty of the developmental processes studied. The authors provide very convincing proof-of-concept that their dataset can recover known features of gastruloid development, including NMP differentiation and endothelial development. However, further analysis and/or investigation would be required to discover new principles of gastruloid development and patterning.

      We agree that the developmental processes studied here are not inherently novel, and we hope that by showing sufficient overlap with different, less highly resolved methods we have created a convincing document that highlights the potential for this technique to be used to analyze other systems. We appreciate the reviewer’s comments and that the manuscript is improved after making the suggested changes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) S1.2 plates shown individually, but unclear from which experiment.

      The reviewer was right to point out this oversight — we have updated the figure (now S1.1a) with labels for the individual experiments:

      (2) Figure 2 could include a clearer indication of the types of triples/ doublets to make it even more informative.

      We thank the reviewer for pointing out that the types of triples were not clear — we’ve added a color key and more explanatory text to this figure in order which explicitly explains the type of triplets considered.

      (3) Figures should be presented in order. Figure 3c is before 3b, etc.

      We appreciate the reviewer’s attention to detail and have swapped these two panels so that their order in the figure reflects the order they are referenced in the text.

      (4) Figure 3 is interesting, and the L-metric appears useful to pinpoint crucial genes that, when expressed, indicate a type transition has occurred. It would be great to test this with another set of cell types besides NMPs/presomitic/spinal cord.

      We thank the reviewer for their interest in this biological transition, and agree that testing on another transition would be really interesting. There isn’t another set of cell types expected at this stage of gastruloid development that are predicted to have the same type of bifurcating differentiation. However, in an effort to address the spirit of this comment (that looking at other sets of cell types with the L-score would be interesting), we have used our analytical framework on a non-spatial, single-cell dataset from van den Brink et al 2020.

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (5) Figures 4 and 5 felt exploratory, and I would recommend combining them into a single figure highlighting the usefulness of L-metric and its spatial version.

      We thank the reviewer for the suggestion to merge the content of Figures 4 and 5. While we agree that they are thematically related, given the size of the heatmaps generated in the analyses, we were unable to combine them in a way that preserved the readability of the figure and stayed within the space constraints of the page size; thus, we have chosen to keep these as separate figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Quantification methods require clearer formalization and justification

      A key limitation is that the manuscript relies on several spatial metrics whose definitions are not sufficiently formalized.

      (a) To evaluate biological interpretations, the reader needs a precise description of:

      - How each metric is calculated (mixing score, exposure index, scL-metric, spatial L-metric).

      - Why specific choices were made (normalizations, parameter values, distance thresholds).

      - What the expected ranges and interpretations are.

      We agree that these elements are crucial to interpreting quantitative metrics and thank the reviewer for their close read of the work. The reviewer had many comments on the exposure/mixing values calculated for Figures 1 and 2, and for the L-metric (now called L-score) values calculated for Figures 3, 4, and 5. To address the reviewers concerns we have done the following:

      (1) Created a new, unified framework for calculating cell type exposure and mixing.

      (2) Rewritten the methods section for this section with a particular emphasis on including elements the reviewer suggested, including specifically outlining normalizations and what they account for, distances chosen and the biological rationale behind them, and the range of values expected for each measure:

      “Quantification of Cell Type Spatial Relationships: Exposure index

      To characterize the spatial organization of cell types, we computed two related metrics: a pairwise exposure index matrix capturing type-specific spatial relationships, and a scalar mixing index summarizing overall spatial integration. For each cell, we identified neighbours as all cells whose centroids fell within a specified radius r of the focal cell's centroid. We chose a value of r of 16 μm, which gave an average of 5-6 neighbors per cell. We chose this value as it captures local interactions, which was the primary goal of these analyses. For each ordered pair of cell types (s, t), we calculated the exposure index as the proportion of type s cells' neighbors that are type t:

      where Ns→t denotes the count of neighbour pairs in which the focal cell is type ‘s’ and the neighbour is type ‘t’, and Ns denotes the total number of neighbours across all type ‘s’ cells. Each row of the resulting exposure matrix sums to unity and represents a probability distribution over neighbour types for a given source type. To account for differences in cell type abundance, we normalized exposure indices relative to the expectation under random spatial arrangement:

      Where pt is the proportion of cells that are type ‘t’. Normalized values of zero indicate exposure consistent with random mixing, positive values indicate spatial attraction (co-localization), and negative values indicate spatial avoidance, with a minimum of −1 representing complete exclusion.”

      “Quantification of Cell Type Spatial Relationships: Mixing index

      To summarize overall spatial integration across all cell types, we computed a mixing index defined as the fraction of neighbour pairs involving different cell types:

      where N_cross is the number of neighbour pairs involving cells of different types and N_total is the total number of neighbour pairs. For normalization, we compared the observed mixing to the expectation under random spatial arrangement:

      where

      is the expected cross-type interaction rate given cell type proportions. Normalized values of zero indicate random spatial mixing, positive values indicate greater integration than expected (hyper-mixing), and negative values indicate spatial segregation.

      Exclusion of untyped cells. When computing the mixing index, cells lacking confident type assignments were optionally excluded from both the numerator and denominator, ensuring the metric reflects only spatial relationships among typed cells. These cells were retained in the exposure matrix to quantify how typed cells interact with unclassified cells.

      Statistical analysis of variance. To identify cell type pairs whose spatial relationships varied significantly across samples, we computed the variance in exposure indices across samples for each type pair. To account for the expected relationship between mean exposure and variance, we regressed log-variance against log-absolute-mean across all type pairs and computed residuals. Type pairs with residual variance exceeding the 97.5th percentile (two-tailed α = 0.05) were considered significantly variable, indicating spatial relationships that differ across samples beyond what is expected from sampling variation and composition differences.”

      (3) We have also rewritten the methods for how the L-score is calculated, adding emphasis to where we normalize, what ranges of values are expected, and what the interpretation of these values are:

      “Calculating the single-cell L-score

      The single-cell L-score (scL-score) was computed for each ordered gene pair (gene A, gene B) within a single gastruloid. The cell-by-gene expression matrix was filtered to retain only cells with at least 2 detected transcripts for both genes. Cells were sorted in descending order by gene A's expression values; gene A thus serves as the reference distribution, and the score is asymmetric with respect to gene order.

      Three reference distributions were constructed for gene B: (1) perfect coexpression, in which gene B's values were sorted in the same descending order as gene A; (2) perfect mutual exclusivity, in which gene B's values were sorted in ascending order; and (3) independence, in which every cell was assigned the mean expression value of gene B.

      Cumulative sums of expression values were computed for gene A, for the observed expression of gene B, and for each reference distribution. The cumulative sum of each gene B distribution (observed and reference) was then plotted against the cumulative sum of gene A. This cumulative-sum-versus-cumulative-sum representation captures how gene B's expression accumulates relative to gene A's: if gene B's expression is concentrated in the same high-expressing cells as gene A, gene B's cumulative curve rises steeply at first; if concentrated in opposite cells, the curve rises steeply at the end. The area under each curve was calculated using trapezoidal integration and normalized by the product of gene A's and gene B's total expression, yielding four normalized areas: A_observed (observed relationship), A_positive (perfect coexpression), A_negative (perfect mutual exclusivity), and A_uniform (independence). This normalization ensures that scores are comparable across gene pairs with different overall expression levels (Figure S3.2a,b). The scL-score was then defined as follows:

      If A_observed > A_uniform: scL-score = (A_observed − A_uniform) / (A_positive − A_uniform), yielding values in (0, 1].

      If A_observed = A_uniform: scL-score = 0.

      If A_observed < A_uniform: scL-score = −(A_observed − A_uniform) / (A_negative − A_uniform), yielding values in [−1, 0).

      A score of +1 indicates perfect coexpression, −1 indicates perfect mutual exclusivity, and 0 indicates independence.

      The scL-score was computed for all gene pairs in each gastruloid from the 05/07/2025 dataset (n = 18 gastruloids). The other two datasets (n = 8 gastruloids) were excluded due to lower transcript detection quality. To generate an average scL-score matrix, the analysis was restricted to a common set of 202 genes well-detected across all 18 gastruloids, and per-gastruloid matrices were averaged. Unless otherwise noted, a symmetrized scL-score was used: scL-score_sym(A, B) = [scL-score(A, B) + scL-score(B, A)] / 2.

      Calculating the spatial L-score

      The spatial L-score extends the scL-score to spatial regions. For each gene, a kernel density estimate (KDE) was fitted over all detected transcript spots and evaluated on a regular square grid spanning the gastruloid. Bin side length was set to twice the median nearest-neighbour distance between detected spots, calculated separately for each gastruloid. Spatial bins were ranked by KDE-derived density and processed identically to the scL-score calculation. Low-density bins were not filtered, as KDE smoothing produced non-zero density values throughout the imaging area. The spatial L-score was symmetrized as for the scL-score, except when displaying asymmetric heatmaps.

      Hierarchical clustering of L-score matrices

      Gene-gene distances were defined as Euclidean distances between L-score vectors. Agglomerative hierarchical clustering was performed using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean'). This approach operates on L-score vector differences rather than on pairwise L-score values directly, and therefore does not require the L-score itself to satisfy the properties of a mathematical distance metric; the Euclidean distance between L-score vectors is non-negative and symmetric by construction, satisfying the requirements of Ward's method. Heatmaps display pairwise L-score values, not vector distances.

      We applied this clustering procedure to the following gene sets:

      (1) A subset of NMP, presomitic mesoderm, and spinal cord marker genes in one gastruloid (n = 36 genes; Figure 3g).

      (2) All well-detected genes excluding cell cycle genes, averaged across all gastruloids (n = 166 genes; Figure 4 and Figure S4.3a).

      (3) All well-detected genes common to all gastruloids, averaged across gastruloids (n = 202 genes; Figures S4.1a, S4.3b, S5.1a).

      (4) All well-detected genes excluding cell cycle genes in one example gastruloid (n = 171 genes; Figures 5b, S5.2a).

      (5) All well-detected genes common between our seqFISH panel and those that were detected in > 3 cells in scRNA-seq data from [XXX] (n=207 genes; Figure S4.5a).

      (6) All genes detected in > 3 cells in scRNA-seq data from [XXX] (n=19075 genes, Figure S4.5c-f).

      In Figure S4.2a,b a transformation of the L-metric values was used to cluster genes. The scL-scores were averaged across gastruloids as described above, and then each pairwise scL-score was transformed to a distance-like value with(1 - scL)/2. The matrix was then symmetrized as described previously. Agglomerative hierarchical clustering was performed directly on this transformed gene-gene distance matrix using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean').”

      (4) We have added an illustrative figure about how the L-score is calculated which defines expected behaviour for several cases, gives a visual explanation of the process, and shows several extreme behaviours and their biological interpretation (see Revised Figure S3.2).

      (b) For example:

      - Mixing score: The normalization is unclear. Why only 1-nearest neighbor instead of k-NN? Why not consider existing spatial-autocorrelation metrics such as Moran's I, which would also apply to gene-level mixing?

      - Exposure index: The normalization makes the metric unbounded (e.g., exposure > 1 when local frequency > global frequency). It is unclear whether this behavior is intended. Since exposure to self is meaningful, the same metric could replace the mixing score and simplify the framework. The choice of k = 5 is not justified; parameter-free approaches like Delaunay triangulation could avoid arbitrary cutoffs. If k-nn is preferred, then the robustness of the score to change the k value should be studied.

      These issues make it difficult to interpret the magnitude of reported effects or compare them across studies.

      The reviewer makes an excellent point — we have completely overhauled this analysis in the following way to address the issues raised:

      (1) Created one unified metric (see points 1 and 2 above) which considers for every cell, the identity of its neighbors in a 16 um radius (on average 5 or 6 neighbors for each cell in each gastruloid). This value was chosen so that in most cases, the cells in the immediate vicinity of a cell were considered, but not those further out (i.e. the measure is sensitive to close interactions rather than far ones). We made a matrix of all interaction pairs for a given gastruloid, with diagonal elements representing within-type interactions and off-diagonal elements representing cross-type interactions. We have replaced the previous description with the following:

      “Several studies of gene expression in gastruloids have used pooled measurements to infer the AP axis-location of genes and cell types [3,7,13,24,27] and our data are largely consistent with these lower-resolution findings (Figure S1.3a). Yet it is obvious from individual gene staining [1,4,8,28] and our detailed 2D maps of cell identity and location that gastruloid organization is much more complex than the average order of cells along the AP axis. We thus needed an analytical method for quantifying spatial organization beyond distributions along the AP axis. To further characterize spatial organization, we sought to quantify the degree to which cells were mixed in each gastruloid, and how that mixing might vary between gastruloids. For each cell in each gastruloid, we counted the interactions between that cell and all its neighbours within a 16 μm radius (on average 5-6 neighbors per cell), and summarized all these interactions for all cells in the gastruloid in a matrix, normalizing each element by the frequency of the cell type considered to be the ‘neighbour’ in the interaction.”

      (2) To quantify overall mixing, we calculate the sum across types of the frequency of self interactions (normalized to the total interactions) and then take the inverse (1-M). We then normalize this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of that cell type).

      Because the density of cells is fairly consistent across the gastruloids, w is very close to p (the proportion of that type).

      is the expectation of cross-type rate with random mixing.

      Mixing ranges from -1 (totally segregated) to +1 (more mixed than random, i.e. there is attraction between unlike types). 0 is completely random, and negative values indicate that cell types within that gastruloid tend to cluster. We added the following to the text to explain this:

      “To quantify overall mixing, we calculated the sum (across types) of the frequency of self interactions, normalized to the total interactions) and then took the inverse. We normalized this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of the neighbouring cell type). This gave us, for each gastruloid, a value that we call the mixing index that ranged from -1 (totally segregated) to +1 (totally mixed with less frequent self-interactions than expected from chance). A mixing index of 0 indicates a random distribution, i.e., neighbour frequency is exactly what would be predicted by that cell type’s frequency alone.”

      We also edited the following description of the overall distribution of mixing indices:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). All gastruloids had a negative mixing index, indicating that they all, on average, had more like-cell type interactions than would be expected given random mixing of types. However, we note that there is a ~14% difference in the mixing index across gastruloids, meaning some variation in overall mixing is present.”

      (3) The exposure index for a given pair can be found from the off-diagonal elements of the interaction matrix, and the normalization means it represents relative overexposure/clustering (positive values) or underexposure/avoidance (negative values). The minimum value is -1 and the maximum is (1-pt)/pt. We changed the description of how the exposure index is calculated to reflect this unified method of quantification:

      We noted that the off-diagonal elements of the matrix we used to calculate the mixing index were informative about cell type-cell type interactions. Specifically they quantify the degree to which each cell type (source) is exposed to another cell type (neighbours). To assess the overall frequency of cell type-cell type interactions, we first pooled the data from all gastruloids together into one interaction matrix (Figure 2b). The measure can range from -1 (no interactions at all), with higher values indicating a greater frequency of being found in close proximity. It is asymmetric in that the exposure of cell type A to B may not be the same as the exposure of cell type B to A.

      We have updated all of the quantification in Figures 1 and 2 with these new measures.

      (2) L-metric: unclear justification and interoperability

      (a) First of all, even if it is not a major issue, the L-metric is not a "metric" at least in the mathematical sense of a metric since a metric is always positive. The L-metric is central to several major conclusions (gene exclusivity, modules, spatial organization), but its conceptual basis and computational steps need more justification.

      We thank the reviewer for pointing this out and have changed “L-metric” to “L-score” throughout. We have also endeavoured to clarify the conceptual basis and have fleshed out the various computational steps as outlined in more detail in our responses below.

      (b) Several steps (ranking, cumulative curves, area under the curve) are difficult to interpret biologically It is unclear why each transformation is required and how it responds to common scenarios (highly expressed genes, correlated vs mutually exclusive patterns).

      Since the metric is rank-based, two genes that are both highly expressed in all cells may show low L-metric despite being truly correlated.

      We appreciate the reviewer’s comments about both the interpretation of the scL-score calculation and how it behaves in common expression scenarios, particularly for genes that are broadly expressed across many cells. To address these points, we generated a set of simulated examples spanning five scenarios: ubiquitously expressed genes with similarly high average expression, ubiquitously expressed genes with differing average expression, ubiquitously expressed genes with similarly low average expression, genes coexpressed across a subset of cells rather than all cells, and genes generally expressed in opposite subsets of cells. For the first three simulations, we independently sampled two genes across 50 cells using Poisson distributions with mean expression set to 100 or 50 (to simulate a gene with high or low average expression, respectively), without any expression bias towards any subsets of cells. For the latter two simulations of dependent expression relationships, we first sampled gene 1 across 50 cells using a Poisson distribution with mean expression set to 4, then generated gene 2 from gene 1 by sampling from cell-specific Poisson distributions fitted to either generally match or oppose the transcript count obtained for gene 1 in that cell. These simulations show that genes can independently appear broadly coexpressed at the population level (regardless of average expression) simply by being ubiquitously expressed, yet still receive low scL-score values. The simulations of dependent coexpression or mutually exclusive expression relationships receive scL-score values near +1 and -1, respectively. These results align with the reviewer’s prediction, but they reflect why we designed the L-metric to follow a rank-based methodology since they preserve the specificity of the metric’s upper bound (+1) for detecting non-random coexpression relationships rather than chance coexpression relationships resulting from independently ubiquitous expression. The results of these simulations are depicted (See Revised Figure 3.3).

      We have made the following edits to the text to specifically address the case the reviewer raised about highly expressed genes:

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3.”

      (c) Interpretation of L-metric values is ambiguous

      What does 0 represent? Randoms? Co-expression? Is 1 the strongest exclusivity? The manuscript currently mixes "co-expression" and "mutual exclusivity" scales.

      We agree with the reviewer that clearly defining what values of the L-score mean is critical to understanding the text. We have added a more explicit discussion of this in the text (excerpted from the response to 2c):

      “A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes.”

      And made a figure representing visually how the L-score is calculated which shows the behaviour and biological interpretation of several extreme values and an example of how the L-score is calculated. See new Figure S3.2:

      We also edited figure captions where we referred to plots as ‘co-expression’ plots, since in some cases the plots showed genes that were mutually exclusive or not expressed together in most cells. Figure S3.1:

      “c. Spatial distribution of the expression of Nkx1-2 and Rfx4 in an example gastruloid.”

      In all other cases we checked, we used the term “co-expression” to mean the opposite of mutually exclusive, as outlined in the definition above.

      (d) Additional issues also limit interpretability - Benchmarking is missing.

      - No tests on synthetic datasets, negative controls, or curated examples.

      - Prior exclusivity methods (EEI, COTAN) routinely benchmark against ground truth; this is now standard.

      - The code for the L-metric seems to be missing in the repository.

      We appreciate this suggestion offered by the reviewer as benchmarking against prior exclusivity-oriented methods provides an important comparison for clarifying both where the scL-score agrees with existing approaches and where it offers distinct advantages. To address this, we explicitly compared the scL-score to the Exclusively Expressed Index (EEI), which is bounded below by 0 and increases with mutual exclusivity, and to the signed coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework, in which positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. We performed this comparison using seven simulated scenarios as well as four representative gene pairs from one gastruloid sample (2025-05-07_roi2). In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (high_high, high_low, low_low), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained low, and COEX was also positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.

      We then applied the same comparison to four gene pairs from one gastruloid sample (2025-05-07_roi2). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types. Pax6-Eogt, Rfx4-Eogt, and Nkx1-2-Rfx4 all showed opposing expression by scL-score, with values of -0.572 and -0.526, -0.970 and -0.955, and -0.537 and -0.615, respectively. EEI detected exclusivity most strongly for Rfx4-Eogt (0.171), but gave values of 0 or approximately 0 for the other two pairs, while COEX was negative for all three pairs (Pax6-Eogt: -0.235, Rfx4-Eogt: -0.350, and Nkx1-2-Rfx4: -0.106), consistent with opposing expression, with strongest signal for Rfx4-Eogt. By contrast, Cdx4-Cdx2 showed moderate levels of coexpression by scL-score (0.402 and 0.424) and EEI (0), but COEX indicated that they were not coexpressed (-0.331).

      These comparisons also clarify the practical advantage of the L-metric over the EEI and COTAN frameworks. EEI is based on binary zero/non-zero quantification and is therefore designed specifically to measure exclusivity rather than coexpression. COEX provides a signed value and, in our simulations, tracked both exclusivity and coexpression; however, on representative gene pairs from one gastruloid sample, scL and COEX diverged in magnitude for highly exclusive expression relationships (Rfx4-Eogt) and sign for a coexpression relationship (Cdx4-Cdx2), motivating our introduction of a signed measure based on the mutual exclusivity of expression with the scL-score (see New Figures S3.3 and S3.4).

      We have updated the text to address the reviewer’s concerns: we benchmark using simulations of commonly occurring scenarios (such as varying expression levels, degree of mutual exclusivity, and amount of noise present in the relationship between the two genes in question) as the reviewer suggested. We also provided a direct comparison to an earlier exclusivity measure (EEI):

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3 and Figure 3.4.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We compared the scL-score to two existing measures of exclusivity. The Exclusively Expressed Index (EEI) (Nakajima et al. 2021) is bounded below by 0 and increases with mutual exclusivity. EEI is computed from binary zero/non-zero quantification and is designed specifically to measure exclusivity but not coexpression. The coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework (Galfrè et al. 2021) can also be used to quantify relationships between genes: positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (but with varying relative expression levels), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained close to 0, and COEX was positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We calculated the scL-score, EEI, and COEX for four gene pairs from one gastruloid sample (2025-05-07_roi2). The scL-score consistently delineated gene pairs possessing opposing expression profiles, while EEI was not always able to measure those exclusivity patterns (Pax6-Eogt: scL-score=-0.572 and -0.526, EEI=0; Rfx4-Eogt: scL-score=-0.970 and -0.955, EEI=0.171; Nkx1-2-Rfx4: scL-score=-0.537 and -0.615, EEI~0). Only the scL-score was able to detect the coexpression pattern present between the positively associated expression profiles of Cdx4 and Cdx2 (scL-score=0.413, EEI=0, COEX -0.331) (shown visually in Figure S3.4a). Thus, while EEI was informative for measuring gene expression relationships characterized by mutual exclusivity, the scL-score more clearly separated positive, random, and mutually exclusive relationships on a single signed bounded scale. The COEX value trended in the opposite direction than expected, but this may be due to the fact that it cannot be calculated on single-gene pairs and necessarily uses information from the entire count table, which here only consisted of 6 genes. These comparisons combined with the simulations in Figure S3.3, clarify a conceptual advantage of the L-metric over the EEI and COTAN frameworks. In contrast, by leveraging the ranked structure of transcript counts across cells, the L-metric framework does not binarize expression and does not require fitting a parametric distribution. It can be calculated on single gene pairs, and is more sensitive to mutual exclusivity.”

      We also amended the Data and Code Availability section to include a specific reference to the L-metric package that was previously missing:

      “All code used to process the raw data and generate figures, as well as the processed data and figures can be found at the following link :

      https://www.dropbox.com/scl/fo/bchkqlbcjb8ub9m606did/AIudcWZaC566toXzb2L-jXc?rlkey=u0wgtkq8j erxoqb5ump599oip&dl=0

      Additional custom scripts used to process the raw seqFISH data can be found on GitHub:

      https://github.com/arjunrajlaboratory/NimbusImage/

      The code for calculating the L-score can be found on GitHub:

      https://github.com/arjunrajlaboratory/l-metric

      Images of all gastruloids generated for this study, as well as single-channel seqFISH images with segmentation and annotations are available here:

      https://app.nimbusimage.com/#/project/69d3fa8f1be4701f5fab6359 Raw seqFISH images are available upon request.”

      (e) Clustering using L-metric vectors

      In this study, the authors use hierarchical clustering to group genes according to the L-metric. This choice is reasonable: hierarchical clustering provides a natural representation of similarity relationships across multiple scales, and the L-metric captures a form of signed dissimilarity between genes. However, this approach raises an important issue. Standard hierarchical clustering methods typically assume a non-negative metric, whereas, as noted earlier, the L-metric can take negative values, meaning it does not strictly satisfy the requirements of a metric in the mathematical sense.

      To the best of our understanding from both the text and the source code, the authors address this issue by defining the distance between two genes A and B as the Euclidean distance between two vectors: the L-metric values from A to all other genes, and from B to all other genes. Although this procedure is mentioned in the manuscript, it is neither justified nor accompanied by any discussion of how such a distance should be interpreted. It is not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes. These two gene distances are not without overlap, but they are not the same.

      This choice magnifies the interpretability issue: readers must understand two layers of transformations. If the [-1,1] range poses problems for hierarchical clustering, simple transformations (e.g., 1 − L) or alternative clustering methods could avoid these issues.

      Given that the L-metric underlies major biological inferences (novel gene modules, spatial subclusters, endothelial states), clearer justification and benchmarking are essential. We think this can lead to more consistency in the spatial metrics.

      We really appreciate that the reviewer took the time to understand our proposed method thoroughly, and apologize for any confusion resulting from a lack of clarity in how it is calculated, and the language used to describe it. The reviewer is absolutely correct that it is not a metric in the mathematical sense; we have replaced the word ‘metric’ with the word ‘score’ throughout the text.

      The reviewer also raised concern about how the hierarchical clustering was performed, and they were absolutely correct about what the vectors represent—they are, as the reviewer states, “not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes”. The heatmaps in figures 3, 4, and 5 are intended to cluster genes that have similar expression patterns, i.e. similar L-score values with all other genes. The reviewer pointed out that this transformation wasn’t clear, so we have added an explicit explanation of what the vectors represent, as well as an explanation of why we were interested in how these vectors, which represent a ‘fingerprint’ of how the gene interacts with all other genes, clustered (because this is in the section discussing NMP differentiation we focus on a specific subset of genes here, but later apply to the entire panel):

      “For each gene annotated as belonging to any of the three cell types (NMP, PSM, or spinal cord), we calculated a vector of scL-score values with all other genes. Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      We also appreciate the reviewer’s suggestion that alternative transformations of the scL-score may improve clustering interpretability. We tried using the reviewer’s suggestion of doing a simple transform: we averaged the scL-score matrices across the 18 gastruloids using the shared gene panels, transformed each scL-score from the original [-1,1] scale to a [0,1] scale using (1-scL)/2, symmetrized the resulting matrix so that each gene pair was represented by a single value, and then performed hierarchical clustering directly on this gene-by-gene distance matrix using average linkage. We used this analysis to test whether a more direct distance-based approach would change the gene groupings recovered by our original clustering method.

      This alternative approach largely recovered the same cell type-associated groupings, but the separation between branches in the dendrogram was smaller, making fine-scale ordering harder to interpret. We measured this by evaluating the average cell type dispersion, measured in terms of additive branch length, which was 0.682 and 0.671 for the 166-gene and 202-gene panels, respectively. Since the branch separation on our original dendrograms was greater and thus representative of more robust groupings, we chose to continue using hierarchical clustering based on Euclidean distances between scL-score vectors.

      We comment on this alternative transformation we tried in the next section, when considering the clustering of the entire gene panel. We feel this is appropriate as the motivation for looking at the difference vector was derived from expected behaviour of a smaller set of genes, and as the reviewer pointed out it is not clear that that expectation should or would hold for the entire panel.

      “To generate the heatmap shown in Figure 4a and S4.1a, we used the same clustering method as described earlier with the Euclidean distance between scL-score vectors. However, we also tried clustering directly on the scL-scores themselves, by transforming each scL-score from the original [-1,1] scale to a [0,1] distance-like scale using a (1-scL)/2 mapping (Figure S4.3a,b). The results were largely consistent, however the cophenetic distance scale was relatively compressed when the transformed values were used (Figure S4.3a,b). This shallow structure implies that many branches are separated by only modest distances, so fine-scale ordering within the dendrogram should be interpreted more cautiously than the larger-scale cell type block structure. We chose to continue using Euclidean distance-based hierarchical clustering of scL-score profiles, where cell type grouping is observed alongside larger cophenetic separations between clusters” (See New Figure S4.3).

      (3) Claims of "remarkably consistent" spatial organization are stronger than the data currently support

      (a) The manuscript emphasizes reproducible organization across gastruloids, but several factors complicate this interpretation.

      We agree that the distinction between what is reproducible/invariant between gastruloids and what varies was not clear in the original manuscript. To address this, we have updated the text to more explicitly distinguish between the two categories. The other suggestions made by the reviewer to sharpen the quantitative measures to strengthen these claims was very helpful and we appreciate the thought put into them, and have used that framework (emphasizing the statistically significant variations and consistencies) in summarizing our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrate several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior:posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (b) Possible selection bias. Only elongated, QC-passing gastruloids were retained; 18/26 datasets remain. seqFISH runs with uneven housekeeping signals were excluded.

      We agree with the reviewer that our data are elongated, QC-passing gastruloids, although these represent two sources of variation (biological and technical respectively). Our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category, and we have attempted to signal this to readers by consistently including language like “morphologically normal” and “elongated”. We have further updated the language in the manuscript to emphasize this point:

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. To ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].”

      In regards to the exclusion of datasets, the only time 18 out of 26 were used was when calculating the averaged L scores for all genes in Figures 4 and 5. In this case we used all 18 gastruloids from the seqFISH run performed on 4/4/2025; this dataset had the highest spot counts due to protocol improvement between runs, and integrating the datasets with very different spot counts was problematic because a minimum expression level is needed to calculate L scores. We used all 26 samples for the spatial metrics calculated in Figures 1 and 2 (Figures 3 and 6 focus on specific gastruloids). We have added additional labels in Figures 1, 2, 3, and 6 to make clear when all 26 datasets are used and when only a subset is used.

      (c) Gastruloids are known to be variable; restricting to morphologically "normal" samples could inflate apparent regularity.

      We thank the reviewer for this observation and agree that the degree of variability among gastruloids is an important consideration. As stated in response to b), our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category. The rate of occurrence of ‘normal’ gastruloids in our hands is ~80% (Figure S1.1a). We agree that it would be interesting to consider how variations from this baseline affect cell type composition and arrangement, and while we make no claims about it in this paper, we have updated the introduction to highlight this point:

      “To address these gaps, and to create a systematic, high-resolution dataset of gene expression in gastruloids considered to be morphologically normal, we developed a spatially resolved, single-cell molecular map of the location, identity, and gene expression of cells within 26 individual gastruloids with normal morphologies. We found that despite some morphological variability within the qualitative category of elongated and polarized, “normal” gastruloids had largely reproducible cell type composition.”

      (d) Partial lack of statistical validation. The manuscript shows descriptive consistency but no formal tests across runs or batches (e.g., mixed-effects models, ICCs, leave-one-run-out validation).

      We agree with the reviewer that a quantitative comparison between batches is important. We have added the following to the text to address this point:

      “To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (See Revised Figure S1.1a-c).

      (e) 2D sampling limitations. Spatial metrics rely on a single imaging plane chosen as the "midplane," but z-position varies between gastruloids. AP projections, mixing, and triplet analyses could all be sensitive to z-plane choice. Prior work shows that 2D slices can underestimate distances and contacts by large margins (https://pmc.ncbi.nlm.nih.gov/articles/PMC5522766). These limitations should be acknowledged explicitly.

      We agree that sampling in 2D can limit the interpretation of our findings and we thank the reviewer for bringing up this important point. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      “Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.”

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2017].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (4) Uncertainty in cell-type assignment is not incorporated into spatial metrics

      Many spatial measurements depend directly on cell-type calls (exposure, triplets, mixing). However:

      (a) Anterior cell types have higher entropy in their marker-based scores (Figure S1.1a).

      We thank the reviewer for pointing out that several of the cell types in the anterior have high entropy — specifically cardiac mesoderm and paraxial mesoderm. However, we think there is nuance to this point; two of the other prominent anterior cell types (somite and endothelial) have low entropy scores overall, and that spinal cord/neural precursor cells, which are mostly posterior, have somewhat higher entropy; higher entropy scores are not exclusive to the anterior, nor is low entropy exclusive to the posterior. Inspired by the reviewer’s comments, we have re-analyzed our data to include this nuance (see response to point d) below.

      (b) Uncertain labels inflate apparent "mixing" or "disorder," because misclassifications randomly create mixed neighbors and triplets.

      We agree that uncertainty in labels could affect the interpretation of mixing. We appreciate these comments and the reviewer’s suggestions, and we have followed them in our response to point d) below.

      (c) Posterior cell types have low entropy, so comparisons between anterior vs posterior mixing may partly reflect label uncertainty, not biology.

      We thank the reviewer for bringing up this important caveat to our findings. We incorporated discussion of this in our text edits (see point d) below.

      (d) The authors should incorporate confidence measures (e.g., probability-weighted neighbors, entropy filtering, bootstrapping) to confirm that patterns hold independently of classification noise.

      We appreciate these suggestions and have chosen to use entropy filtering to assess whether the spatial organization we observe is highly sensitive to what values are considered ‘low’ entropy. We have updated the text (see below), and added Figure S2.2 to address the reviewer’s comments:

      “The contrast between organized posterior clustering and disorganized anterior mixing matches expectations based on literature that shows that self-organization mechanisms in gastruloids in the anterior vs. posterior are differentially sensitive to culture conditions, with somitic patterning requiring external matrix support [1,4,8,24], distinguishing it from the seemingly more autonomous organization observed in posterior cell types.”

      “One potential caveat to this finding is that differences in uncertainty in cell typing could be the primary driver of mixing and cell type interaction differences, both between the anterior and the posterior within an individual gastruloid, or overall between gastruloid. To control for this, we applied an entropy filter to our dataset. We filtered out cells that had entropy > 1.5 (see plot below for cutoff), which was chosen based on the distribution of entropy values for ‘none’ type cells, which effectively describe the upper limit of random transcript assignment (99.7% of ‘none’ type cells are removed with this filter, and about 50% of cardiac mesoderm cells and paraxial mesoderm cells, see Figure S2.2a). We first examined overall mixing; there was strong correlation between the per-gastruloid mixing indices before and after entropy filtering (Pearson r = 0.809, Figure S2.2b). Globally, mixing indices decreased with filtering, meaning that overall the cell types were more clustered. When we compared the absolute value of the change in mixing index pre and post-filtering to the proportion of each cell type, the only significant correlation was with cardiac mesoderm (Figure S2.2c). Exposure indices were overall quite similar after filtering, although the strength of somite-somite and somite-paraxial mesoderm interactions increased (Figure S2.2d).”

      “We also examined how entropy filtering might affect the exposure index, given that the mixing index is calculated from the exposure index of across cell types. In general, the magnitude of the exposure index values increased when more uncertain cells were excluded, but the directionality and relative ordering was not affected. Although the magnitude of change in the posterior cells types was less than the anterior cell types, the cross-cell type exposure values, particularly between paraxial mesoderm/endothelium and somites, doubled. From these results we conclude that mixing in the posterior is driven mainly by NMP/presomitic mesoderm interactions, and is overall lower than mixing in the anterior, which is driven by rarer cell types like cardiac mesoderm, paraxial mesoderm, and endothelium, being interspersed within somite cells” (See Figure S2.2).

      (e) This leads to reviewing the claim on endothelial heterogeneity, which strongly depend on spatial adjacency and gene exclusivity metrics.

      We have extensively considered claims of endothelial cell heterogeneity, and these are discussed in detail in response to the reviewer’s next point. We have also copied them here for the reviewer’s convenience:

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in the Author response image 3:

      Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although many of the somite-associated genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0). See Author response image 4.

      Author response image 4.

      (5) Endothelial "spatially dependent" gene expression may reflect spillover rather than intrinsic state

      The comparison between anterior-associated and posterior-associated endothelial nuclei suggests two transcriptional states. However, spatial adjacency confounds the interpretation:

      (a) seqFISH assigns transcripts to nuclei in dense tissue; partial-volume effects can mix RNA from neighboring endodermal or somitic cells.

      We thank the reviewer for their attention to detail and agree that a careful consideration of these points is important. We also note that given that some of our differentially expressed genes in endothelial cells are endoderm or somite genes, there indeed may be some transcript misassignment.

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript misassignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for none-typed cells at any dilation considered.

      (b) Endoderm and endothelium are closely intermixed (Figure S6.1d), and their gene expression co-localizes in KDE maps (Figure 5b-c).

      We agree with the reviewer and thank them for their close reading of the manuscript. We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      (c) Without explicitly quantifying spillover, differential expression between these two endothelial subsets cannot be confidently attributed to cell-intrinsic differences.

      (d) You may control for spatial proximity with any of the following:

      - Include adjacency index as a covariate in DE models.

      - Use scL-metric to test the mutual exclusivity of endothelial vs endoderm genes within the same nucleus.

      - Apply local permutation nulls: shuffle transcripts within local windows and recompute DE.

      - Restrict analysis to gastruloids that contain both endothelial subsets.

      We thank the reviewer for bringing up this important point, which we were eager to address. We agree with the reviewer’s point that by only considering gastruloids that contain both subsets of endothelial cells is the correct way to do the analysis. We were already only considering this case (specifically the values calculated in Figure 6 and for 1 gastruloid pictured in Figure 6a). We added an n=1 label to revise Figure 6d to emphasize this point.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although some genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0) (See Author response image 4).

      We have also included in the supplement larger images of some of the top differentially expressed genes, which more intuitively show the differential expression results (Endoderm enriched and Somite enriched).

      Finally, we have referenced several previously-reported instances in the literature where distinct subsets of endothelial precursors with unique gene expression programs were identified. Although in these cases 1) the embryo models were different (in [Rossi 2021, Rossi 2022] gastruloids made with a different protocol and treated with factors designed to promote blood development, and in [Veenlveit 2020] trunk-like structures) and 2) the methods were different (IF and 10x single-cell sequencing) this at least establishes a precedent for the observation of multiple types of endothelial precursors. In the case of [Veenlveit 2020] the authors specifically note that one subset is associated with somites, and we have updated the text to reflect these new results:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b,c). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Pecam1 also clustered uniquely in our L-metric analysis (Figure 4c), suggesting this differential expression is conserved across gastruloids. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behaviour is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6)

      (6) Interpretation of gene-program modules may be overstated

      Claims that the L-metric reveals "novel gene programs" should be softened:

      (a) The seqFISH panel is an approx. 200-gene marker-enriched panel, already biased toward known cell-type markers.

      This is true and we appreciate that this came through in the text since it’s important for the reader to understand the approach we took in this study.

      (b) Strong blocks in Figure 4a and S4.1a may reflect panel design rather than newly discovered programs.

      We agree with the reviewer that the panel design was not sufficiently highlighted, so we have made the following changes to the text to emphasize which patterns would be expected due to the genes we are probing for, and which findings were surprising given the known functional role of the gene.

      At the end of the section titled ‘The L-metric captures the spatial distribution of gene expression despite being calculated without spatial information’:

      “These analyses demonstrate that information contained within the hierarchical relationships between genes, determined by scL-score can reveal novel information about cell states within cell types, although we acknowledge that since cell type is determined by a limited panel of marker genes, results should be further functionally verified. scL-score analysis can also identify distinct spatial locations of cells in this cell state, all without explicit encoding of spatial information, but rather quantifying and clustering the degree to which genes are mutually exclusively expressed with one another.”

      At the end of the section titled ‘Clustering scL-metric vectors clearly resolves cell types and reveals novel genetic interactions’:

      “Finally, although the strong blocks we find in the heatmaps in Figures 4a and S4.1a largely reflect cell types, as is consistent with our panel design, we discovered some novel functions of genes in the panel through their location in the scL-score tree: although Tgfβ was initially included in our panel to generally detect inflammatory and growth signaling, clustering by expression patterns revealed its unique association with endothelial precursors.”

      To further address the concern that the generality of clustering is due to gene selection, we performed random gene drop-out and assessed how well cell types clustered as a function of the number of genes removed:

      Author response image 5.

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      Finally, we performed scL-score analysis on an unbiased, scRNA-seq dataset without any panel selection, and were able to show similar groupings of cell type markers:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (c) Cluster robustness is not assessed (bootstrap, stability).

      We appreciate the reviewer’s suggestion that cluster robustness should be assessed. To address this point, we performed a clustering-stability analysis on the common 202-gene panel that included cell cycle genes by asking to what extent hierarchical clustering could recapitulate cell type-based groupings of genes as the number of genes used for clustering was varied across progressively larger, randomly sampled panel subsets. For each resulting tree, we averaged cell types’ dispersion of genes across the tree using the framework described in our response to suggestion 6m and compared the observed value to a permutation-based distribution generated on the same tree.

      This analysis showed that the observed cell-type dispersion remained consistently lower than the corresponding permutation distribution across all subset sizes examined. In other words, genes assigned to the same annotated cell type remained closer together in the dendrogram than expected by chance even when clustering was performed on reduced random subsets of the panel. We also observed that dispersion values increased as larger gene subsets were included, which is expected as the clustering problem becomes more complex with increasing panel size; however, the separation between the permuted distribution of average cell type dispersion and the observed dispersion value increased as the panel subset size increased. Taken together, these results indicate that the cell type-resolved organization captured by the scL-score derived hierarchy is not dependent on one particular subset of genes, but is instead a stable property of the broader gene panel. New Figure S4.2 addressing cluster stability:

      We have added the following explanation in the text:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      (d) Agreement with cNMF (claimed in text) is not quantified (ARI, Jaccard, hypergeometric overlap).

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See Revised Figure S4.2 (now S4.4)).

      We have updated the text to reflect these quantitative comparisons:

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (e) Testing the scL-metric on larger, unbiased scRNA-seq datasets would help demonstrate generality.

      We agree with the reviewer that this would demonstrate generality, so we applied scL-score analysis to the scRNA-seq dataset from van den Brink 2020 — several gastruloids at the same stage of development as those used in this paper were pooled and sequenced. We added a figure, new Figure S4.5 with the results of this analysis, and a new section in the paper describing them:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (7) Minor Comments

      (a) We suggest including representative raw seqFISH images. The manuscript does not show raw images, which makes it difficult to evaluate the quality of the underlying data that all spatial analyses depend on. A figure showing raw fluorescence channels, detected spots, and nuclei segmentation masks for at least one anterior region, one posterior region, and one dense interface (e.g., endoderm-endothelial) would allow readers to assess spot intensity and background levels, signal-to-noise ratio, segmentation accuracy, and potential over/under-segmentation, channel cross-talk, and transcript crowding or dropouts in dense tissues. A small panel of raw images would improve the transferability to the spatial metrics.

      This is an excellent suggestion and we have included examples of the raw images for a representative gene for all hybridizations, the spots as determined by the spot-finding algorithm distributed with the seqFISH instrument, deconvolved spots, and nuclear segmentation. These images can be found in Revised Figure S1.1d.

      (b) The Introduction mentions 3D organization, which can confuse readers into thinking the seqFISH dataset is volumetric. The data shown and analyzed come from a single 2D plane per gastruloid, not from full 3D z-stacks. Since all spatial metrics rely on true adjacency, the manuscript should explicitly state early on that the dataset is 2D and briefly justify why a single plane is sufficient for the analyses.

      We appreciate the reviewer bringing up this point and we have removed references to 3D so as not to confuse readers.

      We also have included an extensive discussion of the 2D nature of the data. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2018].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (c) Results, first paragraph: "good agreement" along the AP axis should be quantified or defined.

      In the first paragraph we compare the peak in gene expression along the (length-normalized) AP axis of each gene with a similar but orthogonally measured dataset from another group (van den Brink 2020). The text specifically reads:

      “When we compared how gene expression varies along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section.”

      We have revised Figure S1.2 with a summary plot showing the distribution of correlation coefficients for all genes and for the Hox genes in our panel (which are known to be expressed sequentially along the AP axis).

      We have updated the text as follows:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      (d) Clarify what "greater cell type distinction" means when using marker panels, and what metric demonstrates improvement?

      By “greater cell type distinction” we meant that when using traditional clustering methods we were not able to individually resolve some cell types: there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together. Cluster labeling was performed by considering which genes showed up as being differentially expressed in each cluster using the same associations as were used to perform the cell type scoring with the marker gene panel. While we acknowledge that these results are subjective to clustering parameters, this is a general problem with clustering and not specific to this study. We have added additional clarification in the text and removed the phrase “greater cell type resolution” since it wasn’t clear what we were comparing to:

      “To profile the spatial organization of individual gastruloids, we assigned a cell type to each nucleus using a cell type scoring method with known marker genes. We compared these results to those obtained with unsupervised clustering. We found that although clustering did produce clusters, they were not strongly separated and we were not able to individually resolve some cell types we expected to find: specifically, there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together (Figure S3.6a). Therefore, we proceeded with the scoring-based method; see Methods for additional details. A representative gallery of typed gastruloids is shown in Figure 1b; the full dataset is in Figure S1.3. Because each cell received a cell type score for each type, we could use the entropy of the cell type score probability distribution to assess confidence of our cell type assignment: a cell that received a similar score for multiple cell types would have a high entropy distribution, while one which scored highly for one type and low for the rest would have low entropy. On average, the cell type entropies for most cells in a given cell type were low, with the exception of paraxial mesoderm and cardiac mesoderm cells, which had intermediate values (see additional discussion below). The generally low entropies indicate that most of our cell type assignments were high-confidence (Figure S1.4a).”

      (e) The variation in "none-typed" cells across gastruloids should be expanded and shown quantitatively.

      We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      (f) Figure S1.1b: explain the criteria used to decide which tissues "varied significantly." Showing the proportion of "None" per gastruloid would help.

      We thank the reviewer for pointing this out. We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      To justify the use of the phrase ‘varied significantly’ we have calculated the degree to which cell type pairs significantly co-vary (adjusted p value < 0.05) in their proportions and added the plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We address said variation in the text:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids profiled. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.32a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.41dc).”

      “The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.41dc).”

      (g) In Figure 1c, showing the variability of "None" cells is important.

      We made this adjustment and added the plot to Figure S1.3c

      (h) For Figure 1e, consider normalizing to T-expression proportion, not only AP length. This could clarify multimodality in the endoderm and spinal cord.

      We thank the reviewer for this suggestion to normalize by molecular as well as physical features. We agree this could be a useful projection of the data, since expression of T is used in many contexts to define the posterior of gastruloids.

      We incorporated T expression into the length normalization in the following way: we generated a cumulative distribution of all T spots along the AP axis, and when this value exceeded a threshold (specifically 90% of all spots) we defined this as the midpoint of the gastruloid, and linearly normalized space before and after it from 0-0.5 and 0.5-1 respectively. We then re-projected nuclei for each gastruloid onto this new coordinate system, and visualized in the same way as the main text figure (See Figure 1f).

      To assess whether this increased or decreased variability in position, we calculated how much the location of the peak of each cell type in each gastruloid differed from the peak position of that cell type in all samples pooled together. We compared this difference to a bootstrapped null drawn from the pooled distribution. Interestingly, we found that nearly all cell types had statistically significant variation (meaning the average distance from the mean fell outside the bootstrapped distribution in the positive direction), however the effect size for most cell types was very small.

      Notably, as the reviewer suggested, normalizing to T expression decreased the effect size of this variability in all cases but one, and particularly decreased spinal cord variability (despite not being a marker for spinal cord):

      We interpret these findings in the following way — that most of the variation observed in cell type arrangement along the AP axis is due to morphological and molecular variability (likely due to stochasticity in initial cell number and differences in developmental timing between gastruloids), and that once these factors are taken into account, for most cell types the effect size of variability is very small. However for some cell types, notably the location of the differentiation front, the position varies among gastruloids. We hypothesize that this may be due to the rhythmic nature of somite development. We have updated the text to reflect these changes.

      “Given that the proportions of cell types within each gastruloid were fairly consistent, to what degree did their spatial organization vary? We first projected each cell’s expression onto the AP axis and looked at the distribution of AP axis locations across gastruloids (Figure 1f, left-hand side). The most posterior cell types (NMP and presomitic mesoderm) showed wide distributions from 0-30% of the axis. Centred around 30%, spinal cord precursors and endoderm had distinctive peaks (clusters), the exact location of which varied between gastruloids. Using a threshold of T expression to define the midpoint of each gastruloid and uniformly length-normalizing each half caused these cell types to collapse into a single peak (Figure 1f, right-hand side). The differentiation front was similarly located in one peak (between 30-40% of the AP axis), and the variation in the location of this peak decreased with T-expression normalization. To quantify this variability, we bootstrapped a null distribution of cell type locations by pooling each cell type together across samples and using the resulting distribution to define a reference mean. When we compared how much peak variation there was among samples randomly drawn from this null to our observed data, we found that although nearly every cell type (excepting paraxial mesoderm) had more variability than expected by chance, the effect size of this variation was small (Figure S1.5a). Normalizing location to T-expression decreased variability in most cell types, most notably for differentiation front and spinal cord (Figure S1.5b).

      We also asked how the order of cell types along the AP axis varied between gastruloids. When we ranked the peaks shown in Figure 1f per gastruloid, we found that Kendall’s W, an overall measure of rank coherence across independent samples that spans from 0 (no agreement) to 1 (complete agreement), was 0.834 (Figure S1.5c). We found that the cell types most likely to swap rank order were spinal cord and endoderm, and paraxial mesoderm and endothelium (Figure S1.5d).”

      (i) Clarify "mixing coefficient values range from 0.29-0.58" (incomplete sentence). Section "Cell types' arrangement is consistent across gastruloids, but varies by type" second paragraph.

      We appreciate that there was some confusion here and we thank the reviewer for pointing it out. We have updated the text with the new values and a more specific interpretation to aid the reader and address the reviewer’s comment:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). While all values are negative, the range was large. This observation led us to conclude that while in all the gastruloids profiled cell types tended to cluster together, there was variation between gastruloids in the degree of coherent clustering between types.”

      (j) The paragraph claiming organization consistency may be too strong, given the lack of statistical validation. "Overall, these data speak to the consistency of gastruloid organization...".

      We agree that quantification of variability, which we only assessed qualitatively, would help readers better evaluate claims of organizational consistency, and we appreciate the reviewer pointing this out.

      We have taken several steps to add quantification, including:

      (1) Calculating the coefficient of variation for the proportion of each cell type

      (2) Quantifying co-variation and assessing statistical significance

      (3) Quantifying the variation in AP-axis location, including comparison to a bootstrapped null, significance testing, and a new form of normalization which decreases some of the variability.

      (4) Quantifying order along the AP axis and calculating Kendall’s W to quantify how concordant this ordering is across gastruloids

      (5) Overhauling the methods used to quantify local patterning and mixing into one unified metric

      (6) Calculating significance for variation in exposure index across samples

      To address the question of whether claims of organization consistency are appropriate, we have added the following summary paragraph at the end of the results for the first two figures, and have taken the reviewer’s suggestion of only discussing statistically significant or quantified results in listing both consistent and variable features. Now rather than making an argument about whether gastruloids are consistent or not, we merely provide the readers with our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (k) Figure 3c: gene set sizes differ substantially; proportion-based normalization may be more appropriate.

      We thank the reviewer for carefully noting the gene set differences. While the NMP-only and spinal cord-only gene sets each have 10 genes, PSM-only and NMP+PSM have 6 genes, and NMP+spinal cord has 5 genes. Given the relatively small N, we feel that normalization would likely introduce a layer of abstraction that would be more confusing for the reader, especially given the qualitative nature of the claims made about the shape of the plots in question. However, we agree this point is important, so we have now indicated the size of the gene sets on the plot (revised Figure 3b, previously c).

      (l) The purpose of the first two graphs in Figure 3c is unclear.

      We thank the reviewer for pointing out that this is unclear. The purpose of this visualization is to demonstrate correlation between genes shared between cell types and genes exclusive to only one of the cell types. The text reads:

      “Figure 3b shows the total expression of each gene group versus the NMP genes for all cells in all gastruloids that we typed as NMP, presomitic mesoderm, or spinal cord. As expected, there was a clear correlation between the mixed categories and NMP genes, supporting the notion of a continuous differentiation process.”

      We have added a correlation line to revised Figure 3b (former Figure 3c) to emphasize this point.

      (m) Clustering of all genes: quantify whether clusters align with cell types.

      We appreciate this suggestion offered by the reviewer as this analysis will allow readers to quantitatively assess the overlap between cell type-based groupings of genes and our scL-score-determined clusters for this dataset. To address this, we have used a bootstrapping approach to evaluate how well our scL-score-based hierarchies capture cell type-based groupings. Specifically, following hierarchical clustering of genes using the scL-score, for each cell type, we computed the average of the minimum cophenetic distance between each pair of genes associated with that cell type. After obtaining each cell type’s average, we took the mean of these averages which we refer to as the cell type dispersion for that tree. We note that the absolute value depends on the topology of the tree. To establish a reference for this measure, we bootstrapped a null distribution by preserving the same tree topology and randomly assigning genes to leaves and calculating the resulting dispersion. We performed this 10,000 times to create a reference null distribution. We compared the true dispersion value to this null, and computed a bootstrapped p-value. In both cases (with and without cell cycle genes), the observed average cell type dispersion is less than that for all permutations (p-value=0.0001, bootstrapping), indicating that genes associated with the same cell type were significantly more likely to cluster together on the observed scL-score hierarchy than would be expected under random clustering.

      We have updated the text to reflect these quantitative comparisons:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type-specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b).

      Author response image 6.

      (n) Clarify how the distance between genes is computed for clustering; the current method is hard to interpret.

      We appreciate the reviewer’s suggestion to clarify how distances between genes were computed for hierarchical clustering. To clarify this point, we have added the following to the text:

      “Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      (o) Quantify similarity between cNMF and L-metric clusters.

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2, now in updated Figure S4.4c). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See updated Figure S4.4 (formerly Figure S4.2)).

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (p) In Figure 3e, clarify what the orange lines represent.

      We have updated the text:

      “Per-cell expression scatterplots of the two pairs of genes shown in b). The y-axis of each is the per-cell expression of Eogt. The x-axis is the per-cell expression of Pax6 (left) or Rfx4 (right). R is Pearson’s r, scL is scL-score. Count data is shown in black; smoothed 2D densities are shown in orange.”

      (q) Ensure figures and panels follow the text order. For example, Figure S1.3a is referenced earlier than 1.1 and 1.2.

      We appreciate the reviewer’s attention to detail. We have changed the order of these figures so that their reference in the text follows their numeric order.

      (r) Typo in "by covariation with any other cell type (Figure S1.1e)" did you mean S1.1.c?

      We have updated the text with this change.

      (s) For circularity (Figure 6), justify the convex-hull-based measure; thin protrusions can distort interpretation. Consider the volume difference between the convex hull and the original shape.

      We thank the reviewer for this helpful suggestion. We tested the difference method suggested by the reviewer, as well as several other methods of clustering and calculating circularity. We determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and text.

      Summary

      This work delivers a rich spatial dataset and introduces creative computational tools. The main limitations lie not in the data but in the clarity, justification, and validation of the quantitative methods. We suggest strengthening these aspects by adding formal definitions, parameter justification, benchmarking, robustness tests, and controlled interpretations. This will improve the manuscript's impact and reproducibility of the methods described, besides making it easier to understand for the readers.

      We thank the reviewer for their kind assessment and also for their many insightful comments for improvement. We feel the revised manuscript is greatly improved because of them.

      Reviewer #3 (Recommendations for the authors):

      In my view, this manuscript is well-designed and clearly written, and supports all of the claims made. I have no suggestions for major revisions for this manuscript; rather, I would suggest the following as outstanding questions for future investigation:

      We thank the reviewer for a careful reading of our manuscript and the several interesting suggestions and useful references in the literature. Including the discussion of these ideas in the text (see below) has improved the flow and scope of the manuscript.

      (1) On the NMP fate, bifurcation has been studied extensively in gastruloids and related structures (see, for example, Underhill et al 2023; Bolondi et al 2024). Can this dataset from Triandafillou and colleagues reveal new regulatory hierarchies in this process? This seems possible in principle, but I was not able to reach this interpretation (for example, it seems the authors interpret the clustering in Figure 3H as reflecting spatial patterns rather than a regulatory hierarchy).

      We agree that this is an exciting implication of the work, but we feel that with the current panel (which was originally chosen primarily to type cells and not to infer regulatory structure) we would not be able to comment on this. However we feel that with a larger set of genes such inference may be possible, so we took your suggestion in point 3) and applied the L-metric to a previously existing single-cell dataset. We looked for possible regulatory structure, and found at least one interesting case where a transcription factor showed strong co-expression with two previously unconnected genes. While more specific analyses and experimental validation would be required to establish a direct relationship, we feel that this suggestion by the reviewer represents an important potential future application of this methodology, so we’ve included a new figure and explanatory text to address this possibility:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.

      (2) On the formation of endothelial clusters ('blood islands'), this process has been observed in gastruloids (Rossi et al 2022). The observation of endoderm-associated endothelium in gastruloids is interesting, but it was also not clear how the authors interpret this finding. Are there two separate endothelial differentiation paths captured here? Or is there one path, and only some migrate towards the endoderm? The authors seem to raise possibilities, and it was slightly unclear on my reading how they interpreted their findings.

      We thank the reviewer for pointing us to this paper. We think it’s especially interesting that we also see close association between endoderm and endothelial precursors, especially given that the protocol used in the referenced paper was designed to generate blood precursors (treatment with VEGF, bFGF and ascorbic acid). In light of the comments of other we re-did this analysis — Author response image 7 shows the updated list of differentially expressed genes.

      Author response image 7.

      Some of the spatially differentially expressed genes are linked to signalling, and likely reflect overall signalling differences between the anterior (where the somite-associated endothelial cells are) and the posterior (where the endoderm-associated endothelial cells are). For example, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA is higher in the anterior). Pecam1 is involved in adhesion and cell morphology, and perhaps is higher in cells interacting with endoderm due to tighter packing/association; the same could also be true of Cdh5.

      While none of these answer the reviewer’s questions about the origin of the cells (and whether it is common), we found another instance of differential endothelial populations in embryo models: in Veenvliet 2020, they find that some endothelial cells have a mesodermal (somitic) origin. Thus we may be seeing a similar phenomenon in our samples. We have updated the text to reflect these additional lines of evidence and to clarify how we think the two populations may differ (while acknowledging that we lack the tools to confidently assign cell of origin or functional differences with this technique):

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.”

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      (3) Finally, a broader question about the analytical framework. The authors emphasize that the L-metric is parameter-free; however, much of their analysis still appears to rely on baked-in priors about known marker genes for cell type assignment. Is there a way to extend their analysis to infer cell types directly from the structure of the L-metric? The hierarchical clustering in e.g., Figure 3G suggests something like this: the hierarchy mostly (but not exactly) follows the marker gene annotation. Does this suggest that the cell type labeling should be revisited?

      Once of the initial motivations behind the creation of the L-metric (now L-score) was to have a more reliable and specific way of quantifying the interaction between genes that are known in the literature to be cell type markers — with more sophisticated (and noisy) methods of analysis like single-cell RNA sequencing, we found that there was substantial variation in the specificity and ubiquity of so-called ‘marker genes’, and that their usefulness often depended on context. We appreciate that the reviewer raises this point as well, and we think that scL-score analysis, such as that exemplified in Figure S4.4, can help identify new marker genes, or at the very least distinguish the biological context in which a marker gene is useful.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      We thank the two reviewers for their constructive comments regarding our manuscript. Below is our point-by-point response (un-bold) to each reviewer’s comments together with the experiments we propose to carry-out in order to strengthen our main conclusions.

      Reviewer #1

      In this study, the authors investigate the effects of pharmacological inhibition of the spliceosome using the SF3B1 inhibitor pladienolide B in models of platinum-resistant non-small cell lung cancer (NSCLC). Using a combination of cell lines, platinum-resistant derivatives, and patient-derived xenograft (PDX) models, the authors show that spliceosome inhibition sensitizes platinum-resistant tumors to treatment and leads to increased DNA damage accumulation and impaired DNA damage response signaling. Transcriptomic analyses indicate that transcripts encoding DNA damage regulators are particularly sensitive to alternative splicing perturbations, and selected mechanistic experiments suggest involvement of specific regulators such as MLH3. The study further explores links between splicing inhibition, transcriptional activity, and cell cycle progression. Overall, the manuscript presents extensive datasets across multiple experimental systems and provides strong evidence that spliceosome inhibition can sensitize platinum-resistant tumors to DNA damage. However, several aspects of the mechanistic interpretation, data consistency, and presentation require clarification or strengthening to fully support the central claims.

      We thank the reviewer for his/her positive comments on our work and we agree that the manuscript requires clarification to support our main claims.

      Major comments* *

      Conceptual clarity and synthesis of mechanistic model

      The manuscript presents multiple mechanistic observations-including altered splicing of DNA repair genes, increased DNA damage accumulation, transcriptional perturbation, and cell cycle changes-but these are not integrated into a coherent conceptual framework. While it is reasonable that not all mechanistic details are fully resolved, the current presentation leaves the reader uncertain about the relative contributions of these processes. A clearer synthesis of the proposed mechanism, possibly including a summary model figure, would substantially improve the conceptual clarity of the study.

      __We thank the reviewer for his/her comment. We acknowledge that our study provides multiple mechanistic observations. However, all these observations converge towards a more global mechanism by which pladienolide B induces cell death in NSCLC. Hence, we demonstrate that pladienolide B, which targets the SF3B1 protein, a core component of the spliceosome machinery, induces a massive shutdown of the DNA Damage Response signaling pathway, both at the transcriptional and splicing levels, which leads to enhanced genomic instability and cell death notably in NSCLC cells with acquired resistance to platinum salts. More specifically, we identified ATR, DNA-PKcs and MLH3 as novel targets of SF3B1 in NSCLC cell lines, as well as more importantly, in NSCLC Patient-Derived Xenografts. To our knowledge, and as also highlighted by the reviewer, this is the first evidence that pladienolide B slows-down tumor growth by negatively impacting DNA Damage Response in patient-derived xenografts. We agree with the reviewer that providing a summary model figure would improve the clarity of the study. We will provide such a graphical abstract in our revised manuscript. __

      Biological specificity of platinum-resistant cell sensitivity

      A central premise of the study is that platinum-resistant cells exhibit enhanced sensitivity to spliceosome inhibition. However, in several experiments (e.g., cell cycle analysis in Figure 2A), similar responses to pladienolide B appear to occur in both platinum-sensitive and resistant cells. This observation complicates the interpretation that resistant cells exhibit uniquely distinct vulnerability. The authors should clarify how these findings align with the proposed model and more explicitly distinguish shared versus resistance-specific responses.

      __We thank the reviewer for this remark. In this study, we did not want to claim that platinum-salts resistant cells exhibit unique vulnerability to pladienolide B as we fully agree with the reviewer that pladienolide B also exhibits cytotoxic effects in parental sensitive cells, although with a delayed kinetic. We acknowledge that the way we introduced our results section in the former manuscript could have contributed to such a misunderstanding. Rather, we propose a model in which pladienolide B induces cell death in NSCLC cells by massively impairing the expression of key components of the DNA damage and repair signaling pathways, at the transcriptional and/or splicing level. As NSCLC cells with acquired resistance to platinum salts are likely more dependent to these pathways for their survival than parental cells, this explains enhanced susceptibility of these resistant cells to pladienolide B-induced cell death. We think that these results highlight spliceosome targeting compounds as an alternative therapeutic strategy in NSCLC patients who escape chemotherapy. According to the remark of the reviewer, we will modify the way we introduce our results and we will clarify all these points in the discussion section of the revised manuscript. __

      Heterogeneity in PDX responses and lack of platinum-sensitive controls

      The PDX experiments represent a major strength of the study. However, resistant tumors display heterogeneous responses to pladienolide treatment, suggesting the presence of additional determinants of sensitivity.

      __We thank the rewiever for this remark. In this study, we used seven distinct NSCLC PDXs We initially selected these PDXs based on their low responsive rate to cisplatin rather than their mutational status. As discussed in the discussion section, we did not find any common mutation(s) that could predict the differential response of these PDXs to pladienolide B. We only noticed that the LCIM10 PDX, which is the most responsive to pladienolide B, exhibits ATRX mutation. ATRX has been shown to protect stalled replication forks from collapsing. As we found that pladienolide B induces early replicative stress in NSCLC cells, it is tempting to speculate that ATRX mutation might interfere with the replicative stress response and potentiates pladienolide B’s cytotoxic effects. Noteworthy, LCIM10 PDX also displayed higher basal levels of both P-DNA-PKcs(Ser2056) and P-ATR(Thr1989) proteins as compared to LCF26, ML1 and LCIM1 PDXs that were less responsive to pladienolide B (data to be added in the revised version of the manuscript as supplementary data). Therefore, and although this remains to be further clarified, this suggests that NSCLC patients with higher basal level of replicative stress, such as those who escape chemotherapy, might be more susceptible to SF3B1 inhibition. __

      Including platinum-sensitive PDX tumors, if available, would provide valuable baseline comparison and strengthen interpretation of resistance-specific effects. If not feasible, the limitations should be acknowledged and discussed.

      To our knowledge, and as also mentioned by the reviewer, our study provides the first demonstration of the effects of pladienolide B on the growth of NSCLC PDXs. We think that these results pave the way for further investigations in additional NSCLC PDXs and we agree with the reviewer that adding platinum-sensitive PDX tumors would be interesting. However, due to cost limitations and time constraints, we will be unable to repeat them for this specific study. Based on the results we already obtained in 7 platinum salts-resistant PDXs as regard to their response heterogeneity, one might also speculate that the comparisons between numerous sensitive and resistant PDXs should be complicated as sensitive PDXs might also display distinct mutational status. According to the remark of the reviewer, we will discuss these aspects in the discussion section of the revised manuscript.

      Consistency between pharmacological inhibition and genetic depletion

      In Figure 4, the authors compare pladienolide treatment with SF3B1 knockdown to demonstrate target specificity. However, the effects observed with the two perturbations are not entirely consistent-for example, pladienolide affects phosphorylation of DNA-PKcs, while SF3B1 knockdown appears to produce broader effects at both protein and mRNA levels. Additionally, differences are observed between resistant cell lines in the response to SF3B1 knockdown. These discrepancies should be addressed and discussed, as they may reflect mechanistic differences between acute pharmacological inhibition and genetic depletion.

      __We thank the reviewer for this remark and we agree that acute pharmacological inhibition using pladienolide B and genetic depletion using SF3B1 siRNA might produce distinct effects as they do not exhibit the same mechanism of action. Hence, pharmacological inhibitors target the protein while siRNA targets mRNA with effects depending on the basal mRNA level/stability. This could explain why we did not exactly observe the same effects on ATR/DNA-PKcs mRNA and protein levels using both approachs in H460/A549 parental and resistant cells (Fig 3-4). For pladienolide B treatment, the effects were analyzed at “early” timepoint [i.e. after 4-6 hours treatment (Fig 3a-d)] or at a “later timepoint” [i.e. after 24-48 hours treatment (Fig 3e-f)]. At early timepoint, pladienolide B induced replicative stress that correlated with DNA-PKcs phosphorylation, while at later timepoints it decreased ATR and DNA-PKcs mRNA/protein levels. The biological consequences of SF3B1 knock-down were analyzed after 72 hours of transfection in resistant cells. This could explain why SF3B1 knock-down appears to produce broader effects at both protein and mRNA levels. However, we acknowledge the existence of differences between SF3B1 knocked-down-H460 and -A549 cells as regard to the downregulation of ATR and DNA-PKcs that occurs at the mRNA and/or protein level depending on the cell line. Again, this could depend on the time as well as the efficiency of SF3B1 knock-down which was more prominent in H460 resistant cells compared to A549 cells. Nevertheless, and despite these mechanistic differences in cell lines and between pladienolide B and SF3B1 siRNA, our results identify ATR and DNA-PKcs as novel targets of SF3B1 in NSCLC cell lines, including cells with acquired resistance to cisplatin, as well as more importantly in NSCLC PDXs (Fig 9c). To our knowledge, this is the first evidence that pladienolide B or SF3B1 knock-down negatively targets ATR or DNA-PKcs in solid tumors. Owing to the crucial role played by both kinases in the maintenance of genomic stability in cancer cells, we think that this result is of importance. Nevertheless, as suggested by the reviewer, we propose to acknowledge and discuss more in details the discrepancies between cellular models and pladienolide B / SF3B1 knock-down in the revised version of the manuscript. __

      Selection and interpretation of splicing-sensitive transcripts

      Transcriptomic analyses in Figure 5 identify both shared and differential splicing changes between sensitive and resistant cells. However, much of the analysis focuses on transcripts that are commonly affected in both conditions, rather than those uniquely altered in resistant cells. Given that the central phenotype is resistance-specific sensitivity, transcripts uniquely mis-spliced in resistant cells may represent more informative candidates. The authors should clarify the rationale behind focusing on shared events and discuss the implications of resistance-specific versus common splicing changes.

      We thank the reviewer for this remark. Indeed, after obtaining RNA-Seq data, we initially looked for genes which differential expression and/or splicing upon pladienolide B treatment could only be observed in resistant cells but not parental ones. We focused first on the 121 genes belonging to the DNA repair pathways full network (WikiPathway WP4946) since Gene-Ontology analyses based on RNA-Seq data demonstrated enrichment of genes involved in DNA metabolic process, which includes DNA repair, among genes down-regulated after pladienolide B treatment (Fig S3). The focus on DNA repair was also justified by our observation showing accumulation of DNA double strand breaks upon pladienolide B treatment in H460 resistant cells (Fig 2d-f). Doing this comparison, we showed that pladienolide B regulates the expression of 41 (33%) and 47 (39%) genes of this network in H460S and H460R cells respectively, ____which were mostly down-regulated in both H460S (33/41) and H460R (33/47) cells (Table 3). Fifteen genes involved in all DNA repair processes were found to be specifically down-regulated in H460R cells upon pladienolide B treatment, including PARP-1. As a whole, these results demonstrated that pladienolide B down-regulates the expression of numerous DNA repair genes in both NSCLC parental and resistant cells. As discussed above, we propose that the enhanced sensitivity of resistant cells to pladienolide B is related to their increased dependency for survival to functional DNA repair pathways.

      When differentially spliced genes were considered, and focusing on exon skipping events, as they were the more prominent (Fig 5e), we found that pladienolide B regulates the splicing of 107 and 87 genes of the WikiPathway WP4946 in H460 parental and resistant cells, respectively (Table 5). Forty five genes were predicted to be regulated in both cell lines. Only 3 genes, namely DCLRE1C, POLD3 and PNKP, were predicted to be differentially spliced upon pladienolide B treatment in H460R cells only. Trying to increase the number of genes to study, we extended our analysis to genes belonging to another DNA repair database (Human DNA Repair Genes, Resources from Wood laboratory, UT MD Anderson), and we found six additional genes, namely MLH3, MSH5, RAD54L, EME1, SETMAR and SMC6, that were also predicted to be differentially spliced upon pladienolide B treatment in H460R cells only. However, four of these exon skipping events (i.e. PNKP-Ex9, DCLRE1C-Ex11, SETMAR-Ex2, SMC6-Ex6) were not validated and we did not observe clear difference between H460 resistant and parental cells for the others (Fig 5g and Fig S6). Skipping of MLH3-Ex8 was the sole event displaying a slight difference between both cell lines, mainly in term of kinetic of recovery. This is why we decided to further analyze this specific splicing event. Noteworthy, we focused only on genes involved in DNA damage and repair signaling pathways. Therefore, we cannot exclude that genes involved in other biological processes might be differentially transcribed or spliced in response to pladienolide B in H460 parental and resistant cells.

      Transient nature of splicing effects

      The authors report transient alternative splicing effects upon prolonged pladienolide treatment. This observation is counterintuitive, as continued spliceosome inhibition might be expected to produce cumulative splicing defects. While the authors reference studies showing that transient inhibition can produce lasting effects, the current observations involve continuous exposure. This apparent discrepancy should be clarified and discussed.

      We thank the reviewer for this remark. We agree with him/her that the transient effect of pladienolide B on most of the splicing events we studied was unexpected as pladienolide B treatment was prolonged. We do not have a clear explanation for that. One possibility is that pladienolide B is not stable in the cell culture supernatant and is degraded rapidly. Another, not exclusive, possibility relies on the structure of studied transcripts. As discussed in the discussion section, it is possible that the nature of the transcripts involved in DNA damage response/DNA repair which have a long length, a large number of small exons per transcript, and an elevated number of introns could explain why they recover very rapidly from pladienolide B inhibition. Alternatively, upon SF3B1 inhibition, compensatory regulations by other splicing factors might occur. We will discuss these aspects in the revised version of the manuscript.

      Use of unrelated cell lines in reporter assays

      The DNA damage reporter assays (Figure 6A-D) appear to be performed in cell lines not directly linked to platinum sensitivity or resistance. Given the central importance of resistance-specific responses, repeating key reporter assays in both sensitive and resistant paired models would strengthen the conclusions.

      __We initially engineered these cellular models to assess the role of SRSF2, a splicing factor, in DNA repair ____(Khalife M. et al., NAR Cancer, 2025). In these models derived of either H1299 or A549 NSCLC cell lines, DNA double strand breaks (DSBs) are produced after cleavage by the SceI enzyme and the efficiency of repair is assessed based on the expression of either GFP (for homologous recombination) or CD4 (for c-NHEJ). These cellular models were difficult to engineer and to work with as they require a first stable transfection with either PBL174 pDR-GFP (for HR analysis) or PBL230 (for c-NHEJ) plasmid followed by a second transient transfection with the PLBL133 plasmid that encodes SceI enzyme. As an example, we tried to generate stable A549 cells with PBL174 plasmid but we never succeeded. The reverse was true for H1299 cells and transfection with PBL230. In addition, during the time course of this project, we also tried to transiently transfect plasmid encoding MLH3 or MLH3 protein devoid of exon 8-encoding amino acids in H460 or A549 resistant cells but we never obtained good efficiency of transfection, although we tested several transfection reagents. So, it looks like that resistant cells are hardly transfectable. This is why repeating these experiments in resistant models will not be possible. However, and although we agree with the reviewer that the cell lines we used were not directly linked to platinum sensitivity or resistance, the idea behind these experiments was to test whether pladienolide B could have a general impact on DNA repair by homologous recombination or c-NHEJ using these SceI-induced DSBs systems that are already widely used in the DNA repair field. As shown in Fig 6b and 6d, pladienolide B prevented DNA repair in these cellular models, confirming that it widely negatively impacts DNA repair pathways. __

      Interpretation of MLH3 splicing results

      In Figure 7, differences between pharmacological inhibition and SF3B1 knockdown in MLH3 exon 8 regulation are not entirely consistent.

      We do not strictly agree with this comment. __As also discussed above, the differences seen between pharmacological inhibition of SF3B1 using pladienolide B and SF3B1 knock-down in term of MLH3-exon 8 regulation could be related to differences in term of mechanism of action and/or the fact that the effects of SF3B1 knockdown were analyzed after 72 hours treatment (as we obtained the best knock-down efficiency at this time point) while those of pladienolide B were studied between 6 to 48 hours treatment. However, as illustrated in Fig 7a and 7c (right panel), we showed that both pladienolide B and SF3B1 knock-down promote MLH3-exon 8 exclusion in H460 and A549 resistant cells. The effects of SF3B1 knock-down were less pronounced in A549R cells which could be consistent with the decreased efficiency of SF3B1 knockdown as depicted in Fig 7c (left panel). Pladienolide B also promoted MLH3-Ex8 exclusion in H460 and A549 parental cells but the recovery was faster in parental cells as compared to resistant cells (Fig 7a). __

      Furthermore, inclusion levels of regulated and non-regulated exons appear similarly correlated with SF3B1 expression, potentially weakening the argument for exon-specific regulation. These observations should be clarified.

      __As regard to MLH3-exon 5 exclusion, we observed its exclusion in A549 parental and resistant cells upon pladienolide B treatment, while this was not observed in H460 parental and resistant cells (Fig 7a-b). In SF3B1 knocked-down H460R and A549R cells, the exclusion of MLH3-exon 5 was seen in A549R cells. When analyzing MLH3 exon 8 or exon 5 usage in lung adenocarcinoma patients (Fig 7e), we agree with the reviewer that there was a significant correlation between SF3B1 mRNA level and MLH3 exons 5 and 8 usage. Therefore, these and our data indicate that both MLH3 exons 5 and 8 could be regulated by SF3B1 in NSCLC although differences might occur depending on the cell line. To make this point clearer, we will clarify the text of the results for Figure 7 and our conclusion. __

      Combination treatment logic

      In Figure 8, co-treatment experiments with pladienolide B and cisplatin are performed primarily in resistant cells. Performing similar experiments in platinum-sensitive cells would provide an important reference point to distinguish additive versus resistance-specific effects.

      __We thank the reviewer for this important remark. The objective of figure 8 was to investigate whether pladienolide B that induces a shutdown of numerous DNA damage response-related genes could resensitize NSCLC cells with acquired resistance to cisplatin-induced apoptosis. The results presented in Figures 8a and 8b show that this is indeed the case. However, and according to the remark of the reviewer, we propose to illustrate in the revised version of the manuscript the results of the co-treatment experiments in sensitive parental cells also. Indeed, we already had the results for H460 parental cells. They did not show any additive or synergistic effects of the combination in these cells. Rather adding pladienolide B to cisplatin tended to decrease apoptosis as compared to cisplatin alone. We will reiterate these experiments in A549 cells. If confirmed, and to a translational point of view, these results would support the idea that treating NSCLC patients who relapse from chemotherapy with a combination of platinum salts and pladienolide B could provide therapeutic benefits, whereas NSCLC patients that primary respond to platinum salts could less benefit from this combination. __

      Minor comments

      • In Figure 1, pladienolide treatment in platinum-sensitive cells appears to plateau at approximately 50% cell killing. Extending the concentration range may help clarify whether maximal efficacy was reached.

      __ We agree with this remark. We will reiterate our MTS experiments increasing the dose of pladienolide B. __

      In Figure 1H, the difference between 2.5 mg/kg and 5 mg/kg pladienolide in PDX models appears disproportionately large relative to the dose change. This should be discussed or experimentally clarified.

      We agree with the reviewer but these are the results we obtained. In Figure 1h, we illustrated the probability of progression based on Relative Tumor Volume (RTV) = 2. The difference between the two doses was less, although remaining significant, when considering RTV = 4 as a marker of progression (Fig S2d).

      In Figure 5, differential expression and splicing analyses are presented using multiple cutoffs (e.g., log₂FC > 0.4 and >1; ΔPSI thresholds). This introduces redundancy and may obscure key findings. A single well-justified cutoff would improve clarity.

      We agree with the reviewer that in initial Figure 5 we provided graphs illustrating different analysis thresholds based on our transcriptomic analyses. We will select one cut-off for Differentially Expressed Genes [absolute Log2Fold Change ≥ 0.4 and p ≤ 0.05 (Fig 5a)] and Differentially Spliced Genes [absolute percent splice in (PSI) ≥ 0.2 and p ≤ 0.05 (Fig 5e)]. We will remove Fig 5b and 5d for more clarity.

      Several isoform-specific RT-PCR gels are difficult to interpret due to low image clarity. Improving gel presentation or focusing on key timepoints would strengthen data readability.

      We will improve gel presentation.

      Some figure panels appear redundant, showing similar datasets under different analysis thresholds.

      We agree with the reviewer that in Figure 5 we provided graphs illustrating different analysis thresholds based on our transcriptomic analyses. We will select one cut-off for Differentially Expressed Genes [absolute Log2Fold Change ≥ 0.4 and p ≤ 0.05] and for Differentially Spliced Genes [absolute percent splice in (PSI) ≥ 0.2 and p ≤ 0.05]. We will remove Fig 5b and 5d for more clarity.

      A graphical summary model illustrating the proposed mechanism would improve reader comprehension.

      We agree with the reviewer. We will provide such a graphical abstract in the revised version of our manuscript.

      In Figure 3D, representative images of γH2AX foci appear visually similar across conditions, whereas quantification shows large differences. The authors should ensure that representative images accurately reflect quantified trends and clarify selection criteria for displayed images.

      We agree with the reviewer. We will select additional images illustrating more the differences we highlighted after quantification of more than 500 nuclei (Figure 3D, right panel).

      Reviewer #1 (Significance (Required)):

      This study represents a comprehensive investigation of spliceosome inhibition as a therapeutic strategy to overcome platinum resistance in NSCLC. The use of resistant cell lines, PDX models, and functional reporters provides strong experimental depth. The most compelling aspects of the study include the demonstration that spliceosome inhibition enhances DNA damage accumulation and sensitizes resistant tumors to platinum-based therapies. However, several mechanistic interpretations require clarification, and data presentation could be streamlined to improve logical coherence.

      We thank the reviewer for his/her constructive remarks. We hope that our answers and the additional works we now propose to carry-out in order to revise the manuscript will get agreement to him/her and will strengthen our main claims

      Advance

      The work provides evidence that targeting spliceosome function-specifically via SF3B1 inhibition-can sensitize platinum-resistant tumors to DNA-damaging agents. To my knowledge, this represents one of the first comprehensive demonstrations that spliceosome-targeting compounds can effectively overcome acquired platinum resistance in solid tumor models. The study also contributes to the emerging understanding that DNA damage response transcripts may represent particularly sensitive targets of splicing perturbation.

      Audience

      The study will be of interest to researchers in RNA biology, cancer therapeutics, DNA damage response, and translational oncology. It is particularly relevant to scientists investigating therapeutic vulnerabilities in drug-resistant cancers and those exploring RNA processing as a therapeutic target.

      Expertise I have expertise in RNA biology, alternative splicing and cancer models. My expertise is more limited in pharmacological dosing strategies and some aspects of in vivo xenograft modeling.

      Keywords: RNA biology, alternative splicing, spliceosome function, cancer biology, DNA damage response, transcriptomics

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      Summary

      In the manuscript by Jamal-El-Hussein et al., the authors demonstrated that pladienolide B, an inhibitor of splicing factor 3B subunit 1 (SF3B1), inhibits cell viability in NSCLC cells and PDXs with resistance to platinum-based chemotherapy. Mechanistically, they identified that pladienolide B regulates splicing events of genes associated with DNA damage signaling and repair, specifically through exon skipping of MLH3, thus increasing vulnerability to chemotherapy. This study provides therapeutic insights into combining pladienolide B with chemotherapy for overcoming therapy resistance.

      __Major comments __

      The authors showed that both H460R and A549R cell lines are sensitive to pladienolide B; however, they exhibit distinct molecular responses. For example, SF3B1 knockdown downregulates the mRNA expression of ATR and PRKDC in H460R cells, while the expression of these genes remains unchanged in A549R cells (Fig. 4b). Furthermore, exon 8 skipping of MLH3 is not observed in A549R cells with pladienolide B treatment (Fig. 7a). These discrepancies suggest that the two cell lines respond to pladienolide B through different mechanisms. The authors should also perform a bulk RNA-seq on A549 cell line to better understand the different behavior of the two cell lines.

      We disagree with the remark regarding Fig 7a as this figure demonstrated MLH3 exon 8 skipping in A549R cells also after 6 and 24 hours pladienolide B treatment. __Although we agree with the reviewer that some differences exist in term of molecular mechanisms regulated by pladienolide B or SF3B1 knock-down in H460R and A549R cell lines, we identified in this study ATR, DNA-PKcs and MLH3 as novel targets of SF3B1 in both cell lines as well as more importantly in NSCLC PDXs. To our knowledge, this is the first evidence that SF3B1 inhibition negatively impacts ATR or DNA-PKcs and regulates MLH3 alternative splicing in solid tumors. In addition, we found a significant correlation between SF3B1 and DNA-PKCs or ATR protein levels in 77 NSCLC cell lines (Fig 4c) which supports a close relationship between these proteins in lung cancer. As also discussed in the response to reviewer 1, the differences between pladienolide B and SF3B1 knock-down could be related to the distinct timepoints at which we analyzed their effects as well as to the different mechanisms of action between pharmacological and siRNA inhibition. Nevertheless, we propose to acknowledge and discuss more in details the discrepancies between cellular models and pladienolide B or SF3B1 knock-down in the revised version of the manuscript. __

      The authors' central claim is that platinum-based chemotherapy-resistant cells are more sensitive to pladienolide B treatment.

      __We thank the reviewer for this remark. In this study, we did not want to claim that platinum-salts resistant cells exhibit unique vulnerability to pladienolide B as we fully agree with the reviewer that pladienolide B also exhibits cytotoxic effects in parental sensitive cells, although with a delayed kinetic. We acknowledge that the way we introduced our results section in the former manuscript could have contributed to such a misunderstanding. Rather, we propose a model in which pladienolide B induces cell death in NSCLC cells by massively impairing the expression of key components of the DNA damage and repair signaling pathways, at the transcriptional and/or splicing level. As NSCLC cells with acquired resistance to platinum salts are likely more “addict” to these pathways for their survival than parental cells, this might explain enhanced susceptibility of these resistant cells to pladienolide B-induced cell death. We think that these results highlight spliceosome targeting compounds as an alternative therapeutic strategy in NSCLC patients who escape chemotherapy. According to the remark of the reviewer, we will modify the way we introduce our results and we will clarify all these points in the discussion section of the revised manuscript. __

      In their bulk RNA-seq analysis, the authors selected genes commonly regulated by pladienolide B in both parental and resistant cells for further validation. This approach raises a critical concern: the observed sensitivity to pladienolide B may already be present in parental cells rather than representing a mechanism uniquely acquired by the resistant cells. To substantiate their central claim, the authors should analyze the differentially expressed genes between parental and resistant cells to identify resistance-specific molecular alterations that may confer enhanced sensitivity to pladienolide B, thereby providing a more mechanistically rigorous basis for their conclusions.

      We thank the reviewer for this remark and we agree with him/her. Indeed, after obtaining RNA-Seq data, we initially looked for genes which differential expression and/or splicing upon pladienolide B treatment could be only observed in resistant cells but not parental ones. We initially focused on the 121 genes belonging to the DNA repair pathways full network (WikiPathway WP4946) since Gene-Ontology analyses demonstrated enrichment of genes involved in DNA metabolic process, which includes DNA repair, among genes down-regulated after pladienolide B treatment (Fig S3). The focus on DNA repair was also justified by our observation showing accumulation of DNA double strand breaks upon pladienolide B treatment in H460 resistant cells (Fig 2d-f). Doing this comparison, we showed that pladienolide B regulates the expression of 41 (33%) and 47 (39%) genes of this network in H460S and H460R cells respectively, ____which were mostly down-regulated in both H460S (33/41) and H460R (33/47) cells (Table 3). Fifteen genes involved in all DNA repair processes were found to be specifically down-regulated in H460R cells upon pladienolide B treatment, including PARP-1. As a whole, these results demonstrated that pladienolide B down-regulates the expression of numerous DNA repair genes in both NSCLC parental and resistant cells. However, and as discussed above, we propose that the enhanced sensitivity of resistant cells to pladienolide B is related to their increased dependency for their survival to functional DNA repair pathways.

      __When differentially spliced genes were considered, and focusing on exon skipping events, as they were the more prominent (Fig 5e), we found that pladienolide B regulates the splicing of 107 and 87 genes of the WikiPathway WP4946 in H460 parental and resistant cells, respectively (Table 5). Forty five genes were predicted to be regulated in both cell lines. Only 3 genes, namely DCLRE1C, POLD3 and PNKP, were predicted to be differentially spliced upon pladienolide B treatment in H460R cells only. Trying to increase the number of genes to study, we extended our analysis to genes belonging to another DNA repair database (Human DNA Repair Genes, Resources from Wood laboratory, UT MD Anderson), and we found six additional genes, namely MLH3, MSH5, RAD54L, EME1, SETMAR and SMC6, that were also predicted to be differentially spliced in H460R cells only. However, four of these exon skipping events (i.e. PNKP-Ex9, DCLRE1C-Ex11, SETMAR-Ex2, SMC6-Ex6) were not validated and we did not observe clear difference between H460 resistant and parental cells for the others (Fig 5g and Fig S6). Skipping of MLH3-Ex8 was the sole event displaying a slight difference between both cell lines, mainly in term of kinetic of recovery. This is why we decided to further analyze this specific splicing event. Noteworthy, we focused only on genes involved in DNA damage and repair signaling pathways. Therefore, we cannot exclude that genes involved in other biological processes might be differentially transcribed or spliced in response to pladienolide B in H460 parental and resistant cells. __

      The authors conclude that pladienolide B treatment correlates with activation of DNA-PKcs signaling followed by a shutdown of ATR and DNA-PKcs pathways. However, the data presented do not fully support this interpretation. P-ATR levels are already elevated in resistant cells and remain unchanged following pladienolide B treatment (Fig. 3a). However, prolonged pladienolide B treatment leads to decreased total ATR protein and mRNA expression (Fig. 3e-f), suggesting that pladienolide B maintains an initial constitutive ATR activation followed by transcriptional downregulation. Since pladienolide B is a splicing inhibitor, the authors should determine whether ATR and DNA-PKcs mRNA downregulation is a direct consequence of aberrant splicing of their transcripts, or a non-specific effect of prolonged cellular toxicity. To strengthen their mechanistic conclusions, the authors should perform time-course experiments to establish the temporal relationship between these signaling events, analyze splicing changes specifically in ATR and DNA-PKcs transcripts (are they in the differential genes from the bulk RNA-seq analysis?), and also check the downstream targets of the ATR and DNA-PKcs signaling.

      __We thank the reviewer for his/her comment. In our RNA-Seq analyses, we did not recover PRKDC among the differential genes expressed or spliced upon pladienolide B treatment whatever the cell line. However, PRKDC was also not in the full RNA-Seq data list of not significant genes. Therefore, it remains unclear whether PRKDC splicing could account for the decrease of PRKDC mRNA level upon pladienolide B treatment. Concerning ATR, it was not in the RNA-Seq data list of the genes significantly up- or down-regulated upon pladienolide B treatment in either H460 parental or resistant cells. However, transcriptomic analyses were performed after 8 hours pladienolide B treatment while the decrease of ATR mRNA was observed after 24 hours (Fig 3f). Regarding splicing, ATR was predicted to be spliced, skipping of exon 30, upon pladienolide B treatment in both H460 parental and resistant cells. We validated this splicing event after 8 hours treatment with pladienolide B in both H460 cellular models but we did not analyze this splicing event at later timepoints, nor in the A549 parental or resistant cells. Therefore, and according to the remarks of the reviewer, we propose to deepen the temporal relationships between all these signaling events by performing time-course experiments for analysis of ATR exon 30 splicing by RT-PCR, ATR and PRKDC mRNA levels by RT-qPCR, and expression of downstream targets of ATR and DNA-PKcs, such as P-CHK1(Ser345) or P-RPA32(Ser4/8) by immunoblotting. __

      The authors report that skipping of exon 8 of MLH3 leads to the complete absence of the protein (Fig. 7d). However, this observation needs further clarification, as at least two alternative explanations exist. First, the antibody used to detect MLH3 may specifically recognize an epitope encoded by exon 8 or downstream exons, in which case the loss of signal would reflect antibody incompatibility rather than true protein absence.

      __We thank the reviewer for this remark. The anti-MLH3 antibody recognizes the C-terminal part of the MLH3 full-length protein (between amino acids 1228-1453). MLH3 exon 8 is 72 base pair and does not encode for the amino acids recognized by the anti-MLH3 antibody. In ENSEMBL, the MLH3-201 transcript encodes for the full-length protein (1453 amino acids) and the MLH3-202 transcript encodes for a MLH3 protein (1429 amino acids) devoid of the amino acids encoded by exon 8. Nevertheless, the two products have the same C-terminus recognized by the anti-MLH3 antibody used in this study. So, the loss of the signal depicted in Figure 7d is not due to antibody incompatibility. __

      Second, skipping of exon 8 may introduce a premature stop codon, triggering nonsense-mediated mRNA decay (NMD) and consequent loss of the transcript. To distinguish between these possibilities, the authors should perform qPCR using primers targeting sequences both upstream and downstream of the skipped exon, as well as consider NMD inhibition experiments, to clarify whether the observed protein loss occurs at the transcriptional or translational level.

      __As shown in Figure 8f, we demonstrated by RT-qPCR that pladienolide B alone or the combination of pladienolide B with cisplatin does not negatively impact MLH3 mRNA level in both H460R and A549R cells. These results were confirmed in a time-course experiment of pladienolide B treatment performed in H460 and A549 parental and resistant cells, as well as in NSCLC PDXs. Similar results were obtained in H460R or A549R cells deprived of SF3B1. These new data will be added in the revised version of the manuscript. The couple of primers we used for MLH3 amplification was located downstream of exon 8, respectively on constitutive MLH3-exon 9 (forward primer) and MLH3-exon 10 (reverse primer). These results indicate that pladienolide B regulates MLH3 splicing but not MLH3 total mRNA level. Considering NMD, in the FASTER DB database, none of the MLH3 transcripts devoid of exon 8 are predicted to be degraded by NMD. So we do not think that pladienolide B-induced MLH3 exon 8 skipping promotes the synthesis of transcripts recognized by the NMD machinery. As discussed above, the anti-MLH3 antibody does not allow to distinguish between the full length MLH3 protein and the MLH3 product encoded by transcript devoid of exon 8. In addition, only 24 amino acids (around 2-3KDa) differentiate both products which could render difficult their specific detection in SDS-PAGE. So, we speculate that the decrease of MLH3 signal detected by immunoblotting in pladienolide B-treated and SF3B1 knocked-down cells is mostly related to the decrease of MLH3 full-length protein due to the decreased level of MLH3 transcript retaining exon 8 and encoding MLH3 full-length protein. Alternatively, and not exclusively, MLH3 product devoid of exon 8 might also be less stable. __

      The difference shown in Fig. 6b after pladienolide B treatment decreases from 2.2% to 1%. This raises concern about whether the observed difference reflects a true biological effect or is confounded by technical limitations such as low transfection efficiency. The authors should consider optimizing their transfection conditions to achieve a more robust and convincing result.

      We agree with the reviewer’s comment. However, using this SceI-inducible system to analyze DNA double strand breaks repair by homologous recombination, it is very frequent to have only a very low percentage of cells able to perform homologous recombination thereby expressing the GFP protein ____(as examples: Yoshino Y et al., Sci Reports, 2019; Brustel et al., Sci Rep., 2018; Croglio et al., Oncotarget, 2016; Mamouni et al., Mol Cell Biol., 2014). This is why the difference is low between each condition but it is significant. Indeed, these engineered cellular models are not easy to manipulate as we first need to obtain stable clones having incorporated the PBL174 pDR-GFP-plasmid and then to transiently transfect them using a second plasmid encoding the SceI enzyme which creates DNA Double Strand Breaks. The efficiency of the second round of transfection might therefore be decreased as the cells already experienced a first round of transfection.

      __Minor comments __

      The abbreviation "S" in H460S and A549S cells is not defined in the manuscript. As this designation is used throughout the text, the authors should clarify what "S" denotes upon its first appearance.

      __We thank the reviewer for this remark. We will correct the text. __

      The current presentation of Figure S1a does not clearly demonstrate that different NSCLC cell lines exhibit differential sensitivity to pladienolide B. The authors should calculate and report IC50 values for each cell line to enable a more rigorous and quantitative comparison of their respective dose-response relationships.

      We thank the reviewer for this remark and agree with it. We will calculate and report in the revised version of Fig S1a the IC50 for pladienolide B for each cell line.

      In Figure 4c, two dashed black lines are present in the plot but are not described or explained in the figure legend or the main text.

      We thank the reviewer for this remark. The two dashed black lines represent the upper and lower boundaries of the 95% confidence interval for the fitted linear regression line. We have now clarified their meaning in the revised figure legend and indicated section of the main text also.

      In Figure 5a, the authors combine the downregulated and upregulated genes in a single Venn diagram. This approach may obscure biologically meaningful differences, as overlapping genes between conditions could reflect opposing directions of regulation. The authors should separate upregulated and downregulated genes into distinct Venn diagrams to provide a more accurate and interpretable comparison.

      We thank the reviewer for this remark and agree with it. Hence, in Fig 5a-b and Fig S3, we already highlighted the number of genes down-regulated or up-regulated upon pladienolide B treatment in either H460 parental and resistant cells using bar graphs. We will provide new Venn diagrams separating up-regulated and down-regulated genes for both cell lines.

      In the figure legend of Fig. 6a, the panel is incorrectly described as a "quantification." As the panel depicts a schematic representation of the experimental construct rather than numerical data, the term "illustration" or "schematic" would be more accurate and should be used instead.

      We thank the reviewer for this remark and we agree with it. We will modify the legend of Figure 6a accordingly.

      Reviewer #2 (Significance (Required)):

      This manuscript provides evidence that pladienolide B can overcome chemotherapy resistance in NSCLC by modulating splicing events of genes associated with DNA damage signaling and repair. Although the underlying mechanism requires further elucidation, this study offers valuable mechanistic insights into how aberrant splicing regulates therapy resistance, with potential implications for the development of novel therapeutic strategies targeting splicing factors in chemotherapy-resistant cancers.

      My research field is in tumor heterogeneity and tumor microenvironment.

    1. When we redefine libraries as means rather than as physical places—as conduits of knowledge rather than as physical buildings filled with physical books—we may think that the new, more “visionary,” more megatrendy definition embraces the old, but in fact it doesn’t: the removal of the concrete word “books” from the library’s statement of purpose is exactly the act that allows misguided administrators to work out their hostility toward printed history while the rest of us sleep.
    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1:

      Major comments:

      1. Lines 103-116 (first paragraph of the results section) describe mainly published data that is more suitable for the introduction section. It is annoying to refer to different published articles in the Results section to strengthen the results instead of showing them. The same goes for paragraphs two and three. Why mention those data in the Results section if they are already published and known?

      We have reorganized this material by moving some background information to the Introduction. Our intention was not to incorporate published data to strengthen our results, but rather to provide essential context for interpreting our findings. We have therefore left some of this foundational information in the results section to create a clear narrative flow, enabling readers to understand the basis for our experimental design and interpretations without needing to recall details from earlier paragraphs in the Introduction. For example, we considered it crucial to restate the earlier report of the BiP:sfGFP:HDEL phenotype in Atlastin mutants, since our results supporting luminal ER protein displacement contradict the previous fragmentation model.

      The following concept was in line 103 in the Results section, and is now in the introduction in lines 82-92: "Conventional light microscopy, commonly used in studies of neuronal ER structure, lacks the resolution necessary to visualize individual ER tubules in small structures, such as presynaptic terminals. The ER is highly sensitive to fixation, and live imaging experiments in neurons in vivo have been conducted on upright microscopes using water dipping objectives with a typical axial resolution limit of >300 nm, which cannot distinguish the densely packed ER tubules at presynaptic terminals (3,8,21,28,39-42). Electron microscopy offers higher resolution, but cannot be used in live samples and has typically been limited to thin 2D sampling (in which it is difficult to distinguish ER cross-sections from synaptic vesicles) (8,20,22)."

      Figure Legends-(in all Figures): The number of experimental repeats must be mentioned in the figure legends.

      This information is provided in Supplementary Table 1, which contains detailed information about the genotype, statistical analysis, and number of larvae and NMJs analyzed. If the journal requires this information in figure legends, we can move it.

      The way the figures are labeled is worrisome; supplementary figures are not ordered numerically.

      We will be happy to rename supplementary figures according to journal guidelines.

      The tubule extension in Figure 2D is not convincing. Is there a movie showing those changes? Better images are needed. It is essential to show which supplemental movie corresponds to which panel.

      We have now included a corresponding video of the same neuron used as an example of tubule extension. We also added another frame to the figure to provide further information on the tubule event we captured. (Figure 2D, Movie S10)

      This is unnecessary in the results section: "To investigate the relationship between ER structure and function at synapses, we examined mutants of Atlastin, a GTPase that regulates ER tubule fusion. Drosophila has a single homolog while mammals have three Atlastin homologs, with Atlastin-1 enriched in the brain (Rismanchi et al., 2008)."

      This information was moved to the introduction.

      "This reduction in ER membrane marker intensity has also been observed in other HSP mutants, suggesting this is a common feature of ER shaping mutants and could indicate changes in ER membrane composition, integrity, or tubule thickness (Perez-Moreno et al., 2023)." This comparison is important and should be shown in the same settings as for the Atlastin mutant rather than referring to published data.

      We agree with the reviewer that it is important to determine whether other ER-shaping proteins, besides Atlastin, also show a decrease in tdTomato:Sec61b to support our claim that this could be a common feature among ER-shaping mutants. To do this, we examined mutants of another ER-shaping protein, Reticulon 1, which regulates membrane bending and stabilization in ER tubules. These loss-of-function mutants were a gift from Dr. Cahir O'Kane at the University of Cambridge and were used in his lab's Pérez-Moreno et al., 2023 publication. We found that in our hands tdTomato:Sec61b levels were reduced in Reticulon 1 mutants, consistent with the results reported by Pérez-Moreno et al. (2023). These results are in Figure 3E-F. We also examined the synaptic distribution of the luminal ER marker, BiP:sfGFP:HDEL, in Reticulon 1 mutants to see if it is displaced to the cytosol. Notably, it remained ER-associated, unlike in Atlastin mutants. These results are in Figure 6F-G, results lines 267-270, and discussion lines 542-545.

      Does the distribution of the luminal ER marker in Figure 6F diffuse due to mislocalization or reflux after being localized to the ER and then refluxed to the cytosol as was previously shown for the ER to Cytosol signaling (ERCYS) mechanism? Could you assess other ER-luminal protein localization biochemically? It is highly recommended to look at another soluble ER-protein localization in the Atlastin mutant without overexpression, which can be an artifact.

      ER stressors can induce ERCYS, in which some luminal proteins, including PDIA3, DNAJB11, ERp29, and an eroGFP reporter, reflux by 30-70% to the cytoplasm without subsequent degradation (unlike ERAD (ER-associated degradation). This phenomenon has only previously been observed in yeast and glioblastoma tumor cells from mice and human . We believe that our work provide the first suggestion that this may occur in neurons, and particularly in a neurological disease model.

      We do not believe that the reflux phenotype for BiP:sfGFP:HDEL is due to its overexpression for two reasons: (1) we observe reflux in our neuronal Atlastin knockdown experiments, even when the levels of BiP:sfGFP:HDEL are significantly reduced artificially because of titration of the GAL4 between the RNAi and the reporter (Figure 7A), and (2) BiP:sfGFP:HDEL overexpression somewhat suppresses endogenous BiP upregulation ((Figure 10 and see Reviewer 1.10), arguing that the transgene does not induce ER stress). We included a new "limitations of the study" section to be transparent about the caveats of the BiP:sfGFP:HDEL reporter (lines 639-664).

      Identifying potential endogenous neuronal ERCYS substrates in our in vivo preparation poses several challenges. First, biochemical approaches, such as fractionation, are not possible in our complex in vivo sample because neuronal ER proteins would mix with ER from other tissues upon homogenization. Second, detecting endogenous proteins with antibodies requires fixation and permeabilization, which notoriously disrupts ER structure and even causes our reporter BiP:sfGFP:HDEL to collapse from a smooth distribution, as visualized by live imaging and FRAP, to a punctate distribution. Third, using antibodies rather than neuronally restricted transgenes makes it challenging to determine whether the signal originates from the neuron or from dense ER structures in the surrounding muscle. Fourth, some ER luminal proteins can displace as little as 30% in the ERCYS examples cited above, and the sensitivity of our imaging assays may limit our ability to detect these small changes. Finally, the limited availability of tagged transgenes and antibodies specific to Drosophila luminal ER proteins (see next paragraph) poses additional challenges. These limitations highlight the need for future studies to develop novel tools and techniques to more definitively test whether we are indeed observing ERCYS. We have included a paragraph on these future challenges in our discussion in lines 639-664. Identifying endogenous targets of ERCYS in fly neurons is a worthwhile goal, but beyond the scope of the current study. These next steps will particularly benefit from identifying the machinery involved in the reflux of our BiP:sfGFP:HDEL reporter.

      Tools we tested: We investigated several options: (1) a tagged PDI transgene (a gift from Karen Hibbard), which was not detectable at presynaptic terminals, (2) a tagged BiP (FlyORF; F000956) that did not localize to the ER, and (3) full-length endogenous BiP detected by antibody staining. We did not detect obvious reflux of endogenous BiP to the cytoplasm (Figure 9), with the caveat that in fixed samples, the BiP signal was not tightly co-localized with the ER marker even under control conditions. However, we did use this antibody to detect an increase in BiP in Atlastin mutant presynaptic terminals, indicating ER stress (see Reviewer 1.10).

      Though we have not identified endogenous targets, we believe that our studies with the exogenous reporter will be of great interest to the field, as they clarify the previously reported Atlastin phenotype and provide the first report of a new defect in a human disease animal model.

      In comparison to Summerville et al. (2016) in Figure 7, the experiment was not done in the same way. It is important to keep the same settings for comparison

      In Figure 7D-E, we compare the distribution of BiP:sfGFP:HDEL in cell bodies, axons, and muscles between controls and Atlastin mutants. To clarify the experimental approach relative to Summerville et al. (2016): while both our studies examined the same cellular compartments (cell bodies, axons and nerve terminals) using the BiP:sfGFP:HDEL reporter, we employed super-resolution Airyscan microscopy. This enhanced resolution was critical for definitively demonstrating that this is a functional rather than a structural phenotype and that ER displacement is progressive, and repeating this experiment at lower resolution as previously reported does not provide any new information. We identified two distinct distribution phenotypes in Atlastin mutants expressing BiP:sfGFP:HDEL, which were not described in the Summerville et al., 2016 paper. From our manuscript (lines 249-251): "We identified two distinct ER network phenotypes in Atlastin mutants expressing BiP:sfGFP:HDEL: "Partial loss" NMJs retained both diffuse signal and identifiable ER network structures, while "Complete loss" NMJs showed no visible ER network structures. Note that the "Complete loss" phenotype in Atlastin mutants reflects the absence of detectable luminal marker signal in organized ER structures, but not the complete absence of ER membranes, as demonstrated by our ER membrane marker tdTomato:Sec61β results."

      Does the Atlastin mutant induce the unfolded protein response and stress within the ER? It is necessary to look for UPR markers in those settings. It was shown previously that ER stress leads to protein reflux from the ER to the cytosol. Is there a difference in the ER stress markers in the presynaptic terminal?

      The reviewer suggested that Atlastin mutant synapses may exhibit ER stress. To address this, we examined levels of the ER chaperone BiP, a well-established ER stress marker whose expression increases during UPR activation. We first validated that our BiP antibody can detect changes in ER stress by feeding control larvae with 50mM DTT for 24 hours. These results are in the new Figure 10A. Note that we were unable to test sensitivity to ER stress in this way in Atlastin mutant larvae because they did not consume the DTT-treated food, as assessed by blue food coloring in the larvae's guts.

      Using this antibody, we measured baseline BiP levels at NMJs of Atlastin mutants on normal food, and found they were slightly increased compared to controls. We conclude from these experiments that Atlastin mutant synapses have mild ER stress. Notably however, Atlastin mutants co-expressing UAS-BiP:sfGFP:HDEL or UAS-tdTomato:Sec61b did not show significantly increased endogenous BiP levels, suggesting that transgene expression at least partly suppresses the mild ER stress response, even though there is extensive cytosolic displacement. These results argue (1) that the mild ER stress in Atl mutants does not strictly correlate with the reflux phenotype, and (2) that the reflux phenotype is not an artifact of overexpression-induced stress. These results are described on lines 430-436 in the results section and shown in Figure 10B-E, and their implications discussed on lines 585-598.

      We also explored another strategy to detect ER stress by assessing eIF2α phosphorylation, a key event in the Unfolded Protein Response (UPR) pathway. We obtained a phospho-eIF2α antibody (Cell Signaling; #3597) that was reported to work in Drosophila. However, when we tested this antibody by Western blot, we were unable to detect a band at the expected molecular weight for phosphorylated eIF2α, even in positive-control samples treated with DTT to induce ER stress. We therefore concluded that this antibody is not suitable for reliably detecting ER stress in our experimental system. The failure of this antibody highlights the challenges of finding robust tools to measure ER stress in Drosophila.

      It is important to add biochemical experiments to show that no fragmentation of the ER membrane occurred. It can be simply demonstrated by looking at the redox state of the ER, which would change if it were mixed with the reducing cytosol. Moreover, this can be shown by using an ER-targeted redox-sensitive fluorescent protein that is tethered to the ER membrane to follow changes in the redox state of the ER.

      The reviewer asked us to test whether the redox state of the ER is disrupted, which could indicate exchange between the cytosol and ER due to membrane rupture. As noted above, biochemical approaches such as fractionation are not possible in this in vivo sample. We attempted to address this concern by creating a UAS-Sec61β:roGFP construct, using the roGFP sequence from Igbaria et al. (2019) to monitor the ER lumen redox environment in Atlastin mutants. Since Sec61β is membrane-tethered, it should remain in the ER and not undergo reflux, making it an ideal sensor for detecting any mixing between the reducing cytosolic environment and the oxidizing ER lumen that would occur if membrane fragmentation and/or ruptures were present. We tested this approach in wild-type Drosophila S2 cells and used the Gal4-UAS binary expression system to co-express Actin-Gal4 (to drive expression of UAS constructs), UAS-Sec61β:roGFP (redox sensor), and UAS-BiP:Halo:HDEL (as a control reporter insensitive to DTT treatment).

      Our experiments showed no detectable changes in the fluorescent properties of UAS-Sec61β:roGFP following 30 min 10mM DTT treatment compared to DMSO vehicle control, including no increase in 405-nm excitation fluorescence or changes in 488nm/405nm excitation ratios. These results suggest that either the roGFP sensor requires further optimization for sensitivity in this cellular system or that additional controls and calibration steps are needed to establish the dynamic range of the assay. We believe this experiment falls beyond the scope of the current study, given the extensive optimization required. However, it represents an important future direction for testing membrane fragmentation as a mechanism underlying the phenotypes observed in Atlastin mutants. The possibility of ER integrity defects is mentioned in the discussion on lines 547-559.

      Minor comments:

      1. It is important to call figures by order. Figure 2C is called before 2A-B. Figure 2B is called before Figure 2A.

      The revised manuscript has all figures in order of appearance in the text.

      Figure legends (Figure 2): "The same control dataset used in E-G was used in Figure 5 and Figure 5_Supplement." Why is this relevant?

      We wanted to be transparent about reusing the same control dataset across multiple figures to avoid any appearance of data duplication. This notation clarifies that, although the data appear in different contexts (Figures 2 and 5. This version does not contain a Figure 5_Supplement), it represents the same biological samples analyzed for different parameters, ensuring readers understand that these are not independent datasets.

      Figure 4F is called before Figure-4D-E which are not called.

      We revised our manuscript and reorganized Figure 4 to ensure that all figure panels are referenced in sequential order and that panels 4D-E, which were previously not cited in the text, are now properly referenced when discussing their corresponding results.

      Figure 5B is called before the previous ones. Same for Figure 5A supplement.

      We referenced Figure 5A in lines 211-212, which precedes our discussion of Figure 5B. To clarify the figure order, we removed the early references to Figures 2D-G and Movies 7-14, which were mentioned only to indicate that we were analyzing the same dataset in different ways.

      The revised manuscript has all figures in order of appearance in the text.

      Referees cross-commenting

      I agree with the comments raised by reviewer2 and 3. Basically it is highly important to validate those data by genetic rescue. Moreover, it is essential to know the source of the displaced luminal marker to the cytosol. Is it mislocalization or it is a reflux of pre-existing protein to the cytosol after insertion to the ER. It is also recommended by me and the reviewers and me to test the endogenous protein rather than overexpression.

      We have addressed these points in our responses to the following reviewer questions:

      • Genetic rescue: Please see our responses to Reviewer 1/Question #10 and Reviewer 2/Question #1.
      • Source of displaced luminal marker: We provide some evidence addressing this in our response to Reviewer 3/Question #1.
      • Endogenous protein localization: We have examined this and detailed our findings in our responses to Reviewer 1/Question #7 and Reviewer 2/Question #6.

        Reviewer #1 (Significance (Required)):

      General assessment: This interesting paper shows that proteins can escape the ER under special conditions. However, the authors need more evidence to show that and rely less on the overexpression system, especially of BIP-GFP, which can cause proteostasis stress within the ER. Advance: The results have been oversimplified in their explanations, and some points and complexities of the study need to be addressed further to make the most of them. These are often some of the more interesting concepts in the paper. I think many points can be addressed in the text by the authors being clear and concise with their reporting. At the same time, other experiments would turn this paper from an observational one into a very interesting mechanistic one. This paper is based on previously published articles from the group and other groups, and it is a nice progression. However, as mentioned, this paper depends primarily on published data, and the novelty is somehow lost between all the comparisons to other published data instead of emphasizing that. Without a substantial mechanistic improvement, the paper would remain observatory.

      Audience: The microscopy tools can be great addition to researchers in the field to monitor protein trafficking especially Cell biologists (basic research)

      My expertise: ER homeostasis, protein trafficking, cell biology

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      Summary The endoplasmic reticulum (ER) is a continuous organelle that extends throughout neurons to regulate fundamental processes. The analysis of ER dynamics at synaptic terminals is limited by the challenge of imaging these structures at high resolution. In this manuscript, the authors use super-resolution (~170 nm) live imaging and a combination of membrane and luminal ER markers at the Drosophila larval NMJ, an important model synapse, to investigate dynamic ER architecture in vivo. They report a detailed characterization of the presynaptic ER organization and dynamics at wild-type and GTPase Atlastin mutant NMJs. Their analysis using the ER membrane marker tdTomato:Sec61b reveals the presence of an intact ER network in Atlastin mutants. This contrasts with the apparent ER fragmentation phenotype previously reported and replicated here when using a luminal marker. Their findings instead point to the progressive displacement of luminal proteins to the cytosol in Atlastin mutants specifically at synapses. The authors propose that the disruption of ER protein dynamics at synapses is a compartment-specific ER stress response. The manuscript is well written, results are clearly presented, and experiments are technically rigorous.

      Major comments

      1. The baseline ER phenotypes in Atlastin mutants are mild with complete loss of ER network only observed in terminal boutons. This interesting and unexpected result should be further confirmed by genetic rescue. The authors can use a UAS rescue line previously reported in PMID: 19341724.

      We tested the UAS-Atl-myc rescue line and unfortunately found that even in wild-type neurons, overexpression of Atlastin produced strong ER organization defects that precluded the rescue experiment. Instead, to confirm the cell autonomy of the phenotype and to test it wth an independent tool, we performed a presynaptic knockdown of Atlastin by RNAi and found that BiP:sfGFP:HDEL is displaced, as observed in the Atlastin null mutant. These results are in now shown in Figure 7A-C.

      Lines 204-7: It's not clear how a greater coefficient of variation indicates that the marker is more concentrated in subsynaptic structures or what is meant by 'subsynaptic structures.'

      We added the following text to explain, in lines 181-183: "A higher CoV indicates an uneven distribution of tdTomato:Sec61β within the presynaptic terminal, with some areas showing higher concentrations than others (in contrast to the uniform, diffuse signal expected from fragmentation)." To avoid confusion with postsynaptic structures called the subsynaptic reticulum, we have removed the term "subsynaptic". The intended meaning is distinct structures found within the presynaptic terminal.

      There's a mistake in Figure 6C and the associated text. The summed percentage of the three phenotypic categories adds up to 110% for Atlastin mutants.

      The reviewer noted that the summed percentage of the three phenotypic categories in Figure 6C adds up to 110% for Atlastin mutants, which appears to be a mathematical error. However, this is not an error, but rather a reflection of our quantification methodology, in which a single bouton can exhibit more than one type of ER dynamics per movie recorded. Our quantification counts each phenotype independently, so boutons displaying multiple phenotypes contribute to more than one category. This approach provides a more comprehensive view of the range of ER dynamics present in Atlastin mutants, as restricting the analysis to mutually exclusive categories would underrepresent the complexity of the phenotypes observed. To make this point clear, we made the following change to the text in lines 257-259: "We note that the sum of these percentages exceeds 100% because one NMJ exhibited multiple phenotypes: one branch had a complete loss, while the other branch had no phenotype. These phenotypes were counted separately."

      Figure 8: the ER looks fragmented in 1st instar controls and mutants. The authors should address this difference from more mature NMJs.

      We would like to clarify that the bulk of experiments in this manuscript (including all ER dynamics, luminal marker redistribution, and membrane marker analyses discussed throughout the Results) were performed in 3rd instar larvae, which are more mature larval NMJ preparations standard in the field. Figure 8 was included specifically to test whether the Atlastin mutant phenotype we describe throughout the paper is also detectable at an earlier developmental stage, not to replace or reinterpret our primary findings.

      Regarding the specific observation that the ER appears more fragmented in Figure 7F-H relative to the more mature NMJs shown elsewhere: this fragmentation, observed similarly in both control and Atlastin mutant 1st instar larvae, likely reflects technical challenges associated with dissecting these smaller, more delicate early-stage specimens rather than a genotype-specific effect. Because fragmentation occurred similarly in both genotypes, we could still reliably assess the redistribution of BiP:sfGFP:HDEL as our primary phenotypic readout in this experiment. We have added the following text (lines 306-309) to clarify this point: "Note that in 1st instar larvae, both normal networks in controls and residual networks in Atlastin mutants appeared more fragmented than in 3rd instar preparations, likely due to the technical challenges of dissecting these smaller, more delicate specimens. Since ER fragmentation occurred similarly in both genotypes, we could still reliably assess the redistribution of BiP:sfGFP:HDEL as our primary phenotypic readout.

      The images in figure 9B do not seem representative of the quantification in Figure 9D. Specifically, the partial loss Atlastin NMJ appears to have recovered as fully as the complete loss Atlastin NMJ.

      The images showed FRAP recovery across the entire bouton, but we photobleached only a small region within each bouton and quantified only this region. We have now added outlines to clearly delineate the specific FRAP regions that were analyzed in each image, which clarify that the partial loss Atlastin showed less recovery than the overall bouton. We have also reordered the figures to more clearly convey our message (Figure 9 is now Figure 8).

      We also made a few changes to the paragraph on lines 347-350 to clarify our experimental reasoning: "We photobleached en passant boutons using a defined region of 6.8 x 7.8 microns (dashed box in Figure 8D) to ensure that BiP:sfGFP:HDEL could recover from the ER networks surrounding the FRAP region (Movies S20-S23)."

      We also added this sentence to the figure legends of Figure 8: "The dashed boxes in (D) indicate areas that were photobleached and analyzed for recovery quantification in (E-F)."

      Optional: An overexpressed luminal marker is displaced to the cytoplasm in Atlastin mutants. It would be interesting to know and increase the significance of the findings if the same is true of endogenous luminal proteins under biological stress conditions.

      As noted in our response to Reviewer #1 suggested that Atlastin mutant synapses may exhibit ER stress. To address this, we examined levels of the ER chaperone BiP, a well-established ER stress marker whose expression increases during UPR activation. We first validated that our BiP antibody can detect changes in ER stress by feeding control larvae with 50mM DTT for 24 hours. We were unable to perform this experiment in Atlastin mutant larvae because they did not consume the DTT-treated food, as assessed by blue food coloring in the larvae's guts. These results are in Figure 10A. In the future, it will be of interest to establish a protocol to examine Atlastin mutants by feeding or treating larval fillets with DTT.

      We measured BiP levels at NMJs of Atlastin mutants and found they were slightly increased compared to controls. Atlastin mutants co-expressing UAS-BiP:sfGFP:HDEL or UAS-tdTomato:Sec61b did not show significantly increased endogenous BiP levels, suggesting that transgene expression suppresses the mild ER stress response. We conclude from these experiments that Atlastin mutant synapses have mild ER stress. These results are in Figure 10B-E).

      Optional: Applying this approach in stimulated conditions (high potassium, increased temperature) might reveal a greater activity-dependent role for Atlastin at synaptic terminals.

      This is a very interesting idea, as we have only examined synapses at rest. However, this is beyond the scope of this paper.

      Minor Comments

      1. Line 16: Atlastin should be italicized.

      Thank you for catching this typo. We have fixed it.

      Figure 5A: Based on the relative intensities, it appears that control and mutant images are not contrast matched but this isn't stated.

      Thank you for catching this omission. We added to the figure legend: "Control and Atlastin mutant images are not contrast matched."

      Line 822: The number of static Atlastin mutant boutons used for analysis is missing.

      Thank you for catching this omission. We have fixed this supplementary table.

      Figure 9: The blue arrows are not annotated in the figure legend.

      Thank you for catching this omission. We have fixed this figure legend.

      Reviewer #2 (Significance (Required)):

      Atlastin is linked to Hereditary Spastic Paraplegia (HSP) and this study changes our understanding of the compartment-specific impacts of its loss. This study reveals the importance of using both membrane and luminal ER markers to accurately interpret phenotypes as well as the importance of considering compartment-specific effects on ER. These findings represent significant mechanistic and conceptual advances. The lack of genetic rescue is a limitation and adding an investigation of an endogenous luminal protein under basal and stress conditions would add significantly to our understanding of Atlastin dysfunction in HSP. Notably, the in vivo imaging approach introduced here can be adapted broadly for live imaging of Drosophila larvae. Thus, this work will be of interest to both neuronal cell biologists and the wider Drosophila community. This review is based on our expertise in neuronal cell biology.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      In this manuscript, the authors investigate the structural dynamics of the endoplasmic reticulum (ER) in Drosophila neurons and examine the role of the ER-shaping protein Atlastin in ER morphology. Their discovery on the neuromuscular junction (NMJ)-specific contribution of Atlastin to ER integrity is intriguing and may provide valuable insights into the pathological mechanisms underlying Atlastin mutations associated with hereditary spastic paraplegia (HSP) and hereditary sensory neuropathy. The key observation on ER protein showing an aberrant cytoplasmic localisation in mutant cells appears convincing. Though this phenomenon's characterisation stays at the point of primary observation with its mechanics unclarified, establishing this new and unexpected functional rather than structural Atl effect is important and useful for the field. The observation that ER is structurally preserved in this mutant with absolute lack of Atl are also extremely useful.

      It is unclear if the cytoplasmic localisation affects an exogenous overexpressed ER marker or endogenous protein would also appear in cytoplams, the authors should consider adding an immunostaining data to test that.

      Authors offer speculations on potential reasons for the cyto localisation of the ER marker suggesting that relocation at the cell periphery specifically combined with slow clearance there is the most likely explanation (still unclear what stops the marker from spreading through the entire cell). They suggest that decrease in cotranslational translocation is unlikely as this would result in somatic accumulation of the marker. However, if the clearance in the periphery is less efficient than in soma, the accumulation there might reflect a compromised translocation. Any clarifying experiments, if practical, to directly demonstrate how ER proteins in relocates to the cytoplasm in atl mutant would help understanding better the phenomenon. For example, would proteasomal inhibition make the marker accumulate more across the cell? Authors also suggest links to ER stress. Would stress induction phenocopy the mutant?

      Reviewer #3 asked whether defective proteasomal clearance underlies the cytosolic accumulation of BiP:sfGFP:HDEL in Atlastin mutants. We addressed this directly. First, proteasome function appears intact in the mutants: baseline ubiquitinated protein levels (FK1 antibody) were comparable between control and Atlastin mutants, and MG132 treatment produced a similar increase in ubiquitination in both genotypes, confirming both antibody specificity and normal proteasome activity. We then examined BiP:sfGFP:HDEL directly. In controls, MG132 caused the marker to accumulate at axons and presynaptic terminals, showing that it is normally cleared from these compartments by the proteasome. Critically, this accumulated marker remained associated with intact ER networks: MG132 did not induce diffuse cytosolic BiP:sfGFP:HDEL in any compartment (cell bodies, axons, or presynaptic terminals), even where levels rose substantially. Thus, blocking proteasomal clearance raises ER-localized marker but does not generate the cytosolic pool seen in Atlastin mutants, indicating that impaired clearance is not sufficient to cause the displacement phenotype. We separately noted that BiP:sfGFP:HDEL was already elevated in Atlastin mutant axons without MG132, paralleling the axonal tdTomato:Sec61β accumulation in Figure 4, consistent with reduced baseline clearance specifically in mutant axons, but this does not lead to cytosolic displacement. This experiment is now shown in Figure 11, described in Results (lines 445-475), and discussed in lines 576-581.

      Minor comments:

      Line 146:

      "fast dynamics (Thank you for catching this mistake. We have corrected it.

      Fig. 2D: The data representation of "Tubule displacement" image is unclear. The ER tubule indicated by the red arrow does not seem to show any changes over time (like static). time 0 in stamp appears behind the image.

      Thank you for catching the typo. We have fixed it. Additionally, we added black arrows to highlight a tubule that is not moving, allowing the reader to compare it with the moving tubule. We also included a video of all types of ER tubule dynamics to ensure the reader can also look at the raw data (Movies S9-11).

      Line 157-158 (and relevant method sections):

      The definition of static and dynamic boutons is ambiguous. The author should describe in more detail this point including how long they observed the structure to define the changes in ER tubule dynamics.

      We provide in the methods (lines 779-791) a detailed explanation of how we categorized boutons as dynamic or static. In addition, we added the following to explain in the results section how we defined static vs dynamic:

      Old sentence: We qualitatively categorized boutons as "static" if we observed no change in ER network structure or "dynamic" if we observed at least one change.

      New sentence in lines 143-147: "We imaged boutons for 40 sec at 0.92 sec intervals to capture ER dynamics over this observation period. Boutons were qualitatively categorized as "static" if we observed no detectable changes in ER network structure throughout the entire 40 sec imaging session, or "dynamic" if we observed at least one of the three defined dynamic events during this time window."

      Fig. 2E: What n=75 and n=29 represent is unclear, are these the number of boutons in en passant and terminal subjected for qualitative analysis?

      We removed these n values from the figure and added this information to the Supplementary Table 1, which contains detailed information about the genotype, statistical analysis, and number of larvae and NMJs analyzed.

      Fig. 2: What the qualitative analysis represents is unclear, are the points pulled from different experiments?

      The data in Fig. 2 E-F comes from movies acquired in the same experiment. The number of independent animals and NMJs imaged is described in Table 1.

      * *Line 231: Regarding "...we found a small but significant reduction in dynamic boutons in Atlastin mutants (76%), ...", how do the authors assess significance. If proportion of static/dynamic ER in boutons was obtained from multiple experiments, it should be presented e.g. as in average {plus minus} standard deviation, or clarify that the proportion is representative of x independent experiments.

      The videos used for this figure were acquired from a single experiment. We use a chi-square test to determine significance relative to the "expected" distribution of dynamics types from controls, as these are categorical rather than continuous data (see PMID 31145670). Information regarding genotype, statistical analysis and number of larvae and NMJs can also be found in Supplementary Table 1.

      Line 267-269 and Fig. 6B: The author's conclusion that "Complete loss of ER network structure in NMJ of BiP:sfGFP:HDEL overexpressing Atl mutant" seem to be based on the lack of signal from luminal marker, which may be undetectable due to changes to tubular volume or marker loss to the cytoplasm, as suggested by the authors, while the membranous ER structure is intact. It would be useful to discuss this point and potentially add ER membrane-stained control.

      We agree with the reviewer that Atlastin mutants categorized as 'complete loss mutants' do not actually lack ER at synapses. We think this is an important point so we added the following to the results in lines 251-254: "Note that the "Complete loss" phenotype in Atlastin mutants reflects the absence of detectable luminal marker signal in organized ER structures, not the complete absence of ER membranes, as demonstrated by our ER membrane marker tdTomato:Sec61β results."

      We attempted to co-label the ER membrane and ER lumen, but these crosses yielded very few live larvae (in either controls or Atlastin mutants, and those that survived had severely deformed NMJs. We added Figure 6-Supplement showing the results of this experiment, and described them on lines 270-273.

      Fig. 6C: In Atl mutant, why does the total of the proportion exceed 100% (10 + 45 + 55)?

      The reviewer noted that the summed percentage of the three phenotypic categories in Figure 6C adds up to 110% for Atlastin mutants. This is not an error, but rather a reflection of our quantification methodology because a single bouton can exhibit more than one type of ER dynamics per movie recorded. Our quantification counts each phenotype independently, so boutons displaying multiple phenotypes contribute to more than one category. This approach provides a more comprehensive view of the range of ER dynamics present in Atlastin mutants, as restricting the analysis to mutually exclusive categories would underrepresent the complexity of the phenotypes observed. To make this point clear, we made the following change to the text in lines 257-259: "We note that the sum of these percentages exceeds 100% because one NMJ exhibited multiple phenotypes: one branch had a complete loss, while the other branch had no phenotype. These phenotypes were counted separately."

      Fig. 9C, line 342-344: In FRAP experiment using CD8, it seems that the Partial loss Atl mutant shows slower recovery that control. There seems to be a mismatch in triangle symbols of Partial loss Atl mutant between legend and plot (one is filled and the other is empty). This should be clarified.

      Thank you for catching this mistake. We have fixed the figure.

      fig. 10 is a clever way to verify the cytoplasmic localization of the ER marker; however, its description and annotation can be improved, and it would be stronger if 4 curves in F for mutant and controls with the trap and normal were shown.

      The reviewer suggested merging our graphs but we believe that keeping them separate is clearer.

      Line 495: Drosophila have ReepA and ReepB, but not Reep1-4. If the authors discuss their speculation based on their observation (using Drosophila), the gene names should be unified in the same species, and explain the corresponding genes to mammalian cells.

      We made the following changes to address the reviewer's concern about gene nomenclature consistency (lines 502-506): "These ER-derived vesicles are likely to involve ReepA and ReepB, the Drosophila orthologs of mammalian REEP1-4, which regulate ER vesicle formation in mammalian cells (67). Notably, while overexpression of Atlastin can regulate REEP vesicle fusion in mammalian systems (67), it is not essential for vesicle formation, suggesting similar regulatory relationships may exist between Atlastin and Reep genes in Drosophila."

      Line 548; should UPR be Unfolded Protein Response?

      Thank you for catching the typo. We have fixed it.

      Reviewer #3 (Significance (Required)):

      This study advances the understanding of how ER morphogens affect neuronal cells specifically, the lack of which limits researchers ability to comprehend the neuronal pathologies associated with ER structure-function. The observation on ER content aberrant localisation caused by the lack of key structural protein should be of a great interest for cell and neuronal biologists and researchers of the associated diseases and shows the field a new direction. Though, mechanistic details remain to be unraveled, it constitutes a fundamental, conceptual advance.

    2. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1:

      Major comments:

      1. Lines 103-116 (first paragraph of the results section) describe mainly published data that is more suitable for the introduction section. It is annoying to refer to different published articles in the Results section to strengthen the results instead of showing them. The same goes for paragraphs two and three. Why mention those data in the Results section if they are already published and known?

      We have reorganized this material by moving some background information to the Introduction. Our intention was not to incorporate published data to strengthen our results, but rather to provide essential context for interpreting our findings. We have therefore left some of this foundational information in the results section to create a clear narrative flow, enabling readers to understand the basis for our experimental design and interpretations without needing to recall details from earlier paragraphs in the Introduction. For example, we considered it crucial to restate the earlier report of the BiP:sfGFP:HDEL phenotype in Atlastin mutants, since our results supporting luminal ER protein displacement contradict the previous fragmentation model.

      The following concept was in line 103 in the Results section, and is now in the introduction in lines 82-92: "Conventional light microscopy, commonly used in studies of neuronal ER structure, lacks the resolution necessary to visualize individual ER tubules in small structures, such as presynaptic terminals. The ER is highly sensitive to fixation, and live imaging experiments in neurons in vivo have been conducted on upright microscopes using water dipping objectives with a typical axial resolution limit of >300 nm, which cannot distinguish the densely packed ER tubules at presynaptic terminals (3,8,21,28,39-42). Electron microscopy offers higher resolution, but cannot be used in live samples and has typically been limited to thin 2D sampling (in which it is difficult to distinguish ER cross-sections from synaptic vesicles) (8,20,22)."

      Figure Legends-(in all Figures): The number of experimental repeats must be mentioned in the figure legends.

      This information is provided in Supplementary Table 1, which contains detailed information about the genotype, statistical analysis, and number of larvae and NMJs analyzed. If the journal requires this information in figure legends, we can move it.

      The way the figures are labeled is worrisome; supplementary figures are not ordered numerically.

      We will be happy to rename supplementary figures according to journal guidelines.

      The tubule extension in Figure 2D is not convincing. Is there a movie showing those changes? Better images are needed. It is essential to show which supplemental movie corresponds to which panel.

      We have now included a corresponding video of the same neuron used as an example of tubule extension. We also added another frame to the figure to provide further information on the tubule event we captured. (Figure 2D, Movie S10)

      This is unnecessary in the results section: "To investigate the relationship between ER structure and function at synapses, we examined mutants of Atlastin, a GTPase that regulates ER tubule fusion. Drosophila has a single homolog while mammals have three Atlastin homologs, with Atlastin-1 enriched in the brain (Rismanchi et al., 2008)."

      This information was moved to the introduction.

      "This reduction in ER membrane marker intensity has also been observed in other HSP mutants, suggesting this is a common feature of ER shaping mutants and could indicate changes in ER membrane composition, integrity, or tubule thickness (Perez-Moreno et al., 2023)." This comparison is important and should be shown in the same settings as for the Atlastin mutant rather than referring to published data.

      We agree with the reviewer that it is important to determine whether other ER-shaping proteins, besides Atlastin, also show a decrease in tdTomato:Sec61b to support our claim that this could be a common feature among ER-shaping mutants. To do this, we examined mutants of another ER-shaping protein, Reticulon 1, which regulates membrane bending and stabilization in ER tubules. These loss-of-function mutants were a gift from Dr. Cahir O'Kane at the University of Cambridge and were used in his lab's Pérez-Moreno et al., 2023 publication. We found that in our hands tdTomato:Sec61b levels were reduced in Reticulon 1 mutants, consistent with the results reported by Pérez-Moreno et al. (2023). These results are in Figure 3E-F. We also examined the synaptic distribution of the luminal ER marker, BiP:sfGFP:HDEL, in Reticulon 1 mutants to see if it is displaced to the cytosol. Notably, it remained ER-associated, unlike in Atlastin mutants. These results are in Figure 6F-G, results lines 267-270, and discussion lines 542-545.

      Does the distribution of the luminal ER marker in Figure 6F diffuse due to mislocalization or reflux after being localized to the ER and then refluxed to the cytosol as was previously shown for the ER to Cytosol signaling (ERCYS) mechanism? Could you assess other ER-luminal protein localization biochemically? It is highly recommended to look at another soluble ER-protein localization in the Atlastin mutant without overexpression, which can be an artifact.

      ER stressors can induce ERCYS, in which some luminal proteins, including PDIA3, DNAJB11, ERp29, and an eroGFP reporter, reflux by 30-70% to the cytoplasm without subsequent degradation (unlike ERAD (ER-associated degradation). This phenomenon has only previously been observed in yeast and glioblastoma tumor cells from mice and human . We believe that our work provide the first suggestion that this may occur in neurons, and particularly in a neurological disease model.

      We do not believe that the reflux phenotype for BiP:sfGFP:HDEL is due to its overexpression for two reasons: (1) we observe reflux in our neuronal Atlastin knockdown experiments, even when the levels of BiP:sfGFP:HDEL are significantly reduced artificially because of titration of the GAL4 between the RNAi and the reporter (Figure 7A), and (2) BiP:sfGFP:HDEL overexpression somewhat suppresses endogenous BiP upregulation ((Figure 10 and see Reviewer 1.10), arguing that the transgene does not induce ER stress). We included a new "limitations of the study" section to be transparent about the caveats of the BiP:sfGFP:HDEL reporter (lines 639-664).

      Identifying potential endogenous neuronal ERCYS substrates in our in vivo preparation poses several challenges. First, biochemical approaches, such as fractionation, are not possible in our complex in vivo sample because neuronal ER proteins would mix with ER from other tissues upon homogenization. Second, detecting endogenous proteins with antibodies requires fixation and permeabilization, which notoriously disrupts ER structure and even causes our reporter BiP:sfGFP:HDEL to collapse from a smooth distribution, as visualized by live imaging and FRAP, to a punctate distribution. Third, using antibodies rather than neuronally restricted transgenes makes it challenging to determine whether the signal originates from the neuron or from dense ER structures in the surrounding muscle. Fourth, some ER luminal proteins can displace as little as 30% in the ERCYS examples cited above, and the sensitivity of our imaging assays may limit our ability to detect these small changes. Finally, the limited availability of tagged transgenes and antibodies specific to Drosophila luminal ER proteins (see next paragraph) poses additional challenges. These limitations highlight the need for future studies to develop novel tools and techniques to more definitively test whether we are indeed observing ERCYS. We have included a paragraph on these future challenges in our discussion in lines 639-664. Identifying endogenous targets of ERCYS in fly neurons is a worthwhile goal, but beyond the scope of the current study. These next steps will particularly benefit from identifying the machinery involved in the reflux of our BiP:sfGFP:HDEL reporter.

      Tools we tested: We investigated several options: (1) a tagged PDI transgene (a gift from Karen Hibbard), which was not detectable at presynaptic terminals, (2) a tagged BiP (FlyORF; F000956) that did not localize to the ER, and (3) full-length endogenous BiP detected by antibody staining. We did not detect obvious reflux of endogenous BiP to the cytoplasm (Figure 9), with the caveat that in fixed samples, the BiP signal was not tightly co-localized with the ER marker even under control conditions. However, we did use this antibody to detect an increase in BiP in Atlastin mutant presynaptic terminals, indicating ER stress (see Reviewer 1.10).

      Though we have not identified endogenous targets, we believe that our studies with the exogenous reporter will be of great interest to the field, as they clarify the previously reported Atlastin phenotype and provide the first report of a new defect in a human disease animal model.

      In comparison to Summerville et al. (2016) in Figure 7, the experiment was not done in the same way. It is important to keep the same settings for comparison

      In Figure 7D-E, we compare the distribution of BiP:sfGFP:HDEL in cell bodies, axons, and muscles between controls and Atlastin mutants. To clarify the experimental approach relative to Summerville et al. (2016): while both our studies examined the same cellular compartments (cell bodies, axons and nerve terminals) using the BiP:sfGFP:HDEL reporter, we employed super-resolution Airyscan microscopy. This enhanced resolution was critical for definitively demonstrating that this is a functional rather than a structural phenotype and that ER displacement is progressive, and repeating this experiment at lower resolution as previously reported does not provide any new information. We identified two distinct distribution phenotypes in Atlastin mutants expressing BiP:sfGFP:HDEL, which were not described in the Summerville et al., 2016 paper. From our manuscript (lines 249-251): "We identified two distinct ER network phenotypes in Atlastin mutants expressing BiP:sfGFP:HDEL: "Partial loss" NMJs retained both diffuse signal and identifiable ER network structures, while "Complete loss" NMJs showed no visible ER network structures. Note that the "Complete loss" phenotype in Atlastin mutants reflects the absence of detectable luminal marker signal in organized ER structures, but not the complete absence of ER membranes, as demonstrated by our ER membrane marker tdTomato:Sec61β results."

      Does the Atlastin mutant induce the unfolded protein response and stress within the ER? It is necessary to look for UPR markers in those settings. It was shown previously that ER stress leads to protein reflux from the ER to the cytosol. Is there a difference in the ER stress markers in the presynaptic terminal?

      The reviewer suggested that Atlastin mutant synapses may exhibit ER stress. To address this, we examined levels of the ER chaperone BiP, a well-established ER stress marker whose expression increases during UPR activation. We first validated that our BiP antibody can detect changes in ER stress by feeding control larvae with 50mM DTT for 24 hours. These results are in the new Figure 10A. Note that we were unable to test sensitivity to ER stress in this way in Atlastin mutant larvae because they did not consume the DTT-treated food, as assessed by blue food coloring in the larvae's guts.

      Using this antibody, we measured baseline BiP levels at NMJs of Atlastin mutants on normal food, and found they were slightly increased compared to controls. We conclude from these experiments that Atlastin mutant synapses have mild ER stress. Notably however, Atlastin mutants co-expressing UAS-BiP:sfGFP:HDEL or UAS-tdTomato:Sec61b did not show significantly increased endogenous BiP levels, suggesting that transgene expression at least partly suppresses the mild ER stress response, even though there is extensive cytosolic displacement. These results argue (1) that the mild ER stress in Atl mutants does not strictly correlate with the reflux phenotype, and (2) that the reflux phenotype is not an artifact of overexpression-induced stress. These results are described on lines 430-436 in the results section and shown in Figure 10B-E, and their implications discussed on lines 585-598.

      We also explored another strategy to detect ER stress by assessing eIF2α phosphorylation, a key event in the Unfolded Protein Response (UPR) pathway. We obtained a phospho-eIF2α antibody (Cell Signaling; #3597) that was reported to work in Drosophila. However, when we tested this antibody by Western blot, we were unable to detect a band at the expected molecular weight for phosphorylated eIF2α, even in positive-control samples treated with DTT to induce ER stress. We therefore concluded that this antibody is not suitable for reliably detecting ER stress in our experimental system. The failure of this antibody highlights the challenges of finding robust tools to measure ER stress in Drosophila.

      It is important to add biochemical experiments to show that no fragmentation of the ER membrane occurred. It can be simply demonstrated by looking at the redox state of the ER, which would change if it were mixed with the reducing cytosol. Moreover, this can be shown by using an ER-targeted redox-sensitive fluorescent protein that is tethered to the ER membrane to follow changes in the redox state of the ER.

      The reviewer asked us to test whether the redox state of the ER is disrupted, which could indicate exchange between the cytosol and ER due to membrane rupture. As noted above, biochemical approaches such as fractionation are not possible in this in vivo sample. We attempted to address this concern by creating a UAS-Sec61β:roGFP construct, using the roGFP sequence from Igbaria et al. (2019) to monitor the ER lumen redox environment in Atlastin mutants. Since Sec61β is membrane-tethered, it should remain in the ER and not undergo reflux, making it an ideal sensor for detecting any mixing between the reducing cytosolic environment and the oxidizing ER lumen that would occur if membrane fragmentation and/or ruptures were present. We tested this approach in wild-type Drosophila S2 cells and used the Gal4-UAS binary expression system to co-express Actin-Gal4 (to drive expression of UAS constructs), UAS-Sec61β:roGFP (redox sensor), and UAS-BiP:Halo:HDEL (as a control reporter insensitive to DTT treatment).

      Our experiments showed no detectable changes in the fluorescent properties of UAS-Sec61β:roGFP following 30 min 10mM DTT treatment compared to DMSO vehicle control, including no increase in 405-nm excitation fluorescence or changes in 488nm/405nm excitation ratios. These results suggest that either the roGFP sensor requires further optimization for sensitivity in this cellular system or that additional controls and calibration steps are needed to establish the dynamic range of the assay. We believe this experiment falls beyond the scope of the current study, given the extensive optimization required. However, it represents an important future direction for testing membrane fragmentation as a mechanism underlying the phenotypes observed in Atlastin mutants. The possibility of ER integrity defects is mentioned in the discussion on lines 547-559.

      Minor comments:

      1. It is important to call figures by order. Figure 2C is called before 2A-B. Figure 2B is called before Figure 2A.

      The revised manuscript has all figures in order of appearance in the text.

      Figure legends (Figure 2): "The same control dataset used in E-G was used in Figure 5 and Figure 5_Supplement." Why is this relevant?

      We wanted to be transparent about reusing the same control dataset across multiple figures to avoid any appearance of data duplication. This notation clarifies that, although the data appear in different contexts (Figures 2 and 5. This version does not contain a Figure 5_Supplement), it represents the same biological samples analyzed for different parameters, ensuring readers understand that these are not independent datasets.

      Figure 4F is called before Figure-4D-E which are not called.

      We revised our manuscript and reorganized Figure 4 to ensure that all figure panels are referenced in sequential order and that panels 4D-E, which were previously not cited in the text, are now properly referenced when discussing their corresponding results.

      Figure 5B is called before the previous ones. Same for Figure 5A supplement.

      We referenced Figure 5A in lines 211-212, which precedes our discussion of Figure 5B. To clarify the figure order, we removed the early references to Figures 2D-G and Movies 7-14, which were mentioned only to indicate that we were analyzing the same dataset in different ways.

      The revised manuscript has all figures in order of appearance in the text.

      Referees cross-commenting

      I agree with the comments raised by reviewer2 and 3. Basically it is highly important to validate those data by genetic rescue. Moreover, it is essential to know the source of the displaced luminal marker to the cytosol. Is it mislocalization or it is a reflux of pre-existing protein to the cytosol after insertion to the ER. It is also recommended by me and the reviewers and me to test the endogenous protein rather than overexpression.

      We have addressed these points in our responses to the following reviewer questions:

      • Genetic rescue: Please see our responses to Reviewer 1/Question #10 and Reviewer 2/Question #1.
      • Source of displaced luminal marker: We provide some evidence addressing this in our response to Reviewer 3/Question #1.
      • Endogenous protein localization: We have examined this and detailed our findings in our responses to Reviewer 1/Question #7 and Reviewer 2/Question #6.

        Reviewer #1 (Significance (Required)):

      General assessment: This interesting paper shows that proteins can escape the ER under special conditions. However, the authors need more evidence to show that and rely less on the overexpression system, especially of BIP-GFP, which can cause proteostasis stress within the ER. Advance: The results have been oversimplified in their explanations, and some points and complexities of the study need to be addressed further to make the most of them. These are often some of the more interesting concepts in the paper. I think many points can be addressed in the text by the authors being clear and concise with their reporting. At the same time, other experiments would turn this paper from an observational one into a very interesting mechanistic one. This paper is based on previously published articles from the group and other groups, and it is a nice progression. However, as mentioned, this paper depends primarily on published data, and the novelty is somehow lost between all the comparisons to other published data instead of emphasizing that. Without a substantial mechanistic improvement, the paper would remain observatory.

      Audience: The microscopy tools can be great addition to researchers in the field to monitor protein trafficking especially Cell biologists (basic research)

      My expertise: ER homeostasis, protein trafficking, cell biology

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      Summary The endoplasmic reticulum (ER) is a continuous organelle that extends throughout neurons to regulate fundamental processes. The analysis of ER dynamics at synaptic terminals is limited by the challenge of imaging these structures at high resolution. In this manuscript, the authors use super-resolution (~170 nm) live imaging and a combination of membrane and luminal ER markers at the Drosophila larval NMJ, an important model synapse, to investigate dynamic ER architecture in vivo. They report a detailed characterization of the presynaptic ER organization and dynamics at wild-type and GTPase Atlastin mutant NMJs. Their analysis using the ER membrane marker tdTomato:Sec61b reveals the presence of an intact ER network in Atlastin mutants. This contrasts with the apparent ER fragmentation phenotype previously reported and replicated here when using a luminal marker. Their findings instead point to the progressive displacement of luminal proteins to the cytosol in Atlastin mutants specifically at synapses. The authors propose that the disruption of ER protein dynamics at synapses is a compartment-specific ER stress response. The manuscript is well written, results are clearly presented, and experiments are technically rigorous.

      Major comments

      1. The baseline ER phenotypes in Atlastin mutants are mild with complete loss of ER network only observed in terminal boutons. This interesting and unexpected result should be further confirmed by genetic rescue. The authors can use a UAS rescue line previously reported in PMID: 19341724.

      We tested the UAS-Atl-myc rescue line and unfortunately found that even in wild-type neurons, overexpression of Atlastin produced strong ER organization defects that precluded the rescue experiment. Instead, to confirm the cell autonomy of the phenotype and to test it wth an independent tool, we performed a presynaptic knockdown of Atlastin by RNAi and found that BiP:sfGFP:HDEL is displaced, as observed in the Atlastin null mutant. These results are in now shown in Figure 7A-C.

      Lines 204-7: It's not clear how a greater coefficient of variation indicates that the marker is more concentrated in subsynaptic structures or what is meant by 'subsynaptic structures.'

      We added the following text to explain, in lines 181-183: "A higher CoV indicates an uneven distribution of tdTomato:Sec61β within the presynaptic terminal, with some areas showing higher concentrations than others (in contrast to the uniform, diffuse signal expected from fragmentation)." To avoid confusion with postsynaptic structures called the subsynaptic reticulum, we have removed the term "subsynaptic". The intended meaning is distinct structures found within the presynaptic terminal.

      There's a mistake in Figure 6C and the associated text. The summed percentage of the three phenotypic categories adds up to 110% for Atlastin mutants. *

      The reviewer noted that the summed percentage of the three phenotypic categories in Figure 6C adds up to 110% for Atlastin mutants, which appears to be a mathematical error. However, this is not an error, but rather a reflection of our quantification methodology, in which a single bouton can exhibit more than one type of ER dynamics per movie recorded. Our quantification counts each phenotype independently, so boutons displaying multiple phenotypes contribute to more than one category. This approach provides a more comprehensive view of the range of ER dynamics present in Atlastin mutants, as restricting the analysis to mutually exclusive categories would underrepresent the complexity of the phenotypes observed. To make this point clear, we made the following change to the text in lines 257-259: "We note that the sum of these percentages exceeds 100% because one NMJ exhibited multiple phenotypes: one branch had a complete loss, while the other branch had no phenotype. These phenotypes were counted separately."

      Figure 8: the ER looks fragmented in 1st instar controls and mutants. The authors should address this difference from more mature NMJs.

      We would like to clarify that the bulk of experiments in this manuscript (including all ER dynamics, luminal marker redistribution, and membrane marker analyses discussed throughout the Results) were performed in 3rd instar larvae, which are more mature larval NMJ preparations standard in the field. Figure 8 was included specifically to test whether the Atlastin mutant phenotype we describe throughout the paper is also detectable at an earlier developmental stage, not to replace or reinterpret our primary findings.

      Regarding the specific observation that the ER appears more fragmented in Figure 7F-H relative to the more mature NMJs shown elsewhere: this fragmentation, observed similarly in both control and Atlastin mutant 1st instar larvae, likely reflects technical challenges associated with dissecting these smaller, more delicate early-stage specimens rather than a genotype-specific effect. Because fragmentation occurred similarly in both genotypes, we could still reliably assess the redistribution of BiP:sfGFP:HDEL as our primary phenotypic readout in this experiment. We have added the following text (lines 306-309) to clarify this point: "Note that in 1st instar larvae, both normal networks in controls and residual networks in Atlastin mutants appeared more fragmented than in 3rd instar preparations, likely due to the technical challenges of dissecting these smaller, more delicate specimens. Since ER fragmentation occurred similarly in both genotypes, we could still reliably assess the redistribution of BiP:sfGFP:HDEL as our primary phenotypic readout.

      The images in figure 9B do not seem representative of the quantification in Figure 9D. Specifically, the partial loss Atlastin NMJ appears to have recovered as fully as the complete loss Atlastin NMJ.

      The images showed FRAP recovery across the entire bouton, but we photobleached only a small region within each bouton and quantified only this region. We have now added outlines to clearly delineate the specific FRAP regions that were analyzed in each image, which clarify that the partial loss Atlastin showed less recovery than the overall bouton. We have also reordered the figures to more clearly convey our message (Figure 9 is now Figure 8).

      We also made a few changes to the paragraph on lines 347-350 to clarify our experimental reasoning: "We photobleached en passant boutons using a defined region of 6.8 x 7.8 microns (dashed box in Figure 8D) to ensure that BiP:sfGFP:HDEL could recover from the ER networks surrounding the FRAP region (Movies S20-S23)."

      We also added this sentence to the figure legends of Figure 8: "The dashed boxes in (D) indicate areas that were photobleached and analyzed for recovery quantification in (E-F)."

      Optional: An overexpressed luminal marker is displaced to the cytoplasm in Atlastin mutants. It would be interesting to know and increase the significance of the findings if the same is true of endogenous luminal proteins under biological stress conditions.

      As noted in our response to Reviewer #1 suggested that Atlastin mutant synapses may exhibit ER stress. To address this, we examined levels of the ER chaperone BiP, a well-established ER stress marker whose expression increases during UPR activation. We first validated that our BiP antibody can detect changes in ER stress by feeding control larvae with 50mM DTT for 24 hours. We were unable to perform this experiment in Atlastin mutant larvae because they did not consume the DTT-treated food, as assessed by blue food coloring in the larvae's guts. These results are in Figure 10A. In the future, it will be of interest to establish a protocol to examine Atlastin mutants by feeding or treating larval fillets with DTT.

      We measured BiP levels at NMJs of Atlastin mutants and found they were slightly increased compared to controls. Atlastin mutants co-expressing UAS-BiP:sfGFP:HDEL or UAS-tdTomato:Sec61b did not show significantly increased endogenous BiP levels, suggesting that transgene expression suppresses the mild ER stress response. We conclude from these experiments that Atlastin mutant synapses have mild ER stress. These results are in Figure 10B-E).

      Optional: Applying this approach in stimulated conditions (high potassium, increased temperature) might reveal a greater activity-dependent role for Atlastin at synaptic terminals.

      This is a very interesting idea, as we have only examined synapses at rest. However, this is beyond the scope of this paper.

      Minor Comments

      1. Line 16: Atlastin should be italicized.

      Thank you for catching this typo. We have fixed it.

      Figure 5A: Based on the relative intensities, it appears that control and mutant images are not contrast matched but this isn't stated.

      Thank you for catching this omission. We added to the figure legend: "Control and Atlastin mutant images are not contrast matched."

      Line 822: The number of static Atlastin mutant boutons used for analysis is missing.

      Thank you for catching this omission. We have fixed this supplementary table.

      Figure 9: The blue arrows are not annotated in the figure legend.

      Thank you for catching this omission. We have fixed this figure legend.

      Reviewer #2 (Significance (Required)):

      Atlastin is linked to Hereditary Spastic Paraplegia (HSP) and this study changes our understanding of the compartment-specific impacts of its loss. This study reveals the importance of using both membrane and luminal ER markers to accurately interpret phenotypes as well as the importance of considering compartment-specific effects on ER. These findings represent significant mechanistic and conceptual advances. The lack of genetic rescue is a limitation and adding an investigation of an endogenous luminal protein under basal and stress conditions would add significantly to our understanding of Atlastin dysfunction in HSP. Notably, the in vivo imaging approach introduced here can be adapted broadly for live imaging of Drosophila larvae. Thus, this work will be of interest to both neuronal cell biologists and the wider Drosophila community. This review is based on our expertise in neuronal cell biology.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      In this manuscript, the authors investigate the structural dynamics of the endoplasmic reticulum (ER) in Drosophila neurons and examine the role of the ER-shaping protein Atlastin in ER morphology. Their discovery on the neuromuscular junction (NMJ)-specific contribution of Atlastin to ER integrity is intriguing and may provide valuable insights into the pathological mechanisms underlying Atlastin mutations associated with hereditary spastic paraplegia (HSP) and hereditary sensory neuropathy. The key observation on ER protein showing an aberrant cytoplasmic localisation in mutant cells appears convincing. Though this phenomenon's characterisation stays at the point of primary observation with its mechanics unclarified, establishing this new and unexpected functional rather than structural Atl effect is important and useful for the field. The observation that ER is structurally preserved in this mutant with absolute lack of Atl are also extremely useful.

      It is unclear if the cytoplasmic localisation affects an exogenous overexpressed ER marker or endogenous protein would also appear in cytoplams, the authors should consider adding an immunostaining data to test that.

      Authors offer speculations on potential reasons for the cyto localisation of the ER marker suggesting that relocation at the cell periphery specifically combined with slow clearance there is the most likely explanation (still unclear what stops the marker from spreading through the entire cell). They suggest that decrease in cotranslational translocation is unlikely as this would result in somatic accumulation of the marker. However, if the clearance in the periphery is less efficient than in soma, the accumulation there might reflect a compromised translocation. Any clarifying experiments, if practical, to directly demonstrate how ER proteins in relocates to the cytoplasm in atl mutant would help understanding better the phenomenon. For example, would proteasomal inhibition make the marker accumulate more across the cell? Authors also suggest links to ER stress. Would stress induction phenocopy the mutant?

      Reviewer #3 asked whether defective proteasomal clearance underlies the cytosolic accumulation of BiP:sfGFP:HDEL in Atlastin mutants. We addressed this directly. First, proteasome function appears intact in the mutants: baseline ubiquitinated protein levels (FK1 antibody) were comparable between control and Atlastin mutants, and MG132 treatment produced a similar increase in ubiquitination in both genotypes, confirming both antibody specificity and normal proteasome activity. We then examined BiP:sfGFP:HDEL directly. In controls, MG132 caused the marker to accumulate at axons and presynaptic terminals, showing that it is normally cleared from these compartments by the proteasome. Critically, this accumulated marker remained associated with intact ER networks: MG132 did not induce diffuse cytosolic BiP:sfGFP:HDEL in any compartment (cell bodies, axons, or presynaptic terminals), even where levels rose substantially. Thus, blocking proteasomal clearance raises ER-localized marker but does not generate the cytosolic pool seen in Atlastin mutants, indicating that impaired clearance is not sufficient to cause the displacement phenotype. We separately noted that BiP:sfGFP:HDEL was already elevated in Atlastin mutant axons without MG132, paralleling the axonal tdTomato:Sec61β accumulation in Figure 4, consistent with reduced baseline clearance specifically in mutant axons, but this does not lead to cytosolic displacement. This experiment is now shown in Figure 11, described in Results (lines 445-475), and discussed in lines 576-581.

      Minor comments:

      Line 146:

      "fast dynamics (Thank you for catching this mistake. We have corrected it.

      Fig. 2D: The data representation of "Tubule displacement" image is unclear. The ER tubule indicated by the red arrow does not seem to show any changes over time (like static). time 0 in stamp appears behind the image.

      Thank you for catching the typo. We have fixed it. Additionally, we added black arrows to highlight a tubule that is not moving, allowing the reader to compare it with the moving tubule. We also included a video of all types of ER tubule dynamics to ensure the reader can also look at the raw data (Movies S9-11).

        • Line 157-158 (and relevant method sections):

      The definition of static and dynamic boutons is ambiguous. The author should describe in more detail this point including how long they observed the structure to define the changes in ER tubule dynamics.

      We provide in the methods (lines 779-791) a detailed explanation of how we categorized boutons as dynamic or static. In addition, we added the following to explain in the results section how we defined static vs dynamic:

      Old sentence: We qualitatively categorized boutons as "static" if we observed no change in ER network structure or "dynamic" if we observed at least one change.

      New sentence in lines 143-147: "We imaged boutons for 40 sec at 0.92 sec intervals to capture ER dynamics over this observation period. Boutons were qualitatively categorized as "static" if we observed no detectable changes in ER network structure throughout the entire 40 sec imaging session, or "dynamic" if we observed at least one of the three defined dynamic events during this time window."

      Fig. 2E: What n=75 and n=29 represent is unclear, are these the number of boutons in en passant and terminal subjected for qualitative analysis?

      We removed these n values from the figure and added this information to the Supplementary Table 1, which contains detailed information about the genotype, statistical analysis, and number of larvae and NMJs analyzed.

      Fig. 2: What the qualitative analysis represents is unclear, are the points pulled from different experiments?

      The data in Fig. 2 E-F comes from movies acquired in the same experiment. The number of independent animals and NMJs imaged is described in Table 1.

      * *Line 231: Regarding "...we found a small but significant reduction in dynamic boutons in Atlastin mutants (76%), ...", how do the authors assess significance. If proportion of static/dynamic ER in boutons was obtained from multiple experiments, it should be presented e.g. as in average {plus minus} standard deviation, or clarify that the proportion is representative of x independent experiments.

      The videos used for this figure were acquired from a single experiment. We use a chi-square test to determine significance relative to the "expected" distribution of dynamics types from controls, as these are categorical rather than continuous data (see PMID 31145670). Information regarding genotype, statistical analysis and number of larvae and NMJs can also be found in Supplementary Table 1.

      Line 267-269 and Fig. 6B: The author's conclusion that "Complete loss of ER network structure in NMJ of BiP:sfGFP:HDEL overexpressing Atl mutant" seem to be based on the lack of signal from luminal marker, which may be undetectable due to changes to tubular volume or marker loss to the cytoplasm, as suggested by the authors, while the membranous ER structure is intact. It would be useful to discuss this point and potentially add ER membrane-stained control.

      We agree with the reviewer that Atlastin mutants categorized as 'complete loss mutants' do not actually lack ER at synapses. We think this is an important point so we added the following to the results in lines 251-254: "Note that the "Complete loss" phenotype in Atlastin mutants reflects the absence of detectable luminal marker signal in organized ER structures, not the complete absence of ER membranes, as demonstrated by our ER membrane marker tdTomato:Sec61β results."

      We attempted to co-label the ER membrane and ER lumen, but these crosses yielded very few live larvae (in either controls or Atlastin mutants, and those that survived had severely deformed NMJs. We added Figure 6-Supplement showing the results of this experiment, and described them on lines 270-273.

      Fig. 6C: In Atl mutant, why does the total of the proportion exceed 100% (10 + 45 + 55)?

      The reviewer noted that the summed percentage of the three phenotypic categories in Figure 6C adds up to 110% for Atlastin mutants. This is not an error, but rather a reflection of our quantification methodology because a single bouton can exhibit more than one type of ER dynamics per movie recorded. Our quantification counts each phenotype independently, so boutons displaying multiple phenotypes contribute to more than one category. This approach provides a more comprehensive view of the range of ER dynamics present in Atlastin mutants, as restricting the analysis to mutually exclusive categories would underrepresent the complexity of the phenotypes observed. To make this point clear, we made the following change to the text in lines 257-259: "We note that the sum of these percentages exceeds 100% because one NMJ exhibited multiple phenotypes: one branch had a complete loss, while the other branch had no phenotype. These phenotypes were counted separately."

      Fig. 9C, line 342-344: In FRAP experiment using CD8, it seems that the Partial loss Atl mutant shows slower recovery that control. There seems to be a mismatch in triangle symbols of Partial loss Atl mutant between legend and plot (one is filled and the other is empty). This should be clarified.

      Thank you for catching this mistake. We have fixed the figure.

      fig. 10 is a clever way to verify the cytoplasmic localization of the ER marker; however, its description and annotation can be improved, and it would be stronger if 4 curves in F for mutant and controls with the trap and normal were shown.

      The reviewer suggested merging our graphs but we believe that keeping them separate is clearer.

      Line 495: Drosophila have ReepA and ReepB, but not Reep1-4. If the authors discuss their speculation based on their observation (using Drosophila), the gene names should be unified in the same species, and explain the corresponding genes to mammalian cells.

      We made the following changes to address the reviewer's concern about gene nomenclature consistency (lines 502-506): "These ER-derived vesicles are likely to involve ReepA and ReepB, the Drosophila orthologs of mammalian REEP1-4, which regulate ER vesicle formation in mammalian cells (67). Notably, while overexpression of Atlastin can regulate REEP vesicle fusion in mammalian systems (67), it is not essential for vesicle formation, suggesting similar regulatory relationships may exist between Atlastin and Reep genes in Drosophila."

      Line 548; should UPR be Unfolded Protein Response?

      Thank you for catching the typo. We have fixed it.

      Reviewer #3 (Significance (Required)):

      This study advances the understanding of how ER morphogens affect neuronal cells specifically, the lack of which limits researchers ability to comprehend the neuronal pathologies associated with ER structure-function. The observation on ER content aberrant localisation caused by the lack of key structural protein should be of a great interest for cell and neuronal biologists and researchers of the associated diseases and shows the field a new direction. Though, mechanistic details remain to be unraveled, it constitutes a fundamental, conceptual advance.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used, and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore a useful and broadly applicable statistical tool is well placed in this journal.

      Strengths:

      The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

      The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

      The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

      The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

      We appreciate the reviewer’s summary of our study and its strengths.

      Weaknesses:

      A few changes could be made if the manuscript is revised. I would consider all of these points minor, but the paper could be improved if these points were addressed.

      (1) The caption of Figure 2 doesn't seem to mention panel D. Figure A-2 also does not mention C in the caption.

      We apologize for this error, and thank you for catching it! The figure legends had missing or incorrect panel labels. This error has been corrected.

      (2) Figure 2D is a little hard to follow. First, the definition of "Power" is not clear, and I couldn't find the precise definition in the text. Second, the legend for the different lines in 2D is only given in Figure A-2. Perhaps a portion of the caption for Figure 2 is missing?

      We have added a definition of power in the main text:

      “Although the permutation test, simultaneous permute-match test, and sequential permute-match test are all valid, they vary in power – the probability of detecting true dependence.”

      We have clarified the use of “power” in legend of Fig 2 and clarified that the color key for Fig 2D is in Fig 2A. The relevant excerpt of the Fig 2 legend is copied here:

      “(D) Statistical power for the permutation test and various permute-match tests as a function of the replicate number n, significance level α, and strength of dependence r<sub>X, Y</sub>. Power was estimated as the proportion of simulations in which dependence was detected, calculated from 5000 simulations at each value of r<sub>X, Y</sub> between r<sub>X, Y</sub> = 0 and 0.54 in steps of size 0.01. At r<sub>X, Y</sub> = 0, there is no dependence, so the curve at that point indicates the false positive rate rather than power. We chose the Pearson correlation coefficient as our correlation function ρ. See (A) for the color legend.”

      We have also added dotted lines connecting the legend in panel A to the curves in panel D.

      (3) The concept of circular variance for the fish data was heard to understand/visualize. The equation on line 326 did not help much. If there is a very simple picture that could be added near line 326 that helps to explain Ct and theta, that could be a big help for some readers who do not work on related systems. The analysis performed is understandable, the reader just has to accept that circular variance captions the degree of alignment of the fish.

      We have replaced references to circular concentration with “mean resultant length”, which is the standard jargon for this term in circular statistics, and we have added an illustration.

      (4) For the data discussed in Figure 3, I wasn’t 100% sure how the time windows were selected. In the caption, it says “time series to different lengths starting from the first frame”. So the 20 s time window was from t=0 to t= 20 s. Would a different result be obtained if a different 20 s window was chosen (from t = 4 min to t = 4 min 20 s just to give a specific example). I suppose by chance one of the time windows would give a pvalue less than the target 0.05, that wouldn’t be surprising. Maybe a random time window should be selected (although I am not indicating what was reported was incorrect)? A little more discussion on this aspect of the study may be helpful.

      As suggested by the reviewer, we have redone the analysis of Figure 3D with random segments. This provides a more complete picture of how the chance of detecting a significant correlation varies with segment length. The main conclusion is unchanged: Perfect match tests reliably detect dependence across a wider range of segment lengths than the naive parametric alternative.

      The relevant panel and an excerpt from the legend text are copied below.

      “(D) Permute-match tests detected a significant correlation between speed and alignment more consistently than the parametric test. For a grid of lengths between 20 and 600 seconds we sampled 500 random segments of each length, each drawn from the first 600 seconds, and determined for each segment whether the parametric test and/or the two possible permute-match tests detected a significant (p ≤ 0.05) correlation. In the edge case of the maximum 600-second length, all 500 “random” segments were identical.”

      Reviewer #2 (Public review):

      Summary:

      This paper presented a hypothesis testing procedure for the independence of two timeseries that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

      Strengths:

      The test offers good flexibility for different kinds of dependence (through adjusting \rho), and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

      We appreciate the reviewer’s summary of key aspects of our manuscript.

      Weaknesses:

      (1) The size of the test is not guaranteed to (asymptotically) equal \alpha, which may damage the power.

      We thank the reviewer for raising the issue of test size and power. We agree that a conservative test (one whose size can fall below alpha) may sacrifice power.

      Our objective is distribution-free false-positive rate (FPR) control. That is, we wish to keep the FPR at or below alpha for every distribution of X and Y, because in our regime (nonstationary time series with few independent replicates) the scientist often cannot verify distributional assumptions. Inspired by the reviewer’s comment, we now show (new Proposition 14) that the perfect match probability can be made arbitrarily close to 1/n<sup>!</sup>. As a consequence, any reported perfect match p-value below 1/n<sup>!</sup> would break the distribution-free validity of the test.

      A test that exploits distributional structure could likely access lower p-values; we have now explored how the empirical FPR of the permute-match test varies with the data-generating process (see our response to reviewer 2's recommendation 1 below).

      (2) The computational time can be an issue for a moderately large sample size when calculating the X / Y-perfect match. It will be beneficial to include discussions on the implementations of the test.

      We agree this is an important consideration. We have added the following text to the Discussion:

      “The test appears computationally tractable for relevant sample sizes: Our implementation of the permute match procedure completed a single test of dependence in the setting of Fig 2 with an average runtime of 3 seconds when n = 10 on a 2023 14-inch MacBook Pro with an M2 Pro processor and 16 GB RAM (see Source data 1). For n > 10, a standard permutation test already can report a p-value below 3 × 10<sup>−8</sup> so the perfect match test is likely unnecessary for typical applications.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A few more minor notes/comments:

      As a personal preference, I like it when figures printed in grayscale retain their meaning (when possible). Just FYI, Figure 3D in grayscale is uninterpretable. Not saying a change is needed, just pointing it out.

      We changed Fig 3D to address a comment above, and think the new version better distinguishes between the permute-match and permutation test results in greyscale.

      On line 304, is that a lower bound or an upper bound? Maybe the issue is the probability mentioned on line 305 is not clear.

      As this point is not the main focus of the investigation, we have rephrased it to make it less technical and eliminate the issue of which bound is in question.

      “It seems likely that the tests could be further modified to report an even lower p-value when an X- and Y -perfect match occur simultaneously, as in Fig 3B. However, we have not investigated further and this problem is left for future efforts.”

      I am not sure if this is a weakness, but the p-value changing depending on the choice of whether to apply the X-perfect match or Y-perfect match test first is fascinating. The authors did discuss this very issue at several points in the manuscript. It is slightly unsettling to me that there isn't an exact p-value for a given set of data. This one point gives me a new perspective on statistics.

      In full transparency, I don't believe I have to background to thoroughly review the appendix. I did read through it and did not notice any errors, but I couldn't confidently say there are not any small mathematical errors or any logical flaws in the proofs. Some sections were not easy to follow (my own shortcomings, the writing appeared sufficient for more of an expert to understand).

      We greatly appreciate the reviewer’s time and effort tackling an appendix outside their comfort zone.

      Reviewer #2 (Recommendations for the authors):

      (1) In the numerical experiment session, the authors should include the null situation, i.e., the performance of the test when X and Y are independent. This helps assess the size of the test.

      We have added a section on size to our results section, copied below:

      “The permute-match test’s false positive rate depends on the process tested. The permute-match test is conservative – meaning that its false positive rate can fall below the significance level – because both the permutation test and perfect match test are conservative. As discussed elsewhere [29], the permutation test is conservative when α is not one of its possible p-values and when ties may occur between the original correlation and shuffled correlations. Checking for a perfect match is similarly conservative. The actual probability of a false-alarm perfect match event can vary depending on the process being tested. To see this consider the permute-match test in the setting where n = 3, where α = 0.05, and where r<sub>X,Y</sub>= 0 (independent X and Y). Note that in this case, obtaining a Y -perfect match (and thus p = 1/n<sup>n</sup>) is necessary and sufficient to detect dependence since α is too low for detection by either the permutation test or the p = 2/n<sup>n</sup> leg of the permute-match test. In the linear system of Fig 2, we observed among 5000 simulations a detection rate of 0.0148, significantly below the upper bound of 1/3<sup>3</sup> (Figure 2 - Source data 1; one-tailed exact binomial test, p < 10<sup>−20</sup>). Conversely, in the nonlinear system of Fig S2, this same event (Y -perfect match under n = 3 and r<sub>X,Y</sub> = 0) occurs with a detection rate of 0.0328, not significantly 9 below the upper bound of 1/3<sup>3</sup> (Figure S2 - Source data 1; one-tailed exact binomial test, p = 0.059). Thus, depending on the underlying process studied, the actual chance of a perfect match happening under independence may be near or significantly below the theoretical upper bound.”

      (2) Some insights regarding the choice of rho should be provided. Especially, are there any examples that the classical test, such as the Pearson correlation or Granger causality test does not work?

      We have redone the example of Appendix 4 with Pearson correlation, showing that Pearson correlation has substantially lower power than cross-map skill in this case (compare figures S2 and S3).

      (3) Line 34 - 35, page 2: Correlations and causality should be separately considered. This sentence talks more about causality rather than correlation.

      We appreciate the reviewer’s perspective and agree that correlation and causality are distinct.

      We feel that pointing out the issue of spurious correlations is helpful to orient our readers, especially those from a broad scientific audience. In the text, we define “correlation” as a descriptive statistic (rather than normalized covariance), and later distinguish it from “dependence”, which has causal implications due to Reichenbach’s common cause principle. We believe this distinction provides a useful backdrop for practitioners who use statistical methods but are perhaps new to thinking deeply about statistical dependence.

      (4) Please add some discussions on the situation that X_i depends on Y_{i - j} for some j > 0, which is associated with the setting of Granger causality test.

      We have added the following to the discussion:

      “No distributional assumptions are required, and the correlation function ρ can be completely arbitrary. For instance, ρ could include a lag to detect delayed dependence, or even evaluate the correlation strength at several lags and report the strongest among them [39].”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Seegren and colleagues demonstrate that in a mouse model of neonatal E. coli meningitis, loss of endothelial toll-like receptor 4 (TLR4) leads to a marked decrease in transcriptional dysregulation across multiple leptomeningeal cell types, a decrease in vascular permeability, and a decrease in macrophage abundance. In contrast, loss of macrophage TLR4 had less pronounced effects. Using cultured wild-type and TLR4knockout endothelial cells, the authors further demonstrate that TLR4-NF-κB signaling leads to reversible internalization of the tight junction protein claudin-5, establishing a potential mechanism of increased vascular permeability. Finally, the authors use RNA sequencing of wild-type and TLR4-knockout endothelial cells to define the TLR4dependent cell-autonomous transcriptional response to E. coli.

      Strengths:

      (1) The authors address an important, well-motivated hypothesis related to the cellular and molecular mechanisms of leptomeningeal inflammation.

      (2) The authors use model systems (mouse conditional knockouts and cultured endothelial cells) that are appropriate to address their hypotheses. The data are of high quality.

      Weaknesses:

      (1) The authors perform single-nucleus RNA-seq on dissected leptomeninges from control and E. coli-infected mice across three genotypes (WT, Tlr4MKO, and Tlr4ECKO). A major discovery from this experiment, as summarized by the authors, is: "Tlr4ECKO mice exhibited a global attenuation of infection-induced transcriptional responses across all major leptomeningeal cell types, as judged by the positions of cell clusters in the UMAP." This conclusion could be considerably strengthened by improving the qualitative and quantitative analysis.

      Thank you for this comment. We agree that the UMAP-based interpretation would benefit from additional qualitative and quantitative support. We have expanded the snRNA-seq analysis with additional images and supplemental figures (Figure 1 – figure supplement 3, Figure 1 – figure supplement 5, and Figure 1 – figure supplement 6). The first and third of these new supplemental figures show dot plots for each major leptomeningeal cell type, for each genotype, for the two experimental conditions (infected vs. uninfected), and for individual genes in three immune-related gene sets (NF-kB and TNF-α, JAK-STAT, and IFN-ɣ), providing a more explicit comparison of infection-induced transcriptional responses across genotypes. The second of these new supplemental figure shows principal component analysis (PCA) of the individual snRNA-seq datasets (one mouse per dataset) for each genotype and experimental condition, demonstrating that the observed transcriptional shifts are consistent across biological replicates. Finally, Figure 1 – figure supplement 7, which was included in the original submission, shows changes in the most up- and down-regulated genes (based on adjusted p-value or fold change) in endothelial and myeloid cells across individual mice and genotypes/conditions, further supporting the genotype-dependent effects at the level of individual animals.

      (2) The authors interpret E. coli infection-induced increases in leptomeningeal sulfo-NHSbiotin as evidence of compromised BBB integrity (i.e., extravasation from the vasculature) (Results, page 7), but another possible route in this context is sulfo-NHS-biotin entry from the dura across a compromised arachnoid barrier. The complete rescue in Tlr4ECKOs is strongly suggestive that the vascular route dominates, but it would strengthen the work if the authors could assess arachnoid barrier fidelity (e.g. via immunohistochemistry). At a minimum, authors should mention that the sulfo-NHS-biotin signal in this context may represent both vascular and arachnoid barrier extravasation.

      Thank you for this comment. We agree that our data cannot rule out leakage across the arachnoid barrier during infection. While the rescue observed in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice strongly supports a dominant vascular contribution, we acknowledge that the sulfo-NHS-biotin signal may reflect permeability at both the vascular and arachnoid barriers. We do not think there is a clear way to directly test this possibility functionally, since the arachnoid barrier appears intact by confocal microscopy. Subtle differences in barrier cell morphology might be detectable by electron microscopy, but this would not definitively address whether infection permits molecular passage across the arachnoid barrier. We have followed the reviewer’s suggestion and revised the Results section to reflect this interpretation. Specifically, we added the following: “The simplest interpretation of these data is that the site of sulfo-NHS biotin leakage is primarily vascular. However, we cannot exclude some contribution from increased arachnoid barrier permeability.”

      (3) The authors state that "deletion of TLR4 prevented both NF-κB nuclear translocation and Cldn5 internalization in response to E. coli (Figure 4A-D)" (Results, page 9). In Figures 4C and D, however, there is no indicator of a statistical test directly comparing the two genotypes. A comparison of within-genotype P-values should not be used to support a genotype difference (PMID: 34726155).

      Thank you for pointing out this omission. We have updated the figures so that the between-genotype p-values are shown for those panels (including this panel) that had not previously shown them.

      (4) In the first paragraph of the Results, the authors summarize the meningeal layers as (1) pia, (2) subarachnoid space, (3) arachnoid, and (4) dura, and then state "The second and third layers constitute the leptomeninges." This definition of leptomeninges seems to omit the pia, which is widely considered part of the leptomeninges (PMID: 37776854).

      Thank you for pointing out this error, which has now been corrected.

      (5) The Cdh5-CreER/+;Tlr4 fl/- mouse lacks TLR4 in all endothelial cells (i.e., in peripheral organs as well as CNS/leptomeninges), and, as the authors note, the periphery is exposed to E. coli. It would be helpful if the authors could comment in the Discussion on the possibility that peripheral effects (e.g., peripheral endothelial cytokine production, changes to blood composition as a result of changes to peripheral endothelial permeability) may contribute to the observed leptomeningeal phenotypes.

      Thank you for raising this point. We agree that peripheral responses could contribute to the observed leptomeningeal phenotypes in this model. We have added two sentences to the second paragraph of the Discussion to address this: “We note that these experiments do not distinguish between local vs. distal anatomic sources of LPS or downstream effector molecules, such as cytokines, that activate the leptomeningeal inflammatory response (Huang et al., 2021). Histologic observations of RFP-expressing E. coli in the brain, liver, and lungs, together with positive blood cultures, indicate substantial systemic dissemination in this model. Thus, the inflammatory responses of leptomeningeal cells likely reflect exposure to bacterial products and inflammatory mediators derived from both local meningeal and peripheral sources.”

      Reviewer #2 (Public review):

      Summary:

      The authors use a postnatal mouse model of E. coli bacterial meningitis and a mouse brain endothelioma cell line combined with cell-type-specific gene deletion to study the function of endothelial TLR4, a cell surface receptor that recognizes gram positive bacterial wall components, in the local leptomeningeal (LPM) response with a focus on endothelial barrier breakdown mediated by TLR4. Single-cell transcriptional profiling and imaging studies using whole-mount preps of the LPM support that LPM endothelial, CD206+ local macrophage and LPM fibroblast and arachnoid barrier cell inflammatory response and is abrogated in endothelial-specific KO of TLR4, pointing to a role for endothelial TLR4 in local LPM response. Culture studies using Bend3.1 cells (a mouse brain endothelioma cell line) support a direct role for TLR4 in the bacteria-mediated inflammatory response and in internalization of Cldn5 via the endosomal-lysosomal pathway, resulting in loss of barrier integrity

      Strengths:

      The local LPM cell response in meningitis and the role of specific LPM cells in inflammation and CNS barrier breakdown have not been extensively studied, despite ample evidence for primary immune response in the meninges in human patients and in animal models. The authors employ a robust, multi-model approach using both in vivo and in vitro models with cell-type-specific knockout to study the function of TLR4 in brain endothelial cell response. The authors nicely combine functional barrier assays with IF for junctional localization in their experimental design, and they delve into potential mechanisms of Cldn5 internalization using markers of endosomal-lysosomal pathway localization. The authors also describe a new type of barrier assay using a streptavidin-coated plate upon which barrier-forming cell cultures can be placted, this could be a very useful alternative or complement to other size-selective barrier assays and presumably could work for other barrier forming cells types, likely epithelial cells.

      Weaknesses:

      (1) There are no measures of bacterial burden in peripheral organs, blood, in the LPM or brain in the TLR4 endothelial cKO mice. Lack of TLR4 in endothelial cells could prevent bacterial 'access' into the LPM and brain, essentially preventing meningitis and leading to a lack of inflammatory responses in the LPM-located cells simply because there is no bacteria present. Bacteremia may also be reduced, as might inflammatory responses in peripheral organs with TLR4-deficient peripheral endothelium. Bacterial counts and inflammatory measures in peripheral organs and blood are important to better understand the mechanism(s) underlying the reduced inflammatory profile in LPM cells and no LPM endothelial breakdown in the Tlr4 endothelial cKO mice. In other words, does deleting TLR4 in EC protect against the development of meningitis by somehow blocking bacteria access to the LPM (this would be supported by low or no CFU counts in infected Tlr4 endothelial cKO) or is it what the authors appear to propose in Figure 1J that TLF4 in EC is the only cell responding to the bacteria to trigger the immune cascade in the LPM? More data is needed to resolve this, as this is a major claim of the paper.

      Thank you for this comment. We agree that it is important to distinguish whether the reduced inflammatory response in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice reflects altered bacterial burden versus altered host sensing. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4CKO mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs of infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice in E. coli burden; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and the time of sacrifice 24 hours later (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid leptomeningeal cells does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (2) The authors look at the underlying cortical response (cerebral vasculature for ICAM and immune cells) but do not use markers that could identify microglia (Iba1), the primary resident immune cell (CD206 is not useful, at this stage, in perivascular macrophages that are extremely sparse in the postnatal brain). This would be important to better study the impact on CNS resident immune cell morphological activation.

      Thank you for this comment. In response, we have analyzed Iba1 staining in the cortex in infected vs. uninfected mice. This is shown in Figure 2 – figure supplement 3. These data demonstrate a several-fold increase in Iba1 immunostaining in infected compared to uninfected cortex, consistent with increased microglial activation in response to infection. There is no statistically significant difference between infected WT and infected Cdh5-CreER; Tlr4CKO mice in Iba1 staining in cortex.

      (3) The authors suggest that Cldn5 junctional localization is selectively disrupted upon bacterial exposure, mediated by TLR4 - they suggest this based on studying PECAM, GLUT1, ZO-1 and B-catenin (all normally junction or cell surface located in cultured Bend3.1) in relationship to Cldn5 localization (normally high) - it is possibly these are also impact by bacteria exposure (maybe through different mechanisms?) - a better measure would be to use the similar cyto/PM measure they do for Cldn5 in Fig. 4D and to evaluate this or to use intensity measurements.

      Thank you for this comment. As the reviewer noted, the analysis of Cldn5 localization with vs. without E. coli exposure and in WT vs. Tlr4KO bEnd.3 cells (shown in Figure 4B and D) – uses Cell Trace to partition the image into cytoplasmic vs. plasma membrane territories. For the analyses in Figure 5, we wanted to compare the localization (and potentially re-localization) behaviors of a variety of subcellular markers with the localization and re-localization of Cldn5 following E. coli exposure. By directly measuring the % overlap of the two immunostains, we get that data. We note that the goal of this analysis is to assess relative co-localization with Cldn5 rather than absolute subcellular partitioning of each marker. While this analysis could have been extended to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories, it is clear by visual inspection of Figure 5A-C that beta-catenin, ZO-1, and PECAM1 remain plasma membrane-associated with E. coli exposure, and GLUT1 goes from the part of the plasma membrane not involved in cell-cell contact without E coli exposure to cytoplasmic with E. coli exposure (as judged by the appearance of a nuclear “shadow” after E. coli exposure). Thus, we do not believe that additional cytoplasmic vs. plasma membrane quantification for these markers would alter the interpretation. The main reason that we did not extend this analysis to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories is because that would introduce the Cell Trace localization as an additional variable.

      (4) The discussion could benefit from delving more into the prior literature on E coli mediated breakdown of junctions in cultured human microvascular brain endothelial cell model and critical host-pathogen interactions of the bacteria with ECs (PMID: 14593586), and how this might involve TLR4.

      Thank you for this comment. Two paragraphs addressing the prior literature have now been added to the discussion.

      (5) It would be important to discuss how their results relate to earlier studies on TLR4-/- and TLR2-/- global knockout mice and protection vs vulnerability to development of meningitis (see PMCID: PMC3524395) - this paper showed that TLR4 global KO mice have increased susceptibility to die from meningitis and have much higher CFU counts in the CNS. In this manuscript and their prior work (Wang et al., 2023), this group shown that both global TLR4-/- mutants and their EC-specific KO have reduced barrier permeability, but we don't have any information about CFU or susceptibility to death from meningitis in their models.

      Thank you for these comments. The model we use – subcutaneous injection of E. coli (a clinical isolate from an infant with meningitis) at postnatal day (P)5 – results in the death of the infected mouse within 2 days (shown in Figure 1 – figure supplement 3 in Wang et al. 2023). Our analyses of infected mice were conducted 24 hours after infection. As noted in the reply to comment #1, in the revised manuscript we present a clinical assessment of WT vs. Cdh5-CreER; Tlr4CKO mice 24 hours after infection based on (1) a quantitative microscopic analysis of E. coli burden in the brain (visualized based on RFP fluorescence in the E. coli used here), (2) quantifying CFUs in blood and (3) mouse weights at P5 and P6, a sensitive indicator of overall health since this is a time when mice are normally gaining weight rapidly (~25% weight gain per day). These data (shown in Figure 2 figure supplement 4) indicate that bacterial burden and disease severity are similar between genotypes in our model. In Wang et al., 2023, we did not conduct a quantitative clinical assessment of WT vs. Tlr4-/- mice following infection, but by visual inspection, infected WT and Tlr4-/- mice appeared to have similar downhill clinical trajectories. We have expanded the Discussion to relate these findings to prior studies of global TLR4 and TLR2 knockout mice, noting that differences in experimental models and the distinction between global versus VECadCreER-specific deletion may account for the differing outcomes reported.

      Comment on the paper listed by the reviewer (PMCID: PMC3524395).

      The cited study demonstrates that global TLR4 deficiency leads to increased bacterial burden and mortality, indicating an essential role for TLR4 in host defense and bacterial clearance. In our study of Cdh5-CreER; Tlr4CKO mice, bacterial burden and disease severity at 24 hours post-infection are similar between WT and Cdh5-CreER; Tlr4CKO mice, indicating that Cdh5-CreER; Tlr4CKO does not alter the clinical course at this time point. This difference is noted in the Discussion section.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the molecular underpinnings of immune responses in the leptomeninges in neonatal bacterial meningitis. Bacterial meningitis is a major disease burden, particularly for neonates, and it has previously been noted that the meningeal immune environment in infants is permissive to opportunistic infection (Kim et al., Sci Immunol, 2023). There is less known about the contribution of the stromal compartment to meningeal immune responses. Seegren et al. interrogate the role of leptomeningeal endothelium in host defence in E. coli infected neonatal mice using mouse genetic tools to delete the LPS receptor Tlr4 from either endothelial cells (using Cdh5-CreER) or macrophages (using LysM-Cre). The authors use snRNAseq, cleared cortical mounts, and in vitro work to define the impact of E. coli infection on leptomeningeal endothelial cells. This study uses a range of innovative techniques to probe the role of the stromal compartment in meningitis.

      Strengths:

      This study makes excellent use of cleared cortical mounts to examine the biology of the leptomeninges, in particular, changes to the endothelium, with unprecedented detail. In combination with high-quality sequencing data provide new insights into the impact of meningitis on the leptomeninges. The data presented by the authors is of very high quality.

      Weaknesses:

      The weaknesses of the study were in terms of interpretation and perhaps study design.

      (1) Most importantly, the authors need to provide additional validation of their conditional knockout models. The authors need to confirm that the Cdh5-CreER does not impact leptomeningeal fibroblasts and to confirm gene deletion in macrophages.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclearlocalized GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite correct. The Results section of the revised manuscript includes an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      (2) The authors could also strengthen the paper by providing data on the impact of these conditional knockout models on the course of meningitis and bacterial burden.

      Thank you for this comment. We agree that these additional analyses strengthen the manuscript. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences in bacterial burden between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) Finally, it is perhaps not surprising that Tlr4 is required for meningitis responses with E. coli. However, it is unclear if these findings can be generalised to other, more common, meningitis infections (streptococcal/pneumococcal).

      At present, it is an open question whether TLR4 plays as a large a role in meningitis caused by other gram-negative bacteria and whether TLR2 plays a similarly large role in meningitis caused by gram-positive bacteria. In the Discussion, the last two sentences under “Limitations of the study” summarize this point: “Finally, the present study focused on E. coli K1, the dominant Gram-negative neonatal pathogen. Future work could assess TLR signaling in response to other bacterial pathogens, such as Group B Streptococcus.”

      (4) There are additional minor issues; for instance, the arachnoid fibroblast 2 population appears to closely resemble dural border cells.

      Thank you for this comment. That is correct, and we have changed the nomenclature to “dural border cells”.

      (5) The cell line model (bEnd.3) is a relatively low-fidelity model of BBB endothelial cells, and this should be acknowledged.

      Thank you for this comment. That is correct. Despite being brain-derived, bEnd.3 cells have lost many BBB-specific attributes. Their responses might best be considered as generic endothelial responses rather than brain-specific endothelial responses. This is now stated in the Results section: “Although they are brain-derived, bEnd.3 cells lack many BBB-specific attributes and, therefore, they likely exhibit generalized endothelial responses rather than brain-specific responses to bacterial exposure.”

      With these caveats, it is difficult to be certain that the endothelium alone is the driver of meningeal immune responses in meningitis, and what the impact of these is.

      We agree with this critique. As noted above, the expression of Cdh5-CreER in essentially all endothelial cells and in a subset of dural border cells and leptomeningeal fibroblasts means that the comparison of TLR4 CKO with Cdh5-CreER vs. Lyz2-Cre is assessing phenotypes driven by TLR4 signaling in endothelial plus a subset of other non-myeloid cells vs. TLR4 signaling in myeloid cells. We have revised the text to reflect this more precise understanding of Cdh5-CreER specificity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Transcriptomic analysis: The analysis and display of the single-nucleus RNA-seq data should be improved. The authors could perform a more granular, unbiased clustering of each cell class in the combined dataset and then compare the proportion of each experimental group (genotype x control/infected) in each cluster. At present, it appears the differentially-expressed genes (DEGs) shown in Figure 1 were identified using the Seurat FindMarkers function with default parameters (Methods). This considers each cell as an independent experimental unit and is therefore not appropriate for a comparison of control versus infected groups (see e.g., PMID 34584091, 35880426. The authors should implement a statistical analysis strategy that considers true biological replicates (mice, as shown in Supplementary File 1).

      We do not fully agree with this critique. We agree that biological replication at the level of individual mice is important for interpreting these data, but within each mouse, the characteristics of individual cells is also of interest, including the degree of heterogeneity, the sample size for a given cell cluster, and the statistical significance of any observed changes in transcript abundance. As requested, we have prepared a new supplemental figure (Figure 1 – figure supplement 5) showing a principal component analysis of the scRNA-seq data for each mouse (one mouse was used for each snRNA-seq dataset) and for each of the principal leptomeningeal cell types. This analysis shows, for example, that the three infected Cdh5-Cre; Tlr4flox/- mice have transcriptomes for each of the six cell clusters that are very similar to the transcriptomes of the two uninfected WT and the two uninfected Cdh5-CreER; Tlr4flox/- mice. Thus, the genotype- and condition-dependent effects are consistent across biological replicates. At the most granular level, Figure 1 – figure supplement 7, which was part of the original submission, shows for the most up- and down-regulated genes (based on adjusted p-value or based on fold-change) in endothelial cells and in myeloid cells how individual transcript abundances change for each mouse and for each genotype/condition.

      (2) The authors use immunohistochemistry to assess claudin-5 "disorganization and redistribution" (Results, pages 7-8 and Figures 3A-B). They state that "Tlr4ECKO mice showed minimal changes in the distribution of Cldn5, implying that cell autonomous endothelial TLR4 signaling regulates tight-junction organization." It is not clear, however, that the quantified parameter (Cldn5+ area relative to total area) would be an accurate readout of claudin-5 organization/distribution (i.e., subcellular localization) as it would also be sensitive to claudin-5 expression, vascular density, and vessel diameter. The authors use a similar assessment of ZO-1 to suggest that changes to claudin-5 are not due to a "generalized disassembly of TJs" and could also use this to argue that the above potential confounds (vascular density, vessel diameter) do not change, but the data in Figure 3 - Figure Supplement 1B, lower panel, show that infection does cause an increase in ZO-1 area relative to total area (P = 0.0004). Thus, the statement in the results "Zonula Occludens-1 (ZO-1) [...] remained unchanged during infection (Figure 3 - figure supplement 1)" is not accurate. The authors should revise this section to ensure their conclusions are aligned with the presented data.

      Thank you for this comment. The reviewer is correct that the Cldn5 area measurement is unable to deconvolve the various factors that might contribute to it (vessel density and diameter, and Cldn5 distribution). This part has been rewritten. “Consistent with prior findings (Wang et al., 2023), both WT and Tlr4<sup>MKO</sup> mice showed an increase in the area occupied by Cldn5 in the leptomeninges following infection, likely referable to both increased vessel diameter and a redistribution of Cldn5 within ECs (Figure 3A-B; Figure 3 – figure supplement 1C).”

      The reviewer is also correct about our initial description of the ZO-1 data. What we meant to write and what the revised manuscript now shows is: “The area occupied by Zonula Occludens-1 (ZO-1), a tight junction scaffold protein, showed a modest but statistically significant increase in WT leptomeningeal vessels but no significant change in Tlr4<sup>VEKO</sup> leptomeningeal vessels during infection (Figure 3 – figure supplement 1A and B).”

      (3) In the Methods, under Mouse Models and E. coli Infection, the authors state, "The Cdh5-CreER line (Monvoisin et al., 2023) was the same line used in Wang et al. (2023)." However, there is no mention of Cdh5-CreER in Wang et al. (2023). Could authors please clarify? Also, because this appears to be an inducible Cre, the authors must include details on the dose and timing of tamoxifen or 4-OHT used in this study.

      Thank you for catching that error. We meant to reference Wang et al (2025), not Wang et al (2023). [Wang et al (2025) is: Wang Y, Rattner A, Li Z, Smallwood PM, Nathans J. (2025) Vascular endothelial-specific loss of TGF-beta signaling as a model for choroidal neovascularization and central nervous system vascular inflammation. Elife 14:RP107018.] This has now been corrected.

      We have now included the details related to 4HT injection in the Methods section “Mouse Models and E. coli Infection”. These are intraperitoneal injection at P2 with 40- 50 µL of 2 mg/ml 4HT.

      (4) The legend for Figure 1A is "Schematic of the leptomeninges", but the figure shows the entire brain-skull interface, including underlying cortex, leptomeninges, dura, and skull.

      Thank you. Corrected.

      (5) Page 7, typo: "In the brain, CD206+ cell were too sparse ..." Should be "cells".

      Thank you. Corrected.

      (6) Page 14, typo: "... could represents a double-edged ..." Should be "represent".

      Thank you. Corrected.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform CFU counts from LPM, dura, brain, peripheral organs (liver) in infected v mock mice from control v TLR4 EC-cKO.

      Thank you for this comment, with which we agree. We have addressed this by quantifying bacterial burden and assessing disease severity in WT and Cdh5CreER; Tlr4floxed mice. Specifically, we performed CFU measurements in blood, monitored mouse weights at P5 and P6, and histologically surveyed the E. coli-RFP signal (i.e., E. coli burden) in brain, liver, and lung. These analyses show that bacterial burden and disease progression are comparable between WT and Cdh5-CreER; Tlr4floxed mice at 24 hours post-infection. These data are presented in Figure 2 – figure supplement 4 and described in the Results.

      (2) Lyz2Cre/+ is used to delete TLR4 from macrophages, but recombination efficiency (in LPM BAMs) is described as only partial, suggesting that TLR4-response in LPM BAMs (and potentially macrophages in the dura) is at least partially intact. It undercuts conclusions that can be made using this line.

      Thank you for this comment. We have conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The Results section text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Also, as noted in the reply to comment 4 below, a direct analysis of Tlr4 recombination efficiency is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (3) Inflammatory responses [qPCR] from peripheral organs and also physiological measures in the pups [weight post-infection, time to moribund or death curves] in control v TLR4 EC-cKO and TLR4 mac-cKO.

      Thank you for this comment. We have not conducted a qPCR analysis of inflammatory gene expression in peripheral organs because (1) the dramatic upregulation of these transcripts in the leptomeninges, (2) the presence of E. coli in blood and peripheral organs, and (3) the clinical assessment (cessation of weight gain) all predict that such an analysis would reveal a large up-regulation of inflammatory gene expression throughout the body. More specifically, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (4) The conditional macrophage line is problematic due to the partial recombination. I question the utility of including this unless they can come up with a way resolve the response of recombined TLR4 macrophages vs ones that are not (could they use the single cell data to pick this a part? Are TLR4-null cells and TLR4 'wt' cells transcriptionally similar in the infected condition, suggesting TLR4 is not doing much in the macs, potentially due to alternate TLRs?). There are good BAM Cre lines that have been described [Lyve1-cre would be good for LPM BAMS, the other is Pf4-cre, see https://pmc.ncbi.nlm.nih.gov/articles/PMC7375817/ - just as an FYI for the future].

      Thank you for this comment. As noted in the reply to point 2 (above), we have conducted a more in-depth analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      We agree that, based on Figure 6 in the cited paper [McKinsey et al (2020) A new genetic strategy for targeting microglia in development and disease eLife 9:e54590], the Pf4-Cre line may be superior to the Lyz2<sup>Cre</sup> line that we used for recombination in leptomeningeal myeloid cells. Unfortunately, we missed this paper in our literature searches, probably because it focuses on a microglial CreER line, P2ry12-CreER, and the Pf4-Cre line is not mentioned in the title or abstract. Our decision to use the Lyz2<sup>Cre</sup> line was based on an extensive comparison among myeloid Cre lines showing that Lyz2<sup>Cre</sup> was the most efficient [Abram CL, Roberge GL, Hu Y, Lowell CA. 2014. Comparative analysis of the efficiency and specificity of myeloid-Cre deleting strains using ROSA-EYFP reporter mice. J Immunol Methods 408:89-100.] However, the Abram et al study did not look at the leptomeninges. Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (5) Figure 1 - Figure Supplement 2 - the authors nicely break down the pathway response [NFKB and TNF] in EC and macs, it would be great to have similar information for the fibroblasts (in the main figure or the supplement). Does their inflammatory response show a similar pattern?

      Thank you for this suggestion. We have now done that analysis and present it in Figure 1 – figure supplement 3. For completeness, we also performed the same type of analyses for JAK-STAT signaling and IFN-gamma response and these are shown in Figure 1 – figure supplement 6. The principal conclusion is that across all major leptomeningeal cell types, the Cdh5-CreER; Tlr4floxed samples (i.e., Tlr4 KO’d in non-myeloid cells) show much reduced transcriptome changes with infection.

      (6) What is ASC and Cd206 quantification measuring, and how does this relate to 'activation' - is this the intensity of signal or a morphological change? What is the precedence for using ASC (citations)? In their prior work, they showed no change in CD206 number, so a significant increase upon infection here, it's confusing exactly what is being studied. Also, loss of Lyve1 is a well-accepted measure of activation that they have previously used, adding that it could be helpful. This is not a major issue since they have robust data that the macrophages are not transcriptionally activated. Clarification of what exactly is being measured would be sufficient (in the text).

      CD206 immunostaining, which reveals myeloid cell morphology, shows that, with E. coli infection, myeloid cells convert from a more compact morphology to a more expanded morphology. This is now explained more fully in the Results section.

      Regarding ASC, changes in the state of ASC aggregation and ASC subcellular localization have been used by others to monitor immune cell responses to inflammatory signals (Sester et al., 2016; Franklin et al., 2018). While this change in subcellular localization may explain part of the increase in immunostained area in myeloid cells in the infected mice (Figure 2D), the increase in the area of ASC immunostaining largely reflects a shift of myeloid cells from a compact to a more extended morphology. This is now explained more fully in the Results section. We have also added two references (Sester et al., 2016; Franklin et al., 2018) that described how ASC distribution changes with inflammation.

      Regarding LYVE1, we observe a decrease in LYVE1 transcript abundance in myeloid cells with infection, as predicted. Given the large amount of other data that document myeloid activation with infection, we have elected not to include this.

      (7) The authors suggest the internalization of Cldn5 is not due to NFKB downstream signaling that includes transcriptional mechanisms because it happens as early as 1 hour, prior to NFKB localization to the nucleus. However, a lot of their experiments, including on endosomal-lysosomal protein co-localization are done at 4 hours, when their RNAseq data show robust NFKB-mediated gene upregulation and (though not tested) potentially protein production of factors that can act back on the cells, including to impact endo-lysosomal processing. Without studies at earlier timepoints post-bacteria exposure, separating these two mechanisms is difficult.

      Thank you for this comment. We have explored this question by looking at Cldn5 internalization in bEnd.3 cells at 1 hour after E. coli exposure, and the data clearly show that internalization occurs within 1 hour. Additionally, we have conducted this experiment in the presence of 1 uM ACHP, an IKK inhibitor that blocks NF-кB migration to the nucleus. ACHP treatment shows no effect on the rapid internalization of Cldn5, implying a mechanism independent of NF-кB control of gene expression. These data are shown in a new figure (Figure 6) in the revised manuscript.

      (8) Figure 2 - CD206 are quite sparse however, Iba1 would work well to look at microglial activation.

      Thank you for this suggestion, which we have followed. To assess microglial activation, we have immunostained for Iba1 and quantified the data. These are now included in Figure 2 – figure supplement 3. The data show that there is an increase in Iba1 immunostaining following E. coli infection in both WT and Cdh5-CreER; Tlr4floxed mice, with more in the former than the latter, but the difference is not statistically significant.

      (9) Suggest performing the LAMP+ co-localization experiment at <1hr, prior to NFKB nuclear localization and transcriptional changes. This would better support it, this is (or is not) independent of the NFKB. Could also test this with an NFKB inhibitor, do they still see the CLDN5 internalization when NFKB is blocked?

      Thank you for these suggestions. We have done both of these analyses, and the results are presented in Figure 6. The results show that (1) Cldn5 is internalized within 1 hour and (2) its internalization is independent of NF-кB signaling inhibition by 1 uM ACHP. Since ACHP treatment shows no effect on the rapid internalization of Cldn5, that implies a mechanism independent of NF-кB control for gene expression.

      Reviewer #3 (Recommendations for the authors):

      Major points

      (1) The most important caveat is that the Cdh5-CreER model is known to recombine in leptomeningeal fibroblasts (10.1038/s41586-023-06993-7, 10.1101/2025.05.13.653681), and Cdh5 expression in these populations is now well described (10.1038/s41467-02341580-4, 10.1016/j.neuron.2023.09.002). Although the authors did not observe recombination in their reporter (details of the tamoxifen injection protocol should be provided), it is imperative to validate the specificity of their model to Tlr4 in endothelial cells, leveraging their sequencing data and providing additional IHC or ISH to confirm this. Alternatively, Tlr4 could be deleted in a more specific model, e.g., the Pdgfb-iCreERT2 or Slco1c1-CreERT2. It is also important to do the same with the LysM model, to confirm that the lack of impact of macrophage Tlr4 is not due to failure to delete the gene. This is again important to the interpretation of the study, since the authors propose that the endothelium, specifically, is the driver of the meningitis response.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclear-localised GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite right. The Results section of the revised manuscript has an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      We have also conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup>-directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text in the Results section now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. The phenotype of Cdh5-CreER; Tlr4floxed mice – a dramatically reduced infection-associated transcriptional response – argues that the floxed Tlr4 target was recombined at appreciable efficiency in those mice (Figure 1D and 1E). For Lyz2<sup>Cre</sup>; Tlr4floxed mice the principal phenotype is an up-regulation of infection-associated transcripts in a subset of dural border cells in the absence of infection; the transcriptional response to infection was largely unaffected in all leptomeningeal cell types (Figure 1D and 1E). We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4VEKO and Tlr4MKO mice.”

      (2) The authors did not examine the consequences of Tlr4 cKO on the course of meningitis or bacterial burden. Knowing the impact of this would strengthen the paper and allow us to determine if the endothelial responses are helpful or harmful in meningitis progression.

      For the revised manuscript, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) TLR4 is a known receptor for LPS. It is unsurprising (especially in the in vitro experiments) that Tlr4 knockout reduces NF-kB signalling and other downstream changes to endothelial cells. Furthermore, it is uncertain if the infection was left to continue, similar changes to the endothelium would nonetheless occur through other mediators such as IL1B and TNFa.

      We agree that it makes logical sense that Tlr4 KO decreases NF-кB signaling. The interesting next question is: what are the mechanistic underpinnings of the responses that are downstream of TLR4 and NF-кB? The cell culture experiments with WT vs. Tlr4KO bEnd.3 cells identify one set of cell biological responses related to Cldn5 and junctional integrity, and the NF-кB inhibition experiment (Figure 6) implies that rapid internalization of Cldn5 occurs in the absence of NF-кB mediated transcriptional changes. Regarding the possibility that other mediators such as IL1B or TNFα might, at least partially, make up for the lack of TLR4 signaling later in the infection, that is an open question at present.

      (3) The arachnoid fibroblast 2 cluster should be renamed to dural border cells based on their high expression of Slc4a10, Adamtsl3, Tmeff2, etc which are all highly enriched in dural border cells. I suspect this cluster is also highly enriched for Slc47a1, probably the most specific marker for these cells (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002).

      Thank you for this comment. The reviewer is correct. These are dural border cells and they express Slc47a1, as seen in a new supplemental Figure 1 – figure supplement 4, which shows UMAP plots for many leptomeningeal cell type-specific genes. We have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (4) It would be helpful to provide higher resolution images of Cldn5 in the leptomeningeal mounts. At the current resolution, it is difficult to tell if there is a similar internalisation/disruption phenotype to what is observed in vitro. Notably, this finding is similar to another recent publication on Cldn5 recycling (in the context of stroke) (10.1186/s40478-025-02125-6).

      Higher resolution images of Cldn5 in leptomeningeal vessels without or with E. coli infection are now shown in Figure 3 - figure supplement 1C. There is a visual impression of greater area occupied by Cldn5, which is confirmed by quantification (Figure 3A and B). This effect appears to be due to both an average increase in vessel diameter and a redistribution of some of the Cldn5 away from plasma membrane junctions. Thank you for pointing out the interesting and relevant Cottarelli et al (2025) paper, which we had not read. This is now referenced.

      Minor points

      (1) Typo: prominant should be spelled prominent.

      Thank you for catching that one. It is now corrected.

      (2) Strictly speaking, the arachnoid layer is not epithelial (despite Cdh1 expression). They are fibroblasts that acquire barrier-forming properties.

      Thank you for that comment. That appears to be the consensus view, and we will go along with it.

      (3) Notably, LyzM Cre will also recombine in other myeloid populations, so I wouldn't describe it as a macrophage.

      Thank you for this comment. We agree, and we have therefore changed the text and figure labels from “macrophage” to “myeloid”.

      (4) It is interesting and notable that ICAM1 expression is observed in nonendothelial populations, in the IHC, too, perhaps.

      We agree. ICAM1 may be a broader marker/mediator of inflammation than is generally recognized.

      (5) In F1B, your labels on the right image to the arachnoid barrier and pial surface are presumably meant to refer to the image on the left with DPP4 and laminin labelling? The subarachnoid should be between the laminin and DPP4 layers (although it will be collapsed in your preparations).

      Thank you for catching this error. The vertical bars were sized erroneously, and the labels were also placed erroneously. These have now been corrected.

      (6) I would reference the papers that defined leptomeningeal cell type markers (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002) when you define your cell types.

      Thank you. We have done that, and we have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (7) I would change references to the subarachnoid space in your figures to the leptomeninges (which include the SAS, but extend either side of it).

      Thank you. The labels have been changed to “leptomeninges”.

      (8) In Figure 2 - Supplement 1A, it looks like the populations are mislabelled.

      Thank you. This has been corrected to be consistent with the assignments in Figure 1B

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Zonation definition under injury has been shown to be sustained broadly, but is not sufficiently validated and quantified, especially considering the resolution of the 10x Visium system and the potential variation of outcomes based on how to define zones.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g.,Alb, Mup20, Cyp2f2,Pck1,Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2,Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g.,Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) The model is built entirely in APAP injury, which specifically targets pericentral hepatocytes. It remains unclear whether the proposed mechanism applies to other liver injuries (e.g., partial hepatectomy, CCl4).

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Baseline proliferation appears higher than expected in homeostasis (Figure 1B), and fold change analysis (not absolute counts) may be needed to assess zonal proliferation suppression (Figure 1D).

      We thank the reviewer for this insightful comment. The baseline proliferation rates observed in our study are consistent with previously reported zonal distributions (PMID: 33632817; PMID: 33632818), with approximately 70% of proliferating hepatocytes located in zone 2, 20% in zone 3, and 10% in zone 1 under homeostatic conditions. To further address the reviewer’s concern, we performed a fold-change analysis of Ki-67<sup>+</sup>hepatocytes across different zones. This analysis revealed that only the mid (zone 2) and pericentral regions exhibited significant changes, whereas no statistically significant differences were observed in the other zones (as shown in Author response image 1). Importantly, when considered together with the absolute cell counts, these results indicate that the apparent suppression of proliferation is most pronounced in the mid zone, likely due to its relatively higher baseline proliferation under homeostatic conditions. In contrast, this effect is less evident in the fold-change analysis, as zones with low baseline proliferation show limited dynamic range for detecting relative changes.

      Author response image 1.

      Fold changes of Ki67-positive cells across liver zones (PC, Mid, PP) at 0, 3, 6, 12 and 24 h post-APAP. Fold change was the number of Ki67-positive cells in the three regions at each time point after APAP treatment divided by the number of positive cells in each region at 0 hour post-APAP. (a) denotes significance between PC and Mid regions, (b) denotes significance between PC and PP regions, and (c) denotes significance between Mid and PP regions.

      (4) AAV-based overexpression raises potential confounds (altered CYP activity before injury) and shows incomplete penetrance that is not quantified (Figure 5 - Figure 6).

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (5) The functional link between proliferation suppression and improved survival is inferred, but direct survival /injury readouts are limited.

      We thank the reviewer for this insightful comment. To more directly evaluate the functional link between proliferation control and liver injury, we manipulated Btg2, a downstream effector of the Atf4–Chop axis and a known inhibitor of cell proliferation. Knockdown of Btg2 using AAV8–CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial knockdown efficiency, we observed a clear exacerbation of liver injury, as evidenced by an approximately 2-fold increase in serum ALT levels and a ~1.5-fold expansion of necrotic areas. In parallel, hepatocyte proliferation was significantly increased (~1.8-fold increase in Ki67⁺ hepatocytes) compared to control mice (Revised Figure 6K–N). Conversely, Btg2 overexpression produced the opposite phenotype, markedly attenuating liver injury while suppressing hepatocyte proliferation (Revised Figure 6G–J). Together, these gain- and loss-of-function data provide direct evidence linking proliferation control to injury severity, thereby supporting a causal relationship between suppressed proliferation and improved liver outcomes (Revised manuscript, page 14, lines 399–406).

      Reviewer #2 (Public Review):

      (1) Starting with the basics, one wonders why midlobular hepatocytes manage to mount a defensive response to APAP but pericentral hepatocytes don't. Is this because midlobular hepatocytes express the relevant Cyps (2e1, but also 1a2 and 3a11) at lower levels, which mitigates toxicity and buys them time? This would be supported by F2A but not by F3B, at least not for the most important Cyp2e1. A moderate difference is shown for Cyp1a2 expression in F3D, but is that enough to explain the different fates? Or are additional post-transcriptional effects on these Cyps at work?

      We thank the reviewer for this important question. We fully agree that the differential susceptibility between mid‑zone and pericentral (PC) hepatocytes is likely rooted in the zonal gradient of cytochrome P450 expression. Our spatial transcriptomics data (Revised Figure 2A) show that mid‑zone hepatocytes express Cyp2e1, Cyp1a2, and Cyp3a11 at levels intermediate between PC and periportal (PP) zones. This intermediate expression may generate sufficient NAPQI to activate stress signaling but not so much as to cause immediate mitochondrial collapse, thus “buying time” for adaptive responses. We also appreciate the reviewer’s observation that Cyp2e1 mRNA levels remain highest in the PC zone even after APAP (Revised Figure 3B). However, mRNA abundance does not necessarily reflect functional protein level. In the PC zone, massive necrosis rapidly compromises cellular integrity; as shown in Revised Figure 3D, Cyp1a2 protein declines sharply around the central vein, and we observed similar degradation for Cyp2e1 (data not shown). Consequently, despite sustained Cyp2e1 transcripts, PC hepatocytes are unable to mount an effective stress response because they are already undergoing cell death. By contrast, mid‑zone hepatocytes retain sufficient metabolic capacity to activate the Atf4‑Chop axis while preserving cellular function.

      (2) The evidence presented in support of cell cycle arrest of midlobular hepatocytes is not fully convincing: there is no overt difference in S and G2/M gene scores in F2F; the marker genes used for S phase and G1 to S progression in F2G are unusual. Along these lines, one wonders if spatial transcriptomics confirmed the Ki67 immunostaining results in F1 also for specific zones, not only overall, as shown in F2E?

      We thank the reviewer for these important observations. We agree that the current spatial transcriptomics (ST) data alone do not provide sufficiently strong support for this conclusion. The limited sensitivity of ST for detecting rare proliferative events further constrains its utility in this context. At baseline, only ~1% of ST spots are Ki67-positive (Revised Figure S1I), and this fraction becomes even lower during the early phase following APAP injury. As a result, there are insufficient Ki67+ spots to robustly assess zonal distribution using ST, which precludes a reliable spatial validation of proliferation patterns at this resolution. For this reason, our primary evidence for zonal proliferation dynamics relies on Ki67 immunohistochemistry (Revised Figure 1), which provides single-cell resolution and higher sensitivity. These data show a marked reduction in Ki67+ hepatocytes specifically in the midlobular zone at 3-6 hours post-APAP, supporting a transient suppression of proliferation in this region. In addition, we agree that the transcriptional evidence for cell cycle arrest was not strong the S and G2/M scores showed no overt difference, and the gene sets used were suboptimal. We have therefore moved these analyses to the supplement and toned down the claims. We have also clarified this limitation in the manuscript (Revised manuscript, page 18, line 518-524)

      (3) The authors conclude in line 364 that halting of proliferation by Btg2 favors survival, which raises the question of whether Btg2 knockout causes death in midlobular hepatocytes in F6K. Data addressing this question, that is, the localization and extent of tissue necrosis and ALT levels after APAP, are missing. The efficiency of the knockout of Btg2 is also not given.

      We thank the reviewer for this insightful comment. We have included the missing data. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup>hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N, Revised manuscript, page 14, line 399-406).

      (4) Related to the previous question, the BTG2 immunostaining in F6F is not convincing when compared to F6D. One also wonders if it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3?

      We thank the reviewer for this insightful comment. We have included an inset of the original image to better show BTG2 staining in revised Figure 6F. During our study, we tested BTG2 expression in mice transduced with AAV‑TBG‑EGFP or AAV‑TBG‑BTG2 for three weeks without APAP challenge. We observed that BTG2 in these non‑injured livers was predominantly cytoplasmic (Author response image 2), contrasting with the nuclear localization seen after APAP treatment (Figure 6F). Regarding whether it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3, we think Ddit3 promotes BTG2 expression (as shown in revised Figure F6F), but APAP is necessary for its nuclear translocation.

      Author response image 2.

      Immunohistochemical detection of Btg2 in liver tissue from mice transduced with AAV-TBG-EGFP or AAV-TBG-Btg2 for 3 weeks without APAP treatment.

      (5) Related to the previous question, the proposed Atf4-Ddit3 axis is challenged by the lack of midlobular induction of Atf4 in the APAP scRNA-seq data published by another group, presented in S4F and G. Further analysis of AAV-Atf4 samples generated for F5 could address whether it is really Atf4 that acts on Ddit3 in APAP toxicity.

      We thank the reviewer for this insightful comment. We agree that Atf4 was not among the top 30 active transcription factors in our initial analysis; however, when we extended the list to the top 50, Atf4 was included. We have therefore updated Revised Figures S4F and G to show the top 50 transcription factors. We also appreciate the reviewer’s suggestion to further investigate whether Atf4 directly acts on Ddit3 in the context of APAP toxicity. While this still shows a less pronounced midlobular enrichment for Atf4 compared with Ddit3, we sought additional evidence for a functional Atf4-Ddit3 link. In primary hepatocytes treated with APAP, we observed nuclear co‑localization of Atf4 and Ddit3 (Author response image 3A) and increased nuclear protein levels of both factors (Author response image 3B), supporting their potential cooperative role. We agree that direct analysis of AAV‑Atf4 samples generated for Figure 5 would provide more definitive evidence; unfortunately, co‑staining for Atf4 and Ddit3 on those tissue sections didn’t work well.

      Author response image 3.

      Subcellular localization of Atf4 and Chop in primary hepatocytes following APAP treatment. (A) Immunofluorescence staining of Atf4 and Chop in primary hepatocytes treated with 10 mM APAP for 6 hours or left untreated (UT). Nuclei were counterstained with DAPI. Scale bar as indicated. (B) Primary hepatocytes were treated with 0, 5, or 10 mM APAP for 6 hours. Cytoplasmic and nuclear fractions were isolated and analyzed by western blot. Lamin B1 and α-Tubulin were used as markers for the nucleus and cytoplasm, respectively

      (6) Related to the previous question, the ATF4 immunostaining in F5A doesn't look convincing, with many brown pigments appearing to be outside of the nucleus.

      We thank the reviewer for this helpful comment. To better demonstrate ATF4 nuclear localization, we have added enlarged insets of the original representative images in revised Figure 5A. These magnified views more clearly show nuclear ATF4 staining after APAP treatment, addressing the concern about extranuclear signal.

      (7) It is not ruled out that AAV expression of Atf4 or Btg2 reduces hepatocyte sensitivity to APAP by affecting the expression of the Cyps needed for activation. In other words, does AAV-Atf4 or AAV-Btg2 change the expression of any of the Cyps relevant to APAP in the 3 weeks before APAP application (F5B)?

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      (8) It is laudable that the authors tried to extend their findings to humans by using snRNA-seq data from a published study (line 391), but it is unclear why they didn't analyze all 10 patients in that study but instead focused on 2 and stated that this small sample number prevented drawing definitive conclusions and could therefore only be mentioned in the discussion.

      We thank the reviewer for this clarification. The analysis mentioned in line 391 originally referred to spatial transcriptomics (ST) data from two ALF patients, not snRNA-seq. For the snRNA-seq dataset, we analyzed all 10 patients, but snRNA-seq lacks spatial resolution and cannot reliably assign zonal identity. We stipulate that snRNA-seq requires viable cells and thus likely excludes necrotic/peri-necrotic areas. Therefore, direct zonal comparison with our ST data was not possible. We have now clarified this in the revised manuscript (Revised manuscript, page 18, line 510-519).

      Reviewer #3 (Public Review):

      The main concern is that the overexpression of ATF4 and DDIT3 is causing reduced cell death and damage by APAP. This makes it harder to understand if these genes are truly increasing survival or if they are just reducing the injury caused by APAP. It may be better to perform overexpression immediately after, or at the same time as APAP delivery. Alternatively, loss-of-function experiments using AAV-shRNAs against these targets could be useful.

      We thank the reviewer for raising this important point. We agree that overexpression prior to APAP administration leaves open the question of whether the observed protection reflects true cytoprotection or simply reduced initiation of injury. To address this, we pursued loss‑of‑function approaches. Due to their very low basal expression, AAV‑shRNA‑mediated knockdown of endogenous Atf4 and Ddit3 proved inefficient. We therefore targeted Btg2, a downstream mediator of Ddit3 that inhibits proliferation. Knockdown of Btg2 resulted in a significant increase in APAP‑induced liver injury, as evidenced by elevated ALT levels and expanded necrotic areas (Revised Figure 6K-N). These results indicate that the ATF4‑DDIT3‑BTG2 axis limits hepatocellular damage, consistent with a protective role. We have clarified this point in the revised manuscript (page 15, line 407-414)

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Clarify how zones were defined when necrosis disrupted pericentral areas. Provide marker validation across time and whether necrotic spots are excluded or not from zonal analysis.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g., Alb, Mup20, Cyp2f2, Pck1, Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2, Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g., Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) Test whether the ISR-Btg2 program applies in other models; even targeted validation via qPCR and IF would be valuable.

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Proliferation quantification in liver sections in Figure 1: how to define the zones and why, at the basal level, there is a high proliferation rate in the mid zone? From Figure 1B-C, all three zones showed decreased hepatocyte proliferation, although the mid zone had a higher baseline. Will the mid-zone stand out by converting to the fold change of Ki-67+ hepatocytes decrease?

      We thank the reviewer for these insightful comments. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x<sup>2</sup> + z<sup>2</sup> – y<sup>2</sup>) / (2z<sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885). Consistent with previous reports (PMID: 33632817; PMID: 33632818), we observed a higher baseline proliferation rate in the mid zone, where approximately 70% of proliferating hepatocytes reside under basal conditions, compared to 10% in zone 1 and 20% in zone 3. However, when analyzing the fold change in Ki-67+ hepatocytes, only Mid and PC region showed significant difference in Ki-67+ hepatocytes, other zones showed no significant differences (as shown in the fold-change results in Author response image 1), indicating that the mid zone does not stand out in the fold change analysis. See Author response image 1.

      (4) The authors need to strengthen the causal chain with rescue experiments, e.g., Atf4/Chop overexpression and Btg2 knockdown. Link proliferation suppression to survival/ALT directly.

      We thank the reviewer for these constructive comments. Besides existing data from Figure 5 (Atf4 overexpression), we included Btg2 knockdown data in the revised Figure. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup> hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N) (Revised manuscript, page 14, lines 399–406).

      (5) Transduction efficiency, distribution, and expression levels via the AAV overexpression need to be quantified. Key CYP genes in the APAP metabolic pathway need to be assessed to exclude confounds.

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (6) The authors claim that the requirement of the Atf4/Chop at the early stage of APAP injury protects hepatocytes from proliferation for survival. What is the consequence if we remove the protective mechanism?

      We thank the reviewer for this insightful question. In our model, early induction of Atf4 and Chop functions as a cell survival checkpoint. Removal of this protective mechanism is predicted to result in two deleterious outcomes: (1) Acute exacerbation of necrosis due to the inability of hepatocytes to manage stress-induced bioenergetic demands, and (2) Impaired long-term regeneration due to depletion of the surviving cell pool. We directly tested the acute prediction (< 24 h) in Author response image 4. We deleted Ddit3 specifically in hepatocytes. Initial attempts using AAV-CasRx failed due to negligible baseline Atf4/Chop expression in healthy liver, preventing effective knockdown. We therefore generated hepatocyte-specific Ddit3 knockout mice (Alb<sup>∆Ddit3</sup>; Author response image 4B). Immunohistochemistry confirmed APAP-induced Chop induction occurs primarily in the centrilobular zone by 6 h (Author response image 4A). Following a two-dose APAP regimen (Author response image 4C), Alb<sup>∆Ddit3</sup> mice displayed significantly larger areas of centrilobular necrosis compared to Ddit3<sup>fl/fl</sup> controls (Author response image 4D; **p < 0.01). Thus, hepatocyte-intrinsic Chop limits acute APAP injury, consistent with its proposed early protective role.

      Author response image 4.

      Hepatocyte-specific deletion of Ddit3 exacerbates APAP-induced liver injury. (A) Immunohistochemical staining of Chop in liver sections at 0,3 and 6 h post-APAP. Red arrows indicate Chop-positive hepatocytes. Scale bar = 50μm. Quantification of zonal distribution of Chop-positive cells in liver sections at 6 h post-APAP is conducted . The statistic is the percentage of Chop-positive hepatocytes in each layer over the total number of Chop-positive hepatocytes. n=3 mice. (B)The construction, genotyping strategy and genotyping results of Alb<sup>∆Ddit3</sup> mice. P: positive control; WT: Wild-type; Neg: Blank control(ddH<sub>2</sub>O). (C) Schematic figure illustrating the experimental strategy for the administration of two doses of APAP to Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice. (D) H&E staining showing liver morphology from Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice at 6 h post-second dose of APAP. Injured area is outlined by black dashed lines. Scale bars = 200 μm. The percentage of injury area is quantified. n = 3- 4 mice/group. Data are represented as means ± SD; *p < 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001; ns, not significant.

      (7) Is there any human relevance to the sensitivity of APAP injury regarding the Atf4/Chop axis?

      We thank the reviewer for this insightful comment. During our study, we analyzed a spatial transcriptomics dataset from APAP patients. In one of two analyzed patients, mid-zone hepatocytes exhibited transcriptional signatures remarkably consistent with our murine findings, including: (1) upregulation of Atf4-Chop pathways, and (2) downregulation of cell proliferation genes (Author response image 5). This suggests that this axis may also be involved in the response to APAP injury in humans. However, given the limited sample size, definitive conclusions cannot be drawn at this stage. We have now included this point in the Discussion section (Revised manuscript, page 18, line 510-519).

      Author response image 5.

      Spatial transcriptomics (GSE223561) reveals zonal gene expression changes in APAP patients. Heatmap of ISR, cell death, and cell cycle gene expression across zonal regions in healthy versus APAP‑treated human livers. 

      (8) Several IHC stainings have a weak signal and need inserts to zoom in for a clear view of the positive signals. Figure 5A, E, G, and Figure 6D, F.

      We thank the reviewer for this observation. We agree that the immunostaining signals for several target genes are relatively weak, which reflects their low endogenous expression levels. To address this, we have included higher-magnification insets in the indicated panels (Revised Figure 5A, E, G and Figure 6D, F) to show the positive signals.

      Reviewer #2 (Recommendations for the authors):

      (1) What is the functional classification of DEG in F2A based on? GO terms?

      We thank the reviewer for this constructive question. The functional classification of differentially expressed genes (DEGs) in F2A is based on Gene Ontology (GO) terms. For each DEG, we retrieved its associated GO annotations across the three main categories (biological process, cellular component, molecular function). In cases where a gene was assigned multiple GO terms, we prioritized the most representative or significantly enriched term for functional interpretation. This clarification has been incorporated into the revised figure legend and the according GO number has been included in the figure.

      (3) The rationale for focusing on CHOP is not clear because Ddit3 is not shown in the spatial transcriptomics in F2A and is not significant in F2B, contradicting what is stated in line 206.

      We thank the reviewer for raising this important point. We apologize that Ddit3 was missing from the original figure. In the revised manuscript, we have included an updated version of Figure 2A, which now shows that Ddit3 is indeed one of the differentially expressed genes (DEGs) in the Mid zone at both 3 and 6 hours post-APAP. We agree with the reviewer that, as shown in Figure S1G (previous Figure 2B), Ddit3 did not reach statistical significance, due to its relatively low expression level in that analysis. Nevertheless, when we examined transcription factor (TF) activity in the Mid zone during early AILI, Ddit3 and Atf3 ranked as the top two most highly expressed TFs among the top ten with the highest activity, whereas Atf4 ranked seventh (Revised Figure 4B and Figure S3B). Given that Ddit3 frequently co-worked with Atf4 and that the Atf4–Ddit3 axis plays a well-established role in cellular stress adaptation, we considered this pathway to be biologically relevant and worthy of further investigation.

      (3) The term "redistribution" used in line 197 to describe the expression of Cyp2e1 and other Cyps in the midlobular zone seems inappropriate, considering that they just continue to be expressed there, whereas pericentral hepatocytes are dying in F3B; the same applies to "Gene Expression Shift" in F3H.

      We thank the reviewer for this important clarification. We have revised the text (Revised manuscript, page 9, line 234-236) to state that selective loss of Cyp‑expressing pericentral hepatocytes leads to the mid‑zone becoming the primary site of residual Cyp activity. The figure label has been changed from “Gene Expression Shift” to “Peri‑necrotic Cyp retention” and the legend now explicitly notes that this is an apparent zonal shift due to necrosis, not active redistribution.

      Reviewer #3 (Recommendations for the authors):

      (1) Please do not use abbreviations like AILI. This makes the paper more difficult to read.

      We thank the reviewer for pointing this out. We have replaced AILI with the full term “APAP-induced liver injury” to ensure easiness for readers.

      (2) It will be important to clarify how pericentral, mid, and periportal were defined. In Figure 1, it appears that some of the pericentral hepatocytes that are Ki67 positive are quite mid-zonal. It would be important to have rigorous definitions for the location determination.

      We thank the reviewer for this constructive comment. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x <sup>2</sup> + z <sup>2</sup> – y <sup>2</sup>) / (2z <sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885).

      We thank the reviewers for their rigorous critique again. We thank eLife for fostering an environment of fairness and transparency that enables authors to communicate openly and present their data honestly.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Manuscript number: RC-2026-03460

      Corresponding author(s): Louise, Walport

      1. General Statements

      We thank the reviewers for their critical reading and insightful comments on the manuscript and for highlighting the relevance of our data to developmental biology audiences.

      Below we have included a detailed point-by-point response to each reviewer comment, divided into those that are linked with revisions in the manuscript, and those that are not. As well as this we have written two general statements regarding the recently published high-resolution structures of the CPLs and our findings linking the CPLs to other oocyte structures, both of which address multiple reviewer comments.

      High resolution CPL structure papers

      While this manuscript was under review, several papers were published in Nature or deposited on bioRxiv detailing high-resolution structures of the CPLs (DOIs: 10.1038/s41586-026-10360-7 ,10.1038/s41586-026-10513-8, 10.1038/s41586-026-10442-6, 10.64898/2026.03.22.713481). These structural studies validate the key scaffolding function of PADI6 in CPL formation, as well as the association of ubiquitination machinery (including the SCF complex) to the CPLs, which we also determined in this work through mass spectrometry. Our work is complementary to these studies. The published structures provide detailed insight into how the CPLs form and function. However, while the new papers identify the structure and composition of the core CPL fiber, due to the requirement for particle averaging, they cannot resolve the identity of proteins beyond this core that interact or are stored on the CPLs sub-stoichiometrically. This limitation is addressed by our sub-cellular single oocyte proteomic workflow which pools all CPLs within each oocyte, identifying proteins which become soluble upon CPL dissolution, whether due to presence in the core fibre or to association with the fibres. Our manuscript therefore provides a complementary atlas of the CPL-associated proteome as a resource for further studying the function of CPLs in early development. To place our work in the context of the new manuscripts we have added the following to the discussion:

      • “During preparation of this manuscript, several high-resolution structures of the CPLs were reported80–83. Our proteomic workflow identified all CPL proteins resolved in these published structures as CPL-associated, including the newly identified SCF complex components and several F-box proteins. Notably, whilst highly informative, the structural studies cannot resolve the precise identity of individual proteins from families of structurally homologous proteins that form the CPL core, or of proteins which associate with the CPLs sub-stoichiometrically. One key example is the F-box proteins which offer substrate specificity to the SCF complex. Our proteomic analysis identified 10 different F-box proteins to have CPL-association. By contrast each published structure reports only a single or a few F-box proteins as part of the core complex. Notably different F-box proteins are reported in different structures. Our data provides an explanation for this variation, suggesting that even the core CPL fibres are non-homogeneous in the cell with different F-box proteins present in different locations (Fig 6B). Similarly, the close homology of a/b-tubulin proteins prevents identification of exact isoforms in the reported structures. Our data suggests that many different isoforms contribute to the CPLs (Fig 5E). Additionally, it is proposed that the CPLs act as storage hubs for a much wider range of proteins30. By allowing identification of proteins associated with the CPLs even at low occupancy our CPL-associated proteome provides a list of potential CPL-associated proteins for further structural and functional characterisation of CPL function.” CPLs and other oocyte structures

      The presence of proteins from other large cellular structures in the CPL-enriched protein dataset was highlighted by reviewers 1 and 3, in particular reviewer 3: “the physical meshwork of CPLs may prevent loss of CPL-associated proteins as well as cytoplasmic protein complexes or organelles that are too large to escape the cytoplasm through the CPLs. This is particularly a concern for the authors' conclusions regarding high amounts of ELVA-associated and mitochondrial and mitochondria-associated proteins associated with the CPLs given the large size of these organelles/structures. Is there any evidence by an alternative method of direct association of ELVAs or mitochondria with CPLs? Others have not detected mitochondrial proteins associated with CPLs”.

      Functional cross-talk with ELVA:

      Regarding an association with the ELVA, we should clarify that we believe that the relation between CPLs and the ELVA is most likely not a direct association but a previously unknown functional interaction and we have revised the text to more clearly reflect this view. We identify proteins involved in protein degradation, most notably the SCF complex, in our CPL-enriched dataset. In support of their presence being likely due to direct CPL association, rather than indirect trapping of the large ELVA, the recently published CPL structures similarly identify numerous proteins involved in protein degradation as associated with the CPLs, suggesting these proteins are sequestered and inactivated on the CPLs. As we do, these other works also propose that the CPLs are critical hubs for proteostasis in the oocyte and early embryo, similar to the proposed function of the ELVAs. However, as discussed below in the context of the mitochondria, we acknowledge that it is possible that the high level of these proteins observed in our dataset could alternatively be due to reduced cytoplasmic escape and we have updated the limitations section of the manuscript to reflect this caveat in our interpretation.

      • “Whilst we interpret proteome differences in Triton X-100 treated Padi6-KO oocytes to indicate an association with the CPLs, in the case of proteins associated with other large cellular structures such as mitochondria or the ELVA it is possible that this difference is instead due to physical entrapment of these structures by the CPLs during the precipitation step.” Interestingly, we identify RUFY1 in the CPL-enriched fraction. RUFY1 is a key marker of the ELVA where it functions as a scaffolding protein. In the manuscript we discussed the possibility that RUFY1 could therefore perform a similar function on the CPLs. In a recent work, it was reported that the number and size of RUFY1 compartments was increased in Padi6-null oocytes (DOI: 10.1038/s41594-026-01758-y). As RUFY1 is the key marker of the ELVAs this shows that their morphology is likely affected in the absence of Padi6. Together these data point towards a functional crosstalk between the two, potentially via the CPLs storing protein degradation machinery required by the ELVAs, as well as also storing/sequestering RUFY1, which could explain the increase in RUFY1 compartments in the absence of the CPLs.

      Mitochondria:

      Whilst it has been reported that there are defects in mitochondrial localisation in the absence of PADI6 and the CPLs (10.1016/j.ydbio.2010.11.033), we agree with the reviewers that it is possible these are not due to direct interactions between the CPLs and mitochondria, but rather from two separate roles of PADI6 and that our observation of mitochondrial proteins in the CPL-enriched set is due to increased physical entrapment of the large mitochondria by the CPLs rather than association. As additional links between the mitochondria and CPLs are not present in the literature unlike protein degradation machinery, we have removed the section entitled ‘The CPLs are associated with the oocyte mitochondria’, along with supplementary Figures 6E and 6F and added to the text the caveat that proteins associated with large cellular structures in our datasets could be due to physical entrapment. As we did not further discuss links between the CPLs and mitochondria in either the abstract or discussion, we do not believe the removal of this section significantly affects either the findings or novelty of this work.

      PADI6 Catalytic Activity

      Both reviewers 1 and 2 ask that we experimentally validate the loss of catalytic activity in the PADI6-C663A mutant. Along similar lines, reviewer 3 questions the rationale for the Padi6-C663A mouse line based on the lack of in vitro catalytic activity. We would like to reiterate that prior to this work there has been contradictory evidence regarding the catalytic activity of PADI6. Early work into PADI6 (10.1016/j.mce.2007.05.005) detected citrulline by IHC specifically in the oocytes of ovarian sections, which was absent in PADI6 knock-out ovaries. Later, it was shown by IF using anti-Citrulline antibodies that there is nuclear citrulline staining in 2-cell and 4-cell embryos that is ablated by a PADI inhibitor, in that study attributed to PADI1 (PADI1 presence detected by antibody-based techniques such as IF, 10.1038/srep38727). Together these manuscripts all posit that there is an active deiminase in oocytes and early embryos. We do not detect transcripts or protein for any PADI aside from PADI6 at the 2-cell stage and before suggesting any citrullination in the 2-cell embryo or before could only be a result of deimination by active PADI6 (Figure S4E).

      However, we and others, have confirmed that wild-type PADI6 is not active in vitro under the same conditions as the other PADIs (DOIs: 10.1016/j.csbj.2024.08.019, 10.4236/abb.2011.24044). Whilst this may be interpreted as overall lack of catalytic activity, an alternative explanation is that incorrect assay conditions have been used. For in vitro assays, supra-physiologically high calcium concentrations are required to activate catalysis by PADIs 1 to 4. We have previously shown that these calcium binding residues are not conserved in PADI6 and PADI6 does not bind calcium. It is therefore possible that PADI6 has evolved to be activated by other as yet unknown activating signals so as to not have its function disrupted during the large calcium transient post-fertilization. In the absence of the correct activating signals, in vitro activity would not be expected, even if the enzyme is catalytically active in vivo.

      We believe that this, together with the conservation of the catalytic tetrad residues, and the contradictory evidence regarding citrullination in vivo demonstrated that based on the current literature a catalytic function of PADI6 cannot be ruled out in vivo. This was our rationale for developing the Padi6-C663A mouse model in this work, to disentangle the dramatic phenotypes observed from full knockout of PADI6 from a possible catalytic function. Based on our previous structural characterization of human PADI6 (DOI: 10.1016/j.csbj.2024.08.019), if PADI6 were to have catalytic activity it would be through cysteine 663. Our data conclusively shows that catalytic activity of PADI6 through cysteine 663 is not required for murine female fertility, but that mice with this mutation are also not fully wildtype. Given the lack of large-scale structural damage loss of this cysteine imparts on the protein (Supplementary Figure 1), the observed phenotypes point towards either the disruption of a previously unidentified non-essential catalytic function or to more subtle changes to function from this mutation, for example through alterating the binding affinities of proteins that interact with PADI6 lacking this cysteine.

      2. Point-by-point description of the revisions

      Description of revisions incorporated into the manuscript along with discussion of reviewer comments

      Reviewer 1:

      • Major Comment 1: “1- The methods for differential gene expression analysis are insufficiently described. Were these calculated using Scanpy? If so, the rationale for this choice should be provided, as the number of replicates and the nature of the data appear more suited to standard bulk differential expression frameworks such as DESeq2 or edgeR.”
      • Response and Incorporated Revision: We chose the Scanpy framework for analysis because it allows for quality control, PCA, normalization, and differential expression within a single Python environment. To calculate differential gene expression, we performed Student's t-test on log-normalized counts for each gene followed by correction for multiple testing using the Benjamini-Hochberg procedure. We have added this further clarification to the methods section:

      “To identify dysregulated genes, for each stage, mean log2CPM ratios (test vs. wild-type) and q-values were calculated for each gene using a custom Python script (see Data and Code Availability), extracted and plotted using GraphPad Prism. In brief, an independent Student’s t-test was applied to the normalized expression values and p-values were corrected for multiple testing using the Benjamini-Hochberg procedure.”

      • Major Comment 2:

      *“2- In Figure 4C, it is unclear what correlation is being calculated. Additionally, the methods state that differential protein abundance was determined using the same approach as the scRNA-seq analysis. Given the lack of methodological clarity noted above, it is difficult to evaluate these results. The authors should justify whether the chosen method is compatible with the normalization approach used for the proteomics data.”

      *

      • Response and Incorporated Revision: We apologize for the lack of clarity on the calculated correlation and to improve this have amended the Figure legend for 4C to the following: “(C) Pearson correlation values for protein abundances in wild-type, Padi6-C663A and Padi6-KO GV oocyte and 2-cell embryo replicates when compared to all other replicates of the same condition. Samples of the same genotype and stage show high correlation in the abundance of individual proteins.”

      Regarding the compatibility, dysregulated proteins and transcripts were determined using the same statistical method, independent Student’s t-tests corrected for multiple testing with the Benjamini-Hochberg procedure.

      • Major Comment 4:

      “4- A notable result is the complete lack of correspondence between differential RNA-seq and proteomics in 2-cell embryos. The authors state that "This is consistent with data showing that minimal translation occurs in the 2-cell embryo, with the first large translational wave only occurring in the morula," citing Israel et al. 2019. This statement is factually incorrect. The cited study did not measure translation directly. Recent work that specifically measured translation during preimplantation development has demonstrated that translation is highly dynamic throughout these stages (Ozadam et al. 2023, Nature). This claim should be corrected and the results discussed in the context of these more recent findings.”

      • Response and Incorporated Revision: We have revised the claim, discussing our results in the context of the more recent Ozadam et al. 2023, Nature paper. The amended text reads: “This is consistent with data showing that protein abundance does not correlate with RNA level, but with ribosome occupancy and translation efficiency in the zygote58,59. Our results show that this lack of correlation in RNA and protein is maintained in the 2-cell embryo with RNA dysregulation and protein dysregulation uncoupled in Padi6-KO embryos, suggesting that transcriptomic and proteomic dysregulation need not be linked at early developmental stages.”

      • Major Comment 6:

      *“6- A large fraction of the manuscript interprets changes in the PADI6 knockout proteomics data as direct evidence that CPLs are associated with specific protein complexes (e.g., ribosomes) or cellular structures (e.g., mitochondria). However, these results are indirect and could alternatively be explained by roles of PADI6 that are independent of CPLs. These conclusions should be tempered, or this limitation should be explicitly acknowledged.”

      *

      • Response and Incorporated Revision:

      We have addressed this comment in the General Statements section discussing links between the CPLs and the mitochondria. Given the previously identified links between protein degradation machinery and ribosomes with the CPLs, we believe the changes are due to the loss of CPLs. We have however included the following statement acknowledging that these defects could be due to alternate function of PADI6 in the Limitations of the study section: “…It is also possible that some differences are due to an alternate defect that alters protein solubility caused by the absence of PADI6 independent of CPL formation.”

      • Major Comment 7:

      “7- Relatedly, the authors conclude that CPLs are associated with mitochondria based on the CPL proteomics data. Several questions arise: Do the EM images show mitochondria in close proximity to CPLs? Is mitochondrial morphology affected in Padi6-null oocytes? Given that respiratory chain complexes are large, membrane-bound assemblies, how might they associate with CPLs? Could the Triton-insoluble CPL fraction be contaminated with other oocyte-specific superstructures, as the authors themselves allude to? Do the EM images show mitochondria in close proximity to CPLs? Is mitochondrial morphology affected in Padi6-null oocytes?

      Reviewer 2:

      • Major Comment 5: “5) The authors found that CPLs are associated with mitochondria. How do CPLs in the cytosol interact with the components of the electron transport chain in the mitochondria? Do CPLs directly interact with these proteins in the cytosol or attach to the outer mitochondrial membrane? Does the loss of PADI6 affect the morphology, membrane potential, and ROS production of mitochondria in oocytes?”
      • Combined response to reviewer 1 major comment 7 and reviewer 2 major comment 5 and Incorporated Revision: As stated in the general statement, whilst it has been reported that there are defects in mitochondrial localisation in the absence of PADI6 and the CPLs (10.1016/j.ydbio.2010.11.033), it is possible these are not due to direct interactions. Presence of mitochondrial proteins in the CPL-enriched proteome could instead be caused by physical entrapment of the mitochondria by the CPLs. Whilst mitochondrial localization is known to be disrupted in PADI6 knockout oocytes (DOI: 10.1016/j.ydbio.2010.11.033), as additional links between the mitochondria and CPLs are not present in the literature unlike protein degradation machinery, we agree further work would be required to support this claim and have therefore removed the section entitled ‘The CPLs are associated with the oocyte mitochondria’, along with supplementary Figures 6E and 6F. As we did not further discuss links between the CPLs and Mitochondria in either the abstract or discussion, we do not believe the removal of this section significantly affects both the findings and novelty of this work.

      • Major Comment 8:

      “8-The authors should discuss their findings in the context of Liu et al. 2026 (Nature), which elucidates the structural basis of PADI6 in CPL formation. This comparison would be particularly informative given the overlapping scope of the two studies.”

      • Response and Incorporated Revision: As discussed in the general statement, we have included a new discussion of how the published CPL structure papers compare to our manuscript in the general comments section above. We have incorporated the following in the Discussion section of our manuscript: “During preparation of this manuscript, several high-resolution structures of the CPLs were reported80–83. Our proteomic workflow identified all CPL proteins resolved in these published structures as CPL-associated, including the newly identified SCF complex components and several F-box proteins. Notably, whilst highly informative, the structural studies cannot resolve the precise identify of individual proteins from families of structurally homologous proteins that form the CPL core, or of proteins which associate with the CPLs sub-stoichiometrically. One key example is the F-box proteins which offer substrate specificity to the SCF complex. Our proteomic analysis identified 10 different F-box proteins to have CPL-association. By contrast each published structure reports only a single or a few F-box proteins as part of the core complex. Notably different F-box proteins are reported in different structures. Our data provides an explanation for this variation, suggesting that even the core CPL fibres are non-homogeneous in the cell with different F-box proteins present in different locations (Fig 6B). Similarly, the close homology of a/b-tubulin proteins prevents identification of exact isoforms in the reported structures. Our data suggests that many different isoforms contribute to the CPLs (Fig 5E). Additionally, it is proposed that the CPLs act as storage hubs for a much wider range of proteins30. By allowing identification of proteins associated with the CPLs even at low occupancy our CPL-associated proteome provides a list of potential CPL-associated proteins for further structural and functional characterisation of CPL function.”

      • Minor Comment 1:

      “1- For the scRNA-seq analysis, please expand the methodology for PCA. Specifically, how are read counts normalized prior to PCA? We assume some form of log-normalization was applied, but this is not described in the methods or figure legends.”

      • Incorporated Revision: We apologize for this lack of clarity; log normalization was applied. The methods have been updated to the following to describe normalization prior to PCA: “PCA analysis of log normalized CPM values was performed using ScanPy on the full dataset, as well as oocyte, zygote, and 2-cell split data.”

      • Minor Comment 2:

      “2- Please clarify the z-score calculation for results shown in Figure 3. While the general approach can be inferred from context, the exact calculation is not provided.”

      • Incorporated Revision: Z-scores were calculated with the ScanPy scale function, the following has been added into the methods to clarify this: “Transcript Z-scores were calculated using the pp.scale function in the Python package ScanPy…”

      • Minor Comment 6:

      “6- It is unclear why the authors conclude that PADI6 regulates UHRF1 at the protein level (Figure S5). The observation that UHRF1 levels are reduced in Padi6-null oocytes could simply reflect reduced maternal deposition. The evidence does not appear sufficient to support a specific claim of protein-level regulation by PADI6.”

      • Incorporated Revision: We agree that it is possible that the reduced UHRF1 levels could be due to reduced maternal deposition, for example by reduced protein translation levels of UHRF1, and therefore we have amended our conclusion to the following: “…suggesting PADI6 regulates UHRF1 at the protein level or conceivably at the translational level during maternal deposition.”

      Reviewer 2:

      • Major Comment 4: “4) The authors also found that several proteasome subunits, as well as LAMP1 and RUFY1, were enriched in CPLs. These proteins are known to localize to ELVAs in GV and MII oocytes (Zaffagnini et al., 2024), suggesting functional crosstalk between CPLs and ELVAs. The authors should confirm that these components localize to the CPLs as well as ELVAs by immunostaining or using fluorescently labeled proteins. Are the morphology and localization of ELVAs affected by the loss of PADI6? Do CPLs colocalize or interact with ELVAs during oocyte maturation? It was reported that ELVAs were disassembled when RUFY1 was removed by Trim-Away in oocytes. Does the loss of RUFY1 affect CPL formation?”
      • Response and Incorporated Revision: Unfortunately, the only way to directly validate protein localisation to the CPLs is by expansion microscopy, which we do not have the technical capacity to do and therefore cannot perform these experiments in an informative manner. Reviewer #3 agrees with this conclusion: “Although some of the suggestions by Reviewer #2 could provide interesting information, the immunofluorescence experiments suggested in my view are not likely to provide definitive information regarding association of specific proteins or structures with CPLs. Instead, higher resolution technologies such as proximity ligation or immuno-EM might be required. These experiments seem like good ways to extend the findings beyond the current manuscript, but I think are not essential for the major take home points.”

      Regarding the morphology and localization of ELVAs in the absence of PADI6, it was recently reported that the number and size of RUFY1 compartments was increased in Padi6-null oocytes (DOI: 10.1038/s41594-026-01758-y). As RUFY1 is the key marker of the ELVAs this suggests that their morphology is likely affected, further pointing towards a functional crosstalk between the two. To highlight this we have added the following sentence in the discussion: “Additionally, recent work identified an increase in the number and size of RUFY1 and ProteoStat positive compartments in Padi6-null oocytes, further pointing towards a functional crosstalk between the ELVA and CPLs.” However, testing whether RUFY1 loss affects CPL formation is beyond the scope of this work investigating the functions of PADI6.

      Reviewer 3:

      • Major Comment 4: ‘4- Proteomics analysis: The authors carried out proteomics analysis on oocytes treated with Triton X-100 so that they would retain only cytoskeleton-associated proteins. As a control, Padi6-null oocytes (lacking CPLs) were used, and the authors interpret the proteins identified in the WT and not in the Padi6-null as CPL-associated proteins. It is not clear to me that this is a reasonable interpretation of the results. My concern is that the physical meshwork of CPLs may prevent loss of CPL-associated proteins as well as cytoplasmic protein complexes or organelles that are too large to escape the cytoplasm through the CPLs. This is particularly a concern for the authors' conclusions regarding high amounts of ELVA-associated and mitochondrial and mitochondria-associated proteins associated with the CPLs given the large size of these organelles/structures. Is there any evidence by an alternative method of direct association of ELVAs or mitochondria with CPLs? Others have not detected mitochondrial proteins associated with CPLs (see J. Li et al., doi 10.1038/s41594-026-01758-y).”
      • Response and Incorporated Revision: We have addressed this comment in the General Statements section entitled “CPLs and other oocyte structures”. We believe links between the CPLs, and protein degradation machinery and the ELVA are well supported by both our data, and the recent and past literature covering the CPLs (DOIs: 10.1038/s41594-026-01758-y, 10.1038/s41586-026-10360-7 ,10.1038/s41586-026-10513-8, 10.1038/s41586-026-10442-6, 10.64898/2026.03.22.713481, 10.1016/j.cell.2024.01.031, 10.1016/j.cell.2023.10.003). As additional links between the mitochondria and CPLs are not present in the literature unlike protein degradation machinery, we have removed the section entitled ‘The CPLs are associated with the oocyte mitochondria’, along with supplementary Figures 6E and 6F. As we did not further discuss links between the CPLs and Mitochondria in either the abstract or discussion, we do not believe the removal of this section significantly affects both the findings and novelty of this work.

      Responses to other reviewer comments including analyses that the authors prefer not to carry out

      Reviewer 1:

      • Major Comment 3: “3-The single-embryo proteomic measurements are an important aspect of the paper. However, additional quality control data are needed to assess data quality. In particular, a more systematic comparison to Ye et al. would strengthen confidence in these measurements.”
      • Response: Unfortunately, the Ye et al. work has not released a list of the proteins identified in oocytes and early embryos, or their intensities, therefore we cannot compare in this manner. In terms of overall number of proteins identified the two approaches are comparable, however they differ in both their sample preparation and MS acquisition methods.

      • Major Comment 5:

      “5- The manuscript refers to the C663A mutation as a "catalytic mutant." While the structural and homology-based rationale is compelling, the entire paper's conclusions depend on this interpretation. Experimental validation of the inferred loss of catalytic activity would substantially strengthen the study.”

      • Reviewer 2 Major Comment 1: “1) There is insufficient evidence to conclude that the C663A mutant is catalytically inactive. The authors should conduct an in vitro citrullination assay to show whether wild-type PADI6 has peptidyl arginine deiminase activity, but the C663A mutant loses it.”
      • Combined response to reviewer 1 major comment 5 and reviewer 2 major comment 1: We are pleased the Reviewer 1 finds our structural and homology-based rationale for the design of the C663A mutation compelling. As discussed in the general statements, prior to this work no catalytic activity of PADI6 had been observed in vitro, despite contradictory in vivo data regarding its activity. It was this challenge in replicating in vivo conditions that might be required for protein activation in an in vitro assay that directly led us to develop our in vivo mouse model in this work. It is therefore not possible to experimentally validate any change in activity following the C663A mutation as wild-type PADI6 is also inactive under the in vitro assay conditions used for other PADI isozymes. Except when discussing the design of the mouse itself we have been careful to always describe the mouse based on its mutation rather than as a catalytic mutant and have used terms such as “putative” or “potential” in the manuscript to make clear that there is no direct evidence for any catalytic activity that could then be abolished.

      • Minor Comment 3:

      “3- MII oocytes are used for scRNA-seq experiments and GV oocytes for proteomics. Please provide a rationale for the use of two different developmental stages.”

      • Response: The reasons for using MII oocytes over GV oocytes in the RNA-seq experiments was due to availability and sample number requirements. GV oocytes were used in place of MII oocytes for proteomic experiments as it was possible to gather many more GV oocytes per mouse than MII oocytes. Therefore, to increase the number of replicates in the single oocyte/embryo proteomics workflow developed in this work, we chose GV oocytes to increase confidence and show reproducibility.
      • Minor Comment 4:

      “4- While the data support the conclusion that maternal RNA and minor EGA mRNA degradation is defective, none of the experiments directly measure mRNA degradation. Direct experimental validation, even for a few select targets using standard decay assays, would strengthen this claim.”

      • Response: Whilst our data does not directly measure mRNA degradation, we believe our data is sufficient evidence to state that mRNA degradation is defective in the absence of PADI6, in line with other work (DOI: 10.1101/gad.351238.123 and consequently that it would not be appropriate to use further mice for these experiments in line with the 3Rs.
      • Minor Comment 5:

      *“5- The statement "Together these results indicate that we have established a powerful sub-cellular proteomic workflow from single mouse oocytes capable of identifying proteins associated with the CPLs" overstates the findings. The results are consistent with this interpretation, but the approach described is not a sub-cellular proteomic workflow in the spatial proteomics sense. This language should be revised.”

      *

      • Response: We agree that our workflow is not a sub-cellular proteomic workflow in the spatial proteomics sense, but we believe our statement and discussion does not claim that our workflow is a spatial proteomic workflow at any point. We therefore do not believe any revision of language is necessary.
      • Minor Comment 7:

      “7- Many ribosomal, proteasomal, and mitochondrial proteins appear to associate with CPLs in a PADI6-dependent manner. Could an alternative explanation be that maternal deposition of these proteins is globally reduced in Padi6-null oocytes, rather than their association with CPLs being specifically affected?”

      • Response: A global reduction in the maternal deposition of CPL-associated proteins would be reflected by a decrease in the levels of these proteins in intact oocytes. As the protein levels of the majority of CPL-associated proteins are not reduced in Padi6-KO oocytes (Figure 6A-B and Figure S6B), and for those that are reduced the effect is generally subtle, this discounts a global reduction in their maternal deposition.

        Reviewer 2:

      • Major Comment 2: “2) The author found that the levels of key CPL scaffolding proteins from the SCMC (OOEP, TLE6, NALP5, and KHDC3) were not affected by the loss of PADI6. It should be examined whether the subcellular localization of these proteins is affected in Padi6-deficient and C663A mutant embryos using immunostaining.”

      • Major Comment 3: “3) The authors found that CPLs contain components of the SKP1-CUL1-F-box protein (SCF) ubiquitin ligase complex. They also found that hPADI6 interacted with CUL1 when it was transiently transfected into HEK-293T cells. It is important to examine whether the stability or subcellular localization of these proteins is affected by the loss of PADI6 in oocytes.”

      • Combined response to reviewer 2 major comment 2 and 3: From our intact GV oocyte proteomics experiments, we know that the stability of the SKP1-CUL-F-box proteins is not affected by the loss of PADI6, similar to the key CPL scaffolding proteins highlighted in major comment 2. Regarding localization, it was reported by Jentoft et al. in 2023 that due to the cytoplasmic abundance of CPL proteins, their true cellular distribution can only be measured by IF using a Halo-tag knock-in line to circumvent the use of secondary antibodies which aggregate at the subcortex. Alternatively, expansion microscopy could be used to determine differences in localization. Unfortunately, we do not have the capacity to generate these lines or perform expansion microscopy and therefore we are not able to conduct these experiments in an informative manner. Reviewer #3 agrees with this conclusion: “Although some of the suggestions by Reviewer #2 could provide interesting information, the immunofluorescence experiments suggested in my view are not likely to provide definitive information regarding association of specific proteins or structures with CPLs. Instead, higher resolution technologies such as proximity ligation or immuno-EM might be required. These experiments seem like good ways to extend the findings beyond the current manuscript, but I think are not essential for the major take home points.”
      • Major Comment 6:

      “6) Figure 6C, G, and I.

      Statistical analysis should be done.”

      • Response: We have performed statistical analysis of Figure 6G and demonstrate that the PADI6-N598S variant binding to UHRF1 is statistically significantly impaired (see below). However, the immunoprecipitation assay performed is a largely qualitative assay and we don’t believe that detailed quantification of it is appropriate. Similarly for Figures 6C and 6I we don’t think quantification is necessary as the assay represents presence or absence of a protein in a sample.

      Reviewer 3:

      • Major Comment 1: “Padi6 catalytic activity: Given that PADI6 was previously shown not to have catalytic activity in vitro, the rationale for doing the experiment mutating Padi6 function is weak. The authors claim that a homologous mutation in human PADI6 does not "significantly damage the folded state of PADI6", but this conclusion does not necessarily mean that a scaffolding or protein interaction function could not be affected by the mutation. The authors provide zero evidence of catalytic activity in the WT oocytes (which I agree would be technically quite challenging given the poor quality/specificity of antibodies that recognize citrullinated proteins and low amount of protein available for mass spec analysis) but still include a full paragraph in the Discussion regarding the potential catalytic activity and why it might be important. This focus implies that the underlying data support the concept, even though the authors do frame the paragraph carefully.”
      • Response: We have primarily addressed this comment in the General Statements section of this document. Regarding scaffolding or interactional functions, as shown in our published X-Ray crystal structure of PADI6 (DOI: 10.1016/j.csbj.2024.08.019), the proposed catalytic cysteine is buried, not surface exposed, and doesn’t appear to be involved in structural interactions or disulphide bonds. We cannot however rule out subtle structural changes around the active caused by the C663A mutation resulting in altered protein-protein interaction binding affinities at proteins interacting near to the proposed PADI6 active site. To account for this possibility the following sentences have been added/amended in the discussion to read:

      “Given the lack of large-scale structural damage the loss of C663 imparts on PADI6, the possibility of a non-essential catalytic function of PADI6 in oogenesis and early embryo development cannot be ruled out, potentially in the epigenetic regulation of transcription similar to PADI4. Alternatively, it is possible that the C663A substitution alters protein binding affinities for interactions on or near to the proposed PADI6 active site.”

      • Major Comment 2:

      *“2. Padi6 mutant embryo development: The Padi6 mutant embryo development findings are minimally different from WT controls. The embryos were all cultured in vitro and it is unclear if they would have developed fine in vivo, which is suggested from the lack of a difference in litter sizes, which if anything were slightly higher in the Padi6 mutant females. In the absence of additional useful information regarding why the development was slightly lower, this experiment does not seem to add to the conclusions of the paper but seems more like an incomplete side note that should be more deeply investigated.”

      *

      • Response: An explanation for the lack of difference in litter sizes but difference in developmental potential has been discussed in the text: “This phenomenon (significant decrease in early embryo numbers in one mouse line over another despite litter sizes remaining comparable) has been observed previously and is attributed to mice, and other species, producing greater numbers of eggs and pre-implantation embryos than the uterus can accommodate, with excess embryos lost during the pre-implantation stage50–52.” The difference in developmental potential is statistically significant for Padi6-C663A embryos. Without an impaired function of PADI6 we do not see how in vitro culture of the embryos would result in such a difference in developmental potential of the mutant compared to the wild type embryos as they were cultured under the same conditions. We agree with the reviewer that we haven’t yet determined the underlying cause for this difference but we think nonetheless that it is an important finding to highlight - that a single cysteine mutation in the active site of PADI6, which does not affect protein structure significantly impacts the development of early-stage embryos.

      • Major Comment 3:

      “3. Padi6 mutant 2C embryo EGA timing: The altered transcription in the Padi6 mutant 2C embryos appears to indicate that they are ahead in development relative to the WT based on the PCA plot and the relative downregulation of minor ZGA genes and upregulation of major ZGA genes. The 1-cell embryos were collected from spontaneously ovulating mice and the time of development was not controlled in any way. Mouse embryos are quite variable in their exact timing of development, even across different embryos in the same mouse. I find these changes in transcription likely to be explained by differences in developmental timing and I don't think the authors have robustly shown "dysregulation of EGA". Similarly, the delay in development of the Padi6-null embryos from zygote to 2C (Figure 1D) explains why the maternal mRNAs are upregulated in the Padi6-null mice - they are simply delayed in development.

      • Response: When harvesting 2-cell embryos, samples from all four mice were harvested on different days at the same time of day. If the differences between samples were only due to differences in developmental timing we would anticipate as significant differences between the embryos from mice with the same genotype as between those from different genotypes which is not what we observe. Given the significant developmental defects observed in these embryos (Figure 1), it is highly likely these are associated with defects on the transcriptional level. Furthermore, we observed a small but significant delay in Padi6-C663A embryos reaching the 2-cell stage (Figure 1D) which we believe makes it highly unlikely that the embryos from both C663A females were further along in development compared to those from both wild-type females.

      The reviewer states that the upregulation of major EGA genes and downregulation of minor EGA genes further points toward an advancement in development. If this is the case, then it would be expected that there would also be increased degradation of maternal transcripts which decrease between the zygote and 2-cell stage in wild-type embryos. We do not see a further decrease of these transcripts in Padi6-C663A embryos. Finally, the reviewer notes the PCA plot as a reason for being advanced in development – PCA only measures differences between samples, not developmental time.

      Taking into account the above, we do not agree with the reviewer’s interpretation of our data. Regarding Padi6-null mice, disrupted EGA has been reported in Padi6-null mice in other work (DOI: 10.1101/gad.351238.123.).

      • Minor Comment 1:

      “1. What was the point of splitting up the 2C embryo blastomeres rather than treating them as single embryos?”

      • Response: 2C embryos were split up to investigate whether defects in Padi6-null 2-cell embryos were due to asymmetric inheritance of transcripts given the significant mis-localization of various oocyte structures in the absence of PADI6. As this was not clear in the text, we have added the following sentence clarifying the rationale and referencing Figure S4C-D where the transcriptomic correlation between blastomeres is shown: “The transcriptomes of separated blastomeres of the same 2-cell embryo showed high levels of correlation for embryos of each genotype suggesting asymmetric inheritance of transcripts is not a cause of PADI6 associated developmental defects (Figure S4C-D).”.

      • Minor Comment 2:

      “2. Proteomics - Because PADI6 makes up a significant fraction of total oocyte protein, and the Padi6-null oocytes don't have any PADI6, does this artificially increase the relative amount of the remaining proteins?”

      • Response: Any such effect would have been corrected during data normalization.
    2. Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.

      Learn more at Review Commons


      Referee #3

      Evidence, reproducibility and clarity

      Summary:

      PADI6 is a highly abundant oocyte-specific protein and component of cytoplasmic lattices (CPLs). The current study attempts to better characterize mouse PADI6 function and CPL-associated proteins using a novel Padi6 mutant allele (predicted to be catalytically dead), RNA sequencing, and single oocyte proteomics approaches. Key findings reported are:

      1. Validation of known Padi6-null female mouse phenotypes including infertility, disruption of CPL formation, impaired embryonic genome activation, and abnormal maternal mRNA degradation.
      2. Mice homozygous for the Padi6-mutant allele, when mated to WT males, have a slight impairment in preimplantation embryo development in vitro but have normal fertility as indicated by average litter sizes.
      3. CPLs are enriched in ribosomal proteins, proteins involved in protein degradation including ubiquitination proteins, and mitochondria. Several of the protein interactions with PADI6 were validated by showing co-immunoprecipitation of epitope-tagged proteins expressed in HEK-293 cells (UHRF1, UBE2D2, and CUL1).

      Major comments:

      • Are the key conclusions convincing?

      • Padi6 catalytic activity: Given that PADI6 was previously shown not to have catalytic activity in vitro, the rationale for doing the experiment mutating Padi6 function is weak. The authors claim that a homologous mutation in human PADI6 does not "significantly damage the folded state of PADI6", but this conclusion does not necessarily mean that a scaffolding or protein interaction function could not be affected by the mutation. The authors provide zero evidence of catalytic activity in the WT oocytes (which I agree would be technically quite challenging given the poor quality/specificity of antibodies that recognize citrullinated proteins and low amount of protein available for mass spec analysis) but still include a full paragraph in the Discussion regarding the potential catalytic activity and why it might be important. This focus implies that the underlying data support the concept, even though the authors do frame the paragraph carefully.

      • Padi6 mutant embryo development: The Padi6 mutant embryo development findings are minimally different from WT controls. The embryos were all cultured in vitro and it is unclear if they would have developed fine in vivo, which is suggested from the lack of a difference in litter sizes, which if anything were slightly higher in the Padi6 mutant females. In the absence of additional useful information regarding why the development was slightly lower, this experiment does not seem to add to the conclusions of the paper but seems more like an incomplete side note that should be more deeply investigated.
      • Padi6 mutant 2C embryo EGA timing: The altered transcription in the Padi6 mutant 2C embryos appears to indicate that they are ahead in development relative to the WT based on the PCA plot and the relative downregulation of minor ZGA genes and upregulation of major ZGA genes. The 1-cell embryos were collected from spontaneously ovulating mice and the time of development was not controlled in any way. Mouse embryos are quite variable in their exact timing of development, even across different embryos in the same mouse. I find these changes in transcription likely to be explained by differences in developmental timing and I don't think the authors have robustly shown "dysregulation of EGA". Similarly, the delay in development of the Padi6-null embryos from zygote to 2C (Figure 1D) explains why the maternal mRNAs are upregulated in the Padi6-null mice - they are simply delayed in development.
      • Proteomics analysis: The authors carried out proteomics analysis on oocytes treated with Triton X-100 so that they would retain only cytoskeleton-associated proteins. As a control, Padi6-null oocytes (lacking CPLs) were used, and the authors interpret the proteins identified in the WT and not in the Padi6-null as CPL-associated proteins. It is not clear to me that this is a reasonable interpretation of the results. My concern is that the physical meshwork of CPLs may prevent loss of CPL-associated proteins as well as cytoplasmic protein complexes or organelles that are too large to escape the cytoplasm through the CPLs. This is particularly a concern for the authors' conclusions regarding high amounts of ELVA-associated and mitochondrial and mitochondria-associated proteins associated with the CPLs given the large size of these organelles/structures. Is there any evidence by an alternative method of direct association of ELVAs or mitochondria with CPLs? Others have not detected mitochondrial proteins associated with CPLs (see J. Li et al., doi 10.1038/s41594-026-01758-y).
      • Are the data and the methods presented in such a way that they can be reproduced?

      The authors should be congratulated on the clarity and thoroughness of the Methods and Results descriptions in this manuscript. I have rarely seen this done so nicely. - Are the experiments adequately replicated and statistical analysis adequate?

      The experiments were replicated and analyzed appropriately.

      Minor comments:

      Points for clarification:

      1. What was the point of splitting up the 2C embryo blastomeres rather than treating them as single embryos?
      2. Proteomics - Because PADI6 makes up a significant fraction of total oocyte protein, and the Padi6-null oocytes don't have any PADI6, does this artificially increase the relative amount of the remaining proteins?

      Referees cross-commenting

      It was interesting to read the additional reviews on this manuscript. Regarding Reviewer #1's points, we seem to be in agreement, in particular regarding major comment 7 regarding the possibility that the Triton-insoluble CPL fraction may be contaminated with other oocyte-specific superstructures. We all had concerns regarding the lack of evidence that PADI6 has catalytic activity and the point mutant does not, which impacts interpretation of many of the embryo development experiments.

      Although some of the suggestions by Reviewer #2 could provide interesting information, the immunofluorescence experiments suggested in my view are not likely to provide definitive information regarding association of specific proteins or structures with CPLs. Instead, higher resolution technologies such as proximity ligation or immuno-EM might be required. These experiments seem like good ways to extend the findings beyond the current manuscript, but I think are not essential for the major take home points.

      Significance

      There are several very recent publications on CPLs, including two recently accepted papers (March 2026) reporting the structure of CPLs and associated proteins detected in situ or after CPL purification using cryo-EM and mass spec analysis (Chi et al., doi 10.1038/s41586-026-10442-6; Liu et al., doi 10.1038/s41586-026-10360-7). The associated proteins include UHRF1, tubulin proteins, and ubiquitin ligase components, similar to what was shown in the current manuscript. A third paper, also published in March 2026 (J. Li et al., doi 10.1038/s41594-026-01758-y), used cryo-EM and IP-mass spec combined with mutagenesis to draw similar conclusions. This manuscript specifically comments on the lack of associated mitochondrial proteins, conflicting with the current manuscript. Finally, a preprint (Y. Li et al, doi 10.64898/2026.03.30.715190) on the same topic was posted on bioRxiv April 1, 2026; similar methods were used and similar conclusions were drawn. The current manuscript stands out for using Padi6-null oocytes as a control, which in theory could improve the mass spec results to improve information regarding the extent and identity of associated proteins over what was done in these manuscripts (though see above for concerns related to this point). Further, if the authors were able to conclusively show that PADI6 catalytic activity played a role in preimplantation embryo development, it would be quite distinct from these papers that are focused on CPL structure and associated proteins.

      • State what audience might be interested in and influenced by the reported findings.

      The audience for this paper would be reproductive biologists or clinical infertility scientists, given the relevance to human Padi mutations. Should the Padi6 catalytic activity be defined and relevant to oocytes and early embryos, interest would be broadened to more general epigenetics of development and nuclear reprogramming. - Define your field of expertise with a few keywords to help the authors contextualize your point of view. Indicate if there are any parts of the paper that you do not have sufficient expertise to evaluate.

      Field of expertise: Reproductive biology, oocyte physiology, embryonic genome activation, epigenetics Lack of expertise: Proteomics

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      In this study, the authors have performed tissue-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior vs posterior cells of the C. elegans intestine. They have performed this analysis in fed and fasted states of the animal. The data generated will be very useful to the C. elegans community, and the role of pyruvate shown in this study will result in interesting follow-up investigations.

      However, several strong claims made in the study are solely based on in silico predictions and are not supported by experimental evidence.

      Strengths:

      Several studies in the past have predicted different functions of the anterior (INT1) vs posterior (INT2-9) epithelial cells of the C. elegans intestine based on their anatomy and ultrastructure, but detailed characterization of differences in gene expression between these cell types (and whether indeed these are different 'cell types') was lacking prior to this study. The genes and drivers identified to be exclusively expressed in the anterior vs posterior segments of the intestine will be very helpful to selectively modulate different parts of the C. elegans intestine in future studies.

      Another strength of this study is the careful experimental design to test how the anterior vs posterior cell types of the intestine respond differently to food deprivation and recovery after return to food. These comparisons between 'states' of a cell in different physiological conditions are difficult to pick up in single-cell analyses due to low sequencing depth, which can fail to identify subtle modulation of gene expression.

      The TRAP-associated bulk RNA-seq approach used in this study is more suitable for such comparisons and provides additional information on post-transcriptional regulation during metabolic stress.

      A key finding of this study is that pyruvate levels modulate the translation state of anterior intestinal cells during fasting. Characterization of pyruvate metabolism genes, especially of the enzymes involved in its mitochondrial breakdown, provides novel insights into how gut epithelial cells respond to the acute absence of food.

      Weaknesses:

      Unlike previous TRAP-seq studies (PMID: 30580965, 36044259, 36977417) that reported sequencing data for both input and IP samples, this study only reports the sequencing data for IP samples. Since biochemical pulldowns are variable across replicates, it is difficult to know if the observed differences between different conditions are due to biological factors or differences in IP efficiency. More importantly, since two different TRAP lines were utilized in this study and a large proportion of the results focus on the differences between the translational profiles of INT1 vs INT2-9 cells, it is essential to know if the IP worked with similar efficiency for both TRAP strains that likely have different expression levels of the HA-tagged ribosomal protein. One way to estimate this would be to perform qRT-PCR of genes that are known to be enriched in all intestinal cells and determine whether their fold-enrichment over housekeeping genes (normalized to input) is similar in INT1 vs INT2-9 TRAP strains and across the fed vs fasted conditions. The authors, in fact, mention variability across biological replicates, due to which certain replicates were excluded from their WGCNA analysis.

      We appreciate the reviewer's comments. We agree that the lack of matched input sequencing libraries limits our ability to directly assess IP efficiency across replicates, conditions, and TRAP strains. However, several features of the dataset support the conclusion that the major differences reported here reflect biological rather than purely technical variation. First, the RPL-22-3xHA construct was integrated into each line to improve consistency across experiments. Second, although the INT2-9 TRAP strain yielded more RNA than the INT1 strain, as expected given the larger number of labeled cells, downstream analyses were performed on normalized count data rather than raw counts. Third, principal component analysis showed robust separation by promoter identity across all conditions, and expected INT1-enriched genes such as ins-7 were recovered in the INT1 dataset. Together, these observations support the interpretation that the TRAP datasets capture reproducible, cell-type-specific differences in ribosome-associated transcripts. Nonetheless, we agree that direct input-normalized measurements would further strengthen the study, and we will explicitly note this as an important limitation.

      It appears that GFP expression is also detectable in INT2 (in addition to strong expression in INT1 in Fig.1A). Compared to INT3-9, which looks red, INT2 cells appear yellow, suggesting that the expression patterns of the two TRAP drivers are not mutually exclusive, which changes the interpretation of many of the results described in the study.

      We agree that the Pges-1ΔB promoter is not absolutely restricted to INT1 and that weak GFP expression can also be detected in INT2. Because Pges-1ΔB is an engineered promoter derived from the intestine-specific Pges-11 promoter, this low-level INT2 expression is not unexpected. However, we note that the expression level in INT1 is substantially higher than in INT2. Thus, although the expression patterns of the two TRAP drivers are not completely mutually exclusive, Pges-1ΔB still provides the most selective available tool for enriching the INT1 translatome in the context of the current study.

      Some parts of the study overemphasize the differences between the INT1 vs INT2-9 cell types, which is a biased representation of the results. For example, the authors specifically point out that 270 genes are differentially expressed in opposite directions in INT1 vs INT2-9 cell types during acute (30 min) fasting without mentioning the 1,268 genes that are differentially expressed in the same direction. They also do not mention here that 96% of the genes are differentially expressed in the same direction in INT1 and INT2-9 cell types after prolonged (180 min) fasting, suggesting that the divergent translational responses of these cell types are only observed in the first 30 minutes of food deprivation. Similar results have also been reported for the effect of fasting on locomotory and feeding behaviors, where 30 min of fasting produces more variable effects, which become more consistent after longer periods of fasting (PMID: 36083280). Hence, the effects of brief food deprivation should be interpreted with caution.

      The intestine functions as a discrete and cohesive organ, so the expected result is that there would be no differences across the different cell types. For us, the surprise was that, in fact, there are differences between these cells at all. However, the point is well taken, and we have added a statement in the text to reflect that many genes change similarly in INT1 and INT2-9, while the differences reflect important functional divergence between these cell types.

      Many of the interpretations of this study primarily rely on pathway enrichment analyses, which are based on the known function of genes. The function of uncharacterized genes that were found to be differentially expressed in INT1 vs INT2-9 cell types, e.g., the ShKT proteins, was not explored in this study. In addition, overreliance on pathway enrichment tools (instead of functional validation) has resulted in several conflicting findings. For example, one of the main messages of this study is that INT1 cells specialize in immune and stress response in response to fasting, which relies on pathway analysis in Figs 5E and 5F. However, pathway analysis at a different time point (shown in Figure S5A) indicates that INT2-9 cells show a much stronger increase in translation of stress and pathogen-responsive genes compared to INT1 cells. Hence, some of the results should be interpreted as different translational effects in INT1 vs INT2-9 cells after different lengths of food deprivation, without making broad claims about selective pathways being affected only in specific cell types.

      We agree that some interpretations in the manuscript relied heavily on pathway enrichment analyses and should be stated more cautiously. In particular, we agree that the current data are most consistent with state-dependent differences in translational responses between INT1 and INT2-9 cells across different durations of food deprivation, rather than with the strongest version of a claim that specific pathways are selectively engaged only in one intestinal subset. We also agree that uncharacterized genes, including the ShKT family, were not mechanistically explored in the present study and should be presented as important candidates for future investigation.

      The authors have compared their TRAP-seq results with genes enriched in the anterior and posterior intestine clusters from a previously published whole-animal adult scRNA dataset (PMID: 37352352). They claim that their TRAP-seq results are in agreement with the findings of the scRNA study. However, among the 10 genes from the 'posterior intestine' scRNA cluster in Fig.S1E, six are downregulated in the INT1 vs INT2-9 comparison, while four are upregulated. Hence, there is no clear agreement between the two studies in terms of the top enriched genes in the anterior vs posterior intestine, which should be considered for cross-study comparisons in the future.

      We have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. We instead assessed the expression levels of our INT1 up-regulated genes in their intestinal cluster and found that they have higher expression in the anterior intestine cluster (new Figure S1C). These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9.

      The authors describe in the manuscript that they have performed INT1-specific RNAi for two C-type lectin genes that are upregulated during fasting. Due to a recent expansion of C-type lectin genes in C. elegans, there is a high chance of off-target effects of RNAi that is designed for members of this gene family. More trustworthy results could have been obtained using CRISPR-based loss-of-function alleles for these genes, one of which is publicly available. Also, the authors do not provide any explanation for why knockdown of these stress-response genes, which are activated in INT1 cells in response to food deprivation, results in improved resistance to pathogens. This, in fact, suggests a role of INT1 cells in increasing pathogen susceptibility, and not pathogen resistance, during food deprivation.

      We agree that RNAi targeting C-type lectin family members may be susceptible to off-target effects, and that validation with CRISPR null alleles, where available, would strengthen these findings. In the current study, we used INT1-specific RNAi as a cell-specific first-pass approach to test candidate gene function. We also agree that the pathogen phenotype requires cautious interpretation. Specifically, the finding that knockdown of fasting-induced INT1 lectin genes improves pathogen resistance does not support a simple protective model for these genes. Instead, it suggests that INT1-expressed stress-response genes modulate host susceptibility or host-pathogen interactions.

      Many of the studies in this field (e.g., references 2-4 in this article) have investigated the effects of food deprivation ranging from 4 hr to 24 hr, which results in activation of starvation responses in C. elegans. In contrast, the authors have used shorter time periods of fasting (30 min and 180 min), and most of their follow-up experiments have used 30 min of food deprivation. Previous work has shown that the effects of food deprivation can either accumulate over time (i.e., the effect gets stronger with longer food deprivation) or can be transient (i.e., only observed briefly after removal of food and not observed during long-term food deprivation). Starvation-induced transcription factors such as DAF-16/FoxO and HLH-30 show strong translocation to the nucleus only after 30 min of fasting. Though gene expression changes in all stages of food deprivation are of biological relevance, the authors have missed the opportunity to explore whether increased INS-7 secretion from the anterior intestine is dependent on these starvation-induced transcription factors (which can be easily tested using loss-of-function alleles) or is due to other fast-acting regulatory mechanisms induced due to the absence of food contents in the gut lumen. A previous study (PMID: 40991693) has shown that DAF-16 activation during prolonged starvation shuts down insulin peptide secretion from the intestinal epithelial cells. Hence, it is not clear if increased INS-7 secretion is only a feature of short-term food deprivation or is also a signature of long-term starvation (e.g., at 8 hr or 16 hr timepoints). Since most of the INS-7 secretion data in this study are for 30 min of fasting, it remains unknown whether the discovered regulators of INS-7 secretion can be generalized for extended food deprivation that triggers major metabolic changes, such as fat loss (e.g., conditions shown in Figure 1D).

      We agree that short-term food deprivation and prolonged starvation likely engage distinct regulatory mechanisms, and that our study primarily addresses an early phase of food deprivation rather than the full spectrum of starvation responses described in prior work. We selected the 30 min fasting condition because our previous study showed that INS-7 secretion is induced within this interval and returns to baseline upon refeeding, even before detectable intestinal fat loss. We also included a 180 min fasting condition to capture a later state associated with metabolic changes. However, we agree that the present study does not determine whether the regulators of INS-7 secretion identified here also govern secretion during more prolonged starvation (for example, 8 hr or 16 hr), nor does it test whether starvation-responsive transcription factors such as DAF-16 or HLH-30 contribute to this regulation. We appreciate that determining how this response transitions during prolonged starvation will be an important direction for future work.

      Two previous studies (PMID: 18025456, 40991693) have shown a strong reduction in the expression of ins-7 in the anterior intestine using GFP-based reporters (both promoter fusions and endogenous CRISPR-generated) and in whole-animal RNA-seq data from starved animals. These results are in contrast to the increased INS-7 secretion from INT1 cells during fasting that is reported in this study. The authors here have reported that INS-7 translation is higher in INT1 compared to INT2-9 during fed, acute fasted, and chronic fasted conditions, but they have not shown whether INS-7 translation is upregulated during acute and chronic fasting in INT1 cells in their TRAP-seq analysis. Knowing whether increased INS-7 secretion during acute fasting is due to increased transcription, translation, or secretion of INS-7 is crucial to resolve the discrepancy between these studies.

      In our dataset, INS-7 translation in INT1 tended to increase during acute fasting relative to the fed state (log<sub>2</sub>FC = 0.69), although this effect did not reach statistical significance after adjustment for multiple comparisons. Consistent with this trend, our secretion assay showed that INS-7 release from INT1 increases during fasting. However, we agree that the current data do not distinguish whether this increase in secretion is driven by enhanced synthesis, regulated release of pre-existing peptide stores, or a combination of both.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to understand whether the discrete segments of the C.elegans intestine were specialized to carry out distinct functions during an animal's exposure and adaptation to a fast-changing nutrient environment. To achieve this, the authors used a method called Translating ribosome affinity purification (TRAP), which provides a snapshot of what genes are being translated into proteins (and therefore functionally prioritized by the animal) under different fasting and re-feeding conditions. By expressing the TRAP constructs in two distinct segments of the intestine (INT1) and (INT2-9), the authors were able to identify how these segments responded to changing nutrient availability.

      Already under steady state nutrient conditions, the authors found that INT1 and INT2-9 appeared to have different 'tasks', with INT1 expressing more immune- and stress-response related genes. Exposing animals to different regimens of starvation and refeeding also showed marked differences between the intestinal segments, and the gene expression patterns in INT1 were consistent with INT1 cells playing an integrative role in linking nutrient cues to the secretion of insulin molecules that regulate fat metabolism with food intake. In summary, the data presented catalogue, for the first time, gene expression differences between two areas of the intestine, suspected to play different roles, and through clever experiments, links these gene expression changes to responses to nutrient availability.

      Strengths:

      The data presented catalogue - for the first time and in a careful manner - gene expression differences between two areas of the intestine. They strongly support the presence of intriguing differences between two areas of the intestine in immune, metabolic, and stress-response regulation, and link these gene expression changes to the responses of these regions to nutrient availability.

      Weaknesses:

      The conclusions of this paper are mostly well-supported by data, but the relevance of the changing gene expression patterns could be better clarified and extended in the discussion.

      We thank the reviewer for this constructive comment. In the revised manuscript, we have now expanded the discussion to more clearly interpret these dynamic translatomic changes in the context of intestinal subset specialization. The most pronounced difference between INT1 and INT2-9 cells is the enrichment of stress-response genes in INT1. Based on the present findings, together with our previous work identifying INS-7 as an INT1-secreted signal (PMID: 39127676), we propose that INT1 cells are sentinel enteroendocrine cells that integrate information from the luminal environment and the metabolic state of intestinal cells.

      Reviewer #3 (Public review):

      Summary:

      In this study, Liu and colleagues utilize TRAP-seq to profile the repertoire of actively translated mRNAs in different intestinal cell types (anterior INT1 vs. posterior INT2-9 cells) in C. elegans. A key goal of this study was to identify transcripts differentially expressed/translated between these intestinal cell subtypes in the context of animals being well fed or subjected to acute (30 minutes) or chronic (3 hours) starvation, followed by refeeding.

      The authors identify a number of differentially expressed genes across all of the conditions tested. They then provide an initial survey of the landscape of translatome changes through Weighted Gene Network Correlation Analysis (WGNA), and some high-level functional surveys via Gene Ontology (GO) term analysis and protein domain analysis. The authors validate the enriched expression patterns of some of their identified candidate genes using fluorescent promoter fusion reporters, confirming INT1-specific expression. The authors further implicate the role of several other candidate genes in pathogen avoidance and in response to nutritional cues by knocking them down specifically in INT1 cells by RNAi. Finally, the authors identify pyruvate as a major nutrient signal coming from the bacterial diet that suppresses the release of a key insulin peptide (INS-7), and identify some of the genes expressed in INT1 that are required for this response.

      Strengths:

      (1) Good use of and justification for TRAP-seq, because scRNA-seq would be difficult under the varied conditions used (starvation, refeeding).

      (2) The manuscript is generally clear to read, and the data are generally well-presented with good supporting data that includes replicates, sample sizes, error measurements, and associated statistics.

      (3) The dataset will be an interesting resource to mine for future studies focusing on mechanisms of how particular intestinal cell types respond to different environmental signals.

      Weaknesses:

      (1) A limitation of TRAP-seq, although powerful, is that only relative comparisons can be made between genotypes/conditions to identify differentially-expressed genes, rather than assessing whether a given gene is expressed at a certain level in a cell type under a certain condition. This limitation is due to the non-specific association of sticky RNA species with the beads during the immunoprecipitation step. This is a minor point, however, and the authors do a nice job of focusing their analysis on differentially expressed transcripts in the current study.

      We agree that a limitation of TRAP-seq is that it is best suited for relative comparisons across cell types or conditions, rather than for determining the absolute expression level of a given transcript in a specific cell type. As the reviewer notes, this limitation arises in part from nonspecific recovery of background or sticky RNAs during the immunoprecipitation step, complicating the interpretation of absolute expression levels. For this reason, our analysis was designed to focus primarily on differentially enriched transcripts between INT1 and INT2-9 cells and across feeding states, rather than on assigning absolute expression levels to individual genes. We appreciate the reviewer’s recognition of this point. Our study uses TRAP-seq specifically to define relative translatomic differences between intestinal subsets and physiological states, which is well aligned with the strengths of this approach.

      (2) Another limitation of the current study is that the experiments testing the role of candidate genes identified by their profiling experiments do not delve a bit deeper into providing a mechanistic understanding of the phenotypes being studied. At present, the results are thus viewed more as a genomics-based screen with some limited follow-up on interesting hits. However, this reviewer appreciates that when placed in the context of the work presented, a presentation of the profiling data along with some validation is an excellent starting point for future mechanistic studies elaborating on these interesting candidates.

      We agree that the current study does not fully resolve the molecular mechanisms by which the candidate genes identified by TRAP-seq regulate the phenotypes examined here. Our primary goal was to generate a spatially resolved translatomic framework for INT1 and INT2-9 cells across feeding states, and to perform focused validation of selected candidates to establish the physiological relevance of the profiling results. We therefore view the current functional analyses as an initial validation and proof of principle, rather than a comprehensive mechanistic dissection of the molecular pathways for each candidate. We appreciate the reviewer’s recognition that these findings provide an excellent starting point for future studies.

      Appraisal of whether the authors achieved their aims, and whether the results support their conclusions:

      The main goal of the study was to survey the dynamic responses at the level of actively translated mRNAs of the INT1 vs INT2-9 cells in response to metabolic challenge.

      Overall, the authors use established methods to perform their genome-wide analysis, and the set of differentially regulated genes is enriched for expected molecular functions and forms coherent networks in anticipated pathways.

      The validation experiments (promoter::GFP fusion reporters, INT1-specific knockdowns of highly regulated genes) further corroborate the quality of the TRAP-seq datasets generated.

      I have a few points for the authors that would further strengthen this work:

      (1) The authors rightfully focus on the top differentially-regulated candidates, but it's unclear at present how far down their fold change list would lead to expression pattern validations. It would be useful to test a few more promoter::GFP fusion reporters at different enrichment/fold-change/statistical cutoffs.

      Testing additional promoter::mNeonGreen reporters across a wider range of fold-change and statistical thresholds could be somewhat useful for calibrating ranked TRAP-seq candidate genes. However, given the variation in strains bearing extrachromosomal arrays, we did not consider this a stringent enough test, given that the sensitivity and dynamic range of RNA-seq far outpaces genetic fluorescence-based reporters. For these reasons, we focused on the top differentially enriched candidates to provide not only an initial validation of the dataset, but also to determine whether these candidates regulate biological functions in INT1 cells and thus serve as potentially useful biological readouts in future efforts.

      (2) Although the INT1-specific RNAi provides a convenient strategy for rapidly perturbing and testing genes of interest for phenotypes, independently validating the knockdowns with genetic mutants, or alternatively (if genes are essential), degron alleles.

      We agree that validating the INT1-specific RNAi phenotypes with independent genetic approaches, including null-allele or degron-based alleles for essential genes, would further strengthen the conclusions. In the current study, we used INT1-specific RNAi as a rapid and spatially restricted strategy to functionally test candidates identified by TRAP-seq and to determine whether these genes contribute to the specialized physiological functions of INT1 cells. We consider these experiments an initial validation of candidate function rather than a complete genetic dissection, which could be conducted in future efforts to study other aspects of INT1 function.

      Impact:

      The TRAP-seq data and list of differentially-expressed candidate genes will form an interesting set of high-priority candidates to study for their role in the reception and transduction of nutritional cues in response to food status and pathogens. This data will thus benefit the C. elegans community of researchers studying the mechanisms governing these phenomena.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments:

      (1) The authors need to describe the fasting method used in detail. Was fasting performed on unseeded NGM plates or in liquid (M9 buffer)? Were the animals washed with buffer prior to starvation? If yes, how many times? These details are critical for any researcher to follow up on their results.

      We have clarified the fasting/refeeding procedure in the revised Methods section. Briefly, worms were washed off OP50-seeded NGM plates with M9 buffer, washed three times in M9 buffer, and then transferred to unseeded NGM plates for fasting. For refeeding, worms were collected from the unseeded NGM plates with M9 buffer and transferred back to OP50-seeded NGM plates.

      (2) The authors claim that "INT1 and INT2-9 cells maintain fundamentally different molecular identities independent of any and all acute or chronic conditions", which they primarily based on Principal Component Analysis (PCA). The circles shown in Figure 2A are arbitrary, and many such circles can be drawn in the 2D space to separate the samples in different ways. The authors should show this comparison in a translatome-wide similarity heatmap with hierarchical clustering (similar to Fig.2C, but with all the experimental conditions and their replicates on both x- and y-axes).

      The ellipses shown in Figure 2A were generated using the stat_ellipse() function in ggplot2, which calculates the mean and covariance of the PC1 and PC2 for each line and draws ellipses corresponding to the 95% confidence level. The separation between lines is primarily driven by PC2, which accounts for 14% of the variance in the translatomic dataset. Although this difference is less pronounced when considering the full translatome, samples from the same line nevertheless cluster together, supporting line-specific differences in translatomic profile.

      (3) Figure 4 of the study shows 18 Venn diagrams for genes that are differentially expressed between INT1 and INT2-9 cell types in fed, fasted, and refed conditions. In the absence of any statistical comparisons, it is difficult to interpret whether the extents of overlap (higher or lower than expected) are significant. Ideally, P values for hypergeometric tests should be provided for the overlap regions.

      We appreciate the reviewer’s suggestion. We explored the use of hypergeometric testing, implemented through the SuperExactTest package in R, to assess the statistical significance of the overlaps shown in the Venn diagrams. However, this analysis yielded significant P values for essentially all overlap regions, including cases in which the degree of overlap was not especially informative and did not align with the interpretation presented in the text. This outcome likely reflects the dependence of the test on the size of the input gene sets and background universe, which can make statistical significance difficult to interpret meaningfully in this context.

      (4) The claims made in lines 242-244 (Figure 5D) need to be supported by P values from hypergeometric tests.

      Similar to the previous point.

      (5) The interpretation of Figures 6E and 6F described in lines 296-297 needs to be supported by statistical analyses. The authors claim that the undulating pattern of expression of the turquoise module genes is stronger in INT1 compared to INT2-9. However, based on Figures 2C and 6F, it appears that the expression change is not necessarily weaker in INT2-9, but instead is different, i.e., the expression of turquoise module genes goes up during fasting in INT1 and goes down after refeeding, while their expression goes up during fasting and stays up after refeeding in INT2-9 cells.

      We appreciate the reviewer’s point and agree that the turquoise module shows dynamic regulation in both cell populations. The key difference is not the presence versus absence of an undulating pattern, but rather the magnitude of that change, which is greater in INT1. Because of the limited number of biological replicates in some conditions, particularly the fasting group, we interpreted these results cautiously and used a nonparametric approach to assess differences in average module expression between states. This analysis indicated that the turquoise module changes significantly in both lines, but with a larger effect size in INT1. We have included the corresponding statistical analysis and effect size in the revised manuscript.

      (6) Since the INS-7 coelomocyte uptake assay was used extensively in this study, some representative microscopy images should be included to complement the quantification.

      We have added a new Figure 6G showing representative images corresponding to the quantification presented in Figure 6H.

      (7) The authors claim that INT1-specific fmo-2 RNAi results in reduced basal INS-7 secretion, but they do not have the direct statistical comparison for this. Were experiments shown in Figures 6G and 6K done on the same day?

      In the original Figure 6K (now Figure 6L), the data are presented as the percentage of normalized INS-7::mCherry fluorescence intensity relative to fed animals treated with vector RNAi. A statistical comparison between fed animals treated with INT1-specific fmo-2 RNAi and fed vector RNAi controls was performed and was significant. We have also clarified that the experiments shown in the original Figures 6G and 6I (now Figures 6H and 6J) were performed on the same day.

      (8) It is not clear why blocking the mitochondrial breakdown of pyruvate (Figures 7E and 7F) does not mimic the fasted state in terms of increased INS-7 secretion from INT1 cells. Doesn't this contradict the proposed model in this study? Can the authors speculate why this is the case?

      We do not interpret inhibition of pyruvate dehydrogenase or pyruvate carboxylase as equivalent to the fasted state. Rather, our model is that fasting induces INS-7 secretion by lowering intracellular pyruvate in INT1 cells. Under this framework, blocking mitochondrial pyruvate breakdown would be expected to reduce pyruvate utilization and thus maintain intracellular pyruvate, preventing the drop in pyruvate that normally occurs during fasting. This would explain why these manipulations suppress fasting-induced INS-7 secretion. To directly examine this possibility, we performed the experiment in Figure 7G, which tests whether maintaining pyruvate levels in INT1 cells during fasting is sufficient to suppress INS-7 secretion. The results are consistent with this interpretation and further support a model in which decreased intracellular pyruvate is a key determinant of fasting-induced INS-7 secretion.

      Minor comments:

      (1) Figures 2E and 2G are very similar and represent the same result in two different ways (unbiased vs guided comparison). One of these should be moved to the supplementary figures.

      Although these figures show similar patterns, they were derived from two independent analytical approaches, WGCNA and differential expression analysis. We therefore interpret the concordance between these independent methods as strengthening the robustness of the association and increasing confidence in the biological relevance of the observed pattern.

      (2) In Figures S1C, S1D, and S1E, a more significant P-value is shown with a smaller circle, and a less significant P-value is shown with a larger circle. This is confusing to the reader and should be inverted.

      We appreciate the reviewer’s comment and have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes used in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. Our further examination of marker expression across the intestinal clusters in the Ghaddar et al. (PMID: 37352352). dataset confirmed this limitation. These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9. We also note that the spatial identities of the intestinal clusters in Ghaddar et al. (PMID: 37352352) were not clearly established in the text or by spatial transcriptomic evidence, making it difficult to assign the annotated anterior, middle, and posterior clusters to specific intestinal cells. We have revised the manuscript accordingly and replaced the original figure panels.

      (3) Line 131 mentions the comprehensive characterization of the translatomic differences between INT1 and INT2-9 cells under each acute and chronic condition. However, the paragraph only discusses the differences in the fed condition. This is confusing, and the authors should mention the comparison between these cell types under acute and chronic conditions in subsequent sections where it is described.

      In this paragraph, we indeed discuss the ‘fed’ condition, but in subsequent sections we follow with details analyses of acute versus chronic, as well as regional differences across the intestine. We have clarified this in the opening sentence of the referenced paragraph.

      (4) The Venn diagrams in Figure 4 look very similar, and it is hard to differentiate between how 4A is different from 4I, how 4B is different from 4J, etc. The authors should include the labels for 'acute' or 'chronic' above each Venn diagram to guide the reader through these panels.

      We have added labels indicating the acute and chronic conditions to the left side of each Venn diagram in the revised Figure 4.

      (5) It is not clear in the figure legends how Figure 5E is different from Figure S6A, and how Figure 5F is different from Figure S7A. This should be better described in the figure legends.

      We have added a sentence to better describe this in the figure legend.

      (6) Figure 6C: Survival parameters such as median lifespan, number of animals for each condition, etc., should be reported for the different conditions.

      We have revised Figure 6C to indicate the number of animals analyzed in each condition, and the median survival for each group is now reported in the corresponding figure legend.

      (7) The colors used for control RNAi and clec-160 RNAi are very similar in Fig.6C. Easily distinguishable colors should be used.

      We have changed the colors as suggested.

      (8) The INT1-specific RNAi strain should be first described in line 285.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (9) Line 304: 'REF' should be replaced with the reference.

      We have replaced the “REF” with the reference (PMID: 39127676)

      (10) The P value for statistical comparison between the fed and 30 min refed states should be shown in Figures 6G, 6I, and 6K.

      We have now included the p value for the comparison as suggested. Figures 6G, 6I, and 6K are now labeled as 6H, 6J, and 6L, respectively.

      (11) In Figure 7, the authors should consider replacing the 'refed' label with 'recovery' because the pyruvate treatment was done in the absence of 'feeding' (= bacteria consumption).

      We appreciate the reviewer’s point. However, we chose to retain the label “refed” in Figure 7 to maintain consistency across the set of conditions examined, including 2% glucose and OP50 supernatant, which likewise do not involve bacterial consumption despite not showing effect on the refeeding response of INS-7 secretion.

      (12) The full form of DISN should be mentioned in the figure legend of Figure 7.

      We have included the full form of D1SN in the figure legend of Figure 7A.

      (13) Line 367: 'normalization' should be replaced with 'return to basal levels'. 'Normalization of INS-7 secretion' might also mean normalization of INS-7::mCherry signal to CLM::GFP signal.

      We have revised the wording per the reviewer's suggestion.

      (14) The methods section has a quantitative RT-PCR section, but it is not clear if RT-PCR data are reported in any of the figures. Also, no qPCR primers are listed in Table S3.

      We have removed the quantitative RT-PCR part from the methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors describe that the RPL-22-3xHA constructs are not integrated, at the very end, in the section "Limitations of the data". An earlier mention of this caveat would have been useful. In addition, it would help if the authors could provide their defense (which I think is very valid) of using non-integrated strains in the results section, as they describe the experimental setup. Also, some details were missing, which left me wanting to know: Were there expression differences? How were they accounted for? Was expression normalized between these two constructs, and if so, how?

      The RPL-22-3xHA construct was integrated into each line to ensure more consistent transgene expression across experiments. Because the INT2-9 construct is expressed in a larger number of cells than the INT1 construct, the INT2–9 samples yielded greater amounts of pulled-down nascent RNA, as reflected in the supplemental table and in the higher aligned RNA counts observed for the INT2-9 samples. To account for these differences, differential expression analysis was performed using DESeq2, which corrects for library size by estimating sample-specific size factors with the median-of-ratios method. Raw counts are then normalized using these size factors, thereby accounting for differences in sequencing depth and minimizing confounding effects due to variation in library size. Such differences are common in RNA-seq experiments, particularly when comparing samples derived from distinct input populations.

      (2) The 'acute' and 'chronic' exposures are thought through and carefully defined. The question I do have is whether the 3-hour fasting can be considered chronic fasting, given how surprisingly fast the animals lose their fat content. Could these kinetics indicate that the 30-minute fasting is reflective of mechanisms during which senses change in food availability, whereas the 30 minutes represents acute fasting (with chronic fasting - meaning fasting, during which the animal activated alternative pathways - occurring later)? While this may appear to be pure semantics, it could influence how the authors interpret their results. One method to more objectively separate an 'acute' from a 'chronic' stage may be to conduct a time course of fat loss-does fat loss plateau after 3 hours? The timing when the rate of decrease levels off could be more indicative of the beginning of a chronic phase.

      We appreciate this important point and agree that it should be more clearly discussed. We interpret acute fasting as a pre-fat-loss state, since it is 30 minutes off food and no difference in fat levels are detectable at this stage (Fig 1B). The translatomic changes observed under acute fasting therefore likely reflect food-sensing mechanisms and early preparatory responses that promote subsequent fat mobilization. In contrast, chronic fasting (180 minutes off food – see Fig 1D) appears to represent a post-fat-loss state, in which fat stores have already been depleted, and the corresponding translatomic changes likely reflect the effects of sustained metabolic stress.

      (3) The age of the animals used has to be more explicitly stated. Were these animals egg-laying? Or L4/young adults? This is likely to impact the changes that the animals undergo.

      Day 1 young adults were subjected to the fasting. Great care was taken to ensure consistency across biological replicates.

      (4) What is the rationale, in the authors' view, that stress response genes are apparently more enriched than metabolic or mitochondrial enzymes, and membrane receptor changes? Are the latter mostly regulated by PTMs/localization changes, etc?

      Based on our current data, we cannot exclude the possibility that metabolic or mitochondrial enzymes, as well as membrane receptors, are regulated in INT1 and INT2–9 cells through mechanisms not captured at the translatome level, including post-translational modification or changes in subcellular localization under different fasting and refeeding conditions.

      (5) The refeeding experiment with latex beads and killed OP50 is very clever. Details on when INS-7 was evaluated in the caoelomocytes would help the reader better understand and interpret these results.

      INS-7mCherry signal was evaluated in the coelomocytes immediately after refeeding; we included this information in the methods section and referenced our previous paper.

      (6) In the Discussion, I was looking for a more detailed context for how to think about the differences and similarities in the RNA-seq data between the two segments, and perhaps a discussion of whether there were any indications that the two segments communicated with each other.

      The data show that the most pronounced difference between INT1 and the rest of the intestine at the RNAseq level, is the expression of stress response genes in INT1. Although there are some nuanced differences, the prevalence of stress response genes persists across feeding and fasting conditions. This difference, combined with the evidence that INT1 cells secrete the enteroendocrine peptide INS-7 (Fig 6 and PMID: 39127676) is strongly reminiscent of the mammalian enteroendocrine cells, which also secrete peptides and show strong expression of stress response genes (PMID: 37626258 and 27148273). We suggest that this category term reflects not only a canonical stress response, but also a broader response to shifts in the luminal environment, which INT1 cells are anatomically poised to detect well before the absorption of nutrients has begun further down the intestine (INT2-9). Thus, we believe INT1 cells are a newly defined enteroendocrine cell type within the C. elegans intestine.

      Regarding communication between INT1 and INT2-9 – this is an intriguing possibility that we have considered, given that peptide genes and receptors are found in the RNAseq datasets. The extent to which the expression of these genes leads to functional effects is the subject of future investigation.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1A - It would be better to also show single fluorescent protein channels to assess the specificity of the expression patterns. A schematic or labels of where the INT1 vs. INT2-9 boundaries are located would be helpful to non-experts.

      (2) Figure 4 - At present, the Venn Diagrams are a very complicated way to visualize all of the comparisons/conditions. I would recommend that the authors consider using UpSet plots to better summarize the relevant comparisons they would like to make. The same consideration applies to Figure 5D.

      (3) Line 301 - Description of the INT1-specific RNAi strategy. I think it would be better to bring this information earlier, close to line 285, where the authors first mention performing INT1-specific RNAi experiments.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (4) Line 304 - I think the authors meant to cite a reference where the REF placeholder text is found

      We have replaced the “REF” with the reference.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors of this study developed a method to quantify calvarial bone marrow from MRI head scans, enabling the study of its composition in large datasets of adults, usually collected to study the brain. Bone marrow intensity can be semi-quantitatively measured in T1-weighted MRI scans due to the greater signal intensity of fat than watery red marrow. This is an ingenious use of the MRI-produced information for other important phenotypes, such as bone structure and marrow content. Different head types were tested for complying with the model, which is notable.

      The model was also successfully validated using several publicly available MRI resources - real data - in (1) a dataset consisting of 30 individuals that were scanned 10 times each at 3-day intervals, and (2) the monozygotic (MZ) twin data from the Human Connectome Project cohort. Then the authors applied this validated method to head-MRI scans from the UK Biobank (n=33,042) to extract information on the spatial distribution of bone marrow adiposity (BMA) in the calvaria, allowing a GWAS to identify associated genes.

      The authors revealed high heritability and identified 41 genetic loci significantly associated with the BMA trait, including six sex-specific loci. Of note, statistics estimate that 99% of BMA trait-influencing variants are shared with BMD (497 of 500 variants), which may mean these results demonstrate the biological relevance to bone health. Some of the BMA genes were found related to the Wnt pathway, including WNT16, WNT4, NXN; this is a "positive control", since the Wnt/β-catenin signaling pathway was suggested as an important determinant of BMA. Also, associations in genes (BMP4, DLX5, LGR4, LRP4, SFRP4) that are known to specifically influence adiposity, are encouraging. Integrating mapped genes with bone marrow single-cell RNA-seq data revealed patterns of adipogenic lineage differentiation and lipid loading.

      With regards to the reviewer’s comment on the overlap between BMA and BMD trait-influencing variants, we would like to add that the correlation of effect sizes within the overlap is -0.95 as we would expect from bifurcating differentiation of mesenchymal stem cells into osteoblasts or adipocytes: the underlying common biology driving these traits results in shared traits with a negative correlation in the effects.

      The study also investigated the genetic overlap between BMA and twelve (or 13) "brain and body" traits and identified significant genetic correlations with BMI, cognitive ability, and Parkinson's disease.

      In sum, since MRI head scans present a hitherto unexplored opportunity to address unresolved aspects of bone marrow biology, this study is both timely and innovative.

      There are, however, some assumptions, findings, and their interpretation, which require more critical focus.

      Sex-specificity is well described and studied here. Men have higher BMA than women, but post-menopausal women catch up in the BMA values. The authors believe that calvarial marrow has a number of features that make it particularly well-suited to the study of BMA process - which is clinically important in other bone sites. It has a simple "sandwiched" structure that they are able to model. This is true only to some extent: a condition called "Hyperostosis frontalis interna", of unknown etiology (described by Smith & Hemphill in 1956) - is characterized by irregular overgrowth of the inner table of the frontal bone (symmetric/bilateral). Although not of clinical significance, typically benign, studies report a prevalence of 12%; However, it's most common in postmenopausal women - where prevalences up to 49% in women over the age of 65 - have been reported. Thus, sexual dimorphism is obvious and the effect of estrogen is likely shared with whichever bone - and marrow - age-related pathology. So, for women not using HRT, this new layer of the bone might interfere with the calvarial BMA readings and in turn, affect the BMA-related analyses.

      Thank you for bringing the "Hyperostosis frontalis interna" condition to our attention. It is particularly interesting to hear that the etiology is unknown and one may suspect that some kind of calvarial bone marrow dysregulation may be part of the cause. Our model for bone marrow location was trained on simulated data which included variation in the thickness of all anatomical layers (including the inner table), so it will be robust to some thickening of the inner table. It might not be robust to the most extreme cases of inner table thickening (as described in some case reports), but these are rare. Further, it should also be noted that the other calvarial bones, representing a much greater fraction of the calvarial surface, remain largely unaffected by the thickening and would therefore yield correct localisation of the bone marrow layers. In summary, although the severe cases of hyperostosis frontalis interna have the potential to affect our identification of the bone marrow layer, the low frequency of such cases and the restriction of the phenotype to the frontal bone means that the potential for bias is very limited.

      It would be interesting to develop a method for detection of thickened inner bone so that the condition’s prevalence can be quantified in a large sample like the UK Biobank and its genetic architecture be determined. This could help elucidate the etiology.

      The authors suspect that the effect of BMA on BMD may be biased in women; they should comment on those "with low BMD and high BMA" given that hyperostosis frontalis might be an issue. A strong effect of SNPs in the ESR1 chromosomal region might be akin to the above concern.

      Thank you for raising this point, which we have followed up with a new analysis.

      According to ICD-10 data in UK Biobank there are only N=105 individuals with an M85.2-diagnosed disorder. Given the total sample size of N=446,814 individuals with ICD-10 data, this would translate to a prevalence of 0.02%, which speaks for an underdiagnosis in this sample such that we cannot simply remove diagnosed individuals to control for a potential diagnostic confound.

      We have therefore taken a different approach to investigate this potential issue: As you elaborated in your previous comment, the prevalence of hyperostosis frontalis increases with age in females. The literature also suggests that prevalence rates do not differ between males and females in young age / prior to menopause. Therefore, we have studied the association between BMD and BMA for males and females separately, and in two age groups based on a median split of our sample (left plot: younger than 65, right plot: subjects older than 65). In these plots, the relatively large shift in female BMA and BMD is visible with the large yellow cloud at low BMA and high BMD in the left plot disappearing in the right plot. Despite this, we observe:

      (1) Associations in both males (blue) and females (yellow), suggesting that the associations were not driven only by females.

      (2) BMA-BMD association is largely similar across the two age groups.

      If we consider that the old age group is likely to contain more cases of hyperostosis frontalis than the young group, and if we consider that old-aged females are more likely to be in this condition than men of any age, then we would expect an impact of hyperostosis frontalis on our measures to result in observable differences in Author response image 1. This is not the case. We see global age-related shifts in BMA in women, yet the association with BMD remains similar across age groups.

      The technical properties of our neural network (trained on simulated data) makes it unlikely that frontal bone will contaminate the bone marrow detection globally (description above) and these results show that hyperostosis frontalis is not a considerable issue in our analysis.

      Author response image 1.

      Then, there is a perfect overlap of the BMA SNPs that are shared with BMD (497 of 500 variants), which may prove a "face validity" of the MRI-derived BMA. However, the BMD in the study was heel-derived eBMD - which is a good proxy for osteoporosis and is mostly driven by trabecular bone. Thus, there might be a concern that the BMA metrics capture some trabecular BMD.

      The reviewer is correct in pointing out that the BMA causal variants are a near-perfect subset of the BMD causal variants. The reviewer raises the concern that the BMA measurements may capture some trabecular BMD, however it should be noted that the correlation of effect sizes for the BMA/BMD overlapping causal SNPs is negative (-0.95). If our measure of BMA had been erroneously capturing trabecular BMD then we would expect to see a positive correlation of effect sizes for the BMA/BMD overlapping causal SNPs, not a negative one.

      Next, integrating mapped genes with existing bone marrow single-cell RNA-sequencing data revealed patterns of adipogenic lineage differentiation and lipid loading. The problem here is that the scRNAseq studies of the Bone Marrow niche are overwhelmingly mouse. The authors might wish to justify why they are relevant to humans (in the absence of the human-specific scRNAseq).

      We thank the reviewer for pointing this out. We noticed that, although Figure 4 and the Methods do explicitly state that the scRNAseq data is from mouse, it is not stated in the text of the Results. This is now corrected.

      The mouse is commonly used as the model organism for in vivo investigation of human phenotypes and bone marrow adiposity is no exception because, although mice have lower bone marrow adiposity than humans, the timing and sequence in bone marrow adiposity development are similar. BMA research makes extensive use of mouse models literature as exemplified by this review of research within the field (https://www.frontiersin.org/journals/endocrinology/articles/10.3389/fendo.2016.00127/full) and this article recent article (Koh et al. 2024. “Adult skull bone marrow is an expanding and resilient haematopoietic reservoir”. https://www.nature.com/articles/s41586-024-08163-9)

      We updated the results section (line 279):

      “Mesenchymal stem cells of the BM niche commit to either the adipogenic or the osteogenic lineage (Figure 4A) and both the number committing to the adipogenic lineage and their level of lipid-loading influences the total level of BMA. This aspect of BM biology is shared between humans and mice (29), so we made use of an existing mouse scRNAseq dataset of BM mesenchymal lineage cells (30) to study variation in the expression of BMA-associated genes as cells differentiate (Figure 4B).”

      For genetic correlation analysis, the authors selected 7 body and 6 brain traits. The latter traits reflect cognition (general cognitive ability and educational attainment) and brain-related disorders. This selection might seem arbitrary. The interpretation of genetic correlation with cognitive ability, education, and Parkinson's disease was attributed to the recently discovered vascular channels that link calvarial bone marrow to the meninges. This is a fascinating hypothesis, which requires functional proof. However, there might be simpler explanations. Thus, the diploe and the inner table of the calvarium are drained by the same veins as the dura. From the anatomy textbook, we know that diploic veins connect the pericranial and endocranial venous system through the skull.

      Whilst it is true that we did not systematically compare the results of the BMA GWAS to all potentially relevant brain and body phenotypes, we did use criteria to select the phenotypes we compared to. As stated in the manuscript (line 304):

      “We selected body traits (BMD, BMI, waist-to-hip ratio, systolic and diastolic blood pressure, type-2 diabetes, coronary artery disease) that have a logical connection to BMA given the mesenchymal stem cells origin of BM adipocytes and their role in bone, fat, and vasculature (29). For the brain, we selected traits reflecting cognition (general cognitive ability and educational attainment) and disorders that are prevalent in adulthood (insomnia, multiple sclerosis, Parkinson's disease, Alzheimer's disease) since it is primarily in adulthood that the adiposity of calvarial BM experiences a substantial change”

      We entirely agree that the suggestion that the genetic correlation between BMA and cerebral traits may be mediated by the vascular channels linking calvarial bone marrow to the meninges is merely a hypothesis. We have therefore updated the text of the Discussion (line 470):

      “We tentatively speculate that calvarial MALPs may be involved in sensing perivascular flows of CSF from the meninges to the BM and in influencing the BM’s hematopoietic response, and that this might be the basis of the observed genetic overlap between BMA and some cerebral traits. However, more conventional anatomical pathways may also be relevant, as the diploë and inner table communicate with meningeal and dural venous systems through diploic veins.”

      Reviewer #2 (Public review):

      Summary:

      This study develops a new artificial intelligence method for high-throughput analysis of skull bone marrow from MRI data, which may be useful for large-scale biological analyses. Using this method, the authors then attempt to estimate skull bone marrow adiposity (BMA) using T1-weighted signal intensity from MRI scans of ~33,000 people, followed by genome-wide association analysis; however, the approach is inadequate because T1-weighted signal intensity is not validated for measurement of bone marrow adiposity. If it could be validated, the study would be an important advance in understanding of bone marrow adiposity and skeletal biology.

      Strengths:

      This paper is well-written, and the figures are nicely presented. The neural network method used for analysing skull bone marrow is innovative, and the authors validate this through several approaches. Therefore, the authors have achieved the aim of developing a method for large-scale analysis of skull bone marrow from MRI data.

      The GWAS is reasonably well-powered and addresses potential ethnicity differences, with one GWAS done across white males and females, and a separate GWAS in non-white participants. The methodology also conforms to common GWAS standards, including for mapping genetic variants to candidate genes. Moreover, the study further investigates the biological roles of these genes by analysing their expression in single-cell RNA sequencing data.

      Weaknesses:

      The fundamental weakness is that T1-weighted MRI signal intensity (T1W) is used as an estimate of BMA, but it has never been validated for this. The authors show that this T1W parameter measures something that is heritable and can be compared between subjects, but they don't show that it actually measures (or even estimates) calvarial BMA. There is an attempt to do so by comparing the T1W parameter with data from quantitative T1 images: the authors show a reasonable correlation with some of the quantitative T1 image data. However, this still does not show that the parameter is measuring BMA; it could be measuring some other biological characteristic, but this remains unclear. So, there is a need to validate the T1W parameter against an established measure of BMA, such as the bone marrow fat-fraction or proton density fat fraction measured from multi-echo MRI analysis.

      Without validating this BMA measurement method, it is not possible to interpret the GWAS or other findings reported in the study.

      We reject this criticism.

      Although T1-weighted has not been validated as a quantitative measure of fat-fraction, there are several studies showing that it is a semi-quantitative measure of fat content (e.g. Loevner et al 2002, Shen et al 2013, Zhang et al 2020) and we also provide data that support this (figures S9-11).

      Semi-quantitative measures are used in many biomedical GWASes for instance even highly heritable neuropsychiatric disorders (such as schizophrenia and bipolar disorder) involve assessment by clinicians where the test-retest kappas are in the range 0.4-0.6.

      Further, we would suggest that the shortcoming of the imperfect correlation of T1w signal intensity with fat content is more than outweighed by our precision in identifying the calvarial BM cavity and the fact that the flat calvarial bone marrow has a wide range of adiposity in middle-aged and elderly individuals (compared to other bones). This lies at the root of why:

      We clearly recapitulate the known sex and age profiles, as well as the effect of HRT.

      We estimate high BMA heritabilities (43% in males and 23% in females)

      We find clear sex differences (which is a known feature of BMA biology)

      We identify a large number of the genes already known to affect BMA from earlier animal and cell work

      A noisy measurement of an entity with strong biological signal (a well-defined bone marrow cavity with variation in BMA across subjects) will often be more informative than a highly precise measurement of a poorly defined entity with little signal.

      A less critical weakness is that the GWAS has been done only on a single cohort, without replicating the findings in a follow-up cohort. For example, the authors could repeat their analysis on the remaining ~50,000 UK Biobank imaging participants for whom MRI data is now available. However, this would be pointless without knowing what biological characteristic(s) the T1W parameter is actually reflecting.

      We disagree with this comment. We separated the UKB data into discovery and replication sets prior to running the GWAS, so these datasets are independent:

      (1) Further, we ran the discovery (white british individuals) GWAS separately for males and females (prior to combining) and reported in the results section: “We found them to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      (2) We performed our replication GWAS in non-white British males and females. As noted in the results section: “Out of the 168 significant discovery SNPs, 62% replicated at P<.05, and 39% of the 41 lead SNPs replicated at P<.05 (Table S6). One locus replicated at genome-wide significance (P<5e-8). Furthermore, 92.7% of the lead SNPs of the discovery sample showed same effect direction in the replication sample (Table S5).”

      Reviewer #3 (Public review):

      Summary:

      This manuscript, "Estimating bone marrow adiposity from head MRI and identifying its genetic 2 architecture", brings together the groups of Drs. Kaufmann and Hughes in a tour de force work to develop an artificial neural network that localizes calvaria bone marrow in T1-weighted MRI head scans, with the goal of studying its composition in several large MRI datasets, and to model sex-dimorphic age trajectories, including the effect of menopause.

      Strengths:

      Bone marrow adiposity is a very active tissue with far-reaching implications for tissue crosstalk and human health than we had initially recognized. Although MRI has been used to measure BM, studies such as the one by these two groups are still lacking whereas very large datasets are analyzed using advanced AI machine learning tools coupled with genetic studies and a specific pathology. The groups had to develop new methods and new AI machine-learning tools for the imaging analyses.

      Weaknesses:

      Some aspects of the work that authors could add additional clarification.

      (1) Imaging Limitations: The authors provide an excellent overview and references supporting the use of MRI as a method for assessing marrow fat, particularly with some specific modifications. However, MRI images can be affected by various factors, including the presence of other tissues as well as specific MRI settings, which are much harder to precisely control when using different datasets.

      We thank the reviewer for his positive assessment of our review of methods.

      Regarding MRI settings: We agree with the reviewer that differences in scan protocols can create substantial differences in the resulting images between samples. Different tools exist for harmonization of imaging data across sites, but they usually operate on tabulated data and there is no one-size-fits-all approach yet [1]. Here, we took a different approach to prevent confounding bias: We generated a large set of simulated data for training of the neural network. The simulations circumvented potential issues emerging from confound biases in training sets that we might have seen had we had combined multiple samples with different scan protocols. Nevertheless, applied to real data the models may still face confound issues, such as better BMA estimates for some scan protocols over others. We have addressed these issues as follows: (1) Validation analyses (10 repeat scans of 30 individuals and twin pairs, figure 1d and 1e) are based fully on data that was acquired on the same scanner with the same protocol. (2) Analysis in UK Biobank included data from different scan sites albeit harmonized protocols. Here we accounted for scan site in all statistical models (including GWAS).

      (1) Dominik Kraft, Gloria Matte Bon, Édith Breton, Philipp Seidel, Tobias Kaufmann; Removing scanner effects with a multivariate latent approach: A RELIEF for the ABCD imaging data?. Imaging Neuroscience 2024; 2 1–7. doi: https://doi.org/10.1162/imag_a_00157

      Regarding the presence of other tissues: We recognise in the existing text of the results section that sometimes inner or outer table voxels are wrongly identified as bone marrow, but we show that this does not have a major impact on the correct identification of the bone marrow cavity. The existing text reads:

      “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap. However, this did not result in a corresponding fall in the ratio of the predicted intensity of BM to its true intensity, because the typical BM intensity was only marginally higher than the neighbouring bone intensity. This property of the typical relative intensities of these anatomic structures also explains why the intensity ratio at high overlap is not centred on 1: any misidentification of cortical bone as BM, will typically result in an underestimate of true BM intensity (Figure S2). The neural network performed well and intensity ratios were in the range 0.9-1.1 for the vast majority of head types (Figure S3).”

      Also note that we implement a number of QC measures to exclude scans where there is evidence that we may have failed to correctly identify the bone marrow cavity. The existing text reads:

      “We used two additional QC metrics to filter out calvaria where BM location was likely to have failed. First, we set an upper limit of 30 on the standard deviation of the intensity of the outer table as scans with higher values were clear outliers and were probably cases where the location of both outer table and BM has failed (Figure S6). Second, for each calvarium, we computed the Mahalonobis distance for all vertices in the two dimensions “first layer of the BM” and “BM intensity” (Figure S7). By manual inspection we found that data points with MD > 25 often had errors in BM layer identification, typically where the network had erroneously predicted a higher and more intense layer to be the BM. We considered a calvarium as failing this QC criterium if more than 0.5% of vertices have MD > 25. This criterium is very strict as errors on only 0.5% of data points in a calvarium would not significantly affect the average BM intensity for a calvarium.”

      (2) The specific density of cranial bones as it relates to the types of bone marrow: Cranial bones are extremely dense structures, which naturally interfere with MRI imaging. While it is thought that cranial bones have mostly "red bone marrow", this is only true for a short time in humans. How sensitive is their system in differentiating between red and yellow BM?

      We implemented several measures to ensure that our method would be robust to anatomical variation between individuals. As noted in the current version of the Methods section: “In order to train the neural network model, we generated a large synthetic dataset of intensity arrays, with known boundaries between anatomical structures, by simulating the thickness and intensity of the different structures located between the outer skin and the subarachnoid space. The simulation incorporated the following real-world complexities:

      Different anatomical architectures (skin, subcutaneous fat, aponeurosis, outer table, BM, inner table, dura mater, arachnoid space), including when a structure is not present throughout the calvarium

      Variation in thickness and intensity between vertices (on the same calvarium)

      A wide variety of different calvarium types with different combinations of levels of BM adiposity, bone thickness, and subcutaneous adiposity.”

      Further, as noted in the Results section:

      “We evaluated the performance of the neural network on simulated data using two metrics (Figure 1B): 1. the overlap between the predicted and the true BM location, and 2. the ratio between the predicted intensity of the BM and the true intensity of the BM”. The accuracy in localising the bone marrow layers was good, with the only exception being: “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap”

      We also validated our procedure on real data (see Results section, subsection “Procedure validation on real data and heritability estimate”). Briefly, we checked the accuracy of our method using a dataset from the Consortium for Reliability and Reproducibility, a twin dataset from the Human Connectome Project and by manually checking many hundreds of UKBiobank scans.

      We are thus confident that we accurately identify the bone marrow cavity irrespective of whether the bone marrow is red (low adiposity) or yellow (high adiposity).

      (3) Both items above are further complicated by aging, but aging is not a linear event as we have learned. There are specific bursts of aging in humans around the age of 45 and early 60s. How do the system and model predict or incorporate these peaks of aging? It seems from the data shown that aging is reflected more as a linear phenomenon. Is this because additional aging datasets are needed?

      We agree with the reviewer that ageing probably occurs in bursts rather than being a linear process. We do see a non-linear relationship between age and BMA in our data (see figure 2B), with a more rapid rise in BMA between the ages of 45 and 65, than later in life (in women). As a result of this, when we model BMA using regression, we use orthogonal polynomials of degree 2 which allows for a non-linear relationship. However, we cannot observe bursts of BMA increase in our data because it is cross-sectional. Longitudinal data would be required to obtain information on the nature and timing of any bursts in bone marrow adiposity.

      (4) The authors describe in richness of detail their AI learning programming and how it extracted the data from datasets. The authors also show some important correlations with specific genes, SNPs. What is not clear is how conditions such as anemia for example. An expected finding would be that patients with chronic anemia have lower bone marrow (BM) signal intensity on MRI scans than healthy people. This is because the signal intensity of BM depends on the fat-to-cell ratio in the tissue.

      We agree with the reviewer that conditions affecting the bone marrow niche have a potential to affect and be affected by bone marrow adiposity, with leukemia being a known example. This is why we believe that a method, such as the one we present here, has the potential to be useful in several biomedical fields (hematology and osteology).

      Furthermore, patients with a host of musculoskeletal disorders ranging from osteopenia to osteoporosis, sarcopenia, and osteosarcopenia will also have altered MRI scans. When using such large datasets how did the authors control or exclude these pathological conditions, or were all these conditions likely present?

      We did not exclude specific pathologies. We were careful to train our NN model on a wide variety of skull thicknesses, bone marrow adiposity, and subcutaneous adiposity and to evaluate the performance of the model on simulated and real datasets (see answer to your point 2).

      Reviewer 1 raised the issue of individuals displaying Hyperostosis frontalis interna (thickening of the inner table) and we recognize that in extreme case of this condition, where there is a major change in the anatomy of the calvarial bone, our method would probably not correctly localise the bone marrow. However, such extreme cases are rare and thus would not have a major impact on our results derived from over thirty thousand individuals. We demonstrate this with an extra analysis performed in response to the point about hyperostosis frontalis interna made by reviewer 1.

      (5) Some of the genes and SNPs although significant showed very small correlations. What is their likely physiological significance?

      Bone marrow adiposity is a polygenic trait and we have identified 41 statistically significant loci. We had a discovery sample of approximately 30k individuals which is modest for a GWAS study, so these 41 loci are a lower bound on the number of genes influencing the BMA trait. When a large number of genes influence a trait, the effect size of an individual gene is typically relatively small. However, the SNP heritability estimates of 31.5% indicates that we are able to explain approximately one third of the phenotypic variation with the effect sizes estimated by our GWAS: this is quite a high fraction relative to many other GWASs of biomedical traits.

      (6) The authors could use this excellent manuscript to expand their discussion to include the need for studies like theirs to be also complemented by multi-OMICS studies that will include proteomics and lipidomics of BM, bones, and muscles.

      We agree with the reviewer and hope that such studies will be undertaken in the future. We attempted to point in this direction in the last sentence of the Discussion (line 517): “Future studies can build on our developments to further validate the proposed measure of bone marrow composition and to study its effect on bone, blood, and brain”. Word count limits prevented us from further expanding on the specific kinds of studies that should be performed.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      More moderate concerns include:

      (1) In the "Genetic correlation and overlap" part, it is unclear why a high effect correlation with high SNP overlap is suggestive of vertical pleiotropy, while "moderate overlap and a low correlation of effect sizes ... are more indicative of horizontal pleiotropy". This is not intuitive.

      In vertical pleiotropy (genetics > phenotype A > phenotype B): genetics drive phenotype A and phenotype A drives phenotype B (B has few direct genetic drivers of its own). If we perform a GWAS of phenotype A and a GWAS of phenotype B, one would expect to see a high overlap in the causal variants (because phenotype B is largely indirectly determined by the genetics of phenotype A) and the correlation should be high (either negative or positive) because there is a cause-effect relationship between A and B (whether phenotype A has a positive or negative effect on phenotype B).

      In horizontal pleiotropy: the A and B phenotypes share the same genetic loci (but have no phenotypic influence on each other). In this case, we would observe high overlap in the associated loci, but one would not expect to see a high correlation in effect sizes (across loci) because there is no a priori reason to expect that genes associated with both phenotype A and phenotype B would have a consistent (negative or positive) effect across loci. For example, gene X may increase A and B, whereas gene Y may increase A but decrease B.

      We have made a small update to the relevant part of the Results section and otherwise rely on the explanation above (which will be publicly available along with the manuscript):

      Line 325: “These patterns of high overlap and high effect correlation within this overlap are suggestive of vertical pleiotropy i.e. a molecular mechanism influencing one trait, that in turn influences a second trait, such that most of the variants driving the first trait either have the same or the opposite direction of effect on the second trait.”

      (2) "possible causal effect of BMA on cognition" asks for a formal analysis, like Mendelian randomization.

      Given the current wording, the reviewer is justified in asking for a formal analysis. Since we did not perform this analysis, we have changed the wording:

      Line 434: “Since these are two highly correlated traits [36], a high overlap and correlation of genetic effects for both traits with BMA may be consistent with the hypothesis that BMA could have a causal effect on cognition (Table 1).”

      (3) ll. 132-134: please reword this sentence for clarity: "Using the network-estimated location of the BM within ... averaged these across all datapoints...". Please define threshold of desirable overlap between the predicted BM and the true BM (=0.7?).

      Background: The model predicts the BM localisation for a datapoint (which interval of layers of the 50 layers is bone marrow). We tested the model on a wide variety of simulated data and aim for the overlap to be as close to 1 as possible, but some error is inevitable. We found that the average overlap between the true and predicted bone marrow was only below 0.7 when the bone layer is only 4 mm thick (meaning that the bone marrow is only 1-2 mm thick). This demonstrates the high accuracy of our method in identifying a very small anatomical feature.

      When applying our method to real data, we do not know the truth and therefore cannot compute the overlap between the predicted and true value. It is therefore not possible to identify datapoints where the overlap is poor (e.g. lower than 0.7) and filter them out.

      Given the above, we struggle to understand in what way an overlap threshold is relevant to how we compute the signal intensity for a datapoint. Nevertheless, we recognize that the sentence pointed to by the reviewer is poorly formulated and have tried to make it clearer:

      Line 130-133: “To obtain the BM signal intensity for an individual datapoint of the calvarium, we used the network model to estimate the location of the BM within the datapoint’s intensity array and averaged these BM intensities to get the BM intensity for that datapoint. Then, we averaged these datapoint intensities across the calvarium to produce the global BMA measure for the scan.”

      In the GWAS Results, please clarify the phrases - what was "significantly genetically correlated (Rg=.94)" (also, l. 402, "genetic correlation between the sexes" - in what?).

      Genetic correlation is a statistical measure that quantifies the extent to which two traits (or the same trait in two different cohorts) are influenced by the same genetic factors. Simply put, it is the effect sizes of the SNPs in the two GWASs of interest that are correlated (after correcting for confounding effects, such as linkage desequilibrium). When comparing two GWASs, the standard formulation is to refer to their “genetic correlation”. We made a modification to the text to clarify this:

      Line 239: “We found the male and female GWASs to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      "a more than two-fold difference between the sexes" - in which metric?

      We feel that what is being compared is stated clearly in the original sentence:

      Line 262: “A comparison of the male and female effect sizes of the top lead SNPs of each locus revealed 6 loci in which there is a more than two-fold difference between the sexes (loci 10, 18, 26, 30, 32, 37 in Table S5)”.

      (4) Also In GWAS Results, a locus Dlx5 is called "SHFM" in the Supplementary Table.

      Background:

      We identified 41 genome-wide significant loci and named the locus after the gene closest to the top lead SNP (bold in Figure 3C). Other genes in each locus for which genome-wide significant SNPs were eQTLs, are listed below the closest gene in normal font (Figure 3C).

      In table S5, we report details of the top lead SNP for all 41 loci. We report only the nearest gene to the top lead SNP.

      For locus 14, SHFM1 is the closest gene to the top lead SNP whereas DLX5 and DLX6 are genes in the locus for which genome-wide significant SNPs were eQTLs. This explains why DLX5 appears under SHFM1 in Figure 3C, but does not appear in Table S5.

      (5) Please reword MRI jargon - "Dixon method", vertix - should be introduced, as well as abbreviation "KDE".

      We had recognised that the word “vertex” would be confusing and had replaced it by datapoint, but had unfortunately missed one occurrence in the text. This is now corrected.

      Thank you for pointing out the lack of introduction of the term “KDE”. This was only explained in the supplementary materials, but has now been added to the main text:

      Line 500-508: “To ensure between-subject comparability, we used the intensity normalised nu.mgz volume output by FreeSurfer. We validated this approach through comparison with well-established intensity normalization methods; Kernel Density Estimation (KDE), WhiteStripe (WS), Gaussian Mixture Model (GMM), Fuzzy C-Means (FCM), and Z-score normalization (ZS). We found the highest test-retest reliability with our approach (Figure S9), and, together with KDE (based on reference signal intensity in WM), the highest correlation with quantitative T1 relaxation maps (Figure S10).”

      Reviewer #2 (Recommendations for the authors):

      (1) This would be an extremely useful advance for the bone and BMA fields if only it could be confirmed that the T1W signal intensity is actually measuring BMA in some meaningful way. Or, even if not BMA, to confirm what other biological characteristic(s) it is in fact capturing. This is essential for interpreting the findings.

      (2) I note that you have compared the normalized T1W parameter with quantitative T1 data (e.g. Figure S10). However, this doesn't address the fundamental issue, because even these quantitative T1 data (e.g. from MP2RAGE) may not be measuring calvarial BMA. T1W sequences have been used to estimate BM cellularity (if not BMA directly) but are not nearly as precise as water-fat imaging. For example, one study found a reasonable correlation (0.71) between T1 relaxation times and BM fat (https://www.nature.com/articles/s41598-019-57030-5). So, if your normalized T1W parameter shows a correlation of -0.44 with the T1 MP2RAGE MRI signal (Figure S10), what does this mean in terms of how well your parameter reflects the actual BMA adiposity? We can't know this, because we also don't know if the T1 MP2RAGE signal reflects calvarial BMA.

      (3) I think my recommendations are clear from the public review. Ideally, you would be able to compare the skull BM normalized T1W parameter with PDFF data that have T2* correction (since the skull BM cavity is quite small and so may suffer from T2* effects relating to tissue inhomogeneity). But even if you had only dual-echo BMFF data, this would still be much more informative than relying only on T1 data. I hope this can be done so that the findings of the study can be properly interpreted.

      As explained above, we reject this reviewer’s claim that T1-weighted signal intensity cannot be used to perform a GWAS of BMA: other studies have shown that T1-weighted signal intensity is a semi-quantitative measure of fat fraction, we have performed extra analyses that confirm this, and our results further demonstrate this. For further detail on why we reject this criticism, see our response to this reviewer’s comments.

    1. balance may be shifting

      I don't think this is the "burn it to the ground" we've been looking for. It feels more like a paradigm swing back to early 19th C. higher learning. If most (read smaller and less well funded) universities become more like skills training factories, who gets the best education? If LLMs market knowledge as accessible (and transactional), whose knowledge is favoured and lauded? The dominant hegemony. You can go to university, but we don't have any programs to teach you how to think. You'll think what we tell you to think in the way we want you to think it.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

      Weaknesses:

      (1) The conceptual distinction between replication and translation could be more clearly emphasized.

      (2) Interpretation of results is dense and can be challenging to follow without a clear and summarized.

      (3) Some simulation parameters (effect sizes, heterogeneity, and number of animal studies) require more substantial justification.

      (4) Practical recommendations could be more explicit to guide applied researchers.

      We thank Reviewer 1 for the general positive assessment of our study and for the constructive feedback. We have addressed all of the four identified weaknesses in the revised manuscript. Specifically,

      (1) Conceptual distinction between replication and translation. We have reinforced this distinction at multiple points in the manuscript: in Section 2.7 (just before introducing the translation success metrics), in the Discussion, and in a new working definition of translation success added to the Introduction. We further explicitly acknowledge that statistical translation success, as defined here, is narrower than biological translation.

      (2) The dense result section. We have added a summary Table (Table 3) at the end of the Results section that compares all metrics on key properties (strengths and weaknesses, overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch). We also direct readers to this table early in Section 3.2, so that readers less interested in the technical details can obtain the key take-home messages without reading the full section.

      (3) Further justification of simulation parameters. We have substantially extended the rationale for our parameter choices in Section 2.4 and the Limitations section. We explain that our parameters are grounded in an empirical meta-analytic dataset, contextualise the large effect size and heterogeneity value against published benchmarks from preclinical research, and clarify that our main goal was to explore directional trends rather than absolute performance under specific values. We have also added an invitation for others to explore alternative parameter spaces using our openly available code.

      (4) Practical recommendations. We have extended the Recommendations section (pages 21–22) with more explicit scenario-specific guidance, supported by the new summary table.

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodate for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

      We thank Reviewer 2 for the thoughtful summary and for recognising the strengths of our study. We believe that both weaknesses were addressed in the revised version of our manuscript. Specifically,

      (1) Heterogeneity in animal studies. We agree that high heterogeneity in animal studies is an important limitation, and we address it directly in our simulation design by including a wide range of heterogeneity values (including very high levels, as observed in animal studies). Our results show clearly how heterogeneity affects the performance of each metric, and we highlight this in both the new summary Table (Table 3) and the Recommendations section which was extended. We also caution applied researchers to pay special attention to heterogeneity when interpreting translation results.

      (2) Focus on the statistical component of translation. We fully agree that statistical translation success is only one aspect of a broader decision-making process. We have elaborated on this in the revised manuscript, both in a new working definition of translation success in the Introduction (which explicitly distinguishes statistical from biological translation) and in a new paragraph in the Discussion section where we situate our metrics within translational decision-making frameworks such as PATH. They make it clear that progression to human trials depends on a suite of evidence of which statistical translation is only one part.

      Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Weaknesses:

      The weaknesses of the study are the limited focus on specific metrics, the assumptions, particularly in the limited number of human study variables, and the less-than-ideal approachable summary of findings for a non-technical audience.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

      We thank Reviewer 3 for the positive assessment and for the constructive suggestions.

      We have addressed the identified weaknesses as follows:

      (1) The assumptions around human study variables. We acknowledge these as inherent constraints of the simulation design. We have added a note in the Limitations section about the fixed human sample size (N = 107 per group), clarifying that while this value is grounded in a power analysis as per regulatory standards, it represents one particular scenario and may not generalise to all contexts. Further, we have contextualised and motivated the other simulation parameters better. We also invite readers to explore alternative conditions using our openly available code.

      (2) Approachability of the summary of findings for a non-technical audience. We have added a summary Table (Table 3) at the end of the Results section, comparing the metrics on key properties including overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch. We direct readers to this table early in Section 3.2 so that those less interested in the technical details can obtain the main take-home messages without reading the full section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) Conceptual framing: clearer distinction between replication vs translation

      The Introduction correctly points out the conceptual difference between replication and translation (animal to human), but this distinction needs to be reinforced repeatedly, especially when interpreting metric performance. For instance, several metrics (e.g., meta-analysis, replication BF) inherently assume exchangeability of findings, which is rarely justified in translation because species differ biologically.

      The manuscript should explicitly state why treating animal findings as "original studies" and human findings as "replications" can be misleading. Add a subsection in the Discussion: Why replication metrics behave differently in translation settings. This will help provide a more straightforward interpretation of the results beyond the numerical findings.

      Thank you for your feedback. While the purpose of our study is to assess the applicability of the replication success metrics in the translation context, we agree that the reader should be reminded that these two concepts differ and metrics’ assumptions might not always hold. We have reiterated the difference between replication and translation in Section 2.7, just before we introduce the translation success metrics (see end of page 7). We also reiterate it in the Discussion (see end of page 20). Here, we emphasise that while some of the metrics assume that both studies investigate the same effect, this is unlikely to be the case in translation, leading to some of the metric’s assumptions being violated which impacts the performance of the metrics.

      (2) Stronger justification of the simulation parameters is needed

      The simulation factors are comprehensively presented (Table 1), but certain choices appear arbitrary or oversimplified.

      - Effect sizes: The three levels (0, −4.44, −24.37) are derived from the motivating dataset, but the paper should explain that these represent extremely large effects in many biomedical contexts.

      - Heterogeneity values: τ<sup>2</sup> = 291.1 is enormous; adding context about real-world heterogeneity distributions would help.

      - Number of animal studies (k): Using only 2-5 studies may not reflect reality; many preclinical fields have >30 studies before clinical translation.

      These choices should be more explicitly defended in Section 4 (Limitations), beyond the brief mention already there. Provide a sensitivity analysis, or explain why extrapolation beyond this parameter space is reasonable.

      We agree that the choice of the parameter values might sometimes appear arbitrary. However, instead of arbitrarily choosing parameter values, we base our choice on data from a meta-analysis. This particular meta-analysis might not be representative of all of pre-clinical and clinical research, but because Terstappen included both animal and human studies investigating the same research question it was particularly well suited. They further used an outcome (maternal blood pressure) that is comparable between rats and humans, which is quite rare. We have specified this further in Section 2.4 (Motivating dataset, page 5). In the Limitations section, we acknowledge any possibly unrealistic simulation conditions again, and emphasize that our main goal was to explore trends in the metrics’ behavior as the conditions changed rather than their absolute performance under specific values. Further, the effect sizes (0, −4.44, −24.37 mmHg) span a meaningful range on the unstandardized mean difference scale for blood pressure measurements: from no effect to a modest but clinically relevant reduction to a large effect typical of animal studies. The large heterogeneity value corresponds to a relative heterogeneity of I^2 of 95.33% in the animal meta-analysis. While this appears high, it is frequently observed in preclinical research: Hooijmans et al (2022) showed that 55% of animal study meta-analyses using mean differences as effect size measure have I^2>75%. We also added a footnote reiterating the fact that such high effect sizes (on the raw mean difference scale) are indeed common in animal studies (on page 6). Regarding k, we acknowledge that pooling only 2 to 5 animal studies may not reflect common practice. However, the directional trends in type 1 error and power are clearly visible in our Figures. Larger k decreases the type 1 error of the animal studies, while the power is increased unless there is high heterogeneity between animal studies and there is only a small effect. Extending the range further is unlikely to change the conclusions. Moreover, in practice, the decision to advance to human trials considers evidence well beyond the statistical considerations we simulate. All of the above is now emphasized more explicitly in both the methods, where we have substantially extended the reasoning for choosing the simulation conditions, and the limitations section. Finally, we added an invitation to others to use our open material (i.e., code) and explore the behaviour of the metrics under other conditions (see top of page 21).

      (3) Decision criteria (strict/lenient/no criterion) need a clearer rationale

      The three continuation rules are a strength of the study, but:

      - The lenient criterion (any negative estimate is considered "beneficial") is unrealistic and should be reframed.

      - The strict criterion (p < 0.025) heavily inflates effect sizes (in Figure 1b) and may distort interpretation.

      It would be helpful to provide a table showing, for each criterion, its real-world analogue (e.g., regulatory requirement, exploratory progression, mechanistic plausibility).

      We have followed your suggestion and added a Table (Table 2) with the description of the criterion and a description of its real-world analogue. No criterion represents an important reference scenario used to evaluate metric behaviour independent of progression decisions. The strict criterion is the closest to regulatory-style evidence. It is also highly selective and therefore might induce biases (e.g., inflated effect sizes). We link lenient to an exploratory decision-making where efficacy evidence is considered in addition to other factors (e.g., safety), but not intended to represent a certain regulatory standard.

      (4) Interpretation of simulation results needs more focus

      The Results section is extremely detailed, making it challenging to identify the central take-home messages. The authors should consider adding a concise summary table comparing metrics on key properties:

      - T1E control robustness.

      - Sensitivity to heterogeneity.

      - Dependence on animal sample size.

      - Dependence on k.

      - Bias under asymmetric effects.

      Moving some nested-loop plot descriptions to the Supplement. Right now, descriptions are technically correct but cognitively heavy.

      We agree with your comment and have attempted to implement it in our summary Table 3, at the end of the results section. We also point readers early on to the Table, so that they can skip the more technical and detailed description if they want (see first paragraph section 3.2, page 12). After some trial and error, we agreed that the chosen columns are the most useful for an applied researcher to get a quick overview. Our table now summarises for each metric its main strengths and weaknesses, its behaviour with increasing heterogeneity, its sensitivity to more animal data (i.e., larger k and larger animal sample size), and its behaviour under effect mismatch (i.e., when the true effect in the animal and human study are dissimilar).

      (5) The discussion should provide explicit recommendations.

      The authors provide high-level recommendations, but the recommendations lack specific guidance. When heterogeneity is low, controlled sceptical p-value works well. When effect sizes differ: weighted Edgington is stable. The authors should avoid using replication BF when the animal effect ≠ human effect. Meta-analysis should not be used when human heterogeneity is high, because of inflated T1E.

      We agree that explicit recommendations would be helpful to the applied researcher. As mentioned in the reply to the previous comment, we have added a summary table which lists the strengths and weaknesses of each metric. We also extended the paragraph in the Recommendations section (on page 21 and 22) to give some examples of scenarios in which certain metrics would be recommended over others.

      (6) Recommendations for applied researchers

      The study is missing an explicit definition of "translation success". The manuscript implicitly defines translation success as: "Both animal and human results show a beneficial treatment effect according to metric X". But this is different from biological translation, which concerns underlying mechanisms. The authors briefly mention this conceptual challenge, but this should be elaborated, as it is central to interpretation.

      Thank you for this comment. We agree that “translation success” was not explicitly defined. We have now added a working definition in the Introduction, clarifying that, in this paper, translation success is defined statistically, and depends on the metric. We now explicitly acknowledge that this is a narrower definition than biological translation. We also elaborate on this distinction in the Discussion where we note that the appropriate metric and interpretation of translation success depends on the translation goal and that statistical translation is distinct from biological translation.

      Minor points:

      (1) The abstract could include a direct sentence on the main conclusion. For example, no metric was uniformly optimal; controlled sceptical p-value and weighted Edgington performed most consistently.

      Our abstract already included main conclusions. We added the word “However” to emphasize the sentence “no metric was uniformly optimal” a bit more.

      (2) The figures are informative, but nested loop plots are very dense. Consider providing a guided example in the figure caption explaining how to read them (as partially done in Figure 1a, but repeat for all).

      We agree that the Figures can be very overwhelming at first. We did not want to add specific helping elements as we did in Figure 1 to not make the figures even busier. The goal was to introduce the reader gently to the nested loop plots via Figure 1 before having them look at the remaining figures. We hope that with the added summary Table and the more detailed recommendations, applied researchers less interested in the statistical details will still find the information most relevant for them easily.

      (3) Methods: Section 2.4 could clearly state that effect sizes are in units of mmHg (blood pressure) from the dataset.

      Thank you for pointing this out. This has been added.

      (4) Results: This section is long; consider adding a brief summary paragraph at the end of 3.2.

      We added a summary table, allowing interested readers to skip the long section entirely.

      (5) Limitations: Add a note about publication bias in animal studies (you mention it in the Introduction, but not in Limitations). Add a statement about effect direction consistency (i.e., animal effect negative but human positive), which is not explored in the simulation grid.

      Thank you for pointing out this inconsistency. A note about publication bias in animal studies was added to the Limitations section (that this was not investigated). A note about opposite animal and human effects was added to Section 2.5 (Simulation conditions) under “Animal and human effect sizes”.

      Reviewer #2 (Recommendations for the authors):

      Animal studies are typically highly controlled, using animal models that are either outbred to provide higher genetic variability or inbred with very little genetic variability and with a specific phenotype. Additionally, many rodent models are incomplete models of the overall human phenotype and are typically used to investigate only one aspect of the condition/disease. Some of the rat animal models that the Terstappen et al. (2020) systematic review used as the simulation parameters for the study included outbred (Sprague-Dawley, Wistar) and inbred Spontaneous Hypertensive Rats (SHR), which have different mechanisms in which hypertensive onset can occur, especially if inducing preeclampsia in outbred animals. Is it feasible to reduce heterogeneity in the animal results if only outbred or only SHR are considered instead? I realize this may reduce the sample size even further.

      You raise an important point differentiating biological (rather than statistical) translation. We have added a sentence about differences between rat models and humans to the new paragraph in the Limitations section (bottom page 20 and top page 21) on the distinction between biological and statistical translation. As for reducing heterogeneity in the animal results by focusing on one type of rats, we agree focusing on one type of rats might reduce heterogeneity. We however consider this reduction to be very small (because the results of the study with SHR are actually comparable to the results with Wistar and SD rats). Therefore, rerunning the simulation would not yield results that differ in any meaningful way from those already reported and the substantial computational effort required to do so is not warranted.

      Reviewer #3 (Recommendations for the authors):

      Overall, I found this a very detailed study. However, my recommendation is to provide a more approachable overview of the results to reach a wider audience. Currently, the article is much more technical and statistically focused. I think two additions could help.

      (1) A summary table of each of the metrics and their strengths and weaknesses under the various conditions (e.g., animal and human study characteristics). Currently, this is done via text, but I think a high-level summary via a table could be a compelling way to make the simulations more approachable for a non-technical audience.

      As requested also by reviewer 1, we have added a summary table.

      (2) Contextualize the findings within the decision-making process a little more. The authors have a well-written limitations section that acknowledges this; however, I think the discussion (and maybe the introduction) could be enriched by putting the simulation findings into context. For example, this paper suggests a framework that includes replication as part of the decision-making process for human trials (https://www.cell.com/med/fulltext/S2666-6340(24)00296-4).

      We agree that situating our metrics within existing translational decision-making frameworks adds important context. We have added a paragraph in the Discussion (before the Limitations section on page 21) clarifying that the metrics evaluated here should not be viewed as standalone decision rules for progression from animal studies to human trials. Several frameworks have recently emerged precisely to guide such decisions in a more structured, multidimensional way. We refer to PATH and also to the GALENOS approach [DOI: 10.1186/s12874-026-02891-4]. Within such frameworks, translation success metrics of the kind evaluated here may provide a quantitative assessment of the consistency between animal and human efficacy findings, thereby informing one component of a broader translational evidence assessment. We have also briefly mentioned at the end of the Introduction (page 4) that frameworks for structuring the use of preclinical evidence in translational decisions are being developed, further motivating the need for quantitative tools such as those evaluated here.

      Below are some additional minor comments for the authors to consider:

      (1) In the abstract (4th line), there is an extra 'l' in failure.

      Thank you for the detailed review. We have fixed this.

      (2) I think since the study is completed, the objectives in the introduction should be past tense, not future.

      We have fixed this.

      (3) The limitations section should include the fixed human sample size. N=107 per group is grounded in the literature, but this varies widely based on the effect size of interest. Again, not material to the point of translation under simulated conditions (of which this would have increased the simulations well above the 648 already included), but given the impact this has on insights, this limits this investigation to a degree and should be acknowledged.

      Thank you for your comment. We have added a note about the human sample size to the paragraph about the simulation conditions in the Limitations section. The human sample size was computed via power analysis as per regulations, but we realize this could change depending on the effect size.

      (4) I appreciate how shrinkage was calculated. Though it is worth noting that the Reproducibility Project: Cancer Biology found much higher rates, which are similar to reports from biotech and pharma (e.g., 11% and 20-25% for Amgen and Bayer).

      We already mentioned the high rates of shrinkage in the Replication Project Cancer Biology (see page 10). We have now also emphasised that one could adapt these levels further depending on the situation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two higher frequencies within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state-of-the-art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts are in line with recent findings and theories. In sum, this paper makes a very nice contribution. I still have a few comments that I outline below, regarding the data visualization, some methodological aspects, and a couple of theoretical points.

      The authors put a lot of effort into the figure design. For instance, I really like Figure 1, which conveys a lot of information in a nice way. Figures 3 and 4, however, seem over engineered, and it takes a lot of time to distill the contents from them. The fact that they have a supplementary figure explaining the composition of these figures already indicates that the authors realized this is not particularly intuitive. First of all, the ordering of the conditions is not really intuitive. Second, the indication of significance through saturation does not really work; I have a hard time discerning the more and less saturated colors. And finally, the white dots do not really help either. I don't fully understand why they are placed where they are placed (e.g., in Figure 3). My suggestion would be to get rid of one of the factors (I think the voxel selection threshold could go: the authors could run with one of the stricter ones, and the rest could go into the supplement?) and then turn this into a few line plots. That would be so much easier to digest.

      We thank the reviewer for their insightful comments. Below we will address each point separately and highlight the changes made to the manuscript. In agreement with the reviewer we have recompiled Figures 4 and 5 (previously Figures 3 and 4). The new figures only present results for the 10% voxel selection threshold (with 5% and 25% moved to Supplementary Figures, see Figures S1-S9). Instead of the radially arranged layout, we opted for a more traditional figure layout, which significantly improved readability.

      (2) The division between high- and low-frequency alpha in the feature-specific signal correspondence is very interesting. I am wondering whether there is an opposite effect in the feature-unspecific signal correspondence. Would the high-frequency alpha show less of a feature-unspecific correlation with the BOLD?

      Following the reviewer’s interesting suggestion, we added the low/high frequency alpha analysis to the feature-unspecific analysis. Indeed, we have found a significant interaction between the sign of the signal change for selected voxel (positive vs. negative BOLD) and alpha sub-band (low vs high frequency alpha). An analysis of simple effects did not reveal any significant effects, however we found a trend level difference (p=0.097) between low and high-frequency alpha for the positive voxel sub-selection. This indicates a stronger negative relationship between upper alpha and the positive BOLD signal as compared to lower alpha. We interpret this result as partial evidence for a feature-related contribution of the upper alpha band. “Active” cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. The negative relationship between alpha and negative BOLD could thus be interpreted as an indirect effect, resulting from a reduced alpha decrease in cortical patches responding to non-attended receptive field locations. However, the involvement of attention-related processes remains speculative, since attention was not explicitly manipulated as part of the experiment.

      We have added Figure 4 B.

      We have also added this section to the Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub><0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      And the following section of the Discussion was extended:

      “The significant interaction between the sign of the BOLD signal deflection and upper or lower α bands (see Figure 4 B) further indicates that multiple α-related processes contribute differentially to positive or negative BOLD. "Active" cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. This hypothesis receives additional support from the trend-level difference in α sub-bands for positive BOLD, indicating that lower α is less related to the active, possibly feature-related processes. The absence of this difference for negative BOLD again indicates a broader, more general process. Future experiments manipulating visual features and attention might reveal a differential upper and lower α response to attended visual features and a more general relationship between α (and possibly superficial layer cortical activity) for suppressed (unattended) receptive fields.”

      (3) In the discussion (line 330 onwards), the authors mention that low-frequency alpha is predominantly related to superficial layers, referencing Figure 4A. I have a hard time appreciating this pattern there. Can the authors provide some more information on where to look?

      We thank the reviewer for pointing out the lack of clarity of this section in the Discussion. We have now rephrased the Discussion, focusing more on the laminar difference and keeping the frequency difference to a separate paragraph. Our main argument for possibly multiple alpha-related processes are twofold: a difference in alpha frequency depending on the underlying analysis (low vs high frequency alpha) and a different layer distribution (superficial layers vs. superficial and deep layers, depending on the analysis). The respective section in the Discussion now focuses on the laminar difference only. We find a negative relationship between alpha and the BOLD signal most prominently in superficial layers (feature-unspecific contrast for the BOLD signal with negative t-values; Figure 4). In addition, we find a superficial and deep layer contribution for the feature-specific contrast (congruent - incongruent; Figure 5A). While the here presented experiment was set out to investigate feature-specific processes, the meaning of the feature-unspecific results are of speculative nature. Future experiments should target the laminar difference between feature-specific and unspecific processes with respect to alpha frequency and layer distribution directly. 

      We have modified the respective sections in the Discussion:

      “Furthermore, we observed that the relationship between the feature-specific BOLD signal and α is predominantly linked to frequencies above 11 Hz (see Figure 5A). An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies. Since individual frequency variations have been included as a random slope in the linear mixed-effects model, these effects cannot be explained by a subset of participants driving lower or upper α separately. Specifically our findings on upper α indicate that α is not exclusively linked to global signal modulations, which has been the traditional perspective [...]”

      “Not only did we find a dissociation in the frequency domain between the relationship of α and the BOLD signal, but furthermore found that the laminar activation patterns provide further evidence for potentially multiple α-related processes. The association between α and the BOLD signal was strongest in superficial layers for negative BOLD activity and feature-specific activity (see Figure 4A and 5A). However, deep layer-related α effects were limited to feature-specific processes only (see Figure 5 A Co-Inco). These findings suggest that superficial layer α reflects are broader, more general process, while deep layer α operates more narrowly, linked to the processing of the visual features themselves. Previous findings using laminar fMRI (which did not include the investigation of oscillatory activity), indicate that superficial layer activity might be more related to the modulation of attention [...]”

      (4) How did the authors deal with the signal-to-noise ratio (SNR) across layers, where the presence of larger drain veins typically increases BOLD (and thereby SNR) in superficial layers? This may explain the pattern of feature-unspecific effects in the alpha (Figure 3). Can the authors perform some type of SNR estimate (e.g., split-half reliability of voxel activations or similar) across layers to check whether SNR plays a role in this general pattern?

      We agree with the reviewer that the vascular draining effect typically leads to increased signal change in superficial layers, the effect on (t)SNR however might be less straightforward. We did not include any counteracting measures, because we were not interested in the amplitude of the signal change, but now include an estimate of tSNR (See Figure S10 in Supplementary Figures). We found that in fact the signal-to-noise ratio is higher in deep layers. Most importantly however, the tSNR layer profiles we identified do not reflect the correlation layer result patterns of the combined EEG-fMRI analysis. This indicates that our results are most likely not the result of tSNR differences. In order to confirm our tSNR pattern we have also conducted a second layer analysis based on the LAYNII toolbox, which assigns voxels between pial and white matter to distinct layers (as compared to our fraction-based approach) and found a similar profile as with our initial analysis. However, absolute tSNR values were found to be higher for our weighted layer analysis. We speculate that while functionally relevant components of the BOLD signal drain towards superficial layers, physiological noise components will drain towards superficial layers as well.

      It is furthermore worth pointing out that for the contrast (congruent - incongruent), the vascular draining effect would cancel out between the conditions. Our findings on superficial and deep layers for those contrasts can hence not be explained by vascular draining at all.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      (5) The GLM used for modelling the fMRI data included lots of regressors, and the scanning was intermittent. How much data was available in the end for sensibly estimating the baseline? This was not really clear to me from the methods (or I might have missed it). This seems relevant here, as the sign of the beta estimates plays a major role in interpreting the results here.

      This is a very important remark and we would like to apologise for the confusion. It was not clear in the manuscript that the GLM was computed on z-transformed fMRI data. We have not specifically collected any “baseline volumes”. A positive beta value would indicate that the sign of the predictor matches the sign of the BOLD signal deflection (and vice versa).

      We have added or modified the following sections in Results and Methods respectively:

      “Before the GLM was computed, the fMRI data was z-transformed across time, separately for each block and voxel.”

      “A general linear model (GLM) has been computed with predictors for each TF bin separately for all voxels in V1 that later have been sub-selected according to the respective condition. Time courses for each voxel have been z-transformed before the GLM was computed for each voxel and experimental block separately. Afterwards, each of the resulting regression coefficients (β values) were multiplied with the voxel-specific layer weights that have been obtained as described above.”

      (6) Some recent research suggests that gamma activity, much in contrast to the prevailing view of the mechanism for feedforward information propagation, relates to the feedback process (e.g., Vinck et al., 2025, TiCS). This view kind of fits with the localization of gamma to the deep layer here?

      (7) Another recent review (Stecher et al., 2025, TiNS) discusses feature-specific codes in visual alpha rhythms quite a bit, and it might be worth discussing how your results align with the results reported there.

      We would like to thank the reviewer for pointing out these papers. Yes, we believe that those could be very related to the effects reported here. At the time of writing the initial manuscript we were not aware of the mentioned publications. 

      We have now included these papers in the Discussion:

      “Recent publications on the information exchange within and between primary visual cortex areas of macaques also reported deep layer γ band activity depending on the stimulus material (Gieselmann et al., 2022; Ferro et al., 2021). Those publications challenge the feed-forward exclusivity of γ altogether by revealing intra-area feedback communication in V1 from layer 5 to layer 6 and layer 6 to supra-granular layers. Possibly, the relationship between γ and deep layer BOLD we observed is also related to similar processes (Vinck et al., 2025).”

      “Similarly, in a recent opinion article, Stecher et al. (2025) promote the idea of "content-aware" α-oscillations. In agreement with our results, the authors argue that α-oscillations are related to content-specific feedback signals, reflected in increased decoding performance based on α power of top-down related processes, even prior to the onset of the stimulus (Hetenyi et al., 2025).. Accordingly, we interpret the lower α effect [...]”

      Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence, but most seem more like a suggestive trend. Nonetheless, the paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings. Further investigation will be important to strengthen and expand upon these insights.

      One of the main strengths of the paper lies in the use of a well-established and straightforward experimental paradigm (the visual oddball task). As a result, the behavioural effects reported were largely expected and reassuring to see replicated. The acquisition technique used is very novel, and while this may introduce challenges for data analysis, the authors appear to have addressed these appropriately.

      Later findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. However, the layer weight calculation is lacking in the manuscript. While it is discussed in the methods, it would help to briefly explain in the results how these weights are calculated, so that the reader can better follow what is being interpreted.

      Line 104 states there is one virtual channel per hemisphere for low and high frequencies. It may be helpful to include the number of channels (n=4) in the results section, as specified in the methods. Also, this raises the question of whether a single virtual channel (i.e., voxel) provides sufficient information for reproducibility.

      We thank the reviewer for encouraging us to clarify the virtual channel selection and we agree that the current description could be misleading. Indeed, we selected 4 virtual channels in total: 1 for each frequency band (alpha/gamma), for each hemisphere separately. The main goal of this selection was to find the clearest response of that frequency band to the task. Previous publications used a supervised (ICA-based) approach to extract those responses. To increase reproducibility, we have chosen an unsupervised beamformer-based approach. The reconstruction of time or frequency-resolved sources in the brain typically yields spatially highly correlated results. Publications focusing on this type of analyses report a spatial extent of typically multiple centimetres, which here is the case as well (see Figure 3A of the updated manuscript). As such, the single voxel selection boils down to selecting the peak response within a large patch of very similarly responding voxels. Using this approach we were able to select the frequency response with the highest possible SNR. We do not however claim that the respective single voxel is exclusively carrying this information. In addition we have added a short explanation to the Discussion, since we believe that virtual channel selection with a different objective (e.g. maximising the difference between conditions or maximising cross-frequency coupling, etc.) could indeed profoundly impact the EEG-fMRI correlation, which would open up opportunities for interesting analyses that are however beyond the scope of this project.

      We have added the following section to the Discussion:

      “Future work might also vary the exact virtual channel selection for obtaining EEG-based regressors. Here, we focused on the grid points (voxel locations) with the strongest α or γ response for each frequency band in each hemisphere, derived from the average frequency response to maximise SNR. However, selecting the respective virtual channels based on the response to specific stimulus features or the interaction between high and low frequency bands are possibilities worth exploring in future work.”

      One area that would benefit from further clarification is the interpretation of gamma oscillations. The evidence for gamma involvement in the observed effects appears somewhat limited. For example, no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 2). Significant effects emerged only when the analysis was restricted to positively responding voxels, and even then, only for the contrast between EEG-coherent and EEG-incoherent conditions in the feature-specific BOLD response. It remains unclear how to interpret this selective emergence of gamma-related effects. Given previous literature linking gamma to feedforward processing, one might expect more robust involvement in broader, feature-unspecific contrasts. The current discussion presents the gamma-related findings with some confidence, and the manuscript would benefit from a more nuanced reflection on why these effects may not have appeared more broadly. The explanation provided in line 230, that restricting the analysis to positively responding voxels may have increased the SNR, is reasonable, but it may not fully account for the absence of gamma effects in V1's feature-unspecific response. Including the actual beta values from Figure 4 in the legend or main text would also help readers better assess the strength and specificity of the reported effects.

      We agree with the reviewer that the missing gamma-band response for the feature-unspecific signal, as well as the limitation of the effect solely to the feature-specific contrast for positive voxel selections only was unexpected. In fact, based on previous literature, we were expecting a feature-unspecific effect in the gamma band as well. However, the literature on laminar level EEG-fMRI is sparse and previous experiments used tasks that did not allow for the separation into distinct features (here left or right-oriented gratings). While we cannot fully explain the absence of the gamma effect for the feature-unspecific condition, we reasoned that our stimuli evoked weaker gamma band responses compared to previous literature. 

      The fact that we only see a significant gamma band response for the contrast for positive voxel selections can be interpreted twofold: First, previous experiments limit their analyses to positive BOLD responses only, for which we find an effect as well. Second, the fact that a significant effect could only be obtained for the contrast, might indicate that gamma band activity is related to the actual features themselves. A cortical column responding to left-oriented gratings would then be related to a gamma band response linked to that orientation. If this response to a single orientation could not be fully captured due to SNR-related issues, we would not see this effect in the congruent-only condition and also not in the feature-unspecific condition (because this boils down to both congruent conditions combined). If gamma-band oscillations are actually reflecting the response of a column to a certain orientation, then the lowest possible response would be found for the exact orthogonal orientation (here the incongruent condition). The contrast between most preferred and most not-preferred orientation might have helped to overcome the inherently low SNR, explaining the results for the contrast.

      Lastly, we did not include actual beta values in the main text, because those might be misleading. We compute the relationship between EEG power and the BOLD signal for every voxel separately, then weighted the result with the respective layer weight and lastly aggregated across voxels.This means that the beta values express the strength of the association between EEG and fMRI for an average voxel. For this reason the values are tiny and the values themselves are less meaningful than “typical” beta values.

      We have added or modified the following sections in the Discussion or Methods respectively:

      “Based on previous literature, we expected a γ band effect for the congruent condition of the feature-specific analysis (Scheeringa et al., 2016), which we did not observe. A possible explanation could be the used stimulus material in our experiment as compared to Scheeringa et al., (2016). Muthukumaraswamy et al., (2013) found that stationary gratings evoke a weaker γ band response as compared to moving annular stimuli that have been used by Scheeringa and colleagues. If γ is related to the processing of the actual features themselves (e.g. to a column preferably responding to left-oriented gratings), then contrasting congruent and incongruent voxel selections provides the largest possible contrast-to-noise ratio (CNR). In turn annular stimuli as previously used might have activated all possible orientations and thus might have greatly boosted γ SNR.”

      “The described procedure of computing a GLM based on z-transformed data using z-transformed predictors yields β-coefficients that reflect the average relationship of a single voxel's BOLD response for a given layer (fraction of the single voxel's β) with EEG power changes of a specified frequency.”

      Relating to behavioural findings for underlying neural activity, could the authors test on a trial-by-trial basis how behavioural performance relates to the BOLD signal / oscillatory activity change? Line 305 states that "Since behavioural performance in the present study was consistently high at 94% on average and participants were instructed to respond quickly to potential oddball stimuli, a higher alpha frequency might reflect a more successful stimulus encoding and hence faster and more accurate behavioural performance." Also, this might help to relate the findings to the lower vs upper alpha functionality difference.

      This is a very interesting suggestion. We now include an exploratory analysis of the relationship between frequency and behavioural performance in the Supplementary Figures (see Figure S12). We did not perform a correlation between behavioural performance and alpha over trials because of the low numbers of oddball trials (N=40) and very limited number of false responses (94% response accuracy on average). However, we computed a correlation across participants. After averaging the alpha time-frequency spectrum across non-oddball trials, the individual alpha frequency was determined by the frequency where the alpha decrease (between 0.1 and 0.8 s post-stimulus) was largest. The correlation between alpha frequency and either reaction times and d’ (as a measure for accuracy), yields a significantly positive relationship between d’ and alpha frequency. This indicates that alpha frequency is related to task performance. We interpret those exploratory findings such that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power. 

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      In Figure 4, the EEG alpha specificity plot shows relatively large error bars, and there is visible overlap between the lower and upper alpha in both congruent and incongruent conditions. While upper alpha shows a positive slope across conditions and lower alpha remains flat, the interaction appears to be driven by the change from congruent to incongruent in upper alpha. It is worth clarifying whether the simple effects (e.g., lower vs upper within each condition) were tested, given the visual similarity at the incongruent condition. Overall, the significant interaction (p < 0.001, FDR-corrected) is consistent with diverging trends, but a breakdown of simple effects would help interpret the result more clearly. Was there a significant difference between lower and upper alpha in congruent or incongruent conditions?

      We thank the reviewer for this important remark and have added a simple effects analysis (see Figures 4 b and 5 b, e). We found that the main driver for the interaction between congruence condition and alpha frequency is upper alpha. Specifically the negative relationship between upper alpha and the BOLD signal is significantly stronger for the congruent condition and weaker for the incongruent condition. This indicates the upper alpha indeed is related to the processing of visual features.

      We have added a simple effects analysis (See Figures 4 and 5).

      We have added or modified the following in Results, Discussion and Methods respectively:

      In Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub> < 0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      “After correcting for multiple comparisons, we found a significant interaction (p<sub>FDR</sub> < 0.001). This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis. We found a significantly stronger negative relationship of upper α and the BOLD signal for congruent selections (p<sub>FDR</sub> < 0.01) and the reverse for the incongruent condition (p<sub>FDR</sub> < 0.01), as well as a significantly stronger negative relationship within the upper α sub-band for congruent over incongruent voxel selections (p<sub>FDR</sub> < 0.01).”

      “This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis, which revealed a significantly stronger negative relationship of upper α and the BOLD signal for congruent over incongruent selections (p<sub>FDR</sub> < 0.001).”

      In Discussion:

      “An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies.”

      In Methods:

      “Significant interactions were decomposed into simple effects using Wald tests on the model coefficients, ensuring that post-hoc comparisons were derived from the same statistical global variance as the primary interaction.”

      Overall, this study provides a valuable contribution to the literature on oscillatory dynamics and laminar fMRI, though some interpretations would benefit from further clarification or qualification.

      Reviewer #3 (Public review):

      Summary:

      Clausner et al. investigate the relationship between cortical oscillations in the alpha and gamma bands and the feature-specific and feature-unspecific BOLD signals across cortical layers. Using a well-designed stimulus and GLM, they show a method by which different BOLD signals can be differentiated and investigated alongside multiple cortical oscillatory frequencies. In addition to the previously reported positive relationship between gamma and BOLD signals in superficial layers, they show a relationship between gamma and feature-specific BOLD in the deeper layers. Alpha-band power is shown to have a negative relationship with the negative BOLD response for both feature-specific and feature-unspecific contrasts. When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency, and can therefore be interpreted as more feature-specific than lower frequency alpha.

      Strengths:

      The use of interleaved EEG-fMRI has provided a rich dataset that can be used to evaluate the relationship of cortical layer BOLD signals with multiple EEG frequencies. The EEG data were of sufficient quality to see the modulation of both alpha-band and gamma-band oscillations in the group mean VE-channel TFS. The good EEG data quality is backed up with a highly technical analysis pipeline that ultimately enables the interpretation of the cortical layer relationship of the BOLD signal with a range of frequencies in the alpha and gamma bands. The stimulus design allowed for the generation of multiple contrasts for the BOLD signal and the alpha/gamma oscillations in the GLM analysis. Feature-specific and unspecific BOLD contrasts are used with congruently or incongruently selected EEG power regressors to delineate between local and global alpha modulations. A transparent approach is used for the selection of voxels contributing to the final layer profiles, for which statistical analysis is comprehensive but uses an alternative statistical test, which I have not seen in previous layer-fMRI literature.

      A significant negative relationship between alpha-band power and the BOLD signal was seen in congruently (EEGco) selected voxels (predominantly in superficial layers) and in feature-contrast (EEGco-inco) selected (superficial and deep layers). When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency than lower frequency alpha. This is interpreted as a frequency dissociation in the alpha-BOLD relationship, with upper frequency alpha being feature-specific and lower frequency alpha corresponding to general modulation. These results are a valuable addition to the current literature and improve our current understanding of the role of cortical alpha oscillations.

      There is not much work in the literature on the relationship between alpha power and the negative BOLD response (NBR), so the data provided here are particularly valuable. The negative relationship between the NBR and alpha power shown here suggests that there is a reduction in alpha power, linked to locally reduced BOLD activity, which is in line with the previously hypothesized inhibitory nature of alpha.

      Weaknesses:

      It is not entirely clear how the draining vein effect seen in GE-BOLD layer-fMRI data has been accounted for in the analysis. For the contrast of congruent-incongruent, it is assumed that the underlying draining effect will be the same for both conditions, and so should be cancelled out. However, for the other contrasts, it is unclear how the final layer profiles aren't confounded by the bias in BOLD signal towards the superficial layers. Many of the profiles in Figure 3 and Figure 4A show an increased negative correlation between alpha power and the BOLD signal towards the superficial layers.

      We thank the reviewer for this important remark. Reviewer 1 raised a similar concern and I would like to refer you to our response to Reviewer 1, point 4. The veinal draining typically results in a higher signal change closer to the cortical surface. We did not take any measures to counteract this effect, but provide an analysis of tSNR in Supplementary Figures (see Figure S10). Possibly due to the drainage of physiological noise towards the surface, we found the highest tSNR in deep, followed by middle and superficial layers. To verify those results we computed the same analysis using a second layering algorithm, which resulted in the same profile, but overall less tSNR. Crucially the tSNR profile is not reflected in our EEG-fMRI results.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      When investigating if high alpha (8-10 Hz) and low alpha (11-13 Hz) are two different sources of alpha, it would be beneficial to show if this effect is only seen at the group level or can be seen in any single subjects. Inter-subject variability in peak alpha power could result in some subjects having a single low alpha peak and some a single high alpha peak rather than two peaks from different sources.

      We agree with the reviewer that a bias in a subset of participants to generally higher or lower alpha frequencies could potentially skew the presented results. While the initially computed model included a random intercept for the frequencies, we have now added the random slope as well. This ensures that the difference between low and high frequency alpha is indeed only driven by the difference in condition and not the result of individual differences across conditions themselves.

      In order to verify that not a small subset of participants is driving the result pattern, we also computed the fraction of participants that either show the dual alpha pattern (i.e. follow the exact pattern of the group average), contribute to the group average with a single peak or contradict the pattern entirely. Thereby, 40.4% of all participants show a dual alpha pattern, 38.4% a single alpha pattern in the direction of the group average and 21.2% contradict the group average. See Author response image 1:

      Author response image 1.

      Alpha Response Patterns with Example Subjects: V1 Feature Specific Contrast

      We would also like to highlight our added exploratory analysis of the relationship between alpha frequency and behavioural performance, which was requested by Reviewer 2, point 3. We find a significant positive correlation between alpha frequency and task performance on a group level. This indicates that higher alpha frequencies might be related to better discrimination of visual features. We speculate that participants with better task performance are capable of modulating their upper alpha more than participants with worse performance.

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      The figure layout used to present the main findings throughout is an innovative way to present so much information, but it is difficult to decipher the main findings described in the text. The readability would be improved if the example (Appendix 0 - Figure 1) in the supplementary material is included as a second panel inside Figure 3, or, if this is not possible, the example (Appendix 0 - Figure 1) should be clearly referred to in the figure caption. 

      Since Reviewer 1 suggested using an entirely different figure layout, we now opted to remove some information from the main text figures (we only show the 10% threshold, but 5% and 25% is in Supplementary Figures) and chose a more common figure layout. See Figures 4 and 5.

      Recommendations for authors:

      Reviewer #2 (Recommendations for the authors):

      The contrasts used in the analysis are not clearly introduced in the main text. While the methods section explains them more thoroughly, some of this explanation would be better placed in the results section, where the contrasts are first used. Specifically, the concepts of "feature-specific" vs. "feature-unspecific" BOLD signals are introduced with a very brief definition, which could be confusing for readers. The same applies to the terms EEG co and EEG inco; it would help to briefly explain these when they are first mentioned in the results. The supplementary figures and legends are helpful, so it is clear that the authors were prioritising clarity overall.

      The respective analyses are now also explained in the Results section:

      “During each trial either a left or a right-oriented grating was presented, from which two types of analyses have been derived: feature-unspecific BOLD activation (i.e. the response to any stimulus orientation), and feature-specific BOLD activation (i.e. the response to a specific stimulus orientation or the contrast between them). Thereby, fMRI data and EEG-based regressors could either be combined congruently (Co) by combining the BOLD signal of orientation-selective voxels with EEG-based regressors built from the same orientation trials, or incongruently (Inco), by combining the orientation-specific BOLD signal with EEG-based regressors built from the other orientation trials. Finally, those two congruency conditions have been contrasted (Co-Inco).”

      Figures are overall clear and illustrative of the results. For Figure 4, however, the use of dotted elements makes it somewhat harder to interpret what's being shown. While the supplementary figure clarifies the findings, rephrasing the figure legend to explain what the dotted lines represent would be helpful.

      Figures 4 and 5 have been replaced with a new layout and legends have been improved.

      The reported ranges overlap (e.g., alpha: 2-32 Hz; gamma: 20-120 Hz). It would be helpful to explain why such overlapping bands were chosen.

      Both frequency bands of interest differ slightly in their later time-frequency analysis (i.e. number of tapers and filter type). The overlap itself is not meaningful per se and results from the selection of a wide band for each respective sub-band. This wide selection was chosen to avoid filter artefacts. For the alpha sub-band, we also wanted to ensure that the beta spectrum is covered which also includes the alpha harmonic and for the gamma band that the full range of high-frequency activity is captured (e.g. EMG activity).

      Only a single time point was used for baseline correction of the low alpha band. Is this typical? The authors note that due to the gradient artefact arising in the pre-stimulus period, the baseline correction is somewhat difficult, although further clarification would be useful here.

      Relatedly, was pilot scanning conducted? If so, was the presence of strong gradient artefacts unexpected? More details about this would strengthen the methodological transparency.

      Indeed only a single time bin was used as the baseline for the alpha sub-band. After the piloting phase a slight adjustment to the final fMRI sequence has been made which was not expected to introduce gradient artefacts so close to the onset of the stimulus. Unexpectedly, those artefacts were visible until 300 ms before the onset of the stimulus. Similarly, a pre-stimulus alpha was observed (starting 250 ms before the onset of the stimulus), which we also aimed to exclude from the baseline period. In the end only the time bin centered at 300 ms prior to stimulus onset was chosen. However, this time bin contains 400 ms of data (the width of the window for the time frequency analysis). Thus, the term time point was misleading, because the actual time window that made up the baseline is 500 ms to 100 ms prior to the onset of the stimulus. 

      We have adjusted our wording in Methods to make this more clear:

      “For this reason, the low frequency baseline period comprised only a single 400 ms time bin centred around -0.3 s, because a pre-stimulus α decrease was expected starting around 0.25 s prior to stimulus onset.”

      Including a one-sentence explanation of the AROS test in the main text for clarity. As line 796 in the methods: "Each significant cluster has been further processed by means of an auto-regressive rank order similarity (aros) test (Clausner and Gentili, 2022). The fundamental idea behind the AROS test is whether group averages (i.e. averages of the signal of cortical layer in the present case), can be ranked such that the rank order is explained significantly better by the data than it would if the average data could not be meaningfully sorted (i.e. is shuffled)."

      An explanation has been added to the Results section:

      “Each significant cluster was then averaged along the frequency dimension at the widest point to enable an auto-regressive rank order similarity (aros) test Clausner & Gentili (2022), testing the laminar activation profile. The aros test transforms the layer averages into a rank order and tests - using a permutation procedure - if the rank order of the layer averages explains the data better than a random rank order (shuffled layer labels) would.”

      Line 223: "In fact, an analysis of the relationship between the EEG signal and the BOLD signal that focused on the feature contrast only (L - R; independent of the comparison to baseline) revealed a trend-level result with an even stronger deep layer contribution as compared to superficial layers." Could you point to which figure represents this finding - Figure 4B?

      This refers to Figure 4A in the old manuscript, for the 25% threshold for the gamma band. Since now the new figures do not include the 25% threshold anymore, it refers to Figure S4i.

      The number of participants is missing from the main text. Including this in the results section would improve clarity.

      The description of our sample has been moved from Methods to Results.

      Given the complexity of the data acquisition and analysis, the well-designed and easy-to-follow analysis pipeline figure (currently in the supplement) would be better placed in the main text.

      The mentioned Figure has been moved to the main text (now Figure 2).

      Also, simply out of curiosity, what do the authors think about the theta blob around 200ms post-stimulus?

      The theta blob most likely reflects the post-stimulus ERP as often observed in response to visual stimuli. We hypothesise that it is stronger in the middle and superficial layers, but we did not want to extend too much the scope of this paper. Additional analyses could be performed in the future on this evoked activity.

      Reviewer #3 (Recommendations for the authors):

      (1) Minor Corrections to the text and figures:

      We would like to thank the reviewer for the very valuable recommendations. Below we shortly describe how each suggestion has been implemented.

      We have made the white box more clear (see Figure 3 B).

      (b) Page 10: Top of 2nd paragraph - 'The full experimental protocol comprised a high resolution anatomical T1 scan lasting for 8 min'. The methods state this scan is 6 min 31 sec.

      The confusion results from the fact that the T1 scan was recorded during a short practice block that the participants performed inside the scanner. This block lasted 8min during which the 6 min 31 sec T1 scan was recorded. We have made this more clear:

      “Once prepared, the participant was placed inside the scanner and performed an 8 min practice block. A T1-weighted scan was acquired during this time in the sagittal orientation using a 3D MPRAGE sequence Brant-Zawadzki et al., (1992) with the following parameters: TR/TI = 2.2/1.1 s, 11° flip angle, FOV 256 x 256 x 180 mm and an 0.8 mm isotropic resolution. Parallel imaging (iPAT = 2) was used to accelerate the acquisition, resulting in an acquisition time of 6 min and 31s.”

      (c) Page 10: 'Stimulus presentation' paragraph - 'Stimuli were projected onto a screen behind the subject's head using'. The use of 'subject' should be replaced with 'participant' throughout.

      We have corrected the phrasing.

      (d) Page 14: Figures 2A and 2B are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (e) Figure 5 caption: 'Regressors are build for each time-frequency bin separately.' should be 'built'

      We have corrected the mistake.

      (f) Page 16, final paragraph: 'Afterwards, each of the resulting regression coefficients (B coefficients) was multiplied with the voxel specific layer weights that have been obtained as described above.' Should be 'were multiplied'

      We have corrected the mistake.

      (g) Page 17: 'Subsequently, separate analyses were done for two frequency of interest (FOI) ranges centerd around' - typo

      We have corrected the mistake.

      (h) Page 17 - 'Within these frequency ranges inferential statistics based a cluster level' - missing word. Should be 'based on a cluster level'

      We have corrected the mistake.

      (i) Page 14 Figure 2B and 5D are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (2) fMRI data pre-processing:

      Please provide a comment on the EEG-fMRI data quality - e.g. tSNR of EPI data. Perhaps example EPI data could be shown in the supplementary information.

      We included the below Figure S11 in Supplementary Figures showing an example EPI. We have also included an illustration of the result of our layering approach. Furthermore, we included a tSNR analysis (see Response to Reviewer 1, point 4).

      On a practical note - with 14-minute long runs whilst wearing an EEG cap, I would expect participant motion to be a concern. Could you provide some metrics on perhaps the average of the mean and maximum per subject displacement/rotation?

      We ensured that participants receive tactile feedback for their respective head motion from a strip of tape span across their foreheads. This resulted in overall manageable motion during each experimental block. During the main experiment, the average framewise displacement was 0.3 mm, with an average total translation of 1.6 mm and an average total rotation of 1.6 deg within each block. 

      We have added Figure S13 to Supplementary Figures.

      We have added a section to Methods:

      “Subject motion per block was low, with a mean (SD) frame-wise displacement Power et al. (2012) of 0.34 mm (0.24 mm) for the main experiment and 0.23 mm (0.22 mm) for the retinotopy (see also Figure S13 in Supplementary Figures).”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation - thereby impairing β-catenin degradation - the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Weaknesses:

      (1) A key caveat is that the identified tankyrase degrader also targets GSPT1 for degradation. This raises the possibility that GSPT1 degradation may contribute to the observed β-catenin and tumor growth inhibition.

      (2) The authors address this concern reasonably by showing that DLD1 cells resistant to GSPT1 degradation remain sensitive to the tankyrase degraded.

      (3) To further strengthen this point, the authors might consider generating TNKS1/2 double knockout cells (e.g., in DLD1 or SW480 backgrounds) and demonstrating that the degrader loses its growth-inhibitory effect in these models. However, given the technical challenges of creating double knockouts in cancer cell lines, such experiments could be considered optional.

      We thank the Reviewer for the favorable feedback. The major concern is the collateral degradation of GSPT1. As the Reviewer noted, IWR1-POMA was able to suppress colony formation in DLD-1 cells resistant to a GSPT1/2 degrader (DLD-1R, Figure 6B and S9F), suggesting that TNKS but not GSPT degradation is responsible for growth inhibition.

      We also appreciate that the Reviewer brought it to our attention an important early observation of the TNKS scaffolding effects. Cong reported in 2009 that overexpression of TNKS induced AXIN puncta formation in a SAM but not PARP domain-dependent manner (PMID: 19759537, Ref. 12). We have added this reference to the introduction of TNKS scaffolding in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Weaknesses:

      Although the results are promising and mostly compelling, the claim that the PROTACs provide "a deeper suppression of the WNT/β-catenin pathway activity" requires further corroboration, particularly at endogenous tankyrase levels.

      We thank the Reviewer for the encouraging and insightful comments. The major critique concerns whether TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels. IWR1-POMA reduced the level of cytosolic β-catenin more effectively than IWR1 in Wnt3A-stimulated HEK293 cells without protein overexpression (Figure 1D). IWR1POMA also suppressed STF activity more effectively than IWR1 in DLD-1 cells (Figure S8C) and reduced the expression levels of several WNT/β-catenin targets more effectively than IWR1 (Figure 1G and S8D). These results support that TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels.

      There are also some other points that, if considered, would further improve the manuscript, as detailed below.

      (1) Abstract and line 62: Many catalytic tankyrase inhibitors tend to display toxicity, which is likely on-target (e.g., 10.1177/0192623315621192; 10.1158/0008-5472). This constitutes the main limiting factor for these compounds. An incomplete inhibition of Wnt/beta-catenin signalling may contribute to the challenges, but this does not appear to be the dominant problem. A more prominent introduction to this important challenge is probably expected by the field.

      A previous study showed that G007-LK, a selective TNKS inhibitor, exhibited weak efficacy and dose-limiting toxicity at 5‒30 mg/kg BID or 10‒60 mg/kg QD in various mouse xenograft models (PMID: 23539443, Ref. 28). Similarly, G-631, another TNKS inhibitor, also showed dose-limiting toxicity without significant efficacy at 25‒100 mg/kg QD in mice (PMID: 26692561, Ref. 60). However, other studies showed that G007-LK was well-tolerated at 200 mg/kg QD over 3 weeks in mice (PMID: 29316982, Ref. 61), and treating mice with G007-LK at 10 mg/kg QD over 6 months also improved glucose tolerance without notable toxicity (PMID: 26631215, Ref. 62). Importantly, basroparib, a selective TNKS inhibitor, was well tolerated in a recent clinical trial (PMID: 40964966, Ref. 64), and constitutive silencing of both TNKS1 and TNKS2 for 150 days in APC-null mice prevented tumorigenesis without damaging the intestines (PMID: 31337618, Ref. 8). We have included some discussion of the toxicity issue associated with TNKS targeting at the end of the Discussion section.

      (2) The authors do a good job in setting the scene for the need for tankyrase degraders. Their observations relating to the formation of puncta (degradasomes) being tankyrase-dependent are compatible with a previous study by Martino-Echarri et al. 2016 (10.1371/journal.pone.0150484): simultaneous silencing of TNKS and TNKS2 by RNAi abolishes degradasome formation. The paper is cited as reference 17, but only in passing, and deserves more prominence. (It includes an entire paragraph titled "Expression of tankyrases 1 and 2 is required for TNKSi-induced formation of axin puncta").

      Indeed, Henderson’s 2016 paper (PMID: 26930278, previously Ref. 17, now Ref. 18) shed important light on the role of TNKS scaffolding in the DC. However, whereas this study demonstrated that knocking down both TNKS1 and TNKS2 by siRNA prevented G007-LK to induce AXIN puncta, it concluded that “puncta formation requires both the expression and the inactivation of TNKS,” which is inconsistent with our observations that accumulation of either catalytically active or inactive TNKS can promote AXIN puncta formation. The function roles of TNKS scaffolding in the DC also remained unaddressed. We have included additional discussion of Henderson’s findings in the first paragraph the Discussion section.

      (3) Moreover, the scaffolding concept has been discussed comprehensively in other studies: 10.1111/bph.14038 and more recently 10.1042/BCJ20230230. There are also a few studies that focus on targeting the ankyrin repeat clusters of tankyrase to disengage substrates (10.1038/s41598-020-69229-y; 10.1038/s41598-019-55240-5) that illustrate the concept of blocking the scaffolding function. In that sense, the hypotheses are mature, and it is interesting to see some of them supported in this study. The authors could improve how they set their work into the context of these other efforts and proposals.

      Indeed, Guettler demonstrated in 2016 that TNKS scaffolding could promote WNT/β-catenin signaling, which forms the basis of the current work. Meanwhile, whereas there have been efforts to target the SAM or ARC domain to address TNKS scaffolding by Guettler and Lehtiö, our approach of targeting TNKS for degradation is complementary. We have included in the last paragraph of the Discussion section information on efforts to target the ARC or SAM domains as an alternative approach to suppress WNT/β-catenin signaling without promoting TNKS oligomerization (PMID: 31836723 and 32704068, Ref. 66 and 67).

      (4) In several places in the manuscript, the DC is referred to as "biomolecular condensate", at times even as a "classic example", implying that it operates through phase separation. This has not been demonstrated. In fact, super-resolution microscopy indicates that the puncta are not droplet-like (10.7554/eLife.08022), which would argue against the condensate hypothesis.

      Biomolecular condensates are membraneless cellular compartments formed by phase separation of biomolecules, regardless of their physical/material properties (PMID: 28935776 and 28225081, Ref. 22 and 23). Super-resolution microscopy studies by Stenmark (PMID: 26124443, Ref. 17) showed that AXIN, APC, TNKS, and β-catenin interacted with each other to assemble into membraneless complexes, wherein AXIN and APC formed filaments throughout the DC. Peifer has also summarized evidence that supports the condensate nature of the DC (PMID: 30782412, Ref. 9; see also PMID: 26393419). However, we acknowledge that testing the physical properties of reconstituted DC (for example, PMID: 34352208) with TNKS will provide a better understanding of the nature, for example liquid vs. gel, of these condensates.

      (5) It is beautiful to be able to use IWR1 and IWR1-POMA at identical concentrations for direct comparisons. However, this requires the two compounds to bind to tankyrase similarly well and reach the target to a comparable extent. How sure are authors that target engagement is comparable? Has this been evaluated?

      Using a BRET assay, we have confirmed that IWR1-POMA binds to TNKS1 with affinity comparable to that of IWR1. Details of this study is now included in the Results sections, and the data are presented in the Supplementary Information (Fig. S3E–G).

      (6) Figure 1F: It is not immediately apparent how IWR1-POMA shows more complete containment of Wnt/beta-catenin signalling. Most Wnt/beta-catenin targets lie close to the perfect diagonal, so I do not see how the statement "that IWR1-POMA controlled WNT/β-catenin signaling more effectively than IWR1" (in the legend of Figure 1F) is supported. Minimally, an expanded explanation would benefit the reader. Providing the colour-coding legend directly in the figure would help improve clarity. Also, the panel is very small and may benefit from a different presentation in the figure.

      We have updated Fig. 1F to include an inset of Quadrant III for improved clarity and readability. We have also moved Fig. S7C to the main text as Fig. 1G and added an expanded explanation for these figures.

      (7) Figure 2: The conclusion of a "deeper suppression" of signalling relies on overexpression of tankyrase in an otherwise tankyrase-null background. Have the authors attempted to measure reporter activity or endogenous gene expression without tankyrase overexpression, in Wnt3a-stimulated cells (in the context of a normal Wnt/beta-catenin pathway) or CRC cells at the basal level? Non-catalytic activity in a similar assay has previously been observed upon tankyrase overexpression (10.1016/j.molcel.2016.06.019). Whether or not there is a substantial scaffolding effect at endogenous tankyrase levels after tankyrase inhibition remains unconfirmed, and the PROTAC is a valuable tool to address this important question. The findings presented in Figure S7C and D go some way towards answering this question - these data could be presented more prominently, and similar assays could be performed in other cell systems.

      IWR1-POMA suppressed STF activity more effectively than IWR1 in APC-mut DLD-1 and SW480 CRC cells without TNKS overexpression (Fig. S8C). Similarly, IWR1-POMA provided a deeper suppression of STF signals in HeLa cells transfected with AXIN1 and β-catenin while expressing endogenous TNKS (Fig. 4G). These results suggest that inhibitor-induced TNKS scaffolding plays a significant role at endogenous TNKS expression levels. Following the reviewer’s suggestion, Fig. S7C is now Fig. 1G.

      (8) Line 237/238: "TNKS accumulation negatively impacts the catalytic activity of the DC (Figure 5D)" - the data do not show this. Beta-catenin levels are a surrogate readout for DC function (phosphorylation and ubiquitylation). Minimally, this requires rewording, with reference to beta-catenin levels.

      We have rephrased "TNKS accumulation negatively impacts the catalytic activity of the DC" as "TNKS accumulation negatively impacts the exchange of β-catenin in the DC."

      (9) Line 303-304: Beta-catenin is thought to exchange at beta-catenin degradasomes; this is clear from previous FRAP assays and the observation that phospho-beta-catenin accumulates in degradasomes upon proteasome inhibition (10.1158/1541-7786.MCR-15-0125). However, degradasome size hasn't, to my knowledge, been related to activity. Can this be clarified, please?

      We apologize for confusing β-catenin phosphorylation with β-catenin abundance. Here, we refer the catalytic activity of the DC to as the ability of the DC to promote β-catenin degradation rather than the kinetics of β-catenin phosphorylation. It is commonly observed that AXIN stabilization by TNKS inhibitors increases the DC size and reduces the β-catenin levels. As such, the induction of AXIN puncta by TNKS inhibitors is frequently used as an indicator of WNT/β-catenin pathway inhibition. However, we have found that, TNKS inhibition drives TNKS accumulation, which reduces the ability of the DC to promote β-catenin degradation. We agree that the DC only primes β-catenin but does not catalyze its degradation. We have revised our manuscript as follows: "increasing the local concentration of the DC components improves its 'effective activity'[50,51]."

      (10) There are previous hypotheses/proposals that the sensitivity of CRC cells to tankyrase inhibition correlates with APC truncation or PIK3CA status (10.1158/1535-7163.MCT-16-0578; 10.1038/s41416-023-02484-8). Have the authors considered expanding their cell line panel (Figure S7) to sample a wider range of cell lines, including some that are wild-type with regard to APC or Wnt/beta-catenin signalling in general? This would be a valuable addition to the work. Quantitated colony formation data could be moved to the main body of the manuscript.

      We have so far tested the effects of IWR1-POMA on the proliferation of DLD-1, SW480, HT-29, HCT116, and RKO cells (Fig. 6A and 6B). While a heterozygous Ser45 deletion in CTNNB1 confers resistance to IWR1-POMA, we did not observe sensitivity associated with APC or PIK3CA status. The ability of IWR1-POMA to suppress the growth of RKO cells expressing wild-type APC is consistent with a previous report that knockdown of both TNKS1 and TNKS2 stabilized PTEN to suppress cell proliferation and glycolysis in vitro and tumor growth in vivo (PMID: 25547115, Ref. 48) independently of the β-catenin pathway. We have added this new information as well as quantification of the colony growth results (Fig. S8A, S8B, S9A, S9F, and S9G) to the revised manuscript.

      (11) The manuscript only mentions toxicity (i.e., therapeutic window) in the last sentence of the Discussion section. As this is THE main challenge with tankyrase inhibitors (as mentioned above), can the authors expand their discussion of this aspect? Is there an expectation that PROTACs may be less toxic?

      As discussed above, evidence for on-target toxicity of WNT/β-catenin inhibition is mixed. Yet, the absence of dose-limiting toxicity for basroparib at doses up to 360 mg QD in human (PMID: 40964966, Ref. 64) is encouraging. PROTAC works by catalyzing target degradation, which is different from traditional catalytic inhibitors that require continuous target occupancy at a high level. It remains unclear whether the observed on-target toxicity of TNKSi is associated with TNKS accumulation at high doses, akin to the cytotoxicity induced by PARP1-trapping upon catalytic inhibition. We have included a brief discussion of the toxicity issue in the final paragraph of the Discussion section.

      (12) Figures 3, 4, 5A: For fluorescence microscopy experiments, can these be quantified, and can repeat data be included?

      We have included quantification data and replicate information for Fig. 3–5.

      (13) Figure 4, S6: An additional channel illustrating the distribution of cells (e.g., nuclei, cytoskeleton, or membrane) would be helpful for orientation and context for the AXIN1 signal.

      We have included cell outlines or nuclear staining for Fig. 3, 4, S6, and S7.

      (14) How were cytosolic fractions of cells prepared to assess cytosolic beta-catenin levels? This detail is missing from the methods.

      We have updated the Methods section to include additional details on the preparation of the cytosolic fractions of cells.

      Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

      We thank the Reviewer for the helpful suggestions.

      I have some recommendations listed below for consideration by the authors to strengthen their study:

      (1) The title is slightly misleading, as it is already known that the scaffolding function of TNKS is important within the DC. The authors should consider incorporating the PROTAC targeting aspect in the title (e.g., PROTAC-mediated targeting of tankyrase leads to increased inhibition of betacat signaling and CRC growth inhibition).

      We have modified the title accordingly to "Targeting tankyrase scaffolding in the β-catenin destruction complex by PROTAC overcomes the limitation of catalytic inhibitors in cancer."

      (2) The authors should comment in the manuscript on the bell-shaped curve obtained with treatment of cells with the PROTACs (Figure S2C). This likely indicates tittering of the targets within a bifunctional molecule with increasing concentration (and likely reveals the auto-inhibition conferred by the catalytic inhibition alone).

      As suggested by the Reviewer, the bell-shaped dose-response likely originated from the formation of non-productive binary protein-ligand complexes at high PROTAC concentrations. We have added a sentence to clarify this unique behavior of PROTAC molecules.

      (3) The authors comment that using G007-LK as warehead was unsuccessful, but they do not show data. Do the authors know why this was the case?

      The structure-activity relationship of PROTACs is often unpredictable, as both the kinetics and thermodynamics of target and E3 ligase binding play important roles in promoting efficient target degradation. We have include data on G007-LK based PROTACs (Fig. S2D) in the revised manuscript.

      (4) Throughout the manuscript, the authors need to do a better job at quantifying their results (i.e., the western blots and the IF). For example, the degradation of TNKS1/2 in Figure 1D is not overly convincing. Similarly, the IF data in Figure 3 needs to be quantified in some ways. Along the same lines, the effect of IWR1-POMA treatments on the proliferation of cells and organoids should be quantified using viability assays... There is also no indication of how many times these experiments were performed and whether the blots shown are representative experiments. The quantification should include all experiments.

      We have included quantification of the immunofluorescence images, colony formation data, and Western blots in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) For clarity, can the authors use the official gene names, TNKS and TNKS2?

      We favor using TNKS1 and TNKS2 when referring to the protein for clarity and use TNKS for simplicity when referring to both proteins.

      (2) Line 92: The authors refer to TNKS2 "induction" - it remains unclear what is meant by "induction".

      We have changed "without induction" to "under basal conditions".

      (3) Can the authors please display molecular weight markers for Western blots throughout?

      (4) Line 144: The description "significantly more effectively" refers to Figure S5A, which shows a single, non-quantified Western blot. I don't think significance has been tested, and this statement should be reworded, or quantified aggregate data provided.

      We have added a Supplementary Information file showing molecular weight markers and quantification of Western blots.

      (5) Line 226: "plateaued at a much lower level" - can this be expressed more quantitatively in the text?

      We have included more quantitative information on the FRAP results.

      (6) Line 249: Can the authors repeat the cross-reference to Figure S7A here?

      We have repeated the cross-reference to the figures.

      (7) Line 266: The description of the experiment using the GSPT1/2 degrader CC-90009 would benefit from a brief recap of the purpose as not every reader will be familiar with this common PROTAC off-target. This is a very thorough analysis, though, and commendable.

      We have added background information on GSPT1 degradation to the revised manuscript.

      (8) Figure 1A: Can the number of repeats and the type of repeats be indicated, please?

      (9) Figure 2: Does n refer to biological or technical repeats?

      (10) Figure 5B, D: How many separate experiments are the data based on?

      (12) Figure S3D, S9A, D: number and types of repeats and the nature of the displayed data and error bars need to be included, please.

      (13) Figure S6B, S7B: I can see three data points, but it would still be helpful to state the number and type of repeats in the legend.

      (14) Figures S9A, S9D: There is value in showing the cumulative data from several repeats in the main figure (Figure 6, which currently is only qualitative) rather than the supplementary material.

      (15) Where single Western blots are shown, can the authors indicate how many experiments they are representative of?

      We have included the number of biological repeats for all data.

      (11) Figure S2C: For most graphs, the main response of interest occurs at low compound concentrations. The y-axis scale does not always help the reader to appreciate the effects, as the response seems small against the magnitude of the hook effect. Interrupting the y-axis as in the final panel may help, with y-axis scales consistent over all panels in the figure.

      We have updated Fig. S2C to emphasize on the degradation efficacy.

      (16) The authors may want to give further method details for some of their assays to facilitate replication of their experiments in the future. For example, the STF assay description is currently quite minimalistic. I assume the assay is fairly robust, though. Other details include cell media (general media details and specific additives and their concentrations in the 3D spheroid formation assay), etc. A general look at the methods section will likely be beneficial.

      We have updated the Methods section to provide more detailed experimental information.

      Reviewer #3 (Recommendations for the authors):

      (1) In Figure 2A, one of the most important findings of the manuscript is that IWR1-POMA induced promoted deeper suppression of beta-catenin-mediated transcription. This seems to be the case only at 3.2uM. Is it statistically significant? What are the data points on this graph? What are the error bars?

      We have included statistical analysis as Fig. S5G.

      (2) On Figure 2C and 2D, do the authors know why the TNKS20M1054V mutant is much better at promoting signaling than the TNKS1-PD ? Is it expression levels?

      It is indeed interesting that TNKS2-M1054V promoted significantly stronger WNT signaling than TNKS1-PD. The basis for its strong scaffolding effect is unclear.

      (3) In Figure 4C, the authors claim that when cells are treated with IWR1-POMA, AXIN1 is distributed diffusely throughout the cytoplasm. It appears that small punctae are visible.

      Quantitative analysis (Fig. 4F) suggest that the size of AXIN1 puncta upon IWR1-POMA is rather insignificant.

      (4) Label on Figure 1D has a spelling error TNKS1/2.

      Corrected.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

      Weaknesses:

      Many aspects of this study have been described in other Down syndrome mouse models, though there are certainly aspects that are new. It would be useful if the authors could do a direct critique and comparison with previous publications in the area, utilizing the same Down Syndrome mouse model. There are also a few limitations in the number of animals used and the interpretation of the data that should be acknowledged.

      We have cited all relevant publications using Down syndrome mouse models. Regarding the Dp16 model, we have cited and discussed the only other study addressing metabolic aspects beyond body weight (Reference #138; PMID: 39803786). While that study reported glucose intolerance, insulin resistance, and defective insulin secretion, we did not measure pancreatic insulin content in our mice. Crucially, while the previous study found no sexual dimorphism, our study observed extensive sexual dimorphism in body weight gain, tissue-specific gene expression, and serum and liver metabolite changes.

      Regarding sample size, we used 6 mice per genotype per sex for transcriptomic and metabolomic analyses; this is constrained by the cost of performing these omics-type analyses. For mitochondrial respiration assays, we used 9–10 mice, and for most other in vivo and ex vivo assays, we utilized 12–15 mice, with some assays exceeding 20. We believe these sample sizes are robust and appropriate for this study.

      Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

      Weaknesses:

      (1) Assessing metabolism using dynamic testing is a strength. ITT, GTT and LTTs are included.

      (2) The dosing for GTTs, ITTs and LTTs was performed per body weight. But the mice under chow and HFD had different body weights. This may compromise the interpretation of the data. Further, ITTs are presented as percentage change, and this can be heavily influenced by baseline glucose measures. The changes appear quite dramatic, so can the authors plot the raw data instead?

      We have updated the ITT data plots to show raw glucose values instead of percentage change. Regarding the dosing, we believe basing it on body weight is an appropriate approach. This method is consistent with nearly all published rodent studies, as blood volume and metabolic tissues such as skeletal muscle and adipose tissue scale with body weight. Adjusting for weight prevents potentially erroneous conclusions. As for the diet groups, we compared WT and Dp16 mice only within the same diet group (Chow or HFD) rather than across different diets. We believe this ensures a valid and appropriate comparison for our study.

      (3) In addition, throughout the manuscript, it is not clear which tissues are the most dominant in disrupting metabolism. The ITT and GTT are composite measures across tissues. Tissue-specific analyses using a clamp technique or isolated tissues may provide more clarity here.

      Our data suggest a systemic metabolic deficit across multiple tissues, supported by tolerance tests, pan-tissue transcriptomic analyses, and liver and serum metabolite profiling. This is consistent with the triplication of genes in Down syndrome, several of which have known metabolic roles as highlighted in our discussion. We do not have evidence to support the role of a dominant tissue that contributes to the systemic metabolic dysfunction.

      Regarding the suggestion to use a clamp technique, we agree this would effectively determine whether insulin resistance is localized in the liver or skeletal muscle. However, we do not currently have the necessary equipment at Johns Hopkins University to perform these experiments. Conducting this work would require sending separate cohorts of WT and Dp16 male and female mice (on both chow and HFD) to an NIH-funded Mouse Metabolic Phenotyping Centre (MMPC). While we appreciate the value of this approach, we believe such labor-intensive experimentation falls beyond the scope of the present study.

      (4) One of the aims of the study was "to help inform the selection of mouse model that best reflects the metabolic profile of DS". The discussion does not contain a comparison between the previous work on different strains and relative to known human data.

      We chose not to include a comparison of different mouse models in the "Discussion" section because we previously highlighted the widely used Down syndrome models (Ts65Dn, Tc1, and TcMAC21) and their associated caveats in the "Introduction." Given the significant limitations of those models such as hypermetabolism in TcMAC21 and the presence of 41 triplicated protein-coding genes unrelated to human chromosome 21 we focused our in-depth metabolic analyses on the Dp16 model, which does not share these issues. We felt that restating this information in the "Discussion" would be unnecessarily repetitive.

      (5) Data availability. Raw metabolomic data should be made available.

      We have uploaded all metabolomics data, along with details regarding sample processing and data analysis, to the Metabolomics Workbench, an NIH-funded public repository. We have updated the "Methods" and "Data Availability" sections of the manuscript to include this information and the corresponding access link.

      Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

      Weaknesses:

      It might have been helpful to have included blood pressure measurements, given the differences in 19-Nor-deoxycorticosterone. The discussion references several articles that describe sex-dependent differences in metabolic phenotypes in humans with Down syndrome, and it might have been helpful to state more explicitly whether these differences correlate with those observed here in mice.

      We appreciate the suggestion of blood pressure measurements. While we agree this is an important metric, given the metabolic focus of the present study and the significant volume of data already presented, we feel that blood pressure analysis is beyond the current scope and better suited for a follow-up study.

      Our study highlights sex differences in metabolic phenotypes in individuals with Down syndrome. While most published human studies focus on a limited set of parameters such as body weight, adiposity, serum lipoprotein profile, and fasting lipid/glucose levels our mouse data remain generally concordant with these findings. Beyond these standard measurements, we also observed substantial sex differences in pan-tissue transcriptomes as well as serum and liver metabolites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A major question is how these findings compare to data that have previously been published. For example, Lamantia et al. Bone 2024 and Dard et al. European Journal of Pharmacology 2025 both report no changes in body weight using the same Dp(16)1Yey Down syndrome mouse model? There is also a recent publication on liver dysfunction in Down Syndrome using the same mouse model. It would be useful to understand some of the similarities and differences of what is being reported by Dunn et al. Cell Rep 2026. In this assessment, there is an in-serum alanine transaminase (ALT) level, which was not the case in Dunn et al?

      For the Lamantia et al. Bone 2024 study, the authors only measured the body weights of Dp16 mice at 6 weeks of age. Our findings at 6 weeks align with Lamantia et al., showing no weight differences between Dp16 and WT mice of either sex (Fig. 2A and C). For the Dard et al. 2025 study, the authors only measured the body weights of Dp16 mice at 12 weeks old (P90) and observed no differences in body weights between genotype of either sex. At 12 weeks of age, we also did not observe body weight differences between Dp16 male mice and WT littermates (Fig. 2A). However, at 12 weeks of age, the Dp16 female mice clearly gained more weight compared to WT littermates (Fig. 2C). Our study tracked weights weekly from 6 to 16 weeks, revealing that while Dp16 females start at weights similar to WT littermates, the groups diverge over time. The reason for the difference between our findings and the single-point measurement by Dard et al. is unclear. Notable variables include:

      Mouse Sourcing: We obtained all cohorts and littermate controls from Jackson Laboratory, while Dard et al. bred their mice in-house.

      Diet: We used Envigo standard chow (catalogue # 2018SX). Dard et al. did not specify the chow used in their study.

      It remains uncertain whether these or other environmental factors contribute to the observed weight differences in female mice.

      In the Dunn et al study (Cell Rep 2026), they also performed metabolic analyses on serum and liver tissue in Dp16 mice. Consistent with their metabolic analyses of serum and liver tissue in Dp16 mice, we also observed the upregulation of multiple bile acids, including taurochenodeoxycholic, tauromuricholic, taurolithocholic, and lithocholic acids. Furthermore, our findings align with theirs regarding the transcriptomic and biochemical signatures of hepatic inflammation and fibrosis. However, there are two notable differences between our studies:

      (1) Liver Injury Markers: We observed an elevation in serum ALT, whereas the Dunn et al. study did not.

      (2) Sex Differences: We identified significant sex differences in the Dp16 transcriptome and metabolome. In contrast, Dunn et al. reported minimal to no sex differences and consequently combined male and female data for all analyses.

      Because Dunn et al. combined male and female data, a sex-stratified comparison between our results (separated by sex) and theirs was not feasible.

      (2) It would be important to understand trends in wild-type animals compared to Dp16 mice. For example, the sex specific and non-specific features - are any of these described in obesogenic wild-type animals fed on a high-fat diet? I.e., are the same features at play and just exacerbated in Dp16, or is this a Dp16-specific feature of systemic metabolism?

      Published literature indicates that WT females typically gain significantly less weight on a high-fat diet (HFD) than WT males. However, our data suggest that the weight gain patterns observed in Figure 6A and C are specific to the Dp16 genotype. Dp16 females gained substantially more weight during the first six weeks of HFD before WT females caught up. In contrast, Dp16 males showed robust initial weight gain comparable to WT controls, but their weight plateaued after seven weeks while WT controls continued to gain, leading to a clear divergence (Fig. 6A).

      Other metabolic parameters also appear specific to the Dp16 model. On a standard chow diet, WT mice of both sexes generally do not exhibit glucose intolerance, insulin resistance, dysregulated lipoprotein profiles (VLDL-TG), or an impaired capacity to handle lipid loads. We observed all of these features in our Dp16 male and female mice (Fig. 3). Furthermore, transcriptomic analyses of Dp16 mice on standard chow revealed gene signatures of inflammation, fibrosis, and oxidative stress that are absent in WT mice.

      When challenged with HFD, while WT mice typically develop glucose intolerance and insulin resistance, the triplicated genes in Dp16 mice significantly exacerbated this metabolic deterioration. This is reflected in the worsening of glucose control and insulin sensitivity observed in our tolerance tests.

      In summary, most of these metabolic features are specific to Dp16 mice on a standard chow diet and are further exacerbated when combined with a high-fat diet.

      (3) Food intake data is difficult to interpret when weight has already diverged, as bigger animals will eat more food. Hence, the higher food may be a consequence rather than a cause of the weight gain (data in Figure 1).

      The reviewer makes a valid point. Since physical activity and energy expenditure do not differ significantly between Dp16 females and WT controls (Fig. 2F), the observed increase in food intake may indeed contribute to the higher body weights in Dp16 female mice.

      To rigorously confirm this, food intake would need to be measured between 6 and 8 weeks of age, prior to the divergence in body weight. Unfortunately, we did not measure food intake at that earlier time point.

      (4) The n numbers seem to vary significantly. For example, the use of n=6 for metabolic studies is generally rather small and underpowered. For the seahorse data, another concern is the snap freezing of samples before Seahorse assessment. For example, snap freezing of samples has been shown to increase certain metabolites. Freeze-thaw tissues often show a significant reduction in optical redox ratio.

      Regarding the transcriptomics and metabolomics studies, we utilized N=6 mice per tissue per sex. While we agree that a larger sample size is always preferable, the high cost of OMICS analyses covering 144 RNA-seq and 48 metabolomics samples limited our capacity to increase this number. However, N=6 remains a robust and standard approach for these specific assays. For the majority of our other in vivo and ex vivo data, we employed a higher sample size of 12-15 mice per genotype per sex to ensure statistical rigour. For a few assays, we have sample size of over 20.

      Regarding the respirometry analysis, we acknowledge the limitations of using frozen tissue. We chose this method because it allowed us to perform Seahorse assays on multiple tissues from 9-10 mice, which is a significant sample size for this type of analysis. The alternative isolating mitochondria from fresh tissue would have restricted our ability to process multiple tissues from a large number of animals on the same day due to the length of the protocol. We believe this trade-off was necessary to maintain a high sample size across various tissues.

      (5) For oestradiol measurements, were the samples taken at the same times within the estrous cycle? This may affect the comparability of female Dp16 and WT mice?

      Regarding our protocol, blood samples were collected between 11:00 AM and noon, with food removed two hours prior. While we did not specifically monitor the oestrous cycle of the female mice, serum samples for both the Dp16 females and WT littermates were collected on the same day and at the same time to ensure comparability across the groups.

      (6) Body weight reduction and organ size reduction on an HFD are especially interesting. Could enhanced inflammation and fibrosis be the root cause of this? Are there other mouse models where this is the reason?

      On a high-fat diet, we observed a reduction in iWAT and gWAT fat depot weights in both male and female Dp16 mice, which is consistent with their lower overall body weights (Fig. 6 - figure supplement 3). Conversely, Dp16 females fed a high-fat diet showed increased heart and kidney weights. Despite their lower adiposity, the Dp16 mice on this diet exhibited greater insulin resistance and glucose intolerance (Fig. 7). This suggests that the worsening of glucose control is independent of obesity. While we observed signatures of inflammation and fibrosis, we do not yet have direct mechanistic evidence demonstrating that these factors causally impaired glucose and lipid metabolism.

      (7) The authors are circumspect throughout to avoid over-claiming, as the majority of data is observational. One exception: "Many bile acids serve as ligands for nuclear hormone receptors (e.g., FRX and TGR5) that control various aspects of glucose and lipid metabolism (74, 75), and extensive changes in circulating bile acids are contributing, at least in part, to the systemic metabolic phenotypes in Dp16 mice." The authors have not shown a direct link between bile acids and metabolism in this model. Please edit.

      We have edited the text accordingly.

      Minor:

      (1)"Most human studies at the whole-body level are limited to assessing the impact of trisomy 21 on food intake, adiposity, physical activity level, and energy expenditure in adolescents or adults with DS"

      While we were uncertain of the reviewer's specific intent regarding the suggested changes, we have rephrased the sentence for clarity.

      (2) It is somewhat surprising that T3 is elevated, although there are reports of T3 elevation in visceral obesity in humans (e.g., Sun Nam et al., Obes Res Clin Pract, 2010).

      We observed that T3 levels did not differ by genotype in mice of either sex when fed a standard chow (Fig. 2 - figure supplement 5). However, we noted elevated T3 levels in both male and female Dp16 mice on a high-fat diet (Fig. 6 - figure supplement 2). While increased T3 levels correlated with higher physical activity and a modest increase in metabolic rate in Dp16 females, this was not observed in males (Fig. 6). We do not currently have a clear explanation for these findings. Given that individuals with Down syndrome often present with hypothyroidism and lower T3 levels, this discrepancy may reflect a species-specific difference between humans and mice.

      (3) Please can the authors clarify the percentage gene coverage, as this is quoted as ~58% of Hsa21 gene orthologs or ~65% of the Hsa21 gene orthologs, where the same reference is used.

      We apologize for the confusion. The number of triplicated genes in Dp16 mice corresponds to ~58% of Hsa21 genes (PMID: 26765563). We have corrected the typographical error in the text.

      (4) "segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs" for this given percentage majority sounds too strong, and the use of percentage is recommended.

      We have modified the text accordingly.

      (5) It is puzzling that in female gWAT with 7 triplicated Hsa21 gene orthologs (Rbm11, Chodl, Cldn8, Sh3bgr, Igsf5, Itgb2l, and Tmprss2). Could this be a technical issue? Was the reduced expression quantified by RT-Q-PCR?

      We have examined the normalized counts in the RNA-seq data for the seven genes in question, and the results do not appear to be an artifact. The sample size for this data is six mice per tissue per sex. In general, we prefer utilizing raw and normalized counts from RNA sequencing because there is a linear relationship between transcript amount and raw counts that is independent of housekeeping genes. In contrast, RT-qPCR involves mRNA amplification and requires expression to be normalized by one or more housekeeping genes (such as GAPDH, β-actin, 36B4, or ubiquitin) under the assumption that their levels remain constant.

      (6) The difference in body temperature is of interest. In male Dp16 mice, there is an increase in core temperature and a lowering of body temperature in females. In female Dp16 mice, higher estradiol levels have been stated by the authors to contribute to lower body temperature and higher physical activity (69-72). I am uncertain if the references are all relevant, as some relate to ovariectomized animals. No explanation is given for males.

      We currently do not have an explanation for why Dp16 males on a chow diet exhibit higher core body temperature, while Dp16 females show lower body temperatures. Although elevated T3 levels can increase body temperature, we have ruled this out; our data indicates there are no significant differences in T3 levels between genotypes for either sex on a chow diet.

      (7) The authors find a higher percentage heart weight in Dp16 mice on HFD and comment in the discussion that this is in keeping with "high-fat diet-induced cardiac hypertrophy". From what I can see, no histology has been performed to justify this statement. Furthermore, it would be useful to understand which animals had congenital heart disease in the first instance.

      We have modified the text accordingly. Unfortunately, we do not have histology data on the heart to inform us on whether some of our mice had congenital heart disease.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should comment on the dosing method of glucose/insulin/lipid in the tolerance tests to acknowledge that differences in body weight may affect these tests. In addition, I encourage the authors to present ITT data as raw data, and not % change.

      In response to the reviewer’s comments, we have updated the ITT data plots to show raw data rather than percentage change. Regarding the dosing methodology, we maintain that basing dosage on body weight is appropriate. This approach is consistent with the vast majority of published rodent studies, as blood volume and metabolic tissues—such as skeletal muscle and adipose tissue—scale with body weight. Standardizing dose independently of body weight could lead to erroneous conclusions.

      (2) It would be useful for the authors to include a discussion on the likely specific tissue involvement in the whole-body metabolic disturbance. From my reading of the manuscript, there seems to be data suggesting functional and transcriptional dysfunction across most tissues, but do the authors suggest there is a dominant tissue in this regard?

      Due to the triplication of large number of genes on human chromosome 21, people with Down syndrome exhibit deficits across most organ systems (PMID: 32029743). Metabolic homeostasis also involves multiple tissues and cell types (adipose tissues, liver, skeletal muscle, pancreas, gut, hypothalamus, and immune cells). Most of the triplicated genes do express across these tissues. Our data indicate metabolic dysregulation across adipose tissues (white and brown), liver, skeletal muscle, and hypothalamus. Given the complex genetic perturbations of the Down syndrome mouse model, we do not think that there is a dominant tissue that contributes disproportionately to the systemic metabolic dysfunction phenotypes we observed in the Dp16 mice. Rather, we think that the metabolic phenotype is due to the combined deficits across multiple organs and tissues. As we do not have data to support the disproportionate contribution of any one tissue, we therefore did not speculate on the dominant contribution of any single tissue in the Discussion.

      (3) Related to this, muscle lipid is thought to be a major driver of muscle insulin resistance. Do the authors have measures of muscle lipid accumulation? This might be particularly interesting in the HFD models.

      Unfortunately, we did not measure lipid content in the skeletal muscle during this study. For the chow-fed mice, the entire gastrocnemius muscle was used for RNA isolation to perform RNA sequencing, and no tissue remains for additional analysis. Regarding the HFD-fed group, skeletal muscle was not collected at the termination of the study. As a result, we are unable to provide the requested lipid analysis data.

      (4) For mitochondrial analyses - do the authors have measures of total tissue mitochondria, and might changes in mitochondria abundance be driving some of these differences?

      For all our mitochondrial respiration analyses, we normalized the data to mitochondrial content as quantified by the MTDR assay (PMID: 32432379; PMID: 39704485). These results indicate that for a given amount of mitochondrial content, respiration as measured by the Seahorse assay is reduced in Dp16 mouse tissues, specifically in the BAT and liver.

      (5) To broaden the scope and interest, can the authors compare the transcriptional or metabolomic data to what has been found in non-DS insulin resistance (humans or mice), for example? This may help to highlight the key changes in metabolism that are causal for specific phenotypes.

      Overall, this is a comprehensive assessment of metabolism in a DS model.

      We appreciate the reviewer’s suggestion. However, given the vast number of published datasets on non-DS insulin resistance in both humans and mice, comparisons would yield varying results depending on the specific datasets selected. Consequently, we feel that such an analysis is beyond the scope of this study. We would like to highlight that many of the processes dysregulated in Dp16 mice as identified through our pan-tissue transcriptomes and metabolomes align with those frequently observed in non-DS insulin resistance. These include signatures of chronic low-grade inflammation, fibrosis, ER and oxidative stress, and impaired glucose and lipid metabolism.

      Reviewer #3 (Recommendations for the authors):

      It is slightly disconcerting that Figure 5 - Figure Supplements 2-5 are referred to in the text before the data in Figure 5 are discussed. It might make sense to indicate that the data are discussed further below (assuming that the authors do not wish to renumber these figures).

      We have fixed this issue raised by the reviewer.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      We thank the reviewer for their positive assessment and strive to address the weaknesses they have constructively raised below.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we will include a new supplementary figure (Fig. S2) providing detailed 3D views of the copper adapter, microscope objective, and plastic cap. The figure will show that the rim surrounding the front lens of the objective is covered by the plastic cap to provide thermal insulation between the objective housing and the cryogenic environment (Fig. S2b). We will also clarify that the front surface of the cap is levelled with the front objective lens to maintain the full working distance of the objective while allowing for axial movement of the z-stage. Finally, we will explicitly state that the front objective lens is exposed to the cryogenic environment (cold nitrogen gas).

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      To minimise vibrations arising from the nitrogen refilling pumps, the cryostat and liquid nitrogen tubing are mechanically decoupled from the microscope cage system, objective, translation stages, and sample. Specifically, the cryostat and nitrogen tubing are supported independently on a laboratory jack and surround the cage system without rigid mechanical contact. In the revised manuscript, we will update Fig. S1a, b to illustrate the liquid nitrogen tubing and refilling system more clearly. In addition, we will include a new supplementary figure (Fig. S3) to show the detailed cryostat design, refilling tubing, and temperature sensor position.

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      We agree that the purpose of the rectangular aperture should be made clearer. In line with the suggestion from the reviewer, in the revised manuscript, we will briefly explain that the aperture is intended for selective illumination, such as in applications to cryo-FIB lamellae, and will cite our recent preprint describing this approach.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      We thank the reviewer for this suggestion. In the revised manuscript, we will include an additional supplementary figure (Fig. S4) showing the axial drift measured without focus-lock correction to illustrate the intrinsic mechanical stability of the microscope. We will also clarify that the periodic lateral displacement observed along the x-direction (approximately 300 nm amplitude with a period of ~22 minutes) arises from slight lateral repositioning accompanying z-stage stepping during focus-lock operation, likely due to mechanical coupling between the axes of the translation stage. We will revise the lower panel of Fig. 2d by reducing the y-axis range to improve data visibility.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      We agree that this would be useful. In the revised manuscript, we will update Fig. 3 to include two additional low and intermediate-magnification cryo-EM images showing a representative grid square and a zoomed-in region of it, including a few grid holes. These images provide an overview of the ice contamination across a substantially larger field of view.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      The y-axis in Fig. 4b represents the localization rate (number of detected localizations per frame) within the selected ROI in Fig.4c and was not normalized. The values were calculated in SMAP by binning the localization frames into 100 temporal bins and dividing the number of localizations in each bin by the corresponding bin width, resulting in units of localizations per frame. Therefore, values close to 1 indicate approximately one localization detected per frame at that time point. To avoid potential confusion regarding the interpretation of this representation, we will replace this plot in the revised manuscript with a more explicit visualization showing the number of detected localizations per defined number of frames as a function of time (frame number) for the specific ROI shown in Fig. 4c.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

      In the revised manuscript, we will add fluorescence intensity traces corresponding to the single-molecule events shown in Fig. 4c and Fig. 4f to further demonstrate their single-molecule emission characteristics.

      Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      We thank the reviewer for their critical assessment, and for their suggestions below which we have used to improve the manuscript.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      We thank the reviewer for highlighting this point. We agree that focus stabilization strategies for cryo-FIB lamellae are likely to differ from those used for the individual bacterial cell samples. For lateral drift correction, the presence of a single continuously detectable bright feature within the field of view is sufficient. Importantly, this feature does not need to be a fluorescent bead; any stable signal that can be continuously detected by the camera can serve as a suitable reference for drift correction. We will expand the Discussion to describe potential strategies for stable cryo-SMLM imaging, including the use of intrinsic sample or lamella features for autofocus, minimal fiducial-based approaches, and the practical implications of dedicating the far-red channel to focus stabilization.

      Furthermore, in the revised manuscript, we will include a new supplementary figure (Fig. S4) demonstrating the intrinsic axial stability of the microscope in the absence of active focus-lock correction. These measurements show that the system remains within the objective's depth of focus for a relatively long time, providing adequate stability for experiments in which far-red fluorescent fiducial beads are unavailable, such as cryo-FIB lamella imaging.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      We thank the reviewer for this thoughtful comment. We agree that a correlative example displaying more pronounced structural features would further illustrate the capabilities of cryo-SMLM. However, the primary aim of the present work is the development and characterization of a robust cryogenic super-resolution microscope for reliable cryo-SMLM and correlative cryo-CLEM, rather than the demonstration of new biological applications. The utility of correlative cryo-SMLM/cryo-ET for resolving cellular structures has already been established in previous studies, including those employing rsEGFP2-labelled targets.

      The correlative dataset presented here is intended to demonstrate the compatibility of the microscope with cryo-CLEM workflows rather than to provide detailed biological insight. Moreover, the use of intact E. coli cells imposes inherent limitations on the ultrastructural information accessible by cryo-electron tomography; overcoming these limitations would typically require specimen thinning, for example, by cryo-focused ion beam (cryo-FIB) milling, which is beyond the scope of the present work.

      Regarding the concern about auto-fluorescence, elevated background fluorescence is not unique to bacterial samples or green fluorescent proteins but is a general consideration in cryo-SMLM that depends on the specimen and imaging conditions. While auto-fluorescence may reduce image contrast, it does not affect the conclusions of this work, which focuses on the design and performance of the microscope.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

      Thank you for this useful suggestion. In the revised manuscript, we will make available the complete SolidWorks CAD files for all custom-designed components, together with a comprehensive parts list and the full assembly corresponding to Fig. S1 as supplementary materials.

    1. Author response:

      Reviewer 1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we will also explicitly state this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we will mention and describe alternative approaches for assessing parasite viability. This will also include the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We will revise the text to explicitly mention the use of dual straining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al, 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer 2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We will clarify this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We will add this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We will revise the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      Reviewer 3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1): The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We will add more explanations to the Discussion including the strengths and limitations.

      (2) Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      (3) Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      (4) Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank you for the time you took to review our work and for your feedback! The main changes to the manuscript are:

      We added a paragraph to the Discussion addressing differences in visuomotor mismatch responses recorded over frontal and occipital electrodes, and their possible interpretation.

      We added time-frequency power and phase-locking analysis as supplementary figures to the manuscript.

      We added a statement in the Discussion emphasizing the importance of performing these experiments with denser EEG channel coverage.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference can not be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active. I have a couple of questions/comments.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      The reviewer is correct. But one should keep in mind that our result here stands in the context of a considerable amount of work on mouse cortex investigating responses to very similar visuomotor mismatches. There we can we have much additional evidence to argue that the cortical response to a visuomotor mismatch is a prediction error. We would argue, it is the best one can do in human experiments.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      We have added a paragraph to the Discussion addressing the differences between signals recorded over frontal and occipital electrodes, as suggested.

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      We are not sure what the reviewer is referring to here. We certainly do not think (or are aware of having argued) that a visuomotor prediction error is the only possibly interpretation of the response. In the last paragraph of our response to the reviewers point 3 in the last revision, we explicitly discuss that the interpretation of the responses as a prediction error is only one possible interpretation. Our argument is that it is the most likely given the evidence.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      We would be happy to add additional considerations to other processes. If the reviewer has anything specific in mind, we can add that, but it would need to be somewhat concrete with some theoretical basis. We share the reviewer’s intuition, but unless this can be formalized to the point of being experimentally testable, we do not see any value in discussing it in the manuscript.

      Regarding the reason for a difference in visual responses in walking vs sitting state is, this is not entirely clear to us. Predictive processing would provide one possible explanation. Assuming the precision weighting of predictions is higher during walking, the sudden appearance of a visual stimulus might lead to stronger stimulus history prediction errors than when just sitting. But this is rather speculative.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information ‘from perception’. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Please excuse, the reviewer is correct, the orthodoxy of an idea is not relevant. For dubious reasons, we chose to euphemize what we actually meant to say here. With regards to circuit implementations of predictive processing (we cannot and do not intend to speak to interpretations of predictive processing that relate to conscious perception much of V1 activity is likely not consciously perceived – we assume this is what the reviewer is referring to by “we have one percept”) – the reviewers interpretation was not unorthodox, but rather incorrect (which is what we should have said). The statement that “the brain predictively ‘cancels’ expected action outcomes from perception” is incorrect in the context of sensory processing – based on both theoretical models of predictive processing, and more importantly physiological evidence. If the point was only in regards to conscious perception, we also suspect the statement is wrong, but even if it were correct, don’t see how it pertains to our work.

      Similarly a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      We believe it is absolutely essential for the progress of the field that we start to emphasize the differences between something that is “predictable in principle” and “predicted by the brain”. There is likely indeed a hierarchy of predictability that looks something like this:

      (1) Sensorimotor coupling

      (2) Laws of physics

      (3) Behavior of other living things

      (4) Artificial, human-made statistical relationships

      Almost all published experiments are based on the fourth type of prediction. Indeed, why not use physics simulations instead of oddballs and MMN? We absolutely should! But the field tends to revert to artificial couplings. As a direct consequence of this, the number of papers appearing recently (from both human and mouse fields), that are built on the following premise:

      (1) Expose an animal or human to an artificial coupling between A and B (e.g. an oddball, or a global oddball, or any of a myriad other constructions).

      (2) Probe for prediction error responses to the violation of the artificial coupling.

      (3) Find no prediction error responses and conclude predictive processing is wrong.

      Is utterly baffling. The fallacy here is of course the assumption that if something is predictable in principle, the brain must predict it. Thus, we are, and will continue to be strong on this point, and we think it is essential that we – as a field – are.

      Hope these comments are helpful.

      Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      We have added a discussion of this.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice, and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      The reference electrode placed at Cz makes it is difficult to interpret relative differences between frontal and occipital electrode responses, as the occipital electrodes are placed farther away from the Cz reference than the frontal electrodes. Similarly, signal occuring cortically near the Cz reference might only appear as though it is occipitally distributed in this montage. It is common in EEG research to remontage the data to an averaged common reference in order to better interpret the scalp distributions. As the electrode coverage was sparse for some subjects, this could be challenging, and this reviewer does not feel that it is necessary to do this analysis step, or even to drastically rewrite the body of the paper. We only request that some discussion, however brief, is included in the discussion section or the methods that recommend more dense electrode coverage in the future to better interpret scalp distributions and potential meso-scale sources.

      We have added a discussion of this as suggested.

      This is just a suggestion. The authors are encouraged to analyse (and report) time-frequency power and phase locking for these mismatch responses, as is common in much of the literature (see Roach et al 2008 Schizophrenia Bulletin). This is not to say that doing so will yield insights into oscillations per se, but converting the data to the time-frequency domain provides another perspective that has some advantages. fosters translations to rodent models, as ERP peaks do not map well between species, but e.g. delta-theta power does (see Lee et al 2018 Neuropsychopharmacology; Javitt et all 2018 Schizophrenia research; Gallimore et al 2023 Cereb Ctx). Further, ERP peaks can be influenced by the actual neuroanatomy of an individual (especially for quantifying V1 responses). Time frequency analyses may aid in interpreting the "early negative deflection with a peak latency of 48 ms " finding as well. As it stands, the report is complete, and it would be acceptable if the authors chose to save this type of analysis for a future publication.

      We have added this as suggested.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed most of my concerns by providing additional analyses, partly based on new data. The volume conduction issue is partly addressed based on the result showing latency differences, however to confidently assign responses to visual regions, one would need to perform recordings with a larger number of electrodes, sufficient to perform source localization. Nevertheless, the manuscript is now more solid than the previous version.

      We have now added this point to the Discussion.

      Reviewer #3 (Recommendations for the authors):

      The reviewer appreciates that the authors have carried out time-frequency analyses, and are ok with them leaving this out of this paper.

      We have now added this to the manuscript.

      Finally, in response to the participant quote "are you printing this? hi mom!" - this reviewer concedes that it does not significantly detract from the report, and, in the interest of amusement and joy, would abide its reinstatement.

      We greatly appreciate the reviewers entertaining our attempts at humor but will leave it out as originally suggested.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weaknesses:

      A control for non-specific binding is missing from the in vitro binding assay. The figures suffer sometimes from the very small text in the labels, which obscures understanding. Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

      We thank the reviewer for highlighting the strengths of our imaging pipelines and protein-protein interaction data. We have now included appropriate controls for the in vitro binding assays, which demonstrate that the observed interactions are specific (Fig. 2H). We have also revised all figures to improve clarity, including increasing font sizes to enhance visibility across panels. We agree that mapping the specific phosphorylation sites regulated by these septin kinases would provide valuable mechanistic insights. However, only a few studies have addressed this direction so far (Mortensen et al., 2002; Asano et al., 2006 and Marquardt et al., 2024) [1-3]. The current study focuses on the interplay among septin-associated kinases and their role in regulation of the actomyosin machinery (AMR). In this context, we highlight several key findings:

      (i) a molecular link between the septin kinase network and AMR through physical interaction between the KA1 domain of Gin4 and F-BAR domain of Hof1 (Fig. 2F-2H and S3H);

      (ii) a kinase-independent role for Gin4 in coordinating septin organization and AMR dynamics (Fig. 3A-3E, S3F & S3G, S3I and 4F-4H);

      (iii) a novel role for Hsl1 in regulating septins and the AMR downstream of Gin4 and Elm1, potentially through plasma-membrane binding (Fig. 4I-4K, 5A-5F, 8A-8G and S9H-S9J); and

      (iv) crosstalk between Gin4 and Hsl1 that is independent of their role in the morphogenetic checkpoint (Fig. 6A-6D and S6A-S6C).

      We have clarified this scope in the Discussion section and explicitly stated that mapping these phosphorylation sites will be an important direction for future work (Lines 623-626).

      Reviewer #2 (Public review):

      Summary:

      In this paper, Bhojappa et al. provide insights into the function of septin-related kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and actomyosin ring (AMR) structure and constriction. Their findings are both corroborative of and complementary to previous related studies.

      First, the authors provide a comparative analysis of the dynamic localization of these kinases at the bud neck, as well as a comparative analysis of defects in septin localization, splitting dynamics, AMR constriction rates, and cell morphology in kinase deficient cells. They find that septin localization and splitting kinetics, as well as AMR constriction rates, are significantly perturbed in elm1∆ and gin4∆ mutants but remain largely unaffected in hsl1∆ and kcc4∆. A similar trend is observed in terms of cell morphology and viability.

      Next, the authors focus on elm1∆ and gin4∆ cells, demonstrating that the residence time of the F-BAR protein Hof1 is significantly increased and defective in these mutants. Using yeast two-hybrid (Y2H) and in vitro binding assays, they show that the KA1 domain of Gin4 interacts with the F-BAR domain of Hof1, which may explain the cytokinesis-related functions of Elm1 and Gin4. Supporting this, they find that Gin4's role in septin localization, AMR constriction kinetics, and Hof1 bud neck localization is kinase-independent.

      The authors then conduct a series of artificial tethering experiments given their bud neck localization is mostly interdependent. They first demonstrate that artificially tethering Gin4 to the bud neck rescues the morphology defects of elm1∆ cells, with the strongest rescue observed when Gin4 was forced to interact with Hsl1-an effect that was also kinase-independent. Additionally, artificial tethering of Hsl1 to the bud neck restores the morphology of elm1∆ cells in a KA1 domain-dependent manner, suggesting that Hsl1 functions downstream of Elm1 to maintain normal cell morphology. Consistently, artificial tethering of Elm1 to the bud neck in gin4∆ cells rescues morphology defects, as well as defects in Myo1 localization and AMR constriction, but only in the presence of full-length Hsl1. The rescue fails in the absence of Hsl1 or when using a version of Hsl1 lacking the KA1 domain, which supports the role of Hsl1 downstream to Elm1 in cytokinesis.

      Strengths:

      Altogether, this study offers valuable insights into the mode of cytokinesis regulation mediated by the septin-related kinases, mainly Elm1, Gin4, and Hsl1, and would be an important contribution to the field of septins and cytokinesis after addressing current weaknesses.

      We thank the reviewer for the detailed summary and for highlighting the novel findings of our study.

      Weaknesses:

      (1) When assessing rescue of the elm1∆ phenotype, it needs to become clearer whether only morphology or also cytokinesis and septin organization are rescued.

      To clarify the extent of rescue observed in elm1Δ cells, we extended our analysis beyond morphological parameters by quantifying septin organization and AMR constriction dynamics. These analyses now show that artificial tethering partially restores septin organization and AMR constriction kinetics in elm1Δ cells in addition to improving cell morphology. These results are now described in detail in Fig. 5 and lines 373-406.

      (2) The quantification of the microscopy data does not always match up with the example images, and it's not always clear how the authors quantitatively analyzed their data.

      We revised the manuscript to clearly outline the quantification methods used for microscopy data analysis and specified the statistical tests, number of cells analyzed, and number of experimental replicates in the figure legends. We also clarified the criteria used for phenotype scoring and quantification in the Materials and Methods section. In addition, We replaced representative images where necessary to accurately reflect the quantified data throughout the revised manuscript.

      (3) The forced tethering data are key to the paper, but the lack of a summarizing table makes it difficult to grasp the full picture.

      We agree with the reviewer and have now included a new summary table (Table 1) that compiles the results of all artificial tethering experiments presented in this study, including the percentage of rescue in cellular morphology observed upon forced tethering of these kinases to the bud neck, thereby providing a clearer overview of these experiments.

      (4) Novel results and those confirming earlier results could be better distinguished.

      We have improved the overall clarity of the manuscript to distinguish novel findings from the results that corroborate previous studies, and have cited the appropriate literature throughout the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p, and Kcc4p. Elm1 and Gin4 show strong knock-out phenotypes, whereas Hsl1p and Kcc4p show weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Strengths:

      The combination of genetics, cell imaging, and biochemical characterization of proteinprotein interactions is attractive.

      We thank the reviewer for recognizing the significance of our findings and for the helpful suggestions.

      Weaknesses:

      (1) Imaging and data analysis is the main weakness of this manuscript. The authors must avoid manual counting and selection when easy analysis software can be used to limit bias. Instead of presenting unclear statistics of "percentage phenotypes", they need to define clear metrics to offer meaningful phenotype analysis.

      We agree that improving the quantitative rigour of the image analysis is essential for this study. Accordingly, we implemented a semi-automated image analysis workflow in the revised manuscript that defines reproducible metrics, such as aspect ratio, and reduces reliance on subjective phenotypic scoring. The inclusion of these parametric measurements enables clearer and more objective comparison of the rescued phenotypes.

      (2) This manuscript examines a very complex mechanism with four kinases of overlapping function using new data and existing literature. A clearer picture/model at the end of the manuscript that synthesizes the current knowledge would be beneficial:

      We incorporated a new representative model (Fig. 9) that integrates current knowledge in the field with our findings. This model highlights crosstalk among Elm1, Gin4, and Hsl1 as a key mechanism coordinating septin architectural transitions with AMR constriction during cytokinesis and is discussed in lines 520-536 of the revised manuscript.

      We sincerely thank all the reviewers for their insightful comments. We incorporated new results in Fig. S3A, S3B, S4A-S4F, 5A-5F, 6A-6D, S6A-S6C, 8E, 8F, S10B and 9, along with Table 1 summarizing the artificial tethering experiments in the revised manuscript. We believe that these revisions have improved the rigor of our analyses and enhanced the overall clarity of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 70: " abnormal cytokinesis defects": this language is redundant, either "abnormal" or "defects" would suffice:

      We thank the reviewer for identifying the redundant wording. We have rephrased the final paragraph of the Introduction section.

      Please refer to line numbers 86-97.

      (2) Currently, the final paragraph of the Introduction is an extensive, detailed summary of the results. This is unnecessary, as the Abstract and the Results sections summarize the results. Better would be a short statement of the questions addressed in the manuscript:

      We thank the reviewer for this suggestion. We have revised the final paragraph of the Introduction to remove the detailed summary of results and instead outline the key questions addressed in this study. The revised paragraph now emphasizes the knowledge gap regarding how septin-associated kinases regulate septin organization and coordinate cytokinesis. Please refer to line numbers 86-97.

      (3) In the Results, the wording of the heading associated with the first section is confusing. "and their defects during the cell cycle": it is not expected that the kinases themselves will have defects; defects may be observed in cells upon mutation of the kinases, but that is not clear from this wording:

      We thank the reviewer for this comment. The heading has been revised from “and their defects during the cell cycle” to “and defects associated with their deletions” for improved clarity.

      Please refer to line numbers 99-100.

      (4) Lines 107-109 (" they may act as molecular signals for the transition and a trigger for crosstalk of cytokinesis") are redundant with earlier lines 104-105 (" suggesting them to be a possible trigger for septin remodelling"):

      We have rephrased the text to improve clarity and avoid redundancy.

      Please refer to line number 123-125.

      (5) Lines 121-122: "Septin-associated kinases are believed to play an essential regulatory role": is the role essential, or is it regulatory? Since these kinases are not individually essential for cytokinesis, "essential" doesn't seem appropriate here:

      This has been corrected in the revised manuscript.

      (6) Figure 1 panels C and F: the font size is exceedingly small for the labels and should be greatly increased. This is true for multiple panels in Figures 2-5 and the Supplemental Figures as well:

      We thank the reviewer for highlighting this issue. We have significantly increased the font sizes of the text and the x- and y-axis labels across all figures in the manuscript.

      (7) The timing of mitotic spindle breakdown was used as a timepoint for comparison but it is not made clear in the manuscript how this was determined. Presumably, the mRuby2-Tub1 marker was visualized and the timepoint when the mitotic spindle separated into two discrete entities was called the "breakpoint" timepoint, but it would be important to better describe (and ideally show an example) of how that timepoint was determined. The only images I find with labelled tubulin are Tub1-GFP and these do not show spindle breakdown:

      We thank the reviewer for raising this point. We used GFP-Tub1 (pAFS125-GFPTUB1) and mRuby2-Tub1 (pHIS3p:mRuby2-Tub1+3′UTR::URA3) plasmids to visualize spindle dynamics across the cell cycle, and in both cases defined the spindle breakpoint as time zero. As suggested, we have now included time-lapse images of Cdc3-mCherry and GFP-Tub1 in wild-type cells in Fig. S2D to illustrate the spindle breakpoint event used for temporal alignment. Corresponding changes have also been made in the Materials and Methods section to explicitly describe this analysis.

      Please refer to line numbers 760-761.

      (8) Line 165: "F-BAR protein Hof1, which senses and induces membrane curvature" and lines 170-172 "the F-BAR protein Hof1, which is known to be associated with septin hourglass and transit to AMR during split ring trigger". It is awkward to introduce the same protein twice, in two different ways, within a few lines of each:

      We thank the reviewer for noting the redundant description. We have rephrased the text to introduce the F-BAR protein Hof1 in a more concise and streamlined manner while retaining the relevant functional information.

      Please refer to line numbers 206-208.

      (9) Throughout, it would be helpful to introduce more line breaks and organize the Results into smaller paragraphs:

      We thank the reviewer for this suggestion. We have revised the layout of the Results section and introduced additional line breaks to improve readability.

      (10) This is somewhat of a personal preference, but in the interest of transparency (and with the understanding that the 0.05 value is entirely arbitrary), would authors be willing to show actual P values in the figure panels rather than "**" or "ns", for example? Some readers may wish to apply a different standard of significance than 0.05, and not showing the P values makes this impossible. Furthermore, some readers (like this one) may interpret differently a P value of 0.044 versus "*" or 0.051 vs "ns":

      We agree with the reviewer’s suggestion. To improve the transparency of the quantitative analyses, we have modified the graphs across all main and supplementary figures to display the “actual p-values” and specified the corresponding statistical tests in the figure legends, with significance indicated by asterisks. An example is provided in the attached image showing the residence time of Inn1-mNG in wild-type and gin4Δ cells complemented with kinase-active and kinase-dead Gin4 constructs (Fig. 3E in the revised manuscript). For this analysis, significance was assessed using the nonparametric Kruskal-Wallis statistical test.

      (11) It is written that the authors "performed a Yeast Two Hybrid screen to find novel interacting partners at the bud neck" but I do not find evidence anywhere of a "screen" being performed, i.e., an unbiased search of many proteins to find a few interactors. Instead, it appears that the authors performed a two-hybrid assay to visualize interactions between a small, specific set of proteins. "Screen" should be replaced with "assay", as is currently the case in the Methods section. It is probably also valuable to point out in the text that using a yeast two-hybrid assay to assess interactions between yeast proteins has the caveat that any interactions observed could be indirect, as they may be "bridged" by endogenous yeast proteins.

      We thank the reviewer for raising this point. Our Yeast Two-Hybrid experiments were performed using a small subset of septin-associated proteins rather than as an unbiased screen. Accordingly, we have replaced the term “screen” with “assay” throughout the revised manuscript and in the Materials and Methods section.

      Please refer to line 241.

      We also agree with the limitations inherent to this assay and have now explicitly stated it in the revised manuscript, please see lines 248-255.

      (12) Panel 2H: I do not understand what the middle (as opposed to the top and bottom) blot segment represents. It is labelled "anti-HIS", like the one above it, but it is not associated with any molecular weight/ladder marker and I do not know what other species in the binding reaction in that lane would be recognized by the anti-HIS antibody. Perhaps the top segment is an "input" sample, and below it (in the middle segment) is what was bound to the beads. The figure legend is uninformative in this regard. Also, panel I in this figure is unnecessary to show, assuming that the bands shown in H are what I think they are. The blot makes the point without the need for quantification.

      As suggested by the reviewer, we have added molecular weight markers for each blot panel showing the input and bead-bound fractions. The figure legend has also been updated accordingly to clearly describe the different blot segments and experimental conditions.

      Please refer to figure legend 2H.

      In addition, as suggested by the reviewer, we have removed the quantification graph corresponding to the in-vitro binding assay from the revised manuscript.

      (13) The results in Figure 2H demonstrate that the Gin4-KA1 fragment is not non-specifically "sticky", because it does not bind GST alone, but there is no demonstration that binding by the Hof1 fragment is specific because there is no equivalent negative control for binding:

      We thank the reviewer for this suggestion. We repeated the in-vitro binding assay using 6His-bdSUMO as a negative control alongside 6His-bdSUMO-Gin4<sup>KA1</sup> to demonstrate binding specificity of the Hof1 fragment. The Hof1 N-terminal F-BAR fragment did not pull down the control 6His-bdSUMO fragment but specifically pulled down 6His-bdSUMO-Gin4<sup>KA1</sup> under identical experimental conditions (Fig. 2H), confirming the specificity of the interaction between Hof1 F-BAR domain and Gin4KA1.

      The corresponding text and results have been updated in the revised manuscript.

      Please refer to Fig. 2H and lines 259-263.

      (14) Line 257-258: "Localisation via Hsl1 is necessary to rescue the morphological defects exhibited by Δelm1 cells partially": what does "partially" refer to here? To the rescue, or the defects?

      The term “partially” refers to the extent of rescue. Specifically, elongated cell morphology was rescued in 63.75% of the elm1Δ cell population, rather than in all cells, upon artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP.

      (15) Lines 292-293: "can restore the morphological defects": this wording is unclear. "Restore" means "return to a former condition", which in this case would be normal cellular morphology, not defective cellular morphology. Similarly, see line 304: "While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering": presumably the proper localization was restored, not the mislocalization:

      In Lines 292-293, by phrase “can restore the morphological defects” was intended to indicate rescue of the elongated/clumped morphology associated with gin4Δ cells upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP. We have now rephrased this sentence as: “can rescue the elongated/clumped phenotype exhibited by gin4Δ cells”.

      Please refer to line numbers 481-482.

      Similarly, in Line 304, the statement “While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering” referred to rescue of the Myo1 mislocalization phenotype observed in gin4Δ cells. We have rephrased this sentence as: “However, Myo1-3xmCherry localization was restored to the bud neck upon artificial tethering of Elm1-GFP in gin4Δ cells”.

      Please refer to lines 506-508 in the revised manuscript.

      (16) Lines 295-296: "We find that Elm1-GFP tethering via Shs1-GBP, Bud4-GBP, and Hsl1-GBP" this should be "or", not "and":

      Thank you for pointing this out. We have now corrected the text accordingly.

      Please refer to line number 485.

      (17) Lines 311 and 312 refer to "Inn1-3xmCherry lifetime" but previously Inn1 residence time was measured. Since fluorescence lifetime is a distinct kind of measurement/ assay, it seems important to clarify here what kind of experimental data are being referred to:

      The term “Inn1-3xmCherry lifetime” was intended to describe the residence time of Inn1 at the cell division site, defined as the interval between the initial appearance of the Inn1 fluorescence signal and its complete disappearance during cytokinesis. For clarity and consistency, we have replaced the term “lifetime” with “residence time” in the revised manuscript.

      Please refer to the line numbers 511 and 512.

      (18) Discussion: "Septins are considered as the fourth cytoskeletal elements due to their extensive structural and functional diversity." This sentence is confusing, as it seems to imply that what defines a protein as being "cytoskeletal" is structural and functional diversity rather than anything to do with forming filaments, etc. The rest of this first paragraph of the Discussion also sounds like a summary of the background, is quite redundant with the Introduction, and should be shortened:

      We thank the reviewer for this suggestion. We have rewritten the first paragraph of the Discussion to improve clarity, reduce redundancy with the Introduction, and better emphasize the main findings of the study.

      Please refer to the line numbers 538-552 in the Discussion section.

      (19) Lines 365-366: "Gin4 and Hof1 are synthetic lethal": this should be revised to "gin4∆ and hof1∆ are synthetic lethal". This is a good place to point out that standard yeast nomenclature inserts the ∆ symbol after the gene name, not before (as is done in E. coli genetics, for example):

      We thank the reviewer for this suggestion. We have replaced “Gin4 and Hof1 are synthetic lethal” with “gin4Δ and hof1Δ are synthetic lethal” in the revised manuscript.

      Please refer to line number 576.

      In addition, we have now consistently placed the Δ symbol after the gene throughout the manuscript in accordance with standard yeast nomenclature.

      (20) Line 399: "We also performed an extensive GFP-GBP screens": again, here "screen" implies that a large collection of genes/proteins were assayed, perhaps in an unbiased way, which does not accurately portray what was actually done, which was an extensive tethering study using GFP-GBP:

      We thank the reviewer for this suggestion. We have replaced the term “GFP-GBP tethering screen” with “GFP-GBP tethering assay” throughout the revised manuscript. In addition, we have included a brief description of the specificity and functionality of GBP nanobody and its application in the GFP-GBP tethering strategy, extensively used in this study.

      Please refer to line numbers 326-336.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Analysis of morphological defects of elm1∆ does not directly reflect the defects in cytokinesis and septin organization. For example, the deletion of SWE1 rescues the morphological defects of elm1∆ cells but not the cytokinesis or septin mislocalization (Bouquin et al., 2000). Considering this, the authors should address whether artificial tethering of Hsl1 to Gin4 in elm1∆ cells rescues the septin and cytokinesis defects or just the morphology. Is the role of Hsl1 in cytokinesis dependent on its role in the morphogenesis checkpoint? Can the authors comment on how much the defects observed by Hsl1 tethering to the bud neck may be a result of bypassing the morphogenesis checkpoint?:

      We thank the reviewer for this important point. We performed time-lapse imaging of Cdc3-mCherry in strains where Gin4-GFP partially rescued the elongated phenotype of elm1Δ cells (63.75%) when tethered to the bud neck via Hsl1-GBP. Under these conditions, 64.29% of cells showed rescue of Cdc3-mCherry mislocalization, and Gin4 localization itself was restored to the bud neck in 57.85% of cells. We also examined Myo1-ymScarletI dynamics, while 76.92% of untethered elm1Δ cells displayed Myo1 mislocalization to the bud cortex, this was reduced to 11.36% upon Gin4-GFP tethering via Hsl1-GBP. Together, these results indicate that the morphological rescue observed in elm1Δ cells is accompanied by restoration of normal septin organization and AMR dynamics.

      Previous work (Bouquin et al., 2000) [4] showed that Swe1 deletion rescues cell elongation in elm1Δ cells without restoring septin organization. Consistent with this, we found that 60.66% of elm1Δ swe1Δ cells exhibited a round morphology, but tethering of Gin4-GFP to the bud neck via Hsl1-GBP in elm1Δ swe1Δ background did not further enhance morphological rescue. These results suggest that the rescue of cellular morphology observed in our tethering experiments may, atleast in part, depend on Hsl1-mediated regulation of the morphogenesis checkpoint.

      Importantly, despite the lack of additional morphological rescue, a clear restoration of septin localization was observed when Gin4-GFP was tethered to the bud neck via Hsl1-GBP in elm1Δ swe1Δ cells. Overall, these results suggest that while Hsl1-dependent morphogenesis checkpoint regulation may contribute to cell shape rescue, the restoration of septin organization is independent of Hsl1’s function in morphogenesis checkpoint and instead reflects a direct requirement for Gin4 and Hsl1 at the bud neck.

      Please refer to Figures 5, 6, and S6 of the revised manuscript for these additional data.

      (2) As suggested by the authors, the interaction of the Gin4-KA1 domain with the FBAR domain of Hof1 may explain the cytokinesis-related functions of Gin4. As an orthogonal approach, how does KA1 domain deletion of Gin4 affect cytokinesis and Hof1 bud neck localization?

      We thank the reviewer for this suggestion. We first examined the bud neck localization of Gin4-ka1Δ-GFP in comparison with full-length Gin4-GFP. We observed that the localization kinetics of Gin4-ka1Δ-GFP were significantly altered relative to the full-length protein, with reduced recruitment and earlier removal from the bud neck. We also analyzed the localization kinetics of Hof1-mNG in both gin4-ka1Δ and gin4Δ cells. Our results show that Hof1-mNG displays increased residence time and altered accumulation kinetics during cytokinesis in both genetic backgrounds. Thus, loss of the KA1 domain phenocopies loss of the full-length Gin4 and is consistent with disruption of the physical interaction between Gin4 and Hof1.

      Please refer to Figure S4 for these results in the revised manuscript.

      (3) The authors state that "Elm1 and Kcc4 were present at lower abundance at the bud neck (Fig S1A-D) compared to the higher abundance of Gin4 and Hsl1, as observed in their fluorescence intensities (Figures S1B-C) ". This is not evident in the figures. The authors should show a quantification of how they judged abundance at the bud neck:

      We thank the reviewer for this question. Quantification of septin kinase fluorescence intensity at the bud neck was performed using the established protocol for measuring protein accumulation kinetics described by Okada et. al. 2020 [5]. Time-lapse imaging for kinetic analysis of septin-associated kinases during bud emergence shown in Fig. S1A-S1D was carried out using a point-scanning confocal microscope with a 100×oilimmersion objective. Different laser intensities were required because the fluorescence signals of Kcc4 and Elm1 were comparatively weak and not readily detectable above cellular background under the imaging conditions used for Gin4 and Hsl1. The images shown in Fig. S1A-S1D are therefore displayed using differential contrast settings to facilitate visualization. We have now explicitly clarified this in the figure legend.

      A more direct comparison of septin kinase abundance at the bud neck is now provided in Fig. S1E-F, where localization kinetics during the HDR transition/septin remodelling stage were captured using a laser-scanning spinning-disk microscope under similar imaging conditions. We have additionally included raw fluorescence intensity profiles during the HDR transition to better illustrate the relative abundance of these kinases at the bud neck during cytokinesis.

      Please refer to Figures S1E and S1F in the revised manuscript for the updated images and quantitative analyses.

      (4) How did the authors determine G1 and M-phase in the experiments shown in Figures S1A-D? Can the authors mark these phases on the timelapse images? Also, How do the authors explain the different behaviour of Cdc3 in graphs S1A-D among different strains?

      We thank the reviewer for this comment. Cell cycle stages were initially inferred based on bud size, where kinase accumulation at the bud neck correspond to bud emergence (small bud, G1), and kinase disappearance coincided with septin splitting (large bud, M phase). However, we agree that accurate assignment of cell cycle stages would require specific cell cycle markers. To avoid confusion, we have removed the cellcycle-specific stage assignments from the Results section and describe the kinetics relative to t=0 (bud emergence).

      The differential dynamics observed in the Cdc3-mCherry kinetic profiles likely reflect heterogeneity within the the cellular population. To address this, we combined the normalized fluorescence intensity profiles of Cdc3-mCherry from the strains expressing GFP-tagged septin kinases during bud emergence and have included this data as reference (Author response image 1).

      Author response image 1.

      Plot showing spatiotemporal kinetics of Cdc3-mCherry in strains expressing either Elm1-GFP, or Gin4-GFP, or Hsl1-GFP, or Kcc4-GFP.

      (5) Forced tethering based experiments are one of the key sets of experiments for this work, but it is difficult to have a comprehensive understanding of all the data considering how large the data set is and how dispersed it is in the supplemental and main figures (Figures 3-4-5 and Figures S4-S5-S6). It would be helpful to provide a table summarizing the tested forced tethering’s and the phenotypic outcome in the tested yeast strains (Wt/mutant):

      We thank the reviewer for recognising the extensive dataset generated from the artificial tethering experiments and for suggesting the inclusion of a summary table. We have now added Table 1, which summarizes the proteins used in the GFP-GBP artificial tethering experiments, their genetic backgrounds, the total number of cells quantified across three independent replicates, and the phenotypic outcomes associated with bud neck tethering under each condition.

      Please refer to Table 1 and lines 342, 361, 451, 453, 456 and 492 in the revised manuscript.

      (6) In the introduction section, it would help the reader to provide more information on already known molecular roles of septin-associated kinases in septin organization and AMR. Later in the results section (i.e. Figure S1, S2, and S6), the authors extensively explain and show data that independently corroborate some earlier findings, which makes it difficult for the reader to distinguish novel findings from the repeated findings. I suggest shortening the text for the corroborative results, which will help to put more emphasis on their novel findings:

      We thank the reviewer for this suggestion. We have now included the canonical roles of these four septin-associated kinases in the Introduction section of the revised manuscript.

      Please refer to line numbers 79-85.

      We have also revised sections describing corroborative findings and explicitly cited previous studies wherever relevant in the Results section to better distinguish previously established observations from the novel findings presented in this work.

      Other minor comments:

      (1) In Figure 1D-1F, also show the data for hls1∆ and kcc4∆ - which are shown in S3AC in the current version:

      We have now included the Inn1-mNG residence time in hsl1Δ and kcc4Δ cells, alongside elm1Δ and gin4Δ in Fig. 1E.

      Please refer to Fig. 1E in the revised manuscript.

      (2) The authors should be more careful in interpreting their negative Y2H data in Figure 2F.

      We have now explicitly discussed the caveats and inherent limitations of the Yeast Two-Hybrid assay in the manuscript.

      Please refer to line numbers 248-255.

      (3) Please provide quantification for Figure 5A, Figure S6I:

      We have added quantitative analyses showing rescue of Cdc3-mCherry and Myo13xmCherry mislocalization in gin4Δ and gin4Δ hsl1Δ strains upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP.

      Please refer to Figures 8E, 8F, and S10B in the revised manuscript.

      (4) In Figure S6: label is missing "∆" in front of hsl1:

      We thank the reviewer for pointing out this error. We have corrected the labels accordingly.

      (5) As common consensus on yeast gene nomenclature, I suggest the use of "gene∆" instead of "∆gene":

      We thank the reviewer for this suggestion. We have now consistently placed the Δ symbol after deleted gene names throughout the manuscript in accordance with standard yeast nomenclature.

      (6) Lines (535-536): min(distribution) and max(distribution) in the formula is confusing. Clarify it or if possible use "minimum value", "maximum value" instead:

      We thank the reviewer for this suggestion. We have replaced the term “distribution” with “value” in the formula for protein accumulation kinetics analysis.

      Please refer to the updated formula in the Materials and Methods section (Lines 752753).

      (7) In line 86, "Dynamics of Septin-associated kinases and their defects during the cell cycle": Change the title as it is not clear what is meant by "their defects" given the discussed results under this title:

      We thank the reviewer for this suggestion. We have revised the section heading from “and their defects during the cell cycle” to “and defects associated with their deletions”.

      Please refer to line numbers 99-100.

      (8) On the Hof1-mNG image (Fig2C), show the line used for the line scan profile. Additionally, a similar line-scan profile could be useful in Figure S3G:

      As suggested by Reviewer 3, we removed the line-scan analysis from Figure 2 in the revised manuscript because Hof1 ring organization showed substantial heterogeneity across cells, making line-scan analysis difficult to interpret reliably.

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The % phenotype units are terrible. With no explanation, we do not really know whether they represent the percentage of cells that have a particular phenotype, or whether they correspond to a metric that measures some deviation between normal and extreme phenotypes. I would strongly recommend using precise quantitative metrics systematically (i.e. intensities, aspect ratios, division times, etc.) to properly quantify phenotypes:

      We thank the reviewer for suggesting the inclusion of precise quantitative metrics to assess phenotypic differences in the GFP-GBP tethering experiments. In the revised manuscript, we adopted quantification workflows that have been extensively validated and widely used in the literature, including those reported by Marquardt et al., 2024 (Fig. 7B and 7D) [2] from the Bi Lab. In response to the reviewer’s suggestion, we have now incorporated additional quantitative measurements, including cell area and aspect ratio (defined as the ratio of the cell’s major axis to the minor axis), for the experimental datasets presented in the manuscript.

      In the main figures, we now include aspect ratio quantification, while additional parameters are provided for the reviewer’s reference. We also quantified the fluorescence intensity of tethered proteins at the large bud neck and present these data together with the aspect ratio analysis in Figure 4 for elm1Δ cells in which Gin4GFP is artificially tethered to the bud neck via Hsl1-GBP. These quantitative analysis corroborates our qualitative observations and further strengthens our conclusions. Please refer to Figures 4D and 4E as representative examples.

      We have also changed the y-axis labels throughout the revised manuscript. For example, the y-axis in Fig. 7B is now labelled as “Cells exhibiting round morphology (%)”. Please refer to Fig. 4H, 4K, 7C, 8D, S5G, S6C, S7C, S9G and S9J for inclusion of aspect ratio quantification. Statistical analyses for the represented graphs were performed using Kruskal-Wallis nonparametric test, (N=3, n>150 cells/strain) (*: p<0.05, **: p<0.01, ****: p<0.0001, ns: p>0.05).

      Author response image 2.

      (2) Some quantitative analyses were performed manually where simple automated analysis should be performed to provide unbiased, accurate quantification:

      We fully agree with the reviewer that automated image analysis approaches, such as segmentation-based methods, are generally preferred for minimizing bias in morphological quantification. However, elm1Δ and gin4Δ cells exhibit severe phenotypes, including pronounced elongation and clumping, which makes reliable automated segmentation technically challenging for accurate quantification of parameters such as aspect ratio and cell area. For this reason, we used manual annotation for these analyses, as this approach enabled accurate delineation of individual cell and reliable measurements of morphological parameters such as cell area and size across the datasets despite being more time-consuming.

      Please refer lines 769-775 in the Materials and Methods section.

      (3) The "tethering" data also lack clear quantification. The authors should properly quantify the average intensity of Hsl1-GFP at the bud neck in each condition and correlate the results with cell aspect ratios or any other relevant parameters. For example, when comparing elm1null and elm1null Bud4-GBP with elm1null Kcc4-GBP, the visual impression is that as much Hsl1-GFP protein is recruited to the bud neck, whereas the phenotypes are dramatically different:

      We thank the reviewer for pointing this out. We have revised the image representation to facilitate clearer interpretation of the tethering experiments. In addition, we performed the key GFP-GBP tethering experiments using GBP-ymScarletI constructs, allowing direct visualisation of both the GFP-tagged protein and the GBP-tagged partner at the bud neck following tethering. We have included quantitative analyses of cellular morphology, raw fluorescence intensities of GFP-tagged proteins at the large bud neck, and the corresponding aspect ratio measurements for these updated datasets (Author response images 3, 4, 5). We have also included a summary table (Author response table 1) compiling these quantitative results for easier comparision.

      (4) I am very confused by Figure 3 which shows normal localization of Gin4-GFP in elm1null cells and seems to contradict other claims in the manuscript. This is very problematic for the interpretation of most of the "tethering" data:

      We understand the reviewer’s concern and have replaced the representative images of Gin4-GFP in elm1Δ cells in Figure 4B. Although Gin4-GFP is initially recruited to the presumptive bud neck during bud emergence in elm1Δ cells, it subsequently becomes mislocalized to the bud cortex during early cell cycle stages, resulting in reduced bud neck localization. As the cell cycle progresses, the Gin4-GFP signal at the bud neck decreases substantially in elm1Δ cells while remaining stable in wild-type cells until its departure prior to septin HDR remodelling (Fig. S5A-S5D).

      (5) Figure 6 is neither explained in the text nor in its legend. Could the authors explain the model and offer a comprehensive picture of the current knowledge?

      We have simplified the representative model to more clearly distinguish previously established knowledge from the findings presented in this study. Based on our results, We propose that Hsl1 functions both downstream of and in coordination with Elm1 and Gin4 to regulate septin stability and the timely execution of cytokinesis. Deletion of Elm1 disrupts the normal localization and crosstalk between Gin4 and Hsl1 at the bud neck, leading to septin mislocalization and misregulation of AMR dynamics, thereby revealing a previously uncharacterized role for Hsl1 in cytokinesis. The Results section has also been updated to reflect the revised model.

      Please refer to Figure 9 and lines 520-536.

      (6) Could the authors provide information about the double/triple mutant kinase phenotypes to clarify the overlap of functions among them?:

      Barral et al., 1999 [6] reported that individual deletions of Hsl1 and Gin4 result in mild cytokinetic defects, whereas deletion of Kcc4 does not produce any striking phenotype compared to wild-type cells. In contrast, the hsl1Δ gin4Δ kcc4Δ triple mutant remains viable but exhibits severe morphological abnormalities, including branched chains of elongated cells with defective cell separation. These mutants also display aberrant septin organization at the bud neck, characterized by irregular patch-like structures. Analysis of double mutants (hsl1Δ gin4Δ, gin4Δ kcc4Δ, and hsl1Δ kcc4Δ) revealed intermediate phenotypes between the corresponding single and triple mutants, with the hsl1Δ gin4Δ combination showing the strongest defects. Together, these findings suggest that the Nim1-related kinases function redundantly to regulate Swe1 activity and maintain septin architecture at the bud neck.

      Further supporting this model, Bouquin et al., 2000 [4] demonstrated that Elm1 operates independently of the Nim1-related kinases in controlling septin organization. The hsl1Δ gin4Δ kcc4Δ elm1Δ quadruple mutant exhibits severe septin localization defects and strong growth defects, in contrast to the elm1Δ single mutant, which primarily displays septin mislocalization from the bud neck to the bud cortex. These findings indicate that the combined activity of these kinases is essential for proper septin anchorage at the division plane and for assembly of the septin ring.

      Importantly, the progressively stronger phenotypes observed in double, triple and quadruple mutants also suggest that these kinases retain partially specialized functions at the bud neck. Consistent with this framework, our results support a model in which Nim1-related kinases function redundantly to regulate septin architecture and cytokinesis, likely through modulation of the AMR machinery. Because the localization of these kinases appears interdependent, as reported previously (Marquardt et al., 2020; Marquardt et al., 2024) [2,7] and corroborated by our findings, interpretation of mutant phenotypes remains complex and future studies will be required to delineate their individual contributions more precisely.

      Minor points:

      (1) Abbreviations are not defined in the manuscript.

      We have now expanded and defined all abbreviations throughout the manuscript.

      (2) Some of the writing in the figures is too small. Please make sure that a minimal size of letters/numbers is respected:

      We thank the reviewer for raising this issue. We have enlarged the figure labels and axis labels throughout the revised manuscript to improve readability.

      (3) Figure 2D. I am not sure that the line scans bring any useful information as the rings in mutant cells are quite heterogenous. This panel is also not cited in the text. Please make sure that every panel is cited at least once:

      We agree with the reviewer regarding the heterogeneity observed in the Hof1 ring organization in mutant cells and have therefore removed the line-scan analysis from the revised manuscript.

      (4) Knocked-out genes are written incorrectly. Please use the usual yeast nomenclature:

      We thank the reviewer for this suggestion. We have corrected the nomenclature for all deleted genes throughout the manuscript in accordance with standard yeast nomenclature.

      Author response image 3.

      Artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Gin4-GFP via Shs1-GBP-ymScarletI, Hsl1-GBP-ymScarletI, Bud4-GBPymScarletI and Bni5-GBP-ymScarletI in elm1Δ cells. DC*=Differential contrast. Scale bar5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple-comparison test (**: p<0.01, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=397, elm1Δ: n=408, elm1Δ-Shs1-GBPymScarletI: n=467, elm1Δ-Hsl1-GBP-ymScarletI: n=462, elm1Δ-Bud4-GBP-ymScarletI: n=401 and elm1Δ-Bni5-GBP-ymScarletI: n=317 cells). (C) Quantification of aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, n>170 cells/strain). (D) Graph depicting the raw fluorescence intensity of Gin4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=166, elm1Δ: n=177, elm1Δ-Shs1-GBP-YmScarletI: n=176, elm1Δ-Hsl1-GBPymScarletI: n=188, elm1Δ-Bud4-GBP-ymScarletI: n=185 and elm1Δ-Bni5-GBP-ymScarletI: n=151 cells).

      Author response image 4.

      Artificial tethering of Hsl1-GFP to the bud neck via septins or Nim1-related kinases rescues cellular morphology in elm1Δ cells. (A) Representative images showing the relocalization of Hsl1-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI and Gin4GBP-ymScarletI. Scale bar-5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=508, elm1Δ: n=418, elm1ΔShs1-GBP-ymScarletI: n=535 and elm1Δ-Gin4-GBP-ymScarletI: n=482 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (*: p<0.05, ****: p<0.0001), (N=3, n>165 cells/strain). (D) Quantification of raw fluorescence intensity of Hsl1-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=156, elm1Δ: n=166, elm1Δ-Shs1-GBP-ymScarletI: n=169 and elm1Δ-Gin4-GBP-ymScarletI: n=165 cells.

      Author response image 5.

      Targeted localization of Kcc4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Kcc4-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI, Hsl1-GBPymScarletI and Gin4-GBP-ymScarletI. Scale bar-5µm. (B) Quantitative analysis representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), oneway ANOVA Tukey’s multiple-comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=443, elm1Δ: n=489, elm1Δ-Shs1-GBP-ymScarletI: n=312, elm1Δ-Hsl1-GBP-ymScarletI: n=563 and elm1Δ-Gin4-GBP-ymScarletI: n=337 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (****: p<0.0001, ns: p>0.05), (N=3, n>165 cells/strain). (D) Quantification for the raw fluorescence intensity of Kcc4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (**: p<0.01, ***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=163, elm1Δ: n=168 elm1Δ-Shs1-GBP-ymScarletI: n=161, elm1Δ-Hsl1-GBP-ymScarletI: n=172 and elm1Δ-Gin4-GBP-ymScarletI: n=150 cells).

      Author response table 1.

      Summary table showing rescue of elongated morphology in elm1Δ cells upon forced recruitment of Nim1-related kinases via septins or its related kinases tagged with GBPymScarletI.

      Additional changes:

      The graph in Fig. S3D (revised preprint) has been updated to reflect a slight increase in the Chs2-mNG residence time in both the elm1Δ and gin4Δ strains, whereas our previous version indicated a delay only in the gin4Δ strain. Because the elm1Δ strain exhibited a more pronounced phenotype than the gin4Δ strain, we re-examined the analysis. The revised results show that the residence time of Chs2 during cytokinesis is modestly prolonged by approximately 2 minutes in both backgrounds. Accordingly, the graph and statistical analyses have been updated.

      References:

      (1) Mortensen, E.M., McDonald, H., Yates, J., and Kellogg, D.R. (2002). Cell Cycle-dependent Assembly of a Gin4-Septin Complex. Molecular Biology of the Cell 13, 2091-2105. 10.1091/mbc.01-10-0500.

      (2) Marquardt, J., Chen, X., and Bi, E. (2024). Reciprocal regulation by Elm1 and Gin4 controls septin hourglass assembly and remodeling. J Cell Biol 223. 10.1083/jcb.202308143.

      (3) Asano, S., Park, J.E., Yu, L.R., Zhou, M., Sakchaisri, K., Park, C.J., Kang, Y.H., Thorner, J., Veenstra, T.D., and Lee, K.S. (2006). Direct phosphorylation and activation of a Nim1-related kinase Gin4 by Elm1 in budding yeast. J Biol Chem 281, 2709027098. 10.1074/jbc.M601483200.

      (4) Bouquin, N., Barral, Y., Courbeyrette, R., Blondel, M., Snyder, M., and Mann, C. (2000). Regulation of cytokinesis by the Elm1 protein kinase in Saccharomyces cerevisiae. Journal of Cell Science 113, 1435-1445. 10.1242/jcs.113.8.1435.

      (5) Okada, H., MacTaggart, B., and Bi, E. (2021). Analysis of local protein accumulation kinetics by live-cell imaging in yeast systems. STAR Protoc 2, 100733. 10.1016/j.xpro.2021.100733.

      (6) Barral, Y., Parra, M., Bidlingmaier, S., and Snyder, M. (1999 Jan 15). Nim1-related kinases coordinate cell cycle progression with the organization of the peripheral cytoskeleton in yeast. Genes & Development 13. 10.1101/gad.13.2.176.

      (7) Marquardt, J., Yao, L.L., Okada, H., Svitkina, T., and Bi, E. (2020). The LKB1-like Kinase Elm1 Controls Septin Hourglass Assembly and Stability by Regulating Filament Pairing. Curr Biol 30, 2386-2394 e2384. 10.1016/j.cub.2020.04.035.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The extent to which P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood-stage exported proteins tested in liver stages were not exported. An exception is LISP2, which is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, an incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum, it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages, but in liver stages, no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studies the localization of LSA3 in blood stages and uses a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also shows that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 is thorough and mostly overcame adverse issues such as a cross-reactive antibody and the negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable, as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      We thank the reviewer for their perspective on the strengths of the study.

      Weaknesses:

      The cross-reactivity of the antibody, together with the co-infection strategy, prevents reliable assessment of LSA3 localization in liver stages. Despite this, it seems LSA3 is not exported in liver stages, and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. Given that a previous publication found some inhibitory effect of LSA3 antibodies on blood stage growth, a comparison of the growth of the LSA3 disruption clones with the parent would have been very welcome and easy to do. At this point, LSA3 is one more of many proteins exported in blood stages for which the function remains unclear.

      It might be possible to refine some of the conclusions. The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remains largely unknown. The co-infection (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver, but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. The authors write that they do not know if the cross-reactive protein is expressed at that stage. But this should be immediately evident from the mixed WT/mutant infection. If all cells are positive for LSA3, there is a cross-reaction. If about half of the cells are negative, there isn't. In the latter case, the localization shown in the paper is indeed LSA3, and morphological differences between WT and LSA3 disruption could be assessed without additional experiments.

      We thank the reviewer for their comments. While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Significance:

      The conclusion from the paper that "our study presents just the second PEXEL protein so far identified as important for normal P. falciparum liver-stage development and confirms the hypothesized potential of exported proteins as malaria vaccine candidates" is partially misleading. Neither LISP2 nor LSA3 seems to be exported in P. falciparum liver stages, and we can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

      We thank the reviewer for the comment. We would like to emphasize the possibility that proteins localized at the PVM may be considered exported ‘if’ part or all of the protein (eg, a domain) faces the host cell lumen from the hepatocyte. We have not shown this to be the case for LSA3 or LISP2 but that possibility remains open. Nonetheless, LISP2 is exported (by P. berghei liver stages) and LSA3 is exported (by P. falciparum blood stages); both are exported proteins.

      Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif, but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al, PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood stage LSA3 undergoes PEXEL processing, and a portion of the protein is exported into the erythrocyte, where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepsin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also react with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage, where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      This study provides the first direct analysis of LSA3 function by reverse genetics, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage, but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectable export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      We thank the reviewer for their feedback regarding the strengths of the paper. 

      Weaknesses:

      A previous study reported that anti-LSA3 antibodies inhibit blood-stage growth, suggesting a role for LSA3 during erythrocyte infection. While the authors carefully evaluate the LSA3 mutant in mosquito and liver stages, the impact on blood stage fitness is not tested. While the knockout shows LSA3 is not essential in the blood stage, its importance during erythrocyte infection remains unclear.

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage, and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3C reactivity at this stage could not be determined, and the localization of LSA3 in the liver stage remains somewhat unclear.

      We thank the reviewer for their comments. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      We thank the reviewer for their comments.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      We understand the reviewer’s points. We agree that there is no evidence provided that LSA3 is targeted beyond the PVM; whether any of the protein faces the hepatocyte cytosol is unknown (and challenging to conduct) but this possibility remains plausible. LISP2- and LSA3deficient liver stages are less fit than parental controls and thus we stand by the conclusion that they are the two so far identified P. falciparum PEXEL proteins that are important for liver-stage development.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 163 says: "Altogether, this demonstrates that LSA3 is important but not critical for blood stage growth of P. falciparum": this is based on the cited Morita et al., 2017. However, previously LSA3 was considered dispensable based on a knock out in 3D7 (Maier et al., 2008; PMID: 18614010). Given that the authors generated a mutant for this work, it would be straightforward to test growth and clarify the importance of LSA3 in blood stages. If important, the analysis of the location and transport of LSA3 in blood stages would immediately become more relevant.  Maybe the data for this is already in the paper: the number of stage V gams was similar between mutant and control (Figure 4A). If this was calculated from the total number of asexual starting parasitemia, it includes blood stage growth, and it can be assumed that there is no growth defect in the mutant in the blood stages. If the number of stage 5 gams was calculated from the number of committed schizonts/rings, nothing can be said about blood stage growth, and asexual blood stage growth should be tested in specific experiments.

      We thank the reviewer for raising the function of LSA3 in blood stages and agree it was an obvious omission, though for good reason - a separate, collaborative study was underway. While this eLife preprint was in revision, our accompanying manuscript on the blood stage was published, showing the characterization of our NF54 DLSA3 mutant during blood-stage growth (PMID:41135800). The findings are now summarized and the citation included in the revised version of this preprint.

      Manuscript line 105: "although, notably, functional characterization of lsa3 deletion mutants has not yet been reported to confirm an important function": at least in blood stages, it was reported to be dispensable, see above. The corresponding study (Maier et al., 2008, PMID: 18614010) could be cited in that context. 

      The citation of PMID18614010 and 39913589 have now been added and we thank the reviewer.

      (2) Some questions central to the conclusions of this paper remain because it was unclear whether the serum did indeed detect LSA3 in the liver or not. It would be easy to check if all cells from the WT/Mutant mix experiment show LSA3 signal (this would mean it cross-reacts) or if only about half are positive (the mutants would be negative if there is no cross-reaction). This would be important to mention for Figure 5 because, at present, it is not known that what is labeled by the LSA3-C antibody in these images is (only) LSA3. 

      We thank the reviewer for this point and completely understand. We did check this via microscopy of liver sections co-infected with LSA3 mutant and control liver-stage parasites as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5-day post-infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (3) It is also unclear which parasites were imaged in Figure 5. The text of the results states that NF54 liver stages were used, but later: "As we employed a co-infection strategy to assess the essentiality of LSA3 versus NF54 in mice, we could not perform IFAs on individually infected mice in this study to validate the specificity of LSA3-C at the liver-stage". The legend says NF54 sporozoites on day 5 post-infection were used. I suspect it was a WT/mutant mix, in which case the above applies, and in the absence of cross-reactivity, half of the cells should be LSA3-C negative. If this is not the case, the localization in the liver becomes dubious.

      We apologize for the confusion and have corrected this. In Figure 5, we utilized liver sections from NF54-infected humanized mice that were stored at -80 C from a previously published study (McConville et al, PNAS 2024). Ideally, we would validate the specificity of LSA3 antibodies at the liver-stage using liver sections containing only DLSA3 parasites however the number of mice available was limited and the samples available to us also contained the Control line for qRT-PCR analyses (the co-infection strategy). As mentioned above, we couldn’t distinguish between these two strains by IFA at the time point analysed and this precluded us unequivocally validating the LSA3-C specificity in the liver-stage; however it cannot be excluded that the signal observed at the PVM is indeed LSA3. We are currently focusing research efforts on obtaining more humanised mice to answer this.

      Minor:

      (1) Introduction: Before the part on the PEXEL motifs, there are almost no references; please add references for all statements.

      We have added references.

      (2) Figure 1B is unclear regarding which part of the gene was deleted. The system used would permit a complete gene deletion, but the homology flanks seem to be within LSA3. If parts of the gene are left, the 75 kDa on the western blots might be a degradation product arising from both the truncated and the full-length protein. Please clarify in the sketch exactly where the homology flanks are, with respect to the start and stop of the gene. 

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (3) Line 161: Replace was with were.

      Corrected.

      (4) Figure 2, 224: Why do the authors think LSA3 must be in the luminal leaflet of the PVM as opposed to the outer leaflet of the plasma membrane?

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte.

      (5) Line 245: GFP core "derived from digestion of the reporter in the food vacuole, which confirmed it was secreted from the parasite". I wonder if the amount of GFP "core" really can be used as evidence for secretion, and its amount can be compared between experiments. Did the author quantify this for the full-length protein to get a proportion per sample?

      Use of GFP core to measure defects in P. falciparum GFP reporter secretion has been described previously (for example PMID:23387285 and 35906227). The comparison the reviewer asked for is an interesting and important question: however the control would be to compare the ratio of GFP core to uncleaved in the control lanes as well, which is not possible to do since the full-length protein is digested by plasmepsin V in the native PEXEL versions of the experiments (mLSA3-GFP in the first blot, Vehicle in the second blot) leaving no full-length protein to compare to. It stands to reason that inhibition of N-terminal processing results in less protein removal from the membrane (ER or COPII vesicle or PM) resulting in less secretion out of the parasite for retrograde transport to the food vacuole with cytostomal vacuoles (analogous to plasmepsin II). In the food vacuole, the chimeras are in normal cases digested by proteases back to the GFP core that is resistant to cleavage and evident as GFP core on the immunoblots (PMID:10775264 and 14709539 and 19055692 and 20130643). 

      (6) Figure 3 has the word plasmid in two lanes. In Figure 3E, amend the labelling of the blots.

      We apologize for the formatting error in converting the figures to PDF during the original submission and thank the reviewer for the suggestion. This has now been corrected.

      (7) Lines 266/271/284: "live IFAs", live immunofluorescence assay. Does this mean an antibody was given to living   parasites?

      The correct term is live microscopy and this has been corrected.

      (8) Does Figure 6A fit with the data in Figure 6B? It seems 6B has a milder phenotype than 6A.

      We thank the reviewer for the question. Yes the data directly correspond to each other and are represented in two ways: Panel A shows the qRT-PCR raw data for liver load of each parasite strain per humanized mouse using a scientific scale on the y-axis. Panel B shows that magnitude of the DLSA3 defect as a percentage of the total liver load per mouse:

      % total parasite liver load  = ( strain 1 or strain 2 liver load ) x100

      sum of strain 1 + strain 2 liver loads

      The intent of showing both data is to convey the correct magnitude of the difference in two ways to assist the reader in understanding the true defect, both are accurate and both are statistically significant. In revision we detected mislabelling of humanized mouse 2 and 3 in the original graphs that has now been corrected and we sincerely thank the reviewer for helping us identify this error.

      (9) Line 482: Please add references for this debate. 

      These have been added.

      Reviewer #2 (Recommendations for the authors):

      Major Comments: 

      (1) In general, the authors have taken care not to overstate conclusions from their study. Nonetheless, while not technically inaccurate, the title might misleadingly suggest LSA3 is exported in the liver stage (this was my initial impression on reading it until I looked at the data). I suggest the authors revise the title to avoid confusion by clarifying that export was only observed in the blood stage.

      We sincerely appreciate the reviewer’s point. As this article was posted as a preprint that has now been cited several times, we have carefully weighed the comment and in the end decided to retain the current title for the above reason.

      (2) While the ability to generate the ∆LSA3 parasites clearly shows that the protein is not essential in the blood stage, the impact on parasite fitness is never tested but simply assumed (for instance, in lines 163-164: "...this demonstrates that LSA3 is important...for blood-stage growth..."). Do the ∆LSA3 parasites have a fitness defect in the blood stage consistent with the previous GIA data that would support this claim? Since the rabbit anti-LSA3-C antibodies produced by Morita et al did not have GIA activity against the blood stage, it is possible that the GIA observed with the human and mouse antibodies might have been due to reactivity with a different protein. If ∆LSA3 does cause a fitness defect, it would be interesting to know if the endogenous GFP-tagged line, which alters protein trafficking/membrane association, also produces this effect.

      We agree with the reviewer and would like to clarify that this omission was not intended to create confusion but was by design, due to a separate collaborative study that was underway to address such questions. While this eLife preprint was in revision, our accompanying manuscript on characterising NF54 DLSA3 at the blood stage was published (PMID:41135800). The findings are now summarized and the citation included in the revised version of this eLife preprint. In sum, LSA3 is not critical for erythrocyte invasion but its deletion perturbs the rate and efficiency of merozoite invasion, at the step(s) of resealing of the PVM/host cell, resulting in aberrant accole forms that protrude from the infected erythrocyte.

      (2) Figure 1D: While the images are compelling and I don't doubt the claim that LSA3 is exported in the blood stage (also supported by the fractionation/Pk experiments), the authors should provide quantification of the difference in exported signal between the WT and ∆LSA3 parasites in these IFAs to rigorously support this conclusion. Also, please include details about how many independent experiments are represented by the microscopy data throughout the manuscript (Figures 1, 2, 3, and 5).

      We understand the reviewer’s request and wish to indicate that the export signal was absent in all cells infected with DLSA3 that was imaged. The microscopy performed was from n=2-3 experiments except for Figure 5 which was from n=1 humanized mouse per time point in which multiple EEFs from the liver were imaged. This has been indicated in the figure legends. 

      (3) Careful inspection of the z-series images in Figure 5A shows that most of the LSA3-C signal seen outside the PVM (beyond the boundary delineated by EXP1) is closely associated with DAPI puncta, suggesting these are merozoites. Together with the prominent gap in the EXP1 signal, this suggests the schizont has already ruptured. Thus, anti-LSA3-C signal beyond the PV seems best explained as coming from merozoites or other material released by PV rupture, not from export across the PVM, and this should be added to the text in place of comments about localization to PV extensions or potential export (lines 358-359, 422-423).

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have added the comment as requested.

      Minor Comments:

      (1) The authors may want to denote the disordered repeat region in the LSA3 schematic in Figure 1A that is mentioned in the text.

      We have added the residue boundaries of the predicted domain from AlphaFold into both the schematic and the text and included a link to the LSA3 pages in PlasmoDB and

      AlphaFold in the Methods section.

      (2) The authors use rabbit anti-LSA3-C antibodies previously generated by Morita et al. These polyclonal antibodies were raised against a recombinant C-terminal region of LSA3 (residues 750-1433), but the schematic in Figure 1A indicates the antibodies recognize a smaller region between residues 1154-1433. Please adjust the figure accordingly, or if this is not the same antiLSA3-C antibody reported by Morita, please provide details about its production.

      The figure is corrected.

      (3) The authors use Alphafold to identify a region of LSA3 with similarity to the substrate binding domain of DnaK, but the data is not shown. Please include the Alphafold prediction in supplementary figures and provide information about how the predicted structural homology was determined.

      We have added a link to the AlphaFold page for PF3D7_0220000 in the methods.

      (4) The schematic in Figure 1B indicates that the DHFR cassette was inserted at an internal site within the lsa3 gene. If this is the case, it seems possible that an N-terminal portion of the protein is still expressed, but I was unable to find details about the boundaries of the homology flanks to determine the precise insertion site. Please clarify the knockout strategy and indicate the specific insertion site.

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (5) Line 162: I think this should read "antibodies that react with LSA3 were...".

      Corrected.

      (6) Figure 1D: The merge with the transmitted light channel is missing for the third panel in the ∆LSA3 IFAs. Also, please define the scale bar length in the legend.

      Corrected.

      (7) Lines 744-746: The IFA fixation panel order description (top, bottom) in the Figure 2A legend is reversed from what is shown in the actual figure. Also, please define the scale bar length. 

      Corrected.

      (8) Lines 184-186: Since the fractionation/PK protection assays suggest most of LSA3 is in the PV, it would be interesting to know if the strong peripheral/PV signal observed in the PFA-fixed IFAs in Figure 2A is also present in the ∆LSA3 parasites, or is this non-specific? 

      Thank you for the suggestion. We agree this would be an interesting result to know but do not have the capacity at the present time.

      (9) Lines 219-225: It is unclear to me why these results are interpreted to suggest that the majority of LSA3 is peripherally associated with the luminal leaflet of the PVM. Wouldn't an integral membrane configuration in the PVM (with the C-terminus facing the host cytosol) or PPM (with the C-terminus facing the parasite cytosol) also account for the data? Adding a carbonate extraction would help clarify this point.

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM-associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte. If the question is whether LSA3 is an integral PVM protein, we agree that use of carbonate in the future would answer that question.

      (10) Figures 3D and E: There are some problems with some of the text wrapping in these panels.

      We apologise, this was a formatting issue as the manuscript was converted to PDF.

      We have corrected this error.

      (11) Line 422-423: In fact, the Z-sections shown in Figure 5 appear to indicate that the LSA3-C signal is predominantly located within the parasite, not at the PVM.

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have corrected the final conclusion to be more accommodating of this.

      (12) Lines 468-470: Since cross reactivity of anti-LSA3-C is substantial in the blood stage but was not defined in the liver stage by analysis of unmixed infections, how do the authors know that they were not observing ∆LSA3 parasites in their IFAs? I think what they mean here is that parasites lacking anti-LSA3-C reactivity were not observed, which is an important distinction.

      The reviewer is correct and this has been corrected.

      (13) Lines 478-479: The authors should also mention that the P. berghei PEXEL proteins evaluated in Fougere et al are exported in the blood stage, similar to LSA3. Moreover, other studies have shown something similar for additional endogenous PEXEL proteins or reporters in P. berghei (PMIDs 22329949, 26347246, 34956312).

      We have added the additional text regarding export into the infected erythrocyte and the reference to IBIS1.

      (14) Line 491: The data here don't support that LSA3 is "required" for liver stage development, only that it is important to it. Since the authors have not defined the cross-reactivity of anti-LSA3C in unmixed infections, it is not clear that ∆LSA3 parasites are arrested early in the liver stage, only that they show a reduced number of genome copies relative to the parental control. 

      We have amended the sentence to “required for normal liver stage development”.

      (15) Line 530: I think NGF54 should be NF54.

      Corrected.

      Reviewer #3 (Recommendations for the authors):

      (1) Antibody specificity in liver stage IFA experiments:

      The specificity of the anti-LSA3 antiserum (LSA3-C) used in liver stage IFA is not fully convincing. While the KO parasites were used effectively to validate specificity in blood stages, the same is not true for liver stages. 

      (a) It is essential to repeat IFA with ΔLSA3 parasites in liver stage infections to determine whether the observed PVM staining is truly specific.

      We appreciate the reviewer’s point, however at a cost of over $5000 per humanized mouse, we do not have the capacity to conduct this experiment at the present time. We highlight that, as the blood stage IFAs confirmed the specificity of LSA3-C for LSA3, the possibility remains open that LSA3 is specifically recognized at the PVM.

      (b) If the antibody is the same polyclonal serum used in Morita et al. (2017), why did the authors not employ a monoclonal antibody, which they presumably have access to and which would provide greater specificity? 

      We have included new data confirming that LSA3 is exported using LSA3-T, in addition to LSA3-C.

      (c) Given that rabbit antisera often show non-specific staining at the PVM in liver stage parasites, co-localization with PVM markers is not sufficient. Inclusion of the ΔLSA3 parasites in liver stage IFA is critical. It will also show whether there is any cross-reaction of the antiserum in liver stage parasites, as seen by IFA for blood stage parasites. 

      We thank the reviewer for their feedback.

      (d) To validate the serum further, the authors should infect HC-04 cells in vitro with GFP-LSA3 parasites and stain with LSA3-C to confirm overlap between the tagged protein and the antibody signal.

      We thank the reviewer for their feedback.

      (e) For higher-resolution co-localization, expansion microscopy - now commonly used even in malaria research - would substantially improve the analysis. 

      We thank the reviewer for their feedback.

      (2) The localization of LSA3 in this study differs notably from Morita et al. 2017, who reported localization to dense granules in merozoites and staining in ring-stage parasites at the PVM. 

      (a) The authors confirm DG localization, but they do not examine ring-stage parasites. They should include the IFA of ring stages to clarify whether they can replicate the previous findings.

      We thank the reviewer for their feedback.

      (b) Additionally, the differences in Western blot banding patterns between the two studies should be addressed. Do the authors have an explanation for these discrepancies? 

      We thank the reviewer for their feedback.

      (3) The authors report a ~40% reduction in liver parasite load using qPCR, which is statistically significant. However, this phenotype is modest and should not be interpreted as showing that LSA3 is essential.

      (a) Please avoid terms like "required" or "essential" and instead describe the protein as "contributing to normal development" or "influencing fitness."

      We have used the term “required for normal liver stage development”.

      (b) Since the authors generated liver sections, they should take advantage of these to quantify the number and size of liver stage parasites, which would help determine whether the phenotype reflects fewer infected cells or reduced parasite growth.

      We did check this via microscopy of liver sections, but all mice were co-infected with LSA3 mutant and control liver-stage parasites, as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5 day post infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (c) It would also be valuable to include IFA from singly infected ΔLSA3 livers (rather than co-infected), and possibly at earlier timepoints, to identify the developmental window affected.

      We agree it would be valuable.

      (4) The manuscript suggests that LSA3 may be exported beyond the PVM into the hepatocyte, based on a small number of peripheral puncta.

      (a) This claim is not convincingly supported by the data. The punctate signals shown in Figure 5 are weak and may rather reflect PVM extensions or TVN. In fact, one punctum even overlaps with the DAPI signal (figure 5, middle panel), which raises further doubt about the localization.

      We appreciate the reviewer’s careful eye and caution and have added the comment regarding DAPI.

      (b) Given the lack of KO controls in these liver stage IFAs, the authors should not describe LSA3 as "exported beyond the PVM". The language should be revised to reflect that the protein localizes predominantly to the PVM, and any extra-PVM signal remains unconfirmed and could be non-specific. 

      (c) This is especially important given the well-known tendency of rabbit antisera to produce background PVM staining in liver stage parasites. 

      Corrected.

      (e) In an earlier report (McConville et al, 2024, PNAS), they clearly state that LSA3 is NOT exported beyond the PVM. Actually, the staining in the previous report looks quite different from the images provided for Figure 5. The authors might wish to comment on this. 

      We thank the reviewer for their feedback.

      Minor comments:

      In some sections, the manuscript uses "exported" to refer to trafficking to the PVM. This terminology should be used more carefully and consistently, since "export" often implies translocation into the host cytosol

      We understand that export involves a protein localizing within the host cell and so protrusion through the PVM may also be considered exported, however, we have not confirmed this for LSA3 in liver stages.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. A prevailing model that has been developed over the last five years by work from many laboratories using a variety of biochemical, structural, and microscopic approaches is that HS acts a co-receptor for SARS-CoV-2; its binding to SARS-CoV-2 both concentrates virus on the surface of target cells and allosterically alters the spike protein to promote an "up/open" RBD conformation that enables engagement of the proteinaceous receptor human ACE2 on the cell surface (PMID: 32970989, 35926454, 38055954, 39401361, 40548749). These two events enable plasma membrane fusion (after a cleavage event promoted by plasma membrane TMPSS2) or endocytosis and subsequent pH-dependent fusion (which requires a cathepsin L-mediated cleavage of the spike).

      The authors in this study used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions "downstream" of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. Blocking HS binding with pixantrone, a drug under clinical evaluation for cancer (due to its anti-topoisomerase II activity), inhibited SARS-CoV-2 Omicron JN.1 variant from attaching to and infecting human airway cells. The authors conclude that their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the interesting and technically well-executed study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and the use of primary airway cells. Nonetheless, there are issues that need to be addressed to buttress the proposed model compared to earlier ones. These include: (a) the distinction between macropinocytosis and receptor-mediated endocytosis and what this might mean for productive SARS-CoV-2 infection; (b) the need to account for TMPRSS2 expression and plasma membrane fusion; (c) addition of genetic studies in which hACE2 is expressed in cells lacking HS; (d) an unclear picture of exactly where downstream hACE2 functions; and (e) and a need for comparative/additional study of earlier SARS-CoV-2 variants, which preferentially fuse at the plasma membrane.

      We thank the reviewer for the strong support of this manuscript. We addressed the reviewer’s concerns in the Recommendations to the authors. We did not distinguish whether the endocytic route is macropinocytosis or receptor-mediated endocytosis, because it is a separate study beyond the scope of the present work. We did not examine earlier SARS-CoV-2 variants because we considered it a study beyond the scope of the present work, but a good idea that we may work on in the future. For detail on how we address the remaining concerns, please see our response to the reviewer’s Recommendations for the authors.

      Reviewer #2 (Public review):

      In this manuscript by Han et al, the authors assess the binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that the SARS-CoV-2 spike (in the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface, which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well-described, but the authors attempt to make the claim here that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears to be of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. The manuscript is imprecise and far overstates the actual findings shown by the data. Additional controls would be of great benefit.

      Further, it is this reviewer's opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. While the manuscript provides new insight into the details of this process, the manuscript attempts to oversell this finding by applying new words rather than new molecular details. The authors would be better served by presenting a more balanced and nuanced view of their interesting data. In this reviewer's opinion, the salesmanship significantly detracts from the data and manuscript.

      We thank the reviewer for pointing out that our manuscript is of interest to the field. However, we do not think that we oversell our data. hACE2 has been widely considered the receptor (or the binding partner) that mediates SARS-CoV-2 cell-surface attachment, whereas HS is considered only an attachment factor that facilitates SARS-CoV-2 binding with hACE2 at the cell surface. In the present work, we found that HS, but not hACE2, is the cell-surface attachment receptor (or binding partner), whereas hACE2 is not essential for attachment, but acts downstream of virus endocytosis to facilitate viral genome expression. This finding suggests significant modification of the current model by replacing the attachment receptor (or binding partner) from hACE2 to HS, treating HS as a primary receptor rather than an attachment factor, and relocating the hACE2 action site from the cell surface to the endosome. For these reasons, we do not consider these statements overselling our data. However, as the reviewer suggested in his/her specific comments, we revised the manuscript to ensure that we did not overgeneralize our findings (see our responses to the reviewer’s Recommendations to the authors).

      Major Comments:

      The authors need to rigorously define a "receptor" vs an "attachment factor." They also should avoid ambiguous terms such as "receptor underlying ...attachment" and "attachment receptor" (or at least clearly define them). Much of their argument hinges on the specific definition of these terms. This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm), while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is neither necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this fairly textbook definition. This is proven in Figure 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959, but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. The authors need to discuss this important information and reconcile it with their data and model if they want to claim that HS is broadly important.

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction, which likely won't excite a drug company. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide little discussion of the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2-deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used, as protease expression is a key determinant of the site of fusion.

      The claim that "SARS-CoV2 JN.1 variant binds to heparan sulfate, not hACE2, in primary human airway cells" is extraordinary and thus requires extraordinary evidence.

      First, PIX reduces attachment by 5-fold, which is not the same as "nearly abolished." Also, anti-ACE2 "nearly abolished" entry in 7D, while PIX did not. If the authors want to make these claims, an alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays is TMPRSS2 vs Cathepsin L dependent (using E64d and camostat, for instance) as mentioned above.

      Each figure should clearly state how many independent experiments and replicates per experiment were performed. What does "3 experiments" mean? Are these three independent experiments or three wells on one day?

      In the well-accepted current model, hACE2 is considered the receptor mediating SARS-CoV-2 cell-surface attachment, entry into cells, and infection, whereas HS is an attachment factor that facilitates SARS-CoV-2 binding to hACE2 at the cell surface. The present work revises this view: HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, with ACE2 acting downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined.

      We made this point clearer throughout the newly revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docks at HS clusters.

      The cited CRISPR-screen literature supports context-dependent host-factor usage. However, the absence of HS biosynthesis genes from a given screen does not prove that HS is irrelevant in that cell type; it only indicates that HS biosynthesis was not detected as a genetic dependency under that assay’s conditions. Such negative results can reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency. In the revised manuscript, we included the following in the Discussion:

      “While some studies using genome-wide CRISPR screening to identify genes involved in SARS-CoV-2 reveal genes for HS biosynthesis, others do not (45-50). The negative result, which might reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency, needs to be verified with specific gene knockout.”

      The ~5-fold reduction is likely due to the inhibitor not completely abolishing HS-virus binding. We revised the Discussion to strengthen the suggestion that targeting the virus cell-surface attachment by interfering HS binding is a therapeutic strategy to prevent and treat COVID-19, as in the following:

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding. However, this strategy has not been the focus for developing methods to prevent and treat COVID-19, likely because HS is considered only a regulator that is not essential for SARS-COV-2 entry. Our finding that HS is the attachment receptor re-emphasizes the importance of perturbing virus-HS binding, the first step of the viral entry, to efficiently block SARS-CoV-2 infection. Further supporting this view, inhibition of HS binding with a clinically used HS-binding agent, pixantrone, inhibits authentic SARS-CoV-2 JN.1 subvariant binding with HS on the cell surface and infection in primary human airway cells (Figs. 6, 7). These results suggest a combinatorial anti-SARS-CoV-2 strategy: early HS blockade to prevent attachment combined with ACE2 targeting to inhibit post-attachment steps”

      We include a sentence in the Discussion that our suggestions are limited to the cells we examined as below.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      Three experiments refer to three independent experiments. We added “independent” accordingly throughout the manuscript.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors define a new paradigm for the attachment and endocytosis of SARS-CoV-2 in which cell surface heparan sulfate (HS) is the primary receptor, with ACE2 having a downstream role within endocytic vesicles. This has implications for the importance of targeting virion-HS interactions as a therapeutic strategy.

      Strengths:

      The authors show that viruses are internalized via dynamin-dependent endocytosis and that endocytic internalization is the major pathway for pseudotyped SARS-CoV-2 genome expression. They show that HS-mediated viral attachment is a critical step preceding viral endocytosis and also subsequent genome expression. Further, they show that hACE2 acts downstream of endocytosis to promote viral infection, and may be co-internalised with virions after HS attachment. Pseudotyped virus and authentic SARS-CoV-2 provide similar results. In addition, the authors demonstrate that remarkable clusters of multiple HS chains exist on the cell surface, visualised by a number of elegant microscopy methods, and that these represent the docking sites for virions. These visualisations are an important general contribution in themselves to understanding the nanoscale interactions of HS at the cell surface.

      The use of a complementary range of methods, virus constructs, and cell models is a strength, and the results clearly support the conclusions.

      Overall, the results convincingly demonstrate a different model to the currently accepted mechanism in which the ACE2 protein is regarded as the cell surface receptor for SARS-CoV-2. Here, the authors provide compelling evidence that cell surface clusters of HS are the primary docking site, with ACE2 interactions occurring later, after endocytosis (whilst still being essential for viral genome expression). This is an exciting and important landmark evidence which supports the view that HS-virion interactions should be viewed as a key site for anti-viral drug targeting, likely in strategies that also target the downstream ACE2-based mechanism of viral entry within endosomes.

      We thank the reviewer for the strong support of the present work.

      Weaknesses:

      This reviewer identified only minor points regarding citing and discussing other studies and typos, which can be corrected.

      We have addressed these points in the revised manuscript. For detail, please see our response to the reviewer’s Recommendations to the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Pathway of internalization.

      The authors show clearly that labeled SARS-CoV-2 (pseudovirus or authentic virus) can become internalized in cells lacking hACE2, and this process depends on HS. However, they also show that this pathway is non-productive with regard to infection. Are the entry vesicles mediated by HS alone, HS + hACE2, and hACE2 alone the same? Or does the combination of co-receptor (HS + hACE2) drive SARS-CoV-2 into endocytic vesicles, whereas HS alone promotes macro- or micropinocytosis (lines 361-362). If HS alone directed SARS-CoV-2 into a non-productive entry vesicle, then hACE2 likely would be acting concurrently with HS and not downstream. A more detailed analysis of the different entry vesicles/pathways that occur with HS alone, HS + hACE2, and hACE2 alone is needed.

      During endocytosis, we did not detect a difference in the size distribution of virus-containing vesicles between BHK (HS alone) and BHK<sub>hACE2</sub> cells (HS+hACE2) (Fig. 2D). The similarity in the vesicle size suggests a similar endocytic path with HS alone or with HS + hACE2. In the revised manuscript, we added the following sentence.

      “Third, 3D-STED imaging showed that A490-labeled vesicle’s full-width-at-half-maximum (W<sub>H</sub>) was 363 ± 17 nm (n = 55) in BHK cells, similar to that (333 ± 13 nm, n = 70) in BHK<sub>hACE2</sub> cells (Fig. 2B-D), supporting a similar endocytic path regardless of hACE2 presence or not.”

      (2) TMPRSS2 and plasma membrane fusion.

      Although the authors allude to membrane fusion as an alternate mechanism of entry, their mechanistic experiments do not address the roles of HS and hACE2 in this process, possibly because their BHK and other cells do not co-express significant levels of TMPRSS2. While many Omicron variants preferentially enter cells via endocytosis (relative to antecedent strains in the pandemic) because of spike mutations that reduce cleavage by TMPRSS2 (PMID: 35104837, 36625591, 35145066), plasma membrane fusion can still occur. The authors should add experiments with co-expression of TMPRSS2/hACE2 [with or without HS] and earlier SARS-CoV-2 variants to establish the role of HS in plasma membrane fusion. Also, are there differences in entry pathways if viruses are prepared in cells expressing TMPRSS2?

      We thank the reviewer for this important comment and agree that our mechanistic experiments were not designed to address TMPRSS2-supported plasma membrane fusion. The reviewer’s suggestion for direct testing of HS function in TMPRSS2-supported plasma membrane fusion, including hACE2/TMPRSS2 co-expression and comparison with earlier SARS-CoV-2 variants, will require dedicated experiments and detection of the fusion pathway that we have not yet designed. It is beyond the scope of the present work. In the revised manuscript, we clarify that the observed ACE2-independent uptake and the predominance of endocytic entry refer to the tested cell systems and do not exclude TMPRSS2-dependent plasma membrane fusion in other cell types, as in the following.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Experiments with hACE2 in cells lacking HS.

      Apart from drug treatment (heparinases or pixantrone) studies shown, the current studies do not directly address whether expression of hACE2 on human cells can allow for endocytosis and productive infection in the complete and genetic absence of HS. The only experiments that use genetically deficient cells are the CHO [hamster] cell studies, and these cells lack hACE2 expression. The authors should knock out a key HS biosynthesis gene (e.g., B4GALT7) in more relevant human cells (e.g., A549-hACE2; ideally with sorted subpopulations having different levels of surface hACE2 expression) and assess endocytosis and infection. This is important given studies in the literature by others suggesting that KO of HS expression reduces but does not abrogate SARS-CoV-2 infection.

      We thank the reviewer for these comments. We showed that virus endocytosis is independent of hACE2 (Fig. 1). The reviewer’s question is whether hACE2 alone can allow for endocytosis of viruses. We have shown that in either BHK (without hACE2) or BHK<sub>hACE2</sub> cells (BHK cells expressed with hACE2), heparinase I/II/III mixture (HPRase) nearly abolished cell-surface immunolabelled HS (Fig. 3D), reduced cell-surface virus attachment by ~83-85% (Fig. 3E), reduced viral uptake by ~80% (Fig. 3F, 3G). These results suggest that hACE2 is not essential for viral attachment and endocytosis. We did not test whether hACE2 alone (without HS) plays a minor role for viral attachment and endocytosis, because to our knowledge, HS is present in nearly every cell. Under this physiological condition, it is HS, not hACE2, that plays an essential role in viral cell-surface attachment and endocytosis. In the revised manuscript, we added a sentence admitting that we did not test whether hACE2 alone is sufficient to support viral uptake and productive infection, as in the following.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis.

      (4) hACE2 function in entry.

      In many places, the authors suggest that hACE2-spike functional interaction occurs "downstream" of HS-dependent binding and endocytosis (e.g., lines 25, 33, 48, 210, 309, 312, 318, 333, 346). However, in their model, it is not clear where exactly this interaction occurs. Are the authors suggesting that this spike binds hACE2 on the cell surface, but this has nothing to do with endocytosis, or that the interaction with hACE2 is occurring at a post-entry step? Can they experimentally demonstrate the stage at which hACE2 is functioning? Is it the same or different in cells lacking HS? What about when TMPRSS2 is present?

      We showed that viral attachment and endocytosis are independent of hACE2, whereas entry as determined by viral gene expression, depends on hACE2. We also showed that most virions bind to HS, not hACE on the cell surface. Based on these results, we propose a model that hACE2 functions downstream of virion endocytosis. We cannot rule out the possibility that a small subset of viruses can also bind to hACE2 after their binding with HS at the cell surface.

      We have not been able to design an experiment to visualize hACE2 mediated virion fusion in endosomes, where hACE2 may facilitate virus fusion and delivery of viral genomes to the cytosol. Productive infection requires only a limited number of successful virion–hACE2 engagement events. While many internalized virions can be visualized, the specific virion or vesicle that ultimately gives rise to productive infection cannot be identified from the present imaging data. This makes it difficult to trace the precise stage or compartment in which the functionally relevant spike–hACE2 interaction occurs. In the revised manuscript, we added a paragraph discussing this limitation as below.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis. Although our data suggest that hACE2 functions downstream of endocytosis to facilitate viral fusion at the endosome for genome delivery to the cytosol, we do not know whether hACE2 binding with the virus occurs at the cell surface or endosomes. The binding may occur in both places, but not essential for virus attachment and endocytosis.”

      (5) Other comments.

      (a) Figure 1A. "Antibody" is misspelled.

      Corrected. Thank you.

      (b) The imaging experiments with pseudoviruses and authentic viruses lack any information on the multiplicity of infection or the number of virions added per cell. If this is particularly high and non-physiological (e.g., >100), is it possible that such conditions might enable viruses to enter [dominantly] through secondary [non-infectious] pathways?

      To address the reviewer’s concern, we used flow cytometry to measure cell-associated VSV-S signal as we diluted the virus by ~600-fold. We found that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHK<sub>hACE2</sub> cells over a ~600-fold dilution of the virus (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations. In the revised manuscript, we included the following sentence and Fig. S7 (Supplementary Information).

      “Flow cytometry also showed that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHKhACE2 cells over a ~600-fold dilution of the virus concentration (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations.”

      (c) Figure 1C and elsewhere. Most of the internalization studies rely on various imaging modalities to demonstrate the pseudovirus or virus on or in the cell. The experiments would be strengthened by inclusion of data from orthogonal binding/internalization assays that measuring virion-associated viral RNA on the surface [4oC binding assay] or inside the cell [after a 37oC temperature shift and exogenous proteinase K and RNAse A treatment]) - such assays can be performed at much lower MOI (e.g., <1, addressed comment #2 above) an also allow more objective quantitation and kinetic analyses of virus internalization (e.g., 0, 5, 15, 30 min at 37oC).

      We demonstrate virion attachment and uptake using multiple approaches, including confocal, STED, and EM analysis, showing virions with the expected morphology at the cell surface and in the cytosol. Furthermore, flow cytometric analysis provides population-level quantitation supporting the same overall conclusion. Thus, while we appreciate and agree that an RNA-based binding/internalization assay would provide additional information, we do not consider it essential to the main conclusion of this work.

      (d) Figure 2. (i) Is there any indication of which vesicles the bath dye is in? Is most of this fluid taken up by micropinocytosis? Are these the same vesicles where the virus that is destined for productive infection (HS/hACE2 engaging) transits? (ii) In all panels, can the authors clearly indicate/label which cells are being used (BHK or BHK-hACE2)? (iii) For the studies with dynasore or dominant-negative dynamin-2-K44A, the readout is at 24 h, a late timepoint, which also could affect virus egress and spread. Can the studies be repeated at much earlier time points (e.g., 15 min to 2 h) to demonstrate that viruses are internalized via dynamin-dependent endocytosis in these cells?

      (i) The bath dye A490 was used as a fluid-phase marker for endocytic uptake, rather than as a marker for a specific vesicle class or intracellular compartment. In principle, any vesicle that takes up extracellular fluid could become labelled by this approach. Since nearly all viruses are in the A490-containing vesicles, productive virus infection must come from some of these vesicles.

      (ii) In the revised Fig. 2 legends, we explicitly indicate which cells are used for each panel.

      (iii) To address the reviewer’s concern, we examined earlier time points for dynasore treatment and found that the virus uptake and genome expression were already markedly reduced at 1 h and 8 h after virus incubation. In the revised manuscript, we described these results as below and in Fig. S5.

      “Fourth, dynasore or dominant-negative dynamin 2-K44A overexpression, which inhibits fission of dynamin-dependent endocytosis [28-30], substantially reduced V-A647 internalized 1-24 h after viral incubation (Figs. 2F-G, S5).

      In addition to inhibiting V-A647 endocytosis, dynasore or dynamin 2-K44A inhibited V-EGFP expression 8-24 h after virus incubation by ~66-77% (Figs. 2F-G, S5), suggesting that endocytosis is the main route for viral genome expression.”

      (e) Line 225. "Envelop" should be "envelope".

      Corrected, thank you.

      (f) Line 235. The authors should clarify that they conclude that the "Omicron variant" of SARS-CoV-2 enters "BHK" cells indistinguishably from VSV-S.

      Thank you for pointing this out. We have rephrased the conclusion as “…omicron variant of SARS-CoV-2 enters BHK cells indistinguishably to VSV-S.”

      (g) Line 278. What happens to virus binding if the authors ectopically express hACE2 in CHO-K1 WT and CHO-pgsA-745 cells?

      We did not perform this experiment (see also our response to major comment 3 above).

      (h) Lines 280-281 and elsewhere (line 635). The authors state "pixantrone (PIX), a drug under clinical trial that binds HS to inhibit HS binding with proteins...." The authors should clarify that the drug is under clinical evaluation for cancer treatment because of its DNA intercalating activity (and not its HS binding activity) and cite any relevant ongoing trials. Also, in line 635, is reference #46 correct?

      As suggested, we modified this sentence as “pixantrone (PIX), a drug under clinical trial for cancer treatment due to its DNA intercalating activity, which can bind HS to inhibit HS binding with proteins”

      (i) Line 281-282. The authors should confirm in a Supplementary Figure that the anti-hACE2 antibody used blocks SARS-CoV-2-JN.1 binding to ACE2.

      In Figure 7D, we showed that PIX and anti-hACE2 antibody block SARS-CoV-2-JN.1 infection, suggesting that anti-hACE2 blocks SARS-CoV-2-JN.1 binding with hACE2.

      (j) Figure 7B. Can hACE2 co-localization be added to this panel?

      We did not perform this experiment. We addressed the role of ACE2 in these airway cells in subsequent panels of Fig. 7.

      (k) Figure 7C. The quantitative data show a 50% reduction in binding signal with pixantrone, whereas the microscopy images appear to show a much greater effect. Can more representative images be shown so that the data better corresponds?

      A ~50% effect is not as visually obvious as the current Fig. 7C. Therefore, we chose not to change the images. However, the statistics in Fig. 7C (right) clearly indicate an average effect of about 50%, as the reviewer pointed out.

      (l) In the Discussion, it is not necessary to use Figure callouts (as done in the Results). Please remove, with the exception of reference to the model.

      We prefer to call out Figures in the Discussion so that we can remind the readers where to find the data. The readers may choose to neglect these callouts. But some readers may read most the discussion part without going through the results carefully. In this case, the figure callouts may help these readers.

      (m) Please delete all references to "new" or "novel" models. It is unnecessary.

      As the reviewer suggested, we deleted “new” and “novel” throughout the revised manuscript.

      (n) Figure legends. Please make sure each panel indicates the # of independent experiments performed. This is included for some but not all panels. Also, a few panels use an unpaired t-test where an ANOVA with multiple comparisons is required (e.g., Figure 1G and S1).

      We agree that, for the three-group sub-comparisons shown within Fig. 1G and Fig. S1, the relevant analyses should account for multiple comparisons. In the revised manuscript, we therefore analyzed these predefined three-group subsets using ordinary one-way ANOVA followed by Dunnett’s multiple-comparisons test, with BHK or Vero used as the reference group as appropriate. The two-group comparisons were analyzed using unpaired two-tailed t-tests.

      Reviewer #2 (Recommendations for the authors):

      (1) It is well established that ACE2 is the receptor for SARS-CoV-2. The authors should not downplay this by saying it is "widely assumed", "typically thought", etc. The specific molecular details at various stages of entry (i.e, the role of HS) remain a bit unclear, but it is disingenuous to imply ACE2 is not the bona fide receptor by any conventional definition.

      The present work does not challenge the well-established view that ACE2 is the receptor for SARS-CoV-2 entry/infection, but suggests that HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, whereas ACE2 acts downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined. We made this point clearer throughout the revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docked at the HS clusters.

      As the reviewer suggested, we removed “assumed” and “typical” and clarify that our findings do not challenge this concept. For example, we modified the abstract

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely assumed as the receptor.”

      as

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely considered the receptor for cell-surface attachment and subsequent cell entry.”

      (2) When the authors state pseudovirus internalization is independent of ACE2, they should clarify that this is the case in cells not expressing TMPRSS2. Most physiologically relevant cell types express TMPRSS2, which will facilitate entry at the plasma membrane.

      As the reviewer suggested, we included the following sentence in the Discussion section: “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Line 130: "Endocytic internalization is the main viral infection pathway" and Line 180-181 is not precise and should be rephrased to include the cell types described in the figure. This may be true in BHK-ACE2 cells, but the evidence in this section does not show that this is universally or broadly true.

      We agree and have revised these sentences to limit the conclusions to the experimental context directly supported by our data. Specifically, our results support endocytic uptake as the major route leading to pseudovirus genome expression in the pseudovirus assays and cell types examined here, rather than as a universal entry mechanism for SARS-CoV-2 across cell types. We have therefore modified the subsection title and the relevant sentence in the Results to explicitly refer to the tested cells/assays.

      Across the revised manuscript, we have accordingly revised the text to distinguish initial virion docking/attachment from productive entry, to acknowledge ACE2 as the established receptor for productive infection, and to limit our mechanistic conclusions to the cellular systems directly tested here.

      (4) All bar plots should show individual dots (i.e., Figure 1G) to better reveal the variance of each dataset.

      While we respect the reviewer’s suggestion, this is not required in the journal style. We prefer plotting bar graphs without individual data points, which often makes it difficult to see the mean values.

      (5) Line 57: This is not accurate. HIV uses a receptor and a co-receptor, for instance.

      We thank the reviewer for noting this inaccuracy. We agree that viral entry frequently involves coordinated engagement of multiple host factors rather than a single receptor, for example, HIV requires both a primary receptor and a co-receptor. We have revised the statement in the Introduction (Line 57–58) to reflect that entry can involve receptors together with co-receptors and/or attachment factors, which collectively facilitate membrane fusion or endocytic uptake.

      In the Introduction (Line 57), we replaced the sentence with “Viral entry is often initiated by engagement of host receptors and associated co-factors that together facilitate subsequent viral membrane penetration.”

      (6) Line 60: "most" --> "many"

      As suggested, we have changed “most” to “many”.

      (7) Remove "clinically relevant" in reference JN.1, as JN.1 is not circulating currently. A more appropriate term could be "full-length" or "authentic", or "wild-type".

      As suggested, we changed it to “authentic”.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors omit to mention the work of Zhang et al, 2023 Nature Comms. "Host heparan sulfate promotes ACE2 super-cluster assembly and enhances SARS-CoV-2-associated syncytium formation". These authors also use PIXN and MTN compounds and define different mechanisms based on ACE2 clustering for virus entry. The authors should mention this work in the Discussion and try to reconcile the different findings.

      As suggested, we include the following discussion in the revised manuscript.

      “Consistent with this possibility, HS may promote spike-dependent ACE2 super-cluster assembly at the cell surface and enhance SARS-CoV-2–associated syncytium formation, suggesting that HS may organize ACE2 nanoscale architecture in a cell–cell fusion context [43].”

      (2) The authors should strengthen their case for the validity of HS-virion interactions as a therapeutic target by mentioning studies showing effectiveness of interference with HS-Covid interactions by heparin and other investigational drugs eg. first study to demonstrate heparin inhibition of SARS CoV2 attachment, Mycroft-West et al, Thromb Haemostatis, 2020; recent report of successful clinical trials of nebulized heparin, The Lancet, Sept 2025; and the superior efficacy of HS mimetic Pixatimod/PG545 compared to heparin (Guimond et al 2022 ACS Chemical Sciences).

      We thank the reviewer for this suggestion and add the following paragraph with citations the reviewer mentioned in the Discussion section.

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding.”

      (3) Figure 1a: incorrect label for antibody.

      Corrected, thank you.

      (4) Some misspellings noted in the manuscript, e.g., MINFLLUX, so please recheck the manuscript for typos.

      We have rechecked the manuscript and corrected the typos.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      There are two main criticisms:

      (1) It is not clear how much the factors uncovered here are true beyond B6 mice. B6 mice, compared to humans, are known to be very Th1-skewed, and Tbet is a strong inhibitor of Th17-specific T cells. Many people make IL-17-producing T cells in response to Mtb infection.

      We appreciate the point that not all findings in mice are directly translatable to humans. The B6 mouse is widely used as a model organism for tuberculosis due to its tractability and the wealth of genetic tools available for this strain. While it is true that many individuals do produce Th17 cells after infection with Mtb, humans are still very Th1-dominant, and not all infected individuals produce Th17 cells. We can speculate that the mechanisms outlined in this paper may contribute to the reasons that Th17 responses are not more robust in humans, a finding that may be useful in guiding vaccine design in the future.

      (2) Very few novel insights are mechanistically revealed about how Th17 induction is restricted by Mtb. Tbet induction is known to restrict Th17 development, and this is a T-cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. How, if at all, are these factors related to each other in restricting Th17 induction? Also, the conclusion that it is not a result of attenuation is not completely convincing.

      While it is established that Th1 differentiation can inhibit Th17 differentiation, we believe that rigorously demonstrating this genetically in the context of Mtb infection remains important. Moreover, it cannot be assumed that IL-17 elicited by dampening the Th1 response can lead to enhanced control of infection. We view addressing this as a significant contribution. Furthermore, we also show that the ESX-1 and PDIM virulence factors are functionally linked by suppression of IL-17 responses. The effect is unlikely to be simply due to attenuation of the strains as an equally attenuated control strain does not elicit Th17 cells. We believe that these insights are both novel and important for understanding immune responses to Mtb.

      Other points:

      (1) The authors show that mice infected with a deficiency in ESX-1 have more IL-17-producing CD4 T cells in response to stimulation with an ESAT-6 peptide pool (Figure 3B). Because ESAT-6 is encoded by ESX-1, why do mice infected with this Mtb mutant have any ESAT-6-specific T cells? Is it an incomplete knockdown?

      The ESX-1 knock-out M. tuberculosis Erdman strain is a ΔEccC1 mutant. This strain can produce Esat-6 but cannot secrete Esat-6 out of the bacterial cell. Thus Esat-6 protein is present and able to be processed for MHC-II presentation. We also use Ag85b peptide pool stimulation and report similar effects as Esat-6 peptide pool stimulation.

      (2) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." I don't think the last sentence is necessarily true. I can imagine a scenario in which the induction of the Th17s is, in fact, due to the attenuation, and the Th17 induction still contributes to protection.

      We tested another attenuated M. tuberculosis strain with no known relationship with ESX-1 or PDIM, ΔMmpL4. This attenuated mutant fails to induce IL-17A–producing CD4 T cells to the same extent as observed in mice infected with ESX-1-deficient or PDIM-deficient strains, which is strong evidence that simple attenuation of virulence does not result in higher numbers of Th17 cells being elicited.

      (3) ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but what about the LN? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction.

      We acknowledge that this is a formal possibility, however we maintain that the phenotype is specific to ESX and PDIM mutants, rather than MmpL4 mutants. Even if this phenotype arises from a tissue-specific attenuation of ESX/PDIM mutants, it remains a specific phenotype of these mutants, and not all attenuated mutants, albeit less directly. More importantly, the observation that these mutants induce higher levels of the Th17-polarizing cytokine IL-23 from infected cells ex vivo suggests that this is not an indirect phenomenon.

      (4) Do LN cDC1 and high levels of IL-12 p35 in mice infected with the mmpl4 mutant? Likewise, LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb)? If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than those infected with ESX-1 and PDIM mutants.

      Because ΔMmpL4 and complemented strains resulted in T cell profiles that were not different from the wild-type, we did not measure mediastinal lymph node dendritic cell expression of IL-12 p35 and IL-23 p19 in infections with these mutants.

      (5) ESX-1 and PDIM are very different virulence factors - a protein secretory pathway and cell wall lipid, respectively? Mechanistically, how would mutants in these pathways give very similar outcomes regarding Th17 cells unless it was simply as an aspect of their attenuation? Perhaps, mmpl4 mutants simply differ in some aspects of their attenuation, such as bacterial burdens in LNs, or their interaction with cDCs?

      We are not the first to link phenotypes of ESX-1 and PDIM. Both systems have both been shown to be important for M. tuberculosis permeabilization of the host cell phagosome after phagocytosis, and for suppression of type I IFN responses, among other responses. Thus, these seemingly different virulence factors clearly work together to support specific virulence traits during infection. The exact mechanism of how ESX-1 and PDIM interact is not completely understood and is an area for future investigation.

      Reviewer #2 (Public review):

      The following conclusions and interpretations should be revisited, rephrased, and re-evaluated:

      (1) The manuscript neglects to analyze T cell responses in the dLN, which is the critical site where these responses are initiated (only DC cytokine production is measured in the dLN). The differences in the lungs could reflect trafficking of T cells to the lungs, local lung T cell responses, or durability of the T cell responses in the lungs. The authors state in the last results section that "These results indicate that the ESX-1 and PDIM virulence factors impact naïve T cell differentiation at the draining mediastinal lymph node..." but T cell responses are never measured in the dLN.

      Due to the limited size of the mediastinal lymph node at 3 weeks post infection, we were unable to obtain enough cells for both myeloid cell analysis and T cell analysis, as we perform staining for these panels separately due to the decrease in viability of myeloid cells observed during T cell restimulation. In addition, because T cells in the lung are the population of cells most critical for mediating the outcome of infection, we believe analyzing the T cell response in the lymph nodes though interesting, is not crucial for this study. We have edited the manuscript to be clearer, as suggested by the reviewer.

      (2) Figure 2: The authors state that "Importantly, IFN-γ deficient mice did not exhibit elevated levels of IL-17A producing CD4 T cells demonstrating that IFN-γ production is not the mechanism by which Th1 T cells limit a Th17 response during Mtb infection", but the difference is significantly different and even more obvious in Panel B. In fact, if the Panel D y-axis was on a log scale, the Ifng-/- would likely look more like Tbet-/- than WT. Based on this data, it seems like IFNg is having an effect and should not be completely discounted. Does the deletion of Ifng affect the number of Tbet+ T cells?

      We agree that the IFN-γ<sup>-/-</sup> have only 5x more IL-17 producing CD4 T cells than WT mice while Tbet<sup>-/-</sup>mice exhibit a 25-fold increase compared to WT. We have added this information to the text, and now point out that IFN-γ production is not the sole mechanism by which Th1 T cells limit a Th17 response during Mtb infection.

      In addition, the deletion of Tbet results in an increased number of IFNg+IL-17+ double positive T cells (Figure 2B), in addition to a sizable IFNg single positive T cell population maintained in the Tbet-/- mice (10x the negative control of Ifng-/-). Is this why Tbet deletion is not as severe as Ifng deletion, because T cells are still making IFNg?

      It is possible that the residual IFN-γ produced by T-bet-deficient animals contributes to their relatively modest susceptibility to infection. However, our data show that deletion of IL-17 in this background renders T-bet–deficient mice nearly as susceptible as IFN-γ deficient mice, arguing that the remaining IFN-γ is not a major protective factor.

      Along these lines, the statement in the text that, "Tbet-/-Il17a-/- mice completely lacked both IFN-γ producing...." T cells is not supported by the data in Figure 2C. Tbet-/-Il17a-/- mice look to have more gamma-producing T cells than Tbet-/- mice (which is already 10x the negative control of Ifng-/- in panel 2B if one includes the gamma single positive and IFNg/IL-17 double positive).

      We have amended the language in the text to be more consistent with the data.

      (3) In the Results sections describing Figures 3, 4, and 5, the authors equate IL-17 production by T cells with TH17 responses and IFNg expression with TH1, but Tbet and RORgt expression in the T cells should be measured to make conclusions about TH1 and TH17. Or the authors can rephrase their findings to specifically state the observations as IFNg or IL-17 expressing CD4+ T cells.

      We believe that calling a CD4 T cell in the lung that is producing IFN-γ (and not IL-17) a Th1 cell is appropriate. Potentially confounding cells include those which also produce IL17, which we have ruled out, or T<sub>FH</sub> cells that may be common in lymph nodes but are not common in lungs at this time point and under these conditions.

      (4) Conceptually, do the authors think that ESX1/PDIM promotes TH1 responses and this blocks TH17 or are ESX1/PDIM blocking TH17 responses directly, allowing for increased TH1 responses? It would be helpful to clarify the model in this regard, describe how the data supports one model or the other, and then make sure the language is consistent throughout. Can these effects on T cell responses be tested and recapitulated in vitro using infected APC and T cell co-cultures?

      While it is possible that PDIM and ESAT-6 suppress Th17 through promotion of Th1 differentiation, we do not have data to support this model currently. However, we have added a comment making this point to the discussion.

      Reviewer #3 (Public review):

      Weaknesses:

      (1) The authors should acknowledge and reference key findings from the literature that have identified suppression of Th17 differentiation as an Mtb virulence mechanism, e.g., the role of the Hip1 protease and CD40 signaling (Madan-Lala JI 2014, Sia Plos Path 2017, Enriquez iScience 2022) and Khader JI 2005, showing the requirement of IL-23 for Th17 responses in vivo in a TB mouse model.

      We thank the reviewer for pointing these references out and have added them to the discussion section of the manuscript.

      (2) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may be due to a difference in timing, since only 3-week data are shown; earlier and later time points would provide better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, data on neutrophils are important for clarifying the mechanism for the CFU outcomes.

      We agree that it is surprising that, in the context of TB, Th17 responses are protective whereas excessive neutrophil recruitment is detrimental to the host. In IFN-γ–deficient mice, neutrophils are recruited and contribute to the increased susceptibility of this strain (PMID: 21967766). In separate work from our lab, we have shown that the phenotype of neutrophils recruited to the lungs during Mtb infection influences disease outcome (PMID: 40937719). It is possible that differences in the host environment and the timing of the response shape the effects of neutrophils on the host; these and the other questions raised by the reviewer will be the subject of future studies.

      (3) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Figure 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6, etc.) in the Tbet KO experiments.

      While these are interesting points, investigating mechanisms of Tbet-dependent suppression of IL-17 is beyond the scope of this study.

      (4) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Figure 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants. Additionally, what was the level of Type I IFNs elicited by these mutants?

      We included the MmpL4 knockout Mtb Erdman strain as a control to ensure that attenuation of mutants is not the cause of the increase in IL-17. We also showed that eliminating type I IFN signaling by deleting its receptor has minimal impact on Th17 differentiation, even in the context of a host that produces excess type I IFN. Therefore we do not believe that type I IFN elicited by these mutants is explanatory for the phenotype.

      (5) Since macrophages have been implicated in the reduced cytokines seen in the ESX-1 mutant, IL-23 and other cytokine data on lung macrophages would complement the DC data.

      Because dendritic cells are primarily responsible for priming CD4 T cell responses, we believe that this result in macrophages would not substantially alter our conclusions. That said, it was demonstrated previously that macrophages infected with ESX-1 mutants produce less IL-12p40, a subunit of IL-23.

      (6) Figure 5. There are many fewer DCs overall in the eccC1 and fadD28 mutant groups, which could account for the increased % IL-23p19 in DCs (5D). What were the levels of IL-23 in DC1s?

      The amount of IL-23 p19+ in type I conventional dendritic cells (cDC1s) was near zero as shown in supplementary figure 6A. cDC1s are known to not express IL-23 p19 in mice.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What do the authors mean by the alternative secretion part of "ESX1 Type VII alternative secretion system" that they refer to?

      Bacterial alternative secretion systems facilitate the export of proteins from the bacterial cell independent of the canonical Sec-dependent secretion system required for export of most secreted bacterial proteins across the inner membrane. However to avoid confusion, we have removed the word alternative.

      (2) Not sure naïve fits in this sentence at the end of the introduction: "Furthermore, we observe a strong Th17 response during infection with ΔESX-1 or PDIM lacking Mtb in naïve mice....".

      We have removed the word Naïve.

      (3) Figure legend for 1A-C says analysis performed at 21 dpi, but the figure shows the time course.

      We have corrected this error.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1 should show the non-stimulated flow plot.

      We have added the unstimulated samples.

      (2) The % IL-17 in the flow plots is not consistent across Figures 1, 2, and 3. Not sure why the scales for the Y-axis for IL-17 differ so much between Figures 1 and 2/3. IS there a technical issue with compensation?

      We did not experience any difficulties with compensations. These experiments were done over several years of work. For every experiment, new single-color controls were used and gating was done with the FMO gating strategy. Minor variation such as we see here is not surprising.

      (3) Discuss Yeh et al J Neuroimmunol 2014- show that IFNγ inhibits Th17 differentiation and function via Tbet-dependent and Tbet-independent mechanisms.

      We have added this reference to the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      We thank the reviewer for the kind comments.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      We have added more details on the annotation process and have included the list of instructions sent to the annotators. We have also included mmpose configs with the code provided. The following files include the relevant details:

      File including the list of instructions sent to the annotators:

      OpenMonkeyWild Photograph Rubric.pdf

      Mmpose configs:

      i) TopDownOAPDataset.py

      ii) animal_oap_dataset.py

      iii) init.py

      iv) hrnet_w48_oap_256x192_full.py

      Anaconda environment files:

      i) OpenApePose.yml

      ii) requirements.txt

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths:

      Open source dataset with excellent annotations on the format, as well as example code provided for working with it.

      Properties of the dataset are mostly well described.

      Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones?

      Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out).

      The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      We thank the reviewer for the kind comments.

      Weaknesses:

      More details on training hyperparameters used (preferably full config if trained via mmpose).

      We have now included mmpose configs and anaconda environment files that allow researchers to use the dataset with specific versions of mmpose and other packages we trained our models with. The list of files is provided above.

      Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010).

      We have included a datasheet for our dataset in the appendix lines 621-764.

      Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here.

      We have included the list of instructions sent to the Hive annotators in the supplementary materials. File: OpenMonkeyWild Photograph Rubric.pdf

      Should include model cards, as described in Mitchell et al (arXiv:1810.03993).

      We have included a model card for the included model in the results section line 359. See Author response image 1:

      Author response image 1.

      It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year.

      We agree that the source could introduce structural biases. This is why we included images from so many different sources and captured images at different times from the same source—in hopes that a large variety of background and lighting conditions are represented. However, doing so limits our ability to document each source background and lighting condition separately.

      Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically.

      The focus of this work is on overall keypoint localization accuracy and hence we wanted a metric that is easy to interpret and implement, in this case we made use of PCK (Percentage of Correct Keypoints). PCK is a simple and widely used metric that measures the percentage of correctly localized keypoints within a certain distance threshold from their corresponding groundtruth keypoints.

      A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

      We have now included a histogram of unnormalized bounding boxes in the manuscript, see Author response image 2:

      Author response image 2.

      Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential.

      We thank the reviewer for the kind comments.

      However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      We thank the reviewer for pointing this out. Inter-observer reliability is important for ensuring the quality of the annotations. We first used Amazon MTurk to crowd source annotations and found that the inter-observer reliability and the annotation quality was poor. This was the reason for choosing a commercial service such as Hive AI. As the crowd sourcing and quality control are managed by Hive through their internal procedures, we do not have access to data that can allow us to assess inter-observer reliability. However, the annotation quality was assessed by first author ND through manual inspections of the annotations visualized on all of the images the database. Additionally, our ablation experiments with high out of sample performances further vaildate the quality of the annotations.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      Our goal was to obtain as many images as possible from the most commonly studied ape species. In order to ensure a large enough database, we focused only on the species and combined images from as many sources as possible to reach our goal of ~10,000 images per species. With the wide range of people involved in obtaining the images, we could not ensure that all the photographers had the necessary expertise to differentiate individuals and subspecies of the subjects they were photographing. We could only ensure that the right species was being photographed. Hence, we cannot include more detailed information.

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      We created the training set to reflect the species composition of the whole dataset, but used test sets balanced by species. This was done to give a sense of the performance of a model that could be trained with the entire dataset, that does not have the species fully balanced. We believe that researchers interested in training models using this dataset for behavior tracking applications would use the entire dataset to fully leverage the variation in the dataset. However, for those interested in training models with balanced species, we provide an annotation file with all the images included, which would allow researchers to create their own training and test sets that meet their specific needs. We have added this justification in the manuscript to guide the other users with different needs. Lines 530-534: “We did not balance our training set for the species as we wanted to utilize the full variation in the dataset and assess models trained with the proportion of species as reflected in the dataset. We provide annotations including the entire dataset to allow others to make create their own training/validation/test sets that suit their needs.”

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

      We thank the reviewer for this important suggestion. Since we can separate a video into its constituent frames, one can indeed use the provided model or other models trained using this dataset for inference on the frames, thus allowing video tracking applications. We now include a short video clip of a chimpanzee with inferences from the provided model visualized in the supplementary materials.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Please provide a more thorough description of the annotation procedure (i.e., the instructions given to crowd workers)! See public review for reference on dataset annotation reporting cards.

      We have included the list of instructions for Hive annotators in the supplementary materials.

      An estimate of the crowd worker accuracy and variability would be super valuable!

      While we agree that this is useful, we do not have access to Hive internal data on crowd worker IDs that could allow us to estimate these metrics. Furthermore, we assessed each image manually to ensure good annotation quality.

      In the methods section it is reported that images were discarded because they were either too blurry, small, or highly occluded. Further quantification could be provided. How many images were discarded per species?

      It’s not really clear to us why this is interesting or important. We used a large number of photographers and annotators, some of whom gave a high ratio of great images; some of whom gave a poor ratio. But it’s not clear what those ratios tell us.

      Placing the numerical values at the end of the bars would make the graphs more readable in Figures 4 and 5.

      We thank the reviewer for this suggestion. While we agree that this can help, we do not have space to include the number in a font size that would be readable. Smaller font sizes that are likely to fit may not be readable for all readers. We have included the numerical values in the main text in the results section for those interested and hope that the figures provide a qualitative sense of the results to the readers.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fueled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. This work will be of interest to scientists in the field of aging and metabolism. Some clarifications in the text would address the following concerns, which would increase the impact of the study:

      (1) What does activation of AMPK (via PGDP-Sak1 expression) do to the replicative lifespan? How many bud scars, in general, do the subpopulations that are older - yet have less Tom70 (increased mitochondrial fitness) - have, after the 48 hrs timepoint that they are examining? How many divisions occurred in this 48hr time period - i.e. is it long enough to have all cells reach the end of their replicative lifespan? This information is important to rule out that a subset of the mutant cells just divided faster and hence had more divisions within 48 hrs (growing faster and living longer are different things). Having identical growth curves doesn't indicate per se that they all divide at the same rate, as there may be a subpopulation that divides faster and a subpopulation that doesn't grow so well.

      Increasing AMPK activity increases replicative lifespan [PMID: 25869125], but given our finding that AMPK activation splits the population, such replicative lifespan assays are hard to interpret. Bud scar counts have a similar issue. Hence we restricted the lifespan and bud scar analyses to wt and A2A which are more homogenous (Figures S2 B and E). A2A cells at 48 h have ~25% more bud scars than wt cells. Yes, by 48 h most of the cells have lost viability (Figure 2E). The reviewer is correct that you can't properly compare the lifespan curves if the cells divide at different rates, hence our follow-up test of wt at 48 h vs A2A at 40 h viability after we had confirmed that these time points captured cells at equivalent replicative ages (Figure 2D, E). This shows that viability of A2A is slightly lower than wt at matched age, indicating a slightly shorter lifespan. 

      (2) A2A cells do not have an extended replicative lifespan (RLS) but show an increase in the "low senescence" population (Figure 2). If the cells are not becoming senescent, why don't they have longer RLS? Not having a longer lifespan seems inconsistent with the statement that "bud scar counting confirmed that A2A cells reach a higher age than wild type", which comes back to how many times the cells can divide in the 48hr timepoint studied and their rate of cell division? Also, the lifespan curve shown is plotted against time, not cell division number, which does not take into account different division times of cells within the population (described above). It would be much more useful to show standard lifespan curves showing cell division numbers per lifespan per cell.

      Our observation that cells can reach the end of life without senescing is consistent with other studies that have studied the life course of individual cells by microscopy [PMID: 31291577, 32675375]. These studies always highlight some proportion of the cells that reach the end of life with no or minimal senescence, though this fraction varies with the experimental system. The question of why cells lose viability without senescing is a complete unknown in the field, and reflects a wider lack of consensus as to why yeast lose viability with replicative age.

      In liquid culture we can only assess viability over time, not cell division number, which we agree is not optimal and we are wary about making strong statements on lifespan for exactly the reasons the reviewer notes. Unfortunately, it is clear from the comparison of liquid and solid media lifespans performed by the Gottschling lab [PMID: 19652178] that culture system has a huge effect on lifespan, with cells in classical plate-based microdissection assays living far longer than the same strains do in liquid. This means that lifespans determined by microdissection-based assays are of questionable relevance to ageing studies performed in liquid culture. Senescence cannot be assayed on plates, while microfluidic systems lack the throughput necessary and preclude key techniques like RNA-seq, so liquid culture assays were the only option for this work. We agree that this leaves an unsatisfactory approximation for lifespan measurements, but we consider it critical that everything is measured in the same system. We therefore restricted our conclusion on lifespan to simply say that lifespan of A2A cells is not extended which our data in Figures 2D, E, S2B does support (see also answer to Q1), and therefore with the majority of A2A cells showing low senescence marks and high fitness at 48 h we can conclude that lifespan and fitness loss must be separable.

      We have added a note of these limitations of lifespan measurements in the materials and methods section of the manuscript.

      (3) Increased "fitness" of the old cells is implied from the increased size of the colonies that the old cells can make. However, this is a measure of the fitness of the daughters per se, not the old mother cells. Are the old mothers just passing on healthier mitochondria and more lipids to the daughters, such that they can divide more times? If the aged cells have an "increased fitness", why don't they divide more times themselves (i.e. live longer?).

      Yes, colony growth speed is defined by daughter cell replication, but as long as the daughters and subsequent generations divide at the same rate irrespective of whether they come from a young or old mothers then the size of the colony after 24 hours varies based on the time it took the initial mother to produce a daughter. This is what the assay really measures. We note that aged wildtype mothers often do not divide at all in the first 24 hours after being put on an agar plate (hence the tiny reported colony size), even though they do eventually produce a daughter which then forms a colony, whereas A2A cells tend to produce the first daughter rapidly whether young or old. It is known that daughters of aged wildtype mothers also divide slower, as to some extent do grand-daughters (PMID: 2644196), which will also contribute to differences in colony size, and this may well result from a lipid and/or mitochondrial contribution, but the primary driver of colony size in 24 hours is the time the mother took to initially divide. We have added this detail to the materials and methods section of the manuscript.

      As noted above, the mechanistic basis of lifespan is unknown, but although senescence can shorten lifespan, our work and that of others shows that lifespan is still limited in the absence of senescence.

      (4) The statement is made that "these experiments define two classes of aging cells with distinct metabolic needs, coherent with the model of two aging trajectories previously proposed (referencing Nan Hao's work)". However, the big difference here is that in Nan Hao's work, their two aging trajectories influenced the length of lifespan, but that does not appear to be the case here. That distinction should be made clear. Perhaps the authors could also speculate as to why the A2A yeast stops dividing after presumably the same number of cell divisions, even though they have an activated AMPK and activated fatty acid synthesis pathway.

      Yes, this is a good point and we have added this distinction to the Discussion:

      “Here we have characterised two classes of ageing cells seemingly differentiated by high and low availability of cytosolic Acetyl-CoA, consistent with a previous demonstration that ageing follows two trajectories in yeast though it should be noted that in this previous report, the two trajectories also differed in replicative lifespan (6).”

      We would love to speculate on why the A2A cells don't have an extended lifespan, but at this point we don't have a strong hypothesis. We have come up with many theories for this, but none that we haven’t managed to disprove experimentally. One thing worth considering is that many cells which lose replicative viability in liquid culture and probably in plate assays remain intact – for example, DNA and RNA integrity is not compromised over 24- 48 h – so those cells are probably not dead per se. But we also detect apoptosis-sized DNA fragments, which must come from dead cells, so there is clearly not a single mechanism defining the end of replicative lifespan.

      (5) I am a bit confused by the use of the word "senescence" by this lab here and in their previous growth on galactose studies. If yeast don't senesce, which is usually defined as an irreversible arrest of the cell cycle where cells stop dividing, shouldn't the yeast that do not senesce still be dividing and hence have a longer lifespan? Should a different term be used rather than senescence? Such as "fitness late in life". The authors giving their definition of senescence may help reduce this apparent contradiction.

      We completely agree, this is confusing and noted this distinction in the Introduction. Use of the term senescence to mean a loss of fitness late in life in yeast stems from the classical definition of senescence as applied to whole organisms. However, the term senescence as applied to cells has a more specific meaning in terms of the cell cycle as the reviewer notes. As an individual S. cerevisiae is both a cell and an organism, the terminology clashes. However, the marker we largely employ (Tom70-GFP) which in our hands is a very good proxy for fitness was originally defined as marking the senescence entry point (SEP), so overall we feel we can't avoid the term.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells.

      Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metab to olic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Weaknesses:

      (1) While the manuscript focuses on mitochondrial transport and fatty acid synthesis, cytosolic acetyl-CoA is also a key regulator of histone acetylation and chromatin silencing. It would strengthen the study to consider whether acetyl-CoA depletion contributes to improved fitness through enhanced rDNA silencing. Given the well-established role of rDNA instability in yeast aging, additional experiments examining rDNA silencing and stability would be valuable. For example, monitoring rDNA copy number changes (not necessarily ERCs) under AMPK activation, oleic acid supplementation, and in the A2A strain, similar to approaches used in the authors' prior work, would help clarify whether chromatin regulation contributes to the observed phenotypes.

      We have added data addressing these points to the manuscript and Supplemental Figures 2, 3 and 4, though the outcomes are complex. Histone acetylation changes chromatin accessibility and could therefore alter global gene expression; in accord with this, RNA-seq shows that P<sub>GPD</sub>-SAK1 reduces known age-linked gene expression dysregulation. However, A2A does not further reduce the effect, meaning either that another driver exists in addition to cytosolic acetyl-CoA, or that age-linked gene expression dysregulation is unrelated to cytosolic acetyl-CoA. Oleic acid has little effect on age-linked gene expression dysregulation despite rescuing fitness. With regard to rDNA silencing, transcription of the rDNA intergenic spacer non-coding RNAs promotes ERC formation; we have added data showing that ERC accumulation is not reduced in A2A but slightly higher coherent with the higher replicative age of A2A at 48 h, which suggests silencing is not better in A2A. By RNA-seq, these intergenic spacer transcripts are massively upregulated with age, but this will be a consequence of the increased genomic copy number on ERCs; the upregulation is less in A2A than other conditions, but this arises because the log phase spacer transcript levels are higher and so does not reflect better rDNA silencing. We have previously assayed for heritable changes in rDNA copy number arising during ageing and found (to our surprise) absolutely nothing, so we don't expect any changes under these conditions. The upregulation of transcripts from Sir2-repressed telomeric and MAT loci with age is decreased in P<sub>GPD</sub>-SAK1 and A2A, but the effect size is not different from any other low-expressed genes so we do not think there is a particular effect at loci subject to chromatin silencing (see our previous study Zylstra et al PMID 37643194 for evidence that Sir2-mediated gene silencing is not affected by age). We have added our conclusions from these experiments to the Discussion.

      (2) The current data do not fully distinguish whether AMPK activation and oleic acid supplementation act on distinct subpopulations of aging cells. An alternative explanation is that oleic acid supplementation enhances mitochondrial function and acts additively with AMPK activation, thereby increasing the fraction of cells in the "low senescence" state. Since this distinction is not central to the main conclusions, I suggest softening the language around subpopulation specificity. Emphasizing instead that the A2A strain coordinately modulates multiple branches of acetyl-CoA metabolism to improve late-life fitness would maintain the strength of the central message without over interpretation.

      We respectfully disagree with the reviewer on this point. We show that P<sub>GPD</sub>-SAK1 rescues senescence in ~half the population by a Cat2/Mls1 dependent mechanism (Figure 1F). We then show that in A2A, which rescues most cells, deletion of CAT2/MLS1 restores senescence in ~half the cells (Figure 3F/G). This cannot be explained by an additive mechanism as this would either result in all cells being partially rescued in the P<sub>GPD</sub>-SAK1 and in the A2A cat2Δ mls1Δ mutants, which is definitely not the case either by Tom70-GFP or fitness. Instead the population splits into high/low senescence and fit/unfit cells in the different assays.

      On the specific point of whether lipid synthesis additively increases mitochondrial function, we have added oxygen consumption rate data showing that A2A cells respire more than P<sub>GPD</sub>-SAK1 at 48h but only by a relatively small amount (Figure S3D), so there is indeed an additive improvement in mitochondrial function, but too little to explain the difference in population fitness in our opinion.

      We realise that the reviewer is asking more specifically about oleic acid, but again in the flow data, Figure 4C, what changes with oleic acid or P<sub>GPD</sub>-SAK1 is the proportion of cells in the low Tom70 / high WGA sector. Under an additive effect model, oleic acid or P<sub>GPD</sub>-SAK1 individually would partially reduce Tom70 and partially increase WGA, but the population in the low Tom70 / high WGA sector has the same average Tom70/WGA values in oleic acid, P<sub>GPD</sub>-SAK1 or P<sub>GPD</sub>-SAK1+oleic acid. It is the proportion of cells in this population that changes. Furthermore, under an additive model, wildtype cells aged with oleic acid would not have highest fitness than P<sub>GPD</sub>-SAK1 or A2A (Figure 4D) as these individual cells would lack the mitochondrial upregulation from P<sub>GPD</sub>-SAK1.

      (3) The manuscript proposes that lipid starvation and excess acetyl-CoA are major drivers of senescence in distinct subpopulations of wild-type aging cells. This conclusion is not yet fully supported by the presented data. Direct measurements of age-dependent divergence in acetyl-CoA and fatty acid levels at the single-cell level would be needed to substantiate this model. Based on the current evidence, a more conservative interpretation would be that aging cells exhibit differential sensitivity to perturbations in acetyl-CoA and lipid metabolism. Accordingly, I recommend revising the statement in the Abstract ("We further implicate lipid starvation and excess acetyl coenzyme A availability as major drivers of senescence...") and the corresponding discussion text to better align with the data.

      We agree and have adjusted the abstract to make it clearer that the lipid starvation / excess acetyl-coA interpretation is a model.

      “Our findings support a model in which lipid starvation and excess acetyl-coenzyme A availability are major drivers of senescence in replicatively aged wild-type yeast.”

      Reviewer #3 (Public review):

      Summary:

      These findings suggest that PGPD-SAK1 yeast show a subpopulation with lowered TOM70-GFP expression in high bud scar staining aged cells. Deletion of CAT2 or MLS1 reduces this effect. A PGPD-SAK1 acc1S1157A double mutant (called "A2A" here) shows an even larger effect of lowered tom70 expression in high bud scar staining aged cells. Utilization of various additional mutants involved in acetyl-CoA transport, carnitine shuttle, respiration, etc., leads the authors to conclude that these shifts in TOM70-GFP in aged cells are linked to the AMPK-fatty acid metabolic regulatory system.

      Strengths:

      These extensive and clearly described experiments reveal interesting changes in TOM70-GFP intensity in subsets of aged yeast in several mutants eventually identified as linked to the AMPK-fatty acid metabolic regulatory system.

      Weaknesses:

      (1) 3 biological replicates for mRNASeq is low.

      Thank you for pointing this out. We performed another replicate after posting the initial preprint to confirm the finding but didn’t update the figure in the eLife-reviewed version. We have added this to the scatter plots and analysis in Figure 1, there are minor changes but the set of genes we followed up are still highly significant. For ageing experiments, we sequence to n=3 as a first pass which is sufficient to detect widespread age-linked gene expression effects, and add more replicates if required to solidify findings for specific sets of genes. Hence, the additional RNAseq experiments we have added to the manuscript to Address Reviewer 2’s comments on widespread gene expression effects are also n=3-4.

      (2) While "Traditional conceptions of ageing implicate a progressive accumulation of damage leading to systemic degradation in performance until death, with evolutionary pressures acting to maximise early life fitness and fecundity at the expense of ageing health." is tangential perhaps to the data and conclusions of the study, both claims of this sentence are at best controversial, and the manuscript is no weaker for their omission.

      We would prefer not to remove this sentence, which we see as important to a major message of the manuscript: that ageing does not have to involve a loss of fitness before death. Outside the ageing biology field, ageing is often described as the progressive wearing out of components leading to decline and death (‘like an old car’ is a common analogy); in the ageing field this is certainly controversial, but outside the field it remains the normal understanding. This is what we mean by traditional conceptions, and it is important to consider the contradiction between this widely held viewpoint and our findings (and of course those of many others in the ageing field).

      The second part of the sentence about evolutionary pressures alludes to antagonistic pleiotropy, which we have now made explicit. Antagonistic pleiotropies as a driving mechanism for ageing, while not universally accepted, are as far as we can tell the most widely accepted type of theory in the ageing field. Our interpretation that yeast are bet-hedging as a population growth strategy and this drives ageing in the long term is a classic antagonistic pleiotropy and we need to raise this concept in the introduction.

      (3) The statement that "Here, we determine the basis of senescence and fitness loss in replicatively ageing yeast" is a bit strong as a summary of the present careful work presented here. If the authors had created yeast mutants that retained fitness indefinitely, this would be a more appropriate strength of claim to summarize the work.

      We agree and have moderated this sentence:

      “Here, we show that senescence and fitness loss in replicatively ageing yeast can be almost completely avoided without extension of lifespan by rewiring the conserved AMPK-fatty acid metabolic regulatory system.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The labelling of Figure 3G horizontal axis needs to be realigned with the data.

      Fixed – thank you.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 3G, the x-axis labels appear misaligned and should be corrected for clarity.

      Fixed – thank you.

      (2) Figures S3B and S3C appear to be mislabeled and should be revised.

      Fixed – thank you.

      (3) On page 6 (3rd paragraph), the statement that the beneficial impact arises from acetyl-CoA removal "rather than a benefit of respiration" may be overstated. The data support a role for acetyl-CoA removal but do not fully exclude a contribution from respiration. A more balanced phrasing would improve accuracy.

      We have revised this sentence and also added data:

      “Working in sip2Δ to avoid an increase in AMPK activity due to reduced Acetyl-CoA availability, we observed that ald6Δ increased the low senescence population through decreasing Tom70-GFP (S3C), and therefore the beneficial impact of PGPD-SAK1 on this pathway arises primarily through Acetyl-CoA removal. It is possible that respiration is adding to this benefit, and we detect a significant increase in Oxygen Consumption Rate in aged PGPD-SAK1 cells, but the further increase in A2A is smaller and we consider that this cannot fully explain the effect of acc1S1157A.”

      Reviewer #3 (Recommendations for the authors):

      This manuscript is clearly written, and the data are clearly presented. While 3 biological replicates is inadvisably low for mRNASeq, the subsequent experiments motivated by the genes identified there nevertheless stand on their own as presented.

      Thank you.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Noirot-Gros et. al. presents a herculean effort to map the protein-protein interactome of the c-di-GMP signaling network in Pseudomonas fluorescens (Pf). C-di-GMP, the key driver of biofilm formation in bacteria, is controlled by a highly complex network of synthesis, degradation and effector proteins. Pf is no exception as it encodes dozens of such proteins. The authors use a Yeast Two-Hybrid approach genome-wide screen with 10 diguanylate cyclase (DGC) enzymes as bait to assess protein-protein interactions in this network. The results identify over one hundred such interactions with several different hubs, including c-di-GMP signaling, other signaling systems, membrane proteins, etc. The authors then explore the original bait proteins as well as identify interactors on biofilm formation-related phenotypes and swarming using a high-throughput CRISPRi expression knockdown approach. The amount of data generated is quite impressive. Much of the manuscript uses statistical-based network analysis to group different proteins based on their interactions or impact on phenotypes, which is a high-level analysis that can catalyze further study into this system. The authors chose three specific proteins to assess their impact on cell morphology, DNA repair, and protein localization. Overall, in my view, this is perhaps the best analysis of a c-di-GMP protein-protein interactome, and it provides a multitude of hypotheses to be tested. However, therein lies the weakness of the manuscript in that very few of these hypotheses are actually tested. But such is not the goal of this network analysis type of approach. Overall, I think the work will be highly impactful to those in the c-di-GMP field, and it provides a template for others attempting such analyses of protein-protein interactions.

      Strengths:

      The manuscript is impressive in the sheer scale of the protein-protein interactions identified, network analysis, and phenotypic analysis of specific proteins in the network. It is an impressive amount of work that could be very useful to the field. It is also statistically rigorous in its analysis of significant interactions or network nodes.

      Weaknesses:

      The weakness of the manuscript is that, with three exceptions, very few of the hypotheses are actually tested. For example, BifA is shown to be a network hub protein that interacts with many other diguanylate cyclases, and this is hypothesized to be through GGDEF heterodimerization. I appreciate that experimentally testing such a hypothesis is probably another entire manuscript, but some early forays into such ideas could be undertaken using AlphaFold structural modeling of protein-protein interactions compared with GGDEFs that don't form heterodimers. Also, an inherent weakness is that such detailed analyses of a c-di-GMP signaling network, in which each diguanylate cyclase and phosphodiesterase may respond to a unique cue, is that the network identified and the conclusions made are highly specific to the experimental conditions in which the work was done. Therefore, it is unclear how broadly these conclusions (i.e. BifA is the central regulator of c-di-GMP signaling) apply to other conditions. But it is impossible to get around such a limitation, and this work can lead to testing the robustness of the identified network in other environments.

      We would like to thank the reviewer sincerely for their positive comments on our manuscript and for their constructive feedback. We recognize the limitations arising from the lack of extensive knowledge regarding the environmental cues that trigger the regulation of all CDG activities in P. fluorescens. We hypothesize that DipA acts as a central local hub that positively or negatively regulates the activity of its interacting CDG partners throughout the cell life cycle, lifestyle transitions and environmental signals. Testing this hypothesis would indeed require extensive biochemical and omics approaches. However, strengthening the significance of DipA complexes in silico using AlphaFold is a very appealing proposition and we are currently considering including this analysis in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Noirot-Gros and coworkers investigated the network of c-di-GMP associated protein complexes in Pseudomonas fluorescens. They did so by using a genome-wide yeast two-hybrid screen, and that was further probed by phenotypic screening that focused on biofilm and motility phenotypes. From this network map, they discovered that the phosphodiesterase DipA interacts with the GGDEF domains of many c-di-GMP-binding proteins.

      Strengths:

      (1) Broadness of screen led to identification of new interactions: The genome-wide yeast two-hybrid screening approach permitted broad investigation of c-di-GMP-associated protein-protein interactions. These interactions included some previously validated interactions as well as newly discovered interactions.

      (2) Complementary experimental validation: The proposed network was experimentally validated, including by using a CRISPRi-based approach in which the expression of genes encoding proteins identified in the network was systematically suppressed, and then the impact on the biofilm and motility phenotypes was assessed.

      Weaknesses:

      The findings would have been strengthened by further biochemical analysis, but this is likely beyond the scope of the paper.

      We would like to express our gratitude to the reviewer for their positive evaluation assessment, and for taking into account the limitations of the study's scope.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Noirot-Gross et al take an open-ended approach to elucidate the c-diGMP-associated protein complexes in Pseudomonas fluorescens. Starting with 10 cyclic d-GMP putative proteins, they use a combination of genome-wide two-hybrid system followed by CRISPRi-mediated exploration of phenotypes to describe the cyclic di-GMP-associated regulation of biofilm formation, and how it relates to other functions. Overall, this work presents an excellent example of how genome annotations can be further confirmed with the use of integrated functional genomic approaches. Some areas of improvement can be applied to this manuscript to enhance readability and provide a clearer distinction between confirmatory results and new findings, which are provided below:

      Strengths:

      (1) The authors have explored their findings extensively and provide a comprehensive view of the topic.

      (2) The combination of genome-wide explorations of protein-protein interactions with the more focused phenotypic exploration of the interactions found provides a solid framework for the work presented.

      Weaknesses:

      (1) Overall goal of the work:

      While articles that describe open-ended approaches can be comprehensive and descriptive in nature, the authors should have a main overall goal, which can guide the reader through the main and most compelling findings at the end. As written, the overall goal is not clear. The network perspective is interesting, and the focus on biofilm formation appears in the title. Why P. fluorescens? How is cyclic di-GMP-mediated regulation of biofilm formation in P. fluorescens different from P. aeruginosa? Why would it be studied? (Positive or negative regulation of biofilm formation?)

      We would like to express our appreciation to the reviewer for their thorough evaluation of our manuscript and for the constructive feedback they provided. The overall goal of this study will be further refined, and outlined in the introduction in the revised version of the manuscript.

      (2) Abstract:

      The abstract is very well written and guides the reader to the DipA as a hub protein in the network. From further reading, the article could clarify whether this finding is confirmatory or novel (does DipA play a similar role in P. aeruginosa?) It would be appropriate to mention the role of DipA in other Pseudomonas species from the beginning, and not only in the discussion session.

      (3) Introduction:

      The introduction is nicely written. An area of improvement could be giving more attention to protein interactions as relevant to c-di-GMP. The authors could consider an independent paragraph starting with line 84-85 "Protein-protein interactions involving DGCs, PDEs, and target effectors are crucial in establishing localized signalling through the generation of local pools of c-di-GMP", expanding on this particular aspect with an example of localized signal, after explaining that localization could help decipher specific function within the network of DGCs and PDEs. Then go into connecting biofilms with c-di-GMP and protein-protein interactions, using the example of GcbC and LapD.

      We propose highlighting the example to the local signalling cascade formed by the tripartite system YdaM, YciR and MlrA. This will be addressed in the revised version of the manuscript.

      (4) The rationale of choosing 10 PDEs could be clarified. The nice diagrams shown in the supplementary table could be used as part of Figure 1, so the reader understands why these proteins were used, and what is known about them (for example, add them as Figure 1a).

      We propose to include a specific section in the supplementary file to explain the whole rationale behind choosing these CDGs. These proteins were selected based on their involvement in different steps of biofilm formation in Pseudomonas, as well as their role in the ability of P. fluorescens strains to colonize plant roots.

      (5) Figures 1b and 2 convey the same information as in Figure 1a. They could be removed without affecting the understanding of the article.

      Figure 2 will be transferred in Supplementary as part of the Figure S1

      (6) CRISPRi and Figure 3. Figure 3 shows the methodology of CRICPR phenotypic screening. A diagram showing the CRISPRi system in P. fluorescens could help the non-expert reader. While the choice of 23 proteins related to the emerging hub DipA is clear, the choice of the other 33 genes could be better explained. Are these proteins already related to biofilm formation? Where are they part of the network detected? How about the other 14 SBW25 genes? The authors could clarify the rationale of the choices. Figure 4 could be combined with Figure 3 or moved to the supplementary material.

      A better description of the rationale behind the choice of tested interacting protein partners will be provided. We also agree to combine Figure 4 with Figure 3.

      (7) Figures 5, 6 and 7 represent solid network analysis of the findings. Still, they could be improved in clarity on the main findings. The authors conclude at the end of section 3.2.3 that there are networks that exert a "positive role" and a "negative role". The authors could show that in the figures, explaining what those roles are: more biofilm structural coding genes? positive or negative regulation of biofilm formation?)

    1. Even though secular archivists may feel uncomfortable thinking of this workas part of anything other than fulfilling professional responsibilities, we mustunderstand that such work is also a component of a contemporary ritual.

      I find the tone of this article quite interesting, as it grants archivists a role almost akin to a secularized priest, administering sacred rituals to the dying, the dead, and their survivors. I had never considered this angle on archival studies, but it is a thought-provoking way of framing this aspect of archival work. I think that it properly conveys the gravity and responsibility of working with donors near the end of their lives.

    1. representation

      I need to push back on this comment. If tokenizing provides queer people a space for representation of their community, but the "token queer" only represents a very small part of the community that is considered "acceptable" or "palatable" by others or makes heterosexual, cisgender people "less uncomfortable," then I don't think that we can really call this "representation" of the community. It is really just a perpetuation of stereotype (as the author mentions), and not at all an authentic space for representation. This may do more harm than good.

    2. Furthermore, the responsibility ofeducating students on queer issues fallson individuals who are invested in thetopic

      This idea of 'being invested in the topic' makes me question our education system deeply. The system encourages an individual to funnel into a specific topic as they move ahead in their education - elementary, secondary, undergraduate, graduate and PhD. A vegan may not think about caste while they promote the beef ban. I may advocate for one marginalized group and not the other because I am 'invested in that topic'. Transdisciplinary attempts should be encouraged throughout education so that we can carry this responsibility better. This article comes under Education Leadership, but also Social, Political and Cultural Contexts of education.

    3. make so muchnois

      I have seen this with several marginalized groups - we are forced to overcompensate just to be heard. However, sometimes the very nature of "making so much noise" perpetuates the discrimination these groups experience. I think about friends of mine who are activists combating anti-black racism, and they can be really difficult to approach and learn from. I have also had experiences with women in male-dominated spaces who are extra tough because they feel they need to be (e.g., female border officers who are intentionally mean or aggressive as an overcompensation in response to their marginalization or the assumption that they can't be tough or mean). I am not in any way saying that I'm against activist or speaking out -- I am commenting on how unfair it is that marginalized groups are "forced to make so much noise" and that, in doing so, they may not even achieve their desired outcome.

    1. Author response:

      General Statements:

      We appreciate the reviewers for the critical review of the manuscript and the valuable comments. We have carefully considered the reviewer’s comments and have revised our manuscript accordingly.

      Point-by-point description of the revisions:

      Reviewer #1 (Evidence, reproducibility and clarity):

      Major comments

      (1) This study leaves out lipid metabolism as a major energy metabolism pathway relevant to AD. The authors themselves cite the significance of acylcarnitines and CPT1A in AD (pg. 3, lines 32-33, pg. 4, lines 1-2). Lipid metabolism and homeostasis is known to be disrupted in AD1. Fatty acid oxidation is a known energy source in the prefrontal cortex2 and will also generate acetyl coA, which this study reveals is a significant decreased metabolite in AD. Furthermore, sphingomyelin emerges as one of the major decreased DEMs as well. Thus, lipid metabolism should be highlighted in Figure 3 and discussed throughout the manuscript; otherwise its omission should be clearly stated and justified.

      We appreciate the reviewer’s insightful comment regarding a critical role of lipid metabolism in AD. We recognize that lipid metabolism is a metabolic pathway deeply involved in AD pathology (Baloni et al., 2022, 2020; Varma et al., 2021). Accordingly, we have revised the Limitations section to more strongly emphasize its role as a vital energy source (pg. 13, lines 15-17). Regarding the visualization of lipid metabolism, we extracted lipid-related pathway from the trans-omic network but found that the regulatory relationships among DEPs and DEMs were excessively complex and interconnected. Thus, interpreting this regulatory network seemed to be more challenging compared to the other energy production pathways presented in our manuscript. Therefore, we have concluded that the pathway analysis in our trans-omic network may not be suitable for deeply elucidating the lipid dysregulation in AD. We have added a statement acknowledging this as a limitation of our current methodology in the revised manuscript (pg. 13, lines 13-22).

      (2) The covariates used for differential analysis should be discussed and justified. Notably, age is used as a covariate for transcriptomic analysis but not proteomic and metabolomic analysis, with no justification. Additionally, given the known importance of lipid metabolism in AD and the putative role of APOE in lipid homeostasis3, APOE genetic status should be considered as a covariate, or its omission should be justified.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, such as age at death and RIN, is that these data were not available for each sample. Thus, we referred to the original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, and years of education. Regarding the proteomic dataset, in the original article (Johnson et al., 2020), age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (3) The authors make a conclusion statement that suggests intervention: "Collectively, our data suggests that preserving or improving the ability to produce ATP and early intervention in the process of nitrogen metabolism are candidates for the prevention and treatment of dementia" (pg. 12, lines 12-14). This claim is not well-supported by the evidence provided in the study. There are a few limitations: (a) This was an observational, not interventional study; (b) The study did not establish whether the metabolic disruptions are causes or effects in AD; and (c) ATP or other bioenergetic indicators were not directly measured. Therefore, any statements about potential interventions should be removed or qualified as highly speculative.

      We agree with the reviewer that the statement regarding potential interventions was not sufficiently supported by our analyses. Accordingly, we have removed the sentence regarding prevention and treatment from the revised manuscript (e.g., we have deleted final paragraph of the previous manuscript).

      (4) In conjunction with the last point, the main conclusion of the study is that energy production is down in AD. The data presented in Figure 3 are consistent with this conclusion, but it is far from definitive due to limitations stated above in comments 3a and 3b. The authors should offer additional support for this conclusion: experimental follow-up, flux modeling, analysis of alternative datasets with ATP measurement, causal inference.

      We sincerely thank the reviewer for this valuable and constructive suggestion. Regarding flux modeling, we agree that metabolic flux analysis could provide important mechanistic insight. Indeed, previous studies have applied flux modeling in the context of lipid metabolism in Alzheimer’s disease (Baloni et al., 2022). We also attempted to perform flux modeling focusing on energy metabolism. However, we found it difficult to obtain biologically meaningful and robust results and therefore decided not to include these analyses in the current manuscript.

      With respect to ATP measurements, we fully agree that direct evidence of altered ATP levels would further strengthen our conclusion. However, to the best of our knowledge, there are currently no publicly available large-scale datasets that directly measure ATP levels in human postmortem brain tissues. This limitation makes it challenging to incorporate validation in the present study.

      Regarding experimental follow-up, we agree that functional validation is essential to confirm the mechanistic implications of our findings. We are actively considering follow-up experimental studies. However, we consider the present work to be a multi-omic integrative analysis aimed at identifying key molecular alterations and generating biologically important hypotheses. We have revised the Limitation section to more clearly position this manuscript as an observational systems-level analysis (pg. 13, lines 20-22).

      (5) The validation analysis did not sufficiently show the generalizability of this study's results. The authors demonstrated a correlation of 0.53 to the MSBB transcriptomics data and 0.60 to the AMP-AD DiverseCohorts proteomics data. Beyond these correlation coefficients, no meaningful comparison between the datasets is offered. How concordant are the differentially expressed features (or pathways) between the datasets? How robust would the trans-omic network be if incorporating the alternate datasets? Is the main conclusion (energy metabolism is down in AD) supported by the validation datasets? We think this analysis should be expanded and described in the main text.

      Although the results for external metabolomics datasets are reported in Fig S2C, correlation coefficients with the external data are not reported. The authors state, "Note that each study used different definitions for AD and CT groups, had variations in measurement methods and brain regions analyzed." We appreciate these limitations. However, the external data should be re-analyzed using the same definitions of AD and CT, if possible. The limitations and results (which DEMs are shared between datasets) should be discussed in the main text.

      We thank the reviewer for this important comment regarding the generalizability of our findings. In the revised manuscript, we have expanded the validation analyses and summarized the results in Figure S2. First, at the transcriptomic level, Figure S2B and S2C show the overlap between up- and downregulated genes in AD identified in our ROSMAP-derived analyses and those reported in a previously published large-scale meta-analysis of 2,114 postmortem samples across seven brain regions (Wan et al., 2020). A substantial proportion of DEGs were shared, supporting cross-cohort and cross-region robustness to some extent. At the proteomic level, Figure S2E shows a comparison between the ROSMAP and the AMP-AD DiverseCohorts datasets. We highlighted the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3 and calculated a separate correlation coefficient for this subset (Pearson coefficient = 0.86, p-value = 1.5e-7), further supporting our main conclusion. In addition, to assess the concordance between the two datasets in a threshold-independent manner, we additionally performed Rank-Rank Hypergeometric Overlap (RRHO) analysis (Figure S2E). RRHO analysis (Cahill et al., 2018; Plaisier et al., 2010) enables the comparison of ranked protein lists without relying on arbitrary differential expression cutoffs and has been used for cross-dataset comparison in several previous studies (Fröhlich et al., 2024; Maitra et al., 2023). The RRHO heatmaps demonstrated significant enrichment in the concordant quadrants, confirming systematic agreement between datasets beyond simple correlation coefficients. For metabolomics, Figure S2G shows RRHO analyses comparing the ROSMAP metabolomic data with other datasets measured by the same UPLC-MS/MS platform (Batra et al., 2024; Novotny et al., 2023), demonstrating significant concordance in ranked metabolite changes in AD.

      (6) The glycolysis analysis and discussion needs more development. Glycolysis and gluconeogenesis share many of the same enzymes, but they are not the same pathway and should not be discussed as such. To make a claim about the overall influence of enzyme and metabolite levels on glycolysis, the authors should focus on the energetically committing steps of glycolysis (hexokinase, phosphofructokinase, pyruvate kinase) in Figure 3A, and include the full/current version of the figure in the supplement. Gluconeogenesis-specific enzymes (pyruvate carboxylase, PEPCK) are not mentioned at all - are they among the DEPs/DEGs?

      We appreciate the reviewer’s comment regarding the distinction between glycolysis and gluconeogenesis pathway. Among the gluconeogenesis-specific enzyme proteins, G6PC1, FBP1, PC, and PCK2 were measured in our dataset, but none of them were identified as DEPs. In addition, gluconeogenesis is a process that occurs primarily in the liver and kidney rather than the brain. Given this biological context and the lack of significant changes in relevant enzymes, we have revised the terminology throughout the manuscript, replacing “glycolysis/gluconeogenesis pathway” with “glycolysis pathway” in the revised version.

      (7) Given that there wasn't good concordance between the DEGs and DEPs, did including the mRNA and transcription factor layers in the network really add anything useful? It seems like the main conclusions of the manuscript were driven by the protein and metabolite layers only. How many of the DE metabolic enzymes were coregulated at the transcript and protein level? It would be useful to include the 5-layer trans-omic network in the supplement to display these results. Given your network, at what level does it appear that energy metabolism is regulated?

      It is true that our primary conclusion regarding the regulation of energy metabolism is driven by the changes in protein and metabolite abundance. However, we consider the low concordance between mRNA and protein expression itself to be an important feature of AD pathology, as also reported in previous studies (Johnson et al., 2022; Tasaki et al., 2022). Although we did not perform a further analysis of this discordance, we believe that including the TF and mRNA layers into the metabolic trans-omic network strengthens a system-wide view of metabolic dysregulation in AD.

      Regarding the mRNA changes corresponding to the DEP enzymes, please refer to Figure S7A.

      (8) Comment further on the results from Figure 2D. What can be learned from identifying metabolites with the greatest degree centrality? What pathways other than energy metabolism are highlighted by the trans-omic network?

      We assume that some energetic indicators, including AMP and acetyl-CoA, and nitrogen metabolism-related metabolites, Glu, 2-oxoglutarate, and urea, can be potential key regulators of dysregulated metabolism in AD.

      (9) (Suggestion) We suggest the authors leverage their trans-omic network in additional ways beyond giving a snapshot of a few energy metabolism pathways. The analysis of top DEMs could go further. What pathways are impacted beyond energy metabolism? Among the metabolic reactions allosterically regulated by top DEMs, what metabolic pathways are enriched?

      We identified the enriched metabolic pathways that were allosterically regulated by DEMs in AD using Fisher’s exact test. Alanine, aspartate, and glutamate metabolism pathways were significantly enriched in 2-oxoglutarate, glutarate, alanine, and glutamate-regulating metabolic reactions. Arginine and proline metabolism pathway was enriched in N-methyl-L-arginine and putrescine-regulating metabolic reactions. Arginine biosynthesis pathway was enriched in arginine-regulating metabolic reactions. Glycerophospholipid metabolism pathway was enriched in CDP-ethanolamine-regulating metabolic reactions. Glycine, serine, and threonine metabolism pathway was enriched in serine-regulating metabolic reactions. Purine metabolism pathway was enriched in AMP-regulating metabolic reactions. Pyrimidine metabolism pathway was enriched in deoxyuridine and thymidine-regulating metabolic reactions. Sphingolipid metabolism pathway was enriched in sphingosine-regulating metabolic reactions. However, this analysis did not yield sufficiently valuable insights into the regulatory relationships among biomolecules in AD. Thus, we did not include these results in the revised manuscript.

      (10) (Suggestion) Figure 3 shows that most differential signal in AD points to lower energy production due to the combination of differentially expressed metabolites and enzymes, but we are not given much context about the strength of these among all the differential signals. We would suggest including volcano plots where the features of interest, i.e. DE enzymes and metabolites, are colored differently (or a similar figure).

      We thank the reviewer for this constructive suggestion. To provide better context regarding the importance of the differential signals, we have added volcano plots for mRNAs, proteins, and metabolites in Figure S4A, B, and C.

      (11) (Suggestion) The PPI network could be better leveraged to understand metabolic changes in AD. If nodes are grouped into subnetworks (e.g. by Louvain / Leiden clustering) and tested for pathway enrichment, could you find functional subnetworks of coordinately up- and down- regulated metabolic enzymes? This could yield some pathways of interest beyond the energy metabolism pathways already highlighted.

      We appreciate the reviewer’s suggestion to utilize the PPI network for subnetwork analysis. However, it is important to note that the proteomic dataset analyzed in this study is derived from the original work of (Johnson et al., 2020). In that paper, the authors already performed a Weighted Gene Co-expression Network Analysis (WGCNA) across several datasets to identify co-expressed modules and functional pathways.

      Given this, we assumed that applying additional clustering methods to the same dataset would be unlikely to yield significant biological insights beyond the established findings.

      Minor comments

      (1). "All genes" and "all metabolites" should not be the background for the proteomic and metabolic pathway enrichment analysis by Metascape and MetaboAnalyst. The background should be limited to the proteins and metabolites that were measured.

      We fully agree with the reviewer that using “all gene” or “all metabolites” as a background is not suitable for enrichment analyses. As suggested, we have revised the enrichment analyses using the measured proteins and metabolites as a background in both Metascape and MetaboAnalyst (Fig. S4D).

      (2) Highlight the metabolic enzymes in Fig S2B. Calculate a separate correlation coefficient for the enzymes extracted in the energy metabolism analysis from Fig 3.

      We appreciate the reviewer’s suggestion to refine the correlation analysis. As requested, we have revised Fig. S2D to explicitly highlight the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3. We calculated a separate correlation coefficient for the subset (Pearson coefficient = 0.86, p-value = 1.5e-7).

      (3) Use a multiple hypothesis adjusted p-value or q-value in Figure S3.

      We agree with the reviewer regarding the necessity of correcting for multiple comparisons. Accordingly, we have revised Fig. S4D using q-values.

      (4) Describe the methods used to calculate the logFC values from the validation dataset.

      We have revised the Methods to include a detailed description of the procedure used to calculate the log2FC values for the validation datasets (pg. 21, lines 13-15).

      (5) It is difficult to read Figure 3. We would recommend really emphasizing to the reader to refer to Fig S7B as a "key" to this figure. The description of the red/blue arrows and nodes in the methods section (pg. 24, lines 21-36, pg 25, lines 1-4) were also helpful, but very lengthy. We recommend putting an abridged version of this description into the Fig S7 figure legend.

      We appreciate the feedback regarding the readability of Fig. 3. As recommended, we have revised the manuscript to explicitly direct readers to Fig. S8B as an essential “key” for interpreting the network visualization (pg. 8, lines 28). Furthermore, we have added an abridged description of the network elements to the legend of Fig. S8B.

      (6) The S7 figure legend should refer to panels A and B, not E and F.

      We apologize for this oversight. We have corrected the legend of Fig. S8.

      (7) (Suggestion) Are any of the differentially expressed metabolites allosteric regulators of the DE transcription factors? This could be interesting to discuss.

      We appreciate the reviewer’s insightful suggestion about the potential allosteric regulation of the DETFs by DEMs. We conducted an extensive literature search to identify any reports related to this perspective. However, to the best of our knowledge, no such direct interactions have been reported to date.

      Reviewer #1 (Significance):

      The study's strength lies in leveraging three omics modalities across large patient cohorts (n ~ 150-240) to identify coherent signals between transcriptomics, proteomics, and metabolomics in postmortem DLPFC tissue. It was encouraging to see that the main result, showing downregulation for TCA, oxidative phosphorylation, and ketone body metabolism, emerged from consistent signals across both proteomics and metabolomics. This result was consistent with previous findings in other models cited by the author4,5 and other studies 6,7 demonstrating deficiency in energy-producing pathways in AD.

      Another strength of the study is the application of thoughtful methodology to connect differentially expressed proteins and metabolites via an intermediate data layer of metabolic reactions. The authors leverage the KEGG and BRENDA databases and apply sound logic to estimate the effects of enzyme level and metabolite level on pathway activity, with metabolites serving as substrate, product, or allosteric regulator for reactions. This trans-omic network methodology was developed in previous studies cited by the author8,9.

      However, as written, this study is limited in its contribution of new knowledge to the AD research field. The main conclusion (energy production is down in AD, due to regulatory disruption of energy metabolism) is not strongly supported (see comments 1, 3, and 4 for elaboration). The evidence could be improved by orthogonal approaches: further experimentation, further integration of external datasets, causal modeling, or flux modeling. Alternatively, even in the absence of new experimental and computational approaches, the story could be made more complete by further leveraging the trans-omic network to provide insights into (a) the regulation of energy metabolism; and (b) the impacts of key disrupted metabolites (see comments 7-9).

      The study is also limited in its demonstrating the power of these methodologies to provide integrative insights. As mentioned above, the integration of enzyme levels and metabolite levels is clearly useful (Figure 3). In contrast, the utility of the mRNA and transcription factor layers was not evident. The study did not appear to improve or expand upon trans-omic network methodology described in the previous works. Finally, the various analyses (analyzing the trans-omic network for nodes with the highest degree centrality, the PPI analysis, and viewing the energy metabolism pathways in the network) provided disparate results that were only tenuously connected in the discussion section.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary

      This manuscript integrates public transcriptomic, proteomic, and metabolomic datasets from ROSMAP DLPFC samples to construct a multi-layer metabolic trans-omic network in Alzheimer's disease. By linking transcription factors, enzyme mRNAs, proteins, metabolic reactions, and metabolites, the authors report coordinated downregulation of the TCA cycle, oxidative phosphorylation, and ketone body metabolism, along with mixed regulatory signals in glycolysis/gluconeogenesis. They interpret these patterns as indicative of broad energetic dysfunction and alterations in amino-acid/nitrogen metabolism in AD. While the framework is conceptually appealing, much of the analysis remains descriptive, and several biological interpretations extend beyond what the data can robustly support. The reliance on bulk tissue without accounting for cell-type composition, limited covariate adjustment, and the absence of validation or sensitivity analyses reduce confidence in the mechanistic conclusions. Overall, the study provides a preliminary systems-level overview, but additional rigor is needed before the proposed trans-omic regulatory insights can be considered convincing.

      Major Comments

      (1) Interpretation requires more cautious phrasing, and validation is essential. The manuscript frequently asserts that specific pathways are "inhibited" or that energetic deficits are "compensated," but these conclusions extend beyond what the descriptive, bulk-level data can support. Because no metabolic flux, causality, or direct functional measurements are included, the results should be framed as putative regulatory shifts, not confirmed impairments. Critically, key claims about pathway inhibition would require flux modeling, perturbation analyses, or experimental validation to be convincing. Without such validation, the mechanistic interpretations remain speculative.

      We thank the reviewer for this crucial comment. We fully agree that, given the descriptive and bulk-level nature of our analysis, mechanistic interpretations must be made with caution. In the absence of direct metabolic flux measurements or experimental validation, our findings should be interpreted as putative regulatory shifts rather than confirmed functional impairments. Accordingly, we have revised the manuscript to temper mechanistic claims. We have replaced definitive statements with more speculative phrasing (e.g., “Our analysis revealed a putative coordinated downregulation …” instead of “Our analysis revealed a coordinated downregulation …” in Abstract section; “we demonstrate the systems-level view of the potential dysregulated energy production …” instead of “we demonstrate the systems-level view of the dysregulated energy production …” in pg. 10, lines 25-26).

      (2) Although the authors acknowledge this in the limitations, bulk-level differences may primarily reflect altered proportions of neurons, astrocytes, microglia, and oligodendrocytes rather than true within-cell-type regulation. Incorporating a cell-type deconvolution or performing a sensitivity analysis would substantially improve interpretability. This issue also impacts the trans-omic network: if the molecules included originate from different cell types, the inferred regulatory relationships may not reflect true intracellular processes.

      We appreciate the reviewer’s point that bulk-level differences can reflect altered proportions of different brain cell types, subsequently affecting the inferred trans-omic network analysis. To assess the changes in cell type proportions of the samples that we used in our study, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglias, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two group. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      (3) Differential analysis covariates. For the differential expression analyses, only gender and PMI were included as covariates. Additional variables, such as age at death, RIN, neuropathological measures, and comorbidities, can strongly influence molecular profiles and should be considered to ensure that the observed differences reflect AD-related biology rather than confounding pathological or technical factors.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, including age at death and RIN, is that these data for each sample were not available. Thus, we referred to original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, or education. Regarding the proteomic dataset, in the original article, age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (4) Network stability and sample non-overlap. Proteomic, transcriptomic, and metabolomic data come from partially overlapping individuals. The authors should test whether the reconstructed network is robust to: different significance thresholds, restricting analyses to overlapping samples and alternative definitions of AD vs control.

      We appreciate the reviewer’s comment for the trans-omic network stability. In our study, the number of individuals for whom all omic modalities were measured was relatively small (n=25 in CT and n=35 in AD). This limited overlap reduces statistical power and can affect the downstream network construction. We have acknowledged this limitation in the revised manuscript and clarified that the reconstructed networks should be interpreted with caution regarding reproducibility and generalizability (pg. 13, lines 13-23).

      Minor Comments

      (1) Some TF enrichment and regulatory inferences lack explicit mention of multiple-testing correction.

      We apologize for the lack of clarity in our original description. We have corrected for multiple-testing for the TF inference. Thus, we have revised the Methods section to explicitly describe the correction method used and the threshold applied (pg. 23, lines 23-24).

      (2) The limitations section is strong but should explicitly discuss the influence of postmortem interval on metabolite levels.

      We appreciate the reviewer’s comment about the effect of postmortem interval on changes in metabolite levels. Accordingly, we have added the description of this perspective in our revised manuscript (pg. 13, lines 1-5).

      Reviewer #2 (Significance):

      The study extends a trans-omic integration framework, originally applied to metabolic disease, into the context of Alzheimer's pathology. Although the biological findings largely confirm known alterations in mitochondrial and energy metabolism, the network-based approach offers a structured way to view cross-layer regulatory changes. Its main advance is conceptual rather than biological, providing a unified framework rather than uncovering fundamentally new mechanisms. This work will primarily interest researchers in neurodegeneration and systems biology, as well as computational groups developing multi-omics integration methods.

      Reviewer #3 (Evidence, reproducibility and clarity):

      This study leverages existing transcriptomic, metabalomic and proteomic datasets from prefrontal cortex (PFC) to assess metabolic dysregulation in Alzheimer's disease (AD). They found a downregulation of multiple metabolic pathways, including TCA cycle, oxidative phosphorylation, and ketone metabolism, that may explain bioenergetic alterations in AD.

      The study used matching ROSMAP omics datasets from the DLPFC that have allowed more robust data integration. However, the datasets are all generated using bulk tissue, which makes data interpretation difficult. For example, the AD changes they observed may be due to shifts in cell type proportion with disease (e.g. cell death, neuron inflammation). Did the authors account for any potential shifts in cell type proportion in their analysis?

      If the assumption is that the changes in AD are cell intrinsic, which cell types are likely to be impacted? Can the authors integrate any existing single-cell analysis to infer which cell types may be driving the signals they detect, and whether this accounts for some of the antagonistic regulatory effects that were detected?

      We thank the reviewer for their insightful comments. We agree that the use of bulk tissue datasets cannot account for cell-type heterogeneity. As noted in our Limitations section (pg. 12, lines 24-27), we recognize that previous studies have found that the Braak stage is correlated positively with microglia and astrocyte proportions and negatively with oligodendrocyte proportion (Hannon et al., 2024; Shireby et al., 2022). Regarding the integration of single-cell analysis, we have referenced recent snRNA-seq findings (Mathys et al., 2024) in our Limitations section (pg. 12, lines 28-32) to deconvolve our bulk signatures.

      Furthermore, in our revised manuscript, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglia, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two groups. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      Reviewer #3 (Significance):

      The manuscript provides multimodal insight into metabolic dysregulation in AD in the PFC. Given that metabolic dysfunction is likely to play a major in disease pathogenesis, this is a study of importance. However, the findings lack granularity at the cell type level, which limits the impact of the study.

      Reference

      (1) Baloni, P., Arnold, M., Buitrago, L., Nho, K., Moreno, H., Huynh, K., Brauner, B., Louie, G., Kueider-Paisley, A., Suhre, K., Saykin, A. J., Ekroos, K., Meikle, P. J., Hood, L., Price, N. D., Alzheimer’s Disease Metabolomics Consortium, Doraiswamy, P. M., Funk, C. C., Hernández, A. I., … Kaddurah-Daouk, R. (2022). Multi-Omic analyses characterize the ceramide/sphingomyelin pathway as a therapeutic target in Alzheimer’s disease. Communications Biology, 5(1), 1074.

      (2) Baloni, P., Funk, C. C., Yan, J., Yurkovich, J. T., Kueider-Paisley, A., Nho, K., Heinken, A., Jia, W., Mahmoudiandehkordi, S., Louie, G., Saykin, A. J., Arnold, M., Kastenmüller, G., Griffiths, W. J., Thiele, I., Alzheimer’s Disease Metabolomics Consortium, Kaddurah-Daouk, R., & Price, N. D. (2020). Metabolic Network Analysis Reveals Altered Bile Acid Synthesis and Metabolism in Alzheimer’s Disease. Cell Reports. Medicine, 1(8), 100138.

      (3) Batra, R., Arnold, M., Wörheide, M. A., Allen, M., Wang, X., Blach, C., Levey, A. I., Seyfried, N. T., Ertekin-Taner, N., Bennett, D. A., Kastenmüller, G., Kaddurah-Daouk, R. F., Krumsiek, J., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2023). The landscape of metabolic brain alterations in Alzheimer’s disease. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(3), 980–998.

      (4) Batra, R., Krumsiek, J., Wang, X., Allen, M., Blach, C., Kastenmüller, G., Arnold, M., Ertekin-Taner, N., Kaddurah-Daouk, R., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2024). Comparative brain metabolomics reveals shared and distinct metabolic alterations in Alzheimer’s disease and progressive supranuclear palsy. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 20(12), 8294–8307.

      (5) Cahill, K. M., Huo, Z., Tseng, G. C., Logan, R. W., & Seney, M. L. (2018). Improved identification of concordant and discordant gene expression signatures using an updated rank-rank hypergeometric overlap approach. Scientific Reports, 8(1), 9588.

      (6) Fröhlich, A. S., Gerstner, N., Gagliardi, M., Ködel, M., Yusupov, N., Matosin, N., Czamara, D., Sauer, S., Roeh, S., Murek, V., Chatzinakos, C., Daskalakis, N. P., Knauer-Arloth, J., Ziller, M. J., & Binder, E. B. (2024). Single-nucleus transcriptomic profiling of human orbitofrontal cortex reveals convergent effects of aging and psychiatric disease. Nature Neuroscience, 27(10), 2021–2032.

      (7) Green, G. S., Fujita, M., Yang, H.-S., Taga, M., Cain, A., McCabe, C., Comandante-Lou, N., White, C. C., Schmidtner, A. K., Zeng, L., Sigalov, A., Wang, Y., Regev, A., Klein, H.-U., Menon, V., Bennett, D. A., Habib, N., & De Jager, P. L. (2024). Cellular communities reveal trajectories of brain ageing and Alzheimer’s disease. Nature, 633(8030), 634–645.

      (8) Hannon, E., Dempster, E. L., Davies, J. P., Chioza, B., Blake, G. E. T., Burrage, J., Policicchio, S., Franklin, A., Walker, E. M., Bamford, R. A., Schalkwyk, L. C., & Mill, J. (2024). Quantifying the proportion of different cell types in the human cortex using DNA methylation profiles. BMC Biology, 22(1), 17.

      (9) Johnson, E. C. B., Carter, E. K., Dammer, E. B., Duong, D. M., Gerasimov, E. S., Liu, Y., Liu, J., Betarbet, R., Ping, L., Yin, L., Serrano, G. E., Beach, T. G., Peng, J., De Jager, P. L., Haroutunian, V., Zhang, B., Gaiteri, C., Bennett, D. A., Gearing, M., … Seyfried, N. T. (2022). Large-scale deep multi-layer analysis of Alzheimer’s disease brain reveals strong proteomic disease-related changes not observed at the RNA level. Nature Neuroscience, 25(2), 213–225.

      (10) Johnson, E. C. B., Dammer, E. B., Duong, D. M., Ping, L., Zhou, M., Yin, L., Higginbotham, L. A., Guajardo, A., White, B., Troncoso, J. C., Thambisetty, M., Montine, T. J., Lee, E. B., Trojanowski, J. Q., Beach, T. G., Reiman, E. M., Haroutunian, V., Wang, M., Schadt, E., … Seyfried, N. T. (2020). Large-scale proteomic analysis of Alzheimer’s disease brain and cerebrospinal fluid reveals early changes in energy metabolism associated with microglia and astrocyte activation. Nature Medicine, 26(5), 769–780.

      (11) Maitra, M., Mitsuhashi, H., Rahimian, R., Chawla, A., Yang, J., Fiori, L. M., Davoli, M. A., Perlman, K., Aouabed, Z., Mash, D. C., Suderman, M., Mechawar, N., Turecki, G., & Nagy, C. (2023). Cell type specific transcriptomic differences in depression show similar patterns between males and females but implicate distinct cell types and genes. Nature Communications, 14(1), 2912.

      (12) Mathys, H., Boix, C. A., Akay, L. A., Xia, Z., Davila-Velderrain, J., Ng, A. P., Jiang, X., Abdelhady, G., Galani, K., Mantero, J., Band, N., James, B. T., Babu, S., Galiana-Melendez, F., Louderback, K., Prokopenko, D., Tanzi, R. E., Bennett, D. A., Tsai, L.-H., & Kellis, M. (2024). Single-cell multiregion dissection of Alzheimer’s disease. Nature, 632(8026), 858–868.

      (13) Novotny, B. C., Fernandez, M. V., Wang, C., Budde, J. P., Bergmann, K., Eteleeb, A. M., Bradley, J., Webster, C., Ebl, C., Norton, J., Gentsch, J., Dube, U., Wang, F., Morris, J. C., Bateman, R. J., Perrin, R. J., McDade, E., Xiong, C., Chhatwal, J., … Harari, O. (2023). Metabolomic and lipidomic signatures in autosomal dominant and late-onset Alzheimer’s disease brains. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(5), 1785–1799.

      (14) Plaisier, S. B., Taschereau, R., Wong, J. A., & Graeber, T. G. (2010). Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures. Nucleic Acids Research, 38(17), e169.

      (15) Shireby, G., Dempster, E. L., Policicchio, S., Smith, R. G., Pishva, E., Chioza, B., Davies, J. P., Burrage, J., Lunnon, K., Seiler Vellame, D., Love, S., Thomas, A., Brookes, K., Morgan, K., Francis, P., Hannon, E., & Mill, J. (2022). DNA methylation signatures of Alzheimer’s disease neuropathology in the cortex are primarily driven by variation in non-neuronal cell-types. Nature Communications, 13(1), 5620.

      (16) Tasaki, S., Xu, J., Avey, D. R., Johnson, L., Petyuk, V. A., Dawe, R. J., Bennett, D. A., Wang, Y., & Gaiteri, C. (2022). Inferring protein expression changes from mRNA in Alzheimer’s dementia using deep neural networks. Nature Communications, 13(1), 655.

      (17) Varma, V. R., Wang, Y., An, Y., Varma, S., Bilgel, M., Doshi, J., Legido-Quigley, C., Delgado, J. C., Oommen, A. M., Roberts, J. A., Wong, D. F., Davatzikos, C., Resnick, S. M., Troncoso, J. C., Pletnikova, O., O’Brien, R., Hak, E., Baak, B. N., Pfeiffer, R., … Thambisetty, M. (2021). Bile acid synthesis, modulation, and dementia: A metabolomic, transcriptomic, and pharmacoepidemiologic study. PLoS Medicine, 18(5), e1003615.

      (18) Wan, Y.-W., Al-Ouran, R., Mangleburg, C. G., Perumal, T. M., Lee, T. V., Allison, K., Swarup, V., Funk, C. C., Gaiteri, C., Allen, M., Wang, M., Neuner, S. M., Kaczorowski, C. C., Philip, V. M., Howell, G. R., Martini-Stoica, H., Zheng, H., Mei, H., Zhong, X., … Logsdon, B. A. (2020). Meta-Analysis of the Alzheimer’s Disease Human Brain Transcriptome and Functional Dissection in Mouse Models. Cell Reports, 32(2), 107908.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial; by combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength. However, the evidence supporting key conclusions is incomplete, with several claims insufficiently substantiated by the data presented. Improvements in data presentation (e.g., quantification of qualitative statements, statistical estimates, and clearer description of results) would substantially strengthen the paper.

      We thank the editor and reviewers for their positive comments. We have revised the manuscript in response to the points raised, as described below.

      First of all, we would like to summarise what we have changed regarding the presentation of summary statistics throughout the text. We agree that many of the statements in the original submission tended towards being qualitative. This was the result of shying away from presenting two separate estimates, with different ways of quantifying uncertainty, in the text. The phylogenetics could be presented as mean and confidence interval, while the IBM would need some measure of centrality (mean or median) and the highest density interval for a summary statistic (e.g. the mean age gap) as it varies over the posterior. These are not directly comparable. We have now changed this to present both where appropriate, with cautionary note about the difference between the CIs and HDIs (lines 257-260).

      We also were somewhat arbitrary regarding where we chose to summarise the posterior in the IBM or look at the best-fitting single simulation, and where we presented the mean as opposed to the median. We have done a considerable overhaul of what is presented in this revision:

      (1) We always present the posterior summary unless the level of detail is such that summarising uncertainty over the posterior is not feasible (e.g. in figures 3, 4 and 5). In the latter case we still use the best-fitting IBM replicate.

      (2) In the main text we always present the mean. For the phylogenetics the summary statistics are mean and confidence interval. For the IBM this is the posterior mean, and 95% HDI, of the mean of a particular statistic as calculated in each of the 1000 IBM replicates. For example, each replicate will have its own distribution of male source ages which have a mean value. These means also vary over the posterior, and a mean of them is calculated, as well as the HDI interval to represent posterior uncertainty. This “mean of means” may be a slightly confusing piece of terminology at first glance, but it allows us to properly capture posterior uncertainty in a way we mostly avoided in the first submission.

      One result of 1) above is a change to figure 6. It is now summarised over the posterior, with the result that time trends that were previously not evident become clear. This changes our conclusions slightly (lines 529-537) but it should be noted that the magnitudes of the trends remain small.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia. The current manuscript draft is methodologically strong, but needs revision to strengthen the take-home messages. As written, there are many possible take-away conclusions. For example, the agreement between IBM and phylogenetic analysis is noteworthy and provides a methodological focus. The revealed age patterns of transmission could be a focus. The effects of the PopART intervention and the consequences of a 1-year disruption could be a focus. It is important, though, that any main messages summarized by the authors are substantiated by the evidence provided and do not extrapolate beyond the data that have been generated. I recommend that the authors think deeply about what the most important, well-supported messages are and reframe the discussion and abstract accordingly.

      We have rewritten the abstract, and also made changes to the discussion in order to centre our message around the contribution of particular of demographic groups to transmission, and how, with that contribution revealed, such groups can be selected for specialised interventions.

      Strengths/weaknesses by section:

      (1) ABSTRACT

      The Abstract summarizes qualitative findings nicely, but the authors should incorporate quantitative results for all of the qualitative findings statements.

      The abstract in the revision is extensively revised, and contains quantitative estimates throughout, from both methodologies where appropriate.

      The ending claim is not substantiated by the modeling scenarios that have been run: "targeted interventions for demographic groups such as under-35 men may be the key to finally ending HIV." It is straightforward to run this specific scenario in the model to determine whether or not this is true.

      Our modelling framework is not set up to model the “last mile” of HIV elimination, notably as it has no component for MSM or FSW transmission, and we do not feel that we could confidently present results regarding it. As a result, this statement has been greatly softened in the new abstract (lines 75-78).

      The authors should add confidence intervals to the quantitative metrics, such as the 93.8% and 62.1% incidence reduction.

      These have been added.

      (2) RESULTS

      The authors should check the Results section for any qualitative claims not substantiated by the analyses performed, and ensure the corresponding analyses are presented to support the claims.

      The Results and Methods describe the model's implementation of the PopART intervention differently. The Methods describes it as including VMMC, TB, and STI services, while the Results only mentions intensified HIV testing and linkage.

      This is a slight misreading of the text. That paragraph in the Methods is describing the trial itself, not the modelling framework.

      A limitation of the model is that HIV disease progression is based on the ATHENA cohort in the Netherlands, which is a different HIV subtype (B) than the one in the research setting (C). The model should be configured using subtype C progression data, which have been published, or at least a sensitivity analysis should be conducted with respect to disease progression assumptions.

      The available literature does not suggest a significant difference in progression between subtypes B and C, and we have added text and citations to this effect (lines 699-701).

      In Table 2, the authors should consider adding a p-value to establish whether or not IBM and phylogenetics estimates are different.

      We have done this; the appropriate test was a posterior predictive check. See lines 261-263, 575-579 and 805-814.

      (3) DISCUSSION

      The literature review and comparison of study results to previously published phylogenetic studies is very nice. The authors could strengthen this by providing quantitative estimates with CIs for a more scientific comparison of the study results vs. prior studies, perhaps as a table or figure.

      We have expanded the discussion on this point (lines 504-527). We considered adding a table, but the existing literature that directly answers the questions we ask is quite limited and fragmentary. For example, Monod et al do not present a complete treatment of age gaps. The literature using regression analyses to identify predictors of HIV prevalence or incidence related to partner age is extensive, but those results are not directly comparable to ours.

      The authors state that due to "the narrow geographical catchment area... The results should not be automatically extrapolated to apply to other SSA settings." The authors should exercise this caution when comparing the results to studies in South Africa and elsewhere.

      We have made more explicit acknowledgements of these limitations (lines 598-600).

      There are many other limitations to the analysis, including some mentioned above, that are not acknowledged. The authors should think carefully about what the most important limitations are and acknowledge them honestly at the end of the Discussion section.

      The limitations paragraph has been revised (lines 598-605).

      Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex-specific heterosexual HIV transmission dynamics in Zambia, with the goal of allocating resources.

      Strengths:

      Important analysis to hone in on the key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches were used, and while the phylogenetic data were markedly more limited, they mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses. The authors did a nice job of providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce

      Weaknesses:

      To increase the impact and utility of this work, it would be helpful to parse the analysis just a bit further to estimate the roles of undiagnosed vs diagnosed and untreated subpopulations on this transmission. PopART is a multifaceted intervention, but the cost, effort, and approach to reengagement in care vs testing/treatment can be quite different.

      We have now provided stratified results by diagnosed and non-diagnosed status of the source, as well as an overall summary of the proportion of undiagnosed sources by age and sex. See lines 305-310, 539-547, and table 3.

      Recommendations for the authors:

      Reviewing Editor:

      We commend you for conducting a rigorous and comprehensive study titled "The age and sex dynamics of heterosexual HIV transmission in Zambia: an HPTN 071 (PopART) phylogenetic and modelling study" that significantly advances the understanding of HIV transmission dynamics in sub-Saharan Africa. The study utilizes an innovative dual-methodology approach integrating individual-based mathematical modelling (IBM) and pathogen phylogenetics to characterize heterosexual HIV transmission patterns by age and sex during the PopART trial in Zambia.

      This manuscript reports on HIV transmission dynamics in Zambia using data from the PopART study, combining individual-based modelling and phylogenetic analysis. The use of two independent methodologies enhances confidence in the consistency of the findings and enables robust cross-validation. The work addresses an important topic in HIV prevention, particularly in settings where resources may become more constrained, and offers insight into potential demographic targets for intervention.

      However, several aspects of the manuscript limit its current impact. The main take-home messages are diffuse and not clearly presented. Some conclusions in the abstract and discussion appear to go beyond the scope of the presented data. For instance, the claim that targeting under-35 men may be key to ending HIV is not directly tested in the modelling scenarios and should be reframed or removed unless supported by new analyses. Furthermore, important quantitative details, such as confidence intervals, p-values, and precise age group estimates, are lacking in key sections (e.g., the Abstract and Results).

      The authors are encouraged to clearly identify and communicate their central findings, ensure all claims are fully supported by their analyses, and make the data more accessible to readers by adding detailed, quantitative summaries where needed.

      The following are our recommendations to the Authors:

      (1) Clarify Study Objectives and Central Messages

      Reframe the abstract and discussion to highlight a clear, well-supported set of main findings.

      Avoid overgeneralized or unsubstantiated claims, especially those not directly tested by your model (e.g., the effectiveness of targeting under-35 men).

      As stated above, we have revised this text accordingly.

      (2) Support Qualitative Claims with Quantitative Data

      Provide numerical results, including effect sizes and confidence intervals, wherever qualitative trends are mentioned.

      For example, restate: "The largest gaps for female recipients were among the youngest" as "... in the age group XX-YY with OR = Z.Z (95% CI: A.A-B. B)."

      As mentioned at the top of the review, we have overhauled the treatment of summary statistics extensively, and now give confidence or highest density intervals throughout the text.

      (3) Improve the Results Section

      Check that all claims are supported by the analyses, and ensure figure references are accurate.

      The statements that went beyond what was supported, notably about ending the epidemic by targeting young men, have been removed. The typo in table references has been fixed.

      Annotate Figure 6 with trendline coefficients and p-values where applicable.

      The takeaway message of figure 6 has now changed and we no longer see no trend, just a minor one.

      Revise Figure 4 for clarity or consider replacing it with a tabular format.

      We would prefer to keep the current figure 4, as we have not found any clearer way to illustrate the patterns, which are the consequence of the phenomenon observed in figure 5. We have put more explicit descriptive text in the discussion, linking the two figures (lines 470-476).

      (4) Address Potential Bias and Model Assumptions More Rigorously

      Explain sampling bias in IBM and phylogenetics (e.g., how the 355 high-confidence phylogenetic pairs were selected).

      The reviewer comment regarding the 355 pairs was based on a misapprehension; we used all the pairs we found using the phyloscanner pipeline. There are no sampling bias issues involved in the IBM as every individual in the simulations is considered. Appendix 2 includes some sensitivity analysis results if the procedure used to find the 355 is changed.

      Discuss how the use of subtype B disease progression data from the ATHENA cohort may impact results in a subtype C setting. A sensitivity analysis would strengthen this.

      Subtype B progression data was used in the absence of any appropriate data from subtype C, but the literature does not suggest any major difference between the two (lines 699-701).

      (5) Include More Detail on Undiagnosed Populations and ART Effects

      Estimate the roles of undiagnosed and untreated subpopulations in driving transmission.

      As mentioned above, this analysis has been added.

      Clarify mechanistically how ART might influence age gaps in transmission dynamics.

      This now is clarified in the introduction (lines 127-129).

      (6) General Improvements

      Provide p-values where comparisons are made (e.g., in Table 2).

      Use consistent terminology and definitions across Methods and Results.

      Add more discussion on limitations, especially regarding generalizability to other SSA settings.

      All of these have been inserted as previously mentioned.

      By addressing these points, the manuscript would present a more coherent narrative and a stronger, evidence-based contribution to the field. We appreciate you all for your fantastic effort and hope you will reflect the feedback in your final paper.

      Reviewer #1 (Recommendations for the authors):

      Thank you for the opportunity to review this interesting manuscript.

      In the public review, I have recommended that the authors should incorporate quantitative results for all of the qualitative findings statements. As one example, I would recommend that "We found the largest gaps for female recipients were among the youngest of those recipients" is re-written as "The largest gaps for female recipients were in the age group XXX-YYY with OR=ZZZ (XXX-YYY)." such as odds ratios, and specific outcome definitions including ages. To give one more example: "immediate increase in the average age at transmission of both sources and recipients" could be rephrased as "increase in the average age at transmission by XXX (YYY-ZZZ) years for sources and XXX (YYY-ZZZ) for recipients over [TIME PERIOD]."

      We hope the revisions we have made to the statistical presentation are satisfactory as a response to this request.

      Again in the public review, I recommended checking the Results section for any qualitative claims not substantiated by the analyses performed, and ensuring the corresponding analyses are presented to support the claims. An example is: "Trends are minor or non-existent in the former two variables." - please annotate Figure 6 (assuming the authors meant to reference Figure 6 and not 7 here?) to show over what period trendlines were fit and provide the coefficient and CI. To support the stated claim even more strongly, a p-value might be apt with a null hypothesis of a slope of zero.

      Please check the numbering on all figure references in the text, as some appear to be misnumbered. E.g., where the text refers to Figure 7, I believe the authors meant to reference Figure 6.

      The change to how we handled the statistics has changed the message of figure 6 (which is now figure 7) and rendered this somewhat moot. We have checked that all figure and table references are now correct.

      Figure 3 is very nice, but if the axes were flipped on one panel, it would make them easier to compare, and then adding some statistics to assess whether the patterns are the same or different when a man vs woman is the source.

      We have flipped the axes here.

      Figure 4 was too complicated for me. I could not follow the Sankey flows because there is too much going on and overlapping. Consider revising to make it easier to digest... perhaps to table format?

      As mentioned above, we would prefer to keep this figure, but we have situated it better in the text.

      Reviewer #2 (Recommendations for the authors):

      A few points that would improve the clarity and the strength of the manuscript

      (1) There is a need to clarify more about how the IBM and phylogenetic data does not suffer from sampling bias. For e.g.,

      Line 205: What proportion of the transmissions modeled in the IBM from Zambia?

      All of them. We confined the analysis of the IBM to the Zambian communities from which phylogenetic data was acquired (lines 755-758).

      Line 217: What proportion of the phylogenetic pairs (cherries) suggesting transmission were the 355 that had high confidence in directionality. How do these pairs compare to the others

      There was no identification of “cherries” involved in picking these pairs; the phyloscanner procedure does not use that step. We confined our analysis solely to the pairs for which we did identify a direction of transmission; that is the 355. Appendix 2 includes a sensitivity analysis involving varying the parameters by which these were identified.

      (2) I appreciate the authors noting that MSM transmissions are unlikely to be playing a role in this cohort, as noted in previous work by the group. However, systematic undersampling of men is common in other study cohorts of HIV. While the MSM and heterosexual networks may be relatively distinct, undersampled men who are bridging the networks could impact the estimates. Can the authors use the time to diagnosis analysis (HIV phyloTSI) to estimate rates of undiagnosed men and women?

      We feel that this is beyond the scope of this work. The phylogenetics dataset in its totality could be used for this purpose (although it is probably highly biased towards undiagnosed individuals due to the considerable majority of samples coming from the healthcare facilities). However, we concentrate here solely on the subset involved in our probable transmission pairs, which is fairly small. Extending the scope to an exploration of the full dataset would seem like a separate study, which we do have plans to do.

      We have used the IBM for this question instead (lines 303-321), however, as MSM transmission was not modelled, it is also not ideal for answering this question. Ultimately we feel that the way these studies were implemented makes it an unsatisfactory tool for answering the MSM question, important as it is.

      (3) Expanding on the point above, in other settings, transmission to young men has been associated with partnerships with older men, and if these young men then transmitted to young women, would we see a similar effect as noted in these models (assuming the young men were less well sampled).

      Our previous work (Hall et al., 2024) suggested no excess of identified male-male pairs in the phylogenetics dataset which might suggest cryptic male-to-male transmission. The age disparities would be worth exploring had this been found, but is curtailed by the lack of it.

      (4) Related to the point above, is there an estimate of the populations (age and sex) that are undiagnosed in the IBM model? Can this be teased out... is transmission from men to women more likely 2/2 lack of diagnosis... or lack of engagement in care?

      We have explored results by diagnostic status as it pertains to age and sex, but we feel that moving on to a more general exploration of the role of diagnosis and lack of engagement in care is again going beyond the scope of what is already a long paper.

      (5) I'm still not fully clear as to why ART might affect age gaps. Can this be explained in more detail?

      See lines 127-129.

    1. AbstractRegeneration relies on precise spatiotemporal gene expression and cellular responses to establish tissue identity and body patterning. Using high-resolution Stereo-seq (715 nm) on 353 sections from 16 whole animals at 8 regeneration timepoints, we constructed a 4D spatiotemporal transcriptomic map of planarian regeneration. Our analysis captured 36 refined cell types from 3,508,004 segmented cells, enabling genome-wide transcriptional imputation of gene expression dynamics across body axes at cellular, tissue, and organismal scales. We identified dynamic positional gradients and distinct spatially distributed cell types during regeneration, including an injury-induced Anterior Regenerative Zone (ARZ). The ARZ exhibited enriched positional signals in epidermal, muscle, and neural cells and was regulated by Mediator 8, which is crucial for polarity remodeling and blastema formation. This study provides a comprehensive spatial molecular and cellular map of regenerative processes, highlighting injury-induced spatial domains and key regulatory factors in planarian regeneration. We also provide an interactive web portal, offering a valuable resource for exploring and analyzing regeneration mechanisms in a spatiotemporal context.

      This work has been peer reviewed in GigaScience (see https://doi.org/10.1093/gigascience/giag064), which carries out single-anonymized peer review. These reviews are published under a CC-BY 4.0 license and were as follows:

      Reviewer 2:

      In the manuscript '4D single-cell spatial transcriptomics reveals dynamic morphogenetic gradients and regenerative domains in planarians,' Han and colleagues generate a truly stunning spatial transcriptomics dataset of planarian regeneration from the species Schmidtea mediterranea. The authors' dataset includes whole 3D reconstructions of two regenerating planarian fragments at 8 different timepoints during regeneration, a fantastic accomplishment and resource of broad interest to the regenerative biology community. The authors analysis of the dataset includes characterization of spatially biased genes (SBGs) and exploration of an anterior regenerative zone (ARZ) and the role of the gene med8 in its' regulation. While the authors' dataset is remarkable and their analysis of spatially biased genes and med8 function is interesting, I'm not yet convinced that their conclusions are fully tested by the included experiments. In addition, I think that the authors have not included sufficient quality control metrics for their spatial dataset, which makes determining the limitations or caveats of their analysis and conclusions more difficult. However, my concerns could be addressed by additional analysis and minor experiments, or by softening the conclusions of the authors to include alternative models. I've detailed the areas of analysis/discussion that I believe require improvement below:

      Major Criticisms: 1. Stereo-seq resolution and capture efficiency: The authors assert that their spatial approach is high enough resolution to resolve cell types and they claim to have characterized 36 cell types in their abstract. However, the 'cell type' in their dataset that they choose to focus on - Clu.31 - has gene markers expressed in three different cell types that have been shown to be distinct in the literature and prior planarian atlases. The authors should analyze gene expression signatures of other stereo-seq 'cell types' to determine if they also show mixed expression signatures. In addition, I am curious if stereo-seq is more likely to capture highly expressed genes (like those expressed in parenchymal cell types) than more lowly expressed genes (like the transcription factors expressed in stem cells). If it exists, this bias could influence annotation of cell types in highly heterogeneous regions of the worm like the parenchyma or parapharyngeal region. Finally, there is very little QC data in the supplementary materials (Size/volume of segmented cells, UMIs and features per cell, variability in features/UMIs per section, per replicate, and per cell type, etc.) I think this analysis would be highly valuable for the reader to interpret the data and the 36 identified 'cell types'.

      1. Dynamics of spatially biased genes: The authors analysis on the dynamics of spatially biased genes (SBGs) is very interesting, but the 'oscillations' the authors referred to were not clear to me in the data across all or even most of the pattern clusters in Figure 2A. In general, it seemed more like the pattern cluster was 'noisy' or more broad before stabilizing to its final location. In addition, the PCA analysis in Figure 2B seems to show that Intact and 14dpa transcriptomics is very similar, but 0h, 12h, and 36h timepoints are very distinct from 3, 5, 7, and 10 day fragments. This would suggest that early wound response gene expression is highly distinct (even opposing) the gene expression programs active during late in regeneration. More exploration of this idea, as well as clarified language on exactly what the author means by 'oscillations' and which gene groups follow this pattern would greatly improve this section and better support the author's conclusions.

      2. The Cellular/Functional identity of Clu.31: The authors state throughout the manuscript that Clu.31 (the ARZ) is an injury-induced anterior state enriched for SBGs and regulating polarity establishment. However, it is also possible that this spatial state represents the anterior peripheral nervous system (numerous sensory neurons and surface epithelial cells that help sense mechanical and chemical cues). SBGs could be enriched because this combination of cell types is only present in the anterior of the animal. Indeed, the authors show that the ARZ is localized to the anterior in intact animals in the absence of an injury (Figure 3) and enriched genes (S4Aii) strongly indicate that Clu.31 contains gabrg+ mechanosensory neurons. If Clu.31 is regenerating nervous system, this would also explain its ventral bias and expression of tgs-1 and other nb2 genes, since nb2 neoblasts have been suggested to be both an amputation responsive neoblast subset (Zeng et al. Cell) and a neural progenitor state (Raz et al. Cell Stem Cell). Clarifying how the composition of the tri-lineage region changes during regeneration may help distinguish if Clu.31 is truly an injury induced region vs. the regenerating sensory nervous system. For example, it is known that agat-1+ cells transcriptionally responsive and enriched at the wound site a 2-4 days post amputation, but less so at later timepoints (Benham-Pyle et al Nature Cell Biology, Kent et al. Developmental Biology). This shift in composition should be observable in Clu.31 since it contains agat+ epidermal cells. Such a shift in composition or the identification of a regeneration-specific marker expressed in Clu.31 would add support to the author's conclusions. Regardless of the outcome of these experiments/analyses, the discussion and interpretation of the data could be modified to address the hypothesis that Clu.31 represents the cellular neighborhood created when the peripheral nervous system intercalates with the anterior DV boundary epithelium and body wall muscle, which needs to be regenerated in amputated worms. As is, the comparison to the apical epithelial cap considered in the discussion (Line 438) may be pre-mature.

      3. Med8 function: Med8 produces a clear phenotype in the authors' experiments, and their data indicates that it is required for ARZ formation. However, I am not sure that the authors data supports the claim that Med8 is directly regulating blastema and PCG expression, as opposed to regeneration of the nervous system (which is highly interconnected with formation of the anterior pole and the size of the anterior blastema) and stem cell function more broadly. The fact that Med8 RNAi also leads to head degeneration in intact worms (Figure S6F) strongly suggests a more fundamental defect in neural differentiation or stem cell function. The strongest evidence presented by the authors supporting a broader function in polarity establishment is the disruption of posterior Wnt expression, (Figure 5F and G), but these in situs are single representative images with no quantitation and could also be explained by a stem cell defect. Additional data could be provided (e.g. visualization of wound-induced gene expression, quantitation of anterior or posterior stem cell numbers and proliferation rates at 2dpa) to support regulation of PCGs or blastema formation. The authors could also leverage their single cell sequencing to determine if Med8 RNAi impacts neural progenitor abundance more than other progenitor cell types. Together, these experiments would determine if Med8 is important for amputation-induced blastema formation and polarity re-establishment vs. stem cell function and neural differentiation more broadly.

      Minor Criticism/Feedback: 1. In Figure 1I, the authors show DEGs enriched in each cluster/region. In the blastema regions, I was surprised by the number of DEGs for each time point. It appears that there are ~10K upregulated and 10K downregulated DEGs by the later time points, which suggests that 2/3 of the transcriptome is differentially expressed… The authors should clarify in the text or methods what cutoff they used for the DEGs and how significant the DEGs are in this figure. 2. For readability, I really think that all figures should be on a white background. 3. How do gene expression profiles from the stereo-seq compare to bulk rnaseq at similar timepoints? 4. It is very interesting that there are some cell types that appear to contract and then expand during regeneration (Cluster 0, 23) or that aggregate/become more targeted during regeneration (pharynx pouch, cluster 29). Molecular differences between early and late cells within these cell types would be particularly interesting for understanding different phases of regeneration, but this may be beyond the scope of the current study. 5. The authors frequently reference Han et al. submitted, but this manuscript would need to be pre-printed or published in order for this work to reference it. 6. The Y axis of Figure 2E should be labeled

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths:

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (4) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

      We will add two recent in-press Cortex papers to the Discussion. One provides lesion-based double-dissociation evidence against V1 as a necessary causal substrate of visual imagery. The other shows that aphantasic individuals can display visualizer-like oculomotor patterns during mental map exploration despite reporting little or no imagery vividness. Together, these studies help clarify our interpretation of our null V1 findings and structural effects in higher-order brain regions, which are consistent with aphantasia involving altered integration or access rather than a primary V1-dependent imagery deficit.

      Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

      We thank the reviewer for this thoughtful and constructive comment. Lack of introspective report of voluntary imagery is arguably the defining signature of aphantasia. This motivated us to primarily interpret our anatomical findings in a broader cognitive context of higher-order control, internal monitoring, and conscious access in aphantasia. We expect that a reliable behavioural test measuring imagery sensitivity and accessibility would allow us to direct link these findings to individual imagery ability. Nevertheless, to our best knowledge, this kind of test on imagery is still missing. Instead, our findings point to some plausible structural signature or brain regions that may be related to conscious imagery, which motivate future studies to examine their direct or causal roles. We agree with the reviewer, future studies should test the relationship between these anatomical structures and the accessibility to internal representation, together with related aspects of internal monitoring. We will therefore add a dedicated paragraph to discuss the plausible cognitive mechanisms during the revision.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

      We are grateful to the Reviewer for the constructive and thoughtful assessment of our manuscript. In response to the reviewer’s comments, we will revise the manuscript to clarify the dependency between diffusion-derived analysis streams, to state more explicitly the biological limits of diffusion MRI metrics, to add a null-network sensitivity analysis for the clustering coefficient findings, to include confidence intervals for reported effect sizes, and to temper the interpretation of the cortical thickness result. We will also revise the Abstract and Discussion to better reflect the exploratory nature of the study and to frame the findings as hypotheses requiring replication in larger independent samples. We believe that these revisions will make the manuscript more balanced, transparent, and appropriately cautious, while preserving the central conclusion that congenital aphantasia is associated with structural differences centered on higher-order frontotemporal and cingulate systems.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

      In particular, the "saccadic execution" parameter appears far too long and too variable to reflect merely execution; instead, it likely includes decisional components. This would make more sense since manual and saccadic planning essentially rely on distinct brain areas, hence it seems unrealistic that crossing a single threshold would trigger both manual and saccadic execution. Similarly, recovered manual non-decision times are substantially longer (though not more variable) than expected motor execution durations for button presses. These patterns suggest that parts of what the model treats as non-decision time are likely decisional in nature, although perhaps related to "action decision" rather than the "value-based decision" of interest to the authors. To what extent these two processes neatly follow each other or overlap could be usefully considered.

      We have added a paragraph to the Discussion explaining how our model’s estimates of sensory and motor latencies relate to corresponding values inferred from physiology or behavioral manipulations (e.g., Bompas et al., 2024). Specifically, we write:

      “The key assumption of the PDG model is that there is a delay between the moment a choice is internally committed and the moment it is externally reported with a key press. Because eye movements are typically faster than manual responses (𝜏<sub>e</sub> < 𝜏<sub>m</sub> in our simulations), this delay creates a window during which gaze can already be directed toward the covertly chosen item before the response is formally registered. We do not interpret these non-decision latencies as irreducible physiological minima for moving the eyes or pressing a button (Bompas et al., 2025). Rather, they are inferred indirectly by fitting an additive non-decision-time parameter to the behavioral data, which we decompose into a sensory delay (𝜏<sub>s</sub>) and a manual execution delay (𝜏<sub>m</sub>). Values of 𝜏<sub>e</sub> are then chosen so that the model reproduces the observed magnitude of the behavioral effects. This estimation procedure has important limitations. Some participants show relatively “flat” chronometric functions: response times vary little with value despite otherwise normal psychometric performance. Such patterns likely reflect processes not explicitly represented in the model, including procrastination, reduced motivation, task-unrelated thought, or noise in item ratings. Within a drift-diffusion framework, however, these cases are accommodated by assigning a long non-decision time together with a short evidence-accumulation period (Table S1). Consequently, some estimated non-decision times are substantially longer than would be expected if they represented only sensory and motor delays. A further limitation is conceptual. We model non-decision time as occurring either before or after evidence accumulation, whereas in reality decisional and non-decisional components are likely temporally interleaved (Graziano et al., 2011). This simplification may also inflate the recovered latency estimates. With these caveats in mind, sensory and oculomotor delays on the order of 300 ms remain broadly plausible, although they likely lie near the upper end of a realistic range. The estimated eye-movement latency is especially long. For instance, in monkeys trained to report simple perceptual decisions with a saccade, roughly 100 ms elapses between the threshold-crossing signal in parietal cortex (or the superior colliculus) and the executed eye movement (Roitman and Shadlen, 2002; Stine et al., 2023). Crucially, however, varying the assumed non-decision latencies across a reasonable range does not alter the qualitative predictions of the model (Fig. 8).”

      Further, we have added a parameter sensitivity analysis. Importantly, although the magnitude of the predicted effects depend on the non-decision latencies, the qualitative aspect of these predictions do not (new Figure 8). Specifically, (i) the increasing tendency to look at the ultimately chosen item as time elapses (new Fig. 8A), (ii) the lack of an interaction between the last-fixation bias and overall value (Fig. 8B), and (iii) the absence of an effect of choice consistency on Δdwell (Fig. 8C) are all findings that are independent of 𝜏<sub>e</sub>.

      Reviewer #2 (Public review):

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Following this suggestion (and the reviewer’s elaboration in the private comments to the authors), we have substantially restructured the manuscript. Both aDDM predictions are now presented together (new Fig. 2), and Figs. 3–4 test these predictions across multiple food-choice datasets. In doing so, we no longer treat the data from Krajbich et al. (2010) separately, and we extend the analysis of the last-fixation–choice association (MELFB) to additional datasets. We note that the same datasets could not be used in both Figs. 3 and 4, as some lack information on the final fixation required for the MELFB analysis. Nevertheless, results are highly consistent across datasets and align with findings from a recent study by Ting & Gluth (2025), which independently identified and examined one of our key predictions; this work is now cited in the revised manuscript. Finally, to reduce redundancy, we have consolidated all aDDM variants and optimal models into a single figure (new Fig. 10).

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

      The key predictions of the model depend on the difference between the manual (𝜏<sub>m</sub>) and eye-movement-related (𝜏<sub>e</sub>) latencies. We have now added a parameter-sensitivity analysis to show how the model predictions depend on this difference. The new analysis shows that while the quantitative predictions do depend on the precise latency values, the results are qualitatively similar across values of 𝜏<sub>e</sub> (new Figure 8).

      Reviewer #3 (Public review):

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

      Thank you for this suggestion. We added a new paragraph to the discussion (paragraph #2), where we argue that it is sensible for a decision maker to direct the gaze to the chosen item once a covert choice commitment has been made, as the benefits of attending to a stimulus do not end with the decision itself. Specifically we now write:

      “Instead, these observations are better explained by a post-decision account of the gaze-choice association that is, one in which gaze shifts to the selected item after a covert commitment to a choice. We argue that directing gaze to the chosen item after a covert choice commitment is sensible, as the benefits of attending to a stimulus do not end with the decision itself. In naturalistic settings, for instance, selecting a food item is typically followed by the action of reaching toward it, where visual attention supports spatial localization and motor planning for the upcoming action. Although participants in our computerized task did not physically act on their choices, these sensorimotor processes are likely highly automatized and may still be engaged by default, even when not strictly required. Beyond motor preparation, post-decisional attention may also serve additional functions, such as facilitating sensory anticipation of the reward, supporting metacognitive evaluation of the decision, and contributing to value updating for future choices. From this perspective, a degree of attentional “stickiness” whereby the chosen item remains preferentially attended after commitment could emerge as an effectively optimal policy once these post-decisional processes are taken into account. Moreover, a specific feature of the task design may further reinforce this tendency: in the snacks paradigm, the unchosen item typically disappears from the screen immediately after a response is registered. It is therefore plausible that directing gaze to the chosen item after commitment partly reflects anticipation of the imminent disappearance of the unchosen option. To disentangle these mechanisms, it would be interesting for future work to test whether this attentional bias persists when the chosen item, rather than the unchosen one, is the stimulus that disappears upon response.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments:

      (1) Framing of the modelling approach

      The manuscript would benefit from acknowledging the known limitations of DDM-based frameworks, especially given that the entire study is conducted within these constraints. The introduction highlights successes of the DDM, but the manuscript does not mention any of its conceptual or empirical limitations.

      We are unsure about what specific limitations the reviewer has in mind, but we have added a paragraph to discussion mentioning some limitations, like the inflation of the non-decision times and the difficulty of interpreting the fit parameters (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (2) Dependence on non-decision time assumptions

      The alternative model's explanatory power appears to rely heavily on assumptions regarding the decomposition of non-decision time: fixed visual encoding (𝜏<sub>s</sub>= 0.3 s), manual non-decision time (𝜏<sub>m</sub>; two free parameters), and saccadic execution (𝜏<sub>e</sub>; fixed parameters μ<sub>e</sub> = 0.35, σ<sub>e</sub> = 0.11).

      - 𝜏<sub>e</sub> is substantially longer and more variable than typical saccadic execution times, suggesting it likely incorporates decisional components.

      - Estimated 𝜏<sub>m</sub> values are approximately twice as long as known manual execution durations.

      - σnd is more plausible, implying that variability is captured correctly but mean durations are not.

      Together, these points raise the possibility that portions of what the model treats as non-decision time are in fact part of a (action) decision process. Only then does it make sense to assume that Tm is usually larger than Te. If Tm and Te were truly execution delays, then Tm would always be larger than Te.

      You may find it helpful to consider the framework in Bompas et al. Psych Review (2024), which discusses in detail what non-decision time is likely to comprise across effectors.

      Thank you we have added (i) a sensitivity analysis showing that our results are robust to changes in the specific value used for the eye movement related latencies (new Fig. 8), and (ii) a new paragraph in Discussion addressing the issue of the mismatch between our parameter estimates and the manual and saccadic execution times (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (3) Code availability.

      The authors should consider sharing all relevant code and data publicly.

      We agree, we now share the code and data on GitHub and indicate so in the revised manuscript.

      Minor Comments:

      (1) Lines 74-77. These are not worded as predictions but as questions; one tests predictions, but answers questions. I feel it would be clearer to stick to predictions (like in the abstract), and the introduction could benefit from explaining these predictions in a bit more detail (I found it difficult to get my head around these predictions from the intro text only).

      We rewrote the section in the introduction where we provide a gist of the model predictions (last paragraph of Introduction). We agree with the reviewer that the previous explanation was not clear.

      (2) It is confusing that panel B appears to the left of panel A in Figure 2.

      We agree. We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed and the panels follow a more logical order.

      (3) Figure 3C - remove MATLAB toggles.

      Yes, thanks.

      (4) Figure 5A shows the proportion of left choices, but the text and legend refer to right choices.

      Good catch, thank you.

      Reviewer #2 (Recommendations for the authors):

      This may appear self-serving, but the authors seem to be unaware of some highly relevant work from our group. Most importantly, in a recent publication (Ting & Gluth, 2024, JEP General), we have already looked at the dependency of the last- (or final-) fixation bias on overall value in value-based (VB) and perceptual (P) decisions. In VB, we found a negative effect; in P we did not find a significant effect. This is largely consistent with the current results, showing a negative but not significant trend. Another relevant work is Gluth et al. (2020, Nat Hum Behav), where we extended the aDDM by assuming that the probability to fixate on an option is a function of the accumulated evidence for that option. It would be interesting to know whether this assumption changes the predictions of the aDDM. Finally, we just published a new theory on how people search for information to make efficient value-based decisions (Gluth et al., in press, Psychol Rev; https://osf.io/preprints/psyarxiv/3qzak_v2). Although this theory focuses on multi-attribute choices, it can be applied to "simple" choices, too (by assuming that there is only one attribute = value). Interestingly, while the model also mispredicts a (slight) increase of the last-fixation bias with overall value, it correctly predicts the independency of the dwell-time advantage effect on choice consistency as well as the small increase of the effect with RT (attached here is a figure to show this: [https://elife-rp.msubmit.net/elife-rp_files/2026/01/22/00149589/00/149589_0_attach_9_477122. pdf], and the match with the empirical data shown in Figure 3B and 12 is striking). In general, the model shares many features of the Callaway and Jang models, but does not need to assume a biased value prior, which the authors suggest is responsible for the misprediction of the second effect. I leave it up to the authors to discuss this new theory, but I wanted to point this out.

      Thank you for pointing this out; these are all relevant points and studies.

      We now note that the first of our predictions has recently been identified and tested by Ting and Gluth (2025).

      We also considered extending the manuscript with a variant of the model proposed by Gluth et al. (Psychological Review, 2026). In fact, we attempted to fit this model to the Krajbich et al. (2010) dataset under the assumption that the duration of each sampling epoch is a free parameter. We find this model very interesting. However, in our current implementation it appears to make the same qualitative prediction as the aDDM, namely that ΔDwell depends on choice consistency (see Author response image 1).

      Given this, we have decided not to include these results in the manuscript. It remains possible that with further development particularly with a more realistic specification of fixation durations (e.g., allowing them to depend on value) the model could account for the full set of observed effects. We think this would be best addressed in a separate study.

      That said, we do find the model promising, as it provides a better account than most of the alternative models we explored for the patterns shown in panels D, H, and I.

      Author response image 1.

      Fits of a variant of the MACS model (Gluth et al. 2026) to the data of Krajbich et al. (2010).

      The paper would benefit substantially from restructuring. The aDDM's predictions are provided first, together with the empirical data, and then the optimal models are discussed. But Figure 2 shows all of this together. Later, the new (PDG) model is elaborated, and its predictions are shown. Towards the end of the results, variations of the aDDM and combinations of aDDM and PDG are shown in a series of figures (8-11), followed by a last figure showing one of the tested effects in other datasets. All of this feels pretty much thrown together without a clear structure. For instance, the aDDM and the optimal models could be described together (or the optimal models get a separate figure). The additive variants could be described earlier. And some figures could be put into the supplement. And the empirical results of the different studies could be shown together.

      We fully agree with this suggestion. We have now restructured the manuscript along the lines proposed by the reviewer (see the more detailed explanation of the restructuring in our response to the public comments).

      I strongly suggest avoiding the term "influence" in the y-axis of Figure 2, upper row, as it implies causality. Similarly, in line 182, the term "causal influence" is used in the context of the Callaway model, but as far as I know, this is not what the model assumes.

      We replaced the y-axis label with “Association of last dwell with choice (β)”

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 2 - Panel labels for A and B are reversed?

      We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed.

      (2) Does 3C include a .pdf screenshot?

      Thank you, it’s a Matlab bug on Mac. I guess they want us to switch to Python -:)

      (3) Figure 4 - It would be helpful if the green line were defined in the figure legend.

      Added

      (4) The effect size in 5B looks much more dramatic than in 2B(A?) - Is this for one example subject as opposed to all subjects? Please clarify what is different about the data.

      We are no longer showing the psychometric functions in Figure 2.

      (5) Line 252 - they say they compared the probability of choosing the right item (Fig. 5B) by the y-labels of that figure, which are all p(choose left).

      Yes, corrected now.

      (6) In general, they reference the subpanels of Figure 5 out of order, which causes the reader to jump around. They might consider reordering the panels of the figure so they follow the ordering of descriptions in the text.

      We agree, we have rearranged the figure panels to follow the ordering of the descriptions in the text.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) Alternative mechanisms for performance differences.

      The authors assume that the difference in performance between the low-switch (LS) and high-switch (HS) frequency conditions is explained by a change in the "leakiness" of integration. However, several other mechanisms could potentially explain this effect:

      (1) Temporal Uncertainty: Integration might start later in the HS condition, leading to lower performance.

      (2) Reduced Efficiency: Integration could be less efficient in the HS condition (i.e., lower signal-to-noise ratio) without a change in the leak parameter itself.

      (3) Evidence Contamination: Motion information from the adapting stimulus in the HS condition may be integrated rather than ignored, which might be the case since the transition from the adapting to the test stimulus is not externally cued.

      To distinguish between these alternatives, I suggest two possible analyses. First, a formal model comparison could be performed, though I acknowledge this may be inconclusive in the absence of response-time data. Second, an analysis of motion energy kernels could be revealing; the leak hypothesis makes the specific prediction that for long test stimuli, early samples should contribute more to the choice in the LS condition than in the HS condition, relative to late samples.

      We thank the reviewer for raising these important points. We agree that we cannot definitively identify the algorithmic underpinnings of the behavioral effects we report and have made substantial revisions to the manuscript to be clearer about what is supported and what is speculative in our claims. Most importantly, we agree that we do not know if the context-dependent differences in how accuracy depends on viewing time are based on adjustments to a leak or to something else (e.g., a saturating non-linearity, as we identified in Glaze et al, 2015, that is separate from the leak itself), which we cannot resolve with this dataset, even with more formal model comparisons. We therefore:

      Changed the wording throughout the manuscript to refer to changes in leakiness as just one of several possible sources of the behavioral differences. We also added this point to the list of “limitations” (and possible future directions, including using motion-energy kernels, which would require us to use lower-coherence test stimuli) in the Discussion (L487-493).

      Added a new figure panel (Fig. 2D), a new Extended Data figure (Extended Data Fig. 3), and additional explanatory text (L168-175) that collectively describe the behavior in more detail, including quantifying a “crossover” dynamic similar to what we reported previously (Glaze et al, 2015).

      Added new explanations (L152-163) and analyses (Extended Data Fig. 9) indicating that the monkeys used some information from the end of the adapting stimulus to inform their decisions, which accounts for the patterns of choices at the shortest viewing durations.

      Indicate that the context-dependent differences in the slopes of the psychometric functions (and complementary analyses based on “raw” accuracy measures as a function of binned viewing duration) rule out the temporal uncertainty and evidence contamination explanations, but are consistent with effects on the temporal dynamics of the decision process (L175-179).

      (2) Independence of neural and pupil-linked signals.

      The authors take the lack of session-wise correlation between context-dependent contributions from neural and pupil terms as evidence that these two signals provide independent contributions to the behavioral effect. However, could this lack of correlation simply be a result of high variability or noise in these estimates? The data shown in Figure 7B suggests that measurements are very noisy, which might obscure a potential relationship.

      We agree that the lack of session-wise correlation between neural and pupil terms cannot be taken as definitive evidence of independence. We have both softened the language around the claim (L368) and added a sentence to the Discussion (L464-468) acknowledging that this lack of correlation may reflect underlying noise and/or variability rather than true independence of the underlying mechanisms.

      Reviewer #1 (Recommendations for the authors):

      (3) The neural data analyses rely fundamentally on "switch" trials (Figures 3-5). It might be informative to also examine "non-switch" trials to see if there are specific neural markers indicating the exact moment the motion stimulus becomes behaviorally relevant. Given that this may fall outside the primary focus of the paper, it is up to the authors whether to pursue this line of inquiry.

      We thank the reviewer for this suggestion. We agree and have added new analyses of data from non-switch trials (Extended Data Fig. 9), which show some effects of stimulus information from the adapting epoch on the monkeys’ choices, as we detail below in response to related comments from the other reviewers.

      Reviewer #2 (Public review):

      Aspects of the behavioral analysis would benefit from a tighter connection between theoretical claims about evidence accumulation and the empirical features of the psychometric functions. For example, the rightward shifts observed across adapting conditions are interpreted as consistent with a reset of accumulation on switch trials, but similar patterns could also arise from failures to detect the test stimulus on a subset of trials, leading responses to default to the final adaptor direction. Likewise, changes in psychometric slope and asymptote are attributed to differences in evidence accumulation without explicit modelling or consideration of alternative explanations.

      Clarifying how specific features of the psychometric functions map onto distinct components of the decision process will strengthen the link between the theoretical framework and the behavioral data.

      We agree and have made substantial revisions to address these important points. Specifically, we added a new figure panel (Fig. 2D), new Extended Data Figures (3 and 9), and several lines of explanatory text (L152-179) that collectively describe the behavior in more detail, including clarifying that: 1) for the shortest viewing durations, the monkeys’ decisions were informed by information from the adapting stimulus, which accounts for generally lower accuracy on LSF (longer exposure to the final adapting direction, thus more accumulated evidence for that direction before processing the switch) vs. HSF (shorter exposure to the final adapting direction, thus less accumulated evidence for that direction before processing the switch) switch trials; and 2) as viewing duration increased, the rate of rise of accuracy versus viewing duration was higher for LSF vs. HSF trials, implying differences in the process of evidence accumulation. As detailed in our response to a similar comment from Reviewer 1, above, we are now careful to temper our claims about the specific computational basis (e.g., a leak or other form of nonlinearity) for these differences.

      We also de-emphasized our treatment of the asymptotes of the psychometric functions. In principle, these regimes could give insights into leakiness (which can limit the total amount of information that can be accumulated) and lapses (which are measured at the asymptotes). In practice, however, the long-duration trials that constitute the asymptotes were relatively under sampled (to promote the unpredictability of the offset of the stimulus, which we believed was the more important consideration when designing the experiment), yielding unreliable estimates.

      A slight concern is the lack of a consistent analytical approach for relating behavioral changes to neural and pupil-linked measures. Different sections of the manuscript rely on different behavioral metrics-such as differences in accuracy within a selected stimulus-duration range (e.g., Figure 5C) or psychometric slope differences (Figure 6C) without clear justification for these choices. The analytical approach likewise varies between simple correlational analyses (Figure 5C, Figure 6C), pseudo-experimental group comparisons (Figures 5D, E), and the inclusion of neural or pupil terms in the behavioral psychometric regression model (Figure 7B). While each metric and approach may be defensible in isolation, adopting a more consistent framework will help convince readers that the reported effects are robust and not contingent on the selective choice of metric or analysis.

      We thank the reviewer for this thoughtful critique and agree that the rationale for our choice of behavioral metrics and analytical approaches could be stated more clearly. We have added text to the relevant sections of the Results (L247-251) clarifying these choices. In particular:

      The neural analyses (Figures 3D-E, Figure 4, Figure 5D-E) focused on preferred-motion switch trials, because: 1) low switch-frequency non-switch trials provide an additional 800 ms of exposure to the final adapting-stimulus motion direction relative to high switch-frequency non-switch trials, which confounds comparisons of context-dependent evidence encoding between conditions, and 2) MT neurons exhibit minimal responses to null motion (although note that we also included analyses based on ROC area, which is computed from both preferred- and null-motion switch trials, to account for possible contributions of null-motion responses; Figure 5A-C). Thus, to ensure a meaningful comparison between neural and behavioral measures, we used behavioral accuracy on switch trials as the relevant metric in Figure 5C-E, rather than psychometric slope, which is estimated across both switch and non-switch trials.

      The pupil analyses (Figure 6) focused on a time window preceding test-stimulus onset, representing the arousal state around when the decision process started, and included both switch and non-switch trials. Thus, for these analyses we used psychometric slope, which is estimated across both switch and non-switch trials.

      We used several different analyses to compare and contrast the neural-behavioral and pupil-behavioral relationships because they provide complementary and useful insights. The correlational analyses in Figures 5C and 6C characterize session-level relationships between neural/pupil signals and behavior. The group comparisons in Figures 5D–E provide a complementary visualization of the same relationship. The model-based approach in Figure 7 then allows direct quantification of the trial-wise contributions of each signal to behavior within a common framework. Importantly, the conclusions drawn from each approach converge on the same interpretation, which we believe speaks to the robustness of the reported effects.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 2 legend. Description of 'running average (5-trial window)' is unclear - presumably this is a running average in stimulus space rather than across trials.

      We thank the reviewer for flagging this ambiguity. We have updated the legend (L136-137) to clarify that the running average is computed across trials sorted by test-stimulus duration.

      (2) L158. Difficult to establish an asymptotic performance level for HSF conditions within the stimulus duration range tested.

      We have removed the reference to asymptotic performance and replaced it with a discussion of performance on longer-duration switch trials in the context of the newly added Figure 2D.

      (3) L515 Equation 1. While this is a standard formulation of lapse rate in psychometric functions, the construction here in terms of switch probability is not standard. Given the task and training, it seems more likely that on lapse trials, the animal will respond according to the last adapted direction (rather than randomly switch/stay with equal probability).

      We thank the reviewer for this point. We agree that it is possible that on at least some of the “lapse” trials the monkeys may respond according to the final adapting-stimulus direction rather than choosing randomly. However, we cannot distinguish those alternatives using this task design. We include a statement to this effect in Methods (L569-571).

      To explore the idea further, we refit the behavioral data using separate upper and lower asymptotes corresponding to lapse rates on switch and non-switch trials, respectively. Across monkeys, there were no significant differences between upper and lower lapse rates for either low (Wilcoxon signed-rank test for equal medians: p = 0.15, Cohen's d = -0.13) or high switchfrequency (p = 0.07, Cohen's d = -0.16) conditions. So, at the very least, there was no evidence for lapse-like errors driven by switch- (or non-switch-) specific defaults to the final adapting direction.

      (4) L256. Statistical significance of attenuation is not directly tested here.

      We have replaced "were attenuated" with "we did not identify any reliable context-stability differences" (L297) to accurately reflect what was directly tested without implying a statistical comparison between groups of sessions that was not performed.

      (5) L429. Does the increase in explanatory power warrant the increased complexity of the model here?

      We thank the reviewer for raising this important point. We used Tjur's pseudo-R<sup>2</sup> because it does not increase by default with added model complexity, making it more conservative than other R<sup>2</sup> measures in this respect. Tjur's pseudo-R<sup>2</sup> is a coefficient of discrimination, and as such its value increases only when additional terms improve the model's ability to separate predicted probabilities across response outcomes. Thus, the observed increases in explanatory power when adding neural or pupil terms reflect real improvements in discriminability rather than an artifact of model complexity. We have added a brief clarification of this point to the Methods (L662-664).

      Reviewer #3 (Public review):

      The task design may not be optimal. While the amount of time the monkey is exposed to each motion direction during the adapting stimulus is matched, it's hard to know if the reduced MT responses to the test stimulus are truly due to the greater frequency of switches during the HSF adapting stimulus or because the monkeys have been exposed to more repetitions of the stimulus. It's increased sensory adaptation in either case, but it makes it problematic to interpret this as temporal context-dependent adaptation specifically. I think this could potentially be partially addressed by an analysis that is in the paper, but could potentially be emphasized/fleshed out more, specifically the results shown in Figure 4D that seem to show that most of the reduction in neural response for adapting units occurs between the first and second stimuli.

      The reviewer raises an important point. The number of stimulus repetitions and switch frequency are confounded in the experimental design, making it difficult to attribute context-dependent differences in MT responses to the temporal pattern of switches rather than to accumulated repetitions. We also note, as the reviewer acknowledges, the observed differences reflect sensory adaptation either way. Figure 4D does offer relevant evidence, suggesting that a majority of the change in neural response occurred with just one stimulus repetition. This finding complicates an interpretation where adaptation scales with the number of stimulus repetitions. We have added several lines to the Results about these points (L231-233).

      The pupillometric analysis seems to be an indirect way of assessing whether the accumulator itself might be modulated by temporal context, but the link could be made clearer. The authors show that context-dependent behavior is related to pupil size, which is related to arousal/neuromodulation, but it would be helpful to have some idea of what neural mechanisms underlying adaptive decision-making are actually impacted by this neuromodulation. Lacking neural data to address this question (e.g., from a brain region proposed to be involved in the accumulation process), at least more discussion of this would be helpful. Essentially, I'm unsure of how to interpret the pupil results: the argument that temporal context affects instantaneous evidence encoding in MT that then drives the accumulator is very clear, but I am a bit confused about what, mechanistically, I should think about the effect of neuromodulation doing.

      We thank the reviewer for this thoughtful comment and agree that the mechanistic interpretation of the pupil results could be made clearer. We acknowledge that we cannot directly identify the neural mechanisms underlying the arousal-related contributions to adaptive evidence accumulation from pupil data alone, given that pupil size is an indirect and imperfect proxy for neural (e.g., LC-NE system) activity. However, we can offer some informed conjecture and have added to the Discussion (L469-482) in an effort to elaborate on possible mechanisms.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract could be retooled - does not emphasize the pupillometry/arousal results very much, and they are presented more as a control than an independent result.

      We agree and have revised the Abstract accordingly.

      (2) Do all neural/pupil analyses use only switch trials? Sometimes the figure captions do specify only switch trials, but not everywhere. It would be helpful to specify either in the Methods or at the beginning of each figure caption that all subplots show switch trial results. Also, if you do always use switch trials, it would be useful to see in the Supplement how the non-switch trial results differ from switch trials. It seems like they may in interesting ways based on the behavioral results (supporting a reset of evidence accumulation on switch but not non-switch trials).

      We thank the reviewer for flagging these important points. We have added a justification for switch trials (L186-190) as well as clarification about which trial types were used for which analyses (L246-249) and information about trial types to relevant figure captions. We have also added a new Extended Data figure (Extended Data Fig. 9) examining relationships between neural activity and behavior on non-switch trials. As inferred by the reviewer, behavior on non-switch trials is consistent with the use of information from the adapting stimulus.

      (3) In Figure 3C, 5B, etc, when computing firing rate for the test stimulus (50-500 ms), are differently sized windows used to compute the rate for different test stimulus durations (since some will be <500 ms)? Or are only trials where the test stimulus duration is > 500 ms used for this analysis?

      We thank the reviewer for raising this point. To clarify, the 50–500 ms window does not reflect a fixed window applicable for all trials. Rather, neural activity from 50 ms after test-stimulus onset through test-stimulus offset was included for each trial, with 500 ms serving as the upper bound for trials with longer durations (> 500 ms). We have clarified this in the Methods (L607-610) to avoid ambiguity.

      (4) I think it might be better to be consistent with the time windows used for analysis; specifically, to choose either the 50-500 ms window used in Figures 3, 4, and 5B, or the 200- 400 ms window used for the remaining analyses in Figure 5.

      We agree that using the same window for all of the analyses would improve consistency, but not doing so provides advantages that we believe take precedent and now describe in more detail. The broader 50–500 ms window used for Figures 3, 4, and 5B was chosen to characterize MT neural activity over a relatively large a time window, ensuring that every trial contributes to each estimate. Because test-stimulus durations were drawn from a truncated exponential distribution (100–1200 ms), restricting these analyses to the 200–400 ms window would have excluded the substantial proportion of trials with durations <200 ms (but would yield similar figures and conclusions). The narrower window used in subsequent analyses allows us to focus on the conditions that exhibited the biggest modulations of neural activity when comparing them to behavior.

      (5) Similarly, provide justification for using only trials ending 375-600 ms after test stimulus onset for the behavioral correlations. It seems reasonable to choose a subset of test stimulus durations where the monkeys' behavior is greater than chance but less than ceiling, but it would be good to specify this so that it doesn't seem arbitrary.

      We agree and have added text to make this important point (L249-251).

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. My major concerns from the initial manuscript, especially regarding cherry picking and circularity have been addressed with revised analytical approaches. I have some remaining minor comments.

      (1) The comparison between two wild-type samples versus wild-type-mutant samples is helpful - I think this could be added to the manuscript.

      As suggested, we added the figure comparing two wild-type samples against wild-type mutant samples as Supplementary Figure 2. 

      (2) For results of time-lapse imaging - please spell out in the results section the direction of change (lines 270 - 277).

      As suggested, we added the direction of change (an increase in the turnover rate) to the text (page 12, lines 270-271).

      (3) Using linear mixed effect models for statistical analysis is a significant improvement. While a sample size (n) of mice = 3 is not ideal, I think given the multiple different mouse lines used and intensity of analysis, this is probably the best that can be done, although further validation in larger samples eventually is to be hoped for.

      We appreciate the reviewer for recognizing the effort required to collect data across multiple mouse lines.

      (4) The revised text is much improved, but I still think the authors should be upfront somewhere in the text that the schizophrenia-associated genes can only confer biased risk for schizophrenia (and that the clinical phenotype can also include autism). As I said before, I think this is the best we can do and I agree with their choices, but it is important not to overstate the link. The differences they see make it clear that these are still relevant distinctions.

      As suggested by the reviewer, we further modified the discussion related to the comparison between ASD- and schizophrenia-associated mouse models (pages 23-24, lines 508-522).

      “The nanoscale features of dendritic spines in mouse models of Nlgn3<sup>R451C/(y or R451C)</sup>, Syngap1<sup>+/−</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, which we classified as being related to ASD, are highly heterogeneous. This heterogeneity may reflect the broad clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. Accordingly, these four mouse models may represent distinct subgroups characterized by different degrees or forms of hippocampal dysfunction. Notably, among the ASD-related models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those found in the 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> mouse models. Although we classified 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> as schizophrenia-related models, both 22q11.2 deletion syndrome and Setd1a haploinsufficiency in humans are also associated with ASD, suggesting substantial overlap in the genetic risk factors underlying ASD and schizophrenia. Further systematic analyses linking rare genetic variants to synaptic phenotypes in mouse models may provide important insights into the mechanisms underlying both shared and disorder-specific synaptic alterations in neurodevelopmental and psychiatric disorders.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I would suggest that it might be preferable to use the word 'neuropsychiatric' rather than 'mental' in the title.

      As suggested, we modified the manuscript title.

      (2) I think it would be clearer to say that DEGs are listed if present 'in three or more models' rather than >2 (I appreciate the latter is mathematically clear, but can easily be read as 2 or more if reading fast). This is changed in the figure legend, but I suggest it is also changed in the main text (line 352-3)

      As suggested, we changed the main text to incorporate "in three or more models" (page 16, line 352).

      (3) Please add to Methods (line 557) that 'control cultures were prepared from littermate embryos....'

      As suggested, we added the phrase "control cultures were prepared from littermate embryos" (page 26, line 559).

      (4) Sorry to add something, but please could the authors add a definition of how they calculate spine turnover (and add units to the y axis of Figure 5A-C)?

      As suggested, we modified the y-axis of Figure 5A-C (% as unit) and added the method of calculating spine turnover rate in the text (page 36, lines 808-811).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource-limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. While the model is sensible from a mechanistic perspective, the experimental evidence that supports how each model's rate is affected by the cell metabolic state is weak, as only ratios of these rates can be inferred from the data.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phages that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well-designed, and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication, and the paper shows that host physiology has to be taken into account at all steps of phage infection.

      Weaknesses:

      There are some concerns about the experimental setup and which conclusions can be drawn from it:

      Before phage infection, bacterial cultures are grown to exponential growth, washed, and then resuspended with glucose or arsenate-azide for 10min. It is however, questionable that 10 minutes is enough to simulate high and low metabolic states realistically. 10 minutes seems to be quite short to go from exponential growth to a low metabolic state, given the transcriptional memory of previous environments. It seems more likely that the population will be quite heterogeneous, with cells in various states of transition towards low metabolic states.

      While we agree with the reviewer that during metabolic transitions there may be a period in which the population is heterogeneous, with cells in different stages of transition toward a low metabolic state, the 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyper diffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027). We have also corrected the DOI for this reference in the manuscript. Furthermore, the ATP pool of log-phase E. coli turns over several times per second (Holms et al., Arch. Mikrobiol. 1972, http://dx.doi.org/10.1007/BF00425016). We therefore assumed the bacteria were energy depleted after 10 minutes. We have clarified this point in the revised manuscript.

      Given that arsenate and azide inhibit cellular metabolism, i.e., have antimicrobial effects, cells might not just downregulate metabolism but also activate the stress response, and this causes some of the observed effects on phage adsorption. Therefore, the 'low metabolic state' of the cells in this paper could mean that cells are starved or that they are stressed or both.

      The reviewer is correct. We don’t exclude indirect effects. However, as nutrients were removed from the bacteria by washing and energy metabolism was inhibited by the addition of arsenate and azide, we assumed a stress response requiring biosynthesis would be unlikely to occur.

      The abundance of receptors could change between the high and low metabolic media conditions and contribute to the observed differences in adsorption, while the authors seem to assume in their model that the initial adsorption rate always remains the same.

      We do not think that the observed differences in adsorption are explained by a change in receptor abundance. In a previous study using the same experimental protocol as in the present work, phage λ was compared to the metabolically insensitive mutant λh (Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119). If the lower adsorption in the low-metabolic condition were caused by a reduced number of receptors, then λh should also have shown a lower adsorption rate under the same condition. Instead, no measurable effect on λh adsorption rate was observed. We therefore conclude that the effect is not explained by changes in receptor number on the timescale of the experiment. We have clarified this point in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, ɸ80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy-depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      Strengths:

      The data presented by the authors clearly supported the principal conclusion of the study ("Viral commitment to infection depends on host metabolism"). The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully used a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Weaknesses:

      (1) The authors isolated and measured the numbers of free phages in the medium after infection of bacteria under different treatments. These measurements were analyzed in two different ways: (1) simply as ratios (corrected/normalized using different controls), and (2) fitted using a simple mathematical model. I have concerns regarding both analyses.

      (1.1) For the first method, having different time points at which the sample of each phage is collected critically complicates data interpretation. As one incubates the phage-bacteria mixture for a longer time, more infection occurs, and the number of phages collected from the mixture decreases. Therefore, the different incubation time forfeits the goal of "a systematic and quantitative comparison across different phages [...]", just as the authors self-criticized. Conceivably, the authors could have used the shortest measurement time for all phages (i.e., 10 minutes, as for phage λ). Alternatively, the authors could have applied a systematic criterion such as half (or any other fraction) of the latent period of each phage, which would still "maximize the incubation period while ensuring that manipulations were completed before the first infection cycle concluded". In my view, the seemingly arbitrary measurement time for each phage renders the entire first analysis very challenging to interpret. It also goes against the author's proposition that the protocol was "standardized" or "consistent". It is not clear what the readers are supposed to take away from this first analysis, or rather, which evidence, finding, or conclusion the manuscript would lose if the authors only presented the modeling-based analysis.

      (1.2) The second method of analysis sought to remove the dependence of the measurements on time. I completely agree with this goal, and the findings extracted from this analysis significantly contributed to the merits of this manuscript. However, the authors achieved this goal using a single time point for each phage to calculate the infection rate (η). As shown in Figure S3, each of the phage depletion curves is anchored by only one data point (note that the P(t)/P(0) = 1 at t = 0 is assumed, not measured). This goes against the typical way this collision model is used in the literature, where a time series is measured and used to fit the model (e.g., DOI 10.1007/978-1-60327-164-6 18, or more recently, PMID 39700139). This practice in the current manuscript reduced the robustness of the inferred η values. This problem is exacerbated by assumptions used by the authors in formulating this model. For instance, the authors used a constant value for the bacterial concentration, B, because "bacterial growth and lysis were negligible" (lines 135-136). However, considering that the bacteria were cultured at 37oC in a very rich medium (first in YT broth, then in 2% glucose), the measurement times of 20, 30, and 55 minutes are most likely one or a few generations of bacterial growth and division.

      Related note: I suggest that one of the panels in Figure S3 should be moved to the main text, since it is critical to the second method of analysis.

      We would like to clarify that the manuscript does not present two separate methods, but rather one method presented in two steps: a first step with results that are directly tied to the experimental measurements and show whether the effect is present for each phage, followed by a second, analytical step that makes the results comparable across phages.

      The first step presents the ratios because they directly reflect the measurements performed in the experiment and allow the reader to see the effect of the metabolic state for each phage in contrast to its control. We agree that these ratios are time-dependent and therefore not suitable for quantitative comparison between phages. Their purpose is to illustrate the experimental outcome and to show that the effect is present (or absent) on a per-phage basis not to compare magnitudes across phages.

      We then follow this with the second step, allowing the reader to follow the logic of the analysis. The analytical step that follows does not represent a second method, but a continuation of the same analysis. Here, we remove the time-dependence specifically in order to make comparison of the effect across phages possible, by connecting our results to standard measures such as the adsorption rate η. Importantly, P(0) is measured for every phage in every experiment. The only modeling assumption used (a standard one in the field) is the exponential form for the decay in free phage number, which naturally yields P(t)/P(0) = 1 at t = 0.

      Regarding the reviewer’s concern that bacterial growth may not have been negligible over the relevant time window, we note that recent work on rich-to-minimal growth lags in E. coli reports substantial delays before growth resumes after nutrient downshift. One 2023 study (Wu et al., Nature Microbiology 2023, https://doi.org/10.1038/s41564-022-01310-w) considering wild-type E. coli shows in Fig. 2c a lag of up to about 2 hours after a shift from MOPS minimal medium with 0.2% glucose plus 18 amino acids to the same medium without amino acids. Another 2023 study (Zhu and Dai, Nature Communications 2023, https://doi.org/10.1038/s41467-023-36254-0) examining both rel+ and rel− strains reports a growth lag of about 49 minutes for rel+ and more than 5 hours for the relA deletion strain. While these conditions are not identical to ours, they support the general point that growth does not immediately resume after such shifts. We therefore think it is unlikely that, following transfer from YT, the cells underwent one or a few full generations during the time window of our adsorption measurements.

      On the related note: Following the comments of all reviewers on Figure S3, we have decided to remove it to avoid confusion.

      (2) The data were able to distinguish phages that successfully infected bacteria and those that remained free in the medium, and the authors appropriately interpreted the data as such throughout the Results section. However, in the Discussion (starting from the very first sentence, line 172), the authors used terms that include "adsorption" and "entry" more interchangeably (for example, see the three sentences in lines 310-313, for "viral entry efficiency is shaped by [...]", then "adsorption kinetics modeling"). I do not see how the authors' data could distinguish between adsorption (the phage particles attaching to the outside of the cell) and entry (the phage DNA being injected into the cell). Conceivably, any phage particles that irreversibly attach to a cell but do not yet inject their genome into the cell would still be removed from the medium and therefore not quantified. Another example: in lines 189-191, the authors interpreted that "[...] when the bacterium is in a low metabolic state, the phage does not bind irreversibly to the host", but how do the authors eliminate the case of no phage binding (i.e., the reversible step) to begin with?

      We agree with the reviewer that our use of the terms adsorption, entry, and infection should have been more careful. Our experiment can only identify the irreversible commitment of phage to a host cell. We have therefore revised the text to refer consistently to phage commitment.

      Similarly, in lines 283-293, how do the authors delineate whether energy depletion would increase the k_off term or decrease the k_inj term, because either would result in more free phages in the medium as observed in the data? I believe that the writing of the Discussion, as it stands now, is doing a disservice to the conclusions presented in the Results section.

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: if energy depletion leads to reduced commitment, this can arise either because k_off increases, because k_inj decreases, or because both change, as long as k_off/k_inj becomes larger in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also leads to the trade-off now discussed in the revised manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become significantly larger in the inactive case. Depending on whether this is achieved through changes in k_off or k_inj, the cost of discrimination appears either as slower commitment or as additional energy dissipation. We agree that the previous wording overstated the mechanistic interpretation, and we have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (3) The authors presented an argument that performing infection of all five phages in the same condition is an advantage, allowing for comparison across different phages. While this goal is a completely valid one, it is difficult to reconcile that with the fact that different phages require different optimal conditions for successful infection. For instance, phage T5 famously requires Ca2+ for successful infection into the host bacterium (and later successful replication); see PMID 13174489. However, all infections were performed in TMG, which lacks Ca2+. Perhaps the absence of T5 dependence on the host metabolism is because the infection condition used by the authors was not optimal for T5 to begin with? Similar arguments could be made for other phages.

      Our study alone cannot eliminate that possibility. However, we have cited multiple previous studies, for example references citing Braun et al., showing that T5 remains insensitive to the host metabolic state under different buffer conditions. We therefore believe it is unlikely that the lack of metabolic dependence we observe for T5 is simply due to suboptimal infection conditions.

      (4) Whereas the manuscript examined five coliphages, only phage T5 and phage λ were discussed extensively. I believe some discussion points for these two phages need clarification.

      We focused our discussion on the phages T5, λ and φ80 because these are the phages for which similar effects have been reported previously in the literature. This allowed us to connect our findings directly to existing work and to discuss mechanistic hypotheses in a meaningful comparative framework. For the remaining phages, to our knowledge no prior studies have examined their behavior under comparable metabolic conditions, and therefore a similarly detailed discussion would have been speculative. Nevertheless, all five phages are treated equally in the presentation of the experimental results and in the quantitative comparison of adsorption rates.

      (4.1) Phage T5: The data obtained by the authors show that the infection rate of phage T5 is not impacted by the metabolic state of the host cell. Considering that the authors used the terms "infection", "adsorption", and "entry" interchangeably to refer to the irreversible commitment of a phage to a host cell (see point 2), this discussion regarding phage T5 lacks one critical literature context: DNA entry of phage T5 is known to occur in two phases (first-step transfer and second-step transfer). Critically, the second step can only occur if phage proteins encoded by the phage DNA transferred in the first step are expressed (see PMID 10577483 and the cited papers therein). In that context, metabolic poisoning of the host bacteria should have impeded T5 infection. The authors should comment on this point.

      As the reviewer pointed out, our usage of the terms infection, adsorption, and entry should have been more careful. Our experiment can only identify irreversible commitment of phage to a host cell. For T5, we expect that this irreversible commitment already occurs upon first-step transfer of phage DNA. As a result, even if second-step transfer is impeded under metabolic poisoning, our method would not resolve that effect. We have added this clarification to the revised manuscript.

      (4.2) Phage λ: The experiment using phage λ in this current study shares many resemblances to that in Brown et al. 2022. That feature alone is not a problem, but at many places in the text, the writing is ambiguous as to whether it is discussing the results in Brown et al. 2022 or in the current manuscript. I am giving three examples below, but this is not exhaustive: (i) Lines 67-69, there is no Brown et al. 2022 reference immediately after "a mutant phage variant (λh) could bypass this dependency [...]" (not just in the previous sentence); (ii) Line 228 should clearly say "Our previous findings suggested that phage λ is capable of [...]", since it concerns Brown et al., 2022, not the current study; and (iii) Lines 245-246, there is no Brown et al., 2022 reference immediately after "we observed that a mutant variant [...] even energy-depleted host" (without a reference, it reads like the authors "observed" that finding in this current manuscript).

      The reviewer is right. In those places, the text was ambiguous as to whether it referred to the present study or to Brown et al. (2022). We have now inserted the reference at the relevant points and revised the wording where needed to make this distinction explicit.

      Also, regarding phage λ: The discussion between line 230 and line 249 is very interesting, but since it concerns the differences between λ PaPa and Ur-λ, the authors should consider mentioning and discussing a very relevant recent study, PMCID: PMC6312755.

      We agree that the study by Guan et al. is very relevant and interesting. However, our point in this part of the Discussion is only to clarify that we used λ PaPa and not the originally isolated λ strain. We have therefore limited the discussion here to that distinction.

      (5) Control experiments, or references to prior studies, are needed to support that the As/Az treatment at this concentration and duration (at least 10 minutes) is sufficient to deplete the metabolic state of the cell. For instance, this can be shown by impeded or null cell growth, arrested motility (using a standard swimming assay), or a fluorescent reporter for the energetic state of the cell.

      The 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyperdiffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027) where the effect was assessed by monitoring the rate of movement of the λ receptor on the bacterial surface. We have clarified this point in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As mentioned earlier, I found the paper interesting and addressed an important and significant knowledge gap.

      My biggest concern is about the interpretation of the experimental data in light of the two-step model. In particular, around line 286, it is stated "k_inj is more sensitive to metabolic state than k_off". Assuming k does not depend on metabolic state, which is a fair assumption, the equation for eta only depends on the ratio between k_inj and k_off and not on the individual parameters separately. Consequently, there is no way of saying which one of the two is more affected by metabolic state, unless the model already assumes that k_off is not influenced by metabolic state. The results could equally be explained by k_inj decreasing in metabolically depleted cells, or k_off increasing in such cells. If this is an assumption of the model, this should be clearly stated and not reported as a consequence of the data, as it is at the moment. Also, how does this mathematical model connect to the fitting function used in Figure 2b?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: discrimination requires that k_off/k_inj be larger for inactive hosts than for active hosts, such that commitment is specifically reduced in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This introduces a trade-off: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination requires this ratio to become significantly larger for inactive hosts. If this is achieved through changes in k_off, discrimination comes at the cost of slower commitment by allowing more time to leave; if it is achieved through changes in k_inj, it can preserve fast commitment to active hosts but requires additional energy dissipation in order to actively modulate commitment. We have therefore revised the text accordingly to frame the argument in terms of this trade-off, rather than attributing the effect specifically to k_inj. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      I have a related experimental criticism. The kinetic model presented assumes an exponential decay of free phage, which is a commonly used assumption in the phage literature. Given that the phage types used in this study lyse relatively slowly, it would be good to actually see adsorption curves, in which free phage is measured at different time points between inoculation and lysis. This data would not only provide useful evidence for the kinetic model, but it should also replace what is now in Figure S3, which consists of fitting one experimental point with one line. As it currently stands, Figure S3 is not useful actually misleading.

      We appreciate the reviewer’s point. We agree that adsorption curves, in which free phage is measured at different time points between inoculation and lysis, would provide a stronger basis for evaluating the kinetic model. However, we do not have the resources to perform these additional experiments within the scope of the present study. Following the comments of all reviewers on this point, we have therefore decided to remove Figure S3 to avoid confusion.

      Finally, it is not clear to me why the quantity "Ratio" has been chosen to be presented in Figure 1, rather than the ratio of estimated adsorption rates eta'/eta, which is much more intuitive for a phage study and contains the same information. I would recommend switching to this choice, unless there is a clear rationale for why the quantity "Ratio" is more useful/effective. Showing eta'/eta would also increase the readability of Figure 1, as it would move the y-axis to a logarithmic scale and better visualize values around 1.

      We used “Ratio” in Figure 1 to illustrate the experimental design, controls, and measured quantities directly, as it more transparently reflects the data collected. In the second part of the analysis, where we compare time-independent adsorption rate estimates, we have presented the corresponding values of η′/η as suggested.

      Minor comments:

      (1) Introduction

      Line 31: "... such as nutrient limitation, fluctuating temperatures, and variable energy availability" - if drawing a distinction between energy availability and nutrient limitation, please make explicit what this distinction is. Energy availability seems like a natural consequence of nutrient availability.

      While energy and nutrient availability are often linked in E. coli, they represent distinct physiological constraints. Nutrient limitation refers to the lack of essential biosynthetic precursors such as nitrogen, phosphorus, or amino acids. Energy availability, in contrast, reflects the cell’s ability to generate ATP and reducing equivalents through metabolic processes. For example, under anaerobic conditions, E. coli may have ample nutrients but limited energy production due to the lower efficiency of fermentation compared to aerobic respiration. Thus, energy limitation can occur independently of nutrient limitation.

      (2) Results

      (a) Whole Section: Please label equations.

      All equations have now been labelled in the revised manuscript.

      (b) Lines 105 to 114: As stated in Major Comments, I think the clarity of the paper would be improved by introducing the relative adsorption rate here and dropping the concept of Ratio entirely. However, if the authors wish to use Ratio, I would recommend the following:

      Lines 105 to 109 are confusing to read because of the number of connectives: "... ratio of free viruses from permissive AND resistant hosts respectively TO the free viruses in buffer under energy-depleted AND energy competent conditions". This would be clearer if each quantity were given an algebraic symbol, and RP, RR, and Ratio were defined through formal algebra, rather than mixed mathematical and sentence notation.

      This section has been rewritten for clarity. We now introduce explicit algebraic symbols and define the quantities formally, which removes the ambiguity present in the sentence-only description while retaining the intended meaning.

      The chemical names "arsenate" and "azide" should appear in the body of the text before they appear abbreviated in an equation. Please state at this point that these are both metabolic inhibitors, as it is not immediately clear what role they play or why you are using them.

      The text has been updated to introduce arsenate and azide by name before the abbreviations are used, and we now explicitly note that they act as metabolic inhibitors.

      On line 114, the authors helpfully provide an interpretation of Ratio = 1. It would be useful to provide at the same time interpretations of Ratio >1 and <1, perhaps 2 and 0.5 specifically?

      We have added brief explanations illustrating the interpretation of Ratio values greater than and less than 1, including examples of 2 and 0.5.

      I would consider giving this quantity a more interpretable name than Ratio. This quantity represents how much a bacteriophage preferentially adsorbs to metabolically active cells, so perhaps "Selectivity" or "Adsorption Bias"?

      We intentionally retained the generic term “Ratio”, as this quantity reflects an intermediate experimental measure used to describe the process rather than a newly defined metric. Its purpose is to bridge the experimental observations and the subsequent quantification of effects on the adsorption rate (η).

      (c) Lines 117 to 122: the authors sometimes refer to ratios explicitly, "average ratio of around 1.6" and other times say e.g., "a greater than 3 times increase in viral particles". Using more consistent language (saying "Ratio" every time) would be clearer.

      We have standardized the terminology in this section and now refer to all fold-changes consistently using “Ratio” to avoid ambiguity.

      (d) Figure 1

      Phages λ and T6 look like they have ratios less than 1 for resistant cells? If this is true / if the ratio is statistically significantly below 1, please comment.

      The ratios for λ and T6 are not statistically different from 1. The apparent deviation is within the standard error of the mean. To make this clearer, we have added the corresponding p-values to Table S2 in the Supplementary Information.

      Ratios near 1 are difficult to distinguish from 1, especially in panels A and D. Using a logarithmic scale on the y-axis would make the plots more readable.

      Because the values in these panels are not statistically different from 1, changing to a logarithmic scale would not alter the interpretation. We therefore retained the current axis scaling to reflect that there is no meaningful deviation from 1 in these cases.

      The data corresponding to individual experiments have no error bars. Given that the number of free virions was determined by plaque assay, which carries an intrinsic sampling error, this uncertainty should be reflected in the plots.

      We thank the reviewer for this important comment. Because plaque assays have compound sources of stochastic variation, assigning a per-measurement error bar would risk implying false precision. For this reason, we present the values from each biological replicate directly, and the uncertainty is represented in the statistical summary across replicates. Specifically, for each phage and condition we show the three independent experimental measurements and report the mean along with the standard error of the mean. This approach allows us to represent biological variability without implying a precision that cannot be accurately quantified at the level of single plaque counts.

      Similarly, the average value does show error bars, but it is not stated what these error bars correspond to: standard error in the mean, standard deviation of the sample, or combined uncertainty?

      The caption has been updated to state that the error bars represent the standard error of the mean.

      The resistant bacteria seemed to have ratios close to 1 in all cases. Is this because very few virions adsorbed under both energy conditions?

      Resistance is commonly associated with a lack of a surface receptor for the phage (or generally an entry pathway). We use the resistant bacteria as a control group for the effect of the conditions on adsorption. For resistant bacteria, the Ratio should be 1 since virions do not adsorb under both energy conditions. Any slight variations from 1 should come from sampling errors or small heterogeneity in the population.

      (e) Figure 2

      Please comment on what the error bars here represent. Error bars in Figure 2 A seem to permit negative (or at least zero) values of relative adsorption rate for phages m13 and T6, possibly implying an overestimate of the error? If it is the case that multiple values used to calculate the mean are far apart, possibly showing the values individually through a superimposed swarm plot would be clearer.

      This point is now addressed in the Supplementary Information, where we clarify how the error bars were calculated.

      (3) Discussion

      (a) Line 189: "high metabolic state" is imprecise. Say "energy-competent" to be consistent with earlier language.

      To maintain continuity with earlier terminology, we now include “energy-competent” in parentheses alongside “high metabolic state,” while retaining the original phrasing for readability.

      (b) Figure 3, population level

      Show adsorbed virions physically attached to bacteria, rather than removing them completely from the image, as currently, the implication is that at a high metabolic state, there are fewer virions total, not fewer virions remaining in solution because more are adsorbed. You could go as far as to add a third "after centrifuging" row, showing the adsorbed phages stuck in the pellet and the unadsorbed phages remaining in solution.

      Thank you for this suggestion. Figure 3 has been updated to depict adsorbed virions attached to bacterial cells, clarifying that the decrease represents adsorption rather than loss of total particles. This change improves the accuracy and interpretability of the schematic.

      (4) Methods and Materials

      (a) Figure 5

      The step "estimate cell numbers from OD" appears to follow incubating plates overnight. If the cells you are counting come from the pellet produced by centrifuging 3 steps prior, you could add a fork into the black line connecting the steps, with one branch corresponding to the supernatant and phages, and the other to the pellet and cells?

      Thank you for pointing this out. The order in the figure has been corrected: cell numbers are estimated from OD before overnight incubation. This resolves the confusion without the need for branching in the workflow diagram.

      (a) Line 332

      You allow as much time as possible for adsorption without the possibility of lysis. Did you determine the lysis times / latent periods of these phages through one-step-growth-curves, or use published results, in which case please cite? Having obtained the lysis time by either method, what fraction of the lysis time did you allow for adsorption? Also, please add supplementary tables with lysis times used for the different phages.

      We thank the reviewer for this comment. We used published latent-period values as guides and verified compatibility with our own system when selecting incubation times. We have clarified this in the text and added the relevant citations. We did not use a common fixed fraction of the lysis time for all phages; instead, incubation times were chosen to allow sufficient time for adsorption but not for completion of the first lytic cycle. For λ, productive lytic development was blocked in the host background used, as in Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119. For ϕ80 and T5, we used published latent-period values as guides and verified their compatibility with our own system (De Paepe and Taddei, PLoS Biology 2006, http://dx.doi.org/10.1371/journal.pbio.0040193). M13 is a chronic filamentous phage and therefore does not have a standard lytic latent period; in our host–phage combination, it required more than 1 h before phage release. For T6, we relied primarily on the kinetics observed in our own system, since adsorption was unusually slow for this phage–host pair under our assay conditions. Although literature reports describe shorter T6 latent periods under specific assay conditions (Foster and Johnson, Journal of General Physiology 1951, http://dx.doi.org/10.1085/jgp.34.5.529), this is consistent with published work showing that adsorption and infection kinetics can vary substantially with host background, surface structure, and experimental conditions (Heller and Braun, Journal of Bacteriology 1979, http://dx.doi.org/10.1128/jb.139.1.32-38.1979; Storms et al., Biochemical Engineering Journal 2012, http://dx.doi.org/10.1016/j.bej.2012.02.010).

      (5) Supplementary

      Figure S1

      This data is useful in understanding the main body of the paper, and I think this should form part of a main figure (possibly with the individual experimental data points superimposed over the bars). This could come before or as part of Figure 1?

      We thank the reviewer for this suggestion. We have explored including these data directly in the main figure but found that doing so substantially reduced the readability of the figure, as the underlying table is visually dense. For this reason, we chose to summarize the results in Figure 1 and present the detailed data separately in Figure S1 of the Supplementary Material, along with the Ratio analysis, which more effectively conveys the trends without overloading the main figure.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) L16-18: This sentence could be made more accessible as 'error correction' is not an intuitive term in the phage field.

      We have updated the overall theory section including the terminology. Instead of error correction, we now refer to it as a discrimination process.

      (2) L96-98: Does this potentially indicate a trade-off where evolution for stronger binding cannot evolve at the same time as responsiveness to metabolic activity?

      We agree that this sentence made a stronger evolutionary claim than our data support. Since we only tested four laboratory phages, we cannot conclude that there is an evolutionary trade-off between stronger binding and responsiveness to host metabolic activity. We have therefore removed this sentence to avoid making an unsupported evolutionary interpretation.

      (3) L102: What does 'post-cellular' mean?

      Postcellular supernatant is simply the liquid that remains after cells have been removed. During centrifugation, the cells pellet at the bottom, and the liquid above (which can contain viruses) is the postcellular supernatant.

      (4) L105-107: Worth splitting into two sentences as it is a bit unclear if ratios are built between permissible and resistant hosts or between buffers or both.

      Thank you for the suggestion. We have rewritten this section into two sentences to clarify how the ratios are constructed, and we hope the revised wording improves readability.

      (5) L110-122: Figures S1 and S2 could be referenced here.

      References to Figures S1 and S2 have now been added in this section.

      (6) L137: As P(0) is the viral concentration in buffer, I am assuming that the phage lysate has been diluted in buffer and phages have been added to cultures from the same dilution tube to guarantee equal starting numbers, but I couldn't find this in the methods.

      This clarification has been added to the Methods and Media section of the Supplementary Information.

      (7) L243: It would be worth defining what 'hyperdiffusion' means.

      We have added a brief definition of “hyperdiffusion”.

      (8) L253-256: I do not entirely follow this explanation.

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #3. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (9) L284: Why is k_inj necessarily more sensitive to the metabolic state than k_off? Could membrane changes under stress increase k_off?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: reduced commitment in inactive cells can arise through an increase in k_off, a decrease in k_inj, or both, as long as k_off/k_inj becomes larger in the inactive case. What matters is therefore not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also underlies the trade-off now discussed in the manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become much larger in the inactive case. We have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (10) Figure 1: There seems to be more variation between replicates in phage Lambda than in other phages. Is this caused by receptor number heterogeneity in the population?

      Unfortunately we do not have a way to compare receptor number heterogeneity across the different phage receptors in our experiments. We therefore cannot conclude that the larger variation observed for phage λ is caused by receptor number heterogeneity in the population.

      (11) Figure S1: There seems to be a significant difference between phage Lambda viability in the two buffers - do the authors have an idea where this comes from?

      There is no difference in λ viability between the two buffers. The apparent difference in the figure is due to sampling variability.

      (12) Figure S3: Last sentence of the legend probably shouldn't say 'upper'.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      Reviewer #3 (Recommendations for the authors):

      (1) The text reads as incomplete in some places. Can the authors please provide clarifications on the following points?

      (1.1) Lines 235-256: How do the authors draw a conclusion that "a phage can detect host metabolic status" from a study that used purified LamB receptors (i.e., no live cells with any metabolism) extracted from two different bacterial species (i.e., not a difference in metabolic states)?

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #2. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (1.2) Line 270, in the abstract, and in the caption of Figure 4: The authors described the model using terms such as "an error-correction mechanism" or "standard error correction", but there is little explanation. Can the authors clarify what kind of "error" is discussed here, and how it is "corrected"? In the "standard error correction" model, what determines which method of correction is "standard"? If "error correction" is a standard term in phage-bacterial interaction modeling, please provide references.

      We agree with the reviewer that our use of the term error correction was not appropriate in this context. The proper term is discrimination process rather than error correction. We have now corrected this terminology throughout the manuscript and clarified the underlying logic in the relevant sections.

      (1.3) Line 301: The authors speculated that phage T5 is "better suited to ecological niches", but I am not sure how that is consistent with their data showing T5 is more rampant, that they infect both energy-competent and energy-depleted cells, not just depleted cells. Why "niches", and why are T5 better suited to environments "where energy-limited cells dominate", not just any environment?

      We agree that this point was not stated clearly enough. What we intended to convey is that T5 would be at a net disadvantage in a niche containing a mixture of energy-competent and energy-deficient hosts. We have updated the main text accordingly.

      (1.4) Line 303, and related to point 6.3. above: Phage λ can also infect and replicate in "starved bacterial cells" (shown in Kourilsky 1974 and Geng et al. 2024, both of which were cited in this manuscript). How do the authors reconcile these reports with the discussion point in line 303, and their data that only phage T5, but not λ, shows insensitivity to the host metabolic state?

      Our data do not imply that phage λ is unable to infect starved bacteria. As shown in Kourilsky (1974) and Geng et al. (2024), λ can indeed infect and replicate in nutrient-limited cells. Our results specifically indicate that λ infection under starvation proceeds with a reduced adsorption rate, while T5 maintains the same adsorption rate even when the host is starved. Thus, our conclusion is that T5 is insensitive to the host metabolic state at the level of adsorption, whereas λ is not. We acknowledge that the wording in line 303 may have unintentionally led to confusion, and we have revised this part of the text to avoid that.

      (2) The following comments relate to the text and figures in the manuscript. There are many places in the manuscript that could use fine proofreading and copy-editing for clarity and consistency. For example:

      (2.1) If I understand it correctly, the equation in between lines 109 and 110 should be clarified using terms such as "Free viral particles after mixing with bacteria in Arsenate and Azide" and "Free viral particles in bacteria-free buffer with Arsenate and Azide". As it stands, it is not clear which terms correspond to conditions where bacteria are present.

      The equation has been updated to explicitly indicate which terms refer to mixtures containing bacteria and which refer to bacteria-free controls, so that the correspondence between conditions is now clear.

      (2.2) Equations in between line 276 and 283, and elsewhere: Some concentration terms are enclosed in brackets ("[BP]"), while most are not.

      This notation has been clarified. We now use “[PB]” specifically to denote the transient phage–bacterium complex, distinguishing it from the product P⋅B. All other concentration terms are written without brackets for consistency.

      (2.3) Figure 4 and in equations: "BP" or "PB"?

      The notation has been made consistent throughout; we now use “PB” exclusively to denote the phage–bacterium complex.

      (2.4) Line 284 and line 286: The "inj" in "k_inj" is sometimes italicized, sometimes not.

      The notation has been standardized so that k_inj is now formatted consistently throughout the manuscript, without italicizing “inj.” Also we have replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (2.5) Figure 5: Was the step "Estimate cell numbers from OD" really performed on the next day after the experiment (i.e., >12 hours after infection and phage plating), not immediately after cell washing?

      Thank you for pointing this out. The figure has been updated to reflect the correct order of steps: cell numbers are estimated from OD immediately after washing, followed by overnight incubation of the plates.

      (2.6) Figure S1: As it stands now, the x-axis of each panel can be read either as "Permissive, Resistant bacteria, Buffer" (missing "bacteria" for the first pair of bars), or "Permissive (bacteria), Resistant (bacteria), Buffer (bacteria)" (extra "bacteria" for the last pair of bars).

      The intended interpretation is the second one (permissive bacteria, resistant bacteria, buffer).

      (2.7) Figure S3: The panel letters "A" and "B" are missing in the figure. Also, it is not clear why the legend for the five phages and the legend for the measurement times are not combined.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      (2.8) Strain table in the Methods and Materials: Please write genotypes with italicization, and consistently indicate mutations and deletions with the minus sign superscript or the Δ prefix. Also, for the S3222 strain: Is it really the entire Mal regulon mutated ("Mal-"), or just lamB-? In Brown et al. 2022, it was only the latter.

      Genotypes have been reformatted with consistent notation. For S3222, the correct designation is Mal-, as in the SI of Brown et al. 2022. In this case, Mal- is intended as a phenotypic designation rather than a specific genotype, and we have therefore formatted it accordingly, i.e. neither italicized nor written in lower case.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      1. __ General Statements__ We thank the reviewers for their thoughtful and constructive evaluations of our work. We are particularly encouraged that both recognize the value of this study as a scalable and systematic framework for the functional exploration of the human KZFP family and agree that the resource generated here will be of broad interest to the KZFP, transposable element, and genome regulation communities. Reviewer 1 explicitly notes that "the screening framework itself represents a potentially useful resource for prioritizing candidate KZFPs for downstream study" and that "the study may nonetheless serve as a useful starting point for future investigations into KZFP biology and transcriptional regulation." Reviewer 2 similarly emphasizes that "the authors provide an efficient and valuable screening platform that can identify promising candidates for further investigation" and that "the methodological advance represents the primary contribution of the work."

      We can only concur with these assessments. The principal goal of this study was not to elucidate the physiological roles of all or even a subset of individual KZFPs, but rather to provide a scalable framework that enables their systematic prioritization and generates experimentally testable hypotheses regarding their functions. To support our argument, we ventured into some mechanistic analyses, but these could not pretend to be complete and definitive. In that respect, we hear the reviewers when they note that the original manuscript does not always sufficiently distinguish candidate discovery from mechanistic validation. In its revised version, we will therefore more clearly frame the inducible K562 overexpression assay as a standardized and sensitive readout of regulatory potency rather than as a direct surrogate of physiological function. Within this framework, K562 fitness defects are interpreted as a quantitative measure of the extent to which ectopic KZFP expression perturbs transcriptional homeostasis in a controlled cellular context, while the direct targets and transcriptional networks identified through our integrative analyses are presented as hypotheses to be tested in more physiologically relevant systems. Accordingly, the revised manuscript preserves the broad scope and resource aspect of the study while incorporating additional experimental validation, expanded methodological descriptions, and a more cautious interpretation of the proposed biological functions of the selected KZFPs.

      __Although this document is submitted as a Revision Plan, we have already incorporated a substantial number of revisions into the transferred manuscript. In particular, we have implemented most of the presentation, methodological, and conceptual modifications requested by the reviewers, including clarification of the scope of the study, extensive revisions of the Results and Discussion, expanded Materials and Methods, and numerous figure and text corrections. These revisions are detailed in Section 3 ("Description of the revisions that have already been incorporated into the transferred manuscript"). __

      The remaining points requiring additional experimentation or more extensive analyses are described in Section 2 ("Description of the planned revisions").

      __ Description of the planned revisions__

      Reviewer 1 Major comment 1

      “Finally, several aspects of the data presentation are currently difficult to reconcile. In Fig. 1D, the meaning of the purple category is unclear, and the percentage scaling on the x-axis is difficult to reconcile with the cumulative values displayed. For instance, the sum of all the bars would not reach 100%, as the values of the bars span percentages up to 4% at most (for 105 MYO KZFPs) according to this plot. Similarly, the reported numbers of TE-binding KZFPs in Fig. 1E-F and Fig. S1D appear internally inconsistent and should be clarified. Specifically, 53+14=67 KZFPs are reported to bind TEs in total, yet a larger number of KZFPs appears associated with individual TE families (e.g., 86 for LTR.ERV1). If the values shown correspond to percentages rather than absolute counts, this should be explicitly clarified in both the figure and legend. In addition, Fig. S1D appears inconsistent with the counts reported in Fig. 1E-F, as only 5 out of the 53 toxic KZFPs displayed in the plot show no enrichment for any of the highlighted TE families.”

      We thank the reviewer for this insightful comment, which has helped us identify areas where the presentation of our data can be substantially improved. We agree that the current presentation of the TE-binding analyses could be clearer and that revising these figures will improve both their readability and the overall consistency of the manuscript. In the revised manuscript, we will clarify the apparent inconsistencies in the presentation of the TE-binding KZFP analyses and revise the corresponding figures and legends accordingly. Importantly, these inconsistencies do not arise from errors in the underlying data but rather from an insufficient explanation of the statistical enrichment analyses and the way the results are represented. We will therefore redesign the relevant figures and expand their legends to more clearly describe the analytical approach, the enrichment criteria, and the interpretation of the results. We believe that these revisions will improve the clarity, transparency, and internal consistency of the manuscript, allowing readers to more readily interpret the TE-binding analyses. Minor comments of the reviewer 1 were extremely useful to detect mistakes and we are grateful for that. All the modifications that were asked see below were included in the manuscript.

      Reviewer 1 Major comment 2

      “Finally, while the proteomics results aimed at identifying SCAN-dependent interactors are of interest, several aspects of the experimental design and data analysis remain unclear. In particular, it is not specified whether the experiment was performed in biological replicates or as a single measurement. This is important, as it directly affects how the data can be interpreted and how stringent downstream filtering can be. In the Results section, the authors state that "we identified a set of SCAN-dependent interactors, i.e., proteins that co-immunoprecipitated with the full-length construct but were absent in controls and lost upon deletion of the SCAN domain," which suggests a relatively binary, "presence/absence" filtering strategy. However, this description does not specify whether any quantitative threshold (e.g., enrichment ratio) was applied when comparing full-length constructs to deletion mutants. In contrast, the Methods section states that "proteins lacking signal above background were excluded and proteins were additionally required to show stronger signal in at least one bait condition than in GFP controls based on heatmap clustering (see script)," which instead suggests that a threshold-based criterion was used to define enrichment relative to controls and deletion mutants. If this is the case, the exact criteria and thresholds used for filtering should be clearly stated and consistently reported between the Results and Methods sections. If replicate measurements were not performed, this should be explicitly acknowledged, as peptide-level variability may substantially influence the identification of high-confidence interactors, particularly if the applied cutoffs are not highly stringent.”

      We agree that a more detailed description of the experimental design and analysis strategy, together with additional validation, will strengthen the interpretation of the proteomic data. In the revised manuscript, we expanded the Results and Materials and Methods sections to provide a clearer and more quantitative description of the filtering strategy, including the enrichment criteria and thresholds used to define SCAN-dependent interactors. To further strengthen these findings, we propose to perform an independent biological replicate of the co-immunoprecipitation mass spectrometry experiment. This additional experiment will increase confidence in the identified SCAN-dependent interactors and further support the conclusions drawn from the proteomic analysis.

      Reviewer 1 Minor comments

      • In Fig. 2A, readability could be improved by adjusting the layering of points, as the darker dots (in particular the red ones) are currently obscured by lighter ones. Alternatively, removing the outline of the points (which is not transparent) may also improve visibility, but in that case the legend for point size would need to be updated accordingly.

      Thank you for this helpful suggestion. We will revise Figure 2A to improve its readability by reworking the layering of the points in accordance with the reviewer's recommendation. We will also evaluate the point outlines and, if appropriate, remove them and update the point-size legend accordingly to ensure the figure is clear and easy to interpret.

      Reviewer 2 – Major comment

      “- The authors looked at available chromatin data in either K562 cells or HEK293 cells, which I think is a very good way of utilizing publicly available data. Since the authors showed that different KZFPs might be functionally relevant in different cell types/tissues, I was wondering if they checked if there is available ChIP Seq or CUT&RUN data in those specific cell types/tissues. If yes, that data should be included in the manuscript.”

      We agree that integrating KZFP binding data generated in biologically relevant cell types or tissues would further strengthen the proposed regulatory models. As described in the revised manuscript, we have already adopted this approach for ZNF43 by integrating chromatin landscape data from thymus and liver, where suitable datasets were available.

      To further address this point, we propose to systematically explore publicly available ChIP-seq, CUT&RUN, CUT&Tag, and related chromatin profiling datasets for the other KZFPs investigated in this study. Where suitable datasets are available, these analyses will be incorporated into the revised manuscript to further support the proposed tissue-specific regulatory models and provide additional biological context for the identified target genes.

      __ Description of the revisions that have already been incorporated in the transferred manuscript__

      Reviewer 1 Major comment 1

      “The large-scale overexpression screen represents the foundation of the manuscript and provides a potentially valuable resource for prioritizing candidate KZFPs for downstream study. However, several aspects of the experimental setup and data presentation currently limit the interpretation of the reported proliferation defects. First, key details regarding the screening workflow remain unclear. While the Methods section describes the overall procedure, it is difficult to determine when cells were seeded relative to doxycycline induction, in which plate format the cells were maintained throughout the experiment, and whether medium exchange was performed during the 9-day assay. These points are particularly relevant given the use of suspension K562 cells (which can complicate medium exchange in a 96-well plate format and make long-term culture more difficult to control) and a metabolic viability readout (PrestoBlue), as differences in nutrient depletion or overgrowth could also influence the signal independently of reduced proliferation or toxicity. Additional clarification regarding seeding density, timing of induction, plate format, culture handling throughout the assay, and whether cell morphology/density was visually monitored would substantially improve interpretability and reproducibility. Second, it is unclear whether the observed proliferation phenotypes may be influenced by differences in transgene expression levels or integration effects. Were all constructs validated for comparable expression following induction? In the absence of such controls, it remains difficult to determine whether the reported phenotypes reflect specific KZFP activities or differences in overexpression efficiency. While it may not be possible to conclusively distinguish KZFP-specific effects from toxicity associated with high transgene expression levels, this limitation should at least be acknowledged. In addition, the possibility that some phenotypes may be influenced by transgene integration effects should also be considered. Unless independent transductions were validated for the KZFPs classified as toxic, it remains difficult to exclude integration-site-specific contributions to the observed proliferation defects. Third, the normalization strategy would benefit from additional clarification. In Fig. S1A, the LacZ control appears variably affected by doxycycline treatment across plates, whereas the GFP control appears more stable. Since normalization relies on the mean behavior of both controls within each batch and condition, the authors should clarify whether this variability could influence hit calling.”

      We agree that additional methodological details improve the clarity and reproducibility of the screening assay. Accordingly, we substantially expanded the Materials and Methods section to describe the experimental workflow, quality controls, data normalization, and hit-calling criteria. The revised paragraph is reproduced below.

      Arrayed overexpression screen

      To systematically assess the effect of human KZFP overexpression on cellular fitness, K562 cells were individually transduced with doxycycline-inducible lentiviral vectors encoding 366 human KZFPs. Lentiviral particles were produced as described above and used to transduce cells without MOI calculation. __Instead, a fixed volume of viral supernatant (200µL per × 104 cells in 48 well plate filled with 200ul of RPMI) was used for all transductions to ensure comparable experimental conditions. Transduced cells were selected with puromycin before doxycycline induction. Following puromycin selection (1µg/mL for 3 days), cells were seeded at 20 000 cells per well in 24-well plates filled with 1ml of medium in technical triplicate for each KZFP. Following puromycin selection and prior to doxycycline induction, cell survival was visually assessed as a quality control metric for each KZFP construct (Supp __Table 2____). Doxycycline (1µg/mL) was added immediately after cell seeding to induce expression of the HA-tagged KZFPs. At each time point, metabolic activity was measured using PrestoBlue™ reagent according to the manufacturer's instructions (10µL reagent added to 100µL culture medium, incubated for 3h in a 96 plates). Absorbance was recorded at 570 nm and 600 nm using a plate reader (Hidex Sense Microplate Reader), GFP- and LacZ-expressing control wells were included on every plate to account for plate-to-plate and batch-to-batch variability. Peripheral wells were filled with culture medium to minimize evaporation-induced edge effects. Cells were maintained in RPMI supplemented with 10% fetal bovine serum (FBS) and 1× penicillin–streptomycin, and splited (1/10) with aspiration of the surface medium every three days throughout the assay while maintaining doxycycline at 1µg/mL. Cell proliferation was assessed after 4, 7, and 9 days of induction. and the A570/A600 ratio was used as a surrogate measure of viable cell number and proliferative capacity. For computational normalization, raw A570/A600 values were first background-corrected by subtracting the signal from medium-only controls and then normalized in two steps. First, each value was divided by the mean signal obtained from the GFP and LacZ control wells from the corresponding batch and induction condition to correct for inter-batch variability. Second, the resulting value was normalized to the corresponding −Dox condition for the same KZFP and time point to correct for seeding variability, yielding a relative proliferation score that reflects the effect of KZFP induction. KZFPs with a normalized proliferation score ≤ 0.85 at day 9 were arbitrarily classified as proliferation-impairing hits in this screening framework.

      After doxycycline induction, dot blot analysis using anti-HA and anti-actin antibodies was systematically performed to assess KZFP expression and sample loading, respectively (Supplementary DotBlot.pdf). The HA signal following doxycycline induction (HA_Dox) and actin signal following doxycycline induction (Actin_Dox) were visually scored from the dot blot signals (__Supp __Table 2).

      In addition, to strengthen the methodological description and address these concerns more directly, we will:

      1/ Include a supplementary table summarizing our experimental observations for each individual KZFP throughout the screening process (See preliminary Supp Table 2). -> See header here:

      2/ Perform and include Dot Blot analyses, to assess and compare transgene expression levels across KZFP constructs. (Supplementary File DotBlot.pdf____). Generation of these files is in progress, with a few missing dot blots still being completed (we have done 303 over 366 already). However, preliminary versions have already been submitted. -> See header of the .pdf here:

      In addition, we agree that a more explicit discussion of the limitations of our screening approach improves the interpretation of our findings. Accordingly, we expanded the Discussion to address the limitations associated with variable transgene integration, heterogeneous transgene expression, potential toxicity due to ectopic KZFP overexpression, and the use of K562 cells as a standardized rather than physiological cellular model.

      “Several methodological considerations should be taken into account when interpreting these results. As with any lentiviral overexpression screen, three potential sources of technical variability may influence the observed phenotypes: differences in transgene integration sites, heterogeneity in transgene expression levels, and non-specific toxicity resulting from ectopic overexpression. Variable integration sites are unlikely to represent a major source of bias in the present study because all analyses were performed on polyclonal populations of transduced cells rather than individual clones, thereby averaging integration-site effects across many independent events. In contrast, heterogeneity in transgene expression levels is expected, as the abundance of each KZFP depends not only on transduction efficiency but also on intrinsic differences in mRNA stability, translational efficiency, and protein stability. To minimize these sources of variability, all constructs underwent systematic quality control, including assessment of cell survival following puromycin selection and evaluation of transgene expression by HA dot blot after doxycycline induction. Although transgene expression levels varied across KZFPs (Supplementary File DotBlot.pdf), this variability showed no systematic relationship with the proliferation phenotypes, suggesting that differences in overexpression efficiency are unlikely to be the primary determinant of toxicity. Nevertheless, ectopic expression exposes cells to supraphysiological concentrations of KZFPs capable of generating non-physiological interactions or regulatory effects. Therefore, while the screening strategy is well suited for identifying candidate functional regulators, independent validation under endogenous expression conditions remains essential to confirm KZFP-specific functions.”

      Reviewer 1 Major comment 2:

      “A central conceptual issue throughout the manuscript is that the downstream functional analyses of the selected KZFPs remain largely disconnected from the original screening phenotype. The four candidates were prioritized based on proliferation defects observed upon overexpression in K562 cells; however, the subsequent analyses (with the only exception being a more in-depth experimental analysis of ZNF498 in ciliogenesis, which stands out as comparatively more directly supported by experimental evidence) primarily rely on correlative expression patterns and KZFP ChIP-seq datasets to infer potential biological functions in unrelated cellular contexts. As a result, it remains unclear whether the proposed transcriptional programs are mechanistically linked to the proliferation phenotypes that motivated candidate selection in the first place. This issue is evident across multiple sections of the manuscript. For example, the proposed role of ZNF43 in regulating fatty acid metabolism and detoxification pathways is primarily inferred from tissue-level expression correlations. While these analyses focus on genes identified as potential ZNF43 targets, the underlying ChIP-seq datasets were themselves generated under ZNF43 overexpression conditions. Therefore, the current analyses do not establish whether ZNF43 regulates these pathways under physiological expression levels or within a relevant cellular context, nor how such regulation relates to the proliferation defect observed in K562 cells. Moreover, several proposed target genes remain substantially expressed in tissues where ZNF43 expression is not particularly low (e.g., kidney and heart muscle), suggesting that additional regulators are likely involved. Similarly, the proposed model of ZNF257-mediated regulation of MAGEA genes during spermatogenesis is intriguing but does not fully account for the expression behavior of all MAGEA family members, particularly MAGEA2B, which displays strong expression in spermatocytes despite high ZNF257 expression. This expression pattern should be acknowledged in the main text and reflected in Fig. 3K. In addition, the labels for MAGEA6 and MAGEA2B in Fig. 3C appear to be inverted. More broadly, the proposed regulatory model is difficult to reconcile with the generally restricted expression pattern of MAGEA genes across adult tissues, as their expression does not appear to consistently correlate with ZNF257 levels outside the germline context. Related concerns also apply to the analyses of ZNF498 and ZNF18, where the proposed functions in cilium formation and sperm maturation remain disconnected from the proliferation defects identified in the initial screen.”

      We agree that this comment raises an important conceptual point and has helped us clarify the scope of the study and the interpretation of our findings. In the revised manuscript, we explicitly distinguish hypothesis generation from mechanistic validation by clarifying that the proliferation phenotype observed in K562 cells reflects the regulatory potential of ectopically expressed KZFPs rather than their physiological functions. We also adopted a more cautious interpretation of the functional analyses, emphasizing that the proposed regulatory networks are hypothesis-generating and that individual KZFPs are unlikely to act as sole regulators. More broadly, we emphasize that the primary objective of this study is to establish a scalable screening platform for prioritizing KZFPs and identifying biologically relevant contexts for future investigation, rather than to provide a comprehensive functional characterization of individual KZFPs. We agree that this comment highlights an important limitation of our proposed regulatory model. In the revised manuscript, we adopted a more nuanced interpretation by presenting ZNF257 as a contributor to, rather than the sole regulator of, the MAGEA transcriptional program, and by explicitly discussing the exceptions identified by the reviewer.

      Modification in the revised manuscript:

      1/

      “Integrative transcriptomic, chromatin and proteomic analyses reveal diverse mechanisms, including transposable element–linked repression (ZNF43), promoter-proximal regulation (ZNF257), and SCAN domain–dependent transcriptional activation (ZNF498/ZSCAN25 and ZNF18).”

      Is now:

      “Integrative transcriptomic, chromatin and proteomic analyses identify distinct regulatory properties and generate testable hypotheses regarding diverse mechanisms, including transposable element-associated repression (ZNF43), promoter-proximal regulation (ZNF257), and SCAN domain-dependent transcriptional activation (ZNF498/ZSCAN25 and ZNF18).”

      2/

      “Detailed follow-up of four such candidates, ZNF43, ZNF257, ZNF498 and ZNF18, revealed as hypothesized distinct modes of action, ranging from TE-linked transcriptional repression to promoter-proximal gene silencing and SCAN domain-mediated transcriptional activation. These findings reinforce the view that KZFPs, while often viewed as a homogeneous family of TE-repressive TFs, are rather functionally diverse regulators with wide-ranging impacts on human biology.”

      Is now:

      “Detailed follow-up of four such candidates, ZNF43, ZNF257, ZNF498 and ZNF18, identified distinct regulatory properties and generated hypotheses regarding their physiological functions. By integrating overexpression-induced transcriptional responses, chromatin occupancy, proteomic analyses and tissue-specific expression data, we propose candidate biological contexts in which these KZFPs may operate. These hypotheses now provide a framework for future mechanistic studies performed under physiological conditions. Together, these findings reinforce the view that KZFPs, while often viewed as a homogeneous family of TE-repressive transcription factors, comprise functionally diverse regulators with broad potential roles in human biology.”

      3/

      “We conclude from these data that ZNF43 regulates a transcriptional program related to fatty acid metabolism and detoxification, allowing for the preferential expression of its effectors in the liver (Fig. 2G). Interestingly, neither expression nor chromatin state followed the same pattern at the functionally unrelated DNAI4 locus, indicating that this gene is subjected to other dominant regulators.”

      Is now:

      “Together, these observations identify a small set of candidates ZNF43 target genes involved in fatty acid metabolism and detoxification and suggest that ZNF43 may contribute to the regulation of these transcriptional programmes in appropriate physiological contexts (Fig. 2G). However, these conclusions are derived from overexpression-based datasets and tissue-level expression analyses and should therefore be considered hypothesis-generating. Interestingly, neither expression nor chromatin state followed the same pattern at the functionally unrelated DNAI4 locus, indicating that additional regulatory mechanisms contribute to the control of these genes.”

      4/

      “It strongly suggests that ZNF257 contributes to initiating the transcriptional repression of these two MAGEA genes during early spermiogenesis, after which their silencing may be stabilized through stable epigenetic mechanisms such as DNA methylation.”

      Is now:

      “These observations suggest that ZNF257 may contribute to the initiation of transcriptional repression of a subset of MAGEA genes during the spermatogonia-to-spermatocyte transition, after which their silencing may be stabilized through epigenetic mechanisms such as DNA methylation.”

      5/

      “Together, these results identify ZNF498 as a transcriptional activator of gene modules controlling cytoskeleton-dependent processes and suggest that this TF may act as a regulator of neuronal cytoskeletal architecture, warranting investigation in relevant neural models.”

      Is now:

      “Together, these results indicate that ZNF498 functions as a transcriptional activator in our overexpression system and support the hypothesis that it contributes to transcriptional programmes controlling cytoskeleton-dependent processes in physiologically relevant neural contexts, warranting further investigation in dedicated neural models.”

      6/

      “The co-expression of ZNF18 and its target genes at the spermatid stage suggests that ZNF18 activates a transcriptional program supporting these processes.”

      Is now:

      “The co-expression of ZNF18 and its candidate target genes at the spermatid stage is consistent with the hypothesis that ZNF18 contributes to transcriptional programmes supporting these processes.”

      7/

      “The four KZFPs characterised here illustrate this diversity. ZNF43 represses a coherent set of genes involved in fatty acid metabolism and detoxification through binding to nearby LTR/ERV1 integrants, with its expression anticorrelating that of its targets: i.e., highly expressed in thymus and bone marrow, where these metabolic genes are silent, and lowly expressed in liver, where they are most active. This represents a clear example of host genomes coopting TE-derived sequences and shaping their regulatory activities in a cell-type specific manner by the differential expression of KZFPs. ZNF257, by contrast, acts as a promoter-proximal repressor whose targets show accelerated sequence evolution at their promoters, consistent with integration into a KZFP-orchestrated GRN through rapid promoter diversification, a feature previously described for KZFPs (Farmiloe et al., 2023). Its regulation of the MAGEA gene cluster exemplifies a distinct evolutionary mechanism: an ancestral intronic binding site, present in MAGEA6 gene body, before ZNF257 emerged, was propagated across the cluster through tandem duplication, enabling coordinated regulation of multiple paralogs. Temporal expression analysis during spermatogenesis further suggests that ZNF257 initiates MAGEA repression at the spermatogonia-to-spermatocyte transition, after which silencing may be maintained through epigenetic mechanisms such as DNA methylation. ZNF498 and ZNF18, both SCAN-containing KZFPs with variant KRAB domains, on the other hand acted as transcriptional activators. ZNF498 activates a programme centred on microtubule cytoskeleton organisation, as demonstrated by the disruption of ciliogenesis upon its overexpression, and both ZNF498 and its targets are broadly expressed in the central nervous system, particularly in excitatory neurons where microtubule dynamics are essential for axonal architecture. ZNF18 similarly activates genes involved in chromatin remodelling and cytoskeletal reorganisation at the spermatid stage, processes that are hallmarks of spermiogenesis. Together, these case studies demonstrate that even within a single screen, KZFPs with fundamentally different regulatory logics can be identified through a single unifying phenotype and then mechanistically dissected to uncover their unique properties.”

      Is now:

      “The four KZFPs characterized here illustrate the functional diversity that can be uncovered using this screening strategy. For ZNF43, integration of overexpression transcriptomics with ChIP-exo binding data identified a small set of candidate direct target genes located near LTR/ERV1 elements. Their tissue-specific expression patterns are consistent with the hypothesis that ZNF43 contributes to transcriptional programmes associated with fatty acid metabolism and detoxification, although these analyses, which rely on overexpression-derived datasets and tissue-wide correlations, do not establish physiological regulation or causality. Rather, they identify a candidate regulatory network whose functional relevance will require investigation in appropriate biological models. More generally, these observations support the concept that host genomes may exploit TE-derived regulatory sequences in a tissue-specific manner through differential KZFP expression, while recognizing that additional transcription factors almost certainly participate in controlling these gene expression programmes. Similarly, ZNF257 emerged as a promoter-associated transcriptional repressor in our overexpression system. Evolutionary analyses suggest that tandem duplication propagated an ancestral ZNF257-binding sequence across the MAGEA locus, generating the hypothesis that ZNF257 may contribute to coordinated regulation of this gene cluster during spermatogenesis. The temporal expression profiles of ZNF257 and the MAGEA genes are compatible with such a model but remain correlative and therefore require direct functional validation. ZNF498 and ZNF18, two SCAN-containing KZFPs with variant KRAB domains, displayed transcriptional activation rather than repression following overexpression. For ZNF498, the integration of transcriptomic analyses with expression profiling pointed to microtubule cytoskeleton organization as a candidate biological process, a prediction that was further supported experimentally by the marked impairment of ciliogenesis following ZNF498 overexpression in hTERT-RPE1 cells. This represents the strongest functional validation presented in this study and supports the biological relevance of the analytical framework developed here. For ZNF18, the co-expression of the KZFP and its candidate target genes during spermatogenesis is consistent with the hypothesis that it contributes to transcriptional programmes involved in chromatin remodelling and cytoskeletal reorganization during spermatid differentiation. Together, these case studies illustrate how a standardized overexpression screen can identify KZFPs with distinct regulatory properties and generate biologically coherent hypotheses regarding their physiological functions. Rather than establishing definitive functions for individual KZFPs, this framework prioritizes candidates, proposes relevant cellular contexts, and provides a foundation for future mechanistic studies performed under physiological conditions.”

      “In addition, interpretation of the SCAN-deletion experiments is complicated by the reduced expression levels of the deletion constructs relative to the corresponding full-length proteins, making it difficult to determine whether the observed proliferation phenotypes are pathway-specific or partially driven by differential expression.”

      We thank the reviewer for this important observation and agree that differences in expression levels between the full-length and ΔSCAN constructs could complicate the interpretation of the observed phenotypes. To address this concern, we performed a quantitative comparison of the expression levels of full-length and ΔSCAN proteins using both western blotting and transgene expression using RNAseq, while accounting for differences in transgene length. This result are now added in (Fig S6C, D).

      With modification of the legend:

      • HA signal after OE of HA-tagged ZNF18, ZNF18∆SCAN, ZNF498, ZNF498∆SCAN or GFP in K562 cells. Actin as control.
      • Quantification of ZNF18, ZNF18∆SCAN, ZNF498, ZNF498∆SCAN It appears that the difference is small (Minor comments of the reviewer 1

      “- In the Abstract and in the "Limitations of the study" section, the term "annotation" is used. It would be preferable to specify "functional characterization" instead of "annotation".

      Done as suggested by the reviewer.

      • In the Introduction, there may be a minor citation confusion. Following the sentence: "Characterized by an N-terminal KRAB domain and a C-terminal tandem array of C2H2 zinc fingers, KZFPs primarily target transposable element (TE)-embedded sequences," the cited references are predominantly experimental studies supporting this statement. However, the inclusion of the review "Bruno, Mahgoub and Macfarlan, 2019" appears less appropriate in this context, as it does not directly present ChIP-seq data supporting this claim. More relevant primary studies from the same research area include "Wolf et al. 2020" and "Bruno et al. 2025.".

      Done as suggested by the reviewer.

      • In Fig. 1A, "D10" appears inconsistent with the text and other figures (Fig. 1B, 1G, 1H), which refer to 9 days post-induction.

      Done as suggested by the reviewer.

      • In Fig. S1, there may be a mismatch in the highlighted plate: the zoomed image appears to correspond to the first plate from the top. The correct plate should be highlighted for consistency.

      Done as suggested by the reviewer.

      • In Fig. 1B, there is a typographical error ("K ZFPs" instead of "KZFPs").

      Done as suggested by the reviewer.

      • In Fig. S1E, it is unclear what "other" refers to. Please clarify whether this represents the mean of all remaining KZFPs or a defined subset, ideally in the figure description.

      Done as suggested by the reviewer.

      • In Fig. S2E, "SetDB1" should be corrected to "SETDB1".

      Done as suggested by the reviewer.

      • In Fig. 3B, it is unclear what distinguishes the upper and lower "Diverse REs". A brief clarification in the figure legend would improve interpretability, particularly regarding the transposable element families included.

      Done as suggested by the reviewer.

      • In Fig. S3C, the x-axis labels appear slightly misaligned and shifted to the right.

      Done as suggested by the reviewer.

      • In Fig. 3C, the labels for MAGEA6 and MAGEA2B appear to be inverted.

      Done as suggested by the reviewer.

      • In Fig. 3K, "MAGE3" should be corrected to "MAGEA3".

      Done as suggested by the reviewer.

      • In the ZNF498 section, line 4, the punctuation should be corrected so that the period appears after the figure reference ("promoters (Fig. S1E).").

      Done as suggested by the reviewer.

      • In the final sentence of the ZNF498 section, a noun appears to be missing after "cytoskeleton-dependent," possibly "processes".

      Done as suggested by the reviewer.

      • In the last section of the Results and corresponding figures and their descriptions, "SCAN dependant" should be corrected to "SCAN-dependent".”

      Done as suggested by the reviewer.

      Major comments of the reviewer 2

      “- The authors chose four KZFPs to study in detail, but why they chose these 4 candidates is unlcear to me. It would be nice to add a more detailed description of the process by which they chose the four candidates.”

      We agree that the rationale for selecting the four KZFPs should be presented more explicitly. Accordingly, we revised the manuscript to clarify the selection criteria.

      “However, a modest correlation was noted between the number of transcription start sites (TSS) bound by KZFPs and the drop in PrestoBlue signal induced by their overexpression (Fig. 1G), and SCAN-containing KZFPs (SKZFPs) tended to induce proliferation defects more frequently than family members lacking this domain (Fig. 1H).”

      Is now:

      “However, a modest correlation was noted between the number of transcription start sites (TSS) bound by KZFPs and the drop in PrestoBlue signal induced by their overexpression (Fig. 1G), and SCAN-containing KZFPs (SKZFPs) tended to induce proliferation defects more frequently than family members lacking this domain (Fig. 1H). These observations indicated that KZFPs affecting proliferation do not constitute a homogeneous functional group, prompting us to select representative candidates spanning the evolutionary, structural, and genomic diversity of the KZFP family for mechanistic characterization.____”

      “- The materials and methods part of the manuscript is not detailed enough for other researchers to reproduce the study. They should add more details to both experiments and data analysis part of this section. Below I highlight some examples for sake of clarity, but the authors should revise the whole materials and methods section and add more details keeping these examples in mind:

      • The authors do not state the titer of lentiviral vectors they generate nor the MOI or amount of virus they use to transduce the cells

      • In many cases, the specific softwares and the software version is not stated e.g., the analysis of the Gene Ontology Biological Processes

      • It would be beneficial for the readers to get more details about the construct they used, for example a map of the plasmid.

      • It is unclear how many cells were used for RNA extraction

      • It is unclear which microscopes were used for imaging.

      • The concentration of antibodies used for staining and the product number, and provider of the antibody is not always depicted.”

      We agree that the additional methodological details requested by the reviewer will improve the reproducibility and transparency of the study. Accordingly, we have expanded the Methods section to provide a more detailed description of the experimental procedures and data analysis workflow.

      “Lentiviral particles were produced in HEK293T cells by transient co-transfection of transfer, packaging and envelope plasmids. Cells were transfected at approximately 70–80% confluence using a standard lipid-based transfection reagent. Viral supernatants were collected 48 h after transfection, cleared by centrifugation, filtered through 0.22-µm membranes, and used fresh or stored appropriately until use. Recipient K562 or hTERT-RPE1 cells were transduced under conditions optimized for efficient gene delivery.”

      Is now:

      “Lentiviral particles were produced in HEK293T cells. 105 cells were seeded in 24 well plates filled with 1ml DMEM the day before transfection. Cells were co-transfected individually with 0.15ug of each plasmids encoding KZFPs tagged with HA (pTRE-KZFPX-HA-PGK-puro), 0.1ug of the packaging plasmid (pR8.74) and 0.07ug of the envelope plasmid (pMD2G) using TransIT®-LT1 Transfection Reagent (MIR 2306), according to the manufacturer's instructions. Viral supernatants were harvested 24h after transfection, clarified by centrifugation, filtered through 0.45-µm filters and used immediately.”

      “Coding sequences were cloned into doxycycline-inducible lentiviral transfer vectors designed to express N-terminally HA-tagged proteins.”

      Is now:

      “Coding sequences corresponding to 366 human KZFP open reading frames were codon-optimized for human expression and cloned into doxycycline-inducible lentiviral transfer vectors expressing C-terminal HA-tagged proteins under the control of a tetracycline-responsive promoter pTRE-KZFPX-HA-PGK-puro. All expression constructs used in the primary overexpression screen have been deposited and are publicly available (De Tribolet et al., 2023). A schematic representation of the lentiviral expression cassette, including the promoter, HA tag, cloning site, antibiotic resistance cassette, and regulatory elements, is provided in Supplementary file. Selected constructs encoding ZNF43, ZNF257, ZNF498 and ZNF18 were used for follow-up mechanistic studies. For SCAN-domain functional analyses, deletion constructs lacking the SCAN domain (ΔSCAN) were generated for ZNF18 and ZNF498 in the same lentiviral backbone. Deletion were done using In-Fusion cloning with specific primers. PCR was performed with high-fidelity polymerase, followed by gel purification and recombination with the linearized plasmid using the In-Fusion HD Cloning Kit (Takara Bio©) according to the manufacturer’s protocol. The product was transformed into HB101 Escherichia coli cells, and colonies were screened by PCR. Positive clones were verified by Sanger sequencing, and confirmed plasmids were propagated and purified for further use.”

      “Total RNA was extracted...”

      Is now:

      “For each biological replicate, approximately 1 × 10⁶ K562 cells were harvested 72 h after doxycycline induction. Total RNA was extracted…”

      “Images were acquired by fluorescence microscopy under identical conditions across samples.”

      Is now:

      “Images were acquired using a confocal microscope Leica-SP8 (Leica Biosystems) with an objective HC PL APO 63x/1.40 and a pinhole size of 1 AU, using identical acquisition settings for all conditions. Images were processed using Fiji/ImageJ (version 2.9.0) without nonlinear intensity adjustments.”

      “Cells were fixed and stained with antibodies against ciliary markers (ARL13B)”

      Is now:

      “Cells were fixed in 4% paraformaldehyde, permeabilized with 0.1% Triton X-100, blocked with 2% BSA, and incubated with rabbit anti-ARL13B (Proteintech, Cat. No. 17711-1-AP, 1:200) followed by Alexa Fluor 568-conjugated donkey anti-rabbit IgG (Thermo Fisher Scientific, Cat. No. A-10042, 1:1000). Nuclei were stained with Hoechst (1 µg/mL).”

      “- The authors mention that KZFPs are usually expressed at a low level in the K562 cell line they use, but there is no figure showing the expression level of KZFPs in this cell type. It would be important to see the baseline KZFP expression in these cells, the level of overexpression and compare it to the endogenous expression levels they show in different cell types/tissues, at least for the four candidates studied more in depth. This would help to understand whether this level of activity is something that could occur naturally in a physiologically relevant context.”

      We thank the reviewer for this insightful suggestion and fully agree that providing additional context regarding endogenous and ectopic KZFP expression levels will help readers better assess the physiological relevance of our findings. As suggested, we included data showing the baseline expression levels of the four selected KZFPs in K562 cells together with the expression levels achieved following doxycycline-induced overexpression. We also compared these values with publicly available transcriptomic data from cell lines. Importantly, only cell lines are assessed as we need ground through (K562) to estimate transgene expression. We modified Fig. S2, Fig. S3, Fig. S4 and Fig. S5 to add the results of these analysis. Here is ZNF43 as an example:

      With the following legend:

      “(C) Distribution of endogenous expression levels, (using GFP control cells), of all expressed genes (light grey) and all KZFPs (dark grey) in K562 cells. The solid red line indicates endogenous ZNF43 expression in GFP control cells, whereas the dashed red line indicates the corrected transgene expression following doxycycline induction.

      (D) Endogenous ZNF43 expression across Human Protein Atlas cell lines, (https://www.proteinatlas.org/about/download#cell_line), following normalization to the local RNA-seq dataset. K562 cells are highlighted in red. The dashed red line indicates the corrected transgene level measured following doxycycline-induced overexpression in K562 cells overexpressing ZNF43.”

      Modified the result section:

      “ZNF43 is a ~43-million-year-old KZFP with a canonical TRIM28-recruiting KRAB domain and 19 zinc fingers that preferentially recognize an LTR/ERV1-embedded sequence (Fig. S1F). We first verified that ZNF43 overexpression impaired the growth of K562 cells (Fig. S2A, B). Endogenous ZNF43 expression was readily detectable in K562 cells and across human cell lines (Fig. S2C, D). Following doxycycline induction, transcript abundance markedly increased and exceeded the highest endogenous expression level observed among the analyzed cell lines (Fig. S2C, D).”

      We also updated the Methods section:

      Quantification of endogenous and transgene expression levels

      Endogenous KZFP expression in K562 cells was estimated from GFP control RNA-seq samples using normalized mean expression values obtained from the differential expression analyses. For ZNF18, whose transgene sequence is identical to the endogenous coding sequence (i.e., not codon-optimized), transgene-derived expression was estimated directly by subtracting the endogenous transcript abundance measured in GFP controls from the total transcript abundance measured following doxycycline induction (OE − GFP). For ZNF43, ZNF257 and ZNF498, the overexpression constructs were synthesized using codon-optimized coding sequences. RNA-seq reads were therefore additionally aligned against the codon-optimized transgene reference sequences to specifically quantify exogenous transcripts without interference from endogenous reads. Because these codon-specific counts are generated through an independent alignment strategy, they are not directly comparable to the endogenous RNA-seq expression values. To calibrate these measurements, a scaling factor was derived from the ZNF18 dataset by comparing the codon-specific read counts with the transgene abundance estimated from the differential expression analysis (OE − GFP). This empirically determined correction factor was subsequently applied to all codon-optimized constructs, thereby expressing transgene abundance on the same scale as the endogenous RNA-seq measurements. Corrected transgene expression values were then used for all downstream comparisons. To compare endogenous expression across physiological contexts, publicly available RNA-seq datasets from the Human Protein Atlas (cell lines) were downloaded and normalized to the local RNA-seq scale. A normalization factor was calculated from the median expression ratio of KZFPs detected in both the Human Protein Atlas K562 dataset and the local K562 GFP control RNA-seq dataset, and subsequently applied uniformly to all Human Protein Atlas datasets. This normalization enabled direct comparison of endogenous expression across biological contexts with the corrected transgene expression values. Global KZFP expression was calculated as the median normalized expression of all annotated KZFPs within each biological context. For the four KZFPs selected for detailed characterization, endogenous expression across Human Protein Atlas cell lines was compared with corrected transgene expression following doxycycline induction. Expression distributions of all genes and KZFPs were visualized using ranked expression plots and density histograms. All analyses were performed in R using the tidyverse package.”

      We fully acknowledge that the overexpression system used in this study was primarily designed as a discovery platform to identify candidate functions, targets, and interaction partners of KZFPs that are otherwise expressed at lower levels in K562 cells. As the reviewer correctly points out, determining whether these regulatory effects occur at endogenous expression levels in physiologically relevant cellular contexts represents an important next step. We Thereby also clarified this in the “Limitations to this study” paragraph:

      “To better place our experimental system into a physiological context, we compared endogenous KZFP expression in K562 cells with publicly available transcriptomic datasets from the Human Protein Atlas. These analyses showed that K562 cells do not exhibit unusually low global KZFP expression compared with other human cell lines. However, consistent with the restricted expression patterns of this protein family, KZFPs as a whole are expressed at substantially lower levels than the average human gene. For the four KZFPs characterized in detail, doxycycline induction produced transcript levels that exceeded the highest endogenous expression observed across the analyzed human cell lines. Accordingly, the overexpression system used in this study was not designed to recapitulate physiological expression levels but rather to maximize the identification of candidate target genes, interacting partners, and regulatory pathways for KZFPs that are otherwise expressed at low endogenous levels. Consequently, the molecular interactions identified here should be considered as hypotheses requiring validation under endogenous expression conditions in physiologically relevant cellular models.”

      “- RNA seq analysis: It is unclear how many cells were used in the RNA seq analysis, I would like to ask the authors to clarify that. Moreover, from my understanding the RNA seq analysis was done on day 3, while the Presto Blue analysis was done on days 4, 7 and 9. I would like to kindly ask the authors to motivate their choice for the day of the RNA sequencing analysis.”

      We agree that this information required clarification. The Methods section has been revised to specify the number of cells used for RNA-seq library preparation and to explain the rationale for performing RNA-seq after 3 days of doxycycline induction, before measurable proliferation defects emerge, in order to capture primary transcriptional responses to KZFP overexpression. The corresponding modification has also been added to the Results section when introducing the RNA-seq analyses.

      “For transcriptome profiling, K562 cells expressing the indicated inducible constructs were treated with doxycycline for 72 h before harvest. Total RNA was extracted using the NucleoSpin RNA plus kit (Macherey-Nagel) according to the manufacturer’s recommendations. RNA quantity and purity were assessed by spectrophotometry, and RNA integrity was evaluated before library preparation.”

      Is now:

      “For transcriptome profiling, 1 × 10⁶ K562 cells expressing the indicated inducible constructs were treated with doxycycline for 72 h before harvest. RNA was collected after 3 days of induction to capture the primary transcriptional responses to KZFP overexpression before substantial differences in proliferation became apparent. This early time point was chosen to minimize secondary transcriptional changes resulting from altered cell growth, cell-cycle distribution, or cellular stress, which become detectable in the proliferation assays performed after 4, 7, and 9 days of induction. Total RNA was extracted using the NucleoSpin RNA plus kit (Macherey-Nagel) according to the manufacturer’s recommendations. RNA quantity and purity were assessed by spectrophotometry, and RNA integrity was evaluated before library preparation.”

      “We then profiled the transcriptome of K562 cells overexpressing ZNF43 by deep RNA sequencing (RNA-seq)”

      Is now:

      “We then profiled the transcriptome of K562 cells overexpressing ZNF43 by deep RNA sequencing (RNA-seq) after 3 days of doxycycline induction, a time point selected to capture primary transcriptional responses before the onset of measurable proliferation defects.”

      Minor comments of the reviewer 2

      “- Figure S1D is not mentioned in the text before figure S1E. The order of the panels should be changed in the figure.

      Done as suggested by the reviewer.

      • "We selected genes that were downregulated upon ZNF43 overexpression and harboured a ZNF43 binding site within 10kb of their TSS (Fig. 1A) - don't the authors mean Fig. 2A?

      Done as suggested by the reviewer.

      • In Figure 4D, the GO terms cannot be read, as the sentences seem to be cut.

      Done as suggested by the reviewer.

      • All figures and figure legends need to be revised. In some cases, the letter size is too small, or the legend and explanation of colours is missing. Please see some examples below: Fig. S6C, Fig 6C, Fig S4C, Fig S5C (letter size too small) Fig S6G, Fig 4E (label/scale is missing)”

      Homogenized to Arial 6 by default as requested by most of journal guidelines

      __ Description of analyses that authors prefer not to carry out__

      We think that by proceeding as described above we will have addressed all major conceptual issues raised by the reviewers.

    2. Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.

      Learn more at Review Commons


      Referee #2

      Evidence, reproducibility and clarity

      Summary

      Foley et al establishes a scalable framework to probe KZFP function. They performed an array of inducible overexpression screen of 366 human KZFPs in K562 cells. This screen, together with the analysis of transcriptomic and available chromatin and proteomic datasets revealed that KZFPs regulate many different mechanisms, highlighting the functional diversity of KZFPs. Understanding this functional diversity is a very interesting, timely and relevant question, but it is also challenging to study. Therefore, the approach the authors develop is promising. While the quality of the experiments and data analysis is high, the weakness I see in the manuscript is the lack of major biological insights in relevant model systems. Please see my detailed comment in the significance part.

      Major comments

      • The authors chose four KZFPs to study in detail, but why they chose these 4 candidates is unlcear to me. It would be nice to add a more detailed description of the process by which they chose the four candidates.
      • The materials and methods part of the manuscript is not detailed enough for other researchers to reproduce the study. They should add more details to both experiments and data analysis part of this section. Below I highlight some examples for sake of clarity, but the authors should revise the whole materials and methods section and add more details keeping these examples in mind:
        • The authors do not state the titer of lentiviral vectors they generate nor the MOI or amount of virus they use to transduce the cells
        • In many cases, the specific softwares and the software version is not stated e.g. the analysis of the Gene Ontology Biological Processes
        • It would be beneficial for the readers to get more details about the construct they used, for example a map of the plasmid.
        • It is unclear how many cells were used for RNA extraction
        • It is unclear which microscopes were used for imaging.
        • The concentration of antibodies used for staining and the product number, and provider of the antibody is not always depicted.
      • The authors looked at available chromatin data in either K562 cells or HEK293 cells, which I think is a very good way of utilizing publicly available data. Since the authors showed that different KZFPs might be functionally relevant in different cell types/tissues, I was wondering if they checked if there is available ChIP Seq or CUT&RUN data in those specific cell types/tissues. If yes, that data should be included in the manuscript.
      • The authors mention that KZFPs are usually expressed at a low level in the K562 cell line they use, but there is no figure showing the expression level of KZFPs in this cell type. It would be important to see the baseline KZFP expression in these cells, the level of overexpression and compare it to the endogenous expression levels they show in different cell types/tissues, at least for the four candidates studied more in depth. This would help to understand whether this level of activity is something that could occur naturally in a physiologically relevant context.
      • RNA seq analysis: It is unclear how many cells were used in the RNA seq analysis, I would like to ask the authors to clarify that. Moreover, from my understanding the RNA seq analysis was done on day 3, while the Presto Blue analysis was done on days 4, 7 and 10. I would like to kindly ask the authors to motivate their choice for the day of the RNA sequencing analysis.

      Minor comments

      • Figure S1D is not mentioned in the text before figure S1E. The order of the panels should be changed in the figure.
      • "We selected genes that were downregulated upon ZNF43 overexpression and harboured a ZNF43 binding site within 10kb of their TSS (Fig. 1A) - don't the authors mean Fig. 2A?
      • In Figure 4D, the GO terms cannot be read, as the sentences seem to be cut.
      • All figures and figure legends need to be revised. In some cases, the letter size is too small, or the legend and explanation of colours is missing. Please see some examples below: Fig. S6C, Fig 6C, Fig S4C, Fig S5C (letter size too small) Fig S6G, Fig 4E (label/scale is missing)

      Significance

      Understanding the diverse roles of KZFPs is an important and interesting research question. However, studying KZFPs is challenging, as many KZFP-mediated effects appear to be highly cell type- and tissue-specific. This complexity is also highlighted by the findings of the current manuscript.

      A major strength of this study is the development of a scalable system that enables the simultaneous investigation of the entire KZFP family. Performing such analyses on an individual basis would be extremely time-consuming. Therefore, the authors provide an efficient and valuable screening platform that can identify promising candidates for further investigation. In this regard, the methodological advance represents the primary contribution of the work.

      At the same time, the study lacks a clear biological conclusion. While the screen identifies KZFPs with potential functional effects, it would substantially increase the impact of the manuscript if the authors selected at least one candidate for in-depth characterization in a biologically relevant cellular context. The current study is still of high quality and importance without these experiments, but such follow-up analyses would greatly strengthen the biological significance of the findings.

      Another limitation is that the experiments were performed in a cell type in which many of the investigated KZFPs are not normally expressed. As a result, the forced overexpression strategy may not accurately reflect physiological conditions and could potentially generate false-positive results. This concern is particularly relevant in light of the authors' statement that "KZFPs with sufficient regulatory potency to perturb cellular fitness outside of their normal setting are strong candidates for playing important roles within it." While this may indeed be true for some KZFPs, it is also possible that certain observed phenotypes simply arise from ectopic expression in an inappropriate cellular environment.

      More generally, the observation that KZFPs can have functions beyond TE repression is already established in the literature. Therefore, the manuscript provides limited new biological insight into this concept. The authors could potentially strengthen the novelty of the study by placing greater emphasis on specific KZFP subfamilies, such as SCAN-containing zinc finger proteins, which are a novel direction and have been implicated in non-canonical regulatory roles.

    1. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodents vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

    2. Author response:

      eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

      We thank the Editor and the Reviewers for taking the time to carefully review our manuscript and for providing constructive comments and suggestions, as well as the opportunity to revise our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

      We sincerely thank the reviewer for the positive evaluation of our study and for highlighting the significance of examining synaptic tagging and capture following theta-burst stimulation (TBS) in rodents and non-human primates.

      We agree that TBS is a physiologically relevant induction paradigm and that differences in inhibitory circuit dynamics may also contribute to the species-specific effects observed in our study. As highlighted by Larson and Munkácsy (2015), repeated bursts delivered at theta frequency (~5 Hz) can transiently suppress feed-forward inhibition through GABAB receptor-mediated mechanisms, thereby enhancing postsynaptic depolarization and facilitating NMDA receptor activation. We therefore agree that species differences in inhibitory regulation and burst-evoked depolarization may contribute to the distinct expression of synaptic tagging and capture observed between rats and non-human primates.

      We further agree that analysis of the “depolarization envelope” during sequential bursts may provide additional insight into the relative strengths of early induction mechanisms. We will therefore perform these analyses using the existing recordings and compare the depolarization envelope between rodents and NHPs in the revised manuscript. Following the reviewer’s suggestion, we will expand the Discussion section to acknowledge the potential contribution of inhibitory circuit dynamics and depolarization envelope differences during sequential bursts.

      Importantly, however, we believe that differences in downstream molecular maintenance mechanisms also contribute to these species-specific effects. In support of this, our molecular analyses revealed enhanced recruitment of plasticity-related proteins and transcriptional pathways in NHP hippocampus following TBS, including increased expression of BDNF and PKCζ. These findings suggest that both induction-related network properties and downstream molecular stabilization mechanisms may collectively contribute to the enhanced associative plasticity observed in NHPs.

      We also thank the reviewer for the important point regarding PKMζ antisense experiments and the study by Mei et al. (2011). We agree that the absence of an effect of PKMζ antisense oligodeoxynucleotides does not necessarily exclude a role for persistently elevated PKMζ in the maintenance of theta-burst late-LTP. As demonstrated by Mei et al., BDNF together with theta-burst stimulation can maintain late-LTP in the absence of protein synthesis, potentially through stabilization of PKMζ protein levels by reducing degradation rather than through de novo synthesis. However, these findings are not directly comparable to our study, since our experiments involved theta-burst stimulation alone without exogenous BDNF application. Interestingly, our results suggest species-specific differences in the interaction between BDNF and PKMζ signaling pathways. In rats, TrkB/Fc-mediated blockade of BDNF impaired TBS-LTP maintenance, whereas PKMζ inhibition alone had no significant effect. In contrast, in NHP hippocampal slices, inhibition of either BDNF signaling or PKMζ alone failed to abolish late-LTP, whereas simultaneous inhibition of both pathways disrupted LTP maintenance.

      These findings suggest that endogenous BDNF signaling and PKMζ may operate through partially redundant or compensatory mechanisms, particularly in the primate hippocampus. Therefore, although our findings indicate that de novo PKMζ synthesis may not be strictly required under the present experimental conditions, we cannot fully exclude the possibility that protein synthesis-independent stabilization or maintenance of PKMζ contributes to theta-burst late-LTP maintenance in rodents or NHPs. We will now clarify this point in the revised Discussion section.

      Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      We thank the reviewer for this important and thoughtful comment regarding statistical interpretation and biological replication. We agree that, particularly for electrophysiological experiments where multiple slices may originate from the same animal, the effective sample size for species-level conclusions should be considered at the animal level rather than solely at the slice level.

      In the revised manuscript, we will clearly indicate the number of biological replicates (animals) together with the number of slices contributing to each electrophysiological experiment, as well as the biological replicates used for qPCR and Western blot analyses. We will also clarify whether multiple slices from the same NHP/rat contributed to the same experimental condition. These details will be incorporated into the figures and figure legends wherever appropriate.

      In addition, we will perform animal-level analyses by averaging slice responses within each animal prior to statistical comparison and, where appropriate, apply hierarchical or mixed-effects statistical models to account for the nested structure of slices within animals.

      We acknowledge that the number of non-human primates (NHPs) available for this study was inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with primate electrophysiology and tissue collection. Consequently, achieving sample sizes comparable to rodent studies is often not feasible in NHP research. Nevertheless, to further strengthen the biological robustness of the findings, we are currently in the process of obtaining additional NHP brain samples and plan to repeat key experiments in an additional 3-4 animals. We believe these revisions and additional experiments will substantially strengthen the statistical rigor and overall interpretation of the study.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      We thank the reviewer for this insightful comment regarding the mechanistic interpretation of the STC findings. In the present study, we selected the 30 min interval based on well-established classical STC paradigms in rodents, where this interval reliably falls within the effective tagging and capture window. Using this experimentally validated interval allowed us to directly compare whether TBS is sufficient to support STC in primates versus rats under equivalent experimental conditions. Accordingly, the primary objective of this study was to determine whether TBS-induced STC varies across species, rather than to comprehensively define the temporal dynamics of the tagging window.

      We agree, however, that the current experiments do not distinguish whether the primate-specific effect reflects prolonged tag persistence, enhanced plasticity-related protein (PRP) synthesis, altered capture efficiency, or a shifted temporal window. Addressing these possibilities would indeed require systematic temporal interval analyses (e.g., ±15, ±30, ±60, and ±90 min), which represent important future directions. Such experiments are particularly challenging in non-human primates because the availability of primate tissue and experimental resources for large-scale electrophysiological studies remains limited and is currently beyond our experimental capacity due to substantial ethical, logistical, financial, and technical constraints.

      Nevertheless, we fully agree with the reviewer that these experiments are important for advancing the mechanistic interpretation of the findings. Similar temporal analyses have recently proven informative in our rodent studies (Chong YS, Ang SR, Sajikumar S. Commun Biol. 2025;8:553). Importantly, we are currently in the process of obtaining additional non-human primate samples and plan to extend the present work by examining an additional 60 min temporal interval to further characterize the temporal properties of synaptic tagging and capture in non-human primates.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and post-mortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      We thank the reviewer for raising this important and insightful point. We agree that differences in developmental stage between the experimental groups represent an important consideration when interpreting potential species-dependent effects. In the present study, rat experiments were performed in 5-7 week-old animals, whereas non-human primate (NHP) tissues were obtained from 5-7-year-old monkeys. This difference largely reflects the practical, ethical, and logistical constraints associated with NHP research and tissue availability. We acknowledge that these ages are not developmentally equivalent and that maturation state may influence BDNF signaling, protein synthesis capacity, synaptic plasticity thresholds, and transcriptional responses relevant to late-LTP and STC mechanisms.

      We also recognize that differences in euthanasia procedures, tissue extraction, slice preparation, and postmortem handling between rodent and primate tissues may influence tissue physiology and electrophysiological properties. Although extensive care was taken to optimize tissue viability and maintain stable recordings within each species, these variables cannot be completely excluded as contributing factors to the observed differences.

      Accordingly, we will revise the Discussion section to more explicitly acknowledge these limitations and clarify that our findings support potential species-dependent differences under the present experimental conditions, rather than definitive intrinsic species-specific mechanisms. Nevertheless, despite the inherent challenges associated with NHP electrophysiological studies, we believe that the present findings provide an important initial framework for understanding the translational relevance of synaptic tagging and capture mechanisms across species.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      We thank the reviewer for this important comment regarding statistical reporting and interpretation. We agree that the repeated occurrence of identical exact p-values in several nonparametric analyses reflects the relatively small sample sizes and the discrete nature of the statistical distributions. This issue is particularly relevant for the NHP experiments, where biological replication is inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with obtaining and processing primate tissue.

      In the revised manuscript, we will provide exact n values for all comparisons, including the number of biological replicates (animals) and slices where applicable. We will also include additional statistical details, including effect sizes and confidence intervals where appropriate, to improve transparency and facilitate interpretation of the reported findings. Furthermore, we are currently in the process of obtaining additional NHP samples and will attempt to include more biological replicates in the revised version to further strengthen the robustness of the analyses.

      We also agree that the issue of multiple testing should be addressed more explicitly, particularly because multiple genes and proteins were examined. In the revised manuscript, we will clearly state the statistical correction methods applied for multiple comparisons where appropriate. For analyses in which corrections were not applied, we will provide justification, noting that several experiments were based on hypothesis-driven candidate targets rather than exploratory large-scale screening analyses. These statistical considerations will be clarified in the Methods and Results sections.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

      We thank the reviewer for this thoughtful and balanced assessment of our work. We agree that the present data primarily support the conclusion that, under the specific experimental conditions examined, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than that observed in rat slices. We also agree that broader interpretations regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and potentially redundant BDNF/PKMζ-related mechanisms require additional mechanistic investigation and experimental validation.

      Accordingly, we will moderate these interpretations throughout the revised manuscript and clearly state that these conclusions remain preliminary. We will further emphasize that additional experiments, including increased biological replication, expanded temporal analyses, and further mechanistic investigations, will be necessary to more conclusively define the basis of the observed species-dependent differences. Within our current experimental capacity, we are actively working to obtain additional non-human primate samples and plan to incorporate additional biological replicates and key follow-up experiments in the revised version to further strengthen the robustness of the findings.

      At the same time, we believe the present study provides an important initial contribution to an understudied area by directly examining synaptic tagging and capture mechanisms in the primate hippocampus. Given the limited availability of non-human primate electrophysiological data in the field, these findings may offer a valuable framework for future studies investigating the translational and evolutionary relevance of associative synaptic plasticity mechanisms across species.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodent’s vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

      We thank the reviewer for this thoughtful and insightful comment, as well as for the encouraging appreciation of our long-duration plasticity recordings and associative plasticity experiments, which are both technically demanding and time-intensive. We fully agree that interpretation of cross-species differences in synaptic plasticity requires careful consideration of multiple biological and environmental variables, including circadian state, enrichment conditions, strain differences, diet, lighting conditions, and species-specific behavioral ecology.

      Regarding the specific concern related to circadian phase and sleep-wake state, the reviewer raises an important point. Rats are nocturnal animals, whereas macaques are diurnal, and hippocampal plasticity mechanisms are known to be influenced by circadian rhythms and sleep-dependent regulation of synaptic proteins and signaling pathways. Previous studies have demonstrated modulation of LTP, synaptic tagging and capture and protein synthesis in rats across normal sleep-wake cycles. We therefore agree that these factors may influence plasticity outcomes and should be carefully considered in comparative studies.

      Studies have further shown that theta frequency is highly sensitive to sleep-related manipulations. Specifically, theta frequency decreases immediately after sleep, remains elevated during sleep deprivation, and rapidly declines following recovery sleep. In aged animals, these effects appear comparatively attenuated, suggesting reduced sleep-dependent modulation of theta dynamics with aging. Therefore, disruption of normal circadian or sleep-wake patterns may significantly alter theta activity and associated plasticity mechanisms within a species and may not accurately reflect physiological baseline states (Utku Kaya et al., 2026).

      In our experiments, recordings from rats and macaques were performed during their respective active phases under standardized laboratory housing conditions, and we will further clarify these details in the revised Methods section. Nevertheless, we acknowledge that circadian state and related physiological variables cannot be completely excluded as contributing factors to the observed differences between species.

      More broadly, we agree with the reviewer that the present study does not permit definitive conclusions regarding universal “rodent versus primate” rules of synaptic plasticity. Our intention was not to propose a generalized dichotomy between rodents and primates, but rather to report that, under the experimental conditions used here, SC-CA1 TBS-LTP and associated synaptic tagging mechanisms differed between rats and macaques. We agree that broader evolutionary or cognitive interpretations would require systematic comparative analyses across multiple species, including both nocturnal and diurnal rodents as well as diverse primate species. Such studies would provide a stronger framework for distinguishing conserved versus species-specific mechanisms of plasticity.

      At the same time, we believe the present findings remain important because they provide one of the first direct experimental comparisons of SC-CA1 TBS-LTP-associated plasticity mechanisms between rodents and non-human primates under controlled ex vivo conditions. Although the interpretation should be done cautiously, the observed differences raise the possibility that certain metaplastic or protein synthesis-dependent mechanisms may not be fully conserved across species. Accordingly, we will revise the Discussion section to better emphasize the exploratory and comparative nature of the study, while explicitly acknowledging the limitations and potential confounding factors highlighted by the reviewer.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      We thank the reviewer for the comments.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      We can try to simplify the description of protocols at specific points for example, by providing an overarching description of the study design in the beginning of the Methods, rather than citing our previous eLife paper (Amaral et al., 2019), as suggested below. The methods are indeed quite extensive, but the this may be inevitable in a large-scale project such as this and we note that Reviewer #2 thought that part of the supplementary material should be incorporated back in the main text, which is a suggestion in the opposite direction. It may thus be hard to strike a balance between readability and comprehensibility that can address both reviewers’ opinions.

      (2) The article appears to oscillate between:

      (i) a description of the approach and the inherent challenges of such a multicenter replication program

      (ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      There is a bit of redundancy between tables and text, but this was intentional to make both of them self-explanatory. We also think stating the results in the text can allow us to make each of the replication criteria clearer, a concern that was also mentioned by the reviewer.

      As for requiring particular expertise in statistics for understanding, we mostly disagree. The main results (Tables 1 and 2, Figure 2) are expressed as percentages, and the only statistical concepts needed for interpreting these results are understanding prediction and confidence intervals. For this, we could provide a bit more guidance on their interpretation in the Methods section. Beyond that, most of the secondary results (e.g. Figure 3 and Figure 4) involve linear correlations, which is about as simple as statistical analysis gets.

      Of the results presented in the main manuscript, only Table 3 contains anything beyond percentages and correlations. We do agree that the meaning of each ratio in this table could be more clearly described, but there are essentially no expert-level statistics involved in their calculations.

      Other than that, the main statistical issues are the ideal way to aggregate the results from different replications for which we use different strategies for robustness purposes. However, all of these results are already in the supplementary material, so we don’t feel they interfere to much with the readability of the main manuscript.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections.

      This is indeed a good idea, and we plan to include an initial overarching description of the project in the Methods section of the revised manuscript.

      Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      Again, this is the opposite of what was suggested by Reviewer #2, so we would rather keep the Methods section more or less at its current level of detail.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      The reviewer is correct in his interpretation. Evaluating the main findings of articles or cleaning a field of wrong statements was never a goal of our study (and we were clear about this from the start). Our aim with the project was metascientific (i.e. evaluate the reproducibility of biomedical experiments with a set of common methods) rather than driven by a particular interest in the findings themselves. This is reflected by our choice of selecting experiments from a random sample of articles from multiple fields, rather than filtering by area of interest or importance. It also underlies our choice to evaluate experiments rather than claims, as this was more statistically tractable and potentially more objective as a meta-research goal.

      To be clear, we don’t feel this approach is inherently better or worse than evaluating claims in the literature, as in the Drosophila immunity article case (i.e. Westlake et al., 2026), which is also an important goal. They are merely approaches that answer different questions. Ultimately, we probably made our choice based on (a) our expertise/interest in meta-research rather than in the fields the replications stemmed from and (b) an attempt to engage Brazilian researchers in the project in a way that was non-confrontational and minimized backlash from their peers. We feel this was valuable for many of the lessons learned, although it also meant learning less about the research findings in question.

      Even though this was not a goal of the study, there is some knowledge obtained about the findings that is indeed largely absent from the current manuscript. We do not feel the current format allows for much discussion of 45 different findings, but we do have plans to address these in future articles (as outlined in our response to point 5). In the meantime, qualitative descriptions of each experiment can be found at https://osf.io/w5z9a. This is already mentioned in the Methods but could be reiterated in the results as well.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      We did not go too deep into that finding because we are publishing a separate article focused on the prediction project, which should look into factors that correlate with prediction accuracy, both at the level of predictors (e.g. research field, career level) and of individual predictions (e.g. information taken into account for each answer). We also feel that, given the multiplicity of predictors in the prediction analyses, these findings are a bit tentative, as the strongest predictors may be subject to effect size inflation from the “winner’s curse” effect (as outlined by Reviewer #2). We can try to emphasize it a little more in the discussion (although it already merits a whole paragraph on pages 23-24), but we feel we would be able to discuss it more critically in a follow-up article.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers.

      Indeed, some of our results did not fit this overarching analysis and were left for future publications. One of them is already available as a preprint, while the others are currently in preparation. Specifically, other results from the project should be spread about across five different articles.

      (a) A narrative article focused on challenges and lessons learned with the project, already published as a preprint at https://osf.io/preprints/metaarxiv/8y3tg_v1 (Amaral et al., 2026).

      (b) An article analyzing the prediction survey and markets results in detail (following the pre-analysis plan detailed in https://osf.io/6av7k/files/pjhgd and adding some exploratory analyses on prediction rationales).

      (c) Three articles describing the results of specific experiments with each experimental method (MTT, PCR, elevated plus maze) along with a discussion of aspects inherent to the method that seem to influence reproducibility.

      We can add this information more explicitly to the Methods section, including the links to the papers that have already been published at the time the manuscript is revised.

      Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was underrepresented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Thanks!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      We once more thank the reviewer for the compliments.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      We do acknowledge that the article currently includes a lot of supplementary material. This includes both supplementary figures/tables relating to the paper and many supplementary methods files (mostly hosted at the Open Science Framework). However, we also note that this is already a rather long paper as it stands and that Reviewer #1 has made the opposite suggestion of simplifying it. Thus, it may be hard to strike a balance that will suit all preferences, and we feel that maybe our attempt has landed somewhere in the middle of both reviewers’ ideal versions of the paper.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      We can try to engage with some of the above-mentioned literature in more depth in particular replication studies from other fields (some of which have appeared after our preprint (e.g. Tyner et al., 2026) and with the risk of bias and transparency literature (e.g. Serghiou et al., 2021). That said, we note once more that the article (and the Discussion section) are already quite long, and that analyzing each of these articles in depth is likely to be unfeasible.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Our original plan was to use p values as a predictor (see protocol at https://osf.io/9rnuj), but we later realized this was inadequate as it did not account for effect direction (i.e. significant effects in the opposite direction as the original may yield low p values, but this should not count as replication success). We thus switched to t values to be able to assign positive and negative signs depending on effect size direction. We note that, as we are using non-parametric Spearman coefficients (in which the module of t correlates negatively with the p value), the two approaches are effectively equivalent when original and replication effects have the same direction. This change was accounted for and justified in our list of protocol deviations at https://osf.io/9hj7t.

      Effect size (in relative terms) is already being used in the second predictor in the analysis (i.e. effect size decrease), as our idea was to use one significance-based predictor and one effect size-based predictor, to match what was done for the replication rates). We feel that using relative effects (e.g. response ratios) by themselves may not be as adequate, as for experimental methods with large coefficients of variation and/or low sample sizes (especially PCR ones), one can find large relative effects that are nevertheless far from statistical significance. This also makes relative effects not very commensurable between methods.

      We do believe there is a fair argument, however, to use standardized effect sizes as an alternative to t values (i.e. difference measured in standard errors of the mean) to measure significance/evidence strength. As some replications ended up underpowered, low t values may sometimes be due to insufficient statistical power/low sample size rather than replication failures. Using standardized effect sizes is not devoid of pitfalls (e.g. they can be quite variable when sample size is low), but it is worth doing as a robustness analysis.

      That said, there are a few statistical issues to be decided on how to calculate this (e.g. whether studies should be meta-analyzed using standardized mean differences rather than relative ones for this purpose, or whether an analog of the standardized effect size should be calculated for the log ratio of means). We would have to look more carefully into the multiple possibilities to decide on the best approach (and we do accept suggestions!).

      In the meantime, we note that running the prediction analysis using only experiments with ≥80% power yields a slightly higher correlation of t scores with researcher predictions (ρ = 0.49, p = 0.005), so we do not think that these underpowered experiments affect the trend too much. If anything, they could be masking a higher correlation between researcher predictions and replicability.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      We agree that we should make the description more precise (e.g. “reaching the same results when analyzing a set of data in the same way” for reproducibility and “finding similar results with new data collected under similar conditions” for replicability). We will update these definitions in the revised manuscript.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      We would argue that the reader would be correct in this case: the argument is a bit speculative. It does go in the direction of what is generally accepted within the field (i.e. that publication pressure can lead to lower reproducibility for a range of factors), but we’re not sure this connection has been demonstrated empirically, except for indirect evidence (such as the lower reproducibility in papers stemming from top institutions and “trophy journals” in, the higher frequency of positive results in US states with more researchers in Fanelli, 2010, or the higher number of problematic images for highly productive researchers in some countries in Fanelli et al., 2022. We could cite this evidence in the introduction and make the speculated connection more explicit, perhaps adding modeling work as well (e.g. Ioannidis, 2005; Smaldino & McElreath, 2016) to explain why this could be the case. But essentially, our opinion is that the connection remains a speculation.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      We can offer a more detailed description of the recruitment process (e.g. number and distribution of lectures, social media strategy used, etc.), although we would rather do this in a supplementary document so as not to make the Methods section even lengthier. We note, however, that we never aimed to recruit a “representative sample” of labs from the country: we were busy enough trying to get enough labs for the project to happen, and aware that the call would be inevitably biased by our own communication capabilities and personal networks.

      That said, the response rates for different regions of Brazil do generally match the distribution of research labs and graduate programs within the country (with some distortions likely caused by our personal networks, such as the large number of labs in Rio de Janeiro state), and seem to indicate a rather wide dissemination of the call. One way to visualize this would be to present the distribution of corresponding articles from the original studies selected for the replication (or even from the whole sample of articles obtained for experimental selection) along with the distribution of labs at different stages of the project in Figure S3, which generally show similar patterns. This would actually lend support to our statement that “the population of labs that performed replications was largely similar to the one that produced the original results” in the discussion.

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      We tried to find the combination of methods that would maximize the number of labs that would be included in the project. This is explicitly stated in our Methods Selection document at https://osf.io/qxdjt, but could be stated more explicitly in the paper as well.

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      We initially used single screening by three different reviewers (see https://osf.io/6av7k/files/u5zdq for criteria), as we were merely looking for a sample of experiments; thus, comprehensive inclusion of all eligible studies was not a priority. After this initial screening step, inclusions were confirmed in a consensus meeting with the three reviewers involved.

      Data extraction was also done by a single individual, but the resulting data led to a protocol that was later checked by two reviewers who had access to the paper and were explicitly oriented to judge whether the protocol consisted in a valid replication. Thus, discrepancies between what was in the paper and what was included in the protocol could potentially be flagged at these stages (as they were in many cases). We do note, however, that this is likely not as effective to prevent errors as having data extracted independently, as reviewers may overlook mistakes more easily when comparing two documents rather than extracting data anew. We did find that some errors in extraction slipped by, such as an MTT experiment where treatment concentration was inadvertently changed from mM to μM in a particular protocol step; this was picked up and corrected by 2 out of the 3 labs, but not by the third one, leading the latter replication to be invalidated.

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      As stated in the manuscript, we initially capped experiments at a predicted cost of R$ 5.000 (around USD 1336 at that time), considering reagent cost alone (as equipment and labor was provided by labs), as mentioned in the manuscript. Exclusion rates for that reason were 12/74 (16%) for MTT experiments, 36/132 (27%) for PCR ones and 4/40 (10%) for EPM ones. This is stated at

      This turned out to be an underestimation in many cases, especially as it did not account for pilot experiments, need for repetition, etc; thus, many experiments ended up costing considerably more than that ceiling. As we had included a contingency fund for those cases which we expected would occur , we avoided removing experiments from the sample for this reason as much as possible. Nevertheless, one elevated plus maze experiment ended up not being replicated for cost reasons, as the necessary rat strain was provided by a single facility in the country, meaning that a large number of rats would have to be acquired and transported to all labs at a cost that we were not able to cover.

      As these costs were covered by the coordinating team, we do not feel that this is likely to underlie the reduction in geographical coverage. Other reasons related to lab structure could have led to labs in less well-resourced regions to leave the project, but they probably has nothing to do with the experiments selected.

      That said, the cost cap does mean that the selection of experiments is not completely representative of the literature, but is enriched in relatively cheap and simple experiments which were able to perform (which was our next step for selecting the final sample of experiments. Exclusion rates due to lack of lab expertise and/or infrastructure to perform the experiment were 21/56 (37%) for MTT experiments, 67/89 (75%) for PCR ones and 7/34 (21%) for EPM experiments.

      We will try adding some of this information to the flowchart in Figure 1, as we agree it provides more context on the representativeness of the selected experiments.

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      The scale ranged from 1 (No relevant differences) to 5 (Very relevant differences that prevent considering the study as a direct replication). This scale was used for both the lab and the validation committee scores, and is described at https://osf.io/xgth2 (debriefing protocol) and https://osf.io/e3fjg (validation protocol).

      For the validation committee, we did use a threshold (any score of 4 or a sum of scores of 10 or more among 3 evaluators) to decide what had to be discussed to decide on inclusion, as mentioned on Page 7 of the Methods. For the labs, we used no threshold labs answered the protocol deviation question as a scale, but the decision of whether to consider the study a valid replication or not was not tied to this score.

      We can make both of these points (meaning of the scale and connection to lab’s decision to consider the replication valid) clearer in the Methods section.

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      For the initial analysis of justifications, one reviewer read all answers and flagged those that seemed to concern reproducibility of the methods (e.g. “we replicated the protocol exactly as planned”) rather than results reproducibility (e.g. “effects went in the opposite direction”). We then revised these answers among the whole coordinating team to decide whether we should contact the lab asking them to revise them. We can add this information to the Methods section.

      For classifications of the justification into categories (i.e. Table S7), justifications were classified by two independent reviewers based on categories created after an initial inspection of the data, and discrepancies were resolved by consensus. We can add this information to the table legend.

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      After we extracted data from the lab spreadsheets and summarized the results by code, labs received the results by e-mail and were asked to fill in a form on whether the results were in agreement with what they had found (see details at https://osf.io/nfr6y). Discrepancies in results at least 1 experiment were noted by 36% of the 53 (out of 56) labs that responded. Many of these stemmed from the coordinating team misunderstanding issues such as group identity or experimental unit identification in the spreadsheet. Others had to do with different ways to perform calculations (e.g. relative gene expression or % time spent in open arms). In some cases, simple errors in data transcription or typos caused the discrepancy.

      We were also surprised (and concerned) by the number of experiments in which we later found data errors that were not detected by this process (e.g. 18% of total). Our best understanding of this is that not every lab checked the results with the necessary care, as some errors were quite obvious, as in experiments in which sample size was different, or in which group labels were reversed. Ultimately, agreeing with a form that says “did you find any discrepancies?” may have been performed as a box-ticking exercise with little attention, and was probably not the ideal way to check data which led us to start reviewing results in live meetings afterwards. This is discussed in more detail in our challenges article (Amaral et al., 2026)

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX).

      R 4.5.1 was used for the analysis. We can add this information (which was present in the data repository in the R session info.txt file) and provide the R reference in the manuscript as well.

      In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this.

      If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Yes, we did use the escalc() function for this calculation (for both the replications and the original effect sizes). We can mention this in the manuscript.

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications.

      We calculated the coefficients of variation as the pooled SD divided by the mean of both group means. The reviewer is correct about the possibility of small-sample effects in this case (which we were not aware of). We will thus look into the possibility of implementing this via the escalc () function in the analysis of the revised manuscript.

      We also acknowledge that this could be a source of bias in the comparisons between original and replication CVs (albeit likely a minor one). That said, we note that sample sizes are not always larger in the replication for some experiments with large original effects, power calculations sometimes yielded lower sample sizes in the individual replication, albeit infrequently. On average, though, replication sample sizes were indeed larger.

      I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean.

      This is indeed the case; that said, the CV of the original effect also has random error relative to the true population CV and in that case, there is no way to estimate the uncertainty, as we have a single measure of that parameter. So there is probably no way around ignoring uncertainty in this case.

      We also note that we are looking for evidence of systematic CV inflation across all experiments (rather than for a statistically robust comparison between the CVs of any individual replication). For the sake of measuring this systematic inflation, the use of multiple experiments does allow us to estimate variability at the experiment level which should incorporate the lower-level variability between individual replications if this is not included in the model. Thus, we do not feel that our procedure introduced a systematic bias in the analysis at the experiment-level (although one could argue that it may lead to less precision).

      An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      We thank the reviewer for this suggestion, which indeed seems like an option in this case. We will look into this possibility, although we cannot guarantee at the moment that we will implement it, as we were not previously familiar with the method and will have to study it in more detail.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Yes, but there are likely different reasons for that. Labs leaving after being included may have been due to those in less privileged regions of Brazil (e.g. the northern and western regions of Brazil, generally speaking) having more difficulty in persisting in the project. That said, most of the “disappearance” happens between registration and inclusion which usually has to do with the labs not working with the methods that were ultimately included in the project. We also note that most of the states that lose representation were those that had a single lab to begin with, which may make the visual pattern more striking than the actual trend (as states in the South/Southeast also lose labs, but don’t disappear from the map).

      We note again that we never planned to achieve geographical representativeness when recruiting the labs on the contrary, we were aiming to maximize the number of available labs to run the project. That said, we do agree that for the sake of examining whether the population of labs is similar to the one that generated the original experiments (a claim that we do make in the discussion), this representativeness is important to assess. Once more, to allow the reader to evaluate this, we plan to add an additional map to Figure S3 to describe the Brazilian states where the original experiments came from (based on corresponding author affiliations) in which a similar bias towards the South and Southeast Region can be observed.

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      We agree that this would be useful information, and can experiment with the possibility, but our feeling is that the figure will likely become too noisy in cases where the 95% CIs overlap (which are quite frequent). If this is indeed the case, an option to allow the reader to examine this would be better to add an explicit link to the forest plots for each individual experiment (https://osf.io/sx9gv) in the figure legend.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      We selected the highest correlation values for each continuous outcome (t score and lnRR) and presented these separately in the text. This is a systematic way to perform the selection, but is obviously subject to the “winner’s curse” effect. We agree that adding both metrics for each predictor would be a fair way to keep this in perspective for the reader, but we would have to think about how to do this without sounding too confusing (as results for the two main outcomes are quite different).

      We do note, however, that the outcomes are indeed different and are expected to vary independently in some cases. For the correlation with replication probability predictions, for example, the effects in opposite directions would likely be expected, as larger original effect sizes will likely lead to larger probabilities to be assigned, but also to a higher possibility of effect size decrease. This low correlation between outcomes is probably something that should be pointed out and discussed in the revised manuscript.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made.

      We had given this some thought when writing the manuscript – but ultimately opted not to include confidence intervals for our replication percentages and to use the replication rates as descriptive measures only (as done in other replication studies such as (Errington et al., 2021).

      Even though we aimed for our sample of original experiments to be as systematic as possible, it is ultimately constrained by many factors (the choice of methods, the particular expertise of the labs, etc.) thus, adding confidence intervals represents the uncertainty around the replication rate of a very specific population of experiments, which is not directly comparable to those included in other replication efforts in any case.

      We will reconsider whether we should include confidence intervals for replication rates: although doing this for every replication rate in Table 1 and Table 2 may end up being too much information, it could probably be done at least for the replication rates of the main analysis in the text. We note that calculating confidence intervals for percentages is straightforward, requiring only the numbers that are in the table thus, any reader that wants to estimate uncertainty for those rates should be able to do it easily.

      We will also point out the uncertainty around the percentages mentioned in the discussion when comparing our replication rates with those of other studies, which we agree is an important issue to touch on.

      In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Both correlations in that case are non-parametric (e.g. Spearman’s ρ), so they cannot be directly transformed into Pearson’s r without making assumptions about the distribution (which we would probably avoid doing given the very marked outlier in our own). We can calculate a non-parametric confidence interval for our own correlation coefficient by resampling, but we will have to investigate whether this can be done using the available data from (Errington et al., 2021) (which is probably the case if effect sizes for all experiments have been shared).

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      It is indeed interesting, and seems to confirm an intuition that has long been present in the reproducibility field, but actually has little evidence to support it: if anything, there is evidence in the opposite direction in psychology (Youyou et al., 2023), although they looked at cumulative publication number, while we used number of publications in a fixed interval.

      We can expand a bit further on that finding: that said, we do note that the correlation is relatively weak and has a p value of 0.04. Thus, given the multiplicity of predictors would not be that unlikely to occur by chance, even though it seems intuitive. Thus, even though the relationship seems intuitive, we think it should be considered tentative at best and would refrain from discussing it in too much detail.

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Yes, that is what we mean by “incorrect labeling of error bars” (as can be grasped from the cited references).

      We can perform this regression, which seems relatively straightforward to do. That said, we note that another likely cause for outliers at least for cell line studies would be the use of different (and eventually inadequate) experimental units (e.g. having error bars that represent technical replicates of the same measurement rather than truly independent experiments). We suspect that this may have an even greater effect in terms of causing error bars not to express the same thing and the regression will not help in differentiating the two causes.

      We should also note that different types of experiments may be expected to have very different SDs, so the regression is likely to have a lot of error associated with it. In particular, it’s probably worth doing separate regressions for each method, to account for the likely difference in CVs between animal and cell line experiments, for example. This could also help tease apart the two causes above, as the experimental unit problem mentioned above will likely only be observed for cell experiments.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

      Thanks!

      Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      We thank the reviewer for the compliments. Again, a more extensive list of insights can be found in our challenges article (Amaral et al., 2026), which we will cite in the revised version.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

      We acknowledge the missing replications as a weakness, and we hope we have made that point clear in the discussion.

      Concerning the Brazilian research ecosystem, we could try to explore this in more detail in the introduction. In particular, we believe that a better understanding of the Brazilian academic system, including its regional disparities and the general composition of its workforce (which is largely composed of undergraduate and graduate students), can be useful in interpreting some of the findings.

      We can try to provide a bit more context at the end of the introduction (perhaps between the last 2 paragraphs, which would also address a point made by Reviewer #1), and also in different points of the discussion including those comparing replication rates with other studies or discussing infrastructural difficulties, some of which may be specific to the Brazilian context (such as difficulties in acquiring specific reagents or licenses). Still, we reiterate that, due to the lack of studies with comparable samples in other regions, we cannot tease apart the factors that are specific to Brazil from those affecting lab biology as a whole from the data alone.

      References:

      Amaral OB, Neves K, Wasilewska-Sampaio AP, Carneiro CF. 2019. The Brazilian Reproducibility Initiative. eLife 8:e41602. DOI: https://doi.org/10.7554/eLife.41602

      Amaral OB, Valério B, Carneiro CFD, Mota GPS, Neves K, Abreu M, Tan PB. 2026. Challenges for building up confirmatory science in lab biology: lessons learned from the Brazilian Reproducibility Initiative. MetaArXiv, DOI: https://doi.org/10.31222/osf.io/8y3tg_v1

      Errington TM, Mathur M, Soderberg CK, Denis A, Perfito N, Iorns E, Nosek BA. 2021. Investigating the replicability of preclinical cancer biology. eLife 10:e71601. DOI: https://doi.org/10.7554/eLife.71601

      Fanelli D. 2010. Do pressures to publish increase scientists’ bias? An empirical support from US states data. PLoS One 5:e10271. DOI: https://doi.org/10.1371/journal.pone.0010271

      Fanelli D, Schleicher M, Fang FC, Casadevall A, Bik EM. 2022. Do individual and institutional predictors of misconduct vary by country? Results of a matched-control analysis of problematic image duplications. PLoS One 17:e0255334. DOI: https://doi.org/10.1371/journal.pone.0255334

      Ioannidis jpa. 2005. why Most Published Research Findings Are False. PLoS Medicine 2. DOI: https://doi.org/10.1371/journal.pmed.0020124

      Serghiou S, Contopoulos-Ioannidis DG, Boyack KW, Riedel N, Wallach JD, Ioannidis JPA. 2021. Assessment of transparency indicators across the biomedical literature: How open is open? PLOS Biology 19:e3001107. DOI: https://doi.org/10.1371/journal.pbio.3001107

      Smaldino PE, McElreath R. 2016. The natural selection of bad science. R Soc Open Sci 3:160384. DOI: https://doi.org/10.1098/rsos.160384, PMID: 27703703

      Tyner AH, Abatayo AL, Daley M, Field S, Fox N, Haber NA, Hahn KM, Struhl MK, Mawhinney B, Miske O, Silverstein P, Soderberg CK, Stankov T, Abbasi A, Aberson CL, Aczel B, Adamkovič M, Albayrak N, Allen PJ, Andreychik M, Awtrey E, Axxe E, Azevedo F, Bader MD, Bago B, Bailey J, Bakker M, Banik G, Banks GC, Baskin E, Batruch A, Beatteay A, Behr SM, Berente N, Berry Z, Białkowski J, Bodroža B, Boeschoten L, Bognar M, Bokhove C, Bonfiglio D, Bouwman R, Brady TF, Braithwaite SR, Briceño Jiménez G, Brick C, Bricka T, Briker R, Brown AN, Brown GDA, van Aert RCM, Caldwell K, Capitan S, Capitán T, Chandler J, Charles T, Chartier CR, Chawdhary R, Cheng KJ, Chopik WJ, Clark B, Colvin VE, Comer CC, Costantini G, Coupé T, Cummins J, Czernatowicz-Kukuczka A, de Leeuw J, Dobolyi D, Druckman JN, Duan J, Dujmović M, Dunleavy DJ, Durkee PK, Emery C, Esterling KM, Evans TR, Fedor A, Fernández-Castilla B, Fiala N, Field JG, Fong N, Fonseca MA, Freeman ALJ, Freese J, Geiger SJ, Geng J, Getz LM, Geven LM, Gleibs IH, Gonzales DP, Gooty J, Gourdon-Kanhukamwe A, Greculescu C, Griffin SM, Grigoryan L, Grunow M, Gunby N, Hall B, Hanel PHP, Hannon EE, Harper S, Held MJ, Hickman L, Higgins NC, Hippel S, Hoeppner S, Hong S, Hostler TJ, Inzlicht M, Izydorczak K, Jaeger B, Jankowsky K, Jarke-Neuert J, Jensen M, Jokić B, Jolles D, Jolly P, Jones AM, Juanchich M, Kačmár P, Kapoor H, Keljanovic A, Koirala S, Kołczyńska M, Kouroupaki D, Kühnen U, Landgrave M, Larson MJ, Laulié L, Lawrence ACE, Le Forestier JM, Leahy KE, Lee S, Leslie J, Lewis SC, Limnios C, Lin H, Liu A-C, Lloyd JW, Ludvig EA, Lynott D, MacDonald J, Mallik P, Mallinson DJ, Marinazzo D, Martarelli CS, Matacotta J, McBride A, McHugh C, McMillan G, Méndez E, Metzger M, Michaelides MP, Michalak J, Micheli L, Miller JK, Milyavskaya M, Molden DC, Monjaras AG, Moreau D, Morrow A, Moya C, Mudrik L, Mulder LB, Munt KA, Nandi A, Nason K, Nast C, Nave G, Nax HH, Neubauer F, Nguyen PLL, Nichols AL, Nilsonne G, O’Boyle E, Oettinghaus J, Oh J, Oshana A, Ostermann T, Ostrowski RP, Oyebanjo A, Panczak R, Patrianakos J, Pavez I, Pavlov YG, Persson S, Perugini M, Peters K, Pieters C, Ponizovskiy V, Porter ND, Prenoveau JM, Purić D, Purol MF, Puthillam A, Quinn KA, Ramljak M, Reed WR, Ritchie M, Ritzau M, Roche SP, Rodela R, Röer JP, Ropovik I, Rothschild J, Saal J, Safadi H, Samaha J, Sanchez M, Sankaran S, Santos D, Sargent AC, Sauter M, Schmidt K, Schnabel L, Schroeder AN, Schuetz SW, Schuetze BA, Schulte-Mecklenbeck M, Schütz A, Sevigny EL, Shackleton E, Shafranek RM, Shaki S, Shakya S, Sirota M, Sisco MR, Sitnikov MM, Slevc LR, Smalarz L, Smith CT, Snyder JS, Sommet N, Sonmez F, Spellman BA, Stanulewicz-Buckley N, Stock G, Street CNH, Strømland E, Sundelin T, Syed M, Szabelska A, Szaszi B, Szumowska E, Tagat A, Täuber S, Tay L, Thapa S, Thatcher J, Tsaklakidou D, Tummers L, Turkovich E, Tutor MV, Urbanska K, van ’t Veer AE, van Assen M, van de Ven N, van den Goorbergh R, Vargo EJ, Vaughn LA, Vazire S, Vermeulen JM, Vo DTH, Volkman V, Wagenmakers E-J, Wagner D, Walasek L, Walter F, Warmelink L, Wei L, Weißflog MI, Weller N, Wichman AL, Wilbiks J, Williams JR, Wolfe K, Wort F, Wright R, Wulff JN, Xue X, Yan VX, Yang Y, Yoon S, Žeželj I, Zhang Y, Ziano I, Zogmaister C, Zupan Z, Zwaan RA, Nosek BA, Errington TM. 2026. Investigating the replicability of the social and behavioural sciences. Nature 652:143–150. DOI: https://doi.org/10.1038/s41586-025-10078-y

      Westlake H, David F, Tian Y, Krakovic K, Dolgikh A, Juravlev L, Bournonville TE de, Carboni A, Melcarne C, Shan T, Wang Y, Mu Y, Kotwal A, Pirko N, Boquete JP, Schüpfer F, Rommelaere S, Poidevin M, Liu Z, Kondo S, Ratnaparkhi GS, Chakrabarti S, Liu G, Masson F, Xiaoxue L, Hanson MA, Jiang H, Cara FD, Kurant E, Lemaitre B. 2026. Reproducibility of scientific claims in Drosophila immunity: A retrospective analysis of 400 publications. eLife 15. DOI: https://doi.org/10.7554/eLife.108404.1

      Youyou W, Yang Y, Uzzi B. 2023. A discipline-wide investigation of the replicability of Psychology papers over the past two decades. Proceedings of the National Academy of Sciences 120:e2208863120. DOI: https://doi.org/10.1073/pnas.2208863120

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      This interesting manuscript uses single cell RNAseq of developing C. elegans larvae to identify temporal pulses or oscillations in gene expression within glia and many other epithelial cell types - mostly in genes related to cuticle synthesis or remodeling. It identifies different sets of genes that oscillate within different cell types, and identifies many apparent oscillatory genes that were missed in prior studies because they are expressed in smaller populations of cells (whereas bulk data mainly report on oscillations within the major hypodermis).

      A second major contribution of this manuscript is to pioneer analysis methods for detecting oscillatory gene expression in scRNAseq datasets. That said, it's important to state that the methods for estimating phase coherence, GAM, perplexity, etc. make sense to me intuitively but I can't assess the math and other details, which are outside of my expertise.

      Most of my comments are minor ones about suggested clarifications to the text or figures. Some may require additional analyses, but none should require additional data collection.

      1. The manuscript focuses much of its analysis on one specific glial cell type (ILso), yet the authors tell us almost nothing about this cell type or why they would care about it. It would be helpful to include just a little more background on glial biology and the epithelial-like characteristics of socket glia.

      We added the following to the second paragraph of Results:

      "To this end, we used C. elegans strains expressing GFP specifically in ILso glia or in all glia (grl-18pro::GFP or mir-228pro::GFP, respectively). In C. elegans, all glia are found in sense organs. Most sense organs consist of one or more sensory neurons – each of which is specialized to detect different types of stimuli – and exactly two glia, called the sheath and socket. The sheath and socket glia form an epithelial tube continuous with the skin, through which the ciliated dendritic endings of sensory neurons protrude to sense cues in the external environment. In prior work we found that, in some sense organs, the socket glia produce cuticle specializations around specific sensory neuron cilia, but how these are coordinated with general cuticle synthesis was unknown (Fung et al. 2023)."

      Many transcriptomic studies of epithelia (including the Purice et al study of adult glia) use single NUCLEI RNAseq rather than single cells because of the challenges in separating cells connected by tight junctions. In C. elegans there are also various epithelial syncytia to contend with. In text or Methods, the authors should comment on why they think cells were appropriate to look at in this instance, and whether there are certain cell types that were missed or could only be obtained as cell fragments based on that choice.

      We added the following to the Methods section:

      "Presumably, fine cellular projections such as axons or glial processes are lost during cell dissociation, leaving mainly cell soma with nuclei. There is a risk that some cell types could be undersampled in cell sorting, as compared to sorting isolated nuclei, due to differences in how readily they undergo dissociation. On the other hand, retention of cytoplasmic material in this approach may better represent the total mRNA complement of the cell"

      1. Related to above, the authors do not mention any detection or exclusion of likely doublets. Is there reason to think that doublets were not present in any substantial numbers? I'm not super concerned about this since doublets containing hyp7 fragments should have worked against them in detecting glia-specific oscillations, but I do think the issue should be addressed in the text or Methods.

      We added the following in Methods:

      "Ambient RNA was subtracted using SoupX (Young and Behjati 2020). Potential doublets were assessed using DoubletFinder (McGinnis et al. 2019), but no cells were excluded on this basis. "

      1. p. 4 "previously unappreciated local differences in cuticle patterning." This statement should be tempered since many stage- or tissue-specific differences in cuticle patterning have been described previously (including in papers from the Heiman lab and others that are cited here). This study uncovers many additional examples but it's not a completely new finding.

      We have revised this:

      "Surprisingly, most pulsatile genes are specific to small sets of cell types, suggesting that previously unappreciated local differences in cuticle patterning are more widespread than previously recognized."

      1. Table 1 and text: the distinction between pulsatile and oscillatory should be explained more at the outset. These terms sometimes seem to be used interchangeably, but then Table 1 seems to make a distinction, not discussed until the final "limitations" section.

      We added the following definition to the Introduction:

      "Within a single larval stage, oscillatory genes display a characteristic sharp single peak of expression and we define rigorous metrics for identifying this signature, which we call "pulsatile expression."

      • *

      We also added a further clarification in the Results section under "De novo identification of pulsatile genes":

      "We reasoned that for individual genes, if gene expression in a given cell type were plotted as a function of pseudotime, oscillatory genes would display a distinct peak because they are expressed at a particular pseudotime (Fig. 4A). We refer to this transcriptional signature as "pulsatile" when viewed in a single developmental stage; genes with pulsatile expression are predicted to be oscillatory when viewed across all of larval development, but there may be important exceptions (see "Limitations of the study")."

      1. Figure 1 and Figure 3A,B. These UMAPs look very unusual, with no discernable individual dots. Is this just a resolution issue? Or, if relevant, please add info to legend and/or Methods explaining what data smoothing was done here to make them look this way and why.

      We have reduced the size of the dots (to point size 1 from point size 2) in the UMAPs in Fig. 1 and Fig. 3 to make individual dots more apparent. The noted effect is due only to the size of the dots; the UMAPs are plotted in the conventional way. The effect of different point sizes on the Fig. 3 UMAP is shown below [IMAGE CANNOT BE ATTACHED HERE]

      1. Figure 2C and Figure 6B. In the pseudotime plots, it would be natural for readers to assume that 0 is the beginning of the larval stage and 360 is the end, but that is not actually the way the Meeuse 2020 phase angles work - instead the beginning of the larval stage falls around 160. Please make sure this is made clear, especially when referring to "early and late groups" of TF targets. In Fig 6B, Early and Late categories appear reversed because of the way the data are plotted.

      We have replotted Fig. 6B using percent of larval stage progression rather than phase angles in degrees, with 0% corresponding to the peak of dpy-6 expression, to make the timing more intuitive. We have revised the description of the early and late groups in the Discussion.

      As Fig. 2C compares our data directly with the phases defined by Meeuse et al., we prefer to keep it consistent with that publication.

      1. Figure 3B and Figure 5D-G. The authors group many unidentified clusters into the catchall "skin" category but don't clearly define it in the main text. Table S2 suggests this category includes anterior and posterior skin cells but possibly also other cuticle-lined tubular epithelia that aren't properly referred to as skin (e.g. vulva cells, excretory socket or pore cells). It may also include things like rectum, buccal cavity, excretory duct. Please define your criteria for "skin" more precisely in the main text (any cuticle-lined cell type that is not glia?), and perhaps a more general term such as external epithelia would be more appropriate.

      We have changed this in the text to "skin-related cell types" to clarify that it includes hyp, seam, and some unidentified skin-related clusters (which may include some of the cell types you mention, for example the "skin_5" cluster may include vulval epithelia or their precursors as shown in Table S2).

      1. Also related to cluster assignments: please specify if "excretory" category includes canal, duct, pore, gland all together, or only a subset of these. Only the duct and pore are cuticle lined and therefore expected to have oscillatory matrix gene expression.

      We have changed this to "excretory cell" (or "exc cell") for clarity. We did not examine markers for the excretory duct, pore, or gland.

      1. Figure 5. This figure feels disjointed and could be broken up into two figures (panels A-C and panels D-G). The first 3 panels seem more related to Figures 3 & 4 - identifying which cell types have strong pulsatile gene expression - whereas the later panels get into the degree of cell type specificity in matrix gene expression.

      We appreciate the merit of this point and in fact we strongly considered splitting up this figure (in various ways) while writing. While we agree that the figure covers a lot of ground in this format, we feel that the subparts do not hold up as their own independent figures on equal footing with the other figures in the manuscript.

      1. Figure 5D-E. The very low degree of sharing is fascinating but could be an underestimate that depends on the thresholds chosen for calling a gene "pulsatile". It may be helpful to test a range of thresholds to see how much this matters. For those ~2,500 genes that appear pulsatile in just one cell type, are they called expressed but non-pulsatile in other cell types? That would seem odd to me biologically and most likely a threshold artifact.

      We have added the following caveat to the Results:

      "Put another way, 45% (2,390 of 5,268) of the genes we identified were expressed and pulsatile exclusively in a single cell type while only 17% (915 of 5,268) were pulsatile in five or more cell types (Fig. 5E). A potential caveat to this conclusion would be if some genes are not categorized as pulsatile in particular cell types due to lower expression (e.g., falling into Cluster 7 with high peak amplitude in one cell type, and Cluster 8 with low peak amplitude in other cell types; see Fig. 4B-C). However, if this occurs, it affects only a minority of cases: among genes categorized as pulsatile in only one cell type, 82% are not detected as expressed in any of the other oscillatory cell types, indicating that the apparent specificity most likely reflects cell-type-specific gene expression rather than thresholding effects."

      1. Figure 5F and p.21 Methods, the authors analyze only 140 collagen genes and 38 ZP domain genes retrieved from InterPro, but there are at least 173 cuticle collagen genes and 43 ZP domain genes described in the literature. Therefore, their lists are incomplete and the Methods should say so.

      Thank you for pointing this out. We changed the gene list to the cuticular collagens listed in Teuscher (2019). This did not affect the figure in a major way. We retrieved 44 ZP domain genes from InterPro, which match the ones listed by Cohen (2019) with the addition of cutl-19.

      1. If most oscillatory gene expression is truly a function of the molt cycle, as suggested by the matrix gene families in Figure 5, then one might expect that most of the detected oscillatory genes would no longer be expressed in adults, or at least wouldn't appear "pulsatile" in adults. Is this true? There now are a variety of published adult data sets, including the Purice et al data on glia, that could be examined to address this.

      PCA of adult cells does not exhibit the circular structure necessary to assess pulsatile expression. Previous work showed that most oscillatory genes are not expressed in adults, as expected (Meeuse et al. 2020).

      **Referees cross-commenting**

      I agree with the other reviewers' critiques, including the point of Reviewer #2 that orthogonal confirmation methods (such as by imaging) could have been nice but are not necessary. The question of Reviewer #3 about tissue synchrony/asynchrony is a very important one but I am not confident it can be addressed with these types of data.

      Reviewer #1 (Significance (Required)):

      As a molting invertebrate, C. elegans must build and shed its protective cuticle at multiple times across its life cycle, and this requires temporal control of many genes involved in matrix structure and processing. Although temporal oscillations were already well documented from bulk RNAseq data, this manuscript extends those prior findings by showing that different sets of genes oscillate within different cell types (including sensory glia), and by identifying many apparent oscillatory genes that were missed in prior studies because they are expressed in smaller populations of cells (whereas bulk data mainly report on oscillations within the major hypodermis). These data about cell-type specific temporal programs and gene sets emphasize the exquisite specificity of apical matrix and will be broadly useful to researchers in the C. elegans community.

      A second major contribution of this manuscript is to pioneer analysis methods for detecting oscillatory gene expression in scRNAseq datasets, even where bulk temporal data may not exist. This will be valuable for others doing sRNAseq studies in nematodes but also in other systems where cells may have molt cycle- or circadian-regulated oscillations.

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      SECTION A - Evidence, reproducibility and clarity

      Summary:

      Provide a short summary of the findings and key conclusions (including methodology and model system(s) where appropriate).

      The authors use single cell sequencing (scRNA-Seq ) of cells obtained from larval stages of C elegans -- primarily the L4 stage, but also the L2. Worms are disrupted and individual cells are sorted by expression of fluorescent markers specific to glial cells, a cell type that is relatively rare in the population, and of particular interest to the focus of the study. In this fashion, the samples were enriched for glial cells but, bedcause the soting is not perfect, also contain representations of other cell populations, including hypodermal (skin) and other epithelial cells, muscles, and neurons. 2D representation of the scRNA-Seq data in principle component (PC) space reveals sets of cells of the same cell type (for example glia or skin cells) arranged in roughly circular patterns, indicative of rhythmic gene expression in those cell types. Much of this circular PC behavior is shown to be driven by genes that had previously been shown, by bulk RNAseq of staged larvae, to undergo rhythmic gene expression in conjunction with the larval stages and the molts that punctuate the larval stages. Based on the previously published relative timing of expression of these cycling genes, and the pattern of peak expression of each gene in the scRNA-Seq 2D PC space, the authors could calculate a phase angle of expression of each gene relative to its peak in each cell, and thereby calculate an average phase angle of all cycling gene for each cell, and place that metric in register with the roughly circular pattern of cell types in the 2D PC plot.

      The authors show that many of these rhythmic genes encode extracellular matrix (ECM) proteins or other proteins related to cuticle synthesis and assembly, or molting. Cell types exhibiting cyclic gene expression included skin, pharyngeal epithelial, as well as several types of glia, notably socket glia, which synthesize a specialized ECM that surrounds and protects sensory neurons.

      Finally, the authors analyze the patterns of cyclic gene expression in several cell types with respect to the expression of transcription factors (TFs) that are expressed in the cell type, including TFs that appear to likewise cycle, and whose predicted targets are enriched for cycling genes. From this computational analysis, the authors derive sets of hypothetical transcriptional regulatory circuits underlying phased expression of cycling genes.

      Major comments:

      • Are the key conclusions convincing?

      1) Yes, the data support the conclusion that the authors' approach and methodology can take a list of genes known to cycle in expression level at larval stages and identify the cycling gene expression profiles of those genes in single cell sequencing datasets. It is also convincing that the authors' data analysis methods can identify cycling genes from the scRNA-Seq data that had not been previously identified as cycling from bulk RNAseq. Furthermore, the enrichment of genes encoding collagens and other ECM components is clear from the data.

      2) The above being said, it is noteworthy that the conclusions of the manuscript - including the sets of predicted novel cycling genes, and the predicted transcription factor-target circuits -- were not confirmed experimentally using independent samples or orthogonal methodology. I think it is OK for the authors to leave these predictions for later experimental confirmation, but it would be appropriate for the authors to discuss this caveat about the need for strategic experimental tests to confirm the more novel findings presented here, while at the same time pointing out predictions from their analysis that fit with previous experimental findings (for example cases such as NHR-85 and NHR-23 where previous studies support that the relevant TF is involved in regulating molting-associated transcriptional activity.)

      We have added the following sentence to "Limitations of the study":

      "Further, while our results are consistent with other studies (Meeuse et al. 2020; Gaidatzis et al. 2025) and successfully identify known regulators such as NHR-23 and NHR-85, it will be important in future work to test expression of the novel oscillatory genes and the roles of novel regulators we have predicted."

      3) There is an issue of concern that is perhaps about terminology, and not necessarily conceptual: Throughout the manuscript the authors variously use the terms, "oscillatory", "transient", and "pulsatile" to refer to cyclic gene expression. It seems that each of these terms could have distinct meanings, based on their English usage: The term "oscillatory" gene expression would seem to be a general term for gene expression that varies in a regular, rhythmic fashion. "Transient" gene expression seems like a general term for ON/OFF dynamics, albeit not necessarily oscillatory. "Pulsatile" gene expression implies oscillatory dynamics where the rise and fall of gene expression is relatively abrupt and might also imply ON/Off dynamics (between zero to some positive value). These terms are used seemingly interchangeably in the early parts of the manuscript, and then later, "pulsatile" is used increasingly, so the reader starts to wonder why. The authors should define these terms precisely and use the terminology deliberately and consistently.

      We have clarified the important point about oscillatory vs. pulsatile in the text. Please see our response to Reviewer 1, Point 5. Additionally, we have removed the use of "transient" except in the context of the phrase "transient aECM" that has been established in the literature.

      4) Related to the above, the authors should address how lowly-expressed genes behave in scRNA-Seq data, where the transcriptome is not fully sampled in each cell, and how that phenomenon could affect the apparent variation of gene expression within a population. My understanding is that if the expression level of a gene goes below some threshold percentage of the total transcriptome, it may not show up at all in the reads from that cell, even though the gene may still be expressed. Therefore, a gene can display apparent on/off behavior within a population of cells whilst the underlying variation in mRNA levels for that gene may be far less abrupt. How might this phenomenon affect the interpretation of a gene's dynamics as "pulsatile"?

      We added the following to clarify that sampling variation among cells was mitigated by applying a smoothing function based on each cell's five nearest neighbors in PCA space:

      "We then fitted the expression pattern of each gene with a generalized additive model (GAM) to obtain smoothed expression profiles. Because the GAM is fitted across many cells ordered along pseudotime, it captures the underlying expression trend even when individual cells show zero counts due to incomplete transcriptome sampling (e.g., Fig. 4A, black dots at y = 0)."

      As further described in Methods, our pipeline also incorporates several features that mitigate this valid concern:

      • First, before fitting gene-level dynamics, we retain only genes detected in at least 20 cells and in at least 5% of cells of a given cell type (Methods). While this filter may exclude some genuinely low-expressed oscillating genes, it ensures that pulsatile calls are made on genes where expression is reliably measurable.
      • Second, we apply two levels of smoothing. Prior to PCA, k-nearest-neighbor smoothing ensures that each cell's expression profile reflects a local average of transcriptionally similar cells rather than a single noisy measurement. When modeling gene expression along pseudotime, we fit a generalized additive model (GAM) with cyclic cubic splines, pooling information across many cells. The curves we score as pulsatile therefore reflect averaged expression across neighborhoods of cells, rather than raw per-cell counts subject to dropout.
      • Critically, dropouts arising from incomplete transcriptome sampling are independent of pseudotime (e.g., see dnj-1 in Fig. S3A). Our pulsatility criterion explicitly requires a low baseline combined with a well-shaped, high-amplitude peak in a specific pseudotime window, which dropout noise alone cannot generate. Indeed, as shown in Fig. 4A, the method readily identifies peaks even when many individual cells have zero detected reads (black dots at y = 0), demonstrating that the smoothed fit recovers the underlying dynamics from sparse data.
      • Finally, during development we also tested a logistic GAM that models the probability of detecting at least one read per cell, rather than read counts directly, which produced comparable results, though it saturated for highly expressed genes.
        • Should the authors qualify some of their claims as preliminary or speculative, or remove them altogether?

      5) Page 12: "Taken together, our results suggest that cuticle formation is the main commonality among pulsatile genes, and that distinct cell types use very different gene expression programs during this process. Thus, while cuticle aECM is typically perceived as a single homogeneous meshwork, our results suggest that the cuticle is actually a patchwork matrix with different patterning and composition contributed by distinct cell types."

      It is not necessarily surprising that the cuticle made by skin cells could have composition non-identical to the cuticle made by glial cells or pharyngeal cells. But by describing the cuticle as a 'patchwork' elicits in the reader's mind an image of the skin of the animal (seam + Hyp) being mosaic for distinct cuticle compositions. Is that what the authors intend to say? It would be interesting if there were differences in composition of cuticle between skin cell types, and so it would be helpful if the authors could comment on how the transcript profiles compare for hypodermal seam cells vs multinucleate Hyp cells.

      We have expanded on this idea:

      "Taken together, our results suggest that cuticle formation is the main commonality among pulsatile genes, and that distinct cell types use very different gene expression programs during this process. Classical work showed that the cuticle exhibits regionalized specializations – for example, alae are present only over seam cells; annuli and struts are present over hyp7 but not near the nose; the vulval cuticle is thought to present structural or chemical signatures for recognition during mating; and the pharyngeal cuticle exhibits three short projections in the buccal cavity, sieve-like fingers between the metacorpus and isthmus, and grinder elements in the posterior bulb. However, the extent to which these structural differences correspond to distinct molecular composition was not known. Thus, while cuticle aECM is typically perceived as a single homogeneous meshwork, our Our results suggest that the cuticle is actually a patchwork matrix with different patterning and molecular composition contributed by distinct cell types."

      • Would additional experiments be essential to support the claims of the paper?

      6) The data here are mostly from L4 stage larvae, with a possible (but unknown) contribution from L2 larvae. It would be helpful, in terms of broader understanding of their roles in larval progression, if some of the oscillatory genes identified here (especially the novel ones) were tested by orthogonal methodology (such as fluorescent protein tagging) for oscillatory expression at other stages. However, these experiments are arguably beyond the scope of this paper, and as long as the authors note the importance of such confirmatory experiments in their Discussion, I don't think that further experimentation is critical for this paper.

      We agree about the importance of these confirmatory experiments, and have added a comment in the Discussion (see response to Point 2 above).

      • Are the data and the methods presented in such a way that they can be reproduced?

      7) In general, yes.

      • Are the experiments adequately replicated and statistical analysis adequate?

      8) Yes.

      Minor comments:

      • Specific experimental issues that are easily addressable.

      9) Page 4: Regarding the single cell sequencing approach, the authors should comment on the extent to which mRNAs are efficiently recovered from hypodermal syncytial cells (Hyp), which are multinucleate. Could the data from Hyp be chiefly from nuclear transcripts? If so, how might that affect the interpretation of the data?

      We have added a comment in Methods related to caveats of cell sorting vs. nuclei sorting (see response to Reviewer 1, Point 2). As the proportion of immature (unspliced) mRNA and reads corresponding to the mitochondrial genome are not noticeably different in the hypodermal cells than in other cell types, we do not think the data are chiefly from nuclear transcripts.

      10) It is confusing that Table S1 lists male-enriched samples that were apparently sequenced, but only hermaphrodite data were analyzed for the paper. To prevent confusion, the male samples should not be listed.

      We have clarified in the Methods that these samples are included in Table S1 because we wanted to share the datasets with the community:

      "(Supp. Table S1; note this table includes related samples that were not used in the present analysis but that are deposited in the Gene Expression Omnibus (GEO) repository as a public resource)"

      11) Page 5, bottom: The following analysis requires clarification (at least for this reader): "To test if such oscillatory gene expression is present in ILso glia, we computed the average phase of each cell (Fig. 2B). Specifically, for each cell, we computed a weighted circular average of the peak phases of oscillating genes (derived from the previous bulk RNA-Seq data), using the gene expression levels in that cell as weights."

      In reading this part of the main text, this reader struggled to understand how one can compute the phase angle for a given gene in a cell by comparing its level of expression in that cell to measurement of the level of that gene in previous bulk RNA-Seq data. Of course, there is far more to the analysis than that, which the Methods and Materials section on page 18 describes in more detail, where one learns that the level if each gene in each cell is scaled to its maximum expression across all the cells analyzed, and that the previous bulk sequence analysis is used to simply provide a phase angle for the gene's peak expression relative to an arbitrary framework (which corresponds to a larval stage, one assumes). The presentation of this analysis on Page 5 in the main text should be revised to include a full description of what was done so the reader can follow along and understand it without having to read the Methods section. But moreover, the Methods section treatment of this analysis is still not entirely clear; for example, certain variables (W, s, and c) are not defined. The presentation of the mathematics should be clarified so that the reader can understand the analysis without having to look up scTransform-normalization.

      We have expanded and clarified this section:

      "To test if such oscillatory gene expression is present in ILso glia, we computed the average phase of each cell (Fig. 2B). Specifically, for each cell in our dataset, we considered its expression level of each of the 3,739 previously described oscillatory genes (Meeuse et al. 2020). To avoid biasing towards inherently highly-expressed genes (e.g., those encoding structural proteins), the expression of each gene in a given cell was scaled to its maximum expression across all cells. We then computed the average phase of each cell by taking the known phase for each gene (Meeuse et al. 2020) and calculating a weighted circular average, using the scaled expression of each gene in that cell as weights (Fig. 2B; each colored line represents one oscillatory gene with its angle representing its known phase and its length representing its scaled expression in that cell). we computed a weighted circular average of the peak phases of oscillating genes (derived from the previous bulk RNA-Seq data), using the gene expression levels in that cell as weights. This average results in a vector whose direction reflects the average phase of genes expressed in that cell, and whose length reflects how consistently the genes’ peak times align in that cell (Fig. 2B, black arrow)."

      In the Methods, we have moved the definitions of W, s, and c so that they precede the formula for the average angle .

      • Are prior studies referenced appropriately?

      12) yes

      • Are the text and figures clear and accurate? - Do you have suggestions that would help the authors improve the presentation of their data and conclusions?

      13) Figure S3 Panel A: What does the green line mean? Figure S3 Panel C: The use of the "predictors" tem is confusing because on page 8, the part of the narrative referring to Figure S3, the term used is "descriptors". Is that an meningful switch in terminology?

      We have expanded and clarified the Supp. Fig. S3 legend. For simplicity, we now use the term "metrics" to refer to the parameters used for hierarchical clustering of pulsatile expression (peak amplitude, baseline, fit, shape). This replaces our previous uses of "predictors" and "descriptors".

      14) The legend to Figure S3 requires more details to enable the reader understand the Figure. The same critique applies to most of the Supplemental Figure legends, where more details are required to allow the reader to understand each Figure without having to refer back to the main text.

      Thank you for pointing this out. We have revised and expanded all of the Supplemental Figure legends.

      15) Page 19: "The cells were grouped by cell type independent of the stage of collection (L2 or L4), and each cell type was processed individually."

      Why were L2s and L4s pooled? How does this affect the analysis and/or the outcomes? Could there be confounding effects from pooling the samples that could affect the analysis or the conclusions?

      We added the following clarification in the main text:

      "Because we found that cells clustered together based on their cell type rather than developmental stage, L2 and L4 cells of the same cell type were pooled for all downstream analyses (see Methods)."

      as well as the following explanation in the Methods:

      "The cells corresponding to the same cell type at different stages were then merged for subsequent analysis. After annotation, cells of the same cell type from L2 and L4 datasets were pooled for downstream analysis, such that each cell type is represented as a single combined cluster across stages. This provides two advantages: it increases statistical power by increasing the number of cells, and it favors genes that are oscillating in both larval stages. Because L2 representation is more limited (Table S1), the pooled pseudotime is dominated by L4 dynamics, ensuring that L2 cells are anchored on the L4-defined trajectory."

      **Referees cross-commenting**

      There is substantial agreement amongst all three reviewers, regarding the signifcance of the findings and that the conclusions are well enough supported by the data such that no additional experiments are required. We all recommend revisions to clarify or expand the description of the experiments and/or analysis. Many comments are reiterated by more than one Reviewer. I agree with all the other reviewers' critiques.

      Reviewer #2 (Significance (Required)):

      SECTION B - Significance

      • Describe the nature and significance of the advance (e.g. conceptual, technical, clinical) for the field.

      The finding of oscillatory and molting-related gene expression patterns in glial cells emphasizes the importance of molting-related ECM/cuticle production by these sensory-accessory cells and will serve as a platform for further studies and further understanding the structural and molecular basis of glial cell support functions, especially in the context of changing roles for sensory neurons during developmental progression.

      The methodology and data analysis of C. elegans scRNA-Seq data presented here offers several significant advances, especially since it had been known that thousands of genes cycle in rhythm with the C. elegans molting cycle, yet that was based on previous bulk sequencing, so it was not possible to resolve cell-type specific expression. This paper presents methods for analysis of cycling gene expression in specific cell types.

      The manuscript derives hypothetical TF-target regulatory interactions that are proposed do underly cyclic gene expression in specific cell types. This is a significant resource for future work to explore and delineate upstream oscillator mechanisms, and answer questions such as, Is there a central oscillator for all of larval stage rhythmic gene expression? and, How are different genes expressed with different phases of the larval stages? etc.

      • Place the work in the context of the existing literature (provide references, where appropriate).

      The authors cite important previous studeis that used scRNA-Seq to profile gene expression in specific C. elegans cell type in various developmental and physiological settings, and previous studies that used bulk RNAseq to identify genes whose transcripts cycle along with the larval stages. This manuscript reports the first study to examine cyclic gene express in C. elegans on the single-cell level.

      • State what audience might be interested in and influenced by the reported findings.

      Moderately broad audience of biologists interested in biological oscillators; developmental biologists interested in gene regulatory control of developmental cell fate timing and reiterative developmental processes; neurobiologists interested in glial cell function in developmental contexts.

      • Define your field of expertise with a few keywords to help the authors contextualize your point of view.

      • C. elegans larval development; temporal control of cell fate progression. Are there are any parts of the paper that you do not have sufficient expertise to evaluate.

      Honestly, some of the mathematical analysis is beyond my ability to judge whether the chosen approach is the best choice for the particular setting.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      This study describes both new scRNA-seq data from C. elegans, targeting glia/epidermal cell types and especially the ILso glial cell, and analytical approaches to identify periodically expressed genes in the dataset. Overall the data appear of high quality so have value as a resource, and the analysis provides a substantial improvement in our understanding of how different cell types vary in their cyclic expression across the molt cycles. While I have make many suggestions, overall this is a very nice study as is and definitely seems likely to be an impactful publication.

      Major

      Figure 2:

      • The dataset includes cells from multiple stages (L2 and L4 mentioned in the text, adult as well listed in Supplemental Table 1). There is a superficial display in Figure S1 which seems to imply that whether the same cell type clusters together across stages vs making stage specific clusters might be complex. But this isn't really discussed at all in the paper. For Figure 2 specifically it seems critical to know whether stages were pooled or separated for this analysis, and the question of whether the cyclic program varies across stages (at least for well sampled cells like ILso) is important.

      Please see our response to Reviewer 2, Point 15.

      Figure 3

      • This approach overall is good for cells that cycle in a way their signature comes up in Meeuse but would it detect rare cell type cycles? Maybe the PCA space velocity approach could be a way to screen for cell types that cycle in a way that isn't detected in the whole organism time course data (or rule out the presence of large cycling gene sets)? For example "pharyngeal gland" seems to have a weak cycling signature using the Meeuse gene set (Fig. 3D) but fairly clear "circular UMAP" structure (Fig 3B).

      We added the following to emphasize that our rationale for developing the perplexity metric is to identify oscillatory cell types de novo, i.e. without relying on the Meeuse et al. dataset:

      "We hypothesized that this [perplexity] metric would distinguish between pulsatile cell types (corresponding to relatively high perplexity) and non-pulsatile cell types (corresponding to low perplexity), without relying on prior bulk annotations that may be insensitive to rare cell types."

      Consistent with the reviewer's intuition, this approach does identify cell types missed by the Meeuse-based local phase coherence analysis, specifically coelomocytes and PHsh glia, which have perplexity >30 but did not reach significance by phase coherence (Fig. 5C).

      Regarding the pharyngeal gland, this cell type has intermediate scores by both metrics and falls below our conservative thresholds (Fig. 5C). It is possible that it has a genuine but weak oscillatory program that our methods are underpowered to detect given the number of cells recovered for this cell type.

      We considered using RNA velocity, but we have not succeeded in developing a satisfying quantitative score; therefore, the perplexity metric serves this role in our current analysis.

      • Do these data say anything about the question of whether all cell types in an organism are synchronized in the same phase as each other, or whether some might be systematically earlier or later in the cycle at a given time? It seems like if individual samples have enough stage bias (as an illustrative but made up example, if sample "230421_AM" has mostly early L4 while "230421_PM" has mostly late L4), then these data could be used to see if for example ILso cells tend to have earlier or later phases in the same sample compared to hyp cells. In my view, this is an important enough general question in the field to be worth addressing if the data are sufficient. And it could also provide an independent way to support/refute the presence of additional cycling cells (see Fig. 5 comments)

      This is an interesting idea, but unfortunately our synchronization at the population level is not sufficiently precise, and each sample contains cells spanning nearly the full phase range (shown below [IMAGE CANNOT BE ATTACHED HERE]). Under these conditions, between-sample phase offsets are dominated by within-sample dispersion, and we cannot reliably estimate systematic phase differences between cell types.

      Figure 5

      • This seems a nice approach to address the earlier question about detecting cell type specific oscillations. But then only results for the cell types previously identified as oscillatory are reported. It seems important to report the potential cycling genes for the other cell types (PHsh, Coelomocytes, maybe Pharyngeal Gland) so their cycling status could be tested by others in the future.

      We revised the text to highlight that these gene lists are in Supp. Tables S3 and S4.:

      "We limit our subsequent analyses to the 17 high-confidence cell types that appear oscillatory using both approaches (local phase coherence and perplexity), which together contain 5,268 pulsatile genes. Pulsatile gene lists for these and other cell types are provided in Supp. Tables S3 and S4 to facilitate independent assessment of their cycling status. A summary of pulsatile genes in these 17 cell types is shown in Table 1."

      • Regarding the last section ("only 17% were pulsatile in five or more cell types", "only 10 genes were pulsatile in all 17 oscillatory cell types" etc) - thresholding of a dataset like this can lead to false negatives resulting from incomplete (and cell type specific differences in power) which is a common source of technical non-overlap in this type of comparison. Indeed it is notable that the highest overlap was with ILso (specific sort target, likely to be especially well powered) and CEPso. There are various approaches to estimate not just confident overlap but also confident non-overlap, for example the "irreproducible discovery rate" (IDR) approach commonly used for ChIP-seq data. While clearly based on gene set enrichment there is cell type specificity, I'd suggest toning down the interpretation of the fractional overlap in the text if this can't be resolved.

      We toned down the interpretation of the fractional overlap:

      "To what extent are the same sets of pulsatile genes shared between cell types? To address this question, we examined the overlap between the pulsatile genes we identified in each cell type (Fig. 5D), noting that because power to detect pulsatile expression varies across cell types, the overlap values we report are likely to underestimate the true sharing between cell types.

      […]

      Put another way, 45% (2,390 of 5,268) of the genes we identified were detected as expressed and pulsatile exclusively in a single cell type while only 17% (915 of 5,268) were pulsatile in five or more cell types (Fig. 5E)."

      Minor

      Figure 1:

      • I was a little unclear about the coloring in Fig 1C (are the colors by annotated tissue or something else like clusters?) - suggest specifying in the legend.

      We updated the legend: "UMAP of the same cells as in B, with each cell colored by its annotated tissue identity."

      • More details on clustering and annotations approaches in the methods would be useful.

      We have substantially expanded the corresponding section.

      • I have mixed feelings about the word "skin" in the figure panels - while more accessible to a broad audience, hypodermis or hyp subset labels (hyp 7 etc) might be more precise.

      We have changed many of these to "skin-related." We cannot use the anatomical terms because we cannot confidently distinguish, for example, hyp1 vs hyp2, due to the lack of known markers for each cell type. We therefore refer to skin-related cluster 3 as "skin 3," because calling it "hyp 3" would lead to confusion with the anatomical term.

      • Table S2 would benefit from including the number of cells annotated with each cell type name

      We have added the number of cells per cell type to Supp. Table S2, with separate columns for L2, L4, and the total.

      Figure 2

      • Fig 2B is nice - clearly shows the difference in expression of phase specific genes in the two example cells and conceptual framework for averaging. I was struck by the relatively broad range of phase values though (For example the bottom cell has highly expressed genes with phases ranging from ~100 degrees to ~280). It seems this could reflect technical noise in the single cell data or imprecision in the phase calls in Meeuse. But there is also the interesting possibility that there is biological flexibility in the order/expression of phased genes at this single cell level. Not sure if there is an obvious way to address this or whether it should be in the scope of this work but maybe at least worthy of a mention

      A parsimonious explanation for the broad range of phase values in a single cell is the shape of the peak: examining the data from Meeuse et al, oscillatory genes do not generally display a sharp peak, but rather elevated expression over a span of ~3h (out of a larval stage of ~8h), which would correspond to expression over 100°. Indeed, the decentered genes in Fig. 2B correspond to the genes F53F4.2 and cutl-10 which have their peak expression at ~135° (26 h of larval development in the Meeuse dataset) but are still expressed at ~180° (28 h of larval development in the Meeuse dataset). Importantly, expression peaks tend to be roughly symmetric around the cell's true phase and therefore reduce the length of the phase vector but do not affect the average phase itself.

      Figure 3

      • The class Alter et al SVD paper https://www.pnas.org/doi/full/10.1073/pnas.97.18.10101 was the first use case of SVD/PCA in genome wide expression data and used (cell cycle) periodic expression as the main use case. The plots in Figures 2 and 3 are very similar to that approach, which basically used the relevant (~sin and ~cos correlated) principle components to define phases of both samples and genes. I mention this mostly in case it is useful to see how they approached the question and maybe as a relevant citation.

      We added the citation.

      Figure 4

      • Minor method clarification - how was DTW adapted to deal with circular data, specifically to identify cases where the peak is centered at pseudotime == 0/1? It seems from the figures that maybe some approach was used to center the raw data on the peak but I didn't see a description of how this was done (apologies if I missed it)

      We edited the Methods to make the connection with the previous section more explicit:

      "We used the trained model to predict expression of each gene along a grid of 128 regularly spaced pseudotime values, resulting in a smoothed expression profile. Further, for each gene, we shifted the pseudotime values to center the maximal expression value, and fitted a GAM as described above. To facilitate comparison of profile shapes across genes with different peak times, we additionally produced a centered version of each profile. For each gene, we identified the pseudotime at which the uncentered profile reached its maximum, then circularly shifted the pseudotime values so that this maximum fell at the center of the range (pseudotime 0.5). A new GAM was fit on the shifted data as described above, and used to predict expression along the same regular grid. This yielded a centered, smoothed expression profile for each gene in which all genes have their peak at the center of the pseudotime axis. These centered profiles were used in the subsequent section to compute both the baseline and shape metrics of each gene.

      […]

      We then scaled the curve by its maximum value and centered it around its maximal value. We then scaled the centered smoothed expression profile (defined in the previous section) by its maximum value. The Dynamic Time Warp distance between the scaled and centered expression and an ideal sharp peak defined as the density of a normal distribution of mean 0.5 and standard deviation 0.01 was computed with the dtw package."

      • The 2-PC view of ILso seems to align well with phase, but some of the other cell types (such as Seam in Figure 3C) are more complex - and also it seems possible that there could be cell types where the phase information is in e.g. PC2 and 3 instead of 1 and 2; how customizable is the approach and how dependent is it on a clean circular pattern in the PC plot?
      • *

      We have expanded the Discussion to include this point:

      "This could indicate either a genuine absence of oscillatory programs; the presence of oscillations driven by only a few genes that are insufficient to shape PCA structure; or oscillations that are present but reside in higher principal components dominated in PCs 1-2 by other sources of cell-to-cell variation."

      By using an Elastic Principal Cycle (ElPiGraph) to fit pseudotime rather than relying on angle from the origin (as is common for this type of data), we accommodate trajectories within PCs 1-2 that deviate from perfect circularity, including elongated or asymmetric shapes such as in seam cells (Fig. 3C). However, when phase information resides in higher-order PCs, in the absence of an independent timing reference there is no principled way to identify which PCs carry oscillatory signal versus other gradients of cell-to-cell variation. Recovering oscillations in such cell types would therefore require complementary approaches, such as synchronized time-course sampling, rather than a modification of the current pipeline.

      • It would be useful to annotate the examples (Fig 4A, lower panels) with whether they were newly identified or known from the bulk time course. And consider a larger supplemental figure with a sampling of newly identified genes in a similar format across a range of amplitudes etc

      We added Supplementary Figure S4C with examples across a range of amplitudes.

      • The examples are all relatively tight peaks (width We developed an approach to quantify the width of peaks in the Meeuse data (Methods); we display the distribution of peak width for genes expressed in ILso and seam cells in the new Supplementary Figure S4A. Our approach did not display a systematic bias to detect narrow or wide peaks. We added the following in the Results:

      “More generally, our classification captures a range of expression profile morphologies without apparent bias (Supp. Fig. S4).”

      Figure 5

      • There are important caveats in the interpretation of perplexity. For example a cell type that oscillates but with the vast majority of genes expressed uniformly or at one specific phase, would get a low perplexity, while a cell with multiple distinct states that don't cycle (this may be why body muscle has a modestly elevated sore) might achieve high perplexity. Worth addressing at some point.

      We added these caveats:

      "Potential caveats are that some non-oscillatory cell types might have high perplexity, for example if there are other sources of complex transcriptional heterogeneity among cells, while some oscillatory cell types might have low perplexity, for example if oscillating genes do not dominate the PCA structure."

      • Fig 5C raises the question of how power (for each of these metrics) relates to number of cycling genes in a cell type and the density of its sampling across time (for example is glia 4 just a poorly sampled cell type, or is it qualitatively different in what fraction of its transcriptome is cycling?). Just recoloring this plot by number of single cells per annotation might touch on this, or could try a subsampling approach.

      We added this with the new Supp. Fig. S5C:

      "To test whether differences in perplexity could be explained by differences in the number of sampled cells, we recomputed perplexity after subsampling to progressively smaller numbers of cells for several representative cell types. Perplexity values were largely stable across subsample sizes, indicating that the classification of cell types as oscillatory or non-oscillatory is not driven by differences in statistical power (Supp. Fig. S5C)."

      • The identification of pharyngeal muscle and epithelial oscillatory genes is a nice resource aspect of this paper given past work by EM showing these cells changing across the life cycle; it appears these cells have distinct enrichments (Fig 5G) and I think talking about these differences more explicitly could add to the closing paragraph in this section about aECM heterogeneity

      We have added the following:

      "In addition, the nematode astacin (NAS) metalloproteases appeared enriched in pulsatile genes in glia and pharynx, but not hypodermis (Fig. 5G, Supp. Table S6), consistent with ultrastructural observations that the pharyngeal muscle becomes secretory during molts and that the protease NAS-6 is required to digest the old pharyngeal cuticle (Sparacio et al., 2020)."

      Figure 6

      • This section is great and a very useful resource for future work. A detailed analysis may be beyond the scope of this work, but for Fig 6B I wondered whether the TF oscillation phase matched/preceded the timing of its predicted targets (for the subset of TFs that were themselves oscillatory in that cell type)? Even a qualitative analysis of this would be informative.

      We added a new supplementary Figure S7 and commented on it in the text:

      "__For the subset of TFs that are themselves pulsatile, we asked whether their peak expression coincides with or precedes that of their predicted targets. We found that pulsatile targets are modestly enriched in a temporal window around the TF's own peak (Supp. Fig S7), consistent with near-simultaneous expression of TFs and their targets, as previously observed for nhr-23 (Johnson et al., 2023). This temporal enrichment was most consistent for nhr-23 and nhr-25, which showed significant enrichment across most cell types, while other TFs showed more variable patterns (Supp. Fig. S7)."__

      Open-ended/discretionary

      A general challenge in single cell data analysis is that standard methods like clustering can give misleading or hard to interpret results when multiple processes occur simultaneously. For example, cells can have signatures of cell fate and cell cycle and depending on the genes used for clustering and strengths of those signals, naïve clustering may cause them to group by either fate of cell cycle phase. This is a long-winded way to say an application of the approach reported here would be to identify cycling genes shared between cell types that could be removed from the "variably expressed genes" lists prior to clustering to improve cell type separation, or used exclusively to allow clustering by phase rather than cell type. (definitely discretionary to consider this but could be mentioned in Discussion as a possible application)

      **Referees cross-commenting**

      I agree with all of this, including Reviewer #1 that asynchrony may be hard to address with current data, and with both reviewers that the dataset stands on its own.

      Reviewer #3 (Significance (Required)):

      This paper addresses the problem of how to identify cycling genes in single cell data, using the C. elegans larval/molt cycle as a model system. The system has emerged as a powerful model for understanding regulation of periodic gene expression, with past bulk RNA-seq time course have identified 1000s of cycling genes. However, how cyclic gene expression varies across cell types was not known. This study uses single cell RNA-seq and develops new analysis approaches to identify cycling genes across dozens of C. elegans cell types. Strengths are the generation of a new single cell data enriched for larval glia, identification of cyclic gene expression across many C. elegans cell types, an improved analytical framework for identifying cycling genes that could be applied in other datasets, and substantial analysis of pathways and regulators involved. Weaknesses are limited, and include minor overinterpretations of the data and missed opportunities for additional analyses. The work should be of interest to a broad audience including not just C. elegans researchers but also the single cell and chronobiology communities.

    2. Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.

      Learn more at Review Commons


      Referee #2

      Evidence, reproducibility and clarity

      Summary:

      Provide a short summary of the findings and key conclusions (including methodology and model system(s) where appropriate).

      The authors use single cell sequencing (scRNA-Seq ) of cells obtained from larval stages of C elegans -- primarily the L4 stage, but also the L2. Worms are disrupted and individual cells are sorted by expression of fluorescent markers specific to glial cells, a cell type that is relatively rare in the population, and of particular interest to the focus of the study. In this fashion, the samples were enriched for glial cells but, bedcause the soting is not perfect, also contain representations of other cell populations, including hypodermal (skin) and other epithelial cells, muscles, and neurons. 2D representation of the scRNA-Seq data in principle component (PC) space reveals sets of cells of the same cell type (for example glia or skin cells) arranged in roughly circular patterns, indicative of rhythmic gene expression in those cell types. Much of this circular PC behavior is shown to be driven by genes that had previously been shown, by bulk RNAseq of staged larvae, to undergo rhythmic gene expression in conjunction with the larval stages and the molts that punctuate the larval stages. Based on the previously published relative timing of expression of these cycling genes, and the pattern of peak expression of each gene in the scRNA-Seq 2D PC space, the authors could calculate a phase angle of expression of each gene relative to its peak in each cell, and thereby calculate an average phase angle of all cycling gene for each cell, and place that metric in register with the roughly circular pattern of cell types in the 2D PC plot.

      The authors show that many of these rhythmic genes encode extracellular matrix (ECM) proteins or other proteins related to cuticle synthesis and assembly, or molting. Cell types exhibiting cyclic gene expression included skin, pharyngeal epithelial, as well as several types of glia, notably socket glia, which synthesize a specialized ECM that surrounds and protects sensory neurons.

      Finally, the authors analyze the patterns of cyclic gene expression in several cell types with respect to the expression of transcription factors (TFs) that are expressed in the cell type, including TFs that appear to likewise cycle, and whose predicted targets are enriched for cycling genes. From this computational analysis, the authors derive sets of hypothetical transcriptional regulatory circuits underlying phased expression of cycling genes.

      Major comments:

      • Are the key conclusions convincing?

      1) Yes, the data support the conclusion that the authors' approach and methodology can take a list of genes known to cycle in expression level at larval stages and identify the cycling gene expression profiles of those genes in single cell sequencing datasets. It is also convincing that the authors' data analysis methods can identify cycling genes from the scRNA-Seq data that had not been previously identified as cycling from bulk RNAseq. Furthermore, the enrichment of genes encoding collagens and other ECM components is clear from the data.

      2) The above being said, it is noteworthy that the conclusions of the manuscript - including the sets of predicted novel cycling genes, and the predicted transcription factor-target circuits -- were not confirmed experimentally using independent samples or orthogonal methodology. I think it is OK for the authors to leave these predictions for later experimental confirmation, but it would be appropriate for the authors to discuss this caveat about the need for strategic experimental tests to confirm the more novel findings presented here, while at the same time pointing out predictions from their analysis that fit with previous experimental findings (for example cases such as NHR-85 and NHR-23 where previous studies support that the relevant TF is involved in regulating molting-associated transcriptional activity.)

      3) There is an issue of concern that is perhaps about terminology, and not necessarily conceptual: Throughout the manuscript the authors variously use the terms, "oscillatory", "transient", and "pulsatile" to refer to cyclic gene expression. It seems that each of these terms could have distinct meanings, based on their English usage: The term "oscillatory" gene expression would seem to be a general term for gene expression that varies in a regular, rhythmic fashion. "Transient" gene expression seems like a general term for ON/OFF dynamics, albeit not necessarily oscillatory. "Pulsatile" gene expression implies oscillatory dynamics where the rise and fall of gene expression is relatively abrupt and might also imply ON/Off dynamics (between zero to some positive value). These terms are used seemingly interchangeably in the early parts of the manuscript, and then later, "pulsatile" is used increasingly, so the reader starts to wonder why. The authors should define these terms precisely and use the terminology deliberately and consistently.

      4) Related to the above, the authors should address how lowly-expressed genes behave in scRNA-Seq data, where the transcriptome is not fully sampled in each cell, and how that phenomenon could affect the apparent variation of gene expression within a population. My understanding is that if the expression level of a gene goes below some threshold percentage of the total transcriptome, it may not show up at all in the reads from that cell, even though the gene may still be expressed. Therefore, a gene can display apparent on/off behavior within a population of cells whilst the underlying variation in mRNA levels for that gene may be far less abrupt. How might this phenomenon affect the interpretation of a gene's dynamics as "pulsatile"? - Should the authors qualify some of their claims as preliminary or speculative, or remove them altogether?

      5) Page 12: "Taken together, our results suggest that cuticle formation is the main commonality among pulsatile genes, and that distinct cell types use very different gene expression programs during this process. Thus, while cuticle aECM is typically perceived as a single homogeneous meshwork, our results suggest that the cuticle is actually a patchwork matrix with different patterning and composition contributed by distinct cell types."

      It is not necessarily surprising that the cuticle made by skin cells could have composition non-identical to the cuticle made by glial cells or pharyngeal cells. But by describing the cuticle as a 'patchwork' elicits in the reader's mind an image of the skin of the animal (seam + Hyp) being mosaic for distinct cuticle compositions. Is that what the authors intend to say? It would be interesting if there were differences in composition of cuticle between skin cell types, and so it would be helpful if the authors could comment on how the transcript profiles compare for hypodermal seam cells vs multinucleate Hyp cells. - Would additional experiments be essential to support the claims of the paper?

      6) The data here are mostly from L4 stage larvae, with a possible (but unknown) contribution from L2 larvae. It would be helpful, in terms of broader understanding of their roles in larval progression, if some of the oscillatory genes identified here (especially the novel ones) were tested by orthogonal methodology (such as fluorescent protein tagging) for oscillatory expression at other stages. However, these experiments are arguably beyond the scope of this paper, and as long as the authors note the importance of such confirmatory experiments in their Discussion, I don't think that further experimentation is critical for this paper. - Are the data and the methods presented in such a way that they can be reproduced?

      7) In general, yes. - Are the experiments adequately replicated and statistical analysis adequate?

      8) Yes.

      Minor comments:

      • Specific experimental issues that are easily addressable.

      9) Page 4: Regarding the single cell sequencing approach, the authors should comment on the extent to which mRNAs are efficiently recovered from hypodermal syncytial cells (Hyp), which are multinucleate. Could the data from Hyp be chiefly from nuclear transcripts? If so, how might that affect the interpretation of the data?

      10) It is confusing that Table S1 lists male-enriched samples that were apparently sequenced, but only hermaphrodite data were analyzed for the paper. To prevent confusion, the male samples should not be listed.

      11) Page 5, bottom: The following analysis requires clarification (at least for this reader): "To test if such oscillatory gene expression is present in ILso glia, we computed the average phase of each cell (Fig. 2B). Specifically, for each cell, we computed a weighted circular average of the peak phases of oscillating genes (derived from the previous bulk RNA-Seq data), using the gene expression levels in that cell as weights."

      In reading this part of the main text, this reader struggled to understand how one can compute the phase angle for a given gene in a cell by comparing its level of expression in that cell to measurement of the level of that gene in previous bulk RNA-Seq data. Of course, there is far more to the analysis than that, which the Methods and Materials section on page 18 describes in more detail, where one learns that the level if each gene in each cell is scaled to its maximum expression across all the cells analyzed, and that the previous bulk sequence analysis is used to simply provide a phase angle for the gene's peak expression relative to an arbitrary framework (which corresponds to a larval stage, one assumes). The presentation of this analysis on Page 5 in the main text should be revised to include a full description of what was done so the reader can follow along and understand it without having to read the Methods section. But moreover, the Methods section treatment of this analysis is still not entirely clear; for example, certain variables (W, s, and c) are not defined. The presentation of the mathematics should be clarified so that the reader can understand the analysis without having to look up scTransform-normalization. - Are prior studies referenced appropriately?

      12) yes - Are the text and figures clear and accurate? - Do you have suggestions that would help the authors improve the presentation of their data and conclusions?

      13) Figure S3 Panel A: What does the green line mean? Figure S3 Panel C: The use of the "predictors" tem is confusing because on page 8, the part of the narrative referring to Figure S3, the term used is "descriptors". Is that an meningful switch in terminology?

      14) The legend to Figure S3 requires more details to enable the reader understand the Figure. The same critique applies to most of the Supplemental Figure legends, where more details are required to allow the reader to understand each Figure without having to refer back to the main text.

      15) Page 19: "The cells were grouped by cell type independent of the stage of collection (L2 or L4), and each cell type was processed individually."

      Why were L2s and L4s pooled? How does this affect the analysis and/or the outcomes? Could there be confounding effects from pooling the samples that could affect the analysis or the conclusions?

      Referees cross-commenting

      There is substantial agreement amongst all three reviewers, regarding the signifcance of the findings and that the conclusions are well enough supported by the data such that no additional experiments are required. We all recommend revisions to clarify or expand the description of the experiments and/or analysis. Many comments are reiterated by more than one Reviewer. I agree with all the other reviewers' critiques.

      Significance

      • Describe the nature and significance of the advance (e.g. conceptual, technical, clinical) for the field.

      The finding of oscillatory and molting-related gene expression patterns in glial cells emphasizes the importance of molting-related ECM/cuticle production by these sensory-accessory cells and will serve as a platform for further studies and further understanding the structural and molecular basis of glial cell support functions, especially in the context of changing roles for sensory neurons during developmental progression.

      The methodology and data analysis of C. elegans scRNA-Seq data presented here offers several significant advances, especially since it had been known that thousands of genes cycle in rhythm with the C. elegans molting cycle, yet that was based on previous bulk sequencing, so it was not possible to resolve cell-type specific expression. This paper presents methods for analysis of cycling gene expression in specific cell types.

      The manuscript derives hypothetical TF-target regulatory interactions that are proposed do underly cyclic gene expression in specific cell types. This is a significant resource for future work to explore and delineate upstream oscillator mechanisms, and answer questions such as, Is there a central oscillator for all of larval stage rhythmic gene expression? and, How are different genes expressed with different phases of the larval stages? etc. - Place the work in the context of the existing literature (provide references, where appropriate).

      The authors cite important previous studeis that used scRNA-Seq to profile gene expression in specific C. elegans cell type in various developmental and physiological settings, and previous studies that used bulk RNAseq to identify genes whose transcripts cycle along with the larval stages. This manuscript reports the first study to examine cyclic gene express in C. elegans on the single-cell level. - State what audience might be interested in and influenced by the reported findings.

      Moderately broad audience of biologists interested in biological oscillators; developmental biologists interested in gene regulatory control of developmental cell fate timing and reiterative developmental processes; neurobiologists interested in glial cell function in developmental contexts. - Define your field of expertise with a few keywords to help the authors contextualize your point of view.

      C. elegans larval development; temporal control of cell fate progression.

      Are there are any parts of the paper that you do not have sufficient expertise to evaluate.

      Honestly, some of the mathematical analysis is beyond my ability to judge whether the chosen approach is the best choice for the particular setting.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully addressed the insightful comments provided by the reviewers which thoroughly increased our comprehension of the dynamics of centriole amplification. The manuscript has been revised accordingly and put in the context of the two papers we published since our last submission, showing that MCC differentiation is a genuine cell cycle variant. A point by point answer to all reviewer comments is provided below.

      Briefly:

      We have streamlined terminology and nomenclature in text and figures / better define experimental conditions with nocodazole

      We have tested the role of dyneins in the dynamics of centriole amplification

      We have done correlative light and electron microscopy on the early stages of centriole amplification

      We have analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors

      Collectively, this allowed us to make a clearer parallel with what occurs during centriole duplication and to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles.

      We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Public Reviews:

      Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      We have streamlined terminology and nomenclature, clarified the description of the complex events, and test the role of dyneins in centriole amplification. Microtubules density in MCC does not allow to extract information from imaging. In addition, we have done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Altogether, our new data allowed to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles. We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Reviewer #2 (Public Review):

      This important work will be of interest to centriole and cilia cell biologists. It describes in detail how microtubules control multiple aspects of centriole amplification in brain multiciliated cells. This study provides a greater time-resolved and molecular proteomic mapping of the different steps involved, with or without microtubule disruption. Boudjema et al. show that microtubules are important throughout the centriole amplification process, from the early stages, where the procentrioles emerge from a pericentriolar "nest", through the growth stage where microtubules maintain the perinuclear localisation, to the detachment stage, where microtubules assist in perinuclear disengagement and apical migration. The results are generally well supported by the evidence, but the manuscript would benefit significantly from some heavy editing to introduce more niche terms, standardize abbreviations in text, and labels on figures to help bring the readers, especially non-specialists, along with them - increasing the accessibility of their work.

      We thank the reviewer for his/her enthusiasm. We have streamlined terminology and nomenclature and clarified the description of the complex events to increase the accessibility of our work. We also replied points by points to his/her specific comments.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical-basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, and MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the CEN2-GFP; mRuby-DEUP1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described CEN2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of DEUP1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM, and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells (treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules in centriole amplification. Addressing the role of microtubules during different stages of centriole amplification required the development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive. This comprehensive analysis advances our understanding of MCC biology.

      Weaknesses:

      The role of microtubules and other molecular players during different stages of centriole amplification in brain MCCs can be further studied and strengthened using the tools developed in the manuscript. A more quantitative description of some of the analysis performed in the manuscript is required to strengthen the conclusions.

      We thank the reviewer for his/her enthusiasm. We have tested the role of dyneins in the dynamics of centriole amplification, done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Recommendations for the authors:

      As you will see, all reviewers felt that the analyses of the involvement of microtubules should be strengthened by including controls and additional experiments. Also, they agree that significant text editing would help to improve the manuscript's accessibility and readability.

      Specifically, they would suggest (1) streamline terminology and nomenclature in text and figures; (2) better define experimental conditions with nocodazole (concentrations used, effect on microtubules, effect on canonical centriole duplication); and (3), in the absence of other complementary genetic perturbation experiments, add a limitations paragraph in the discussion about conclusions drawn from nocodazole treatment alone.

      Reviewer #1 (Recommendations For The Authors):

      Main issues:

      (1) The authors use variable terminology to describe the same or similar events/structures. For example, in Figure 1 they refer to "centrosome stage" where they observe a pericentrin "cloud", which they later refer to as a "nest". In all other figures the first stage is not referred to as the "centrosome stage" but as the "cloud stage". Again, they also describe the "cloud" as a "nest" occasionally, but not always. In the cartoon, the nest is termed "centrosome cradle". The variable and inconsistent use of terms is confusing and the authors do not provide any explanation for the use of one vs. another.

      The text is now corrected. The centrosome stage corresponds to the stage preceding the beginning of centriole amplification in MCC progenitor. The pericentrosomal cloud of centriole and deuterosome elements forms later on, during the amplification A-stage. The formation of this cloud marks the beginning of A-stage, and persists up to G-stage where it dissolves. When we show that the cloud hosts the first stages of centriole biogenesis, we defined it as a “nest”. We do not use anymore the term craddle.

      (2) What prompted the authors to use the term "nest"? It gives the impression that they describe aspecific physical entity/structure (also depicted in this way in Figure 3P, with microtubules outside of this structure), but what is the evidence for this?

      The cloud is the spatial entity and the term “nest” is used to define a function of this transient compartment. We decided to keep the term “nest” as we now identified it with correlative light and electron microscopy, in addition to U-ExM, and show that the accumulation of centriole and deuterosome elements is accompanied by the formation of immature procentrioles, deprived of MT walls, as well as immature and empty deuterosomes. The scheme with MT outside the cloud/nest is misleading as we see MT organized by the mother centriole. We have now changed this.

      (3) The "nest" may simply be a dynamic accumulation of precursor particles around the centrosome, similar to what has been described for centriolar satellites. Rather than proposing a new entity, I suggest testing whether the "nest" particles may colocalize with PCM1 and thus may be related to centriolar satellites. Based on the data, the nest would simply be the centrosomal MTOC that organizes a radial microtubule array on which particles move around its center. In the absence of other evidence, I am not convinced that a new term is needed.

      We totally agree with the reviewer: the centrosome, as MTOC, concentrates centriolar and deuterosome components. This cloud is consistently dissolved when MT are depolymerized or dyneins inhibited. So, the physical entity is a “cloud”. We used the term “nest” to propose one function for this cloud which is to form deuterosomes and centrioles, before they move away for maturation. In fact, deuterosome and centriole formation are hindered when the cloud is dissolved. We have tried to edit the text all over the manuscript to make it clearer.

      (4) Role of MTs: are microtubules required or do they just facilitate some of the investigated events?

      The reason why the role of MT has not been tested yet during centriole amplification is probably because MT not only constitute the cell cytoskeleton on which molecular motors ride to transport cargos or distribute forces, they are also the core component of the structures we are studying. This is why we have tested a range of nocodazole concentrations and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (Fig. 4 Supplementary 1A-B). This may lead to an underestimation of the role of MT but we cannot study the role of MT on centriole amplification if centrioles cannot be formed.

      Does multi-ciliation in these models eventually occur normally under the concentrations and treatment conditions used here? This should be tested and discussed in the context of whether microtubules are indeed required and at what step of the entire process (amplification, migration, ciliogenesis) they may be critical.

      We did both chronic and acute treatments.

      Chronic treatments were done to test the overall efficiency of centriole amplification when MT (or dyneins) are perturbed. Chronic treatments were used to assess the role of MT (or dyneins) on the global efficiency of centriole and deuterosome formation (number of cells able to amplify, number/size/loading of deuterosomes, final number of centrioles (Fig. 4H-I, Fig. 4 Supplementary 2 B-D). In these chronic treatment, we focused on centriole amplification and not ciliation since it was the scope of this study. Also, we did not take ciliation as a readout of amplification because ciliation is relying on MT polymerization.

      Then, we also did acute treatments to test the role of MT (or dyneins) at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration; Fig. 4, 5, 7, 8 and associated supplementary figures). Since one stage is dependent on the precedent one, this enabled us to decipher the direct role of MT (or dyneins) on each single stage. We have now edited text, methods, legends and pictograms to be clear on whether acute or chronic treatment was done.

      (5) Can the authors include control (non-amplifying) progenitors in their analyses? It would be useful to know what the signal and distribution of each specific marker are before differentiation begins (before the cloud stage).

      Non amplifying progenitors are analyzed and constitute the so-called “centrosome stage”. We have now precised it and called it the “progenitor stage”.

      (6) Figure 2: Again, the terminology is confusing, since the authors describe that DEUP1 forms a "cloud" with centrin during the A stage.

      Corrections have been done as explained in point 1.

      (7) Description Figure 3: the authors introduce yet another term: "halo" A-stage. Is this the early A stage? Again, this is not explained and confusing. More systematic and consistent description is needed.

      Corrections have been done as explained in point 1. The term halos is used un the lab as it was the first term we used in our Nature paper in 2014 in reference to the halo described by Erich Nigg when they overexpressed Plk4. It was an error to use it in the manuscript.

      (8) Nocodazole treatments: the used concentrations are quite high.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely (Fig. 4 Supplementary 1A-B).

      (a) To avoid non-specific effects the authors should test what the minimal concentration is that completely depolymerizes microtubules in their cell model and perform analyses at this concentration.

      We have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A-B), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). In case it was not clear, we refer to this now several time and more clearly in the text and methods.

      (b) They should demonstrate depolymerization of microtubules by microtubule staining in the acute and chronic noc treatments and at the different noc concentrations used.

      This is, and was, in supplementary material (same, Fig. 4 supplementary 1A).

      (c) The authors should demonstrate that the used nocodazole concentrations do not impair normal centriole biogenesis during the cell cycle in these cells; if so, impaired assembly of centriole wall MTs may contribute to the observed effects in Figure 4.

      As mentioned in point 8b, we have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). The ability of the cells to form centrioles during chronic treatments were always assessed using immunostainings of SAS6 and/or CEN2-GFP signals (now exemplified in Fig. 4 Supplementary 1B). We also did EM analysis on cells treated with the highest doses of nocodazole (Nocodazole 10 uM for 24h) and this showed that centrioles can form with, what seems to be MT walls, in cells totally deprived of cytoplasmic MT fibers (Fig. 4 Supplementary 3-4). However, this does not show that all the cells can, because the number of cells that can be analyzed by EM are not sufficient to conclude. Also, one cannot assess whether MT walls are properly polymerized. However, the absence of MT walls should not change the results of the Figure 4, which are based on DEUP1, SAS6 or CEN2-GFP signals for deuterosomes and centrioles. Also MT depolymerization affects the formation of deuterosomes, which should not be altered by MT wall defects as it is not affected, even when centriole formation is blocked (LoMastro et al., 2024). Last but not least, we now show that blocking dyneins, as a comparable and even greater effect, on the formation of the cloud, deuterosomes and centrioles (Fig. 4C-I and Supplementary Fig. 4), which confirms that MTOC function, rather that MT wall formation, explain the centriole biogenesis alteration shown in Figure 4.

      (9) The authors repeatedly refer to the centriole-to-centrosome conversion of amplified centrioles and how this resembles centriole-to-centrosome conversion during the cell cycle. However, they incorrectly claim that this occurs at the G2/M transition. PLK1-dependent modification occurs at this stage, but conversion and PCM recruitment only occur after mitosis (see original work by the Tsou lab, which needs to be cited here).

      We agree with the reviewer. We have now added additional data to show clearly that centriole biogenesis, which requires two cell cycles to proceed in cycling cells, is accelerated during the MCC cell cycle variant where the elongation and maturation cycles are superimposed. This is now clearly shown in Fig. 3, 5, 9 and discussed.

      (10) Figure 6H-J: the authors claim that at low noc concentration, more D-stage cells showed incomplete disengagement than in controls, but the effect is shown only for the highest 10 µM concentration. Do any eof the phenotypes in Figure 6 also occur at the lowest noc concentration (assuming it depolymerizes MTs)? Again, it is crucial to demonstrate this, to exclude unspecific effects not linked to MT depolymerization.

      An error was made on the figure (but not in the legend). In Figure 6, chronic treatments are at 1 or 5 µM. Only acute treatments were done using 10 µM. In both cases, MT are not entirely depolymerized in these experiments (Fig. 4 supplementary 1A).

      (11) Disengagement, Figure 7: The authors describe that DEUP1 signal spreads all over the cytoplasm and becomes diffuse during this process, but one cannot see a diffusive signal throughout cells in the figures.

      We pushed the contrast to make it clearer but the deuterosomes are still bright at this stage and it is difficult to have both signal clear (now in Fig. 6B). We have also changed the example in video (now video 19) to show it more clearly with DEUP1 channel alone.

      (12) Figure 7: localization of disengaged centrioles at microtubule "nodes" is not clear from the images. There are many centrioles and random colocalization may be expected simply based on the high number. Higher resolution and/or magnification and quantification would be needed.

      We have edited and now say that centrioles “colocalize” with MT which, since centrioles nucleate MT, seems normal. We agree that it could be random, but given the density of MT, and the number of centrioles, it does not seem opportune to us to quantify. We can just say that we never see centrioles is regions that are deprived of MT.

      (13) The term "diffusive" to describe slow centriole movements in Figure 8 suggests that it is not motor or force-dependent, but there is no evidence for that. Movement based on opposing forces could produce a similar result, but would not be considered diffusive.

      We agree. We have changed “diffusive” by “diffusive-like”.

      (14) The manuscript would greatly benefit from the analysis of some candidate motor activities that may drive the movement and migrations of centrioles in this system. This would support the importance of the microtubule network for the specific steps in these processes, and better define its role beyond "being required". Dynein may be a candidate or minus end-directed kinesins. Since chemical inhibitors are available, these types of experiments would be straightforward.

      We formerly tested ciliobrevin but had hard time because of the small stability of the drug. Since our submission to eLife, we tested dynapyrazol and dynarestin and found dynapyrazol very efficient in dissolving the Golgi, a good readout of dynein inhibition. We sought to test the role of dyneins, using dynapyrazol, on (i) the formation of the pericentrosomal cloud in A-stage, (ii) the oscillation of DEUP1+ structures during A-stage, (iii) the number, size, loading of deuterosome, (iv) the final number of centrioles, (v) the migration to the nuclear membrane and (vi) the final apical migration of centrioles. The results are now inserted in main and associated Fig. 4, 5, 7, 8, 9.

      (15) Discussion:

      "the role microtubules" lacks "of"

      This is now edited.

      "This lack is..." Lack of what?

      This is now edited.

      "reflexive link" - meaning of "reflexive" is not clear in this context

      We have removed it.

      In my opinion, the study does not identify a nest composed of DEUP1, PCNT, and Centrin2; it only shows that these components accumulate as particles around the centrosome, which functions as MTOC. Consequently, it seems that the "nest" does not exist when MT is depolymerized. One could consider the center of the centrosomal MT array as a nest in this context, but there is no evidence of a specific new structure as suggested by the way the term is used in the manuscript.

      This is what we want to say: the center of the MT array become a nest in this context. We do not state that there is a specific new structure. We just say that MT and dynein dependent concentration of centriole and deuterosome components exists and that this region nests the birth of centrioles and deuterosomes. Also, this compartment is restricted in time and space, which justifies to use a specific term. The MTOC exists in the progenitor cell, while this compartment, marked by DEUP1, Centrin, PCNT accumulation, appears at the beginning of amplification and grows during A-stage to be dissolved at G-stage when all the deuterosomes and centrioles have moved away.

      What is the evidence that "DEUP1 is a centrosomal protein before building deuterosome structures"? It would be good to refer to the specific experiment. Does DEUP1 localize at centrioles also in the absence of microtubules? If not, I would not consider it a centrosomal protein.

      We have removed this statement to avoid misinterpretation.

      "This reminds the centriole-to-centrosome conversion..." the sentence is missing an "of"; also, again the authors confuse the order of events during the cell cycle, where centrosome conversion occurs after completion of mitosis, not at G2/M transition.

      We have removed this statement to avoid misinterpretation. Also, see Point 9.

      "microtubule dependent nuclear migration" should be rephrased; it sounds as if the nucleus migrates.

      This has been changed

      The following discussion of disengagement being linked to association with the nuclear envelope and resembling the process in cycling cells is misleading. In cycling cells movement of centrioles along the nuclear envelope occurs at G2/M and drives centrosome separation (separation of centriole pairs) in preparation for mitosis, not centriole disengagement.

      We are now clearer. We compare centriole-loaded deuterosome organization around the nuclear membrane to the migration of new centrosomes during early prophase (Fig. 5F-H, Fig. 5 Supplementary 2G-K).

      Regarding the possibility that forces by microtubules generated by the daughter centriole drive disengagement also in cycling cells, I would argue that this is unlikely since the daughter centriole can only nucleate microtubules after disengagement has occurred (and conversion to centrosome/PCM recruitment). Once this happens, it may physically separate the disengaged centrioles, which is a different type of activity. Indeed, originally the term "disengagement" was coined to specifically describe the loss of the perpendicular engagement of daughter centrioles with their mothers (Tsou and Stearns, Nature, 2006).

      We have removed this statement to avoid misinterpretation. The perpendicular engagement is difficult to assess on deuterosomes but we do see by live imaging, that attachment changes during D-stage, before centrioles detach clearly from deuterosomes.

      "high resolutive" should be "high resolution"

      Edit done.

      "splitted" should be "split"

      Edit done.

      "Consistently, when the mitotic oscillator is dis-inhibited and cells enter pseudo-mitotic events, centrioles show clear and rapid cell-cycle like clustering" This sentence is not understandable without further explanation; what does mitotic oscillator refer to? What are pseudo-mitotic events? What is cell cycle-like clustering?

      We have removed this statement.

      Minor:

      (1) Abstract: "Centriole number must be restricted to two..." Since cells are born with two centrioles and have 4 centrioles (2 pairs) when they enter mitosis, this sentence is inaccurate.

      The sentence has changed.

      (2) Abstract: "reflexive link"; I am not sure what the term "reflexive" refers to?

      We have removed this statement to avoid misinterpretation.

      (3) Figure 1C, D: it should be described better that the larger magnification panels represent overlays of many cells and what marker they show. This is not obvious since the smaller single-cell panels always show two different markers. Also, it would be more useful to show also single cells in the magnified view. The overlay does not allow us to see if a marker forms a cloud or a single dot, which is as important as the cell-to-cell variation in distribution.

      We have clarified this in the text and the legend. The cell-to-cell variation cannot be estimated with the overlay, but the projection from several cells (number precised) allows to see that the signal is confined in a restricted region. Or not. Which is what we wanted to analyze.

      Related to the above, the authors say that pericentrin forms a cloud at the top left in panel D, but there is only one confined centrosomal dot in the single-cell panel.

      The sentence has changed.

      (4) Results, Figure 2F; video 4: The authors claim connection and disconnection of DEUP1 aggregates with centrosomal centrioles; can the authors comment on the spatial resolution including in z in this movie to support this claim? Can they exclude that the structures are in proximity of each other rather than "connected"?

      This is a single z-section of 500nm. The resolution in xy is 128nm/pixel. Given the sizes of deuterosomes and a mature centriole, and given the fact that we observed this dynamics in several cells in live, we can state that the structures are connected. This is consistent with deuterosomes frequently observed “kissing” the daughter centriole by EM in the present manuscript (Fig. 2D, Fig. 2 supplementary 3 and 4 and Fig. 4 Supplementary 3-4). One has to look carefully at the daughter centriole (marked “dc”) and span in on the serial sections to see the connected deuterosome (marked by a star): this is at very early stage and therefore it is small. We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both structures are frequently sticked to each others on tens of nanometers.

      (5) The term "dynamics" as used in the manuscript should be plural.

      It has been used plural, except when for “dynamic microtubules” and “dynamic attachment to the nucleus”, which we think is ok? We have not found any other singular uses in our manuscript.

      (6) Figure 5: what does "YL1/2 procentriole intensity" refer to in panel F? This should be the intensity of microtubule asters.

      This has been modified.

      (7) Figure 6 - supplement 1B: contrary to the claim in the text, one cannot see tight colocalization with the nuclear pore marker. This seems to be a very small subset of particles and even in those cases colocalization is not tight. Also, what is the relevance of nuclear pore colocalization?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. What we want to say is that there is a tight connection with the nuclear envelope as shown by the localization of NPC on the same z-section as centrioles. This is why we present a single z, to show that centrioles and NPC are on the same z-plane of 500nm. NPC are stained to outline the nuclear membrane. This is also clearly visible for G-stage centrioles in the XY plane. We have now added an entire z-stack on video 18.

      Reviewer #2 (Recommendations For The Authors):

      To improve accessibility of their manuscript, we would suggest making the following edits:

      (1) Define 'specialist' or 'niche' terms each time you introduce them, such as 'pericentrosomal nest', or 'flower-like structures'.

      This has been clarified.

      (2) Have a think about abbreviations, again ones that work for people outside the project- this paper uses 'PC' for 'procentriole' but for many 'PC' is 'Parental centriole' or Figure 6J talks about 'D total' or 'D partial', leaves readers confused.

      This has been clarified.

      (3) Standardize your abbreviations throughout particularly for your treatments- sometimes Noco sometimes, NOCO, or your imaging experiments sometimes Cen-GFP, sometime CEN2-GFP (Figure 7A, D vs. Figure 6) or DEUP1- mRuby, DEUP1-mRuby3 or mRuby3-DEUP1?

      We now use Nocodazole or Noco in the text and the figure respectively, CEN2-GFP and mRubyDEUP1.

      (4) About 10% of the population, including several key figures in this field, are red-green color blind. Although 4 colour fluorescence is difficult to get right for everyone, choosing palettes (especially for two colour panels) is inclusive. More so, greyscale or inverted monochrome images make it easier for everyone to visualize changes in localization, size, and intensity. Red on black small foci is particularly difficult to discern. For example, Figure 3 - more individual channels in grayscale with arrows to mc, dc, and cilia would be helpful - difficult to distinguish stainings.

      We thank the reviewer for this comment and for this recommendation of being more inclusive. We have done the changes.

      To improve the conclusions drawn, we suggest some revisions below:

      (1) Since the paper really hangs on it, a clearer description of the rationale for when, how long and how much nocodazole treatment was done is needed. The logic currently is difficult to follow seemingly random jumps 10x concentration are used. Microtubules control many aspects of cell biology and could be impacted. For example, I particularly found Figures 6D and H difficult to follow i.e. the timing for 6H seems off.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely. We have of course tested a range of nocodazole concentrations at the beginning of the study and shown the extent of MT depolymerization under each treatment. We used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4 reviewer 1). The level of perturbation of MT and consequences on centriole formation at the different timings and doses were done for each experiment and are exemplified in Fig. 4 supplementary 1A-B. This figure was already present in the first version of the manuscript but we have now edited text, methods and pictograms to clarify this.

      (2) Perhaps an extension of this point- in general how interdependent are the processes? If there is a defect at the nest stage, how much are the later defects secondary to this, or do MTs genuinely play direct roles at all stages or are these knock-on effects? How do the authors rule this out? Defects in the nest, lead to smaller and more DEUP1+ foci, with defects in concentrating procentriole factors and centrin, which lead to... For example, Figure 4B looks like centrin is reduced upon noco treatment? Does noco treatment affect Cetn2GFP levels globally? Individual channels grayscale would help visualise this better.

      See also our answer to reviewer 1 point 8c.

      The stages are indeed interdependent. This is why we did both chronic and acute treatments. Chronic treatments were done to test the overall efficiency of centriole amplification when MT are perturbed. We typically used low dose of 1µM because nocodazole remains 48h in the culture medium. Acute treatments were done to test the role of MT at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration). Most of the acute treatments were done live and nocodazole was applied after the first time point of live monitoring. We used 10µM to have a rapid effect, and because nocodazole remains only several hours in the culture medium. This allowed to monitor the stage “n”, in cells where the stage “n-1” was completed without any drug which allowed to analyze a stage without having perturbed the precedent one.

      We now also test the consequences of dynein inhibition using both acute and chronic dynapyrazole treatments. We show that except for centriole migration, dynein inhibition phenocopies MT depolymerization (centriole number, perinuclear organization and disengagement as well as deuterosome number/loading/size).

      Nocodazole chronic treatments do affect intensity of CEN2-GFP at G-stage centrioles suggesting an altered A-to-G transition. In D-stage, CEN2-GFP signal seems normal. We now mention this in the text and in the Fig. 4 Supplementary 1B.

      (3) The authors nicely show the importance of MTs in the structure of the nest from which procentrioles and DEUP1 positive structures emerge. They suggest this nest may be what supports procentriole generation in the absence of DEUP1 and parental centrioles. Firstly how does this nest look in the absence of DEUP1 and/or parental centrioles (centrinone treatment)? This may be what they are trying to show in Figure 5 Supplement 1 but it currently is very difficult to digest what it is showing relative to controls and whether this is significant in the way it is plotted.

      The nest is conserved in the DEUP1KO with or without centrosomal centrioles, as shown by accumulation of Centrin and PCNT at the center of the self-organised MT network (Mercey et al., 2019). This is in fact what motivated our study on the role of MT in centriole amplification. We have edited the legend to precise the quantification done, which is not related to this question. In this quantification, we show that the increased propensity to accumulate PCNT by centriole-loaded deuterosomes between A and G-stage is maintained in the absence of deuterosomes, indicating that centrioles themselves accumulate/recruit PCNT.

      (4) Can you do CLEM on DEUP1-Ruby and these early foci at the cloud stage to see if they are visible at the ultrastructural level, relative to procentrioles, microtubules, and other electron-dense structures?

      We thank the reviewer for this question. We have done CLEM on the pericentrosomal cloud during very early steps of centriole amplification. This showed that DEUP1 early accumulation at the centrosome corresponds to a region rich in fibro granular aggregates, suggesting that DEUP1 may be translated here, through locally concentrated centriolar sattelites, known to be involved in local translation. Then, small deuterosomes and immature centrioles are formed, within this cloud of sattelites, confirming that the pericentrosomal cloud is a nest for centriole biogenesis (Fig. 2C-D + Fig. 2 Supplementary 2-6 for control and Fig. 4 Supplementary 3-4 for nocodazole treated cells). This also shows that immature deuterosomes are not necessarily round shaped, and can be deprived of centriole loading.

      (5) Check the scale bars- see Fig 4E. Check throughout.

      Done.

      (6) Figure 3 Supplement 1 and 2 don't match the legend and are likely reversed - which one is right?

      Done.

      (7) Technical issue - I couldn't play videos 6 or 16? Check these work.

      Done.

      (8) Nomenclature mammalian proteins- mouse or human- should be all caps DEUP1, PLK4, SAS6,etc. Watch your units- space between number and unit.

      This has been done.

      (9) Many of the graphs involve three biological replicates but why not plot the mean of each of the three experiments and do stats? The number of events measured may conflate the significance. Try using Superplots.

      Here is how we proceed: we count the number of occurrence of the phenotype we monitor, and the total number of cells. We apply a X<sup>2</sup> to test whether there is a significative difference between our replicates in each condition. If not, we pool the number of occurrence of the phenotype we monitor and the total number of cells for the 3 replicates, and for each condition. Finally we apply a X<sup>2</sup> between the different conditions. This is how we usually proceed to avoid comparing a mean of percentages. This is now explained in the methods.

      Minor points:

      (1) "DEUP1 is a centrosomal protein and assembles deuterosomes in the pericentrosomal region in brain MCC". I am not sure you have evidence that DEUP1 is a centrosomal protein. You don't seem to study the relationship between centrosomes and DEUP1? Rewrite this title and tone down this claim.

      This has been modified.

      (2) Why the crossbow micropattern (versus some other shape) - seems very specific but not discussed?

      We wanted a shape where centrosome is not localized at the center of mass of the nucleus. Among the corresponding patterns, the crossbow was the one where differentiating cells had less propensity to detach.

      (3) Figure 2 - are the foci of DEUP1 at the cloud stage smaller than at A stage? How do they grow? Measure the diameter at cloud stage, just after they leave the cloud and then once they move away from centrosomal cloud and each other. If so, and they do indeed grow in size from the cloud stage to the growth stage which I think your images suggest - do you envision this happening with the gradual addition of DEUP1 rather than fusion?

      Early deuterosomes are not easy to detect by light microscopy, because of accumulation of DEUP1 in the cloud. We did CLEM on the cloud of early A-stage cells to resolve the earliest deuterosomes which are often very small (see Fig. 2D, Fig. 2 Supplementary 2-6) suggesting that they grow, either by fusion, which we never observe in our movies at later A-stage, or by accretion of DEUP1. However, by light microscopy, we can detect very early but big deuterosomes, which we see splitting later on into smaller ones. So, we cannot conclude on the mechanism that regulate deuterosome size. This is now discussed in the discussion of the manuscript.

      You say in the discussion:

      "Consistently, we never observed fusion events of DEUP1 condensates in our time-lapse experiments. More importantly, we did FRAP experiments on endogenously tagged mRuby-DEUP1 in cells at the different stages of centriole amplification, and did not find significant recovery, supporting that centrosomal DEUP1+ foci and deuterosomes are not liquid-like structures (Figure 8 Supplementary 2)." How do you prove there is no fusion of deuterosomes?

      It is always difficult to prove the absence of something, we agree! But we did tens of movies with high temporal resolution and never observed fusion events. But, as we say in the previous question, the very early deuterosomes can be very small and we do not distinguish them from the DEUP1+ cloud by live imaging. So at this stage, we cannot say. But later on, during A- or G-stage and when deuterosomes are outside the cloud to be easily observed, we very often observe deuterosomes bumping into each others and stay in close contact for minutes, but then moving away. This, for us, supports the lack of fusion properties. But the question remains open. We now explain this in the manuscript and have added an example in video 28.

      If they are getting bigger as I think your imaging suggests from cloud to growth stage, then how is this happening?

      MT depolymerisation and dynein inhibition leads to the formation of very small deuterosomes. Dynein inhibition can even lead to a block in the formation of new deuterosomes suggesting that DEUP1 concentration is a crucial parameter for condensation into deuterosomes. Deuterosome growth may happen through oligomerization of DEUP1 molecules allowed by their dyne-independent concentration. Sorokin in 1968 proposed that a supersaturation of deuterosome components may lead to their solid crystallization into deuterosomes. Deuterosome size can also be regulated by a more complex molecular cascade, involving post-translational modifications of DEUP1 or PCM, such as phosphorylations driven by the cell cycle machinery. This would be consistent with the fact that deuterosomes are very big in the absence of CCNO, a cyclin required for entering the MCC cell cycle variant. This will need further investigations.

      I'm not sure FRAP actually proves fusion doesn't happen.

      Agreed, this is not what we wanted to say, we clarified. The FRAP experiment just suggests that it is not liquid-like.

      It is technically difficult to laser ablate individual or only subsets of deuterosomes...

      This is what was done but anyway, FRAP does not firmly show that deuterosome compartments are not liquid-like as we now precise.

      (4) How do you fix your cells for expansion as you have no preservation of cytoplasmic microtubules? You are saying that there is a "nest" of MTs but beta tubulin ONLY stains the cilia and centriole - why is this? Tyrosinated tubulin on regular confocal shows strong cytoplasmic staining. See Figure 3.

      Cytoplasmic microtubules do not preserve well through the expansion process. We did try a few different fixations and pre-extraction methods but they come at a trade-off to preserving centrioles. i.e. we could either preserve cytoplasmic tubes or centrioles but not both with the same processing method.

      (5) "PCNT puncta partially overlap with centrin (Figure 3 Supplementary 2C). At this stage, PLK4, the master regulatory kinase, and SAS6, one of the first centriolar components are either absent or present as small foci within the cloud, often on the wall of the parent centrioles (Figure 3B-C)." some arrows to highlight this would be useful - difficult to see?

      We have tried to make arrows on what is now Fig. 3 Supplementary 1 G, but there is to many CENTRIN colocalizing with PCNT. We have enhanced the contrast of the merge to make it more visible.

      (6) Figure 3I legend - what are the arrows pointing at? Yellow and white on inserts? ". Around the same time as tubulin, centrin is also recruited to procentrioles (Figure 3I). This stage is probably the stage that we previously documented as A"

      However you see centrin at DEUP1 foci in D, and you don't show any eg. SAS6 or PLK4 positive DEUP1+ structures lacking centrin specifically, centrin seems to be present on all the procentrioles in Figure 3I. Did I miss it where you show centrin negative procentrioles in the cloud?

      Fig. 3I (now Fig. Supplementary 1J), yellow arrows are pointing at centrioles with non-acetylated MT while white arrows point at acetylated MT. This is now indicated in the legend.

      Regarding CENTRIN, it is present as a diffuse staining around the centrosome since the very beginning of amplification (now in Fig. 3 Supplementary 1A with different contrasts), in addition to compose the parental centrioles. This staining can therefore overlap with DEUP1 staining when DEUP1 appears (Fig. 3 Supplementary 1B, E) but not necessarily. In live we observe that CENTRIN and DEUP1 foci can move independently at early stages (Fig. 2 Supplementary 1B, video 2). This is later on, as shown now in Fig. 3 Supplementary 1J (previously Fig. 3I), that procentrioles are all strongly positive for CENTRIN.

      A new paper (Laporte et al., Cell 2024) recently showed that the recruitment of CENTRIN on duplicating procentrioles first occurs at the distal end, visible by a small dot, and then appears gradually at the level of the inner scaffold when procentriole reach 160nm, the stage where POC5 appears, which corresponds to the A-to-G transition in our MCC progenitors (Al Jord et al., 2014). One can therefore consider that the same is happening in our cells, and that, with the CENTRIN cloud, we have difficulties to detect the distal CENTRIN dot. We have changed the text to add this reference and discuss CENTRIN apparition in MCC procentrioles.

      (7) " The DEUP1 asymmetry previously described at the centrosomal daughter centriole (Al Jord etal., 2014) becomes visible in some cells during the cloud stage (Figure 3B, N; Figure 3 Supplementary 2B) and in a majority of cells" difficult to see - maybe enlarge and single channel from Figure 3F-H in the supplemental Figure 3 to emphasise this?

      We have either changed the pictures or the contrast to be more representative with the quantifications. This is visible in Fig. 3A, D, E, G; Fig3. Supplementary 1E and now using correlative light and EM in Fig. 2 Supplementary 2, 3, 4 and Fig. 4 Supplementary 3-4. One has to look carefully at the daughter centriole (marked “dc”). We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both strutures are frequently sticked to each others on tens of nanometers.

      (8) Do you have videos of DEUP1 oscillations with nocodazole to show a lack of oscillations?

      We have now added videos of DEUP1 oscillations under nocodazole and dynapyrazole treatments.

      (9) "In addition, co-staining of centrioles and nuclear pore proteins show a tight colocalization(Figure 6 Supplementary 1B)." I see the colocalisation in panel 1 but less obvious with panel 2 maybe have some more zoomed in panels and some quantification of the colocalization? Is it more striking at the G stage than the D stage?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. There is too many centrioles and NPC, they cannot do otherwise than colocalize… What we want to say is that there is a tight connexion with the nuclear envelope. This is why we present a single z, to show that centrioles and NPC are on the same z-plane. This is also clearly visible for centrioles that are loaded on deuterosomes that are around the nuclear membrane in the XY plane. We also added a video to show an entire z-stack of this kind of staining.

      (10) "Indeed, SAS6 normally disappears from procentrioles when centrioles are docked, just beforeciliation (Al Jord et al., 2014). This suggests that centrioles were able to degrade SAS6, a process also dependent on APC/C (Strnad et al., 2007), but failed to disengage from deuterosomes." Figure 6 Supplement 1E-F - are you sure it wasn't that Sas6 wasn't loaded correctly at the earlier stage and so is reduced recruitment rather than premature disengagement of Sas6? If it is indeed premature disengagement of Sas-6 - what about CP110 - does the CP110 get loaded and is it still present in noco treated cells arrested in the D phase?

      We do not observe SAS6-negative procentrioles on deuterosomes at G-stage but only on deuterosomes in D-stage cells (cells with partly disengaged procentrioles). This is why we hypothesize that, because of the long duration of D-stage and knowing that SAS6 is finally degraded at the end of amplification (Al Jord et al., 2014), we are in the presence of cells where SAS6 has been degraded but where centrioles did not manage to disengage. This is now clarified in the text.

      (11) Can you track deuterostome splitting live? Maybe not enough spatial or time resolution?

      One has to monitor in 3D (multiple z because deuterosomes move a lot), 2 colors, high temporal resolution (dt=2-5’; to be able to track a single deuterosome), and long duration (deuterosomes are sometimes touching each other and then moving away, giving the impression that they split). This eventually leads to the bleaching of the mRuby fusion protein… We have put an example of what we think is a deuterosome splitting in Fig. 6E (former Fig. 7D). But we decided to finally monitor with low temporal resolution (dt=40’) to avoid photobleaching, and analyze numerous deuterosomes and cells to quantify the number and size of deuterosomes over time in single cells.

      (12) The MT nodes - can you segment the tyrosinated MTs and define nodes and then quantify theDEUP1 presence on them?

      Please see answer to reviewer 1 regarding this point.

      (13) Figure 8 supp 1 (E): Representative XY distribution of CEN2-GFP+ centrioles at the end of migration (Sas6 negative) in brain MCCs treated with DMSO, Nocodazole 1µM and 5µM (48h). Scale bar, 5µm Bit more detail on how you define fully migrated vs still migrating centrioles in z. You say you are using Sas-6 negativity to define fully migrated cells in the legend, yet you say noco treatment leads to premature sas-6 negativity, and yet the apical migration takes longer upon noco treatment?

      Nocodazole does not lead to premature SAS6 negativity but to a partial disengagement which lead to SAS6 negative “mature” centrioles being still connected to deuterosomes. We define complete migration when all the centrioles are on the apical side of the nucleus. We now clearly define what “apical” migration stands for in the main text and changed the pictograms in Fig. 8G to clarify this.

      (14) Figure 8H and video 18 - it isn't obviously clear to me that the noco-treated cells are "more erratic" or how you decide what counts as apically migrated successfully. How do you control for drift in z? Can you track individual centrioles as you did in untreated and define what is "erratic about their movement?

      Erratic means that the centrioles are moving away from each others, and back, in a non-predictable way, instead of migrating up and gathering. The drift in z of the whole cell is visible because there is always some centrioles, that are apically located at the beginning, that remains on the apical membrane, probably because they are already docked.

      We have indeed followed the centrioles individually in the nocodazole condition. However, in the control, the XYZ coordinates of one of the centrioles of the centrosome, which normally don’t move, are substracted to the coordinates of all the other centrioles as explained in the method section. This allows to have a subcellular reference, and to circumvent the movements of the cell, which are non-negligible at all at this timescale. In the nocodazole treated cells, the centrosomal centrioles share the erratic movements of the other centrioles and can migrate up and down, which exclude them as a reference. Since the nucleus is also moving a lot, we were left with no reference point.

      (15) Figure 8 supplement 1E can you quantify the final area of centriole patch in XY upon noco treatment?

      It was in main Fig. 8J and is now in Fig. 8 Supplementary 1F.

      (16) Figure 8J legend- MBB is never defined as an acronym.

      Thank you for pointing this.

      (17) Define what is the frequency and how is it calculated - Figure 8J.

      This is the MBB patch area in µm<sup>2</sup>

      Text edits:

      (1) "Altogether, these results suggest that, in this non-tissue-specific proxy of MCC progenitors, microtubules organize the onset of centriole amplification in the pericentrosomal region."

      Sentences have changed.

      (2) "Increasing the temporal resolution to 5-15s reveals that DEUP1+ foci observe an exhibit oscillatory dynamics to at the centrosome (Figure 2E, colored arrows, Video 3, 5/10 cells observed for 1-4min)."

      Sentences have changed.

      (3) "stage procentrioles were involved in this perinuclear migration and distribution. In fact, this dynamic is reminiscent of the centrosome migration that occurs during the G2-to-M progression in cycling cells in preparation for mitotic spindle organization. In cycling cells, this" Grammar - maybe change to "stage procentrioles were involved in this perinuclear migration and distribution. This is reminiscent of the centrosome migration that occurs during the G2-to-M".

      Sentences have changed.

      (4) "We then wondered whether these microtubule-dependent dynamics was were required for an efficient subsequent centriole disengagement during the following D-stage."

      Sentences have changed.

      (5) "Then, monitoring tens of disengagement movies, we identified a transient stage during which disengaging procentrioles redistribute isotropically in the 3 dimensions, along the nuclear membrane (Figure 6A, 4:30, Video 7) before losing its contact to migrate to the apical surface (Figure 6A, 6:30 to 14:00)."

      Sentences have changed.

      (6) Discussion: "Since pioneer electron microscopy studies on basal body production in quail oviduct MCC 35 years ago (Boisvieux-Ulrich et al., 1987, 1990; Boisvieux-Ulrich et al., 1989), this work is the first to assess the role of microtubules in the now finely described centriole amplification process. This"

      Sentences have changed.

      (7) "Using live imaging on brain MCC, we highlight the existence of a nest composed of DEUP1, PCNT and Centrin2, pre-assembled before the onset of centriole amplification onset."

      Sentences have changed.

      (8) "Recently, formation of DEUP1 pure condensates in solution as well as FRAP experiments after overexpression of DEUP1 in MCC progenitors suggested that deuterosomes where are not liquidlike structures (Yamamoto & Kitagawa, 2019). Consistently, we never observed fusion events of DEUP1."

      Sentences have changed.

      (9) "This reminds is reminiscent of the centriole-to-centrosome conversion occurring at the G2-M transition followed by the associated microtubule dependent nuclear migration of new centrosomes at mitosis onset (Agircan et al., 2014)."

      Sentences have changed.

      (10) "Following individual trajectories requires high resolutive resolution spatio-temporal live imaging while avoiding excessive light exposure which disturbs centriole migration (Boudjema et al., 2024)."

      Sentences have changed.

      (11) "Using high temporal resolution microscopy, we further identify that individual dynamics is are complex and can be splitted between divided into the baso-apical migration, where centrioles move in a processive and more..."

      Sentences have changed.

      Reviewer #3 (Recommendations For The Authors):

      (1) Growing MEF-MCCs on micropatterns has successfully mimicked the dynamics of centriole amplification in brain MCCs, allowing the authors to study the spatial origin of procentrioles. Since this is a powerful system, a more quantitative description of the system will be informative and beneficial for future studies. For example: What is the efficiency of this system? Do the cilia that form in MEF-MCCs motile?

      The system of MEF-MCCs has been described in a previous paper from the Kintner lab. It seems that growing the MEF-MCCs on micropatterns did not ameliorate the ciliation which is partial, probably due to the absence of an apico-basal polarity.

      (2) Figure 2: The analogy drawn by the authors between DEUP1 oscillatory dynamics and centriolar satellites is intriguing. In early amplifying cells within the cloud, do these DEUP1 structures co-localize with the satellite marker PCM1?

      We have added immuno stainings of PCM1 in mRuby-DEUP1 / CEN2-GFP cells in Fig. Supplementary 2E. Within the centrosomal cloud, DEUP1 colocalizes with PCM1. Interestingly, this PCM1 concentration at the centrosome is dependent, at least in part, on dyneins. Then, PCM1 can localize around the deuterosomes, but it is never colocalized with deuterosomes (not shown). This is also showed by immuno-EM in Zhao et al., 2019. Although it was shown that PCM1 is a proximity interactor of DEUP1 (called ccdc67 at that time) by Firat-Karalar et al., 2014., absence of PCM1 staining on deuterosomes does not favor the hypothesis of PCM1 and DEUP1 being part of the same entities. One could hypothesizes that DEUP1 is transcribed locally within the satellites, explaining the colocalization of the 2 proteins and the + BioID results, and then form PCM1negative deuterosomes.

      (3) The authors propose a physical link between deuterosomes and centrosomes based on their oscillatory behavior. How are the oscillatory dynamics of DEUP1 affected by nocodazole treatment or inhibition of microtubule motors (i.e ciliobrevin treatment)?

      These oscillations are inhibited by nocodazole (Fig. 4D). They are also inhibited by dynapyrazole (Fig. 4D). We never succeeded in having a nice disruption of the Golgi apparatus with ciliobrevin and therefore we did not used it.

      (4) In addition to nocodazole treatment, it would be important to determine the consequences of microtubule stabilization by taxol and inhibition of microtubule motors during critical stages of centriole amplification where microtubules are reported to play a role for the first time in this manuscript. Another interesting area of investigation will be to study the extent to which microtubule PTMs contribute to these processes.

      We now blocks dyneins during the different stages of amplification. The results are in main and associated Fig. 4, 5, 7, 8. The role of microtubule PTM, is not in the scope of this manuscript.

      (5) Describing microtubule dynamics along with Centrin/DEUP1 dynamics will be informative in assessing whether these structures associate and/or move along microtubules? Have the authors performed their imaging experiments with SIR tubulin?

      Yes, we have tried hard! But we have encountered different obstacles:

      3-color video microscopy is phototoxic,

      siRTubulin is bleaching very rapidly

      The density of microtubules in MCC makes the observation hardly informative

      (6) Figure 5: The role of PLK1 in centriole-centrosome conversion and generation of multiple MTOCs can be tested with a PLK1 inhibitor for further confirmation.

      We have also tried but inhibiting Plk1 blocks the A-to-G and G-to-D transitions so it was not possible to uncouple the role of Plk1 in stage transitions versus centriole maturation.

      (7) Figure 6: The tight co-localization of nuclear pore proteins with centrioles poses questions about the role of nuclear pore proteins or other nuclear proteins that are associated with centrioles during centriole disengagement and migration. Considering the existing literature on centrosome-nucleus attachments, can there be a way to test this question within the scope of this manuscript?

      We have tried to deplete Nup133 but it’s killing the cells. Our additional experiments now show that the nuclear migration of centrioles during G-stage is dynein dependent, reinforcing the parallel with centrosome migration in prophase. We also added results from our scRNA sequencing (Fig. 5 Supplementary 1) showing that some key players of centriole migration to the nuclear membrane are conserved in the MCC cell cycle variant, and expressed with a comparable dynamics as to the canonical cell cycle.

      (8) Figure 8: Manually tracking a subset of migrating centrioles to define their dynamics during centriole migration and docking provides valuable analysis for determining the molecular mechanism of these processes. In addition to microtubules, does actin contribute to this process? Since centrioles eventually migrate to the apical side in nocodazole-treated cells, there should be other molecular players involved in this process.

      We did block actin polymerization but we found that the different stages were affected and that it would be better to dedicate a whole manuscript on the role of actin during each stage of amplification. We discuss the migration mechanism, and the putative role of actin, in the discussion.

      (9) The legends for Supplementary Figures 1 and 2 in Figure 3 are mixed and need correction.

      Figures have been remodelled.

      (10) In Figure 3P, the term "PLK4+" is labeled in bright green, which is not clearly visible. It maybe beneficial to change the color of this label for better visibility.

      We have tried to correct this.

      (11) Figure 6F quantifies "% tethered flowers" on the nuclear membrane. When quantifying, is the3D localization of DEUP1 flowers in both DMSO- and Noc-treated cells considered? A flower may appear to be on the nucleus in 2D, but it could be detached from the membrane in a 3D view.

      The quantifications are done in 3D. However, flowers that are below or above the nucleus are not quantified since the space is confined and the resolution in z to small to see whether they are connected or not. This is now precised in the legend.

      Before the editors proceed with an updated assessment, they've requested that we pass on some of the comments that have arisen as part of the evaluation of your revised manuscript. They feel that these concerns should be addressed before we proceed with issuing a formal assessment and publishing the revised Reviewed Preprint:

      We thank the reviewers and the editors for the corrections and insighfull comments. We apologize for our delayed answer and hope our corrections in the main text and some of the figures will give them satisfaction.

      The revised manuscript is greatly improved with nice new data regarding the role of microtubules. It also has changed quite a bit including the title. The new focus is on the cell and centriole cycle variants in MCC. While this helped to focus the study, there remains an important issue related to the interpretation of the data and the proposed 2-in-1 cycle model. Before providing the final updated assessment, we ask you to address the following points (which were raised already in the first round of review): The manuscript still contains statements that are not aligned with published work and the current view in the field regarding the timing of events during canonical centriole biogenesis. These timings are in conflict with your model that 2 centriole cycles are "superposed" in the MCC cell cycle variant, as currently presented. An alternative straightforward interpretation would be that multiciliogenesis uses an accelerated centriole duplication cycle where key steps occur concomitantly or in short succession instead of being separated by mitotic divisions as in the canonical cycle.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we think we have proposed. When correcting our confusions as regard to centriole-to-centrosome conversion (as explained below) and putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (corrected Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a superposition of events that; although driven by the same molecular machinery, are normally occuring in two consecutive cell cycle. We explain ourself briefly in two paragraphs, before answering point by point to the questions of the reviewers.

      As regard to centriole-to-centrosome conversion:

      We thank the reviewer for pointing out that we used “MTOC conversion” for what is normally called “centrosome maturation”. We have removed the term “centriole-to-centrosome conversion” during the first round of revision but we now realize that “MTOC conversion” leads to the same misinterpretation as regard to the literature on centriole duplication.

      The reviewer asks us to refer to the work of the Tsou lab (Wang 2011, reference now added in the manuscript) showing that daughter centrioles are “modified” (e.g. recruit PCM, become competent for MT nucleation and duplication) during late M/early G1. This “centriole-to-centrosome conversion” can’t occur for our procentrioles at this stage since they are not even born during the mitosis that precedes MCC differentiation. Also, in our cells, such modification does not include the capacity to become competent for duplication since we know that procentrioles become basal bodies without making any round of duplication (Al Jord et al., 2014).

      Also, we have not done the experiments to tackle the question on when our centriole become “modified-like”. What we can say is that during A-stage, they become progressively positive for PCM (Fig. 5 Supplementary 2) and a weak signal shows that some MT are seen emerging from them (Fig. 5 and Fig. 5 Supplementary 2, and see point by point answer).

      What we do see is that, at the A-to-G transition, they increase their PCM recruitment, show clear and strong MTOC ability (sometimes as strong as the centrosomal centrioles), and that this is associated with migration and separation of centrosome/deuterosomes around the nuclear membrane (Fig. 5). We therefore connect this to what occurs at the G2/M transition which is an increased recruitment of PCM protein, an increased ability to nucleate MT, associated with centrosome migration and separation at the nuclear membrane. Since this process in the canonical cell cycle is called “centrosome maturation”, we therefore should refer to this term in our study. However, centrioles in the MCC variants are not organized in centrosomes, so we now compare what we see to the “centrosome maturation” of the canonical cell cycle with an associated reference (Joukov et al., 2018), but name it “centriole maturation”.

      We have modified the text (track changes visibles) and the schemes (Fig. 5, Fig. 5 Supplementary 1 and 2, Fig. 9, Fig. 9 Supplementary S1; new versions uploaded) accordingly.

      As regard to 1.5 or 2 cell cycles

      Except for the “MTOC conversion” that we have now changed, as explained above, we think our work does suggest (depicted on Fig. 9) what the reviewer states for centriole duplication: “In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs. Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium)”.

      We feel that going from early S to a G1 phase, after 2 mitosis, is what one can call “2 cell cycles”. One of the paper that inspired us a lot when studying how the cell cycle machinery can drive centriole amplification in MCC is a paper from Jadranka Loncarek team (Kong et al., 2014) where they also state that “nascent centrioles gradually mature through 2 cell cycles”. Very interestingly, in this study they show that when they enhance Plk1 activation, they could erase centriole age and new procentrioles are able to recruit PCM and appendages within only 1 cell cycle, without mitotic progression, like what we see in MCC. We have added the reference in our discussion.

      Point by point answer

      (1) Original work on canonical centriole disengagement and centriole-to-centrosome conversion should be cited (e.g. PMID: 16862117, PMID: 21576395)

      As explained earlier, we used the wrong term since the begining. We do not speak about the centriole-to-centrosome (nor MTOC) conversion since we do not test when centriole modification (Wang et al., 2011) occurs in the MCC cell cycle variant. We know that PCNT is present on the procentrioles during A-stage (as shown in Fig. 5 Supplementary 2B), but we do not know when it is recruited (UExM did not work properly with this antibody). We quantify a weak MT staining in regrowth experiment during A-stage and see that procentrioles can be connected to MT in both brain MCC and MEFs (as shown in Fig. 5D, E for brain MCC and Fig. 5 Supplementary 2F for MEFs) , but we do not know when during A-stage they become competent for nucleation. We therefore did not speak about this process that we do not document. What we clearly document/quantify is the enhanced MT nucleation capacities at the A-to-G transition, concomitent with the nuclear migration (easily defined with Cen2-GFP or GT335 stainings) and that we compare to centrosome maturation occuring at the canonical G2/M transition.

      (2) The authors state in several places that canonical centriole formation and maturation takes two iterations of the canonical cell cycle. This is imprecise. Based on the above work and work by others, the broadly accepted view is that it takes 1.5 cell cycles. This difference matters for the final proposed model (see below). Reviewed e.g. here: PMID: 20869612; PMID: 30601682

      Our answer is in the preamble.

      (3) "Centriole maturation cycle superposes with centriole elongation cycle in the MCC cell cycle variant": Your description of the canonical cycle differs from the current view in the field. In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). All this occurs in 0.5 cycles. Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs (total of 1.5 cell cycles). Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium).

      (4) Fig 5A, B and Fig. 9

      (a) Are 2 separate figures needed for the model? They seem redundant.

      We find it easier not to wait Fig. 9 to have the first part depicted.

      (b) The model shows loss of SAS6 throughout G1, but this already occurs during M/early G1

      Thanks. It was already ok in Fig. 9, we have modified for Fig. 5.

      The model shows "MTOC capacity/conversion" during S phase, but this occurs during early G1

      Thanks a lot, as explained earlier, we used the term MTOC conversion occurring in G1 for what is normally called centrosome maturation occurring in G2/M, as explained earlier. We do not speak anymore of MTOC conversion since we have not tackled this question (explained above). We have therefore removed MTOC conversion in the texts and the schemes and replaced it by “centrosome maturation” for the duplication cycle, and by “enhanced MT nucleation capacity” for the MCC cycle. To be clearer and schematize that procentrioles are competent for MT nucleation before G2/M or A/G transitions, we have added some MT nucleated from G1 procentrioles during the canonical cycle, and from late A-stage procentrioles during the MCC cycle.

      The model shows disengagement only in the second M phase, but this occurs already at the first M phase, directly following centriole biogenesis, right before centosome conversion.

      This is a big edition error in both Fig. 5 and 9. Of course the daughter centriole disengage during the first M-phase. This has been changed. Thanks a lot for spotting it. This, however does not contradict the hypothesis of superposition.

      We also added the acquisition of distal appendage which was written in Fig. 5 but not in Fig.

      9 for duplication during the second M-phase.

      When the correct timings are incorporated in the figure, the proposed superposition of two cycles is not an accurate description of the events. Instead, your data seem consistent with a model where MCC incorporates all steps in one cell cycle variant that lacks mitoses, so that disengagement and MTOC conversion occur together with centriole elongation, followed immediately by acquisition of DAs and SDAs.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we tried to propose. When putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a super opposition of events that; although driven by the same molecular machinery, are normally occurring in two consecutive cell cycle. This is notably consistent with the findings of Kong et al., 2014 cited previously.

      (5) While all reviewers felt that there was no need to introduce the new term "nest", they leave it to the authors to keep it. However, the authors may want to consider that the term is still not introduced and explained properly, which may confuse readers. For example, while this section reads like an introduction to the term: "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC", the term is already used two times before without explanation. The first mentioning is at the beginning of the results section and is followed by citations, which gives the impression that these studies describe the nest, which is not the case.

      The first mention of “nest” is in the end of introduction resuming the findings of the paper where the term is in the following context: “we found that centriole amplification emerges in a pericentrosomal “nest” concentrating core centriole/deuterosome elements”. We looked at nest definition in the Collins Dictionnary : “a structure or other place where creatures, esp. birds, give birth or leave their eggs to develop”, we felt this was clear. We added quotation marks around the term nest.

      Then, the result section opens with this sentence: “The origin of amplified centrioles in MCC remains controversial. Some live imaging experiments and electron microscopy suggest that the centrosome could constitute a nest for centriole and deuterosome biogenesis (Al Jord et al., 2014; Kalnins et al., 1972; Mori et al., 2017), but others have proposed that procentriole-loaded deuterosomes emerge independently from the centrosome location, all over the cytoplasm (Nanjundappa et al., 2019; Sorokin, 1968; Zhao et al., 2013, 2019).”. Here, the term nest is again used as a place of birth for centrioles and deuterosomes which is what is actually proposed in these papers. First, Kalnins el al., in 1969 (we made an error on the reference date, this has been changed), resume in their abstract “This observation suggests that all of the clusters may form initially in close association with the diplosomal centrioles”. Then, not to mention Al Jord 2014 which comes from our lab, the title of Mori et al. is “Cytoplasmic E2f4 forms organizing centres for initiation of centriole amplification during multiciliogenesis”, and in the paper, they show that E2F4 accumulates at the centrosome. This is now also proposed by collaborators for MCIDAS (Lu et al., 2025). We feel that these references, which are often omitted, are appropriated at this location.

      Then we continue with: “To test whether microtubules drive the organization of a centrosomal nest from which procentrioles emerge”, which keeps the notion of the place of birth.

      Then the title "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC" arrives. In this section we first speak about a pericentriosomal cloud on which we zoom in using CLEM, to then conclude at the end of the section “Altogether live imaging mRuby-DEUP1/CEN2-GFP during early A-stage suggests that core deuterosome and centriole components are concentrated in a primordial cloud around the centrosome, which constitutes a nest where centrioles and deuterosomes concomitantly form before they move away from the centrosomal region (Fig. 2F)”.

      Finally, we begin the discussion section regarding the nest by: “We named this transitory compartment a “nest” since deuterosomes and procentrioles emerge specifically in this region and grow while moving away from it.”

      During the first revision, we tried to make it clearer. If this is still not the case after and the reviewer has another proposition of definitions/phrasing, we will be glad to consider it.

      As replied to the other reviewer, the term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the transitory region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      The following comments from Reviewer #3 may also provide further context regarding the editors' remaining concerns:

      The authors have done an excellent job addressing the points I raised overall, and the revision is substantially improved in focus and clarity. That said, some concerns raised by other reviewers, particularly regarding terminology and statistical analysis, could have been addressed more fully. One issue remains insufficiently resolved. Several quantitative analyses (for example Fig. 5C and 5E) still appear to rely on pooled single-event measurements collected across three independent experiments. This approach can overstate statistical significance. The authors indicate in their rebuttal that they use chi-square tests to compare proportions and to justify pooling across replicates. However, I am not convinced this addresses the issue for the intensity-based and single event distributions shown in the panels specified above. I recommend that these key analyses be represented with biological replicates shown explicitly (superplot-style, with replicates distinguished).

      Our reply was for the comparison of proportions and not the intensity-based and single event distributions shown in the panels Fig. 5C and Fig. 5E. We have now changed our plots to represent biological replicates explicitly (superplot-style, with replicates distinguished). As for the statistical analysis: we evaluated differences in marker intensity between A-stage and G-stage samples using a linear regression model, with stages as the main effect and replicate as a fixed covariate, to account for batch variation. Statistical significance was assessed using Type II ANOVA.

      Separately, I continue to feel that some newly introduced terminology (for example, the "nest") may not be necessary at this stage. It may be sufficient to describe these structures and focus on their spatiotemporal behavior, composition, and measurable features, rather than assigning new names. Having read the authors' response, I understand that they would like to retain this terminology, which is acceptable; however, it may not be readily adopted by the field.

      The term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      Minor correction (remove "in MCCs" part from the following sentence):

      In MCC, PCM1 depletion alters deuterosome formation and centriole production in brain and airway MCC (Hall et al., 2023; Zhao et al., 2021).

      Done

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1

      1. The inverse relationship between PGCLC and DE efficiency is intriguing but under-explored. The observation that lines efficient for PGCLCs (Podx1, Kolf2) are poor at DE differentiation, and vice versa, is one of the key findings. Yet this is presented almost in passing. It would strengthen the paper considerably if the authors discussed whether their Polycomb-regulated gene set predicts DE efficiency with an inverse sign, and whether the logistic regression model can be tested on the DE data directly.

      As suggested by the reviewer, we have enhanced our analysis of definitive endoderm (DE) differentiation efficiency and discussed it more prominently in the manuscript in the section “A subset of Polycomb targets is predictive of differentiation properties”. In particular, we now examine the correlation between gene expression from RNAseq and DE differentiation. For this purpose, we took the genes used to predict PGCLC efficiency (all of which are regulated by H3K27me3), and examined the correlation between their expression and the efficiencies of PGCLC and DE differentiation. We found that that differentiation is not a binary outcome for DEs, with many intermediate cases observed. Thus, instead of using a logistic regression, we employed a sigmoidal regression scheme for DEs. For this analysis, we used the absolute difference between observed and predicted efficiency, resulting in a mean absolute error of 12% [95% CI: 8-22%].

      We have now included an extra panel (Figure 7C) showing the correlations between expression of the H3K27me3 genes and the differentiation efficiencies in both PGCLC and DE fates. As anticipated by the reviewer, this plot reveals an inverse correlation, which we highlight in the main manuscript. Further, we now mention that these genes can also be used to predict DE differentiation efficiency, with satisfactory accuracy (although the confidence interval is wide due to the small sample size, as in the case of PGCLCs).

      The mathematical model is elegant but the choice to vary parameter E across cell lines needs stronger justification. The model assumes that inter-line differences are driven by variation in the overall rate of H3K27 methylation (parameter E). This is a reasonable starting assumption, but the authors should discuss alternative scenarios more explicitly. Could variation in demethylase activity, PRC2 recruitment strength, or replication timing equally well explain the data? The fact that EPOP is differentially expressed is mentioned as a potential mechanistic candidate for modulating E, which is compelling, but the link remains correlative. The authors should be more cautious in their language here, stating that EPOP "may be sufficient to completely switch the transcriptional regulation" goes beyond what the data show.

      We thank the Reviewer for these suggestions and have tightened our discussion of these points. It is correct that variation in demethylase activity can also explain our data. We now explicitly point this out in the section “The behaviour of H3K27me3 can be explained using a simple mathematical model”. However, this possibility does not fit as well with the RNA expression data. While we found differential expression of the PcG gene EPOP, we did not detect any differential expression of the KDM6 histone demethylases (KDM6A-KDM6C). Therefore, we still favour our original suggestion of variation in the methylation rates over this possibility. In addition, we have moderated our wording on EPOP, stating in the Discussion that changes in EPOP expression “may be sufficient to alter the transcriptional regulation of specific target genes.”

      The predictive model for differentiation efficiency is promising but the PGCLC training set is too small for confident generalisation claims.The authors acknowledge this (109 features, ~21 data points), and the L2 regularisation is appropriate. However, the claim of 91% accuracy with a 95% CI of 78-100% on the PGCLC data should be presented more cautiously. With such a small dataset, the confidence interval is very wide. The more convincing validation comes from the DN data (143 lines from Jerber et al.), where the model trained on PGCLC data performs comparably to the full-transcriptome model. This cross-fate generalisation is quite strong and should be emphasised more prominently as the primary evidence for the validity of the model.

      We have followed the reviewer’s guidance and revised our language, when discussing the PGCLC case in the section “A subset of Polycomb targets is predictive of differentiation properties”. We have also emphasised more clearly the successful validation of the DN data.

      1. The claim of "epigenetic memory" during differentiation (iPSC to pre-ME) is suggestive but would benefit from additional analysis. The authors show that 60% of pre-ME DEGs overlap with iPSC DEGs, and that H3K27me3-cluster genes maintain their expression patterns. However, 60% overlap could partly reflect genes that are simply not regulated during the short 12-hour pre-ME induction. To strengthen this claim, the authors should compare the overlap rate for H3K27me3-cluster genes specifically versus other clusters. If Polycomb targets show significantly higher overlap than, for example, K4&ATAC genes, this would more convincingly support a Polycomb-specific memory mechanism.

      We have now performed this analysis, examining the persistence of DEGs into the pre-ME state (i.e., whether a gene that was differentially expressed in hiPSCs remains differentially expressed in pre-ME). Excluding the H3K9me3 cluster (the smallest cluster containing fewer than 25 genes), the K27 cluster is the most persistent cluster in terms of fraction of genes per cluster. When the clusters were pooled into the different variables involved (ignoring H3K9me3), H3K27me3 again emerged as the most persistent chromatin feature. Unfortunately, however, these results were not statistically significant, so we are unable to include them in the manuscript.

      Lack of genetic background analysis. Ten lines from nine donors will harbour substantial genetic variation. The authors note that genetic variation has been linked to iPSC heterogeneity but do not analyse whether the three "outlier" lines (Kucg2, Sojd3, Yoch6) share genetic features. For instance, common variants at PRC2 component loci, EPOP regulatory variants, or structural variants that might alter H3K27me3 domain boundaries. The HipSci consortium provides genotyping data for these lines. A targeted analysis of variants at Polycomb-related loci would be feasible and could either strengthen the epigenetic interpretation or reveal a genetic confounder.

      We thank the Reviewer for raising this important point. To investigate potential confounding effects due to genetic variation between the hiPSC lines in our panel, we performed a targeted analysis of genetic variation across Polycomb-related loci (H3K27me3 occupied loci and Polycomb group genes) in all ten cell lines (using whole genome sequencing data from the HipSci consortium). This analysis specifically tested whether the three “compromised” lines (Yoch6, Sojd3 and Kucg2) share consistent genetic variants relative to the seven “normal” lines. We identified 15 indels (out of 4115) that satisfied this criterium. However, all are located in non-coding regions and none overlap with ATAC-seq peaks. Hence, they are unlikely to function as gene regulatory elements (e.g., enhancers), but we cannot exclude the possibility that they affect gene expression in other ways. We have added a new Results section “Genetic variants shared between differentiation-compromised hiPSC lines” to discuss these points, as well as adding new text to the Discussion and Methods.

      Minor Comments

      The promoter definition ({plus minus}1 kb from gene start) is non-standard; most studies use a window upstream of the TSS rather than gene start. The authors mention they confirmed robustness to an alternative definition (-1 kb to gene start) but do not show this data. It should be included in the supplement.

      We now show the data for the alternative promoter definition in Supplementary Fig. 5B and Supplementary Fig. 7C. These results demonstrate that our conclusions are robust to different promoter definitions.

      For CUT&Tag, no spike-in normalisation is mentioned. Given that the key conclusions is based on quantitative comparisons of H3K27me3 levels across cell lines, the absence of spike-in controls is a potential concern. The authors should discuss whether technical variation between CUT&Tag libraries could contribute to the observed bimodality. At minimum, the correlation between replicates for H3K27me3 should be shown (presumably it is high, but this should be documented).

      We thank the Reviewer for this suggestion. As now shown in Supplementary Fig. 4C, the correlation between our H3K27me3 replicates is indeed high (R between 0.93 and 0.96). Hence, technical variation between CUT&Tag libraries is unlikely to contribute to the observed bimodality.

      The statistical test for the PGCLC/H3K27me3 overlap (p We thank the reviewer for noticing this. Indeed, this is the case. The test assumes independence of lines, which is in general a reasonable assumption, but may not always hold. Specifically, the Kolf2 and Kolf3 lines are derived from the same donor, which implies they are not completely independent. However, for all other lines, we still think independence is a reasonable assumption and, thus, the overall result of the test should be a good approximation. We have added this caveat to the manuscript.

      Figure 6A: the heatmaps for H3K4, ATAC and H3K27 are shown side by side but at apparently different scales; this should be clarified or made consistent.

      Indeed, the scales in all heatmaps are the same. We have clarified this in the captions of the figures.

      Reviewer #2

      1.) Figure 2B. Are all GO terms shown in the figure or are these just the top terms? If this is a suset then all terms should be provided as a supplemental table. If this is all significant terms, this is relitavely modest considering the number of DEGs (712) and is probably due to the fact that DEGs are derived from all comparisons and so could be diluted by the presence of multiple opposing effects. If this is the case, you could identify DEGs that define the PCA groupings and then re-run the GO analysis to potentially provide a better definition of the functional differences between groups of cell lines.

      The GO terms previously displayed were the top hits. We have now included all the significant terms in Supplementary Files 4 and 5 (for the Molecular Function and the Biological Process ontologies, respectively).

      Chromatin accessibility at gene promoters is a poor predictor of transcription, but it is likely that accessibility at distal regions (e.g putative enhancers) might be a better predictor. Did the authors look at this? This possibility should at least be mentioned when discussing the ATC-seq data and the lack of correlation with transcription.

      • *

      We thank the reviewer for this suggestion. To locate additional regulatory regions, we downloaded tracks for the enhancer-associated marks H3K4me1 and H3K27ac for the ten cell lines from Todd and colleagues (Todd et al., Genome Biology, 2025; https://genomebiology.biomedcentral.com/articles/10.1186/s13059-025-03658-8). We then intersected the ATAC-seq peaks with the H3K4me1 peaks in each cell line to identify putative enhancers. For each protein coding gene, we then identified the closest ATAC and H3K4me1 positive peak (among all cell lines), which we assumed was the most likely enhancer for that gene. We then evaluated the ATAC, H3K27ac and H3K27me3 signal within these enhancers for each cell line. With this information, we tried using a version of our SVM-based pipeline to improve our understanding of transcriptional regulation in genes within the ‘origin’ cluster (for which we failed to get significant insights from our standard SVM approach). Thus, we used seven variables as an input for the SVM: The four of the standard approach and three additional variables from the ATAC/H3K27ac/H3K27me3 signal at the nearest enhancer. However, for genes with an enhancer closer than 100kb, the performance of the SVM with enhancer variables was similar to the standard SVM (or slightly worse). If we focused on genes with enhancers 10kb or closer to the TSS (75 genes), then the SVM with the enhancer signal did modestly improve the prediction. However, when analysing the results more closely, it was only for a handful of genes (around 10) where the usage of the enhancer data was beneficial, and, even then, it was mostly down to the H3K27me3 signal rather than the more standard enhancer marks, such as H3K27ac or chromatin accessibility. This lack of improvement in the accuracy is probably due to our inability to identify the correct enhancers, as distance on the linear genome scale is often a poor predictor of enhancer-promoter interactions.

      Ultimately, because the improvement is for such a small number of genes, we have not included this analysis in the manuscript. However, we do now mention in the manuscript in section “Chromatin accessibility does not always correlate with transcription” that we tried to include distal enhancers but that this approach was not successful.

      2.) Fig 1C. Statistic overview at end of legend should be moved under section describing panel C in the legend.

      We have now made this change.

      3.) 'Furthermore, the transition value of 30% enables repression to be stably maintained even after DNA replication, when, on average, histone modification levels will be transiently halved'. Whilst this is potentially true and a plausible interpretation, you cannot exclude that the signal is not derived from different cell populations in the culture due to cellular heterogeneity such as cell cycle or spontaneous differentiation. This possibility should be noted in the text.

      We thank the Reviewer for this suggestion. Due to the possible alternative explanations pointed out by the reviewer, and to minimise any possible misunderstandings, we decided to drop this sentence from the manuscript, which is not required for any of our main conclusions.

      4.) 'Higher values indicate stronger correlation or anticorrelation and, thus, stronger differences between cell lines.' I don't believe this makes sense as written. Do the authors mean stronger partitioning of different iPSC lines into clusters?

      Indeed, this sentence wasn’t very clear -- we have now rewritten it to improve clarity: “Because absolute correlation values were used, high values indicate that expression profiles between two cell lines are either highly correlated or highly anticorrelated. Across all pairwise comparisons, high values suggest strong partitioning of cell lines with highly similar or markedly different transcriptional profiles.”

      5.) 'We found that 60% of the DEGs in pre-ME were also DEGs in hiPSCs'. This needs to be made clearer. Do the authors mean DEGs between iPSCs following differentiation or DEGs between undifferentiated iPSCs and their differentiated derivatives? The former suggests that the iPSCs are already partially differentiated and that differentiation in promoted or constrained by this starting state whilst the latter would suggest that some lines are skewed towards the mesendoderm.

      We mean that of the genes that are differentially expressed between the 10 lines in pre-ME, 60% were also differentially expressed between the 10 lines in iPSCs (prior to differentiation). We have reworded this sentence to make it clearer.

      6.) 'Finally, histone marks in the iPSC state were also predictive of expression in the pre-ME state, albeit with slightly lower accuracy than for the iPSC state (Supplementary Fig. 8C, D), which may indicate the existence of an epigenetic memory system that is maintained during differentiation.' Or the retention of an epigenetic signature that failed to be erased during the initial generation of the iPSCs.

      We agree with the reviewer that this is entirely possible: our point is that memory states may persist from iPSCs to pre-ME. The memory state may of course predate the initial generation of the iPSCs. We have amended the section “Pre-ME transcriptomes suggest inheritance along the developmental trajectory” to include this possibility.

      7.) 'To minimise the risk of overfitting, only reliable targets were retained'. Whilst this is outlined in the methods as stated, a summary of what this means should be included in the body text.

      We thank the Reviewer for this suggestion. We have included the required extra text in the section “A subset of Polycomb targets is predictive of differentiation properties”. We have also revised the performance metrics so that they are strictly comparable with the results of Jerber and colleagues (which implies, in some cases, removing error bars, as in the results of Jerber et al., 2021). The reviewer may notice differences in the values reported but all our claims remain valid.

      Reviewer #3

      The major claim that among histone modifications that have been profiled in this manuscript, H3K27me3 is the most predictive for expression is supported by the analysis. However the analysis may be skewed because the RNAseq and the H3K27me3 difference are driven by the extreme skewing of the 3 cell lines Yoch6, Sojd3 and Kucg (Fig 2A, 2C and 6A). Two of these lines cannot form EBs at all, a major failure in their pluripotent characteristics.

      We thank the reviewer for raising this fundamental point. Our aim for this study was to use iPSC lines that have passed existing standards and could easily be chosen from a panel of lines by an unsuspecting user. Indeed, the differentiation-compromised lines in our study are indistinguishable from other PSCs from a validated source that extensively characterises the distributed material (HipSci resource, https://www.hipsci.org). This source categorises these cell lines as correctly reprogrammed and fully pluripotent. In addition, we now present PluriTest data (doi: 10.1038/nmeth.1580) from all normal lines available from the HipSci resource (835 lines) and highlight the ten cell lines used in this study (see Supplementary Fig. 1A). All cell lines in our panel have pluripotency scores over 20, and all but one (Bima1 – which notably differentiates efficiently into PGCLCs and DNs) have novelty scores below 1.67; these values have been empirically determined as pluripotency signature thresholds (Müller et al., 2011). This analysis clearly demonstrates that the cell lines in our study are not outliers, an important fact which we have now added to section “Marked differences in the developmental efficiency of hiPSC lines”.

      Furthermore, one of the key advances of our study is that we identify a chromatin and transcription signature that will enable researchers in the stem cell community to identify iPSC lines with compromised differentiation potential early on. We also note that compromised differentiation potential is widespread among human PSCs. For example, Jerber et al. report that 48 out of 183 hiPSC lines could not be differentiated successfully into dopaminergic neurons (doi:10.1038/s41588-021-00801-6). Thus, our study addresses an important and widespread issue in the stem cell field, a point we now emphasise in the introduction of the manuscript.

      Further, one of the lines that can form EBs, fails to make PGCLCs but can differentiate into DE, Letw5 has neither the RNA profile nor the H3K27me3 profile of the skewed iPSC lines. Therefore, whether H3K27me3 truly influences phenotype at least in terms of PGCLC and DE differentiation of iPSCs is not supported by the analysis in the manuscript.

      We agree that the behaviour of Letw5 is interesting, and we discuss its properties extensively in section “Marked differences in the developmental efficiency of hiPSC lines” and Fig. 1E. As we state, comparing Letw5 with Kucg2, “These findings suggest that Kucg2 hiPSCs have limited developmental competence to generate PGCLCs, while Letw5 hiPSCs are capable of PGCLC specification but fail to sustain the germ cell fate, pointing to a defect in fate maintenance rather than in initial developmental capacity.” Hence, the evidence points towards Letw5 having a separate defect which is unrelated to the impaired Polycomb regulation identified in the other three problematic lines. We also emphasise this point in section " A major role for H3K27me3 in hiPSC transcriptional heterogeneity", where we state that "[...] in this case [Letw5], a distinct mechanism, independent of H3K27me3 dysregulation, may result in impaired germ cell development."

      1. What are the predictions from applying SVM to data from only the 6 cell lines Podx, Kolf2, Kolf3, Bima 1, Qolg1, Wibj2. The DE differentiation potential will also have to be measured for each of these cell lines.

      Following the reviewer’s suggestion, we applied the SVM only to data from those six cell lines (which do not include any of the defective cell lines), see section “Linking variation in chromatin features with transcriptional output using SVMs”. Given that the SVM only takes as input data from differentially expressed genes, the set of genes used decreased markedly as there are fewer genes differentially expressed among these cell lines (125 DEGs). Nevertheless, for this subset of genes, the SVM still retains satisfactory accuracy (both AUROC and overall accuracy in the 70% to 75% range; now shown in Supplementary Fig. 6H). This result is particularly remarkable given that the SVM is operating with very little data (five datapoints for training and one for testing, per gene) and that the cell lines are very similar to each other. As the reviewer points out, we hope these results might encourage other researchers to pursue similar analysis approaches.

      For DE differentiation, we previously included data (Supplementary Fig. 3B, C) for the following lines: Podx1, Kolf2, Kucg2, Letw5, Sojd3, and Yoch6. Only Kolf3, Bima1, Qolg1 and Wibj2 were missing. We have also now measured DE differentiation in three remaining lines (Kolf3, Qolg1, and Wibj2).

      The above analysis may also shed light on howextreme the input parameters must be for SVM to be a good classifier? Such an analysis may also assist future users of the method to assess whether SVM would be useful for their datasets.

      Please see our previous answer. We argue that the results presented above for six similar cell lines imply that this type of computational approach can have general applicability and does not require extreme inputs. We have followed the Reviewer’s suggestion and now incorporate this finding in section “Linking variation in chromatin features with transcriptional output using SVMs”: “Furthermore, the SVM does not require extreme values or outliers, and hence the overall approach could be of rather general applicability. As a performance verification, we applied the SVM to a dataset containing only the cell lines that could generate PGCLCs with high or intermediate efficiency, and while the performance is slightly reduced, it remains satisfactory (accuracy 75%; Supplementary Fig. 6H).”

      If the SVM on the 6 lines does not predict a binary switch in H3K27me3 to be predictive could the authors incorporate DNA methylation and H3K4me1 from the same publication as the chromatin accessibility. Such an analysis may also assist future users of the SVM method to assess the number of parameters required to separate closely related phenotypes.

      See previous answer. We note that DNA methylation data for our hiPSC panel is not available; it is not part of the study that the reviewer mentions (https://link.springer.com/article/10.1186/s13059-025-03658-8). Although H3K4me1 data is available in Todd et al., we did not find that this data improved the ability of our model to make successful predictions (see reply to Reviewer #2, point 1).

      Most gene regulation occurs at the level of the enhancer, restricting analysis to promoter associated histone modifications is limiting.

      We thank the Reviewer for raising this very valid point. Please see response to Reviewer #2, point 1.

      One puzzling piece of data is the very high 60% of PGCLCs on day 1 of differentiation (Fig 1E) in the competent cell lines. BLIMP1 is expressed in hiPSCs, calling into question whether the initial differentiation into pre-ME was successful.

      We think there is a misunderstanding regarding the experimental timeline. Day 1 of differentiation in Fig. 1E refers to one day after PGCLC induction from the pre-ME stage following the addition of BMP4, SCF, LIF, and EGF (see schematic in Fig. 1A). We have revised the text to make this clearer. Furthermore, BLIMP1 (PRDM1) is not expressed in hiPSCs. To demonstrate this, we now show the expression levels of BLIMP1 (PRDM1), B2M (low to mid-level expression in most human cell types), SOX2 (highly expressed pluripotency marker), and HOXC10 (differentiation marker that is not expressed in PSCs) across our cell line panel. At this scale, BLIMP1/PRDM1 expression is not detectable. When SOX2 is omitted from this bar plot, the very low expression levels of BLIMP1/PRDM1 become apparent, as it is close to the levels for the differentiation marker HOXC10. We conclude that BLIMP1/PRDM1 is expressed at extremely low levels across our ten hiPSC lines.

      The H3K27me3 and H3K9me3 signals are integrated over the entire gene as inputs into the SVM, however PCA analysis to separate the cell lines is only shown for the promoter

      This is not quite correct. For the PCA analysis for the histone marks and ATAC-seq, we used both the promoter region (Fig. 2C, Supplementary Fig. 5B) and the gene body (Supplementary Fig. 5A), with similar results. For the SVM, for H3K27me3 and H3K9me3, we primarily used the entire gene region, but we also tested other regions (Supplementary Fig. 6A), with similar or slightly inferior results.

      SVMs have been used to predict enhancers from epigenomic data PMID: 22328731 and to classify cancers PMID: 11120680. Applying SVM as classifier for gene expression prediction is not very novel.

      We thank the Reviewer for raising this point. We did not claim that the use of SVMs was itself novel. It has certainly been used in other contexts, as the reviewer points out, to predict enhancers, for cancer classification, and to predict expression patterns. In fact, SVMs had already been used to predict gene expression from chromatin features (Cheng et al, 2011; already cited in our manuscript). What is novel in our work is the reverse-engineering of the method to extract mechanistic information about each gene (i.e., assign a chromatin feature set relevant to the changes in expression). This computational methodology, in conjunction with the rich experimental dataset produced, allows us to classify differentially expressed genes in terms of the chromatin features that enable prediction of transcription. This highlights the differences between cell lines and enables further downstream analysis such as, mechanistic models of histone modification dynamics and the prediction of iPSC differentiation efficiency. We have rewritten the Introduction to the manuscript to better emphasise these points.

      The biological insights are limited. For example, the observation that " a variety of forms of transcriptional regulation" Fig 4B. It is well known that H3K27me3 decorates lineage specifying genes and is part of the bivalent domain with H3K4me3. The anti-ATAC category could represent locations where a repressor is bound DNA which would also result in increased accessibility and is not a surprising result.

      We believe our work does offer significant biological insights. While we agree that it is well known that H3K27me3 decorates lineage specifying genes, it was not previously known that digital Polycomb dysregulation at specific loci was a key feature controlling the ability of pluripotent cell lines to differentiate properly. In addition, we have been able to identify a core set of genes whose H3K27me3 profiles are highly informative for differentiation efficiency. Moreover, we are able to explain the variation in H3K27me3 levels by simple, quantitative, mathematical model.

      Finally, the anti-ATAC category is a minor finding and not one of the central conclusions of this paper. Nevertheless, we appreciate the Reviewer’s suggestion and have incorporated this possible interpretation into section “Chromatin accessibility does not always correlate with transcription”.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this valuable study, the authors developed long-term imaging tools to simultaneously monitor the temporal and spatial dynamics of excitatory and inhibitory synapses and reported that excitatory and inhibitory synapses need to develop synergistically during synaptogenesis to maintain balance. While the analysis and quantification of the imaging data are incomplete, there is convincing evidence that the developed tools are feasible. If these tools can function stably in vivo, their applications will be much broader.

      We have completely overhauled our analysis and quantification methods and generated custom-made drift correction and tracking pipelines. Also, we have tested these tools ex vivo.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      We thank the Reviewer for their positive feedback and careful review of our study.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      We also thank the Reviewer for their insights and suggestions. We have added discussion of this important point to the Discussion section. Furthermore, we have tested the applicability of our tools ex vivo (new Figures 1, 4, and 6). While using these tools in vivo for live imaging is the eventual goal, we started in a reduced culture system given the relative simplicity. Our current study now provides a framework for future experiments applying these approaches in more complex in vivo systems.

      Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomatoGephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HAHomer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      We thank the Reviewer for their positive assessment of our study.

      Weaknesses:

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Minor weaknesses of the manuscript include:

      (1) The lack of a characterization of endogenous Homer1-positive excitatory synapses using TKIT.

      We attempted to perform live imaging of endogenous Homer1-positive synapses using the TKIT approach by tagging endogenous Homer1 with mClover3 but encountered low signal/noise while live imaging. This prompted us to focus our current study on live imaging endogenous Gephyrin. Future studies using more robust tags (e.g. StayGold, HaloTag) for TKIT tagging of endogenous Homer1 will likely help circumvent this issue.

      (2) Discussion about other approaches to study excitatory and inhibitory synapses using endogenous proteins (e.g., intrabodies - FingR or nanobodies) should be included.

      This important point was also raised by other Reviewers. We have now significantly expanded the Discussion section, including discussion of this point.

      (3) The activity state of a neuron and/or a synapse might alter the dynamic properties (formation, maintenance, and/or elimination). A discussion on whether the overexpression of Homer1 and/or gephyrin might alter synapse/neuron activity would provide greater interpretability of the results. A discussion of the potential limitations and benefits of the reporter and TKIT approaches would be beneficial.

      We agree and have added discussion of these points to the Discussion section.

      (4) A description and interpretation of the computational approach to calculate particle tracking would be helpful. I found that particle tracking figures, while elegant, are difficult to interpret.

      As discussed in more detail below, we have generated drift correction and particle tracking approaches for the revised manuscript. We now elaborate on these new approaches in the paper.

      We thank the Reviewer again for their very helpful input and suggestions.

      Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      We thank the Reviewer for their positive assessment of our study.

      However, a number of aspects remain to be addressed in order for the study to support the claims made by the authors. First, the novelty aspect of the development of the fluorescently tagged synaptic proteins is unclear, since reporters of this nature are in routine use in many labs. Second, the analysis of the acquired images often seems incomplete, with only example images but no quantification shown, or the distinction between spatial and temporal dynamics appearing unclear. Third, given this incomplete analysis, the interpretations of the authors are not always convincingly supported by the data presented. In conclusion, substantial improvements are required to render the main messages of the study clear and compelling.

      We agree and have incorporated all of the Reviewer’s suggestions in the revised manuscript (please see below).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This is an interesting study. This reviewer has the following questions/comments for the authors:

      (1) Please provide evidence that the gRNAs targeting each gene of synaptic protein have no offtarget effects.

      We now include analysis of off-target effects for the TKIT tools (new Figure S6).

      (2) While structural E/I balance is shown, functional electrophysiological validation (e.g., mEPSC/mIPSC ratios) is absent. It is interesting to know whether the balanced functional structural changes translate to functional?

      We thank the Reviewer for this insightful suggestion and now include these recordings in the revised paper (new Figure 8).

      (3) In lines 217-218, please define thresholds for "stable" vs. "dynamic" puncta (e.g., temporal and spatial criteria).

      We more clearly define our categorization parameters (e.g. new Figure 2).

      (4) In Figure 5B: The low co-localization between endogenously tagged Bassoon and antibodystained Bassoon is likely due to the low TKIT efficiency. Quite a few HA-tagged Basson signals are insensitive to Basson-antibody. The authors are suggested to explain those.

      We thank the Reviewer for identifying this and add discussion to the Results section.

      (5) For the data analysis. If each n represents an independent neuronal culture, should the authors are suggested to provide the number of neurons/dendrites analyzed for each independent culture?

      We have added these important details to the manuscript.

      (6) Regarding the title, the author used the term "coordinated dynamics". This reviewer finds it is a bit over-claim because the stable ratios of the number of excitatory synapses and inhibitory synapses are likely an association, not actively "coordinated". I suggest that the authors rephrase this.

      We agree that we cannot argue that excitatory and inhibitory synapses are causally coordinated in our current study. Their levels are likely associated by either association or direct coupling, which we now discuss further in the first paragraph of the Discussion. We have rephrased the title accordingly.

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions that I think will improve the manuscript:

      (1) Please define Syn1/2 on line 129.

      We have defined this in the revised paper.

      (2) For Figures 2B, C, and 4B, C: are the puncta in panel C from the dendrites in panels B? If so, it would be helpful to identify the ROIs selected in panels C.

      We now include this in new Figure 2.

      (3) For the particle tracking figures, while the ability to track all synaptic puncta is very impressive, it is sometimes difficult to clearly track the lifespan of a synaptic puncta from the current figures. I believe that it would be helpful if the authors selected specific examples of synapses formed, maintained, and eliminated.

      We agree and now include more examples.

      (4) I believe that more detail about the computational approach and analysis for the particle tracking (Figs 2E and 4E) would help the interpretability of the figure.

      This important point was also raised by the other Reviewers. We generated custom tools during the revision that significantly expand the capabilities of our tracking approaches and more clearly describe them in the revised manuscript.

      (5) Similar to the rigorous gephyrin TKIT analysis (Fig. 6), did the authors perform a similar analysis for Homer1c TKIT? This might be valuable to confirm that overexpression of the Homer1 reporter does not indirectly alter synapse dynamics.

      We attempted to perform live imaging of mClover3 TKIT-tagged endogenous Homer1 but encountered low signal/noise with live imaging. We now add discussion that optimization of more robust tags (e.g. StayGold, HaloTag) will likely be necessary for live imaging of different target proteins.

      (6) The tools developed by Garbett et al. have the potential to be broadly utilized in the field to provide new insight into the coordination of excitatory and inhibitory synapses. It would thus be helpful for the authors to include a discussion about the strengths and limitations of the reporter and TKIT methods relative to other approaches used to live image synapses (e.g., intrabodies (FingR and nanobodies)).

      We have now significantly expanded the Discussion to include these important points.

      (7) In the discussion, can the authors elaborate on whether it is experimentally feasible to apply their TKIT labeling of gephyrin and Homer1c in the same neuron to assess the endogenous excitatory and inhibitory synapse dynamics from the same neuron?

      We have added discussion of this point and also proof-of-concept data supporting tagging of two postsynaptic targets within the same neuron (new Figure S5D).

      Reviewer #3 (Recommendations for the authors):

      (1) While the new tools described in the current manuscript can undoubtedly be used for the described purposes, the novelty of these tools is unclear to me. Viral vectors expressing fluorescently tagged versions of Homer1, synaptobrevin, and gephyrin are commercially available, e.g., via Addgene, and they are in routine use in many labs. CRISPR-mediated strategies for this purpose have also been previously reported (e.g., Willems et al. 2020, PLOS Biology; Fang et al. 2021, eLife). It is not clear to me how the tools reported here present a significant improvement over existing resources, other than that they use different fluorescent tags. If this aspect is a central part of the current manuscript, it should be expanded on in the discussion, including a direct comparison with available tools to highlight the novel aspects.

      We agree and have significantly expanded the Discussion to include these important points. Also, rather than argue that our tools are superior to pre-existing approaches, we adjust the text to argue that our tools and analytical approaches have been designed and optimized for the purposes we apply them to.

      (2) In addition to generating new tagged constructs, the authors also state that they have developed new imaging and analysis strategies to facilitate long-term assessment of synaptic dynamics. However, in many figures, they present only sample images, with little quantification to allow assessment of the wider relevance of the imaged synapses. For example, in Figures 2C and 4C, they present one example each of, e.g., a stable, nascent, transient, or eliminated synapse. However, they do not provide any quantification on how frequently any of these events occur, or whether they can be reliably quantified at all. These quantifications (i.e., percentage of each event type across a large population of synapses) would be necessary and should be added to demonstrate that this tool can be used for more than single example images.

      We have generated custom-made drift correction and particle tracking approaches for the revised manuscript. Based on the reviewer’s suggestion, we have quantified the relative frequencies of stable, nascent, transient, and eliminated synapses (Fig 2B-G, Fig3A-F, Fig 5A-F, Fig 7B-C). These metrics greatly enhance the biological interpretation of our results. We have also added a supplemental movie with an example image with corresponding categorized tracks for each puncta type (Movie S3)

      (3) The authors do present an automated visual representation of spatial track length across the neuron, e.g., in Figure 2E and 4E, although this is also not quantified. Moreover, the track lengths appear surprisingly short, despite the authors' claims that their analyses 'highlight the dynamic nature of excitatory synapses over these timescales'. It is not clear to me whether these short tracks are more than just jitter, either in the synapses themselves or in the images due to technical limitations. E.g., in panel 2E, I see very few examples in which the track is not simply centered around one point, but actually expands over a distance. Quantification of the distance between start and end points of the tracks would be important to support the claim that these synapses are dynamic in terms of spatial translocation (if that is what the authors meant). Or if the 'dynamic nature' of the synapses referred to temporal dynamics, it is unclear to me how this information can be gained from the represented tracks.

      We thank the reviewer for these excellent points. To accurately access spatial motion, we drift-corrected our images with a custom correction algorithm to eliminate stage or microscope drift as a source of contaminating motion (See Methods, Movie S2), in addition to collecting time-lapse imaging with Nikon perfect focus. We noticed heterogeneity in our cultures such that some areas contained very mobile neurites, while other remained stationary (Fig. S1). We binned movies into either moving or still neurites and assessed spatial metrics as suggested (Fig. S1A). Consistent with our binning, puncta on moving neurites showed larger net displacement (distance between start and end points), but puncta on still neurites also showed ~1 µm net displacement (Fig. S1D). We also quantified puncta speed and found that puncta on moving neurites generally moved faster (Fig. S1C). We appreciate the reviewer’s insight that track length were surprisingly short, and after employing our drift correction and revised tracking methods, we now see substantially longer track lengths (Fig 2E, Fig 3C & F, Fig S2B & C). We additionally see a large fraction of tracks that persist throughout the imaging session (Fig 2E, Fig S2B & C).

      (4) In Figure 3, the authors now quantify track length, but in this case in the unit 'minutes', from which I would interpret that this is now meant to assess the temporal dynamics rather than the spatial dynamics. The lack of a clear distinction between spatial dynamics and temporal dynamics is very confusing to me, since these are entirely independent measures. 'Track length' to me indicates spatial dynamics, and I would expect the units to be a measure of distance. 'Track duration', which the authors also use in some places, but inconsistently as far as I can tell, makes sense to me for the assessment of temporal dynamics, with the units being a measure of time. I would strongly recommend being very clear about this distinction, since the current representation of the data is very difficult to follow and interpret.

      In addition to new spatial metrics, we have clarified in the text when we are referring to spatial dynamics (distance) versus temporal dynamics (time). As suggested, we use duration when referring to time, and speed or distance when referring to spatial metrics.

      (5) The images from the newly generated CRISPR-based tags in Figures 5-7 are striking and very compelling - these will be very useful tools. However, here too, it seems that the interpretation of the data does not really match the results. All quantification indicates that there is very little change in synapse density or other assessed parameters over the time course of the imaging, and yet the authors emphasize the dynamic nature of visualized synapses. More compelling quantification would be needed to support this claim.

      We have quantified spatial and temporal metrics for live neuron culture imaging for all tools developed including CRISPR-based tags (Figure 7).

      (6) The discussion is extremely short and provides almost no integration of the results of the study into the framework of existing knowledge. Instead, it focuses almost exclusively on unanswered questions and future perspectives, which are also important, but not helpful in interpreting the findings from the current study. The latter aspects should be added to provide essential context for the current findings.

      We agree and have added additional discussion of our current findings to help contextualize their significance.

      We thank the Reviewers again for their positive feedback and insightful input, which has undoubtedly strengthened our study.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance understanding of the neural bases of voice perception and the evolutionary roots of voice sensitivity in the human brain.

      Weaknesses:

      (1) Acoustic-phylogenetic confound: The design does not fully disentangle acoustic similarity from phylogenetic proximity, as species co-vary along both dimensions. A promising way to address this would be to include an additional model focusing on the acoustic features that specifically differentiate bonobo from chimpanzee calls, which share equal phylogenetic distance to humans.

      (2) Selectivity vs. sensitivity: Without non-vocal control sounds, the study cannot determine whether TVA responses reflect true selectivity for primate vocalizations or general auditory sensitivity.

      (3) Task demands: The use of an active categorization task may engage additional cognitive processes beyond auditory perception; a passive listening condition would help clarify the contribution of attention and task performance.

      (4) Figures and presentation: Some results are partially redundant; keeping only the most representative model figure in the main text and moving others to the Supplementary Material would improve clarity.

      We thank the reviewer for contributing to the improvement of the present study and for the extremely constructive criticism. Concerning the identified weaknesses of our work, we provide here some general answers while the detailed review (below) addresses point-by-point the reviews in high detail.

      (1) We totally agree that acoustics and phylogeny cannot be disentangled in our study, which is a limitation. We now provide the suggested analysis on the acoustic specificities of chimpanzee and bonobo calls.

      (2) This point on selectivity vs. specificity is indeed crucial, and we now provide a more careful viewpoint and phrasing on this aspect, since our study can only provide partial arguments for this important distinction.

      (3) Task demand following species categorization might rightfully yield to the engagement of distinct brain network compared to merely listening to the stimuli. We discuss this aspect and put forward the argument that, while we cannot control for this aspect, our attentional control study performed by an independent sample, N=28 provides clear evidence that no species triggered an attention bias. In other words, task demand might play a role, but at least in the study we know that attentional resources were not biased towards one species in particular since no effects were observed.

      (4) We agree that results were not articulated in a clear fashion and that figures were redundant. We addressed this aspect and regrouped the figures where appropriate while we include the rest in the supplementary material now.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding - that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations - suggests a potential neural sensitivity to calls from phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weaknesses:

      The interpretation of the findings in this paper regarding the evolutionary continuity of voice processing lacks sufficient evidence. A simple explanation is that the observed effects can be attributed to the similarity in low-level acoustic features, rather than effects specific to phylogenetically close species. The authors only tested vocalizations from three non-human primate species, other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      We want to thank the reviewer for the constructive criticism and for evaluating the manuscript.

      Concerning the principal weakness highlighted, we provide new analyses behavioral, acoustics, model-based fMRI that improve our understanding of the influence of both phylogeny and bioacoustics in our data. We argue that the explanation proposed by the reviewer cannot explain our results, as also observed in several other research from us and others. We discuss this aspect and emphasize that including stimuli from more species would greatly improve the understanding of phylogeny and bioacoustics in this context.

      Reviewer #3 (Public review):

      Summary:

      Ceravolo et al. employed functional magnetic resonance imaging (fMRI) to examine how the temporal voice areas (TVA) in the human brain respond to vocalizations from different nonhuman primate species. Their findings reveal that the human TVA is not only responsible for human vocalizations but also exhibits sensitivity to the vocalizations of other primates, particularly chimpanzee vocalizations sharing acoustic similarities with human voices, which offers compelling evidence for cross-species vocal processing in the human auditory system. Overall, the study presents intellectually stimulating hypotheses and demonstrates methodological originality. However, the current findings are not yet solid enough to fully support the proposed claims, and the presentation could be enhanced for clarity and impact.

      Strengths:

      The study presents intellectually stimulating hypotheses and demonstrates methodological originality.

      Weaknesses:

      (1) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy.

      We thank the reviewer for evaluating our manuscript and for the constructive criticism as well as the many suggestions. Concerning the weaknesses of the study, we provide here some quick answers while more detailed responses can be found below.

      (1) We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator.

      (2) We totally agree that figure redundancy was a problem and we now reduced confusion by combining congruent aspects while pushing other results to the supplementary material.

      Recommendations for the authors:

      Reviewing Editor Comments:

      With additional analyses and discussions, the work has the potential to offer important insight into the evolutionary continuity of voice processing.

      We thank the Reviewing Editor for this additional motivation and for offering us the possibility to revise our manuscript. We will now provide our point-by-point reviewing, referring to manuscript modifications by section and/or line number(s). All modifications are also highlighted in light grey in the text.

      Reviewer #1 (Recommendations for the authors):

      The manuscript is clearly written and addresses an important comparative question about the specificity of human TVA responses. The acoustic analyses are well designed, and the imaging work is careful and thorough. However, several conceptual and methodological issues need clarification or tempering of claims, particularly regarding (i) the distinction between sensitivity and selectivity, (ii) the confounding of acoustic and phylogenetic factors, and (iii) the interpretation of "chimpanzee-specific" TVA activity.

      (1) Introduction

      Line 48: cite more recent infant EEG evidence for early voice sensitivity (Calce, Curr Biol).

      The reference and explanation were added, lines 46-48.

      Line 53: mention recent data on voice processing in marmosets (Jafari, Cell Rep; Dureux, Curr Biol).

      We added the references and the mention of these interesting studies on common marmosets, lines 53-54.

      Line 59: Fecteau et al. (2004) already explored cross-species selectivity; please integrate and discuss.

      We now mention here the work from Fecteau and colleagues and its relevance, see lines 57-59.

      Line 70: clarify that in [27] (Bodin et al., 2021) human TVA responded similarly to human nonverbal vocalizations and macaque coos, likely due to acoustic similarity.

      We added this important aspect, thank you for this precision. See lines 71-72.

      Clarify why an active species-categorization task was chosen instead of passive listening, which is standard in TVA research. Were participants familiarized with stimuli beforehand?

      We added a sentence on this aspect, but basically to summarize it here: we wanted to be able to test human recognition of nonhuman primate species’ calls. From the start, we wanted to test the frontal mechanisms related to decision-based processes of humans when categorizing non-human primate calls hence the 2023 article we published. See lines 75-77 and we also added information on familiarization to the stimuli in the Methods, lines 679-682.

      The 16 acoustic features mentioned should be briefly defined earlier, as they are central.

      We feel like describing 16 acoustic parameters in the introduction would be heavy on the reader, so we instead added a reference to the supplementary table (Table S1) in which these are named and described. See line 80.

      Explain why only chimpanzees and bonobos were selected among the great apes, and discuss the value of including both, given their equal phylogenetic proximity but largely dissimilar acoustics.

      The stimuli were obtained by Thibaud Gruber and his team and through collaborations with Katie Slocombe and Zanna Clay. Unfortunately, at the time we could only use chimpanzee and bonobo calls for the great apes. Therefore, it was mainly a material constraint rather than a deliberate choice to exclude other great apes. We now discuss this aspect and present the absence of other great apes as a limitation (lines 587-591).

      Rephrase references to "recruitment" of TVA - this term implies general activation, while the key question concerns selectivity (stronger responses to voices vs. non-vocal controls).

      We rephrased throughout the manuscript, thank you for this suggestion.

      The hypothesis section should more clearly separate the acoustic and phylogenetic predictions, and clarify which earlier data motivate each.

      We now explicitly categorize the hypotheses according to either Bioacoustics or Phylogeny to clarify. We also added references motivating each hypothesis. See lines 114-120.

      (2) Methods

      Clarify whether stimuli were RMS-normalized or otherwise balanced for energy (line 128).

      Sound pressure level was kept constant but the stimuli were not normalized, specifically to avoid a negative impact on their naturality. We added a sentence (lines 131-132) including a reference on this aspect.

      The task design could benefit from reporting accuracy in addition to reaction times for the 4AFC species classification task.

      We agree this aspect was missing. We now report accuracy data (controlled for reaction times and acoustics of Model 3) for the species categorization task (lines 147-165; Fig.1B), and in the Methods (lines 769-786). The fitted regression values of this analysis are also used for a new fMRI model (Model 4), to uncover within-TVA correlates of the probability of correct species categorization (lines 309-325; Fig.4).

      Please note that previously, the behavioral data of the species categorization task were completely absent (N=23), and the reaction times data previously part of Fig.1 were for the species attentional bias task (independent sample of N=28). Since this aspect was not clear at all (same remark by all reviewers—apologies for that), we now include a clear separation in Fig.1, with newly added panels D & E part of a distinct figure area named: “Control task: Testing for Species attentional bias (N=28)”. Panel D illustrates the control task paradigm (each species as exogenous cue; “dot-probe” paradigm) while panel E shows the results (target sine wave tone or “bip” detection reaction times), showing that no species triggered more attentional capture than the others (Species effect non-significant).

      The acoustic parameters used in Models 2 and 3 should be explicitly listed in the Methods (even if already published elsewhere).

      In addition to their description in Table S1, we now include the 16 acoustic parameters used to calculate acoustic distance between the species in the Methods, see lines 828-844.

      Consider simplifying the presentation of the three models: a figure summarizing their relationships would help.

      We now include only one figure (Fig.2) for Model 3, and we pushed model 1&2 to the supplementary material. We also simplified Fig.3 for a clearer view of the overlaps between the 3 models within the TVA.

      The description of “systematic and thorough control of phylogeny” (line 119) is overstated, given that only three nonhuman species were included.

      We agree with the reviewer and we suppressed both “systematic” and “thorough” from the sentence.

      Provide rationale for not including a nonvocal control category (e.g., scrambled vocalizations or environmental sounds) to assess TVA selectivity.

      The main objective of the study was to uncover whether human participants could recognize the vocalizations from nonhuman primates—from both great apes and monkeys—as compared to the human voice. We therefore did not include nonvocal or noise stimuli. We added this point as a limitation in the Discussion (lines 593-596 and 609-611).

      Even though we did not include such stimuli for the reason mentioned above, the delineation of subtypes of nonvocal material within the TVA of our participants (Fig.2) are, in our opinion, clarifying the message: chimpanzee-selective activations are fully within ‘voice vs. animal’ and ‘voice vs. nature’ TVA subareas, while it is not the case in ‘voice vs. music’ and ‘voice vs. noise’ TVA subareas.

      Clarify if participants were trained or had a practice session to recognize the four species before scanning.

      The participants were indeed trained on 3 stimuli per species before entering the MRI scanner. These stimuli were discarded from the species categorization task. We added a sentence about this aspect, see lines 131-132.

      Specify what is meant by "no good or bad response" in the attentional control task (line 724).

      We suppressed this wording as it was highly confusing.

      (3) Results

      Behavioral accuracy should be reported to complement reaction times.

      We now added behavioral data for the species categorization task as well as the neural correlates of accurate species categorization. See our previous response above (‘‘‘).

      Figures 2-4 largely overlap; consider merging or simplifying to reduce redundancy.

      We agree and this point was raised by the other reviewers as well. Task-based results are now presented only for Model 3 as Fig.2, while Fig.3 (previously Fig.5) summarizes the overlap between the three models. Figures for Models 1 & 2, previously labelled Fig.3 and Fig.4, were moved to the supplementary material.

      Figure 2: Please indicate more clearly where "chimp-selective" areas are located (perhaps with zooms).

      We agree, we now modified Fig.2 with zoomed-in panels and a clearer outline of chimp-selective areas (solid blue outline). This outline is also referenced in the text (lines 236-237).

      Correction for multiple contrasts: With many pairwise tests, adjustments (Bonferroni or FDR) should be mentioned explicitly.

      We now specify ‘FDR correction at the voxel level’ at the beginning of the Results section (lines 195-198) as well as in each figure.

      Replace "specific to chimpanzee" with "selective for chimpanzee" to avoid implying exclusivity.

      We made the suggested replacement throughout the manuscript.

      Discuss whether the small macaque-related clusters might simply reflect acoustic overlap rather than true category selectivity.

      We added a section on this important aspect, including results that support the role of mid-STG/STS regions for more noise-like stimuli, including the use of macaque coos. See lines 450-461.

      (4) Discussion

      The discussion overstates claims of "chimpanzee-selectivity" in TVA. The evidence shows relative preference, not absolute selectivity.

      We now specify from the start of the Discussion that we are not interpreting the results as absolute selectivity but rather as more relative preference, see lines 371-373.

      The authors repeatedly conflate acoustic and phylogenetic factors; this should be explicitly acknowledged as a limitation.

      We agree, and we completed the limitations section already dedicated to this aspect by a more explicit account of the confound, see lines 609-611.

      Clarify what is meant by "recruitment" and "selectivity" (lines 411-419, 577). TVA activity often reflects enhanced responses to voices compared to non-vocal sounds, not exclusive activation.

      We clarified this wording in the Discussion (lines 377-378) and replaced another instance by “activated the […]” to make it clearer what we imply, namely enhanced activity triggered by chimpanzee calls within human TVA.

      The lack of non-vocal control conditions should be discussed as a major interpretive limitation.

      We added this point as a limitation in the Discussion (lines 593-596).

      The statement that "chimpanzee-selective activity" arose in humans who have never been exposed to chimp calls (line 450) invites evolutionary speculation but should be more cautiously phrased.

      We agree, and we rephrased by: “[…] with chimpanzee calls triggering responses in the anterior STG/TVA of our human participants […]”. See lines 432-433.

      The comparison to recent macaque data (Giamundo et al., 2024 PNAS) is crucial: these findings of human-voice-selective neurons in macaques directly parallel the present human-chimp result.

      We agree with the reviewer, and we are hopeful to read similar results for other apes/great apes in the future.

      Reviewer #2 (Recommendations for the authors):

      (1) The primate vocalizations used in this study were recorded in diverse social and emotional contexts, which may have contributed to the observed differences in TVA activation. Since the temporal voice areas are known to be sensitive to affective and socially relevant cues, these contextual differences could confound the interpretation of species-specific neural responses. Therefore, I suggest that the authors conduct a post-hoc analysis to quantify and compare the affective valence, arousal levels, and social contexts associated with each stimulus set.

      We agree that the TVA are sensitive to social—or socially relevant—cues, motivating the very thorough work of the expert reserve personnel on-site to accurately categorize the calls according to the very specific context they were produced in. If the reviewer meant presenting these stimuli to non-expert participants and asking them to categorize the context or valence, we think it would make no sense since the ratings would be completely below chance level and therefore uninformative. The newly added behavior—and model-based fmri—data include this crucial point, a factor that we named ‘Context’ in our analyses. In fact, for each species’ 18 stimuli, we control for agonistic and affiliative production context—split evenly, per species. Also, computing an additional posthoc analysis by splitting the stimuli according to Context would result in too few trials to get sensible and reliable fMRI results.

      That being said, our study targets this specific aspect by extracting the acoustic features that characterize our stimulus set the best, across context-species-valence-arousal, which is exactly what we want. Through the three types of modeling we used—from more simplistic to more elaborate the results converge only for one species: chimpanzee calls.

      We think the addition of behavioral data, model-based fMRI data, and the specific analysis on acoustic differences between chimpanzee and bonobo calls strengthens the message and the validity of our findings.

      (2) Although the author mentioned that the behavioral effects triggered by these vocalizations have been reported previously, the behavioral responses of the participants in the current study are also crucial for our understanding of the results. If the MRI data can be combined with the participants' behavioral responses for comprehensive analysis, the conclusions of this study will be more compelling.

      We agree with the reviewer, and we added the behavioral data—controlling for reaction times, production context and acoustics of interest—and we also included a model-based fMRI modeling of the probability of correct species categorization as Model 4, Fig.4. See, respectively: lines 147-165, Fig.1B; Methods, lines 769-786; Neuroimaging results, lines 309-325.

      (3) I am still not convinced that phylogenetic proximity drives the observed neural selectivity. While chimpanzee vocalizations do elicit stronger responses in anterior STG, the claim that this reflects evolutionary relatedness lacks evidence. If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree with the reviewer that generalizing our results in terms of phylogenetic proximity alone is not a viable option. Including many more primate species including other great apes would be necessary, and we mention this crucial aspect in the limitations section. We also insist in the Discussion on the interdependence between phylogeny and acoustics in our data, since: 1) we cannot fully disentangle these factors here, 2) we cannot attribute our results to either one or the other. See lines 387-390, 410-411, 473-477, 587-591.

      If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree, and nobody could disagree: if an auditory object is extremely similar to the human voice in terms of acoustics, it would therefore potentially activate the TVA. This is exactly our message: in the natural ‘auditory world’, the calls from chimpanzees seem to be among the very few animal auditory signals that are sufficiently close, acoustically, to the human voice and therefore trigger TVA activity. They also happen to be the calls from a species which is phylogenetically the closest to humans with minimal differences with other great apes. Our results are in that sense very aligned with work from the laboratory of Pascal Belin, namely on ‘voice patches’ in the primate brain located in the (anterior) TVA, cited in our manuscript.

      We therefore think our interpretation does not exclude that in the near future, similar results within the TVA could be observed for other auditory objects, and if animal, from a species potentially much more distant phylogenetically or from vocal signals of other great apes.

      We added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as scrambled or spectrum shifted per-species stimuli, which would have made the interpretation clearer identical acoustics but alteration/destruction of the species auditory object. See lines 593-596 and 609-611.

      Reviewer #3 (Recommendations for the authors):

      While the manuscript presents intriguing results, several concerns are raised for further consideration, detailed below.

      We thank the reviewer for evaluating the manuscript and for the constructive criticism and suggestions.

      Major concerns:

      (1) This study claims that bilateral anterior superior temporal gyrus (aSTG) in humans can be specifically activated by chimpanzee vocalizations rather than all other primate species after regressing out relevant acoustic parameters using three distinct analyses. I am wondering if a control stimulus (e.g., scrambled chimpanzee vocalizations) were presented, would the activation patterns in these same temporal voice areas (TVA) exhibit significant differences compared to the natural chimpanzee vocalizations?

      We completely agree with the reviewer, and this point was also raised by the other reviewers. We therefore added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as per-species scrambled or spectrum shifted stimuli, which would have made the interpretation clearer—identical acoustics but alteration/destruction of the species auditory object. See lines 609-611.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy. E.g:

      Figure 1C is the same as Figure S1. In addition, Figure 1C lacks a figure legend and descriptive label.

      The scatter plots in Figures 2D, 2H, 3D, 3H, and 4D, 4H are same as those in Figures S2, S3, and S4. However, some of these duplicate plots even have inconsistent axis labels.

      In several panels, the main figures appear to be summaries derived from the supplementary figures. The authors should organize these figures well to eliminate redundancy.

      Please double-check all the figures to make sure of accuracy.

      We agree that the figures were badly organized and were too crowded and redundant. We now suppressed the redundancy between Fig.1 and Fig.S1, and we reduced fMRI results to one figure for statistical Model 3 while the other models are in the supplementary data—we also justify this decision in the text by highlighting that model 3 is the most elaborate and sensitive one. Fig.3 (previously ‘Fig.5’) shows the overlaps between models and was simplified and clarified as well.

      (3) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task. It is possible that processing vocalizations from certain species requires more cognitive effort or induces higher decision uncertainty. Could the observed neural effects be confounded by the decision-making process itself?

      We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator. We now display these results in Fig.4 and we introduce the motivation factor for including a categorization task rather than more traditional passive listening (lines 75-77), as well as limitations, lines 595-596.

      (4) One interesting attempt of this study is to dissociate biologically salient information in animal vocalizations from their low-level acoustic properties. This presents a fundamental conceptual challenge: how to rigorously disentangle a vocalization's species-specific attributes from its inherent acoustic correlates. More precisely, what essential biological information persists in a species' vocal signal after statistically accounting for all quantifiable acoustic features? I recommend that the authors address it in the discussion.

      We thank the reviewer for this very important comment, and for suggesting we discuss it in the manuscript. We completely agree: we cannot fully orthogonalize species and acoustics, and this aspect relates also more broadly to cognitive and affective neuroscience studies involving vocal material. Namely: “What is an auditory object without acoustics?”

      We included a full paragraph on this aspect, see Discussion, lines 570-584.

      (5) If a brain region, such as TVA, is responsive to both acoustic parameters and biological meanings of animal vocalizations, the method used in this study might be inadequate by setting covariates to zero. It is possible that species information is embedded within a specific acoustic pattern. The current modeling approach may not capture such complex information and could potentially introduce bias when estimating the species effect. I recommend that the authors address this issue in the discussion.

      We thank the reviewer for this point once again, we addressed it in the Discussion, lines 581-584, and also in the section dedicated to study limitations, lines 609-613.

      (6) In the discussion, non-human primate vocalizations are "unreadable" to humans. If this is the case, what is the fundamental perceptual difference between these vocalizations and those from the other animal species? An alternative and highly plausible explanation for the findings is the differential familiarity of the participants with the various species, driven by media exposure (e.g., documentaries) or zoo visits and interactions. The authors need to provide a stronger justification for their control stimuli and directly address, either through discussion or additional analysis, how the factor of familiarity might explain their results better than the proposed "evolutionary distance" hypothesis.

      We now discuss this important aspect, see lines 560-569.

      We thought about doing additional analyses on this aspect but we concluded that we did not have any reliable indicators of familiarity for our participants, and additionally they were all recruited for being ‘unfamiliar’ with great apes or old-world monkeys’ vocalized communication.

      Also, frequent mismatches in the media between images of apes and the associated vocal signals (for instance, the depiction of a chimpanzee but with background audio of macaque coos) are not helping this cause.

      Minor:

      (1) No figure legend and result description for Figure 1.

      Figure 1 has a legend, maybe it was cut out during the uploading process, but it is present and verified now.

      (2) In the main text, three statistical models were referenced. Was the data used in each subsequent statistical model derived from the processed data of the preceding model? Please clearly explain this in the main text.

      We now specify this aspect in the Methods and the Results section to clarify that each model is independent from the others (lines 964-966 and 189-191, respectively).

      (3) In Figure 5, the two dashed lines representing Model 1 and Model 2 are confusing for readers.

      We modified the figure (now Fig.3) and simplified it by removing some outlines and clarifying the colors, therefore improving readability.

      (4) Lack of reaction times in the species categorization task.

      We clarified behavioral data, including the results for the species categorization task and for the control, exogenous cueing task, see modified Fig.1 and behavioral results section of the Results.

      (5) Figures 2, 3, 4, 5, Please keep the font size of the figure title consistent.

      Figure title font size were uniformized.

      (6) Line 201, Line 224, and so on, (EFG) → (E, F, G).

      We modified this aspect in every figure legend, including the supplementary material.

    1. To understand the educational benefit of home language literacy development, imagine an ELL whose home language is Arabic. This student arrives in his 3rd-grade classroom already reading at a 3rd-grade level in Arabic. He will need to learn that books in English are read from left to right and that what he considers to be the front of a book is the back of a book written in English. Through literacy instruction, he learns that both English and Arabic print depend on word order and progression for meaning and that letters and words in both languages represent sounds (although the student may not have heard or articulated some of the sounds in English).

      This section made me stop and think because I had never considered that a student could already be a strong reader in another language but still struggle in English. I've only worked with one ELL student, and this helped me realize that just because a student is still learning English doesn't mean they don't already have strong academic skills. As teachers, I think we need to recognize those strengths and use them to help students continue learning.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Response to Reviewer's Comments

      We thank the reviewers for their careful, constructive, and encouraging assessment of our manuscript. As described in detail in the point-by-point response below, we have extensively revised the manuscript and Supplementary Information. Together, these changes provide further support for the role of Rlig1 in neural function and visually guided behaviour during zebrafish development.

      Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      Summary: Provide a short summary of the findings and key conclusions (including methodology and model system(s) where appropriate).

      This study characterizes the function of RNA ligase 1 (Rlig1) in the vertebrate model zebrafish. Rlig1 is one of only two known RNA ligases in vertebrates, and its biological roles remain poorly understood. The authors combine gene expression analysis, loss-of-function approaches, transcriptomic profiling, calcium imaging, and behavioral assays to investigate its function during development. They show that loss of rlig1 (including maternal-zygotic loss) has no major effects on development or morphology, but that it leads to impairments in visually-guided behavior and altered neuronal activity in response to visual stimuli. Transcriptomic analyses reveal widespread dysregulation across multiple developmental stages, nominating genes that may underly the observed neural phenotypes. Together, the findings support a role for Rlig1 in neural development and function in vertebrates.

      We thank the reviewer for this accurate and positive summary of our study and for recognising the complementary, multi-level approaches used to examine the in vivo role of Rlig1.

      Major comments: - Are the key conclusions convincing?

      The key conclusion of this study is that Rlig1 plays an important role in the development and function of vertebrate neural circuits. Overall, this overarching conclusion, as well as the individual conclusions from each set of experiments, are well supported by the data presented. The combination of tissue-specific expression of rlig1, robust behavioral phenotypes in mutants, transcriptomic changes across multiple developmental stages, and circuit differences observed through calcium imaging provides a coherent, multi-faceted argument for the importance of this enzyme in brain development and function. While the precise RNA substrates of Rlig1 and the mechanistic link between transcriptomic changes and neural phenotypes remain to be defined, the authors clearly acknowledge these next steps and limitations. This study is a critical foundation for those future experiments.

      We appreciate the reviewer’s positive assessment of the strength and coherence of the evidence.

      • Should the authors qualify some of their claims as preliminary or speculative, or remove them altogether?

      The claims in the manuscript are generally well-supported. The authors clearly acknowledge limitations and future experiments to further dissect mechanism in the Discussion section.

      • Would additional experiments be essential to support the claims of the paper? Request additional experiments only where necessary for the paper as it is, and do not ask authors to open new lines of experimentation.

      No major additional experiments appear essential for supporting the current claims.

      • Are the suggested experiments realistic in terms of time and resources? It would help if you could add an estimated cost and time investment for substantial experiments.

      No experiments are required for the current claims of the manuscript.

      We thank the reviewer for this assessment.

      • Are the data and the methods presented in such a way that they can be reproduced?

      The methods are generally well described. I would suggest that the "raw images, data, and source code for custom scripts used in this work" be made accessible without having to request from the authors. Zenodo provides up to 50 GB of storage, which is likely sufficient for the data presented in this manuscript. In particular, I think it is important to share the behavior analysis, calcium imaging pipeline, and transcriptomics analysis. Even if all the data is too large, a sample dataset and analysis scripts should be publicly available.

      We agree and thank the reviewer for this important suggestion. To ensure that the study can be reproduced without the need to contact the authors, we have made the underlying data and custom analysis code publicly accessible. The RNA-seq data have been deposited in the GEO repository under accession number GSE308510 and are available at https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE308510.

      In addition, the raw imaging data, behavioural and calcium-imaging datasets, processed data, and custom scripts used for the behavioural, calcium-imaging, as well as the tRNA and rRNA sequencing data have been deposited on KonDATA (DOI: 10.48606/vpwgm69277srrgaj) – together more than 190 GB – and can be accessed using this link: https://kondata.uni-konstanz.de/radar/en/dataset/vpwgm69277srrgaj?token=gLEaYEENHjmHBhjhHUHK.

      We have revised the Data and code availability statement in the manuscript accordingly.

      • Are the experiments adequately replicated and statistical analysis adequate?

      The experiments appear adequately replicated, and statistical analyses are appropriate for the types of data presented.

      We thank the reviewer for this positive assessment. To further improve transparency, we have revised the figure legends and Methods to define sample-size notation consistently throughout the manuscript. As suggested by Reviewer 3, we now distinguish biological replicates or independent experiments (N) from individual embryos, larvae, cells, imaging planes, or trials (n), as appropriate.

      Minor comments: - Specific experimental issues that are easily addressable. - Are prior studies referenced appropriately? - Are the text and figures clear and accurate? - Do you have suggestions that would help the authors improve the presentation of their data and conclusions?

      Throughout the manuscript: use the prime symbol for 5/3 DNA/RNA instead of an apostrophe. The prime symbol is present in a small number of sentences, but mostly the apostrophe is used.

      We thank the reviewer for noting this. We have replaced apostrophes with prime symbols throughout the manuscript to ensure consistent notation of 5′ and 3′ RNA/DNA termini.

      Line 227: "Next, we compared the total number of neurons". The elavl3 driver labels brain cells in addition to neurons. - The authors compared to the total number of brain cells, but can they make any comments on the size of the brain across the various areas? I imagine this data is also accessible by analyzing the imaging already collected.

      The elavl3 promoter is widely used as a pan-neuronal driver in zebrafish. Our calcium-imaging experiments used the Tg(elavl3:H2B-GCaMP8s) line, in which nuclear-localised GCaMP8s is expressed under the control of the elavl3 regulatory region. This established configuration enables brain-wide functional imaging of neuronal activity in larval zebrafish.

      To assess whether differences in regional brain size might contribute to the observed phenotype, we quantified brain dimensions in 5 dpf larvae using the existing imaging data. Measurements were performed manually in Fiji in a blinded manner, with genotypes assigned only after completion of the analysis. We quantified tectum width, hindbrain width, and tectum length, as illustrated in the new Supplementary Figure 6.

      MZrlig1 larvae showed a modest reduction in tectum width (MZrlig1: 299 ± 14 µm; WT: 312 ± 10 µm; one-sided t-test, p = 0.00125) and tectum length (MZrlig1: 122 ± 5 µm; WT: 134 ± 9 µm; one-sided t-test, p = 1.03 × 10⁻⁵). In contrast, hindbrain width did not differ between genotypes (MZrlig1: 164 ± 10 µm; WT: 164 ± 10 µm; one-sided t-test, p = 0.52). Following assessment of data distribution, statistical significance was evaluated using one-sided t-tests with Bonferroni correction for three comparisons (n = 18 MZrlig1 and n = 20 WT larvae).

      Importantly, the unchanged hindbrain width indicates that the reduced number of motion-responsive hindbrain neurons in MZrlig1 larvae is unlikely to be explained by a gross difference in hindbrain size. These findings therefore support our interpretation that Rlig1 loss is associated with reduced neuronal responsiveness in the hindbrain.

      Given that there is already a mouse mutant for this gene and transcriptomics, can the authors do a more thorough job comparing the transcriptomics from that study with their own?

      We thank the reviewer for this helpful suggestion. When we applied the differential-expression thresholds used in our zebrafish analysis (absolute log₂ fold change ≥ 1.5 and adjusted p value ≤ 0.05) to the genes reported in the mouse study, only flg2 met these criteria. Thus, the available mouse dataset provides limited scope for a direct gene-by-gene comparison with our data.

      To extend our analysis beyond poly(A)-enriched mRNA sequencing, we additionally performed tRNA and rRNA sequencing using total RNA from 5 dpf WT and MZrlig1 larvae. The tRNA analysis identified 17 significantly altered tRNAs in MZrlig1 larvae, including seven upregulated and ten downregulated species (Figure 5i; Supplementary Tables 8–9). Notably, the affected tRNAs include tRNA-Lys-CTT, which was previously identified among RNAs enriched in human Rlig1 immunoprecipitates, and tRNA-Thr-CGT, which was reported to be increased in female rlig1 knockout mouse brains. Although the direction of change is not fully conserved across these studies, these overlaps further support the possibility that Rlig1 influences tRNA homeostasis.

      In parallel, rRNA sequencing revealed differential abundance of 122 5S rRNA transcripts, with 86 upregulated and 36 downregulated in MZrlig1 larvae (Figure 5h; Supplementary Tables 10–11). Together, these new analyses show that loss of Rlig1 is associated with altered abundance of both tRNA and rRNA species, consistent with previous evidence linking Rlig1 to RNA homeostasis. At the same time, we explicitly state that these data do not identify direct enzymatic substrates of Rlig1, but provide a resource and rationale for future mechanistic studies.

      A clearer statement on the similarities and differences of Rlig1 and RtcB would be helpful. Is it possible RtcB is compensating at all?

      We thank the reviewer for this comment. We have clarified the similarities and differences between Rlig1 and RtcB in the Introduction and Discussion. Although both enzymes catalyse RNA ligation, they act on distinct end chemistries. RtcB mediates 3′–5′ ligation of RNA ends generated during canonical tRNA splicing, joining a 5′-hydroxyl end to a 2′,3′-cyclic phosphate or 3′-phosphate end. In contrast, Rlig1 catalyses 5′–3′ ligation of RNA fragments bearing a 5′-phosphate and a 3′-hydroxyl group.

      These distinct substrate requirements make direct functional compensation by RtcB unlikely. RNA ends generated for ligation by Rlig1 would first require end processing to generate termini compatible with RtcB-mediated ligation. Nevertheless, indirect compensation or partial functional overlap after such processing cannot be excluded.

      We sought to address this question experimentally by obtaining rtcb mutants from the European Zebrafish Resource Center. However, subsequent genotyping showed that the supplied sperm did not contain the intended rtcb mutant alleles, precluding analysis in the present study. We have therefore explicitly acknowledged that the extent to which RtcB may compensate for loss of Rlig1 remains unresolved and will require analysis of validated rtcb mutant lines in future work.

      I examined the DEG tables, and I did not notice an obvious substantial enrichment of genes on chromosome 25 (White et al., 2022, https://doi.org/10.7554/eLife.72825). Were the different samples from different clutches or the same clutch? I may have missed it. Regardless, I would carefully check the DEGs that are important for conclusions and check that they are not on the same chromosome as rlig1. It is likely worth rerunning all of the GO/GSEA with genes on chromosome 25 excluded.

      We thank the reviewer for raising this potential confound. The RNA-seq samples were derived from independent clutches. To determine whether the observed transcriptional changes could be influenced by local effects associated with the rlig1 locus on chromosome 25, we performed two complementary analyses.

      First, we examined the chromosomal distribution of differentially expressed genes (DEGs) at each developmental stage. The chromosomal distribution was assessed using the original DEG analysis presented in the manuscript (no pre-filtering before DESeq2; DEGs defined as padj 1). Chromosome 25 contains 806 of 25,254 annotated protein-coding genes in the zebrafish genome, corresponding to 3.2% of all coding genes. Across developmental stages, the proportion of DEGs located on chromosome 25 ranged from 1.4% to 4.1% (cleavage: 12/419; sphere: 17/; shield: 37/892; bud: 26/781; 1 dpf: 3/216; 5 dpf: 8/587). Relative to the genomic expectation, this corresponds to enrichment values between 0.43- and 1.30-fold. Only the shield stage showed a modest increase in the proportion of chromosome 25 DEGs (1.30-fold), whereas all other stages were at or below the genomic expectation. Thus, genes on chromosome 25 are not globally overrepresented among the DEGs in the rlig1 mutant dataset.

      Second, we repeated the complete differential-expression analysis for each developmental stage after excluding all chromosome 25 genes before DESeq2 normalisation, size-factor estimation, and dispersion modelling. This re-analysis was performed using an updated workflow, including removal of genes with zero total counts prior to DESeq2, which changes the number of genes entering Benjamini–Hochberg correction and consequently the total number of detected DEGs; all other analysis parameters were identical to the original analysis. This approach ensured that chromosome 25 genes could not influence either normalisation or statistical inference for genes on other chromosomes. Using the same DEG thresholds as in the original analysis (padj 1), exclusion of chromosome 25 had only minimal effects on the remaining DEG sets.

      Stage

      Full DEGs

      Non-Chr25 DEGs

      Lost (Chr25)

      Lost (non-Chr25)

      Gained

      1 (4-cell)

      419

      415

      5

      0

      1

      2 (Sphere)

      913

      879

      34

      5

      5

      3 (Shield)

      592

      553

      37

      5

      3

      4 (Bud)

      349

      329

      20

      0

      0

      5 (1 dpf)

      7

      6

      1

      0

      0

      6 (5 dpf)

      168

      164

      4

      0

      0

      Across all six developmental stages, only ten non-chromosome-25 genes lost significance and nine genes gained significance. These minor changes were confined largely to the sphere and shield stages, which also showed the highest relative representation of chromosome 25 DEGs. At the 4-cell, bud, 1 dpf, and 5 dpf stages, no non-chromosome-25 genes lost significance after chromosome 25 was excluded.

      We also repeated the GO and GSEA analyses after excluding chromosome 25 genes. As expected, a small number of individual terms changed; however, the principal enrichment patterns and overall biological interpretation remained unchanged. Together, these analyses indicate that the transcriptomic phenotype is not substantially driven by chromosome 25-linked DEGs or by local effects associated with the edited rlig1 locus. While this analysis cannot exclude effects on individual linked genes, it shows that such effects do not substantially affect the main transcriptional or pathway-level conclusions of the study.

      **Referees cross-commenting**

      I missed the point about the RNA-seq samples being cousin-matched. While I am optimistic that the results won't change, I agree with Reviewer #3 that some confirmation is necessary. It was unclear to me whether the samples were from the same or different clutches - if they are from different clutches and share overlapping genes, that would also add support to the results. I think that detail was missing from the methods, and I had pointed it out. Either additional RNA-seq or even qPCR of some top genes from a heterozygous incross is a reasonable request.

      We thank the reviewer for raising this point and apologise that the breeding design for the transcriptomic experiments was not described sufficiently clearly. The developmental RNA-seq samples were not cousin-matched. Rather, WT and MZrlig1 embryos were collected from separate group matings and therefore originated from different clutches. Independent pooled samples were analysed at each developmental stage, as now described explicitly in the revised Methods.

      We agree that independent validation in a sibling-controlled genetic setting is important. We therefore performed RT-qPCR for eight genes selected from the 5 dpf mRNA-seq dataset using sibling-matched zygotic rlig1 mutants and WT larvae generated by heterozygous incrosses. For each genotype, three independent biological replicates were analysed, with four larvae per sample. Six of the eight selected genes showed changes in the same direction as in the original MZrlig1 RNA-seq dataset: cyp2p9, itln3, sult3st4, fabp7b, hamp, and rlig1 itself. In particular, itln3 remained strongly upregulated, whereas rlig1 expression was markedly reduced in the sibling-matched zygotic mutants. In contrast, gdf3 and gstp1.1 did not show the same directional change in this validation experiment.

      These results provide independent support that several of the transcriptional changes identified in the MZrlig1 RNA-seq dataset are also observed in sibling-matched zygotic mutants. At the same time, the incomplete concordance of individual genes is consistent with the fact that maternal-zygotic and zygotic mutants represent biologically distinct conditions and may differ in both effect size and molecular consequences. We have added these validation data as Supplementary Figure 7 and revised the Results and Methods accordingly.

      Reviewer #1 (Significance (Required)):

      • Describe the nature and significance of the advance (e.g. conceptual, technical, clinical) for the field.

      This study provides a conceptual and biological advance by identifying a role for a vertebrate RNA ligase in brain development, behavior, and transcriptional regulation.

      • Place the work in the context of the existing literature (provide references, where appropriate).

      Although RNA ligases from single-cell organisms and phage are well-characterized, the roles of RNA ligases in vertebrates are relatively understudied. There are only two, including the one the one that is the focus of this manuscript. This study demonstrates an in vivo function for Rlig1, linking molecular changes to neural development and function. The Rlig1 enzyme was only very recently discovered (2023), making this work timely and an important addition to an area with relatively few studies.

      A major strength of the study is its multi-level approach, integrating diverse techniques to coherently link this gene to organism-level phenotypes. This work provides a strong conceptual and functional advance by demonstrating a role for Rlig1 in vertebrate neural circuit function and behavior. A remaining mechanistic gap is that the direct RNA substrates of Rlig1 are not identified, and the observed transcriptomic changes in mRNA are likely downstream consequences of its loss. However, these points are clearly acknowledged in the discussion, making the study a well-balanced contribution. Given the existence of a mouse knockout model, further discussion comparing the zebrafish transcriptomic results and phenotypes to those observed in mouse would help place this work in the context of prior studies. Overall, the main conclusions are well supported, and the limitations do not undermine them. This study represents an important contribution that establishes a foundation for future mechanistic work linking Rlig1 substrates to the observed phenotypes.

      We thank the reviewer for this thoughtful and encouraging assessment.

      • State what audience might be interested in and influenced by the reported findings.

      Zebrafish basic science researchers, particuarly those studying how genes lead to altered neural circuits and behavior, are the most direct target audience. However, the work is of more broad interest to those in the fields of neurodevelopment, gene regulation, and RNA biology / processing.

      • Define your field of expertise with a few keywords to help the authors contextualize your point of view. Indicate if there are any parts of the paper that you do not have sufficient expertise to evaluate.

      I am comfortable evaluating zebrafish mutants, transcriptomics, and behavioral assay design. I have more limited experiment in neural circuit anaysis and interpretation of calcium imaging data, though this part of the manuscript was also clearly presented and understandable.

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      Summary Klusman et al have investigated the function of the RNA ligase rlig1 in zebrafish. They first document expression of the gene, by quantitative RT-PCR and HCR-fluorescent in situ hybridization. They then test ligase activity of the Rlig1 protein in vitro. They next generate a null mutant and test function of the visual system using behaviour as well as calcium imaging. The data indicate that rlig1 is broadly expressed and capable of ligating RNA; loss of rlig1 has mild effects on overall development and pronounced effects on behavioural and neuronal response to visual stimuli. Finally, the authors use bulk transcriptome analysis to identify changes in gene expression in the mutants.

      We thank the reviewer for this accurate summary of our study and for recognising that the behavioural and calcium-imaging results together support a role for rlig1 in visual processing and visually guided behaviour.

      **Referees cross-commenting**

      I agree that more details are required about the crosses would be useful.

      We also agree that further detail on the breeding schemes is important. We have therefore expanded the Methods and figure legends to describe the crosses used for each experiment, including the relationship between mutant and control animals and whether samples were sibling- or cousin-matched.

      Reviewer #2 (Significance (Required)):

      Overall, the conclusions that rlig1 is required for normal development of the embryo, especially of a fully functioning visual system, are well supported. The optomotor response experiments have high power and, together with functional imaging, show a clear difference between mutant and wildtype.

      One limitation of this manuscript is in the characterization of gene expression. The gene expression database in Zfin contains one image of rlig1 (https://zfin.org/ZDB-IMAGE-060710-1925#image), which shows broad expression in cells of the embryo and larvae and no expression in the yolk. The images here, with the exception of the mutant in Figure 3C, show expression in the yolk. This would suggest that the yolk signal is not autofluorescence, which is inconsistent with the Thisses' data. Additonally, Figure S1 indicates a variable level of non-specific signal, especially in panel g. Thus, the distribution of rlig1 mRNA is unclear.

      We agree that the yolk-associated signal should not be interpreted as specific rlig1 expression.

      rlig1 transcripts are completely absent from the RNA-seq datasets of MZrlig1 mutants at all developmental stages analysed. Thus, the variable fluorescence observed in the yolk and in the no-probe controls (Supplementary Figure 1) cannot represent residual rlig1 expression, but must reflect non-specific background signal and/or autofluorescence. We have clarified this point in the revised manuscript.

      The transcriptome analysis identified changes in gene expression in the mutant. This establishes a role for rlig1 in development, and identifies several processes that are disrupted by loss of rlig1. However, the molecular analysis sheds little light on direct targets of the ligase. Given the established effects on tRNA, for example, it is unclear why RNA was analysed only by short reads on poly(A) RNA. The reader is left wondering whether zebrafish tRNA contains introns that require Rlig1 for processing. In this context, it would be useful for the authors to provide more background on tRNA splicing in vertebrates, including a mention of tricRNA, and potentially the role of TSEN complex in brain development.

      We have expanded the Introduction as suggested to provide additional context on tRNA splicing in vertebrates. We now explain that canonical tRNA splicing is initiated by the TSEN complex and completed by RtcB, which ligates RNA ends with chemistries distinct from those used by Rlig1. We also discuss that excised tRNA introns can form stable tRNA intronic circular RNAs (tricRNAs), and that defects in TSEN complex components are associated with neurodevelopmental disorders, underscoring the importance of RNA processing for nervous-system development.

      We agree that our poly(A)-enriched RNA-seq data do not identify direct RNA substrates of Rlig1. We have clarified throughout the manuscript that these experiments were designed to characterise downstream transcriptional consequences of rlig1 loss.

      We have additionally analysed tRNA and rRNA abundance in total RNA from 5 dpf WT and MZrlig1 larvae. These analyses identified altered levels of specific tRNA and 5S rRNA species in MZrlig1 larvae (Figure 5h,i; Supplementary Tables 8–11), supporting an association between Rlig1 loss and altered RNA homeostasis.

      To summarize, this manuscript extends work in the mouse and in cell lines that demonstrate a requirement for rlig1. It does not shed light on direct targets of Rlig1, but provides a strong foundation for future work on the role of RNA ligation in vertebrate development and brain function.

      This paper is expected to be of interest to a specialised audience.

      Minor points: The images showing gene expression in Figure 2 are not easy to see, due to the LUT used and low intensity of the signal. To aid the reader, the HCR channel should be shown in grayscale, possibly with the contrast enhanced (to the same extent in all images).

      To improve the visibility and interpretation of the HCR signal, we have added a new Supplementary Figure 2 showing the rlig1 channel in greyscale. Within comparable developmental-stage panels, identical contrast settings were applied to all images.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      Summary This paper provides good evidence that a newly described enzyme that catalyzes 5'-3' RNA ligation - rlig1 - plays some role in early vertebrate neurodevelopment. Using embryonic and larval zebrafish as a model, they found that, while rlig1 mRNA is highly maternally deposited and ubiquitously expressed early on, expression later in development localizes to the brain and eyes. They generated a stable CRISPR/Cas9 large deletion mutant spanning from upstream the 5'UTR past the start codon. By comparing wild type and maternal-zygotic (MZ) rlig1 mutants, the authors found that animals developed overtly normally but did show reduced behavioral responsiveness to a visual stimulus experimental paradigm. By combining calcium imaging and poly(A)-enriched RNA-sequencing transcriptomic analyses, they found that there was decreased neuronal activity in regions needed for visual processing, and that there was dysregulation of neural-related gene networks and metabolic and translational pathways.

      We thank the reviewer for this detailed and accurate summary of our study and for recognising the convergent evidence linking Rlig1 loss to altered neural activity and visually guided behaviour in developing zebrafish.

      Major comments 1) My main major comment is that, because there is so much inherent variability in behavior and even development across different clutches, this study relies on comparing (cousin-matched) WT and maternal-zygotic rlig1 mutant animals. In most reliable peer-reviewed papers, this is not a fair comparison. While I appreciate that authors stated that they used parents that were siblings (so, offspring would be cousin-matched), I do not consider this scientifically rigorous enough for the claims presented. a. I do not consider it a reasonable request to ignore the massive amount of work that went into this paper using WT and MZrlig1 comparisons. However, at minimum, authors should consider performing essential behavior and RNA-seq (see point b) experiments with heterozygous incrosses of single-pair matings, and genotyping the animals post-hoc. Including this critical data in a main figure, as the basis for using MZ animals for the rest of the paper, would induce some confidence that the phenotypes and claims presented are not a result of inherent variability. If the authors already have adult heterozygous animals of mating age, I estimate that these experiments may be completed very reasonably within 3-4 weeks; if new animals need to be generated, this request would take ~4 months. Typically, these kinds of experiments would not be considered a financial burden to perform.

      Our central genetic condition was maternal-zygotic loss of rlig1, motivated by the strong maternal deposition of rlig1 mRNA during cleavage stages. A heterozygous incross would produce zygotic mutants that still receive maternal rlig1 transcript and protein, and would therefore test a related but biologically distinct condition. For the maternal–zygotic experiments, we used cousin-matched WT controls derived from the same parental family to minimise genetic-background differences, and we performed the behavioural assays with substantial numbers of larvae across independent experiments.

      We nevertheless repeated the behavioural analysis as suggested using zygotic rlig1 mutants and WT sibling controls obtained from heterozygous incrosses. This analysis revealed a qualitatively similar, although less pronounced, reduction in visually guided behaviour in zygotic mutants (new Supplementary Figure 4). We speculate that the reduced effect size is consistent with partial compensation by maternally supplied rlig1 transcript or protein in zygotic mutants.

      b. For transcriptomic analyses, I have two main points: i) again, it is difficult to statistically rigorously compare transcriptomes of nonsibling-matched animals with such low numbers of single 5 dpf brains. In line with point a, it would be essential to pool at least a few WT and rlig1 mutant siblings for at least 3 biological replicates per samples and compare those analyses with the results from MZ animals. ii) Typically this would not be a major concern, however given the nature of the gene of interest and published in vitro findings, I do consider that the rlig1 enzyme catalyzes 5'-3' RNA ligation, has been shown to be implicated in rRNA integrity and tRNA targeting, and is broadly essential for repair, splicing, and editing of RNAs. Thus, while the poly(A)-enriched RNA sequencing can provide context about gene networks that are affected (either primarily or secondarily), sequencing that enriches for tRNAs, polysome profiling or ribosome profiling, or some more targeted sequencing approach would be more appropriate to more rigorously support the claims in the paper. Depending on readiness of mating-age animals, this experiment and analyses may reasonably take up to 3 months; this approach may be considered a financial burden. Alternatively, with the current mRNA sequencing, the authors could delve into whether they can identify altered splicing or RNA editing dynamics in different RNA modules. I estimate that this alternative analysis approach may take up to one month to develop and interpret.

      We would like to clarify that the poly(A)-enriched RNA-seq was not performed on single 5 dpf brains, but on independent pools of 8–10 age- and genotype-matched whole embryos or larvae collected across six developmental stages. We have also validated eight selected 5 dpf RNA-seq candidates by RT-qPCR using sibling-matched zygotic rlig1 mutants and WT larvae generated by heterozygous incrosses. For each genotype, we analysed three independent biological replicates, each comprising a pool of four larvae. Six of the eight tested genes showed changes in the same direction as in the original MZrlig1 RNA-seq dataset, including cyp2p9, itln3, fabp7b, hamp, sult3st4, and rlig1 (new Supplementary Figure 7). Although zygotic mutants are not equivalent to maternal–zygotic mutants because they retain maternally supplied rlig1 transcript and protein, these results provide independent support for a substantial subset of the transcriptional changes identified in the MZrlig1 dataset. We have revised the Methods, Results, and Discussion to describe the breeding schemes and this limitation more explicitly.

      We also agree that poly(A)-enriched RNA-seq alone cannot identify direct Rlig1 substrates or adequately assess non-polyadenylated RNA classes. We therefore added targeted analyses of tRNA and rRNA abundance from total RNA isolated from 5 dpf WT and MZrlig1 larvae. The tRNA analysis identified seven tRNAs with increased and ten with decreased abundance in MZrlig1 larvae, including tRNA-Lys-CTT, previously found among RNAs enriched in human Rlig1 immunoprecipitates, and tRNA-Thr-CGT, which was reported to be increased in female rlig1 knockout mouse brains (Figure 5i; Supplementary Tables 8–9). In parallel, the rRNA analysis identified altered abundance of 122 5S rRNA species, with 86 increased and 36 decreased in MZrlig1 larvae (Figure 5h; Supplementary Tables 10–11).

      These new data provide additional evidence that loss of Rlig1 is associated with altered tRNA and rRNA homeostasis. At the same time, we explicitly state that neither the mRNA-, tRNA-, nor rRNA-seq datasets establish direct enzymatic substrates of Rlig1 or demonstrate altered tRNA splicing, RNA editing, or translation. Direct substrate mapping and analyses such as ribosome profiling will be important directions for future work. The revised manuscript frames the transcriptomic analyses accordingly.

      o The experiments as documented are adequately replicated and statistical analyses adequate (minus the nonsibling-matched point 1). I note that labels should more clearly state or denote individual (n) or experimental (N) numbers, some of which I provide in Minor comments below.

      We agree and have revised the figure legends accordingly. We now distinguish N for independent experiments or biological replicates from n for individual embryos, larvae, imaging planes, segmented cells or trials. Where pooled samples were used, the legends and Methods now state the number of embryos or larvae per pool and the number of independent pools or experiments.

      Minor comments Comments on figures or figure legends: 1) Figure 1e, align the "#" labels better, they look diagonal.

      Thank you. We corrected the alignment of the labels in Figure 1e.

      2) For 1f, consider labeling independent replicates directly on the graph instead of just the label, otherwise not very clear to the reader.

      We have revised Figure 1f to make the independent replicates more transparent. The figure now clearly indicates the number of independent replicates used for quantification. Every replicate has a different colour now, and N = 3 is indicated in the figure.

      3) Figure 2a, consider adding the reference gene (eef1a) in the legend.

      We have added eef1a to the Figure 2a legend and clarified that relative rlig1 mRNA levels were calculated using eef1a as the reference gene.

      4) Figure 2a - if I understand the experiment correctly, the current label n=3 (which would mean 3 individual embryos/larvae) should read N=3 (three independent experiments of x number of embryos/larvae per run)

      Thank you very much for this suggestion. We have corrected the sample-size notation in Figure 2a. The label now uses N for independent experiments and specifies the number of embryos or larvae used per experiment where appropriate.

      5) Supplementary Figure 1 was very unconvincing comparing WT to MZ mutants, I'm sorry to say I really could not tell much difference. When compared to Figure 3c, they look quite different. The DRAQ7 labeling also appeared uneven in Supplementary Figure 1. Consider optimizing the imaging strategy and providing more interpretably images. A separate, aesthetic comment - magenta was very difficult for me to see against a black background, consider switching the rlig1 channel to grayscale or flip the colors so that rlig1 mRNA is cyan, for example.

      We thank the reviewer for this comment and apologise that the purpose of Supplementary Figure 1 was not sufficiently clear. This figure shows no-probe control samples imaged in the rlig1 detection channel to document stage-dependent background and autofluorescence. Because no rlig1 probe was applied, no genotype-dependent difference between WT and MZrlig1 samples is expected in these images. The variable signal, including the yolk-associated fluorescence, therefore represents background rather than specific rlig1 mRNA detection.

      In contrast, Figure 3c shows samples processed with the rlig1 HCR probe set. The marked reduction of punctate signal in MZrlig1 larvae in this experiment is therefore attributable to the absence of rlig1 transcripts, consistent with the RNA-seq and RT-qPCR data. We have clarified this distinction in the revised text and figure legends.

      The apparently uneven DRAQ7 signal in some no-probe control images reflects differences in embryo orientation and imaging planes rather than genotype-specific staining differences. To improve the visibility and interpretability of the HCR data, we have additionally included a new Supplementary Figure 2 showing the rlig1 channel in greyscale, with matched contrast settings within comparable developmental-stage panels.

      6) Calcium imaging - related to Major comments above, consider performing this experiment in sibling-matched animals, especially with only one copy of the transgene. If WT vs. sibling mutant results look similar to the WT vs MZ mutant results, this would be more convincing.

      We agree that calcium imaging in sibling-matched zygotic mutants would provide a valuable complementary dataset. However, zygotic mutants retain maternally supplied rlig1 transcript and protein and therefore represent a biologically distinct condition from the maternal–zygotic mutants examined in our principal imaging experiments. Consistent with this distinction, the behavioural phenotype in sibling-matched zygotic mutants was qualitatively similar but less pronounced than in maternal–zygotic mutants.

      A sufficiently powered brain-wide calcium-imaging analysis in sibling-matched animals would require generation, imaging, and analysis of a substantial additional cohort, while the expected smaller effect size would limit its ability to directly test the maternal–zygotic phenotype reported here. We therefore believe that this experiment extends beyond the scope of the present study.

      **Referees cross-commenting**

      I agree with Reviewer #1 that at least the raw code is uploaded to GitHub or Zenodo, and raw data to be uploaded to Zenodo.

      We agree and thank the reviewer for this important suggestion. To ensure that the study can be reproduced without the need to contact the authors, we have made the underlying data and custom analysis code publicly accessible. The RNA-seq data have been deposited in the GEO repository under accession number GSE308510 and are available at https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE308510.

      In addition, the raw imaging data, behavioural and calcium-imaging datasets, processed data, and custom scripts used for the behavioural, calcium-imaging, as well as the tRNA and rRNA sequencing data have been deposited on KonDATA (DOI: 10.48606/vpwgm69277srrgaj) – together more than 190 GB – and can be accessed using this link: https://kondata.uni-konstanz.de/radar/en/dataset/vpwgm69277srrgaj?token=gLEaYEENHjmHBhjhHUHK.

      We have revised the Data and code availability statement in the manuscript accordingly (also see the response to Reviewer #1).

      I agree with Reviewer #1 that brain size can and should also be assessed, presumably using the same images already collected. For example, in Figure 5b, number of neural cells (even when normalized) could be lower if brain size is small. Reasonable control analysis.

      As suggested, we have quantified tectum width, tectum length, and hindbrain width from the existing calcium-imaging datasets in a blinded manner. Although MZrlig1 larvae showed modest reductions in tectum width and length, hindbrain width did not differ between genotypes. Thus, the reduced number of motion-responsive hindbrain cells is unlikely to be explained by a gross difference in hindbrain size. These control analyses are presented in the new Supplementary Figure 6 (also see the response to Reviewer #1).

      I agree with Reviewer #2 that addressing, either by writing or experimentally, a bit more about direct targets of the ligase (including tRNAs and rRNAs) will strengthen the manuscript significantly.

      We thank the reviewer for this helpful suggestion. To address this point, we have added new analyses of rRNA and tRNA abundance in 5 dpf WT and MZrlig1 larvae, together with an expanded discussion of their interpretation. These data provide additional evidence that loss of Rlig1 is associated with altered RNA homeostasis, while we distinguish such effects from the direct RNA substrates of the ligase, which remain to be identified (also see the response to Reviewer #2).

      I agree with Reviewer #1 first comment (last sentence) that, if RNA-seq (or other appropriate sequencing) of sibling-matched samples is financially prohibitive, then at least qPCR of some top genes would be acceptable.

      We have performed RT-qPCR validation of selected top differentially expressed genes using sibling-matched WT and zygotic rlig1 mutant larvae generated by heterozygous incrosses. These data provide independent support for the altered expression of several genes identified in the maternal–zygotic rlig1 RNA-seq dataset and are presented in new Supplementary Figure 9 (also see the response to Reviewer #1).

      I agree with the additional comment from Reviewer #1 - the manuscript details cousin-matched samples in lines 666-667, but I'd like to add a suggestion that the authors include details about "single-pair" versus "group-mating". For behavior and all analyses in these kinds of zebrafish experiments, it is very important that multiple replicates of single-pair (one female crossed to one male), sibling-matched groups are used.

      We appreciate the reviewer’s helpful suggestion. We agree that further detail on the breeding schemes is important. We have therefore expanded the Methods to specify, for each experiment, whether embryos or larvae were obtained from single-pair or group matings, the number of independent crosses or clutches, and whether mutant and control animals were sibling- or cousin-matched.

      Reviewer #3 (Significance (Required)):

      This study provides a good increase in our knowledge about a newly described RNA ligase enzyme - rlig1 - in vivo. The authors integrate their results across organismal behavior, brain cell activity, and transcriptomes using a newly generated stable genetic mutant to uncover a new link between neuronal RNA processing, development, and sensory-motor computation. Given that the human orthologue of this gene has been associated with neurological and cognitive conditions, including neurodevelopmental and neuroinflammatory disorders and Alzheimer's disease, the generation and characterization of this stable mutant line proves valuable. There are important technical limitations, specifically related to the comparison of wild type and maternal-zygotic mutant animals, that may not faithfully represent statistical differences compared to sibling-matched animals. Basic biological audiences, including in neurodevelopment, genetics, and RNA biology, would be interested in this research.

      We thank the reviewer for recognising the value of the stable rlig1 mutant line and for highlighting the importance of the breeding design. We agree that comparisons between cousin-matched WT and maternal–zygotic (MZ) mutant larvae require careful interpretation. However, a fully sibling-matched WT versus MZrlig1 comparison is not genetically possible. Maternal–zygotic mutants must be produced by homozygous mutant mothers, whereas WT siblings can only be obtained from a different maternal genotype. Thus, the maternal genotype and, critically, the presence or absence of maternally deposited rlig1 RNA and protein – necessarily differs between these conditions. This is not merely a technical limitation of the experimental design, but an intrinsic feature of testing maternal–zygotic gene function. A heterozygous incross instead produces sibling-matched zygotic mutants, which retain maternal rlig1 products and therefore represent a biologically distinct genetic condition rather than a direct replacement for the MZ comparison.

      For the MZ experiments, we minimised genetic-background differences by using cousin-matched controls derived from the same parental family and by analysing independent experimental replicates. Importantly, the principal behavioural finding was independently supported in sibling-matched zygotic mutants generated by heterozygous incrosses. These larvae showed a qualitatively similar reduction in visually guided behaviour, although with a smaller effect size (new Supplementary Figure 4). We also validated selected transcriptional changes in sibling-matched zygotic mutants by RT-qPCR (new Supplementary Figure 9). The weaker phenotype in zygotic mutants is consistent with partial buffering by maternal rlig1 transcript or protein. Future studies will be valuable to further separate how maternal and zygotic Rlig1 affects gene expression and visually guided behaviour.

      Insufficient expertise to evaluate: While I understand the first part of Figure 1, I do not have expertise in these sorts of assays. The rest of the experiments I do have sufficient expertise to evaluate. And thank you to the authors for providing direct DOI links to references.

      We are grateful for the reviewers’ detailed comments, which substantially improved the manuscript. We hope that the revised text and additional analyses address the central concerns and make the study more transparent and useful to the field.

    2. Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.

      Learn more at Review Commons


      Referee #1

      Evidence, reproducibility and clarity

      Summary:

      Provide a short summary of the findings and key conclusions (including methodology and model system(s) where appropriate).

      This study characterizes the function of RNA ligase 1 (Rlig1) in the vertebrate model zebrafish. Rlig1 is one of only two known RNA ligases in vertebrates, and its biological roles remain poorly understood. The authors combine gene expression analysis, loss-of-function approaches, transcriptomic profiling, calcium imaging, and behavioral assays to investigate its function during development. They show that loss of rlig1 (including maternal-zygotic loss) has no major effects on development or morphology, but that it leads to impairments in visually-guided behavior and altered neuronal activity in response to visual stimuli. Transcriptomic analyses reveal widespread dysregulation across multiple developmental stages, nominating genes that may underly the observed neural phenotypes. Together, the findings support a role for Rlig1 in neural development and function in vertebrates.

      Major comments:

      • Are the key conclusions convincing?

      The key conclusion of this study is that Rlig1 plays an important role in the development and function of vertebrate neural circuits. Overall, this overarching conclusion, as well as the individual conclusions from each set of experiments, are well supported by the data presented. The combination of tissue-specific expression of rlig1, robust behavioral phenotypes in mutants, transcriptomic changes across multiple developmental stages, and circuit differences observed through calcium imaging provides a coherent, multi-faceted argument for the importance of this enzyme in brain development and function. While the precise RNA substrates of Rlig1 and the mechanistic link between transcriptomic changes and neural phenotypes remain to be defined, the authors clearly acknowledge these next steps and limitations. This study is a critical foundation for those future experiments. - Should the authors qualify some of their claims as preliminary or speculative, or remove them altogether?

      The claims in the manuscript are generally well-supported. The authors clearly acknowledge limitations and future experiments to further dissect mechanism in the Discussion section. - Would additional experiments be essential to support the claims of the paper? Request additional experiments only where necessary for the paper as it is, and do not ask authors to open new lines of experimentation.

      No major additional experiments appear essential for supporting the current claims. - Are the suggested experiments realistic in terms of time and resources? It would help if you could add an estimated cost and time investment for substantial experiments.

      No experiments are required for the current claims of the manuscript. - Are the data and the methods presented in such a way that they can be reproduced?

      The methods are generally well described. I would suggest that the "raw images, data, and source code for custom scripts used in this work" be made accessible without having to request from the authors. Zenodo provides up to 50 GB of storage, which is likely sufficient for the data presented in this manuscript. In particular, I think it is important to share the behavior analysis, calcium imaging pipeline, and transcriptomics analysis. Even if all the data is too large, a sample dataset and analysis scripts should be publicly available. - Are the experiments adequately replicated and statistical analysis adequate?

      The experiments appear adequately replicated, and statistical analyses are appropriate for the types of data presented.

      Minor comments:

      • Specific experimental issues that are easily addressable.
      • Are prior studies referenced appropriately?
      • Are the text and figures clear and accurate?
      • Do you have suggestions that would help the authors improve the presentation of their data and conclusions?
      • Throughout the manuscript: use the prime symbol for 5/3 DNA/RNA instead of an apostrophe. The prime symbol is present in a small number of sentences, but mostly the apostrophe is used.
      • Line 227: "Next, we compared the total number of neurons". The elavl3 driver labels brain cells in addition to neurons.
      • The authors compared to the total number of brain cells, but can they make any comments on the size of the brain across the various areas? I imagine this data is also accessible by analyzing the imaging already collected.
      • Given that there is already a mouse mutant for this gene and transcriptomics, can the authors do a more thorough job comparing the transcriptomics from that study with their own?
      • A clearer statement on the similarities and differences of Rlig1 and RtcB would be helpful. Is it possible RtcB is compensating at all?
      • I examined the DEG tables, and I did not notice an obvious substantial enrichment of genes on chromosome 25 (White et al., 2022, https://doi.org/10.7554/eLife.72825). Were the different samples from different clutches or the same clutch? I may have missed it. Regardless, I would carefully check the DEGs that are important for conclusions and check that they are not on the same chromosome as rlig1. It is likely worth rerunning all of the GO/GSEA with genes on chromosome 25 excluded.

      Referees cross-commenting

      I missed the point about the RNA-seq samples being cousin-matched. While I am optimistic that the results won't change, I agree with Reviewer #3 that some confirmation is necessary. It was unclear to me whether the samples were from the same or different clutches - if they are from different clutches and share overlapping genes, that would also add support to the results. I think that detail was missing from the methods, and I had pointed it out. Either additional RNA-seq or even qPCR of some top genes from a heterozygous incross is a reasonable request.

      Significance

      • Describe the nature and significance of the advance (e.g. conceptual, technical, clinical) for the field.

      This study provides a conceptual and biological advance by identifying a role for a vertebrate RNA ligase in brain development, behavior, and transcriptional regulation. - Place the work in the context of the existing literature (provide references, where appropriate).

      Although RNA ligases from single-cell organisms and phage are well-characterized, the roles of RNA ligases in vertebrates are relatively understudied. There are only two, including the one the one that is the focus of this manuscript. This study demonstrates an in vivo function for Rlig1, linking molecular changes to neural development and function. The Rlig1 enzyme was only very recently discovered (2023), making this work timely and an important addition to an area with relatively few studies.

      A major strength of the study is its multi-level approach, integrating diverse techniques to coherently link this gene to organism-level phenotypes. This work provides a strong conceptual and functional advance by demonstrating a role for Rlig1 in vertebrate neural circuit function and behavior. A remaining mechanistic gap is that the direct RNA substrates of Rlig1 are not identified, and the observed transcriptomic changes in mRNA are likely downstream consequences of its loss. However, these points are clearly acknowledged in the discussion, making the study a well-balanced contribution. Given the existence of a mouse knockout model, further discussion comparing the zebrafish transcriptomic results and phenotypes to those observed in mouse would help place this work in the context of prior studies. Overall, the main conclusions are well supported, and the limitations do not undermine them. This study represents an important contribution that establishes a foundation for future mechanistic work linking Rlig1 substrates to the observed phenotypes. - State what audience might be interested in and influenced by the reported findings.

      Zebrafish basic science researchers, particuarly those studying how genes lead to altered neural circuits and behavior, are the most direct target audience. However, the work is of more broad interest to those in the fields of neurodevelopment, gene regulation, and RNA biology / processing. - Define your field of expertise with a few keywords to help the authors contextualize your point of view. Indicate if there are any parts of the paper that you do not have sufficient expertise to evaluate.

      I am comfortable evaluating zebrafish mutants, transcriptomics, and behavioral assay design. I have more limited experiment in neural circuit anaysis and interpretation of calcium imaging data, though this part of the manuscript was also clearly presented and understandable.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Editors for the positive assessment on our manuscript. We also thank the Reviewers for their positive remarks and constructive comments. Based on the Reviewers’ feedback, we have conducted additional experiments and provided supporting data to address Reviewers’ comments. Particularly, we provided quantitative measurement for rotational polarity of ependymal cells in Agbl5<sup>M1/M1</sup> mutants and assessed the microtubule polarization. We quantified the intensity of apical actin network in ependymal cells to strength the role of CCP5 in organizing actin network. Using scanning electron microscopy, we demonstrated the affected polarity of trachea multicilia in Agbl5<sup>M1/M1</sup>. We co-immunostained ependymal cilia with GT335 and acetylated tubulin to address the effects on their length in cilia in the mutant. We assessed the presence and length of primary cilia in ependymal cell progenitors to identify their potential contribution to the defective polarity in Agbl5<sup>M1/M1</sup> ependymal cells. We feel that these revisions have much strengthened this MS.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights.

      We thank the Reviewer’s positive comments. To address the molecular insights of the dysregulated planar cell polarity (PCP) in Agbl5<sup>M1/M1</sup> ependyma, we have conducted additional experiments to assess the microtubule polarization in ependymal cells (Figure 7O-P). We quantified the intensity of actin networks around BB patches to better understand how it is affected in the ependyma of the mutants and contributes to the dispersion of BBs (Figure 4M-N), (Please see Recommendations for the authors).

      We also assessed trachea multicilia in Agbl5<sup>M1/M1</sup> mutants using SEM and found that the polarity of trachea multicilia was affected as well (Figure S2).

      Reviewer #2 (Public review):

      Summary:

      This study analyzed the consequences of Agbl5 mutation on ependymal cell development and function. The authors first characterize their mutant mouse line reporting a reduced lifespand and severe hydrocephalus. Next, they report a defect in ependymal cell cilia number and motility. They provide evidence for impaired basal body organisation and cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicates Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype is incomplete:

      We thank the Reviewer’s constructive comments. We have performed additional quantitative analysis of the phenotypes in Agbl5<sup>M1/M1</sup> that we feel strengthen this study.

      Figure 3G - the sequence from the movie is not really informative. Providing beating frequencies as quantification of the data would be more informative.

      We have provided the beating frequency as well as the mean vector length of cilia beating directions (that reflects the coordination of cilia) in Figure 3H and 3I respectively in the revised manuscript.

      Figure 3 - the quantification of actin network would strengthen the message.

      We agree with the Reviewers. We have quantified the total intensity of actin around BBs and the actin intensity normalized to signals of the BB marker (CEP164). The data have been provided in Figure 4M and 4N respectively. The quantitative analysis showed that both the total intensity of apical actin network and the intensity of F-actin per BB are reduced in Agbl5<sup>M1/M1</sup> ependymal cells compared to that in wild-type mice, suggesting that CCP5 is involved in organizing actin network around BB. This analysis certainly improves the clarity of this message.

      Lines 219 -220 - the authors conclude «Taken together, in Agbl5M1/M1 ependymal cells, the expression of genes promoting multiciliogenesis were not impaired but certain proteins associated with differentiated ependymal cells are not properly expressed». However, they do not assess gene but protein expression (IF). In addition, their quantification shows differences in the number of FoxJ1 positive cells which indeed is an impaired expression.

      We will clarify this statement and emphasize the number of FoxJ1-positive cells.

      Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting for the authors to check whether the subapical network of microtubules is glutamylated or not during ependymal cell differentiation and how this network is affected in their mutants.

      We thank the Reviewer’s constructive comments. We conducted an immunostaining on whole-mount lateral walls of lateral ventricles for GT335 and Centrin1, the position of the latter being used to localize the subapical layer. While the GT335 signal in multicilia is increased in Agbl5<sup>M1/M1</sup> ependyma (Figure S8E), its signals underneath BBs are not much different between the mutant and wild-type (Please see Figure S8C, D, G, H).

      Showing the data mentioned in the discussion on Cep110 would be a nice addition to the paper.

      These data have been provided in Supplementary Figure S9.

      Line 354: "The latter serves as a component of tissue polarity that is required for asymmetric PCP protein localization in each cell (Boutin et al., 2014; Vladar et al., 2012)." The cited reference did not demonstrate that this microtubule network is required for asymmetric PCP localization.

      We thank the Reviewer for critical reading. The cited reference (Bountin et al., 2014) has been removed.

      Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      Weaknesses:

      The manuscript is well-written but lacks specific interpretations of the results presented. Further experiments are needed to be fully convincing.

      We thank the Reviewer’s comments. We have performed further analysis and conducted additional experiments to strengthen this study.

      (1) We have quantified the intensity of actin staining around BB patches and its intensity relative to the number of BBs to assess to which extent the actin networks in Agbl5<sup>M1/M1</sup> ependymal cells are affected (please refer to the above response to the comments of Reviewer 2#). The results were shown in Figure 4M-N.

      (2) We Co-stained tdTomato with an ependymal cell-specific markers to strengthen the expression of Agbl5 in ependymal cells (please see Figure 6C-E).

      (3) We have conducted co-immunostaining of GT335 and Ac-Tub and compared the length of their signals in ependymal multicilia between WT and Agbl5<sup>M1/M1</sup> mice (please see Figure 6O, P, R, S).

      (4) We quantified the area of ependymal cells in the wild-type and Agbl5<sup>M1/M1</sup> mice. Indeed, the area of ependymal cells is increased in the mutants. However, the primary cilia are present in the ependymal cell progenitors of Agbl5<sup>M1/M1</sup> mice and have similar length with that in the wild-type (Please see Figure 7M, N and our response to this point below).

      (5) We performed additional analysis to address the affected rotational polarity in the Agbl5<sup>M1/M1</sup> mutant mice (please see Figure 3I, Figure 7E).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors showed that the actin networks were severely affected, leading to impaired stability of basal bodies and that the intensity and length of acetylated tubulin signal in the multicilia were dramatically reduced in AGBL5M1/M1mutant mice (Figures 3 and 5). Data also suggested the dysregulation of planar cell polarity. Are expression and localization of other planar cell polarity proteins such as tyrosinated tubulin and Fzd6 affected in mutant mice?

      We thank the Reviewer’s recommendations. We have assessed the expression of tyrosinated tubulins and found they are similarly polarized in ependymal cells from wild-type and Agbl5<sup>M1/M1</sup> mice. The results are presented in Figure 7O, P in the revised MS. We also tried to assess the expression of Fzd6. However, with the antibody we tested, Fzd6 signals were not convincing. Therefore, we prefer to not showing the results and drawing a conclusion on it.

      (2) The phenotype of multiciliated cells in tracheas should also be examined in mutant mice. It is important to elucidate whether AGBL5 commonly functions in multiciliated cells of other organs.

      We thank the Reviewer’s suggestion. We have assessed the multicilia in the tracheas of P30 mice using scanning electron microscopy. Indeed, unlike the multicilia in wild-type mice that orientate to the same direction, those in the tracheas of Agbl5<sup>M1/M1</sup> mice often radiate to different directions in individual cells (Figure S2). Therefore, Agbl5 appears commonly involved in the alignment of multicilia.

      (3) According to Figure 1B, AGBL5 is highly expressed in the brain. Which cells in the brain express it besides ependymal cells?

      Based on the localization of tdTomato tracer engineered in Agbl5 mutant alleles (Figure 5B), Agbl5 is broadly expressed in the brain, including most if not all neurons, but its expression is much weaker in the subventricular zone (Please see Figure 5B). We clarified this in the revised MS.

      (4) From a mechanistic point of view, it is necessary to identify binding proteins with the N-domain of AGBL5 and perform functional analyses.

      We agree with the Reviewer. We feel that identification of the binding partners of CCP5 N-domain and functional analysis may be more suitable to go along with other mechanistic analysis on the function of CCP5 in ependymal cell polarities in our future study.

      Reviewer #2 (Recommendations for the authors):

      (1) Movie 3: The authors could comment on beating direction that seems impaired at the cell scale here, analysis of rotational polarity would be a plus.

      We thank the reviewer’s recommendation. We have analyzed the beating directions of cilia in individual cells and presented their consistency in each cell using mean vector length. These results indeed demonstrated defective rotational polarity in the cell level in Agbl5<sup>M1/M1</sup> mice (please refer to Figure 3I). We also analyzed the beating directions of ependymal multicilia in earlier stage in tissue level (Figure 7E). The mean vector length of cilia beating direction in Agbl5<sup>M1/M1</sup> mice is significantly reduced compared to that in wild-type, suggesting an aberrant rotational polarity in the tissue level in the mutant (Figure 7E).

      (2) Line 166 : ref to Werner et al., 2011 is not correct (no ependymal cells in that paper).

      We thank the reviewer’s critical reading. This reference has been removed.

      (3) Figure S4: B and D look similar picture to me same for C and F.

      We apologize for using the wrong images in this Figure. It has been corrected (Revised Figure S5).

      (4) Line 328: "Therefore, CCP5 apparently contributes to the establishment of both translational and tissue polarities in ependymal cells." Should be rephrased since translational polarity is also a tissue-level parameter which is the coordinated positioning of the ciliary patch. Cf Mirzadeh et al., 2010; Boutin et al., 2014.

      We thank the Reviewer’s comments. The sentence has been rephrased. This concept has been clarified where else needed in the revised manuscript. 

      (5) Line 348: "Planar cell polarity (PCP) pathway is essential for the establishment of rotational and tissue polarities in ependymal cells" Rotational polarity also has a tissular component (ie coordination of beating direction across tissue which is reflected by coordination of basal body polarities across tissue).

      We thank the Reviewer’s comments. We have clarified this point in the revised MS.

      (6) Incomplete bibliography citation (ie Walentek et al. without date).

      We thank the Reviewer’s critical reading. This bibliography citation has been fixed.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 3: The authors assert that the mutant's apical actin networks are significantly disrupted. However, the cell shown in Figure 3Q-R exhibits less compact centrioles than the controls, which could account for the reduction in phalloidin staining. Because centriole dispersion is variable in the mutant, quantifying actin staining in representative cells would be necessary to support such a statement.

      We thank the Reviewer’s comments. To address this concern, we have quantified the total intensity of actin network around BBs as well as the intensity of F-actin signals normalized to the level of immunosignals of BBs ((revised Figure 4M, N) please also refer to our response to Reviewer 1#). The results indicated the intensity of actin signal per BB is reduced in the mutant compared to that of wild-type mice. We feel that this analysis strengthened our statement.

      (2) Figures S3 and 4A-B show that the authors examine tdT expression to show that Agbl5 is expressed in ependymal cells but not in the SVZ. However, the tdT signal intensity is very low, and cells are very dense in this brain region. Double staining with specific markers of ependymal and/or SVZ cells would help convince readers that tdT is not expressed in SVZ cells.

      We agree with the Reviewer that the intensity of tdT signal is low, but broadly detectable in brain. Compared with its expression in ependymal cells, that in SVZ is much lower if any (Figure 4B’). To further confirm the identity of tdT-positive cells along the surface of ventricles, we have co-stained the brain sections of Agbl5<sup>WT/M1</sup> mice for tdT and S100b, a marker of mature ependymal cells (Figure 5C-E). The signal of tdt is colocalized with that of S100b and is much lower in cell layers next to S100b-positive cells.

      (3) Figure 4C-D and S4: The authors demonstrate that the number of FoxJ1+ cells per section increases at P7 (4C-E), while the number of S100β+ cells per mm decreases. Quantifications should be carried out in a similar manner to ensure comparability (number of positive cells per mm). Additionally, it remains unclear how to interpret these results, as S100β and FoxJ1 are two markers of differentiated cells, yet they exhibit opposite trends compared to controls. Is this a direct or indirect effect of Agbl5 mutation? The increase in the number of FoxJ1+ cells is particularly surprising given that the number of GT335 multicilia per mm remains unchanged (Figure 5).

      We agree with the Reviewer that quantifications should be carried out in a similar manner. In the revised MS, the quantification of Foxj1-positive cells is presented in number per mm (Figure 5I). To be noted, the expression of Foxj1 was assessed at P7 when ependymal cells are differentiating. while the expression of S100β was assessed at P17 when ependymal cells are supposed to be fully mature. Although S100b is used as a marker of mature ependymal cells, given its unclear function, we removed the results of S100b-positiving cell counting to avoid confusion in the revised manuscript.

      (4) Figure 5: In this figure, the authors analyze the labeling obtained with GT335, Acetylated Tubulin, and Arl13b antibodies. They show that the area of the cilium labeled by GT335 has increased, while the area labeled by the Acetylated Tubulin antibody has decreased in the knockout (KO) compared to the control. However, the length of the cilia observed through labeling with the Arl13b antibody remains unchanged. These observations are intriguing, but the low-magnification images in Figure 4 do not allow for the differences in ciliary axoneme labeling to be seen. Double GT335/AcTub labeling and higher magnifications are necessary for improved visualization of the differences in labeling along the axonemes.

      We thank the Reviewer comments. We have co-stained the cilia with GT335 and Ac-Tub antibodies, re-quantified cilia length labeled with respective antibodies and provided high magnification images. Please see the revised Figure 6O,P,R,S.

      (5) Figure 6: An analysis of ciliary beats using a high-speed camera shows no difference in ciliary beat frequency between the control and KO groups. At least, 3 animals should be analyzed. According to Figure 5, these findings indicate that the decrease in ciliary acetylation and the increase in ciliary glutamylation do not affect the beat frequency; instead, they disrupt the orientation of the beats. While these results are intriguing, they require further confirmation. Analyzing ciliary beats with a high-speed camera is informative, but at least three animals per genotype should be examined to ensure rigor. Furthermore, if the coordination of ciliary beats is impaired within the cells, this should be validated by double-labeling centrioles and basal feet to demonstrate that the orientation of cilia within the cells is abnormal.

      We thank the Reviewer’s comments. Sections shown in Figure 5 (currently Figure 6) are from P7 mice, while the ciliary beating analysis shown in Figure 6 (currently Figure 7) is from P15 mice. As the PTM changes in cilia were also observed in Agbl5<sup>M2/M2</sup>, we don’t think this is the cause that disrupts the orientation of the beats. The rotational polarity of Agbl5<sup>M1/M1</sup> ependymal cells is affected. Please refer to the analysis in Figure 3I and Figure 7E in the revised manuscript.

      (6) Figure 6F-G: β-Catenin labeling reveals cells of varying sizes in the KO. This phenotype is typical of ciliary mutants that lack primary cilia (Mirzadeh et al., 2010). Hence, it is essential to examine the mutation's impact on the presence, length, and positioning of the primary cilium in ependymal cell progenitors.

      We thank the Reviewer’s constructive comments. We assessed the area of ependymal cells labeled with β-Catenin. Indeed, the ependymal cells in the mutant showed larger area than that of wild-type. The ratio of the area of BB patch over that of cell surface is reduced (please see Figure 7O, P in the revised manuscript). However, primary cilia are present in ependymal cell progenitors in the mutant and exhibit comparable length with those in the wild-type (Figure S8). Due to some technique problems, we were unable to get convincing results from whole-mount ventricle walls for the primary cilium positioning at this time. We speculate that the localization of certain sensory proteins in primary cilia or the positioning of primary cilia might be affected in Agbl5<sup>M1/M1</sup> mice. We discussed this possibility and will certainly systemically assess this intriguing aspect in our future investigation.

      (7) Given the regular beating frequency in the KO at P15, how do the authors explain the complete absence of ciliary beating in the adult? How many animals were analyzed? One would expect ciliary beating to remain unaffected as it was at P15 unless the cilia structure was specifically altered at the adult stage. Is that the case?

      We thank the Reviewer’s critical questions. We do think that the ciliary structure of Agbl5<sup>M1/M1</sup> ependymal cells is likely altered during aging. Given that only Agbl5<sup>M1/M1</sup> but not Agbl5<sup>M2/M2</sup> mice develop hydrocephalus, we speculate the N-domain of CCP5 may contribute to the integrity of ependymal multicilia. We have added this in the Discussion section. For each genotype, 2 mice were analyzed.

      (8) Line 264 of the manuscript: replace intercellular with intracellular.

      It has been revised.

      (9) Indicate the number of animals analyzed in each experiment

      It has been included in figure legends.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained by local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales and is more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor out topographic differences, and I found the relationship between heritability and neural time scales very interesting. The writing is clear, and the results are compelling.

      We thank Reviewer 1 for their kind words and enthusiastic support of our manuscript.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified.

      (1) On page 16, the authors compare heritability in functional connectivity (FC) and response time series, and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and timeseries differences cannot, by definition). This makes me wonder how this connectivity result would change if the authors used inter-subject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript, but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      We agree that investigating the heritability of ISFC (or stimulus-driven functional connectivity) would make for a very interesting future direction. Ultimately, we chose to analyze FC (vs. ISFC) profiles to allow for direct comparison with the sizable existing literature on the heritability of FC (such as in our Movie vs. Rest FC analysis) and decided to refrain from analyzing ISFC data in order to keep the present manuscript focused. ISFC analysis of this dataset will be a focus of future work.

      (2) The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. The authors could even color the data points according to networks, like in Figure 3C. (They also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      We thank R1 for this helpful suggestion- we originally set the y-axis limits to r = 1 in order to facilitate comparison between ISC (Fig. 1B) and FC profile (Fig. 6B) similarity, but we agree that this renders the group differences harder to discern and have updated the plot accordingly (along with thicker lines to enhance readability). We prefer to keep the line plots in the main body as they allow for direct comparison of all three groups on the same plot, but we have included the scatter plot version in Fig. S2 for those who are interested.

      (3) On page 9, if I understand correctly, the authors regress the vector of ISC values across parcels out of the vector of heritability values across parcels, and then plot the residual heritability values. Do they center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can the authors explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could they go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?).

      We indeed included an intercept in this model using MATLAB’s fitlm function. This means that the model estimates the best-fitting line of the following form: heritability<sub>i</sub>=β0+β1ISC<sub>i</sub> +ε<sub>i</sub>. We agree that the interpretation of these ε<sub>i</sub> values and alternative approaches to controlling for ISC should be clarified. As such, we have added the following passages to the text:

      Methods: “Because the heritability of ISC is constrained by the degree of synchronization in a given area, we also sought to identify areas in which BOLD time courses were more/less heritable than would be expected based on ISC alone by fitting a linear model of the form heritability<sub>i</sub>=β0+β1ISC<sub>i</sub>+ε<sub>i</sub> and plotting the residuals. Regarding alternative approaches to controlling for ISC, although the heritability model introduced by Ge et al. allows for the inclusion of covariates defined at the subject level (e.g., age), it does not allow for covariates that are defined at the dyad level (e.g., pairwise ISC).”

      Results: “Here, negative values in the residual map indicate parcels where heritability is lower than expected based on ISC, while positive values indicate higher-than expected heritability.”

      (4) On page 4 (line 155), the authors say "we shuffled dyad labels"- is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure their approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned in line 189?

      Briefly, shuffling the kinship matrix involved permuting the rows and columns of the matrix in the same manner (also known as the quadratic assignment procedure), whereas shuffling the dyad labels involved random permutations of the three group labels (MZ, DZ, unrelated), which could not be done through matrix operations as the age- and gender matching precluded the use of a complete similarity matrix. However, given concerns raised by Reviewer 2, we have removed our significance claims from this (and similar) sections, which we discuss in more detail in response to Reviewer 2’s weakness A.

      (5) I found panel A in Figure 4 to be a little bit misleading because their parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels they are performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      We agree that our efforts to present a simplified depiction of hyperalignment may mislead less familiar readers and have amended Fig. 4A according to this suggestion. We have also added text to the methods section (below) to clarify that the outputs of hyperalignment are time series that reflect linear combinations of other voxels’ time series from that parcel.

      “This approach independently transforms each subject's data within discrete anatomical parcels into the common space, yielding functionally aligned vertex time series that are calculated as weighted linear combinations of the original time series from all other vertices within that same parcel for that subject.”

      (6) I believe the subjects watched all different movies across the two days, however, for a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

      We agree that this would be helpful and have added the following text to the relevant sections:

      “All clips were only viewed once by each subject, with the exception of the brief montage which was included at the end of each of the four runs for test-retest purposes.”

      “To characterize the heritability of brain responses to complex stimuli, we used 7T fMRI data from 178 HCP Young Adult subjects acquired across two days (using two largely non-overlapping sets of movie stimuli, see Methods)…”

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to estimate the heritability of brain activity evoked from a naturalistic fMRI paradigm. No new data were collected; the authors analyzed the publicly available and well-known data from the Human Connectome Project. The paper has 3 main pieces, as described in the Abstract:

      (1) Heritability of movie-evoked brain activity and connectivity patterns across the cortex.

      (2) Decomposition of this heritability into genetic similarity in "where" vs. "how" sensory information is processed.

      (3) Heritability of brain activity patterns, as partially explained by the heritability of neural timescales.

      Strengths:

      The authors investigate a very relevant topic that concerns how heritable patterns of brain activity among individuals subjected to the same kind of naturalistic stimulation are. Notably, the authors complement their analysis of movie-watching data with resting-state data.

      Weaknesses:

      The paper has numerous problems, most of which stem from the statistical analyses. I also note the lack of mapping between the subsections within the Methods section and the subsections within the Results section. We can only assess results after understanding and confirming the methods are valid; here, however, Methods and Results, as written, are not aligned, so we can't always be sure which results are coming from which analysis.

      (A) Intersubject correlation (ISC) (section that starts from line 143): "We used nonparametric permutation testing to quantify average differences in ISC for each parcel in the Schaefer 400 atlas for each day of data collection across three groups: MZ dyads, DZ dyads, and unrelated (UR) dyads, where all UR dyads were matched for gender and age in years." ... "some participants contributed to ISC values for multiple dyads (thus violating independence assumptions)"

      This is an indirect attempt to demonstrate heritability. And it's also incorrect since, as the authors themselves point out, some subjects contribute to more than one dyad.

      Permutation tests don't quantify "average differences", they provide a measure of evidence about whether differences observed are sufficient to reject a hypothesis of no difference.

      Matching subjects is also incorrect as it artificially alters the sample; covarying for age and sex, as done in standard analyses of heritability, would have been appropriate.

      It isn't clear why the authors went through the trouble of implementing their own nonparametric test if HCP recommends using PALM, which already contains the validated and documented methods for permutation tests developed precisely for HCP data.

      The results from this analysis, in their current form, are likely incorrect.

      We appreciate that permutation tests do not quantify average differences and intended to write “We used non-parametric permutation testing to quantify [the significance of] average differences…”. Our intention with this analysis was not to demonstrate heritability, but rather to quantify group differences in ISC in a manner that is interpretable for readers who are unfamiliar with h<sup>2</sup> (e.g., “identical twins’ BOLD time courses were 59% more similar than those from pairs of unrelated individuals”) and motivate the formal heritability analysis used later in the paper. Indeed, all of the heritability analyses in this paper leveraged a validated multidimensional heritability method first introduced by Ge et al. (2016) and used by many other investigators since then. Furthermore, we covaried for age and sex at the subject level in all our heritability analyses, and always tested the significance of these heritability values using a validated permutation procedure (the quadratic assignment procedure; Hubert & Schultz, 1976) that respects the non-independence of dyadic data.

      Regarding the shuffling procedure used for Figure 1, while PALM is the standard for univariate, subject-level GLMs in the HCP pipeline and can accommodate nested designs (i.e., subjects within families), it is not designed to handle the unique relational dependencies of dyadic ISC analysis (i.e., the same subject contributing to multiple dyads). Although the element-wise resampling approach was the most appropriate approach available, it is known to inflate the false positive rate (Chen et al., 2016; doi:10.1016/j.neuroimage.2016.05.023); given that this analysis was simply meant to motivate our later hypothesis testing heritability analyses, we have removed significance claims from this section of the manuscript. Still, we emphasize that this has no bearing on the validity of our conclusions which were supported by our formal heritability analyses; throughout our paper we have correctly used the appropriate methods to back the stated claims.

      (B) Functional connectivity (FC) (section that starts from line 159): Here the authors compute two 400x400 FC matrix for each subject, one for rest, one for movie-watching, then correlate the correlations within each dyad, then compared the average correlation of correlations for MZ, DZ, and UR. In addition to the same problems as the previous analysis, here it is not clear what is meant by "averaging correlations [...] within a network combination". What is a "network combination"? Further, to average correlations, they need to be r-to-z transformed first. As with the above, the results from this analysis in its current form are likely incorrect.

      We regret that R2 had difficulty understanding our analysis and have added the following text to the relevant Methods section to clarify our approach:

      “For example, there are 16 parcels in the Kong et al. Auditory network and 17 parcels in the Language network, so the FC profile for a given subject’s Auditory-Language network combination consists of the (16 * 17 =) 272 correlation coefficients between all unique pairs of one parcel from each network.”

      As we stated in the previous Methods paragraph, “All Pearson r values in this and all other analyses were Fisher z-transformed before averaging (and converted back to Pearson r for visualization)”. Thus, contrary to the reviewer’s assertion, these analyses were performed correctly. Once again, we emphasize that this analysis was not intended to demonstrate heritability, but rather to describe group differences in FC in familiar units.

      (C) ISC and FC profile heritability analyses (section that starts from line 175): Here, the authors use first a valid method remarkably similar to the old Haseman-Elston approach to compute heritability, complemented by a permutation test. That is fine. But then they proceed with two novel, ill-described, and likely invalid methods to (1) "compare the heritability of movie and rest FC profiles" and (2) to "determine the sample size necessary for stable multidimensional heritability results". For (1), they permute, seemingly under the alternative, rest and movie-watching timeseries, and (2), by dropping subjects and estimating changes in the distribution.

      The (1) might be correct, but there are items that are not clearly described, so the reader cannot be sure of what was done. What are the "153 unique network combinations"? Why do the authors separate by day here, whereas the previous analyses concatenated both days? Were the correlations r-to-z transformed before averaging?

      The (2) is also not well described, and in any case, power can be computed analytically; it isn't clear why the authors needed to resort to this ad hoc approach, the validity of which is unknown. If the issue is the possibility that the multidimensional phenotypic correlation matrix is rank-deficient, it suffices that there are more independent measurements per subject than the number of subjects.

      Regarding (1), we have clarified in section 2.6 that the 153 unique network combinations reflect each unique pair of 17 Kong networks. All of our analyses, including this one, were performed separately for each day of data collection, as we state throughout the paper and visualize in our figures (although we acknowledge that, on some occasions, we [conservatively] performed FDR-correction on a combined set of p-values, as discussed in our response to K). Given that the null hypothesis for this analysis is that rest FC and movie FC are equally heritable, we are not sure why permuting rest and movie FC matrices would be invalid. All Pearson r values were z-transformed before averaging, as we stated in our paper.

      Regarding (2), we included this analysis in response to editorial concerns that our heritability analyses were not sufficiently powered, and we chose this approach because it serves as a simple way to demonstrate the stability of our results at various sample sizes whose validity is self-evident. Furthermore, this sort of subsampling approach has been used many times before in our field (e.g., Marek et al., 2022) and others (e.g., Manyara et al., 2024) to demonstrate the sample-size dependence and stability of statistical effects. We have added text explaining this to the relevant Methods section (2.6).

      (D) Frequency-dependent ISC heritability analysis (from line 216): Here, the authors decompose the timeseries into frequency bands, then repeat earlier analyses, thus bringing here the same earlier problems and questions of non-exchangability in the permutations given the dyads pattern, r-z transforms, and sex/age covariates.

      We did not use dyadic permutation testing for any of the frequency-dependent ISC analyses; rather, we used the jackknife SEMs to compare heritability across frequency bands and have added an explicit description of this to section 2.7. We have addressed the r-z transform and covariate concerns in previous comments.

      (E) FC strength heritability analysis (from line 236): Here, the authors use the univariate FC to compute heritability using valid and well-established methods as implemented in SOLAR. There is no "linkage" being done here (thus, the statement in line 238 is incorrect in this application. SOLAR already produces SEs, so it's unclear why the authors went out of their way to obtain jackknife estimates. If the issue is non-normality, I note that the assumption of normality is present already at the stage in which parameters themselves are estimated, not just the standard errors; for non-normal data, a rank-based inversenormal transformation could have been used. Moreover, typically, r-to-z transformed values tend to be fairly normally distributed. So, while the heritabilities might be correct, the standard errors may not be (the authors don't demonstrate that their jackknife SE estimator is valid). The comparison of h2 between dyads raises the same questions about permutations, age/sex covariates, and r-z transforms as above.

      We used jackknife SEs for these analyses to maintain consistency with the multidimensional heritability package used here, which only outputs jackknife SEs. We note that this jackknife approach (and the corresponding multidimensional heritability analysis) was detailed in prior work (Anderson et al., 2021), and that the leave-one-family-out jackknife has a long history of being used to estimate SEs in heritability studies, especially when working with smaller samples (Knapp et al., 1989). We are also not sure what “the comparison of h2 between dyads” means- heritability cannot be compared “between” dyads; rather, it is defined across dyads.

      (F) Hyperalignment (from line 245): It isn't clear at this point in the manuscript in what way hyperalignment would help to decompose heritability in "where vs. how" (from the Abstract). That information and references are only described much later, from around line 459. The description itself provides no references, and one cannot even try to reproduce what is described here in the Methods section. Regardless, it isn't entirely clear why this analysis was done: by matching functional areas, all heritabilities are going to be reduced because there will be less variance between subjects. Perhaps studying the parameters that drive the alignment (akin to what is done in tensor-based and deformation-based morphometry) could have been more informative. Plus, the alignment process itself may introduce errors, which could also reduce heritability. This could be an alternative explanation for the reduced heritability after hyperalignment and should be discussed. An investigation of hyperaligment parameters, their heritability, and their co-heritability with the BOLD-phenotypes can inform on this.

      To help set up our hyperalignment analyses, we have added text to the introduction explaining how hyperalignment would help to decompose heritability. The description in the Methods section included a reference to Bazeille et al., 2021, in which the hyperalignment method used here is discussed in detail. Still, we have added citations to additional papers (also cited in the Bazeille et al. paper, and elsewhere in our paper) in case that might be helpful. We note that it is not the case that all heritabilities were reduced by hyperalignment- as can be seen in Figs. 4D, 8A, and S15, hyperalignment did increase heritability in some voxels and network combinations. This would be expected under the alternative (albeit unlikely) hypothesis that functional topographies are not heritable, such that topographic variation between related individuals would obscure similarities in their (heritable) topography-independent brain responses. Recognizing that this alternative is unlikely, we believe the main novelty of this analysis comes from the magnitude of the hyperalignment effect (up to 40% of brain-wide heritability) and its spatial pattern (e.g., larger heritability decreases in visual vs. auditory cortex, the opposite of our NT result).

      We agree that we would see lower post-hyperalignment heritability if the alignment process itself introduced errors/noise, but this would be deeply surprising as hyperalignment increases ISC by design (and errors/noise could only decrease ISC). To demonstrate this, we have added Figure S7 which shows that (as expected) ISC across all voxels and subject pairs increases after hyperalignment (and that this increase is larger when hyperalignment is performed in larger parcels). Given that hyperalignment increased ISC, and that it is blind to twin status, we are unsure how it could have introduced errors that would have confounded this result.

      (G) Relationships between parcel area and heritability (from line 270): As under F), how much the results are distorted likely depends on the accuracy of the alignment, and the error variance (vs heritable variance) introduced by this.

      We agree that alignment accuracy could potentially impact parcel-level differences in how much heritability changes following hyperalignment, and we included the frequency dependent h<sup>2</sup><sub>residuals</sub> (controlling for differences in ISC) in Fig. 3 for this reason, as more accurate hyperalignment should result in greater increases in ISC, raising the heritability ceiling. We note that we observe similar relationships between parcel rank and frequency dependent changes in these residualized maps, suggesting that our parcel-level differences are not simply the result of better alignment in more sensory parcels.

      (H) Neural timescale analyses (from line 280): Here, a valid phenotype (NT) is assessed with statistical methods with the same limitations as those previously (exchangability of dyads, age/sex covariates, and r-z transforms). NT values are combined across space and used as covariates in "some multivariate analyses". As a reader, I really wanted to see the results related to NT, something as simple as its heritability, but these aren't clearly shown, only differences between types of dyads.

      We have addressed the exchangeability, covariates, and r-z transform comments above (in A). As we explained for our FC strength analyses, we are underpowered to evaluate the heritability of unidimensional traits (like the heritability of NT magnitude), and the heritability of a closely-related measure (BOLD turnover magnitude) has already been established in a larger sample of HCP subjects (https://doi.org/10.1152/jn.00402.2022). Still, we agree that more results related to the heritability of NTs would be of interest to our readers. As such, we have added an analysis in section 3.4 quantifying the heritability of multivariate NT topographies and used SOLAR to quantify the heritability of NT magnitudes, with the disclaimer that this and similar analyses are underpowered (hence the large difference in day 1 and day 2 heritability effect sizes). We also removed significance claims for the dyadic NT similarity analysis.

      (I) Significance testing for autocorrelated brain maps and FC matrices (from line 310): Here, the authors suddenly bring up something entirely different: reliability of heritability maps, and then never return to the topic of reliability again. As a reader, I find this confusing. In any case, analyses with BrainSMASH with well-behaved, normally distributed data are ok. Whether their data is well behaved or whether they ensured that the data would be well behaved so that BrainSMASH is valid is not described. As to why Spearman correlations are needed here, Mantel tests, or whether the 1000 "surrogate" maps are valid realizations of the data under the null, remains undemonstrated.

      We brought up reliability in this section because we show the reliability of our results across the two days of data collection several times in the paper. R2 is correct to point out that BrainSMASH was validated using normally distributed brain maps, and although some of our brain maps contain normally distributed values, others are right skewed (due largely to the fact that many voxels/parcels exhibit low ISC while visual/auditory areas have very high ISC). In preparing our original manuscript, we visualized BrainSMASH’s variogram outputs for one of the most skewed inputs (vertex-wise BOLD time course heritability) and found that the autocorrelation structures of the empirical and null maps were well-matched. We did not include this in the original manuscript as it is not commonplace in the field to report the variograms, see Author response image 1. Furthermore, our use of Spearman (vs. Pearson) correlations renders these distributional differences less relevant, as the Spearman correlation transforms all inputs to a uniform distribution. To empirically check that these distributional differences do not bias our results, we retested the significance of all brain map associations using the spin test (10.1016/j.neuroimage.2018.05.070), an alternative method that does not assume normally distributed inputs, and obtained identical p-values for all analyses (P<.001 in all cases).

      Author response image 1.

      (J) Global signal was removed, and the authors do not acknowledge that this could be a limitation in their analyses, nor offer a side analysis in which the global signal is preserved.

      Although we agree that GSR is a contentious preprocessing step for certain analyses, it has explicitly been shown to increase ISC signal-to-noise without compromising FC fingerprints (Graff et al., 10.1016/j.dcn.2022.101087), and it is uncommon to perform ISC analyses with and without GSR. Still, we have added additional text to our Methods section explaining our rationale for using GSR and that this could affect our results. We also re-ran our main analysis (BOLD time course heritability) with and without GSR and found that GSR had little impact on our results; we have included this in our manuscript as Fig. S4.

      Specifically, we see that GSR resulted in a slight increase in heritability (average Day 1 h<sup>2</sup> with/without GSR = .064/.060; Day 2: .068/.061) and almost no effect on the spatial pattern of our results (With GSR/without GSR Spearman ρ = .99, P<sub>brainSMASH</sub> < .001 on both Day 1 and Day 2).

      (K) FDR is used to control the error rate, but in many cases, as it's applied to multiple sets of p-values, the amount of false discoveries is only controlled across all tests, but not within each set. The number of errors within any set remains unknown.

      We agree that the FDR usage in our original manuscript was inconsistent, in that for two analyses we FDR-corrected p-values from the two days of data collection together (instead of correcting p-values from each day separately and reporting voxels/parcels/etc. that were significant at q<.05 on both days, as in the rest of our analyses). We note that both approaches are more conservative than reporting significant results at q<.05 separately; regardless, to maintain consistency we have updated all analyses such that FDR correction is always performed separately for each day of data collection.

      (L) Generally, when studying the heritability of a trait, the trait must be defined first. Here, multiple traits are investigated, but are never rigorously defined. Worse, the trait being analyzed changes at every turn.

      Here, we analyze the heritability of movie-evoked BOLD time courses (Figures 1-5) as well as FC profiles (Figures 6-8). We defined FC profiles in our Introduction as an individual’s pattern of pairwise FC strengths (and further detailed how we quantified FC profiles in the relevant Methods section), and believe that “BOLD time course” is a well understood phrase in the field and does not need to be further defined. We also used hyperalignment to decompose the heritability of these traits into topography-dependent and independent portions, and (new to this version) also explicitly quantify the heritability of neural timescales, which we defined as the AUC of the ACF until the first negative ACF value in both the relevant Results and Methods sections.

      To make this clearer, we have modified the last paragraph of our Introduction to begin with:

      In the present work, we address these questions by analyzing 7T fMRI recordings of a twin sample acquired by the Human Connectome Project (Van Essen et al., 2013) to quantify the heritability of two distinct high-dimensional traits—stimulus-evoked BOLD time courses and functional connectivity profiles—across the cortex.

      Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      We believe that this question is getting at the difference between pairwise ISC (i.e., correlating one BOLD time course from one subject with that from another subject) and leave-one-subject-out ISC (i.e., correlating one BOLD time course from one subject with the corresponding average time course across all other subjects). We chose to use the pairwise ISC method because it allows us to capitalize on the information contained in the n<sup>2</sup> pairwise ISC matrix (whereas the other approach averages out meaningful information to yield a n<sup>1</sup> ISC matrix) and leverage a more sophisticated multidimensional heritability approach. Also, the leave-one-subject-out approach introduces additional issues re: handling family-level data (e.g., should we include a subject’s twin in the leave-one-subject-out average? If so, how should we handle subjects who don’t have a twin in the dataset, as averaging data from different numbers of subjects will lead to different ISC magnitudes? etc.).

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      We report p-values for heritability throughout our paper (e.g., stating that BOLD time courses are significantly heritable in 99% of parcels in Figure 2), and we believe that the reliability of our spatial maps across days of data collection (also quantified with p-values) further demonstrates the trustworthiness of our results. Finally, as we demonstrate in Figure S5, our sample size is more than sufficient to reliably detect small effects.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      In addition to driving shared neuronal responses (which are captured in BOLD signal oscillations <.1 Hz or so), movies also elicit shared cardiac, respiratory, and motion responses across participants at higher frequencies. Although we used a relatively conservative denoising approach here, we believe some of these non-neuronal signals are still present in our data; alternatively, it is also possible that these signals reflect “fast” BOLD responses at >.15 Hz (as discussed in 10.1016/j.neuroimage.2021.118658). In any case, the fact that information in this frequency band is considerably less heritable than information in slower frequency bands supports the idea that this band is noisier and suggests that our heritability results are driven by canonical neuronal activity-related BOLD signals.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      Although the decrease of 0.025 is small, we note that this constitutes around ~50% of BOLD time course heritability in some voxels (seen in comparison to Fig. 4C), and the spatial pattern of this result is quite consistent across days of data collection, indicating its reliability. Furthermore, the whole-brain distributions of results shown in Fig. 5B are clearly skewed towards negative values, indicating that controlling for NT partially reduces (or “explains”) BOLD time course heritability. Still, we agree that showing raw h<sup>2</sup> values in addition to the difference maps would be helpful for some readers and have added a corresponding supplementary figure (S12) which shows these.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      We did consider adding standard errors for these heritability estimates, but found that visualizing standard errors for each of the 153 unique network combinations in our heatmaps rendered the visualizations difficult to parse, and given that our hypotheses concerned global (e.g., hyperaligned vs. MSM-aligned) or network-level (e.g., sensory vs. associative) patterns, we focused on calculating standard errors/p-values for these analyses (although we note that dyad-level standard errors can be found in Fig. 6B, where they are clearly marginal compared to the group effects).

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      We agree that this result was relatively under-explored in our Discussion section and have added additional text (lines 851-855) to connect this result to recent work on arousal-dependent uniqueness of FC.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Do the authors have any ideas why we see this hotspot of heritability in pMTG/LOTC? It really jumps out in Figure 1A and Figure 2. The more posterior sensory MT+ area seems to drop when regressing out ISC in Figure 2B, but this pMTG area stays hot. Is there anything special about this kind of multimodal biological motion/action observation / social perception area (Pitcher & Ungerleider, 2021)? I don't think this is necessary to discuss in the manuscript, but I'm curious if the authors have any speculation.

      We are not certain as to why BOLD time courses in this parcel are particularly heritable- although this area is associated with biological motion, that particular function tends to be more right lateralized, and here we see nominally higher heritability in the left hemisphere. Per a Neurosynth review (and consistent with the left lateralization), we believe this may have more to do with speech processing, but a more definitive answer will require further investigation.

      (2) Page 3, line 127: "More information on these clips"-it might be worth saying a little bit more here just to make sure people understand that these are audiovisual clips, they include language, they're long enough to convey meaningful social and narrative information, etc.

      We agree and have added additional details on the clip composition to the relevant methods paragraph.

      (3) Figure 1 caption: can you add a sentence reminding readers what's going on with Day 1 and Day 2?

      We thank R1 for this suggestion and have added a sentence to this effect at this location.

      (4) Page 9, line 379: "although these more associative parcels do not encode a substantial amount of stimulus-specific information"-is this really true? I suspect these association areas still have decent ISCs, even if there are many processing stages downstream of the raw stimulus.

      Although these parcels are not the most synchronized by the stimulus, we agree that it is unfair (and vague) to say that they do not encode a substantial amount of stimulus-specific information. We have edited this sentence to make a more specific claim and highlight the relatively lower ISC in these parcels vs. more unimodal sensory areas.

      (5) Page 9, line 417: Can you unpack a bit more what you mean by "supra-BOLD frequency band"?

      Here, we refer to the fact that BOLD signals resulting from neuronal firing events have frequencies below ~.15 Hz (Josephs and Henson, 1999). We have added additional text and the Josephs and Henson citation to this line to further unpack this point.

      (6) Page 18, line 695: This discussion of how attention and gaze might partly shape response time series reminded me of recent work by Borovska & de Haas (2024)-might be worth citing.

      We are grateful to R1 for alerting us to this very relevant work and have included a reference to it in our discussion.

      (7) Page 19, line 755: I'm not sure I'd describe the hyperalignment results here as a "deleterious effects [on] heritability"-my reading was that hyperalignment allows you to say something more specific about heritability of function by allowing you to effectively factor out heritability effects that reduce to individual differences cortical topography; this seems like a good thing!

      We agree that “deleterious” was a poor word choice given its negative connotation, and have edited this sentence to read:

      “With this in mind, future studies investigating genetic correlations between brain function and behavioral variables may benefit from hyperalignment, as it can factor out individual-specific cortical topography and thus yield more precise estimates of functional heritability.”

      (8) I would love to see a ventral view in some of these plots! Not asking you to recreate the figures, but the ventral temporal cortex is an area of interest for many folks in the movie fMRI space (e.g., Haxby et al., 2011).

      We agree that ventral views would be of interest to some readers and have added the corresponding maps for our main results in supplementary figures S3 and S9.

      References:

      Borovska, P., & de Haas, B. (2024). Individual gaze shapes diverging neural representations. Proceedings of the National Academy of Sciences, 121(36), e2405602121. https://doi.org/10.1073/pnas.2405602121

      Haxby, J. V., Guntupalli, J. S., Connolly, A. C., Halchenko, Y. O., Conroy, B. R., Gobbini, M. I., Hanke, M., & Ramadge, P. J. (2011). A common, high-dimensional model of the representational space in human ventral temporal cortex. Neuron, 72(2), 404416. https://doi.org/10.1016/j.neuron.2011.08.026

      Pitcher, D., & Ungerleider, L. G. (2021). Evidence for a third visual pathway specialized for social perception. Trends in Cognitive Sciences, 25(2), 100-110. https://doi.org/10.1016/j.tics.2020.11.006

      Reviewer #2 (Recommendations for the authors):

      (1) To address the common core analytical problems listed under A), B), C), D), E), and basically throughout the methods:

      (a) Conduct permutations with exchangability restrictions to account for the pattern of dyad-relationships as e.g. implemented in PALM.

      (b) Control for age and sex covariates as covariates (e.g. as in SOLAR), rather than by matching.

      (c) Perform r-to-z transforms when conducting further analyses on correlations that assume normality.

      (d) For all analyses that assume normal distributions, e.g. in SOLAR and BrainSMASH, check that this is the case.

      We have explained how PALM is not suited for the study of effects that are defined at the dyad level (A), that we controlled for age and sex covariates in all our formal heritability analyses in our original submission (B), that we always performed r-to-z transforms when indicated in our original submission (C), and that our spatial permutation results don’t hinge on distributional differences (D).

      (2) Replace SEs derived from kacknife approach with those from SOLAR, or provide a comparison and motivation and/or demonstrate that SEs are correct.

      A more thorough explanation of the block jackknife procedure can be found in prior work introducing the multidimensional heritability method used here (Anderson et al., 2021).

      (3) Given problem (F & G):

      (a) Consider studying the parameters that drive the hyperalignment. They can be included as covariates in heritability analyses, and/or their heritability is of interest to understand the reasons for the heritability reduction post-hyperaligment.

      We agree that this would be interesting but the specific parameters that drive hyperalignment are beyond the scope of this study.

      (b) Include the alternative explanation of hyperalignment-induced noise in the discussion.

      We have added a figure showing that hyperalignment does not increase noise in ISC and explained here why “hyperalignment-induced noise” does not constitute a reasonable alternative explanation for our results.

      (4) Add heritability results for NT phenotypes.

      We have added heritability analyses for NT topography and (global) NT magnitude, as detailed above.

      (5) Motivate global signal removal, and acknowledge this process typically alters results substantially.

      We have added an explanation of our rationale for using GSR and shown in this response that it does not in fact substantially alter the results.

      (6) Rephrase and/or clarify the following:

      (a) "permutations quantify average differences" (under A).

      (b) "network combinations" and related analyses (under B & C).

      (c) why some analyses are separated per visit/day and others not (C).

      (d) methods and reasons for sample size estimation (C).

      We have rephrased or clarified all of the above.

      Reviewer #3 (Recommendations for the authors):

      (1) Participants should be recleared. I know HCP 7T data has 184 subjects. How can the authors have 176 twins and 690 unrelated subjects?

      As we reported in our Methods section, 178 subjects had complete movie-watching datasets, and 176 subjects had complete movie-watching and resting-state datasets. Of the 178 subjects with complete movie-watching data, we identified 690 age- and sex-matched dyads.

      (2) Figure 1. I don't find Figure S1A in Figure S1.

      We thank R3 for catching this error- we have amended this reference to read Fig. S1.

      (3) I could also suggest putting Figure 1 and Figure 2 together.

      We thank R3 for this suggestion- ultimately, we prefer to keep these figures separate to reinforce the difference between our dyadic similarity and formal heritability analyses.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      This study from the Niedergang lab establishes SNAT7 as a host-dependency factor in human macrophages that supports HIV-1 replication. They show a modest increase in SNAT7 levels HIV-1 infected macrophages and suggest that SNAT7 levels are transiently increased. Employing siRNA against SNAT7 they show reduction in HIV-1 protein levels and viral RNAs and claim that there is a block of reverse transcription in SNAT7 KD cells. Focusing on a known HIV-1 restriction factor in macrophages, SAMHD1, they interconnect the SNAT7 depletion with a reduction in phosphorylated, i.e. catalytical inactive SAMHD1 arguing that SNAT7 regulates the phosphorylation and thereby antiviral activity of SAMHD1. Since SNAT7 is a glutamine transporter that provides this AA from lysosomes, they lastly supplement glutamine and this somehow rescues the reduction of HIV-1 production in SNAT7 KD cells.

      Major comments:

      The strength of this manuscript is the clear focus on primary human macrophages that are HIV-1 infected and the interconnection of HIV-1 replication to the SNAT7 siRNA KD experiments in combination with SAMHD1 depletion and lastly glutamine supplementation. This establishes a stringent and coherent story line. The effects reported are modest; high variability is not a problem since using primary hMDM this is expected and can be addressed by testing several donors and applying stringent statistics.

      1. Having said so, I realize that while they give information on the statistical test used, i.e. one-way ANOVA they miss to explain the post-test used to assess significance (i.e. Bonferroni, Fishers LSD, whatsoever). Please add this information.

      We thank the reviewer for this comment. The figure legends have been updated to include more details of all the statistical tests used.

      1. Another issue that might underestimate the effects of HIV-1 infection on SNAT7 levels and vice versa of SNAT7 KD on HIV-1 replication is the non-single cell approach employed, i.e. WBlots. I assume that HIV-1 infection rates in macrophages are not super high, usually not exceeding 20-30%. So indeed the effects the authors observe could be much higher, when checking at the single cell level. I do not know about the SNAT7 ab, but all the other reagents should work via flow cytometry and could hence improve the readout a lot.

      We agree with the reviewer and indeed, in previous studies on HIV-1 infection of human macrophages performed in the lab, we observed via immunofluorescence that the proportion of infected cells ranged from 20 to 40 %. At the time of submission, we did not have the possibility to label the native SNAT7 protein by immunofluorescence, as the commercial antibody used only works for western blotting.

      In the meantime, we have been validating a new antibody (Proteintech) targeting SNAT7 for immunofluorescence. If this is confirmed, we will be able to detect and quantify HIV-1 p24 by immunofluorescence in SNAT7-depleted human macrophages and control cells, thus confirming our results in single-cell analysis.

      Flow cytometry analyses are difficult to perform on primary human macrophages because these cells are highly adherent and must be detached first. The process induces significant cell death and damage. This is why we would prefer to carry out these analyses using immunofluorescence and microscopy on adhered cells. This option will be undoubtedly pursued.

      1. Furthermore the authors never commented about a dose-response effect in terms of HIV-1 infection levels. There is a MOI dependency described for Suppl.Fig.1 C-F, unfortunately the data is missing in the manuscript.

      We apologize for this omission. The figures showing the increase in SNAT7 protein expression following HIV-1 infection at MOIs ranging from 0.05 to 0.5 were added to the new version of the manuscript (Supp. Fig. 1 C-F).

      1. Figure1: specify circulating T lymphocytes. I would expect to see levels of SNAT7 in PHA or CD3/CD28 activated lymphocytes versus resting T cells and a time course of SNAT7 levels upon activation. I think even though SNAT7 levels in T cells might be low, they could also be increased by HIV-1 infection and it is essential that the authors test for this. If not, the result is a valid negative control. For this they should employ HIV-1 primary strains with a tropism for T cells, or at least lab-adapted HIV-1 NL4-3

      We thank the reviewer for this comment. Circulating T lymphocytes isolated from the blood of healthy donors are now referred to resting lymphocytes in the new version of the manuscript, as opposed to activated T lymphocytes stimulated with IL2 and PHA-P for several days (Fig. 1 A-C).

      The expression levels of SNAT7, both at the gene and protein levels, are lower in resting or IL2/PHA-P-activated T cells than in macrophages from the same donors. As suggested, we will perform a kinetic of T-cell activation upon HIV-1 infection to investigate how SNAT7 expression varies in these conditions.

      1. Figure 2 again single cell measurements could reveal much more pronounced effects; it is a bit counterintuitive that siRNA #2 is more efficient in SNAT7 KD but has higher levels of HIV-1 replication in terms of Gag levels. I assume when looking at the stats it is always a comparison to the Ctl treated cells (C-G), but this is not entirely clear. Unify labeling as compared to the stats in Fig.2 I (this also applies for all the other figs).

      We thank the reviewer for this comment. Fig. 2B indeed shows one of the different donors analyzed. However, protein quantification across six different donors shows that SNAT7 is more depleted with siRNA #2 (Fig. 2C), and that Gag Pr55 protein levels are consequently more reduced, than with siRNA #1 (Fig. 2D).

      We use GraphPad Prism software to perform statistical analysis. Depending on the test used, the software automatically plots the comparison bar and displays the p-value above it. We changed the representation of statistics as suggested.

      Figure 3: It is a bit odd that they finally conclude on RT as essential step that is reduced in the absence of SNAT7 and then they fail to provide statistical significance for this (Fig.3 panels F and G). One would expect that RT is much more affected given the huge effects on HIV-1 capsid and particle production shown in Fig.2 F, G and I.

      The reviewer is right in pointing that we observed a stronger effect during the later stages of the viral cycle, from transcription of viral RNAs (Fig. 2I and Supp. Fig. 2G) to the production of viral particles in the supernatant (Fig. 2D-G), than during the earlier stage of reverse transcription (Fig. 3F, G). Also, it is also possible that we might have missed the peak in ERT/LRT production, which is transient.

      It should be noted that SAMHD1 exhibits both dNTPase (Goldstone et al., 2011) and nuclease (Beloglazova et al., 2013) activities. The ability of SAMHD1 to restrict the virus, through dephosphorylation at T592, is mediated by its RNase activity (Ryoo et al., 2014), and not by the dNTPase activity (Welbourn et al., 2013; White et al., 2013).This could explain why SNAT7 exhibit a stronger impact on viral transcription than on reverse transcription.

      Figure 4; again single cell flow measurements of SAMHD1, pSAMHD1 and p24 /SNAT7 might help to more clearly discriminate effects that are specifically induced upon infection or happen in virally infected cells. Maybe alternatively IF?

      We thank the reviewer for this suggestion. As mentioned under comment #2, flow cytometry analyses are difficult to perform on strongly adherent primary human macrophages.

      With regard to immunofluorescence, there is a technical limitation based on the species in which the antibodies are produced. The antibody that targets the native SNAT7 protein, which is currently being validated in our laboratory, is produced in rabbits. An anti-CAp24 antibody produced in goats can be used. It will then be necessary to co-label the cells with anti SAMHD1 and phospho-SAMHD1produced in mouse. We will try to find options to co-label the cells.

      The wblot shown in panel D does not really reflect the point the authors want to make by the quantification in panels G-I. Primary data (D) suggests that SNAT7 KD reduces HIV-1 production even in the absence of SAMHD1. The quantification rather indicates that SNAT7 KD does not affect HIV-1 production in the absence of SAMHD1. This needs clarification/corroboration by orthogonal approaches.

      We respectfully disagree with the reviewer.

      Figure 4D shows a representative blot of the six donors analysed. As mentioned, the depletion of SNAT7 in the absence of SAMHD1 reduces the production of the viral proteins GagPr55 and CAp24 (see Fig. 4D). This is illustrated by the quantifications (Fig. 4G–I). Following treatment with Vpx, GagPr55 protein expression in SNAT7 KD macrophages is reduced by a factor of 2.6 for siRNA #1 (mean = 1.48, light grey bar) and by a factor of 1.83 for siRNA #2 (mean = 2.13, orange bar), compared to the control (mean = 3.9, pink bar) (Fig. 4G). Similarly, CAp24 protein expression was reduced by a factor of 2.2 for siRNA #1 (mean = 2.05, light grey bar) and by a factor of 1.36 for siRNA #2 (mean = 3.34, orange bar), compared to the control (mean = 4.52, pink bar) (Fig. 4H).

      These differences are therefore consistent between the Western blot and the quantifications. However, they are not significantly different to those observed in cells treated with Vpx and depleted with control siRNA, suggesting that the viral restriction observed in SNAT7 KD cells is primarily due to SAMHD1.

      Figure 5: show SAMHD1 and pSAMHD1 levels upon glutamine supplementation.

      We thank the reviewer for this comment, we will perform the suggested experiment.

      1. I think the discussion is very thin, mainly summarizing the results; but fails to give broader context or critically discuss the limitations and further directions.

      We thank the reviewer for this comment. The discussion will be modified further accordingly.

      Looking at the data as a whole, I think the results support a modest functional importance of SNAT7 for HIV-1 production in macrophages. I acknowledge that the experiments in primary macrophages are prone to high variability in different donors and the authors transparently depicted their data. However clearly, I would advice the authors to tune down the extend in which they claim SNAT7-dependency given this huge variability and the sometimes-borderline statistics. We respectfully disagree with the reviewer.

      The cells used here imply greater variability than a cell line, but are also more relevant.

      Indeed, the effects observed in the late stages of HIV-1 production are:

      • ~80 % decrease in viral transcription compared to the control (Fig. 2I),

      • ~85 % decrease in CAp24 protein expression compared to the control, as quantified by western blot (Fig. 2E), or ~90 % by ELISA measurement (Fig. 2F),

      • a reduction of more than 90 % in the release of infectious particles (Fig. 2G).

      These results were all significant across donors, while SNAT7 depletion was always partial (Fig. 2C, between 31 to 62 % of depletion compared to the control in infected cells).

      Therefore, the data were obtained from a mixture of depleted and non-depleted macrophages. This means that the results may be underestimated.

      Together, our results show that SNAT7 is necessary for HIV-1 production.

      However, reading the comments, we realized that our conclusions regarding reverse transcription were too strong. SNAT7 depletion does not affect viral fusion and reverse transcription. The manuscript was modified accordingly.

      On top, there are a lot of optional experiments I am sure the authors are aware of that should be done at least in the future.

      For instance, how does HIV-1 upregulate SNAT7, is a viral accessory protein involved? What is the mechanism of SNAT7 dependent SAMHD1 phosphorylation? Does SNAT7 (or glutamine) regulate the activity of the SAMHD1 associated kinase / phosphatase) If so, does this impact on other targets of these enzymes? We thank the reviewer for these questions.

      To address the role of accessory viral proteins, we have already performed one experiment infecting hMDM with HIV-1 strains deleted for genes such as Nef, Vpr, Vpu and Vif, and have found no clear effect on SNAT7 protein expression compared to WT strains. As an alternative experiment, we could overexpress individual viral genes, such as Nef or Vpr, in HeLa cells and analyze their impact on SNAT7 expression by Western blot.

      It is also possible that SNAT7 expression and recycling of lysosomal glutamine are modulated by the macrophage intrinsic immunity in response to HIV-1 infection.

      The Thr592 motif of the SAMHD1 protein is phosphorylated by Cyclin A2/CDK1 and type 1 IFN in non-cycling cells, such as MDMs (Cribier et al., 2013). For now, the relationship between SNAT7 and SAMHD1 remains unclear. However, (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This has been added to the discussion to explain the relationship between the 3 partners.

      **Referees cross-commenting** I think the comments from the other referees are reasonable and consistent with my assessment

      Reviewer #1 (Significance (Required)):

      Strength and limitations see above;

      Significance: I think this work is of high interest for virologists working in the field of HIV-1 and infection of myeloid cells. In case SNAT7 (and hence glutamine) indeed regulates the phosphorylation of SAMHD1, there could potentially be broad relevance of this work. However unfortunately, this aspect remains underdeveloped and is also not discussed

      Field of expertise: HIV-1, immunology, cell biology

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      In this report, Herit and colleagues describe the role of a HIV-1 dependency factor that promotes virus replication in macrophages. The authors suggest that the lysosomal membrane-associated SNAT7 glutamine transporter is a HIV dependency factor, that promotes virus replication by enhancing reverse transcription and Gag synthesis. The authors use transient knock-down approaches in primary macrophages to identify that SNAT7 depletion does not impact viral entry but inhibits early reverse transcription which was reversed by exogenous glutamine addition. While reverse transcription enhancement was likely due to selective increase in phosho-SAMHD1 expression, mechanisms by which SNAT7 enhanced viral gene expression were not clearly defined. These are well-controlled studies that pinpoint the role of SNAT7 in the early steps of viral life cycle and highlight the intricate interplay between macrophage metabolism and HIV-1 replication. While the question that is addressed is important, and the hypothesis overall sound, the data presented needs to be strengthened to support the conclusions. There are numerous weaknesses in data interpretation as well.

      1. Figure 1: SNAT7 expression was selectively enhanced upon differentiation of monocytes into macrophages but absent in CD4+ T cells. Though there is a claim of enhancement of SNAT7 expression upon HIV-1 infection of macrophages, RT-qPCR analysis shows the opposite trend (Fig 1E) and SNAT7 protein expression changes are modest. Statistical analysis in Fig. 1H needs to be revisited. The number of replicates vary for the lysates harvested at different day post infection, which might have an impact on the statistical test. To determine if SNAT7 expression enhancement is dependent on establishment of virus infection, as the authors imply, control lysates of virus infections in presence of replication inhibitors should be included.

      We thank the reviewer for this comment. Indeed, there is a modest, but statistically significant increase in SNAT7 protein expression upon HIV-1 infection over time (Fig. 1G, H), without any modulation of SNAT7 gene expression (Fig. 1E). This indicates that the regulation of SNAT7 expression in this context is only at the translation level (i.e. increase of translation or stabilization of the SNAT7 protein).

      As mentioned, Fig. 1H aggregates between 3 to 7 independent experiments on different donors depending on the infection time point. SNAT7 protein expression is increased already at 1 day post-infection and until 8 days. The statistical test used here, i.e. 2 way-ANOVA, compared Mock-infected and HIV-1-infected condition for each time point with the same number of donors. In this figure, the comparison is statistically different only at day 6 of the time course (7 donors). We agree that increasing the number of donors of the other time points could help to improve the statistical difference between control and infection condition.

      We thank the reviewer for the suggestion mentioning the use of replication inhibitors in this experiment. We plan to use inhibitors of reverse transcription (Nevirapin) and integration (Dolutegravir).

      The authors rely exclusively on western blot analysis for HIV-1 Gag expression in cell lysates as a measure of effects of SNAT7 on virus replication. Single cell analysis such as intracellular p24gag analysis by FACS should be included; this will provide a better measure of effects of SNAT7 onHIV-1 infection establishment.

      We respectfully disagree with the reviewer for this question. Indeed, to evaluate the effects of SNAT7 on HIV-1 replication, we measured Gag Pr55 and Cap24 using a Western blot approach (Fig. 2B, D and E), but also assessed the quantity of Cap24 in the supernatants and lysates using an ELISA measurement, the quantity of infectious particles using TZM reporter cells, and total viral transcription or more specifically Gag Pr55 transcription using qPCR (Fig. 2F, G and I and Supp. Fig. 2G).

      Regarding the quantification of CAp24 at the cell single level, please refer to comment #2 under Reviewer #1.

      Knockdown of SNAT7 in MDMs was partial at best; only 30-50% decrease in expression (Fig 2C), but the effects on viral gene expression (Fig. 2I), p24 release and infectious particle production is dramatic (Fig. 2F and G). This discrepancy is not addressed. Does SNAT7 knock-down negatively impact virus particle release? Please note that the representative WB in Fig 2B does not correlate with the quantification in Fig. 2D. There are no p55gag or p24gag bands in SNAT7#1 siRNA condition (Fig. 2B)? Data could also be rearranged to follow the logical sequence of virus replication cycle (viral RNa expression followed by Gag expression, and then release).

      We thank the reviewer for this comment. Our samples are indeed a mixture of SNAT7-depleted and non-depleted macrophages and RNA interference in these cells often leads to a decrease of 50 % of the protein expression.

      To determine whether SNAT7 is involved in the release of particles, we quantified Cap24 in cell lysates and in the cell culture medium separately, and normalized the results to the total protein content. The absence of SNAT7 reduced the amount of Cap24 measured by ELISA in both samples to the same extent, showing that there is no storage of Cap24-positive viral particles inside the infected macrophages. These data were initially pooled in one graph (Fig. 2F), but separate graphs are now provided in new Supp. Fig. 2 E, F.

      Regarding the western blot shown in Fig. 2B, please refer to comment #5 under Reviewer #1.

      In the new version of the manuscript, we arranged the figures and placed the later stages of the viral cycle in Fig. 2 and the earlier stages, such as fusion, reverse transcription and transcription, in Fig. 3.

      Data interpretation would be greatly improved by including infection controls (RT or integrase inhibitors) to confirm that measurements of viral RNA and Gag are indeed modulated by SNAT7 expression.

      We thank the reviewer for this suggestion to include inhibitors of viral replication as controls. In our experiments, cells were Mock-infected in parallel as a negative control of viral detection. We provide the results in the new version of the manuscript to show that (i) there is no detection of viral or Gag RNA in the absence of the virus, (ii) the expression of viral genes measured in HIV-1-infected SNAT7-depleted cells is not different from Mock-infected cells, indicating almost complete inhibition of viral transcription (Fig. 3H and Supp. Fig. 3B), also confirmed at the protein level (Fig. 2B, D-F).

      Figure 3: Decrease in SNAT7 expression in macrophages resulted in lower levels of early reverse transcripts. But surprisingly, LRT levels were not as affected by decreases in SNAT7 expression. The authors go on to suggest that decreases in early RT are due to loss of phospho-SAMHD1 and increases in catalytically active form of SAMHD1. Mechanistically this does not make sense: LRT should be similarly affected by increase in catalytically active SAMHD1. dNTP concentrations should be measured to determine if the rescue of RT is dependent on SAMHD1 dNTPase activity.

      We thank the reviewer for this comment. LRT concentrations are very low in human macrophages and more challenging to detect than ERT concentrations. This might explain why the differences observed between the SNAT7-depleted and control conditions appear less pronounced for LRT than for ERT.

      Furthermore, we cannot rule out the possibility that SNAT7 has a cumulative effect throughout the viral cycle. While reverse transcription remains statistically unaltered, and despite the reduced levels of ERT and LRT in SNAT7-depleted macrophages (Fig. 3 F, G), there is a significant impact on the transcription of viral RNAs (Fig. 2I) and Gag (Supp. Fig. 2G). This step may also be altered by the ribonuclease activity of SAMHD1 (Beloglazova et al., 2013; Ryoo et al., 2014).

      Finally, with the help of Dr Baek Kim in Atlanta, we attempted to quantify dNTP concentrations in our human macrophages. Unfortunately, it was not possible to draw any conclusions, as the concentrations of dNTPs extracted from our cells were far too low.

      Furthermore, it should be noted that SAMHD1 viral restriction through its phosphorylation at T592 is not correlated with its dNTPase activity (Welbourn et al., 2013; White et al., 2013), but with its ribonuclease activity (Beloglazova et al., 2013; Ryoo et al., 2014). This is supporting why SNAT7, by modulating the ribonuclease activity of SAMHD1, could have a greater effect on viral transcription than on reverse transcription.

      There is lack of consistency in the data: p24 release upon SNAT7 depletion is highly variable. While there is a dramatic >90-95% decrease in p24 release (Fig. 2G), the effects are much more moderate in Fig. 4H (50-60% attenuation), even though siRNA-mediated depletion was similar across the data sets. The authors should comment on the variability in their findings.

      We thank the reviewer for this comment, but believe that Figure 2E rather than Figure 2G is to be mentioned regarding the quantification of CAp24 by Western blot and to be compared with Figure 4H.

      In Fig. 2E, we observed an average reduction of 85 % in CAp24 expression normalized to Clathrin HC expression across different donors for both siRNAs targeting SNAT7. For Fig. 4H, there was a 73 % reduction in CAp24 levels for siRNA #1 and a 56 % reduction for siRNA #2. In addition, it should be noted that the reduction in Gag levels is greater in Fig. 4G (between 77 % and 83 %) than in Fig. 2D (between 55 % and 72 %).

      Therefore, there is some variation in the results obtained with the different donors, which could be explained by variations in Gag cleavage among donors, but this does not impact the conclusions for both figures.

      SNAT7 is postulated to affect 2 steps in the virus life cycle: reverse transcription and viral transcription. But Vpx-mediated SAMHD1 degradation reversed both. Its not clear to me as to how SAMHD1 degradation impacts the role of SNAT7 in viral transcription. No explanation is provided.

      We thank the reviewer for this comment. As suggested, we will perform experiments to assess the impact of Vpx-mediated SAMHD1 degradation on viral transcription.

      Exogenous addition of glutamine only partially restored Gag synthesis and p24 release, which could be attributed to increased cytoplasmic levels and viral protein synthesis. What about effects on reverse transcription and viral gene expression?

      We thank the reviewer for this comment. We will perform the suggested experiments to assess the impact of glutamine supplementation on viral transcription.

      Reviewer #2 (Significance (Required)):

      This is a novel finding, as there are limited number of studies on amino acid transporters and HIV-1 replication enhancement in macrophages. Most of the previous work has focused on CD4 T cells. These studies on SNAT7 and HIV-1 infection establishment in macrophages might better inform the influences of macrophage metabolism on HIV-1 persistence and inflammatory responses.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      This study investigates the role of the lysosomal glutamine transporter SLC38A7/SNAT7 in HIV‑1 replication in primary human macrophages. The authors demonstrate that SNAT7 is highly expressed in macrophages and upregulated upon HIV‑1 infection. They show that SNAT7 depletion inhibits HIV‑1 production at the reverse transcription step without affecting viral fusion or global cellular translation/transcription. Mechanistically, SNAT7 knockdown reduces the inhibitory phosphorylation of SAMHD1 at T592, and degradation of SAMHD1 by Vpx fully rescues viral replication. Extracellular glutamine supplementation partially restores HIV‑1 production in SNAT7‑deficient cells. Overall, the authors report interesting observations; however, the mechanistic investigation remains preliminary, raising concerns about whether the data fully support all the conclusions drawn. Major Concerns: 1. The mechanistic depth is insufficient. The authors do not elucidate how glutamine regulates SAMHD1 T592 phosphorylation, whether through metabolite‑mediated control of kinases/phosphatases or via indirect effects.

      We thank the reviewer for this comment. It is worth noting that (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity using drugs decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This is now further discussed in the discussion section of the manuscript.

      The authors do not measure intracellular dNTP levels upon SNAT7 knockdown, which is the key functional substrate of SAMHD1. They also do not directly demonstrate that glutamine supplementation restores dNTP pools.

      We thank the reviewer for this comment. Please, refer to comment #5 under Reviewer #2.

      Extracellular glutamine only partially rescues viral production, implying the existence of transport‑independent functions of SNAT7 or additional pathways. This important observation is not discussed.

      We thank the reviewer for this comment. The discussion has been modified accordingly.

      It is suggested that the key findings be validated in immortalized THP‑1 cells differentiated into macrophage‑like cells by PMA.

      We thank the reviewer for this suggestion but don’t really understand why this would strengthen our conclusions. Indeed, despite the known variability between donors and technical limitations to transduce cells, we chose human blood monocyte-derived macrophages as a relevant non-transformed model for HIV-1 infection of macrophages. They also represent to some extent the human diversity.

      The Discussion section should be expanded to include the potential translational implications and limitations of the present study.

      We thank the reviewer for this comment. The discussion points to some elements of potential translation and limitations of the study.

      Reviewer #3 (Significance (Required)):

      General assessment: This study identifies the lysosomal glutamine transporter SLC38A7/SNAT7 as a novel host dependency factor for HIV‑1 replication in primary human macrophages. The major strengths include the use of physiologically relevant primary macrophage models, a well-organized experimental pipeline from expression profiling to functional validation, and the establishment of a link between SNAT7, glutamine metabolism, and the HIV restriction factor SAMHD1.

      Advance: It extends current understanding of HIV‑1 host dependency factors and immunometabolism by revealing a compartment‑specific metabolic pathway that supports viral reverse transcription.

      Audience:This work will primarily interest specialized researchers in HIV‑1 biology, host-virus interactions, restriction factors, and antiviral innate immunity.

      Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      This study from the Niedergang lab establishes SNAT7 as a host-dependency factor in human macrophages that supports HIV-1 replication. They show a modest increase in SNAT7 levels HIV-1 infected macrophages and suggest that SNAT7 levels are transiently increased. Employing siRNA against SNAT7 they show reduction in HIV-1 protein levels and viral RNAs and claim that there is a block of reverse transcription in SNAT7 KD cells. Focusing on a known HIV-1 restriction factor in macrophages, SAMHD1, they interconnect the SNAT7 depletion with a reduction in phosphorylated, i.e. catalytical inactive SAMHD1 arguing that SNAT7 regulates the phosphorylation and thereby antiviral activity of SAMHD1. Since SNAT7 is a glutamine transporter that provides this AA from lysosomes, they lastly supplement glutamine and this somehow rescues the reduction of HIV-1 production in SNAT7 KD cells.

      Major comments:

      The strength of this manuscript is the clear focus on primary human macrophages that are HIV-1 infected and the interconnection of HIV-1 replication to the SNAT7 siRNA KD experiments in combination with SAMHD1 depletion and lastly glutamine supplementation. This establishes a stringent and coherent story line. The effects reported are modest; high variability is not a problem since using primary hMDM this is expected and can be addressed by testing several donors and applying stringent statistics.

      1. Having said so, I realize that while they give information on the statistical test used, i.e. one-way ANOVA they miss to explain the post-test used to assess significance (i.e. Bonferroni, Fishers LSD, whatsoever). Please add this information.

      We thank the reviewer for this comment. The figure legends have been updated to include more details of all the statistical tests used.

      1. Another issue that might underestimate the effects of HIV-1 infection on SNAT7 levels and vice versa of SNAT7 KD on HIV-1 replication is the non-single cell approach employed, i.e. WBlots. I assume that HIV-1 infection rates in macrophages are not super high, usually not exceeding 20-30%. So indeed the effects the authors observe could be much higher, when checking at the single cell level. I do not know about the SNAT7 ab, but all the other reagents should work via flow cytometry and could hence improve the readout a lot.

      We agree with the reviewer and indeed, in previous studies on HIV-1 infection of human macrophages performed in the lab, we observed via immunofluorescence that the proportion of infected cells ranged from 20 to 40 %. At the time of submission, we did not have the possibility to label the native SNAT7 protein by immunofluorescence, as the commercial antibody used only works for western blotting.

      In the meantime, we have been validating a new antibody (Proteintech) targeting SNAT7 for immunofluorescence. If this is confirmed, we will be able to detect and quantify HIV-1 p24 by immunofluorescence in SNAT7-depleted human macrophages and control cells, thus confirming our results in single-cell analysis.

      Flow cytometry analyses are difficult to perform on primary human macrophages because these cells are highly adherent and must be detached first. The process induces significant cell death and damage. This is why we would prefer to carry out these analyses using immunofluorescence and microscopy on adhered cells. This option will be undoubtedly pursued.

      1. Furthermore the authors never commented about a dose-response effect in terms of HIV-1 infection levels. There is a MOI dependency described for Suppl.Fig.1 C-F, unfortunately the data is missing in the manuscript.

      We apologize for this omission. The figures showing the increase in SNAT7 protein expression following HIV-1 infection at MOIs ranging from 0.05 to 0.5 were added to the new version of the manuscript (Supp. Fig. 1 C-F).

      1. Figure1: specify circulating T lymphocytes. I would expect to see levels of SNAT7 in PHA or CD3/CD28 activated lymphocytes versus resting T cells and a time course of SNAT7 levels upon activation. I think even though SNAT7 levels in T cells might be low, they could also be increased by HIV-1 infection and it is essential that the authors test for this. If not, the result is a valid negative control. For this they should employ HIV-1 primary strains with a tropism for T cells, or at least lab-adapted HIV-1 NL4-3

      We thank the reviewer for this comment. Circulating T lymphocytes isolated from the blood of healthy donors are now referred to resting lymphocytes in the new version of the manuscript, as opposed to activated T lymphocytes stimulated with IL2 and PHA-P for several days (Fig. 1 A-C).

      The expression levels of SNAT7, both at the gene and protein levels, are lower in resting or IL2/PHA-P-activated T cells than in macrophages from the same donors. As suggested, we will perform a kinetic of T-cell activation upon HIV-1 infection to investigate how SNAT7 expression varies in these conditions.

      1. Figure 2 again single cell measurements could reveal much more pronounced effects; it is a bit counterintuitive that siRNA #2 is more efficient in SNAT7 KD but has higher levels of HIV-1 replication in terms of Gag levels. I assume when looking at the stats it is always a comparison to the Ctl treated cells (C-G), but this is not entirely clear. Unify labeling as compared to the stats in Fig.2 I (this also applies for all the other figs).

      We thank the reviewer for this comment. Fig. 2B indeed shows one of the different donors analyzed. However, protein quantification across six different donors shows that SNAT7 is more depleted with siRNA #2 (Fig. 2C), and that Gag Pr55 protein levels are consequently more reduced, than with siRNA #1 (Fig. 2D).

      We use GraphPad Prism software to perform statistical analysis. Depending on the test used, the software automatically plots the comparison bar and displays the p-value above it. We changed the representation of statistics as suggested.

      Figure 3: It is a bit odd that they finally conclude on RT as essential step that is reduced in the absence of SNAT7 and then they fail to provide statistical significance for this (Fig.3 panels F and G). One would expect that RT is much more affected given the huge effects on HIV-1 capsid and particle production shown in Fig.2 F, G and I.

      The reviewer is right in pointing that we observed a stronger effect during the later stages of the viral cycle, from transcription of viral RNAs (Fig. 2I and Supp. Fig. 2G) to the production of viral particles in the supernatant (Fig. 2D-G), than during the earlier stage of reverse transcription (Fig. 3F, G). Also, it is also possible that we might have missed the peak in ERT/LRT production, which is transient.

      It should be noted that SAMHD1 exhibits both dNTPase (Goldstone et al., 2011) and nuclease (Beloglazova et al., 2013) activities. The ability of SAMHD1 to restrict the virus, through dephosphorylation at T592, is mediated by its RNase activity (Ryoo et al., 2014), and not by the dNTPase activity (Welbourn et al., 2013; White et al., 2013).This could explain why SNAT7 exhibit a stronger impact on viral transcription than on reverse transcription.

      Figure 4; again single cell flow measurements of SAMHD1, pSAMHD1 and p24 /SNAT7 might help to more clearly discriminate effects that are specifically induced upon infection or happen in virally infected cells. Maybe alternatively IF?

      We thank the reviewer for this suggestion. As mentioned under comment #2, flow cytometry analyses are difficult to perform on strongly adherent primary human macrophages.

      With regard to immunofluorescence, there is a technical limitation based on the species in which the antibodies are produced. The antibody that targets the native SNAT7 protein, which is currently being validated in our laboratory, is produced in rabbits. An anti-CAp24 antibody produced in goats can be used. It will then be necessary to co-label the cells with anti SAMHD1 and phospho-SAMHD1produced in mouse. We will try to find options to co-label the cells.

      The wblot shown in panel D does not really reflect the point the authors want to make by the quantification in panels G-I. Primary data (D) suggests that SNAT7 KD reduces HIV-1 production even in the absence of SAMHD1. The quantification rather indicates that SNAT7 KD does not affect HIV-1 production in the absence of SAMHD1. This needs clarification/corroboration by orthogonal approaches.

      We respectfully disagree with the reviewer.

      Figure 4D shows a representative blot of the six donors analysed. As mentioned, the depletion of SNAT7 in the absence of SAMHD1 reduces the production of the viral proteins GagPr55 and CAp24 (see Fig. 4D). This is illustrated by the quantifications (Fig. 4G–I). Following treatment with Vpx, GagPr55 protein expression in SNAT7 KD macrophages is reduced by a factor of 2.6 for siRNA #1 (mean = 1.48, light grey bar) and by a factor of 1.83 for siRNA #2 (mean = 2.13, orange bar), compared to the control (mean = 3.9, pink bar) (Fig. 4G). Similarly, CAp24 protein expression was reduced by a factor of 2.2 for siRNA #1 (mean = 2.05, light grey bar) and by a factor of 1.36 for siRNA #2 (mean = 3.34, orange bar), compared to the control (mean = 4.52, pink bar) (Fig. 4H).

      These differences are therefore consistent between the Western blot and the quantifications. However, they are not significantly different to those observed in cells treated with Vpx and depleted with control siRNA, suggesting that the viral restriction observed in SNAT7 KD cells is primarily due to SAMHD1.

      Figure 5: show SAMHD1 and pSAMHD1 levels upon glutamine supplementation.

      We thank the reviewer for this comment, we will perform the suggested experiment.

      1. I think the discussion is very thin, mainly summarizing the results; but fails to give broader context or critically discuss the limitations and further directions.

      We thank the reviewer for this comment. The discussion will be modified further accordingly.

      Looking at the data as a whole, I think the results support a modest functional importance of SNAT7 for HIV-1 production in macrophages. I acknowledge that the experiments in primary macrophages are prone to high variability in different donors and the authors transparently depicted their data. However clearly, I would advice the authors to tune down the extend in which they claim SNAT7-dependency given this huge variability and the sometimes-borderline statistics. We respectfully disagree with the reviewer.

      The cells used here imply greater variability than a cell line, but are also more relevant.

      Indeed, the effects observed in the late stages of HIV-1 production are:

      • ~80 % decrease in viral transcription compared to the control (Fig. 2I),

      • ~85 % decrease in CAp24 protein expression compared to the control, as quantified by western blot (Fig. 2E), or ~90 % by ELISA measurement (Fig. 2F),

      • a reduction of more than 90 % in the release of infectious particles (Fig. 2G).

      These results were all significant across donors, while SNAT7 depletion was always partial (Fig. 2C, between 31 to 62 % of depletion compared to the control in infected cells).

      Therefore, the data were obtained from a mixture of depleted and non-depleted macrophages. This means that the results may be underestimated.

      Together, our results show that SNAT7 is necessary for HIV-1 production.

      However, reading the comments, we realized that our conclusions regarding reverse transcription were too strong. SNAT7 depletion does not affect viral fusion and reverse transcription. The manuscript was modified accordingly.

      On top, there are a lot of optional experiments I am sure the authors are aware of that should be done at least in the future.

      For instance, how does HIV-1 upregulate SNAT7, is a viral accessory protein involved? What is the mechanism of SNAT7 dependent SAMHD1 phosphorylation? Does SNAT7 (or glutamine) regulate the activity of the SAMHD1 associated kinase / phosphatase) If so, does this impact on other targets of these enzymes? We thank the reviewer for these questions.

      To address the role of accessory viral proteins, we have already performed one experiment infecting hMDM with HIV-1 strains deleted for genes such as Nef, Vpr, Vpu and Vif, and have found no clear effect on SNAT7 protein expression compared to WT strains. As an alternative experiment, we could overexpress individual viral genes, such as Nef or Vpr, in HeLa cells and analyze their impact on SNAT7 expression by Western blot.

      It is also possible that SNAT7 expression and recycling of lysosomal glutamine are modulated by the macrophage intrinsic immunity in response to HIV-1 infection.

      The Thr592 motif of the SAMHD1 protein is phosphorylated by Cyclin A2/CDK1 and type 1 IFN in non-cycling cells, such as MDMs (Cribier et al., 2013). For now, the relationship between SNAT7 and SAMHD1 remains unclear. However, (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This has been added to the discussion to explain the relationship between the 3 partners.

      **Referees cross-commenting** I think the comments from the other referees are reasonable and consistent with my assessment

      Reviewer #1 (Significance (Required)):

      Strength and limitations see above;

      Significance: I think this work is of high interest for virologists working in the field of HIV-1 and infection of myeloid cells. In case SNAT7 (and hence glutamine) indeed regulates the phosphorylation of SAMHD1, there could potentially be broad relevance of this work. However unfortunately, this aspect remains underdeveloped and is also not discussed

      Field of expertise: HIV-1, immunology, cell biology

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      In this report, Herit and colleagues describe the role of a HIV-1 dependency factor that promotes virus replication in macrophages. The authors suggest that the lysosomal membrane-associated SNAT7 glutamine transporter is a HIV dependency factor, that promotes virus replication by enhancing reverse transcription and Gag synthesis. The authors use transient knock-down approaches in primary macrophages to identify that SNAT7 depletion does not impact viral entry but inhibits early reverse transcription which was reversed by exogenous glutamine addition. While reverse transcription enhancement was likely due to selective increase in phosho-SAMHD1 expression, mechanisms by which SNAT7 enhanced viral gene expression were not clearly defined. These are well-controlled studies that pinpoint the role of SNAT7 in the early steps of viral life cycle and highlight the intricate interplay between macrophage metabolism and HIV-1 replication. While the question that is addressed is important, and the hypothesis overall sound, the data presented needs to be strengthened to support the conclusions. There are numerous weaknesses in data interpretation as well.

      1. Figure 1: SNAT7 expression was selectively enhanced upon differentiation of monocytes into macrophages but absent in CD4+ T cells. Though there is a claim of enhancement of SNAT7 expression upon HIV-1 infection of macrophages, RT-qPCR analysis shows the opposite trend (Fig 1E) and SNAT7 protein expression changes are modest. Statistical analysis in Fig. 1H needs to be revisited. The number of replicates vary for the lysates harvested at different day post infection, which might have an impact on the statistical test. To determine if SNAT7 expression enhancement is dependent on establishment of virus infection, as the authors imply, control lysates of virus infections in presence of replication inhibitors should be included.

      We thank the reviewer for this comment. Indeed, there is a modest, but statistically significant increase in SNAT7 protein expression upon HIV-1 infection over time (Fig. 1G, H), without any modulation of SNAT7 gene expression (Fig. 1E). This indicates that the regulation of SNAT7 expression in this context is only at the translation level (i.e. increase of translation or stabilization of the SNAT7 protein).

      As mentioned, Fig. 1H aggregates between 3 to 7 independent experiments on different donors depending on the infection time point. SNAT7 protein expression is increased already at 1 day post-infection and until 8 days. The statistical test used here, i.e. 2 way-ANOVA, compared Mock-infected and HIV-1-infected condition for each time point with the same number of donors. In this figure, the comparison is statistically different only at day 6 of the time course (7 donors). We agree that increasing the number of donors of the other time points could help to improve the statistical difference between control and infection condition.

      We thank the reviewer for the suggestion mentioning the use of replication inhibitors in this experiment. We plan to use inhibitors of reverse transcription (Nevirapin) and integration (Dolutegravir).

      The authors rely exclusively on western blot analysis for HIV-1 Gag expression in cell lysates as a measure of effects of SNAT7 on virus replication. Single cell analysis such as intracellular p24gag analysis by FACS should be included; this will provide a better measure of effects of SNAT7 onHIV-1 infection establishment.

      We respectfully disagree with the reviewer for this question. Indeed, to evaluate the effects of SNAT7 on HIV-1 replication, we measured Gag Pr55 and Cap24 using a Western blot approach (Fig. 2B, D and E), but also assessed the quantity of Cap24 in the supernatants and lysates using an ELISA measurement, the quantity of infectious particles using TZM reporter cells, and total viral transcription or more specifically Gag Pr55 transcription using qPCR (Fig. 2F, G and I and Supp. Fig. 2G).

      Regarding the quantification of CAp24 at the cell single level, please refer to comment #2 under Reviewer #1.

      Knockdown of SNAT7 in MDMs was partial at best; only 30-50% decrease in expression (Fig 2C), but the effects on viral gene expression (Fig. 2I), p24 release and infectious particle production is dramatic (Fig. 2F and G). This discrepancy is not addressed. Does SNAT7 knock-down negatively impact virus particle release? Please note that the representative WB in Fig 2B does not correlate with the quantification in Fig. 2D. There are no p55gag or p24gag bands in SNAT7#1 siRNA condition (Fig. 2B)? Data could also be rearranged to follow the logical sequence of virus replication cycle (viral RNa expression followed by Gag expression, and then release).

      We thank the reviewer for this comment. Our samples are indeed a mixture of SNAT7-depleted and non-depleted macrophages and RNA interference in these cells often leads to a decrease of 50 % of the protein expression.

      To determine whether SNAT7 is involved in the release of particles, we quantified Cap24 in cell lysates and in the cell culture medium separately, and normalized the results to the total protein content. The absence of SNAT7 reduced the amount of Cap24 measured by ELISA in both samples to the same extent, showing that there is no storage of Cap24-positive viral particles inside the infected macrophages. These data were initially pooled in one graph (Fig. 2F), but separate graphs are now provided in new Supp. Fig. 2 E, F.

      Regarding the western blot shown in Fig. 2B, please refer to comment #5 under Reviewer #1.

      In the new version of the manuscript, we arranged the figures and placed the later stages of the viral cycle in Fig. 2 and the earlier stages, such as fusion, reverse transcription and transcription, in Fig. 3.

      Data interpretation would be greatly improved by including infection controls (RT or integrase inhibitors) to confirm that measurements of viral RNA and Gag are indeed modulated by SNAT7 expression.

      We thank the reviewer for this suggestion to include inhibitors of viral replication as controls. In our experiments, cells were Mock-infected in parallel as a negative control of viral detection. We provide the results in the new version of the manuscript to show that (i) there is no detection of viral or Gag RNA in the absence of the virus, (ii) the expression of viral genes measured in HIV-1-infected SNAT7-depleted cells is not different from Mock-infected cells, indicating almost complete inhibition of viral transcription (Fig. 3H and Supp. Fig. 3B), also confirmed at the protein level (Fig. 2B, D-F).

      Figure 3: Decrease in SNAT7 expression in macrophages resulted in lower levels of early reverse transcripts. But surprisingly, LRT levels were not as affected by decreases in SNAT7 expression. The authors go on to suggest that decreases in early RT are due to loss of phospho-SAMHD1 and increases in catalytically active form of SAMHD1. Mechanistically this does not make sense: LRT should be similarly affected by increase in catalytically active SAMHD1. dNTP concentrations should be measured to determine if the rescue of RT is dependent on SAMHD1 dNTPase activity.

      We thank the reviewer for this comment. LRT concentrations are very low in human macrophages and more challenging to detect than ERT concentrations. This might explain why the differences observed between the SNAT7-depleted and control conditions appear less pronounced for LRT than for ERT.

      Furthermore, we cannot rule out the possibility that SNAT7 has a cumulative effect throughout the viral cycle. While reverse transcription remains statistically unaltered, and despite the reduced levels of ERT and LRT in SNAT7-depleted macrophages (Fig. 3 F, G), there is a significant impact on the transcription of viral RNAs (Fig. 2I) and Gag (Supp. Fig. 2G). This step may also be altered by the ribonuclease activity of SAMHD1 (Beloglazova et al., 2013; Ryoo et al., 2014).

      Finally, with the help of Dr Baek Kim in Atlanta, we attempted to quantify dNTP concentrations in our human macrophages. Unfortunately, it was not possible to draw any conclusions, as the concentrations of dNTPs extracted from our cells were far too low.

      Furthermore, it should be noted that SAMHD1 viral restriction through its phosphorylation at T592 is not correlated with its dNTPase activity (Welbourn et al., 2013; White et al., 2013), but with its ribonuclease activity (Beloglazova et al., 2013; Ryoo et al., 2014). This is supporting why SNAT7, by modulating the ribonuclease activity of SAMHD1, could have a greater effect on viral transcription than on reverse transcription.

      There is lack of consistency in the data: p24 release upon SNAT7 depletion is highly variable. While there is a dramatic >90-95% decrease in p24 release (Fig. 2G), the effects are much more moderate in Fig. 4H (50-60% attenuation), even though siRNA-mediated depletion was similar across the data sets. The authors should comment on the variability in their findings.

      We thank the reviewer for this comment, but believe that Figure 2E rather than Figure 2G is to be mentioned regarding the quantification of CAp24 by Western blot and to be compared with Figure 4H.

      In Fig. 2E, we observed an average reduction of 85 % in CAp24 expression normalized to Clathrin HC expression across different donors for both siRNAs targeting SNAT7. For Fig. 4H, there was a 73 % reduction in CAp24 levels for siRNA #1 and a 56 % reduction for siRNA #2. In addition, it should be noted that the reduction in Gag levels is greater in Fig. 4G (between 77 % and 83 %) than in Fig. 2D (between 55 % and 72 %).

      Therefore, there is some variation in the results obtained with the different donors, which could be explained by variations in Gag cleavage among donors, but this does not impact the conclusions for both figures.

      SNAT7 is postulated to affect 2 steps in the virus life cycle: reverse transcription and viral transcription. But Vpx-mediated SAMHD1 degradation reversed both. Its not clear to me as to how SAMHD1 degradation impacts the role of SNAT7 in viral transcription. No explanation is provided.

      We thank the reviewer for this comment. As suggested, we will perform experiments to assess the impact of Vpx-mediated SAMHD1 degradation on viral transcription.

      Exogenous addition of glutamine only partially restored Gag synthesis and p24 release, which could be attributed to increased cytoplasmic levels and viral protein synthesis. What about effects on reverse transcription and viral gene expression?

      We thank the reviewer for this comment. We will perform the suggested experiments to assess the impact of glutamine supplementation on viral transcription.

      Reviewer #2 (Significance (Required)):

      This is a novel finding, as there are limited number of studies on amino acid transporters and HIV-1 replication enhancement in macrophages. Most of the previous work has focused on CD4 T cells. These studies on SNAT7 and HIV-1 infection establishment in macrophages might better inform the influences of macrophage metabolism on HIV-1 persistence and inflammatory responses.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      This study investigates the role of the lysosomal glutamine transporter SLC38A7/SNAT7 in HIV‑1 replication in primary human macrophages. The authors demonstrate that SNAT7 is highly expressed in macrophages and upregulated upon HIV‑1 infection. They show that SNAT7 depletion inhibits HIV‑1 production at the reverse transcription step without affecting viral fusion or global cellular translation/transcription. Mechanistically, SNAT7 knockdown reduces the inhibitory phosphorylation of SAMHD1 at T592, and degradation of SAMHD1 by Vpx fully rescues viral replication. Extracellular glutamine supplementation partially restores HIV‑1 production in SNAT7‑deficient cells. Overall, the authors report interesting observations; however, the mechanistic investigation remains preliminary, raising concerns about whether the data fully support all the conclusions drawn. Major Concerns: 1. The mechanistic depth is insufficient. The authors do not elucidate how glutamine regulates SAMHD1 T592 phosphorylation, whether through metabolite‑mediated control of kinases/phosphatases or via indirect effects.

      We thank the reviewer for this comment. It is worth noting that (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity using drugs decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This is now further discussed in the discussion section of the manuscript.

      The authors do not measure intracellular dNTP levels upon SNAT7 knockdown, which is the key functional substrate of SAMHD1. They also do not directly demonstrate that glutamine supplementation restores dNTP pools.

      We thank the reviewer for this comment. Please, refer to comment #5 under Reviewer #2.

      Extracellular glutamine only partially rescues viral production, implying the existence of transport‑independent functions of SNAT7 or additional pathways. This important observation is not discussed.

      We thank the reviewer for this comment. The discussion has been modified accordingly.

      It is suggested that the key findings be validated in immortalized THP‑1 cells differentiated into macrophage‑like cells by PMA.

      We thank the reviewer for this suggestion but don’t really understand why this would strengthen our conclusions. Indeed, despite the known variability between donors and technical limitations to transduce cells, we chose human blood monocyte-derived macrophages as a relevant non-transformed model for HIV-1 infection of macrophages. They also represent to some extent the human diversity.

      The Discussion section should be expanded to include the potential translational implications and limitations of the present study.

      We thank the reviewer for this comment. The discussion points to some elements of potential translation and limitations of the study.

      Reviewer #3 (Significance (Required)):

      General assessment: This study identifies the lysosomal glutamine transporter SLC38A7/SNAT7 as a novel host dependency factor for HIV‑1 replication in primary human macrophages. The major strengths include the use of physiologically relevant primary macrophage models, a well-organized experimental pipeline from expression profiling to functional validation, and the establishment of a link between SNAT7, glutamine metabolism, and the HIV restriction factor SAMHD1.

      Advance: It extends current understanding of HIV‑1 host dependency factors and immunometabolism by revealing a compartment‑specific metabolic pathway that supports viral reverse transcription.

      Audience:This work will primarily interest specialized researchers in HIV‑1 biology, host-virus interactions, restriction factors, and antiviral innate immunity.

      Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      This study from the Niedergang lab establishes SNAT7 as a host-dependency factor in human macrophages that supports HIV-1 replication. They show a modest increase in SNAT7 levels HIV-1 infected macrophages and suggest that SNAT7 levels are transiently increased. Employing siRNA against SNAT7 they show reduction in HIV-1 protein levels and viral RNAs and claim that there is a block of reverse transcription in SNAT7 KD cells. Focusing on a known HIV-1 restriction factor in macrophages, SAMHD1, they interconnect the SNAT7 depletion with a reduction in phosphorylated, i.e. catalytical inactive SAMHD1 arguing that SNAT7 regulates the phosphorylation and thereby antiviral activity of SAMHD1. Since SNAT7 is a glutamine transporter that provides this AA from lysosomes, they lastly supplement glutamine and this somehow rescues the reduction of HIV-1 production in SNAT7 KD cells.

      Major comments:

      The strength of this manuscript is the clear focus on primary human macrophages that are HIV-1 infected and the interconnection of HIV-1 replication to the SNAT7 siRNA KD experiments in combination with SAMHD1 depletion and lastly glutamine supplementation. This establishes a stringent and coherent story line. The effects reported are modest; high variability is not a problem since using primary hMDM this is expected and can be addressed by testing several donors and applying stringent statistics.

      1. Having said so, I realize that while they give information on the statistical test used, i.e. one-way ANOVA they miss to explain the post-test used to assess significance (i.e. Bonferroni, Fishers LSD, whatsoever). Please add this information.

      We thank the reviewer for this comment. The figure legends have been updated to include more details of all the statistical tests used.

      1. Another issue that might underestimate the effects of HIV-1 infection on SNAT7 levels and vice versa of SNAT7 KD on HIV-1 replication is the non-single cell approach employed, i.e. WBlots. I assume that HIV-1 infection rates in macrophages are not super high, usually not exceeding 20-30%. So indeed the effects the authors observe could be much higher, when checking at the single cell level. I do not know about the SNAT7 ab, but all the other reagents should work via flow cytometry and could hence improve the readout a lot.

      We agree with the reviewer and indeed, in previous studies on HIV-1 infection of human macrophages performed in the lab, we observed via immunofluorescence that the proportion of infected cells ranged from 20 to 40 %. At the time of submission, we did not have the possibility to label the native SNAT7 protein by immunofluorescence, as the commercial antibody used only works for western blotting.

      In the meantime, we have been validating a new antibody (Proteintech) targeting SNAT7 for immunofluorescence. If this is confirmed, we will be able to detect and quantify HIV-1 p24 by immunofluorescence in SNAT7-depleted human macrophages and control cells, thus confirming our results in single-cell analysis.

      Flow cytometry analyses are difficult to perform on primary human macrophages because these cells are highly adherent and must be detached first. The process induces significant cell death and damage. This is why we would prefer to carry out these analyses using immunofluorescence and microscopy on adhered cells. This option will be undoubtedly pursued.

      1. Furthermore the authors never commented about a dose-response effect in terms of HIV-1 infection levels. There is a MOI dependency described for Suppl.Fig.1 C-F, unfortunately the data is missing in the manuscript.

      We apologize for this omission. The figures showing the increase in SNAT7 protein expression following HIV-1 infection at MOIs ranging from 0.05 to 0.5 were added to the new version of the manuscript (Supp. Fig. 1 C-F).

      1. Figure1: specify circulating T lymphocytes. I would expect to see levels of SNAT7 in PHA or CD3/CD28 activated lymphocytes versus resting T cells and a time course of SNAT7 levels upon activation. I think even though SNAT7 levels in T cells might be low, they could also be increased by HIV-1 infection and it is essential that the authors test for this. If not, the result is a valid negative control. For this they should employ HIV-1 primary strains with a tropism for T cells, or at least lab-adapted HIV-1 NL4-3

      We thank the reviewer for this comment. Circulating T lymphocytes isolated from the blood of healthy donors are now referred to resting lymphocytes in the new version of the manuscript, as opposed to activated T lymphocytes stimulated with IL2 and PHA-P for several days (Fig. 1 A-C).

      The expression levels of SNAT7, both at the gene and protein levels, are lower in resting or IL2/PHA-P-activated T cells than in macrophages from the same donors. As suggested, we will perform a kinetic of T-cell activation upon HIV-1 infection to investigate how SNAT7 expression varies in these conditions.

      1. Figure 2 again single cell measurements could reveal much more pronounced effects; it is a bit counterintuitive that siRNA #2 is more efficient in SNAT7 KD but has higher levels of HIV-1 replication in terms of Gag levels. I assume when looking at the stats it is always a comparison to the Ctl treated cells (C-G), but this is not entirely clear. Unify labeling as compared to the stats in Fig.2 I (this also applies for all the other figs).

      We thank the reviewer for this comment. Fig. 2B indeed shows one of the different donors analyzed. However, protein quantification across six different donors shows that SNAT7 is more depleted with siRNA #2 (Fig. 2C), and that Gag Pr55 protein levels are consequently more reduced, than with siRNA #1 (Fig. 2D).

      We use GraphPad Prism software to perform statistical analysis. Depending on the test used, the software automatically plots the comparison bar and displays the p-value above it. We changed the representation of statistics as suggested.

      Figure 3: It is a bit odd that they finally conclude on RT as essential step that is reduced in the absence of SNAT7 and then they fail to provide statistical significance for this (Fig.3 panels F and G). One would expect that RT is much more affected given the huge effects on HIV-1 capsid and particle production shown in Fig.2 F, G and I.

      The reviewer is right in pointing that we observed a stronger effect during the later stages of the viral cycle, from transcription of viral RNAs (Fig. 2I and Supp. Fig. 2G) to the production of viral particles in the supernatant (Fig. 2D-G), than during the earlier stage of reverse transcription (Fig. 3F, G). Also, it is also possible that we might have missed the peak in ERT/LRT production, which is transient.

      It should be noted that SAMHD1 exhibits both dNTPase (Goldstone et al., 2011) and nuclease (Beloglazova et al., 2013) activities. The ability of SAMHD1 to restrict the virus, through dephosphorylation at T592, is mediated by its RNase activity (Ryoo et al., 2014), and not by the dNTPase activity (Welbourn et al., 2013; White et al., 2013).This could explain why SNAT7 exhibit a stronger impact on viral transcription than on reverse transcription.

      Figure 4; again single cell flow measurements of SAMHD1, pSAMHD1 and p24 /SNAT7 might help to more clearly discriminate effects that are specifically induced upon infection or happen in virally infected cells. Maybe alternatively IF?

      We thank the reviewer for this suggestion. As mentioned under comment #2, flow cytometry analyses are difficult to perform on strongly adherent primary human macrophages.

      With regard to immunofluorescence, there is a technical limitation based on the species in which the antibodies are produced. The antibody that targets the native SNAT7 protein, which is currently being validated in our laboratory, is produced in rabbits. An anti-CAp24 antibody produced in goats can be used. It will then be necessary to co-label the cells with anti SAMHD1 and phospho-SAMHD1produced in mouse. We will try to find options to co-label the cells.

      The wblot shown in panel D does not really reflect the point the authors want to make by the quantification in panels G-I. Primary data (D) suggests that SNAT7 KD reduces HIV-1 production even in the absence of SAMHD1. The quantification rather indicates that SNAT7 KD does not affect HIV-1 production in the absence of SAMHD1. This needs clarification/corroboration by orthogonal approaches.

      We respectfully disagree with the reviewer.

      Figure 4D shows a representative blot of the six donors analysed. As mentioned, the depletion of SNAT7 in the absence of SAMHD1 reduces the production of the viral proteins GagPr55 and CAp24 (see Fig. 4D). This is illustrated by the quantifications (Fig. 4G–I). Following treatment with Vpx, GagPr55 protein expression in SNAT7 KD macrophages is reduced by a factor of 2.6 for siRNA #1 (mean = 1.48, light grey bar) and by a factor of 1.83 for siRNA #2 (mean = 2.13, orange bar), compared to the control (mean = 3.9, pink bar) (Fig. 4G). Similarly, CAp24 protein expression was reduced by a factor of 2.2 for siRNA #1 (mean = 2.05, light grey bar) and by a factor of 1.36 for siRNA #2 (mean = 3.34, orange bar), compared to the control (mean = 4.52, pink bar) (Fig. 4H).

      These differences are therefore consistent between the Western blot and the quantifications. However, they are not significantly different to those observed in cells treated with Vpx and depleted with control siRNA, suggesting that the viral restriction observed in SNAT7 KD cells is primarily due to SAMHD1.

      Figure 5: show SAMHD1 and pSAMHD1 levels upon glutamine supplementation.

      We thank the reviewer for this comment, we will perform the suggested experiment.

      1. I think the discussion is very thin, mainly summarizing the results; but fails to give broader context or critically discuss the limitations and further directions.

      We thank the reviewer for this comment. The discussion will be modified further accordingly.

      Looking at the data as a whole, I think the results support a modest functional importance of SNAT7 for HIV-1 production in macrophages. I acknowledge that the experiments in primary macrophages are prone to high variability in different donors and the authors transparently depicted their data. However clearly, I would advice the authors to tune down the extend in which they claim SNAT7-dependency given this huge variability and the sometimes-borderline statistics. We respectfully disagree with the reviewer.

      The cells used here imply greater variability than a cell line, but are also more relevant.

      Indeed, the effects observed in the late stages of HIV-1 production are:

      • ~80 % decrease in viral transcription compared to the control (Fig. 2I),

      • ~85 % decrease in CAp24 protein expression compared to the control, as quantified by western blot (Fig. 2E), or ~90 % by ELISA measurement (Fig. 2F),

      • a reduction of more than 90 % in the release of infectious particles (Fig. 2G).

      These results were all significant across donors, while SNAT7 depletion was always partial (Fig. 2C, between 31 to 62 % of depletion compared to the control in infected cells).

      Therefore, the data were obtained from a mixture of depleted and non-depleted macrophages. This means that the results may be underestimated.

      Together, our results show that SNAT7 is necessary for HIV-1 production.

      However, reading the comments, we realized that our conclusions regarding reverse transcription were too strong. SNAT7 depletion does not affect viral fusion and reverse transcription. The manuscript was modified accordingly.

      On top, there are a lot of optional experiments I am sure the authors are aware of that should be done at least in the future.

      For instance, how does HIV-1 upregulate SNAT7, is a viral accessory protein involved? What is the mechanism of SNAT7 dependent SAMHD1 phosphorylation? Does SNAT7 (or glutamine) regulate the activity of the SAMHD1 associated kinase / phosphatase) If so, does this impact on other targets of these enzymes? We thank the reviewer for these questions.

      To address the role of accessory viral proteins, we have already performed one experiment infecting hMDM with HIV-1 strains deleted for genes such as Nef, Vpr, Vpu and Vif, and have found no clear effect on SNAT7 protein expression compared to WT strains. As an alternative experiment, we could overexpress individual viral genes, such as Nef or Vpr, in HeLa cells and analyze their impact on SNAT7 expression by Western blot.

      It is also possible that SNAT7 expression and recycling of lysosomal glutamine are modulated by the macrophage intrinsic immunity in response to HIV-1 infection.

      The Thr592 motif of the SAMHD1 protein is phosphorylated by Cyclin A2/CDK1 and type 1 IFN in non-cycling cells, such as MDMs (Cribier et al., 2013). For now, the relationship between SNAT7 and SAMHD1 remains unclear. However, (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This has been added to the discussion to explain the relationship between the 3 partners.

      **Referees cross-commenting** I think the comments from the other referees are reasonable and consistent with my assessment

      Reviewer #1 (Significance (Required)):

      Strength and limitations see above;

      Significance: I think this work is of high interest for virologists working in the field of HIV-1 and infection of myeloid cells. In case SNAT7 (and hence glutamine) indeed regulates the phosphorylation of SAMHD1, there could potentially be broad relevance of this work. However unfortunately, this aspect remains underdeveloped and is also not discussed

      Field of expertise: HIV-1, immunology, cell biology

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      In this report, Herit and colleagues describe the role of a HIV-1 dependency factor that promotes virus replication in macrophages. The authors suggest that the lysosomal membrane-associated SNAT7 glutamine transporter is a HIV dependency factor, that promotes virus replication by enhancing reverse transcription and Gag synthesis. The authors use transient knock-down approaches in primary macrophages to identify that SNAT7 depletion does not impact viral entry but inhibits early reverse transcription which was reversed by exogenous glutamine addition. While reverse transcription enhancement was likely due to selective increase in phosho-SAMHD1 expression, mechanisms by which SNAT7 enhanced viral gene expression were not clearly defined. These are well-controlled studies that pinpoint the role of SNAT7 in the early steps of viral life cycle and highlight the intricate interplay between macrophage metabolism and HIV-1 replication. While the question that is addressed is important, and the hypothesis overall sound, the data presented needs to be strengthened to support the conclusions. There are numerous weaknesses in data interpretation as well.

      1. Figure 1: SNAT7 expression was selectively enhanced upon differentiation of monocytes into macrophages but absent in CD4+ T cells. Though there is a claim of enhancement of SNAT7 expression upon HIV-1 infection of macrophages, RT-qPCR analysis shows the opposite trend (Fig 1E) and SNAT7 protein expression changes are modest. Statistical analysis in Fig. 1H needs to be revisited. The number of replicates vary for the lysates harvested at different day post infection, which might have an impact on the statistical test. To determine if SNAT7 expression enhancement is dependent on establishment of virus infection, as the authors imply, control lysates of virus infections in presence of replication inhibitors should be included.

      We thank the reviewer for this comment. Indeed, there is a modest, but statistically significant increase in SNAT7 protein expression upon HIV-1 infection over time (Fig. 1G, H), without any modulation of SNAT7 gene expression (Fig. 1E). This indicates that the regulation of SNAT7 expression in this context is only at the translation level (i.e. increase of translation or stabilization of the SNAT7 protein).

      As mentioned, Fig. 1H aggregates between 3 to 7 independent experiments on different donors depending on the infection time point. SNAT7 protein expression is increased already at 1 day post-infection and until 8 days. The statistical test used here, i.e. 2 way-ANOVA, compared Mock-infected and HIV-1-infected condition for each time point with the same number of donors. In this figure, the comparison is statistically different only at day 6 of the time course (7 donors). We agree that increasing the number of donors of the other time points could help to improve the statistical difference between control and infection condition.

      We thank the reviewer for the suggestion mentioning the use of replication inhibitors in this experiment. We plan to use inhibitors of reverse transcription (Nevirapin) and integration (Dolutegravir).

      The authors rely exclusively on western blot analysis for HIV-1 Gag expression in cell lysates as a measure of effects of SNAT7 on virus replication. Single cell analysis such as intracellular p24gag analysis by FACS should be included; this will provide a better measure of effects of SNAT7 onHIV-1 infection establishment.

      We respectfully disagree with the reviewer for this question. Indeed, to evaluate the effects of SNAT7 on HIV-1 replication, we measured Gag Pr55 and Cap24 using a Western blot approach (Fig. 2B, D and E), but also assessed the quantity of Cap24 in the supernatants and lysates using an ELISA measurement, the quantity of infectious particles using TZM reporter cells, and total viral transcription or more specifically Gag Pr55 transcription using qPCR (Fig. 2F, G and I and Supp. Fig. 2G).

      Regarding the quantification of CAp24 at the cell single level, please refer to comment #2 under Reviewer #1.

      Knockdown of SNAT7 in MDMs was partial at best; only 30-50% decrease in expression (Fig 2C), but the effects on viral gene expression (Fig. 2I), p24 release and infectious particle production is dramatic (Fig. 2F and G). This discrepancy is not addressed. Does SNAT7 knock-down negatively impact virus particle release? Please note that the representative WB in Fig 2B does not correlate with the quantification in Fig. 2D. There are no p55gag or p24gag bands in SNAT7#1 siRNA condition (Fig. 2B)? Data could also be rearranged to follow the logical sequence of virus replication cycle (viral RNa expression followed by Gag expression, and then release).

      We thank the reviewer for this comment. Our samples are indeed a mixture of SNAT7-depleted and non-depleted macrophages and RNA interference in these cells often leads to a decrease of 50 % of the protein expression.

      To determine whether SNAT7 is involved in the release of particles, we quantified Cap24 in cell lysates and in the cell culture medium separately, and normalized the results to the total protein content. The absence of SNAT7 reduced the amount of Cap24 measured by ELISA in both samples to the same extent, showing that there is no storage of Cap24-positive viral particles inside the infected macrophages. These data were initially pooled in one graph (Fig. 2F), but separate graphs are now provided in new Supp. Fig. 2 E, F.

      Regarding the western blot shown in Fig. 2B, please refer to comment #5 under Reviewer #1.

      In the new version of the manuscript, we arranged the figures and placed the later stages of the viral cycle in Fig. 2 and the earlier stages, such as fusion, reverse transcription and transcription, in Fig. 3.

      Data interpretation would be greatly improved by including infection controls (RT or integrase inhibitors) to confirm that measurements of viral RNA and Gag are indeed modulated by SNAT7 expression.

      We thank the reviewer for this suggestion to include inhibitors of viral replication as controls. In our experiments, cells were Mock-infected in parallel as a negative control of viral detection. We provide the results in the new version of the manuscript to show that (i) there is no detection of viral or Gag RNA in the absence of the virus, (ii) the expression of viral genes measured in HIV-1-infected SNAT7-depleted cells is not different from Mock-infected cells, indicating almost complete inhibition of viral transcription (Fig. 3H and Supp. Fig. 3B), also confirmed at the protein level (Fig. 2B, D-F).

      Figure 3: Decrease in SNAT7 expression in macrophages resulted in lower levels of early reverse transcripts. But surprisingly, LRT levels were not as affected by decreases in SNAT7 expression. The authors go on to suggest that decreases in early RT are due to loss of phospho-SAMHD1 and increases in catalytically active form of SAMHD1. Mechanistically this does not make sense: LRT should be similarly affected by increase in catalytically active SAMHD1. dNTP concentrations should be measured to determine if the rescue of RT is dependent on SAMHD1 dNTPase activity.

      We thank the reviewer for this comment. LRT concentrations are very low in human macrophages and more challenging to detect than ERT concentrations. This might explain why the differences observed between the SNAT7-depleted and control conditions appear less pronounced for LRT than for ERT.

      Furthermore, we cannot rule out the possibility that SNAT7 has a cumulative effect throughout the viral cycle. While reverse transcription remains statistically unaltered, and despite the reduced levels of ERT and LRT in SNAT7-depleted macrophages (Fig. 3 F, G), there is a significant impact on the transcription of viral RNAs (Fig. 2I) and Gag (Supp. Fig. 2G). This step may also be altered by the ribonuclease activity of SAMHD1 (Beloglazova et al., 2013; Ryoo et al., 2014).

      Finally, with the help of Dr Baek Kim in Atlanta, we attempted to quantify dNTP concentrations in our human macrophages. Unfortunately, it was not possible to draw any conclusions, as the concentrations of dNTPs extracted from our cells were far too low.

      Furthermore, it should be noted that SAMHD1 viral restriction through its phosphorylation at T592 is not correlated with its dNTPase activity (Welbourn et al., 2013; White et al., 2013), but with its ribonuclease activity (Beloglazova et al., 2013; Ryoo et al., 2014). This is supporting why SNAT7, by modulating the ribonuclease activity of SAMHD1, could have a greater effect on viral transcription than on reverse transcription.

      There is lack of consistency in the data: p24 release upon SNAT7 depletion is highly variable. While there is a dramatic >90-95% decrease in p24 release (Fig. 2G), the effects are much more moderate in Fig. 4H (50-60% attenuation), even though siRNA-mediated depletion was similar across the data sets. The authors should comment on the variability in their findings.

      We thank the reviewer for this comment, but believe that Figure 2E rather than Figure 2G is to be mentioned regarding the quantification of CAp24 by Western blot and to be compared with Figure 4H.

      In Fig. 2E, we observed an average reduction of 85 % in CAp24 expression normalized to Clathrin HC expression across different donors for both siRNAs targeting SNAT7. For Fig. 4H, there was a 73 % reduction in CAp24 levels for siRNA #1 and a 56 % reduction for siRNA #2. In addition, it should be noted that the reduction in Gag levels is greater in Fig. 4G (between 77 % and 83 %) than in Fig. 2D (between 55 % and 72 %).

      Therefore, there is some variation in the results obtained with the different donors, which could be explained by variations in Gag cleavage among donors, but this does not impact the conclusions for both figures.

      SNAT7 is postulated to affect 2 steps in the virus life cycle: reverse transcription and viral transcription. But Vpx-mediated SAMHD1 degradation reversed both. Its not clear to me as to how SAMHD1 degradation impacts the role of SNAT7 in viral transcription. No explanation is provided.

      We thank the reviewer for this comment. As suggested, we will perform experiments to assess the impact of Vpx-mediated SAMHD1 degradation on viral transcription.

      Exogenous addition of glutamine only partially restored Gag synthesis and p24 release, which could be attributed to increased cytoplasmic levels and viral protein synthesis. What about effects on reverse transcription and viral gene expression?

      We thank the reviewer for this comment. We will perform the suggested experiments to assess the impact of glutamine supplementation on viral transcription.

      Reviewer #2 (Significance (Required)):

      This is a novel finding, as there are limited number of studies on amino acid transporters and HIV-1 replication enhancement in macrophages. Most of the previous work has focused on CD4 T cells. These studies on SNAT7 and HIV-1 infection establishment in macrophages might better inform the influences of macrophage metabolism on HIV-1 persistence and inflammatory responses.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      This study investigates the role of the lysosomal glutamine transporter SLC38A7/SNAT7 in HIV‑1 replication in primary human macrophages. The authors demonstrate that SNAT7 is highly expressed in macrophages and upregulated upon HIV‑1 infection. They show that SNAT7 depletion inhibits HIV‑1 production at the reverse transcription step without affecting viral fusion or global cellular translation/transcription. Mechanistically, SNAT7 knockdown reduces the inhibitory phosphorylation of SAMHD1 at T592, and degradation of SAMHD1 by Vpx fully rescues viral replication. Extracellular glutamine supplementation partially restores HIV‑1 production in SNAT7‑deficient cells. Overall, the authors report interesting observations; however, the mechanistic investigation remains preliminary, raising concerns about whether the data fully support all the conclusions drawn. Major Concerns: 1. The mechanistic depth is insufficient. The authors do not elucidate how glutamine regulates SAMHD1 T592 phosphorylation, whether through metabolite‑mediated control of kinases/phosphatases or via indirect effects.

      We thank the reviewer for this comment. It is worth noting that (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity using drugs decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This is now further discussed in the discussion section of the manuscript.

      The authors do not measure intracellular dNTP levels upon SNAT7 knockdown, which is the key functional substrate of SAMHD1. They also do not directly demonstrate that glutamine supplementation restores dNTP pools.

      We thank the reviewer for this comment. Please, refer to comment #5 under Reviewer #2.

      Extracellular glutamine only partially rescues viral production, implying the existence of transport‑independent functions of SNAT7 or additional pathways. This important observation is not discussed.

      We thank the reviewer for this comment. The discussion has been modified accordingly.

      It is suggested that the key findings be validated in immortalized THP‑1 cells differentiated into macrophage‑like cells by PMA.

      We thank the reviewer for this suggestion but don’t really understand why this would strengthen our conclusions. Indeed, despite the known variability between donors and technical limitations to transduce cells, we chose human blood monocyte-derived macrophages as a relevant non-transformed model for HIV-1 infection of macrophages. They also represent to some extent the human diversity.

      The Discussion section should be expanded to include the potential translational implications and limitations of the present study.

      We thank the reviewer for this comment. The discussion points to some elements of potential translation and limitations of the study.

      Reviewer #3 (Significance (Required)):

      General assessment: This study identifies the lysosomal glutamine transporter SLC38A7/SNAT7 as a novel host dependency factor for HIV‑1 replication in primary human macrophages. The major strengths include the use of physiologically relevant primary macrophage models, a well-organized experimental pipeline from expression profiling to functional validation, and the establishment of a link between SNAT7, glutamine metabolism, and the HIV restriction factor SAMHD1.

      Advance: It extends current understanding of HIV‑1 host dependency factors and immunometabolism by revealing a compartment‑specific metabolic pathway that supports viral reverse transcription.

      Audience:This work will primarily interest specialized researchers in HIV‑1 biology, host-virus interactions, restriction factors, and antiviral innate immunity.

      Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      This study from the Niedergang lab establishes SNAT7 as a host-dependency factor in human macrophages that supports HIV-1 replication. They show a modest increase in SNAT7 levels HIV-1 infected macrophages and suggest that SNAT7 levels are transiently increased. Employing siRNA against SNAT7 they show reduction in HIV-1 protein levels and viral RNAs and claim that there is a block of reverse transcription in SNAT7 KD cells. Focusing on a known HIV-1 restriction factor in macrophages, SAMHD1, they interconnect the SNAT7 depletion with a reduction in phosphorylated, i.e. catalytical inactive SAMHD1 arguing that SNAT7 regulates the phosphorylation and thereby antiviral activity of SAMHD1. Since SNAT7 is a glutamine transporter that provides this AA from lysosomes, they lastly supplement glutamine and this somehow rescues the reduction of HIV-1 production in SNAT7 KD cells.

      Major comments:

      The strength of this manuscript is the clear focus on primary human macrophages that are HIV-1 infected and the interconnection of HIV-1 replication to the SNAT7 siRNA KD experiments in combination with SAMHD1 depletion and lastly glutamine supplementation. This establishes a stringent and coherent story line. The effects reported are modest; high variability is not a problem since using primary hMDM this is expected and can be addressed by testing several donors and applying stringent statistics.

      1. Having said so, I realize that while they give information on the statistical test used, i.e. one-way ANOVA they miss to explain the post-test used to assess significance (i.e. Bonferroni, Fishers LSD, whatsoever). Please add this information.

      We thank the reviewer for this comment. The figure legends have been updated to include more details of all the statistical tests used.

      1. Another issue that might underestimate the effects of HIV-1 infection on SNAT7 levels and vice versa of SNAT7 KD on HIV-1 replication is the non-single cell approach employed, i.e. WBlots. I assume that HIV-1 infection rates in macrophages are not super high, usually not exceeding 20-30%. So indeed the effects the authors observe could be much higher, when checking at the single cell level. I do not know about the SNAT7 ab, but all the other reagents should work via flow cytometry and could hence improve the readout a lot.

      We agree with the reviewer and indeed, in previous studies on HIV-1 infection of human macrophages performed in the lab, we observed via immunofluorescence that the proportion of infected cells ranged from 20 to 40 %. At the time of submission, we did not have the possibility to label the native SNAT7 protein by immunofluorescence, as the commercial antibody used only works for western blotting.

      In the meantime, we have been validating a new antibody (Proteintech) targeting SNAT7 for immunofluorescence. If this is confirmed, we will be able to detect and quantify HIV-1 p24 by immunofluorescence in SNAT7-depleted human macrophages and control cells, thus confirming our results in single-cell analysis.

      Flow cytometry analyses are difficult to perform on primary human macrophages because these cells are highly adherent and must be detached first. The process induces significant cell death and damage. This is why we would prefer to carry out these analyses using immunofluorescence and microscopy on adhered cells. This option will be undoubtedly pursued.

      1. Furthermore the authors never commented about a dose-response effect in terms of HIV-1 infection levels. There is a MOI dependency described for Suppl.Fig.1 C-F, unfortunately the data is missing in the manuscript.

      We apologize for this omission. The figures showing the increase in SNAT7 protein expression following HIV-1 infection at MOIs ranging from 0.05 to 0.5 were added to the new version of the manuscript (Supp. Fig. 1 C-F).

      1. Figure1: specify circulating T lymphocytes. I would expect to see levels of SNAT7 in PHA or CD3/CD28 activated lymphocytes versus resting T cells and a time course of SNAT7 levels upon activation. I think even though SNAT7 levels in T cells might be low, they could also be increased by HIV-1 infection and it is essential that the authors test for this. If not, the result is a valid negative control. For this they should employ HIV-1 primary strains with a tropism for T cells, or at least lab-adapted HIV-1 NL4-3

      We thank the reviewer for this comment. Circulating T lymphocytes isolated from the blood of healthy donors are now referred to resting lymphocytes in the new version of the manuscript, as opposed to activated T lymphocytes stimulated with IL2 and PHA-P for several days (Fig. 1 A-C).

      The expression levels of SNAT7, both at the gene and protein levels, are lower in resting or IL2/PHA-P-activated T cells than in macrophages from the same donors. As suggested, we will perform a kinetic of T-cell activation upon HIV-1 infection to investigate how SNAT7 expression varies in these conditions.

      1. Figure 2 again single cell measurements could reveal much more pronounced effects; it is a bit counterintuitive that siRNA #2 is more efficient in SNAT7 KD but has higher levels of HIV-1 replication in terms of Gag levels. I assume when looking at the stats it is always a comparison to the Ctl treated cells (C-G), but this is not entirely clear. Unify labeling as compared to the stats in Fig.2 I (this also applies for all the other figs).

      We thank the reviewer for this comment. Fig. 2B indeed shows one of the different donors analyzed. However, protein quantification across six different donors shows that SNAT7 is more depleted with siRNA #2 (Fig. 2C), and that Gag Pr55 protein levels are consequently more reduced, than with siRNA #1 (Fig. 2D).

      We use GraphPad Prism software to perform statistical analysis. Depending on the test used, the software automatically plots the comparison bar and displays the p-value above it. We changed the representation of statistics as suggested.

      Figure 3: It is a bit odd that they finally conclude on RT as essential step that is reduced in the absence of SNAT7 and then they fail to provide statistical significance for this (Fig.3 panels F and G). One would expect that RT is much more affected given the huge effects on HIV-1 capsid and particle production shown in Fig.2 F, G and I.

      The reviewer is right in pointing that we observed a stronger effect during the later stages of the viral cycle, from transcription of viral RNAs (Fig. 2I and Supp. Fig. 2G) to the production of viral particles in the supernatant (Fig. 2D-G), than during the earlier stage of reverse transcription (Fig. 3F, G). Also, it is also possible that we might have missed the peak in ERT/LRT production, which is transient.

      It should be noted that SAMHD1 exhibits both dNTPase (Goldstone et al., 2011) and nuclease (Beloglazova et al., 2013) activities. The ability of SAMHD1 to restrict the virus, through dephosphorylation at T592, is mediated by its RNase activity (Ryoo et al., 2014), and not by the dNTPase activity (Welbourn et al., 2013; White et al., 2013).This could explain why SNAT7 exhibit a stronger impact on viral transcription than on reverse transcription.

      Figure 4; again single cell flow measurements of SAMHD1, pSAMHD1 and p24 /SNAT7 might help to more clearly discriminate effects that are specifically induced upon infection or happen in virally infected cells. Maybe alternatively IF?

      We thank the reviewer for this suggestion. As mentioned under comment #2, flow cytometry analyses are difficult to perform on strongly adherent primary human macrophages.

      With regard to immunofluorescence, there is a technical limitation based on the species in which the antibodies are produced. The antibody that targets the native SNAT7 protein, which is currently being validated in our laboratory, is produced in rabbits. An anti-CAp24 antibody produced in goats can be used. It will then be necessary to co-label the cells with anti SAMHD1 and phospho-SAMHD1produced in mouse. We will try to find options to co-label the cells.

      The wblot shown in panel D does not really reflect the point the authors want to make by the quantification in panels G-I. Primary data (D) suggests that SNAT7 KD reduces HIV-1 production even in the absence of SAMHD1. The quantification rather indicates that SNAT7 KD does not affect HIV-1 production in the absence of SAMHD1. This needs clarification/corroboration by orthogonal approaches.

      We respectfully disagree with the reviewer.

      Figure 4D shows a representative blot of the six donors analysed. As mentioned, the depletion of SNAT7 in the absence of SAMHD1 reduces the production of the viral proteins GagPr55 and CAp24 (see Fig. 4D). This is illustrated by the quantifications (Fig. 4G–I). Following treatment with Vpx, GagPr55 protein expression in SNAT7 KD macrophages is reduced by a factor of 2.6 for siRNA #1 (mean = 1.48, light grey bar) and by a factor of 1.83 for siRNA #2 (mean = 2.13, orange bar), compared to the control (mean = 3.9, pink bar) (Fig. 4G). Similarly, CAp24 protein expression was reduced by a factor of 2.2 for siRNA #1 (mean = 2.05, light grey bar) and by a factor of 1.36 for siRNA #2 (mean = 3.34, orange bar), compared to the control (mean = 4.52, pink bar) (Fig. 4H).

      These differences are therefore consistent between the Western blot and the quantifications. However, they are not significantly different to those observed in cells treated with Vpx and depleted with control siRNA, suggesting that the viral restriction observed in SNAT7 KD cells is primarily due to SAMHD1.

      Figure 5: show SAMHD1 and pSAMHD1 levels upon glutamine supplementation.

      We thank the reviewer for this comment, we will perform the suggested experiment.

      1. I think the discussion is very thin, mainly summarizing the results; but fails to give broader context or critically discuss the limitations and further directions.

      We thank the reviewer for this comment. The discussion will be modified further accordingly.

      Looking at the data as a whole, I think the results support a modest functional importance of SNAT7 for HIV-1 production in macrophages. I acknowledge that the experiments in primary macrophages are prone to high variability in different donors and the authors transparently depicted their data. However clearly, I would advice the authors to tune down the extend in which they claim SNAT7-dependency given this huge variability and the sometimes-borderline statistics. We respectfully disagree with the reviewer.

      The cells used here imply greater variability than a cell line, but are also more relevant.

      Indeed, the effects observed in the late stages of HIV-1 production are:

      • ~80 % decrease in viral transcription compared to the control (Fig. 2I),

      • ~85 % decrease in CAp24 protein expression compared to the control, as quantified by western blot (Fig. 2E), or ~90 % by ELISA measurement (Fig. 2F),

      • a reduction of more than 90 % in the release of infectious particles (Fig. 2G).

      These results were all significant across donors, while SNAT7 depletion was always partial (Fig. 2C, between 31 to 62 % of depletion compared to the control in infected cells).

      Therefore, the data were obtained from a mixture of depleted and non-depleted macrophages. This means that the results may be underestimated.

      Together, our results show that SNAT7 is necessary for HIV-1 production.

      However, reading the comments, we realized that our conclusions regarding reverse transcription were too strong. SNAT7 depletion does not affect viral fusion and reverse transcription. The manuscript was modified accordingly.

      On top, there are a lot of optional experiments I am sure the authors are aware of that should be done at least in the future.

      For instance, how does HIV-1 upregulate SNAT7, is a viral accessory protein involved? What is the mechanism of SNAT7 dependent SAMHD1 phosphorylation? Does SNAT7 (or glutamine) regulate the activity of the SAMHD1 associated kinase / phosphatase) If so, does this impact on other targets of these enzymes? We thank the reviewer for these questions.

      To address the role of accessory viral proteins, we have already performed one experiment infecting hMDM with HIV-1 strains deleted for genes such as Nef, Vpr, Vpu and Vif, and have found no clear effect on SNAT7 protein expression compared to WT strains. As an alternative experiment, we could overexpress individual viral genes, such as Nef or Vpr, in HeLa cells and analyze their impact on SNAT7 expression by Western blot.

      It is also possible that SNAT7 expression and recycling of lysosomal glutamine are modulated by the macrophage intrinsic immunity in response to HIV-1 infection.

      The Thr592 motif of the SAMHD1 protein is phosphorylated by Cyclin A2/CDK1 and type 1 IFN in non-cycling cells, such as MDMs (Cribier et al., 2013). For now, the relationship between SNAT7 and SAMHD1 remains unclear. However, (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This has been added to the discussion to explain the relationship between the 3 partners.

      **Referees cross-commenting** I think the comments from the other referees are reasonable and consistent with my assessment

      Reviewer #1 (Significance (Required)):

      Strength and limitations see above;

      Significance: I think this work is of high interest for virologists working in the field of HIV-1 and infection of myeloid cells. In case SNAT7 (and hence glutamine) indeed regulates the phosphorylation of SAMHD1, there could potentially be broad relevance of this work. However unfortunately, this aspect remains underdeveloped and is also not discussed

      Field of expertise: HIV-1, immunology, cell biology

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      In this report, Herit and colleagues describe the role of a HIV-1 dependency factor that promotes virus replication in macrophages. The authors suggest that the lysosomal membrane-associated SNAT7 glutamine transporter is a HIV dependency factor, that promotes virus replication by enhancing reverse transcription and Gag synthesis. The authors use transient knock-down approaches in primary macrophages to identify that SNAT7 depletion does not impact viral entry but inhibits early reverse transcription which was reversed by exogenous glutamine addition. While reverse transcription enhancement was likely due to selective increase in phosho-SAMHD1 expression, mechanisms by which SNAT7 enhanced viral gene expression were not clearly defined. These are well-controlled studies that pinpoint the role of SNAT7 in the early steps of viral life cycle and highlight the intricate interplay between macrophage metabolism and HIV-1 replication. While the question that is addressed is important, and the hypothesis overall sound, the data presented needs to be strengthened to support the conclusions. There are numerous weaknesses in data interpretation as well.

      1. Figure 1: SNAT7 expression was selectively enhanced upon differentiation of monocytes into macrophages but absent in CD4+ T cells. Though there is a claim of enhancement of SNAT7 expression upon HIV-1 infection of macrophages, RT-qPCR analysis shows the opposite trend (Fig 1E) and SNAT7 protein expression changes are modest. Statistical analysis in Fig. 1H needs to be revisited. The number of replicates vary for the lysates harvested at different day post infection, which might have an impact on the statistical test. To determine if SNAT7 expression enhancement is dependent on establishment of virus infection, as the authors imply, control lysates of virus infections in presence of replication inhibitors should be included.

      We thank the reviewer for this comment. Indeed, there is a modest, but statistically significant increase in SNAT7 protein expression upon HIV-1 infection over time (Fig. 1G, H), without any modulation of SNAT7 gene expression (Fig. 1E). This indicates that the regulation of SNAT7 expression in this context is only at the translation level (i.e. increase of translation or stabilization of the SNAT7 protein).

      As mentioned, Fig. 1H aggregates between 3 to 7 independent experiments on different donors depending on the infection time point. SNAT7 protein expression is increased already at 1 day post-infection and until 8 days. The statistical test used here, i.e. 2 way-ANOVA, compared Mock-infected and HIV-1-infected condition for each time point with the same number of donors. In this figure, the comparison is statistically different only at day 6 of the time course (7 donors). We agree that increasing the number of donors of the other time points could help to improve the statistical difference between control and infection condition.

      We thank the reviewer for the suggestion mentioning the use of replication inhibitors in this experiment. We plan to use inhibitors of reverse transcription (Nevirapin) and integration (Dolutegravir).

      The authors rely exclusively on western blot analysis for HIV-1 Gag expression in cell lysates as a measure of effects of SNAT7 on virus replication. Single cell analysis such as intracellular p24gag analysis by FACS should be included; this will provide a better measure of effects of SNAT7 onHIV-1 infection establishment.

      We respectfully disagree with the reviewer for this question. Indeed, to evaluate the effects of SNAT7 on HIV-1 replication, we measured Gag Pr55 and Cap24 using a Western blot approach (Fig. 2B, D and E), but also assessed the quantity of Cap24 in the supernatants and lysates using an ELISA measurement, the quantity of infectious particles using TZM reporter cells, and total viral transcription or more specifically Gag Pr55 transcription using qPCR (Fig. 2F, G and I and Supp. Fig. 2G).

      Regarding the quantification of CAp24 at the cell single level, please refer to comment #2 under Reviewer #1.

      Knockdown of SNAT7 in MDMs was partial at best; only 30-50% decrease in expression (Fig 2C), but the effects on viral gene expression (Fig. 2I), p24 release and infectious particle production is dramatic (Fig. 2F and G). This discrepancy is not addressed. Does SNAT7 knock-down negatively impact virus particle release? Please note that the representative WB in Fig 2B does not correlate with the quantification in Fig. 2D. There are no p55gag or p24gag bands in SNAT7#1 siRNA condition (Fig. 2B)? Data could also be rearranged to follow the logical sequence of virus replication cycle (viral RNa expression followed by Gag expression, and then release).

      We thank the reviewer for this comment. Our samples are indeed a mixture of SNAT7-depleted and non-depleted macrophages and RNA interference in these cells often leads to a decrease of 50 % of the protein expression.

      To determine whether SNAT7 is involved in the release of particles, we quantified Cap24 in cell lysates and in the cell culture medium separately, and normalized the results to the total protein content. The absence of SNAT7 reduced the amount of Cap24 measured by ELISA in both samples to the same extent, showing that there is no storage of Cap24-positive viral particles inside the infected macrophages. These data were initially pooled in one graph (Fig. 2F), but separate graphs are now provided in new Supp. Fig. 2 E, F.

      Regarding the western blot shown in Fig. 2B, please refer to comment #5 under Reviewer #1.

      In the new version of the manuscript, we arranged the figures and placed the later stages of the viral cycle in Fig. 2 and the earlier stages, such as fusion, reverse transcription and transcription, in Fig. 3.

      Data interpretation would be greatly improved by including infection controls (RT or integrase inhibitors) to confirm that measurements of viral RNA and Gag are indeed modulated by SNAT7 expression.

      We thank the reviewer for this suggestion to include inhibitors of viral replication as controls. In our experiments, cells were Mock-infected in parallel as a negative control of viral detection. We provide the results in the new version of the manuscript to show that (i) there is no detection of viral or Gag RNA in the absence of the virus, (ii) the expression of viral genes measured in HIV-1-infected SNAT7-depleted cells is not different from Mock-infected cells, indicating almost complete inhibition of viral transcription (Fig. 3H and Supp. Fig. 3B), also confirmed at the protein level (Fig. 2B, D-F).

      Figure 3: Decrease in SNAT7 expression in macrophages resulted in lower levels of early reverse transcripts. But surprisingly, LRT levels were not as affected by decreases in SNAT7 expression. The authors go on to suggest that decreases in early RT are due to loss of phospho-SAMHD1 and increases in catalytically active form of SAMHD1. Mechanistically this does not make sense: LRT should be similarly affected by increase in catalytically active SAMHD1. dNTP concentrations should be measured to determine if the rescue of RT is dependent on SAMHD1 dNTPase activity.

      We thank the reviewer for this comment. LRT concentrations are very low in human macrophages and more challenging to detect than ERT concentrations. This might explain why the differences observed between the SNAT7-depleted and control conditions appear less pronounced for LRT than for ERT.

      Furthermore, we cannot rule out the possibility that SNAT7 has a cumulative effect throughout the viral cycle. While reverse transcription remains statistically unaltered, and despite the reduced levels of ERT and LRT in SNAT7-depleted macrophages (Fig. 3 F, G), there is a significant impact on the transcription of viral RNAs (Fig. 2I) and Gag (Supp. Fig. 2G). This step may also be altered by the ribonuclease activity of SAMHD1 (Beloglazova et al., 2013; Ryoo et al., 2014).

      Finally, with the help of Dr Baek Kim in Atlanta, we attempted to quantify dNTP concentrations in our human macrophages. Unfortunately, it was not possible to draw any conclusions, as the concentrations of dNTPs extracted from our cells were far too low.

      Furthermore, it should be noted that SAMHD1 viral restriction through its phosphorylation at T592 is not correlated with its dNTPase activity (Welbourn et al., 2013; White et al., 2013), but with its ribonuclease activity (Beloglazova et al., 2013; Ryoo et al., 2014). This is supporting why SNAT7, by modulating the ribonuclease activity of SAMHD1, could have a greater effect on viral transcription than on reverse transcription.

      There is lack of consistency in the data: p24 release upon SNAT7 depletion is highly variable. While there is a dramatic >90-95% decrease in p24 release (Fig. 2G), the effects are much more moderate in Fig. 4H (50-60% attenuation), even though siRNA-mediated depletion was similar across the data sets. The authors should comment on the variability in their findings.

      We thank the reviewer for this comment, but believe that Figure 2E rather than Figure 2G is to be mentioned regarding the quantification of CAp24 by Western blot and to be compared with Figure 4H.

      In Fig. 2E, we observed an average reduction of 85 % in CAp24 expression normalized to Clathrin HC expression across different donors for both siRNAs targeting SNAT7. For Fig. 4H, there was a 73 % reduction in CAp24 levels for siRNA #1 and a 56 % reduction for siRNA #2. In addition, it should be noted that the reduction in Gag levels is greater in Fig. 4G (between 77 % and 83 %) than in Fig. 2D (between 55 % and 72 %).

      Therefore, there is some variation in the results obtained with the different donors, which could be explained by variations in Gag cleavage among donors, but this does not impact the conclusions for both figures.

      SNAT7 is postulated to affect 2 steps in the virus life cycle: reverse transcription and viral transcription. But Vpx-mediated SAMHD1 degradation reversed both. Its not clear to me as to how SAMHD1 degradation impacts the role of SNAT7 in viral transcription. No explanation is provided.

      We thank the reviewer for this comment. As suggested, we will perform experiments to assess the impact of Vpx-mediated SAMHD1 degradation on viral transcription.

      Exogenous addition of glutamine only partially restored Gag synthesis and p24 release, which could be attributed to increased cytoplasmic levels and viral protein synthesis. What about effects on reverse transcription and viral gene expression?

      We thank the reviewer for this comment. We will perform the suggested experiments to assess the impact of glutamine supplementation on viral transcription.

      Reviewer #2 (Significance (Required)):

      This is a novel finding, as there are limited number of studies on amino acid transporters and HIV-1 replication enhancement in macrophages. Most of the previous work has focused on CD4 T cells. These studies on SNAT7 and HIV-1 infection establishment in macrophages might better inform the influences of macrophage metabolism on HIV-1 persistence and inflammatory responses.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      This study investigates the role of the lysosomal glutamine transporter SLC38A7/SNAT7 in HIV‑1 replication in primary human macrophages. The authors demonstrate that SNAT7 is highly expressed in macrophages and upregulated upon HIV‑1 infection. They show that SNAT7 depletion inhibits HIV‑1 production at the reverse transcription step without affecting viral fusion or global cellular translation/transcription. Mechanistically, SNAT7 knockdown reduces the inhibitory phosphorylation of SAMHD1 at T592, and degradation of SAMHD1 by Vpx fully rescues viral replication. Extracellular glutamine supplementation partially restores HIV‑1 production in SNAT7‑deficient cells. Overall, the authors report interesting observations; however, the mechanistic investigation remains preliminary, raising concerns about whether the data fully support all the conclusions drawn. Major Concerns: 1. The mechanistic depth is insufficient. The authors do not elucidate how glutamine regulates SAMHD1 T592 phosphorylation, whether through metabolite‑mediated control of kinases/phosphatases or via indirect effects.

      We thank the reviewer for this comment. It is worth noting that (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity using drugs decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This is now further discussed in the discussion section of the manuscript.

      The authors do not measure intracellular dNTP levels upon SNAT7 knockdown, which is the key functional substrate of SAMHD1. They also do not directly demonstrate that glutamine supplementation restores dNTP pools.

      We thank the reviewer for this comment. Please, refer to comment #5 under Reviewer #2.

      Extracellular glutamine only partially rescues viral production, implying the existence of transport‑independent functions of SNAT7 or additional pathways. This important observation is not discussed.

      We thank the reviewer for this comment. The discussion has been modified accordingly.

      It is suggested that the key findings be validated in immortalized THP‑1 cells differentiated into macrophage‑like cells by PMA.

      We thank the reviewer for this suggestion but don’t really understand why this would strengthen our conclusions. Indeed, despite the known variability between donors and technical limitations to transduce cells, we chose human blood monocyte-derived macrophages as a relevant non-transformed model for HIV-1 infection of macrophages. They also represent to some extent the human diversity.

      The Discussion section should be expanded to include the potential translational implications and limitations of the present study.

      We thank the reviewer for this comment. The discussion points to some elements of potential translation and limitations of the study.

      Reviewer #3 (Significance (Required)):

      General assessment: This study identifies the lysosomal glutamine transporter SLC38A7/SNAT7 as a novel host dependency factor for HIV‑1 replication in primary human macrophages. The major strengths include the use of physiologically relevant primary macrophage models, a well-organized experimental pipeline from expression profiling to functional validation, and the establishment of a link between SNAT7, glutamine metabolism, and the HIV restriction factor SAMHD1.

      Advance: It extends current understanding of HIV‑1 host dependency factors and immunometabolism by revealing a compartment‑specific metabolic pathway that supports viral reverse transcription.

      Audience:This work will primarily interest specialized researchers in HIV‑1 biology, host-virus interactions, restriction factors, and antiviral innate immunity.

      2.15.1.0 Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      This study from the Niedergang lab establishes SNAT7 as a host-dependency factor in human macrophages that supports HIV-1 replication. They show a modest increase in SNAT7 levels HIV-1 infected macrophages and suggest that SNAT7 levels are transiently increased. Employing siRNA against SNAT7 they show reduction in HIV-1 protein levels and viral RNAs and claim that there is a block of reverse transcription in SNAT7 KD cells. Focusing on a known HIV-1 restriction factor in macrophages, SAMHD1, they interconnect the SNAT7 depletion with a reduction in phosphorylated, i.e. catalytical inactive SAMHD1 arguing that SNAT7 regulates the phosphorylation and thereby antiviral activity of SAMHD1. Since SNAT7 is a glutamine transporter that provides this AA from lysosomes, they lastly supplement glutamine and this somehow rescues the reduction of HIV-1 production in SNAT7 KD cells.

      Major comments:

      The strength of this manuscript is the clear focus on primary human macrophages that are HIV-1 infected and the interconnection of HIV-1 replication to the SNAT7 siRNA KD experiments in combination with SAMHD1 depletion and lastly glutamine supplementation. This establishes a stringent and coherent story line. The effects reported are modest; high variability is not a problem since using primary hMDM this is expected and can be addressed by testing several donors and applying stringent statistics.

      1. Having said so, I realize that while they give information on the statistical test used, i.e. one-way ANOVA they miss to explain the post-test used to assess significance (i.e. Bonferroni, Fishers LSD, whatsoever). Please add this information.

      We thank the reviewer for this comment. The figure legends have been updated to include more details of all the statistical tests used.

      1. Another issue that might underestimate the effects of HIV-1 infection on SNAT7 levels and vice versa of SNAT7 KD on HIV-1 replication is the non-single cell approach employed, i.e. WBlots. I assume that HIV-1 infection rates in macrophages are not super high, usually not exceeding 20-30%. So indeed the effects the authors observe could be much higher, when checking at the single cell level. I do not know about the SNAT7 ab, but all the other reagents should work via flow cytometry and could hence improve the readout a lot.

      We agree with the reviewer and indeed, in previous studies on HIV-1 infection of human macrophages performed in the lab, we observed via immunofluorescence that the proportion of infected cells ranged from 20 to 40 %. At the time of submission, we did not have the possibility to label the native SNAT7 protein by immunofluorescence, as the commercial antibody used only works for western blotting.

      In the meantime, we have been validating a new antibody (Proteintech) targeting SNAT7 for immunofluorescence. If this is confirmed, we will be able to detect and quantify HIV-1 p24 by immunofluorescence in SNAT7-depleted human macrophages and control cells, thus confirming our results in single-cell analysis.

      Flow cytometry analyses are difficult to perform on primary human macrophages because these cells are highly adherent and must be detached first. The process induces significant cell death and damage. This is why we would prefer to carry out these analyses using immunofluorescence and microscopy on adhered cells. This option will be undoubtedly pursued.

      1. Furthermore the authors never commented about a dose-response effect in terms of HIV-1 infection levels. There is a MOI dependency described for Suppl.Fig.1 C-F, unfortunately the data is missing in the manuscript.

      We apologize for this omission. The figures showing the increase in SNAT7 protein expression following HIV-1 infection at MOIs ranging from 0.05 to 0.5 were added to the new version of the manuscript (Supp. Fig. 1 C-F).

      1. Figure1: specify circulating T lymphocytes. I would expect to see levels of SNAT7 in PHA or CD3/CD28 activated lymphocytes versus resting T cells and a time course of SNAT7 levels upon activation. I think even though SNAT7 levels in T cells might be low, they could also be increased by HIV-1 infection and it is essential that the authors test for this. If not, the result is a valid negative control. For this they should employ HIV-1 primary strains with a tropism for T cells, or at least lab-adapted HIV-1 NL4-3

      We thank the reviewer for this comment. Circulating T lymphocytes isolated from the blood of healthy donors are now referred to resting lymphocytes in the new version of the manuscript, as opposed to activated T lymphocytes stimulated with IL2 and PHA-P for several days (Fig. 1 A-C).

      The expression levels of SNAT7, both at the gene and protein levels, are lower in resting or IL2/PHA-P-activated T cells than in macrophages from the same donors. As suggested, we will perform a kinetic of T-cell activation upon HIV-1 infection to investigate how SNAT7 expression varies in these conditions.

      1. Figure 2 again single cell measurements could reveal much more pronounced effects; it is a bit counterintuitive that siRNA #2 is more efficient in SNAT7 KD but has higher levels of HIV-1 replication in terms of Gag levels. I assume when looking at the stats it is always a comparison to the Ctl treated cells (C-G), but this is not entirely clear. Unify labeling as compared to the stats in Fig.2 I (this also applies for all the other figs).

      We thank the reviewer for this comment. Fig. 2B indeed shows one of the different donors analyzed. However, protein quantification across six different donors shows that SNAT7 is more depleted with siRNA #2 (Fig. 2C), and that Gag Pr55 protein levels are consequently more reduced, than with siRNA #1 (Fig. 2D).

      We use GraphPad Prism software to perform statistical analysis. Depending on the test used, the software automatically plots the comparison bar and displays the p-value above it. We changed the representation of statistics as suggested.

      Figure 3: It is a bit odd that they finally conclude on RT as essential step that is reduced in the absence of SNAT7 and then they fail to provide statistical significance for this (Fig.3 panels F and G). One would expect that RT is much more affected given the huge effects on HIV-1 capsid and particle production shown in Fig.2 F, G and I.

      The reviewer is right in pointing that we observed a stronger effect during the later stages of the viral cycle, from transcription of viral RNAs (Fig. 2I and Supp. Fig. 2G) to the production of viral particles in the supernatant (Fig. 2D-G), than during the earlier stage of reverse transcription (Fig. 3F, G). Also, it is also possible that we might have missed the peak in ERT/LRT production, which is transient.

      It should be noted that SAMHD1 exhibits both dNTPase (Goldstone et al., 2011) and nuclease (Beloglazova et al., 2013) activities. The ability of SAMHD1 to restrict the virus, through dephosphorylation at T592, is mediated by its RNase activity (Ryoo et al., 2014), and not by the dNTPase activity (Welbourn et al., 2013; White et al., 2013).This could explain why SNAT7 exhibit a stronger impact on viral transcription than on reverse transcription.

      Figure 4; again single cell flow measurements of SAMHD1, pSAMHD1 and p24 /SNAT7 might help to more clearly discriminate effects that are specifically induced upon infection or happen in virally infected cells. Maybe alternatively IF?

      We thank the reviewer for this suggestion. As mentioned under comment #2, flow cytometry analyses are difficult to perform on strongly adherent primary human macrophages.

      With regard to immunofluorescence, there is a technical limitation based on the species in which the antibodies are produced. The antibody that targets the native SNAT7 protein, which is currently being validated in our laboratory, is produced in rabbits. An anti-CAp24 antibody produced in goats can be used. It will then be necessary to co-label the cells with anti SAMHD1 and phospho-SAMHD1produced in mouse. We will try to find options to co-label the cells.

      The wblot shown in panel D does not really reflect the point the authors want to make by the quantification in panels G-I. Primary data (D) suggests that SNAT7 KD reduces HIV-1 production even in the absence of SAMHD1. The quantification rather indicates that SNAT7 KD does not affect HIV-1 production in the absence of SAMHD1. This needs clarification/corroboration by orthogonal approaches.

      We respectfully disagree with the reviewer.

      Figure 4D shows a representative blot of the six donors analysed. As mentioned, the depletion of SNAT7 in the absence of SAMHD1 reduces the production of the viral proteins GagPr55 and CAp24 (see Fig. 4D). This is illustrated by the quantifications (Fig. 4G–I). Following treatment with Vpx, GagPr55 protein expression in SNAT7 KD macrophages is reduced by a factor of 2.6 for siRNA #1 (mean = 1.48, light grey bar) and by a factor of 1.83 for siRNA #2 (mean = 2.13, orange bar), compared to the control (mean = 3.9, pink bar) (Fig. 4G). Similarly, CAp24 protein expression was reduced by a factor of 2.2 for siRNA #1 (mean = 2.05, light grey bar) and by a factor of 1.36 for siRNA #2 (mean = 3.34, orange bar), compared to the control (mean = 4.52, pink bar) (Fig. 4H).

      These differences are therefore consistent between the Western blot and the quantifications. However, they are not significantly different to those observed in cells treated with Vpx and depleted with control siRNA, suggesting that the viral restriction observed in SNAT7 KD cells is primarily due to SAMHD1.

      1. Figure 5: show SAMHD1 and pSAMHD1 levels upon glutamine supplementation.

      We thank the reviewer for this comment, we will perform the suggested experiment.

      1. I think the discussion is very thin, mainly summarizing the results; but fails to give broader context or critically discuss the limitations and further directions.

      We thank the reviewer for this comment. The discussion will be modified further accordingly.

      Looking at the data as a whole, I think the results support a modest functional importance of SNAT7 for HIV-1 production in macrophages. I acknowledge that the experiments in primary macrophages are prone to high variability in different donors and the authors transparently depicted their data. However clearly, I would advice the authors to tune down the extend in which they claim SNAT7-dependency given this huge variability and the sometimes-borderline statistics. We respectfully disagree with the reviewer.

      The cells used here imply greater variability than a cell line, but are also more relevant.

      Indeed, the effects observed in the late stages of HIV-1 production are:

      • ~80 % decrease in viral transcription compared to the control (Fig. 2I),

      • ~85 % decrease in CAp24 protein expression compared to the control, as quantified by western blot (Fig. 2E), or ~90 % by ELISA measurement (Fig. 2F),

      • a reduction of more than 90 % in the release of infectious particles (Fig. 2G).

      These results were all significant across donors, while SNAT7 depletion was always partial (Fig. 2C, between 31 to 62 % of depletion compared to the control in infected cells).

      Therefore, the data were obtained from a mixture of depleted and non-depleted macrophages. This means that the results may be underestimated.

      Together, our results show that SNAT7 is necessary for HIV-1 production.

      However, reading the comments, we realized that our conclusions regarding reverse transcription were too strong. SNAT7 depletion does not affect viral fusion and reverse transcription. The manuscript was modified accordingly.

      On top, there are a lot of optional experiments I am sure the authors are aware of that should be done at least in the future.

      For instance, how does HIV-1 upregulate SNAT7, is a viral accessory protein involved? What is the mechanism of SNAT7 dependent SAMHD1 phosphorylation? Does SNAT7 (or glutamine) regulate the activity of the SAMHD1 associated kinase / phosphatase) If so, does this impact on other targets of these enzymes? We thank the reviewer for these questions.

      To address the role of accessory viral proteins, we have already performed one experiment infecting hMDM with HIV-1 strains deleted for genes such as Nef, Vpr, Vpu and Vif, and have found no clear effect on SNAT7 protein expression compared to WT strains. As an alternative experiment, we could overexpress individual viral genes, such as Nef or Vpr, in HeLa cells and analyze their impact on SNAT7 expression by Western blot.

      It is also possible that SNAT7 expression and recycling of lysosomal glutamine are modulated by the macrophage intrinsic immunity in response to HIV-1 infection.

      The Thr592 motif of the SAMHD1 protein is phosphorylated by Cyclin A2/CDK1 and type 1 IFN in non-cycling cells, such as MDMs (Cribier et al., 2013). For now, the relationship between SNAT7 and SAMHD1 remains unclear. However, (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This has been added to the discussion to explain the relationship between the 3 partners.

      **Referees cross-commenting** I think the comments from the other referees are reasonable and consistent with my assessment

      Reviewer #1 (Significance (Required)):

      Strength and limitations see above;

      Significance: I think this work is of high interest for virologists working in the field of HIV-1 and infection of myeloid cells. In case SNAT7 (and hence glutamine) indeed regulates the phosphorylation of SAMHD1, there could potentially be broad relevance of this work. However unfortunately, this aspect remains underdeveloped and is also not discussed

      Field of expertise: HIV-1, immunology, cell biology

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      In this report, Herit and colleagues describe the role of a HIV-1 dependency factor that promotes virus replication in macrophages. The authors suggest that the lysosomal membrane-associated SNAT7 glutamine transporter is a HIV dependency factor, that promotes virus replication by enhancing reverse transcription and Gag synthesis. The authors use transient knock-down approaches in primary macrophages to identify that SNAT7 depletion does not impact viral entry but inhibits early reverse transcription which was reversed by exogenous glutamine addition. While reverse transcription enhancement was likely due to selective increase in phosho-SAMHD1 expression, mechanisms by which SNAT7 enhanced viral gene expression were not clearly defined. These are well-controlled studies that pinpoint the role of SNAT7 in the early steps of viral life cycle and highlight the intricate interplay between macrophage metabolism and HIV-1 replication. While the question that is addressed is important, and the hypothesis overall sound, the data presented needs to be strengthened to support the conclusions. There are numerous weaknesses in data interpretation as well.

      1. Figure 1: SNAT7 expression was selectively enhanced upon differentiation of monocytes into macrophages but absent in CD4+ T cells. Though there is a claim of enhancement of SNAT7 expression upon HIV-1 infection of macrophages, RT-qPCR analysis shows the opposite trend (Fig 1E) and SNAT7 protein expression changes are modest. Statistical analysis in Fig. 1H needs to be revisited. The number of replicates vary for the lysates harvested at different day post infection, which might have an impact on the statistical test. To determine if SNAT7 expression enhancement is dependent on establishment of virus infection, as the authors imply, control lysates of virus infections in presence of replication inhibitors should be included.

      We thank the reviewer for this comment. Indeed, there is a modest, but statistically significant increase in SNAT7 protein expression upon HIV-1 infection over time (Fig. 1G, H), without any modulation of SNAT7 gene expression (Fig. 1E). This indicates that the regulation of SNAT7 expression in this context is only at the translation level (i.e. increase of translation or stabilization of the SNAT7 protein).

      As mentioned, Fig. 1H aggregates between 3 to 7 independent experiments on different donors depending on the infection time point. SNAT7 protein expression is increased already at 1 day post-infection and until 8 days. The statistical test used here, i.e. 2 way-ANOVA, compared Mock-infected and HIV-1-infected condition for each time point with the same number of donors. In this figure, the comparison is statistically different only at day 6 of the time course (7 donors). We agree that increasing the number of donors of the other time points could help to improve the statistical difference between control and infection condition.

      We thank the reviewer for the suggestion mentioning the use of replication inhibitors in this experiment. We plan to use inhibitors of reverse transcription (Nevirapin) and integration (Dolutegravir).

      The authors rely exclusively on western blot analysis for HIV-1 Gag expression in cell lysates as a measure of effects of SNAT7 on virus replication. Single cell analysis such as intracellular p24gag analysis by FACS should be included; this will provide a better measure of effects of SNAT7 onHIV-1 infection establishment.

      We respectfully disagree with the reviewer for this question. Indeed, to evaluate the effects of SNAT7 on HIV-1 replication, we measured Gag Pr55 and Cap24 using a Western blot approach (Fig. 2B, D and E), but also assessed the quantity of Cap24 in the supernatants and lysates using an ELISA measurement, the quantity of infectious particles using TZM reporter cells, and total viral transcription or more specifically Gag Pr55 transcription using qPCR (Fig. 2F, G and I and Supp. Fig. 2G).

      Regarding the quantification of CAp24 at the cell single level, please refer to comment #2 under Reviewer #1.

      Knockdown of SNAT7 in MDMs was partial at best; only 30-50% decrease in expression (Fig 2C), but the effects on viral gene expression (Fig. 2I), p24 release and infectious particle production is dramatic (Fig. 2F and G). This discrepancy is not addressed. Does SNAT7 knock-down negatively impact virus particle release? Please note that the representative WB in Fig 2B does not correlate with the quantification in Fig. 2D. There are no p55gag or p24gag bands in SNAT7#1 siRNA condition (Fig. 2B)? Data could also be rearranged to follow the logical sequence of virus replication cycle (viral RNa expression followed by Gag expression, and then release).

      We thank the reviewer for this comment. Our samples are indeed a mixture of SNAT7-depleted and non-depleted macrophages and RNA interference in these cells often leads to a decrease of 50 % of the protein expression.

      To determine whether SNAT7 is involved in the release of particles, we quantified Cap24 in cell lysates and in the cell culture medium separately, and normalized the results to the total protein content. The absence of SNAT7 reduced the amount of Cap24 measured by ELISA in both samples to the same extent, showing that there is no storage of Cap24-positive viral particles inside the infected macrophages. These data were initially pooled in one graph (Fig. 2F), but separate graphs are now provided in new Supp. Fig. 2 E, F.

      Regarding the western blot shown in Fig. 2B, please refer to comment #5 under Reviewer #1.

      In the new version of the manuscript, we arranged the figures and placed the later stages of the viral cycle in Fig. 2 and the earlier stages, such as fusion, reverse transcription and transcription, in Fig. 3.

      Data interpretation would be greatly improved by including infection controls (RT or integrase inhibitors) to confirm that measurements of viral RNA and Gag are indeed modulated by SNAT7 expression.

      We thank the reviewer for this suggestion to include inhibitors of viral replication as controls. In our experiments, cells were Mock-infected in parallel as a negative control of viral detection. We provide the results in the new version of the manuscript to show that (i) there is no detection of viral or Gag RNA in the absence of the virus, (ii) the expression of viral genes measured in HIV-1-infected SNAT7-depleted cells is not different from Mock-infected cells, indicating almost complete inhibition of viral transcription (Fig. 3H and Supp. Fig. 3B), also confirmed at the protein level (Fig. 2B, D-F).

      Figure 3: Decrease in SNAT7 expression in macrophages resulted in lower levels of early reverse transcripts. But surprisingly, LRT levels were not as affected by decreases in SNAT7 expression. The authors go on to suggest that decreases in early RT are due to loss of phospho-SAMHD1 and increases in catalytically active form of SAMHD1. Mechanistically this does not make sense: LRT should be similarly affected by increase in catalytically active SAMHD1. dNTP concentrations should be measured to determine if the rescue of RT is dependent on SAMHD1 dNTPase activity.

      We thank the reviewer for this comment. LRT concentrations are very low in human macrophages and more challenging to detect than ERT concentrations. This might explain why the differences observed between the SNAT7-depleted and control conditions appear less pronounced for LRT than for ERT.

      Furthermore, we cannot rule out the possibility that SNAT7 has a cumulative effect throughout the viral cycle. While reverse transcription remains statistically unaltered, and despite the reduced levels of ERT and LRT in SNAT7-depleted macrophages (Fig. 3 F, G), there is a significant impact on the transcription of viral RNAs (Fig. 2I) and Gag (Supp. Fig. 2G). This step may also be altered by the ribonuclease activity of SAMHD1 (Beloglazova et al., 2013; Ryoo et al., 2014).

      Finally, with the help of Dr Baek Kim in Atlanta, we attempted to quantify dNTP concentrations in our human macrophages. Unfortunately, it was not possible to draw any conclusions, as the concentrations of dNTPs extracted from our cells were far too low.

      Furthermore, it should be noted that SAMHD1 viral restriction through its phosphorylation at T592 is not correlated with its dNTPase activity (Welbourn et al., 2013; White et al., 2013), but with its ribonuclease activity (Beloglazova et al., 2013; Ryoo et al., 2014). This is supporting why SNAT7, by modulating the ribonuclease activity of SAMHD1, could have a greater effect on viral transcription than on reverse transcription.

      There is lack of consistency in the data: p24 release upon SNAT7 depletion is highly variable. While there is a dramatic >90-95% decrease in p24 release (Fig. 2G), the effects are much more moderate in Fig. 4H (50-60% attenuation), even though siRNA-mediated depletion was similar across the data sets. The authors should comment on the variability in their findings.

      We thank the reviewer for this comment, but believe that Figure 2E rather than Figure 2G is to be mentioned regarding the quantification of CAp24 by Western blot and to be compared with Figure 4H.

      In Fig. 2E, we observed an average reduction of 85 % in CAp24 expression normalized to Clathrin HC expression across different donors for both siRNAs targeting SNAT7. For Fig. 4H, there was a 73 % reduction in CAp24 levels for siRNA #1 and a 56 % reduction for siRNA #2. In addition, it should be noted that the reduction in Gag levels is greater in Fig. 4G (between 77 % and 83 %) than in Fig. 2D (between 55 % and 72 %).

      Therefore, there is some variation in the results obtained with the different donors, which could be explained by variations in Gag cleavage among donors, but this does not impact the conclusions for both figures.

      SNAT7 is postulated to affect 2 steps in the virus life cycle: reverse transcription and viral transcription. But Vpx-mediated SAMHD1 degradation reversed both. Its not clear to me as to how SAMHD1 degradation impacts the role of SNAT7 in viral transcription. No explanation is provided.

      We thank the reviewer for this comment. As suggested, we will perform experiments to assess the impact of Vpx-mediated SAMHD1 degradation on viral transcription.

      Exogenous addition of glutamine only partially restored Gag synthesis and p24 release, which could be attributed to increased cytoplasmic levels and viral protein synthesis. What about effects on reverse transcription and viral gene expression?

      We thank the reviewer for this comment. We will perform the suggested experiments to assess the impact of glutamine supplementation on viral transcription.

      Reviewer #2 (Significance (Required)):

      This is a novel finding, as there are limited number of studies on amino acid transporters and HIV-1 replication enhancement in macrophages. Most of the previous work has focused on CD4 T cells. These studies on SNAT7 and HIV-1 infection establishment in macrophages might better inform the influences of macrophage metabolism on HIV-1 persistence and inflammatory responses.

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      This study investigates the role of the lysosomal glutamine transporter SLC38A7/SNAT7 in HIV‑1 replication in primary human macrophages. The authors demonstrate that SNAT7 is highly expressed in macrophages and upregulated upon HIV‑1 infection. They show that SNAT7 depletion inhibits HIV‑1 production at the reverse transcription step without affecting viral fusion or global cellular translation/transcription. Mechanistically, SNAT7 knockdown reduces the inhibitory phosphorylation of SAMHD1 at T592, and degradation of SAMHD1 by Vpx fully rescues viral replication. Extracellular glutamine supplementation partially restores HIV‑1 production in SNAT7‑deficient cells. Overall, the authors report interesting observations; however, the mechanistic investigation remains preliminary, raising concerns about whether the data fully support all the conclusions drawn. Major Concerns: 1. The mechanistic depth is insufficient. The authors do not elucidate how glutamine regulates SAMHD1 T592 phosphorylation, whether through metabolite‑mediated control of kinases/phosphatases or via indirect effects.

      We thank the reviewer for this comment. It is worth noting that (Meng et al., 2022) demonstrated that SNAT7 positively regulates mTORC1 activity at the lysosomal membrane through release of lysosomal glutamine, and (Dias et al., 2024) showed that inhibiting mTORC1 activity using drugs decreases SAMHD1 Thr592 phosphorylation in hMDM. Therefore, we could speculate that the absence of SNAT7 down-regulates mTORC1 activity, which then leads to decreased SAMHD1 phosphorylation. This is now further discussed in the discussion section of the manuscript.

      The authors do not measure intracellular dNTP levels upon SNAT7 knockdown, which is the key functional substrate of SAMHD1. They also do not directly demonstrate that glutamine supplementation restores dNTP pools.

      We thank the reviewer for this comment. Please, refer to comment #5 under Reviewer #2.

      Extracellular glutamine only partially rescues viral production, implying the existence of transport‑independent functions of SNAT7 or additional pathways. This important observation is not discussed.

      We thank the reviewer for this comment. The discussion has been modified accordingly.

      It is suggested that the key findings be validated in immortalized THP‑1 cells differentiated into macrophage‑like cells by PMA.

      We thank the reviewer for this suggestion but don’t really understand why this would strengthen our conclusions. Indeed, despite the known variability between donors and technical limitations to transduce cells, we chose human blood monocyte-derived macrophages as a relevant non-transformed model for HIV-1 infection of macrophages. They also represent to some extent the human diversity.

      The Discussion section should be expanded to include the potential translational implications and limitations of the present study.

      We thank the reviewer for this comment. The discussion points to some elements of potential translation and limitations of the study.

      Reviewer #3 (Significance (Required)):

      General assessment: This study identifies the lysosomal glutamine transporter SLC38A7/SNAT7 as a novel host dependency factor for HIV‑1 replication in primary human macrophages. The major strengths include the use of physiologically relevant primary macrophage models, a well-organized experimental pipeline from expression profiling to functional validation, and the establishment of a link between SNAT7, glutamine metabolism, and the HIV restriction factor SAMHD1.

      Advance: It extends current understanding of HIV‑1 host dependency factors and immunometabolism by revealing a compartment‑specific metabolic pathway that supports viral reverse transcription.

      Audience:This work will primarily interest specialized researchers in HIV‑1 biology, host-virus interactions, restriction factors, and antiviral innate immunity.

      2.15.1.0

    1. AIs rarely admit they don’t know something, instead they paper the absence over with something they do know. We may not be able to answer questions about how Indonesians see the world, but LLMs will happily disguise those useful absences with opinions of how Americans imagine Indonesians see the world.

      I think AI literacy or awareness needs to be more popular than it currently is. There is this burden on us as individuals to be vigilant of AI yet it has been embraced so quickly across several aspects of life. AI literacy needs to be a priority.

    1. Reviewer #2 (Public review):

      Summary:

      This review by Dorrell and Whittington covers a number of aspects related to normative modeling of grid cells. They begin by discussing key experimental insights on grid cell phenomenology. Then, they discuss how grid cells can be used to perform path integration and how they size up as efficient codes of space. These two sections then lead the authors to discuss how combining path integration and efficient coding objectives leads to models of axis-aligned grid cells in a single module. Discussion on non-linear objectives leading to multi-modules is presented. The review ends with several outstanding questions and an optimistic outlook of how normative models (particularly, task-optimized RNNs) can be used as tools for advancing understanding in neuroscience.

      Strengths:

      (1) The review is timely and covers an area that has seen a lot of recent activity. This discussion around many of the different results (and kinds of models), I think, will be generally helpful for the field.

      (2) Although I think the story could be a little more coherently made (see below), in general I enjoyed the author's flow from efficient coding -> efficient coding + path integration -> efficient coding + path integration + non-linear objective. This framing supports the specific conclusion the authors arrive at.

      (3) I also really liked the message that the review made of how normative modeling, despite some of its challenges/limitations, can be used effectively in neuroscience. The discussion of cycling between "experimental" modeling (e.g., vanilla RNNs) and theoretically-grounded models was nice, and I think it helps demonstrate the value of this approach.

      (4) Showing how the metric loss could be seen as a bandpass filter (Figure 3C) was nice and a contribution of the review.

      (5) While the focus of P4 (conjunctive HD-grid cells) felt initially a little cast aside, the discussion around "brain and task-optimised RNNs with standard architectural choices use fundamentally different path-integration mechanism" was nice and I think helpful for steering the community to an interesting open problem.

      (6) Identifying how "non-linear functionality" can lead to multi-modules was nice and not something that I have seen as clearly presented before.

      Weaknesses:

      (1) The authors view the experimental evidence for grid cells being linked to path integration as "specific and strong" and that the " key computational feature that defines entorhinal cortex [is] path-integration". I think experimentalists (at least the ones I work with) would push back on that. First, it's hard to isolate path integration in rodent experiments. So while Gil et al. (2018) did about as good a job as you could do, there are still other interpretations of the results that are not purely path integration dependent. And second, as the authors point out later in the review, there is experimental work finding that grid cells are disrupted in large environments and 3D. Path integration certainly happens (to some extent) in these spaces, which begs the question of how it is achieved with weakened grid coding. Thus, I think reducing the claims about how strongly grid cells are experimentally linked to path integration is called for.

      (2) The authors introduce the idea of efficient coding of space and discuss how grid cells are not optimal. It is later clarified (Sec. 5.3) that multi-module codes can be efficient (even if not the most optimal). I was confused reading Section 3, because in Section 2 the multiple modules are discussed, but then in Section 3, they are dropped, and only a single module is being considered. Equation 2 was also a little confusing to me. Alpha is not defined, and I would have thought that it would be x^Tx' - g(x)^T g(x') and not x^Tx' g(x)^T g(x'). Given that there is no page limit here, I think a little more detail in Section 3 would be helpful.

      (3) In Section 3, the authors make use of P2 (translation invariance within a module) to rule out (or, at least, question) certain models/approaches. While this is certainly a standard assumption made in theoretical work, it is not very well supported by experimental findings. In particular, Diehl et al. (2017), Ismakov et al. (2017), and Dunn et al. (2017) all found that individual grid fields systematically vary in their peak firing rate. In addition, Redman et al. (2025) found that, within a given module, there was a small but robust diversity of grid orientations and spacings. These suggest that grid cells within a single module may actually be able to encode properties of local space and give some support to normative models that find efficient space coding with grid cells by finding non-axis-aligned grid fields. I think this is all important to mention because: a) it provides more biological nuance to the question about spatial coding; b) it provides more ways in which to test models. For instance, in Redman et al. (2025), the Sorscher et al. (2022) model was shown to produce variability in grid properties that loosely matched what was found in real data. For tests like this (e.g., how much does a model reproduce variability in grid firing field peak rates), I think it is going to be important for continuing to evaluate models.

      (4) The focus of the review, I know, is grid cells, but of course, grid cells are part of the MEC and the larger hippocampal network. I totally understand, at some level, you have to make a decision of what to model, but it seems that there are other functional classes of neurons (border cells, head direction cells) that all play an important role in path integration. And while the models the authors consider at the end of the review capture properties of grid cells really well, they do so at the cost of not modeling anything else. The authors mention this in the context of the models not capturing conjunctive grid-head direction cells, but I think the point is a deeper one, and more discussion of at what level it makes sense to consider grid cells only is important.

      (5) As I mentioned in the Strengths section, I did enjoy the flow of the paper on how path integration + efficiency is needed to get grid single modules and path integration + efficiency + non-linearity is needed to get multiple grid modules. This creates the story that adding more of these theory-driven constraints helps lead to more "accurate" models of grid cells. But one alternative view is that, if path integration + efficiency is enough to get a single grid module (but only a single grid module), then maybe the utility (or need) of multiple grid modules comes from something else. That is, instead of saying "we need more constraints to get multiple modules", it could be evidence for "we need to re-think whether multiple modules might need a different theory to explain". While I understand this is a big picture question that maybe isn't entirely fair to ask of the authors, I think: 1) the authors do a nice job of positioning their review as a kind of discussion on what normative modeling can provide to neuroscience, so having this discussion on when the failure of a model to capture ALL aspects of the biological features motivates further constraints as opposed to a new approach, would be useful; 2) this question connects with the title of the paper, i.e. "what is the question?"

    2. Author response:

      We thank the reviewers for their time and attention which will significantly improve the paper. Further, we are grateful for their appreciation of our goals and work. In sum, the reviewers point to our overstated discussion of experimental evidence which we will tone down, some slightly confusing points of argumentation which we will clarify, and some discussion points on the role of normative theories that we will add text to address. We believe this will improve the paper significantly and hope you agree!

      Major Concern: Experimental Support for Path-Integration is not as strong as suggested

      The major point raised by all reviewers (reviewer 1 comment 1, reviewer 2 comment 1, reviewer 3’s only weakness) was that our presentation of the experimental perturbation evidence for path-integration is stronger than the reality. On reflection, we agree with this evaluation. We thank the reviewers for raising it; we will moderate our writing and include the sensible caveats raised. In sum, we still think that the convergence of evidence points to path-integration: first, disruptions to grid cells lead to path-integration problems, though these perturbations admittedly aren’t perfectly precise; second, normative theories of path-integration lead to grid cells and predict grid cell behaviour; third, mechanistic models of path-integration match grid cell behaviour and predict connectivity subsequently measured in entorhinal cortex. However, the evidence is not as all-encompassing as we suggested.

      That said, we’d like to further comment on one point. It is argued (reviewer 1, comment 1) that there are other theories of grid cell function, and that we discuss these theories. We discuss efficient-coding only models of grid cells and emphasise strongly why we reject them. We also briefly discuss oscillatory-interference models of path-integration and our reasons for not pursuing them further. As such, the reviewer is correct that our reading of literature strongly points us towards path-integration rather than other theories. We will slightly change the framing of the paper to make it clear that we are making a case. However, we are not aware of other theories the reviewer might be referring to. If the reviewer can point us to the other suggested theories that we do not address we would be happy to evaluate and include them.

      We now turn to the remaining comments, and how we plan to address them.

      Reviewer 1, Comment 2 – There could be multiple roles for grid cells

      The reviewer is indeed right that grid cells might perform multiple functions. This could just mean that the same computational motif (e.g. path-integration) is reused across different computations though that introduces no changes to the required normative theory. A stronger claim would be that grid cells perform both path-integration and some other function. This, according to a normative perspective, would most likely change how grid cells were optimally structured. We use the fact that large parts of the grid cell code can be captured with only path-integration as an argument against additional roles for grid cells. That said, there exist properties of grid cells not well-captured by path-integration which could well be smoking guns for additional roles of grid cells. The review already discusses both discrepancies between grid cells in three and two dimensions, and inhomogeneities in the grid in complex environments, and we will add two more (heading direction and peak-to-peak/angular variability, discussed below) that we are grateful to the reviewers for raising, and we discuss each of these in detail below.

      That said, whether these are necessarily arguments against purely path-integration or a reflection of interesting mappings of the core path-integration mechanism to the measurements we make remains to be seen. We would argue that both 3D grid cells (as explained below: there appear to be 2D slices in which grid cells behave as you’d expect) and spatial inhomogeneities (as explained in the paper: mappings of torus to world can introduce warping) can be explained without reference to additional computational roles of grid cells, which remain to us the most parsimonious explanation. We discuss next the slight update to path-integration only that the heading direction story suggest. But in sum, our view is that these discrepancies are likely not fatal for our path-integration-centric view of grid cells, but may well suggest some very interesting clarifications.

      Reviewer 1, Comment 4 – The system has two heading signals: true & internal, why?

      The reviewer is right to point to the puzzle over true vs. purely internal heading direction and which drives grid cells. We believe recent work from Abraham Vollan has effectively solved this puzzle: there appear to be two parallel circuits, one theta-modulated and following internal heading direction, another theta-unmodulated and aligning more with true heading direction. We will make sure to include discussion of this exciting work in our revised submission. This serves as a good example of an update we concede to the most austere version of the path-integration only view. Rather, it seems there are two parallel path-integrators working with different heading signals. The reasons for this remain unclear, but seem to be related to attention and planning (Vollan et al. 2026).

      Reviewer 2, Comment 3: Real Grid Cells have peak-to-peak variability & Angular variability

      The reviewer is right to point to the discrepancy in peak-to-peak firing rate and angles within a module that we did not adequately address. First, it is Sorscher’s RNN models, not nonnegative PCA that can generate a distribution of grid angles (Redman et al. 2025), which suggests that path-integration and such variability are compatible. We emphasise this point because the non-path-integration results from nonnegative PCA produce grid cells oriented at 30 degree offsets, something not measured even when you’re careful as in Redman et al. 2025. Thus, this becomes an interesting target for future work: perhaps using theories of path-integration up to an error threshold (rather than perfect) such angular diversity would be recovered. We will include this in our discussion. Further, we will include discussion of peak-to-peak variability that, as yet, has no obvious role.

      Reviewer 2, Comment 1: grid cells are inhomogeneous in 3D or complex environments, doesn’t that break the theory?

      Disrupted grid coding in extended or 3D environments indeed deserve more discussion, which we will add. In particular, we will add recent evidence that grid cells in 3D can be understood via the correct sequence of 2D projections(Qi & Yartsev, 2026). These two phenomena seem, to us, consistent with a path-integration only view of grid cells, as discussed above, and we hope to make this position clearer.

      Reviewer 2, Comment 5: Couldn’t there be other reasons for multiple modules?

      We have suggested a consistent normative framework in which multiple modules are explained through their role in non-linear coding. We think this elegant, and the most parsimonious current theory. We could, of course, be wrong. The discrepancies pointed to above might be good clues to follow to work out what else these modules might be doing, but currently these alternative explanations seem not to exist. We will text to clarify this.

      Reviewer 1, Comment 3: The review confuses computational and parameter parts of normative theory

      We disagree with the reviewer’s dichotomisation of normative theory. We view a normative theory as the complete procedure that produces the predictions. Almost all such theories have parameters and hence fitting a theory to data comprises both elements (a) [computational role] and (b) [specific parameters] identified by the reviewer. Occasionally theories have no parameters in the traditional sense, e.g. Rebecca et al.; instead they have heavy assumptions that play an equivalent role. It is true that, as the reviewer says, Sorscher et al.’s work was criticised for producing grid cells only for specific parameter values. We never found this as damning as Schaeffer et al. argued: simply it says that that theory is only correct within the given parameter range. Rather, arbitrating between models, parameters, or assumptions seems the same basic process: see what they predict and keep working with models while they remain useful ways to understand measured phenomena. If a model with very specific parameter values remains useful, that seems okay. In fact, we argued extensively why we think the nonnegative PCA model is not a useful model, but this was for completely different reasons. To us this story just reinforces the importance of hygiene in normative research: perform parameter sweeps and clarify how they constrain the claims you are making, carefully arbitrate what models can capture. Indeed, that is the whole goal of this review. We might be misunderstanding and, if so, we welcome correction.

      Reviewer 2, Comment 4: Normative Models of Cells Beyond Grid Cells

      The reviewer is right that extending these models to other cell types is an interesting area for further work, and that other cell types do seem to be involved in aspects of navigational computations both in RNNs and the brain. We will include a discussion to this effect in the revised manuscript. That said, we think the modularity of grid cells and their tight-linking to path-integration calculations should also be appreciated as a win!

      Reviewer 2, Comment 2: Multi-modularity is not cleanly explained

      We thank the reviewer for the comments, we agree. We will clarify the story regarding multiple modules, and will explain the equation further.

      Reviewer 1, Comment 5: the early introduction of phase-shifted Grid Cells seem the perfect place to normatively argue for Path-integration!

      We agree with the reviewer that this point can be made both normatively (‘oh look! If I try to do this optimally, I get translations!’) or, as we did early in the paper, mechanistically (‘oh look! With these cells I can do this!’). Indeed, a large part of the point of our paper is that path-integration is what is required to normatively derive phase-shifted grid modules, something discussed by Rebecca et al., our earlier work, and RNN studies, and appreciated for two decades. The earlier part of the paper does not discuss these papers as that section is aimed at giving intuition for the solution (mechanism). Later sections then heavily discuss the normative angle. We hope that division of labour makes sense.

      Finally, we will refine our summary of Rebecca et al. The reviewer is right that neurons don’t have to be discrete, we apologise for that error, but our understanding is that the only meaningful role of a neuron in Rebecca et al.’s work is the region in which is active, effectively making every neuron a binary unit, which seems dubious. We will clarify that by “predict velocity from each current and next encoding” we mean that the normative constraint they enforce is axiom 1: sequential activity of sets of neurons i then j can be uniquely interpreted as a trajectory, i.e. a step or velocity. Their work is elegant, and we will try to do more justice to it in the revision.

      To conclude, we thank the reviewers for their extensive comments, and look forward to releasing a version that addresses their concerns.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      WIPI1 is a PROPPIN family protein that has been implicated in Retromer-mediated membrane fission events. Although the cargos that it has been tested to be important for are diverse, one of the cargos that is unaffected is Beta1-Integrin. This leads the authors to assess another PROPPIN family protein - WIPI2, which is a homolog of WIPI1. KD using siRNA is effective and had no consequences on LAMP1, EGFR trafficking or GLUT1 trafficking. Integrin-B1, however, had a large and significant defect in its recycling from the endosome, with a clear endosomal colocalisation. Complementation experiments with WT WIPI2 recovered the phenotype, but various mutant WIPI2 complements resulted in elongated tubules, and there was also a dominant negative effect of the mutant. Integrin is a classic retreiver cargo, so the authors rationalise that WIPI2 may be playing a role with retreiver that WIPI1 plays with retromer. To assess this, they perform a set of immunoprecipitations. SNX17, the retreiver-associated sorting nexin, co-IPs with WIPI2 in a VPS26C-dependent manner. VPS26C but not VPS26 co-IPs with WIPI2, and the reciprocal with WIPI1. These interactions were not present for the FSSS mutation of WIPI2. WIPI2 localises to Rab11 endosomes mainly, as does retriever. Mutations of WIPI2 not only affected WIPI2 localisation, but also VPS35L mutations, indicating that there is a functional relationship between the two.

      On the whole, I find the manuscript compelling. The manuscript is very clearly written, the results are convincing and well performed. The flow of experiments is logical, and although not comprehensive in the subsequent mechanistic understanding, the fundamental findings are important and convincing. My comments below are, on the whole, minor and are intended to support the communication of the findings to the field.

      We are happy that the reviewer has received our work quite positively.

      (1) The IP interaction data were convincing; however, for me and some others, an interaction is only convincing when performed in vitro, and understood at a structural level. I do not suggest the authors do that in this case; however, I think, at a minimum, some sensible moderation of claims would be useful here.

      Indeed, quantitative in vitro data on the affinities would be a nice addition. However, we have significant trouble to recombinantly express and purify well-behaved WIPI2 in sufficient quantities for such studies. We keep working in this direction but are not there yet.

      We have now inserted a phrase into the discussion section highlighting this limitation: "Our immunoprecipitation assays cannot distinguish and more detailed structural and interaction studies with pure compounds will be necessary to elucidate the nature of this interaction". We nevertheless think that the the isoform specificity of the IPs, the effect of the point mutations in WIPI2 on these interactions, and the functional effects in vivo lend signficant support to the notion of a complex even if there is no proof of direct binding of WIPI2 to Retriever.

      (2) I found the final localisation data and its interpretation confusing. My interpretation of that data would not be that the retreiver is relocalised, but rather that there is less of both recruited to the membrane and the remaining localisation distribution is shifted. In addition, I am not quite sure of the model here - is the idea that WIPI2 recruits retreiver, if that is the case, I find it hard to resolve with its role as a mediator of fission. Clarity would be appreciated here.

      We are not quite sure what "final" localisation data the reviewer refers to, but we guess it is Fig. 9. This figure primarily provides in vivo evidence supporting the connection between Retriever and WIPI2. It does this by showing that the S67 substitution shifts both proteins. In WIPI2 wildtype cells, WIPI2 and VPS35L strongly colocalize in Rab11 compartments. S67 substitutions in WIPI2 abolish this localisation; WIPI2 shifts mainly to Rab5 compartments, where VPS35L shows only a moderate increase, and to Rab7 compartments, where VPS35L shows no increase at all.

      We do not understand the reviewer's interpretation that less Retriever would be recruited to the membranes in the S67 variants. VPS35L remains completely associated with punctate, presumably membrane-bounded structures also in the mutants, providing no evidence for a detachment from the membrane. The same is observed in a WIPI2 knockdown. Therefore, we did not claim that WIPI2 is the main factor recruiting Retriever to the membrane, for which our experiments yield no hints. This does not exclude that the interaction of WIPI2 could strengthen membrane recruitment, or that two pools of Retriever exist, one interacting with Snx17 and another interacting with WIPI2, and that both link to each other in a coat. We did not dwell on this in the discussion because our experiments cannot distinguish these possibilities and were not conceived to analyse membrane recruitment of Retriever.

      (3) I am concerned that the repeats being compared for statistical analysis are not biological repeats but technical repeats (cells in the same experiment). I should think the idea of the statistical comparison is to show experimental reproducibility and variability across biological repeats. Therefore, I would expect an appropriate number of biological repeats (3 or more minimum), to be the data compared in the statistical analysis and graphs. I think it is appropriate to average the technical repeats from each biological repeat. I find these to be useful resources https://doi.org/10.1083/jcb.202401074, https://doi.org/10.1083/jcb.200611141

      The repeats being compared are biological repeats from independent experiments. This is described in Methods, where the reviewer may not have seen it. In order to make the independent experiments more evident in the figures, we have now colour coded the individual cell measurements from the three independent experiments. This allows to visualize both the individual data points, the average from each experiment and the variability across the independent experiments.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from De Leo and Mayer presents evidence that the PROPPIN protein, WIPI2, associates with the Retriever complex, and is required for the proper transport of the SNX17-Retriever cargo, beta1-integrin. This finding fits with prior papers from the Mayer lab, which showed that a related PROPPIN, WIPI1, is required for the transport of some SNX27-Retromer cargo, including GLUT1. The retromer and retriever complexes are architecturally similar. Importantly, they act at the same endosomes, and each transports cargo from endosomes to the plasma membrane. Thus, the possibility that each also requires a structurally related PROPPIN is of interest. However, the manuscript is incomplete, and the main claims are only partially supported.

      Strengths:

      The topic that PROPPIN proteins are important for the function of the Retromer and Retriever complexes expands our view of the trafficking complex.

      Weaknesses:

      Many important controls are missing. Several points that are made in the manuscript are only supported through a single approach.

      We made a serious effort and implemented many suggestions of this reviewer, but orthogonal approaches are not always available or accessible.

      Reviewer #3 (Public review):

      Summary:

      The manuscript of Mayer and colleagues analyzes the function of WIPI proteins in mammalian cells. The authors previously identified CROP as a complex consisting of WIPI1 and the retromer complex, primarily in yeast cells. In mammalian cells, both WIPI1 and WIPI2 exist, whereas retromer has a homologous complex termed retriever. They now find that WIPI2 can form a complex with retriever subunits. They named this complex CROP2. Their data further indicate that CROP2 and CROP1 have distinct substrate specificities as knockdown of CROP2 subunits affects beta1 integrin sorting, whereas knockdown of CROP1 affects EGFR and GLUT1. They further identify a similar sequence (FSSS) in both WIPI1 and WIPI2, which is required for their specific binding to retromer and retriever.

      Strengths:

      CROP1 and CROP2 seem to use similar features for their formation, and have different substrates, which is convincingly shown.

      Weaknesses:

      The analysis lacks information that this is a complex as claimed. It can be deduced from the interaction analysis, but was not shown.

      It is of course desirable to obtain a detailed structural and in vitro characterisation of this interaction, which we have not provided because we currently do not have sufficient amounts of well-behaved source material for this. We nevertheless think that the interaction we show, which is strictly isoform-specific and dependent on single amino acid substitutions in a motif that in CROP1 is necessary for the interaction its recombinant subunits, supports that CROP2 is a similar a complex. We don't show a direct interaction but also don't claim in the manuscript that the interaction between WIPI2 and Retriever is direct and independent of additional factors.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, the reviewers generally value the contribution to the field, but they feel that some claims require additional experimental support.

      (1) I have summarized the major points below.

      (a) Both reviewers 1 and 2 agree that the quality of localization data presented in Figure 9 and S5-S7, and the interpretation of the data, could be improved. See comment 2 from reviewer 1 and comments 23, 24 and 25 from reviewer 2. They not only suggest ways to improve the presentation of the data, but additionally suggest improving the staining of the Rab11 marker and additionally explain the lack of co-localization between VPS35 and Rab5, which has been reported in the literature.

      This impression was due to the fact that some figures showed projections of image stacks, which was not indicated clearly in the figure legend. We have changed this and now show single image planes throughout all figures.

      (b) Both reviewers 1 and 3 note that the evidence supporting a functional WIPI2-Retriever complex in vivo is currently weak. We agree that additional biochemical data demonstrating the presence of the CROP1 and CROP2 complexes in vivo would strengthen the central message of the paper and elevate it to a more fundamental discovery.

      We understood that the reviewers did not ask for further in vivo evidence but would welcome structural characterisation of the complex and quantitative binding data in vitro with purified proteins. Structural characterisation is out of scope of our study and in vitro binding studies have remained hampered by the fact that WIPI2 is hard to express and purify and not well behaved in vitro.

      (c) All reviewers agree that the authors should carefully repeat their statistical analysis to account for the number of biological replicates. Reviewer 1 suggests publications that the authors could refer to.

      The reviewers have probably overlooked the respective description in the methods section, where it had been stated that we analysed biological replicates from independent experiments. In graphs showing measurements from individual cells we now make this evident through colour coded dots, in which each colour represents data points stemming from an independent experiment. This makes it evident that the variance from experiment to experiment is low. The means (n = 3) were generally compared using a two-tailed unpaired t-test.

      (d) Reviewer 2 additionally has various minor points that would greatly improve the readability and presentation of the work, and we recommend addressing (comments 1, 2, 3, 4, 12, 15, 17, 20, 27, 28, 29). All reviewers, in general, provide great minor suggestions. It would be great if the CROP1 and 2 complexes could be clearly introduced in each figure. We also agree that the WIPI2 CT labelling is confused and should be changed to "control" or similar.

      Many of the points raised by this reviewer were actually quite minor or questions of personal preference, not major problems as stated in the review. Nevertheless, we found a number of useful suggestions in this review and have addressed these points as detailed in the response to reviewer 2.

      (2) In addition to the major shared concerns laid out in the points above, reviewer 2 has some further minor suggestions:

      (a) Comment 6. Could the author explain the discrepancies between the example blot shown in Figure 1D and the quantification (1E).

      The two have actually been quite consistent. The reviewer might have mistaken the marker lane as the 0 min reference value to arrive at this impression. We have now removed the marker lane to avoid this.

      (b) Comment 9 - could the authors clarify how surface labelling experiments were carried out?

      This had been clearly described in the methods section, where this reviewer has probably not seen it.

      (c) Comment 11 - The reviewer suggests normalizing the surface levels of markers to the cell area and not per cell. This is a reasonable suggestion.

      The analysis had already been performed as proposed. This had been clearly described in the methods section, which the reviewer may not have looked at.

      (d) Comment 19 "In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures." The authors could try a Rab4 or Rab11 overexpression plasmid to show whether these are elongated recycling tubules.

      This has now been added.

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The figures are not colourblind friendly, and should be changed to be so. Additionally, single colour images should be grayscale.

      That was a good learning opportunity. We adapted the colour schemes of the images to make them more colourblind friendly, now using magenta, green, and white for the overlaps. In doing so we have relied on published recommendations, but we have not found a colourblind colleague to check the efficacy of this change.

      (2) WIPI2^CT labels are confusing, as people may think they are a mutant. I suggest changing to "control" or similar.

      These have been changed.

      (3) "The effect was comparable to that of a knockdown of SNX17 (Figure 3 A, B)." On page 6. Based on this sentence, I was expecting to see a comparison to SNX17 KD, but it was not there as far as I can tell.

      This statement referred to a publication by P.Cullen and collaborators. We have changed the wording and inserted the (missing) reference to make this clear.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is modest. In addition, many of the claims should be better supported by the addition of orthogonal data. Moreover, the quality of some of the data presented needs to be improved. Overall, the manuscript requires better descriptions of the methods. In many figures, it was not clear how the experiments were performed.

      The experimental descriptions that the reviewer refers to had been provided in the Methods section, where this reviewer may have overlooked them.

      The paper should also be better organized. Some less important findings are in the main figures, whereas some critical results are in the supplemental figures. In addition, there were multiple issues with the readability of the paper, and the authors should consider using a professional editor to make the paper easier to read.

      We had given the paper to colleagues who found it clear, and also Reviewer 1 has underlined its clarity. Nevertheless, we have re-phrased the manuscript in some parts to optimise it.

      One of the main claims in the paper is that the FSSS motif of WIPI2, as well as a conserved amphipathic helix, is critical for WIPI2 function in the CROP2 complex. It is notable that these are the same regions that are also critical for the role of WIPI2 in autophagy (Gubas et al., 2024 PMID: 39152217). The authors should include this information in the manuscript and cite the paper.

      Indeed. We mention this now in the introduction of the revised version.

      Additional Major Issues:

      While some of the issues raised below are actually minor and/or matters of personal preference, several comments led us to improve and correct the figures and we thank this reviewer for the constructive suggestions.

      (1) In Figure 1, it appears from the representative images that WIPI2 KD cells have higher levels of EGFR (Figure 1A and 1B). Is this correct?

      To some degree. This increase is not systematic. A moderate increase has been observed only in 2 experiments out of 4. Therefore, we did not investigate this.

      (2) Also in Figure 1, the colocalization is difficult to see. The authors should add the separate channels in addition to the merged images. Since the point is supposed to be that there is no impact on EGFR, all of this data could go into the supplement.

      We had considered this already for the original version but dismissed the idea. The overlap is quantified in Fig. 1C, which provides the relevant values from four experiments. Fig. 1A/B provide only sample pictures, which also permit to see overlap (yellow) 0 and 5 min after the induction of degradation, which vanishes at later timepoints. Separating the channels would quadruple the space that this figure occupies, which would not be practical and not change the point to be made.

      (3) The scale bars for each panel differ from each other. To better assess the data, the exact same magnification should be shown for each panel.

      Corrected

      (4) Figure 1C is confusing. The authors should explain which lines correspond to EEA1 and LAMP1.

      Corrected

      (5) In Figure 1D, the authors show different blots for control and WIPI2 KD. Could the authors compare WIPI2 and EGFR in the same blot? Without a comparison on the same blot, it is impossible to know whether the starting levels of EGFR are the same. Moreover, the quantitation in Figure 1E sets the value for each cell line to 100%. Instead, the starting levels in each cell line should be compared. The authors should use the amount of EGFR at zero time in the control cells to define 100%, and then indicate the relative initial EGFR levels in the WIPI2KD cells.

      A new blot is shown now and the quantification has been performed as proposed.

      (6) The quantification in Figure 1E does not match the representative blot shown in Figure 1D. According to the graph, the rate of degradation of EGFR is similar in both cell lines. But the representative blot shows that there are large differences.

      We do not understand this comment. The representative blot shows similar kinetics for both. Perhaps the reviewer got confused by the fact that a marker lane was still present on the left blot and not labelled as such. The new version of the figure corrects this.

      (7) The blot showing the WIP2 knockdown in Figure 1D has a lot of background. However, the blot of the WIPI2 knockdown in Figure S1 looks very good. The authors should make sure that they load enough sample and use a good antibody for the experiments in Figure 1.

      The new blot that we added in response to comment 5 corrects this.

      (8) In Figure 2 and Figure 3A, the cells are too confluent. This is an issue because the cells might not be metabolically active. In addition, the signal is saturated. The authors should make sure that all of the data is collected on cells that are not too confluent.

      The confluency of the culture cannot be judged from single frames, which were selected to show several cells. We had controlled confluency and underlined in the Methods section that “For microscopy, the cells were plated on 18-mm-diameter glass coverslips on 24-well plates and grown for 2 or 3 days according to the protocol of DNA or siRNA transfection by reaching a confluency of 70-80%”. The reviewer may not have seen this.

      (9) One main issue with these figures, especially the non-permeablized cells, is that it is impossible to assess how much of the signal is on the cell surface. The authors should provide the methods that they used to prevent inadvertent permeabilization of the cells. Were these experiments performed at 4 degrees? The authors should include a control of an antibody to a protein that is not found on the cell surface.

      There is an internal control in that the non-permeabilised WIPI2KD cells, which have been treated with the same antibody, show no much less staining than the control cells (Fig. 3A). In WIPI2KD cells, integrin becomes accessible for antibody staining only upon detergent permeabilization. This demonstrates that our procedure does not lead to significant inadvertent permeabilization of the cells.

      (10) The authors should perform surface biotinylation assays as an orthogonal approach to determine GLUT1 levels and beta1-integrin levels at the cell surface, respectively.

      There is a strong, qualitative difference in the surface labelling of beta1-integrin that is not observed for GLUT1. Given that, it is not obvious to us what additional argument would be provided by surface biotinylation or subfractionation experiments.

      (11) In quantifying surface levels of GLUT1 or beta1-integrin by microscopy, the authors should normalize to the cell area, rather than per cell.

      The reviewer has probably not seen that the Methods section states that the cell area has been used for normalisation.

      (12) In Figure 3, the nuclear DAPI stain in the KD cells is much less bright than in the control cells. The authors should make sure to choose representative images.

      The nuclear DAPI signal has been visible in all cells. Depending on the position of the nucleus, is shape and dimension in the z-direction, individual nuclei can show different degrees of staining. The images shown are representative. We have adjusted the settings now to make the nuclei in the WIPI2KD cells easier to spot.

      (13) For the immunofluorescence studies, the authors should be using single z planes rather than maximum projection.

      Images have been exchanged by single planes.

      (14) For the experiments in Figure 3, the authors should check the total levels of EEA1 and LAMP1 by western blot to test whether WIPI2 KD affects the levels of these proteins. If these organelle marker proteins are impacted, this could impact the colocalization measurements shown in Figures 3C and D.

      We have measured the total fluorescence intensity of EEA1 and LAMP1 in the images. It shows no significant difference between control and WIPI2 knockdown cells (new Fig. 3F, H).

      (15) In Figure 4A, the helical representation is rotated in the WIPI2-Sloop; the orientation of the residues that are not mutated should stay the same.

      Yes. Done.

      (16) In Figure 4B and 4C, cells that were not transfected with WIPI2 WT or WIPI2 Sloop should be shown.

      Since the transfection efficiency is limited, the fields contain both non-transfected (lacking green fluorescence) and transfected cells (showing green fluorescence). We have now marked transfected cells with an asterisk.

      (17) The cells in the lower panel of 4B have an unusual morphology and are much more round. The authors should choose cells that are representative of each experimental condition.

      We now provide another field.

      (18) In Figure 4C, it looks like the magnification of the top panels is different from the bottom panels. The same magnification for all the panels should be shown (and the size of the scale bars should be the same.

      Corrected

      (19) In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures.

      We have done this (new Fig. S4). Rab4 is on tubules. Rab5 on the structures from which the tubules emanate.

      (20) In Figure 5A, the top scale bar is missing.

      Corrected.

      (21) In Figure 5B, the confluency is too high.

      See our response above. A single field does not permit to judge this. Confluency was controlled for all cultures. The cultures were not confluent.

      (22) The IP studies shown in Figures 6, 7 and 8, should be accompanied by colocalization studies.

      Colocalization measurments have now been integrated into the manuscript (Figs. S5, S6). They are consistent with the IP data.

      (23) Figure 9 was very confusing and should be broken up into multiple figures. Data showing that localization did not change in any of the cell lines can be put in figures that are distinct from figures that show that localization changed in the various mutants. Figures that show no change can go in the supplement.

      Since every panel of Fig. 9 shows a statistically significant difference we left the figure unchanged.

      (23) Representative figures should be shown in the same figure as the corresponding graph. In addition, the order of the colocalization data shown in the graphs and figures should match the order described in the text.

      We consider the graphs of Fig. 9 as the relevant information. Representative images are just illustration. Integrating them with the graphs would make it necessary to split everything up into multiple figures, making it harder to compare the different combinations. Therefore, we left the figures unchanged.

      (24) In Figure S7, the Rab11 signal looks continuous, which makes the colocalization analysis meaningless. The authors should determine how to take images that can be evaluated. On a more minor note, the zoomed panels should be labeled as well.

      This is a result of having shown a projections of multiple planes. The images have now been replaced by single plane images. Zoomed panels have been labelled and the scale bar added.

      (25) The low colocalization of VPS35L with Rab5 is surprising, as SNX17 has been previously shown to co-localize with early endosomes positive for EEA1. This result may have occurred due to overexpression because the authors chose to utilize plasmids that express a tagged protein. There are antibodies to each of the endogenous proteins, and this is what should be used for this set of experiments.

      This comment made us control the analysis performed for these images, which by mistake had been performed on z-projections rather than on single planes. This distorted the values. The re-analysed data shows a higher colocalisation with Rab5, but it remains inferior to colocalisation with Rab11.

      (26) The authors should determine whether β1-integrin colocalizes with WIPI2 in endosomal compartments.

      This was done. WIPI2 colocalizes with beta-integrin on EEA1-and SNX17-positive strcutures but not positive for LAMP1 (Fig. 3E/F).

      Minor points

      (27) In one of the panels in Figure 1A, "30 min" is duplicated.

      Removed

      (28) In Figures 5C and 5D, the y-axis should indicate that this is surface β1integrin.

      Changed and added “surface”

      (29) In Figure 9 there is a typo in panel A. It is VPS35L and not VPS35.

      Corrected

      Reviewer #3 (Recommendations for the authors):

      This is an overall convincing study, which shows that the two complexes, CROP1 and CROP2 function at different membranes and serve different substrates. While I agree with their localization analysis, I have one key issue. The authors claim that each of the two forms a complex and base this on their specific pull-down and western blot analyses.

      I find it important that they show that both indeed form stable complexes in vivo, using pull-down and mass spectrometry approaches. They have all the necessary tools in hand and could use WIPI1 and WIPI2 to demonstrate the existence of the two complexes. The FSSS mutants of each are good controls for such an analysis.

      The manuscript actually presents the demanded in vivo experiments. Figs. 6 to 8 show pull-downs of WIPI1 and WIPI2 from cells, including also the FSSS mutant. While we haven't analysed this interaction by mass spectrometry, the Western blot analysis confirms the analysis. Cooperation of these proteins is further supported by the in vivo phenotypes, where the S67A substitution in WIPI2 produces a similar phenotype on integrin beta1 localisation as inactivation of Retriever.

      A second aspect is the general presentation. The paper would be a lot more accessible if the subunits of each complex (CROP1 and CROP2) were also introduced in the figures of each part. For readers, a final model is helpful to put the data into context and show where each complex operates in the cell.

      We have introduced a scheme of the respective complexes, including the names of the compunds, in Figs. 6 and 7 to avoid confusion.

      Finally, it is not clear how the statistics compare to repeats in their data. This should be clarified.

      This had been described in methods. Statistics has always been done on biological replicates stemming from independent experiments. We have added a cartoon (Fig. 10) depicting the trafficking pathways affected by CROP1 and CROP2.

    1. Reviewer #1 (Public review):

      Summary:

      This paper reports the findings of a neuroimaging experiment that tested the hypothesis that the cortex, specifically early visual areas, reinstates the content from single events during our lives. The researchers tested this hypothesis by presenting to-be-remembered pictures of objects at spatial locations on the computer screen and then testing subjects with both recall and recognition. They show that during memory testing, the spatial location of the object can be decoded from the pattern of cortical BOLD responses measured with fMRI. They go on to show that the spatial tuning is higher during recognition than recall, that the tuning is correlated with memory retrieval accuracy, and that the retrieved precision is predicted by the encoded precision, particularly in the higher-level visual areas. Thus, the paper finds evidence of cortical reinstatement of details from a single event in a human life.

      Strengths:

      This is a strong manuscript that I have had the luxury of commenting on during a round of review at another prestigious journal. As a result, the authors have already made changes to address previous comments about highlighting the complementary learning systems approach more to motivate the alternative prediction that the cortex should only show evidence of reinstatement after repeated presentations. In addition, the authors have fleshed out the discussion of working memory in this task. They also revised their review of the literature to include citations suggesting spatial locations are normal parts of our episodic representations, likely obligatory in nature, as my group and others have argued in completely unrelated work. I applaud the authors for being responsive to a previous round of review and using the comments to address relatively minor issues with the paper, even though they moved on to a different journal. Thus, I found the paper even stronger than at first approach, and at first blush, the results were intriguing and the paper well written.

      Weaknesses:

      There is a logical perspective in the narrative that seems to unnecessarily weaken the paper. The paper shows evidence consistent with the conclusion that mnemonic representations are contained in early visual cortex, but then argues that those representations are not actually stored therein. For example, the first half of the last sentence of the conclusions (see page 19 of the manuscript). I understand the perspective that subcortical mechanisms must be involved in the act of retrieval, given the neuropsychology and other evidence. But if storage is elsewhere with the same fidelity so as to code this information, then how would such a memory system work? The MTL neurons would need to have the real, precise representation of all the orientations encoded at all the retinotopic locations, a mirror to V1 in terms of precision, because that's the actual memory representation being retrieved, so its fidelity will be limited by what is stored in the file, so to speak. Then, at retrieval, the paper proposes that the brain just reactivates the encoding context in V1 to help with the response output and ensure the precision of the behavioral responses. This must mean that the hippocampus/MTL has cells and networks with tuning functions that match the precision in all the cortical sensory systems that they are integrating context across, given the episodic memory models like Polyn and colleagues (2009, Psych Rev). So, there are little MTL maps that are completely redundant with V1, M1, A1, S1, etc.? Why such redundancy?

      Why not propose that what the subcortical systems do is to encode a unique pattern for that episode, that is separated from others, that just links (or provides pointers to, in computer science jargon) the contextual details stored in the cortical networks themselves? In this way, we can explain why neglected patients also neglect their memories of the town square. This has always been my interpretation of the results of the Polyn et al. (2006, Science) paper and the models tested with those whole-brain results. That is, you see widespread cortical context reinstatement during (one-shot) free recall events that included visual selective cortex for faces when faces were being recalled, but included a broad network, probably V1, and activating sounds in A1, body posture in M1, etc., though the latter three examples did not discriminate between categories of memoranda, in their experiments. Given that you show that activity in V1 during retrieval looks like it is being used, you should propose that the early cortex really participates in memory storage functions. V1 neurons are wired up to neurons of other selectivities in a competitive network with plastic synaptic connections. How would experience be prevented from changing activity in the cortex? Yes, cortical changes slow after the critical periods, as studied in the classic eye suturing experiments to study ocular dominance, but changes in cortical representations do not stop with maturity, with the pinwheel centers looking like they are context sensitive, thus, changing rapidly to events across time (Okamoto, Ikezoe, et al., 2011, Sci Reports). The brain would need a no-plasticity mechanism, and instead, it looks like the cortex can completely rewire even in adulthood (Buonomano & Merzenich, 1998, Annu Rev Neuro).

      I believe that the paper needs to describe the strong/radical interpretation of the current findings; that they are consistent with the view that the entire brain may be a memory structure, with encoding linking representations across sensory cortices. But also activating semantic and lexical systems, emotional networks encoding those aspects of context which we know can sometimes strongly drive effects, a nice prediction that could be made in the discussion/conclusions. Here you are looking at how precise the visual reinstatement is in V1 during retrieval following one exposure. One parsimonious mechanism to explain this effect is that the brain stores details of events using the neurons that do the high-fidelity perception of the event. Given that our goal is to stimulate thinking among fellow scientists so that this paper can be a citation classic, I think the paper should be revised so that it paints a complete picture of the theoretical possibilities of its findings.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      This preprint investigates the molecular mechanism by which warm temperature induces female-to-male sex reversal in the ricefield eel (Monopterus albus), a protogynous hermaphroditic fish of significant aquacultural value in China. The study identifies Trpv4 - a temperature-sensitive Ca²⁺ channel - as a putative thermosensor linking environmental temperature to sex determination. The authors propose that Trpv4 causes Ca²⁺influx, leading to activation of Stat3 (pStat3). pStat3 then transcriptionally upregulates the histone demethylase Kdm6b (aka Jmjd3), leading to increased dmrt1 gene expression and ovo-testes development. This work aims to bridge ecological cues with molecular and epigenetic regulators of sex change and has potential implications for sex control in aquaculture.

      Strengths:

      (1) This study proposes the first mechanistic pathway linking thermal cues to natural sex reversal in adult ricefield eel, extending the temperature-dependent sex determination paradigm beyond embryonic reptiles and saltwater fish

      (2) The findings could have applications for aquaculture, where skewed sex ratios apparently limit breeding efficiency

      Weaknesses:

      Although the revised manuscript represents an improvement over the original version, substantial weaknesses remain.

      We thank you for the critical comments. We have responded to your concerns by a point by point manner, and please see detail below.

      Scientific Concerns

      (1) Western blot normalization and exposure: The loading controls (GAPDH) in Fig. S3C appear overexposed, as do several Foxl2 blots. Because these signals are likely outside the linear range, I am not convinced that normalization is reliable. This raises concerns about the validity of the quantified results.

      We thank you for the concerns. We have repeated the experiments, and new blots were loaded in Fig.S3C.

      (2) Antibody validation and referencing (Line 776): The authors need to refer explicitly to figures demonstrating antibody validation. At present, these data are provided only as a supplementary file that is not cited in the manuscript. In addition, the Sox9a antibody appears to yield indistinguishable signals in control and RNAi conditions, suggesting that it may not recognize eel Sox9a. This issue is not addressed by the authors. Furthermore, antibody validation Western blots should be quantified.

      We thank you for the comments. We have repeated the siRNA experiments to show the specificity of the antibodies used. This file, named as the supplementary file 1, is now cited in “WB analysis” in the Materials and Method part. As required, the antibody validation of WB are uploaded in the supplementary file 1. Antibody validation for WB are now quantified, and please see the new figure 3 and supplementary Figure 3.

      (3) Unclear sample sizes (N values): Sample sizes remain unclear for several figures:

      (a) Fig. 3F - No N value is provided. Each graph shows three data points; does this indicate that only three samples were quantified? If ten samples were collected, why were all not quantified?

      We apologize for the confusion. Three data points were previously used to shown data of 3 replicates. In new figure 3F, 10 randomly selected sections were imaged, and the data are shown. In the revised manuscript, the sample numbers (the N values) are added, and all the information can be found in the figure legend.

      (b) Fig. 4 - No N values are reported.

      Now N values are added. Please see the figure legend.

      (c) Fig. 5A - Again, only three data points are shown per group, despite the apparent availability of twelve samples. The rationale for this discrepancy is not explained.

      We apologize for the wrong data representation. Now all the data points are shown in Figure 5.

      (4) qRT-PCR normalization: The manuscript does not specify the reference gene(s) used for qRT-PCR normalization. Although expression levels are reported as "relative," neither the identity of the reference gene(s) nor the justification for their selection is provided.

      We now have specify the reference gene in “Quantitative real-time PCR (qPCR) experiments” part in the Materials and Methods section.

      (5) Specificity of key antibodies: While the authors have made some effort to validate anti-Amh, anti-Sox9, and anti-Dmrt antibodies, the results remain incomplete. The Amh and Dmrt antibodies detect reduced protein levels following knockdown of their respective targets, which is encouraging. However, the Sox9a antibody shows no difference between control and RNAi conditions, suggesting it does not recognize eel Sox9. This is not acknowledged in the manuscript. In addition, no validation data are presented for Foxl2. Antibody validation data must be clearly referenced in the main text and presented in an interpretable and quantitative manner.

      The antibody specificity is very important. For that reason, we have generated at least two different antibodies for each target protein, using full-length or small peptide as antigen. We have repeated the experiments for key antibodies such as Dmrt1 and Sox9a. IF and WB results clearly showed the specificity of the antibodies.

      Author response image 1.

      Foxl2 antibody has also been reported in ricefield eel (Hu et al. SCIENTIFIC REPORTS | 4: 6884 | DOI: 10.1038/srep06884, Molecular cloning and analysis of gonadal expression of Foxl2 in the ricefield eel Monopterus albus).

      After short term warm temperature exposure, only a small portion of somatic cells in ovary may be induced to express the male markers. As different techniques have different capacity (sensitivity), some techniques were more easy to detect that change. For instance, qPCR and WB are ready to detect it, whereas IF is a little difficult in obtaining good quality data.

      (6) Immunofluorescence data quality: The immunofluorescence images remain difficult to interpret. I strongly encourage the authors to enlarge the image panels and to present monochrome images (white signal on black background). The current presentation severely limits interpretability.

      We thank you for the comments. We think that our IF images are of decent quality. Due to the limits of the Figure space (already busy for Figure 3), enlarging the image panels or presenting additional monochrome images will compromise the quality of other data. Alternatively, if you still concern its quality, we can put it in the supplementary.

      Author response image 2.

      (7) Unreferenced supplementary figure: Fig. S4 is included in the submission but is not referenced anywhere in the manuscript text.

      We now have renamed the supplementary Figures. And we have double checked the text to make sure all Figure information is correctly referenced. Figure S4 is removed, as it is not necessary.

      (8) Fig. 5B image resolution: The micrographs in Fig. 5B are too small to allow meaningful evaluation of the data.

      Now new Figure 5B images with higher resolution were shown.

      (9) Unexplained data inclusion (Fig. 5E): Fig. 5E includes a pERK blot that is not mentioned in the Results section. The rationale for including these data is unclear.

      Previous work have shown that FGF/ERK signaling may play a role in sex change of ricefield eel (in Chinese). We therefore examined the Erk activity to explore whether it is involved in sex reversal. The results showed that pErk was comparable between ovary and ovotestis. At your suggestion, we decided to remove the data.

      (10) Poor blot quality (Fig. S3C): The blots in Fig. S3C exhibit high background and overexposure. I am concerned about the reliability of the quantification shown in panel D.

      The experiments have been repeated at least three times, and similar results were obtained. We now have replaced some of the WB that were of high background or overexposure.

      (11) Poor blot quality (Fig. S5G): The Stat3 blots in Fig. S5G contain numerous white artifacts, raising concerns about their suitability for normalization in panel H.<br />

      We now have repeated the experiments, and uploaded a new representative blot with better quality.

      (12) Missing controls (Fig. 6E): Fig. 6E lacks controls for HO-3867 and Colivelin treatments alone. Without these controls, it is not possible to determine whether the reported effects are meaningful.

      We thank you for the comments. We now have added the data required (with HO-3867 and Colivelin treatments alone).

      (13) Graphical presentation: The use of a light blue-to-pink gradient in bar graphs throughout the manuscript does not aid interpretation. I recommend using more distinct colors (e.g., red, orange, green, blue, purple, gray, black) to improve clarity.

      We thank you for the comments. We now have changed the blue-to-pink gradient to more distinct color system to better present the data. Please see the detail in the revised Figures.

      In summary, the interpretation of the study remains limited by persistent issues related to data presentation, image quality, and reagent specificity.

      We thank you for the critical comments about our data, in particular for antibody specificity and image quality, and the detailed instruction for how to better present the data. Answering your questions have greatly improved the quality of the manuscript. We admit that due to the technique challenging (with different conditions and different doses of small molecules) and higher cost of animal experiments, some of the WB or IF experiments may not be of high standards.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Editorial Concerns

      (1) Overstatement of conclusions: In lines 16-18, the authors state that Trpv4 "mediates" warm temperature-driven sex reversal. This claim is too strong given the data and should be toned down.

      We agree with our editorial comment about the overstatement. Now it reads “Trpv4 links environmental temperature to testicular differentiation in ricefield eel”.

      (2) Misuse of statistical language (Line 213): The term "significant" is used where statistical significance was not measured. The wording should be revised.

      We thank you for the point, and now have replaced “significant” to “marked”.

      (3) Terminology (Line 238): The term "co-expression" is inaccurate in this context. I suggest replacing it with "co-upregulation."

      We thank you for the point, and have changed it accordingly.

      (4) Drug description errors (Lines 241-242): The manuscript incorrectly identifies which drug functions as an agonist and which as an antagonist. This caused considerable confusion and must be corrected.

      We have carefully checked the sentence, and it was correct, as RN1734 and GSK1016790A are known Trpv4 specific antagonist and agonist, respectively.

      (5) Gene examples missing (Lines 247-250): The authors should explicitly name the testis-biased and ovary-biased genes referred to in this section.

      We thank you for the point, and now it reads “warm temperature exposure increased the expression of testicular differentiation genes such as dmrt1 and gsdf, accompanied by moderately decreased expression of ovarian differentiation genes such as cyp19a1a and foxl2”.

      (6) Lack of experimental context (Lines 322-324): Rather than simply listing the drugs used, the authors should briefly explain what each compound inhibits or activates and why it was employed.

      We have described this in the manuscript. The information of pStat3 activator and inhibitor has been described in Lines 305-309, as “HO-3867, a curcumin analogue, is a selective pStat3 inhibitor, which blocks pStat3 activity by directly binding to Stat3 DNA binding domain, and Colivelin is a potent synthetic peptide activator of pStat3, which increases pStat3 levels by acting through the GP130/IL6ST complex”, and the rationale has been stated in lines 32--322 as “To functionally demonstrate that pStat3 signaling is downstream of Trpv4, rescue experiments were performed by injecting into ovaries with individual and combined small molecules”.

      (7) Discussion of evolutionary differences: The Discussion misses an important opportunity to address why Stat3 activates kdm6b in ricefield eel but represses it in turtles. It is difficult to reconcile how the same transcription factor could exert opposite effects on the same gene during sex determination without additional context. A comparison of kdm6b regulation and sequence conservation between turtles and ricefield eel would strengthen this section.

      We have downloaded the promoter sequences of red eared turtle and ricefield eel. Based on the DNA sequences (Author response image 3), the similarity (conservation) was low between the two species.

      Author response image 3.

      It was appeared that DNA around the Stat3 binding sites in turtle are GC rich (CpG island), which may be subjected to DNA methylation modification, whereas the DNA in ricefield eel are not GC rich.The observations imply that the role of pStat3 is to promote the repression of kdm6b in turtle but the activation of kdm6b in ricefield eel.

      Moreover, our unpublished data showed that Trpv4-controlled calcium signaling is required to remove the repressive histone modification H3K27me3 at the kdm6b gene. If pStat3 is downstream of Trpv4 in this case, it supports again that Trpv4-pStat3 axis activate kdm6b in ricefield eel.

      Warm temperature promotes female sex in turtle but male sex in ricefield eel. If pStat3 is mediating Trpv4, it is not surprising that it represses kdm6b in turtle but activate it in ricefield eel.

      Based on above, we have added some sentences in the discussion part, and it reads “We reasoned that a yet-unidentified co-factor may determine whether Stat3 is a transcriptional repressor or activator. A comparison of promoter sequences of kdm6b between turtle and ricefield eel supported this”.

      (8) Supplementary figure formatting: Supplementary figures should be provided in accordance with eLife formatting guidelines.

      We have now formatted the supplementary figures that are in accordance with eLife formatting requirement. Please see the new uploaded supplementary figures.

      In sum, the interpretations are still limited by the above concerns regarding data presentation and reagent specificity.

      We thank our editor for the inspiring comments. We believe we have addressed all the major concerns by our editor.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study combined careful computational modeling, a large patient sample, and replication in an independent general population sample to provide a computational account of a difference in risk-taking between people who have attempted suicide and those who have not. It is proposed that this difference reflects a general change in the approach to risky (high-reward) options and a lower emotional response to certain rewards. Evidence for the specificity of the effect to suicide, however, is incomplete, which would require additional analyses.

      We thank the editors and reviewers for this important assessment. Based on clinical interviews, we included patients with and without suicidality (S<sup>+</sup> and S<sup>-</sup> groups). However, in line with suicidal-related literature (e.g., Tsypes et al., 2024), two groups also differed substantially in the severity of symptoms (see Table 1). To address the request for evidence on specificity to suicidality beyond general symptom severity, we performed separate linear regressions to explain in gambling behaviour, value-insensitive approach parameter (β<sub>gain</sub>), and mood sensitivity to certain rewards (β<sub>CR</sub>) with group as a predictor (1 for S<sup>+</sup> group and 0 for S<sup>-</sup> group) and scores for anxiety and depression as covariates. Results remained significant after controlling anxiety and depression (ps < 0.027; Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) on the clinical questionnaire to extract the orthogonal components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. We then performed linear regressions using these components as covariates to control for anxiety and depression. Our main results remained significant (ps < 0.027; Table S9). We believe that these analyses provide evidence that the main effects on gambling and on mood were specific to suicide.

      Moreover, as Reviewer 3 pointed out, these “absence of evidence” cannot provide insights of “evidence of absence”. Although we median-split patients by the scores of general symptoms (e.g., depression and anxiety-related questionnaires) and verified no significant differences in these severities (Figure S11), we additionally conducted Bayesian statistics in gambling behavior, value-insensitive approach parameter, and mood sensitivity to certain rewards. BF<sub>01</sub> is a Bayes factor comparing the null model (M<sub>0</sub>) to the alternative model (M<sub>1</sub>), where M<sub>0</sub> assumes no group difference. BF<sub>01</sub> > 1 indicates that evidence favors M<sub>0</sub>. As can be seen in Table S7, most results supported null hypothesis, suggesting that general symptoms of anxiety and depression overall did not influence our main results. Overall, we believe that these analyses provide compelling evidence for the specificity of the effect to suicide, above and beyond depression and anxiety.

      Beyond these specific findings, this work highlights the broader utility of computational modelling and mood to better understand behavioral effect, showing how to use both mood and choice data to better comprehend a psychiatric issue.

      Please see Tables S7, S8, S9 and our revisions below:.

      Page 17:

      “Within patients, this group effect on gambling rate remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.024; also see Figure S11, Table S7 and Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, (ps < 0.001), we performed Principal Components Analysis (PCA) to extract main components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. To further control for anxiety and depression, linear regression using these components as covariates revealed that the group effect on gambling rate remained significant (p = 0.024; Table S9).”

      Pages 18-19:

      “Within patients, this group effect on the approach parameter remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.027; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on approach parameter remained significant (p = 0.027; Table S9).”

      Page 21:

      “Within patients, this group effect on βCR remained significant after controlling for gambling rate, earnings, mood-related outcome effect, mood drift effect, sex, illness duration, family history, diagnosis, and various medications use (ps < 0.032), as well as general symptoms (e.g., depression and anxiety; p = 0.001; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on this mood parameter remained significant (p = 0.001; Table S9).”

      Page 27:

      “Beyond these specific findings, this work highlights the broader utility of computational modelling and mood to better understand behavioral effect, showing how to use both mood and choice data to better comprehend a psychiatric issue.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors use a gambling task with momentary mood ratings from Rutledge et al. and compare computational models of choice and mood to identify markers of decisional and affective impairments underlying risk-prone behavior in adolescents with suicidal thoughts and behaviors (STB). The results show that adolescents with STB show enhanced gambling behavior (choosing the gamble rather than the sure amount), and this is driven by a bias towards the largest possible win rather than insensitivity to possible losses. Moreover, this group shows a diminished effect of receiving a certain reward (in the non-gambling trials) on mood. The results were replicated in an undifferentiated online sample where participants were divided into groups with or without STB based on their self-report of suicidal ideation on one question in the Beck Depression Inventory self-report instrument. The authors suggest, therefore, that adolescents with decreased sensitivity to certain rewards may need to be monitored more closely for STB due to their increased propensity to take risky decisions aimed at (expected) gains (such as relief from an unbearable situation through suicide), regardless of the potential losses.

      Strengths:

      (1) The study uses a previously validated task design and replicates previously found results through well-explained model-free and model-based analyses.

      (2) Sampling choice is optimal, with adolescents at high risk; an ideal cohort to target early preventative diagnoses and treatments for suicide.

      (3) Replication of the results in an online cohort increases confidence in the findings.

      (4) The models considered for comparison are thorough and well-motivated. The chosen models allow for teasing apart which decision and mood sensitivity parameters relate to risky decision-making across groups based on their hypotheses.

      (5) Novel finding of mood (in)sensitivity to non-risky rewards and its relationship with risk behavior in STB.

      Weaknesses:

      (1) The sample size of 25 for the S- group was justified based on previous studies (lines 181-183); however, all three papers cited mention that their sample was low powered as a study limitation.

      We thank the Reviewer for rising this concern. We agree that the sample size for S<sup>-</sup> group (n=25) is modest, and the prior studies we cited also acknowledged limited power. We wanted to point out that we obtained a comparable sample size to a prior study. In the revision, we therefore updated the section to justify this sample size in which we acknowledge the limited power of our study in the limitation section. Please see our clarification below:

      Page 32:

      “Third, despite replicating our main results in an independent dataset (n=747), the modest S<sup>-</sup> subgroup size (n=25) has a limited statistical power.”

      (2) Modeling in the mediation analysis focused on predicting risk behavior in this task from the model-derived bias for gains and suicidal symptom scores. However, the prediction of clinical interest is of suicidal behaviors from task parameters/behavior - as a psychiatrist or psychologist, I would want to use this task to potentially determine who is at higher risk of attempting suicide and therefore needs to be more closely watched rather than the other way around (predicting behavior in the task from their symptom profile). Unfortunately, the analyses presented do not show that this prediction can be made using the current task. I was left wondering: is there a correlation between beta_gain and STB? It is also important to test for the same relationships between task parameters and behavior in the healthy control group, or to clarify that the recommendations for potential clinical relevance of these findings apply exclusively to people with a diagnosis of depression or anxiety disorder. Indeed, in line 672, the authors claim their results provide "computational markers for general suicidal tendency among adolescents", but this was not shown here, as there were no models predicting STB within patient groups or across patients and healthy controls.

      Thank you for these thoughtful comments. Our study focuses on why adolescent patients with suicidality have increased risk behavior, aiming to provide a mechanism-based target for suicide prevention. Therefore, our dependent variable in the mediation model was gambling behavior. We also agree that the clinically relevant question is whether suicidality can be predicted from task-derived behavior/parameters. We thus used risky behavior and the potential mental parameters to predict STB. Linear regressions showed that gambling behavior, as well as the value-insensitive approach parameter, can predict suicidal symptom scores among patients (former: β = 9.189, t = 2.004, p = 0.048; latter: β = 5.587, t = 2.890, p = 0.005). In healthy controls, these predictions failed (gambling behavior: β = 1.471, t = 0.825, p = 0.411; approach: β = 0.874, t = 1.178, p = 0.241). These results suggest that clinical relevance of these findings apply exclusively to people with a diagnosis of depression or anxiety disorder. We found same patterns for the mood parameter (mood sensitivity to certain rewards: patients: β = -28.706, t = -2.801, p = 0.006; healthy controls: β = -2.204, t = -0.528, p = 0.599). In sum, we believe that our statement of “computational markers for general suicidal tendency among adolescents” is reasonable now. Please see our revisions below:

      Page 17:

      “Furthermore, linear regression showed that gambling rate can predict the current suicidal ideation score (BSI-C, β = 9.189, t = 2.004, p = 0.048) among patients, but not among HC (β = 1.471, t = 0.825, p = 0.411), suggesting that gambling behavior has patient-specific predictive utility for suicidal symptoms.”

      Page 19:

      “Furthermore, linear regression showed that approach parameter can predict the current suicidal ideation score (β = 5.587, t = 2.890, p = 0.005) among patients, but not among HC (β = 0.874, t = 1.178, p = 0.241), suggesting that value-insensitive approach parameter has patient-specific predictive utility for suicidal symptoms.”

      Page 21:

      “Furthermore, linear regression showed that mood sensitivity to CR can predict the current suicidal ideation score (β = -28.706, t = -2.801, p = 0.006) among patients, but not among HC (β = -2.204, t = 0.528, p = 0.599), suggesting that mood sensitivity to CR has patient-specific predictive utility for suicidal symptoms.”

      (3) The FDR correction for multiple comparisons mentioned briefly in lines 536-538 was not clear. Which analyses were included in the FDR correction? In particular, did the correlations between gambling rate and BSI-C/BSI-W survive such correction? Were there other correlations tested here (e.g., with the TAI score or ERQ-R and ERQ-S) that should be corrected for? Did the mediation model survive FDR correction? Was there a correction for other mediation models (e.g., with BSI-W as a predictor), or was this specific model hypothesized and pre-registered, and therefore no other models were considered? Did the differences in beta_gain across groups survive FDR when including comparisons of all other parameters across groups? Because the results were replicated in the online dataset, it is ok if they did not survive FDR in the patient dataset, but it is important to be clear about this in presenting the findings in the patient dataset.

      Thank you for raising the important issue of multiple testing and for asking us to clarify exactly which tests were covered by the FDR procedure. In the clinical dataset we conducted a large number of inferential tests (χ<sup>2</sup>, t-tests, ANOVAs, regressions) spanning: (i) group differences in demographic/clinical characteristics; (ii) sanity checks (e.g., anxiety/depression questionnaires); (iii) primary hypotheses (e.g., group differences in risky behavior); (iv) model-based analyses (parameter checks and between-group contrasts); and (v) control/sensitivity analyses. Post-hoc t-tests were performed only when the three-group ANOVA was significant. This yielded >150 p-values. FDR was applied using all these p-values. Please see Supplementary Note 8.

      (4) There is a lack of explicit mention when replication analyses differ from the analyses in the patient sample. For instance, the mediation model is different in the two samples: in the patient sample, it is only tested in S+ and S- groups, but not in healthy controls, and the model relates a dimensional measure of suicidal symptoms to gambling in the task, whereas in the online sample, the model includes all participants (including those who are presumably equivalent to healthy controls) and the predictor is a binary measure of S+ versus S- rather than the response to item 9 in the BDI. Indeed, some results did not replicate at all and this needs to be emphasized more as the lack of replication can be interpreted not only as "the link between mood sensitivity to CR and gambling behavior may be specifically observable in suicidal patients" (lines 582-585) - it may also be that this link is not truly there, and without a replication it needs to be interpreted with caution.

      Thank you for these important comments. This study focused on cognitive and affective computational mechanisms underlying increased risky behavior in STB. Accordingly, we compared patients with STB (S<sup>+</sup>) with patients without STB (S<sup>-</sup>) and healthy controls (HC) to examine the effects of STB on risky behavior. Therefore, group comparison, instead of dimensional measure of suicidal symptoms by Beck Scale for Suicidal Ideation, can answer our research questions directly.

      To enhance consistency between the clinical and replication datasets, we included all participants in each dataset when performing the mediation analysis. Given that S<sup>-</sup> and HC did not differ in gambling behavior or the approach parameter in the clinical dataset, we merged these two groups. In the replication dataset, to mirror the S<sup>+</sup> vs. S<sup>-</sup> contrast used clinically, we categorized the general sample into S<sup>+</sup> and S<sup>-</sup> based on BDI item 9. The mediation results remained significant in both datasets (the clinical dataset: a×b = 0.321, 95% CI = [0.070, 0.549], p = 0.016; the replication dataset: a × b = 0.143, 95% CI = [0.016, 0.288], p = 0.031), suggesting that STB is associated with increased risk behavior via stronger approach motivation.

      We also acknowledge the non-replication of the correlation between gambling behavior and mood sensitivity to certain rewards in the online sample. While this pattern might indicate that the link is specific to suicidal patients, it may also reflect sample-specific or unstable effects; thus, we now state this explicitly and interpret the finding with caution. Please see our revisions below:

      Page 15:

      “We next verified our results in an independent dataset, including the same task and BDI questionnaire in 747 general participants (500 females; age: 20.90±2.41)[46]. One item in BDI involves the measurement of STB. In item 9 of BDI, participants chose one option that describes them best: Option 1, “I don't have any thoughts of killing myself.”; Option 2, “I have thoughts of killing myself, but I would not carry them out.”; Option 3, “I would like to kill myself.”; Option 4, “I would kill myself if I had the chance.”. In line with the current definition of S<sup>+</sup>/S<sup>-</sup> in the clinical dataset, we identified S<sup>+</sup> group as choosing Option 2, 3, or 4, while participants selecting Option 1 were categorized as S<sup>-</sup> group.”

      Page 19:

      “Given significant correlations between group, approach parameter, and gambling rate for gain trials (ps < 0.017), we further conducted a mediation analysis with the assumption of the mediating effect of approach motivation of suicidality on the risk behavior. Given that we aimed to test the effect of STB, with S<sup>-</sup> and HC as controls, and given that S<sup>-</sup> and HC did not differ in gambling behavior or in the approach parameter, we merged these two groups for the mediation analysis. Results supported our hypothesis (a×b = 0.321, 95% CI = [0.070, 0.549], p = 0.016; Figure 2C), confirming that suicidal thoughts and behavior increase risk behavior through stronger approach motivation.”

      Page 26:

      “However, we did not observe any significant correlation between mood sensitivity to CR and gambling behavior (ps > 0.389), which suggests that the link between mood sensitivity to CR and gambling behavior may be specifically observable in suicidal patients. Alternatively, this non-replicated result may also reflect sample-specific or unstable effects, which needs to be interpreted with caution.”

      (5) In interpreting their results, the authors use terms such as "motivation" (line 594) or "risk attitude" (line 606) that are not clear. In particular, how was risk attitude operationalized in this task? Is a bias for risky rewards not indicative of risk attitude? I ask because the claim is that "we did not observe a difference in risk attitude per se between STB and controls". However, it seems that participants with STB chose the risky option more often, so why is there no difference in risk attitude between the groups?

      Thank you for pointing out the ambiguity. In our manuscript, “motivation” and “risk attitude” are defined at the computational level. Following prior work with this task Rutledge et al., (2015, 2016), we decompose observed gambling into (i) value-dependent valuation parameters that capture risk attitude (e.g., risk aversion and loss aversion, which scale the subjective value of outcomes), and (ii) value-insensitive, valence-dependent biases that capture approach/avoidance motivation. Accordingly, a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups which is what we observe for S<sup>+</sup> vs. controls. We have clarified this point in the computational modeling section.

      Pages 12-13:

      “Please note that a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups. Risk attitude is indeed conceptualized in economics as the curvature of the utility function (i.e., the subjective value) of the objective outcomes, with concave curves associated with risk aversion, and convex curves associated with risk seeking [54,56]. By contrast, the approach or avoidance bias apply to all the value. A possible interpretation of the approach bias is that participant approach the option with the highest possible gain (the lottery) in the gain frame; the avoidance bias would then reflect a tendency to systematically avoid the highest potential losses (the lottery) in the loss frame.”

      Reviewer #2 (Public review):

      Summary:

      This article addresses a very pertinent question: what are the computational mechanisms underlying risky behaviour in patients who have attempted suicide? In particular, it is impressive how the authors find a broad behavioural effect whose mechanisms they can then explain and refine through computational modeling. This work is important because, currently, beyond previous suicide attempts, there has been a lack of predictive measures. This study is the first step towards that: understanding the cognition on a group level. This is before being able to include it in future predictive studies (based on the cross-sectional data, this study by itself cannot assess the predictive validity of the measure).

      Strengths:

      (1) Large sample size.

      (2) Replication of their own findings.

      (3) Well-controlled task with measures of behaviour and mood + precise and well-validated computational modeling.

      Weaknesses:

      I can't really see any major weakness, but I have a few questions:

      (1) I can see from the parameter recovery that the parameters are very well identified. Is it surprising that this is the case, given how many parameters there are for 90 trials? Could the authors show cross-correlations? I.e., make a correlation matrix with all real parameters and all fitted parameters to show that not only the diagonal (i.e., same data is the scatter plots in S3) are high, but that the off-diagonals are low.

      Thank you for raising these thoughtful concerns. The current task consisted of 90 choices and 36 mood ratings. There were 5 choice parameters and 4 mood parameters. The apparently strong identifiability is not unexpected, as 90 choice trials and 36 mood ratings are comparable to those in prior computational modeling literature (Blain & Rutledge, 2022).

      As suggested, we computed cross-scorrelations between all generating (“true”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery. Please see Supplementary Pages 2-3.

      “Parameter recovery: Figure S3 shows good parameter recovery for both choice and mood winning model (choice: rs > 0.91, ps < 0.001; intraclass coefficients > 0.78; mood: rs > 0.90, ps < 0.001; intraclass coefficients > 0.86). Moreover, we computed cross-correlations between all generating (“true”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery.”

      Page 10:

      “The numbers of choice trials and mood ratings were comparable to those in prior computational modeling studies [34,35].”

      (2) Could the authors clarify the result in Figure 2B of a correlation between gambling rate and suicidal ideation score, is that a different result than they had before with the group main effect? I.e., is your analysis like this: gambling rate ~ suicide ideation + group assignment? (or a partial correlation)? I'm asking because BSI-C is also different between the groups. [same comment for later analyses, e.g. on approach parameter].

      Thank you for pointing out the lack of clarity. We performed group difference analysis and correlation of suicidal ideation analysis, separately. We first performed group difference analysis to test our hypothesis of STB effects. We then conducted correlational analysis to further specify our findings.

      (3) The authors correlate the impact of certain rewards on mood with the % gambling variable. Could there not be a more direct analysis by including mood directly in the choice model?

      Thank you for this insightful suggestion. As suggested, we tried to integrate mood into choice models by adding mood bias component(s) in line with previous literature (Vinckier et al., 2018). The first model (mcM1) assumes that mood biases choice, building on cM3 (the winning choice model). cmM2 further separated the mood bias parameter into two components according to participants’ choices.

      However, model comparison using BIC supported cM3 (Table S6), that is, without consideration of mood in choice modeling. This can be due to the lack of block design in our experimental design unlike e.g., Vinckier et al., (2018) and Eldar & Niv, (2015). Please see Supplementary Note 6.

      (4) In the large online sample, you split all participants into S+ and S-. I would have imagined that instead, you would do analyses that control for other clinical traits. Or, for example, you have in the S- group only participants who also have high depression scores, but low suicide items.

      Thank you for this insightful suggestion. Following prior suicide-related literature (Tsypes et al., 2024), we controlled for depression by including them as covariates. Note that depression scores were derived from our established bifactor model (Wang et al., 2025), which decomposed depression from the anxiety. These results remained largely significant (ps ≤ 0.050), except a marginally significant effect of group on gambling behavior (p = 0.059). Despite a trend, this effect with covariates of depression-related questionnaires is strong in our clinical cohort (p = 0.024; Table S8). This suggests that the link between suicidality and risky behavior persists above and beyond general depressive symptoms.

      Please see our clarifications below:

      Page 26:

      “After controlling for depression severity using our established bifactor model (see ref 60 for details), these results remained significant (ps ≤ 0.050), except a marginally significant effect of group on gambling behavior (p = 0.059). Despite a trend, this effect with covariates of depression-related questionnaires is strong in our clinical cohort (p = 0.024; Table S8). This suggests that the link between suicidality and risky behavior persists above and beyond general depressive symptoms.”

      Reviewer #3 (Public review):

      This manuscript investigates computational mechanisms underlying increased risk-taking behavior in adolescent patients with suicidal thoughts and behaviors. Using a well-established gambling task that incorporates momentary mood ratings and previously established computational modeling approaches, the authors identify particular aspects of choice behavior (which they term approach bias) and mood responsivity (to certain rewards) that differ as a function of suicidality. The authors replicate their findings on both clinical and large-scale non-clinical samples.

      (1) The main problem, however, is that the results do not seem to support a specific conclusion with regard to suicidality. The S+ and S- groups differ substantially in the severity of symptoms, as can be seen by all symptom questionnaires and the baseline and mean mood, where S- is closer to HC than it is to S+. The main analyses control for illness duration and medication but not for symptom severity. The supplementary analysis in Figure S11 is insufficient as it mistakes the absence of evidence (i.e., p > 0.05) for evidence of absence. Therefore, the results do not adequately deconfound suicidality from general symptom severity.

      Thank you for this important comment. Based on clinical interviews, we included patients with and without suicidality (S<sup>+</sup> and S<sup>-</sup> groups). However, in line with suicidal-related literature (e.g., Tsypes et al., 2024), two groups also differed substantially in the severity of symptoms (see Table 1). To address the request for evidence on specificity to suicidality beyond general symptom severity, we performed separate linear regressions to explain in gambling behaviour, value-insensitive approach parameter (β<sub>gain</sub>), and mood sensitivity to certain rewards (β<sub>CR</sub>) with group as a predictor (1 for S<sup>+</sup> group and 0 for S<sup>-</sup> group) and scores for anxiety and depression as covariates. Results remained significant after controlling anxiety and depression (ps < 0.027; Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) on the clinical questionnaire to extract the orthogonal components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. We then performed linear regressions using these components as covariates to control for anxiety and depression. Our main results remained significant (ps < 0.027; Table S9). We believe that these analyses provide evidence that the main effects on gambling and on mood were specific to suicide.

      As pointed out, these “absence of evidence” cannot provide insights of “evidence of absence”. Although we median-split patients by the scores of general symptoms (e.g., depression and anxiety-related questionnaires) and verified no significant differences in these severities (Figure S11), we additionally conducted Bayesian statistics in gambling behavior, value-insensitive approach parameter, and mood sensitivity to certain rewards. BF<sub>01</sub> is a Bayes factor comparing the null model (M<sub>0</sub>) to the alternative model (M<sub>1</sub>), where M<sub>0</sub> assumes no group difference. BF<sub>01</sub> > 1 indicates that evidence favors M<sub>0</sub>. As can be seen in Table S7, most results supported null hypothesis, suggesting that general symptoms of anxiety and depression overall did not influence our main results. Overall, we believe that these analyses provide compelling evidence for the specificity of the effect to suicide, above and beyond depression and anxiety.

      Please see Table S7, S8 &S9 and our revisions below.

      Page 17:

      “Within patients, this group effect on gambling rate remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.024; also see Figure S11, Table S7 and Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) to extract main components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. To further control for anxiety and depression, linear regression using these components as covariates revealed that the group effect on gambling rate remained significant (p = 0.024; Table S9).”

      Pages 18-19:

      “Within patients, this group effect on the approach parameter remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.027; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on approach parameter remained significant (p = 0.027; Table S9).”

      Page 21:

      “Within patients, this group effect on βCR remained significant after controlling for gambling rate, earnings, mood-related outcome effect, mood drift effect, sex, illness duration, family history, diagnosis, and various medications use (ps < 0.032), as well as general symptoms (e.g., depression and anxiety; p = 0.001; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on this mood parameter remained significant (p = 0.001; Table S9).”

      (2) The second main issue is that the relationship between an increased approach bias and decreased mood response to CR is conceptually unclear. In this respect, it would be natural to test whether mood responses influence subsequent gambling choices. This could be done either within the model by having mood moderate the approach bias or outside the model using model-agnostic analyses.

      Thank you for this important suggestion. As suggested, one interesting question was whether mood responses influence subsequent gambling choices and how to model them. First, we median-split mood responses (except the final rating) to compare gambling rate. Results showed a trend for less gambling rate in higher mood (t = -1.971, p = 0.050). However, there was no significant group difference (F = 0.680, p = 0.507). Second, with the assumption that mood biases choice, we constructed mcM1 based on cM3 (the winning choice model). Based on our finding of the negative correlation between mood sensitivity to certain rewards and gambling rate in S<sup>+</sup>, we separated β<sub>Mood</sub> parameter into β<sub>Mood-CR</sub> and β<sub>Mood-GR</sub> (cmM2). Model comparison using BIC supported cM3 (Table S6), that is, without consideration of mood in choice modeling. This can be due to the lack of block design in our experimental design unlike e.g., Vinckier et al., (2018) and Eldar & Niv, (2015). Please see Supplementary Note 6.

      (3) Additionally, there is a conceptual inconsistency between the choice and mood findings that partly results from the analytic strategy. The approach bias is implemented in choice as a categorical value-independent effect, whereas the mood responses always scale linearly with the magnitude of outcomes. One way to make the models more conceptually related would be to include a categorical value-independent mood response to choosing to gamble/not to gamble.

      We apology for the unclear statement. The approach bias is implemented in choice as a continuous value-independent effect, ranging from -1 to 1.

      It was true that the mood responses always scale with the magnitude of outcomes, since mood ratings were request after the outcomes. Therefore, mood parameters and the approach bias were both continuous.

      We also attempted to integrate mood into choice modelling. See Response 2 for Reviewer 3 for details.

      (4) The manuscript requires editing to improve clarity and precision. The use of terms such as "mood" and "approach motivation" is often inaccurate or not sufficiently specific. There are also many grammatical errors throughout the text.

      Thank you for this important suggestion. We have now explained motivation and mood in the Introduction section and the computational modeling section. Please see our clarifications below:

      Pages 3-4:

      “A growing literature indeed shows that risky behavior can be far better explained after adding value-insensitive approach and avoidance components to prospect theory [18,19], that is by including a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. This class of models highlights the important role of value-insensitive motivational components in decision making in addition to risk attitude-driven valuation (e.g., loss/risk aversion) [20].”

      Page 5:

      “Although mood is thought to persist for hours, days, or even weeks [30–33], momentary mood, measured over the timescale in the laboratory setting, represents the accumulation of the impact of multiple events at the scale of minutes [30,32,34–38]. Momentary mood external validity is demonstrated e.g., through its association with depression symptoms [37]. Mood is different from emotions, which reflect immediate affective reactivity and is more transient (e.g., from surprise to fear) [31–33,39].”

      We have corrected grammatical errors throughout the manuscript.

      (5) Claims of clinical relevance should be toned down, given that the findings are based on noisy parameter estimates whose clinical utility for the treatment of an individual patient is doubtful at best.

      Thank you for this comment. We agree that we did not evaluate the noise in our estimate e.g., by assessing the test-retest reliability on the task parameters, which is outside the scope of the study, and it is indeed possible that parameter estimate is somehow noisy. Therefore, we tone down the clinical relevance of our results. Please see our revision below:

      Page 32:

      “Next, we did not evaluate the noise in our estimate e.g., by assessing the test-retest reliability on the task parameters and it is indeed possible that parameter estimate is somehow noisy.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Title: I believe "aberrant mood dynamics" is both too general and overstating the results of this study, which did not measure mood dynamics longitudinally. "Aberrant" is also overly pathologizing. I would suggest sticking more directly to the results, for instance, "Insensitivity of momentary mood to non-risky rewards in adolescent suicidal patients".

      Thank you for this suggestion. We have now corrected it.

      (2) Abstract: in line 61, "Our study uncovers the cognitive and affective mechanisms" suggests that these are the only ones, and you uncovered them. Of course, there could be more mechanisms contributing to risk behavior in STB, so I would suggest removing the word "the" or adding "one of the".

      Thank you for this suggestion. We have now corrected it.

      (3) One major weakness of this study is that suicidal thoughts and behaviors were not assessed via a clinical instrument such as the Columbia Suicide Severity Rating Scale - this should be mentioned upfront.

      Thank you for this comment. According to medical records and information from family and friends by the researcher and psychiatrists, patients with suicidal thoughts and behaviors were categorized as suicidal group (S<sup>+</sup>), while patients without suicidal thoughts and behaviors were identified as control group (S<sup>-</sup>). Note that medical records and information were recorded from clinical interviews where the psychiatrists were vigilant for signs of suicidal ideation and inquired about suicidal-related thoughts and behaviors from both the patients and their families. Therefore, the current group operation was possibly comparable to Columbia Suicide Severity Rating Scale.

      (4) Table 1: female/male are sex, not gender (gender is man/woman/transgender/non-binary).

      Thank you for this suggestion. We have now corrected it.

      (5) Equation 1: It would be good to clarify what happens in gain-only or loss-only trials (the other value is then 0, but this can be clarified as it is not technically a loss or a gain).

      Thank you for this suggestion. We have now corrected it. Please see below for our revision:

      Page 12:

      “Please note that V<sub>gain</sub> is 0 in gain trials and V<sub>loss</sub> is 0 in loss trials.”

      (6) Figure 1E: The model prediction is not informative here. Given the linear regression model, there is no other option except that the mean prediction would overlap with the mean empirical measurement (unless the model was specified incorrectly). The same is true in Figure 2A.

      Thank you for this suggestion. We have now removed plots for model prediction.

      (7) Figure 1G: There was no analysis of the differences between groups in terms of earnings, given that the ANOVA was not significant. Still, if the claim is that risky behavior is sometimes suboptimal in this task, it would be good to show that there is a correlation between, say, symptoms of STB across groups and 1) risky behavior and 2) earnings.

      Thank you for this insightful comment. In the patient cohort, risky behavior (gambling rate)—but not earnings predicted the current suicidal ideation score (BSI-C, β = 9.189, t = 2.004, p = 0.048; earnings, β = 0.001, t = 0.582, p = 0.562). The lack of association for earnings is consistent with the task design, in which there is no stable optimal policy and payouts are only a coarse proxy for decision quality. Future work in learning paradigms, where optimality is well defined, may be better suited to test earning-based links to STB. We have clarified this point below:

      Page 32:

      “Second, although we assumed that increased risky behavior in STB was suboptimal, the current task was not suited to test this, given the task design of random feedback for gambling option. Future work in learning paradigms, where optimality is well defined, may be better suited to test earnings-based links to STB.”

      (8) Line 290: "beta_gain: -1-1" is unclear. I believe you meant beta_gain \in [-1,1].

      Thank you for this suggestion. We have now corrected it to make it clear.

      (9) The gain and loss biases are modeled as minimum and maximum probabilities for choosing the gamble. This is a legitimate choice for value-agnostic biases, but it is not the traditional choice (as far as I know). I wonder if the same results would hold with the more traditional formulation of the bias as an added constant to the utility of the gamble, i.e., p(gamble) = 1/(1+ exp(-mu(U_gamble + beta_gain - U_certain)). I believe in this case, you would also not have to specify different equations for positive or negative biases, or to limit the bias to the range of [-1,1] (indeed, the bias would be in reward-equivalent units).

      Thank you for this suggestion. The winning choice model we used here was consistent with previous literature (Rutledge et al., 2015 & 2016), which decomposed the decision process into risk-attitude-driven valuation (e.g., loss and risk aversion) and value-insensitive motivational components. These approach/avoidance parameters are a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference.

      As suggested, we also compared the traditional bias choice model. Model comparison did not support this. Please see Supplementary Page 4.

      (10) Also, for equations 5-8, it seems that 5-6 are identical to 7-8 except for the use of beta_gain versus beta_loss. You might want to consider simplifying by putting beta in the equations and specifying in the text that, depending on the trial type (loss or gain), the relevant beta is used.

      Thank you for this suggestion. We have now simplified it. Please see our revision below:

      (11) It is not clear what equations are applied to mixed trials in cM3.

      Sorry for the confusion. We have now clarified this point.

      Page 12:

      “Approach/avoidance parameters are not applied to in mixed trials.”

      (12) Model comparison: the mood models are nested within each other (e.g., mM3 can be derived from mM1 by setting beta_EV = beta_RPE). In this case, model comparison can use the likelihood ratio test instead of BIC, which can be too conservative (and therefore does not support the extra beta parameter for RPE, different from previous results in the literature). I wonder if a likelihood ratio test would lead to results more in line with previous findings with this task?

      Thanks for this suggestion. We agree that mM1 (CR+EV+RPE) and mM3 (CR+GR) are nested. However, our model space also included unnested models, such as mM5 (CR+GR<sub>better</sub>+GR<sub>worse</sub>). Therefore, it was not reasonable in our model space to use likelihood ratio tests.

      (13) Line 346: The replication sample is described as "healthy participants," however, their health (or mental health) status was not assessed, and they may as well have mental health concerns. I would suggest calling this a general sample or an undifferentiated sample - but not a healthy sample.

      Sorry for the confusion. We have now corrected this phrase.

      (14) Line 363: "in addition to the replication of previous findings in the validation dataset" is unclear. Are those tests not two-tailed?

      Sorry for the unclear statement. In the replication analyses, we used one-tailed t-tests because the direction of the effect was revealed on the clinical dataset. Please see our clarification below:

      Page 15:

      “For the replication of previous findings in the validation dataset, we used one-tailed tests in line with our clinically motivated directional hypothesis.”

      (15) Line 372: "validating our group manipulation" - the presented work does not have a manipulation. Maybe you meant "validating our grouping of participants"?

      Thank you for this suggestion. We have now corrected it to make it clear.

      (16) Figure 2B: It is not clear how the data were binned for illustration purposes only, and why this binning is necessary (I have not seen it in other papers) - presenting the data from each subject and the correlation line with error margins (as is done here) should be sufficient.

      Thank you for flagging this. For illustration only, we binned the data proportional to group sizes: in the patient sample (S<sup>-</sup> n = 25; S<sup>+</sup> n = 58; ≈1:2), we displayed 3 bins for S<sup>-</sup> and 6 bins for S<sup>+</sup>. We agree that binning is not necessary; all statistics were computed on raw, unbinned data. The binned panel was included solely for visualization, consistent with our prior work (Blain et al., 2023).

      (17) Table 2: delta BIC should be presented per subject (that is, divided by the number of subjects in each group), as the groups are of different sizes, so as presented now, the columns are not comparable across groups.

      Thank you for the helpful suggestion. Our goal in Table 2 is not to compare ΔBIC magnitudes across groups, but to identify the winning model within each group. The ΔBICs are aggregated at the group level solely to rank models for that group. Dividing by the number of participants would rescale each group’s column by a constant and would therefore not affect the within-group ranking or the conclusion that cM3 is the best model in all groups. For this reason, we retain the current presentation and interpret each column within group rather than across groups.

      (18) Line 640 - the effect of expectations and prediction errors on mood was not only shown in healthy people, but also in people with depression (Rutledge et al., 2007, https://pubmed.ncbi.nlm.nih.gov/28678984/)

      Thank you for this comment. Indeed, Rutledge et al., (2017) showed evidence for CR+EV+RPE mood model in adult people with depression. However, our study recruited adolescents with depression or anxiety, given that adolescent period might provide a developmental window for opportunities for early intervention of suicidality. Therefore, it is also possible that the current winning model was specific to adolescents. Please see our clarifications below:

      Page 28:

      “It is also possible that the current winning model was specific to adolescents. Given that Rutledge et al., (2017) supported the “CR-EV-RPE model” in adults with depression, our study with adolescent populations may suggest a developmental change for mood sensitivities.”

      (19) Supplemental material: Is the R2 section about R-squared? Perhaps you can use superscript on the 2 to make that clearer? For Figure S2, how was model recovery determined? Should I interpret the confusion matrix as suggesting that the winning model for each and every simulated subject was the generating model, or was the winning model determined for the whole simulated population in each of the 100 simulations? Traditionally, confusion matrices use the former measure, but the results of 100% recoverability make me suspect the latter was used here. In Figure S3, should we not be looking at simulated parameters and recovered parameters? What are "real parameters" here?

      Thank you for these important comments. We now consistently denote the coefficient of determination as R<sup>2</sup> (with a superscript 2) throughout the manuscript and Supplementary Materials.

      For the model recovery analysis in Figure S2, we have clarified that the confusion matrix is computed at the population level. Specifically, for each of the 100 simulations we generated a full dataset under each candidate model, fit all models to that dataset, and selected the winning model based on group-level model evidence (BIC). Each cell in the confusion matrix therefore reflects the proportion of simulations in which model j was selected as the best-fitting model when the data were generated by model i. This operation was reasonable because the decision of the winning model is made on the population-level dataset rather than on individual subjects.

      In Figure S3, the term “real parameters” referred to the parameters used to generate the simulated data. To avoid confusion, we now relabel these as “simulated (generating) parameters” and explicitly describe the figure as showing the relationship between simulated (generating) parameters and recovered parameters. Please see Supplementary Pages 2-3:

      “Model recovery: We generated 100 simulated datasets for each model (3 choice models and 8 mood models) using the fitted parameters of each model as the ground truth. Each dataset contained 201 trials and included 3 (or 8) sets of simulated data corresponding to the respective models. For each simulated dataset, we then fit all models and determined the winning model at the population level based on group-level BIC, yielding a confusion matrix in which each entry represents the proportion of simulations in which model j was selected as the best-fitting model when the data were generated by model i. As shown in Figure S2, all models are highly identifiable, indicating excellent recovery performance for both the choice and mood models.”

      “Parameter recovery: Figure S3 shows good parameter recovery for both choice and mood winning model (choice: rs > 0.91, ps < 0.001; intraclass coefficients > 0.78; mood: rs > 0.90, ps < 0.001; intraclass coefficients > 0.86). Moreover, we computed cross-correlations between all generating (“generating”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery.”

      Typos:

      (1) Line 90: original → originate

      (2) Line 596-598 - the same phrase is repeated twice.

      (3) Line 616: on the other word → hand.

      Sorry for the mistakes. We have now corrected them throughout the manuscript.

      Reviewer #2 (Recommendations for the authors):

      For people unfamiliar with interpersonal theory or motivational-volitional model, or three-step theory (lines 105-106), could you briefly explain the key idea of mood and suicide before going to the decision-making tasks? And from this, maybe motivate the predictions in your task? In particular, in the abstract and introduction, the phrasing could be a bit more concise and simpler. In the abstract, sentences were sometimes quite long. In the introduction, some paragraphs are somewhat repetitive. In the discussion, there were some typos.

      Thank you for these suggestions. We have now explained the key idea of mood and suicide before going to the decision-making tasks in the introduction, which can be seen below:

      Pages 4-5:

      “Contemporary theories of suicide converge on the idea that STB is initially caused by low mood experience. The interpersonal theory of suicide proposes that suicidal desire arises when people simultaneously feel socially disconnected (“thwarted belongingness”) and like a burden on others (“perceived burdensomeness”), experiences that are tightly linked to chronically low mood [25]. The motivational–volitional model [26] and the three-step theory [27,28] similarly emphasize that when negative mood and feelings of defeat or entrapment are experienced as inescapable, they can give rise to suicidal ideation, and that the progression from ideation to suicide attempts depends on additional factors such as reduced fear of death, increased pain tolerance, and a tendency to act impulsively under intense affect. Some official organizations, e.g., National Institute of Mental Health, have also listed mood problems as warning signals [8]. Interestingly, within the framework of decision making under uncertainty, gambling on lotteries with a revealed outcome has been found to induce high mood variance [29], providing an opportunity to assess the relationship between deficient mood and increased gambling decisions in STB.”

      We have also refined the wording and corrected typos throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Since many readers might only read the abstract, it is important that it is both informative and accurate. I have two suggestions in this respect. First, for the abstract to be more informative, it may be helpful to indicate already there that these are value-insensitive approach-avoidance parameters, in the sense that they favor/disfavor the gamble regardless of the potential outcomes' magnitude or probability. This issue is also present throughout the text, where the phrases "approach and avoidance motivation" are referred to as if they have established and precise computational definitions. In my view, these terms could just as easily be interpreted as parameters that multiply the value of potential gains or losses, which is not what the authors mean. It would be helpful to clarify this terminology.

      Thank you for these suggestions. In line with previous literature (Rutledge et al., 2015 & 2016), approach and avoidance motivation are indeed defined at the computational level, referring to a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. We have cited these papers in the manuscript. We also make it clear to further clarify approach and avoidance parameters in the abstract and introduction. Please see our revisions below:

      Page 2 (Abstract):

      “Using a prospect theory model enhanced with value-insensitive approach-avoidance parameters revealed that this rise in risky behavior resulted only from a heightened approach parameter in S<sup>+</sup>.”

      “Altogether, model-based choice data analysis indicated dysfunction in the approach system in S<sup>+</sup>, leading to greater propensity for gambling in the gain domain regardless of the lottery expected value.”

      Page 3 (Introduction):

      “A growing literature indeed shows that risky behavior can be far better explained after adding value-insensitive approach and avoidance components to prospect theory [18,19], that is by including a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. This class of models highlights the important role of value-insensitive motivational components in decision making in addition to risk attitude-driven valuation (e.g., loss/risk aversion) [20].”

      (2) The statement "our study uncovers the cognitive and affective mechanisms contributing to increased risk behavior in STB" is overstating the findings, as the study may have uncovered some contributing mechanisms, but likely not all of them. Removing the word "the" would fix this issue.

      Thank you for this suggestion. We have now corrected it.

      (3) Since mood is typically defined as lasting hours, it's inappropriate to refer to ratings that only reflect the last few trials as self-reports of mood. To be sure, I view the distinction between emotions and moods as quantitative, not qualitative, so I do not think there is a problem studying the former to understand the latter, but to avoid confusion, the terminology should follow common usage.

      Thank you for this suggestion. We follow previous work and operational definitions regarding mood (Rutledge et al., 2014, Eldar & Niv, 2015, Vinckier et al., 2018). Emotion is usually a very brief response to a specific stimulus (Emanuel & Eldar, 2023), e.g., leading to rapid changes like surprise then fear. In contrast, mood is defined as a diffuse state that is not specific to one stimulus. Here, we operationally and computationally define mood as an affective state reflecting the recent history of safe and gamble outcomes. We now clarify that point in the main text. Please see our revision below:

      Page 5:

      “Although mood is thought to persist for hours, days, or even weeks [30–33], momentary mood, measured over the timescale in the laboratory setting, represents the accumulation of the impact of multiple events at the scale of minutes [30,32,34–38]. Momentary mood external validity is demonstrated e.g., through its association with depression symptoms [37]. Mood is different from emotions, which reflect immediate affective reactivity and is more transient (e.g. from surprise to fear) [31–33,39].”

      (4) Line 78: The phrases "increase in risk attitude", "decrease in loss attitude", and "decrease in value-independent choice biases" are unclear to me in terms of their directionality. An attitude might be avoidant or embracing. If it is the former then increasing it would decrease risk-taking.

      Thank you for pointing out the ambiguity. We have now corrected them throughout the manuscript. Please see our revision below:

      Page 4:

      “We therefore hypothesized that heightened approach motivation, or weakened avoidance motivation, would account for increased risk behavior in STB.”

      (5) Line 125: I was not sure why one would expect the mood response to gamble-related quantities (EV and RPE) to be lower in STB and not higher.

      Sorry for the typo. We hypothesized that mood would respond more strongly to gambling-related quantities expected value (EV) and reward prediction error (RPE)—in adolescents with STB than in controls, given prior evidence that STB is associated with greater risk-taking.

      (6) The text could use proofreading, as there are many typos. These are from the first 100 lines alone:

      (a) Abstract: regardless the lotteries -> regardless of the lotteries'.

      (b) Line 78: it remains whether.

      (c) Line 80: can each -> each can.

      (d) Line 90: may original from.

      Sorry for the mistakes. We have now corrected them throughout the manuscript.

      (7) The rationale for focusing on the S+ group for mood model comparison is incorrect. The purpose is to identify parameters that vary as a function of suicidality, and for that, the S- group is just as important.

      Thank you for this comment. We agree that the S<sup>-</sup> group is as important as the S<sup>+</sup> group. A direct comparison was complicated because the winning mood models differed (S<sup>+</sup>: mM3; S<sup>-</sup>: mM5; Table 3). To ensure comparability, we checked results from both model specifications (mM3 and mM5). The conclusions were convergent: mood sensitivity to certain rewards (CR) was lower in S<sup>+</sup> than in S<sup>-</sup> (see Fig. 3 for mM3 and Fig. S8 for mM5).

      (8) There appears to be a contradiction between the inclusion criteria, which include having experienced suicidal thoughts and behaviors, and the definition of the S- group as not having suicidality.

      Thank you for pointing out this mistake. The corrected version of inclusion criteria can be seen on Page 7:

      “Patients were included if they met the following criteria: 1) both the researcher and psychiatrists agreed on their group classification; 2) they had a current diagnosis of major depressive disorder (MDD; unipolar depression), generalized anxiety disorder (GAD), or bipolar disorder with depressive episodes (BD), confirmed by two experienced psychiatrists using the Structured Clinical Interview for DSM-IV-TR-Patient Edition (SCID-P, 2/2001 revision; see Supplementary Note 1 for details);3) they were between 10 and 19 years of age; 4) they had no organic brain disorders, intellectual disability, or head trauma; 5) they had no history of substance abuse; 6) they had no experience of electroconvulsive therapy.”

      (9) It would be helpful to specify whether mood modeling was based on objective or subjective values, and why.

      Thank you for this helpful suggestion. We have now clarified whether mood modeling was based on objective or subjective values, and why. Specifically, we constructed two model families: one in which mood was driven by objective monetary outcomes (objective values) and one in which mood was driven by subjective values derived from each participant’s fitted choice model (subjective values). We then used the VBA_groupBMC function in the VBA toolbox to perform family-wise model comparison, with 8 candidate mood models within each family. Consistent with previous literature, the objective-value family provided a clearly superior fit to the data (exceedance probability, EP = 1.000). Based on this result and for parsimony, we report and interpret the mood modeling results from the objective-value family in the main text. We have clarified this point in Supplementary Note 9.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents an interesting approach for finding electrophysiological models that match experimental patch-clamp data. The authors develop a new method for deriving optimized current clamp protocols by training a neural network on synthetic data. This optimized current clamp is then used on both computational training data and on experimental data to predict current gating and conductance parameters that correctly reconstruct the electrical phenotype.

      Strengths:

      (1) The fitting of gating variables through an optimized patch clamp protocol is interesting.

      (2) The inclusion of experimental data is important, and the approach is shown to be effective in fitting them.

      Weaknesses:

      (1) Some clarity is necessary on the generation and selection of variable IPSC models. With such a large variation in so many parameters, I would expect some resulting parameters to generate non-realistic phenotypes, quiescent cells, etc. Are all 200,000 or 1,100,000 generated cells viable? Or are they selected somehow for realistic cell properties?

      Thank you for this important point. We agree that broad parameter variation can generate non-physiological model behavior. Indeed, with the +/-40% perturbation range, some simulated cells produced non-realistic outputs, including quiescent behavior, and failure to generate a complete action potential. These cases were excluded from the dataset. As a result, only cells exhibiting physiologically meaningful and numerically stable behavior were retained for further analysis. We have clarified this selection procedure in the Methods section. We applied a large variation to ensure that all possible combinations and morphologies were included in the training and testing data so the model would readily ingest new data and perform robustly.

      (2) The error shown in Figure 4 between different population sizes is not completely explained in the text - there seems to be a minimal difference between a population of 1,000 and 10,000, followed by a very good fit at 200,000. Is there a particular threshold that needs to be crossed where the error drops off? Related, how was the 200,000 number chosen?

      Thank you for this observation. We agree that the decrease in error shows a gradual performance improvement as the population size increases, rather than a strict cutoff. As shown in Figure 4, the difference between 1,000 and 10,000 samples is small, but as we continue to increase and get to around 200,000 samples, we see strong error minimization. This indicates how much training data is needed for optimal model performance. This improvement is due to better coverage of the high-dimensional parameter space, which helps the network learn the nonlinear relationships between the parameters and outputs.

      We tested a range of training data sets and found that above 200,000 training data sets, the model consistently produced low, stable errors and good test-training agreement. The test error decreased with the training error as the population size increased, indicating better generalization and suggesting that the model accurately predicts unseen data rather than overfitting to the training set.

      (3) Related to the point above, the 1,100,000 population for fitting experimental data also needs a more complete explanation: how was this number chosen, and how does the error compare with the other population sizes shown in Figure 4?

      Thank you for this question. We found that at a training data set size of 1,100,000 we were able to cover the large parameter space induced by +/-40% parameter perturbation. iPSC-CM measurements are known to exhibit high variability, and we wanted to capture the full range in the training data set so the model could ingest a wide range of experimental data. It is trivial to generate new training data, for example, to capture different experimental conditions like temperature differences, mutations, drugs, or ionic variability. We view this flexibility as a substantial strength of the approach. But the large perturbations we show in this study (+/-40%) allow the generation of a very broad range of cellular phenotypes while maintaining physiologically realistic ionic current properties and action potential behavior. Consistent with Figure 4, increasing population size reduces prediction error and improves generalization. The larger dataset provided more stable, accurate predictions when fitting experimental data, without evidence of overfitting.

      (4) Why are the optimized current clamp protocols different between panels A and B in Figure 5? Are they somehow informed by experimental data?

      Thank you for this question. The stimulation protocol used in panels A and B is identical. Panels A and B show whole-cell currents recorded under the same stimulation conditions as in Figure 3. The differences reflect variability in the underlying whole-cell ionic currents of the model cells rather than differences in the applied protocol. This is exactly the idea: the exact same protocol will generate different whole-cell currents in individual cells, but the model can find parameter sets for all of them.

      (5) Figure 6D: Is the EAD risk in panel D specific to cell 1, 2, or the pooled variants of both?

      Thank you for this question. We have clarified this point in the revised manuscript. The EAD risk shown in panel D is computed from the pooled variants of both Cell 1 and Cell 2, rather than being specific to either cell individually.

      (6) How sensitive is the fitting to minor parameter variation? Further, if one were to pick, let's say, the next-best-fitting value, would that fall close to the best one? Is the solution found unique, or are there multiple sets with good fits?

      Traditional optimization methods, such as Nelder–Mead, directly fit the model to the observed data by iteratively minimizing the error for each dataset. As a result, the solution can depend on the initial parameter guess and may converge to different local minima. In contrast, our approach trains a deep learning model on synthetic data generated from the baseline model, learning a mapping from whole-cell currents to the corresponding 52-parameter sets by minimizing prediction error. The mean squared error (MSE) decreases from approximately 10⁻² to below 10⁻³, with training and test errors overlapping closely, indicating stable training, good generalization, and accurate reproduction of the observed signals.

      The model achieves very low MSE and reproduces the electrophysiological outputs with high fidelity. However, accurate reproduction of the outputs does not imply a unique parameter solution. This is illustrated in Figure S1, where baseline and predicted parameter values show close agreement overall, yet small deviations persist across parameters. This indicates that different parameter combinations can yield similar whole-cell behaviors due to parameter correlations and compensatory effects. In such cases, the model learns to predict a representative parameter set that is most consistent with the training data and loss function, rather than converging to a single unique solution within a fixed numerical tolerance.

      Reviewer #2 (Public review):

      Summary:

      The authors present a computational framework for generating "cell-specific" digital twins of human iPSC-CMs from a single optimized voltage clamp recording. Using deep learning trained on > 1 million artificial cells, the authors demonstrate that the model can infer 52 biophysical parameters governing 6 major ionic currents, and the resulting digital twins can reproduce experimentally recorded action potentials.

      Strengths:

      The framework has clear potential for understanding cellular heterogeneity in iPSC-CMs, predicting individual drug responses, and reducing the experimental burden of multiple patch clamp protocols.

      Weaknesses:

      There are several concerns about the validation of the model and its clarity. First, the biological variability being modeled in this manuscript is not defined well. It is unclear whether the framework addresses cell-to-cell differences within a single differentiation batch, variability across iPSC lines, or donor-to-donor differences. This ambiguity makes it difficult to interpret what the "digital twin populations" actually represent biologically. Second, the main claim, "the digital twins enable drug testing and arrhythmia prediction that would be impractical experimentally", is not experimentally validated. For example, the E-4031 simulations predict EAD rates, but no direct experimental head-to-head comparison is provided to confirm that these predictions are accurate. Third, technical reproducibility and biological representativeness are not assessed. Single voltage clamp recordings are inherently noisy. Without knowing how much variability comes from the recording process (technical variation) vs true biological differences, it is difficult to judge whether observed "cell-specific" parameter differences are meaningful. In addition, the optimized protocol is claimed to be superior to conventional approaches, but again, no experimental comparison is shown.

      The authors should address these concerns, with particular emphasis on clarifying the biological context and providing direct experimental validation. Below are detailed specific points:

      (1) Ambiguous definition of iPSC-CM heterogeneity. The authors model "typical iPSC-CM heterogeneity" by varying 52 parameters +/- 40% around a baseline model (Figure 1), generating > 1 million synthetic cells. However, the manuscript does not clearly state what biological variability this model is intended to capture. Is this modeling within-line, cell-to-cell variability (e.g., cells from the same dish or differentiation batch that differ due to stochastic gene expression or maturation state)? Or is this modeling between-line or between-donor variability (e.g., genetic background differences, reprogramming efficiency)? This distinction is critical for interpretation. If the goal is to understand why different cells in the same dish behave differently, then training data should reflect that. If the goal is to compare patient lines or disease models, the framework needs validation across multiple donors or lines.

      For example, the experimental validation in Figure 5 uses a single iPSC line (iPS-6-9-9T.B), but how many differentiation batches or dishes were tested, or whether cells came from the same preparation are unclear. Another example is that the wide AP diversity in the training population (Figure 1A) is impressive, but there is no demonstration that real experimental cells actually fall within this assumption range of +/- 40%.

      From a biological perspective, iPSC-CMs are known to be highly heterogeneous within lines (maturation state, metabolic differences, epigenetic variation, spatial differences within the same dish, etc) and between lines (different donor/genetic background). Thus, please explicitly state whether the +/- 40% variation is intended to model within-line or between-line heterogeneity, and justify this choice with wet experiment data (or reference to experimental literature on iPSC-CM variability). Please clarify how many dishes, differentiation batches, and time points post-differentiation were used for experimental recordings (Figures 5-6). If the framework is intended to generalize across lines from different donors, please test the model on multiple independent iPSC lines (from different donors).

      Thank you for this important and insightful comment. The selected ±40% range was chosen to broadly explore all physiologically plausible electrophysiological behaviors, not to match a specific experimental distribution. Our goal was to cover enough behaviors for the model to learn a reliable mapping between responses and ionic parameters.

      We recognize that this approach does not explicitly account for variability between lines or donors. We have a current project focused on extending the framework to include multiple iPSC-CMs from patient donors, but given that the model framework successfully reproduces such a broad range of cell phenotypes, we feel confident that it will readily apply to different genetic backgrounds from patient-specific cells. This study is underway.

      We have updated the manuscript to clarify how the modeled variability is interpreted and added a discussion of these limitations. Furthermore, we clarified the experimental conditions, such as the number of differentiation batches and recording settings, in the revised Methods section.

      (2) Biological representativeness of single-cell measurements.

      The framework generates digital twins from single voltage clamp recordings. The patch clamp recordings in iPSC-CMs are subject to substantial technical variability. The manuscript does not address a fundamental question: "How representative are the measurements from a single cell on the dish (or line)?" In other words, if I measure one cell from a dish of a million cells, does that cell's digital twin tell me something about the dish as a whole, or just about that one cell? The manuscript presents Cell 1 and Cell 2 (Figures 5-6) as distinct individuals, but it's unclear whether these differences reflect true biological heterogeneity or simply sampling variability. I think the authors should perform replicate recordings on multiple cells (e.g., > 10 cells) from the same dish (same differentiation batch) and quantify how much the inferred parameters vary, and then compare between lines.

      Thank you for this important comment. We agree that the representativeness of single-cell measurements and the impact of technical variability are important considerations in interpreting the results. In this study, the framework is designed to generate digital twins that reflect the electrophysiological properties of individual recorded cells, rather than to directly represent the behavior of the entire cell population within a dish.

      As such, differences observed between Cell 1 and Cell 2 are intended to reflect variability at the single-cell level, which may arise from a combination of biological heterogeneity and experimental variability. We agree that systematic replicate recordings across multiple cells are valuable to quantify the relative contributions of biological and technical variability, and to assess the consistency of inferred parameters. However, this is beyond the scope of the current study. We have added clarification in the manuscript to explicitly state this limitation and to outline this as an important direction for future work.

      (3) No experimental validation of the main claim that in silico populations can replace wet experiments.

      The most exciting claim in the manuscript is that digital twins enable drug testing and arrhythmia prediction "at scale" without requiring hundreds of patch clamp experiments. Specifically, the authors show that in silico populations derived from two experimental cells (Figure 6C) predict dose-dependent EAD incidence for the IKr blocker E-4031 (Figure 6D), with ~3% of cells showing EADs at 50 nM.

      However, this prediction is not validated experimentally. If I actually patch 20-30 real iPSC-CMs and apply 50 nM E-4031, will ~3% of them show EADs, as the model predicts? Without this validation, I think the drug testing framework is purely hypothetical. The model may be internally consistent (e.g., Cell 1's twin behaves differently from Cell 2's twin), but there is no evidence that these in silico populations reflect real biological variability in drug response. Please provide experimental validation that justifies the prediction by digital twins.

      Thank you for this important comment. We agree that experimental validation of population-level drug response will be valuable for establishing the quantitative accuracy of the predicted EAD incidence. The E-4031 simulations are intended as a proof-of-concept illustrating how the framework can identify susceptible subpopulations and quantify relative proarrhythmic risk in silico. We agree that direct comparison with large-scale experimental datasets is a key next step, and we are working hard to get the study funded so that we can perform those experiments and bring this technology to scale.

      (4) Experimental validation and head-to-head comparison of optimized protocol.

      The authors claim that their deep learning-optimized voltage clamp protocol (Figure 3, Figure 4A) is superior to conventional approaches, but they have not validated this experimentally by doing a head-to-head comparison. The manuscript does not compare the optimized protocol to any published voltage clamp designs. If the optimized protocol is genuinely easier to implement and more informative than existing approaches, this would be a major practical advance. But without side-by-side comparison, it is impossible to judge whether the optimization made a real difference.

      Thank you for your comment. We agree that comparing directly with traditional voltage-clamp protocols through experiments would be useful. In this study, our main aim was to show that the optimized protocol enhances parameter inference within the modeling framework, not to prove experimental superiority. We have clarified this point in the revised version.

      Reviewer #3 (Public review):

      Summary:

      This work uses a convolutional neural network to optimize a voltage clamp protocol to identify features and parameters from human pluripotent stem cell-derived cardiomyocytes.

      Yang et al. introduce an innovative experimental framework that integrates computational modeling and deep learning to generate a digital twin of human pluripotent stem cell-derived cardiomyocytes (hPSC-CMs).

      Strengths:

      The major strength is the methodology used to bridge in silico prediction of cell behavior and mechanistic insights from the experimental dataset.

      The approach used in this study represents a significant step toward precision medicine by enabling in silico prediction of cellular behavior and mechanistic insight from experimental datasets. The study addresses an important and timely challenge in stem cell-based and personalized medicine, and the authors compellingly leverage state-of-the-art methods alongside strong expertise in computational modeling and cardiac electrophysiology

      Weaknesses:

      While the overall approach is highly compelling and the potential impact is substantial, there are two areas where clarification and refinement, particularly in the phrasing and framing used throughout the manuscript, would further strengthen the work.

      (1) While the overall goal of the study is compelling, the manuscript would benefit from clearer articulation of how the proposed framework is intended to be used in practice. In particular, it is not entirely clear whether the authors envision this approach as:

      (a) a method to extract population-level trends that, when paired with biological data, enhance statistical power and interpretability, or

      (b) a strategy capable of constructing a population-based model from limited single-cell recordings. If the latter is intended, additional guidance on the number of action potentials required per cell and the assumptions underlying this extrapolation would greatly clarify the scope and applicability of the method.

      Thank you for this thoughtful comment. We agree that the intended use of the framework should be more clearly articulated. In this study, we generate a large synthetic population of iPSC-CM models by varying 52 biophysical parameters governing key ionic currents. A neural network is trained on simulated whole-cell current responses to learn a mapping between current profiles and model parameters. Experimental recordings are then used as inputs to this trained model to infer ionic parameters, rather than directly fitting the model to data. This enables individual recordings to be interpreted within a large, physiologically plausible parameter space and supports population-level analysis of electrophysiological variability. The primary goal of the framework is therefore to facilitate mechanistic interpretation of variability and relate experimental observations to underlying ionic currents. But the longer-term intended goal is to develop digital twins from patient-derived cell lines and then use populations constructed from patient-specific digital twins to screen therapeutics and identify arrhythmia marker vulnerability in a very thorough and high-throughput way. We have clarified this in the revised manuscript.

      (2) The manuscript would also benefit from a clearer explanation of how electrophysiological heterogeneity observed in hPSC-CMs is linked to inter-patient variability. Although the authors state that this framework can be generalized to compare patient-specific hiPSC-CM lines, it remains unclear how this generalization is achieved, given the substantial sources of variability intrinsic to hiPSC-CMs (e.g., batch effects, reprogramming strategy, differentiation protocol, and maturation state). As acknowledged by the authors, addressing this level of variability likely requires large datasets; further clarification of how the proposed approach mitigates or accommodates these challenges would strengthen the translational claims.

      Below are my suggestions that could help strengthen the claims in the manuscript:

      (1) Adding a dedicated section describing the electrophysiological phenotype of the hPSC-CMs used in this study would help justify the choice of the underlying ionic model and the selection of the six ion currents analyzed. These currents are not only developmentally regulated but may also vary substantially across different hPSC-CM lines, which has implications for generalizability.

      Thank you for this important suggestion. We agree that providing additional context on the electrophysiological phenotype of the hPSC-CMs strengthens the rationale for both the underlying ionic model and the selection of currents analyzed.

      We have expanded the Methods section to clarify this point. Briefly, the ionic currents were selected based on the Kernik-Clancy iPSC-CM model developed in our prior work, which was specifically designed to capture the range of electrophysiological variability observed within an iPSC-CM cell line using a population-based framework. In this model, variation in key ionic conductances is sufficient to reproduce the diversity of action potential morphologies, spontaneous activity, and repolarization dynamics commonly reported experimentally, while avoiding non-physiological behaviors.

      Accordingly, we focused on six primary ionic currents that are known to play dominant roles in shaping action potential characteristics and variability in iPSC-CMs. This selection reflects a balance between model parsimony and physiological relevance, enabling the framework to capture the expected spectrum of variability within a given cell line. We also note that the framework is extensible, and additional currents or alternative parameterizations can be incorporated to account for differences across cell lines, donors, or experimental conditions in future studies. See updated discussion.

      (2) If feasible, inclusion of patch-clamp data from an additional hPSC-CM line would significantly strengthen the claim that this framework can harmonize and generalize across datasets and cell sources.

      Thank you for this helpful suggestion. We agree that adding data from more hPSC-CM lines would improve the framework's generalizability. In this work, our goal was to show that the digital twin framework is data-driven and can easily be expanded to include more hPSC-CM lines, allowing for cross-line comparisons in future studies. We have clarified this and included a discussion of this limitation in the revised manuscript. We are currently seeking funding for patient-specific lines as well to allow scalability.

      (3) The authors note that the experimental cells exhibited high variability in action potential morphology. This is an important observation that directly supports the motivation for the study and should be explicitly presented, even if only in the supplementary materials.

      Thank you for this suggestion. We agree that explicitly showing the variability in experimental action potential morphology strengthens the motivation for this study. We have now added a section in the discussion discussing this and referencing the many prior studies that focused on iPSC-CM variability, including the studies upon which our initial model (Kernik-Clancy) was based.

      (4) In the hERG-blocker experiments, further clarification is needed regarding the biological relevance of the reported 3% incidence of early after depolarizations (EADs). Additionally, an interrupted sentence in this section makes it unclear whether the goal is to demonstrate that the digital twin can capture rare arrhythmic risk events or whether the digital twin is necessary to determine whether this level of risk is clinically meaningful.

      Thank you for this important comment. We agree that more clarification is needed on the ~3% EAD incidence and the digital-twin role. This analysis aims to show that electrophysiological variability can create a small, susceptible subpopulation under drug effects, not to set a clinical risk threshold. The observed ~3% EAD incidence reflects the emergence of such a susceptible subpopulation under hERG block. While relatively small, this fraction is important because it arises from modest, physiologically plausible variation in ionic properties and would be difficult to capture using single-cell or small-sample approaches. As described in the Discussion, this variability-driven emergence of EADs provides a quantitative measure of proarrhythmic risk at the population level. The digital-twin framework enables systematic identification and quantification of these rare events, linking cell-level variability to population-level responses. We have revised the manuscript to clarify this point.

      (5) The manuscript states that some action potentials were excluded from the experimental dataset. A brief explanation of the exclusion criteria, along with guidance on how to distinguish high-quality from low-quality recordings, would improve transparency and reproducibility.

      Thank you for this comment. We agree that the definition of failed recordings should be clarified. We have now specified the exclusion criteria in the Methods section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It would be helpful if the network cartoon in Figures 2 and 3 were replaced with a simplified sketch of the actual neural network used.

      Thank you. We now have new figures 2 and 3.

      (2) Subsection title for the Introduction has a typo.

      Thank you. We have fixed it.

      Reviewer #2 (Recommendations for the authors):

      (1) Technical quality control criteria are not specified.

      The Methods section states that "any incomplete or failed recordings were excluded," but does not define what constitutes a failed recording. The criteria could be subjective.

      Thank you for pointing this out. We agree that the definition of failed recordings should be clarified. We have now specified the exclusion criteria in the Methods section.

      “Recordings were excluded if they exhibited no spontaneous firing, abnormally slow firing rates, or failed to capture a complete action potential waveform. These criteria were applied consistently across all recordings.”

      (2) "Cell-specific" may overstate the claim.

      The term "cell-specific digital twins" (title, throughout) implies that the inferred parameters reflect the true biological state of each cell. However, parameters are derived only from curve-fitting to electrophysiological data and do not reflect other biological components (e.g., gene expression, contractility, calcium handling, metabolism, etc). Please consider rephrasing to "electrophysiology-based digital twins", "voltage clamp-matched digital twins", etc.

      Thank you for this important comment. We agree that the term “cell-specific” could be interpreted as implying a complete representation of the biological state of each cell. We have also adjusted the wording in relevant sections to avoid over-interpretation.

      Reviewer #3 (Recommendations for the authors):

      (1) I would add the list of the 52 parameters in the method section/SI and not just in the reference. Additional justification of why the perturbation was set as +/- 40% for the 52 parameter or +/- 20% for the EAD population would also help.

      Thank you for this helpful comment. We have included model equations and highlighted the 52 parameters in the Supplementary Information and provided additional justification in the Methods.

      (2) In Figure 1B, might be helpful to add the axis of the Vm instead of the dotted line indicating 0 mV to show differences in the diastolic potential.

      Thank you! We have now updated Figure 1B.

      (3) Figure 1C-I might be more impactful to show traces from the AP shown in Figure B to reinforce the impact of a single current in the AP shape.

      We have now updated Figure 1C-I to include traces from the AP shown in Figure 1B.

    1. In short, you are testing the compatibility of your schemata with the new people you encounter. Although storytelling will continue to play a part in your relational development with these new people, you may be surprised at how quickly you start telling stories with your new friends about things that have happened since you met.

      Story telling is one of my favorite ways to bond with a new friend or coworker. I find it fun not just to share but also to listen. If I think the story is on topic or I think they would find interesting, I like to share. I find this helps build a sense of knowing each other deeper even if we don't get to spend that much time together. As a listener, I get a chance to see what they are like in different scenarios.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental, yet poorly understood biological process.

      Weaknesses:

      A major weakness is that all the authors' conclusions are based on descriptive studies, in which the role of cell fusion is not directly tested. This is particularly important because other models of wound induced polyploidization have demonstrated that another cytoskeletal protein, myosin, was upregulated and dependent on endoreplication, and not cell fusion. Therefore it remains unclear to what extent cell fusion, endoreplication, or both are required to outcompete mononucleated cells as well as pool actin as described in this study.

      We thank the reviewer for appreciating our live imaging and meticulous approach. In this revision we have identified that the gene Atg1 is required for wound-induced fusion in the pupal notum: when Atg1 is knocked down, there is a reduction in wound-induced cell fusions, both border breakdown and cell shrinking. Analysis of Atg1 knockdown shows that the wounds close more slowly. This is a direct test of the role of cell fusion in speeding wound closure, presented in new Fig. 4.

      Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laserinduced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFPpositive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to sycytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Weaknesses:

      The authors provide some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound. The authors suggest that the syncytial cells might be better able to close the wound. However, some genetic studies would need to be done to establish this more convincingly. E.g. Could the authors genetically block syncytia formation and then show that these wounds now heal slower?

      We now present such data in new Fig. 4, which describes knocking down Atg1, previously shown by the Leptin lab to promote wound-induced fusions in larval epidermis. We quantify the resulting reduction in fusion in the pupal notum and show that the leading edge advances more slowly to heal the wound.

      The authors suggest that radial border breakdown reduces the requirement for cell intercalation. While this might be true it also raises the question of how the various syncytia facing the wound border change shape to allow the shrinkage of the first cell row over time to allow wound closure. None of the four movies included in the study shows the whole wound healing process until the later stages, making it hard to assess this. It would be good to include one such movie showing the syncytia in the whole wound and comment on this point.

      In response to the reviewer's request, we now extend Supplemental Video S1 out through 8 hours after wounding (same video as included previously but extended longer). In this video, as in many of the wounds, it is hard to determine the exact moment of closure because a syncytium extends across the wound whereas the nuclei do not. However, during the process of closure, one can clearly observe the large syncytia becoming more wedge-shaped – drastically reducing the section of their perimeter remaining in contact with the wound’s leading edge.

      In addition, we now explore how syncytia reduce the need for intercalation in a computational model, presented in new Fig. 7 and Supplemental Videos S5 and S6. One can observe the modeled syncytia becoming similarly wedge-shaped. The modeling shows that the presence of syncytia and their ability to reshape can speed closure by about 1/3 even if the syncytia have no special properties aside from their relative size.

      In both the experiments and models, some syncytia are also removed from the leading edge by intercalation, but the presence of syncytia reduces the total number of intercalations needed.

      The authors hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. They show convincingly through the fusion of a single Actin-GFP-positive cell in the second cell row with a GFP-negative cell in the first cell row that Actin-GFP spreads in the fused cell and labels the previously unlabelled actomyosin cable. While the hypothesis of resource sharing to improve healing is intriguing and makes sense, this experiment doesn't necessarily prove the benefit of resource sharing. It does show cytoplasmic mixing following fusion, now allowing the GFPlabelled actin to diffuse and be incorporated into the actomyosin cable. In a wild-type condition, fusion would not increase the total concentration of resources, although it would increase the total amount of resources within this bigger fused cell. The question is whether resource sharing without increasing the protein concentration is beneficial and increases the efficiency of certain wound healing mechanisms. There might be a benefit of cell fusion, if for example certain resources were only present in limited amounts or if protein transport could increase the concentration locally. To provide better evidence for the hypothesis that resource sharing improves wound healing, maybe the authors could look at the actomyosin cable in a wounded epithelium (such as in Figure 4E, F), in which all cells express MyoII-GFP. The authors could compare the average intensity of the actomyosin cable at the wound edge in mononucleated cells versus in syncytia. If resource sharing is indeed beneficial, it might be that the actomyosin cable is stronger/brighter in syncytia or it forms quicker.

      We agree with the reviewer that we have not "proved the benefit of resource sharing". Because we cannot inhibit resource sharing while still allowing cell fusion, we can think of no rigorous way to test this hypothesis. We appreciate the reviewer's suggestion of quantifying the myosin at the leading edge cable, but we can imagine too many caveats to the interpretation to make it worthwhile. Rather, we accept the limitation that this is an untested, perhaps untestable, hypothesis -- but nevertheless intriguing.

      We do want to clarify ideas about the concentration of resources after fusion. We agree that the overall concentration of a given resource (mass/volume) throughout a syncytium would be the same as the overall concentration in the unfused progenitor cells; however, a syncytium would have a larger total resource mass to direct subcellularly, allowing for local subcellular concentration to be greater in a syncytium vs. an unfused cell. We demonstrate this subcellular localization of actin in a syncytium twice, in Fig. 7C and E (previously Fig. 6C,E), which we think is evidence for increased local concentration.

      The biggest limitation of this study is that the authors don't address how the formation of these syncytia is regulated. While the manuscript in its current form provides some valuable new insights into syncytial-driven wound closure, it would be much more informative if it also provided some mechanistic details. The authors could test if some of the mechanisms shown to regulate syncytial formation in other types of syncytia-driven wound healing are also involved here. E.g. Yorkie was shown to negatively regulate cell fusion in adult syncytial-driven wound closure (Losick et al 2013). The authors could test for the effect of Yorkie-RNAi in the epithelium on wound closure and syncytia formation. Expression of the dominant negative RacN17 also blocked cell fusion in adult syncytial-driven wound closure (Losick et al 2013).

      Moreover, JNK activation was shown to be needed in larval syncytial-driven wound closure (Galko and Krasnow 2004). The authors could test JNK pathway reporters to assess pathway activation or test if the JNK pathway is needed for syncytial-driven wound closure by expressing a dominantnegative form of Basket JNK in the epithelium.

      Or could syncytia formation be regulated by changes in Integrin-mediated adhesion as shown by the Galko lab in Wang et al 2015? They show that wounding provoked a striking relocalization of PINCH and ILK, indicating the disassembly of functional FA complexes concomitant with syncytium formation. Maybe the authors could investigate some of these.

      We investigated the role of JNK in fusion by expressing bsk<sup>DN</sup> on one side of the wound. Comparing the numbers of border-loss fusion on each side, we did not find a significant difference in our seven-sample cohort (see Author response image 1). If we had increased the sample size, we may have found a significant difference with a small effect size, but because of the small difference in fusions on each side we did not think this was worth pursuing. Instead, we include data that the autophagy gene Atg1 is required for cell fusion in new Fig. 4, which begins to address mechanism, and relates the wound-induced fusion described here in pupae to wound-induced fusion shown in larvae. A complete mechanism for wound-induced fusion is outside the scope of this paper, as we focus on the function of syncytia in healing wounds.

      Author response image 1.

      Another general question that the authors raise but don't address enough is whether syncytia-driven wound closure in proliferation-competent epithelia is any different from the one in post-mitotic, polyploid epithelia. Since the mechanism regulating the former is not known, this remains unclear.

      We now include a paragraph on this question in the discussion.

      Finally, it is not clear, whether syncytia in these proliferation-competent epithelia get resolved after wound healing. Do they get removed and replaced by mononucleated proliferation-competent cells or do the syncytia stay in the epithelium like a scar? The authors should provide some images of wound areas a few hours after wound closure is complete and comment on this.

      To answer the reviewer’s question: some but not all syncytia do get removed during wound closure by remarkable apoptotic/extrusion events. This will be the subject of a future manuscript, as it is outside the scope of this paper focusing on the function of syncytia in promoting wound healing.

      Minor points:

      Figure 3: It would be better to have the microcopy images alongside the quantifications.

      The images in Figs. 1 and 2 show the border breakdown and shrinking cells, and we do not see benefit in adding them in Fig. 3.

      Figure 4A: The syncytium at the wound edge here doesn't look straight but wavy. Does it not form an actomyosin cable that straightens the front? Or are there lamellipodia/filopodia?

      We assume the reviewer is asking about the wavy edge outlined at 400 min after wounding (now Fig. 5A). As shown by Jacinto and colleagues in the first pupal wounding paper (JCB 2013), the actin cable forms quickly, within 15 minutes; much later actin protrusions extend from the leading edge to close the wound. This result is consistent with the wavy edge 400 min after wounding.

      248: The authors suggest an interesting hypothesis that mitochondria or ER could be pooled in fused cells. It would be nice to see some evidence: e.g. by labeling mitochondria and assessing where they are in syncytia versus mononucleated cells and whether they are concentrated around the wound edge.

      Although we don't think that exploring mitochondria or ER is central to this manuscript, we agree it would be an interesting question for the future.

      141-145 (Figure 4B and C) This example is not completely convincing. First, it is hard to see where the wound edge is. Second, it would be good to include an even later time point when the cell is clearly no longer at the wound edge.

      We have revised this figure, now Fig. 5B,C, to include a later image at 360 min after wounding healing, and this additional panel clarifies that the smaller cell leaves the wound edge. As noted in the text, the wound edge is indicated by the cell borders lacking p120ctn.

      Reviewer #3 (Public Review):

      Summary:

      White et al. described laser-induced wound healing of the Drosophila pupal notum. They found that the epithelial monolayer is dynamically induced to form syncytia by cell-cell fusion as an important part of repair. They reveal two processes: cell shrinking and border breakage that occur as part of syncytia formation. Expression of GFP in the cytoplasms of some epithelial cells reveals that cytoplasmic contents mix following injury and the GFP rapidly diffuses between cells. Using live imaging they observe that syncytia expand towards the wound, maintain their positions close to the leading edge, and apparently displace smaller cells. They propose that syncytia redistribute cellular components towards the wound facilitating repair and show that labelled actin becomes concentrated at the leading edge.

      Strengths:

      The manuscript is interesting and on an important and emerging topic of wound healing in a genetically tractable organism. The manuscript is very well written.

      Weaknesses:

      There are three major issues that the authors must address: 1. Is cell-cell fusion sufficient to enhance/facilitate wound healing? 2. Characterization of "border breakdown"; Is this phenomenon disassembly of apical junctions following membrane fusion? 3. Are cells really shrinking or is it only the apical domains that "shrink" as the cells join the syncytium.

      We thank the reviewer for recognizing the importance of this topic. Our responses to the specific weaknesses are below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Major Components:

      (1) For syncytia measurements the nuclei are labeled with histone-GFP which is expressed in all cell types. How do you know the nuclei within the cell junctions are epithelial and not another cell type, such as immune cells recruited to the injury site? It would be helpful to verify the number of nuclei per cell using an epithelial-specific nuclear marker as well. This could be via epithelial Gal4-specific expression of a UAS-nls-GFP.

      This is an interesting point. In response to the reviewer's question, we investigated by doing the converse experiment, labeling immune cells with hml-Gal4, UAS-GFP, and observing what they do after wounding (analyzing six wounded pupae). They do get recruited to the wound, but they remain either in the wound center or at the basal side of the leading edge. Because they are labeled with cytoplasmic GFP, we would be able to ascertain whether they fused with epithelial cells because they would share their GFP with epithelial cells in the epithelial plane, and they did not. Thus we are confident that the many syncytial nuclei are not derived from immune cells. Our live tracking throughout the manuscript, and specifically of GFP-labeled clones, also supports our interpretation that syncytial nuclei derive from epithelial cells.

      (2) The manuscript focuses on cell fusion, but other mechanisms of cell enlargement have been observed to occur during wound healing via endoreplication. To what extent do epithelial cells in pupae notum endocycle or endomitosis post injury? It is unclear if the increase in syncytia size during a 1-2hr period could also be due to endomitosis, which would also increase nuclear number.

      Since the first submission of this manuscript, we published our results demonstrating limited wound-induced endoreplication after this type of explosive laser injury to the pupal notum (White et al, 2024, PMID: 38495588). We chose to publish this work separately because we could not offer the same degree of depth for endoreplication as we could for fusion: our pupal notum injury model is extremely well-suited to analyzing cell fusion and wound closure by live imaging; however, it is not particularly well-suited for analyzing endoreplication in fixed tissue. With respect to reviewer's question about endomitosis -- i.e. nuclear divisions that are not accompanied by cell divisions -- even after many years we have not observed an endomitosis event, which would be visible by live imaging, whereas we frequently and easily observe mitosis of diploid cells.

      (3) One of the major conclusions of this study is that cell fusion is necessary to pool resources at the leading edge. Therefore it is critical that authors identify a mechanism to inhibit cell fusion to test this assumption.

      We now include new Fig. 4, an analysis of the role of Atg1 in promoting wound-induced fusion and wound closure. These results build on the finding of the Leptin lab (Kakanj et al, 2022) that autophagy genes are required for fusion. Our results are consistent with the model that syncytia speed wound closure.

      (4) There is evidence that myosin increases in endoreplicating cells during wound healing hence it is, maybe equally - if not more - probable that the increase in resources (here actin-GFP) at the leading edge is dependent on endoreplication instead of cell fusion.

      Some of the new data we provide for this manuscript is a correlation between cell size and distance traveled, showing that larger cells travel more within the wound (Fig. 4F,G). Endoreplication would certainly be expected to contribute to increasing cell size, and our published 2024 data indicates that there can be one extra S-phase induced by these types of wounds. Doubling the genome is not a significant contribution to cell size compared to the 10s of nuclei we observe in syncytia from fusion. Nevertheless, we do not claim that actin is the only important resource that can be pooled subcelluarly for the benefit of the cell; we use it only as a proof-of-principle. Finally, we discuss the work on myosin in wound-induced endoreplicating cells (Losick and Duhaime, 2021).

      Reviewer #3 (Recommendations For The Authors):

      Major comments

      (1) Can induction of epithelial fusion enhance wound healing?

      Different epithelial cell-cell fusion processes have been well-characterized: i) Trophoblast fusion in the placenta mediated by Syncytins. ii) Viral induced cell-cell fusion mediated by diverse viral glycoproteins (e.g. gp41 from HIV, Hemaglutinin from Influenza, GP from Ebola, and G glycoprotein from VSV). iii) Epidermal, myoepithelial, and other epithelial cell-cell fusion in C. elegans mediated by EFF-1 and AFF-1. iv) Cell-cell fusion in the eye lens (unknown fusogens). The authors may want to compare and discuss the temporal dynamics and intermediates observed in the diverse processes of epithelial cell-cell fusion with the characterization of syncytia formation during wound healing of the Drosophila pupal notum. Since some of these characterized cell-cell fusogens can fuse heterologous cells, including Drosophila S2 cells (Shilagardi et al., 2013; https://pubmed.ncbi.nlm.nih.gov/23470732/), the authors may consider expressing these fusogens in Drosophila pupal notum before, during and after injury. This could determine whether syncytia formation is sufficient to stimulate efficient wound healing.

      We thank the reviewer for the suggestion of comparing and discussing temporal dynamics and intermediates observed in the many types of epithelial fusion that are well understood. Regretfully, we do not think this article is the right venue for such a complex discussion, especially since we have little by way of comparison in our own wound-induced fusion data. As for overexpression of fusogens, it is an intriguing idea to force cell fusion with a heterologous fusogen such as EFF-1 and then investigate any resulting changes in wound healing. However, since half the cells within 70 µm of the wound already fuse even without a heterologous fusogen, it seems unlikely we could meaningfully increase the level of cell fusion unless we expressed the fusogen universally, forcing the fusion of nearly all the epithelial cells as well as other cells throughout the body that express pnr-Gal4. Because the overexpression of EFF-1 in C .elegans results in lethality (PMID: 26854231), a widespread induction of fusion would be expected to cause other types of physiological problems that would interfere with the interpretation of wound closure rates. Further, the conditional expression tools in Drosophila allow excellent spatial control, but temporal control is still somewhat low-resolution, so that we would have difficulty expressing EFF-1 before, during, and after wounding at times that would be relevant to understanding wound healing.

      (2) The phenomenon of "border breakdowns" described here is not clear. The authors are probably studying the disassembly of the apical junctions following the initiation of membrane fusion and pore expansion. This should be clarified by using membrane labels to directly observe membrane fusion. Researchers have used electron microscopy and membrane fluorescent probes to follow cell-cell fusion. For example, GPI-mCherry, FM4-64, lipid-modified-GFPs (e.g. PH-domain fluorescently labeled proteins) DiO, DiI, and many others. See for example: Markosyan et al., 2016; https://pubmed.ncbi.nlm.nih.gov/26730950/; Mohler et al., 1998; https://pubmed.ncbi.nlm.nih.gov/9768364/; Meng et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32668210/.

      We agree completely with the reviewer, that border breakdowns represent the disassembly of apical junctions following initiation of membrane fusion and pore expansion. Direct evidence for this order of events is found in the video stills of Figure 1 panel I and video S2, which show that cytoplasmic GFP is transferred to the fusion partner 14 minutes before there is a visible decrease in the apical adherens junction marker p120ctn. The reproducibility of this order of events is documented in Fig. 3: among 107 GFP-labeled cells, 30 of them first visibly shared GFP with a fusion partner, and then 11/30 displayed border breakdown, 16/30 displayed cell shrinking, and 3/30 did not fuse. This last category is consistent with a fusion pore that closed rather than expanded productively. Although we have obtained TEM images of wound-induced fusion pores, these are included in another manuscript currently in revision and so cannot be included here, and further these EM images do not shed light on border breakdown per se, as only live imaging can establish the relationship between border breakdown and pore formation (GFP-sharing).

      (3) The observation of cell shrinking may be misleading. The process the authors describe as "cell shrinking" may involve shrinking of the apical domain, maintaining the cell volume. To clarify this process, the authors may simultaneously label the apical and basolateral domains. It is possible that fusion pore formation occurs in the basolateral, apical, or both domains. The apical shrinking could reflect the migration of the apical junctions following fusion. A similar process has been described in epidermal and vulval cells of C. elegans and other nematodes (Mohler et al., 1998; https://pubmed.ncbi.nlm.nih.gov/9768364/; Sharma-Kishore et al., 1999; https://pubmed.ncbi.nlm.nih.gov/9895317/; Kolotuev and Podbilewicz 2008; https://pubmed.ncbi.nlm.nih.gov/18031720/).

      We thank the reviewer for pointing out these examples of cell fusion in nematodes, and we now compare our findings to Mohler et al, 1998. In Fig. 2D, we specifically investigated what happened to the cell volume of these shrinking cells, and we hope we have now clarified both the text and the annotations on the figure to make our findings more clear. In the X-Z plane, the entire cell volume of two shrinking cells is visible from cytoplasmic GFP labeling. For both cells, the cytoplasmic volume moves laterally into the neighboring syncytia, appearing to initiate the movement from the basal-most area of the cell so that 150 minutes after wounding, both cells have a reduced apical footprint and only a whisp of apically-oriented cytoplasm, with the remainder of the cytoplasm having moved into the syncytia. These images make it clear that fusion is occuring, and that when the apical area disappears the corresponding cytoplasm has also moved into the territory of the neighboring syncytium. In response to the reviewer's suggestion, we did try labeling basolateral domains, but the fluorescent proteins we examined are not restricted to the basolateral domain and are difficult to interpret.

      Minor comments

      (1) Lines 40-43. Repair of injuries has also been observed in non-proliferative syncytial epidermal cells and involves cell-cell fusogens. The authors may want to include this reference: Meng et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32668210/.

      We thank the reviewer for the suggestion, and we have included this reference in the Discussion paragraph about fusogens.

      (2) Lines 128-130. Is "Shrinking fusion" an "artefact"?

      The apical junction shrinks not the cell. I suggest following basolateral membranes to see whether the cell is indeed shrinking as it fuses. The authors may want to share whether the cell volume is maintained but spills into an existing syncytium; the apical junction shrinks because it disappears/disassembles (see also Major comment 3).

      As discussed in Major comment 3, we do provide evidence that the cell cytoplasm spills into an existing syncytium. Perhaps the reviewer finds the term "shrinking cell" to be misleading, as we all agree that the cell contents do not disappear. We have updated the manuscript to use the term "apical shrinking" throughout.

      (3) Lines 157-159. Are these small cells or instead they are small apical junctions? The interpretation should include basolateral domains of the small cells to determine their size! It is also possible that some small cells have fused with the syncytia but on the basolateral domain without apical junction disassembly.

      We appreciate the reviewer's rigor. As noted above, we were not able to analyze the basolateral domains of these cells. Because our all analyses are live-imaging videos, we are able to identify the cells are undergoing apical shrinking and clearly delineate those from stable diploid cells. We now realize that the term "small cells" is confusing and can be mixed up with apical shrinking. These cells are not "small" but normal sized, small only in comparison with the gigantic syncytia around them. We have removed the term "small" from this description.

      (4) Lines 204-206. Many genes required for myoblast fusion in Drosophila have been shown to play a role in different stages of cell-cell fusion. Do they play roles in epithelia fusion during wound closure in the pupal notum?. For example, actin polymerization? Dynamin? Ig-domain and integrin cell adhesion machineries?

      We now provide a new Fig. 4 that shows that the autophagy gene Atg1 reduces wound-induced cell fusion, as it does in larvae (Kakanj et al, 2022), and importantly these wounds close more slowly. We have not analyzed mutants in actin polymerization because we are confident they would interrupt many aspects of wound healing. The Galko lab has identified that integrins suppress wound-induced cell fusion in larval epidermis, but we have not tested these. We have a manuscript in revision demonstrating a requirement for Dynamin and other endocytosis genes in wound-induced fusion, and without dynamin-mediated fusion, these wounds close more slowly.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This useful study presents an improved protocol for long-term in vitro culture of Schistosoma mansoni that enables progression toward sexually dimorphic stages, representing a meaningful advance for studying parasite development and reducing reliance on animal models. The findings show that host-specific culture conditions support essential developmental and metabolic functions required for parasite maturation, although development remains delayed compared to in vivo conditions. The evidence is solid overall, but limited pairing efficiency and the absence of egg production indicate that the system does not yet fully recapitulate complete reproductive development.

      On behalf of the co-authors, we thank the three reviewers and the editors for their complimentary remarks as well as the major and minor comments/ concerns. Addressing these concerns have led to revisions that improved the manuscript. In particular, further analyses have generated an updated Figures 3 and 4, and Supplementary Tables S1, and S4-S6.

      Public Reviews:

      Reviewer #1 (Public review):

      Pichon, Rémi et al. describe an in vitro method for transforming Schistosoma cercariae into mature adult worms. The authors show that human serum (HS) supports parasite growth and differentiation more effectively than fetal bovine serum (FBS). They also observed differences in parasite growth and activity, with worms cultured in HS efficiently digesting human red blood cells (hRBC). Cultured worms were able to pair with ex vivo adult worms and produce eggs, indicating functional maturation suitable for downstream applications such as drug screening. While the experimental approach is comprehensive and supports the advantage of HS culture conditions, the pairing efficiency was low (≈7%) and required long culture periods (70-80 days), highlighting limitations that may affect reproducibility.

      We acknowledge the reviewer for the positive highlights. Regarding the low in vitro pairing efficiency, we have now edited the manuscript to clarify a misleading statement related to 7%. We decided to remove the value of 7% — which corresponds to the percentage of experiments in which couples were observed, as it does not accurately represent the actual number of observed worm pairs and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff.:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      We also agree with the reviewer that the extended culture periods required to obtain fully sexually dimorphic parasites remain a limitation. As elaborated in Discussion (see below), key factors, probably derived from the host, are missing in the in vitro system explaining both the slow in vitro development and low rate of spontaneous pairing between in vitro developed, sexually dimorphic male and female worms. This was discussed as follows (lines 340-343): “That said, while our system was highly efficient in producing sexually dimorphic worms, spontaneous pairing between male and female parasites was extremely rare, mainly in aged in vitro cultures (from 80 to 100 days in culture) indicating that other factors, e.g., cholesterol, may be missing [35].”

      A major strength of the study, in particular, is that the authors clearly differentiate the effects of FBS versus HS on developmental progression. The conversion rate observed in HS cultures is significant and consistent with previously published data.

      While the study has several strengths, some aspects of the work are not fully explored. In particular, the role of hRBC supplementation requires further clarification. Although HScultured worms were shown to digest hRBC more readily, the implications of this observation remain unclear. Specifically, it would be useful to understand whether hRBC supplementation influences (1) long-term culture stability, (2) molecular pathways associated with development and differentiation, or (3) the pairing capacity of the worms. While addressing these questions may not be the main objective of the study, further discussion of these points would strengthen the manuscript.

      We agree that deciphering the role of the human Red Blood Cells (hRBCs) supplementation is critical. Regarding the influence of hRBCs on the long-term culture stability in parasite development it has been well established for more than four decades that schistosomes do need red blood cells to grow in culture [Basch, P. F. Cultivation of Schistosoma mansoni in vitro. II. production of infertile eggs by worm pairs cultured from cercariae. J Parasitol 67, 186-190 (1981); Basch, P. F. Cultivation of Schistosoma mansoni in vitro. I. Establishment of cultures from cercariae and development until pairing. J. Parasitol. 67, 179-185 (1981)]. The molecular pathways underlying development, sexual differentiation and pairing and modulated by hRBCs in culture is currently being investigated by our team. We decided not to include these data and analyses in the current manuscript, as they fall outside its scope.

      The manuscript is clearly written and represents a valuable contribution to the field. Overall, the experimental approach is sound, and the results support a useful methodological framework for the in vitro culture of Schistosoma worms and the attainment of sexual maturity, particularly for adult male worms.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Reviewer #2 (Public review):

      Summary:

      The authors perform confirmation studies of Paul Basch's seminal schistosome work from 1981, demonstrating the development of transformed schistosomules into sexually dimorphic adult parasites, albeit without successful egg production. In addition to the findings from Basch's earlier work, the authors add some new molecular data in the form of an analysis of proliferative cells in in-vitro-derived animals.

      Strengths:

      The authors successfully confirm experimental results from earlier schistosome researchers, providing a potential new tool for studying schistosome biology without the need for vertebrate hosts.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Weaknesses:

      The display of data from the authors is sometimes difficult to follow/understand where it comes from. For example:

      (1) Line 136: The authors claim that parasites in HS and FBS conditions have substantially different mortality rates (11.3 +/- 2.7 vs 5 +/- 2.3) but a quite high p-value (0.8). Analyzing the raw data myself, I obtained a mean of 8.2 +/- 1.7% vs 4.8% +/- 4.3% with a p-value of 0.15. Either the data are not clearly presented, and I did not follow them, or the data presented in the text do not match the raw data in the supplemental files.

      We thank the reviewer for pointing this out; we have now edited Supplementary Tables S1 and S6 by turning them into a long format for the sake of clarity. Accordingly, Results, Methods sections, and indicated supplementary tables were edited as follows:

      Results, lines 142 ff.:

      “No morphological differences were observed between parasites cultured either in FBS or HS within the first week in culture; in both conditions most parasites were classified as early schistosomula [category 1: 76% ± 30 (average ± SD) in FBS and 73% ± 29 (average ± SD) in HS] with few lung (category 2) and early liver schistosomula (category 3) (Figure 1B, week 1; Supplementary Figure S1). The mean mortality (category 0) at week 1 was slightly higher, but not statistically significant (P= 0.42), in worms cultured in HS [9.75% ± 2.76 (average ± SD)] compared to the mortality registered in FBS-cultured parasites [5.52% ± 5.18 (average ± SD), Supplementary Table S6], consistent with previous findings [39].”

      Methods, lines 463-465:

      “To evaluate differences in mortality between HS- and FBS-cultured parasites, data from 5 experiments were combined and analysed using a Shapiro-Wilk normality test to test normality of the data and a non-parametric Wilcoxon rank sum exact test (Supplementary Tables S1 and S6).”

      Supplementary Tables:

      Supplementary Table S1. “Raw counts of parasites within each developmental stage category. Each row corresponds to a picture of parasites in culture medium containing FBS or HS. Each column corresponds to the raw parasite counts at indicated stage development (categories 0 to 5), time in culture (Time in days - D), and experimental condition.”

      Supplementary Table S6. “Summary of all statistical tests employed in this study. 1. Statistical tests of parasite mortality and the raw data table used for this test. 2. Statistical tests for worm size comparisons (correspond to Figure 2). 3. Statistical tests for worm black gut comparisons (correspond to Figure 3). BG: Black gut. 4. Statistical tests for EdU positive cells comparisons (correspond to Figure 4). Replicate code: E, M and L correspond to day 2, 8 and 15 respectively; R and W correspond to the presence (R) or absence (W) of RBCs added 13 days after transformation.”

      For clarity, below we provide the R script used to perform the statistical tests on the data shown in Supplementary Table S6 (column ‘Raw count of parasite developmental category per image and experiment’)

      Author response image 1.

      (2) Line 187/Figure 4: Though it is not clearly stated, it appears that the authors treat their EdU counts as an ordinal data set of 61 steps (from 0 to >60) rather than a continuous measure of EdU+ cells per animal. In this author's opinion, the graph strongly suggests a continuous data set, and the fact that this reviewer had to dig through poorly-labeled raw data to discover the nature of the data is problematic. The authors should either switch to a continuous data set or make it explicit that the data shown are ordinal. If counting EdU+ cells is too arduous, the authors could consider comparing the amount of EdU+ area to the amount of DAPI+ area in maximum intensity projections of their confocal images, as this would roughly approximate the amount of proliferative cells in the animals.

      As the reviewer correctly pointed out, the data were treated as ordinal because counting worms with more than 60 Edu+ cells became extremely difficult and highly inaccurate. Therefore, we decided to group in a single category, “60 EdU+ cells”, all worms showing more than 60 EdU+ cells. We have now updated Figure 4 where medians are shown instead of media values, Supplementary Table S5 to provide more comprehensive access to the raw counts, and Supplementary Table S6 to indicate the data for EdU+ cells per worm were considered ordinal. Accordingly, we have revised the corresponding sections as follows:

      Results, lines 211 ff:

      “HS-cultured schistosomula showed higher numbers of proliferating stem cells, with a median of >48 and >60 EdU+ cells per worm at days 8 and 15, respectively (Figure 4). On the other hand, most FBS-cultured parasites displayed no more than an average of 20 EdU+ cells per worm (Figure 4).”

      Methods, lines 520 ff:

      “EdU+ cells per parasite were counted for an average of 100 parasites across three independent experiments (Supplementary Table S5). Worms were grouped based on the number of cells per individual, but all those showing ⪰ 60 EdU+ cells were counted in the same group named ‘60 EdU+ cells'. Therefore, the data were considered ordinal data. Statistical analysis was performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 considered significant (Supplementary Table S6).”

      Figure 4 legend, lines 830 ff:

      “A. Violin plots showing the number of Edu+ cells per worm at indicated time points (2, 8, and 15 days post cercarial transformation) in parasites cultured either in Foetal Bovine Serum (FBS, blue) or Human Serum (HS, light brown). Human Red Blood Cells (hRBCs) were added in the culture at day 13 post cercarial transformation. The small black dots indicate individual worms, and the big black point indicates the median of EdU+ cells per worm. All worms showing ⪰ 60 EdU+ cells were counted and clustered together in the group named ‘60 EdU+ cells’. Hence, the data were treated as ordinal and statistical analysis performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 (*) considered significant (Supplementary Tables S5 and S6).”

      We thank the reviewer for the very interesting suggestion to quantify cell proliferation by calculating the ratio between EdU+ area to DAPI+ area in maximum intensity projections images. Measuring the fluorescence area for each worm in maximum projection is an excellent idea; however, due to the number of EdU+ cells present in some samples, we think this technique would not provide additional information or produce more detailed data compared with our analysis when the number of Edu+ cells exceeds 60 per worm. We will certainly consider this approximation for future studies.

      There are some minor issues as well:

      (1) Line 122: It is perhaps incorrect to refer to humans as "the" definitive host of schistosomes, as S. japonicum is primarily considered a zoonotic infection with water buffalo/cows being the primary definitive host.

      We thank the reviewer for pointing this out; we have now replaced ‘schistosomes’ with ‘Schistosoma mansoni’ (current line 131)

      (2) Line 185/298: The authors refer to EdU pulse-chase experiments, but the experiments described here are EdU pulse experiments.

      This is a very good point, we thank the reviewer for bringing this up and have accordingly edited by replacing ‘EdU pulse-chase’ with ‘EdU pulse’ experiments in lines 37, 204, and 321.

      Reviewer #3 (Public review):

      Summary:

      This study is significant as it established a protocol for the long-term culture of Schistosoma mansoni newly transformed cercariae, which developed in vitro into sexually dimorphic forms. The impact of two different sera, Fetal Bovine Serum (FBS) and Human Serum (HS), added to the culture medium supplemented with human red blood cells was evaluated. The authors demonstrated that HS-cultured parasites were able to digest red blood cells, a critical step for long-term parasite development. Furthermore, while most FBS-cultured parasites did not progress beyond an early liver stage, sexual dimorphism was clearly evident in the HS-cultured worms, albeit delayed compared to in vivo development.

      Strengths:

      This study could contribute to further in vitro studies for a better understanding of the unique sexual biology of Schistosoma mansoni and for screening novel schistosomicidal compounds. By increasing parasite development in in vitro studies, this protocol could have a positive impact on the principles of the 3Rs (Replacement, Reduction and Refinement) for animal research.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Weaknesses:

      As the authors mentioned, "pairing between male and female parasites was rare. Pairing was observed in approximately ~7% of the experiments, usually after day ~ 80 in culture. Egg production was also not achieved with this protocol.

      Following the reviewer’s point and to clarify a misleading point, we have now decided to remove the value of 7% - which corresponds to the percentage of experiments in which couples were observed. However, this value does not accurately reflect the actual number of observed worm pairs, and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The manuscript is well-written overall. However, there are some minor revisions that would further improve the clarity and presentation of the data.

      (1) At the beginning of the manuscript, it would be helpful to clearly state three to four specific aims or objectives. This would help readers better understand the expected outcomes and the broader methodological contribution of the study.

      We agree with the reviewer and accordingly have stated the overall goals of the study, as follows:

      Introduction, lines 106 ff:

      “We aimed at optimising a platform to study intra-mammalian schistosomes that supports in vitro sexual dimorphism establishment, consequently leading to an overall positive impact in the 3Rs (Reduction, Replacement, Refinement) for animal research (https://nc3rs.org.uk/) [42]”.

      (2) In the abstract, you highlighted the relevance of the work according to the 3R principles of reduction in animal experimentation. However, this point is not clearly introduced in the Introduction section. Including a short discussion of this aspect would improve continuity and context.

      Following this and previous item raised by the reviewer, we have now clarified the potential impact in the 3Rs by our research outcomes and included that link to the NC3Rs website and a representative reference [Louis-Maerten E, Rodriguez Perez C, Cajiga RM, Persson K and Elger BS (2024). Conceptual foundations for a clarified meaning of the 3Rs principles in animal experimentation. Animal Welfare, 33, e37, 1–11)].

      (3) In line 43, please italicize Schistosoma spp.

      Edited accordingly.

      (4) When discussing the importance of "interfering with sexual development," in line 52, please specify the life cycle stages being referred to.

      Revised accordingly as follows:

      Introduction, lines 54-56:

      “This suggests that interfering with the sexual development of schistosome intra-mammalian stages could potentially restrict human pathology.”

      (5) Between lines 56-58, please rephrase this sentence for clarity.

      We thank the reviewer for this editorial suggestion. The text has been revised as follows:

      Introduction, lines 58 ff :

      “Therefore, novel control strategies are urgently needed, and new targets for drug/ vaccine development became a priority. A better understanding of the mechanisms underlying schistosome development, including sexual dimorphism establishment, will pave the wave to achieve this goal.”

      (6) In lines 66-68 & line 88, please clarify whether the transcriptomic studies cited were performed in vivo, in vitro, or ex vivo, and indicate the developmental stages analyzed.

      We have now included the information suggested by the reviewer as follows:

      Introduction, lines 69-70:

      “Transcriptomic studies, at both bulk [7-11] and single cell [12-1]4 levels for intra mammalian stages in vivo and ex vivo,...”

      (7) Please indicate, in line 110, the day of culture for reference. Without this information, the conversion rates per life cycle stage are difficult to interpret and reproduce. Overall, please try to give an overview in the text of these rates of conversion for context, wherever possible.

      Following the reviewer’s question, we have clearly indicated the in vitro and in vivo timings for ‘conversion’ (understood as sexual dimorphism establishment.) We have written:

      Introduction, lines 117-120:

      “Finally, while most of the FBS-cultured parasites did not progress beyond lung and early liver stage, HS-cultured parasites reached sexually dimorphic stages by week 6, albeit at a slightly delayed rate compared to in vivo development. In the mouse model, parasites become dimorphic by day 21 post-infection (~3 weeks) [12].”

      (8) The section beginning with "Furthermore, phenotypic...cell proliferation" (line 110) may be easier to follow if moved earlier in the Introduction.

      Following the reviewer’s suggestion, we have moved and slightly rewritten the sentence to current line 112, as follows: “First, phenotypic differences between FBS- and HS- cultured parasites became evident as early as 48 hours in culture, with HS-cultured parasites exhibiting higher rates of cell proliferation resulting in larger worms in the HS condition.”

      (9) In line 126, please remove the DOI and add the citation.

      Edited accordingly.

      (10) When referring to 10-week-old parasites, in line 130, please indicate the developmental stage at which they stalled and relate this to the phenotypic scoring shown in Figure 1.

      Based on this suggestion, we have now revised the third paragraph of Results section (‘Sexually dimorphic schistosomes developed entirely in vitro from cercariae’), as follows:

      Results, lines 137 ff.:

      “The development of schistosomula derived from mechanically transformed cercariae was assessed in at least 15 independent experiments, five of which were maintained over a period of at least 10 weeks to assess parasite survival and ability to mate and produce fertile eggs (Figure 1A; Supplementary Table S1).”

      Lines 151 ff.:

      “Differences in parasite development between the two conditions became apparent by week 2 (Figure 1B). At this time point, 14.8% ± 24.9 (average ± SD, excluding dead worms) or 36% ± 33.6 (average ± SD, excluding dead worms) of the parasites cultured in FBS or HS, respectively, have reached category 3, i.e., early liver schistosomulum. Parasites in FBS rarely progressed beyond this stage during the 10-week experiment, with very few parasites (<0.1% ± 0.2, average ± SD) reaching category 4, i.e., late liver schistosomulum. In contrast, worms cultured in HS developed over time across all categories, achieving marked sexual dimorphism by week 6 (13.4% ± 18.6, average ± SD) (Figure 1B; Supplementary Figure S3A), as confirmed by PCR (Supplementary Figure S3B; Supplementary Table S2). No differences in the timing for sexual dimorphism establishment were observed between male and female parasites. The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture. As previously described for the in vivo development of schistosomes [12], in vitro cultured parasites showed developmental asynchrony in agreement with Basch’s observations [33]; however, by week 10 most of the worms in HS (73.7% ± 25.4, average ± SD) acquired an evident sexual dimorphism (Figure 1B).”

      (11) In line 142, please provide a standard deviation value for the reported average of 14.8%, if available. As well as the absolute numbers of these parasites or indicate them in the supplementary. Otherwise, it is difficult to understand the true conversion rate.

      We followed the reviewer’s suggestions and have now rewritten the text (see above, item 10). In addition, Supplementary Table S1 was edited in long format (see answer for item 1, reviewer #2)

      (12) Please explain, IN line 144, why all cultures were maintained for 10 weeks and provide the rationale for this experimental design.

      We thank the reviewer for this opportunity to clarify this point and hence improve the manuscript. The experimental condition stopped at week 10 included only FBS-cultured worms, not HS-cultured parasites. This is relevant as most of the parasites in FBS were dead by this time, unlike the HS-developed schistosomes. Indeed, some experimental groups consisting of parasites cultured in HS were maintained for up to 22 weeks. We have now updated the text to clarify this point, as follows:

      Results, lines 160 ff.:

      “The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture.”

      (13) In lines 146-151, please streamline the timelines of culture conditions and observed outcomes in FBS versus HS media. As the current wording makes interpretation difficult.

      Following the reviewer’s suggestion we have streamlined the culture timelines and observed outcomes, as follows:

      Results, lines 137 ff.:

      “The development of schistosomula derived from mechanically transformed cercariae was assessed in at least 15 independent experiments, five of which were maintained over a period of at least 10 weeks to assess parasite survival and ability to mate and produce fertile eggs (Figure 1A; Supplementary Table S1).”

      Results, lines 151 ff.:

      “Differences in parasite development between the two conditions became apparent by week 2 (Figure 1B). At this time point, 14.8% ± 24.9 (average ± SD, excluding dead worms) or 36% ± 33.6 (average ± SD, excluding dead worms) of the parasites cultured in FBS or HS, respectively, have reached category 3, i.e., early liver schistosomulum. Parasites in FBS rarely progressed beyond this stage during the 10-week experiment, with very few parasites (<0.1% ± 0.2, average ± SD) reaching category 4, i.e., late liver schistosomulum. In contrast, worms cultured in HS developed over time across all categories, achieving marked sexual dimorphism by week 6 (13.4% ± 18.6, average ± SD) (Figure 1B; Supplementary Figure S3A), as confirmed by PCR (Supplementary Figure S3B; Supplementary Table S2). No differences in the timing for sexual dimorphism establishment were observed between male and female parasites. The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture. As previously described for the in vivo development of schistosomes [12], in vitro cultured parasites showed developmental asynchrony in agreement with Basch’s observations [33]; however, by week 10 most of the worms in HS (73.7% ± 25.4, average ± SD) acquired an evident sexual dimorphism (Figure 1B).”

      (14) In lines 153-159, please clarify comparisons between worms cultured in FBS and HS at equivalent time points (e.g., 2 weeks FBS vs 2 weeks HS), rather than comparing only 10 week cultures.

      Following the reviewer’s comment, we have now rewritten the whole third paragraph in Results, under the heading “Sexually dimorphic schistosomes developed entirely in vitro from cercariae” - changes detailed in answers to items 10 and 13 (above).

      (15) It would also be helpful to include information on male versus female development in the context of sexual dimorphism.

      This is a relevant point that we have not clarified in the original submission - we have now indicated in the text that no differences were detected in the timing for male and female dimorphism establishment. New text included as follows:

      Results, lines 159-160:

      “No differences in the timing for sexual dimorphism establishment were observed between male and female parasites.”

      (16) In line 163, please resolve the editing marks and punctuation.

      Resolved accordingly.

      (17) In lines 169 and 172, when referring to stages such as "early liver stage," please indicate the corresponding time in culture (e.g., 3 weeks, 7 weeks + 3 days), or define these stage classifications earlier in the manuscript.

      Following the reviewer’s suggestion we have now included the developmental category after stating ‘early liver stage’, as follows:

      Results, line 187:

      “Even though few parasites in FBS reached the early liver stage (category 3)…”

      (18) Please indicate, in line 173, the developmental stage of worms used when assessing hRBC digestion in HS and FBS cultures. Additionally, here, it would be useful to discuss how hRBC supplementation may influence worm development beyond culture conditions, including possible molecular mechanisms. As a revision, that way maybe you can include data, if already performed or conduct it, to show the effect of adding or not adding hRBC even in HS cultured worms.

      We thank the reviewer for highlighting this important item that warrants further clarification. As stated in Results washed human red blood cells (hRBCs) were added to the culture at day 13. Pilot experiments in which hRBCs were added at different time points had been previously performed; no hemoglobin digestion was apparent when hRBCs were added at days 4, 5 and 6 consistent with previous findings (Correnti JM, Jung E, Freitas TC, Pearce EJ. Transfection of Schistosoma mansoni by electroporation and the description of a new promoter sequence for transgene expression. Int J Parasitol. 2007 Aug;37(10):1107-15. doi: 10.1016/j.ijpara.2007.02.011. Epub 2007 Mar 18. PMID: 17482194.).

      Following this observation, we have added a line to clarify this point, as follows (lines 181187): “Based on both previous reports [45], and pilot experiments in which adding human Red Blood Cells (hRBCs) to the culture before day ~10 did not show obvious haemoglobin digestion, we decided to supplement the culture media with hRBCs at day 13. The addition of hRBCs allowed the parasites to feed and thus continue their development [19]. At this point, they began to swallow and degrade erythrocytes, producing hemozoin, a black pigment derived from host haemoglobin degradation and visible in the worms' intestines.”

      Regarding the specific effect of adding hRBCs in the culture, this is a very good point. First, it has been well established for more than four decades that schistosomes need red blood cells in culture to grow, as example see (Basch, P. F. Cultivation of Schistosoma mansoni in vitro. II. production of infertile eggs by worm pairs cultured from cercariae. J Parasitol 67, 186-190 (1981); Basch, P. F. Cultivation of Schistosoma mansoni in vitro. I. Establishment of cultures from cercariae and development until pairing. J. Parasitol. 67, 179-185 (1981). Second, we are currently analysing transcriptomic data from parasites cultured in different conditions, including in the presence or absence of hRBCs. We decided not to include these data and analyses in the current manuscript, as they fall outside its scope.

      (19) In line 183, please clarify whether the referenced single-cell transcriptomic data were obtained from adult worms.

      We have now clarified this point in the manuscript as follows:

      Results, lines 199 ff:

      “In schistosomes, a complex stem cell system consisting of both somatic and germline stem cells has been described by leveraging recent single cell transcriptomic data across different developmental stages, including schistosomula and adult worms [47].”

      (20) In lines 210 and 213, please indicate the absolute number of worms used for these observations, rather than only percentages. If possible, also report any sex bias in pairing.

      Following this and a similar item raised by reviewer #3 (public review), we decided to remove the mention of 7% given it is misleading. This percentage corresponds to the percentage of experiments in which couples were observed. However, this value does not accurately reflect the actual number of observed worm pairs, and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff.:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      (21) In the final results section, please clarify whether pairing enhances sexual maturation of already mature worms or whether maturation occurs primarily after pairing.

      This is a very relevant point, and we thank the reviewer for giving us the opportunity to clarify it in the manuscript. As described in the manuscript the parasite sexual dimorphism was established in vitro and developed male and female parasites were capable of pairing. Moreover, enlarged oocytes in the ovary’s posterior section of in vitro developed female parasites became apparent after pairing. This observation (Figure 6E, F and Supplementary Video S6) suggests that these female parasites, fully developed in HS-supplemented culture media, were not only capable of pairing, but of starting to fully maturate. We have clarified this aspect in the manuscript as follows:

      Results, lines 243 ff.:

      “Moreover, in vitro developed females coupled with ex vivo collected mature males displayed signs of primordial ovary maturation with larger oocytes towards the posterior region of the ovary (Figure 6E, F; Supplementary Video S6). On the other hand, females developed in vitro but not paired with ex vivo collected males remained immature.”

      (22) Further in the Materials and methods sections, please clarify, isn't 8000 schistosomula/well of a 6-well plate really a confluent culture condition, and does it contribute to NTS mortality in that way, as shown in previous in vitro transformation publications? Please clarify, at least with relative values, percentages of parasite transformation in such a concentrated system.

      No formal titration experiments were carried out but based on empirical observations during pilot experiments we decided to add no more than 8,000 schistosomula per well. This is something to further investigate in the future. We have now added the following sentence in Methods:

      Methods, lines 423-426:

      “The number of parasites cultured per well (~8,000 schistosomula) was determined empirically, as no formal titration experiments were performed. At higher densities (>10,000 per well), more frequent media changes were required, and parasite development appeared to be impaired.”

      (23) Also, what was the rationale of adding hRBCs as early as 13 days post-transformation, when the parasites are in the lung and early liver stage, just forming the guts? Therefore, is it possible that this would have contributed to the observation of lesser parasites disgesting hRBCs? Also, were the hRBC supplemented each time with the media change? This was not clear.

      We thank the reviewer for these questions. The rationale of adding hRBCs at day 13 has been elaborated above (question 18). In addition, in the mouse model, parasites have already migrated through and left the lungs by day 13 post-infection, as described by Nation et al [Nation CS, Da’dara AA, Marchant JK, Skelly PJ (2020) Schistosome migration in the definitive host. PLoS Negl Trop Dis 14(4): e0007951] as follows: “In the mouse, S. mansoni schistosomula begin to arrive in the lungs between 2 and 3 days post-infection, peaking at around day 7 and lasting until around day 11”. Hence, we do not think that adding hRBCs at day 13 contributed to the observation of fewer parasites digesting hemoglobin, because this was only seen in parasites cultured in FBS, not in HS.

      The hRBCs were replaced every two weeks, or sooner if their numbers decreased due to consumption. We have now clarified this point in Methods as follows (lines 427-430): “LTC medium was replaced twice a week and washed human red blood cells (hRBCs) added to a final concentration of 0.02% v/v at 13 days after transformation. Washed hRBCs were replaced every two weeks, or sooner if their numbers decreased due to consumption.”

      (24) In the Discussion, please address the limitations related to the relatively late onset and low frequency of pairing in vitro.

      Following the reviewer’s suggestion and comments from reviewer #1, we have now included a section in Discussion highlighting the limitations of the study and avenues to overcome these in the future.

      Discussion, line 360 ff.:

      “Considering these elements in future experiments will help overcome the limitations encountered in this study, including the low rate of spontaneous pairing between in vitro– developed male and female worms and the requirement for extended culture periods (>70 days). In addition, further research is needed to assess the role of host- and parasite-derived cues in schistosome development.”

      (25) Figure 1: Please consider adding arrows or markers indicating which parasites correspond to the representative developmental stages used for classification.

      We acknowledge the reviewer for the suggestion; however, we respectfully consider this may not be necessary as (1) the images shown in Figure are representative pictures of each time point included for illustrative purposes; (2) Supplementary Figure S1 clearly depicts representative images of worms in each developmental category associated with specific morphological descriptions. For greater clarity we have now added the following text at the end of Figure 1 legend:

      Figure 1 legend, line 810-811:

      “A detailed description of the developmental categories and representative images are provided in Supplementary Figure S1.”

      (26) Figure 2: This plot is somewhat misleading in showing that the HS cultured worms grew significantly more than the FBS worms, where the latter did not grow at all, as also shown by the blue bars all over the plot.

      We appreciate the reviewer’s observation; critically, the data shown in Figure 2 represent measurements of the worm's area, which means that some worms may have become longer but thinner maintaining the same area. Most of the FBS-cultured worms did not develop beyond lung or early liver stages, in which the parasites were long/ thin or shorter/wide, respectively. Therefore, the overall area of these FBS-cultured worms almost did not change (please see the raw data and statistical analyses in Supplementary Tables S3 and S6. We believe that, as presented, Figure 2 is sufficiently clear and self-explanatory. However, we would be happy to consider any suggestions to further clarify this point in the manuscript.

      (27) Figure 3: For panel A, what is the worm percentage corresponding to? The context is missing. Please clarify in the text.

      Following the reviewer’s question and for clarity, we have now (1) modified the axis-legend in Figure 3 as “Percentage of worms displaying or not Black Guts - BG (%)”, and (2) slightly edited the legend as follows:

      Figure 3 legend, lines 820-823:

      “Bar Plot representing the percentage of Human Serum (HS)- or Foetal Bovine Serum (FBS)-cultured schistosomula with (blue bar) or without (light brown bar) black guts (BG) due to the presence of intestinal hemozoin.”

      Reviewer #2 (Recommendations for the authors):

      The authors need to clarify their presentation of data. The raw data needs to be more clearly labeled/explained, and the representation of the data in Figure 4A needs to be explicitly described or changed.

      We acknowledge the reviewer for highlighting this issue related with the data presentation and have decided to follow their advice by editing Figures 3 and 4, and improving the data presentation in Supplementary Tables S1, and S4-S6. In particular:

      Figure 3. We have now modified the axis-legend as “Percentage of worms displaying or not Black Gut - BG (%)”, and slightly edited the legend as follows:

      Figure 3 legend, lines 820-823:

      “Bar Plot representing the percentage of Human Serum (HS)- or Foetal Bovine Serum (FBS)-cultured schistosomula with (blue bar) or without (light brown bar) black guts (BG) due to the presence of intestinal hemozoin.”

      Figure 4. We have edited this figure to show medians instead of media values, and updated the legend as follows: lines 830 ff.:

      “A. Violin plots showing the number of Edu+ cells per worm at indicated time points (2, 8, and 15 days post cercarial transformation) in parasites cultured either in Foetal Bovine Serum (FBS, blue) or Human Serum (HS, light brown). Human Red Blood Cells (hRBCs) were added in the culture at day 13 post cercarial transformation. The small black dots indicate individual worms, and the big black point indicates the median of EdU+ cells per worm. All worms showing ⪰ 60 EdU+ cells were counted and clustered together in the group named ‘60 EdU+ cells’. Hence, the data were treated as ordinal and statistical analysis performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 (*) considered significant (Supplementary Tables S5 and S6).”

      Supplementary Table S1. We have clarified the data presentation by turning it into a long format and updated the legend accordingly as follows (lines 864-867): “Raw counts of parasites within each developmental stage category. Each row corresponds to a picture of parasites in culture medium containing FBS or HS. Each column corresponds to the raw parasite counts at indicated stage development (categories 0 to 5), time in culture (Time in days - D), and experimental condition.”

      Supplementary Table S4. We have clarified the table by turning it into a long format, simplified the data presentation, and updated the legend accordingly as follows (lines 873874): “Percentage of parasites displaying either black positive (hemozoin) or black negative (no hemozoin) intestine.”

      Supplementary Table S5. We have simplified the table by turning it into a long format, and explained the naming for elements in columns C (‘Group’) and D (‘Replicate’). We have updated the legend accordingly as follows (line 876 ff.): “Raw counting of EdU positive cells per parasite for indicated experimental group, replicate and experiment in long format. The worms were classified by group (column C) and replicate (column D), using the following code: E (‘early’), M (‘medium’) and L (‘late’), corresponding to days 2, 8 and 15, respectively. R and W correspond to conditions with (R) or without (W) human red blood cells, and HS and FBS to culture medium employed.”

      Supplementary Table S6. We have incorporated a new section with the statistical analyses for parasite mortality estimation and updated the legend accordingly as follows (lines 882887): “Summary of all statistical tests employed in this study. 1. Statistical tests of parasite mortality and the raw data table used for this test. 2. Statistical tests for worm size comparisons (correspond to Figure 2). 3. Statistical tests for worm black gut comparisons (correspond to Figure 3). BG: Black gut. 4. Statistical tests for EdU positive cells comparisons (correspond to Figure 4). Replicate code: E, M and L correspond to day 2, 8 and 15 respectively; R and W correspond to the presence (R) or absence (W) of RBCs added 13 days after transformation.”

      Reviewer #3 (Recommendations for the authors):

      The study was well conducted, and the data presented clearly support the conclusions. The protocol is well described, making it reproducible. The pairing experiments could be improved.

      Specific Questions.

      (1) "Male and female adult worms that developed in vivo and recovered from mice by portal perfusion on day 42 post-infection were sorted by sex and placed in culture with worms of the opposite sex developed in vitro (>70 days). Within 24 hours of initiating the co-culturing of in vitro developed worms with ex vivo collected worms, couples were observed".

      In the interest of clarity, and considering that stating ‘worms developed in vivo were collected from infected mice’ is redundant, we have now shortened and edited these lines as follows (lines 238- 242): “Male and female adult worms were recovered from mice by portal perfusion on day 42 post-infection, sorted by sex and placed in culture with worms of the opposite sex developed in vitro. Within 24 hours of initiating the co-culturing of in vitrodeveloped worms with ex vivo collected worms, couples were observed (Figures 6C, D; Supplementary Video S5).”

      (2) Have the authors conducted experiments with in vitro female and male parasites under the same experimental conditions as the in vitro/ex vivo pairing experiments? Is it possible that the tissue culture medium used for the development of sexually dimorphic forms is inhibiting pairing?

      The reviewer raises an interesting point that warrants clarification. First, the experimental conditions tested for in vitro developed parasites were the same as for the pairing experiments, as the ex vivo collected worms were washed and placed in HS-supplemented media. Second, as the culture conditions were the same (same culture protocol and medium) between in vitro pairing and in vitro / ex vivo pairing experiments, we do not think that the tissue culture medium used for developing sexually dimorphic parasites inhibited the pairing. As elaborated in Discussion (see below), key factors, probably derived from the host, are missing in the in vitro system explaining the low rate of spontaneous pairing between in vitro developed, sexually dimorphic male and female worms. This was discussed as follows (lines 340-343): “That said, while our system was highly efficient in producing sexually dimorphic worms, spontaneous pairing between male and female parasites was extremely rare, mainly in aged in vitro cultures (from 80 to 100 days in culture) indicating that other factors, e.g., cholesterol, may be missing [35].”

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      General Statements

      We thank the reviewers for thoroughly reading our manuscript and their constructive feedback. We have considered each comment carefully and came up with a revision plan that can be found below.

      1. Description of the planned revisions

      Response to Reviewer #1, #2 and #3 concerning alternative assays to measure the effects of ER stress on ribosome translation:

      Reviewer 1: 4. The modest reduction in the translation upon ER stress induction could be supported by alternative biochemical assays such as polysome profiling and amino acid incorporation.

      Reviewer 2: Is there a difference in the number polysomes in the non-stressed vs stressed yeast cells? A functional assay, such as polysome profiling combined with nascent-chain labeling might help to support the observation that an increased level of hibernating ribosomes exist in DTT or Tm-treated cells. Has this been considered? Perhaps there is evidence available in the literature?

      Reviewer 3: 5. The study mainly provides structural snapshots and population distributions of ribosomal states, without direct functional measurements of translation activity. Could the authors provide orthogonal biochemical or functional evidence supporting reduced translation under these exact stress conditions?

      We thank the reviewers for this suggestion. Other studies have shown a decrease in protein synthesis rate (Geronimo RAC et al., 2025 (PMID: 40959222), Pincus et al., 2014 (PMID: 25275008)) under similar conditions which matches our findings. However, we agree that confirming a reduction in translation for our specific conditions will strengthen our findings.

      In our lab we have previously performed polysome profiling for mammalian cells, we will adopt this protocol to yeast cells and use this to measure changes in monosome abundance and changes in the polysome to monosome ratio. We will perform this experiment for our main strain (Ire1cGFP) upon 0 hr, 45 min. and 4 hr. DTT treatment. Given the increase in hibernating ribosomes upon prolonged ER stress, we hypothesize that polysome profiles will reveal an increase in monosome subunits, assuming that the technique is sensitive enough to pick up moderate changes in translational activity. In case polysome profiling is not sensitive enough to pick up the moderate change in hibernation, we also aim to quantify the decrease in protein synthesis using C-35 labeling as we have done earlier in collaboration (Fedry et al, Mol. Cell 2024 (PMID: 38340715)).

      Reviewer #1 continued:

      1. For identification of hibernating ribosomes, the authors rely mainly on the presence of empty ribosomes along with eEF2, eEF5A, and eEF3. Whether these particles indeed possess known dormancy factors or they are different subclass of empty ribosomes is unclear. Similar analysis in the absence of dormancy factors would strengthen the authors claims.

      While empty 80S ribosomes (lacking tRNAs, elongation or hibernation factors) have been described in purified ribosome samples, those seem to be an in vitro re-association artifact as they are never observed in cells (bacteria: Xue et al. Nature 2022 (PMID: 36171285), yeast: Cheng et al. 2025 (PMID: 39789210), mammals: Xing et al. 2023 (PMID: 37410833), Fedry et al. 2024 (PMID: 38340715), etc.).

      Instead, in cells ribosomes are found in three possible states:

      1. individual subunits (40S and 60S in eukaryotes),
      2. translating 80S ribosomes; featuring a tRNA in the P-site
      3. hibernating 80S ribosomes; lacking a tRNA in the P-site and hence non translating. Those ribosomes are typically bound by eEF2 interacting with the dormancy factor bound in the mRNA channel, as well as possible additional factors (eIF5A, Dap1, SNOR, etc.). Our Hib class is seen in cells without tRNA in the P-site and bound by eEF2; we can therefore unambiguously assign this class as hibernating ribosomes.

      While doing a similar analysis on the translation landscape under ER stress in the absence of hibernating factors may yield some interesting insights, it would also alter the overall stress response and therefore may not help with the interpretation of our current structures. It would require us to repeat our complete workflow with new yeast strains (with Stm1 and/or Lso2 knocked out) and in our opinion the amount of time that these experiments would require do not justify the additional confidence that would be gained from the results. We have therefore decided that these experiments will be beyond the scope of this study. We will add a supplementary figure showing the presence of a density in the mRNA channel further supporting the presence of a dormancy factor interacting with eEF2.

      1. The authors suggest that ER induces modest level of increase in hibernating ribosomes. Adding controls such as glucose deprivation and nitrogen starvation would have provided more strength in relative comparison of these ribosomal sub populations.

      To our knowledge, no cryo-ET study has been done to study the effect of glucose deprivation and nitrogen starvation on the abundance of different translational states, making it an interesting and relevant experiment to do. However, it is unknown what change(s) in translational state abundance(s) these low-nutrient conditions might cause, so we are unsure if they could serve as control conditions. Collecting data on yeast under different stresses would require extensive resources, and would not directly address the translational response to ER stress. Therefore, we consider these suggested experiments beyond the scope of this work. We think this is an interesting future research direction and will comment on this in our discussion.

      We will include the suggested conditions in our polysome profiling experiments (proposed above in the first part of our revision plan) and analyse their monosome to polysome ratios. These can serve as positive controls for strong translation shutdown. We are grateful for the reviewers suggestion.

      1. The authors show the retention of dormant ribosomes on the ER surface. As usual notion of ribosome association with ER membrane to be dependent on nascent translation, retention of dormant ribosomes on ER membrane is interesting and puzzling. Analysis using strains deleted for dormancy factors may provide more insights on this mechanism.

      We agree with the reviewer that the presence of hibernating ribosomes on the ER surface is an interesting observation, but we do not consider it surprising. For yeast and mammals, idle ribosomes bound to Sec61 are well established in vitro (e.g., Becker et al (PMID: 19933108)), indicating that the interaction between these two components is not dependent on active translation. Furthermore, an average of an ER-bound hibernating ribosomes have been found on microsomes derived from human cells, and they become the prevalent form upon DDT-treatment, which strongly suggests that hibernating ribosomes can stay bound to the ER (Gemmer et al., 2023 (PMID: 36697828)). To clarify this point we will refer to these previous findings in our revised manuscript. As our observation is consistent with current knowledge in the field, we do not believe that additional analysis is necessary on this point.

      Reviewer #2 continued:

      The new Dec3 state might be clarified a bit further by zooming in to the corresponding areas in the Dec1 and Dec2 structures. This is a point of novelty in the paper and should be emphasized for future reference. Does an additional classification algorithm, such as cryo-DRGN-ET, verify the various states, especially the new Dec3 state? The structures should of course be uploaded to EMDB or another suitable server.

      We will provide an additional supplemental figure, zooming in on the eIF5a area in Dec1/2/3. Dec3 was found in 3 separate classification runs and we will therefore not perform classification with an alternative algorithm. We will upload the novel structures (Dec3, Hib and the ER-bound ribosome) to EMDB, which will be released upon publication of this manuscript

      There is generally a lack of supporting quantification, which will bother a number of readers. For example, a "high confidence rigid body fit" shows additional density in the hibernating state, but what is the confidence? Even the resolution measures of 7-8 Angstrom are simply stated. Presumably they come from a WARP report. There should be some specification for the evaluation. How many lamellae were used, and how many tomograms? Were they taken from different biological experiments, or all collected from the same grid, for each condition?

      We agree with the reviewer that this additional information is required to properly judge our conclusions. We will provide a confidence score for the eEF3 fit. We will provide FSC curves as supplemental data, specifying where the FSC curve was obtained from. Local resolution estimates are derived from Relion. Table 4 indicates the number of tomograms collected per sample and we will add the number of lamella/grids used.

      Reviewer #3 continued:

      Major comments:

      1. For the analysis of ER-bound ribosomes, the authors applied an ellipsoidal mask during subtomogram averaging. However, this masking strategy may not be sufficient because the relative orientation of ribosomes with respect to the ER membrane can be variable, and membrane density may influence particle alignment. The authors may consider including an additional masking step to exclude membrane density and minimize potential alignment bias.

      We thank the reviewer for pointing out this confusing point in our manuscript. The ellipsoid mask was only used in the image classification step aiming at separating ER-bound ribosomes from soluble ribosomes. The ER-bound ribosomes were subsequently aligned with a mask comprising the large ribosomal subunit and the membrane. This was crucial for the alignment not to go astray. We will clarify this in the text:

      “Alignment and averaging of these particles using a mask comprising the ribosomal large subunit and the membrane region yielded a ribosome with a clear membrane bilayer and an additional density at the exit tunnel”

      The signal coming from the ribosomal RNA is very strong (unlike single particles studies of smaller membrane proteins) and typically much stronger than the signal coming from the ER membrane. This strategy is well established in the field (Pfeffer et al. 2014 (PMID: 24407213), 2015 (PMID: 26411746), Braunger et al. 2018 (PMID: 29519914), Gemmer et al. 2023 (PMID: 36697828)).

      1. Supplementary Figures 1B and 1C appear to suggest that the Ire1i-GFP and Ire1i-NG strains exhibit stronger HAC1 splicing upon DTT treatment. Given this apparent increase in UPR activation, it would be interesting to analyze these strains as well to determine whether they display more pronounced changes in translational states.

      We thank the reviewer for raising this interesting point. All strains display ~25% of hibernating ribosomes under ER stress. The corresponding analysis can be found in Sup Figures 4 and 5. We will clarify this point by adding a sentence about this and the reference to the corresponding Sup figures: “First, we observed an increase in the relative abundance of hibernating ribosomes (from 3 to 25%) at the expense of some of the major elongating states, like Dec2 and Pre (from 22% to 16% and from 38% to 17%, (Fig. 3E-F, Supp. Fig. 4D-e). A similar effect was observed in the Ire1i-GFP and Ire1i-NG (Sup. Fig. 4, 5). This increase in inactive 80S complexes is indicative of a reduced translation activity in the cell, that is typically caused by the inhibition of translation initiation.”

      There is a difference in the magnitude of HAC1 splicing, but all have sufficiently high HAC1 splicing levels to robustly activate ER stress. This can explain why they all show a similar abundance in hibernating ribosomes.

      Minor comments:

      1. "FOV" should be defined as "Field of View" upon first use

      We will correct the corresponding sentence to: “Cells were then imaged with cryo-ET at an intermediate magnification (6.32-7.09 Å/pix, Field of View (FOV): ~9 µm2), allowing us to laterally capture near-complete cellular ultrastructure in each tomogram (Figure 1B-D)“

      1. In Supplementary Figure 7A, the image quality appears insufficient to clearly resolve structures within the autophagic bodies. As a result, it is difficult to determine whether ER-derived membranes are present within these structures. If ER-like membranes are observed, this could suggest induction of ER-phagy under ER stress conditions, consistent with previous reports (e.g., Mizuno et al., PLoS Genetics, 2020).

      The tomogram in supplemental Figure 7A does not contain obvious ER-derived membranes. We have observed membranes in other tomograms but our cryo-ET approach does not allow us to identify their origin (ER or other organelles). Therefore, we refrain from making any claims about ER-phagy in our manuscript and limit our discussion to the more general autophagy.

      1. In the sentence "Using this approach, we identified 7 distinct ribosome states," the authors should clearly specify which strains and treatment conditions were analyzed. Similarly, statements such as "A similar increase in Dec3 was seen for the other conditions" and "Overall, we observed a consistent, stress-independent increase of the Dec3 state at the ER for all Ire1c-GFP conditions" should explicitly define the corresponding conditions in the text.

      We agree that these statements are too vague and we will specify the corresponding conditions in each of these sentences in the revised manuscript.

      1. In the sentence "Finally, like for cytosolic ribosome states, we observed that upon ER stress the abundance of hibernating states at the ER increased over time at the expense of other translating states (Dec2 and Pre)," the authors should explicitly reference the corresponding figures.

      Agreed, we will refer to the Figure 4D for explicit comparison.

      1. In the References section, "Elife" should be corrected to "eLife" for the citation of van Anken et al.

      Agreed, we will correct this citation.

      1. The enrichment of eEF3 on inactive ribosomes leads the authors to propose a possible role for eEF3 in yeast ribosome hibernation improvement or keep. However, this interpretation currently appears speculative because the map resolution for the external density is relatively limited (~9-15 Å). Could the authors strengthen this claim by performing focused refinement/classification of the eEF3 density, testing eEF3 mutants or depletion strains or examining whether eEF3 occupancy changes quantitatively during stress progression?

      Based on the abundance and function of eEF3 we deemed eEF3 the most likely candidate to fit this external density. However, we agree that currently the strength of the claim does not match the strength of the evidence. We will try to improve the local resolution by performing a focused refinement on the eEF3 density. Though, eEF3 is a small density for cryo-ET. In this resolution range we are uncertain whether it will improve the quality of the map in this region. We will quantify the quality of the fit.

      Regarding the mutant/depletion strains, eEF3 is essential in yeast and it is also required for translation elongation. Hence depletion strains or interaction mutants will also perturb the role of eEF3 in translation. This strongly limits the possibilities to specifically investigate the functional importance of eEF3 for ribosome hibernation.

      1. The current study only examines translational states during ongoing stress exposure and does not investigate whether these changes are reversible after stress resolution.

      We thank the reviewer for this interesting point. We hypothesize that the modest translation decrease we observe upon ER stress is most likely reversible. We will check this by including a recovery condition in our polysome profiling experiment.

      2. Description of the revisions that have already been incorporated in the transferred manuscript

      No revisions have been carried out yet.

      3. Description of analyses that authors prefer not to carry out

      Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      1. The authors mainly focus on 80S particles in their analysis for suggesting the different states of ribosomes. However, there is a possibility of free subunits being stored under specific condition. Can the authors comment on free 40S and 60S subunits?

      The reviewer is correct and several factors binding free ribosomal subunits have been proposed to play a role in translation inhibition and ribosomal subunit hibernation (Saba et al. EMBO J 2024 (PMID: 39533057)). While we appreciate this interesting perspective, we believe that it is beyond the scope of the present work and that incorporating it would distract from the central focus of the manuscript, namely the effect of ER stress on translation elongation dynamics.

      However, if the polysome profiling experiments, now planned based on the suggestions of the reviewers, highlight a significant change in 40S and/or 60s subunit abundance relative to 80S we will try to re-analyze our data, focusing on 40S and 60S subunits.

      1. A previous study has reported the storage of dormant ribosomes on the mitochondrial surfaces. Analysis of mitochondria associated dormant ribosomes in S. cerevisiae would shed more light on this phenomenon.

      We thank the reviewer for raising this interesting point. Because of our focus on ER stress, our data collection was targeted at the ER, hence only a few of our tomograms contain mitochondria. On the few mitochondria that we did image, we do not observe lattice-like tethering of ribosomes, as described upon glucose starvation in S. Pombe (Gemin and Gluc et al. 2024 (PMID: 39379376)). This tethering is a novel observation, and its function still needs to be explored. Indeed, earlier experiments also indicated that glucose deprivation can induce ribosome binding to mitochondria in S. cerevisiae spheroplasts (Kellems et al., 1975 (PMID: 1092698)). However, initial experiments should first confirm whether this phenomenon also without conversion to spheroplasts before moving on to ER stress.

      The structural analysis of mitochondria-associated ribosomes upon ER stress would require new sample preparation of lamellae of control and DTT-treated yeast cells and data collection targeted at mitochondria. It is unlikely to be very different from the modest effect we describe in the cytosol and at the ER membrane. Finally it would distract from the main message of our manuscript centered on the impact of ER stress on translation dynamics. Hence, we consider these experiment beyond the scope of our current work.

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      Were the cryo-FM lamellae maps shown in Fig. S1F-I used to target the tomogram acquisitions? A correlation between the FM and the EM could provide hints about where the Ire1p clusters are located. The puncta are curious somehow, although established in the literature. I'm wondering if they appear somewhere in the tomograms. I would not insist on new experiments to find them, but it would make sense to show if they are already present in the data. Fig S1D does not show a lamella, and it is hard to conclude that the puncta are really absent there. As a general/historical comment, is it clear that the GFP does not affect the protein condensation?

      We thank the reviewer for this highly relevant and interesting question. Indeed, the cryo-FM data was used to collect tomograms targeted at Ire1p oligomers. However, none of the conditions (the 3 different cell lines, different timing and different type of stressors) showed detectable clusters in the tomograms.

      The absence of clusters can be explained by at least three possible reasons. First, as pointed out by the reviewer, it is possible that the fusion of a fluorescent protein (GFP or NeonGreen) affects the assembly of Ire1p clusters. We think that this is unlikely as these clusters could be visualized by cryoCLEM using a similar fusion constructs in mammalian cells (Tran et al. Science 2021 (PMID: 34591618)). The second possible explanation is a technical limitation. To detect enough fluorescent signal, our cryo-FM data was collected on ~400 nm thick lamellae prior to polishing the lamellae down to Since it is very challenging to convincingly determine which explanation is correct, and it still remains largely speculative to us, we decided to not elaborate on this part of the research effort.*

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      1. For the description of ER volume changes under ER stress, the manuscript currently presents only tomograms from Ire1c-GFP cells treated with DTT. To strengthen this observation, it would be helpful to also include representative tomograms and corresponding segmentations from additional treatments and strains in the supplementary figures.

      Indeed we have only collected these low magnification tomograms on a single strain. Similar to HAC1 splicing, ER expansion in response to ER stress is a widely accepted phenomenon (Bernales et al. 2006 (PMID: 17132049), Schuck et al. 2009 (PMID: 19948500)). There is likely only limited potential for new insights from reproducing these data, and we do not feel the resource investment required is justified.

      1. While cryo-ET enables structural analysis of small cellular volumes at high resolution, volume EM approaches such as FIB-SEM can provide complementary large-scale ultrastructural information. In particular, samples prepared by high-pressure freezing and freeze substitution generally preserve membrane morphology well and closely resemble native membrane architecture. Incorporating such approaches could further support and complement the cryo-ET observations.

      We think that additional volume EM experiments cannot be justified here as the enlargement of ER volume upon ER stress has already been well established through various volume EM approaches (eg; Sriburi et al. 2004 (PMID: 15466483), Bernales et al. 2006 (PMID: 17132049), Schuck et al. 2009 (PMID: 19948500), Heinz et al. 2025 (PMID: 40795978)). Here we only collected additional low magnification tomogram and quantified the ER volume on these to confirm that our experimental conditions lead to a similar ER stress response as previously described. We will add additional references related to this volume EM work to the text to clarify this.

      Minor Comments

      1. Figure 1: the number of biological replicates (N = 3) is relatively small, particularly considering that yeast samples are generally not difficult to prepare.

      The sample size here was not limited by yeast preparation but by cryo-ET data collection time. Because ER expansion upon ER stress is well established in the literature, we only collected a few low magnification tomograms to confirm this effect in our samples, and dedicated most of our microscope time to the collection of high magnification tomograms for the analysis of translation elongation dynamics.

      1. In addition to ER volume expansion, were there any detectable changes in nuclear size or nuclear envelope morphology? Since the nuclear envelope is continuous with the ER network, this could provide additional insight into the cellular response to ER stress.

      Only a few of our low magnification tomograms contained parts of the nucleus. The few nuclei that we observed may show an increased distance between the nuclear membranes, but we collected too few examples to reliably quantify this effect. Hence we refrained from discussing this in our manuscript, focused on translation elongation.

      1. The discussion of alternative pathways remains underdeveloped. Specifically, the authors briefly mention Gcn2p and PKA signaling as potential contributors. Yet no experiments directly test whether the observed ribosome hibernation depends on these pathways. Could the authors clarify: whether eIF2α phosphorylation was induced under their stress conditions, whether Gcn2-deficient strains alter the hibernation phenotype and how much of the observed effect is truly UPR-specific rather than a generic integrated stress response?

      We touched upon alternative pathways in the discussion to explain that the observed hibernation was plausible. Since the PERK pathway does not exist in yeast, we expect that some readers might be surprised by our findings. We will adjust this paragraph to improve its readability.

      The suggested experiments would show which pathways are activated, however the activation of these pathways has already been described in the literature (Pincus et al., 2014 (PMID: 25275008) ; Patil et al., 2004 (PMID: 15314660)). To really explain which factors directly trigger hibernation, various additional biochemical experiments will have to be performed. This is definitely interesting for future research, but beyond the scope of this cryo-ET focused paper.

      We agree that we cannot directly attribute the observed changes to the UPR, that is why we focus on ER stress instead of the UPR. We will double check that this is done consistently throughout the paper.

    1. Author response:

      Reviewer #1 (Public Review):

      Zeng et al.’s work links several key issues in Cryo Electron Tomography in ways that reinforce each other, inspired by the cycleGAN model, leading to very positive results across several benchmark datasets. The related topics include tomogram cleaning and simulations (two crucial areas in the field), with ”spin-off” outcomes in automatic annotation and the completion of the missing wedge. The manuscript covers nearly all essential topics in Tomography, making it very comprehensive and potentially critical in the field. The generalization capabilities on the SHREC 2021 data set are very interesting, although difficult to quantify. I appreciate the approach, but I have serious concerns about some of the limitations of the results presented by the authors.

      We thank the reviewer for the encouraging assessment of our work and for recognizing the potential importance of integrating tomogram denoising and simulation within a unified unsupervised framework. We appreciate the reviewer’s thoughtful evaluation and the concerns raised regarding the limitations of the current results. We address these concerns in detail below and have revised the manuscript to clarify the scope, evaluation strategy, and practical applicability of DUAL.

      (1) Simplified data versus nowadays challenging tomography data. It is acknowledged the difficulty inmaking general tests. In this work, the method shows excellent results on potentially simple data sets (the SHREC 2021, which was used for a benchmark in ET several years ago, but not much used since then) and, even more, the old Relion data set for picking).

      We appreciate the reviewer raising this important point regarding dataset difficulty and relevance. The SHREC 2021 dataset was selected because it is currently the most widely used benchmark simulated dataset for cryo-electron tomography and originates from the last SHREC contest specifically designed for evaluating cryo-ET analysis methods. It provides standardized simulated tomograms with known ground truth structures, which enables objective and reproducible quantitative comparison between different methods. The RELION ribosome dataset is also a commonly used experimental benchmark for evaluating particle detection performance. Nevertheless, we agree that demonstrating performance on additional recent and challenging datasets will further strengthen the evaluation of the method. In response to this comment, we have expanded the experimental evaluation in the revised manuscript by applying DUAL to additional recent cryo-ET datasets to further demonstrate its effectiveness on recent tomograms with more complex biological structures and imaging conditions.

      Specifically, we added an evaluation on the CZII Cryo-ET Object Identification dataset, a popular competition in 2025 with more than 1,000 participants. This experiment complements the original SHREC 2021 and RELION ribosome benchmark results and shows that DUAL can also be successfully applied to more recent cryo-ET data. The quantitative results and representative visual comparisons (shown above in Figure 1 and 2) are provided in the new section 2.6.

      (2) Reproducibility by the average user. I have found many cases in which a specific software producesexcellent results when run by the authors. Still, the average user is lost with the parameters and cannot reproduce these promising results. I propose that the authors address this issue by involving some experimental colleagues and ask them to repeat the work. This is a general concern that applies not only to this work but to many others. I think this consideration is crucial for a field that is growing very quickly and where method development happens at an extraordinary pace... but are all of them generally useful?

      We fully agree with the reviewer that reproducibility and usability are critically important for computational methods in cryo-ET. In response to this concern, we substantially improved the accessibility and reproducibility of the DUAL framework and revised the accompanying documentation to make the implementation easier to inspect and use, as two experimental colleagues have used and reproduced the results. The updated software repository now includes improved documentation, a clearer README, practical tutorials, a method-to-implementation description, a code reference, and example workflows demonstrating how to reproduce the experiments described in the manuscript. We also provide pretrained models together with the configuration files used to generate the results reported in the paper. In addition, the revised documentation clarifies the data interface, domain convention, training workflow, model outputs, and the interpretation of the trained translators. We believe that these improvements will significantly facilitate reproducibility and make it easier for users to apply the method to their own datasets.

      Reviewer #2 (Public Review):

      This study introduces DUAL (Deep Unsupervised simultAneous denoising and simuLation), an unsupervised deep learning framework that jointly addresses denoising and realistic data simulation for cryo-electron tomography (cryo-ET). By leveraging a cyclic, unpaired learning strategy, DUAL avoids reliance on paired clean ground-truth tomograms, which represents a practical advantage over many existing supervised approaches.

      We thank the reviewer for the positive summary of our work and for recognizing the advantages of the unsupervised framework in avoiding reliance on paired ground-truth data.

      Through extensive quantitative evaluations on benchmark datasets, together with qualitative and downstream analyses on diverse experimental tomograms, the authors show that DUAL performs robustly across both denoising and simulation tasks.

      We appreciate the reviewer’s recognition of the robustness of the framework and the evaluation strategy presented in the manuscript.

      If feasible, a limited quantitative or qualitative comparison with one or more recently published deep learning approaches for cryo-ET denoising or simulation, such as CryoSamba, or DeepDeWedge, would further strengthen the evaluation and help contextualize DUAL’s performance.

      We thank the reviewer for this helpful suggestion. As also recommended by the editor, we extended the experiments to include comparisons with recently proposed methods CryoSamba and DeepDeWedge. These comparisons were performed using the same evaluation metrics used in the current experiments so that the results remain directly comparable. The additional comparisons are added into section 2.6.

      Specifically, DUAL was compared with CryoSamba for denoising and with DeepDeWedge for missing wedge compensation on the CZII Cryo-ET Object Identification dataset, a popular competition in 2025 with more than 1,000 participants. The results are shown above in Figure 1 and 2.

      Reviewer #3 (Public Review):

      The paper is titled “DUAL: Deep Unsupervised Simultaneous Simulation and Denoising for Cryo-Electron Tomography.” The authors provided two closely related code branches: one for denoising and one for missingwedge correction. However, I did not find the simulation component. This is important, as the authors state that “the simulation branch provides learning-based cryo-ET simulation to generate synthetic tomograms indistinguishable from experimental ones.”

      We thank the reviewer for carefully examining the released code and for pointing out this source of confusion. We would like to clarify that, in the DUAL framework, simulation and denoising are the two simultaneous branches that are trained jointly, rather than separate sequential modules. The simulation branch learns the transformation from clean/simulated tomograms to realistic experimental cryo-ET tomograms, while the denoising branch learns the reverse transformation from experimental tomograms to the clean domain. Together, these two translators form the cyclic unsupervised learning framework described in the manuscript.

      In the original repository release, the organization of the code may not have made this relationship sufficiently clear, which likely led to the impression that only denoising and missing-wedge correction components were provided. To address this issue, we have substantially revised the repository structure and documentation. The updated repository now explicitly documents the two simultaneous branches of DUAL, explains how the simulation and denoising translators interact during training, and provides clear instructions for reproducing both functionalities. We have also added a dedicated method-to-implementation guide, code reference, and tutorial examples that describe the usage of the simulation component and its role in generating realistic synthetic tomograms that are statistically and visually consistent with experimental cryo-ET data.

      We believe these revisions clarify the implementation of the simulation branch and make the correspondence between the manuscript and the released code substantially easier to understand and reproduce.

      In addition, no pre-trained models were provided. Given that the authors indicate that all training data are publicly available, sharing trained models together with references to the corresponding datasets would significantly facilitate evaluation of the reported performance.

      We agree with the reviewer that providing pretrained models will greatly facilitate reproducibility and evaluation by other researchers. In the revised release of the repository, we have provided pretrained models corresponding to the experiments described in the manuscript together with clear references to the datasets used for training.

      The provided instructions are quite minimal and do not currently support reproduction of the reported findings.

      We appreciate the reviewer highlighting this issue. We have expanded the documentation substantially and provided detailed instructions describing the full workflow required to reproduce the experiments presented in the manuscript. In the revised repository, we added documentation that more explicitly connects the method described in the manuscript with the released implementation. The README summarizes the repository scope and data interface, the tutorial describes the practical workflow for preparing data and running training, and the method and code reference documents describe the mapping between the DUAL formulation and the main implementation files. We believe these additions will make the workflow clearer for users who wish to reproduce or adapt the experiments.

      After many hours of trial, debugging, and experimentation, I was able to train a model for missing-wedge correction using the default parameters, although the process was slow and memory-intensive.

      We thank the reviewer for investing significant effort to test the software and for reporting this observation. Training large 3D deep learning models on cryo-ET volumes can indeed be computationally demanding. We have clarified the computational requirements in the revised manuscript and provide guidance for efficient training and inference.

      Once these points are addressed, I would return to my original request that the authors provide: 3. A fully solved and functional tutorial based on their updated notebooks with all the intermediate results.

      We agree that a comprehensive tutorial will be extremely helpful for users. In the revised repository we have provided a complete end-to-end tutorial demonstrating the workflow from raw tomograms to the final outputs including simulated tomograms, denoised tomograms, and missing-wedge-corrected tomograms.

      We once again thank the editor and reviewers for their insightful comments and suggestions, which have helped us significantly improve the manuscript and the accompanying software.

  4. Jun 2026
    1. Should the app keep this fully hidden, or are there moments where naming the kind to the user would actually help (for instance, framing a screen as “what happened” versus “what’s still there”)?

      When thinking about this question, my mind goes to pedagogy and ways to cue the user into understanding the new data structure (and even data structures in general!)

      That is: in addition to asking "how do minimize the amount of stuff people in the field have to think about", an interesting parallel question is "how can we offer (optional?) learning opportunities for people to understand how data works in CoMapeo?"

      Part of the reason for suggesting this is that this taxonomy definitely introduces a learning curve! And it may also impact field teams who have to intuit what observations to create in order to set up a good entity causal chain. ("OK, first let me map the creek, so I can then map the contamination in the creek...")

      See my comments in What & why and CoMapeo data model for more of me wrestling with instantiating entity types, for reference

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Strengths:

      This is an ambitious study that provides a quantitative dissociation of the roles of phasic and tonic pain in adaptive behavior, by integrating ecological neuroscience, motivational theory, and computational modeling. The use of immersive VR combined with a freeoperant foraging task offers a more ecologically valid context to study pain-related behavior compared to traditional paradigms. Furthermore, the study employs a multimodal approach by combining behavioral data, computational frameworks, physiological signals, and EEG. In particular, one of the main strengths of the study is the use of sophisticated computational modeling to capture phasic and tonic pain effects. The experiment codes are available on GitHub, increasing reproducibility.

      We appreciate the reviewers’ recognition of the study’s ambition, the integration of ecological and computational approaches, and our efforts to support reproducibility through open code.

      Weaknesses:

      The main limitations of this article are that it provides insufficient detail on VR implementation. The design of the VR environment is, at this stage, under-described. Crucial information is missing, such as the number of pineapples per block, timing precision, details on how motion is mapped to the virtual movement, etc. This aspect strongly limits the reproducibility of the experiments.

      We thank the reviewer for highlighting the importance of detailed reporting to ensure reproducibility. In response to this valuable feedback, we have taken the following steps:

      (1) Open Access to Software and Data: We have now uploaded the full software and hardware specifications used in our study to a public GitHub repository: https://github.com/ShuangyiTong/PineappleStudy2025ReplicationSoftware. This includes the complete VR implementation, allowing readers to directly experience the task using a commercially available VR headset. The repository also contains the raw data and analysis scripts to facilitate full replication of our results. These links have been updated in “Data and Code Availability” section.

      (2) Expanded Methodological Details: We have revised the Methods section to include the specific details requested, such as:

      (a) The number of pineapples presented per block,

      (b) The temporal resolution and precision of the data collection,

      (c) The mapping between physical motion and virtual movement within the VR environment.

      Specifically, the paragraph containing the changes is following: “At the beginning of each one-minute block, a total number of 150 virtual pineapples of varying heights from 0.33 to 1 m were randomly generated in a circle centred around the participant with a diameter of 6.67 m. Five identical baskets were placed within the space. Spatial locations of trees and vegetation were generated using the game engine's default tree painting tool (Unity Technologies, San Francisco, US).”

      We hope these updates address the reviewer’s concerns and significantly improve the transparency and reproducibility of our experimental design.

      A second limitation lies in the lack of clarity regarding the study hypotheses. Although two overarching hypotheses can be inferred, they are not explicitly formulated. To this end, it is unclear which analyses were merely exploratory, especially for physiological and EEG outcomes.

      We thank the reviewer for this constructive feedback. We agree that making the hypotheses more explicit—particularly regarding the computational framework and the role of physiological measures—strengthens the manuscript. We have significantly revised the final section of the Introduction to explicitly formulate our two primary hypotheses and operationalise the associated behavioural and neurophysiological measures.

      (1) Phasic Pain Hypothesis: We hypothesised that phasic pain serves as a discrete valuation signal that updates the state-action value of specific actions. We predicted this would be evidenced behaviourally by reduced choice probability and increased ‘distance bias’ for pain-associated targets. Neurally and physiologically, we predicted that these aversive values would be tracked by skin conductance responses (SCRs) and the amplitude of pain event-related potentials (ERPs), which serve as established markers for the encoding of aversive magnitude and salience.

      (2) Tonic Pain Hypothesis: We hypothesised that tonic pain acts as a coefficient modulating the trade-off between opportunity cost and vigour cost. This was tested by applying tonic pain to the non-dominant (non-task) limb to ensure that any observed changes were motivational rather than mechanical. We predicted a global reduction in motivational vigour, operationalised as decreased movement velocities and foraging rates.

      By framing the study this way, we clarify that the physiological and EEG outcomes were used to quantitatively test whether the brain and body implement the computations (valuation and vigour-regulation) defined by our model. We have updated the text in the Introduction (see below) to reflect these explicit formulations.

      Updated paragraphs: “Our first hypothesis was that phasic pain provides a distinct valuation signal that updates the value of specific actions within complex environments. In our task, this was implemented by associating specific fruit (distinguishable by colour) with a brief electrical stimulus to the grasping hand, emulating thorns. In our computational model, this was defined as an aversive utility term incorporated into the state-action value evaluation process. We predicted that this computational mechanism would manifest behaviourally as a reduction in choice probability for pain-associated targets and an increase in ‘choice distance bias’ (the willingness to travel further for pain-free options). Neurally and physiologically, we predicted that these aversive values would be tracked by skin conductance responses (SCRs) and the amplitude of nociceptive event-related potentials (ERPs), specifically the N1-P2 complex (Favero et al., 2023).

      Second, we hypothesised that tonic pain acts as a coefficient modulating the tradeoff between opportunity cost and vigour cost, thereby serving a recuperative function. To test this in Experiment 2, we delivered continuous tonic pressure to the non-dominant arm via an inflated cuff to emulate a background state of injury. Within our free-operant framework, tonic pain was modelled as a weighting factor that shifts the optimal balance toward reduced energy expenditure. Because the stimulus was applied to the non-task limb, we specifically predicted a global reduction in motivational vigour—operationalised as decreased movement velocities and foraging rates—rather than a direct mechanical impairment. By applying this formal computational approach, we move beyond exploratory observations to provide a rigorous, mechanism-based explanation for how distinct pain states adaptively govern choice and action.”

      In Experiment 2, the reduction in vigor during tonic pain could plausibly reflect attentional load rather than pain per se. As recognized by the authors, there is no control condition involving an innocuous salient stimulus to rule out non-specific effects of distraction. Perhaps a tonic non-painful but salient somatosensory stimulus (e.g., a strong vibrotactile stimulus applied on the same arm) could have been used as a control stimulus.

      We agree that examining the potential role of attentional load on the interaction between tonic and phasic pain is an important area of future investigation. The inclusion of additional control conditions matched for attentional salience with additional experiments is possible but introduces other confounds related to their different qualities (e.g. a salient vibrotactile stimulus might invigorate behaviour). More fundamentally, attentional processes are a core part of pain function, and should not necessarily be viewed as a confound (i.e. the way that pain mediates some of its core functional effects may directly be through its salient attentional nature). This view is formalised in Wall and Melzack’s classical tripartite model of pain, and distinguishes pain from purely sensory systems such as somatosensation, vision and so on.

      Reviewer #1 (Recommendations for the authors):

      (1) Computational models may be difficult to follow without prior familiarity. Including simplified explanations could make the approach more accessible.

      We thank the reviewer for this constructive suggestion. To make the computational framework more accessible to a broader audience, we have added two new schematic diagrams (Figure 2 and Figure 8) that provide a visual overview of the models used in Experiment 1 and Experiment 2, respectively. These figures illustrate the state-action transitions and provide a clear decomposition of the payoff components—including reward, pain, and temporal costs. We believe these additions significantly clarify the modelling logic and help ground the mathematical descriptions in a more intuitive visual context.

      (2) Lines 220-222: I don't think it is possible to talk about "objective measures of pain" as pain is, by definition, subjective. I suggest rephrasing the sentence.

      We thank the reviewer for this thoughtful observation regarding our terminology. We recognise that the phrase ‘objective measures of pain’ may be misintepreted. Our intention was to highlight the distinction between the internal, reported experience and the behavioural manifestations of pain that our computational method reveals.

      To avoid ambiguity and to better align the text with the core focus of our study, which is the motivational function of pain, we have rephrased the sentence as suggested. We have shifted the emphasis from ‘measuring pain’ to quantifying its specific impact on behaviour.

      Original lines 220-222 have been revised as follows:

      "Taken together, this indicates the composite nature of overall aversiveness and highlights the benefit of combining subjective ratings with model-based measures of its motivational impact on behaviour."

      We believe this revision more accurately reflects our approach of using choice and movement as objective indices of the motivational value of pain.

      (3) The explanation for choosing the foraging task is very interesting, but should be provided in the Introduction rather than in the Methods section. In contrast, the Methods section should include the details of the VR implementation.

      We thank the reviewer for these constructive suggestions regarding the manuscript structure.

      Regarding the rationale for the foraging task: We agree that providing the theoretical justification for the task earlier in the manuscript improves the narrative flow. We have revised the Introduction to explicitly outline why a foraging paradigm was chosen by added the following sentences:

      “A foraging paradigm provides a robust, free-operant framework that captures the core components of adaptive behaviour: it is goal-directed, involves complex movement, and requires the learning of an optimal strategy to maximise rewards. This allows us to computationally dissociate how different types of pain influence the control of action.”

      We believe this addition clarifies the link between our computational hypotheses and the experimental design.

      Regarding the VR implementation: We have updated the Methods section to include the specific experimental parameters requested in the reviewer's previous comments (e.g., timing precision, stimulus counts, and motion mapping) to ensure full reproducibility. However, we have opted not to include the exhaustive engineering details of the underlying software architecture and communication protocols. To ensure complete transparency, the full software and firmware source code, which allows for the exact replication of the environment, is available in our public GitHub repository shown in the code and data availability section.

      (4) It is unclear how the sample size was determined. This information should be included.

      We thank the Reviewer for this comment. For the present study, an a priori power analysis was not conducted due to the novelty of the investigation and the complexity of the analyses. Standard power analyses are not commonly conducted for studies where computational modelling is the primary focus, as results would be potentially misleading. Instead, we based our sample size estimate of N ≈ 30 participants on previous studies using computational modelling of neurophysiological data [6], as well as EEG, SCR and pain studies [7, 8] and studies in our group using combined neurophysiological recordings and VR [9]. This approach represented a pragmatic balance which ensured the credibility of our results and the stability of our model estimates while accounting for the high persubject cost and the depth of the data collected from each individual. This has now been described more accurately in the Method section:

      “An a priori power analysis was not conducted due to the novelty of the investigation and the complexity of the analyses. Instead, we based our target sample size (N ≈ 30 per experiment) on previous studies using computational modelling of neurophysiological data (Mahajan et al., 2025), as well as EEG, SCR, and pain studies (Schulz351 et al., 2015; Zhang et al., 2018), and studies from our group using combined neurophysiological recordings and VR (Hewitt et al., 2026). This approach represents a pragmatic balance that ensures the credibility of the results and the stability of model estimates while accounting for the high per-subject cost and depth of data collected from each individual.”

      (5) Please clarify how / when the monetary performance incentive was provided.

      We thank the reviewer for the opportunity to clarify the incentive structure. The monetary performance incentive is detailed below:

      Participants were informed at the start of the study that they would earn a performance-based bonus of up to £10, determined by the points they collected during the foraging task. To ensure that motivation remained consistent across the entire session for all individuals—regardless of their baseline foraging speed—the specific exchange rate between points and currency was not disclosed. This prevented potential 'ceiling effects', where a high-performing subject might stop exertive effort after reaching the maximum bonus early, or 'floor effects', where a subject might perceive the reward for an individual action as too small to be motivating.

      Following the completion of the experimental session, all participants were compensated with the full £10 bonus in addition to their base payment for participation.

      We have updated the Methods section to reflect these details:

      “Participants were informed at the start of the experiment that their total points would be rewarded with a monetary incentive of up to £10. To maintain a constant level of motivation throughout the task, the exact point-to-currency exchange rate was not specified. Upon completion of the session, all participants were awarded the maximum bonus of £10.”

      Reviewer #2 (Public review):

      Strengths:

      Overall, this study aims to address an important topic and is generally well written.

      We thank the Reviewer for the generally positive evaluation of our work.

      Weaknesses:

      First, phasic pain was induced using electrical stimulation, which typically elicits somatosensory evoked potentials (SEPs). These responses may not reflect pain-specific processes and thus complicate interpretation. This issue bears directly on the study's conclusions, especially when discussing interactions between phasic and tonic pain. For example, tonic pain is known to reduce perceived intensity or cortical responses to phasic pain stimuli delivered elsewhere on the body - an effect not expected for SEPs elicited by electrical stimuli.

      We acknowledge the reviewer’s concern regarding the specificity of evoked potentials elicited by electrical stimulation. We agree that traditional SEPs— particularly those evoked by large surface electrodes—primarily reflect activation of non-nociceptive A-beta fibres and thus may not reliably index pain-specific processes or be modulated by tonic pain via descending nociceptive control. However, we would like to clarify that phasic pain was administered in the present study using small-diameter concentric ‘Wasp’ electrodes. These are comparable to intraepidermal electrodes shown to preferentially activate nociceptive A-delta fibres, thereby eliciting ERPs more closely associated with nociceptive processing rather than mixed somatosensory input [1, 2]. Accordingly, our ERP results demonstrated a reliable increase in N1-P2 amplitude with higher phasic pain intensity, suggesting that the evoked responses captured stimulus-evoked nociceptive processing.

      We acknowledge that these ERPs may still reflect mixed sensory processing and thus may not be fully modulated by tonic pain. Previous studies have shown that ERPs elicited by nociceptive electrical stimulation can be attenuated during tonic pain using cold-water immersion in CPM paradigms [3, 4]. However, these studies typically employ passive tasks, whereas our paradigm involved continuous voluntary behaviour during sustained tonic pressure pain. This difference in task context may engage distinct modulatory systems, possibly prioritising behavioural adaptation over sensory gating.

      We have revised the Discussion and Methods sections to explicitly clarify the electrode design and address the lack of ERP modulation by tonic pain in the context of active behaviour:

      Discussion: “Although we utilised concentric ‘Wasp’ electrodes designed to selectively activate nociceptive A-delta fibres, and confirmed that the resulting ERPs (N1-P2) were significantly modulated by phasic intensity (Figure 6E, F), we observed no such attenuation by tonic pain (Fig. 6G, H).”

      Methods: “These electrodes preferentially activate nociceptive A-delta fibres, thereby eliciting ERPs that more accurately reflect nociceptive processing compared to standard bipolar stimulation (Inui et al., 2002; Mørch et al., 2011).”

      Second, additional control experiments are necessary to rule out alternative explanations. For instance, the authors are suggested to deliver phasic pain to the contralateral arm (e.g., at 1-2 Hz), which might also reduce action velocity. Similarly, tonic pain applied to the grasping hand should be tested to disentangle hand-specific effects.

      We thank the reviewer for these suggestions regarding the spatial configuration of stimuli. The decision to deliver phasic pain to the grasping hand and tonic pain to the contralateral arm was a deliberate feature of our experimental design.

      First, delivering phasic pain to the grasping hand ensured spatial congruency between the virtual stimulus (the fruit) and the physical consequence (the pain). This congruency is essential for subjects to form a coherent representation of the 'painful' object; a contralateral delivery would have introduced a sensory-motor mismatch that could complicate the interpretation of the learning and choice data.

      Second, tonic pain was applied to the contralateral arm specifically to avoid mechanical interference with the grasping action. Applying sustained pressure to the ipsilateral limb would likely have impeded the manual dexterity and fine motor control required to operate the controller buttons. This would have introduced a physical confound, making it difficult to determine if changes in behaviour were due to motivational vigour or simply the mechanical difficulty of performing the grasp while the arm was under pressure.

      We agree that exploring the spatial generalisation of these effects is an important future direction, and we have added a paragraph to the Discussion to clarify these design choices:

      “It is also important to consider the spatial configuration of the stimuli used in this study. Phasic pain was delivered to the grasping hand to maintain spatial congruency with the virtual fruit, ensuring a coherent nociceptive feedback signal for the interactive task. Additionally, tonic pain was applied to the contralateral arm to prevent mechanical interference with motor execution, which would have occurred if pressure were applied to the ipsilateral limb used for grasping the controller. Whilst this design promotes spatial congruency and avoids mechanical confounds, future studies might explore how these effects generalise across different body parts, for which VR experiments serve as a promising tool to test relevant hypotheses (Hewitt et al., 2026).”

      Reviewer #2 (Recommendations for the authors):

      (1) First, the abstract mentions only EEG, yet Experiment 1 employed skin conductance response (SCR) measures while Experiment 2 utilized EEG. Also, the rationale for using SCR in Experiment 1 and EEG in Experiment 2 is not provided and should be explicitly stated.

      We thank the reviewer for identifying the discrepancy between the physiological signals reported in Experiment 1 and Experiment 2. We have revised the Abstract and Methods section to clarify the rationale for these measures.

      In Abstract, the following sentence has been revised: This could be explained by a free-operant computational framework that formalises and quantifies the function of tonic and phasic pain in terms of motivational vigour and decision value, and model parameters correlated with EEG “physiological and neural responses.”

      Regarding the rationale for the measurements, the following sentences were inserted into the Methods section: “Experiment 1 was designed to establish the robust behavioural effects of the foraging task while ensuring the collection of reliable physiological data. We chose SCR as it is a well-validated index of autonomic arousal that we were confident would provide a clear peripheral measure of pain-related processing in this novel VR paradigm.”

      For Experiment 2, we aimed to build on these findings by adding EEG. This was intended as a complementary piece of neural evidence to provide insights into the underlying central neural mechanisms of phasic and tonic pain interactions.

      (2) Second, the quality of both SCR (Figure 3A) and EEG/ERP data (Figure 5A-D) appears compromised by low SNR. For instance, ERP signals show baseline drift at low frequencies, potentially due to movement-related artifacts. The authors are encouraged to enhance data quality and provide cleaner, more interpretable results.

      We thank the reviewer for this observation. We acknowledge that our recordings exhibit a lower SNR compared to conventional, stationary EEG studies. This is a recognized characteristic of Mobile Brain-Body Imaging (MoBI), particularly in immersive VR experiments where participants are physically active [10]. However, previous research has demonstrated that it is possible to recover valid, interpretable neural signals in active settings using modern cleaning methods including trained ICA labels which we have adopted for artefacts cleaning [11]. We also believe we should be restrained from over cleaning the EEG data as pointed out by Delorme in the paper ‘EEG is better left alone’ [12]. Therefore, we have added a new paragraph in the Discussion:

      “It is important to acknowledge that the signal-to-noise ratio in both our physiological and neural recordings is lower than that typically observed in conventional, stationary laboratory experiments (Gramann et al., 2011). This is primarily due to the motion artefacts inherent in an immersive and active virtual reality environment. Whilst we utilised robust cleaning and artefact-correction methods (Klug and Gramann, 2021), the elevated noise floor may limit our capacity to detect more subtle neural effects or interactions. These challenges highlight a critical area for future methodological research, particularly in the development of hardware and signal-processing tools designed to isolate neural signals during complex, mobile behavioural tasks.”

      Another factor contributing to the appearance of the raw signal is the "free-operant" nature of our task. Unlike conventional neurophysiological study paradigms with fixed, sufficient intervals between trials, our participants were free to move and interact with fruit at their own pace. This means that neurophysiological signals from successive actions (e.g., picking up one fruit followed quickly by another) can overlap. For the SCR analysis, we addressed this by using a canonical response function (CRF) to model and "unfold" the overlapping signals with GLM to produce our final results [13]. While we did not perform a similar deconvolution for the EEG data, we focused our analysis on the early, salient components (N1-P2 and early time-frequency changes < 500ms) which are less susceptible to overlap from subsequent actions than the much slower SCR.

      In summary, while significant efforts representing the state-of-the-art approach for MoBI analyses have been taken to minimise the contributions of noise to the dataset, residual noise does remain in the final data. We have employed a combination of robust preprocessing and model-based analytical methods to account for the complexities of a free-operant task. We believe these results represent the best possible balance between signal clarity and the ecological validity of an active foraging task, and we have called for future research to continue improving these tools for immersive VR environments.

      (3) Third, although the authors state that time-frequency analysis was conducted on the EEG data, no corresponding results are presented in Figure 8 or elsewhere. Furthermore, the statistical maps shown appear noisy and require further clarification and possible denoising.

      We thank the reviewer for pointing this out. The time-frequency results are indeed presented in Figure 8 (now Figure 10); however, they are depicted as topographic maps of the t-statistics derived from our LMM rather than raw power change plots.

      The application of EEG to a novel, free-operant task represents a significant methodological development in this study. Unlike conventional EEG experiments where variables are strictly controlled and a "clean" pre-stimulus baseline is easily obtained, our task involves continuous participant engagement and movement. In this context, for the decision-making event, a stable baseline is unattainable as multiple variables, most notably head movements, are constantly in effect.

      Therefore, we believe that presenting the LMM statistical maps in the main text is the most appropriate and rigorous interpretation of the time-frequency results, as these maps represent the signal after accounting for these complex fixed and random effects. This approach was also adopted in previous pain studies [7]. We also updated the figure legend and caption specifically saying that the figure represented correlation between band power and variables we were investigating to improve clarity.

      Second, for more salient stimuli like phasic pain stimulation, we can indeed obtain a highly interpretable time-frequency analysis without further LMM analysis. We have added induced oscillatory responses to phasic pain stimuli to the Supplementary Material (section: Induced oscillatory responses to phasic pain stimuli). The results showed that, consistent with our ERP findings, the intensity of phasic pain significantly modulated induced responses, while the background tonic pain state did not significantly alter the induced oscillatory response to the phasic pain stimulus.

      Regarding the SNR and Denoising Strategy, we acknowledge that the statistical maps appear noisier than those from stationary studies. This is a direct consequence of the lower signal-to-noise ratio (SNR) inherent in mobile VR. Moving EEG from strictly controlled laboratory settings to ecologically valid, "real-world" VR scenarios introduces higher levels of noise, which we believe represents a key frontier for future methodology research. Regarding the denoising process, the maps in the main text represent the data after our full pipeline (including ICA-based artifact rejection and high-pass filtering). Regarding further denoising, we have deliberately chosen not to apply excessive spatial or temporal smoothing [12]. Also, it is important to note that the LMM framework itself serves as a powerful statistical "filter." By including head movement velocity as a regressor and accounting for random intercepts across subjects, the model effectively "cleans" the signal by partitioning out noise components not related to the task conditions.

      Reviewer #3 (Public review):

      Strengths:

      The experimental paradigm is highly innovative. Assessing human behaviour in a naturalistic yet highly controlled setting represents a promising approach to pain research. Notably, assessing pain magnitude implicitly, via its motivational value, offers insights about the overall pain experience that are not usually accessible via common pain ratings.

      Weaknesses:

      Despite these strengths, the manuscript would benefit significantly from more precise definitions of key concepts and an overall clearer, more coherent presentation of its main arguments. The writing, in its current form, often presents claims that are too vague or insufficiently connected with the experimental findings. Moreover, certain aspects of the computational modeling and statistical analysis appear flawed or inadequately justified.

      We thank the Reviewer for the generally positive evaluation of the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) The analyses presented in the section

      "Results/Additional cost of effort associated with movement" require clearer explanations. The intention here appears to be to assess the association between moving distances and pain intensity to test the hypothesis that the higher the average pain ratings within blocks, the longer the distances moved (i.e., the higher the effort to avoid pain). It is unclear why and how exactly "egocentric distance differences between painful and non-painful fruits" were computed.

      We thank the reviewer for pointing out the need for a clearer definition of the egocentric distance calculation. As the reviewer correctly identified, this analysis tests the hypothesis that subjects would trade off physical effort (distance) for pain avoidance. To compute this, we used a blockwise approach: for each one-minute block, we calculated the average egocentric distance travelled to pick up non-painful fruits and subtracted the average distance travelled to pick up painful fruits. This difference (labelled as "Choice Distance Bias" in Figure 3B) represents the additional effort subjects were willing to exert to reach a pain-free option. We have clarified the computation method and our motivation for using it in the revised text:

      “As shown in Figure 3B, the vertical axis represents the 'choice distance bias', calculated as the difference between the average egocentric distance to non-painful fruits and the average egocentric distance to painful fruits within each block. The egocentric distance is the fruit distance relative to the participant. This metric was computed to test whether subjects would trade off physical effort for pain avoidance; specifically, a positive bias indicates that subjects were willing to bypass closer painful fruits to reach more distant pain-free ones. As hypothesised, we found that as the pain intensity (VAS) of the aversive fruits increased, this distance bias grew significantly, confirming that subjects exerted greater movement effort to avoid higher levels of pain.”

      We have also updated the text in the beginning of " Avoidance increases with increasing phasic pain intensity" section to emphasize the calculation is analysed at the block level to clarify the computation procedure:

      “For this analysis, both aversive choice probabilities and subjective pain ratings were estimated at the block level.”

      (2) In its current form, the explanation of the first optimality equation lacks precision and transparency. Consider the following improvements:

      (a) Precisely define the features that characterize a state/decision point: e.g., i) memory of available options (= set of 7 fruits that were seen but not picked up) and ii) subject's current position, iii) pain intensity associated with green fruit in the current block.

      (b) Precisely define the set of values the action variable a can assume.

      (c) Precisely define the function u(a) in mathematical notation, including its hyperparameters. The fact that a is likely a categorical variable, while u(a) is later described as a sigmoid function (i.e., as a function of a continuous variable), is confusing. In my understanding (see Figure 2F), u is actually a function of the stimulus intensity associated with a given fruit. Since the stimulus intensity depends on the current state s (and varies from block to block), the phasic pain utility function technically also depends on s.

      (d) Precisely define the function d(a) in mathematical notation, including its hyperparameters.

      (e) Precisely describe how the separate horizontal and vertical components of C_m enter the equation.

      (f) Provide a summary of all parameters and hyperparameters being optimized. Are parameters and hyperparameters optimized jointly? What distinguishes parameters and hyperparameters practically?

      We thank the reviewer for this insightful critique. We agree that the original presentation of the optimality equation was insufficiently formal. We have now added a dedicated subsection, "Experiment 1 model summary", which includes a comprehensive table (Table 2) and supporting text to address these points with mathematical precision.

      Specifically, we have implemented the following clarifications in the revised manuscript:

      State and Action Space (a, b): We have formally defined the state s as an ordered memory list M_s of up to 7 items, governed by a FIFO principle. The action a is now explicitly defined as a one-to-one mapping from these memory items to physical reach trajectories.

      Utility and Cost Functions (c, d, e): We have provided the full mathematical notation for the phasic pain utility u(a) and the effort cost d(a). We have clarified that while the choice of fruit (a) is categorical, it serves as an indicator variable that determines the application of a continuous sigmoid utility function based on the block-level pain intensity (x_stim). We have also explicitly decomposed the effort cost into its horizontal (C_h) and vertical (C_v) egocentric components.

      Parameters and Hyperparameters (f): We have clarified that because our model focuses on steady-state motivational trade-offs rather than online learning, the hyperparameters listed are the only variables subject to optimisation. These are fixed for each subject across the duration of the experiment.

      We believe these additions, centred around the new Table 2, provide the transparency and precision requested.

      Furthermore, we would like to clarify a subtle caveat regarding the assumption of a fixed x_stim for the entirety of a block. While participants were aware that green pineapples were aversive, the specific stimulation intensity for a given block was only fully revealed upon picking up the first green pineapple.

      To ensure our model-fitting remains robust despite this 'information lag', we considered several computational alternatives:

      (1) Prior Estimation Modelling: Modelling a participant’s prior estimation of pain stimulation based on previous blocks. We found this unsuitable due to the independent block design and the limited number of trials available to establish a stable prior.

      (2) Data Trimming: Excluding all decisions made before the first green pineapple pickup. While theoretically 'cleaner', this approach introduces significant data imbalance and ignores blocks where a participant—dissuaded by high pain— only picked up a single green fruit before ceasing (approx. 8.75% of blocks).

      Crucially, we performed a sensitivity analysis by re-running the model-fitting procedure using only the data collected after the first green pineapple was harvested in each block. This analysis yielded the same qualitative statistical results as the full-block model presented in the main text. We have added a detailed discussion of this caveat and the alternative study designs we explored (such as pre-block stimulation or stochastic choice paradigms) to the Supplementary Material (Section Discussion of pain intensity information and model robustness). We believe this confirms that our current approach provides a faithful representation of the underlying motivational trade-offs.

      (3) The statistical method selected for assessing the association between decision values and pain ratings is problematic (Figure 2G): Since there are multiple data points from multiple subjects, which introduces dependence between data points, a multilevel instead of a single-level linear regression should be employed.

      We appreciate the reviewer’s suggestion to utilise a multilevel modelling approach. We agree that a single-level regression does not fully account for the nested structure of our data.

      In response, we re-analysed the association using a linear mixed-effects model with a maximal random effects structure. Specifically, we included both random intercepts and random slopes for Ratings grouped by Subject (in R syntax: PainFunc ~ Ratings + (1 + Ratings | Subject)).

      The results of this mixed effect model are consistent with our original findings, showing a significant relationship between decision values and pain ratings (p = .001). We have updated the Figure caption (now Figure 3G) to reflect these multilevel model statistics. We believe this addition addresses the concern regarding data dependence and provides a more rigorous validation of our conclusions.

      (4) The statistical method selected for assessing how decision values/pain ratings relate to SCR coefficients is problematic (Figures 3B and C): Again, a multilevel regression method should be used.

      We thank the reviewer for this important point. We agree that a multilevel approach is more appropriate for our nested data structure, and that the interpretation of the SCR data required more explicit justification in the context of the divergence between decision values and ratings.

      We have now re-analysed the relationship between SCR coefficients (both fixationevoked and shock-evoked), decision values, and subjective ratings using a multilevel (mixed-effects) regression model. This model included random intercepts and random slopes for each participant to account for individual variability. We have updated Figure 4 (previously Figure 3) caption and the corresponding Results and Discussion sections to reflect these findings (revised text are copied to the response to next comment (5) below. This more rigorous approach provided a clearer and more nuanced picture of the data. Specifically, while the simple regression previously suggested that both measures correlated with fixation-evoked SCR, the multilevel model reveals a dissociation: fixationevoked SCR is significantly associated with decision values, but not with subjective ratings.

      (5) The interpretation of the skin conductance analysis results as evidence of "dissociation between expected and experienced utility" is vague and not well-supported given the presented data and statistical shortcomings. The low R2 in Figure 2G already indicates divergence between decision values and pain ratings. It is unclear what the decision values' differential association with shock-evoked SCR coefficients adds to this insight.

      The reviewer correctly notes that the low R^2 in the correlation between decision values and pain ratings (Figure 3G) already suggests a divergence between these two measures. We agree that this is one of the key findings, as it highlights that decision values provide a dimension of pain assessment that is not fully captured by subjective report. However, we believe the SCR results add crucial physiological evidence to explain why and how these measures diverge. The updated multilevel results provide a more concrete double dissociation that aligns with the distinction between decision utility and experienced utility:

      Experienced Utility (Shock-evoked SCR): This measure of physiological arousal during the painful event was significantly predicted by subjective pain ratings (beta = 0.0154, p = .006) but not by decision values (p = .672). This suggests that ratings are more closely tied to the immediate, experienced aversiveness of the stimulus.

      Decision Utility (Fixation-evoked SCR): In contrast, arousal during the period of evaluation/fixation was a significant predictor of decision values (beta = -0.0739, p = .009) but was not significantly associated with subjective ratings (p = .105).

      By using a more rigorous statistical method, we found that decision values are actually a more robust predictor of anticipatory/evaluative arousal (fixation) than subjective ratings are. This supports our interpretation that decision values and ratings capture different temporal and functional aspects of pain processing— specifically, the evaluation of potential outcomes (decision utility) versus the reaction to the outcome itself (experienced utility). We have revised the Discussion to be more conservative regarding the strength of this evidence while clearly articulating how these physiological results provide a mechanistic grounding for the divergence observed in the behavioural data.

      Summary of changes in the manuscript:

      Figure 4 Caption: Updated to report multilevel regression statistics (beta, 95% CI, t, and p-values) instead of R^2 from simple linear regression.

      Results Section: Updated the text to describe the mixed-effects model results, highlighting the dissociation between fixation-evoked and shock-evoked SCRs. Revised text:

      “Analysis using a multilevel linear mixed-effects model revealed a clear dissociation in the relationship between physiological responses and motivational parameters. Fixation-evoked SCR coefficients were significantly associated with decision values, but not with subjective pain ratings (Fig. 4B). Conversely, shock-evoked SCR coefficients showed a significant association with subjective pain ratings, while the association with decision values was not significant (Fig. 4C). This double dissociation suggests a notable divergence between the physiological correlates of expected utility (at the decision level) and experienced utility (the actual pain experience). Taken together, these findings highlight the composite nature of the overall aversiveness of pain and underscore the benefit of combining subjective ratings with model-based measures to capture its distinct impacts on behaviour.”

      Discussion Section: Revised the paragraph discussing decision versus experienced utility to include the "further hint" provided by the divergent SCR correlations.

      Revised text:

      “In our task we get a further hint of this in the SCR measures in experiment 1, whereby a discrepancy exists between decision values and pain ratings in their respective associations with fixation-evoked SCRs and phasic pain-evoked (shock) SCRs. Taken together, this indicates the composite nature of overall aversiveness of pain, and highlights the benefit of combining subjective ratings with model-based measures of its motivational impact on behaviour.”

      (6) When investigating the effects of tonic pain on the neural processing of phasic pain (Figure 5), why were only ERPs analyzed and not induced oscillatory responses?

      We thank the reviewer for this insightful suggestion. We initially focused our analysis on Event-Related Potentials (ERPs) because the N1-P2 amplitude is an established and robust marker in pain research, providing a clear and reliable metric for comparing phasic pain processing across conditions.

      However, we agree that induced oscillatory responses provide a more comprehensive view of cortical dynamics. Following your suggestion, we have performed a Time-Frequency Representation (TFR) analysis at electrode Cz. These results, now included in the Supplementary Material (Figure S4, S5), are entirely consistent with our ERP findings. Specifically:

      Phasic Modulation: Both ERP amplitudes and induced oscillatory power (notably in the theta and gamma bands) were significantly modulated by the intensity of the phasic pain stimulus.

      Tonic Independence: Consistent with the ERP results, the presence of background tonic pain did not significantly modulate the induced oscillatory responses to phasic stimuli.

      We believe this additional analysis significantly strengthens the manuscript by demonstrating that the observed effects are consistent across both phase-locked and non-phase-locked neural domains. We have amended the ERP results section to reflect the addition of induced oscillatory responses in supplementary materials: “We focused our neural analysis of phasic pain on ERPs as phasic stimuli are well characterised by these time-locked evoked potentials. Nevertheless, to ensure a comprehensive assessment of the neural response, we also examined induced oscillatory responses. These results were consistent with the ERP findings and are detailed in the Supplementary Materials (Fig. S4, S5).”

      (7) The explanation of the second optimality equation (involving motivational vigour) requires substantial clarification. Besides the points mentioned for the previous optimality equation, specific opportunities to improve the explanations include the following:

      - In the provided formula, C_v and C_m appear indistinguishable given they are multiplied together, rendering this an ill-posed optimization problem. This should be clarified.

      - In my understanding, d(a)/V_speed corresponds to the temporal delay associated with picking fruit a. Then, what is tau, and why compute the sum tau + d(a)/V_speed?

      - V* is not introduced properly. Is V*(s') = Q*(s', a, tau)? If so, why introduce V*? Moreover, the notational similarity between V_speed and V* is confusing.

      - Gamma = 0 still holds?

      - Summarize all parameters and hyperparameters that are optimized to model the data and more precisely describe the method used for optimization.

      We thank the reviewer for these insightful comments. We agree that the transition from a standard reinforcement learning framework to one incorporating motivational vigour requires precise definitions to ensure the model is well-posed and interpretable. We have addressed these points as follows:

      (1) Clarification of C_v and C_m: We have clarified C_m and d(a) in the newly added Experiment 1 model summary table. Specifically, C_v is the scalar vigour constant and C_m is a unit vector representing the horizontal and vertical components. Because C_m is a unit vector, the optimization does not suffer from a collinearity issue from the scalar multiplication between C_v and C_m.

      (2) Bridging Theory to Practice (tau and Total Delay): In the theoretical framework of Niv et al. (2007), "delay" is an abstract sum encompassing both waiting and execution. In practice, when fitting to real-world VR data with variable execution times , we must distinguish between the waiting time tau (time spent stationary or searching) and the execution time (||d(a)|| / V_speed). This is necessary because participants take time to look around the forest to search for fruits before deciding to commit to an action. The sum tau + ||d(a)|| / V_speed represents the total delay between two actions, which directly aligns with the notion of opportunity cost of time. We have added a table (Table 3) and added a new Figure 8 to clarify these distinctions.

      (3) V*, Q*, and gamma: The reviewer is correct that V*(s') = max_{a’, tau’} Q*(s', a', tau'). We previously used V* for simplicity. Since the notation of V* and V_speed was confusing, we have updated the term to max_{a’, tau’} Q*(s', a', tau') in the optimality equation. We confirm that gamma = 0 (a greedy policy) still holds for the Experiment 2 framework to maintain focus on steady-state motivational trade-offs. We have added this statement to the method section.

      (4) Summary of Parameters and Optimization: We have summarized the hyperparameters {k, x_0, C_p, C_v, h, v} in the new summary table for Experiment 2.

      (8) It is not clear what the results of the modelling approach presented in Figure 7a+b concretely add to the comparison of movement velocities and collection rates in Figure 6.

      We appreciate the reviewer's comment regarding the relationship between the raw behavioral metrics and the computational results. While both sets of findings support the argument for reduced motivational vigour in the tonic pain condition, we believe the modeling approach provides distinct and essential value:

      (1) Finer-Grained Analysis Tool: The computational model acts as a more sophisticated analysis tool than simple velocity or rate averages. Unlike Figure 9a+b (in the revised manuscript, previously Figure 7), which summarizes overall performance, the model accounts for the trial-by-trial trade-off between opportunity costs, movement effort, and choice values. This allows us to isolate vigour from other confounding components.

      (2) Direct vs. Indirect Measurement: If we assume that motivational vigour in a free-operant task can be quantified through an RL framework, as established in animal studies, then the model's vigour constant (C_v) serves as a direct, concrete estimate of that internal state. In contrast, overall speed and collection rates are indirect markers that can be influenced by multiple factors, such as different choice sets available to the participants as the fruits locations are randomly generated.

      In summary, the computational approach provides a rigorous, parameterized bridge between observable behavior and the underlying neuro-computational mechanisms of recuperative pain. We have updated the Discussion section to more explicitly state how the computational approach provides a controlled measure that is isolated from the other confounders of the task. Added text to the Discussion:

      “Compared to overall speed and collection rate, which can be influenced by multiple factors, such as different choice sets available to participants as the fruit locations are randomly generated, the model's fitted parameters (e.g. vigour constant C_v) in theory serves as a direct, concrete estimate of that internal state.”

      (9) Claims made in the discussion should be more thoroughly and closely linked to the results presented previously. Specifically, experimental outcomes supporting the following claims should be directly referenced:

      - "tonic and phasic pain serve different motivational functions".

      - "phasic pain provides a punishment teaching signal that directs avoidance".

      - "tonic pain reduces motivational vigour".

      - "these two functions [punishment teaching signals and reduction of motivational vigour?] can be formally distinguished and quantified".

      - "We did not see interactions between tonic and phasic pain".

      We have revised the Discussion to more explicitly link these claims to our experimental results. Revised text:

      “The experiments show that tonic and phasic pain serve different motivational functions during adaptive behaviour, in line with ecological and evolutionary theories of pain (Bolles and Fanselow, 1980; Walters and Williams, 2019). Specifically, our findings point towards phasic pain providing a punishment teaching signal that directs avoidance through value-based learning, balancing the cost of future harm alongside potential reward. This is supported by the observation that increasing phasic pain intensity significantly reduced choice probability and increased distance bias between choices, whereby participants were willing to travel further to reach a pain-free fruit. In contrast, we found that tonic pain reduces motivational vigour, which supports energy conservation and recuperation in the context of bodily damage. This claim is directly evidenced by the reduction in taskrelated movement velocities and fruit collection rates during tonic pain blocks. The experiments are the first to show that these two functions can be formally distinguished and quantified during ongoing behaviour. By utilising a free-operant RL computational framework, we were able to dissociate these roles phasic pain was quantified as a generally negative utility term affecting choice values, while tonic pain was formalised as a change in vigour constants that were significantly higher (increasing delays between actions) in tonic pain condition. This illustrates how pain simultaneously acts in different ways to serve self-protection.”

      “One notable aspect of our results is that we did not see interactions between tonic and phasic pain at either the behavioural or neural level. Behaviourally, we observed that average aversive choice probabilities remained similar regardless of the presence of tonic pain, with no significant interaction effect on punishment sensitivity. Furthermore, our model-fitting confirmed that tonic pain did not significantly modulate the fitted phasic pain utility values. There are two contexts in which these might be predicted. First, in `conditioned pain modulation' paradigms (Kennedy et al., 2016), a tonic pain stimulus is sometimes seen to reduce both the perceived intensity and the cortical evoked responses to phasic pain stimuli delivered somewhere else on the body (Hoffken et al., 2017; Enax-Krumova et al., 2020). Although we utilised concentric ‘Wasp’ electrodes designed to selectively activate nociceptive A-delta fibres (Inui et al., 2002), and confirmed that the resulting ERPs (N1-P2) were significantly modulated by phasic intensity, we observed no such attenuation by tonic pain. Indeed, neither subjective pain ratings nor the N1-P2 amplitude showed a significant modulation by the tonic pressure pain stimulus. In contrast, our results were more compatible with a trend in the other direction.”

      (10) The paragraph in the discussion "A concern that is sometimes raised..." (lines 243 - 254) raises interesting points, but its particular relevance to the study at hand is unclear.

      We appreciate the reviewer's feedback. The motivation for including this discussion is to address a common critique we received for the study: whether the observed reduction in vigour under tonic pain is "simply" due to distraction or cognitive load, rather than being a specific functional output of the pain system. We have revised this paragraph to link the concern to our paper’s specific finding.

      Our central argument is that for tonic pain, distraction is not a confounding "sideeffect" but rather the primary mechanism of action. By being inherently "distracting," tonic pain successfully withdraws resources from ongoing tasks (like foraging) to promote the energy conservation required for recuperation.

      (11) The clinical perspective of the methodological framework presented at the end of the discussion is interesting and could be expanded.

      We thank the reviewer for this encouraging comment. We have expanded the final paragraph of the Discussion to more explicitly state the clinical utility of our framework. Specifically, we now contrast our approach with standard clinical assessments such as Quantitative Sensory Testing (QST). We highlight that while QST is a valuable tool, it can lack ecological validity; in contrast, our VR-based task allows for a more realistic, behaviourally sensitive assessment of how pain impacts a patient’s daily functional activities and motivational state. We believe this represents a significant step towards more objective and "real-world" clinical pain phenotyping.

      (12) The statistical analyses part in the methods section should provide a clear definition of dependent and independent variables and clearly state which test was used for which analysis, e.g., by referencing the corresponding subfigure in the main text.

      We agree that a more structured summary of the statistical approach would improve the clarity of the Methods section. We have now included a comprehensive summary table (Table 1) in the Statistical Analysis subsection. This table explicitly defines the dependent and independent variables for each analysis, identifies the specific statistical model used (e.g. Linear Mixed Models or repeated measures ANOVA), and directly maps these to the corresponding figures in the results section.

      Minor comments:

      (1) Introduction:

      (a) The introduction should elaborate more on the advantages of employing an "ecologically meaningful context".

      We thank the reviewer for suggesting further elaboration on the advantages of employing an "ecologically meaningful context". We have updated the introduction to provide additional reasoning of choosing an ecologically valid context for the study:

      “One of the challenges in studying adaptive functions of pain is the difficulty of embedding experiments within ecologically meaningful contexts. To solve this, we designed an immersive foraging task using virtual reality (VR), in which humans search a forest to collect fruits from the low-lying bushes at varying heights. A foraging paradigm provides a robust, free-operant framework that captures the core components of adaptive behaviour: it is goal-directed, involves complex movement, and requires the learning of an optimal strategy to maximise rewards. This allows us to computationally dissociate how different types of pain influence the control of action.”

      (b) It would be helpful to clarify why tonic pain applied to a limb not involved in the task is expected to influence the motivational vigour with respect to the task.

      We thank the reviewer for pointing out additional clarification for applying tonic pain to the non-dominant arm. We have added the following text to the introduction clarifying our hypothesis and why it was applied to the non-task limb:

      “Second, we hypothesised that tonic pain acts as a coefficient modulating the tradeoff between opportunity cost and vigour cost, thereby serving a recuperative function. To test this in Experiment 2, we delivered continuous tonic pressure to the non-dominant arm via an inflated cuff to emulate a background state of injury. Within our free-operant framework, tonic pain was modelled as a weighting factor that shifts the optimal balance toward reduced energy expenditure. Because the stimulus was applied to the non-task limb, we specifically predicted a global reduction in motivational vigour—operationalised as decreased movement velocities and foraging rates—rather than a direct mechanical impairment.”

      (2) Results/Experiment 1:

      (a) How were monetary rewards implemented exactly? How much money per fruit?

      We thank the reviewer for the opportunity to clarify the incentive structure. Participants were informed at the start of the study that they would earn a performance-based bonus of up to £10, determined by the points they collected during the foraging task. To ensure that motivation remained consistent across the entire session for all individuals—regardless of their baseline foraging speed—the specific exchange rate between points and currency was not disclosed. This prevented potential 'ceiling effects', where a high-performing subject might stop exertive effort after reaching the maximum bonus early, or 'floor effects', where a subject might perceive the reward for an individual action as too small to be motivating.

      Following the completion of the experimental session, all participants were compensated with the full £10 bonus in addition to their base payment for participation. We have updated the Methods section to reflect these details:

      “Participants were informed at the start of the experiment that their total points would be rewarded with a monetary incentive of up to £10. To maintain a constant level of motivation throughout the task, the exact point-to-currency exchange rate was not specified. Upon completion of the session, all participants were awarded the maximum bonus of £10.”

      (b) A green pine apple is not ripe and, in a naturalistic context, possesses some aversive value, even in the absence of phasic pain stimuli. Why was the color coding not counterbalanced across individuals? To what degree could this have confounded the results?

      We thank the reviewer for this insightful point. We acknowledge that the lack of counter-balancing for fruit colour (green vs. yellow) is a limitation of the current study design. However, we believe the potential confounding effect of "unripe" green pineapples on the final analysed data is minimal due to the principles of associative learning.

      While a naturalistic heuristic (green = unripe) might establish a weak prior bias, fundamental associative learning [14] and reinforcement learning models [15] demonstrate that extensive training with a highly salient unconditioned stimulus (such as pain) rapidly overrides mild initial priors. The task objective focused strictly on maximizing reward points, and participants underwent extensive training (10 blocks in Experiment 1; 6 blocks in Experiment 2) before the analysed sessions began. During this time, the strong, explicit contingencies (green = pain, yellow = safe) were learned and verbally verified. Therefore, by the time the main experimental data was collected, any weak baseline aversion to green had been overshadowed by the explicit task contingencies, making the learned associative value the primary driver of behaviour. We have added a statement acknowledging this limitation and outlining this theoretical rationale in the Methods section.

      “While the colour association (green for painful, yellow for pain-free) was not counter-balanced across subjects, any inherent aversive value of green pineapples (e.g., as 'unripe' fruit) is expected to have a minimal confounding effect on the analysed data. In associative learning frameworks, while mild prior biases may influence initial value estimations, extensive training with a highly salient unconditioned stimulus (e.g. phasic pain) rapidly updates these values, driving them toward an asymptote determined entirely by the explicit task contingencies (Rescorla & Wagner, 1972; Sutton & Barto, 2018). Because participants underwent extensive training (10 blocks in Experiment 1 and 6 blocks in Experiment 2) to establish the explicit pain associations prior to the analysed sessions, the observed avoidance behaviour was predominantly driven by the learned phasic pain contingencies rather than baseline colour preferences.”

      (c) In the "Avoidance increases with increasing phasic pain intensity" section, clarify upfront that pain ratings and choice probabilities were estimated at the block level. This information is provided only in a later section.

      We agree with the reviewer that this information should be stated earlier for clarity. We have updated the beginning of the "Avoidance increases with increasing phasic pain intensity" section to specify that these metrics were estimated at the block level:

      “For this analysis, both aversive choice probabilities and subjective pain ratings were estimated at the block level.”

      (3) Results/Experiment 2:

      (a) ERP visualizations (Figure 5) should include standard error indicators.

      We have updated Figure 5 (now Figure 6) to include 95% confidence intervals for standard error of the mean across subjects for all ERP traces. This provides a clearer visualization of the variance in the neural response.

      (b) In the section "A unified model...", clarify what is meant by saying that the unified model is "validated by the behavioural data", since behavioral data is what is being modeled in the first place.

      We clarify that "validation" in this context refers to the consistency between the parameters estimated by our generative unified model and the results obtained from the independent, model-free regression analysis of the raw behavioural data. While both approaches use the same source data, the unified model provides a finer-grained analysis of latent internal states (like motivational vigour), whereas the regression provides a direct empirical benchmark (more details were discussed in the response to major comment (8)). We have rephrased this section to better describe this as a consistency check against empirical regression results.

      (c) In the context of Figure 8a, the term "correlations" is misleading if referring to pairwise comparisons.

      We appreciate the opportunity to clarify our terminology. The results presented in Figure 8a (and the associated text) are derived from a Linear Mixed Model (LMM) where the tonic pain condition was treated as a binary independent variable. The term "correlation" was used to describe the statistical association (represented by the t-values) between the presence of tonic pain and EEG band power, accounting for subject-level random effects. It does not refer to simple pairwise comparisons (like t-tests). However, we agree that "correlation" can be ambiguous when applied to a binary predictor. We have revised the text and figure legends to use the terms "associated with" or "predicted by" to more accurately reflect the LMM framework.

      (d) Based on the presented data, there is no evidence for the section headings claim "Neural activities link to vigour".

      We agree with the reviewer that our results primarily provide evidence for a significant neural association with the tonic pain condition rather than a direct, statistically robust correlation with the vigour parameter itself (after Bonferroni correction). While tonic pain is associated with reduced vigour behaviourally, the EEG markers we identified are more accurately described as signatures of the pain state. We have revised the section heading and the corresponding text to focus on the characterisation of the tonic pain state to ensure our claims are strictly supported by the statistical evidence.

      (4) Methods:

      In the supplementary materials, the headings pertaining to different LMMs are confusing and not consistent with the Figure labeling in the manuscript (e.g., 4(ii)b likely corresponds to Figure 4d).

      We thank the reviewer for identifying these inconsistencies in the supplementary material. We apologize for the confusion caused by the labelling errors during reformatting the manuscript. We have now thoroughly audited the supplementary headings and updated them to ensure they correspond directly and consistently with the figure labels in the main manuscript.

      References

      (1) Inui, K., Tran, T. D., Hoshiyama, M., & Kakigi, R. (2002). Preferential stimulation of Adelta fibers by intra-epidermal needle electrode in humans. Pain, 96(3), 247–252. https://doi.org/10.1016/S0304-3959(01)00453-5

      (2) Mørch, C.D., Hennings, K. & Andersen, O.K. Estimating nerve excitation thresholds to cutaneous electrical stimulation by finite element modeling combined with a stochastic branching nerve fiber model. Med Biol Eng Comput 49, 385–395 (2011). https://doi.org/10.1007/s11517-010-0725-8

      (3) Höffken, O., Özgül, Ö.S., Enax-Krumova, E.K. et al. Evoked potentials after painful cutaneous electrical stimulation depict pain relief during a conditioned pain modulation. BMC Neurol 17, 167 (2017). https://doi.org/10.1186/s12883-017-0946-7

      (4) Enax-Krumova, E., Plaga, A.-C., Schmidt, K., Özgül, Ö. S., Eitner, L. B., Tegenthoff, M., & Höffken, O. (2020). Painful Cutaneous Electrical Stimulation vs. Heat Pain as Test Stimuli in Conditioned Pain Modulation . Brain Sciences, 10(10), 684. https://doi.org/10.3390/brainsci10100684

      (5) Enrico Schulz, Elisabeth S. May, Martina Postorino, Laura Tiemann, Moritz M. Nickel, Viktor Witkovsky, Paul Schmidt, Joachim Gross, Markus Ploner, Prefrontal Gamma Oscillations Encode Tonic Pain in Humans, Cerebral Cortex, Volume 25, Issue 11, November 2015, Pages 4407–4414, https://doi.org/10.1093/cercor/bhv043

      (6) Mahajan Pranav, Tong Shuangyi, Lee Sang Wan, Seymour Ben (2024) Balancing safety and efficiency in human decision making eLife 13:RP101371 https://doi.org/10.7554/eLife.101371.2

      (7) Enrico Schulz, Elisabeth S. May, Martina Postorino, Laura Tiemann, Moritz M. Nickel, Viktor Witkovsky, Paul Schmidt, Joachim Gross, Markus Ploner, Prefrontal Gamma Oscillations Encode Tonic Pain in Humans, Cerebral Cortex, Volume 25, Issue 11, November 2015, Pages 4407–4414

      (8) Suyi Zhang, Hiroaki Mano, Michael Lee, Wako Yoshida, Mitsuo Kawato, Trevor W Robbins, Ben Seymour (2018) The control of tonic pain by active relief learning eLife 7:e31949

      (9) Hewitt, D., Tong, S., Schreiber, S., & Seymour, B. (2026). Tonic pain modulates neural correlates of associative phasic pain memories. PAIN. DOI: 10.1097/j.pain.0000000000003917

      (10) Gramann, K., Gwin, J. T., Ferris, D. P., Oie, K., Jung, T.-P., Lin, C.-T., Liao, L.-D., and Makeig, S. (2011). Cognition in action: imaging brain/body dynamics in mobile humans. Reviews in the Neurosciences, 22(6):593–582.

      (11) Klug, M. and Gramann, K. (2021). Identifying key factors for improving ica-based decomposition of eeg data in mobile and stationary experiments. European Journal of Neuroscience, 54(12):8406–8420.

      (12) Delorme, A. EEG is better left alone. Sci Rep 13, 2372 (2023). https://doi.org/10.1038/s41598-023-27528-0

      (13) Bach, D. R., Flandin, G., Friston, K. J., and Dolan, R. J. (2010). Modelling event-related skin conductance responses. International Journal of Psychophysiology, 75(3):349–356.

      (14) Rescorla, R. and Wagner, A. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement, volume Vol. 2

      (15) Sutton, R. S. and Barto, A. G. (2018). Reinforcement learning: An introduction, 2nd ed. Adaptive computation and machine learning. The MIT Press, Cambridge, MA, US.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Li and colleagues examines how defensive responses to visual threats during foraging are modulated by both reward level and social hierarchy. Using a naturalistic paradigm, the authors test how the availability of water or sucrose, with sucrose being more rewarding than water, shapes escape behavior in mice exposed to looming stimuli of different intensities, which are used to probe perceived threat level and defensive responses. In parallel, the study compares dominant and subordinate animals to assess how social rank biases the trade off between reward seeking and threat avoidance. By combining detailed behavioral analyses with computational modeling, the work addresses how reward level and social context jointly influence escape decisions in an ethologically relevant setting.

      Across the different experimental conditions, perceived threat level is the main determinant of behavior. The authors show that looming stimuli associated with higher threat (contrast) consistently elicit faster and more robust escape responses than lower threat stimuli. This effect is particularly evident during early exposures, when animals are highly vigilant and have not yet habituated to the looming stimulus (learned that it is not dangerous). Later they described that as animals gain experience and habituate, behavior becomes more flexible, and reward level begins to exert a graded modulation of the escape response. Importantly, the authors show that under high threat conditions increasing reward value leads to more frequent and faster escape rather than greater reward pursuit. This finding is particularly relevant, as it suggests that highly valued rewards can heighten vigilance and thereby enhance responsiveness to threat, highlighting that reward does not simply compete with defensive behavior but can also reshape it depending on the perceived level of danger, in contrast to low threat conditions, where threat can be more easily outweighed by reward. Thus, an important conceptual contribution of the study is the introduction of vigilance as a useful framework to interpret these effects. Vigilance is treated as a behavioral state reflecting heightened attention to potential danger. In line with what is known from natural foraging, mice initially maintain high vigilance when confronted with an innate threat. This perspective helps clarify a finding that might otherwise appear counterintuitive. One might expect higher rewards to motivate animals to tolerate risk, explore more, and habituate faster in any scenario. Instead, the data suggest that highly rewarding outcomes can elevate vigilance, making animals more responsive to threat and leading to faster or more frequent escape under high threat conditions. In this sense, reward does not simply compete with threat but can also amplify sensitivity to it, depending on the internal state of the animal.

      The social results are particularly interesting in this context as well. Dominant mice consistently prioritize avoidance over reward, showing stronger escape responses and slower habituation than subordinates. This behavior is well captured by the vigilance framework proposed by the authors: dominant animals appear to maintain higher vigilance, which biases decisions toward threat avoidance. The authors further suggest that stable social relationships sustain high vigilance and slow habituation, framing this as an evolutionarily conserved strategy that may enhance survival. This interpretation provides a valuable perspective on how social structure shapes defensive behavior beyond immediate physical interactions. At the same time, there are important limitations to this interpretation. All experiments were conducted in male mice, and it is possible that the relationship between social hierarchy, vigilance, and defensive behavior would differ substantially in females. In addition, the idea that stable social relationships maintain elevated vigilance does not straightforwardly align with broader views of social stability as protective for mental health and as a buffer against anxiety and stress. These points do not undermine the findings but suggest that the social effects described here should be interpreted with caution and within the specific context of the task and sex studied.

      We thank the reviewer for raising this important point. In the context of repeated looming exposure, slower habituation reflects more sustained vigilance over time. Compared to individually housed mice, group-housed mice exhibit slower habituation (Lenz et al., 2022), and pair-housed mice showed even slower habituation in our current work. Importantly, this pattern does not indicate that pair-housed mice have higher overall vigilance than individually housed animals. Although individually housed mice habituate more quickly, they display higher initial vigilance, as reflected by their increased probability of escaping in response to looming stimuli (Lenz et al., 2022). Thus, pairhoused mice exhibited reduced defensive responses compared to individually housed animals, consistent with a social buffering effect.

      Furthermore, in a separate study (Rank- and Threat-Dependent Social Modulation of Innate Defensive Behaviors; Li, Gao, Li, 2026, eLife 15:RP109571), we directly compared responses to looming stimuli when mice were tested alone versus in the presence of a social partner and observed clear evidence of social buffering.

      Another important limitation is that the neural mechanisms underlying these effects remain speculative. The manuscript includes an extensive discussion of candidate circuits, particularly involving the superior colliculus and downstream structures, but this section is necessarily based on prior literature rather than on data presented in the study. Given the complexity of the circuits involved in integrating internal state, reward, social context, and vigilance, the current work should be viewed as providing a strong behavioral and conceptual framework rather than direct insight into underlying neural mechanisms.

      We fully agree that the proposed neural mechanisms remain speculative and that the circuits involved in integrating internal state, reward, and social context are likely far more complex. We have revised the manuscript to acknowledge this limitation.

      Methodologically, the behavioral paradigm is well suited for studying escape decisions in socially housed animals, and the machine learning based classification of defensive responses is a clear strength. The computational model provides a useful formalization of how threat level, reward level, and vigilance interact and may be valuable for other laboratories studying escape, approach avoidance, or conflict situations, particularly as a way to classify behavioral outcomes after pose estimation. More generally, the work will be of interest to the neuroethology community for its detailed characterization of escape behavior under naturalistic conditions.

      Given the ethological nature of the study and the high inter individual variability reported by the authors, clarity and precision in the methods are especially important for reproducibility. While the revised manuscript addresses many earlier concerns, some aspects remain slightly difficult to follow. For example, the main text states that animals were not water deprived to avoid differences in internal state, whereas parts of the methods describe conditions in which animals were water deprived, suggesting that internal state manipulation may differ across experiments. Clearer separation and explanation of these conditions would further strengthen confidence in the work.

      To improve clarity, we have revised the Methods section to clearly distinguish between experimental conditions that involved water deprivation and those that did not.

      Overall, this study provides a rich and thoughtful analysis of how reward level and social hierarchy modulate defensive behavior through changes in vigilance. It offers a useful conceptual advance for thinking about escape behavior in naturalistic settings and lays a solid foundation for future work aimed at linking these behavioral states to underlying neural circuits.

      Reviewer #2 (Public review):

      Zhe Li and colleagues investigate how mice exposed to visual threats and rewards balance their decisions in favour of consuming rewards or engaging in defensive actions. By varying threat intensity and reward value, they first confirm previous findings showing that defensive responses increase with threat intensity and that there is habituation to the threat stimulus. They then find that water-deprived mice have a reduced probability of escaping from low contrast visual looming stimuli when water or sucrose are offered in the environment, but that when the stimulus contrast is high, the presence of sucrose or water increases the probability of escape. By analysing behaviour metrics such as the latency to flee from the threat stimulus, they suggest that this increase in threat sensitivity is due to increased vigilance. Analysis of this behaviour as a function of social hierarchy shows that dominant mice have higher threat sensitivity, which is also interpreted as being due to increased vigilance. These results are captured by a drift diffusion model variant that incorporates threat intensity and reward value.

      The main contribution of this work is quantifying how the presence of water or sucrose in water-deprived mice affects escape behaviour. The differential effects of reward between the low and high contrast conditions are intriguing, but I find the interpretation that vigilance plays a major in this process not supported by the data. The idea that reward value exerts some form of graded modulation of the escape response is also not supported by the data. In addition, there is very limited methodological information, which makes assessing the quality of some of the analyses difficult, and there is no quantification on the quality of the model fits.

      (1) The main measure of vigilance in this work is reaction time. While reaction time can indeed be affected by vigilance, reaction times can vary as a function of many variables, and be different for the same level of vigilance. For example, a primate performing the random dot motion task exhibits differences in reaction times that can be explained entirely by the stimulus strength. Reaction time is therefore not a sound measure of vigilance, and if a goal of this work is to investigate this parameter, then it should be measured. There is some attempt at doing this for a subset of the data in Figure 3H, by looking at differences in the action of monitoring the visual field (presumably a rearing motion, though this is not described) between the first and second trials in the presence of sucrose. I find this an extremely contrived measure. What is the rationale for analysing only the difference between the first and second trials? Also, the results are only statistically significant because the first trial in the sucrose condition happens to have zero up action bouts, in contrast to all other conditions. I am afraid that the statistics are not solid here. When analysing the effects of dominance, a vigilance metric is the time spent in the reward zone. Why is this a measure of vigilance? More generally, measuring vigilance of threats in mice requires monitoring the position of the eyes, which previous work has shown is biased to the upper visual field, consistent with the threat ecology of rodents.

      (2) In both low and high contrast conditions, there are differences in escape behaviour between no reward and water or sucrose presence, but no statistically significant differences between water and sucrose (eg: Figure 3B). I therefore find that statements about reward value are not supported by the data, which only show differences between the presence or absence of reward. Furthermore, there is a confound in these experiments, because according to the methods, mice in the no-reward condition were not water-deprived. It is thus possible that the differences in behaviour arise from differences in the underlying state.

      (3) There is very little methodological information on behavioural quantification. For example, what is hiding latency?

      Is this the same are reaction time? Time to reach the safe zone? What exactly is distance fled? I don't understand how this can vary between 20 and 100cm. Presumably, the 20cm flights don't reach the safe place, since the threat is roughly at the same location for each trial? How is the end of a flight determined? How is duration measured in reward zone measures, e.g., from when to when? How is fleeing onset determined?

      (4) There is little methodological information on how the model was fit (for example, it is surprising that in the no reward condition, the r parameter is exactly 0. What this constrained in any way), and none of the fit parameters have uncertainty measures so it is not possible to assess whether there are actually any differences in parameters that are statistically significant.

      These are the public reviews for the original submission. The corresponding authors responses are provided below.

      (1) We agree that reaction time can be influenced by multiple factors, including stimulus strength. Consistent with this, reaction times (i.e. latencies to flee) were substantially shorter under high-contrast conditions (Figure 3E). However, even under the same high-contrast condition, reaction times were significantly shorter in the water condition compared to the no-reward condition, suggesting that other factors such as vigilance may contribute.

      Upward-directed attention includes rearing, up-stretching, and upward head orientation, which will be clarified in the Method section. To address concerns about statistical validity, we will quantify these behaviors across the first 10 trials rather than limiting the analysis to the first two.

      As for the dominance-related results, we interpret them as reflecting both enhanced vigilance and reduced reward-seeking behavior. Time spent in the reward zone is not a measure of vigilance but an indicator of reward-seeking motivation. We will clarify this in the revised manuscript.

      (2) In Figure 3B, the difference between water and sucrose conditions did not reach statistical significance (p = 0.08). We plan to collect additional data to determine whether this is due to limited statistical power. It is also possible that some behavioral readouts are more sensitive to the differences between water and sucrose conditions. For example, Figure 3F shows that escape speed was significantly higher in the sucrose than in the water condition under high-contrast stimulation.

      Thank you for pointing this out. To control for the potential confounds related to internal state, mice were not water-deprived under any of the three conditions in Figures 3A-3H. We will clarify this in the main text and Methods. For Figures 3I-3M, which compare decision-making under no-reward and water conditions, we will conduct additional experiments using non-deprived mice in the water condition.

      (3) Hiding latency was defined as the time from stimulus onset to the animal’s arrival at the safe zone. Reaction time was quantified as the latency to flee, measured from stimulus onset to the initiation of the first flight state. The flight state was defined as locomotion exceeding 10 cm at a speed greater than 10 cm/s. Distance fled was defined as the distance covered between stimulus onset and offset for all trials. However, in trials classified as no reaction or freezing, this measure does not accurately reflect escape behavior. We will therefore rename it as distance under threat to better capture its meaning. The reward zone was defined as the region within 15 cm of the reward port at the end of the arena. Duration in the reward zone was measured as the time spent within this region during the 20 seconds following stimulus onset. In Figure 4E, the percentage of time spent in the reward zone was calculated relative to the total time the mouse remained in the arena during the 2-hour social session.

      All definitions and additional details on behavioral quantification will be included in the revised Methods section.

      (4) We appreciate the comment and agree that further clarification is needed. We will provide a more detailed description of the model fitting procedure in the revised Methods section. Specifically, the drift rate parameter (r), which reflects the perceived reward value, was constrained to zero in the no-reward condition. To enable statistical comparison across conditions, we will report uncertainty measures for all fit parameters.

      Comments on the revised manuscript:

      The manuscript has been revised and improved significantly by the addition of methodological details and new analysis. I remain, however, unconvinced by the argument that increased vigilance in the presence of reward leads to heightened escape behaviour.

      In response to my criticism that the work does not measure vigilance directly, the authors have included measures of foraging interval and foraging speed, which they state are "two direct behavioral analyses of vigilance". I disagree - like reaction time, foraging speed and foraging interval can be modulated, for example, by changes in threat sensitivity. Increased threat sensitivity comes with diverse behavioral changes that may well include increased vigilance, but foraging interval and foraging speed can certainly change without the animal expressing increased vigilance behaviors. A bigger issue I still have though, is with the conclusion that the presence of reward increases "direct escape behaviors". Comparing the no reward, water and sucrose groups indeed shows a difference (which is now clear after the split into early and late phases), but the issue is that these are different mice. As the text is written, is sounds like introducing reward will acutely increase escape. But if we look at the raw data show in Figure 2C, what I think is happening is that the presence of reward is decreasing habituation to the stimulus. The data for trials 1 and 10 in the three conditions show this - there is habituation with no reward (reaction times are all shifting to the right), a bit less with water and very little with sucrose. This is interesting in its own right and we can speculate why it might be happening, but I think this is conceptually different from what the authors are proposing.

      We agree that vigilance is not directly observable as a single variable. Our intent was not to claim that foraging speed and foraging interval provide a direct measure of vigilance, but rather to suggest that they may serve as indirect behavioral correlates.

      We also considered an alternative interpretation: these two measures could reflect perceived reward value under high-threat conditions across distinct reward types. If that were the case, animals would be expected to exhibit shorter intervals and faster speeds across no reward, water, and sucrose conditions. However, our data do not support this interpretation (Figures 3L and 3M), suggesting that these measures are more likely correlated with vigilance.

      Furthermore, it is unlikely that changes in foraging interval and speed are driven by altered threat sensitivity, as animals could not see the threat during most of the foraging bout and only encountered it at the end.

      Regarding the conclusion that the presence of reward increases direct escape behaviors, our interpretation is that increased reward value reduces habituation, thereby maintaining higher vigilance during the late phase. This was discussed in the second-to-last paragraph of the "Economic and social modulations of innate decision-making under threat" subsection in the Discussion.

      Reviewer #3 (Public review):

      Male mice were tested in a classic behavioral "flee the looming stimulus" paradigm. This is a purely behavioral study; no neural analyses were done. Mice were housed socially, but faced the looming stimulus individually, using an elegant automated tunnel (see videos for clarity).

      The additional changes made to the paper clarify the work done. While there are some limitations (male mice, weird stimulus), the general results are interesting and a valuable addition to the experimental literature. The main claim of the paper is that the different rewards (none, water, sucrose) did not change the escape properties early in learning, but did late, particularly that in the late (already experienced) conditions, reward value (assuming sucrose > water > no reward) interacted with the salience of the looming stimulus (light gray, dark gray). (Panels 3D, 3G, 3K, 3N).

      For readers, I want to note that one of the most interesting results is actually in Figure S2, where they find that a looming stimulus behind the mouse still makes a mouse run to the nest. In these conditions, the mouse runs past the looming stimulus to get to safety! (I also do love the video of the mouse running around the barriers like a snake to get home.)

      I have a few minor clarification questions and a few notes that I think would be useful additions for authors and readers to think about.

      Dominance: What does the mouse social science literature say about the "test tube" test? What can we conclude from this test? This would be useful when trying to understand what is causing the dominance/submissive difference in responses. Figure 4 shows that the dominant mice are more risk-averse than the submissive mice. Is "dominance" in the test-tube actually a measure of risk-seeking? Is the issue that the submissive mice don't think they can get back to the food-site easily, so they are less willing to sacrifice the current (if dangerous) foraging opportunity? Is the issue that the submissive mice can't get back to the nest? As I understand it, the nest was always available to all the mice, so I suspect inability to get to the nest is an unlikely hypotheses. Is the issue that the submissive mice also don't feel safe in the nest?

      The tube test is a widely used assay in the rodent social behavior literature to assess dominance hierarchies, operationally defined by the ability of one animal to force its opponent to retreat from a narrow tube. Importantly, this assay does not directly measure risk-seeking or anxiety-related traits, but rather competitive outcomes during social conflict. Furthermore, our data indicate that the behavioral responses of subordinate mice to looming stimuli are primarily driven by the visual threat itself rather than by social avoidance. This point was elaborated in the second paragraph of the “Social modulation of innate decision-making” subsection in the Results section.

      Limitations of the study: There is an acknowledged limitation to male mice, and the limitations of the small data sets that are typical of such experiments. In addition, however, it is also worth noting the strangeness of the looming stimulus, which is revealed clearly in the videos. The stimulus is a repeating growing circle, growing in a single location within the environment. The stimulus repeats 10 times, once per second. This is not what an attacking hawk or owl would look like. (I now have this image of an owl diving down, and then teleporting up and diving down again.) Note - I am fine with this stimulus. It produces an interesting experiment and interesting results. I do not think the authors need to change anything in their paper, but readers need to recognize that this is not a "looming predator".

      These "limitations" are better seen as "caveats" when folding these results in with the rest of the literature that has gone before and the literature to come. (Generally, I do not believe that science works by studies making discoveries that change how we think about problems - instead, science works by studies adding to the literature that we integrate in with the rest of the literature.) Thus, these caveats should not be taken as problems with the study or as fixes that need to be done. Instead, they are notes for future researchers to notice if differences are found in any future studies.

      Thus, my only suggestion is that I think authors could write a more careful paper by using the past and subjunctive tense appropriately. Experimental observations should be in past tense, as in "the influence of reward was contextdependent and emerged in the late phase" instead of "the influence of reward is context-dependent and emerges in the late phase" - it emerged in the late phase this once - it might not in future experiments, not due to any fault in this experiment nor due to replicability problems, but rather due to unexpected differences between this and those future experiments. At which point, it will be up to those future experiments to determine the difference. Similarly, large conclusions should be in the subjunctive tense, as in "these data suggest that threat intensity is likely to be the primary determinant of decision making" rather than "threat intensity is the primary determinant of decision making", because those are hypotheses not facts.

      We thank the reviewer for the helpful suggestions and have revised the Abstract accordingly.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Figure 5: The points in panel 5G and 5H are unreadable. What are these stars and symbols supposed to mean? They are also too small to see without zooming way in.

      We have increased the symbol size.

      Figure 5: What is the final panel of 5J? I did not understand this panel at all. The first three panels of 5J (threat-based detection, reward-based detection, vigilance-based detection) are, I believe, three patterns we should look for in the data. But then what is the "experimental results" section? It contains all three, but they don't overlap? Shouldn't we have an experimental results section for each condition?

      Panel 5J was to compare three hypothesized decision patterns with the experimentally observed data. To make this distinction explicit, we have revised the panel titles to: “H1: Threat-based decisions,” “H2: Reward-based decisions,” “H3: Vigilance-based decisions,” and “Experimental results.”

      Thank you for including the videos. They made the task construction and the stimulus much clearer.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      While the evidence in favour of the two gradients largely supports the claims, the evidence for a new visual field map cluster in the anterior temporal lobe falls short of the level used historically when identifying visual field maps in the visual cortex and is, at present, not convincing. More specifically, the progressions of polar angle within the putative anterior lobe cluster are highly variable across subjects. Few subjects have convincing polar angle reversals at either the horizontal or vertical meridians. In other cases, a putative border is shown that spans different polar angles, which does not align with the accepted definitions for visual field maps in the cortex.

      We agree with the reviewer that more evidence could be provided in support of retinotopic representations within the anterior temporal lobe. We have performed a number of new analyses to further explicate the receptive field properties of this anterior temporal lobe visual representation. We have pasted updated Figure 2e-i. We have added additional participants, increasing the total number from N=12 to N=21. In panel g, we show that in this larger group, we can still observe pRFs that are about 3x larger than those in early visual cortex, and that the relationship between their size and eccentricity shows the expected steeper slope compared to these early representations. In this new participant group, we also illustrate the visual field coverage of the left and right anterior temporal lobe representations (panel h). As expected, the left hemisphere pRFs largely sample the right visual field, and right hemisphere pRFs largely sample left visual space. One can also see that both the upper and lower visual fields are sample quite evenly, consistent with the hemi-field representation of visual field maps observed in earlier visual cortex. To quantify whether there is a left-right contralateral bias in the sampling of visual space (and to test whether such a bias is significantly different in each hemisphere), we calculated for each pRF a laterality index as previously defined by Sheremata and Silver (2015) according to the equation below:

      Where resulting values of 1 mean the pRF is contralateral, 0.5 is no laterality bias, and 0 is ipsilateral bias. Additionally, we input pRF sigma values that were adjusted for the non-linearity exponent as defined by Kay et al. (2013). For the purposes of visual comparison, we subtracted 0.5 from index values so that resulting laterality scores were relative to 0 to represent the center of the visual field, and then values were inverted with a -1 scalar so that left hemisphere pRF laterality index values are plotted on the right side of space, and the right hemisphere on the left as shown in panel i. The laterality index was calculated for each pRF for a given participant and then averaged within that participant to result in a single mean laterality index for the left hemisphere pRFs and a single index for their right hemisphere pRFs. The histograms illustrated in panel i depict density of participants (kernel smoothed). We find a significant difference between laterality indices with left AT pRFs showing significantly rightward index values compared to right AT pRFs (paired-samples t-test, t(20) = 7.6, p = 2.7 x10<sup>-7</sup>). These data thus offer stronger evidence of a hemifield representation with a contralateral bias, and it should also be noted that there is stronger ipsilateral coverage in these high-level visual pRFs compared to earlier visual field maps like V1, which is consistent visual field maps in latera stages of the visual processing hierarchy as quantified by Mackey et al. (2017).

      Lastly, we note that the progression of polar angle values on the cortical surface is certainly not as strikingly topographic as in visual field maps V1 through hV4. This is perhaps a result of the strong ipsilateral visual field coverage in which pRFs whose centers were near or within the ipsilateral field (especially those near the fovea) are not visualized appropriately when using a contralateral colormap. It is also possible that at this very late stage of visually-responsive cortex within entorhinal cortex that retinotopic topography becomes less clear as is the case in higher stages of the dorsal visual stream. To improve visualization, we have created a new Supplemental Figure 6 using a binary color map that colors lower and upper visual field in separate colors and extends into the ipsilateral visual field (pasted below for convenience). We hope that this color map helps to show the upper and lower visual field coverage. While there is a clear radial eccentricity gradient within these AT pRF clusters, and while most participants do show a polar angle gradient that runs perpendicular to this radial eccentricity gradient as expected for a visual field map, we do agree that it is difficult to observe polar angle traversals as clearly as in earlier visual cortex. Nonetheless, the presence of these pRF clusters which show their own distinct eccentricity representation (i.e., a foveal confluence) and a full sampling of the contralateral visual space is still consistent with our anatomical model’s prediction in which PC2 anchor points predict foveal representations shared by visual field map clusters. While the topographic clarity of these representations on the cortical surface is less than earlier visual cortex, the existence of contralateral representations of visual space with a full eccentricity gradient that spans the upper and lower visual field is strongly supported by the data and consistent with our anatomical model’s prediction that there should have been a distinct eccentricity gradient. These findings are also consistent with work showing that the human hippocampus also shows sensitivity to contralateral visual space (Silson et al., 2021) and suggests the hippocampus may inherit this contralateral bias from this entorhinal visual representation. We have updated the manuscript to incorporate these new findings, and refer to these AT clusters as contralateral visual representations, remaining agnostic to whether or not they can be fully defined as topographic maps which can be the focus of future work using smaller voxel sizes to better capture small topographic gradients.

      We have revised the manuscript to incorporate these points in the following sections.

      Line 466: “We performed pRF mapping on 21 participants with high-contrast, …”

      Line 601-625: “To produce maps of visual field coverage (Figure 2h) similar to previous work, … The histograms illustrated in Figure 2i depict density of participants (kernel smoothed).”

      Line 236-246: “We find that consistent with its high position within the processing hierarchy, … We find a significant difference in laterality indices between left and right AT pRF’s (pairedsamples t-test, t(20) = 7.6, p = 2.7 × 10-7).”

      Line 373-383: “The organization of polar angle in anterior temporal cortex was not as orderly as earlier visual cortex, … in more posterior portions of ventral occipitotemporal cortex.”

      Reviewer #2 (Public review):

      Weaknesses:

      (1) The neurobiological model does not take into consideration present knowledge about the microstructural organization of the visual system. This limits the way the results are interpreted correctly. Critical information on the layer-specific myeloarchitecture and cytoarchitecture (and their relation to cortical thickness), as explored for example by Sereno et al. 2013 Cereb Cortex, is missing. There is no information given with respect to how different visual areas differ in their microstructural profile. It is also not mentioned that cortical parcellation is indeed characterized by sharp boundaries between areas, rather than structural gradients, so it remains unclear why focusing on a gradient is of interest. The authors cite the parcellation atlas by Glasser et al. 2016, but do not discuss the rationale of this publication, which was not the definition of gradients, but the definition of sharp boundaries for cortex parcellation. Indeed (as explained below), the results of the authors seem to a large extent to be driven by cortex parcellation, but instead of acknowledging this fact, the authors write (line 179) that "we hypothesize that these local deviations from the canonical thickness and density of cortex underlie the finer-scale division of visual cortex into categorically distinct regions. That is, does the realization of the cortex into distinct regions involve these regions becoming more distinct from a prototypical cortical sheet (i.e., gradient 1)?" - While the first sentence is reasonable, the second sentence is pure speculation ignoring present knowledge on cortical parcellation of this area according to which there is no "prototypical cortical sheet", but each area has its distinct microstructural profile.

      We thank the reviewer for this important comment. We first want to point out that we believe there is a conceptual misunderstanding on the part of the reviewer, as we address in our lengthy response below. In this response, we explain that our findings capture what we believe is a novel finding—that variation across participants in the cortical sheet is not random across the spatial expanse of cortex but respects its functional boundaries—which we view as a finding that is complimentary to the current knowledge about the microstructure of visual cortex. It was not our intention to ignore or gloss over this present knowledge, but instead show that variation in these cortical microstructures across brains is not random.

      We agree that incorporating current knowledge about the microstructural organization of visual cortex, including its laminar architecture and sharp areal boundaries, is critical for situating our findings within the broader literature. In response, we have added key background information on the relationships among cytoarchitecture, myeloarchitecture, and cortical thickness, as described in previous studies (for example, Maingault et al., 2021; Sereno et al., 2013; Shafee et al., 2015). While our study does not aim to capture layer-specific properties per se, which would require different imaging modalities and higher-resolution data, we focus on spatial properties tangential to the cortical surface.

      We first address a concern that the particular parcellation might be driving effects with an analysis showing that we believe our finding is robust to this concern. As suggested by the overall negative covariance observed between cortical thickness and tissue density, we further confirmed this relationship not only across larger visual ROIs, which could potentially reflect effects of arealization, but also within individual ROIs at a finer spatial scale. To avoid potential circularity in ROI definition, we used a visual ROI atlas derived from population-level retinotopy based on independent datasets (Abdollahi et al., 2014). We found that at the global level, cortical thickness and T1w/T2w ratio showed a strong negative correlation across visual ROIs (Fig. 3, revised Supp. Fig. 3a & b). Although only a portion of the visual cortex is clearly delineated in this atlas, we replicated similar results across the entire visual cortex using the MMP atlas (Glasser et al., 2016). At the within-ROI level, we found robust negative correlations between cortical thickness and T1w/T2w ratio across most visual ROIs in both hemispheres, with the notable exception of V1, V2 and VO1, which exhibited a positive relationship, consistent with prior work (for example, Maingault et al., 2021; Sereno et al., 2013; Shafee et al., 2015). These results highlight both common and distinct microstructural profiles across the visual cortex and provide important context for interpreting our data-driven findings.

      We also want to address what we think is a conceptual misunderstanding by the reviewer, which likely resulted from a lack of clarity on our part. The reviewer’s confusion likely results from the fact that we theoretically “transposed” the typical PCA analysis such that we get a subject-wise contribution (PC loadings) per participant (also see response to next point), which is how we’re able to relate inter-participant variability in their loadings to behavior in Figure 3. This is also why we refer to a “typical” cortex/cortical sheet because the surface maps being visualized for PC2 can be thought of as a map explaining variance of deviation orthogonal to PC1 (which captures the primary relationship between thickness and T1/T2). Thus, because PC2 is orthogonal to PC1, it captures the spatial pattern in which participants deviate from the primary relationship (e.g., the typical relationship). Therefore, if a given participant is far from the PC1 vector and has high PC2 loading, their cortical sheet is either thicker or more myelinated than predicted by the PC1 relationship and is therefore more distinct from the “typical” or “average” cortical sheet values captured by PC1. We want to emphasize that PCA is agnostic to spatial structure across the cortex. Thus, the fact that deviation from the primary thickness-myelination relationship (i.e. PC2) captured by PC1 had any spatial structure at all is interesting. Furthermore, the fact that the spatial structure of PC2 across the cortical sheet seems to separate visual cortex into its constituent processing streams is also interesting. Therefore, we are not speculating but rather describing the PCA model itself whereby a participant’s loading on PC2 describes their deviation or distinctness from the PC1 relationship. The fact that PC2 has spatial structure on the cortical sheet (which did not have to be true) and the fact that this structure seems to capture broad borders between visual processing streams and field maps is what we find interesting and quantify within the paper. We hope this additional explanation clarifies the broader theoretical thrust of the paper. We view these findings as complimentary to the present knowledge of the microstructural organization of the visual system. Our findings suggest that variability in these microstructural features across participants (PC2) don’t occur randomly across cortex but seem to respect the functional borders of the neural populations of the underlying cortical sheet.

      Regarding the concern that our gradient approach may contradict established knowledge of cortical arealization, we would like to clarify that the primary goal of our gradient analysis is not to redefine visual areas, or to go against cortical arealization, but to explore the continuous variation in cortical architecture across brains that may co-exist alongside sharp boundaries which is phenomenon complementary to the arealization. In our study, cortical thickness maps were regressed for curvature before entering any analyses, given the covariance between cortical folding and area borders (Fischl et al., 2008). We acknowledge that cortical parcellation is traditionally characterized by discrete transitions between areas. However, our results suggest that gradients of cortical properties—particularly those shared across participants—may capture supra-areal organizing principles that reflect how distinct regions relate to one another within a broader cortical sheet.

      Finally, we agree with the reviewer that the phrase “prototypical cortical sheet” was speculative and potentially misleading. We have removed this language from the manuscript and revised the corresponding discussion.

      We have revised the manuscript to incorporate these points in the following sections.

      Line 92-94: “Thickness and density maps showed a robust anti-correlation both at the coarse across-area level based on an independent parcellation and at the finer within-area level, except in primary regions (Figure S3a, b).”

      Line 350-353: “The convergence pattern, arising from the negative correlation between thickness and density, is consistent with previous findings and may support the balloon model, whereby cortical thinning is associated with tangential stretching due to myelination.”

      Line 188-189: “That is, does the arealization of cortex into distinct regions involve these regions becoming more distinct from a typical cortical sheet (i.e., gradient 1)?”

      (2) Instead of building on present, detailed knowledge of brain anatomy and in-vivo cortex parcellation of the visual system and its known relation to visual maps, the authors focus on two metrics of cortex architecture (mean T1/T1 over depth and cortical thickness), and conduct a PCA to explore their shared variance. It needs to be clarified if the PCA was conducted correctly. There is no mention of standardizing the variables, which could bias the results. In addition, in a PCA, all possible features are categorized as vector components, and those are scanned through the samples, hence, one such analysis per vertex. But the authors write "in which participants are features and cortical vertices are samples" and "the thickness and tissue density maps were concatenated". This needs clarification. The architecture of the PCA should be visualized better.

      We thank the reviewer for pointing out the need to clarify the PCA methodology. In response, we have revised the Methods section to provide a clearer and more accurate description of our approach.

      We also would like to point the reviewer’s attention to Figure 1a, in which the PCA was illustrated graphically. The reviewer’s confusion likely results from the fact that we theoretically “transposed” the typical PCA analysis such that we get a subject-wise contributions (PC loadings) per participant, which is how we’re able to relate inter-participant variability in their loadings to behavior in Figure 3. This is also why we refer to a “typical” cortex/cortical sheet because the surface maps being visualized for PC2 can be thought of as a map explaining variance of deviation orthogonal to PC1 (which captures the primary relationship between thickness and T1/T2). Thus, because PC2 is orthogonal to PC1, it captures the spatial pattern in which participants deviate from the primary relationship (e.g., the typical relationship).

      We have revised the manuscript in the following sections.

      Line 493-502: “For each hemisphere, individual cortical thickness and T1/T2-weighted ratio maps from all HCP-YA participants—each represented as an M × N matrix, … corresponding participant-wise contributions (i.e., PC loading or individual weights) in pairs.”

      (3) Because the PCA only contains two features, PC1 is driven by the positive relationship between cortical thickness and mean T1/T2, whereas PC2 is driven by their negative relationship. Because in the early visual cortex, cortical thickness and mean T1/T2 correlate positively, it naturally follows that PC1 relates to pRF size (but mediated by the actual cortex parcellation). However, it is unclear why this insight is interesting. I also do not share the view that "these findings demonstrate that gradient 1 acts as a global gradient enveloping the entire visual cortex (...) while gradient 2 acts as a local gradient specific to individual visual streams". I think this relationship between cortical thickness and T1/T2 ratio does not have much to do with local and global gradients. But if so, stronger arguments as to why this should be the case should be presented. What the authors make of this result (particularly the discussion starting line 366) is not clear to me. I cannot follow the line of argumentation, which in my view is too far away from the data.

      We appreciate the reviewer’s thoughtful comments and agree that, in general, cortical thickness and T1w/T2w ratio tend to be negatively correlated, with early visual areas (i.e., V1 and V2) representing a notable exception—an observation we highlight and support with evidence in R2. Given this overall pattern of correlation, it may seem intuitive to interpret PC1 as capturing a convergent relationship across the two metrics, and PC2 as reflecting their divergence. Alternatively, one can think of PC2 as the orthogonal residuals from the linear relationship between thickness and myelin captured by PC1. In this framework, PC2 is not necessarily the inverse correlation, but instead what is left unexplained through a simple linear model. However, it is important to note that PCA is inherently agnostic to spatial structure, as our PCA operates solely on inter-subject variance. As such, the spatial patterns observed in the resulting component maps are not direct or trivial consequences of the input correlations.

      Upon examining the spatial properties of the PCA-derived maps (Fig. 1d), we found that PC1 manifests as a large-scale, low-frequency gradient spanning broad portions of the visual cortex, whereas PC2 exhibits a fine-scale, high-frequency pattern confined to subregions of the visual cortex (quantified in Fig. 1f, g). Our initial use of the terms “global” and “local” may have inadvertently implied functional interpretations beyond our intent. We have revised the manuscript to clarify that these descriptors were intended purely to convey differences in spatial scale based on the observed frequency content of the gradients.

      Motivated by the reviewer’s comment, we performed additional analyses to explicitly test whether the PCA components reflect consistent (i.e., global) or variable (i.e., local) relationships across visual ROIs. Specifically, we examined whether the direction and magnitude of PC1 and PC2 scores within each ROI align with the global relationships between cortical thickness and tissue density. As shown in the revised Supp. Fig. 3e, we found that in most ROIs, vertices with high PC1 scores consistently exhibit high cortical thickness and low T1w/T2w ratios, while those with low PC1 scores show the opposite pattern. This within-ROI consistency mirrors the largescale cross-ROI correlation structure (see Supp. Fig. 3a), supporting the interpretation of PC1 as reflecting a large-scale, cortex-wide organizational principle. In contrast, PC2 shows more heterogeneous profiles across ROIs, with peaks and troughs that differ in the two metrics. This variability suggests that PC2 captures more localized, region-specific features.

      We have incorporated the results of these new analyses into the Results section to strengthen our argument regarding the spatial scale and cross-regional consistency of the PCA-derived gradients:

      Line 102-107: “Within-area analyses further confirmed that PC1/2 represent the consistent/deviating components … while PC2 represents the spatial divergence from this commonality.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      Through collaborative discussions among the reviewers, we first summarised the key recommendations for enhancing the significance and strengthening the evidence of the work - integrating public reviews and recommendations to authors by each reviewer individually. The individual reviewer recommendations can be found below this.

      (1) Modelling component 2

      The geodesic model for component 2 is interesting but we can recommend ways to improve the evidence and interpretation (see Reviewer 1 comments). As the polar angle reversals are inconsistent and boundaries ambiguous, the OTS maps do not meet the standard of evidence required for showing a new map. The 181 pRF maps available for these HCP data would provide an independent more powerful test of the OTS map cluster. To further strengthen the evidence for the proposed correspondence of foveal confluences and gradient 2, why not define the geodesic model anchoring points based on retinotopic measures, e.g., using HCP pRF data? About the current anchoring points for the geodesic model, what were the criteria - were they objective to avoid circularity?

      We appreciate the reviewer’s suggestion to incorporate the HCP 7T retinotopy dataset as an independent test of the proposed geodesic model and its relation to foveal confluences and gradient 2. We agree in principle that such data could provide a valuable validation resource. However, as detailed in the publication accompanying the HCP 7T retinotopy dataset (Benson et al., 2018), the authors recommend a threshold of 9.8% variance explained to distinguish reliable pRF estimates from noise. As illustrated in their Figure 4, this thresholded pRF data shows poor signal coverage in higher-order visual regions, particularly those along the occipitotemporal sulcus (OTS), where gradient 2 effects are most prominent in our data. This lack of reliable pRF signal in these regions limits the utility of the HCP retinotopy data for anchoring the geodesic model or validating the observed spatial gradients.

      To address this limitation, we relied on our in-house data collected using high-contrast, naturalistic images designed to robustly activate high-level visual areas. This approach allowed us to define more complete and consistent topographic patterns in the regions of interest. We have thus expanded the size of this in-house dataset to N=21. We also point the editor’s attention to the response to Reviewer 1’s first comment regarding the visual field maps for a more detailed response to this point. For convenience, we have pasted the Figure 2 e-i panels in which we conduct additional analyses showing that these anterior temporal pRF clusters tile contralateral visual space as one might expect (Fig 2h), and significantly differ across hemispheres in their laterality bias (Fig 2i). We have revised the manuscript accordingly.

      To mitigate the concern of circularity in defining the geodesic model’s anchor points, we conducted a split-half cross-validation. Anchors were defined on one half of the participants and used to predict the PC2 map in the other half. The PC2 maps across the two halves were highly similar (r = 1.00, p < 0.001), indicating strong reliability. Importantly, the cross-predicted geodesic model accounted for a significant portion of variance (r<sup>²</sup> = 0.23) in the held-out PC2 map, suggesting that the geodesic organization is not an artifact of overfitting or circular reasoning. We have revised the manuscript accordingly:

      Line 139-142: “A split-half cross-validation yielded similar results, … underlying the spatial organization of PC2.”

      (2) Speculation about prototypical cortical sheet

      You hypothesise that gradient 1 characterises a global "prototypical cortical sheet" characteristic, with gradient 2 reflecting that regions become more distinct from this prototype. There is an alternative simpler possibility: the data can be explained by the stronger relationship between cortical thickness and T1/T2 ratio in early compared to late sensory areas, as can for example be seen in Glasser et al. 2016 Nature, Figure 4. We recommend omitting or balancing the statement about a "prototypical" cortex, and integrating findings on cortex parcellation and the view that sharp boundaries characterize transitions between high and low T1/T2 and cortical thickness areas.

      Please see R2 for reviewer #2

      (3) Confounds

      We'd like to see more data to understand the contributions of data quality to these results. For the component 1 gradient specifically, could its features be influenced by spatial SNR inhomogeneities? Could the developmental effects for both gradients be explained by lower SNR and other data quality markers in younger and older participant data? We missed appropriate tests that gradients develop differently across age, controlling for such confounds (Reviewer 1 comments).

      Regarding the reviewer’s concern about the component 1 gradient, we believe it is unlikely to be merely a consequence of uneven spatial SNR. Our findings are consistent with previous histological studies demonstrating systematic variations in cortical architecture—specifically, thinner cortex (Wagstyl et al., 2020) and higher myelin content (Dinse et al., 2015) in occipital compared to ventral visual regions. This correspondence between in vivo MRI-derived measures and postmortem histology suggests that the large-scale organization captured by PC1 is grounded in biologically meaningful cortical architecture, and not an artifact of SNR variability.

      To statistically assess whether the two PCs show different developmental trajectories across age, we performed an ANOVA with age, LC, and their interaction as factors on LC’s similarity to PC (i.e., r ~ age + LC + age × LC). Significant age × LC interactions were observed in the developmental (HCPD: F<sub>1,118</sub> = 257.01, p < .001) and aging (HCPA: F<sub>1,132</sub> = 263.85, p < .001) cohorts, but not in the young adult cohort (HCPYA: F<sub>1,202</sub> = 0.02, p = 0.80). These findings indicate that the two gradients show distinct age-related changes during development and aging but remain stable in young adulthood. We have revised the manuscript accordingly:

      Line 313-327: “Examining the correlation between the young adult gradient and LC … F<sub>1,132</sub> = 263.85, p < 0.001).”

      (4) Implementation of PCA

      The manuscript raises questions about the correct implementation of the PCA - please clarify that the variables were first standardised to enable fair weightings, and visualise the PCA matrix in more detail than in Figure 1a to ensure the samples and features are correctly defined (Reviewer 2).

      Please see R3 for reviewer #2

      References

      Abdollahi, R. O., Kolster, H., Glasser, M. F., Robinson, E. C., Coalson, T. S., Dierker, D., Jenkinson, M., Van Essen, D. C., & Orban, G. A. (2014). Correspondences between retinotopic areas and myelin maps in human visual cortex. NeuroImage, 99, 509–524. https://doi.org/10.1016/j.neuroimage.2014.06.042

      Benson, N. C., Jamison, K. W., Arcaro, M. J., Vu, A., Glasser, M. F., Coalson, T. S., Van Essen, D. C., Yacoub, E., Ugurbil, K., Winawer, J., & Kay, K. (2018). The HCP 7T Retinotopy Dataset: Description and pRF Analysis. https://doi.org/10.1101/308247

      Dinse, J., Härtwich, N., Waehnert, M. D., Tardif, C. L., Schäfer, A., Geyer, S., Preim, B., Turner, R., & Bazin, P.-L. (2015). A cytoarchitecture-driven myelin model reveals area-specific signatures in human primary and secondary areas using ultra-high resolution in-vivo brain MRI. NeuroImage, 114, 71–87. https://doi.org/10.1016/j.neuroimage.2015.04.023

      Fischl, B., Rajendran, N., Busa, E., Augustinack, J., Hinds, O., Yeo, B. T. T., Mohlberg, H., Amunts, K., & Zilles, K. (2008). Cortical Folding Patterns and Predicting Cytoarchitecture. Cerebral Cortex, 18(8), 1973–1980. https://doi.org/10.1093/cercor/bhm225

      Glasser, M. F., Coalson, T. S., Robinson, E. C., Hacker, C. D., Harwell, J., Yacoub, E., Ugurbil, K., Andersson, J., Beckmann, C. F., Jenkinson, M., Smith, S. M., & Van Essen, D. C. (2016). A multimodal parcellation of human cerebral cortex. Nature, 536(7615), 171–178. https://doi.org/10.1038/nature18933

      Kay, K. N., Winawer, J., Mezer, A., & Wandell, B. A. (2013). Compressive spatial summation in human visual cortex. Journal of Neurophysiology, 110(2), 481–494. https://doi.org/10.1152/jn.00105.2013

      Mackey, W. E., Winawer, J., & Curtis, C. E. (2017). Visual field map clusters in human frontoparietal cortex. eLife, 6, e22974. https://doi.org/10.7554/eLife.22974

      Maingault, S., Pepe, A., Mazoyer, B., Tzourio-Mazoyer, N., & Crivello, F. (2021). Characterization of late structural maturation with a neuroanatomical marker that considers both cortical thickness and intracortical myelination. https://doi.org/10.1101/2021.02.24.432645

      Sereno, M. I., Lutti, A., Weiskopf, N., & Dick, F. (2013). Mapping the Human Cortical Surface by Combining Quantitative T1 with Retinotopy†. Cerebral Cortex, 23(9), 2261–2268. https://doi.org/10.1093/cercor/bhs213

      Shafee, R., Buckner, R. L., & Fischl, B. (2015). Gray matter myelination of 1555 human brains using partial volume corrected MRI images. NeuroImage, 105, 473–485. https://doi.org/10.1016/j.neuroimage.2014.10.054

      Sheremata, S. L., & Silver, M. A. (2015). Hemisphere-Dependent Attentional Modulation of Human Parietal Visual Field Representations. The Journal of Neuroscience, 35(2), 508–517. https://doi.org/10.1523/JNEUROSCI.2378-14.2015

      Silson, E. H., Zeidman, P., Knapen, T., & Baker, C. I. (2021). Representation of Contralateral Visual Space in the Human Hippocampus. The Journal of Neuroscience, 41(11), 2382–2392. https://doi.org/10.1523/JNEUROSCI.1990-20.2020

      Wagstyl, K., Larocque, S., Cucurull, G., Lepage, C., Cohen, J. P., Bludau, S., Palomero-Gallagher, N., Lewis, L. B., Funck, T., Spitzer, H., Dickscheid, T., Fletcher, P. C., Romero, A., Zilles, K., Amunts, K., Bengio, Y., & Evans, A. C. (2020). BigBrain 3D atlas of cortical layers: Cortical and laminar thickness gradients diverge in sensory and motor cortices. PLOS Biology, 18(4), e3000678. https://doi.org/10.1371/journal.pbio.3000678

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to understand, in the context of leaf shape, how the constraints imposed by development inform evolution. Leaf shape is a good place to study the influence of development on evolution because it is a trait that exhibits a lot of diversity, and the developmental mechanisms that give rise to leaf shapes are apparently rather conserved across angiosperms.

      As part of the motivation for their work, the authors cite a previous study (Geeta et al), which found that in angiosperm phylogenies, transitions from complex to simple leaf shapes occur through evolution more often than transitions in the opposite direction. Is this due to developmental constraints or adaptation?

      The authors undertake two parallel lines of work:

      (1) Extending the study of Geeta et al with more data, consisting of both phylogenies and a shape classification dataset. The conclusion from this line of inquiry is that transitions from lobed to unlobed leaves are more common than transitions away from unlobed leaves.

      (2) The authors conduct evolution simulations in a computational model of leaf development. Here, they look at {\it neutral} mutations and whether simply neutral evolution is sufficient to drive the observed trend.

      The conclusion of the second part of the work is that the driver of the evolution toward simple leaf shape is entropy: there are more ways to make unlobed leaves than to make lobed leaves (at least in terms of gene regulation parameters that will produce the two leaf types). The argument is that random gene regulatory networks are more likely to produce unlobed leaves than lobed leaves; therefore, neutral evolution drives this trend.

      Data Analysis

      Roughly $9000$ images of leaves were classified into 4 categories: unlobed, lobed, dissected, and compound. These labels were applied to the tips of 5 phylogenetic trees of angiosperms (3 resolved at the genus level and 2 at the species level). By fitting a continuous-time Markov chain to the labelled trees, the authors claim that there is a significantly higher rate of transition to the unlobed leaf shape compared to transitions to more complex shapes.

      Simulation

      First, the authors validate a computational model (Runions et al) for leaf growth on an experimental dataset. By changing parameters in the model, they can recapitulate the morphological changes in the shapes of Arabidopsis leaves engendered by expression of two particular genes.

      Then the authors run an evolutionary model (without selection, just random mutations) on top of the computational leaf development model. As the random walk in parameter space reaches a stationary distribution, they look at both the proportions of the leaf categories in the steady state as well as the transition rates between different categories. The result is that transitions to unlobed leaves are more common than from unlobed leaves.

      We thank the reviewer for the helpful and clear summary of our work.

      General Comments

      The authors use angiosperm phylogenies from other works as the basis for the data analysis part of their work. Given the centrality of these phylogenies for their conclusions, more information is needed about how these phylogenies were constructed and what they mean. What is the timescale that they span? What method is used to infer them? What regions of DNA were sequenced in order to build the phylogenies? Also, maybe some more discussion of angiosperm evolution (e.g., when was the most recent common ancestor of all angiosperms?) would help put the study in context.

      We also need a more in-depth discussion of the computational model. What are all the $>100$ parameters doing, and what informs the seemingly strange mutational model that changes parameters by 3 orders of magnitude?

      I am confused about how the rates of transitions were inferred from the phylogeny. Here, one has a phylogeny inferred by some method (which needs to be described in more detail), and just the leaves are labelled. It is stated in the methods that BayesTraits was used to infer the transition rates. I realize this method is probably documented elsewhere, but a bit of a summary of how it works and how to interpret its results would (1) make the paper more selfcontained and (2) if the algorithm is credible, make the results firmer.

      We thank the referee for the suggestion to make the paper more accessible. The tool we use to infer transition rates from the phylogenies, BayesTraits, is standard in the field. However, the referee is right that for an interdisciplinary journal, it may be helpful to more fully flesh out how these methods work. To that end, we have added an additional section "Phylogenetic rate inference" in the supplementary information that includes a longer description of how BayesTraits works, and how we used it to infer transition rates from phylogenies.

      All trees are shown in the supplementary information section "Phylogenetic trees" with scale-bars showing the amount of time or genetic change that the trees span. For a broader discussion of angiosperm evolution, there is supplementary information section "The adaptive significance of leaf shape review".

      Regarding the more in-depth discussion of the computational model, we have added supplementary information section S1 "Leaf model details" to give a more detailed description of the leaf model.

      I am a bit skeptical of the authors' interpretation of the biological trend (of complex to simple leaf shapes) as being driven by neutral evolution. Why does one expect that the mutations generated by the random walk models described in the work are in fact neutral mutations?

      A random walk is a well-established way of modelling the dynamics of neutral evolution in the monomorphic regime, where the population has a narrow diversity of different genotypes. In the higher mutation rate polymorphic regime, where the diversity of genotypes in the population is larger, we also expect that a random walk should still recapitulate the correct average transition rates. The purpose of the simulations is not to model every aspect of population genetics, but to ask whether developmental bias alone is sufficient to generate the observed directional asymmetry. By assigning equal fitness to all viable leaves, we isolate the contribution of development from that of selection. The agreement with the phylogenetic transition rates therefore demonstrates sufficiency rather than exclusivity: selection may also contribute, but it is not required to explain the observed bias We discuss the evidence for the role adaptation in leaf shape further in supplementary information section "The adaptive significance of leaf shape review".

      If the entropy of simple leaf shapes is higher than that of complex leaf shapes, why did we have complex leaves at all? I suspect the authors might argue that this is due to selection. In that case, what allows these complex shapes to become simpler? Wouldn't they be losing the selective advantage that drove them to be more complex in the first place? Or maybe the idea is that the rates are inferred assuming some steady state that generates the phylogeny? I did not understand this point.

      The entropy language is a useful framing. Within that framework, one can view our study as showing that the entropy (defined here as the logarithm of the volume of parameter space mapping to a phenotype) of simple leaf shapes is higher than that of complex leaf shapes. If this entropy were to be ignored, then all states would be equally likely in our simulations, where we do not take fitness differences into account. What we show is that the differences in entropy -- related to differences in volumes of the parameter space that map to different phenotypes -- also affects the rates. The inferred transition rates for both simulation and phylogeny from unlobed to more complex shapes are lower than vice versa but not zero. Therefore, complex leaf shapes arise stochastically through mutation and in this model would eventually reach a steady state proportion, even in the absence of selection.

      Are the rates of transitions between leaf types inferred for the phylogeny assuming that the phylogeny is generated by the steady state of some Markov process? (I think the answer is no: in that case, how does one explain the initial condition?)

      The tool we use to infer transition rates from phylogenies—BayesTraits—allows the initial state at the root of the tree to vary during the numerical optimisation (Pagel, 1994). Therefore, it is not assumed that the initial state is generated by the steady state of the Markov process.

      If I take the mutation model (random walk) seriously, then shouldn't I expect that this steady state obeys detailed balance? In that case I should have $p_i r_{i\to j} = p_j r_{j\to i}$ for each of the occupancies $\{ p_i\}$ and transition rates $r_{i\to j}$ for the shape categories. How close are the rates inferred from the phylogenies to obeying detailed balance? Presumably, the Markov chain fitted to the simulation data obeys detailed balance because the mutation model itself does?

      BayesTraits allows off-diagonal transition rates of the rate matrix to vary freely during numerical optimisation (Pagel, 1994). Therefore, there is no requirement for the detailed balance to hold for the inferred rate matrix. For our simulations, the mutations are symmetric at the parameter level, therefore at this level, the process would be expected to obey the detailed balance.

      I find it hard to take the discussion of development seriously without some consideration of mechanics. Presumably, the mechanics are hidden in the computational leaf development model, but this model is not discussed in enough detail for the reader to know. It seems to me that the interesting question is: what are the {\it mechanical} constraints on development that drive the apparent trend in evolution towards simpler leaf shapes? Maybe it is something about the type of differential growth needed to make complex leaf shapes less robust to mutation. But in this case, I would assume that selection plays a role in the complexity of shape. In any case, a better understanding (or explanation) of the computational model is needed to make this interpretation.

      We thank the referee for the suggestion to make the paper more accessible. We have added a more detailed and pedagogical description of the model from (Runions, Tsiantis and Prusinkiewicz, 2017) in the supplementary information section S1 "Leaf model details". We also note that Fig. 5 in the methods that gives an overview of how the model works, including some mechanical aspects of development and growth.

      More generally, mechanics is one component of the developmental map that determines which parameter combinations produce viable leaf morphologies. Our analysis concerns the geometry of this complete developmental map, irrespective of whether its constraints arise from gene regulation, tissue mechanics, or their interaction.

      On the interesting question of what is causal, perhaps the example in figure 2 is helpful. We focus on two parameters, a morphogen repression strength, and a duration of growth. A key physical process here is called webbing, where cellular growth fills in the gaps between branching veins. This process flattens the leaf structure and creates a continuous, solid leaf blade (lamina). Strong webbing, characterized by a significant resistance to stretching and bending, results in a smoother margin (Runions, Tsiantis and Prusinkiewicz, 2017). The morphogen repression strength affects the physical parameters that determine how strong the webbing is. The duration of growth determines how long the leaf has to grow. Varying these two parameters varies the physical processes that determine leaf shape. The mechanics of growth operate downstream of these parameters that we vary in our evolutionary simulations according to the details of the leaf developmental model.

      Some discussion of timescales is needed, especially when invoking neutral evolutionary arguments. If a neutral mutation occurs, its time to fix in a population of size $N$ is $\sim N$ generations. What are the relevant angiosperm population sizes and the number of mutations that separate branches on the tree? Are timescales remotely consistent with e.g., the age of angiosperms on Earth?

      Neutral processes have a well-established role in key aspects of angiosperm evolution, for example genome complexity (Lynch and Conery, 2003). This would suggest that the relevant time scales and generation times are not completely prohibitive of neutral processes also playing a role in the evolution of angiosperm leaf shape. Effective population sizes in plants are highly variable but estimates span 10^3-10^6. Assuming diploidy (and therefore average fixation time of 4Ne) and generation times of 1-10 years, this gives fixation timescales of 10^3-10^7 years. This is within the timescales of the trees we analyse, which span >150 million years.

      Reviewer #2 (Public review):

      Strengths:

      The paper's underlying question is interesting, extending the authors' prior work on RNA along similar conceptual lines. The paper combines both image analysis of leaves and a computational analysis of a simple model of leaf development.

      Weaknesses:

      The entire paper is based on the Runion model. More intuition about the Runion model would be useful for a broader readership that cares about the evolutionary aspect of this, but may not know the developmental model in question. Obviously, this is prior well-established work, but 2 - 3 sentences highlighting the key structural aspects of such a model would be great. Currently, that intuition is found implicitly in a sentence on page 2 ("complex leaf shapes need more specificity in their GRNs than their simpler unlobed leaf shape"), but the reader is left wondering - is the Runion model a detailed mechanistic one with multiple interacting genes/proteins? If so, how many? Or is it just 2 - 3 genes but with complexity entirely in how long they are each expressed/when they are turned off, etc.

      We thank the referee for the suggestion to make the paper more useful for a broader readership. To that end, we have added a more detailed description of the (Runions, Tsiantis and Prusinkiewicz, 2017) model in supplementary information section S1 "Leaf model details".

      The Runions model has nearly 100 free parameters. Random walks in 100dimensional spaces have generic properties like a tendency to move toward regions of larger volume that have nothing to do with leaf biology. How do you disentangle the geometry of high-dimensional random walks from genuinely biological developmental bias? Would a toy model with 100 parameters and arbitrary phenotype categories also show "bias toward simplicity" if "simple" phenotypes occupy more volume?

      Our argument is largely independent of the number of parameters. While it is true that most of the volume is near the surface in a high-dimensional space, our argument is about the relative volumes of the sets of parameters that map to each of the four phenotypes, an entropic argument if you wish. The basic intuition is that a simple phenotype needs fewer parameters to be fine-tuned, and so a larger volume of parameter space will map to a simpler phenotype.

      The question about a toy-model with arbitrary phenotypes is helpful, because it allows us to clarify that what we are illustrating here with the biologically realistic example of leaf shapes is a much more generic principle. We can say with confidence that if the toy-model generates a many to one set of outputs (phenotypes) through an algorithmic process whose description length does not grow faster than logarithmically with the size of the genotype space, then it should produce a bias towards simplicity regardless of the number of dimensions, see for example Johnston et al. (2022) and Dingle, Camargo and Louis (2018) for a longer discussion of this more general point which is based on arguments from algorithmic information theory (AIT). We don’t use that framing in the current paper because the basic intuition for GRNs that more complex phenotypes need more parameters fine-tuned, and so have relatively smaller volumes, is more straightforward to understand that the more abstract AIT arguments. Our general prediction that this principle should hold more widely for GRNs can be made both by the more formal AIT route, or via the more heuristic fine-tuned parameter route.

      The discussion of Figure 4 (PCA of parameter space) uses "area" loosely when what's actually being measured is bin count in a 2D projection of a highdimensional space. I would think that, in general, PCA projections can be misleading about volume in the full parameter space, but I can't tell if that's an issue in this case. Some comments/thoughts here would be useful.

      The quantitative estimate of phenotype frequencies is computed directly in the full parameter space and does not depend on PCA. Ie. We estimate that the total volume of viable leaves maps to simple unlobed leaves about 80% of the time. However, the volume is extremely high-dimensional, and so hard to visualise. PCA is used solely to provide an interpretable visualization of this otherwise high-dimensional structure. The PCA plots in Fig 4 and Fig S16 are there to be illustrative, not quantitative. Because the volume differences are large, we do not think that the projections of the main PCA components would be misleading on at least the ordering of the sizes of the parameter space components that map to each leaf shape. We provided a similar analysis for other projections -- PC1-PC6 (supplementary information section "PCA occupancy for higher dimensions"), finding the same trend. To make this point clearer, we have now changed the sentence in the Fig. 4 caption slightly “This (reveals that --> illustrates how) unlobed leaves occupy a larger region of model parameter space than more complex shapes and that this larger space also contains the majority of more complex leaves.”

      The classifier validation section is in the Methods section, but it seems critical to the whole story. The < 80% agreement with manual classification could propagate to the rest of the estimates in the paper. Again, some comments/thoughts here would be useful.

      We have repeated the analysis of the agreement between by-eye and automatic morphometric classification. Generating a confusion matrix for the two classification methods shows that the agreement is high for unlobed, dissected and compound, with the main source of disagreement being leaves that were classified as lobed by-eye being classified as either unlobed or dissected by the automatic-morphometric method. The proportion of by-eye lobed leaves classified by the automatic morphometric method as either unlobed (27%) or dissected (23%) is relatively balanced, which we think will help cancel out some error as well. Moreover, we find that the agreement between the automatic-morphometric method and by-eye classification increases to 90.0% when using the categories unlobed and all other categories grouped into one. This is the most important classification for our finding that development and phylogeny are both biased towards unlobed.

      The authors should explain Mut2 and Mut5 in the main paper with a sentence or two, at least schematically, because how you mutate is obviously very relevant to interpreting a paper about biases in variation.

      In the results section we have added a sentence for more detail on the random walk.

      "[We mutated the initial sample using a random walk algorithm with two different mutational schemes, MUT2 (alg. 1) and MUT5 (alg. S2).] These algorithms work by iterating through model parameters one by one and perturbing the value by a small amount. We then [automatically classified the resulting shapes...]"

      Moreover, in methods section C there is already a more detailed description of both algorithms.

      “MUT2 (alg. 1) iterates through the parameters in a random order, and attempts to change the parameter by a value selected at random from an array of numbers randomly generated at 3 different orders of magnitude. MUT5 (alg. S2) is the same as MUT2 except the value each parameter is multiplied by 10% of the range of that parameter within the initial leaves (fig. S1). The aim here was to provide some way of accounting for the biologically relevant sampling range. "

      Moreover, the MUT2 algorithm is described in pseudocode in Algorithm 1 in the main text, and the pseudocode for MUT5 is in supplementary information section S1 C, as algorithm S2.

      The two mutational schemes use additive perturbations to individual parameters. Real mutations presumably affect regulatory networks in more structured ways (e.g., changing binding affinities that affect multiple parameters simultaneously). How sensitive are the results to the assumption of independent single-parameter mutations?

      The referee raises an interesting and well-known issue concerning this widely studied class of GRN models. Without a detailed understanding of how individual genetic mutations map onto model parameters, it is difficult to determine with confidence whether a mutation would produce correlated changes in certain sets of parameters. Our main argument, however, is that the primary source of the observed bias is geometric: the volume of parameter space (or equivalently, the entropy) corresponding to simple leaf morphologies is substantially larger than that corresponding to complex morphologies. As long as mutations explore parameter space approximately symmetrically, even if they involve correlated changes in multiple parameters, larger phenotype regions will tend to be encountered more frequently and retained for longer than smaller regions. We therefore expect the observed bias to be robust to many alternative mutation models, although quantifying this robustness is an interesting direction for future work.

      The connectedness argument is made using a 2D PCA projection. Is there a way to check this statement in the full parameter space or perhaps in higher dimensional projections to test the robustness of this result? Connected components can merge/split under different projections.

      Constructing the nearest neighbour graph for the full dimensional data results in the following no. connected components: unlobed-146, lobed-274, dissected-255, compound-315. This follows the same pattern identified for the PC1-PC2 projection, that unlobed splits into fewer connected components than other leaf shape categories.

      References:

      Dingle, K., Camargo, C.Q. and Louis, A.A. (2018) ‘Input–output maps are strongly biased towards simple outputs’, Nature Communications, 9(1), p. 761. Available at: https://doi.org/10.1038/s41467-018-03101-6.

      Johnston, I.G. et al. (2022) ‘Symmetry and simplicity spontaneously emerge from the algorithmic nature of evolution’, Proceedings of the National Academy of Sciences, 119(11), p. e2113883119. Available at: https://doi.org/10.1073/pnas.2113883119.

      Lynch, M. and Conery, J.S. (2003) ‘The Origins of Genome Complexity’, Science, 302(5649), pp. 1401–1404. Available at: https://doi.org/10.1126/science.1089370.

      Pagel, M. (1994) ‘Detecting correlated evolution on phylogenies: a general method for the comparative analysis of discrete characters’, Proceedings of the Royal Society of London. Series B: Biological Sciences, 255(1342), pp. 37–45. Available at: https://doi.org/10.1098/rspb.1994.0006.

      Runions, A., Tsiantis, M. and Prusinkiewicz, P. (2017) ‘A common developmental program can produce diverse leaf shapes’, New Phytologist, 216(2), pp. 401–418. Available at: https://doi.org/10.1111/nph.14449.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer 1 (Public review):

      Summary:

      This study presents a systematic investigation of parent-of-origin effects on gene expression using trio-based data from the Framingham Heart Study, which is notable for its relatively large number of trios. By combining whole-genome and RNA sequencing data, the authors examined the extent to which gene expression is influenced by whether genetic variants are inherited maternally or paternally.

      The authors report that parent-of-origin eQTLs are widespread, identifying 15,893 eQTLs from 14,733 variants and 1,824 genes that were significant in paternal, maternal, or joint tests but not detected by traditional eQTL approaches. They further classified these associations based on the relative strength and direction of paternal and maternal effects, highlighting a subset with opposing directions. The study also highlighted eGenes linked to known imprinted genes as well as those with opposing parent-specific effects, and observed that paternal eGenes are enriched for drug targets. Finally, the work revisits previous findings in which eQTL studies were used to interpret disease-associated loci, emphasizing that conventional eQTL analyses without testing the parent-of-origin may mislead gene prioritization efforts. The study recommends that future downstream analyses, such as Mendelian randomization, take into account the provided lists of SNPs and eGenes and exclude those with strong parent-of-origin effects when linking genetic regulation to disease risk.

      Strengths:

      The major strength of the study lies in the scale and quality of the dataset, the trio-based design, and the systematic application of statistical tests for parent-of-origin effects. The strengths thoughtfully employed Bayes factors rather than p-values to provide stronger evidence of association, which adds rigor to their analyses. These design choices provide compelling evidence that parent-of-origin effects are widespread and that conventional eQTL analyses miss a substantial fraction of regulatory variation. The results are clearly presented and supported by robust analyses, including the identification of opposing parental effects and the enrichment of paternal eGenes for drug targets. Notably, the two examples demonstrating how these findings can reshape disease gene prioritization highlight the broader impact of the study and encourage further work in the community to incorporate parent-of-origin effects.

      Weaknesses:

      The main limitations of the study are threefold.

      First, there is a lack of replication in independent cohorts, which is understandable given the difficulty of identifying datasets with a comparable number of trios, but replication would help establish the generalizability of the findings.

      We fully agree with the reviewer that replication in an independent cohort is a crucial step for establishing generalizability. As the reviewer notes, the Framingham Heart Study, with its 1,477 trios possessing both WGS and RNA-seq data, represents a uniquely powerful and, to our knowledge, currently unmatched resource for this specific type of parent-of-origin eQTL analysis.

      In the absence of an external cohort of comparable size and data richness, we have taken several steps to ensure the internal validity and robustness of our findings within the current study, which we will clarify and expand upon in the revised manuscript:

      Positive Control Validation: We explicitly used well-established, bona fide imprinted genes (e.g., MEG3, NDN, SNURF, as listed in Table 1 and Figure 1) as positive controls. The fact that our analysis correctly identifies their known parent-of-origin expression patterns (e.g., maternal eQTL for MEG3, paternal eQTL for NDN) serves as a powerful internal validation of our phasing methodology, statistical models, and significance thresholds. This demonstrates that our approach has the power to detect true POE signals.

      Conservative Calling Criteria: As the reviewer suggests, we prioritized specificity. Our definition of eQTL sets (Section 4.6) uses stringent thresholds (e.g., log<sub>10</sub> BF > 4 for primary signals and θ = log<sub>10</sub> 2 for exclusivity). We explored different θ parameters (Supplementary Table S2) and chose the one that minimized the inclusion of false positives, ensuring that our core gene sets (e.g., G<sub>1</sub>,G<sub>0</sub>,G<sub>2</sub>) are high-confidence discoveries.

      Rigorous Analytical Pipeline: As we note in the revised text, our conclusions are supported by a robust analytical pipeline. This includes trio-based phasing validated by simulation (Supplementary Table S1), the use of linear mixed models to control for relatedness and population structure, and the application of Bayes factors which inherently penalize variants with low minor allele frequencies, thereby reducing spurious associations.

      We believe these internal consistency checks and methodological rigor provide strong confidence in our findings. To further facilitate external replication, we will make the full list of POE eQTLs and eGenes available as a comprehensive resource (as noted in the Discussion and Supplementary Materials), enabling other researchers to validate these findings as appropriate datasets become available.

      Second, while Bayes factors are thoughtfully used to assess evidence of association, the paper does not fully explore how the chosen thresholds translate to the expected rate of false positives. For example, a minor allele frequency cutoff of 1% was applied, which seems somewhat arbitrary, and without reporting the allele frequency distribution of the identified eQTLs, it is unclear whether rare variants disproportionately contribute to the signals, potentially affecting the reliability of discoveries.

      We thank the reviewer for raising this important point regarding the calibration of our significance thresholds and the potential role of rare variants. We address this by clarifying the relationship between Bayes factors, prior odds, and false discovery rates, and by providing a more detailed characterization of the variants we identified.

      Bayes Factors and False Discovery: The reviewer is correct that the connection between a Bayes factor threshold and a false positive rate is not direct as it has to take into account of prior odds. As we briefly noted, for a given prior odds of association (e.g., 1 in 100 or 1 in 1000 for a cis-eQTL), a log<sub>10</sub> BF = 4 corresponds to a posterior probability of association (PPA) of 0.99 or 0.90 respectively. Consequently, 1 − PPA can be interpreted as the local false discovery rate (lfdr), as we have now explicitly stated in Section 2.2 (citing Soloff et al., 2024). Our choice of log<sub>10</sub> BF = 4 was therefore chosen to ensure a very low or modest lfdr (depending on the prior odds) for our primary findings.

      Minor Allele Frequency Threshold: The 1% MAF cutoff was indeed a pre-analysis filtering step. It was chosen based on the power afforded by our sample size of 1,477 trios. For variants rarer than 1%, our study is underpowered to detect associations, and any signals would be highly unstable. Importantly, the reviewer’s concern about rare variants disproportionately contributing to signals is further mitigated by our use of Bayes factors. As we note in Section 2.2, the prior used in our Bayes factor computation (with σ = 0.5 in the prior for effect sizes, as described in Section 4.4) inherently penalizes variants with small minor allele frequencies. This is because for a given effect size, the evidence for association is weaker for a rare variant than a common one. Thus, the combination of a pre-analysis MAF filter and the Bayesian analysis itself guards against spurious findings driven by very rare alleles.

      Allele Frequency Distribution: To directly address the reviewer’s request for transparency, in the revised manuscript we include a supplementary figure (e.g., Supplementary Figure S4) showing the distribution of minor allele frequencies (1000 genomes European descents) for the SNPs identified in paternal eQTL set S<sub>P</sub> and maternal eQTL set S<sub>M</sub>. This empirically demonstrate that our findings are not disproportionately driven by low-frequency variants and provide a more complete picture of the genetic architecture underlying these POE signals. We also add a sentence to the Results section (Section 2.5) summarizing this distribution.

      Third, the ancestry background of the study samples is not reported, which could be a confounding factor in the genetic analyses.

      We thank the reviewer for highlighting this omission. In the revised manuscript, we explicitly report the ancestry background of the Framingham Heart Study participants analyzed. Consistent with previous reports on this cohort, the vast majority of samples are of European descent.

      Crucially, as the reviewer suggests, population stratification can be a confounder in genetic studies. To mitigate this, our analysis employed a linear mixed model (Section 4.4) that includes a random effect with a covariance structure defined by the genetic relatedness matrix (GRM). This approach is specifically designed to control for spurious associations due to both subtle population structure and known relatedness among individuals, ensuring that our findings are robust to these potential confounders.

      Reviewer 2 (Public review):

      Summary:

      The authors have used 1477 sequenced trios with available gene expression data in the offspring to discover eQTLs that act in a parent-of-origin specific manner. The classified associated SNPs are tested for enrichment for GWAS hits, drug target genes, etc.

      Strengths:

      The manuscript presents an impressive analysis of a very rich data set of parent-of-origin eQTLs. To my knowledge, it is one of the largest studies of its kind, most analyses are sound, and the results are of interest to many in the field and potentially beyond. The different ideas of follow-up analyses are useful and make sense.

      Weaknesses:

      While in general the analyses are well-conducted, I noticed a major issue with the POE eQTL classification, which puts into question most of the downstream analysis. In light of this problem, most of the analysis would need to be rerun, which represents a major revision of the paper, but is straightforward to repair.

      We appreciate the reviewer’s concern and take it seriously. However, we believe the issue stems from a misunderstanding of our classification framework. We clarify our reasoning below, and we are confident that no re-analysis is necessary. In fact, our Bayesian approach was specifically chosen to avoid the very problem the reviewer raises.

      The major problem with the classification of POEs is that simply having significant maternal, but insignificant paternal effect is not an indicator of POE, this happens widely for SNPs with no POE whatsoever (it can happen by chance even when both maternal and paternal effects are the same and non-zero - the authors can see it via simulations under the null [maternal=paternal effect]).

      The reviewer raises a valid statistical concern: under the null hypothesis of equal maternal and paternal effects (β<sub>0</sub> = β<sub>1</sub>≠ 0), sampling variation could occasionally produce a scenario where one effect appears significant and the other does not. This is indeed a form of Type II error (failing to detect a true non-zero effect for one of the alleles).

      However, this is precisely why we chose Bayes factors over p-values. A key advantage of Bayes factors is that they are not blind to power. P-values are calculated solely under the null hypothesis and do not incorporate any information about the alternative hypothesis or the study’s power to detect it. Consequently, when power is low (e.g., due to minor allele frequency differences between paternal and maternal alleles), p-values can be misleading.

      In contrast, Bayes factors are computed under both the null and alternative hypotheses. They inherently incorporate power through the prior specification. As we note in Section 2.2, “Bayes factors penalize genetic variants with small allele frequencies to reduce false positives.” This means that a SNP where, by chance, one allele appears significant and the other does not—but where power is low due to allele frequency imbalance—will not receive a high Bayes factor, because the evidence is appropriately discounted.

      In order to be able to talk about POE, first, a significant difference between maternal and paternal effects needs to be claimed. Therefore, none of the 4 sets of POE eQTLs are justified. To me, the only relevant criterion to pick POE SNPs is the P-value when comparing the maternal and paternal effects.

      We respectfully disagree with the reviewer’s assertion that our approach to POE eQTL classification are not justified. There are multiple biologically meaningful patterns of parent-of-origin effects, and our classification scheme was designed to capture this diversity:

      (1) Paternal-specific eQTL (β<sub>0</sub> = 0, β<sub>1</sub> ≠ 0)

      (2) Maternal-specific eQTL (β<sub>0</sub> ≠ 0, β<sub>1</sub> = 0)

      (3) Opposing eQTL (β<sub>0</sub> ≠ 0, β<sub>1</sub> ≠ 0,β<sub>0</sub> × β<sub>1</sub> < 0)

      (4) Genotype eQTL (β<sub>0</sub>= β<sub>1</sub> ≠ 0)

      The reviewer’s proposed test (H<sub>0</sub>: β<sub>0</sub> = β<sub>1</sub>) collapses these distinct biological scenarios into a single binary outcome. For example: A purely paternal-specific eQTL (β<sub>0</sub> = 0, β<sub>1</sub> ≠ 0) would indeed show a significant difference, and would be captured by the reviewer’s test. However, a gene like ZNF890P in Table 1, where both effects are significant and in the same direction but of different magnitudes, would also show a significant difference. In the reviewer’s framework, this would be classified as a POE eQTL, yet biologically it behaves more like a genotype eQTL with an allelic imbalance. Our framework correctly separates these cases.

      Moreover, the reviewer’s proposed test is a nested special case of our broader approach. As we note in our response, our paternal-specific test (H<sup>0</sup>: β<sub>0</sub> = β<sub>1</sub> = 0 vs H<sub>1</sub>: β<sub>0</sub> = 0,β<sub>1</sub> ≠ 0) is a more constrained hypothesis that yields a subset of the SNPs that would be identified by the reviewer’s difference test, were it to have sufficient power. Our approach is therefore more conservative for classifying paternal- or maternal-specific eQTLs, not less.

      The definitions of the 4 groups are based on somewhat ad hoc priors, BF thresholds, etc. Also, in Section 4.6, the value of theta is arbitrarily chosen (along with the threshold of 4 to declare POE). In my opinion, the clean treatment of the 4 groups would start with a significant P-value (beta-maternal vs beta-paternal). Within this set, you can then use the original criteria presented in the paper, but only among these associations where there is solid evidence of different parental effects.

      We take strong issue with the characterization of our prior specifications and thresholds as “ad hoc” or “arbitrary.” In Bayesian analysis, prior specification is a principled and transparent modeling choice, not an arbitrary one.

      (1) Choice of log<sub>10</sub> BF = 4 threshold: As stated in Section 2.2, this threshold was chosen based on explicit considerations of prior odds and posterior probability of association. For a prior odds of 1:1000 (a reasonable guess for cis-eQTLs), this BF corresponds to a posterior probability of association of 0.91. If one prefers a more optimistic prior odds of 1:100, the PPA becomes 0.99. The threshold is therefore grounded in decision theory, not whim.

      (2) Choice of θ in Section 4.6: We explicitly state that we explored multiple values of θ(0, log<sub>10</sub> 2, log<sub>10</sub> 3) and chose θ = log<sub>10</sub> 2 because it “produced minimum G<sub>1</sub> and G<sub>0</sub> that contain known imprinted genes.” This is a principled, data-driven calibration step using positive controls, not an arbitrary selection. The transparency of this process is a strength, not a weakness.

      (3) Comparison to p-value thresholds: The reviewer suggests that p-value thresholds are somehow less arbitrary. However, the conventional p-value threshold of 0.05 is itself a historical convention with no universal justification. Moreover, as we note, p-values do not account for power differences across SNPs. A p-value of 5 × 10<sup>−8</sup> from a SNP with 40% MAF is not comparable to the same p-value from a SNP with 1% MAF, because the power to detect the association differs dramatically. Bayes factors automatically adjust for this through the prior, making them more comparable across variants, not less.

      In revision, we added a section in supplementary to review relationships between p-values, Bayes factors, and FDR.

      Recommendations for the authors:

      Reviewer 1 (Recommendations for the authors):

      Here are some suggestions to improve the study:

      (1) Provide information about the ancestry background of participants and consider including ancestry principal components in the eQTL models, as is commonly done, to account for population structure.

      We thank the reviewer for this suggestion. In the revised manuscript, we explicitly state that the participants in the Framingham Heart Study are predominantly of European descent, consistent with previous publications from this cohort. Regarding population structure, we respectfully note that our analysis already employs a linear mixed model (Section 4.4) that includes a random effect with a covariance structure defined by the genetic relatedness matrix (GRM). This approach is widely regarded as more robust than including a limited number of principal components, as it accounts for both fine-scale population stratification and known relatedness simultaneously.

      (2) Conduct sensitivity analyses using different Bayes factor cutoffs to assess the robustness of the findings.

      We appreciate the reviewer’s concern about threshold robustness. In fact, we already conducted a form of sensitivity analysis during the classification step. As described in Section 4.6 and shown in Supplementary Table S2, we explored multiple values of θ (0, log<sub>10</sub> 2, and log<sub>10</sub> 3) and observed how they affected the composition of our gene sets. The choice of log<sub>10</sub> BF = 4 for significance was similarly grounded in posterior probability calculations (Section 2.2). To further address the reviewer’s point, we add a Supplementary Table S3 for counts of eQTL and eGenes under different Bayes factor threshold. This demonstrates that our most significant claim, the abundance of POE eQTL, are not overly sensitive to the specific cutoff.

      (3) In the GWAS examples for KCNQ1 and CDKN1C, the assessment of whether the SNPs act as eQTLs for the two genes is based on a single BF threshold, which may be influenced by differences in gene expression levels. The authors could compare the corresponding effect sizes of these SNPs on both genes to provide a more nuanced investigation. While the limitation of missing data from other tissues is discussed in the paper, it remains possible that KCNQ1 plays a role in tissues more relevant to T2D.

      This is an excellent suggestion for a more nuanced investigation. We re-examined the effect sizes for the SNP rs2237892 in our published results. For gene CDKN1C, the paternal log<sub>10</sub> BF<sub>1</sub> = −0.477 and maternal log<sub>10</sub> BF<sub>0</sub> = 4.94, the normalized maternal effect in joint analysis is −4.86 vs −0.74 for paternal. Unfortunately, the published results has no eQTL for KCNQ1, which according to our selection creteria means maximum log<sub>10</sub> BF < 3 for all tests (genotype, paternal , maternal, joint). The concern for different gene expression level may affect BF is valid. We preempt this pitfall by quantile normalization of gene expression levels after controlling for GC content (as documented in Method Section). We agree with the reviewer that the lack of data from pancreatic tissues is a limitation. We add a sentence in revelant section to acknowledging that while whole blood is a valuable and accessible tissue, replication in T2D-relevant tissues (e.g., pancreas, adipose) would be an important future direction, and our findings provide a hypothesis for such targeted investigations.

      Reviewer 2 (Recommendations for the authors):

      Major comments:

      There are some literature elements missing:

      (1) Hofmeister has a newer and larger study [https://pubmed.ncbi.nlm.nih.gov/40770099/].Please cite that too; it also has POE pQTLs, which is relevant.

      (2) POE in pigs has been explored [https://www.nature.com/articles/s41467-02562243-6], please cite it.

      (3) An insightful review covering the mechanisms of POE for gene expression (https://www.sciencedirect.com/science/article/pii/S2352154618300482) should be cited.

      (4) Further studies on POE in gene expression in social insects (https://royalsocietypublishing.org and in mice (https://www.biorxiv.org/content/10.1101/2023.08.24.554674v1.full) are also relevant.

      We thank the reviewer for bringing these important references to our attention. We incorporated the suggested citations in the revision to provide a more comprehensive context for our work, including the newer POE pQTL study by Hofmeister et al., the findings in pigs, and the mechanistic review.

      While it’s OK to report and rank SNPs by BF, it is necessary to show association P-values as well. It is not explained in the text around the Table how the P-value is obtained in the Table. And it is important to show how their priors translate to FWER control. What is the FWER when picking SNPs at a certain BF value? 1-PPA and local FDR depend on the choice of the prior, but we need a prior-independent measure of FDR/FWER.

      We appreciate the opportunity to clarify. The p-value presented in Table 1 (column “P”) is indeed the frequentist p-value testing the null hypothesis of equal maternal and paternal effects (H<sub>0</sub> : β<sub>0</sub> = β<sub>1</sub>), as described in Section 4.5. We included this to provide a familiar metric for readers, but our discovery framework relies on Bayes factors for the reasons outlined in Section 2.2.

      Regarding error control, the reviewer is correct that 1-PPA is a local FDR that depends on the prior. We chose to control the local rate of false discoveries rather than the Family-Wise Error Rate (FWER) because FWER control (e.g., via Bonferroni) is often excessively conservative for exploratory analyses like eQTL mapping, especially given the correlation among tests due to LD.

      Our Bayesian approach provides a more nuanced measure of evidence at the level of each individual test, which is precisely what is needed for prioritizing SNPs with parent-of-origin effects.

      The demand for a prior-independent measure of FDR is conceptually problematic. Any probabilistic statement about a specific hypothesis being true or false necessarily requires a prior—this is a fundamental consequence of probability theory. Frequentist FDR, while prior-independent in one sense, does not provide a probability that a particular finding is false; it is a long-run error rate over many tests. Methods like q-values, often described as “prior-free,” still depend on implicit assumptions (e.g., the estimate of π<sub>0</sub>, independence of tests, and a mixture of effect sizes).

      In our specific context of cis-eQTL analysis, these assumptions are particularly questionable. LD induces correlation among nearby SNPs, violating the independence required for stable π<sub>0</sub> estimation. Moreover, effect sizes in a region are not randomly mixed—SNPs in high LD tend to have similar effect directions and magnitudes, which can bias the mixture model underlying q-value approaches. Our Bayesian approach, by modeling each SNP individually, avoids these cross-SNP assumptions.

      Importantly, while posterior probabilities depend on the choice of prior (π<sub>0</sub>), we have verified that our conclusions are robust across a wide range of plausible π<sub>0</sub> values (0.9,0.99,0.999). Given our extremely stringent Bayes factor threshold (BF<sub>j</sub> > 10<sup>4</sup>), the posterior probability for a maternal effect exceeds 0.90 for any π<sub>0</sub> < 0.999. Thus, the prior dependence is practically irrelevant for the SNPs we report.

      In revision, we added a section in Supplementary to describe the connections between p-value, Bayes factor, and FDR. We hope this will clarify that a (seemingly) prior independent FDR has a hidden assumption that cis-eQTL analysis is likely to violate.

      The major problem with the classification of POEs is that simply having significant maternal, but insignificant paternal effect is not an indicator of POE, this happens widely for SNPs with no POE whatsoever (it can happen by chance even when both maternal and paternal effects are the same and non-zero - the authors can see it via simulations under the null [maternal=paternal effect]). In order to be able to talk about POE, first, a significant difference between maternal and paternal effects needs to be claimed. Therefore, none of the 4 sets of POE eQTLs are justified. To me, the only relevant criterion to pick POE SNPs is the P-value when comparing the maternal and paternal effects. The definitions of the 4 groups are based on somewhat ad hoc priors, BF thresholds, etc. Also, in Section 4.6, the value of theta is arbitrarily chosen (along with the threshold of 4 to declare POE). In my opinion, the clean treatment of the 4 groups would start with a significant P-value (beta-maternal vs beta-paternal). Within this set, you can then use the original criteria presented in the paper, but only among these associations where there is solid evidence of different parental effects.

      We respectfully disagree with the reviewer’s assertion that a significant difference between maternal and paternal effects is the only valid criterion for defining POE, and we maintain that our classification is statistically sound and biologically meaningful.

      The Problem with the “Difference-Only” Approach: The reviewer’s proposed filter (a significant p-value for β<sub>0</sub> ≠ β<sub>1</sub>) is a single hypothesis test. Our goal was to classify eQTLs into multiple, distinct biological categories (paternal-specific, maternal-specific, opposing, etc.). The “difference-only” test collapses these categories. For example, a purely paternal-specific eQTL (β<sub>0</sub> = 0,β<sub>1</sub> ≠ 0) and a gene like ZNF890P (β<sub>0</sub> ≠ 0, β<sub>1</sub> ≠ 0, β<sub>0</sub> > β<sub>1</sub>) would both show a significant difference. In the reviewer’s framework, they would be lumped together, obscuring the fact that one is an imprinted gene and the other is a standard eQTL with allelic imbalance. Our framework correctly separates them.

      Bayes Factors are Not “Ad Hoc”: The choice of prior (σ = 0.5) follows established literature for linear model Bayes factors (Servin and Stephens, 2007). The threshold of log<sub>10</sub> BF = 4 was chosen based on its relationship to posterior probability (0.91-0.99 given reasonable prior odds), which is a transparent and principled decision rule. The selection of θ in Section 4.6 was calibrated using a positive control set of known imprinted genes, ensuring our definitions were conservative and accurate. This is the opposite of arbitrary.

      The Suggested Procedure Has Low Power: One can run the following simple R code to verify. We simulate maternal alleles xx and maternal alleles yy, then simulate phenotype with β<sub>xx</sub> > 0 and β<sub>yy</sub> = 0 (maternal effect only). We fit the joint model and compute p-values for the null β<sub>xx</sub> = β<sub>yy</sub> as suggested by reviewer. From the joint fit, we also extract p-values based on the null β<sub>xx</sub> = 0 and β<sub>yy</sub> = 0 respectively. The simulation was repeated 1000 times and p-values were stored in a matrix.

      We call positives based on suggested procedure, and compare number of positives called using marginal p-values at two threshold of 1×10<sup>−5</sup> and 1×10<sup>−6</sup> to declare significance. We used threshold of 0.01 to declare insignificance.

      The result demonstrates that the suggested procedure has a much lower power compared to the procedure based on marginal statistics.

      For the above reasons, the follow-up enrichment analysis is somewhat questionable. Most enrichments are non-significant, and it is likely because the SP and SM groups are diluted with SG SNPs. The P1-P9 groups have nothing to do with POE, and although the observation of increased enrichment for GWAS SNPs with increased pleiotropy is interesting, it is irrelevant for POE.

      We will address the dilution concern below. We agree that P1-P9 groups are not directly related to POE. But this is an interesting observation non-theless. As we found such an observation is missing in the literature, we ask to keep it in the paper.

      In the same way, section 2.7 is not supported; the claimed maternal and paternal POEs are heavily diluted by simple marginal associations. The same holds for sections 2.82.10. A striking example is Table 3: for clinical trial targets, paternal/maternal eQTLs behave just like simple marginal eQTLs (G<sub>G</sub>). A similar pattern emerges for combined target enrichment.

      The reviewer’s concern that our S<sub>P</sub> and S<sub>M</sub> sets are “diluted with S<sub>G</sub> SNPs” is precisely the issue our Bayes factor thresholds were designed to prevent. By requiring one effect to be significant and the other to be below a low threshold (θ), we explicitly excluded SNPs where both effects are significant and in the same direction (which defines S<sub>G</sub>).

      Regarding Table 3, the reviewer’s interpretation differs from ours. The fact that paternal eQTLs (G</sub>P</sub>) show significant enrichment for drug targets, while genotype eQTLs (G<sub>G</sub>) also show enrichment, does not imply dilution. Rather, it suggests there is an overlap in the biological importance of these gene sets, which is expected. The key message of the finding is the asymmetry: G<sub>P</sub> is significantly more enriched than G<sub>G</sub> (p=0.035 for combined targets), a pattern that would be washed out if G<sub>P</sub> were merely a diluted version of G<sub>G</sub>. This asymmetry supports the interesting biological hypothesis (Moore and Haig, 1991) we discuss. The non-significance for G<sub>M</sub> further highlights this asymmetry.

      I’m not sure how MR would be biased by POE: MR is conducted only if there is a marginal association, i.e., the average maternal and paternal effects are significant. If the expression is causal for a trait, the POE effect is propagated to the outcome; hence, the SNP effect on the exposure will be equally biased as the SNP effect on the outcome, and these cancel out, and the causal effect remains unbiased. Can the authors propose a concrete example of maternal/paternal effects that demonstrates their claimed bias?

      We thank the reviewer for this insightful question, which allows us to clarify our point with a concrete example from our data.

      Consider a scenario where one wishes to use Mendelian Randomization (MR) to test whether the expression of gene NECAB3 causally influences a particular trait (e.g., obesity). The reviewer is correct that if the causal effect is homogeneous, the average effect might still be captured. However, the bias we caution against arises in stratified analyses or in the interpretation of the genetic instrument itself.

      Take the SNP rs4911348 and its effect on NECAB3 (Figure 2). The genotype model shows no marginal association. Therefore, if a researcher were conducting a standard MR study using this SNP as an instrument for NECAB3 expression, they would discard it as an invalid instrument due to the lack of a marginal association. They would miss the true underlying biology entirely. The causal effect of NECAB3 on the trait would be masked in the full population.

      More subtly, even if a SNP has a marginal association, using it as an instrument while ignoring POE can lead to incorrect effect estimates in population subgroups defined by parent of origin. This is analogous to ignoring effect modification. For instance, if a treatment (exposure) has a different effect depending on which parent it came from (which is impossible, but the genetic propensity for the exposure does), failing to account for this can bias the instrumental variable estimate if the instrument’s strength varies by an unmeasured factor (parental origin).

      Our advice to “check the list of POE SNPs” is a practical caution: if the instrument for an exposure exhibits strong POE, the standard MR assumptions about the homogeneity of the instrument’s effect may be violated, potentially leading to biased estimates or incorrect conclusions about causality.

      Minor comments:

      (1) In Table 1, the last column header should be -log10(P), not ”P”.

      The column labelling is an editorial choice to prevent table overflow. This particularly labelling was explained in the caption.

      (2) While BFg/0/1/j are explained in the text, these notations should be explained in the Table caption as well.

      Added explanation in caption.

      (3) It should also be mentioned in the Table 1 caption how these top 10 SNPs were chosen.

      These are sentinel eQTL for each gene. We think the first paragraph of Section 2.3 explains clearly.

      (4) “may ”acquires” a cis-eQTL through” → ”may ”acquire” a cis-eQTL through”.

      Corrected. Thank you.

      (5) “which retained 16, 969 genes out of total 58103”, I assume the 58103 are transcripts, not genes.

      You are absolutely correct. We added transcripts after 58103.

      (6) In Equation (1), Z is not defined. In this concrete setting, isn’t it simply the identity matrix?

      Yes. Z is the identitity (loading) matrix for human study. We added a sentence to clarify in revision.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (on non-trivial pattern transformations):

      (3) All modelling is confined to one spatial dimension, and the very definition of a "non-trivial" transformation is framed in terms of peak positions along a line, which clearly must be reformulated for higher dimensions. It's well-known that diffusions in 1, 2, and 3 dimensions are also dramatically different, so the relevance of the three-class taxonomy to real multicellular tissues remains unclear, or at least should be explained in more detail.

      Reviewer #2 (on non-trivial pattern transformations):

      (5) The definition of non-trivial pattern formation is provided only in the Supplementary Information, despite its central importance for interpreting the main results. It would significantly improve clarity if this definition were included and explained in the main text. Additionally, it remains unclear how the definition is consistently applied across the different initial conditions. In particular, the authors should clarify how slopebased measures are determined for both the random noise and sharp peak/step function initial states. Furthermore, the authors do not specify how the sign function is evaluated at zero. If the standard mathematical definition sgn(0)=0 is used, then even a simple widening of a peak could fulfill the criterion for non-trivial pattern transformation.

      There was indeed a problem on how we defined non-trivial pattern transformations in the original version. This definition was not clear enough beyond 1D. We now provide a simple clear definition in the main text that applies to all dimensions (“P1” and “P2” in the second page of the introduction).

      As we now explain through the main text, even if the solution of the heat/diffusion equation depends on the dimension of the system, our classification of gene networks (and the mathematical analyses we use) does not depend on the dimensionality of the system. However, some aspects of the specific pattern transformations possible from these networks depend on the dimensionality of the system. In the current version of the article, every time we explain something about the resulting patterns in 1D, we also explain it for the resulting patterns in 2D and 3D. We also have added figures for the 2D cases (in current Fig.1 and Fig.9). We now explicitly explain how the possible resulting patterns in space can depend on the boundaries and shapes of the system (i.e. the distribution of cells in space) (see specially the 5th paragraph of the discussion).

      The criticisms about “slope-based measures” mentioned by reviewer 2, is now addressed in a paragraph at the end of the introduction (here we added it):

      “It is worth noting that these three basic initial patterns correspond to spatially discontinuous functions: in homogeneous with noise initial patterns, white noise is discontinuous by definition; in spike and combined spike-homogeneous initial patterns, there is a concentration discontinuity between cells on the edge of the spike and nearby cells outside the spike. However, once extracellular signal diffusion begins, these sharp boundaries are smoothed into differentiable gradients, where critical points can be properly defined (e.g., at the center of the initial spike).”

      The main concern among these relates to the validity of our linearization of the model equations and the extension of the results obtained for the linear system to the fully nonlinear system. In this regard, the reviewers’ comments are:

      Reviewer #1 (on linearization):

      (2) A central step in the model formulation is the linearisation of the reaction term around a homogeneous steady state; higher-order kinetics, including ubiquitous bimolecular sinks such as A + B → AB, are simply collapsed into the Jacobian without any stated amplitude bound on the perturbations. Because the manuscript never analyses how far this assumption can be relaxed, the robustness of the three-class taxonomy under realistic nonlinear reactions or large spike amplitudes remains uncertain.

      Reviewer #2 (on linearization):

      (2) Most of the proofs presented in the Supplementary Information rely on linearized versions of the governing equations, and it remains unclear how these results extend to the fully nonlinear system. We are concerned that the generality of the conclusions drawn from the linear analysis may be overstated in the main text. For example, in Section S3, the authors introduce the concept of dynamic equivalence of transitive chains (Proposition S3.1) and intracellular transitive M-branching (Proposition S3.2), which pertains to the system's steady-state behavior. However, the proof is based solely on the linearized equations, without additional justification for why the result should hold in the presence of nonlinearities. Moreover, the linearized system is used to analyze the response to a "spike initial pattern of arbitrary height C" (SI Chapter S5.1), yet it is not clear how conclusions derived from the linear regime can be valid for large perturbations, where nonlinear effects are expected to play a significant role. We encourage the authors to clarify the assumptions under which the linearized analysis remains valid and to discuss the potential limitations of applying these results to the nonlinear regime.

      We used three linearizations in the original version of the manuscript. One was to analyze hierarchic networks (in the Hierarchic networks section). In the new version of the article we do not use any linearization to study the hierarchic networks, so this problem is solved.

      The second linearization was in section S3 on transitive chains. We realized that this section is not really necessary at all for the article so we deleted it.

      We keep the third linearization but we now explain why such linearization is useful and valid in a section called “Linear stability analysis”. Thus, through this section we justify this choice (explicitly in its two first paragraphs).

      Regarding Reviewer 2 concerns about large perturbations, we acknowledge that the phrasing using “arbitrary height” may have been confusing. As we now explain in the linear stability analysis section, linear stability analysis assumes perturbations to be small.

      For the homogeneous-with-noise initial pattern, as we explain, these perturbations are assumed to be small because they are actually molecular noise.

      For the spike initial pattern and hierarchic networks the perturbation is not necessarily small. However, by the definition of the spike and combined homogeneous-spike initial patterns, all cells outside the spike start with the same concentration of the extracellular signals that are secreted from the spike (e.g. zero). Thus, even in the case in which extracellular signals concentrations in the spike would be unrealistically high, the amount of extracellular signal diffusing from it can be considered small by simply considering it at a small enough time interval. Thus, right outside the spike the diffusion of extracellular signals from the spike can be treated as a continuous small perturbation for which one can study the stability, as we do in the “Linear stability analysis section”. This we now explain at the end of the introduction and in the “Linear stability analysis” section when we talk about the initial patterns again.

      In the following, we respond to the remaining concerns raised by the reviewers:

      Reviewer #1 (Public review):

      (1) The Results section is difficult to follow. Key logical steps and network configurations are described shortly in prose, which constantly require the reader to address either SI or other parts of the text (see numerous links on the requirements R1-R5 listed at the beginning of the paper) to gain minimal understanding. As a result, a scientifically literate but non-specialist reader may struggle to grasp the argument with a reasonable time invested.

      We acknowledge that the original version of the main text may not be as clear as we intended. Initially, we believed that placing the more technical mathematical passages in the Supplementary Information would make the main text more accessible to readers. We were wrong. We have now moved crucial parts of the supplementary to the main text and adapted the rest of the text accordingly. The most important of those is the new “Linear stability analysis” section and the associated dispersion relation (e.g. Fig.6).

      Reviewer #2 (Public review):

      (1) We have serious concerns regarding the validity of the simulation results presented in the manuscript. Rather than simulating the full nonlinear system described by Equation (1), the authors base their results on a truncated expansion (Equation S.8.2) that captures only the time evolution of small deviations around a spatially homogeneous steady state. However, it remains unclear how this reduced system is derived from the full equations -specifically, which terms are retained or neglected and why- and how the expansion of the nonlinear function can be steady-state independent, as claimed. Additionally, in simulations involving the spike plus homogeneous initial condition, it is not evident -or, where equations are provided, it is not correct- that the assumed global homogeneous background actually corresponds to a steady state of the full dynamics. We elaborate on these concerns in the following:

      We are actually simulating the full nonlinear system described by Equation (1). In the current version we are more explicit about this. As we describe in the introduction and, now, through all the text several times (e.g. in the last paragraph of the model section and in the paragraph before the linear stability section), the aim of the article is to describe necessary requirements for non-trivial pattern transformations. We did not intent to describe all necessary requirements nor sufficient requirements. These requirements are at the level of gene network topology not at the level of f or its parameters. In other words, we just claim that gene networks having specific topological features can lead to some specific types of non-trivial pattern transformations but not to others. We do not say for which specific fs (or its parameters) these pattern transformations are possible, we just say that this can happen for some f, as long as these fulfill our requirements. We do show, however, that without some specific topological requirements there are non-trivial pattern transformations that are not possible, no matter the f (this explicitly stated in the last paragraph of the model section and in the paragraph before the linear stability section). Thus, all the simulations shown in the figures are just examples, with specific fs, of the types of non-trivial pattern transformations possible from each type of gene network topology.

      In all simulations we used the f of the Maini-Miura model. We could have chosen other ones but we happen to chose that f. The presentation of the Maini-Miura model has been revised to improve clarity (equation S6.1 in SI). This model we are simulating fully, we are not doing any linearization for the simulations. That may not have been explained clearly enough in the previous version of the article. We just happen to make a change of variable that may have been confused as a linearization. In the current version, the existence of a homogeneous steady state is parameterized by a tunable g<sup>*</sup>, that can be chosen as for spike initial patterns or g for noise-homogeneous and spike-homogeneous initial patterns. We have also included a proof that the model equations satisfy our conditions R1-5. Indeed, the model is non-linear as long as σ<sub>i</sub>≠0 for some gene product (as we explicitly assume).

      It is assumed that the homogeneous steady states are given by g_i=0 and g_i=c_i, where 1/c_i = \mu_i or \hat{\mu}_i, independently of the specific network structure. However, the basis for this assumption is unclear, especially since some of the functions do not satisfy this condition -for example, f5 as defined below Eq. S8.10.5. Moreover, if g_i=c_i does not correspond to a true steady state, then the time evolution of deviations from this state is not correctly described by Eq. S8.2, as the zeroth-order terms do not vanish in that case.

      In the revised manuscript, homogeneous steady states are parameterized by a tunable g<sup>*</sup>, which can be chosen as for spike initial patterns or g for noise-homogeneous and spike-homogeneous initial pattern. Function f(g) in (S6.1), as well as the specific non-linear entries used in certain simulations, are constructed such that g<sup>*</sup> is indeed a steady state of the system and that conditions R1-R5 are satisfied. We have also corrected some typos in section S6 (previously section S8) of the Supplementary Information, that we believe may have induced the confusion indicated by this reviewer.

      Additionally, the equations used contain only linear terms and a cubic degradation term for each species g_i, while neglecting all quadratic terms and cubic terms involving cross-species interactions (i≠j). An explanation for this selective truncation is not provided, and without knowledge of the full equation (f), it is impossible to assess whether this expansion is mathematically justified. If, as suggested in the Supplementary Information, the linear and cubic terms are derived from f, then at the very least, the Jacobian matrix should depend on the background steady-state concentration. However, the equations for the small deviation around a steady state (including the Jacobian matrix) used in the simulations appear to be independent of the particular steady state concentration.

      As described above we just chose an example f to exemplify the non-trivial pattern transformations possible from each class of gene network topologies. There is no special reason to include, or exclude for that matter, cubic cross-species interactions since the point is just to exemplify the types of possible pattern transformations from each type of gene network topology.

      In addition, we believe that part of the reviewer’s concern may have arisen from a notational ambiguity in the previous version of the manuscript, which has now been corrected: the matrix appearing in f(g) has been renamed from J to W<sup>T</sup>. As stated in the main text, the jacobian of the regulation function f(g) evaluated at the homogeneous steady state must coincide with the transpose of the network weight matrix. With the current equations (S6.1), we have , from which we easily get . Also, it is clear that the Jacobian of f(g) is not independent of g.

      This is why we believe that the differences observed between the spike-only initial condition and the spike superimposed on a homogeneous background are not due to the initial conditions themselves, but rather result from a modified reaction scheme introduced through a questionable cutoff.

      "In simulations with spike initial patterns, the reference value g≡0 represents an actual concentration of 0 and therefore, we must add to (S8.2) a Heaviside function Φ acting of f (i.e., Φ(f(g))=f(g) if f(g)>0 , Φ(f(g))=0 if f(g){less than or equal to}0) to prevent the existence of negative concentrations for any gene product (i.e., g_i<0 for some i)." (SI chapter S8).

      This cutoff alters the dynamics (no inhibition) and introduces a different reaction scheme between the two simulations. The need for this correction may itself reflect either a problem in the original equations (which should fulfill the necessary conditions and prevent negative concentrations (R4 in main text)) or the inappropriateness of using an expanded approximation which assumes independence on the steady state concentration. It is already questionable if the linearized equations with a cubic degradation term are valid for the spike initial conditions (with different background concentration values), as the amplitude of this perturbation seems rather large.

      The Heaviside function does not preclude inhibition, it precludes gene product concentration to be negative. In the current version of the article we do not use the Heaviside function but another similar, but continuous, function. Having this function can indeed affect the dynamics but: 1) does not violate our requirements on f 2) Does not affect which non-trivial pattern transformations are possible from which gene network topology. Without this function non-trivial pattern transformations are still possible from the spike initial pattern through hierarchical networks, in the way we describe in the article. The Heaviside function (and the one we now use) simply allows that to happen more easily, i.e. for a larger range of parameter values. With this function large inhibitions do not lead to negative gene products concentrations while without it, this can happen for some parameter combinations. None of the arguments nor proves in our article requires the Heaviside, or any similar function. Again this is simply because our aim is to identify topological requirements that are necessary, but not sufficient, for non-trivial pattern transformation. So an f that leads to negative gene products concentrations for some parameter combinations but to non-trivial pattern transformations for others, is still valid example of our points (although not the most interesting or realistic example f).

      We distinguish between the spike and combined spike-homogeneous initial patterns simply because they are biologically quite different, i.e. in the former the gene product in the spike is only expressed in the spike and nowhere else. As we describe in the current version the pattern transformations possible from these two different initial patterns are very similar. In the same way, which gene network topologies can lead to which types of non-trivial pattern transformations is not affected by using the Heaviside functions or not (although this can affect the range of parameter values in which this happens).

      Lastly, we note that under the current simulation scheme, it is not possible to meaningfully assess criteria RH2a and RH2b, as they rely on nonlinear interactions that are absent from the implemented dynamics.

      The implementation of nonlinear entries in f(g) whenever they are needed is now made explicit in the corresponding subsection in the main text and in section S6 in the Supplementary Information. This entries also satisfy conditions R1-R5 around the steady state given by g<sup>*</sup>. Again we should insist that the simulated fs are nonlinear (as now explicitly explained in the SI).

      (3) Several statements in the main text are presented without accompanying proof or sufficient explanation, which makes it difficult to assess their validity. In some cases, the lack of justification raises serious doubts about whether the claims are generally true. Examples are:

      "For the purpose of clarity we will explain our results as if these cells have a simple arrangement in space (e.g., a 1D line or a 2D square lattice) but, as we will discuss, our results shall apply with the same logic to any distribution of cells in space." (Main text l.145-l.148).

      The result of which gene network topologies can lead to pattern transformations are based on a linear stability analysis and some logical arguments. As we now explain through the text none of them depends on the number of dimensions nor on the shape of the arrangement of cells. The geometry of the domain can influence the specific form of the resulting patterns, but it does not alter the broader type of resulting patterns (e.g., periodic patterns, peaks emerging around a spike, etc.) that a given gene network topology can produce. We now explicitly discuss these dependencies in the 5th paragraph of the discussion.

      "For any non-trivial pattern transformation (as long as it is symmetric around the initial spike), there exists an H gene network capable of producing it from a spike initial pattern." (Main text l.366f).

      We now provide a more detailed justification of this statement and the limits of its applicability. This is now in section: “The ensemble of possible pattern transformations from spike initial patterns in H networks“. To make this section easier to understand, however, we have also done changes through all the hierarchic networks sections.

      "In 2D there are no peaks but concentric rings of high gene product concentration centered around the spike, while in 3D there are concentric spherical shells." (Main text l. 447ff).

      This result pertains specifically to pattern transformations arising from spike initial patterns. As defined in the text, spike initial patterns are radially symmetric (at least far away from the boundary). Since diffusion preserves radial symmetry, pattern transformations from spike initial patterns in two or three dimensions reduce to effectively one-dimensional transformations along each radial direction. In this framework, each pair of concentration peaks symmetric with respect to the spike in one dimension corresponds to a ridge surrounding the spike in two dimensions, and each ridge in two dimensions becomes a spherical ridge shell around the spike in three dimensions. In the current version we explain what happens in 1D but also, in the same places, what happens in 2D and 3D (and we have added figures to visualize this in 2D, e.g. Fig.1 and Fig.9)).

      (4) The study identifies one-signal networks and examines how combinations of these structures can give rise to minimal pattern-forming subnetworks. However, the analysis of the combinations of these minimal pattern-forming subnetworks remains relatively brief, and the manuscript does not explore how the results might change if the subnetworks were combined in upstream and downstream configurations. In our view, it is not evident that all possible gene regulatory networks can be fully characterized by these categories, nor that the resulting patterns can be reliably predicted. Rather, the approach appears more suited to identifying which known subnetworks are present within a larger network, without necessarily capturing the full dynamics of more complex configurations.

      We acknowledge that our explanation regarding the combination of sub-networks may have been too brief. We now provide a more detailed description in the section “Gene networks combining different classes of subnetworks” and in its sub-sections. There we explore the different ways in which signal subnetworks can be combined (upstream, downstream, in series, in parallel, etc.). However, this section cannot be understood (and that may have been the problem in the original version of the manuscript) without the linear stability analysis section that is now in the main text, and the associated discussion on the dispersion relation and results related to it. These are important because they apply to all gene networks and, thus, constrain the possible gene network topologies and the types of possible pattern transformations. In other words, whichever ways gene networks are combined, they will always be RD-stable (i.e. no pattern transformation) or RD-unstable of the first (periodic resulting patterns) or second kind (other patterns we discuss). In the current version, we combine this fact with other arguments to describe the types of pattern transformations possible by gene networks combining the different classes of subnetworks.

      (6) The manuscript lacks a clear and detailed explanation of the underlying model and its assumptions. In particular, it is not well-defined what constitutes a "cell" in the context of the model, nor is it justified why spatial features of cells -such as their size or boundaries- can be neglected. Furthermore, the concept of the extracellular space in the one-dimensional model remains ambiguous, making it unclear which gene products are assumed to diffuse.

      We now clarify all these points in the first three paragraphs of the “Methods: the Model” section. We have also included a figure for that clarification (Fig.3).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I suggest the following changes for each weakness I mentioned in the Public Review:

      (1) Presentation

      (R1.1) (a) Add a one-page "Key Requirements" table (e.g., immediately after the Model section) that lists every requirement code (R1-R5, I1-I2, RH1-RH2, etc.), its one-line statement, and the SI section where it is proved.

      In the new version of the article each requirement has its own paragraph starting with the requirement label, e.g. R1 (in bold): ….. We introduce each requirement there where they are justified or proven, otherwise the reader may not know where do they come from. We have also hyperlinked all requirements and most equations so that the reader can easily go back to the explanation of each requirement and equation.

      (R.1.2) Provide more figures illustrating the general structure of networks when you describe them; the network sketches could be folded into a single summary figure, so the reader sees all motifs at once. For example, in lines 304-311, it took me a while to understand if the requirement means just A -> k - ... ⊣ j, or it additionally requires A->...->j (through another pathway). It seems that the full requirement is A → k ⊣ j together with an independent positive route A → j. A figure describing the network structure, or at least a schematic "inline" plot in the spirit of what I just wrote, could help. This is just one example, but the text consists of a constant flow of such "diagrams encrypted in prose".

      We have followed the reviewer’s suggestions. Not all fit in a single figure so we have constructed new figures 4 and 5 for that purpose.

      (R.1.3) (b) Also consider supporting the main text with some key formulas and arguments from SI. My overall suggestion here is that it would be great to make the main text less prosaic and more self-consistent, if the journal requirements allow it.

      After the suggestions by both reviewers, and for the sake of clarity, we have actually moved (and clarified) several key parts of the SI into the main text. These include the whole “Linear stability analysis” and “Positive regulatory loops determine the kind of RD-instability” sections. These parts, although quite mathematical, facilitate the understanding of our results.

      (2) Linearisation

      (R.1.5) It's clear that keeping non-linearity is complicated and maybe redundant, but please, discuss the assumption of linearity explicitly, especially in the scope of relevance for the real systems, and explain why it's not important, if so. I guess that relaxing this assumption may affect the argumentation in many places, for example, equation (3) of the main text could break (i.e., if the signaling molecule can be consumed in some reaction of A+B->AB kind).

      We agree that the original version was not explicit enough about the reasons for the linear approximation. The first and last paragraphs of the section “Linear stability analysis” are explicitly devoted to justify this linearization. Moreover, the hierarchical network section is now written without using the linearization.

      We are not sure we understand which is the problem with the A+B→AB reaction. We are not assuming any specific f function, just the ensemble of functions that fulfill our requirements (R1 to R5). It is only for the simulations that we have to use a specific f. The reactions suggested by the reviewer could represent an f of the form d[AB]/dt=fAB([A]*[B])-m*[AB]**n for AB and d[A]/dt=-fAB([AB]) and d[B]/dt=-fAB([AB]), where fA and fB are functions that decrease with their arguments. We see no reason why there cannot be a fAB that fulfills our requirements. For example fAB=[A]*[B]/(K+[A]*[B])-m*[AB]. See also related comments in the public comments file.

      (R.1.6) Please, provide a separate section where you reformulate the definition of "non-trivial pattern transformation" for two- and three-dimensional domains, and summarize in this section why the analysis provided for 1D is relevant for higher-dimensional systems. By now, I'm not convinced.

      There was indeed a problem with the way we described non-triviality beyond 1D in the original version of the article. We have now refined the definition of pattern transformations so that it is understandable in 2D and 3D. This definition is presented in the introduction already (in P1 and P2). We have modified figure 1 accordingly.

      Reviewer #2 (Recommendations for the authors):

      Major Issues

      (1) Mathematical Proofs

      (R2.1) We strongly recommend that the authors revisit the mathematical derivations or provide a clear and rigorous justification for the assumptions made therein. These assumptions currently appear unjustified or overly simplistic, especially in light of the nonlinear dynamics the authors aim to describe. The authors should comment on why they expect their results to generalize to all complex network structures, as claimed, and not only apply to the simplified examples analyzed in the paper.

      The article has now been restructured to that end. Concerning the assumptions, they are now all explicitly described in the “Methods: the model” section. Concerning the derivations they are through all the results section. A major change in this line has been the moving of part of the supplementary into specific sections in the main text (and the consequent adaptation of the rest of the text). There are important points of the derivation that may have been buried into the old supplementary and that are crucial to understand the whole argument in the article. In fact, a large part of the results section is just a long argument to show that there are essentially only three classes of gene network topologies that can lead to non-trivial pattern transformations. These arguments are summed up in the last paragraph of the new section “Positive regulatory loops determine the kind of RD-instability” and in the first paragraph of the discussion. In brief:

      (1) Pattern transformation requires gene networks with extracellular signals

      (2) Applying previous mathematical results we show (given the broad requirements on f we have) that pattern transformation is only possible in gene networks that contain positive regulatory loops.

      (3) Applying previous mathematical results we show that in the gene networks in which these loops are extracellular, the only possible non-trivial pattern transformations lead to periodic resulting patterns.

      (4) Applying previous mathematical results we show that in the gene networks in which these loops are INTRAcellular, the only possible non-trivial pattern transformations do not necessarily lead to periodic resulting patterns.

      (5) Using simple logical arguments we also show that no non-trivial pattern transformations are possible in gene networks without negative interactions.

      (6) All the above points combined shows that there are only three classes of gene networks capable of nontrivial pattern transformations. 1) Those with intracellular positive loops, extracellular signals that do not affect themselves and some negative regulation by those (that we call hierarchic networks) 2) Those with intracellular positive loops and extracellular signals that affect themselves negatively (that we now call over-Turing networks) 3) Those with extracellular positive loops and an extracellular negative loops (that following previous work by others are called Turing networks).

      (7) Following previous research and different developmental arguments we explore the types of patterns transformations each of these three classes of gene networks can lead to. These types are characterized only in broad and potential terms. We say nothing about the parameters values for which any gene network leads to any specific pattern transformation. What we say is which types of pattern transformation may be possible (for some possible parameter combination) and which ones are not possible from gene network topology alone (based on the types of loops and so on).

      (R.2.3) Additional to the examples provided in the Public Review, claims such as "despite the large amount of theoretically possible gene network topologies, all gene network topologies necessary for pattern formation fall into just three fundamental classes and their combinations" (l. 34ff)

      This statement was originally intended as an introduction of the text following after it but it seems now clear that this was not apparent enough. This statement has been deleted but we convey a similar message letter in the text, now once its justification is provided. In fact, the justification for this statement is the summary we just described in the previous point (R.2.2) and it is discussed over the main text and summarized in the last paragraph of section “Positive regulatory loops determine the kind of RD instability”.

      (R.2.4) and "The same applies to the topologies we found not to be able to lead to non-trivial pattern transformation" (S7) are not or inadequately justified and should be either substantiated or significantly toned down.

      The same comments that above apply.

      (R.2.5) (a) We advise the authors to argue why it is enough to prove key results by considering linear dynamics (see S2-S7). While linearization is a common technique, the authors themselves emphasize the importance of nonlinearities in pattern formation throughout the paper.

      In the current version we provide an explicit justification for this in the section “Linear stability analysis”, especially in its first paragraph. Moreover, for the analysis of the hierarchical networks we do longer use any linearization.

      (R.2.6) (b) To make linear analysis meaningful, we suggest restricting the initial conditions to small fluctuations (e.g., small spikes or noise), which would justify using linearization to investigate the onset of non-trivial pattern formation. Alternatively, the authors should attempt to generalize the results to fully nonlinear dynamics, ideally for a broader class of functions f.

      As we now explain, the homogeneous-with-noise initial pattern already correspond to small perturbations around the homogeneous steady state (due to molecular noise). In addition, for the spike and spike–homogeneous initial pattern we now explicitly consider spikes of small amplitude. We acknowledge that the use of larger spikes in the previous version could lead to misunderstandings regarding the validity of the linear approximation, even though it does not contradict the assumptions underlying the analysis. In these initial patterns, pattern formation arises because the signal secreted from the spike diffuses into the surrounding domain, so that cells outside the spike experience only small deviations from the equilibrium concentration.

      Larger spikes may induce stronger deviations in cells located very close to the spike; however, because the spike occupies a region that is very small relative to the total domain size, these local effects do not influence pattern formation in the bulk of the domain. A similar situation occurs with boundary effects in cells located near the domain limits, which likewise do not affect the pattern formation process away from the boundaries. We have clarified this point in the revised manuscript, both in the final sentences of the Introduction and in the description of the initial conditions in the fourth paragraph of the “Linear stability analysis” section, where we explicitly state that each initial pattern can be interpreted as a perturbation of an otherwise homogeneous pattern.

      (R.2.7) (c) The assumptions required for the proofs should be explicitly stated and justified. At present, the logic behind the chosen constraints on f is unclear, and the flow of the argument suffers as a result.

      The actual justification for the requirements (i.e. constraints) on f are biological (and we now explain them more explicitly when we introduce these requirements). Most of the mathematical proofs do not require these requirements except when we explicitly say so.

      (R.2.8) (d) The illustrative functions provided in some of the proofs in the SI (e.g. S5.2.1 "To see this, let us consider, for example, that they are both quadratic monomials of the form f_k(g_A)=B_k g_A^2 and f_j(g_A)=B_j g_A^2") do not satisfy the authors' own stated conditions (e.g., this function violates requirement R4 (l.197 f)). More suitable examples should be selected to ensure consistency between assumptions and illustrations.

      We have changed the whole section (based on the comment R.2.9 from the same reviewer). We now provide arguments in the main text that generally do not rely on specific fs.

      (R.2.9) (e) Currently, all mathematical results are confined to the appendix. We recommend including key insights from the proofs in the main text to improve readability and to allow the main claims to stand on their own. For example, the section on the requirements RH2a and RH2b (l. 320 - l. 335)) would benefit strongly from the insights from S5.2.1

      We agree. We have moved the linear stability analysis and the dispersion relation section to the main text. We have also moved what used to be S5.2.1.

      (2) Simulations

      The simulations raise, as mentioned in the Public Review, several concerns regarding their generality and validity.

      (R.2.10) (a) We recommend validating the simulation results by comparing them with simulations of the full nonlinear equations. The authors should at least provide the equations for the full dynamics and explain how the expansion is performed and why it is valid. This also includes verifying the assumed steady states (g_i=0 and g_i=c_i, where 1/c_i = \mu_i or \hat{\mu}_i).

      We are simulating the whole non-linear equations. Here it is important to stress, as we do now in the main text, that our results apply to any f, as long as it fulfills our R1-R5 requirements. However, for the simulations in the figures we have to use a specific f (since there is an infinite amount of fs that fulfill our requirements). Again the figures are just examples to visualize the types of resulting patterns and gene networks we talk about.

      In the original version we may not have been clear enough about the equations used for the simulations. The presentation of the Maini-Miura model has been revised to improve clarity (equation S6.1 in SI). In particular, the existence of a homogeneous steady state is now parameterized by a tunable g<sup>*</sup>, that can be chosen as for spike initial patterns or for homogeneous-with-noise and spikehomogeneous initial patterns). We have also included a proof that the model equations satisfies our conditions R1-5. Indeed, the model is non-linear as long as σ<sup>i</sup>≠0 for some gene product (as we explicitly assume).

      The derivation of this cubic model from a separate expansion of general reaction-diffusion dynamics can be found in the original paper (Miura & Maini, 2004), with further applications to pattern formation that supporting its validity in subsequent works (Marcon et al., 2016; Diego et al., 2018). Importantly, this expansion is independent of the linearization performed in the main text of our article to derive the dispersion relation. The reference to this separate expansion in the previous version was included solely for contextual purposes; however, we have removed it in the revised manuscript to avoid potential confusion.

      (R.2.11) (b) The use of a Jacobian that is independent of the steady-state contradicts the assumption of nonlinearity (requirement R2 (l. 192f)) of f. We ask the authors to clarify this.

      We believe this concern arises from a notational ambiguity in the previous version of the manuscript, which has now been corrected: the matrix appearing in the regulatory term has been renamed from J to W<sup>T</sup>. As stated in the main text, the jacobian of the regulation function f(g) evaluated at the homogeneous steady state must coincide with the transpose of the network weight matrix. With the current equations (S6.1), we have , from which we easily get . Also, it is clear that the Jacobian of f(g) is not independent of g.

      (R.2.12) (c) In Figure S3 and similar simulations, the implementation of the nonlinear terms is ambiguous. The function f shown does not correspond to the Jacobian, and it remains unclear how these components are ultimately implemented in the simulation code. Additionally, as mentioned, it does not fulfill the necessary conditions for the global steady state.

      The implementation of nonlinear entries in f(g) whenever they are needed is now made explicit in the corresponding subsection of section S6 in the SI. With the new notation it becomes clearer that the fs used can fulfill the necessary conditions for the global steady state.

      (R.2.13) (d) The given function f_8 in S8.10.2 cannot correspond to the mentioned network since the number of gene products does not match the Jacobian and the network.

      This was a typo that has now been corrected.

      (R.2.14) (e) The given parameters for the figures in the SI do not match the figures. Please check and ensure that the correct figure is referenced (e.g., S8.2 Figure 3)

      This was a typo in the numeration of the subsections in the SI that has now been corrected.

      (R.2.15) (f) It is unclear which units are used, and the units used for the non-dimensionalization should be provided so one can relate them to biological systems.

      It is now explicitly stated in the revised version that the model equations are formulated in arbitrary units. This implies that the model dynamics are consistent with the characteristic units of any particular biological system under consideration. No non-dimensionalization of the model equations has been considered.

      (3) Conceptual and Structural Clarity

      The manuscript suffers from a lack of structural clarity, which affects both readability and scientific coherence.

      (R.2.16) (a) In one of the central figures (Figure 4) supporting their main claim, the naming of the network is not consistent with the main text. The network category referred to as "Over-Turing" is never mentioned in the main text. We suspect this should actually be labeled as the "noise-amplifying network."

      Indeed. This has now been corrected. We now use only the term “Over-Turing” in the article.

      (R.2.17) (b) The Supplementary Information includes an analysis of dispersion relations to classify patternforming networks, but this approach is not mentioned or referenced in the main text.

      This part of the SI has been moved to the main text and the dispersion relation has been fully and explicitly integrated in the overall argument of the article.

      (R.2.18) (c) In relation to Figure 6, we found that the concept of "diversity of possible final patterns" would benefit from a clearer definition and explanation. It is not immediately evident how this diversity is measured or what criteria are used to compare different networks. For instance, it is unclear why the Over-Turing network - which generates both periodic and noisy patterns - is considered to exhibit low diversity, whereas the Turing networks, which produce only periodic patterns, are described as having high diversity.

      This was just a large typo. The figure has been corrected. The reasons for this differences are now described in the last three paragraphs of the section “The ensemble of possible pattern transformations from H gene networks and spike initial conditions” for the hierarchical networks and in the last paragraph of the section “Pattern transformations in L- subnetworks from spike-homogeneous initial patterns ”, for the noise amplifying networks and in the seventh paragraph of the section “Pattern transformations in the combination of L+ and L- subnetworks” for the Turing networks.

      (R.2.19) (d) Additionally, the dependence of final patterns on initial conditions is not clearly described. It seems that this relationship is only analyzed for non-trivial pattern formations, but this is not explicitly stated. Clarifying these points in the caption of Figure 6 would greatly help readers understand the interpretation and significance of the results presented in this figure.

      Indeed, we have done nothing for the trivial pattern transformations. We are now more explicit about this already from the introduction. This article is only concerned with non-trivial pattern transformations. For each type of gene network we now provide a more detailed description of how the resulting pattern depends on the initial pattern (in the sections for each gene network).

      (R.2.20) (e) The significance statement is simply a verbatim repetition of parts of the abstract. This defeats its purpose, which is to articulate the broader implications of the work. We urge the authors to rewrite this section with a focus on significance rather than summary.

      We have now corrected this.

      (R.2.21) (f) We suggest including a dedicated figure to illustrate the biological model, depicting cells, intracellular and extracellular compartments, and the presence or absence of boundaries between adjacent cells. Such a figure would significantly enhance readers' understanding of the system being discussed.

      We have now done that. See new figure 3.

      (R.2.22) (g) We encourage the authors to strengthen the 2D and 3D results presented in the paper by adding supporting citations, sharing implementation details, or providing a more in-depth analysis of these systems. If such additions are not feasible, it may be best to remove references to the 2D and 3D systems to maintain clarity and focus.

      In the new version of the article we explain why our results on which gene networks can lead to pattern transformation do not depend on the dimensionality of the system. In fact, none of our proofs or arguments assumes or requires a specific number of dimensions. The networks are the same no matter the number of dimensions. The types of possible patterns can be seen as manifesting themselves differently depending on the number of dimensions. In the current version of the manuscript we explain now, every time we explain a resulting pattern, how the pattern is in 1, 2 and 3 dimensions and why. We have added Figures 1 and 9 for that purpose. As we explain in the text, the resulting patterns that are noisy would be noisy no matter the number of dimensions and the ones that are based on a spike in the initial pattern have necessarily radial symmetry (in any number of dimensions). Similarly the periodic patterns will be periodic no matter the number of dimensions (although some aspects of it will change). Similarly, in the 5th paragraph of the discussion we discuss the effects of the shape of the system and the boundary. There was a problem with the definition of pattern transformation we used, but this has now been corrected, in P1 and P2 in the introduction.

      (R.2.23) (h) The results section lacks a consistent structure. Section titles do not clearly indicate which phenomena or initial conditions are being analyzed, making it hard for readers to track the logical progression of the study.

      Now the results start with some introductory results with the subsections:

      “Basic requirements on gene networks capable of pattern transformation”

      The rest of the results are split into four clearly differentiated sections:

      “Gene network classification”

      “Linear stability Analysis”

      “Positive regulatory loops determine the kind of RD-instability”

      “Hierarchical Networks”

      “Emergent networks”.

      “Gene networks combining different classes of subnetworks”

      The last three sections have several sub-sections inside.

      We think that the titles of the sections are self-explanatory since hierarchical networks contain only H subnetworks while the emergent networks contain L+ or L- subnetworks and the last major sections is about how all these can be combined.

      Minor Issues

      (1) Notation and Terminology

      (R.2.24) (a) Variable naming is inconsistent throughout the paper. Terms like g_A(x) and A(x) (S5.2.1) are used for gene network concentrations without consistent usage. The naming of genes in networks also varies between the main text, SI, and figures. I.e., sometimes genes are labelled with small, sometimes with large letters, and sometimes with numbers.

      This has now been corrected.

      (R.2.25) (b) It would improve clarity to use distinct notations for intracellular vs. extracellular concentrations and gene expressions. Ensure networks and examples are consistent across all figures, captions, and supplementary materials. For example, RH2a and RH2b have different networks in the main text compared to the SI.

      As we now explain in the third paragraph of the “Methods: the model” section we consider, for simplicity, that gene products are either intracellular or extracellular. In that sense there is no possible ambiguity. As explained in that section, again for simplicity, we do not consider the receptor nor the signal transduction pathways of signals. This means that an extracellular gene product can “directly” regulate intracellular gene products. Because of that, we think that using different notations for extracellular and intracellular gene products would make things more confusing. We have corrected the misnaming between main text and figures.

      (R.2.26) (c) We suggest using distinct notation for the gene product itself and for its small deviation from a homogeneous steady state in the SI. This would help clarify whether specific statements apply only within the linearized regime or can be generalized to the full nonlinear dynamics.

      We do that in the new version of the article.

      (R.2.27) (d) Line 327 contains a mistake: g_k = g_j should be expressed as a proportional relationship. The division by g_A also seems unnecessary - please revise.

      This is now explained in a different way so this mistake does not apply.

      (2) Model Description

      (R.2.28) (a) Justify why boundary effects and spatial separation between cells can be neglected in the model.

      This is now discussed in the 5th paragraph of the model section. We do not claim that boundary effects are negligible. We claim, instead, that which are the gene networks that can lead to pattern transformations do not depend on the boundaries. The same occurs for the types of resulting patterns, in the coarse way we use, possible from each gene network and initial pattern.

      As stated in the first two paragraphs of the model section, the spatial separation between cells can be ignored because we assume there are many cells in the system and these are evenly spaced and sized (at least roughly). That is usually the case in animal development, although not always (there are exceptions in the very early stages of many marine invertebrates), and we do not claim to know exactly what happens in those cases: as we stated in the first paragraph of the introduction we assume systems made of many small cells.

      (R.2.29) (b) State explicitly that only extracellular gene products are assumed to diffuse - this is currently only mentioned in the SI.

      This is now explicitly stated early on in the first three paragraphs of the model section and also after the introduction of the model equations (1)-(3).

      (R.2.30) (c) In the Supplementary Information, the authors state that both extracellular and intracellular gene products can exhibit non-zero diffusion, which appears inconsistent with the conceptual framework and probably is a typographical error.

      This was indeed a typographical error. It is now corrected.

      (3) Assumptions and Requirements on f

      (R.2.31) (a) The equation for requirement R5 is incorrect as written in the main text and should be reformulated more rigorously. The condition should be stated for all constant values of g_i (and g_j) to avoid misinterpretation; otherwise, one might assume all matrix elements must have the same sign.

      This has now been corrected.

      (R.2.31) (b) Clarify what restrictions on f prevent pathological nonlinearities like 1/(g_k + \epsilon), which would contradict the assumed behavior at high concentrations.

      We do not understand this criticism. 1/(g_+\epsilon) fulfills our requirements on f and we do not see how is that pathological. We are unsure of what the reviewer means by the assumed behavior at high concentrations.

      (4) Figures and Captions

      (R.2.32) In Figure S3b, the diagram shows gene 5 being activated by gene 4, yet the caption states this is a negative regulation - please correct.

      This has now been corrected.

      (5) Readability and Formatting

      (R.2.33) (a) Improve navigation by hyperlinking references to equations, figures, and requirements throughout the document.

      In the new version we have inserted these hyperlinks.

      (R.2.34) (b) Adding hyperlinks to the requirements would additionally help the reader to keep track of them

      In the new version we have inserted these hyperlinks.

      (We.2.35) (c) Correct inconsistent or mismatched equation numbers and references. E.g. SI S5.1 is not referring to the correct equation (the equation it should be referring to would be Equation 3), and the reference to Figure 7 in part of the dispersion relation is wrong (as far as we see, this should be Figure 5).

      This has all been corrected now.

      (R.2.36) (d) Clarify ambiguous language in the introduction. For instance, the description of spike patterns (lines 136f) as a single cell spike contradicts the stated width (SI) and the visual representation involving 500 cells from the figures.

      This has now been corrected.

      (R.2.36) (e) The discussion of 2D and 3D simulations appears limited to the "noise amplifying" network. It's unclear whether a similar analysis was done for other network types.

      In Figures 1 and 9 and through the text we discuss all types of patterns in 2D and 3D.

      (6) Typos

      (R.2.37) Typos in the text (The following is just a small selection of the typos we came across. Since there are quite a few throughout the manuscript, we may not have caught all of them. We kindly recommend that the authors carefully proofread the full text to ensure consistency and clarity):

      We have corrected all the indicated typos and proofread the whole manuscript and SI.

      Reviewer #3 (Recommendations for the authors):

      Major concern:

      (R.3.1) Pattern formation can be induced by the positional information, and reaction-diffusion/Turing mechanisms is a foundational idea in the field. As in the references the manuscript cited, these paradigms were already clearly articulated and synthesized (e.g., Green & Sharpe's work (2015)). Moreover, the search for minimal network topologies that can generate Turing patterns has been extensively explored in Zheng et al. (2016). The novelty of the present work is unclear. It might offer a fresh perspective on an established problem, but it does not seem to present fundamentally new biological or mathematical advances.

      If the authors wish to strengthen the novelty and impact of the manuscript, they should consider explicitly acknowledging prior work and positioning their contribution as a formal extension or generalization, not discovery. To enhance the practical relevance of their work, the authors could demonstrate how their framework can be used to predict or classify gene network behaviors in pattern formation that are not easily identifiable through experimental approaches alone. For example, they could show how their classification helps distinguish between Turing, hierarchical, and noise-amplifying dynamics in complex or ambiguous biological systems, thereby offering a guiding tool for experimental design or interpretation.

      Indeed, the gene networks we identify have been identified before. We were and we are quite explicit about it, in the discussion, and we do cite the relevant work on that (including the one suggested by the reviewer). The novelty of the work is not identifying these gene networks, nor minimal ones, but showing that these are all the possible ones for pattern transformation (that there is no new type of network), this has not been done before (not even intended) and we are very explicit about that being our results (first paragraphs of the discussion).

      Minor concern:

      The writing style and language usage can be improved for clarity. Some explanations in the results and discussion can benefit from tight editing to eliminate redundancy and improve readability.

      We have corrected all the indicated typos and proofread the whole manuscript and SI.

    1. more advanced skills that were already present in some form in the child

      I love this idea because it made me think more deeply. It reminds me of a seed. A seed already has the potential to become a particular kind of tree, but it needs many different factors such as water, sunlight, nutrients, and time to continue growing. Eventually, it becomes a strong tree that benefits the world by supporting the ecosystem and providing shade for others.

      I think human development is similar. We don't suddenly gain completely new abilities as adults. Instead, we build on skills that already exist in an early form during childhood. For example, when we are five years old, we learn to tie our shoelaces. Later, we learn how to solve problems in school, such as passing an English test. As adults, we may solve much more complex problems, like running a business or leading a team. The skill is still problem-solving, it has simply become more advanced over time. That is why I love the idea of continuous development.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this study, the authors set out to determine how two classes of kinase inhibitors, which stabilise a disease-relevant enzyme in either an active (Type I) or inactive state (Type II), influence its organisation and interactions with microtubule filaments in cells. Using the state-ofthe-art in-cell structural imaging approaches, they examine how these compounds affect the formation of protein filaments and their association with microtubules, and succeed in defining the underlying structural basis for these differences.

      A major strength of the work is the application of in-cell cryo-electron tomography combined with correlative imaging, which enables direct visualisation of protein organisation in a near-native cellular context. The data convincingly demonstrate that the Type I inhibitor compound stabilising the active state promotes extensive LRRK2 filament formation and microtubule bundling, whereas compounds stabilising the inactive state markedly reduce these interactions. The structural analysis further provides insight into how conformational states relate to filament organisation, including modelling of previously unresolved regions of the protein.

      These findings are internally consistent and align well with prior biochemical and structural studies, many of which were performed by the same team.

      There are, however, some limitations that should be noted. The experiments rely on overexpression of the I2020T mutant form of the LRRK2 protein, which is a rare variant, in a single cell type (293T cells), which may not fully reflect endogenous behaviour or wild-type LRRK2 in a physiological context. In addition, while the imaging data are compelling, the functional consequences of the observed filament formation and microtubule association remain unclear.

      The study therefore provides strong descriptive and structural insight, but more limited evidence linking these observations to cellular or disease-relevant outcomes.

      Overall, the authors largely achieve their aims, and the results support their central conclusion that different classes of kinase inhibitors have distinct effects on protein organisation in cells. The work represents an important advance in understanding how small molecules can reshape protein architecture in a cellular environment, with potential implications for therapeutic strategies. The methodological approach will also be of broad interest to the field, as it highlights the power of in-cell structural biology to study dynamic protein assemblies that are difficult to capture using traditional approaches.

      We thank the reviewer for their thoughtful and positive assessment of our work. We appreciate their recognition that in-cell cryo-electron tomography and correlative imaging provide a powerful approach for directly visualizing how small-molecule inhibitors reshape LRRK2 organization in a cellular environment.

      We agree that the use of overexpressed LRRK2I2020T in HEK293T cells represents an important limitation of the present study. This experimental system was selected because it enabled visualization and structural analysis of inhibitor-dependent LRRK2 assemblies in cells. However, the extent to which these observations apply to endogenous LRRK2, wild-type protein, other disease-associated variants, or physiologically relevant cell types remains to be established.

      We also agree that the functional consequences of inhibitor-dependent LRRK2 filament formation and microtubule association remain unresolved. The goal of the present study was to define how type I and type II kinase inhibitors alter the cellular organization and structural state of LRRK2. Our data demonstrate that these inhibitor classes have markedly different effects on LRRK2 filament formation and microtubule association in cells, and provide a structural framework for understanding these differences. Future studies will be required to determine how these assemblies influence LRRK2 signaling, microtubule-based processes, and diseaserelevant cellular phenotypes.

      We thank the reviewer for highlighting both the methodological significance of this work and its potential implications for understanding how therapeutic molecules remodel protein architecture in cells.

      Reviewer #2 (Public review):

      Summary:

      Mutations in Leucine-Rich Repeat Kinase 2 (LRRK2) are a major cause of Parkinson's disease. LRRK2 PD-related mutations all result in increased kinase activity. Therefore, LRRK2 has been the focus of the development of kinase inhibitors. So far, two classes of kinase inhibitors have been identified: type 1 LRRK2-specific inhibitors that stabilize LRRK2 in a closed active-like conformation and broad-range type 2 inhibitors that stabilize LRRK2 in an open inactive-like conformation. Basiashvili et al. used here in cell structural biology to study the effect of both type 1 and type 2 inhibitors on the localization and structural conformation of LRRK2-I2020T.

      Strengths:

      They showed that Type 1 and not Type 2 inhibitors induce LRRK2 filament/ on microtubules.

      Furthermore, they were able to build a structural map of full-length LRRK2 I2020T bound to a Type 1 inhibitor in a closed kinase confirmation. Together, this work thus confirms the data of previous studies that showed that LRRK2 Type 1 and 2 inhibitors differently affect filament formation.

      Weaknesses:

      All conclusions are fully supported by the provided data. However, as the authors indicated themselves, the physiological relevance of LRRK2 microtubule binding is questionable. Furthermore, although the authors used a full-length LRRK2 protein, like in previously published structures, the resolution of the N-terminal domains is rather poor. Therefore, it also remains unclear what we learn from this structure compared to the previously published structures.

      We thank the reviewer for their positive evaluation of our study and for recognizing that our conclusions are supported by the data.

      We agree that the physiological relevance of LRRK2 filament formation and microtubule association remains an important open question. Our study was designed to determine how type I and type II inhibitors affect the cellular organization and structural conformation of LRRK2. We explicitly acknowledge that future studies using endogenous LRRK2, disease-relevant cellular systems, and functional assays will be necessary to determine the biological significance of inhibitor-induced microtubule association.

      We also appreciate the reviewer’s comment regarding the resolution of the N-terminal domains. Although the N-terminal density does not support detailed atomic interpretation, its visualization provides information about the global organization of full-length LRRK2 within an inhibitorinduced, microtubule-associated assembly in cells. Importantly, our study does not claim highresolution structural determination of the N-terminal regions. Rather, the advance is the in-cell structural observation of full-length LRRK2<sup>I2020T</sup> in a type I inhibitor-stabilized, closed-kinase conformation, together with density indicating that the N-terminal repeat regions adopt an organization within the microtubule-associated lattice.

      We have revised the manuscript to clarify this point and to more carefully distinguish the structural information supported by the density from interpretations that would require higherresolution data.

      Reviewer #3 (Public review):

      Summary:

      This paper describes new insights into the effects of type-I and type-II LRRK2 inhibitors on HEK293T cells that over-express GFP-labeled LRRK2-I2020T. Using correlative light microscopy and cryo-electron tomography, a type-I inhibitor leads to the extensive decoration of microtubules with LRRK2, which is not seen for a type-II inhibitor. Subtomogram averaging reveals that LRRK2 binds to the microtubules in a closed-kinase conformation, with density for the N-terminal arms.

      Strengths:

      The paper is well written; the CLEM and cryo-ET appear to be done to a high standard. Consequently, I have only minor comments.

      Weaknesses:

      The resolution of the subtomogram averages is somewhat limited, but the authors have adequately limited the number of degrees of freedom in the fitting of their atomic models by only allowing rigid-body transformations of separate parts of LRRK2.

      The authors should include FSC curves between the rigid-body fitted atomic models and the various sub-tomogram average maps.

      We thank the reviewer for their positive assessment of the manuscript and for recognizing the quality of the correlative imaging and in-cell cryo-electron tomography analyses.

      We also appreciate the reviewer’s recognition that our interpretation of the maps was appropriately constrained by fitting domains as rigid bodies, rather than attempting unsupported high-resolution model refinement.

      We thank the reviewer for highlighting this and apologize for the oversight. We have added all the missing FSC curve plots of subtomogram maps presented in this study in Extended Data Figure 8.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I think the current study is OK as it is, and the authors have taken this as far as they can.

      In future work, for either the authors or others in the field, it will be important to determine whether endogenous LRRK2 can be recruited to microtubules in response to compounds that stabilise the active state, particularly in cell types that are more relevant to Parkinson's disease. Does this cause a roadblock that impacts microtubule-driven transport? Establishing whether such recruitment occurs under physiological expression levels will be critical for assessing the broader relevance of the findings.

      In addition, it would be valuable to evaluate whether these Type 1 compounds have detrimental cellular effects linked to altered endogenous LRRK2-driven microtubule association, and whether inhibitors that stabilise the inactive state offer a potential advantage by avoiding this phenotype.

      We thank the reviewer for insightful recommendations for future studies.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 5: What is map C, and how is it different from the other maps? The authors indicate that the resolution of the N-terminal domains is moderate. How certain are the authors of the fit of these domains? Since map C is not provided in the supplemental, it is not possible to check this.

      We apologize for this oversight. We have updated the text to reflect how the map C was calculated. Now the text reads:

      “Additionally, we performed subtomogram analysis in Dynamo on a larger LRRK2<sup>IT</sup>decorated lattice that contained three layers of LRRK2<sup>IT</sup> density around the microtubule; we refer to this average as map C. Refinement was focused on the central four LRRK2<sup>IT</sup> subunits to better resolve additional protein densities within this larger lattice. In map C (Fig. 5A; Ext. Fig. 7).”

      In addition, we updated the figure 5D-F to demonstrate clear fit of the N-terminal domains into the presented map. We also added an Extended Data Figure 7 to the supplemental materials to highlight the fit of the model in the map and highlight the areas that would correspond to the Nterminal domains of LRRK2. We hope these updates demonstrate a good fit and justify observations highlighted in the paper.

      (2) The authors convincingly confirm that LRRK2 Type 1 and 2 inhibitors differently affect filament formation and that type 1 LRRK2-specific inhibitors stabilize LRRK2 in a closed activelike conformation. However, from the way the paper is written, it is unclear what we learn from this new structural data. How similar is the current structure compared to the previous structures? What is the novelty?

      We thank the reviewer for noting that this is unclear and giving us the opportunity to highlight it in the manuscript. We have added the following sentence in the discussion:

      “However, how the N-terminal repeats of LRRK2 are organized when the protein is in its closedkinase conformation remained unresolved. Stabilization of LRRK2 in a closed-kinase conformation by MLi-2 treatment and microtubule association reduces conformational heterogeneity to permit structure determination of full-length LRRK2<sup>IT</sup> with the N-terminal repeats undocked from the catalytic core. Therefore, the key novelty of this structure is that it captures full-length LRRK2<sup>IT</sup> in a cellular, microtubule-associated closed-kinase state and shows that kinase closure is compatible with an undocked N-terminal architecture. This distinguishes the in situ closed-kinase state from previously described in vitro intermediate active states.”

      Minor comments:

      (1) "Its C-terminal catalytic region is composed of WD40, Roc GTPase, Kinase and COR (RCKW) domains."

      Suggest changing this to Roc GTPase, Cor, Kinase and WD40 (RCKW) domains for clarity/following of abbreviation.

      We have made this change.

      (2) "In the MLi-2 treated cells, LRRK2IT strands were organized around microtubules with a regularly spaced lattice, similar to the LRRK2IT strands in cells not treated without the inhibitor (Fig. 3A-E)"

      Phrasing, correct the underlined portion.

      We have made this change.

      (3) "While average pitch. rise, and handedness of the filaments of the rate GZD-824 treated LRRK2 filaments were similar..."

      Punctuation.

      We have made this change.

      (4) "Our results clarify the relationship between kinase conformation, repeat undocking, and microtubule association. Increased microtubule association observed for I2020T mutant favors repeat undocking, a prerequisite for kinase closure and filament assembly"

      Do the authors mean undocking by the N-terminal repeats or repeatedly undocking of these domains?

      We meant undocking of the domains, and have corrected the sentence to clarify this.

      (5) "Together, these findings provide a structural view of full-length LRRK2 in a closed kinaseconformation and capture a resolved snapshot along its conformational continuum"

      Needs a space.

      We have made this change, and thank the reviewer for pointing it out.

      (6) "Microtubule decoration by LRRK2IT has not been studied in cell types that endogenously express high levels of LRRK2, such as lung epithelial cells and brain-resident immune cells including microglia and macrophages44. Thus, it remains possible that aberrant LRRK2microtubule interactions occur under physiological expression conditions, potentially disrupting homeostatic intracellular transport and being further exacerbated by type I LRRK2 inhibitors, as suggested by in vitro studies23,45."

      Many studies have studied the localization of endogenous LRRK2, however were not able to detect filament localization on microtubules. Moreover, to my knowledge, there is also no clear evidence that type 1 inhibitors disrupt microtubule transport in cells expressing endogenous levels of LRRK2.

      Therefore, I suggest to rephrase or remove this paragraph.

      We agree that the current evidence does not establish that this occurs broadly in cells. However, to our knowledge, cells or tissues with high endogenous LRRK2 expression have not yet been systematically examined in this context. We therefore present sparse decoration of hyperactive LRRK2 on microtubules as a possibility rather than a strong conclusion. We have also previously shown that type I inhibitors disrupt microtubule transport in vitro, but determining whether a similar effect occurs in cells is ongoing work and beyond the scope of the present manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) P4: The first section of the Results refers to LRRK2 localising to microtubules in the presence of the type-I compounds, and to the cytosol with the type-II inhibitor. Aren't microtubules in the cytosol also?

      We meant cytosolic LRRK2, we have revised the text to reflect this. It now reads:

      In cells treated with MLi-2, we observed LRRK2<sup>IT</sup> in extended filaments, puncta, and diffuse in the cytosol (Fig. 1D-E; Ext. Fig 1A-D). In contrast, when cells were treated with GZD-824, LRRK2<sup>IT</sup> was mostly localized to puncta and distributed throughout the cytosol, with reduced filament formation (Fig. 1F-G; Ext. Fig 1E-H), in agreement with our previous work [23,24,40].

      (2) P4: second column, halfway down. I don't understand how the 16 and 8 neighbours are derived from Figure 3J-K. Perhaps indicate this in the figure?

      Thank you for bringing this to our attention. We have added an Extended Data Figure 5 to clarify this point. The Extended data figure 5 highlights and annotates the immediate neighboring LRRK2 densities in the MLi-2- and GZD-824-treated lattices, making clear how the 16 and 8 nearest-neighbor values were assigned from the observed lattice organization.

      (3) P6: first column, halfway down: perhaps make it explicit that only rigid-body fitting was performed because of the limited resolution?

      We have incorporated this useful suggestion. The text now reads:

      “We split this model in three parts: the WD40 and C-lobe of the kinase, the N-lobe of the kinase with ROC and COR domains, and the LRR and ANK domains, aligned and fitted each of these three to our map A (Fig. 4D-F). Given the limited resolution of the map A, we fit the model as three rigid bodies without atomic refinement.”

      (4) P6: same column near the bottom: what is map C? and how was it calculated? Also, it is not clear to me from Figures 5D-F whether the statement "clearly correspond to the LRR-ANK-ARM domains" is justified by the map. From Figure 5D-F, I see a rather poor fit in a low-resolution map. This needs to be toned down or better illustrated.

      We apologize for the oversight. We have updated the text to clarify how the map C was calculated. Now the text reads:

      “Additionally, we performed subtomogram analysis in Dynamo on a larger LRRK2<sup>IT</sup>decorated lattice that contained three layers of LRRK2<sup>IT</sup> density around the microtubule; we refer to this average as map C. Refinement was focused on the central four LRRK2<sup>IT</sup> subunits to better resolve additional protein densities within this larger lattice. In map C (Fig. 5A; Ext. Fig. 7).”

      In addition, we updated the figure 5D-F to better demonstrate the fit of the N-terminal domains into the presented map. We also added an Extended Data Figure 7 to the supplemental materials to further highlight the fit within the map and indicate the areas that correspond to the N-terminal domains of LRRK2. We hope these updates clarify how map C was calculated and better illustrate our interpretation of the additional densities.

    1. Reviewer #1 (Public review):

      Summary:

      The authors investigate the relationship between feedback responses and trial-to-trial learning. In their paradigm, participants were constrained to a channel trial, and a cursor was visually perturbed. Using a channel-perturbation-channel structure, the authors obtain feedback responses to the perturbation and the learning response that ensues. In Experiment 1, the authors demonstrate that temporal dynamics of the learning response (LR) are poorly linked to temporal dynamics of the feedback response (FBR). The LR responses are yoked to the start of the movement, even in cases where the FBR is very delayed. Then, in Experiments 2 and 3, the authors dissect FBR and LR responses into two components: (1) a phasic component that has a peak point mid-movement and then declines, and (2) a tonic component that grows over the movement time course and remains stable during the holding period. The authors provide evidence that LR responses are better predicted from the tonic component of the FBR than the phasic component. The idea that tonic FBR components drive learning over phasic components departs from prior models of error-based learning and provides a new theory to understand sensorimotor adaptation.

      Strengths:

      (1) The paper is well-written, and the contribution is important and timely. The authors provide clear experiments that change the way we conceptualize how trial-to-trial learning is driven by feedback responses to error.

      (2) The paper provides solid evidence to demonstrate that feedback (FBR) and learning (LR) responses are not linked by a fixed delay, in contrast to prior models.

      (3) The paper also introduces the concept that both tonic and phasic components of the FBR differentially influence the learning response. The paper provides solid evidence that the tonic forces maintained during holding still have an impact on the learning that proceeds on the next trial. This has implications for models of sensorimotor adaptation and our understanding of the physiology of learning.

      Weaknesses:

      While some conclusions are strong, I feel that the conclusions regarding FBR and LR relationships need additional analysis. All these concerns are elaborated below. Broadly speaking, there is a concern that some conclusions reached by the authors are linked to the particular phasic/tonic model they use to parse FBR and LR responses. Other models are not considered and could lead to differing results. Furthermore, it is assumed that LRs are scaled FBRs. This assumption excludes the possibility that LRs could be driven by FBRs and other mechanisms, which would alter the way the regression analyses are constructed. As described below, model-free analyses are warranted to corroborate the main findings. Further, the role that phasic-FBR plays in the adaptation process is understated in the Discussion despite evidence to the contrary in Figure 8. Much of the analysis is done on trial-averaged and participant-averaged responses, inflating R2 values. More analysis should be done at the trial level to better examine model performance and accuracy. And while valuable, the authors' experimental approach differs from standard force-field experiments that were initially used to test feedback error learning hypotheses. The paper could benefit from a Limitations section to discuss associated limitations.

      Main Concern 1:

      The decomposition of FBR and LR into phasic/tonic components is based on a specific model (i.e., Equation (1)). The notion that tonic FBR predicts phasic/tonic LR is based on responses estimated from the model. Thus, it is unclear whether critical findings (e.g., LR responses are predicted by tonic FBR) are true of the "data" or true when the "data are analyzed in the context of their model". In other words, had the authors proposed a different model to decompose the LR/FBR into tonic/phasic components, would they obtain different results?

      There are many possible alternatives:

      (A) In Equation (1), the phasic and tonic components are assumed to add linearly at all times to obtain the force profile. But the phasic and tonic components could be applied at separate times. The tonic component could be invoked during holding, and the phasic component could be invoked during moving. This type of model will differ from the current version, especially in how the peak force during the moving period is assigned to the phasic/tonic components.

      (B) Another possibility is that the tonic and phasic components do indeed operate at the same time (like in Equation (1)), but they are separate, independent controllers. In the author's model, the tonic component is dependent on the phasic component.

      (C) Another possibility is that the tonic and phasic components are linked, but not by an integral.

      (D) Another possibility is that the phasic component is not a Gaussian function of time.

      Concern 1-1:

      While it is not possible to explore the entire model space described above, the authors should consider whether other phasic/tonic model classes could lead to qualitatively different results. The authors could also consider other phasic/tonic models if appropriate, and demonstrate that Equation (1) is superior based on an information criterion like AIC or BIC.

      Concern 1-2:

      I recommend that the authors pursue model-free, empirical analyses to support their findings. This would decrease the reliance on the "correctness" of a particular model. One logical choice would seem to be empirically estimating the phasic component as the peak force during the moving period and the tonic component as the average force during the holding period. In this model-free estimation of phasic and tonic commands, is it still the case that tonic FBR alone predicts LR components?

      Concern 1-3:

      Building on Concern 1-2, a clear case where the concern about using a model alone to estimate phasic and tonic components is in the across-subject variability analysis in Figure 7. Here, LR and FBR are compared to one another only in the context of the tonic-phasic model in Experiment 1. The result is that only the tonic FBR predicts the tonic LR. But investigating Figures 7b and 7c, it would appear that the peak force applied during the FBR during the moving period (which should reflect the phasic component in large part as in Figure 4a) would predict the peak (or average) force applied during the LR. Thus, the conclusion that tonic FBR only predicts tonic LR may be driven by how the model estimates tonic/phasic FBR/LR rather than a true property of the data. A model-free analysis, as suggested in Concern 1-2, would be helpful in addressing this concern.

      Main Concern 2:

      Analyses in Figures 4g, 4h, 6c, and 6d are based on relating LR and FBR components with no intercept: y = ax; the LR component is a scaled FBR component. It is unclear if the authors' conclusion would vary had a different model been used. For example, suppose that LR on trial n is partly determined by the FBR and also the sensory error (e) on trial n-1 (where c1 and c2 are constants):<br /> LR(n) = c1 FBR(n-1) + c2 e(n-1)

      Another model could suppose that the LR on trial n is due to the FBR on trial n-1, and also a non-specific adaptive component that is independent of both FBR and the sensory error:<br /> LR(n) = c1 FBR(n-1) + c2

      Concern 2-1:

      For these alternate models, y=ax (i.e., zero intercept) is not an appropriate relationship between LR and FBR components. Had the authors allowed a non-zero intercept in Figs. 4g, 4h, 6c, and 6d, will they still observe that only tonic FBR predicts LR components? In other words, would R2 improve for phasic FBR relationships with a non-zero intercept?

      Concern 2-2:

      Why was a non-zero intercept allowed for the between-subject analyses in Figure 7, but not for similar analyses in Figures 4 and 6?

      Main Concern 3:

      The main results in Figures 4g, 4h, 6c, and 6d are based on an R2 value that is calculated on a linear fit to the mean response averaged across participants and trials. This raises the concern that the R2 value is being inflated, and it also misses the rich trial-to-trial variation and subject-to-subject variation that could be used to examine the model's accuracy. A couple of concerns here:

      Concern 3-1:

      As can be seen from the horizontal and vertical error bars in Figures 4g and 4h, there is considerable variability across participants. While not shown, it is almost certainly the case that there is considerable variability across trials within a participant (as alluded to in the Fig. 8 analyses). The authors should evaluate their model performance and report goodness-of-fit (or error) at the single-trial level. For example, the model could be fit to individual trial data, and the R2 values from the trial fits could be used for comparing the various relationships in Figures 4 and 6. Another idea would be to keep the alpha, beta, T and sigma estimates obtained from the average data, and then apply these parameters to individual trial responses and report the model error. Do phasic FBR commands similarly predict LR components at the trial level, or do trial-level analyses corroborate the current conclusions on tonic FBR superiority?

      Concern 3-2:

      The authors report on Line 200 that the R2 values of 0.635 and 0.698 have modest predictive power. It would be helpful for the authors to statistically compare the R2 values between Figures 4g and 4h. One idea would be to obtain an R2 value for each individual participant. Then the distribution of R2 values across participants could be compared between the different relationships in Figure 4g/4h (e.g., via a t-test). This would help to better support the idea that Figure 4h shows better model fits than Figure 4g. These analyses could also be conducted for the relevant parts of Figure 6 (Experiment 3). The authors should consider allow a y-intercept in this process as they do in Figure 7.

      Main Concern 4:

      The authors compare tonic and phasic FBR predictive power in Figure 4. There are other places where the analyses in Figures 4g and 4h should be repeated:

      Concern 4-1:

      Tonic and phases FBR responses appear to vary in Experiment 1 (Figure 2c), but the authors do not test whether they predict the LR component magnitudes in Figure 2d. Analyses in Figures 4e,4f, 4g, and 4h should be added to the Experiment 1 analysis.

      Concern 4-2:

      While I understand the rationale behind computing differences in Figure 6 to isolate the second-shift effect on FBR/LR, the authors should still perform the primary investigation in Figures 4e, 4f, 4g, and 4h on the FBR and LR responses in Figures 5b-g (without subtracting the "Maintained" component). In other words, before analyzing the contributions of the second shift in Figure 6, the authors should repeat their analysis in Figure 4 applied to the FBR and LR responses in Figure 5 (without subtracting off the maintained response). How well does Equation (1) and y=ax capture the FBR and LR responses in Figures 5b-g?

      Main Concern 5:

      Given current practices in human sensorimotor adaptation, the current n=10 (or n=12) group sizes appear limited in size, raising concerns on statistical power.

      Concern 5-1:

      The authors should consider a power analysis or provide some other justification to support their chosen sample sizes.

      Concern 5-2:

      It is unclear why cross-correlation analyses in Figure 2e, 3d, and 5h have error bars, but no other FBR or LR time courses have error bars. Error bars should be provided in Figures 2b, 2c, 2d, 3b, 3c, 5b, 5c, 5d, 5e, 5f, 5g, 6a, and 6b.

      Concern 5-3:

      The subject counts are reported as n=10 for Experiment 1, n=12 for Experiment 2, and n=12 for Experiment 13, but the subject-to-subject analysis in Figure 7 says n=33.

      Main Concern 6:

      I agree that the author's model suggests that LR responses are most strongly predicted by the tonic FBR component. But I feel the narrative and Discussion surrounding this point are too strong. They paint the picture that only tonic FBR is important in learning. To do this, the role that phasic FBR plays is discounted, and mixed results concerning tonic FBR are overlooked. I feel that the Discussion should be broadened to acknowledge that the authors find evidence that both tonic and phasic FBR appear to influence the learning response, with tonic FBR making the stronger contribution in this task. Here are key areas that require attention:

      Concern 6-1:

      Importantly, the authors downplay their result in Fig. 8h, that the phasic FBR predicts phasic LR in their Results on Line 350. This argues against the idea that only tonic FBR influence LR parameters. On Line 485, the authors state that "trial-by-trial variability in LR amplitude was explained by the tonic component of the FBR, but not by the phasic component (Fig. 8)." This is not correct. Both the tonic and phasic components of the FBR altered LR components in Figure 8.

      Concern 6-2:

      Again, it is stated on Line 502, that the phasic FBR component "had only a modest effect on the LR". This again seems to underplay the result. The authors should amend their Results and Discussion to better acknowledge that their data support a role for both tonic and phasic FBR contributions to LR, but the tonic component appears to make a larger contribution in their model.

      Concern 6-3:

      While the role of phasic FBR in determining LR amplitude appears to be understated, the role of tonic FBR is, on occasion, overstated. The Discussion should mention that there is mixed evidence for the role of tonic FBR in LR parameters. For example, in their between-subjects analysis in Figure 7f, the authors do not find that phasic LR can be predicted by tonic FBR. Thus, across subjects, no component of the FBR appears to predict phasic LR.

      Concern 6-4:

      To better investigate the role that both phasic FBR and tonic FBR may play in adaptation, it would be advisable for the authors to consider this hypothesis. As it stands, tonic LR or phasic LR is regressed only onto tonic FBR or phasic FBR individually. In Figures 1 (Experiment 1), 3 (Experiment 2), and 5 (Experiment 3), the authors could regress tonic LR and phasic LR onto both phasic FBR and tonic FBR simultaneously. Models where LR = c1 phasic-FBR + c2 tonic-FBR could be considered and compared against univariate models, LR = c phasic-FBR and LR = c tonic-FBR using AIC or BIC to determine whether a mixed model that predicts LR with both phasic and tonic FBR is warranted.

      Irrespective of the result, the authors should be careful (Concerns 6-1 and 6-2) to state that when levels of tonic-FBR were controlled in Figure 8 (which is likely the cleanest way to look at the role phasic FBR plays in learning), phasic-FBR showed a clear influence on LR.

      Major Concern 7:

      On Line 577, it states the "hand was automatically returned to the starting position". Does this mean that the robot moved the hand back to the start location? If so, was the hand ever released from a force channel in between the perturbation trial and the following channel trial? A concern is that the holding forces from the perturbation trial could "bleed over" into the forces applied during the subsequent channel trial if the subject always remains in a channel trial in between the trials. Suppose we label the 3-trial structure as Channel 1 (C1) - Perturbation (P) - Channel 2 (C2). The authors should confirm that the holding forces on P are not correlated with baseline force (i.e., the channel force prior to movement onset) in C2. I do not expect there to be a strong correlation given that the learning responses in Figs. 2d, 3c, and 5e-g appear near-zero at t=-400ms, but this should still be verified.

      Major Concern 8:

      In Supplementary Figure 1, there appears to be an error in the "Amplitude of phasic LR (N)". In Supplementary Figure 1f, the phasic LR magnitudes appear in line with Supplementary Figure 1d, but there is a mismatch in the magnitudes for the phasic LR in Supplementary Figures 1e & 1d (the phasic LR magnitudes appear to be too low in Supplementary Figure 1e, peaking at around 0.1N when they should peak at around 0.15N).

      Major Concern 9:

      The authors should provide a Limitations section, highlighting unanswered concerns listed above, mixed results, and differences from prior work. These are touched upon in the Discussion section (particularly in Perspectives for future studies) but should be expanded further. At a minimum, the authors should consider including a discussion of the following points:

      Differences from prior work:

      9-1: There are methodological differences between this work and past studies highlighted by the authors. It could be that there are multiple error-based learning mechanisms that drive the FBR. Here, the authors find that visually-driven FBR responses do not drive LRs at a "common temporal shift". Instead, LRs are broadly expressed at the start of the movement (regardless of when the FBR was timed). However, tasks that have other components (e.g., a proprioceptive error) might invoke different learning mechanisms. For example, proprioceptive-driven FBRs might invoke LRs that have different temporal properties than visually-driven FRBs.

      9-2: As noted by the authors, Reference [10] studied FBR-driven learning in muscle commands, as opposed to forces. Muscle responses may have differing temporal and/or magnitude (for phasic/tonic) components that qualitatively differ from the force-based conclusions made here. Thus, the learning mechanisms at the muscle level may differ from those observed at the force level.

      9-3: While the tonic FBR is a strong predictor of the learning response in this experiment, most of the experimental conditions are done where the cursor remains deviated from the target throughout the trajectory and into the holding period. This differs from past work on feedback error learning, where feedback was veridical, and the cursor (and hand) ended on the target. This persistent displacement from the target during the prolonged holding period may influence the learning process and could enhance the tonic-FBR contribution to learning.

      9-4: The authors state in the present study that subjects were told not to use "explicit strategies" and move as straight as possible to the target. For past work, participants were able to use explicit strategies during feedback and learning responses. It could be that the lack of (or reduction in) explicit responses alters single-trial learning mechanisms relative to past work.

      Alternate models:

      9-5: No alternate models are considered here for the tonic-phasic relationship. Other models could relate these two processes differently, which could lead to different conclusions.

      9-6: It is assumed that both the tonic and phasic controllers are active at the same moment in time and sum linearly to generate the overall force output. Other models could have applied each "controller" to different phases of the reach in a differential manner (e.g., two separate controllers, a moving controller and a holding controller operating at different moments in time).

      9-7: It is assumed here that the LR should be a scaled FBR: y = ax. Conclusions made here could change if the LR is due to multiple processes, FBR-driven learning only being one of them. Other models where the LR is driven by both FBR and the sensory error were not considered here.

      Mixed results:

      9-8: While tonic FBR was a good predictor of phasic LR at the group-level (e.g., 4g), it did not predict phasic LR between subjects (Fig. 7f) and in fact tended toward a negative relationship.

      9-9: Phasic FBR predicts Phasic LR at the trial-level (Figure 8h) but not as well at the subject-level (Figure 7d).

      9-10: Overall, with the exception of Figure 8, most analyses look at the relationship between LR and tonic FBR or phasic FBR separately. In Figures 4c, 4d, 6c, 6d, and 7d-g, the authors look at the marginal effect of tonic or phasic FBR on learning, but do not control for variations in the other FBR component (e.g., they look at phasic FBR on tonic LR, but do not control for tonic FR). The only analysis that controls for the other component is in Figure 8, suggesting that both tonic and phasic FBR contribute to LR.

      Minor concerns

      (10) I'm not sure I follow the cross-correlation analysis in Figure 3. Overall, to me, both the FBR in Figure 3b and the LR in Figure 3c look quite similar in their temporal profiles, irrespective of the shift magnitude. The authors state on Line 158 that their cross-correlation analysis "...revealed that the overall shape of the cross-correlation function changed systematically with error magnitude". However, to me, in Figure 3d, the shape of the many curves looks similar.

      What is confusing to me here is including a phasic movement period and a tonic holding period inside the cross-correlation. The tonic "static" component during the holding period will likely greatly influence how well the cross-correlation is able to match the phasic peaks during the LR/FBR moving periods. In other words, the reach consists of a "movement" and a "holding" period. But the cross-correlation is blending the two together, and thus, I am not sure how reliable this measure will be for truly estimating the temporal shift between conditions. For example, if you look at the shaded gray area in Figure 3b, the "Movement period" looks almost identical in temporal properties. The "peaks" and "troughs" happen at nearly the same moment in time across all conditions. The onset of the FBR at approximately 200 ms is also identical across shift magnitudes. Thus, to me, the temporal properties of the FBR seem very similar during the moving period (where the FBR is responding to the error). But including the holding force (the tonic force after the 600ms period) seems to be causing the cross-correlation function to estimate differences at very high lags. If these differences are being driven solely by the holding forces, I am not sure this is meaningful.

      It seems that the authors might want to repeat this analysis, excluding the holding force period from the calculation of the cross-correlation coefficients.

      (11) It would appear that the authors have a significant main effect of their ANOVA (p=0.028) in Fig. 3f, but no post-hoc tests are reported to indicate which group means differ.

      (12) When plotting FBR, a [0,600]ms period is shaded as the movement period. On Line 580, it says that feedback was provided on peak movement speed. Was any feedback provided as to the movement duration? If not, did participants complete the movement within the 600 ms window labeled as movement speed? Were movements during perturbation trials longer than non-perturbed trials?

      (13) Over what time period is Equation (1) fit to the data? Is it the [-200,700]ms window shown in Figure 4a? A concern is that including too much of the "holding period" in the model fit will cause the model to be biased toward fitting the holding period well and not the moving period. This, in turn, might lead to better estimates for the beta parameter than the alpha parameter. In addition to clarifying the fitting process, the authors should also include R2 values for the moving and holding periods separately.

      (14) The procedure is clear from Figure 1e, but it would be helpful on Line 91 to explain that "collapsing" FBR and LR across rightward and leftward means that the FBR and LR were negated for one of the directions (prior to collapsing).

      (15) Are the "Amplitude of tonic LR (N)" supposed to be negative in Figures 6c and 6d?

      (16) Overall, the parameter distributions in Figures 4e and 4f are similar to those in Supplementary Figures 1c and 1d. The FBR amplitudes look nearly identical. Only the Phasic LR amplitudes in Supplementary Figure 1d appear to be larger than the Phasic LR amplitudes in Figure 4f. Can the authors provide an intuition for why the phasic LR contributions increase when T and sigma parameters are allowed to vary between participants?

      (17) There are two points where the authors should consider softening their language:

      17-1: The authors state at multiple points (e.g., Line 154) that "...the waveforms of LRs remained largely similar across conditions, while their amplitudes showed only modest modulation with cursor shift magnitude". However, in Figure 3c, the LR amplitude for the 0.4 cm shift is approximately 0.2 N, and the LR amplitude for the 3 cm shift is approximately 0.3 N - a 50% increase. The authors should consider softening the language here to appreciate the variations in LR amplitude.

      17-2: On Line 258, it is stated that the FBR during holding "diverged only slightly" for the 16 cm condition in Fig. 5b. This seems too strong a statement. The "Maintained" FBR holding force is about 0.2 N, and the reverse is about 0.1 N. Thus, the "Maintained" condition is doubled. While I agree that the LR diverges more than the FBR (i.e., 5b vs. 5e), I think the language choice here should be more careful.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study analyses correlations between traits of Chinese frog species and their Red List status, finding differences between adults and larvae and thus pointing to the importance of considering different life-cycle stages in this and possibly other animal groups when assessing species extinction risks. The current study is, however, incomplete because of unclear threat categories for tadpoles, the omission of other key species traits, and insufficient statistical analysis.

      Thank you very much. We have revised the manuscript according to the reviewers' comments. The parts highlighted in red in the manuscript are the revised portions.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript shows that different traits of adults and larvae correlate with Red List status. The authors argue that this shows a big gap in the conservation of amphibians and that the traits of all life stages should be taken into account in amphibian conservation. Specifically, amphibian conservation should do more for the habitats where the larvae live.

      The manuscript is well written and easy to understand. The methods are sound.

      While the study will make an interesting contribution to conservation science, there are many things that I disagree with.

      (1) I don't think that amphibian larvae and their requirements are a "blind spot" as the title suggests. When reading the manuscript, I didn't learn how conservation practice should change in response to the results.

      Thank you very much for your suggestions. The description of the 'blind spot' was inappropriate, and we have revised it. Investigating the relationship between life history traits and threat status can help us understand which species are more vulnerable to extinction. Furthermore, we can predict the potential threat severity of species that have not yet been assessed. Because we still lack knowledge about the biodiversity of many taxonomic groups. For example, as of early 2024, over 34% of Chinese anuran species have been described in the last ten years, and 100 - 200 new species are still being discovered globally each year. Under these circumstances, given the current investment in biodiversity conservation, it is nearly impossible to assess the threat status of every species and develop conservation strategies. Therefore, predicting the threat status of species is very important for biodiversity conservation, as it will provide support for the subsequent formulation of specific conservation policies. Among the already described animals species, most have complex life history cycles. Moreover, species face threats not only at the adult stage; those with certain traits at other life stages may also be vulnerable to threats. For example, our study takes amphibians as an example and shows that groups with larger body sizes at the tadpole stage may face more serious threats.

      (2) I wonder whether the relationship between species traits and extinction risk is of great importance for conservation. If a species is Data Deficient on the IUCN Red List, then species traits could be used to predict its Red List category. However, for other conservation projects, I don't see how this would work. How would traits be linked to captive breeding, conservation translocation, pond construction or habitat management in general? In some cases, I can envision a link between species traits and pond hydroperiod.

      Thank you very much for your suggestions. Understanding the relationship between traits and threat status is of great importance for the conservation policies and the allocation of conservation resources, especially when conservation resources are insufficient. As mentioned earlier, the current conservation resources are insufficient to support us in surveying and assessing every Data Deficient (DD) species, not to mention the large number of new species being discovered each year. By predicting threat status, we can identify which groups or species should be prioritized for research, such as population size and distribution range surveys, so that specific conservation strategies can subsequently be developed.

      (3) Species traits are body size and morphological traits. That makes sense. However, one of the species traits was microhabitat. I find it far-fetched to call habitat a species trait. This is standard habitat ecology. It is well known that habitats matter and that different habitat types face different threats, and consequently, the species that live in those habitats. Furthermore, habitat and morphology may be confounded. For example, tadpoles in lentic and lotic habitats have very different morphologies. So is it habitat or morphology?

      Thank you very much for your suggestions. The type of habitat in which a species lives affects the threats it faces. In many studies on the relationship between extinction risk and traits, microhabitat or habitat type is widely used as a predictive variable. For example, in studies on Squamata, whether a species is distributed on islands or peninsulas has also been included as a trait. Following your suggestion, we have revised the sentences to refer to 'morphological traits and microhabitat information'. Many morphological traits of species are related to habitat selection, but not all traits associated with habitat selection have been measured or have sufficient data. Therefore, it is necessary to include microhabitat type as an independent variable. Additionally, we calculated the Variance Inflation Factor (VIF) prior to the regression analysis to ensure that the analysis was not affected by multicollinearity.

      (4) I don't know how the threat status of Chinese amphibians is determined. IUCN has multiple reasons why a species can be Red Listed. One reason is range size, and another reason is population decline. Personally, I don't think they should be pooled in an analysis because they are fundamentally different reasons why a species has a high extinction risk. A reduction in population size of greater than 30% in 10 years or 3 generations is not the same thing as a small distribution range. Another issue is that IUCN developed the Green Status of species. The Green Status shows that even a species which is LC on the Red List may be significantly depleted.

      Thank you very much for your valuable suggestions. The assessment method of the China Biodiversity Red List is the same as that of the IUCN Red List, both of which are based on population size and area of distribution. We fully agree with your point that analyses should be conducted according to specific threat types. Unfortunately, the full report of the latest version of the China Biodiversity Red List, released in 2023, has still not been published. Therefore, we were unable to perform the relevant analyses.

      (5) The species traits in Table 1 are mostly functional/morphological and body size related (and microhabitat). While there may be correlations between traits and Red List status, it is unknown whether this is correlation or causation. In addition, it is difficult to know the conservation interventions that may be necessary now that we know that relative head with and Red List status are correlated.

      Thank you for pointing out the important distinction between correlation and causation. Your comment is very insightful, and we have revised our manuscript to further clarify the scope and limitations of our study. The aim of our study is to identify which traits show statistical associations with extinction risk, thereby providing testable hypotheses for future research. We acknowledge that the mechanisms underlying the associations between certain morphological traits (e.g., head length, tympanum diameter) and extinction risk remain unclear, and these findings cannot yet be directly translated into well-established management measures. Nevertheless, the value of our study lies precisely in generating hypotheses about traits that warrant prioritized investigation of their causal mechanisms, as well as offering clues for the initial allocation of conservation resources. Following your suggestion, we have discussed the limitations of the study in the Discussion section of the manuscript.

      (6) In the discussion, the authors explain why body size and other traits may affect extinction risk and whether there is a causal relationship. I agree that body size may have a direct effect because larger species are harvested more frequently (it was interesting to learn that tadpoles are harvested as well). However, as macroecological studies show, smaller species often have larger populations than larger species. Abundance may matter.

      Thank you very much for your suggestion. Following your advice, we have revised the discussion section regarding body size.

      (7) I found it much harder to understand why relative head length and tympanum size correlated with Red List status. I wasn't convinced by the arguments in the discussion. Typanum size may be related to hearing and anthropogenic noise. Several studies are cited which show that frogs alter their calling behaviour in response to noise. Crucially, however, they describe changes in behaviour or properties of the advertisement call, yet none show that noise has effects on population viability. If some anthropogenic stressor affects individuals, then this does not mean that it will cause a population decline. When IUCN published the second global amphibian assessment, did they list noise as a major threat to amphibians?

      We appreciate your insightful comments and fully agree with your assessment. Indeed, the hypothesis that noise threatened anuran amphibians lacks direct evidence. While relevant studies indicate that anthropogenic noise causes auditory masking in anurans and reduces individual reproductive success, the IUCN has not listed noise as a primary threat to amphibians. Although acoustic communication is vital for amphibian reproduction and is susceptible to noise interference, there is currently no definitive evidence proving that noise extensively impacts amphibian survival. Therefore, in the revised manuscript, we retained it as a hypothesis to be tested and explicitly clarified that current evidence is limited to behavioral changes. Regarding the correlation with relative head length, we acknowledge that the underlying mechanism remains unclear; it may stem from phylogenetic signal residuals or unidentified ecological factors (such as diet or locomotor ability). In the Discussion, we revised this part as a correlation requiring further investigation.

      (8) There are statements that the tadpole stage is the most important stage: "a critical period for amphibian survival" (line 78-79). While there is high mortality in the tadpole stage, tadpole survival is rather unlikely to affect population survival. Many population models show this. See, for example, Biek et al. 2002 in Conservation Biology. Other papers have argued that the postmetamorphic juvenile stage is most important (Petrovan and Schmidt 2009 Biological Conservation).

      We greatly appreciate your comment. We agree that the original statement was overly absolute. The most critical life stage for population persistence can differ across species, and many studies have shown that other stages may be more important. Accordingly, we have revised this sentence as you suggested.

      (9) The authors repeatedly make the statement that amphibian conservation should focus more on the tadpole stage. I don't understand why this statement is made. For example, a major activity in amphibian conservation is the restoration and de novo construction of ponds (see Calhoun et al. 2014 PNAS, Moor et al. 2022 PNAS). Ponds are habitats for tadpoles. Others removed fish from amphibian breeding sites because fish prey on tadpoles (and adults; see Vredenburg 2004 PNAS). Semlitsch (2002 in Conservation Biology) argued that the management of pond hydroperiod is a critical element of amphibian recovery plans. Ponds should be temporary because this effectively removes predators that consume tadpoles. Clearly, the tadpole stage is not a neglected stage in amphibian conservation.

      Thank you for pointing this out. The literature you cited (Calhoun et al., 2014; Moor et al., 2022; Vredenburg, 2004; Semlitsch, 2002) convincingly demonstrates that the tadpole stage has received a certain degree of attention in amphibian conservation practice. Our original statement was indeed problematic. What we intended to convey is that information on the tadpole stage needs to be integrated into conservation assessment frameworks and conservation planning. For example, many studies on the relationship between functional traits and threat extent have not included tadpole-related information. Compared with our knowledge of adult amphibians, we know far less about tadpoles, and for many species, information on the tadpole stage is entirely lacking. Therefore, we call for tadpoles to receive greater attention in future research relative to the current situation.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Conceptual problems:

      (1) Many conservation measures for amphibians target larvae; thus, globally, this is not a blind spot. If this is different in China, it would be important to point this out.

      We thank the reviewer for the thoughtful comment. We recognize that the tadpole stage has indeed received attention in amphibian conservation practice, and our original statement was therefore imprecise. Our intended argument was that tadpole-stage information should be integrated into conservation assessment frameworks and conservation planning. For instance, many studies examining the relationships between functional traits and threat extent have failed to include data on tadpoles. Our understanding of tadpoles remains far more limited than that of adult amphibians, and for a large number of species, no information on the tadpole stage is available. Consequently, we advocate for substantially greater research attention to tadpoles than they currently receive. We have revised the text accordingly.

      (2) While traits may be used to predict Red-List status, it is not clear how they could inform conservation measures. This should be discussed.

      Thank you for your comment. The aim of our study is to identify which traits show statistical associations with extinction risk, thereby providing testable hypotheses for future research. We acknowledge that the mechanisms underlying the associations between certain morphological traits (e.g., head length, tympanum diameter) and extinction risk remain unclear, and these findings cannot yet be directly translated into well-established management measures. Nevertheless, the value of our study lies precisely in generating hypotheses about traits that warrant prioritized investigation of their causal mechanisms, as well as offering clues for the initial allocation of conservation resources. Following your suggestion, we have discussed the limitations of the study in the conclusion section of the manuscript.

      (3) The Red-List categories may not be appropriate to link traits to extinction risk. It would be important to explain how these are defined for China and how this may affect the analysis (e.g. linking larval traits to larval extinction risks would be difficult if Red-List criteria do not consider larvae).

      Thank you very much for your suggestions. The assessment method of the China Biodiversity Red List is the same as that of the IUCN Red List, both of which are based on population size and area of distribution. The assessment process is independent of species' morphological traits. Consequently, analyzing correlations between traits and Red List categories does not constitute circular reasoning or contain any inherent logical contradiction. On the contrary, it is precisely because the two are independent that statistically significant associations between traits and extinction risk can have predictive value and inform conservation actions. In the revised manuscript, we clarified the independence of Red List assessments and rephrase any potentially misleading wording (e.g., changing "threat category of tadpoles" to "threat category of the species (assessed based on adults)").

      Methodological problems:

      (4) Choice of traits. Are morphological traits sufficient (add e.g. fecundity)? Justify the use of habitat traits (also, if additional ones would be included: geographic and altitudinal ranges, habitat specificity).

      Thank you for your suggestion. We fully agree that traits such as geographic range, elevational range, fecundity, and habitat specificity have important effects on extinction risk. The core objective of this study is to compare the stage-specific differences in the associations between extinction risk and morphological and microhabitat traits of adults versus tadpoles. Moreover, spatial traits such as geographic range are inherently highly correlated with the threat status of species, and including them might mask life-stage-specific signals. We will acknowledge this limitation in the discussion and identify the above-mentioned traits as important directions for future research.

      (5) Model choice: models have high uncertainty, thus better use model averaging and AICc instead of AIC. Overall, the statistical analysis and model selection procedure are poorly described; only summary results are presented.

      We greatly appreciate the reviewer's suggestion. Accordingly, we re-analyzed the data following your advice. In addition, the description of the methods has been supplemented.

      (6) Caveats: the data only allow for correlational analysis; causation cannot be derived from observational data. Furthermore, with a limited number of species, the number of predictors should not be too large.

      Thank you for your suggestion. Studying the relationship between traits and species threat status is important in conservation biology. Although such studies can only reveal statistical associations between traits and extinction risk rather than infer causality, they can generate hypotheses to facilitate future research. Additionally, this type of study can help predict the threat severity of unevaluated species, which is highly valuable for developing biodiversity conservation plans. In this study, 299 species were included in the analysis, and nine predictor variables (eight morphological traits plus one microhabitat type) were used. The ratio of sample size to number of variables was approximately 33:1, and variance inflation factor (VIF) tests indicated that multicollinearity was within an acceptable range (VIF < 5). Therefore, the risk of model overfitting is low. We will add this clarification in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) My first major concern is the species threat categories for tadpoles. The authors obtained the extinction risk data from the China Biodiversity Red List or IUCN. However, the assessment of threat categories, whether by the China Biodiversity Red List or IUCN, is based solely on adults. That means that the threat categories for both adults and tadpoles are the same, which can be seen in Figure 1. Since there is no specific assessment of threat categories for tadpoles, I have concerns about whether it is reasonable to relate species traits of tadpoles to the extinction risk for adults. I think it is one of the reasons why there is no study examining the association between functional traits and extinction risk in tadpole stages.

      We thank the reviewer for raising this important point, as it addresses a key prerequisite issue. The Red List assessment evaluates species, not individual life stages. The threat categories of both the IUCN and China Biodiversity Red Lists are determined based on criteria such as population size and geographic range of the species. The assessment process is independent of species' morphological traits. Consequently, analyzing correlations between traits and Red List categories does not constitute circular reasoning or contain any inherent logical contradiction. On the contrary, statistically significant associations between traits and extinction risk can have predictive value and inform conservation actions. In the revised manuscript, we will explicitly clarify the independence of Red List assessments and rephrase any potentially misleading wording (e.g., changing "threat category of tadpoles" to "threat category of the species (assessed based on adults)").

      (2) My second major concern is about the Data Analysis. The authors built and compared three types of models, i.e., PGLS_BM, PGLS_OU, and GLS_no_phylogeny. They claim that the OU-based PGLS model provided the best fit for both adult and tadpole datasets. Although the result seems reasonable, it is not clear how the OU-based PGLS model was obtained and what it exactly means. It seems to be a full model including all the predictor variables. However, since eight morphological traits and one microhabitat data of both adults and tadpoles were collected, there should be 29-1=511 candidate models. Unless the best model has an Akaike weight (wi) > 0.90 in all the OU-based PGLS models, it has substantial model selection uncertainty. If this is the case, the model average should be used, and weighted estimates of regression coefficients and unconditional standard errors that incorporate model selection uncertainty are better statistical methods (Burnham & Anderson, 2002).

      Thank you very much for your suggestion. Species' traits are related to evolutionary relationships, with more closely related species tending to be more similar. In the original manuscript, the three models we compared (PGLS_BM, PGLS_OU, GLS_no_phylogeny) were intended to select the optimal evolutionary covariance structure. Since we were more interested in the differences between adults and tadpoles, after selecting the OU structure, we actually used a single full model that included all traits to estimate the regression coefficients for each factor. Following your advice, we have added a model averaging analysis and revised the manuscript accordingly.

      (3) In addition, the Second-Order Information Criterion AICc, but not AIC, should be used for model selection. You have at least 9 variables (eight morphological traits and one microhabitat data) or 11/13 variables for the parameter estimates (Table 1). However, you have only 299 species included in the analysis (n = 299), which is relatively small compared to the number of variables (n/k << 40). Therefore, the AIC corrected for small sample size (AICc) should be used.

      We greatly appreciate the reviewer's suggestion. Accordingly, we re-analyzed the data following your advice.

      (4) Previous studies found that amphibian species with large body size, restricted geographic and elevational ranges, low fecundity or high habitat specificity are frequently predicted to have higher extinction risk (Cooper et al., 2008; Sodhi et al., 2008; Botts et al., 2013; Lips et al., 2003; Murray & Hose, 2005). The authors only included morphological traits and one microhabitat data point in the analyses. I wonder whether they can collect more trait data associated with extinction risk, such as geographic and elevational ranges, fecundity traits, or diet/habitat specificity, so as to gain more insight into the study.

      Thank you for your suggestion. We fully agree that traits such as geographic range, elevational range, fecundity, and habitat specificity have important effects on extinction risk. The object of this study is to compare the stage-specific differences in the associations between extinction risk and morphological and microhabitat traits of adults versus tadpoles. Moreover, spatial traits such as geographic range are inherently highly correlated with the threat status of species, and including them might mask life-stage-specific signals. In the Methods, we acknowledge this limitation and identify the above-mentioned traits as important directions for future research.

    1. We may dislike people from certain racial or ethnic groups because we frequently see them portrayed in the media as associated with violence, drug use, or terrorism. And we may avoid people with certain physical characteristics simply because they remind us of other people we do not like.

      As it goes for associational learning through the unjustified racial prejudices, I feel television also has an impact on this because some movies or shows will paint a character of color to be a stereotypical role to prove a point to viewers which makes people of the opposite color think about how that is true for tht race since it's a stereotype.

    1. Dr. Cutcha Risling Baldy (Hupa, Yurok and Karuk and an enrolled member of the Hoopa Valley Tribe in Northern California) Dr. Cutcha Risling Baldy (Native American Studies, Humboldt State University) researches Indigenous feminisms, California Indians and decolonization. In her blog post titled “Give It Back: Publishing and Native Sovereignty,” Cutcha writes: I’ve become obsessed with the idea of finding out what would happen if I started mourning loss of land, loss of lives, loss of fish - if my grief was on display. As an academic I’ve internalized the message that somehow the work isn’t supposed to be deeply personal. Like I don’t carry the blood of my ancestors in my veins, blood that has run rivers red as we held on to the bodies of slaughtered children and wailed into the night sky asking ourselves “why” or “what are we supposed to do now?” Like we didn’t sing or dance for all those we lost. Like that song doesn’t come from me now. Like I don’t close my eyes and hug my daughter just a little bit tighter at night because there was a time when they would have ripped her from my arms and sold her. And I would never stop looking for her. I would do anything to find her again. Like my ancestors didn’t search until they couldn’t search any longer. Like we don’t continue to search, or grieve even now. And we live here in this space that they stole from us. This place where we buried our beloved. Where we sing and dance and laugh and love. This place where we cried tears of joy and sadness and from laughing so hard our stomachs hurt and from hurting so hard we thought we’d never laugh again (2020). In just 3 short paragraphs Dr. Cutcha Risling Baldy describes what it is like to be a California Indian woman today. She bids us to think about land theft, loss and destruction. She makes us think about the significance of intergenerational trauma and how violence doesn’t just hurt the victim. Cutcha calls us to think about missing and murdered indigenous women and places that California Native people deem sacred. Her life is place based. She speaks and writes from an internal place that is spiritually, emotionally, and intrinsically connected to where her ancestors and she were raised, and where their creation happened. Colonization may have pushed many of us from our homelands, but we return. See more on Dr. Risling Baldy's work on the Hupa Coming-of-Age Dance as decolonizing praxis under Chapter 8, section 8.6: Transformational Liberation through Love.

      This part stood out too me also because Dr. Cutcha Risling Badly explains how the effects of colonization are still in the mist today. She connects with the loss of the land, family, and the culture to her own life, and this made history feel extremely personal instead of just something that only was a thing of the past. It’s very noticeable how she talks about intergenerational trauma and how the hurt and pain from the past situations continue to affect the family’s generation after generation. In this section I see that even after all that the Native communities battled; they continued to protect their culture, traditions, and connections to the land they came from the land that belong to them.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1 (Evidence, reproducibility and clarity (Required)):

      This work focuses on zebrafish notochord morphogenesis during axial elongation. In particular it dissects the role of YAP signalling on regulating the balance between caudal cell addition with the cell enlargement occurring rostrally through vacuolation.

      The article is timely to the field and includes several important experiments. The overall presentation and written style are good, citations are adequate and there is a clear effort to integrate experiments and mathematical modelling from the outset. The logic behind experiments is sound and the conclusion coherent (even if not totally unexpected given the literature): YAP affects progenitor addition which in turn changes packing, vacuolation and axis length. I just have a few points that could make the article clearer and more persuasive.

      We thank the reviewer for these positive comments about our manuscript. We would like to reiterate the two main unexpected findings based on our results:

      • While YAP mutants display a defective notochord (Kimelman et al., 2017; eLife) it has not been clear what specific role that YAP signalling is playing during notochord development. Therefore, the finding that Yap signalling plays a role in controlling the rate to notochord progenitor addition and represents a novel discovery.
      • The observation that the notochord can buffer its elongation rate against an increased influx of progenitors is novel and counter intuitive. Our current understanding of tissue elongation depends on the central idea that the addition of progenitors directly impacts elongation rate. Here we show for the first time that this has minimal impact at the tissue level using the notochord as an example. Major points

      - Last section of results is difficult and confusing. After analysing vgll4b loss-of-function line, effectively over-activating YAP, the focus is on YAP inhibition using Verteporfin.

      o Concerns on Verteporfin: the molecule has been widely used to module YAP, but there are also plenty of studies suggesting it is non-specific (also degrades YAP, has 14-4-3σ dependency and induces stress). I would consider an alternative: truncated TEAD, LATS over-expression or gain-of-function phosphomimetic versions of YAP.

      o Presentation: regardless of point above, Verteporfin's role on YAP should be verified in the system. As such it is crucial to include: images of 4xGTIIC, noto and YAP stains after treatment. Only then inspect the effects on vacuolation and different treatments.

      As suggested by the reviewer, we have added a supplementary figure validating the verteporfin treatment, including quantification of GFP reduction across the three tissues and quantification of notochord staining. We did not include Yap1 immunostaining data because the signal quality was insufficient for reliable analysis.

      A simple over-expression experiment will not allow the spatial and temporal control required to test our hypothesis. Yap has a known function in gastrulation, so we need experiments that allow us to perturb Yap activity only at posterior body elongation stages. This has been achieved with the vgl4b experiments shown in the manuscript, as this gene is specifically expressed in the tailbud at these stages. In addition to the full verification of verteporfin's impact on YAP activity, we feel this is sufficient evidence to support our conclusions.

      - In Fig 3F, noto HCR staining is taken as evidence for progenitor exhaustion/ faster depletion. Other scenarios would be possible without more direct demonstration. Evidence (either experimental or literature) that YAP is not involved in self-renewal or induction of these progenitors at these stages should be discussed.

      We have concluded that the smaller volume of noto expressing cells is consistent with the faster depletion of the progenitor pool based on the direct observation of increased progenitor addition rate from photo-labelling experiments (Figure 3A,B). As suggested by the reviewer, we have now quantified cell divisions within the midline progenitor population and found no significant differences between mutant and control embryos. These data have now been included in Supplementary figure 3.

      - Individual datapoints in Fig 3C and 4D should be shown.

      These data have now been added to the figures

      Additional justification is needed as to why spinal cord is the best to benchmark displacement. Additionally looking at this with respect to mesoderm migration could capture another set of progenitors and behaviour/ displacements.

      Photolabels within the pre-somitic mesoderm are difficult to interpret as the high amount of cell rearrangement in this tissue leads to a spreading out of the labelled clone in a manner that then makes it difficult to assess tissue displacement (see Figure 2D,E; Thomson et al., (2021) Cells and Development). In contrast, aprevious paper has shown that notochord-spinal cord displacements can be mapped in a reliable manner across the anterior-posterior axis which motivated our choice here (McLaren and Steventon (2021) Development).

      - Plotting vacuole area in Fig.4I vs A-P position (similar to plots 1H, 2F-H) could further strengthen the point of gradual (linear) vacuolation.

      As suggested by the reviewer, we have plotted vacuole area as a function of position for the verteporfin treatment experiments, and these data have now been included in Figure 5.


      Minor points:

      - Scheme of Fig1A could benefit from having the info of zebrafish timeline (hpf)

      The scheme has been modified indicating zebrafish timeline

      - Figure 3B, what was time 0?

      Timepoints have now been included in the text and figure legend

      - The authors should address whether Verteporfin-treated mutants are rescued or whether the compound overwhelms the genetic effect.

      Given that verteporfin will impact Yap signalling in a global manner, whereas the vgl4b have a localised over-activation of Yap signalling, we think this experiment would be difficult to interpret and would likely be non-informative.

      - Cell density is an elegant measure but quite abstract. A plot of cells detected at each AP position would be quite valuable to reinforce more cells are being added to a relatively constant area.

      As suggested by the reviewer, we have now plotted these data for mutant and controls and also for verteporfin treatments. These data have now been included in supplementary figures 3 and 7.

      Reviewer #1 (Significance (Required)):

      Significance included above.

      Reviewer #2 (Evidence, reproducibility and clarity (Required)):

      Summary

      Camacho-Macorra et al. investigate the mechanisms of axis extension in zebrafish embryos, focusing on the notochord and its two key elongation processes: progenitor addition (occurring early and posteriorly) and vacuolization (occurring later and in an anterior to posterior sequence). The authors first develop a mathematical model to predict notochord elongation dynamics by integrating these processes. They demonstrate that the YAP signaling pathway is active in both the notochord and its progenitors during axial extension. Their analysis reveals that vgll4b, an inhibitor of YAP, is expressed in the same regions. Knockdown of vgll4b results in YAP hyperactivation in the notochord and posterior progenitor regions, leading to increased progenitor recruitment into the notochord and a reduction in the progenitor pool. The effects of this mutation on extension are most pronounced during the late phase, which is dominated by vacuolization. The authors observe smaller vacuoles in mutants during this phase. However, early (but not late) YAP inhibition decreases notochord cell density and increases vacuole size, suggesting that YAP primarily regulates notochord progenitor uptake, which indirectly affect vacuolization.

      Major Comments

      The authors propose that YAP activity mediates a long-range feedback mechanism linking posterior progenitor addition to anterior vacuolization. Two lines of evidence are presented to support this idea. First, there appears to be compensation for tissue length during Phase 2, when both progenitor addition and vacuolization occur. Second, temporal YAP inhibition experiments show that early, but not late, YAP inactivation affects both cell addition and vacuolization. While these observations are intriguing, they do not conclusively demonstrate spatial long-range coordination. Instead, the global decrease of vacuole size could be a simple delayed consequence of cell density increase or cell disorganization at the posterior end without involving a long-range feedback along AP axis. Claiming that such long-range feedback is taking place would require a more precise characterization and/or the identification of its nature (chemical, mechanical).

      We would like to thank this reviewer for this point, that we feel requires further clarification. As they suggest, the increased additional rate of posterior progenitors leads to a later impact on vacuolation, once these cells have reached more anterior parts of the body axis- creating an effective long-range feedback mechanism to link the two processes. However, this is not a direct propagation of a signal (mechanical or otherwise) across the length of the notochord, as may have been interpreted to be based on the previous framing of our conclusions. We have modified the title of our manuscript to place less emphasis on the 'long-range feedback', and included an additional discussion paragraph to make this point clearer.

      Furthermore, there are several caveats with the interpretations of the claims cited above. The authors do not show quantification of vacuole area using notochord cell segmentation as described in Fig 1C in vgll4b mutants at stages when progenitor addition is increased.

      This is an important point highlighted by the reviewer. We have now included analysis at 24 hpf, where we do see a significant reduction in vacuole area within the anterior part of the notochord during the buffering phase in vgl4b mutants- consistent with our model that reduced anterior vacuolation compensates for increased progenitor addition rate during this phase of notochord elongation (Figure 4E).

      The slope of internuclear distances in Supplementary Figure 4A at 27 hours post-fertilization suggests that vacuolization is initially normal (similar to wt context in Fig 1H), arguing against an early defect in vacuolation dynamics along the Anterior to Posterior axis that could compensate for extra addition of progenitors.

      We have revised Supplementary Figure 4 to present a direct comparison between mutant and control embryos at each time point analyzed. This analysis shows that within the mid-trunk region of the notochord, differences in cell size first emerge at the developmental stage when vacuolation becomes the primary driver of axis elongation. In addition, we observe a progressive decoupling of the scaling relationship in mutant embryos over time. As mentioned above- there is a significant difference in vacuole size within more anterior regions at 22.5 hpf that is consistent with the model that this is buffering against increase posterior addition.

      Finally, the timing of the analysis of the effect of Verteporfin treatments is unclear. According to the legend of Figure 4F, analyses for Treatment A (16-27 hpf) and Treatment B (27-38 hpf) were done at 24 hpf and 30 hpf, respectively. If this is the case, the 3-hour window for Treatment B may not allow sufficient time to reveal effects on vacuolization.

      We agree that the information regarding the verteporfin experiments was not clearly presented in the original figure, and we have therefore revised the schematic accordingly.

      To strengthen the claim of long-range coupling, the authors could:

      Provide direct measurements of vacuolization A-P dynamics/area during Phase 2, before the effect on notochord length in the mutant, to see if there is indeed a compensatory effect on notochord length for the additional accretion of notochord progenitors in the Vgll4b mutant.

      As suggested by the reviewer, we have added an earlier time point to the A-P area dynamics plot in phase 2, corresponding to a stage at which the effect on notochord length in the mutant is not yet detectable. At this stage, we observed no difference in vacuole area between mutants and controls. We have also included an earlier time point analysis in the anterior region of the axis, which shows a similar cell size difference to that observed later in a more posterior region (Figure 4F; see above response).

      Clarify the analysis timing of Treatment B to confirm that YAP inhibition during the vacuolization phase truly has no effect.

      This has now been clarified.

      Additionally, as a non-specialist, I found the distinction between the two modeling hypotheses difficult to follow. Specifically, it is unclear why the first hypothesis assumes YAP affects vacuolation rate, while the second assumes it affects vacuolation front speed. It is also not intuitive how front speed can be independent of vacuolation rate, as one would expect that if cells form vacuoles more slowly, the front should progress more slowly as well. Therefore, it could be good to clarify these aspects of the modeling part.

      We thank the reviewer for this comment and apologise for the lack of clarity in our description of the model. In our framework, the cell size profile along the AP axis of the notochord is governed by two distinct processes: (i) the addition of progenitors at the posterior tip, and (ii) vacuolation, which increases cell size and proceeds from anterior to posterior. We model the latter as a propagating wave with velocity vf​, such that cells begin to vacuolate when the wave front reaches their position.

      Importantly, in the model these two aspects of vacuolation are decoupled: the front velocity vf​ determines when a given cell starts vacuolating, whereas the vacuolation rate J determines how fast the cell increases in size once the process has started. Biologically, this corresponds to distinguishing between the propagation of a trigger or competence signal along the tissue, and the execution of vacuole growth within each cell. Our reasoning was that they need not be strictly proportional: a signalling wave could propagate at a given speed even if the downstream cellular response is slower or faster.

      This is why we considered two alternative hypotheses: either YAP modulates the propagation of the vacuolation front (affecting vf​), or it modulates the growth dynamics within each cell (affecting J). Our quantitative comparison with the experimental data supports the former scenario. This has now been clarified in the main text.

      Minor Comments

      While the study is technically sound, a few areas could benefit from improved clarity or additional data.

      An intriguing but puzzling finding is the reduction in the noto-expressing progenitor domain in vgll4b mutants, despite elevated YAP activity in progenitors. Intuitively, if YAP promotes progenitor maintenance or expansion, one might expect the noto+ domain to increase, not shrink. This paradox suggests that YAP may not only simply maintain progenitors but instead accelerates their differentiation or migration into the notochord (as stated in the manuscript and graphical abstract). Alternatively, YAP could only deplete the noto+ pool by driving premature entry into the notochord, though the lack of clear YAP upregulation in this domain would imply a non-cell autonomous role of YAP for this interpretation. The authors should discuss these possibilities more explicitly in the Discussion section and could consider including additional markers, such as proliferation assays or apoptosis markers, to clarify whether YAP affects progenitor proliferation, differentiation, or migration.

      As also suggested by the reviewer, we have included a cell proliferation analysis in Supplementary Figure 3 and have revised the Discussion section accordingly.

      In Figure 2B, the YAP activity reporter signal in the posterior floor plate is not immediately obvious. The authors should consider providing higher-magnification insets.

      As suggested by the reviewer, we have included higher-magnification insets in Figure 2

      In Figure 2C, the differences in tail shape between wild-type and mutant embryos are visually striking. If these differences have not been quantified or discussed, a brief comment in the text would be helpful.

      We did not see a consistent impact on the morphology of the posterior body, this has now been clarified in the main text.

      Supplementary Figure 6 describes embryo length differences in mutants but does not include a representative image. Adding one would strengthen the phenotypic description.

      As suggested by the reviewer, we have modified Supplementary Figure 6

      Figure 1C is not cited in the text as not associated with a result, but just a description of the approach that is used later in Fig 4I

      We have modified the text to include the appropriate figure reference.

      Finally, the authors might consider citing Michaud & Pourquié (2025) when presenting the role of hydrostatic pressure in axis elongation in the Introduction.

      We have now modified the text to include this citation which we agree is relevant to this work.

      Reviewer #2 (Significance (Required)):

      This study by Camacho-Macorra et al. presents a fascinating exploration of how YAP signaling and its inhibition by vgll4b coordinate progenitor addition and vacuolization during zebrafish notochord elongation. The work is well executed, with clear results and integration of mathematical modeling and experimental data. The findings shed new light on the molecular and mechanical regulation of axis extension, a fundamental process in vertebrate development. However, while the study is innovative and rigorously conducted, the central claim of "long-range coupling" between progenitor addition and vacuolization requires further substantiation. Addressing the points discussed below will make the study more convincing and accessible to developmental biologists and mechanobiologists alike.

      reviewer expertise: developmental biologist specialised in morphogenesis

      Reviewer #3 (Evidence, reproducibility and clarity (Required)):

      In the studies conducted by Camacho-Macorra et al., the authors examine the extension of the body axis is zebrafish, focusing on the notochord. They specifically compare timepoints where progenitor addition to the notochord and vacuolization are important to drive axis extension. They generate a simple mathematical model of notochord extension and show that it recapitulates observations in vivo where progenitor addition and vacuolation drive tissue elongation. They further perturb the system by showing that YAP activity is localized to the midline progenitors of the notochord where when the competitive inhibitor of YAP vgll4b is perturbed it increases YAP signaling and results in increase progenitor addition to the notochord. They further describe a possible indirect-feedback mechanism linking YAP driven progenitor addition to the notochord with anterior vacuolation which when perturbed (i.e. increased YAP) results in reduced notochord elongation.

      Major Comments:

      NA

      Minor comments:

      1.Figure 1B - please put the model equation in the figure or at least point out what variables of the equation refer to each part of the schematic.

      As suggested by the reviewer, we have modified the scheme in Figure 1

      2.Figure 1F - smooth line is misleading, please include individual embryo measurement points. This comment could be applied to several figures

      We agree with the reviewer that the graphs in the original manuscript could be improved, and we have therefore modified all figures to better represent data dispersion within each group.

      3.Figure 2C/D - To make this manuscript more accessible to individuals who are not familiar with the anatomy of zebrafish tail, please include zoom in panels of the region of interest where arrows are pointing out increased YAP signaling in the floor plate and hypochord.

      As suggested by the reviewer, we have included higher-magnification insets in Figure 2

      4.In discussion - "In vgll4b mutants, increased progenitor incorporation initially does not alter overall notochord length due to a buffering mechanism for natural variation in progenitor addition" - this is not directly tested in terms of buffering for variation and is an assumption. Please either cite a paper or reword

      This point has been clarified in the revised discussion.

      Reviewer #3 (Significance (Required)):

      Overall, the logic and experiments conducted in these studies are well defined. However, the significance of the work is minimal and makes only a small contribution to the advancement of the field of developmental biology. Regardless, the studies are well done and worth publication.

      Strengths:

      -The study does a good job of incorporating and testing a computational model in a way that proves/disproves their hypothesis

      -The manuscript is well written and follows a logical order, making it easy for readers to understand the main findings

      -The study uses multiple routes of YAP inhibition (genetic and drug) to show effect on progenitor addition to the notochord and shortened body axis

      -The discussion does aa very good job of giving the context of the study's results.

      Weaknesses:

      -The study is minimal and fails to illuminate the mechanism that connects progenitor addition to vacuolization, claiming only an indirect relationship with YAP signaling. However, this is admitted by the authors and not overstated

      The study provides a minimal advancement to the field by investigating an unexplored area of zebrafish notochord extension. It provides a small step toward connecting mechanical/morphogenic mechanisms with signalling in zebrafish body axis extension.

      The audience of this work is a specialized basic research group of developmental biology scientists. The research is of particular relevance to individuals studying zebrafish or axis elongation. While the authors make comparisons to other systems, due to the unique nature of the zebrafish body extension, this generates a narrow field of focus for the manuscript.

      We have previously discussed the uniqueness of zebrafish posterior body elongation in light of critical differences in the degree to which posterior growth from self-renewing tailbud progenitor populations contribute to the mechanisms of axis elongation (Sambasivan and Steventon (2021) Frontiers in cell and dev. Biol; Steventon and Martinez Arias (2017) Developmental Biology). Here too, we think zebrafish provide an important system to explore differences in the mechanisms that drive notochord elongation, and we envisage that this study will provoke a similar cross species comparison that takes into account differences in the relative timing of progenitor addition and anterior notochord expansion (that occurs much later in amniotes, for example). It is only by considering these species-specific differences across experimental organisms that we can arrive at the fundamental principles that drive developmental processes, and how evolution has acted upon these to drive change in adult body plans. We therefore respectfully disagree with the review about the scope and importance of this work for these reasons.

      In addition, we feel that the principles by which dynamic processes are coupled across an organ are broadly applicable and will illuminate further research into understanding organ growth control.

    1. Reviewer #2 (Public review):

      I have completed a thorough review of this paper, which seeks to use the large datasets of species occurrences available through GBIF to estimate variation in how large numbers of plant and animal species are associated with urbanization throughout the world, describing what they call the "species urbanness distribution" or SUD. They explore how these SUDs differ between regions and different taxonomic levels. They then calculate a measure of urban tolerance and seek to explore whether organism size predicts variation in tolerance among species and across regions.

      The study is impressive in many respects. Over the course of several papers, Callaghan and coauthors have been leaders in using "big [biodiversity] data" to create metrics of how species' occurrence data are associated with urban environments, and in describing variation in urban tolerance among taxa and regions. This work has been creative, novel, and it has pushed the boundaries of understanding how urbanization affects a wide diversity of taxa. The current paper takes this to a new level by performing analyses on over 94000 observations from >30,000 species of plants and animals, across more than 370 plant and animal taxonomic families. All of these analyses were focused on answering two main questions:<br /> (1) What is the shape of species' urban tolerance distributions within regional communities?<br /> (2) Does body size consistently correlate with species' urban tolerance across taxonomic groups and biogeographic contexts?

      Overall, I think the questions are interesting and important, the size and scope of the data and analyses are impressive, and this paper has a potentially large contribution to make in pushing forward urban macroecology specifically and urban ecology and evolution more generally.

      Despite my enthusiasm for this paper and its potential impact, there are aspects that could be improved, and I believe the paper requires major revision.

      Some of these revisions ideally involve being clearer about the methodology or arguments being made. In other cases, I think their metrics of urban tolerance are flawed and need to be rethought and recalculated, and some of the conclusions are inaccurate. I hope the authors will address these comments carefully and thoroughly. I recognize that there is no obligation for authors to make revisions. However, revising the paper along the lines of the comments made below would increase the impact of the paper and its clarity to a broad readership.

      Major Comments:

      (1) Subrealms

      Where does the concept of "subrealms" come from? No citation is given, and it could be said that this sounds like an idea straight out of Middle Earth. How do subrealms relate to known bioclimatic designations like Koppen Climate classifications, which would arguably be more appropriate? Or are subrealms more socio-ecologically oriented? From what I can tell, each subrealm lumps together climatically diverse areas. It might be better and more tractable to break things in terms of continents, as the rationale for subrealms is unclear, and it makes the analyses and results more confusing. The authors rationalized the use of subrealms to account for potential intraspecific differences in species' response to urbanization, but that is never a core part of the questions or interpretation in the paper, and averaging across subrealms also accounts for intraspecific variation. Another issue with using the subrealm approach is that the authors only included a species if it had 100 observations in a given subrealm, leading to a focus on only the most common species, which may be biased in their SUD distribution. How many more species would be included if they did their analysis at the continental or global scale, and would this change the shape of SUDs?

      (2) Methods - urban score

      The authors describe their "urban score" as being calculated as "the mean of the distribution of VIIRS values as a relative species-specific measure of a response to urban land cover."

      I don't understand how this is a "relative species-specific measure". What is it relative to? Figures S4 and S5 show the mean distribution of VIIRS for various taxa, and this mean looks to be an absolute measure. Mean VIIRS for a given species would be fine and appropriate as an "urban score", but the authors then state in the next sentence: "this urban score represents the relative ranking of that species to other species in response to urban land cover".

      That doesn't follow from the description of how this is calculated. Something is missing here. Please clarify and add an explicit equation for how the urban score is calculated because the text is unclear and confusing.

      (3) Methods - urban tolerance

      How the authors are defining and calculating tolerance is unclear, confusing, and flawed in my opinion.

      Tolerance is a common concept in ecology, evolution, and physiology, typically defined as the ability for an organism to maintain some measure of performance (e.g., fitness, growth, physiological homeostasis) in the presence versus absence of some stressor. As one example, in the herbivory literature, tolerance is often measured as the absolute or relative difference in fitness of plants that are damaged versus undamaged (e.g., https://academic.oup.com/evolut/article/62/9/2429/6853425?login=true).

      On line 309, after describing the calculation of urban scores across subrealms, they write: "Therefore, a species could be represented across multiple subrealms with differing measures of urban tolerance (Fig. S4). Importantly, this continuous metric of urban tolerance is a relative measure of a species' preference, or affinity, to urban areas: it should be interpreted only within each subrealm".

      This is problematic on several fronts. First, the authors never define what they mean by the term "tolerance". Second, they refer to urban tolerance throughout the paper, but don't describe the calculation until lines 315-319, where they write (text in [ ] is from the reviewer):

      "Within each subrealm, we further accounted for the potential of different levels of urbanization by scaling each species' urban score by subtracting the mean VIIRS of all observations in the subrealm (this value is hereafter referred to as urban tolerance). This 'urban tolerance' (Fig. S5) value can be negative - when species under-occupy urban areas [relative to the average across all species] suggesting they actively avoid them-or positive-when species over-occupy urban areas [relative to the average across all species] suggesting they prefer them (i.e., ranging from urban avoiders to urban exploiters, respectively).<br /> They are taking a relativized urban score and then subtracting the mean VIIRS of all observations across species in a subrealm. How exactly one interprets the magnitude isn't clear and they admit this metric is "not interpretative across subrealms".

      This is not a true measure of tolerance, at least not in the conventional sense of how tolerance is typically defined. The problem is that a species distribution isn't being compared to some metric of urbanness, but instead it is relative to other species' urban scores, where species may, on average, be highly urban or highly nonurban in their distribution, and this may vary from subrealm to subrealm. A measure of urban tolerance should be independent of how other species are responding, and should be interpretable across subrealms, continents, and the globe.

      I propose the authors use one of two metrics of urban tolerance:

      (i) Absolute Urban Tolerance = Mean VIIRS of species_i - Mean VIIRS of city centers<br /> Here, the mean VIIRS of city centers could be taken from the center of multiple cities throughout a subrealm, across a continent, or across the world. Here, the units are in the original VIIRS units where 0 would correspond to species being centered on the most extreme urban habitats, and the most extreme negative values would correspond to species that occupy the most non-urban habitats (i.e., no artificial light at night). In essence, this measure of tolerance would quantify how far a species' distribution is shifted relative to the most highly urbanized habitat available.

      (ii) % Urban Tolerance = (Mean VIIRS of species_i - Mean VIIRS of city centers)/MeanVIIRS of city centers * 100%<br /> This metric provides a % change in species mean VIIRS distribution relative to the most urban habitats. This value could theoretically be negative or positive, but will typically be negative, with -100% being completely non-urban, and 0% being completely urban tolerant.

      Both of these metrics can be compared across the world, as it would provide either absolute (equation 1) or relative (equation 2) metrics of urban tolerance that are comparable and easily interpretable in any region.

      In summary, the definition of tolerance should be clear, the metric should be a true measure of tolerance that is comparable across regions, and an equation should be given.

      (4) Figure 1: The figure does not stand alone. For example, what is the hypothesis for thermophily or the temperature-size rule? The authors should expand the legend slightly to make the hypotheses being illustrated clearer.

      (5) SUDs: I don't agree with the conclusion given on line 83 ("pattern was consistent across subrealms and several taxonomic levels") or in the legend of Figure 2 ("there were consistent patterns for kingdoms, classes, and orders, as shown by generally similar density histograms shapes for each of these").

      The shapes of the curves are quite different, especially for the two Kingdoms and the different classes. I agree they are relatively consistent for the different taxonomic Orders of insects.

      Comments on revised version:

      I believe their response is thorough and thoughtful. I still disagree with them on some fundamental points of their methodology. However, I would prefer to let my review and their response stand as is. This will allow engaged readers to see both sides of the arguments and judge for themselves whether they believe the revisions are sufficient and if my concerns are valid.

    2. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides an important assessment of how body size influences the occurrence of macro-organisms in urban areas across the globe. Size in most plants, but only some animal families, was positively associated with urban tolerance. The data set is impressive, but the evidence for broad-scale conclusions is incomplete due to methodological issues that need to be resolved.

      We have substantially revised the manuscript to resolve the methodological issues raised, including clarifying the definition, calculation, and interpretation of urban affinity (formerly named urban tolerance), and tightening the scope of our conclusions to align directly with the evidence presented.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors integrate multiple large databases to test whether body sizes were positively associated with which species tolerate urban areas. In general, many plant families showed a positive association between body size and urban tolerance, whereas a smaller, though still non-trivial, percentage of animal families showed the same pattern. Notably, the authors are careful in the interpretation of their findings and provide helpful context for the ways that this analysis can be generative in shaping new hypotheses and theory around how urbanization influences biodiversity at large. They are careful to discuss how body size is an important trait, but the absence of a relationship between body size and urban tolerance in many families suggests a variety of other traits undergird urban success.

      We appreciate this thoughtful and balanced assessment of our work and fully agree with the reviewer’s interpretation. In particular, we share the view that the heterogeneous and often weak association between body size and urban affinity across many families is an important result in its own right, underscoring that no single trait is likely to explain urban success across the tree of life. As the reviewer notes, our intention was not to present body size as a universal predictor, but rather as a widely available, integrative trait that can help reveal where general patterns do and do not emerge. We view the lack of a consistent relationship in many families as strong motivation for future work that explicitly integrates additional functional traits and ecological contexts, and we have clarified this perspective in the revised manuscript.

      Strengths:

      The authors aggregated a large dataset, but they also applied robust filters to ensure they had an adequate and representative number of detections for a given species, family, geography, etc. The authors also applied their analysis at multiple taxonomic scales (family and order), which allowed for a better interpretation of the patterns in the data and at what taxonomic scale body size might be important.

      We thank the reviewer for highlighting these strengths of the study. Considerable effort went into assembling, harmonizing, and filtering these data across taxa, regions, and taxonomic resolutions, and we were deliberate in applying conservative thresholds to ensure that species-level urban affinity estimates were based on adequate and comparable sampling. We hope that, beyond the specific results presented here, the compiled dataset and analytical framework will serve as a valuable resource for future studies aiming to explore additional traits, taxa, or mechanisms underlying species’ responses to urbanization.

      Weaknesses:

      My main concern is that it is not fully clear how the measure of body size might influence the result. The authors were unable to obtain consistent measures of body size (mean, median, maximum, or sex variation). This, of course, could be very consequential as means and medians can differ quite a bit, and they certainly will differ substantially from a maximum. And of course, sex differences can be marked in multiple directions or absent altogether. The authors do note that they selected the measure that was most common in a family, but it was not clear whether species in that family that did not have that measure were removed or not. This could potentially shape the variability in the dataset and obscure true patterns. This may require additional clarity from the authors and is also a real constraint in compiling large data from disparate sources.

      We appreciate this important point and agree that heterogeneity in how body size is measured (e.g., mean vs. maximum values, sex-specific measures) is a real but unavoidable challenge when compiling organismal trait data across such a broad taxonomic scope. We would like to clarify that our analytical approach was explicitly designed to minimize the influence of this heterogeneity rather than ignore it. Specifically, for each family we retained all species for which at least one body size estimate was available, rather than removing species that lacked a particular measurement type. When multiple body size measures existed for a species, we selected the measurement type that was most commonly available within that family in order to maximize comparability among species while retaining sample size. Importantly, differences among body size measurement types (including units, measurement detail, and whether values reflected means, maxima, or sex-specific estimates) were further accounted for by (i) log-transforming all body size values and (ii) centering and scaling body size values within each measurement type, which was included as a random effect in the hierarchical models. This approach reduces the influence of systematic differences among measurement types on estimated relationships with urban affinity. We have added a sentence to the methods clarifying that species with a single measurement type were not removed from analyses:

      “Importantly, this procedure did not result in the exclusion of species lacking a particular body size measurement type; rather, all species with at least one available body size estimate were retained, with measurement heterogeneity explicitly accounted for through hierarchical modeling.”

      We agree that variation in body size definitions may still contribute residual noise and potentially obscure weak relationships, and we now emphasize this more clearly as a limitation of large-scale trait syntheses. However, because our primary inference focuses on the presence, absence, and direction of size–urban affinity relationships across families, rather than precise effect sizes, we believe our approach provides a robust and conservative test of whether body size consistently predicts urban affinity across taxa. We highlight this point in the limitations section of our manuscript:

      “One important limitation of our synthesis is the heterogeneity in how body size is measured across taxa, including differences among mean, maximum, and sex-specific estimates. While our analytical framework explicitly accounts for this variation through transformation, scaling, and hierarchical modeling with random intercepts (see Methods), residual measurement noise may still obscure weak size–urban affinity relationships. This challenge is inherent to large-scale trait syntheses that integrate data from disparate sources, and highlights the need for continued efforts to standardize trait databases and expand the availability of harmonized organismal trait data across the tree of life.”

      Reviewer #2 (Public review):

      I have completed a thorough review of this paper, which seeks to use the large datasets of species occurrences available through GBIF to estimate variation in how large numbers of plant and animal species are associated with urbanization throughout the world, describing what they call the "species urbanness distribution" or SUD. They explore how these SUDs differ between regions and different taxonomic levels. They then calculate a measure of urban tolerance and seek to explore whether organism size predicts variation in tolerance among species and across regions.

      The study is impressive in many respects. Over the course of several papers, Callaghan and coauthors have been leaders in using "big [biodiversity] data" to create metrics of how species' occurrence data are associated with urban environments, and in describing variation in urban tolerance among taxa and regions. This work has been creative, novel, and it has pushed the boundaries of understanding how urbanization affects a wide diversity of taxa. The current paper takes this to a new level by performing analyses on over 94000 observations from >30,000 species of plants and animals, across more than 370 plant and animal taxonomic families. All of these analyses were focused on answering two main questions:

      (1) What is the shape of species' urban tolerance distributions within regional communities?

      (2) Does body size consistently correlate with species' urban tolerance across taxonomic groups and biogeographic contexts?

      We thank the reviewer for their careful reading of the manuscript and for this generous and accurate summary of the study’s aims, scope, and contributions. We appreciate the recognition of our group’s broader body of work using large biodiversity databases to quantify species’ associations with urban environments, and we are grateful for the reviewer’s acknowledgement that this study extends those efforts to an unprecedented taxonomic and geographic scale. We agree with the reviewer’s articulation of the two core questions motivating the paper, and we have revised the manuscript to ensure that these questions are stated clearly and addressed consistently throughout.

      Overall, I think the questions are interesting and important, the size and scope of the data and analyses are impressive, and this paper has a potentially large contribution to make in pushing forward urban macroecology specifically and urban ecology and evolution more generally.

      Thanks! We see this work as an effort to move beyond species-by-species descriptions of urban responses toward a community- and distribution-level perspective, where the shape of species’ urban associations themselves becomes an object of study. By framing species’ distributions along an urbanization gradient as a collective property of regional species pools, our approach opens a complementary way of thinking about how urbanization filters biodiversity.

      Despite my enthusiasm for this paper and its potential impact, there are aspects that could be improved, and I believe the paper requires major revision.

      Some of these revisions ideally involve being clearer about the methodology or arguments being made. In other cases, I think their metrics of urban tolerance are flawed and need to be rethought and recalculated, and some of the conclusions are inaccurate. I hope the authors will address these comments carefully and thoroughly. I recognize that there is no obligation for authors to make revisions. However, revising the paper along the lines of the comments made below would increase the impact of the paper and its clarity to a broad readership.

      We appreciate the detailed comments provided and have addressed each point in turn - see detailed responses below. We took these concerns seriously and undertook a substantial revision of the manuscript. In summary, we clarified the conceptual framing of “urban tolerance” (now referred to as “urban affinity”), explicitly defined the metric and its interpretation, added equations and a step-by-step methodological roadmap, and expanded justification for our regional stratification. Where appropriate, we refined language in the Results and Discussion to ensure conclusions are tightly aligned with what the metric can and cannot support. We agree that these revisions materially improve the clarity, rigor, and interpretability of the study, and we appreciate the reviewer’s perspective on how doing so strengthens the paper’s contribution and accessibility to a broad readership.

      Major Comments:

      (1) Subrealms

      Where does the concept of "subrealms" come from? No citation is given, and it could be said that this sounds like an idea straight out of Middle Earth. How do subrealms relate to known bioclimatic designations like Koppen Climate classifications, which would arguably be more appropriate? Or are subrealms more socio-ecologically oriented? From what I can tell, each subrealm lumps together climatically diverse areas. It might be better and more tractable to break things in terms of continents, as the rationale for subrealms is unclear, and it makes the analyses and results more confusing. The authors rationalized the use of subrealms to account for potential intraspecific differences in species' response to urbanization, but that is never a core part of the questions or interpretation in the paper, and averaging across subrealms also accounts for intraspecific variation. Another issue with using the subrealm approach is that the authors only included a species if it had 100 observations in a given subrealm, leading to a focus on only the most common species, which may be biased in their SUD distribution. How many more species would be included if they did their analysis at the continental or global scale, and would this change the shape of SUDs?

      We thank the reviewer for raising this point and agree that the rationale for using subrealms required clearer explanation. Next to allowing potential intraspecific differences in urban affinity across regions, our subrealm-based approach also provides a practical way to partition global biodiversity into ecologically meaningful regional assemblages while maintaining sufficient sample sizes for analysis. Urban affinity is likely to vary geographically within species due to differences in climate, habitat availability, urban form, and evolutionary history. By calculating urban affinity within subrealms rather than globally, our approach allows species to exhibit region-specific urban affinities while ensuring that comparisons are made among species co-occurring within the same regional ecological context. We have substantially revised the Methods to explicitly define subrealms, cite their origin, and clarify why this spatial stratification is appropriate for our study:

      “Accounting for geographic context through subrealm stratification

      To account for geographic heterogeneity in both species’ distributions and the baseline levels of urbanization, we stratified our analyses by global biogeographic subrealms (N=52; Fig. S1). Subrealms represent an intermediate hierarchical level within the One Earth [82] (https://www.oneearth.org/bioregions/) bioregionalization framework, grouping the 185 terrestrial bioregions into broader units that reflect shared species pools and ecological contexts while maintaining meaningful regional structure. This scale represents a practical compromise between analyzing data at the finer bioregion level (which would result in many regions with insufficient observations for robust analysis) and broader classifications such as continents or the 14 biogeographic realms, which aggregate ecologically distinct regions and species pools. This regionalization has been widely used in macroecological and biogeographic research to contextualize species–environment relationships because subrealms capture meaningful gradients in biotic assemblages that are not accounted for by climatic classifications alone [83,84].

      This stratification allows species’ associations with urban environments to be interpreted relative to the environments available within the regions they occupy. This is important, as previous work has shown that species’ responses to urbanization are constrained by biogeographic context, because regional species pools reflect shared evolutionary, ecological, and historical filters [23]. Previous work has also shown that urban associations among species are context-dependent, and interpreting species’ responses without accounting for regional baselines conflates availability of urban environments with species’ affinity to them. This distinction is critical because identical levels of urbanization (e.g., VIIRS radiance) can have different ecological meanings across regions with different species pools and land-use histories. It avoids conflating species’ urban affinity with global differences in urban availability.”

      We chose subrealms rather than Köppen climate classifications or continental units because our objective was not to partition species by climatic similarity per se, but to evaluate species’ associations with urban environments relative to the ecological and biogeographic contexts in which they occur. Climatic classifications such as Köppen are highly effective for addressing climate–species relationships, but they do not explicitly capture differences in species pools, evolutionary history, or land-use legacies that strongly shape how species interact with urbanization. Likewise, continents often aggregate ecologically disparate regions and species pools, potentially obscuring meaningful variation in baseline urbanization and species’ realized distributions.

      Importantly, urban affinity in our framework is a relative, context-dependent metric, explicitly interpreted within regions. Identical levels of urbanization (e.g., VIIRS radiance values) can have different ecological meanings across regions with distinct species pools, land-use histories, and settlement patterns. Stratifying analyses by subrealm therefore avoids conflating species’ affinity to urban environments with global or continental differences in the availability and intensity of urban land cover. We have clarified this distinction and motivation in the revised Methods (see responses below).

      Regarding the concern that requiring ≥100 observations per species per subrealm biases analyses toward common species: we agree that this threshold focuses the analysis on well-sampled species. This choice was intentional and follows previous work showing that such cutoffs are necessary to robustly characterize species’ responses to urbanization using occurrence data. While a global or continental analysis would indeed include additional, rarer species, it would also substantially increase uncertainty and conflate species’ responses across ecologically distinct contexts. Our study is therefore best interpreted as a macroecological synthesis of common species, which are also the taxa that disproportionately structure urban communities and drive the shape of Species Urbanness Distributions (SUDs). We now clarify this scope and limitation more explicitly in the introduction:

      “Our aim is to identify broad, cross-taxonomic patterns in species’ urban affinity at a global scale, rather than to resolve the specific causal mechanisms driving urban success or failure within individual taxa or cities.”.

      As well as in the discussion:

      “Our synthesis complements taxon-specific, presence–absence trait studies by identifying broad, cross-taxonomic patterns that can motivate and contextualize more mechanistic analyses [17,23].”

      Finally, while alternative spatial stratifications are possible, the central patterns we report particularly the skewed shape of SUDs—are robust to the use of regional context rather than absolute global metrics. Exploring how SUDs change under different spatial frameworks (e.g., continents, climate zones) is an interesting avenue for future work, but we feel is beyond the scope of the present study.

      (2) Methods - urban score

      The authors describe their "urban score" as being calculated as "the mean of the distribution of VIIRS values as a relative species specific measure of a response to urban land cover."

      I don't understand how this is a "relative species-specific measure". What is it relative to? Figures S4 and S5 show the mean distribution of VIIRS for various taxa, and this mean looks to be an absolute measure. Mean VIIRS for a given species would be fine and appropriate as an "urban score", but the authors then state in the next sentence: "this urban score represents the relative ranking of that species to other species in response to urban land cover".

      We agree that the wording in the original manuscript was unclear and conflated two distinct steps in the workflow. We have now revised the Methods to clearly distinguish between (i) the urban score, which is an absolute, descriptive summary of the mean VIIRS radiance associated with a species’ occurrence locations, and (ii) urban affinity, which is the relative, region-specific metric derived from the urban score. Specifically, we rewrote the methods to have distinct steps as subheadings, as follows: (1) urban score; (2) subrealms and why; (3) urban affinity. In the revised Methods, we explicitly define the urban score:

      “an absolute descriptive summary of the urbanization levels associated with a species’ occurrence locations within a given subrealm”.

      We no longer describe the urban score itself as “relative” or as a ranking among species. Relative comparisons among species arise only in the subsequent step, where species-specific urban scores are expressed relative to the regional background level of urbanization within each subrealm to derive urban affinity.

      We refer the Reviewer to the revised version which we feel is much clearer (lines 428-479)!

      That doesn't follow from the description of how this is calculated. Something is missing here. Please clarify and add an explicit equation for how the urban score is calculated because the text is unclear and confusing.

      The previous response, where we discuss the description, hopefully clarifies this. Further, we have revised the Methods to clearly define the urban score and to include an explicit equation. In the revised manuscript, the urban score for species s is calculated as the mean VIIRS radiance across all occurrence locations of that species:

      where n<sub>s</sub>is the number of GBIF occurrence records for species s, and L<sub>i</sub> is the VIIRS nighttime lights radiance value extracted at the location of occurrence i. We also clarify in the Methods that this urban score is an absolute summary statistic of observed urbanization at species occurrence locations

      (3) Methods - urban tolerance

      How the authors are defining and calculating tolerance is unclear, confusing, and flawed in my opinion.

      Tolerance is a common concept in ecology, evolution, and physiology, typically defined as the ability for an organism to maintain some measure of performance (e.g., fitness, growth, physiological homeostasis) in the presence versus absence of some stressor. As one example, in the herbivory literature, tolerance is often measured as the absolute or relative difference in fitness of plants that are damaged versus undamaged

      (e.g., https://academic.oup.com/evolut/article/62/9/2429/6853425?login=true).

      On line 309, after describing the calculation of urban scores across subrealms, they write: "Therefore, a species could be represented across multiple subrealms with differing measures of urban tolerance (Fig. S4). Importantly, this continuous metric of urban tolerance is a relative measure of a species' preference, or affinity, to urban areas: it should be interpreted only within each subrealm". This is problematic on several fronts. First, the authors never define what they mean by the term "tolerance". Second, they refer to urban tolerance throughout the paper, but don't describe the calculation until, where they write (text in [ ] is from the reviewer): "Within each subrealm, we further accounted for the potential of different levels of urbanization by scaling each species' urban score by subtracting the mean VIIRS of all observations in the subrealm (this value is hereafter referred to as urban tolerance). This 'urban tolerance' (Fig. S5) value can be negative - when species under-occupy urban areas [relative to the average across all species] suggesting they actively avoid them-or positive-when species over-occupy urban areas [relative to the average across all species] suggesting they prefer them (i.e., ranging from urban avoiders to urban exploiters, respectively). They are taking a relativized urban score and then subtracting the mean VIIRS of all observations across species in a subrealm. How exactly one interprets the magnitude isn't clear and they admit this metric is "not interpretative across subrealms".

      This is not a true measure of tolerance, at least not in the conventional sense of how tolerance is typically defined. The problem is that a species distribution isn't being compared to some metric of urbanness, but instead it is relative to other species' urban scores, where species may, on average, be highly urban or highly nonurban in their distribution, and this may vary from subrealm to subrealm. A measure of urban tolerance should be independent of how other species are responding, and should be interpretable across subrealms, continents, and the globe.

      We thank the reviewer for this careful and important critique. We agree that the term “tolerance” is commonly used to describe the ability of an organism to maintain performance (e.g., fitness, growth, physiological homeostasis) in the presence of a stressor, and that our metric does not measure tolerance in this mechanistic or fitness-based sense. To address this directly and unambiguously, we have revised the manuscript to explicitly define the term “urban affinity” as opposed to urban tolerance. 

      In the revised Methods, we also reorganized and clarified the calculation of urban affinity, introduced explicit notation, and provided a formal equation. Specifically, we now define urban affinity for species s in subrealm r as:

      where U<sub>s,r</sub>is the mean VIIRS radiance across all occurrence locations of species s within subrealm r, and Ū<sub>r</sub>is the mean VIIRS radiance across all occurrence records of all species in that subrealm. This transformation centers species’ urban scores on the regional background level of urbanization, yielding a relative measure of spatial association with urban environments.

      We agree with the reviewer that this metric is not interpretable as an absolute measure of affinity, and we now state this explicitly. Urban affinity values are, by construction, relative measures, interpretable only within subrealms, and they quantify whether a species tends to occur in more or less urbanized environments than is typical for that region. The magnitude of the metric therefore reflects deviation from the regional baseline, not a universal or global scale of urbanization, and is not intended to be compared directly across subrealms.

      We respectfully disagree, however, that this makes the metric flawed. Rather, it reflects a deliberate analytical choice aligned with our research questions. Our goal was not to estimate absolute urban exposure or physiological performance, but to compare species’ realized spatial associations with urban environments within shared biogeographic contexts. Because baseline urbanization levels, settlement history, and species pools vary strongly across regions, a globally absolute metric would conflate species’ affinities with regional availability of urban environments. By contrast, a relative, region-centered metric allows meaningful comparisons among species that coexist within the same ecological and biogeographic setting. This approach follows a growing body of macroecological work that infers species’ environmental affinities from spatial distributions rather than direct performance measures (e.g., Callaghan et al. 2020; 2021; 2023), and we now cite these studies explicitly.

      I propose the authors use one of two metrics of urban tolerance:

      (i) Absolute Urban Tolerance = Mean VIIRS of species_i - Mean VIIRS of city centers Here, the mean VIIRS of city centers could be taken from the center of multiple cities throughout a subrealm, across a continent, or across the world. Here, the units are in the original VIIRS units where 0 would correspond to species being centered on the most extreme urban habitats, and the most extreme negative values would correspond to species that occupy the most non-urban habitats (i.e., no artificial light at night). In essence, this measure of tolerance would quantify how far a species' distribution is shifted relative to the most highly urbanized habitat available.

      (ii) % Urban Tolerance = (Mean VIIRS of species_i - Mean VIIRS of city centers)/MeanVIIRS of city centers * 100%

      This metric provides a % change in species mean VIIRS distribution relative to the most urban habitats. This value could theoretically be negative or positive, but will typically be negative, with -100% being completely non-urban, and 0% being completely urban tolerant.

      Both of these metrics can be compared across the world, as it would provide either absolute (equation 1) or relative (equation 2) metrics of urban tolerance that are comparable and easily interpretable in any region.

      In summary, the definition of tolerance should be clear, the metric should be a true measure of tolerance that is comparable across regions, and an equation should be given.

      We thank the reviewer for this thoughtful and constructive suggestion, which raises an important conceptual issue regarding how “urban tolerance” should be defined and quantified. We agree that any such metric must be clearly defined, interpretable, and accompanied by an explicit equation, and we have revised the manuscript accordingly to clarify both our definition and its intended interpretation.

      The alternative metrics proposed by the reviewer anchoring species’ distributions to city centers or to the most highly urbanized habitats represent a valid and intuitive absolute framing of urban tolerance. Indeed, a closely related approach was explored and evaluated in Callaghan et al. (2020; https://doi.org/10.1016/j.ecolind.2020.106905), where species’ occurrence-based urbanness scores derived from VIIRS night-time lights were compared against abundance-based estimates of urban tolerance using explicit urban–non-urban contrasts. That study further demonstrated that urbanness scores depend on the choice of spatial baseline (e.g., regional buffers around cities versus continental extents), and showed that different baselines capture complementary, but not identical, aspects of species–urban associations.

      In the present study, we deliberately adopt a relative, regionally contextualized metric (now referred to as urban affinity), expressing each species’ mean VIIRS association relative to the background urbanization of the biogeographic subrealm in which it occurs. This choice reflects our goal of comparing species’ relative affinities to urban environments within shared ecological and biogeographic contexts. Importantly, identical VIIRS values can correspond to very different ecological conditions across regions, and anchoring all species to city centers or global urban maxima risks conflating species’ affinities with regional differences in urban availability and infrastructure.

      We now make this distinction explicit throughout the manuscript, including by (i) defining urban affinity as a relative, occurrence-based measure of urban affinity (rather than physiological or fitness-based tolerance), (ii) providing an explicit equation for its calculation, and (iii) clarifying that these values are interpretable within, but not across, biogeographic subrealms. We view absolute, city-center–anchored metrics and relative, regionally normalized metrics as complementary approaches, each suited to different questions; the latter is most appropriate for the macroecological, comparative analyses pursued here.

      (4) Figure 1: The figure does not stand alone. For example, what is the hypothesis for thermophily or the temperature-size rule? The authors should expand the legend slightly to make the hypotheses being illustrated clearer.

      We now expanded the legend so that the figure and hypotheses presented can be understood based on just the figure and its legend; we did so by explaining the illustrated hypotheses as requested by the Reviewer. The figure legend now reads as follows:

      “Fig. 1: Conceptual framework illustrating hypothesized mechanisms linking urban affinity to interspecific body-size shifts. These include dispersal and mobility constraints under habitat fragmentation [44,45], thermophily and the temperature–size rule driven by the urban heat island effect [15,30], size-biased competition and survival [94,95], and size-biased human preferences [64]. Urban fragmentation of habitat resources can select for increased mobility (e.g., larger butterflies) or reduced mobility (e.g., larger seeds) depending on isolation severity. Elevated urban temperatures favor thermophily, which often negatively correlates with size as it affects the heat balance via thermal inertia. Similarly, these higher temperatures generally favor smaller-bodied adult ectotherms because they accelerate development and reduce time available for growth (i.e., temperature-size rule). In plants, the increased CO<sub>₂</sub> and nutrient availability associated with anthropogenic environments due to heating- and traffic-related CO2 emissions and eutrophication provides a competitive advantage to larger plant species, and human preferences too may favor larger species (e.g., tree-lined streets), whereas smaller species may be advantaged in colonizing built infrastructure.”

      (5) SUDs: I don't agree with the conclusion given on line 83 ("pattern was consistent across subrealms and several taxonomic levels") or in the legend of Figure 2 ("there were consistent patterns for kingdoms, classes, and orders, as shown by generally similar density histograms shapes for each of these").

      The shapes of the curves are quite different, especially for the two Kingdoms and the different classes. I agree they are relatively consistent for the different taxonomic Orders of insects.

      We agree that our original wording overstated the similarity of distributions across taxa and regions. We have revised the text to clarify that the consistency we refer to pertains primarily to central tendencies rather than identical distributional shapes. To address this directly, we conducted additional analyses comparing urban affinity distributions across subrealms for taxonomic groups with the largest sample sizes. These results, now presented in new Supplementary Figures (Fig. S2-S4), show that while distributional shapes vary among higher taxonomic groups, median values and overall spread are broadly similar within comparable taxonomic levels. We have updated the Results text and the Figure 2 legend accordingly to reflect this more precise interpretation. 

      “These patterns in central tendency were broadly consistent across subrealms and taxonomic levels, although distributional shapes varied among higher taxonomic groups (Fig. 2).”

      “To evaluate this more formally, we compared distributions across subrealms for groups with the largest sample sizes and found that while distributional shapes varied among higher taxa, median values and overall spread were broadly similar within comparable taxonomic levels (Fig. S2–S4).”

      Figure 2 caption: “There were consistent patterns for kingdoms, classes, and orders (B) as shown by similar central tendencies despite variation in distributional shape.”

      We refer the Reviewer to the revised manuscript and supplementary material, but show the kindom level in Fig S2.

      More broadly, our goal in introducing Species Urbanness Distributions (SUDs) is not to argue that their exact shapes are invariant, but rather to provide a generalizable framework for describing how assemblages are structured along an urbanization gradient. In this respect, SUDs are conceptually analogous to Species Abundance Distributions (SADs), where the precise functional form has long been debated, yet the framework itself has proven extremely valuable for ecology. We therefore emphasize the utility of SUDs as a descriptive and comparative tool for quantifying community-level responses to urbanization, rather than as a claim about strict uniformity in distributional shape across taxa or regions.

      Reviewer #3 (Public review):

      Summary:

      This paper reports on an association between body size and the occurrence of species in cities, which is quantified using an 'urban score' that can be visualized as a 'Species Urbanness

      Distribution' for particular taxa. The authors use species records from the Global Biodiversity Information Facility (GBIF) and link the occurrence data to nighttime lighting quantified using satellite data (Visible Infrared Imaging Radiometer Suite-VIIRS). They link the urban score to body size data to find 'heterogeneous relationship between body size and urban tolerance across the tree'. The results are then discussed with reference to potential mechanisms that could possibly produce the observed effects (cf. Figure 1).

      We thank the reviewer for this clear and accurate summary of the study. We agree that the primary contribution of this work lies in the scale and taxonomic breadth of the analysis, and in introducing a framework (Species Urbanness Distributions) for quantifying species’ relative affinities to urban environments using globally available data. We have revised the manuscript to further clarify the scope of inference and the distinction between descriptive macroecological patterns and mechanistic explanations.

      Strengths:

      The novelty of this study lies in the huge number of species analyzed and the comparison of results among animal taxa, rather than in a thorough analysis of what traits allow species to persist under urban conditions. Such analyses have been done using a much more thorough approach that employs presence-absence data as well as a suite of traits by other studies, for example, in (Hahs et al. 2023, Neate-Clegg et al. 2023). The dataset that the authors produced would also be very valuable if these raw data were published, both the cleaned species records as well as the body sizes. The paper could strongly add to our understanding of what species occur in cities when the open questions are addressed.

      We appreciate highlighting the novelty of the taxonomic breadth and scale of our analysis. We agree that our approach is complementary to more detailed, taxon-specific trait studies based on presence–absence data. In response, we have further emphasized this distinction in the Discussion:

      “Our synthesis complements taxon-specific, presence–absence trait studies by identifying broad, cross-taxonomic patterns that can motivate and contextualize more mechanistic analyses17,23.”

      We also agree that the cleaned occurrence data and body size information represent a valuable resource, and all data will be made available, with the exception of some body size datasets which we are not able to make available.

      Weaknesses:

      I value the approach of the authors, but I think the paper needs to be revised.

      In my view, the authors could more carefully validate their approach. Currently, any weakness or biases in the approach are quickly explained away rather than carefully explored. This concerns particularly the use of presence-only data, but also the calculation of the urban score.

      The vast majority of data in GBIF is presence-only data. This produces a strong bias in the analysis presented in the paper. For some taxa, it is likely that occurrences within the city are overrepresented, and for other taxa, the opposite is true (cf. Sweet et al. 2022). I think the authors should try to address this problem.

      We thank the reviewer for raising this important point. We fully agree that GBIF occurrence data are subject to well-known sampling biases, including uneven geographic coverage, observer effort, and taxonomic focus. These limitations are now more explicitly acknowledged in the revised manuscript. At the same time, GBIF currently represents the only global biodiversity database that allows the scope of analysis undertaken here, spanning thousands of species across multiple taxonomic groups and regions. Systematic monitoring datasets that provide presence–absence data are typically restricted to particular taxa (often vertebrates or plants) and are geographically concentrated in the Global North, which would substantially limit the taxonomic and geographic breadth of our analysis.

      Importantly, our objective was not to estimate absolute species-specific responses to urbanization, but rather to examine relative patterns of urban affinity across species and families within comparable regional contexts. To address this, we structured our analyses at the subrealm level, which aggregates observations across large spatial extents and reduces sensitivity to fine-scale sampling biases associated with individual cities or urban–rural gradients. In addition, we restricted analyses to species with ≥100 observations per subrealm to focus on well-sampled taxa and reduce the influence of extremely sparse occurrence records. While these steps cannot fully eliminate sampling biases inherent to occurrence data, they substantially mitigate their influence when examining broad comparative patterns.

      Recent work has also evaluated the performance of GBIF data in urban biodiversity contexts. For example, Sweet et al. (2022) compared GBIF-derived species richness patterns with independent state-level biodiversity databases across cities and surrounding regions, finding that GBIF provided comparable or broader coverage across taxa and spatial extents. Their analysis showed that species richness was consistently higher in the surrounding region than in the city itself, suggesting that GBIF data capture broad urban–regional biodiversity gradients rather than systematically overrepresenting urban occurrences. Although our analysis differs in design, these results support the use of GBIF as a valuable resource for examining large-scale biodiversity patterns.

      More broadly, occurrence databases such as GBIF have become widely used for analyzing species–environment relationships at macroecological scales. While they may be insufficient for estimating precise species-specific environmental tolerances, they are informative for identifying broad patterns across taxa and regions. Our goal here is therefore to identify large-scale comparative patterns in urban affinity and generate hypotheses about trait– urbanization relationships, which can subsequently be tested with more structured monitoring datasets where available.

      Another important consideration is that our analyses focus on comparative differences among species within shared taxonomic and geographic contexts, rather than absolute estimates of urban affinity. Sampling biases in occurrence databases are often structured by observer behaviour (e.g., detectability, accessibility, or taxonomic interest), meaning that species recorded by similar observer communities are likely subject to similar sampling biases. Under these conditions, relative differences among species are expected to be preserved even when absolute occurrence frequencies are biased. This logic is consistent with the widely used target-group background approach in presence-only species distribution modelling, where species recorded by similar observer groups (often within the same taxonomic group) are used to control for shared sampling bias. Previous work by Callaghan et al. (2021; https://doi.org/10.1111/gcb.15670) performed additional validation analysis comparing our distribution-based urban affinity metric with estimates derived from occupancy modelling using well-sampled European butterflies (see Fig. S5 from the Callaghan et al. 2021 paper). The strong positive relationship between these approaches suggests that the broad patterns identified here are unlikely to arise solely from sampling artifacts.

      Finally, in the revised manuscript we now include additional comparisons among well-sampled taxonomic groups (see responses to other comments throughout our response document for details), which show substantial variation in urban affinity even among taxa with extensive sampling. These results suggest that the patterns reported here are unlikely to arise solely from sampling artifacts, but instead reflect meaningful ecological variation in how species interact with urban environments.

      The authors should compare their results to studies focusing on particular taxa where extensive trait-based analyses have already been performed, i.e., plants and birds. In fact, I strongly suggest that the authors should compare their results to previous studies on the relationship between traits, including body size and occurrences along a gradient of urbanisation, to draw conclusions about the validity of the approach used in the current study, which has a number of weaknesses.

      We agree that explicitly situating our findings within the existing trait-based urban ecology literature strengthens both interpretation and validation of our approach. We had already referenced several relevant studies (e.g., Hahs et al. 2023 and others) in the Introduction and Discussion, but we recognize that these comparisons were not sufficiently explicit. We have now added text to the Discussion directly comparing our results with previous trait-based studies across taxa:

      “Our results are broadly consistent with prior taxon-specific trait-based studies (eg., Hahs et al.[17]), but also highlight that relationships between body size and urbanization vary across taxa and analytical frameworks. For example, global syntheses and regional studies have reported positive, negative, or null size–urbanization relationships depending on clade and spatial scale. A recent global analysis that compiled empirical occurrence data for multiple terrestrial faunal taxa across cities worldwide reported broadly similar body-size responses to urbanization [17]. For four of the five groups that overlap with our analysis—amphibians, bats, bees, and birds—the direction of the body-size relationship with urbanization was consistent between studies. The only exception was carabid beetles, which tended to be smaller-bodied in highly urbanized environments in that analysis, whereas we detected no significant size effect for this family. Studies on birds, for example, have found mixed results, including positive associations to urbanization in some regional assemblages [45], no global relationship in others [46] or an overall negative relationship globally [23], and negative relationships in particular clades such as raptors [40]. Such discrepancies likely arise because different studies quantify urbanization differently, focus on different spatial grains, or analyze different components of species responses (e.g., presence– absence, abundance, or occurrence distributions). Additionally, a study on multiple taxa including butterflies and moths found a positive relationship in butterfly and moth community-weighed mean body size with increases in urbanization level, similar to our findings [31]. Researchers have also found that smaller-bodied dung-associated beetles potentially benefit from urban environments, which is similar to the negative association we found between urbanization and body size in beetles [47]. Our approach complements these studies by estimating occurrence-based urban associations across thousands of taxa simultaneously, allowing comparison of how consistently body size predicts urban affinity across taxonomic groupings rather than within a single lineage. In this sense, variation among published results does not contradict our findings but instead reinforces the conclusion that body size is a context-dependent filter whose direction and strength depend on ecological setting, taxonomic scope, and the urbanization metric used.”

      These additions highlight that published relationships between body size and urbanization vary widely across taxa, spatial scales, and analytical approaches. For example, prior studies have reported positive, negative, or null size–urbanization relationships depending on clade, geographic extent, and how urbanization or occurrence is quantified. Even within birds alone, the literature spans positive regional relationships, null global relationships, and negative relationships in particular clades such as raptors. We now explicitly discuss these contrasts and clarify that such discrepancies are expected because different studies measure different components of species’ responses (e.g., presence–absence vs. abundance vs. occurrence distributions), use different spatial grains, or focus on different taxonomic subsets.

      We emphasize that our analysis is not intended to replace taxon-specific trait studies, but rather to complement them by providing a macroecological synthesis across thousands of species simultaneously. Importantly, the heterogeneity we observe among families is itself a key biological result, indicating that body size is not a universal predictor of urban affinity but instead a context-dependent filter whose direction and strength vary across ecological and phylogenetic settings. We now state this interpretation more clearly in the revised manuscript.

      They should be be more careful in coming up with post-hoc explanations of why the pattern found in this study makes sense or suggests a particular mechanism. This reviewer considers that there is no way in which the current study can disentangle the different possible mechanisms without further analyses and data, so I would suggest pointing out carefully how the mechanisms could be studied.

      We agree that our study cannot disentangle the causal mechanisms underlying species’ responses to urbanization. Our intent in discussing potential mechanisms was not to claim definitive explanations, but rather to situate our findings within existing ecological theory and to highlight plausible, non-exclusive pathways that may generate the observed patterns. To make this clearer, we have revised the Discussion to explicitly frame these interpretations as hypotheses rather than conclusions, and to emphasize that testing the underlying mechanisms will require additional data and approaches, such as targeted trait datasets, experimental manipulations, and longitudinal or within-city studies:

      “Because our synthesis is correlative and macroecological in nature, the mechanisms discussed above are best viewed as hypotheses that can be evaluated through future work combining experimental, trait-based, and longitudinal data.”.

      Additionally, we modified our overall goal to make it clear that this is not inherently a mechanistic study per se:

      “Our aim is to identify broad, cross-taxonomic patterns in species’ urban affinity at a global scale, rather than to resolve the specific causal mechanisms driving urban success or failure within individual taxa or cities.”.

      More details should be given about the methodology. The readers should be able to understand the methods without having to read a number of other papers.

      We have substantially revised and expanded the Methods section to ensure that all analytical steps can be understood directly from the manuscript without requiring consultation of prior publications. In particular, we now (i) provide a clear conceptual roadmap of the workflow at the start of the Methods, (ii) define all key metrics explicitly, including equations for both the urban score and urban affinity, and (iii) clarify the interpretation, assumptions, and limitations of each step. We also added text explaining the rationale for subrealm stratification and the intended interpretation of relative values. Together, these revisions make the methodological framework fully transparent and self-contained (see revised Methods and related responses above and below).

      References:

      Hahs, A. K., B. Fournier, M. F. Aronson, C. H. Nilon, A. Herrera-Montes, A. B. Salisbury, C. G. Threlfall, C. C. Rega-Brodsky, C. A. Lepczyk, and F. A. La Sorte. 2023. Urbanisation generates multiple trait syndromes for terrestrial animal taxa worldwide. Nature Communications 14:4751.

      Neate-Clegg, M. H. C., B. A. Tonelli, C. Youngflesh, J. X. Wu, G. A. Montgomery, Ç. H. Şekercioğlu, and M. W. Tingley. 2023. Traits shaping urban tolerance in birds differ around the world. Current Biology 33:1677-1688.

      Sweet, F. S. T., B. Apfelbeck, M. Hanusch, C. Garland Monteagudo, and W. W. Weisser. 2022. Data from public and governmental databases show that a large proportion of the regional animal species pool occur in cities in Germany. Journal of Urban Ecology 8:juac002.

      We have incorporated these (and additional new references) into our revised manuscript.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you see from the general comments above and the specific recommendations below, the reviewers are impressed by your comprehensive data set and the analytic approach. However, they ask you to clarify your measures of organism size, occurrence data (vs. presence/absence and corresponding sample-bias caveats), urbanness (lighting differences between cities and regions?), urban tolerance (measure should not be relative to other species and particular regions), and region ("subrealm" vs. more commonly used defintions of world regions such as continents). They also encourage you to compare your general results with more detailed local studies to better justify using size as the only, easily available trait.

      We thank the Editor for this clear synthesis of the key priorities for revision. We have carefully addressed each point and substantially revised the manuscript to improve clarity, methodological transparency, and interpretability. In particular:

      We clarified how body size data were compiled, harmonized, and modeled, including explicit description of how different measurement types (mean, maximum, sex-specific) were retained and statistically accounted for through scaling and hierarchical modeling. We now state these procedures explicitly in the Methods.

      We expanded the Methods and Discussion to clarify that our analyses rely on occurrence data rather than presence–absence or abundance data, and we now explicitly discuss the implications and limitations of presence-only datasets, including potential sampling biases and how these may influence inference.

      We strengthened justification for using VIIRS night-time lights as a continuous proxy for urbanization, added supporting citations, and clarified that spatial heterogeneity in lighting primarily introduces additional variance rather than systematic bias. We also explicitly describe how urbanization values were calculated and interpreted.

      We substantially revised the manuscript to clearly define urban affinity at the outset (including in the Abstract), distinguish it from physiological definitions of tolerance, and provide explicit equations and step-by-step descriptions of how both urban score and urban affinity are calculated and interpreted. We now emphasize that the metric is a relative, region-contextualized measure of occurrence-based urban affinity.

      We added full justification, citations, and methodological explanation for the use of biogeographic subrealms, clarified how they differ from continents or climate zones, and explained why this stratification is appropriate for the ecological questions addressed. We also clarified the scope of inference and limitations of this approach.

      We expanded the Discussion to explicitly compare our results with prior trait-based urban ecology studies across taxa (including birds and other groups), highlighting where results converge, diverge, and why such variation is expected across spatial scales, taxa, and analytical frameworks.

      Reviewer #1 (Recommendations for authors):

      (1) Abstract

      (a) Please define how tolerance is being used here

      We now use affinity throughout and it is defined in various places (see responses to other comments here).

      (b) The abstract should clarify at what taxonomic scale body size is assessed. It is unclear in the abstract as to whether the reader expects intraspecific measures and interspecific, and at what resolution.

      We have revised the abstract by adding one sentence explicitly stating the scale body size was assessed:

      “We then assessed whether body size, an integrative ecological trait fundamental to space use, mobility, metabolism, and environmental sensitivity, showed consistent associations with urban affinity among species and across 371 taxonomic families. Analyses were conducted at the interspecific level and focused primarily on variation among taxonomic families (provided with this paper is an accompanying application to view results).”

      (2) Results/Discussion

      (a) The species urbanness distribution and comparison with the species abundance distribution is an interesting and conceptually useful contribution to urban ecology and underscores how urbanization functions on biodiversity at scale.

      We thank the reviewer for this positive assessment and are encouraged that they view the Species Urbanness Distribution (SUD) as a conceptually useful contribution to urban ecology. We see SUDs as a flexible framework that can be extended in several important directions, including comparisons across additional traits, cities of differing size and configuration, and temporal analyses that track how urbanness distributions shift with ongoing urban expansion or restoration. More broadly, we hope that SUDs can provide a framework to think about a macroecological understanding of how urbanization filters biodiversity.

      (b) In our Lambert et al. (2023) study that you reference, we suggest that 'exaptation' may be valuable to explore in urban areas. Although body size wasn't the trait we were considering at that time, it may be worth putting your discussion around pre-adaptation in this context.

      We agree that exaptation provides a valuable conceptual lens for interpreting species’ responses to urban environments. We have revised the Discussion to explicitly frame species’ urban success in this context:

      “Such traits “pre-adapted” to urban conditions allow for some species to not only persist but thrive in urban environments where most species cannot. Framing these patterns through the lens of exaptation may be particularly useful, as traits that evolved under non-urban selective pressures may incidentally confer advantages in urban environments without having arisen in response to urbanization per se (sensu Lambert et al.[4]). We therefore speculate that the skewed shape of SUDs may reflect the uneven distribution of exaptive traits across species pools, rather than widespread adaptive evolution to urban conditions. 

      Consistent with this interpretation, if exaptive traits that facilitate urban persistence are unevenly distributed across species pools, most species would be expected to exhibit avoidance rather than affinity of urban environments. Indeed, we found that the median urban affinity is most often below one, indicating widespread avoidance among species.”.

      (c) Given the family-scale effect, it would be helpful to discuss how often species within a family co-occur in a given geographic region, how much other traits covary with size, etc. Do we have an a priori reason to expect family to be the taxonomic resolution at which body size seems to be most varied?

      Our exploratory and preliminary analyses revealed that variation in the body size– urban affinity relationship was strongest at the family level, which prompted us to focus our main analyses at this taxonomic resolution. (But we also present results on order as well). Families represent a biologically meaningful intermediate scale in taxonomy: species within families typically share broad morphological, ecological, and life-history characteristics, yet still exhibit substantial variation in body size and ecological strategies. Indeed, body size is well known to covary with multiple traits—including dispersal ability, metabolism, and space use—making it an integrative trait that captures several ecological dimensions simultaneously within and among families. These correlated traits likely contribute to the heterogeneous responses to urbanization observed among families.

      Using the family level also provides a practical balance between biological relevance and statistical robustness. Many families contain sufficient numbers of species to allow independent model estimation while avoiding the strong data imbalance that would arise at higher taxonomic levels. In addition, family is a commonly used unit in macroecological trait analyses (e.g., Roy et al. 2009; Smith et al. 2004), and it often reflects major morphological and ecological similarities among species, as reflected in taxonomic identification frameworks.

      Regarding co-occurrence, our analytical framework already accounts for geographic context by estimating urban affinity within subrealms. This ensures that species are compared within the same regional species pools and environmental contexts, rather than across globally disparate assemblages. Consequently, family-level effects emerge from comparisons among species that co-occur within shared biogeographic settings rather than from global taxonomic aggregation.

      We have added a short clarification in the manuscript to emphasize that body size functions as an integrative trait that covaries with multiple ecological attributes, and that family-level analyses represent a balance between ecological interpretability and data availability:

      “Because body size covaries with multiple ecological traits (e.g., dispersal ability and metabolic rate), we focused on family-level analyses to capture shared ecological strategies while still allowing sufficient variation among species to detect trait– environment relationships [39]”.

      (d) The result that body size shows a stronger effect in plants perhaps could suggest that plant records in GBIF are more sensitive to potential collection bias, perhaps due to detectability differences or preferences for where botanists and citizen scientists collect plant data? You mention ornamental plants late, but it may be worth discussing this here, too.

      We agree that this is a possible mechanism, which likely conflates detectability and ecological signal. We have expanded this point in the discusssion to better address this:

      “These human-driven preferences may also influence detectability and recording effort, as larger and more conspicuous plant species are more likely to be planted, maintained, and documented in urban environments, and thus be available in GBIF for our analyses. However, we suggest that this is not purely a sampling artifact, but such processes likely interact with ecological filtering to shape the realized size structure of urban plant communities.”.

      (e) I appreciate the additional taxonomic layering to the discussion. Seeing patterns at the family and order levels is helpful for generating new theory and predictions about how urbanization structures biodiversity at different taxonomic scales.

      We agree that examining patterns across multiple taxonomic scales is particularly valuable for generating testable hypotheses about how urbanization structures biodiversity, as different mechanisms may emerge or break down depending on the resolution of analysis. We hope this multi-scale perspective helps stimulate new theory and predictions about the ecological processes shaping urban biodiversity across the tree of life.

      (3) Methods

      (a) The methodology provides a scalable, consistent, and reasonable measure of both urbanness and species-level urban tolerance. The urban tolerance measure will, of course, not be useful for certain types of research (e.g., animal behavior), but it is appropriate for the resolution of this study.

      We agree that the urban affinity metric presented here is intended for broad-scale, comparative analyses and is not designed to capture fine-scale processes such as individual behavior or short-term demographic responses. Our goal was to develop a scalable and consistent measure that enables cross-taxon and cross-region comparisons at a global extent, which we believe is appropriate for addressing the questions posed in this study. We have sought to be explicit about this scope throughout the manuscript (e.g., to better alleviate Reviewer #1 concerns) and emphasize that the framework is complementary to, rather than a replacement for, more mechanistic or organism-focused approaches.

      (b) I'm concerned that the authors were not able to constrain their dataset to mean, median, or maximum, not potentially sex variability in sizes. Later in the methods, the authors state that they selected the measure of size that was most common within a family. Does this mean that species within a given family that didn't have that measure of body size were removed from the analysis?

      We appreciate this important point and agree that heterogeneity in how body size is measured (e.g., mean, maximum, or sex-specific estimates) is a real and unavoidable challenge in large-scale trait syntheses. Our analytical approach was explicitly designed to minimize the influence of this heterogeneity while retaining as many species as possible, rather than excluding species based on inconsistent trait metadata.

      Specifically, species within a family were not removed based on the availability of a particular body size definition. All species with at least one body size estimate were retained. When multiple measures existed for a species, we selected the measurement type that was most commonly available within each family to maximize comparability while preserving sample size. Remaining heterogeneity among measurement types (including units, measurement detail, and whether values reflected means, maxima, or sex-specific estimates) was explicitly accounted for through log-transformation and metadata-aware centering and scaling, with measurement metadata included as random intercepts in the hierarchical models. We have clarified this point in the Methods:

      “Importantly, this procedure did not result in the exclusion of species lacking a particular body size definition; rather, all species with at least one available body size estimate were retained, with measurement heterogeneity explicitly accounted for through metadata-aware scaling and hierarchical modeling.”

      In addition, our taxonomic modeling strategy was intentionally hierarchical. Species belonging to families that did not meet the minimum threshold for family-level modeling (≥10 species) were not discarded; rather, they were included in higher-level taxonomic analyses (e.g., order- or class-level models), ensuring that available information was retained wherever statistically appropriate. This approach reflects our broader goal of maximizing data inclusion while matching inference to the resolution supported by the data.

      Reviewer #2 (Recommendations for the authors):

      (1) Overlap between VIIRS and GBIF data: While it would have been nice for the GBIF records and VIIRS timescales to match, the degree of mismatch isn't overly large (2010-2021 vs 2015-2021), and any bias or inaccuracies should be minimal. I am mainly making this comment as a potential counterpoint to a possible criticism from other reviewers.

      We thank the reviewer for this helpful observation and agree with their assessment. While the temporal coverage of GBIF occurrence records (2010–2021) and VIIRS night-time lights data (2015–2021) does not perfectly overlap, the mismatch is relatively small and unlikely to introduce substantial bias, particularly given our focus on broad, global patterns of urban affinity rather than fine-scale temporal dynamics. We appreciate the reviewer highlighting this point as a potential counterargument to concerns about temporal alignment.

      (2) Line 87: "only a select few species seem to possess traits that enable them to thrive in urban...".

      This seems like an odd statement, given how many of these species have positive urban tolerance measures.

      Agreed that this was oddly worded. We have revised for clarity, focusing on the magnitude of urban affinity:

      “Similarly, much like the skewed distributions observed in SADs [24,26], the skewed shape of SUDs indicates that while many species exhibit some degree of urban affinity, a relatively small subset of species attain high levels of urban affinity and dominate urban environments.”

      (3) Line 81: "skewed shape of SUDs suggests that traits enabling species to tolerate urban environments are both rare and specific".

      Again, based on the shape of some of these curves, I'm not convinced that it is rare, and there is nothing about these curves that suggests it is something "specific". Indeed, urban tolerance could be very multivariate, and the authors' own results suggest this is indeed the case.

      We have revised the sentence to retain a focus on traits while avoiding overinterpretation of adaptation from the distributional patterns alone. The revised wording emphasizes the uneven expression of high urban affinity across species without implying rarity or trait specificity:

      “The skewed shape of SUDs suggests that traits enabling species to tolerate urban environments are unevenly expressed, given that only a handful of species show extreme urban affinity values, but our results suggest this is geographically widespread across taxa.”.

      We also agree with the likelihood that it is multivariate, and return to this in the conclusion in a stronger sense:

      “Although body size emerged as a predictor of urban affinity, we found not only substantial heterogeneity across families and orders, but also that body size filtering alone is unlikely to explain the consistently skewed SUD shape. Taken together, these patterns suggest that urban affinity likely emerges from multiple trait combinations rather than a single, universally advantageous trait, and that strong affinity to urban environments is not uniformly expressed across taxa, despite occurring broadly across regions.”.

      (4) Line 100: "UHI", avoid abbreviations unless absolutely necessary.

      We have removed this abbreviation throughout.

      (5) Body size: focusing on one trait seems like a shot in the dark, and so it isn't too surprising that this didn't reveal a strong or consistent pattern. However, I also recognize that collecting consistent trait data across so many taxa is challenging, and size is a low-hanging fruit that correlates with multiple traits. Perhaps discuss more the range of traits you think are most likely to predict urban tolerance.

      Body size is indeed the ‘easiest’ to collect, but we acknowledge that there are other traits which could be important, and body size correlates with multiple traits. We revised our discussion to be more comprehensive to discuss some of the additional traits, and be explicit about the shortfalls of body size:

      “Ultimately, the heterogeneous and sometimes weak relationships between body size and urban affinity suggests that body size alone cannot explain the emergence of extreme urban exploiters and the skewed shape of SUDs. Focusing on body size as a focal trait necessarily represents a simplification of the multidimensional processes underlying species’ responses to urbanization, driven in part by data availability when conducting a taxonomically-broad synthesis. Instead, urban affinity likely depends on multivariate trait combinations [17,58] that vary among taxa [59] and ecological contexts [60]. Traits that are likely to correlate with urban affinity include dispersal capacity, behavioral flexibility, diet breadth, reproductive strategy, thermoregulatory ability, and, in plants, life history traits such as growth form, clonality, phenology, and seed size. The diversity of trait pathways through which species may persist or thrive in urban environments is consistent with the pronounced taxonomic heterogeneity we observe and helps explain why body size alone does not yield a universal pattern.”

      (6) Figure S2: This figure and analysis appear to 'come out of nowhere'. I think this is distracting and tangential, and it should be removed. I have the same thoughts about Figure S3. While I do think a discussion of other traits to measure is well warranted and needed, the inclusion of "preliminary' results that aren't motivated by clear questions, appropriate context, and rigorous analysis should be discouraged.

      We have removed Figure S2 and Figure S3 in response to this comment.

      I hope the authors find my constructive comments useful in their revision process.

      This was a very thorough and thoughtful review. We are greatly appreciative of the opportunity and guidance to improve our work!

      Reviewer #3 (Recommendations for the authors):

      Here is a list of a number of further points that the authors may want to address:

      (1) Figure 1 somehow misses the fact that humans simply do not want very large animals in the city. We kill large predators if they come too close to cities, and the same for large herbivores such as wild boar or deer.

      We agree that direct human persecution and management of large-bodied species can influence which species occur in urban environments, particularly for large predators and herbivores. Such processes represent important mechanisms shaping urban species assemblages and represent an entire field of socio-ecological dynamics. We have now clarified this point in the Discussion by noting that human–wildlife conflict, management, and persecution could contribute to observed size–urbanization relationships for some taxa, and that disentangling these mechanisms represents an important direction for future research. We added some text to highlight this point):

      “Similarly, human–wildlife conflict and active management of large-bodied animals in cities may influence which species persist in urban environments, potentially constraining the upper end of the body size distribution. Taken together, these examples illustrate the importance of considering the socio-ecological context of urban species assemblages [65]”.

      (2) Line 270. So you removed all data from the grid-based survey?

      We did not remove all data originating from grid-based surveys or gridded products. Rather, we retained GBIF point-occurrence records and applied a standard spatial filtering step, removing only those individual observations with reported coordinate uncertainty greater than 1 km. This was done to ensure reliable alignment between species occurrence points and remotely sensed environmental layers. We have clarified this distinction in the Methods to avoid confusion:

      “Due to uncertainty in matching observations with remotely-sensed products, any GBIF observation with a coordinate uncertainty > 1 km was removed. This filtering step removed individual observations with high spatial uncertainty, rather than excluding entire datasets or survey types.”.

      (3) Line 278. Human population density?

      Yes, we have added ‘human’ here (and elsewhere in this section) to make this clearer to the reader.

      (4) Line 284. What is a pixel?

      We have modified the text to make this clearer:

      “VIIRS Stray Light Corrected Nighttime Day/Night Band Composites product, representing monthly composites, (i.e., this dataset in Google Earth Engine: NOAA/VIIRS/DNB/MONTHLY_V1/VCMSLCFG) with a native resolution of ~500 m<sup>2</sup>. We took the median of all monthly composites for each pixel (i.e., a single grid cell of the night-time lights raster representing a fixed ground area) to calculate a pixel-level urbanization value, measured in average radiance, and used imagery from January 2015 to January 2021 to calculate this median”.

      (5) Line 292. It seems to me that lighting is different in different types of cities with the same level of impervious surface, depending on local customs of how many lights are installed, left switched on, etc. I guess that petrol stations and strongly lit industrial areas both produce high levels of light, while for the industrial areas, there could be lawn or other vegetation?

      We thank the reviewer for this thoughtful observation and agree that night-time lighting can vary across cities with similar levels of impervious surface due to differences in land use, infrastructure, and cultural lighting practices. We do not interpret VIIRS night-time lights as a direct measure of any single urban feature, but rather as a continuous, integrative proxy for urbanization that captures the combined footprint of human activity, infrastructure intensity, and energy use. VIIRS radiance has been repeatedly shown to correlate strongly with human population density, built infrastructure, and urban extent, while being negatively correlated with vegetation cover (e.g., EVI). It is repeatedly used in remote sensing and urban sustainability literature. This approach is widely supported in the literature, for example:

      Panić et al. used night-time lights were to map spatial and temporal patterns of artificial lighting as a proxy for human population distribution and activity, distinguishing areas of urban and rural occupancy.

      (https://www.ceeol.com/search/article-detail?id=1035395)

      Zhou et al. used night-time light observations were to develop a globally consistent time series of annual urban extent, delineating urban clusters and quantifying global urban growth over decades. (https://doi.org/10.1016/j.rse.2018.10.015)

      Chakraborty & Stokes used night-time light time series with machine learning to detect and quantify urban change processes—identifying deviations from expected radiance trends to monitor diverse urban transitions.

      (https://doi.org/10.1016/j.rse.2023.113818)

      Zhao et al. reviewed night-time light remote sensing was for its broad capacity to quantify human activities and socioeconomic dynamics—such as urbanization, economic change, and environmental impacts—across scales.

      (https://doi.org/10.3390/rs11171971)

      Zheng et al. used VIIRS nightime lights across 30 global megacities to produce a classification scheme to disentangle urban land changes into five categories, and assess global urbanization processes. (https://doi.org/10.1016/j.isprsjprs.2021.01.002)

      Zhao et al. argue that nighttime lights provide a consistent dataset to model and interpret urbanization dynamics and use this to track urban dynamics in Southeast Asia. (https://doi.org/10.1016/j.rse.2020.111980)

      While localized mismatches may occur (e.g., brightly lit industrial areas with surrounding vegetation), such heterogeneity is expected to introduce additional variance rather than systematic bias in the measure of urbanization, making our inference conservative. We have clarified this interpretation and added additional supporting references in the Methods:

      “Previous work has shown that VIIRS night-time lights is negatively correlated with greenness measured through the Enhanced Vegetation Index (EVI) and positively correlated with human population density [69,71]. Although night-time light intensity can vary among cities with similar impervious surface due to differences in land use, infrastructure, and cultural lighting practices, at broad spatial scales it functions as an integrative proxy of urbanization [75,76,77,78,79,80], with localized heterogeneity contributing primarily to additional variance rather than systematic bias.”

      (6) Line 295. How did you reconcile the spatial uncertainty of >1km with an urbanization pixel of 150m2? For how many species did you have a higher uncertainty than pixel size? In my experience, your ca. 39m accuracy is a strong assumption for GBIF data.

      We would like to clarify that we do not assume species occurrence accuracy at the scale of the geohash blocks (i.e., tens of meters), and we do not interpret GBIF records as having ca. 39 m positional accuracy. The use of geohash7 (~150 m blocks) reflects a computational indexing choice, not an assumption about biological or observational precision. All GBIF observations with reported coordinate uncertainty greater than 1 km were removed prior to analysis, ensuring that retained occurrences were compatible with the effective spatial resolution of the remotely sensed urbanization data. Importantly, the effective spatial resolution of our urbanization metric remains that of the VIIRS night-time lights product (~500 m). Geohash encoding at a finer resolution was used solely to efficiently associate point occurrences with the appropriate VIIRS pixel while avoiding redundant extraction or averaging across adjacent pixels. This approach does not increase the effective spatial precision of the analysis, nor does it imply sub-pixel inference. We have clarified this in the Methods:

      “The VIIRS night-time lights data, with a native resolution of ~500 m<sup>2</sup>, was then matched to these blocks by assigning each geohash7 block the average VIIRS radiance value that intersects it. We do not assume positional accuracy at the scale of the geohash blocks, but geohash encoding was used solely for computational indexing, while the effective spatial resolution of the urbanization metric is that of the VIIRS data (~500 m). This approach allows us to avoid unnecessary redundancy in the data while maintaining the original VIIRS resolution”.

      (7) Line 296. Why this high resolution in the species data when your light data is 500m2?

      The apparent mismatch in resolution reflects a distinction between data handling resolution and analytical resolution. Species occurrence records were retained at their native point-level precision to avoid premature spatial aggregation and to ensure that each observation could be accurately matched to the appropriate VIIRS night-time lights pixel. The finer-resolution geohash encoding does not imply that species data were analyzed at that scale, nor does it increase the effective spatial resolution of the analysis. We note, however, that the reported spatial uncertainty of some GBIF records may approach or exceed the resolution of the VIIRS data. Retaining such records represents a deliberate trade-off between spatial precision and data coverage, and is necessary to maximize taxonomic and geographic representation in a global analysis of this scope. Importantly, any residual spatial uncertainty is expected to introduce additional noise rather than systematic bias, making our estimates of species–urban affinity relationships conservative.

      (8) If you could show how your results match the results of Hahs et al and others with respect to occurrence and traits, this would strengthen your approach.

      We agree that explicitly comparing our findings with prior trait-based studies strengthens the interpretability of our approach. We have now added text to the Discussion that directly compares our results with published analyses, including Hahs et al. (2023) and other taxon-specific studies. In particular, we highlight where our occurrencebased estimates recover similar body size–urbanization relationships (four of five taxa in Hahs et al.) and where they differ (e.g., carabids), and we discuss how such differences likely arise from variation in spatial grain, response variables, and definitions of urbanization. These additions clarify how our framework aligns with, complements, and extends existing trait-based work rather than replacing it.

      (9) I wonder whether you could run your analysis with simplified data. In the end, you do not talk much about how high the urban score is, so you may also aggregate values to "highly lighted", "lighted", "some light" and "dark" and re-do the analysis, after checking how these scores correlate with e.g. impervious surface in a slightly larger area than what you used (maybe 50x50m).

      Our analytical framework—and the concept of Species Urbanness Distributions (SUDs) in particular—relies on retaining the continuous nature of the underlying urbanization metric. Discretizing night-time light values would necessarily introduce arbitrary thresholds, reduce information content, and obscure subtle but ecologically meaningful variation in species’ relative affinities to urban environments. Because we focus on relative affinity patterns rather than absolute urbanization classes, maintaining a continuous metric is central to both our methodological approach and conceptual contribution. That said, we agree that exploring how continuous urban affinity scores relate to categorical urban classes or alternative urbanization proxies (e.g., impervious surface at different spatial grains) represents a valuable direction for future work. Such analyses could be particularly informative for translating continuous affinity metrics into applied conservation or urban planning contexts.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Reviewer #1

      Evidence, reproducibility and clarity:

      In this paper, Tomasek and colleagues describe a series of experiments illuminating the effects of OM-89, a bacterial lysate taken orally for prevention of recurrent UTI, on intracellular dynamics of UPEC, using cell culture and organoid models. Suggestions for improvement and for clarification of the authors' conclusions and relevance to human UTI (and OM-89 use) are offered below.

      Major points:

      1. The data indicate that OM-89 exposure in the organoids enhances lysosomal degradation pathways and (in mBOs) autophagic flux, and the authors conclude this is a mechanism by which UPEC regrowth after antibiotic treatment (modeling rUTI) is inhibited by OM-89. They also show enhanced cellular uptake of fluorescently labeled antibiotics (ampicillin) in organoids - this leads them to conclude (and state in the paper's title) that increased intracellular antibiotic concentration effects increased killing of UPEC and decreased regrowth. These are two separate proposed mechanisms, and especially with regard to the antibiotics, they have not shown that increased intracellular antibiotic concentration actually kills intracellular UPEC in their model - only that regrowth as measured microscopically is less. In total, a mechanistic connection between the observed lysosomal effect and the intracellular antibiotic uptake, and which one is more important for UPEC control in this model, is incomplete. The precise wording of the paper's title should be reconsidered accordingly.

      We agree with the reviewer that our study does not establish a direct mechanistic connection between OM-89-induced lysosomal remodeling and enhanced intracellular antibiotic accumulation, nor does it definitively determine the relative contribution of each process to intracellular UPEC control. Further studies dissecting the molecular pathways underlying these phenotypes will be required to determine whether they are mechanistically linked or represent parallel epithelial defense responses induced by OM-89.

      Importantly, additional CFU experiments performed during revision (as suggested in point number 4) revealed that OM-89 already reduces intracellular bacterial burden following a classical gentamicin protection assay, prior to prolonged ampicillin exposure. These findings suggest that enhanced intracellular bacterial control cannot be explained solely by increased intracellular antibiotic accumulation and support a direct contribution of epithelial antimicrobial mechanisms, including lysosomal activation, to the observed phenotype. Nevertheless, the relative contribution of lysosomal remodeling and enhanced antibiotic uptake to bacterial clearance remains unresolved and will require further investigation.

      Accordingly, we changed the title to "Targeted lysosomal activation in bladder epithelium enhances clearance of intracellular uropathogenic Escherichia coli." This revised title avoids implying a direct causal link between increased intracellular antibiotic accumulation and bacterial clearance while reflecting the central biological process identified in our study.

      OM-89 is taken orally for rUTI prevention, and some "components" reach the urinary tract (line 81). But it isn't explained how applying OM-89 directly to organoids models how its components may reach the bladder epithelium (from the basolateral side, if the OM-89 is applied outside the organoids) in the whole animal or human. At the least, this limitation should be stated in the Discussion.

      We thank the reviewer for pointing out this limitation. Although advanced in vitro models help to better mimic the in vivo situation, they still do not fully recapitulate all aspects of drug exposure and delivery observed in vivo. We included the following statement of limitation now in the discussion in lines 493-503: “One limitation of our study is that OM-89 was applied directly to epithelial cultures and organoids, whereas in clinical use it is administered orally. Although pharmacokinetic studies have demonstrated systemic distribution and urinary accumulation of OM-89-derived components following oral administration (van Dijk, 1982), our experimental setup does not recapitulate the exact route, kinetics or concentration profiles encountered in vivo. Rather, our models were designed to determine whether bladder epithelial cells are capable of responding directly to OM-89-mediated signals and to identify the intracellular pathways involved. Given the documented systemic exposure following oral administration, direct effects on the urothelium are biologically plausible. However, future studies will be required to determine how the epithelial responses identified here integrate with the complex systemic and immune-mediated effects of OM-89 under physiological administration conditions.”

      In the lysosome studies starting on line 319, the cultured cells are all infected (and either treated with OM-89 or not). What observations regarding number and size of vesicles, etc (all the measures in Fig 6) are evident when cells are treated with OM-89 only? These data should be presented (at least as a supplemental figure) to enable optimal interpretation of the OM-89+UPEC data in Fig 6. As the authors themselves indicate, OM-89 may be having a generalized effect on endocytic and/or autophagic flux by bladder epithelial cells, independent of infection.

      We thank the reviewer for this helpful suggestion and agree that assessing OM-89 treatment in the absence of infection provides important context for interpreting the infection-associated phenotypes as shown in Figure 6.

      Accordingly, we have included additional supplementary data examining the effects of OM-89 alone in both murine and human bladder epithelial cells. Specifically, we added analyses of Lamp1-positive lysosomal vesicles, lysosomal acidification (LysoSensor), and Cathepsin L activity under uninfected conditions (Supplementary Figures 4A, 4G and 7D-F). We comment on these additional findings in the Result section in lines 242-246 and lines 366-370, and in the Discussion section in lines 469-483.

      These experiments, together with the transcriptional data in SI Figure 3D, demonstrate that key features of lysosome-centered remodeling and activation are already induced by OM-89 in the absence of infection, indicating that OM-89 directly modulates epithelial lysosomal pathways rather than merely amplifying infection-driven responses. Inclusion of these data provides additional context for interpreting the infection-associated phenotypes shown in the main figures and further supports the concept of OM-89 as a direct modulator of epithelial antimicrobial function.

      With the organoids, beyond the microscopic quantification of UPEC, can CFUs be measured?

      We appreciate the reviewer’s interest in obtaining orthogonal measurements of bacterial burden. Performing CFU quantification directly from microinjected organoids is technically challenging, as it requires highly reproducible injections into identical numbers of organoids while avoiding bacterial leakage into the surrounding extracellular matrix. Even minor variations or accidental release of bacteria into the Matrigel can substantially affect CFU recovery and compromise interpretation.

      To address the reviewer’s underlying question while avoiding these limitations, we performed intracellular CFU assays using differentiated mouse bladder epithelial monolayers. Following a classical gentamicin protection assay for 1 hour, OM-89-treated cells displayed significantly reduced intracellular bacterial burden compared with PBS controls (new Figure 2C). Addition of ampicillin for 3 hours after the gentamicin protection phase resulted in a similar trend but did not further significantly reduce the bacterial burden (new Figure 2D). We commented on these findings in the Results section in lines 169-182, and in the Discussion section in lines 463-469 and lines 474-483. We also updated the Methods section in lines 637-652 with the intracellular bacterial burden assay description.

      These experiments provide an orthogonal readout of intracellular bacterial burden and are consistent with enhanced epithelial control of intracellular UPEC. In addition, we would like to clarify that the higher-throughput microscopy approach used throughout the organoid experiments does not allow strict discrimination between luminal, intracellular and tissue-associated bacteria. We therefore revised the terminology throughout the manuscript and now consistently refer to the measured signal as “intra-organoid bacterial burden”. To clarify this point, we added the following statement to the Results section (line 115): “Hence, the microscopy data represent the total “intra-organoid” bacterial burden at each experimental stage, without distinguishing the exact localization of the bacteria – which can be luminal, intracellular or tissue-associated.”. Consistent with this clarification, we have replaced the term “antibiotic-mediated killing” throughout the manuscript with the more cautious wording “antibiotic-mediated clearance” or “reduced bacterial burden”, where appropriate.

      Minor points:

      1. In Fig 1A, the "co-application" horizontal line is under the 7-10 hour window, but the text suggests that the application of antibiotics and OM-89 in this experiment is between 4-7 hours.

      We thank the reviewer for pointing this out. Indeed, in the co-application regime, OM-89 is added at the same timepoint as the antibiotic – meaning straight after monitoring the growth phase at 4h post-infection (pi). We now adapted the horizontal line for the “co-application” treatment in Figure 1A accordingly to represent the time-point of OM-89 addition better. Additionally, we added a line for the antibiotic-treatment in order to further facilitate readability.

      How are antibiotics and OM-89 "removed" at the 7-hour mark? This was not detailed in the Methods.

      Although we had specified this in the methods section (now line 682: “For every media exchange (e.g. antibiotic treatment or withdrawal), each well was washed with 9 ml of the respective media before leaving 1 ml in the well.”), we realized the positioning was not optimal as we had mentioned this part under the point “Bacterial injection” in “Injection experiments”. We therefore now separated this part, together with the lid preparation, from the “Bacterial injection” part and created the new subsection “Lid preparation for media changes” (line 668 onwards).

      What time point was used for the transcriptomic profiling of organoids? This is not clear from the relevant Methods or Results sections.

      As stated in the methods section, RNA for transcriptomic profiling from mBOs was extracted at 4h post-infection (pi) (now line 892).

      In showing that OM-89 "attenuated" the magnitude of inflammatory responses (Fig 2C and S3B), it would be helpful to add a panel showing the comparison of OM89+UPEC to PBS alone - this would be expected to convey activity (red) in the infection-related pathways, but to a lower magnitude than seen in UPEC vs PBS.

      Please see our combined response at point 5.

      Similarly, in the results outlined starting on line 196, it would be helpful to add a panel showing OM89+UPEC vs OM89 alone.

      We thank the reviewer for these suggestions. We performed the requested additional analyses and generated Gene Ontology Biological Process (GOBP) enrichment plots comparing (i) PBS+UPEC versus PBS, (ii) OM-89+UPEC versus PBS and (iii) OM-89+UPEC versus OM-89.

      As anticipated by the reviewer, these analyses show that infection-associated pathways remain induced in OM-89-treated infected organoids but with a reduced magnitude compared with infected PBS controls. Specifically, pathways that are strongly enriched in the PBS+UPEC versus PBS comparison display lower enrichment significance and effect size in the OM-89+UPEC versus PBS comparison. Furthermore, many of these pathways are no longer significantly enriched in the direct OM-89+UPEC versus OM-89 comparison, indicating that OM-89 attenuates the transcriptional inflammatory response induced by UPEC infection. These observations are consistent with our original interpretation, concluded from Figure 3C, that OM-89 dampens excessive infection-associated inflammatory signaling while preserving epithelial antimicrobial activity.

      Importantly, we found that the direct comparison between PBS+UPEC and OM-89+UPEC, presented in the original Figure 3C, remains the most informative representation of the OM-89 effect because it controls for infection status while specifically highlighting the transcriptional changes induced by OM-89. By contrast, comparisons against PBS or OM-89 alone involve simultaneous changes in both infection and treatment status, making biological interpretation less straightforward.

      Nevertheless, because the additional analyses directly address the reviewer's request and provide complementary context for interpreting Figure 3C, we have included them in Supplementary Figure 3B.

      In line 236, what is meant by lysosomal "activation"? A more specific term should be chosen here.

      We thank the reviewer for this question and aim to increase readability of this section. With lysosomal activation in the first sentence of the mentioned paragraph, we referred to the observed effect of upregulated lysosomal pathways and enhanced lysosomal function (measured by alterations in lysosomal vesicles) in the previous paragraph. However, to make the connection to the previous paragraph better, and given the comment number two of reviewer number two, we changed the whole first paragraph of this section. Therefore, the first sentence of this paragraph (line 252 onwards) reads now: “To test whether the observed effects on lysosomal pathways could mechanistically, at least in parts, explain OM-89-mediated protection, we first used Genebridge analysis (Li et al, 2019) to examine how the lysosomal gene signature identified in our RNA-seq data relates to host defense programs in the human bladder.”

      In the Abstract (line 25), the phrase "Using bladder organoids..." is a dangling modifier.

      We thank the reviewer for pointing this out and changed the sentence accordingly to “OM-89 promotes lysosomal acidification and increases lysosomal protease activity in bladder organoids and differentiated epithelial monolayers, thereby directing intracellular UPEC toward degradative compartments.” (now line 24)

      Typographical and copyediting:

      We thank the reviewer for identifying typographical errors and have corrected them throughout the manuscript.

      1. Line 74 should read "For instance..."

      2. Line 76 should read "when combined with antibiotic therapy..."

      As this sentence is to emphasize the already observed protective effects of OM-89, and the two studies mentioned were either performed without or in combination with antibiotics, we changed the sentence to “For instance, rodent infection studies have demonstrated protective effects of OM-89 alone (Bosch et al, 1988; Lee et al, 2006) and in combination with antibiotic therapy (Canton et al, 2025; Bessler et al, 2010), although this observed in vivo protection could not be linked to any major quantitative changes in bladder immune cell infiltration (Canton et al, 2025), leaving the underlying molecular mechanism not fully resolved.” for better readability. (now line 71)

      Line 122 should read "...regrowth following antibiotic treatment" or "regrowth post-antibiotic treatment"

      Line 138 should use "regimen" not "regime"

      Line 196 delete comma after "Although"

      Line 244 fully hyphenate "OM-89-mediated"

      Line 374 should read "...significantly enhance antibiotic-mediated killing"

      Significance:

      The paper is very well written and though a lot of data are included, the presentation is excellent and helps the reader to follow the story. The paper makes a strong contribution to the UTI pathogenesis field, and the use of mouse and human bladder organoids is innovative in studying intracellular UPEC. My scientific expertise as a reviewer is in UPEC pathogenesis, directly relevant to the content of this paper.

      Reviewer #2

      Evidence, reproducibility and clarity:

      This study examined the effect of OM-89 on UPEC infection, antibiotic clearance, and resurgence in mouse and human organoid models. The goal of the study was to understand the molecular mechanisms by which OM-89 is effective at preventing rUTI in patients.

      Major comments:

      The manuscript is well-written and the figures are well presented. Adequate background information is provided to give the study context and sufficient experimental details are provided to allow replication by other groups. Experiments contain appropriate controls and sufficient replicates to allow appropriate statistical analyses. The authors are careful to acknowledge the differences they observed between the mouse and human system and provide satisfactory potential explanations for these differences. The conclusions they draw are well supported by their data and none of their claims from their data are overstatements. Below are some, which I believe if addressed could improve the paper.

      1. I think the authors overstate the novelty of the concept that the urothelium is an active targetable determinant of infection and treatment outcomes. This is not an entirely new concept since previous studies have examined antimicrobial peptides and other factors from the urothelium.

      We thank the reviewer for this important point and agree that the urothelium has long been recognized as an active participant in host defense through mechanisms such as antimicrobial peptide production, pathogen sensing and regulation of inflammatory responses. We have therefore revised the manuscript to avoid implying that urothelial involvement in infection outcome is itself a novel concept. Instead, we now emphasize the specific advance of our study: the identification of lysosome-centered epithelial activation as a therapeutically targetable mechanism that enhances intracellular bacterial clearance and potentiates antibiotic efficacy.

      In the abstract we changed: “Our findings position the bladder epithelium from a passive barrier to an active, targetable determinant of treatment outcome and suggest host-directed modulation of epithelial antimicrobial pathways as a promising strategy to enhance intracellular bacterial clearance.” to “Our findings demonstrate that bladder epithelial antimicrobial pathways can be pharmacologically reinforced to influence treatment outcomes by enhancing intracellular bacterial clearance.” in line 29.

      In the introduction we changed: “Together with increased intracellular accumulation of antibiotics across different classes, this leads to improved intracellular killing and reduced bacterial regrowth across diverse UPEC strains.” to “Together with increased intracellular accumulation of antibiotics across different classes, these changes are associated with improved intracellular clearance and reduced bacterial regrowth across diverse UPEC strains.” in line 90 and “Together, these findings reveal a previously unrecognized epithelial lysosome-centered mechanism by which OM-89 enhances intracellular antibiotic performance and repositions the bladder epithelium from a passive reservoir of infection reactivation to an actively transformable antimicrobial compartment influencing treatment outcomes.” to “Together, these findings reveal a previously unrecognized lysosome-centered epithelial mechanism by which OM-89 strengthens bladder epithelial antimicrobial defenses and enhances intracellular bacterial clearance, identifying enhanced lysosomal function as a therapeutically targetable component of host defense.” in line 95.

      In the discussion we changed: “Together, these findings provide a mechanistic framework for the long-observed clinical efficacy of OM-89. Our findings reveal that the urothelium itself can be therapeutically targeted to reduce pathogen regrowth by transforming the epithelial barrier from a passive refuge for UPEC into an active defense site.” to “Together, these findings provide a mechanistic framework for the long-observed clinical efficacy of OM-89 and identify epithelial lysosomal pathways as a therapeutically targetable component of host defense that can be used to improve intracellular bacterial clearance.” in line 421 and “In the face of rising antimicrobial resistance (2024), strengthening epithelial antimicrobial function offers a complementary route to shift the bladder mucosa from a passive niche of bacterial survival and infection reactivation toward an active site of accelerated pathogen clearance.” to “In the face of rising antimicrobial resistance (2024), our findings provide a mechanistic rationale for the clinical use of OM-89 and support epithelial lysosomal pathways as a promising target for host-directed therapeutic strategies that enhance intracellular bacterial clearance and improve the efficacy of existing antibiotics.” in line 513.

      Depending on the target audience, the Module-Module association analysis could need more introduction. I am not a computational biologist and it was not obviously apparent how Figure 4A is generated and what it actually showing. How specifically does this analysis demonstrate a functional link between lysosomal activity and immune defense pathways? Without further explanation, it is my opinion that this figure panel is an unnecessary distraction that is not required for any of the conclusions that the group can already draw from the rest of their data.

      We thank the reviewer for this constructive critique. We agree that the rationale and interpretation of this analysis were not sufficiently explained in the original manuscript. We have therefore expanded the description of the MMAS approach and clarified how these data support the translational relevance of the lysosomal pathways identified in our experimental models.

      Specifically, we now explain that the Module-Module Association Score (MMAS) analysis evaluates transcriptional correlations between the lysosomal gene network and functional biological pathways across eight independent human bladder transcriptomic datasets comprising more than 1,400 clinical samples. We further highlight the strong positive associations observed with host defense modules, including “response to molecule of bacterial origin”, “cell activation involved in immune response”, and “innate immune response”. These additions clarify both the methodology and the rationale for including Figure 5A as a translational bridge between our experimental findings and human bladder biology.

      The revised text (starting at line 251) now reads: “To test whether the observed effects on lysosomal pathways could mechanistically, at least in parts, explain OM-89-mediated protection, we first used Genebridge analysis (Li et al, 2019) to examine how the lysosomal gene signature identified in our RNA-seq data relates to host defense programs in the human bladder. To evaluate the translational relevance of our experimental findings, we used a computational Module-Module Association Score (MMAS) analysis across eight independent human bladder transcriptomic datasets comprising over 1,400 clinical samples. This network-based approach evaluates the transcriptional correlation between the lysosomal gene network and functional biological pathways across diverse human cohorts. Module-Module association analysis performed on these human bladder datasets indicated that the lysosome module has strong positive associations with specific host defense modules, including "response to molecule of bacterial origin", "cell activation involved in immune response", and "innate immune response" (Figure 5A), highlighting a conserved functional link between lysosomal activity and immune defense pathways in the bladder epithelium. Altogether, these positive correlations suggest that enhanced lysosomal function represents a conserved pathway integrated within mucosal immunity across species, rather than an isolated cellular response unique to our experimental models.”

      Significance:

      General assessment: Solid experimental design with appropriate controls. Appropriate statistical rigor. Conclusions justified by the data. Limitations acknowledged. Differences in results between mice and humans acknowledged.

      Advance: Moderate technical advance building on prior organoid models. Significant mechanistic advance because OM-89 has been widely used for a long time without detailed understanding of why it works. Moderate conceptual advance that urothelial cells are a targetable determinant of treatment outcomes.

      Audience: I am a basic science researcher in the field of female urogenital tract microbiome and infections. Other researchers studying UTI will certainly be interested in this study. It also may be of interest to people studying other bladder conditions that involve the urothelium (bladder cancer).

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      Point-by-point response to the reviewers (____blue____)

      Dear Editor,

      Thank you for taking care of our manuscript. We are pleased to see that the reviewers are positive about our manuscript. We have amended our manuscript to address nearly all the reviewer’s comments. See below our point by points answer Although we cannot fully establish the exact function of the serine protease homolog Skanda in the Drosophila immune response, our study that combines both biochemistry and genetic provides important insight on the Toll-PO cascade and its complexity

      With best regards,

      Bruno Lemaitre on the behalf of the authors


      __Review____er #1 (Evidence, reproducibility and clarity (Required)): __

      In the manuscript entitled "The serine protease homolog Skanda modulates Toll-phenoloxidase-mediated immunity in Drosophila," Vasanth et al characterize in detail a previously unstudied component of the insect immune response using first biochemical and then in vivo methods. Using proteins overexpressed and purified from insect cells, the authors provide evidence that Skanda could be a negative regulator of the SP cascade, impacting cleavage of proHayan and proPsh, and consequently Toll pathway and PPO1 activation. This work reaches further by transposing these findings into the D. melanogaster in vivo model. Here, however, the picture becomes more confusing as Skanda at native levels does not appear to regulate either the Toll pathway or the melanization cascade. Only one strong phenotype was identified in that decreased expression of Skanda increased susceptibility to S. aureus infection while increased expression decreased susceptibility. The mechanism for this remains unclear. To their credit, the authors carry out an in-depth analysis to rule out all the obvious possibilities. In the discussion, the authors explore the basis of discrepancies between their biochemical and genetic findings. We would suggest that an additional one to consider is differing roles or behaviors of Skanda in the microenvironments of the local site of injury (where S. aureus may be contained when it is tolerated) and the hemolymph. In summary, this is a valuable analysis of the innate immune component Skanda whose role has become somewhat clearer through these studies, but still remains obscure.

      We thank the reviewer for this general assessment of our article. We agree with his idea that discrepancies between the biochemical and genetic findings arise from differing roles or behaviors of Skanda in the microenvironments of the local site of injury and the hemolymph’. We added the following sentence in the discussion: ‘The presence of Skanda in the hemolymph (Rommelaere et al. 2025) suggests a role in the systemic immune response; however, we cannot exclude that it may be particularly important within the local microenvironments at sites of injury’.

      __Major Comments __ - To assess bimodal distribution of bacterial ds within single flies in Fig 6E, authors should either: increase the sample size to allow for proper statistical assessment of different distributions among genotypes, specifically between w1118 and skanda_d107; or, provide a modelling framework for statistical testing. Otherwise, the present results seem insufficient to conclude that Skanda is playing a role in resistance to S. aureus. We agree with the reviewer that our bacterial count was not enough developed. In the revised version we add a new Figure 6E with two time points 13h and 16h that were chosen before flies start to die from S. aureus. We observe at 13h a significantly higher bacterial count in the Skanda mutants but not at the 16 hours although there is higher proportion of wild-type flies that have clear the bacteria. These observations suggest a role of Skanda to resist, but also tolerate S. aureus. The fast killing induced by systemic injury with a low dose S. aureus made difficult to find a condition that would allow to see a clear load difference. So we have amended our text to highlight that Skanda could also play a role in tolerance.

      We agree with the reviewer but measuring the BLUD with S. aureus is rather challenging as flies die quickly to this bacterium. As mentioned above and following revised figure 6E, we discuss in the revised version that Skanda could be involved in both resistance and tolerance.

      • The error bars on qRT-PCR datasets are large, the data points are not shown so we do not know how many replicates were included in the graphs (Fig 5 B and C, Fig 6C, Fig 7 A and B, and Fig 8B). Bar plots are not the most faithful reproduction of biological datasets, as they can hinder significant information regarding datapoints distribution and variation (Beyond Bar and Line Graphs: Time for a New Data Presentation Paradigm | PLOS Biology). We advise that, particularly in the case of datasets such as qRT-PCR, the final values of fold change are represented with individual dots, with the mean value clearly represented, whether with or without the additional bar graph. Furthermore, no statistical tests were applied to determine significance. Data points should be shown and appropriate statistical tests should be applied. The number of biological replicates should be included in the analysis and the statistical test applied should be noted in the figure legends.

      We have changed the figures related to qRT-PCR to show the individual points and we have added statistics in the revised version.

      • Although there are claims of Skanda conferring resistance to S. aureus infection, only Drs levels are tested. These conclusions could be strengthened by assessing expression levels of additional AMPs.

      In the revised manuscript, we report the expression of BomS1 in wild-type, skanda, and spz mutants following S. aureus infection. As previously observed for Drosomycin, Skanda does not markedly affect BomS1 expression (new Supplementary Figure S3E).

      __Minor Comments __ - Parag. 1: (data not shown) should be removed and if possible AlphaFold prediction of skanda conformation added. Alternatively, remove sentence.

      We have removed (data not shown) and indicated that the information derived from Alphafold.

      • Parg. 3: 1000 mL? why not 1L?

      Corrected.

      • Parag. 5: , in last sentence that should be .

      Corrected.

      • Parag. 6: "a role at the same position..." does not convey the correct messageWe have improved the sentence for ‘Our results indicate that Grass processes Skanda in the Toll–PO SP cascade, consistent with Skanda acting at the same level of the proteolytic cascade as Hayan and Psh’.

      • Figure axes (5D, 5E, 6D, etc...) of melanization assays are wrongly named "% melanisation", with "s"

      We have corrected for “Melanization”.

      • Parag. 21: compound mutants (if correctly interpreted as dataset presented in Fig. 8B) were tested at 6h, 24h and 48h, and not 32h, as written in the text

      Indeed, in figure 8B, we monitored expression at 6, 24 and 32h and not 48h. This has been corrected.

      • Results section "skanda is not mandatory for the activation of the Toll pathway" adopts a literal translation which would probably be better phrased as "is not essential"

      We have corrected accordingly.

      • Discussion parag. 2: "Skanda exhibits..."

      • Discussion last parag: "..., but also underlies..."

      • It has been evidenced that

      This has been corrected.

      Additional comments: - The sentence on page 2 beginning with "Upon binding, these PRRs..." is very long and difficult to follow. This should be rewritten.

      We have split this sentence in two shorter ones for clarity.

      • In many places in the manuscript bacterial "dose" is used in place of bacterial burden. The dose is the amount of a substance or bacterium given to the animal.

      We have changed ‘bacterial dose’ for ‘bacterial burden’ when relevant, and we have kept the term “dose” when we mentioned the OD used to infect flies.

      Page 11: Skanda is described as a placeholder when I think a (competitive) inhibitor would be more appropriate.

      We agree that Skanda functionally resembles a competitive inhibitor, but several key differences set it apart from classical small-molecule inhibitors. First, Skanda is comparable in size and structure to Persephone and Hayan, natural substrates of Grass. Second, Skanda-like SPHs, which have close SP paralogs (e.g., Psh), are common in insects (Cao and Jiang, 2019), indicating that they may constitute a distinct class of negative regulators that warrants its own terminology. Moreover, because amplification in protease cascades typically occurs at the terminal step. Negative regulation by Skanda in an intermediate step could be more stochiometric than the freely reversible inhibition expected for a typical competitive inhibitor. As Skanda’s mechanism remains unclear. the neutral term “placeholder” seems more appropriate than “competitive inhibitor”.

      **Referee cross-commenting**

      I agree with the comments of the other reviewers.

      Reviewer #1 (Significance (Required)):

      Strengths: The authors take a multi-disciplinary biochemical and in vivo approach to understand the molecular interactions among SPs and SPHs and thereby uncover the role of the protein Skanda that might otherwise not have been appreciated. They have made extensive use of novel transgenic fly lines, generated in the context of this study, and have thoroughly tested their specificity and cis-acting potential. These will provide a resource to the field. In addition to the new description of Skanda, these findings strengthen previous knowledge regarding systemic infections with different bacteria (M. luteus, S. aureus) and reproduce the known redundancies of Psh and Hayan modes of action. Moreover, this research is relevant for the expansion of basic knowledge on innate immunity, particularly in the field of insect-pathogen interactions, making use of S. frugiperda cell lines and D. melanogaster adults and larvae. Although not at the focus of this work, the evolutionary conserved nature of these aspects of innate immunity across these two distant species enhance the importance of these findings.

      Weaknesses: Some assays do not include enough biological replicates and others do not have enough information on how many biological replicates were performed. Therefore, the conclusions drawn are difficult to assess. Lack of statistical analysis on the qPCR experiments complicates the interpretation of results.

      We thank the reviewer for his assessment. We have added the number of replicates in the revised version and make visible the variability of our data.

      __Review____er #2 (Evidence, reproducibility and clarity (Required)): __

      Summary In this work the authors identify the SPH skanda as an important player in Drosophila resistance to S. aureus infections independent of Toll and classical melanization. The authors conducted rigorous in vitro assays using recombinant proteins of various SPs in the Drosophila Toll-PO cascade to show that skanda negatively regulates activation cleavage of SPs at the level of and downstream of Psh and hayan, two key SPs that converge on Toll pathway activation with the latter playing a central role in cuticular melanization. In parallel, genetic analysis using mutant flies showed that skanda does not negatively regulate Toll pathway nor melanization. Only skanda over expression in vivo led to a reduction in S. aureus melanization which, in my opinion, is most likely due to the artificial increase in the in vivo concentration of the protein rather than an indication of a potential true function. Altogether this an interesting work as it shows the discrepancies between the biochemical and genetic approaches when it comes to dissecting the insect SP cascades regulating melanization and Toll as highlighted by the authors themselves in the discussion section. All experimental work is well controlled, methodology is robust and results are adequately discussed. I have some comments concerning few experiments and interpretations that in my opinion warrant further discussion.

      We thank the reviewer for the analysis and agree that the result showing than Skanda negatively regulates melanization could be due to over-expression.

      __Major comments: __ 1- It seems that SP48 and Grass can redundantly cleave Skanda although the later cleaves more strongly. (Fig 3B) Can other downstream SPs cleave skanda? Can ModSp alone cleave skanda? (ModSP + skanda lane was absent for Fig 3B). It is important to test these possibilities as the in vitro system may be quite relaxed as to the specificity of these cleavage events and may not reflect what happens in vivo. In fact it has been shown in Anopheles gambiae that SPH can be redundantly cleaved by multiple SP in the protease cascade. Although these are cascades with certain hierarchy, information can still flow in more than one direction along the different branches of these cascades.

      We tested whether ModSP could cleave pro-Skanda and found that it did not (data not shown). This result is consistent with our expectations, as ModSP has a chymoelastase-like specificity and preferentially cleavage after Leu. In contrast, Skanda is cleaved by Grass and cSP48, both of which are trypsin-like proteases.

      At present, there is no straightforward way to assess whether downstream SPs activate pro-Skanda. Obtaining an active downstream SP would require sequential activation of all its upstream enzymes, and it is nearly impossible to completely remove these activating proteases afterward. As a result, it is difficult to distinguish the activity of a downstream SP from that of cSP48 and Grass. We are currently developing a new approach to overcome this limitation.

      2- In Fig 4B and 4C the bands of active forms should be quantified from at least 3 immunoblots for robust results especially in Fig 4C where the differences are minimal.

      As suggested by the reviewer, we quantified the band intensities from four independent blots and presented the data in Fig. 4B and 4C (lower panels).

      3- It is not clear to me why skanda should have a specific role in resisting S. aureus infections despite that S. aureus is not a natural pathogen of Drosophila? Has other Gram-positive and Gram-negative bacteria been tested?

      It is true that S. aureus is unlikely to be a natural pathogen of Drosophila. However, this bacterium has been used in several studies (notably Dudzic 2019) to uncover a specific activity associated with melanization modules that is distinct from cuticular blackening. For this reason, we believe that S. aureus provides a sensitive assay to monitor this particular immune mechanism. We further hypothesize that other bacteria related to S. aureus—possibly members of the Staphylococcus family—may infect Drosophila and could be controlled by Skanda. We chose not to elaborate on this point to avoid overextending the scope of the article.

      4- In Fig 6E more points should be collected for statistical power. It is also better to show these data that are not normally distributed in violin charts or boxes and whiskers which give a better indication as to which quartile the bulk of the data belongs.

      We have addressed this point (see answer to Reviewer 1).

      Minor comments: 5- In Figures 3 and 4, It would be easier to follow the cleavage events if a schematic drawing is provided showing the sequence of activation cleavage events of the tested SPs

      Because the order of the two cleavage events is unclear, we felt it was simpler to include the putative cleavage sites in Fig. 2B and refer interested readers to Fig. S1, Table S1, and Fig. 3 legend.

      6- The fact that PPO1/PPO2 depleted flies exhibit increased Drs expression could be due to increased bacterial proliferation in this mutant background that trigger increased Toll stimulation, rather than a negative feedback mechanism. This increased proliferation is shown in Fig 6E.

      This is a good point. The higher expression of Drs in PO1/PPO2 depleted flies could be associated to higher bacterial load in the mutant, or to negative feedback of the melanization reaction. This higher Toll pathway activation has been further characterized in Liu et al., (Plos pathogen 2025) where it was suggested that it relate to a negative feedback loop between the Toll and the melanization cascade.

      7- In Fig 6E more points should be collected for statistical power. It is also better to show these data that are not normally distributed in violin charts or boxes and whiskers which give a better indication as to which quartile the bulk of the data belongs.

      We have addressed this point. See answer to reviewer 1 for discussion.

      8- A phenotype for skanda in melanization was observed only in over-expression assays which may artificially alter molecular interactions in the cascade.

      We agree with this statement and we have added a comment in the discussion of the revised manuscript about the potential artifactual results due to over-expression.

      9- Page 10 last paragraph "peak expression at 32 hrs or 48 hrs as shown on the figure?"

      This is 32h and has been corrected.

      10- The differences in Drs expression levels in Hayan-pshDef and psh-skandaDef double mutant flies infected with M. luteus and S. aureus is surprising. I wonder whether the observed differences are due to biochemical differences in the microbial surfaces to which these cascades are recruited.

      Drs expression is markedly higher following systemic infection with M. luteus than with S. aureus, consistent with the different bacterial doses used. We deliberately employed a low dose of S. aureus because this condition reveals a pronounced susceptibility in skanda flies. Consequently, direct comparison between these two infection regimes remains challenging.

      11- There are several typos in the manuscript

      We have carefully re-read the manuscript and corrected several typos.

      Reviewer #2 (Significance (Required)):

      The main strength of this work is that it combines biochemistry and genetics in a strong genetic model to characterize the biochemical interactions between SPH and Sp in clip cascades and relate the relevant interactions observed in vitro with potential in vivo functions. This is the first time that such a rigorous combined approach was adopted to the study of these cascades. The results obtained also show the advantages and limitations of each approach. As such i believe this study will be of interest to a broad audience in the field of insect immunity.

      __Review____er #3 (Evidence, reproducibility and clarity (Required)): __

      __Summary: __

      Serine protease cascades are central for activation of immune responses in insects. In Drosophila melanogaster, Toll signaling pathway has been quite extensively studied, and several serine proteases, serpins and serine protease homologs (SPH) with functions in Toll activation have been identified. In this work, the authors characterize a new component of this system, a SPH which they name Skanda. Skanda seems to have multiple roles/points of action, on one hand participating in the regulation of Toll together with the established serine protease in the Toll activation, Psh, and on the other hand controlling the response to a systemic S. aureus infection, via not yet fully specified mechanism.

      __Major comments: __

      Key conclusions made in this work are convincing, and backed up by the data presented. The data and methods are presented in a way that allows reproduction of the experiment. The number of individuals used especially in the infection experiment (20 male flies per a replicate) is on the lower side, but the experiments are adequately replicated and the effects seen are clear.

      While this work contributes to our understanding of the regulatory mechanisms governing Toll signaling, at times the authors' reasoning is difficult to follow. I recognize that this is a complex topic, with multiple upstream branches activating Toll signaling, and the authors do consider various mechanisms that could explain their findings. However, the manuscript would benefit from additional clarification, perhaps through a schematic model illustrating the proposed effects of Skanda, to help readers position Skanda within the broader context of Toll signaling. We have done our best to explain the Toll serine protease and added a figure at the beginning of the manuscript. Since we cannot position Skanda in the Toll-Po cascade yet, we prefer to avoid drawing a model. We believe that this study highlights our ignorance of the complexity of serine protease cascades acting upstream of Spätzle and Melanization.

      Statistical analyses for the Drs expression experiments are lacking.

      The statistical analysis for Drs expression has been added in the revised version.

      __Minor comments: __

      The authors could explain what type of cells the sf9 cells are and why they decided to use them.

      Sf9 cells are an insect ovarian cell line derived from Spodoptera frugiperda and are widely used for baculovirus-mediated expression of eukaryotic proteins. They support proper protein folding, disulfide bond formation, and post-translational processing. This information is now mentioned in the Result section in addition to methods.

      Band intensities could be measured and plotted for the immunoblots. The immunoblot methods should be fully described in the Materials and methods section.

      Thanks for the suggestion. We have done this accordingly and included the results in Fig. 4B and Fig. 4C (lower panels). Brief descriptions of densitometric analyses have been added to the figure legends.

      Protein levels of Skanda in the Skanda mutant could be shown as the mRNA levels remain relatively high (Sup. Fig 3B). If this is not possible, could the authors comment on the remaining expression of Skanda in the Skanda mutants?

      We have added a comment on this point: The skanda mutation is a frameshift mutation that affects the coding sequence. There are still transcripts although not functional. The decreased expression of Skanda in SkandaD107 is probably due to non-sense-mediated RNA decay caused by the frameshift.

      Under the heading "Loss of skanda does not further enhance the cuticular melanization defects caused by the loss of Hayan or psh" the text should refer to figure 5D not 5B.

      We have corrected this mistake in the revised version.

      Figure 6C shows that Drs expression is higher in the Skanda mutant than in controls at 32 h post S. aureus infection (although this has not been statistically tested). The authors don't mention this result in the manuscript, but to me it fits with the idea of Skanda acting as a negative regulator (the effect of which is accumulating and seen only late after infection). Could the authors comment on this? We do not think that the higher expression of Drs in Skanda mutant upon S. aureus systemic infection is due a negative regulation the Toll pathway but rather to higher S. aureus burden. We conclude this because Drs is not higher than the wild-type upon injection of M. luteus and proteases. At this stage, we cannot exclude that there are differences between M. luteus and S. aureus.

      Under the heading "Psh and skanda redundantly regulate Toll signaling", the comparison should likely be between Figures 7A-7B and 5B-C (rather than 5A). When examining the effects of single versus double mutants on Drs expression, the Psh-Skanda double mutant clearly reduces Drs more than the Psh single mutant. However, in the context of microbial proteases, the pattern appears different: there is virtually no difference at 6 hours, while at 48 hours there may be a slight decrease in Drs expression in the double mutant compared to the Psh single mutant, although this difference would likely not reach statistical significance if tested. I don't know what this could mean, but I'd like to hear the authors' take on this. The reviewer is correct and we have revised our manuscript to mention the appropriate figure. Figures 7A-7B and 5B-C.

      The reviewer raised a good point; we believe that the additional effect of Skanda in absence of Psh is less marked upon microbial proteases because Psh already has a strong effect by itself in sensing proteases. In contrast there is higher redundancy between Psh and Hayan upon M. luteus and consequently the double mutant psh, Skanda have a stronger effect.

      __**Referee cross-commenting** __

      I also agree with the comments and points raised by the other reviewers.

      __Review____er #3 (Significance (Required)): __

      Research on the Drosophila immune response has significantly advanced our understanding of (innate) immune responses, both generally and in an evolutionary context. Despite over three decades of study, this work demonstrates that there are aspects of Toll signaling that remain unresolved. The authors identify a novel regulator of the Toll pathway and begin to elucidate its functions. Equally important, their findings underscore the complexity and context-dependency of the regulatory events that shape immune responses.

      We fully agree with the assessment of the reviewer. Our study highlights the complexity (and our ignorance) of this important facet of Drosophila immunity, as mentioned in the last sentence of the discussion.

      My fields of expertise are Drosophila melanogaster, innate immunity, cell-mediated immunity.

      __Review____er #4 (Evidence, reproducibility and clarity (Required)): __

      __Summary __

      In this study, the authors investigate the function of Skanda, a serine protease homolog (SPH) in Drosophila innate immunity using both biochemical and genetical approaches. The reason to focus on this SPH is that it lies at the same locus as two key proteases of Drosophila immune defenses, Hayan and Persephone, all of which are induced by an immune challenge. After having modeled this SPH and shown that the three amino-acid of the serine protease catalytic triad are either mutated or poorly oriented, they report that Skanda may limit the cleavage of proteases downstream of Grass, a key event for their biochemical activation. The study of an isogenized, putatively null, mutant line failed to reveal any impact of skanda on Toll pathway activation nor on melanization, albeit a strong but not moderate overexpression somewhat inhibits the formation of a melanization scab only after "clean" but not septic injury. These results are not in keeping with the biochemical analysis: the mutant would have been expected to display an enhanced immune response. Unexpectedly, skanda mutants are as highly susceptible to a low amount of Staphylococcus aureus injection as flies deleted for the adult-expressed phenoloxidases PPO1 and PPO2, melanization playing a key role in host defense in this infection paradigm. No strong impact on the bacterial load was detected at the sole investigated time point, 24h. Because the analysis of the single skanda mutant did not unambiguously reveal its role in host defense, the authors then studied double or triple mutants of the three protease genes and found a redundant role for Skanda with Persephone for Toll pathway activation after a challenge with a nonpathogenic Gram-positive bacterium or a bacterial protease. In the case of S. aureus infection, a strong induction of the Drosomycin gene, is observed at 48h of infection in the compound mutants, which was not observed with the nonpathogenic challenges. Evidence, reproducibility and clarity

      __Major comments __

      The authors state that "These results are consistent with a role of Skanda in resistance to S. aureus". This conclusion rests on a very fragile experiment that measured the bacterial burden 24h after challenge with a low dose of S. aureus: whereas wild-type control flies exhibit a dual low and high distribution of bacterial loads, skanda flies exhibit only the higher values. However, the bacterial load in skanda appears to be as high in persephone mutant flies that are much less sensitive to S. aureus than skanda flies. This makes it highly unlikely that the high susceptibility of skanda to S. aureus is due solely to resistance. The problem is compounded by the poor description of the experiment: it is not stated anywhere how many times the experiment has been performed, whether pooled data are shown, what each data point represents, pooled or single flies. A fine-grained time course with more biological samples would definitely be needed to convince the reader of a (limited) role in resistance. The authors do not consider the alternative, but not exclusive, possibility that skanda plays also a role in disease tolerance. The determination of the bacterial load upon death of single flies may provide some clues about this alternative function (Duneau et al., eLife, 2017). Another approach might be to determine whether the bacterial supernatant is toxic and whether skanda might protect from this toxicity. As Bomanins play a role in the host defense against S. aureus (this study, but see also Hanson et al., eLife 2019 in which the 55C deficiency susceptibility phenotype was stronger) and given the role of Bomanins in host defense against Gram-positive bacteria or fungal infections both in resistance and disease tolerance (e.g., Clemmons et al. PLoS Pathogens 2015, Lindsay et al., J. Innate Immun, 2018, Xu et al., EMBO Reports 2023, Lou et al., BioRxiv, 2025) and that BomS1 has an optimal Dorsal-related Immune Factor Binding site (Busse et al. EMBO J. , 2007), it may be useful to monitor the expression of several Bom genes in complement to that of the expression of Drosomycin, especially after S. aureus challenge. Furthermore, BomT1 is the only peptide that appears to play a role in resistance against Gram-positive bacteria, namely against E. faecalis. This series of qPCR experiments is rapid to make, provided the authors have kept the cDNAs of their samples.

      To address the reviewer’s comment, we extended the bacterial load analysis of S. aureus in skanda mutants (new figure 6E). Our results support a role for Skanda in both resistance and disease tolerance. This point is now briefly discussed in the Results section, and we have added references highlighting a role of the Toll pathway in disease tolerance. We did not elaborate further, as accurately monitoring S. aureus burden following low-dose infection remains technically challenging given the high pathogenicity of this bacterium.

      In the Discussion, the authors speculate "that Skanda acts at the level of Persephone-Hayan to allow Hayan to activate the Toll pathway. Skanda would skew the activity of the Persephone-Hayan platform to induce Toll signaling and resistance to S. aureus rather than cuticular melanization". This model does not fit with the fact that SPE is only moderately susceptible to S. aureus (Dudzic et al., 2019) and that spätzle mutant flies are either not sensitive at all (Dudzic et al., 2019) or moderately sensitive to it (Hanson et al., eLife, 2019) (see also below). Whether it may apply to host defense against other pathogens remains to be determined. To better understand the function of skanda, considering only S. aureus may be limiting as this bacterium is fundamentally not susceptible to the canonical Toll intracellular signaling cascade (e.g., Bischoff et al, Nat Immunol, 2004, Dudzic et al, Cell Reports, 2019) and to the final part of the Toll-activation proteolytic cascade as discussed above with SPE and Spätzle. The authors appear to have chosen not to display the results they have gained with Enterococcus faecalis (but forgot to remove their mention at two places in the Material and Methods): it would definitely be interesting to know what the outcome of these experiments was and also to investigate the susceptibility and microbial burden of skanda mutants to representative yeast and filamentous fungal pathogens, Aspergillus fumigatus being of special interest since its proliferation is limited through melanization whereas the Toll pathway protects against secreted virulence factors (Xu et al., EMBO Reports, 2023). This series of experiments would likely take some three months and might give additional insights into Skanda function(s).

      We agree with the reviewer that examining the role of Skanda in response to additional bacterial species could further help elucidate its function. However, the most robust phenotype we identified is a strong acute susceptibility to S. aureus, which is dependent on the Psh–Hayan–Skanda axis but independent of the SPE–Spätzle pathway. Because the bacterial strains suggested by the reviewers are primarily controlled by the SPE–Spätzle–Toll pathway, we did not pursue this direction further. However, in the revised version we have added survival analysis with Skanda to Candida albicans and Enterococcus faecalis (new supplement Figure 3F and G). Notably, we also observed an intermediate susceptibility to both Candida albicans and E. faecalis (see below). This indicates that Skanda is not a classical regulator of the Toll-PO cascade such as Grass, ModSP, SPE or Hayan/SPE.

      In general, figure legends are not highly informative and fail to provide key information such as the number of independent experiments, whether the data are representative or pooled, which statistical test was used, e.g., qPCR experiments (the descriptions are available for the analysis of survival and melanization experiments at the end of the Mat. and Meth section). As noted above, critical information is lacking to understand microbial load graphs. It is also difficult to check statements such as: ", while psh[sk1] flies showed a reduced Toll pathway reponse". Indeed, no statistical analysis has been performed to analyze any RTqPCR data. Given the low number of experimental data points, each data point ought to be displayed and not bar graphs, for which in addition the error bars are not defined. The Material and Methods section is incomplete. It does not include a description of all the in vitro synthesized proteins used in this study nor indicate the different tags. The primary and secondary antibodies used for Western blot analysis are not reported, e.g., those that detect cleaved spätzle. This would need to be included in the Table at the beginning of this section.

      In the revised version, we have addressed these points by adding statistical tests to the RT–qPCR analyses, displaying all data points, and improving the microbial load measurement. As discussed in the Material and Methods section, Table S2 provides information for all in vitro synthesized proteins used in this study, including affinity tags and the primary and secondary antibodies. On a more personal note, we first identified the striking susceptibility of Skanda/CG15046 flies more than 10 years ago, and the skanda project subsequently experienced a long period of discontinuation before we decided to reassemble and consolidate the most important findings. Unfortunately, this study did not result in a straightforward narrative with a “happy ending.” Nevertheless, we still consider this work an important step toward a better characterization of this aspect of fly immunity.

      __Minor points __ Introduction: 1. The authors may want to cite Stein, Cho&Stevens, FLY, 2013 when referring to the proteolytic cascade regulating the establishment of dorso-ventral patterning.

      This reference has been added

      The statement "The Toll-PO SP cascade can be DIRECTLY activated at the level of Psh-Hayan, through direct cleavage of the Psh protease bait region by microbial proteases" may be slightly misleading as only subtilisin is able to do this, the other tested proteases producing an inactive cleaved Psh that needed to be secondarily activated by a couple of specific cathepsins (Issa et al., Molecular Cell, 2018).

      Good point. This point has been corrected with the Issa reference added.

      Results 3. The reasoning of the second paragraph is difficult to follow as the reader does not understand how the cleavage sites can be computed. It would be important to state that the recombinant proteins are tagged. It would actually be very helpful to provide a scheme of the various recombinant proteins used in the study as had been done in the Shan et al., Science Advances article.

      We followed the reviewer’s good suggestions, modified the text accordingly, and added Table S2.

      With respect to Western blots, many of the bands are faint, e.g., SPE after the addition of Skanda cannot be detected on a printed version of the figure. It is also difficult to determine whether the reduction in band amount is reproducible as no indications are given in this respect. It is important that the images be quantified in several independent blots so that the observed reduction can be statistically assessed. With respect to PPO1 cleavage, it would be important to also check its cleavage in vivo, which would yield higher confidence on the relevance of in vitro study to the in vivo situation.

      In response to the reviewer’s suggestions, we repeated SDS-PAGE and immunoblot analysis, quantified band intensities, and performed statistical analyses for the samples shown in Fig. 3B and 3C (lower panels). The total number of blots for each representative is 3 to 4. For practical reasons, we are unable to assess PPO1 cleavage in vivo.

      First sentence of the paragraph "skanda mutants are highly susceptible": the authors might also want to cite Hanson et al, eLife 2019.

      We have added the Hanson reference and Ryckebusch et al 2025, which is more appropriate.

      In Dudzic et al., Cell Reports, 2019, the authors did not observe any susceptibility to S. aureus with Hayan[sk3] whereas here they find an intermediate sensitivity phenotype with Hayan[sk6]. Was the former not a null allele of Hayan? With respect to the 55C Bomanin deficiency, Hanson et al., 2019 had reported a stronger phenotype than that shown in Fig. 8A, with some 75% of flies dead within three days. Which study should we trust or does this reflect variations between experiments (hence the question about the representation of survival data: are these pooled data from thre independent experiments; how much variation was there between independent experiments?).

      Both Hayan mutant flies were null. We observed differences along the years with different experimenters; although the main results stand. We also tend to observe a stronger impact of psh than initially reported in response to M. luteus (Figure 5B), although this is consistent with its role in the PRR-Grass-SPE pathway. Considering all the parameters that influence survival experiments (temperature, humidity, time to form the bacterial pellet and sometimes bacterial strains) and possible cryptic infections (Nora infection), we consider these variations as expectable.

      It would be interesting to measure the S. aureus bacterial load upon skanda overexpression to confirm a putative role in resistance.

      This is an interesting suggestion but we did not do it because of the technical challenge that monitoring S. aureus burden represents. We have preferred to focus our attention on monitoring S. aureus in Skanda loss-of-function mutants.

      UAS-skanda: besides Fig. 6B, the authors should also refer the reader to Fig. S4A.

      The link to Fig S4A has been added.

      Genetic dissection of the skanda-psh-hayan gene cluster: the last sentence of the paragraph does not reflect what Fig. S7B is showing: one of the double mutants and the triple mutant displayed a significant intermediate susceptibility to S. aureus.

      This is in fact Ecc15 that we discussed. The reviewer is correct as the triple mutants and hayan,psh double have increased susceptibility to Ecc15.

      Paragraphs Compound mutants are EXTREMELY susceptible to S. aureus. The wording is likely too ...extreme: they do not seem to die much faster than skanda simple mutants, which were HIGHLY susceptible to S. aureus, like PPO1-PPO2 double mutants.

      The reviewer is correct and we have avoided to use the term ‘extremely’ in the revised version (replaced by ‘highly’ or removed).

      Last paragraph: psh mutants should be compared side-by-side with psh-skanda double mutants in the same RTqPCR experiment: it is difficult to judge whether the statement of equivalent Drosomycin expression after S. aureus challenge is true given the low resolution of the figures (Fig. 6C vs. Fig. 7B). Last sentence: it would be more appropriate to mention "host defense" rather than "resistance" since the authors did not check the bacterial burdens of the compound mutants.

      Experiments were done simultaneously on single and double/triple mutant but this represents kinetic with 4 times in 10 different backgrounds! We have preferred to separate the data to simplify the reading. We believe that the reader can compare the data despite display in two different panels. We have changed in all the manuscript host defense instead of resistance as following bacterial counting, we suspect that Skanda may play both in resistance and disease tolerance.

      Fig. 1: the scheme is not up to date and oversimplified. It should take into account the complexity revealed in the Shan et al. Science Advances article.

      We disagree on this point. This schema reflect inference done by genetics. An up-to-date figure is shown in Westlake, Hanson Lemaitre Handbook but would require a broad introduction. In the revised version, we have highlighted that this is simplified model based on genetics.

      Fig. S1: numbering the amino-acids in the sequence would help follow the text from Document S1. What are the residues written in light blue? It may be worth highlighting residue E194. Of note, there is a difference between the sequence for peptide 4 as found in the sequence displayed on Fig. S1: KTDRD YV and the sequence of peptide 4 in Table S1: KTDRE YV; the presence of a potential SNP should be indicated, even though it is not making a major change in terms of charge of the peptide.

      We included an asterisk at every tenth position and a numerical indicator near the end of each line to facilitate counting. Residues highlighted in cyan may represent cleavage sites of cSP48, Grass, or a trypsin-like protease released by Sf9 cells. The peptide (E194R212) appears to undergo cleavage to generate P204LNLPLQP__R212__, which is detected in the secondary MS. The reviewer is correct on peptide 4 that we attribute to a potential SNP. This is now indicated in the legend of Figure S1.

      Document S1: trypsin digestion (just before second call to Fig. S1); should it not be purified proteases instead? The text should be somewhat reworded as it is currently slightly misleading.

      "In lane 8, peptide-1 through -19 were nearly undetectable". Table S1 shows that even though peptides 1, 2, 6, , 7 , and 11 are not expressed to strong enough a level to be displayed Fig. S1 lane 8 given the chosen scale, peptides 1, 2, 6, and 7 are expressed in the same range for slices 8B and 8C, whereas peptide 1 is found with just a two-fold difference in slices 8A and 8C.

      Points taken. To better illustrate the differences in band intensities in the top right panel of Fig. S1, we kept the same scale for bands A and B in line 8 (as well as for bands A-C in the top left and middle panels) and used the second y-axis for band C.

      Fig. S2: the effect of skanda on SP7 cleavage is not detectable when Hayan isoforms are co-incubated. The main text should be modified to take this into account. How do the authors explain that pro-MP1 levels are not different upon co-incubation with Psh or Hayan-PB with or without adding Skanda, even though the active MP1 form is detected only in the absence of Skanda? In contrast, the pro-MP1 band can be detected upon co-incubation with Skanda and Hayan-PA.

      Thanks for the comments. We repeated the experiments and obtained four independent blots for each. After scanning, integrated band densities for all paired bands (i.e., with and with Skanda) were quantified using ImageJ (Fig. S2 and data not shown). In the representative blots, Skanda had little effect on SP7 activation by Hayan-PA (507/527; 96%) or Hayan-PB (15,763/15,828; ~100%), in contrast to Psh (937/7,917; 12%). However, when ratios from all blots were considered, the mean reductions were 56 ± 14% for Psh, 49 ± 19% for Hayan-PA, and 65 ± 18% for Hayan-PB. For MP1, comparison of precursor bands is less reliable because small decreases in precursor intensity are difficult to quantify; therefore, we focused on the MP1 product. MP1 levels were reduced to 58 ± 8% (Psh), 44 ± 3% (Hayan-PA), and 90 ± 30% (Hayan-PB). SPE intensity was reduced to 38 ± 12% (Psh), 43 ± 5% (Hayan-PA), and 23 ± 4% (Hayan-PB). Ser7 intensity was reduced to 9 ± 4% (Psh), 35 ± 1% (Hayan-PA), and 27 ± 13% (Hayan-PB). In general, Skanda suppressed the activation of SP7, SPE, MP1, and Ser7 by Psh, Hayan-PA, or Hayan-PB. We included the information in Fig. S2 legend.

      Fig. S3B, S7A: the three genes of the locus are inducible upon immune challenge. Have any NF-kappaB binding sites been detected at the locus. It might be relevant to repeat the experiment shown in S3B and especially S7A after a challenge with M. luteus. These experiments are definitely not essential.

      We did not look to the presence of NF-kB sites in their promoters but they have been shown to be induced and regulated by the Toll pathway (De Gregorio 2002). We did not extend our manuscript in this direction.

      The mention 'Data not shown" is used twice. Not allReview Commons-affiliated journals accept it.

      These mentions have been removed.

      Reviewer #4 (Significance (Required)): A strength of this work is the dual biochemical and genetic characterization of a SPH, an endeavor that is important to understand further the function of this class of protease-like family of secreted proteins that have been so far imperfectly studied from both perspectives (Kambris et al., CB, 2006, but see Westlake Reproducibility study on BioRxiv, Jin et al. Frontiers Immunol. 2023). Unfortunately, the two approaches fail to provide an integrated view of Skanda's function(s). A weakness is that this study does not unambiguously reveal at this stage what are the functions of Skanda in the host defense against S. aureus, let alone against other pathogens controlled to some extent by the Toll pathway or melanization. The authors have not considered a possible role in disease tolerance to S. aureus. These limitations decrease the conceptual advance of this article.

      In the revised version, we have considered a role of Skanda in resilience. This article will be of interest to investigators working on the innate immunity of insects. This reviewer is an expert in the Drosophila innate immunity field.

    2. Note: This preprint has been reviewed by subject experts for Review Commons. Content has not been altered except for formatting.

      Learn more at Review Commons


      Referee #1

      Evidence, reproducibility and clarity

      In the manuscript entitled "The serine protease homolog Skanda modulates Toll-phenoloxidase-mediated immunity in Drosophila," Vasanth et al characterize in detail a previously unstudied component of the insect immune response using first biochemical and then in vivo methods. Using proteins overexpressed and purified from insect cells, the authors provide evidence that Skanda could be a negative regulator of the SP cascade, impacting cleavage of proHayan and proPsh, and consequently Toll pathway and PPO1 activation. This work reaches further by transposing these findings into the D. melanogaster in vivo model. Here, however, the picture becomes more confusing as Skanda at native levels does not appear to regulate either the Toll pathway or the melanization cascade. Only one strong phenotype was identified in that decreased expression of Skanda increased susceptibility to S. aureus infection while increased expression decreased susceptibility. The mechanism for this remains unclear. To their credit, the authors carry out an in-depth analysis to rule out all the obvious possibilities. In the discussion, the authors explore the basis of discrepancies between their biochemical and genetic findings. We would suggest that an additional one to consider is differing roles or behaviors of Skanda in the microenvironments of the local site of injury (where S. aureus may be contained when it is tolerated) and the hemolymph. In summary, this is a valuable analysis of the innate immune component Skanda whose role has become somewhat clearer through these studies, but still remains obscure.

      Major Comments

      • To assess bimodal distribution of bacterial loads within single flies in Fig 6E, authors should either: increase the sample size to allow for proper statistical assessment of different distributions among genotypes, specifically between w1118 and skanda_d107; or, provide a modelling framework for statistical testing. Otherwise, the present results seem insufficient to conclude that Skanda is playing a role in resistance to S. aureus.
      • Another way to assess a role for tolerance in the Skanda mutant would be to measure BLUDs (https://doi.org/10.7554/eLife.28298 ) and/or transcription of CrebA (https://doi.org/10.1371/journal.ppat.1006847).
      • The error bars on qRT-PCR datasets are large, the data points are not shown so we do not know how many replicates were included in the graphs (Fig 5 B and C, Fig 6C, Fig 7 A and B, and Fig 8B). Bar plots are not the most faithful reproduction of biological datasets, as they can hinder significant information regarding datapoints distribution and variation (Beyond Bar and Line Graphs: Time for a New Data Presentation Paradigm | PLOS Biology). We advise that, particularly in the case of datasets such as qRT-PCR, the final values of fold change are represented with individual dots, with the mean value clearly represented, whether with or without the additional bar graph. Furthermore, no statistical tests were applied to determine significance. Data points should be shown and appropriate statistical tests should be applied. The number of biological replicates should be included in the analysis and the statistical test applied should be noted in the figure legends.
      • Although there are claims of Skanda conferring resistance to S. aureus infection, only Drs levels are tested. These conclusions could be strengthened by assessing expression levels of additional AMPs.

      Minor Comments

      • Parag. 1: (data not shown) should be removed and if possible AlphaFold prediction of skanda conformation added. Alternatively, remove sentence.
      • Parg. 3: 1000 mL? why not 1L?
      • Parag. 5: , in last sentence that should be .
      • Parag. 6: "a role at the same position..." does not convey the correct message< replace with equivalent?
      • Figure axes (5D, 5E, 6D, etc...) of melanization assays are wrongly named "% melanisation", with "s"
      • Parag. 21: compound mutants (if correctly interpreted as dataset presented in Fig. 8B) were tested at 6h, 24h and 48h, and not 32h, as written in the text
      • Results section "skanda is not mandatory for the activation of the Toll pathway" adopts a literal translation which would probably be better phrased as "is not essential"
      • Discussion parag. 2: "Skanda exhibits..."
      • Discussion last parag: "..., but also underlies..."
      • It has been evidenced that

      Additional comments:

      • The sentence on page 2 beginning with "Upon binding, these PRRs..." is very long and difficult to follow. This should be rewritten.
      • In many places in the manuscript bacterial "dose" is used in place of bacterial burden. The dose is the amount of a substance or bacterium given to the animal.
      • Page 11: Skanda is described as a placeholder when I think a (competitive) inhibitor would be more appropriate.

      Referee cross-commenting

      I agree with the comments of the other reviewers.

      Significance

      Strengths: The authors take a multi-disciplinary biochemical and in vivo approach to understand the molecular interactions among SPs and SPHs and thereby uncover the role of the protein Skanda that might otherwise not have been appreciated. They have made extensive use of novel transgenic fly lines, generated in the context of this study, and have thoroughly tested their specificity and cis-acting potential. These will provide a resource to the field. In addition to the new description of Skanda, these findings strengthen previous knowledge regarding systemic infections with different bacteria (M. luteus, S. aureus) and reproduce the known redundancies of Psh and Hayan modes of action. Moreover, this research is relevant for the expansion of basic knowledge on innate immunity, particularly in the field of insect-pathogen interactions, making use of S. frugiperda cell lines and D. melanogaster adults and larvae. Although not at the focus of this work, the evolutionary conserved nature of these aspects of innate immunity across these two distant species enhance the importance of these findings.

      Weaknesses: Some assays do not include enough biological replicates and others do not have enough information on how many biological replicates were performed. Therefore, the conclusions drawn are difficult to assess. Lack of statistical analysis on the qPCR experiments complicates the interpretation of results.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      We thank the reviewers for their insightful comments.Please find below a point-by-point response.

      • As the authors acknowledge in the section at the end of the discussion (Limitations of this study) it is not established that LIN-15A has a cell-autonomous function in Y-to-PDA transdifferentiation. Given that LIN-15A has a cell non-autonomous function in vulval development (Herman and Hedgecock, 1990) it is possible that its function here could also be. The authors have used an egl-5 promoter to rescue lin-15A through expression in rectal cells; however, all these cells are in a neighborhood. The lack of a promoter that is specfic for Y has impeded answering this question (a standard genetic mosaic analysis would be problematic because of the incomplete penetrance of the mutation). Although this issue is addressed in the section at the end of the Discussion, I think most readers would like to see this acknowledged earlier in the presentation, perhaps after describing the egl-5 rescue experiment.

      We thank the reviewer for this comment and agree that our data do not formally demonstrate a Y-cell autonomous role for LIN-15A during Y-to-PDA transdifferentiation, as we discussed in the manuscript. As suggested, we have modified the Results section immediately after the egl-5 rescue experiment to explicitly acknowledge this limitation early-on (see p8, l168-170) and retained the discussion in the "Limitations of the study" section.

      • The experiment shown in Figure S1B is unconvincing. To show that they are detecting a LIN-15A-LIN-56 heterodimer, the authors need to show that antibody tags to both proteins detect the band. Mass spectrometry or biochemical purifications would also be helpful. As it is, they show the protein(s) detected depend genetically on lin-56 and lin-15A. It was also unclear what the other bands were in the mutant backgrounds

      We agree that the experiment shown in Figure S1B does not provide sufficient evidence to conclusively demonstrate the existence of a LIN-15A-LIN-56 heterodimer. While the detected species depend genetically on both lin-15A and lin-56, we agree that additional controls, such as detection through reciprocal tagging, biochemical purification, or mass spectrometry, would be required to firmly establish the molecular nature of the complex and to interpret the additional bands observed in the mutant backgrounds. As this experiment is not essential to the conclusions of the manuscript, we have removed Figure S1B and the associated statements from the revised version.

      • Lines 207-212. The authors are making an argument that LIN-15A and LIN-56 function as "Licensers" not "Drivers" because they are not strictly required but appear to facilitate the process. Could it be that LIN-15A and LIN-56 function as "Drivers" but that in their absence the fidelity of the process is compromise? There are many ways that genetic redundancy can be manifested at biochemical levels and the concern is that there are other interpretations of the data. In this regard, the authors should consider rewriting the Abstract to focus on the genetic results underpinning the work. The current version or the Abstract focuses on an interpretation of the data, not the data itself.

      We thank the reviewer for this thoughtful comment. We agree that the Driver/Licenser terminology represents an interpretation of the genetic data and that additional activities acting alongside LIN-15A may exist: lin-15A alleles used in this study correspond to null alleles - that is total loss of lin-15A activity - and approximately half of the animals still successfully undergo Y-to-PDA transdifferentiation in these null mutants. Thus, lin-15A activity is either not strictly required to facilitate the initiation of the process (e.g., threshold model). Or this may point to other factors (than lin-56) able to somewhat compensate for lin-15A absence and that remain to be identified. In line with this interpretation, while we retrieved several alleles for some of the genes identified in our forward genetic screen (in which lin-15A was identified), that screen may not have been saturated. Note that both these hypotheses are compatible with a role for LIN-15A as a licenser of the initiation of the process.

      Importantly, our distinction between "Drivers" and "Licensers" is not solely based on the incomplete penetrance of lin-15A and lin-56 null mutants. First, the distinction reflects the different biological roles inferred from our genetic analyses. The previously characterized factors CEH-6, SOX-2, SEM-4, EGL-27, EGL-5 and HLH-16 are conserved plasticity factors that promote the initiation of transdifferentiation. Their loss results in a complete, or near-complete, failure of Y-to-PDA initiation, and they act within a common plasticity-promoting network. By contrast, LIN-15A and LIN-56 define a genetically distinct pathway. They are neither upstream nor downstream of the Driver cassette, and display additive interactions with partial Driver mutants. Second, loss of LIN-15A does not affect the fidelity or outcome of transdifferentiation. In all defective animals examined, the Y cell retains its normal position, morphology and rectal markers, indicating a failure to initiate the process rather than the production of an aberrant cell type. Third, the fact that a core Driver set is involved in different transdifferentiation events (ie Y-to-PDA and K-to-DVB) but not lin-15A or lin-56 further argues against LIN-15A acting as a Driver. And finally, and most importantly, lin-15A and lin-56 antagonize SynMuvB chromatin regulators known to safeguard differentiated cell identities, while the Drivers do not. In fact, the transdifferentiation process is mostly restored in some lin-15A; SynMuvB double null mutants, suggesting that LIN-15A main function is to block these genes activities. We therefore favor a model in which the Drivers cassette triggers transdifferentiation, whereas LIN-15A and LIN-56 facilitate the process by alleviating inhibitory constraints imposed by identity-safeguarding mechanisms. We have reformulated this in the manuscript in order to make it clearer and also clarified how we define Drivers and Licencers activitities (see p10, l211-217; p13, l283-289 and p14 l312-315). We have further reformulated the abstract to integrate the reviewer's comments.

      Minor Points

      __ 1. Line 279. The authors state that "LIN-15A becomes dispensable when member of the SynMuvB factors are absent." This statement is not completely accurate as the suppression is incomplete.__

      • Addressed, the statement has been reworded in the revised version (see p13, l285-286)

      2. Line 294. The number in Tagble S1 is 58.8% not 65%

      • Addressed, thank you for spotting this, Table S1 was correct, and the typo in the Results section was corrected (see p15, l319).

      3__. Lines 300-301. I couldn't find the data for lin-40. __

      __- __The data can be found in Fig. 3Bii (which we have more clearly indicated in the text) and SI table 1.

      __. Line 363. Should be "represses cell cycle genes." __

      - Addressed

      __5. Line 862. AJM-1 is not a tight junction component. AJM-1 is best described as a component of apical junctions. __

      __- __Absolutely ! Addressed

      Reviewer 2

      • The conclusions derived from the presented data are generally comprehensible but should be phrased more carefully to grant full legitimacy. The reason is that the central mechanistic claim that LIN-15A licenses Td by antagonizing most of the SynMuvBs chromatin factors, including DREAM, rests on whole-animal ChIP-seq that cannot resolve the Y cell. The authors acknowledge that "it was not technically feasible to purify sufficient Y cells for analysis" and therefore use synchronized unstarved L1 whole-animal lysates. This is certainly legitimate, but demands more tact when using such a conclusion as the headline claim.

      We thank the reviewer for this important comment and agree that the mechanistic conclusions drawn from the ChIP-seq data should be presented more cautiously. As noted by the referee and in the manuscript, it was not technically feasible to isolate sufficient Y cells for chromatin profiling and therefore all ChIP-seq experiments were performed on synchronized whole-animal L1 populations. We agree that these experiments cannot directly establish the mechanism operating in the Y cell. Rather, our genetic analyses demonstrate that LIN-15A antagonizes identity-safeguarding SynMuvB factors during Y-to-PDA transdifferentiation. The ChIP-seq data provide an additional and independent line of evidence suggesting that this antagonism may involve modulation of DREAM chromatin occupancy. We have rephrased to state this more clearly. We thus have revised the Abstract (see p2, l7), Introduction (p6, L110-112), Results (see p18-19, l405-420) and Discussion (see p23-24, l523-542 and p25 l566-575) to more clearly separate the conclusions supported by the genetic analyses from the mechanistic interpretation suggested by the ChIP-seq data. We further clarify that the relevance of this mechanism to the Y cell remains a hypothesis consistent with, but not directly demonstrated by, the available data.

      • Also, in the context of the ChIP-Seq experiments, it is understandable that it could not be conducted in a cell-specific manner, but two duplicates in some ChIP-Seq experiments (as stated in the material and methods) is below standard.

      We thank the reviewer for this comment and agree that two biological replicates represent the lower end of what is generally desirable for ChIP-seq analyses. To clarify, more biological samples were initially generated than are represented in the final analysis. In total, five independent biological preparations were performed for each genotype. However, the experimental design imposed substantial technical constraints. Because the experiments required tightly synchronized fed L1 populations (ie, not using a starvation step), standard synchronization procedures could not be used and animals instead had to be collected through successive hatch pulses, resulting in considerably lower yields. Combined with the mutant backgrounds analyzed, this led to variable ChIP-seq quality across preparations. To ensure robustness, we restricted the final analyses to datasets that passed all predefined quality-control criteria. As a result, some conditions were ultimately represented by only two high-quality biological replicates. We agree that this limitation should be made more explicit and have added this information in the Materials and Methods section (p35 l774-779). Despite the reduced number of replicates retained for some conditions, the genome-wide binding patterns observed for LIN-15B and LIN-35 in wild-type animals closely recapitulated those reported previously by the Ahringer laboratory (Gal et al., 2022; SI table 2), supporting the overall robustness and biological validity of the datasets used in this study. More generally, we have also tempered the interpretation of the ChIP-seq experiments throughout the manuscript. We view these data as supportive evidence consistent with a chromatin-level mechanism, rather than as definitive mechanistic proof, and have revised the text to reflect this more clearly.

      • Regarding the genetic interactions with met-2: as MET-2 works in concert with other SET domain proteins, such as SET-25, and also HPL-2, is there a possibility they may be implicated?

      We also considered the possibility that the interaction observed with MET-2 could reflect a broader involvement of the H3K9 methylation machinery, given the well-established functional relationships between MET-2 and other SET domain proteins. To address this possibility, we tested whether SET-25 and SET-32 losses suppressed the lin-15A phenotype. In contrast to met-2 loss-of-function, neither set-25 nor set-32 mutations modified the transdifferentiation defects observed in lin-15A mutants. These observations suggest that the interaction is not a general property of all MET-2-associated SET domain proteins and may instead reflect a more specific role for MET-2 in this context, although we have not tested triple mutant combinations, such as met-2; set-25; lin-15A or met-2; set-32; lin-15A, and therefore cannot exclude additional contributions from these factors. However, based on the available genetic evidence, our data support a model in which the phenotype is more closely linked to the SynMuvB-centered identity-safeguarding machinery than to the canonical MET-2/SET pathways. We now mention these negative results p14, l290-295 and in the discussion (p22, l510-511) of the revised manuscript. HPL-2 itself was tested alongside the other SynMuvBs, as previously reported to be a SynMuvB (Fig. 4Ci). Loss of HPL-2 had the same effect than loss of the other SynMuvBs. Together these data further suggest that the canonical SynMuvB machinery is at play, including MET-2, but not a generic requirement for all H3K9 methyltransferases, and instead points toward a more specific role of MET-2 within the SynMuvB.

      • The fact that Y-to-PDA in males (which involves a cell division) shows the same lin-15A dependence as in hermaphrodites is informative and a bit underplayed. Since this argues against a cell-cycle-coupled mechanism (an important aspect of the reprogramming field) for LIN-15A, it is worth elaborating on this in the discussion.

      We thank the reviewer for this insightful comment and agree that this result deserves further discussion. One of our initial hypotheses was indeed that LIN-15A might be specifically required in transdifferentiation events that occur without a cell division. Cell division and DNA replication have long been proposed to facilitate cellular reprogramming by promoting the dilution or resetting of identity-safeguarding mechanisms. In this context, it was conceivable that LIN-15A and LIN-56 might compensate for the absence of such a process during hermaphrodite Y-to-PDA transdifferentiation. However, our data do not support this model. We found that LIN-15A and LIN-56 are similarly required for Y-to-PDA transdifferentiation in males, despite the fact that this event occurs through a cell division. Conversely, neither factor is required for the K-to-DVB transdifferentiation, which also occurs in the rectum at a similar developmental stage and likewise involves a cell division. Together, these observations argue that the requirement for LIN-15A is not determined by the presence or absence of cell division. Rather, they suggest that the Licensers activity is context-dependent and linked to specific cellular identities. We agree that this point also strengthens the notion that Licensers are distinct from Driver factors, which function in both Y-to-PDA and K-to-DVB transdifferentiation. We have therefore modified the discussion (see p20 l441-443 and l455-480).

      Minor: __ - in the legend of Figure 1 and other places, it should be "Fisher's exact" instead of "Fisher exact" - line 31; exhibits instead of exhibit - line 85: results instead of result - line 228: involvement instead of involvment - line 293: "of missing" in loss of lin-36 had no effect while loss ... lin-53 further - lines 297 - 299: check sentence; reads not correct - line 395: "with an increase" - line 484: "with regard"__

      • All points were all addressed in the revised version.

        Reviewer 3

      Based on the observation that LIN-15A does not affect SynMuvB expression in Y (figure S4), the authors conclude that antagonism of the SynMuvBs by LIN-15A is not likely mediated by a negative control of their expression, but rather by impacting their activity. However, as suggested by the authors, antagonistic functions on the same targe genes is also a possibility. The classical approach to test this would be through expression profiling. I understand that RNA-seq on single Y cells cannot be carried out for technical reasons and that bulk RNA-seq would not be informative. Importantly, the same reasoning applies to the ChIP-seq data that is presented in support for common regulatory functions of a subset of synMuvs and LIN-15A (Figure 6 and S6), which was obtained from whole animals. The relevance of these results to the Y to PDA Td process is therefore extremely limited, as the claim that LIN-15A restricts lin-35/DREAM binding on a subset of target genes is based on a reported decrease in DREAM binding in lin-15 mutants in bulk chromatin. This is especially true as both DREAM and LIN-15A are widely expressed proteins.

      We agree with the general limitation highlighted here. As the reviewer notes, neither expression profiling nor chromatin profiling can currently be performed specifically in the Y cell due to the extremely small number of cells involved and the lack of suitable purification strategies. Consequently, the ChIP-seq experiments were performed on synchronized - and fed - whole-animal L1 populations. These data do not directly establish the mechanism operating during Y-to-PDA transdifferentiation. Rather, our conclusions are based on two distinct observations. First, the genetic analyses demonstrate an antagonistic relationship between LIN-15A and multiple SynMuvB factors during transdifferentiation. Second, the ChIP-seq experiments provide independent evidence that LIN-15A can influence DREAM chromatin occupancy at the organismal level. We interpreted these observations together as supporting a model in which the genetic antagonism may involve modulation of SynMuvB/DREAM chromatin activity. We agree, however, that the ChIP-seq data do not demonstrate that these chromatin changes occur in the Y cell itself, nor do they identify the relevant target genes involved in Y-to-PDA transdifferentiation. We have therefore revised the manuscript to more clearly distinguish between the conclusions supported directly by the genetic analyses and the mechanistic interpretation suggested by the ChIP-seq experiments. Throughout the revised version, and in the discussion in particular, we present the chromatin-level model as a hypothesis consistent with the available data rather than as a demonstrated mechanism operating in Y (see p2, l7 ; p6, l110-112 ; p18-19, l405-420 ; p23-24, l523-542 and p25 l566-575).

      In addition there are specific issues with Figure 6, which is mislabeled: upregulated and downregulated applies to gene expression, while the numbers refer to binding peaks. Why are some numbers in red (not mentioned in the legend). An example of the corresponding genome browser tracks should be shown in supplementary. Was a spike-in used to normalize data?

      We thank the reviewer for these helpful suggestions. We agree that the terminology "upregulated" and "downregulated" is potentially confusing in the context of ChIP-seq peaks. In the revised manuscript, we have replaced these terms with "up-bound" and "down-bound" in Figure 6. Regarding the red numbers, these were originally highlighted to emphasise the relatively small number of peaks showing decreased occupancy in lin-15A mutants compared to the other genotypes analyzed. However, as this information was not explained in the legend and may be confusing to readers, we have removed the color coding in the revised figure. Following the reviewer's suggestion, we have also added representative genome browser tracks in the Figure S6E to illustrate the binding changes described in Figure 6. No exogenous spike-in controls were used in these experiments. The ChIP-seq workflow was intentionally designed to closely follow that used by Gal et al. (2022), to allow direct comparison with the published LIN-15B and LIN-35 datasets. However, several observations suggest that the patterns reported here are unlikely to result from normalization artifacts alone. First, the genome-wide binding profiles obtained for LIN-15B and LIN-35 in wild-type animals closely recapitulate those reported previously, providing an independent validation of the overall quality of the datasets. In addition, the different mutant backgrounds exhibit distinct peak gain/loss profiles rather than a common directional shift that would be expected from a systematic technical bias. Nevertheless, we acknowledge the absence of spike-in controls as a limitation of the dataset and have clarified this point in the revised manuscript in the Material and Methods section (see p36 l84-805).

      Overall the discussion is highly speculative and could be shortened and refocused on the actual findings reported. For example, the fact that GO terms associated LIN-15B targets are associated with membrane processes (mentioned above) is not sufficient to speculate that LIN-15A could increase the delaminating capacities of Y by alleviating SynMuvB repression of membrane process genes.

      Our intention was to discuss possible mechanisms that could connect the observed genetic interactions to the cellular events underlying Y-to-PDA transdifferentiation. We fully agree that some of these interpretations, such as the impact of the DREAM/LIN-15A antagonisms on membrane remodeling, are purely speculative in nature. We have removed the following sentence : "In brief, the role of the Licensers would be to provide a favorable chromatin context for cellular processes that favor/install a plastic state, possibly through the modulation of membrane processes as suggested by our ChIP-seq analyses (Fig. S6). » and changed it to "In this framework, Td Licensers would facilitate transdifferentiation by alleviating identity-safeguarding chromatin states, thereby creating a permissive context for the Drivers to execute the Td program. », and have removed the paragraph describing Y delamination. More generally, we have substantially shortened and refocused the Discussion section to answer the referee's comment.

      The classical definition of a licensing factor is a protein (or complex) that allows the start of DNA replication from a replication origin. In the field of reprogramming, the term "licenser" has been applied to pioneer factors which 'license' transcriptional reprogramming by accessing chromatin to initiate a series of events, including binding of additional, non-pioneer transcription factors and additional chromatin regulators. Here the authors apply the term 'Licensers' to LIN-15A and LIN-56 as factors that facilitate the Td process. This may lead to confusion (and implications) as to what these factors are actually doing.

      We thank the reviewer for raising this point. We agree that the term "licensing" has been used in several biological contexts, including DNA replication and, more recently, cellular reprogramming, where it is often associated with pioneer factors that initiate chromatin remodeling and transcriptional changes. However, our use of the term "Licenser" is intended to describe a distinct functional concept emerging from the genetic analyses presented here. We introduced this terminology to distinguish a class of factors that facilitate transdifferentiation by alleviating identity-safeguarding mechanisms from the previously identified "Driver" factors that actively promote the cell-fate transition itself. In this framework, LIN-15A and LIN-56 are not proposed to act as pioneer factors or direct initiators of transcriptional reprogramming. Rather, the genetic data support a role in creating a permissive context for transdifferentiation by antagonizing mechanisms that oppose cell-fate change. We agree, however, that this distinction was not sufficiently defined in the original manuscript and may lead to confusion. We have therefore revised the Results and Discussion to explicitly frame it in the context of transdifferentiation ("Td Licenser"), define what we mean, and to clearly distinguish this usage from previous applications of the term in DNA replication and reprogramming studies (for instance, see p10, l211-217; p13, l283-289 and p14 l312-315).

      __Minor comments: __

      __ Abstract: why are Drivers and Licensers in capitals? How is Driver defined? __

      __- __We use capital letters to signal that these represent two conceptual categories. However, this could be changed if that impairs reading. Drivers are defined in this study as plasticity factors whose loss completely prevents Td initiation (see p10, l211-217 and p14 l312-315).

      __Figure Aii: no PDE, ajm-1::GFP positive Y cells. It is not clear how the Y cell is identified-isn't ajm-1 supposed to surround the cell? The difference between the top and bottom ajm-1:egl-5 panels is not clear to a non expert. LIN-26 panel is missing. __

      • The Y cell is identified by its location at the ventral-most position on the anterior side of the rectal slit. AJM-1 is a component of the apical junctions, hence it is expressed at the apical domain of the Y cell. The LIN-26 typo has been corrected, the marker used in this experiment is the rectal-specific gene egl-5 which labels the nucleus of the Y cell.

      2F color scheme : licensers are not in yellow but pink

      • Addressed : they now are yellow in the revised version

      Fig S1: need to provide more details about experimental conditions for WB-stage, conditions (reducing agents?), nature of Q2015 antibody. In the absence of this information hard to substantiate claim of a LIN-15/LIN-56 heterodimer in the text -

      See answer to reviewer #1 : we agree that this experiment is dispensable for the results presented in this manuscript and adds more questions than useful information, and it has been removed from the revised version.

      __Line 130. What is the nature of the LIN-56 protein? This would be useful information __

      • Addressed. We have indicated this early on in the introduction (p5, l95-96) and in the Results section (p7, l132-134). Note that little is known about LIN-56 except its association with LIN-15A in VPC specification and that is equally possesses a THAP-like C2CH motif.

      __ Line 38 yielding__

      • Addressed

      __Line 48 identities suggested by Blau and Baltimore (1991). __

      • Addressed
    1. R0:

      Reviewer #1:

      Definitely a timely article - contributing to the evidence base around HCV self-testing - an important further step in HCV elimination methods.

      Overall, I thought the paper to be well-written and close to publication. However, there were some areas that I thought needed either amendment or possible additional analysis. In particular, I have some real concerns about the description of acceptability analysis.

      Abstract:

      • Clarify that use of HCVST is referring only to antibody testing (this applies to the paper throughout)
      • I would clarify that analysis is only based on 1,995 valid participants.

      Introduction:

      • Perhaps just a quick sentence describing the nature of Nigeria's epidemic - i.e. is it generalised of population specific?

      Methods:

      • IMPORTANT: Needs to be more information about participant recruitment. The paper at times implies that the large sample size equates to high acceptability - but there is no information about the refusal rate. This would actually be a critical aspect of determining acceptability of self-testing. If you recruited 2,000 participants - but an additional 2,000 refused to participate (for example) - this would suggest self-testing is acceptable to only 50% of clients.

      • IMPORTANT: Very little in methods about your qualitative interviewing - this needs expansion.

      • IMPORTANT: There is very little satisfactory information about the acceptability questionnaire - and it's subsequent transformation into an acceptability score using factor analysis. Currently, it reads as if the audience is simply supposed to trust the authors that all methods were satisfactory, and that - therefore - results are valid. You don't specify which questions were included in factor analysis, or if it was appropriate to score each item similarly or to aggregate scores. You also don't indicate if the questions were a previously developed tool or something developed specifically for the study - or that it's even available in the supplementary data. This almost feels like a separate paper, so that your method of transforming the acceptability survey can be more properly peer reviewed.

      • IMPORTANT: For you acceptability analysis, how did you determine that 10% equals 'unacceptable'? Shouldn't people with only a 20% acceptability score be considered broadly non-accepting?? Also, considering all the noted limitations and collinearity issues with site - does it even make sense to compare by site?

      • How did you determine the 2,000 sample size?

      • The selection criteria is difficult to follow and should be clarified - perhaps with dot points?

      • Specify that "blood-based" testing means finger prick?

      • Again, need to clarify that testing is antibody, correct? Presuming also that RA information discussed testing procedures – i.e. what happened if the ST was positive and it’s interpretation?

      • How many of each test did you have? Why did you not have enough tests so that any participant could choose whichever test they wanted for the entirety of the study period?

      • Presuming drug abstention is not a requirement to treatment initiation?

      • For the exposure variables - is 'education' referring to attendance or completion?

      • You explain later in the paper that participants could be part of multiple KP categories - but probably good to explain in the "exposure variables" section.

      Results:

      • Non-considering possible refusal rates - your results suggest that of 2000 participants, essentially 100% completed the self-test? This is a remarkable outcome – particularly for the 849 participants who were unobserved!

      • "Five participants 255 took HCVST kits home but did not return to complete the endline survey, therefore no demographic or clinical information was recorded for those participants, and they were excluded from the analysis." Why wasn't demographic information collected at baseline, as would be usual?

      • Suggest re-doing Figure 2 - removing the column for those 'tested' as this is inherently 100% and that every subsequent column is small by comparison. Add percentages to columns. Could consider comparing clinic types.

      • Following on from methods comments regarding the acceptability data - you don't specify the number of participants falling within the purported "least 10%". The sentence describing qual data “Acceptability appeared to be more associated with clients characteristics than facility type…” - also suggests there are issues with your acceptability analysis methods.

      • Did you consider differences across test type? In particular those who elected to conduct unobserved self-testing? What were the characteristics of these individuals?

      • Sentence on page 18 “Observations about test type choice are thus descriptive only, not inferential” should probably be in the methods – following discussion of how tests ran out.

      Discussion:

      • How much do you think the demographic differences between facility-type clients impacted outcomes? For example – OSS clients are much younger and more males.

      • OSS clients used more ‘blood’ tests. Did this have any impact on acceptability?

      • The statement about PWID ‘surprise’ may be misplaced. How many PWID did you approach, who refused to participate? How does the number recruited compare against the OSS client list? Also, it’s 28% (PWID) compared to (31% MSM) – are you surprised about MSM?

      • “Additionally, we did not assess whether the acceptability scale functioned equivalently across ART and OSS settings, so observed differences should be interpreted with that caveat.” - what does this statement mean? Is this another potential flaw of the acceptability scale?

      • I'd include explanation of the SVR12 limitation somewhere higher in the paper.

      AE:

      This is an important report on controlling viral hepatitis and, in particular, hepatitis C (HCV) in Africa. The report is well written and is also a good example of how to conduct feasibility studies. The authors managed to set a clear and moving background. However, there are a few issues that must be addressed.

      Issues: 1. Please indicate how the 2000 sample size was determined.

      1. Line 132 is a bit unclear. Can it be rewritten?

      2. About the acceptability, it needs to be recognized that this instrument has never been tested in this population before. The EFA and CFA results are good, but it is not a guarantee that this will work forever.

      The continuous acceptability score is used/analysed in two ways. One is the study of distribution using histograms and boxplots. The other one is through a dichotomization of the acceptability and use it as an outcome for the binary logistic regression. • For the study of distribution, as in figure 3, I would suggest a cumulative counts plot. On such plot, you could put counts on the left y-axis, and percentages on the right y-axis. Heights are easier to read than histogram areas. • About the “least 10% of acceptability” concept. Can we call this the lowest decile acceptability score? • Figure 4 is successfully to show the distribution of the acceptability score for a few key variables. However, because it is based on quartiles, it fails to tell us the proportion of observations below the line of the overall lowest decile. I suggest adding a table with such proportions.

      1. About the cascade. Table 2 is accurate and quite informative. Figure 2 is a bit deceptive, because it suggests that the denominators for each step are accurate [which is not true for the last step, for example]. So I suggest i) to keep table 2 and ii) compute accurate proportions (in percentages) and add 95% confidence intervals for the cascade plot.

      R1:

      Reviewer #1:

      • Authors have responded to my comments

      • Given the noted limitation of not being able to demonstrate comprehensive acceptability (without refusal rates) – it’s advised to further review the paper for instances of potential overstatement of findings. For example, on page 20. “…while both…facilities demonstrated high engagement…” – 1,000 clients is definitely a lot, but without the context of how many overall clinic clients there are, and how many may have refused, it would be prudent to temper the statements slightly.

      • Importantly, the additional information provided by the authors about their bespoke acceptability measure gives me greater reservations – considering it has no prior validity testing. This is not addressed in the limitations and definitely should be.

      R2:

      Reviewer #1:

      Thank you to the authors for addressing my comments.

    1. Author response:

      The following is the authors’ response to the original reviews

      Summary of revision for all referees:

      We thank referees for their constructive comments. To address their concerns, we now performed additional statistical analyses integrating both paired and unpaired data, performed positive controls for comparisons between NH- and CI- evoked iEEG measurements, developed tools for measuring and collected new experimental data on forward masking ECAP measurements in CI implanted rats (N=3), and reworked both manuscript text and figures to improve clarity. These most significant changes are summarized here, and a complete list of responses to reviewers and corresponding changes will follow.

      Summary of major changes to revised manuscript:

      (1) Statistical treatment of paired vs unpaired recordings using mixed-effects models (updates to all manuscript figures that compare NH vs CI); this largely confirmed the results reported in our original submission.

      (2) New analysis, controlling for information-theoretic cross-modality comparison (i.e., training with tone- and testing with cochlear implant-evoked iEEG measures, Fig. 8).

      (3) Clarification of methods (Supplemental Fig. 2 & manuscript text)

      (4) Additional experiments testing peripheral tuning of our 8-channel CI rodent model via forward masking ECAP measures across 3 animals (N=3, Supplemental Fig. 1)

      (5) Detailed response addressing robustness of tonotopy in NH and CI animals

      Public Reviews:

      Reviewer #1 (Public Review):

      Strengths:

      The study poses a timely, clinically relevant question with clear implications for CI strategy. The analytical toolkit is appropriate: µECoG captures mesoscale patterns; TCA offers a transparent separation of spatial and temporal structure; and mutual-information decoding provides an interpretable measure of single-trial discriminability. Within-subject recordings in a subset of animals, in principle, help isolate modality effects from inter-animal variability. Where analyses are most direct, the acoustic condition yields higher single-trial decoding accuracy, which is a meaningful and clearly presented result.

      We appreciate the comments on the strengths of our analytic approaches.

      Weaknesses:

      Parts of the statistical treatment do not match the data structure: some comparisons mix paired and unpaired animals but are analysed as fully paired, raising concerns about misestimated uncertainty.

      Please see our response to specific comment #2 above. In short, we agree with this critique of our original analyses, and in our revised manuscript we re-analyzed all NH vs. CI comparisons using linear mixed effects models that incorporate both paired and unpaired observations within a single framework. This allows us to include all animals, account for within-animal dependence for paired experiments (normal hearing and cochlear implant data from the same animal when available), and to align the statistical tests with the data shown in the figures. In almost every case, the mixed effects models confirm our original conclusions. Two comparisons that were previously nonsignificant now reach criterion for statistical significance (Fig. 2E, p=0.048 and Fig. 6F, p=0.027). We updated the manuscript to report these values and to clarify the use of mixed effects modeling in the methods under the section titled, “Linear mixed effects modeling.”

      Methodological reporting is incomplete in places; essential parameters for both acoustic and electrical stimulation, as well as objective verification of implantation and deafening, are not described with sufficient detail to support confident interpretation or replication.

      Please see our response to comment #5 below. We have revised our manuscript to now include this information in the methods.

      Figure-level clarity also undermines the message. In Figure 2, non-significant slopes for CI, repeated identification of a single "best channel," mismatched axes, and unclear distinctions between example and averaged panels make the assertion of spatial organisation unconvincing; importantly, the normal-hearing panels also do not display tonotopy as clearly as expected, which weakens the key contrast the paper seeks to establish.

      This is an important point, thanks- please see responses to comment #1 above. We note that conventional tonotopic maps in auditory cortex are characteristic frequency maps, i.e., maps of topographic organization for responses to lowest-threshold stimuli (often presented around 20-50 dB SPL). Our maps were constructed from stimuli presented at 70 dB SPL, thus blunting crisp tonotopy to some degree. Furthermore, we quantified spatial organization using a previously published method from the Polley lab (Romero & Hight et al. 2020), in which local tonotopic gradient vectors (magnitude and direction) were computed from GCaMP responses at each pixel and projected onto a unit circle. Mean vector strength across all pixels was then compared to a shuffled distribution as a measure of tonotopic organization. We applied the same procedure to our iEEG best-frequency and best-channel maps. Both map types yielded mean vector strengths that were substantially larger than those derived from shuffled maps (p < 10<sup>-10</sup>), indicating that our maps have a consistent tonotopic (for BFs) or cochleotopic (for CI channels) organization that is highly unlikely to arise by chance. This is now included in our revised manuscript.

      Finally, the decoding claims would be strengthened by simple internal controls, such as within modality train/test splits and decoding on raw ERP/high-gamma features to demonstrate that poor cross-modal transfer reflects genuine differences in the underlying responses rather than limitations of the modelling pipeline.

      Please see our response to comment #12 below. In short, we have now included this analysis in revised Figure 8.

      Reviewer #2 (Public Review):

      Strengths:

      The study includes interesting analyses of the sound and cochlear implant representation structure based on decoders.

      We appreciate the comment on how interesting our analyses are, thanks!

      Weaknesses:

      The observation that responses to cochlear implant stimulation (stimulation) are spatially organized is not new (e.g., Adenis et al. 2024).

      We agree that it is not particularly novel to report that there is spatial organization to cochlear implant stimulation. However, we believe that our direct comparisons (when possible, within animal) between normal-hearing and cochlear implant modality maps is unusual in the literature, including asking how decoders based on one set of responses might apply to responses evoked from the other modality. Adenis et al. (2024) is a fantastic study of pulse shape and monopolar vs bipolar stimulation modes with a 6-channel implant in guinea pig, but as far as we can tell this study does also not compare normal hearing maps prior to deafening and implantation to the cochlear implant maps in the same animals.

      The claim that spatial and temporal dimensions contribute information about the sound is also not new; there is a large literature on this topic. Moreover, the results shown here are extremely weak. They show similar levels of information in the spatial and temporal dimensions, and no synergy between the two dimensions. This is however, likely the consequence of high measurement noise leading to poor accuracy in the information estimates, as the authors state.

      Good point, please see our response to comment #1 below.

      The main claim of the study - the mismatch between cochlear implant and sound representation - is not supported. The responses to each modality are measured in different animals. The authors do not show that they actually can compare representations across animals (e.g., for the same sounds). Without this positive control, there is no reason to think that it is possible to decode from one animal with a decoder trained on another, and the negative result shown by the authors is therefore not surprising.

      Good point, thanks- please see our response to comment #2 below, where we describe this new control we have added.

      Reviewer #3 (Public Review):

      Strengths:

      The model combining micro-eCoG and cochlear implantation and the methodology to extract both the Event Related Potentials (ERPs) and High-Gammas (HGs) is very well designed and appropriately analyzed. Likewise, the PCA-LDA and TCA-LDA are powerful tools that take full advantage of the information provided by the cortical ensembles. The overall structure of the paper, with a paced and exhaustive progress through each step and evolution of the decoder, is very appreciable and easy to follow. The exploration of single-trial encoding and stimulus identity through temporal and spatial domains is providing new avenues to characterize the cortical responses to CI stimulations and their central representation. The fact that single trials suffice to decode the stimulus identity regardless of their modality is of great interest and noteworthy. Although the authors confirm that iEEG remains difficult to transpose in the clinic, the insights provided by the study confirm the potential benefit of using central decoders to help in clinic settings… the reviewer wants to reiterate that the study proposed by Hight et al. is well constructed, relevant to the field, and that the overall proposal of improving patient performances and helping their adaptation in the first months of CI use by studying central responses should be pursued as it might help establish new guidelines or create new clinical tools.

      We thank the Reviewer for the positive comments about the thoroughness of our analyses and clear organization of our manuscript.

      Weaknesses:

      The conclusion of the paper, especially the concept of distinct cortical encoding for each modality, is unfortunately partially supported by the results, as the authors did not adequately consider fundamental limitations of CI-related stimulation. First, the reviewer assumed that the authors stimulated in a Monopolar mode, which, albeit being clinically relevant, notoriously generates a high current spread in rodent models.

      Thanks, this is an important potential concern. Please see our response to comment #5 of Referee 1 and responses to comment #3 below. We agree that monopolar stimulation would be expected to be less spatially specific than bipolar or multipolar modes. However, we chose monopolar stimulation because it is the main clinical configuration in human CI users and therefore most relevant for translational purposes. For our revised manuscript, we made new ECAP measurements of peripheral (spatial and temporal) tuning via a forward masking paradigm and demonstrate that monopolar is effectively tuned (Supplemental Fig. 2). Together with additional single-animal maps in Supplementary Figure 3, together with our vector-strength analysis (Response Fig. 2), demonstrate that even under acute monopolar stimulation we observe structured cochleotopic organization in cortex, rather than the extremely low-pass patterns one might expect if monopolar spread was a major contaminant.

      Second, comparing the averaged BF maps for iEEG (Figure 2A, C), BFs ranged from 4 to 16kHz with a predominance of 4kHz BFs. The lack of BFs at higher frequencies hints at a potential location mismatch between the frequency range sampled at the level of the cortex (low to medium frequencies) and the frequency range covered by the CI inserted mostly in the first turn-and-a-half of the cochlea (high to medium frequencies). Looking at Figure 2F (and to some extent 2A), most of the CI electrodes elicited responses around the 4kHz regions, and averaged maps show a predominance of CI-3-4 across the cortex (Figure 2C, H) from areas with 4kHz BF to areas with 16kHz BF. It is doubtful that CI-3-4 are located near the 4kHz region based on Müller's work (1991) on the frequency representation in the rat cochlea.

      Please see our responses to comment #3 below.

      Taken together with the Pearsons correlations being flat, the decoder examples showing a strong ability to identify CI-4 and 3 and the Fig-8D, E presenting a strong prediction of 4kHz and 8kHz for all the CI electrodes when using a pure tone trained decoder, it is possible that current spread ended stimulating indistinctly higher turns of the cochlea or even the modiolus in a non-specific manner, greatly reducing (or smearing) the place-coding/frequency resolution of each electrode, which in turn could explain the coarse topographic (or coarsely tonotopic according to the manuscript) organization of the cortical responses. Thus, the conclusion that there are distinct encodings for each modality is biased, as it might not account for monopolar smearing. To that end, and since it is the study's main message and title, it would have benefited from having a subgroup of animals using bipolar stimulations (or any focused strategy since they provide reduced current spread) to compare the spatial organization of iEEG responses and the performances of the different decoders to dismiss current spread and strengthen their conclusion.

      Please see our responses to comment #4 below as well as our responses related to monopolar vs bipolar stimulation. We agree that for future studies, it will be important to do a heads-on comparison of the differences between bipolar and monopolar stimulation depending on electrode location and stimulation intensity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      We thank the reviewer for commenting on the strengths of our manuscript, including appreciating the power and timeliness of our approach.

      (1a) Figure 2 does not convincingly support the claim that "tone-evoked and CI-evoked iEEG measurements are spatially organized," particularly for CI data: Figure 2C repeatedly highlights the same "best channel," and the slopes in Figures 2B and 2G are non-significant; there are also discrepancies between panels (A vs. C, F vs. H) and mismatched frequency ranges (0-16 kHz vs. up to 32 kHz), which should be clarified as exemplar versus averaged displays and harmonized in scale.

      (First we note that Reviewer 3 also raised related concerns about the robustness of tonotopy in our iEEG data.) We address these by comparing our maps to previously published tonotopic maps, and using an established quantitative analysis of tonotopic strength from Romero & Hight et al. (2020).

      First, to place our tone-evoked iEEG maps in context, we overlaid them on the same spatial scale and orientation as both single-unit tonotopy in rat primary auditory cortex (A1) from Polley et al. (2006) and iEEG maps obtained with the same surface array in Insanally et al. (2016). The rostral–caudal and dorsal–ventral axes and cortical extents are matched across panels. Our best-frequency maps (Figure 2C) qualitatively recapitulate the high-to-low frequency gradient and spatial layout reported in both of these prior studies, supporting our claim that tone-evoked iEEG captures canonical mesoscale tonotopy. We have updated the manuscript results section to directly reference these two studies, “The area and orientations of tone-evoked maps qualitatively match those published from single unit recordings (Polley et al. 2006) and published using similar iEEG arrays (Insanally et al. 2016).”

      Second, to quantify tonotopy in a way that is directly comparable to previous work, we reproduced the analysis of Romero & Hight et al. (2020), who examined tone-evoked GCaMP signals (Romero & Hight et al. (2020)). In that paper, local tonotopic gradient vectors (magnitude and direction) were computed at each pixel and projected onto a unit circle; the mean vector strength across all pixels was then compared to a shuffled distribution as a measure of tonotopic organization. We applied the same procedure to our iEEG best-frequency and best-channel maps (Fig. 2C-E). Both map types yielded mean vector strengths that were substantially larger than those derived from shuffled maps (p < 10<sup>-10</sup>), indicating that our maps have a consistent tonotopic (for BFs) or cochleotopic (for CI channels) organization that is highly unlikely to arise by chance. We cite this paper for these analyses related to Figure 2.

      (1b) Figure 2C repeatedly highlights the same ‘best channel’

      We agree that many CI-evoked maps are dominated by a single channel, as seen in our exemplar and in the additional animals shown in new Supplemental Fig. 3. In Fig. 2C, channel 5 emerges as the dominant best channel, as CI-evoked activity in this animal is broad and is strongest for channel 5 (Fig. 2A). This reflects a feature of iEEG signals rather than a plotting artifact. Biophysically, iEEG reflects spatially summed local field potentials that low-pass filter underlying neural activity; these far-field signals aggregate excitatory and inhibitory processes and are not expected to show the sharp single-neuron tuning seen in spike recordings. As a result, broad peaks centered on the most strongly driven channels are expected. We have added text in the results section discussing these limitations, overall maps reduced from iEEG responses were similar in size and orientation compared to single unit maps, “albeit at coarser gradients likely due to aggregate recordings of excitatory and inhibitory activity and low-pass filtering due to potentials originating far from recording sites.” We also added in the results section the comparison of spatial correlations (Fig. 2B,G) at the extremes of stimulus separation “electrode separations (CI 1 vs ≥5 electrodes, ERP: p=0.01, HG: p=0.04)” as analyzed by linear mixed effects models.

      (1c) Mismatched frequency ranges

      We constricted the range of frequencies plotted in some panels (e.g., Fig. 2C from 1.4-32 kHz to 1.4-16 kHz) to emphasize the compressed range of tonotopic gradients and patterns.

      (1d) The slopes in Figures 2B and 2G are non-significant

      We agree that non-significant group-level slopes indicate that CI-evoked tonotopy is weaker than tone-evoked tonotopy, and we now emphasize this point. At the same time, the data exhibit systematic structure: for both ERP and HG, mean spatial correlations decline monotonically with increasing CI channel separation (Fig. 2B,G). We also directly compared spatial correlations at the extremes of stimulus separations (1 vs. ≥5-channel separation) and found a significant difference. This is updated in the manuscript as: “At the extremes, the spatial correlations were always higher for small vs. large tone separations (NH 0.5 vs ≥3.5 octaves, ERP: p<10<sup>-4</sup>, HG: p<10<sup>-4</sup> Student’s one-tailed t-test) and electrode separations (CI 1 vs ≥5 electrodes, ERP: p=0.01, HG: p=0.04).”. Together with the strong deviation from shuffled maps in the vector-strength analysis (Fig. 2E), we argue that analysis of spatial correlations indicates that CI-evoked maps are not random but reflect a coarse underlying gradient. In addition, as tone-evoked maps exhibit tonotopy, we asked if CI stimulation itself is at least spatially tuned in the periphery. Using ECAPs with a forward-masking paradigm (new Supplemental Fig. 1), we show that probe-evoked ECAPs are significantly more suppressed by adjacent than by distant maskers (N = 3), demonstrating functional spatial tuning of CI electrodes in the cochlea. We have also replotted these results in comparison with the same measurements from a human CI user (Author response image 1). This supports the interpretation that peripheral input is spatially specific and that the weaker cortical cochleotopy likely reflects the properties and resolution of iEEG and acute CI stimulation rather than a complete absence of spatial organization. Overall, the new comparative figures and analyses are intended to make transparent that (i) iEEG robustly captures tonotopy for acoustic tones, and (ii) CI-evoked CI-evoked responses exhibit coarser, but statistically non-random, cochleotopic organization.

      Author response image 1.

      Here, we compare data from the new Supplemental Figure 1C,D with human data (N=1) for spatial & temporal tuning in the periphery, as assessed by forward masking ECAP measurements. A) Spatial tuning functions were averaged across all probe electrodes and 3 animals (left) and 1 human subject (right) (black, mean; gray: s.e.m..; orange, average of individual subjects). B) Temporal tuning functions were averaged across all probe electrodes and 3 animals (left) and 1 human subject (right) (black, mean; gray, s.e.m.; orange, average of individual subjects). Note: human subject is the first-author, a long-term cochlear implant user (>10 years) with significant open set speech perception.

      (2) The statistical approach is inappropriate where pairing is incomplete: a Student's paired two-tailed t-test is used despite not all data being paired; a linear mixed-effects model would be more suitable, whereas an unpaired test risks reduced power.

      We agree with this suggestion. As the reviewer notes (also raised by Reviewer 3), our original analyses did not fully exploit the partially paired structure of the data. In the initial submission we used paired t-tests when animals contributed both normal-hearing (NH) and CI measurements, which meant that animals with only NH or only CI data were excluded from those tests.

      To address this, we have re-analyzed all NH vs. CI comparisons using linear mixed-effects models that incorporate both paired and unpaired observations within a single framework. This approach allows us to (i) include all available animals, (ii) appropriately account for within-animal dependence when both conditions are present, and (iii) align the statistical tests with the data shown in the figures. In nearly all cases, the mixed-effects models confirm our original conclusions. Two comparisons that were previously non-significant are now significant in the positive direction: Fig. 2E (p = 0.048) and Fig. 6F (p = 0.027, linear mixed-effects models). We have updated the manuscript to report these values and to clarify the use of mixed-effects modeling in the methods under the section titled, “Linear mixed effects modeling.”

      (3a) Given the surgical complexity, objective verification of implantation and deafening is needed (e.g., eABRs for implant function and post-deafening ABR thresholds)”

      We agree that objective verification of both implant placement and deafening is critical, particularly given the surgical complexity of multichannel CI implantation in rats. Note that we previously extensively documented deafness in our cochlear implant rats with eABRs, histology of hair cell counts, and behavior (turning the implant off and seeing performance drop to chance). As we argued in Glennon et al. Nature 2023, the primary outcome measure and definition of deafness is behavioral, as anatomical and physiological markers are correlates of functional deafness but ultimately deafness must be defined in terms of behavioral performance. This is described in more detail below.

      We agree that objective verification of both implant placement and deafening is critical, particularly given the surgical complexity of multichannel CI implantation in rats. Note that we previously extensively documented deafness in our cochlear implant rats with eABRs, histology of hair cell counts, and behavior (turning the implant off and seeing performance drop to chance). As we argued in Glennon et al. Nature 2023, the primary outcome measure and definition of deafness is behavioral, as anatomical and physiological markers are correlates of functional deafness but ultimately deafness must be defined in terms of behavioral performance. This is described in more detail below.

      Implant placement: Our primary concern during surgery is to ensure that the CI array is correctly positioned along the cochlear spiral toward the apex. As shown in Author response image 2, once the bulla is opened and the cochleostomy is made at the junction of the temporal bone and the stapedial artery, the orientation of the cochlear spiral is clearly visible under the surgical microscope. We advance the 8-channel array only in the apical direction, and we require that all 8 electrodes pass through the cochleostomy. A complete insertion of all 8 electrodes cannot be achieved with a basal-ward trajectory, so full insertion provides a strong anatomical confirmation that the array is directed apically. The white band on the array, visible just basal to the cochleostomy (Author response image 2), serves as a consistent visual marker of complete insertion. We have added text and this figure to the Methods to clarify these criteria, “We required that all eight electrodes pass through the cochleostomy, confirming that the array was inserted in the direction of the apex.”

      Verification of deafening: We also share the reviewer’s concern about confirming profound hearing loss, particularly because some CI animals were presented acoustic tones to drive individual channels. We used the same mechanical-only deafening procedure described and validated in our previous work (King et al., 2016; Glennon et al., 2023), which was chosen to minimize systemic side-effects and maximize post-surgical survival, validated in three ways:

      - Histology: In N=4 deafened animals, inner hair cell loss was ~50% and outer hair cell loss was near complete at almost 100% in all animals.

      - Physiology: For N=14 rats, acoustic ABRs were substantial before deafening but statistically similar to baseline noise after deafening.

      - Behavior: For N=16 deafened rats, behavioral performance with implant on was d′: 1.7±0.1, but when implant was turned off in a subset of sessions, performance dropped to chance (d′: −0.05±0.1, P < 0.0001).

      Author response image 2.

      Visual confirmation of a successful electrode insertion. The direction of an 8-channel array being implanted toward the apex is clear under microscope. Full insertion of all 8 channels is further confirmed by the white band’s (located after basal electrode) proximity to the cochleostomy.

      This combination of histological, physiological, and behavioral evidence indicates that the mechanical-only deafening protocol produces profound hearing loss, with no functionally relevant residual hearing at intensities equal to or greater than those used in our study (70 dB SPL). Given this prior validation under identical surgical and experimental conditions, we are confident that our CI animals were effectively deafened and that the iEEG responses we report are driven by the implant rather than by residual acoustic hearing. We now clarify this in the Methods and explicitly cite our validation: “(mechanical only, as described and validated in Glennon et al. 2023).

      (3b) One CI animal did not learn the task (Fig. 1C), potentially reflecting implantation efficacy.

      Good point, thanks. For both humans and rats, cochlear implant performance can be highly variable, reflecting a number of factors in terms of device performance, training efficacy and motivation, or other technical or biological sources of heterogeneity. We note however that not all animals included in this study were behaviorally trained, and wanted to show the full range of variable performance for the subset of animals that were trained (N=4 typical hearing and N=3 cochlear implant rats, one of the 4 trained animals lost the implant before it could be re-trained on the cochlear implant version of the task). We now highlight this range of performance variability in the results section and explain why N=4 normal-hearing and N=3 cochlear implant rats.

      (4) The behavioural paradigm and cohort accounting are unclear: Figure 1C shows four NH-trained rats, yet subsequent analyses include only two NH-trained animals, which is confusing.

      We have now clarified the relation between the behavioral cohort and the iEEG cohort in the revised manuscript. The key point is that the animals in Figure 1C are defined by their behavioral training history (NH vs CI training), whereas inclusion in the iEEG analyses is defined by the specific stimuli collected during acute recordings, and these two categorizations are not always the same. In total, four rats underwent both iEEG recordings and behavioral training. Of these four, three were subsequently deafened, implanted with chronic CIs, and trained on the CI-driven task (Fig. 1C). With respect to the acute iEEG experiments, we obtained tone-only iEEG in 1 animal, CI-only iEEG in 2 animals, and both tone- and CI-evoked iEEG in 1 animal.

      Thus, the “NH-trained” label in Figure 1C refers to behavioral training status, not to the stimulus conditions used during iEEG recordings. All iEEG measurements were acute and performed immediately after surgery (for CI animals) or in the normal-hearing condition, before any CI behavioral training. Consequently, the behavioral cohort in Figure 1C is larger than the subset of animals that contributed to specific iEEG contrasts in later figures, which explains why some panels include only two NH animals.

      To clarify this, we have added a new Supplementary Figure 2 that provides a timeline for each animal, indicating when behavioral training occurred, when deafening and implantation occurred, and which stimulus conditions (tones vs CI) were used for each iEEG recording. We kept this figure in the Supplementary section because the focus of the manuscript is on evoked iEEG measurements rather than behavior, but the revised text now explicitly refers to this schematic when describing the cohorts “The combinations of animals that underwent behavioral training and acute iEEG measurements are shown in Supplemental Fig. 2.”

      (5) Methods lack essential details: specify acoustic stimulus types and intensities, CI stimulation parameters (e.g., current/charge per phase, phase width, rate, loudness setting), and the recording state (awake vs. anaesthetised), which is only implied in the discussion.

      We agree that these details are essential, and Reviewer 3 raised similar concerns about methodological clarity. We have now expanded the Methods to specify the acoustic stimuli, CI stimulation parameters, and recording state.

      Acoustic stimuli: We now describe the acoustic stimulus set in the Methods, which references Insanally et al. (2016). Briefly, tones were pure sinusoids spanning frequencies from 1.4 to 32 kHz (half octave spaced), presented at 70 dB SPL with a duration of 50 ms with 2ms cosine-squared ramps and at a pseudorandom sequence of 1.25 Hz. These parameters are now updated in the methods under “Stimulus presentation for cortical sensory mapping in normal hearing rats.”

      CI stimulation parameters: CI stimulation used standard clinical-style monopolar mappings. We now specify in the Methods that pulses were biphasic, charge-balanced, with 8 µs interphase gaps and 25 µs /phase (total pulse width = 58 µs); stimulation rate was 900 pulses per second (pps); and current amplitude (and thus charge per phase) was set individually for each electrode based on its ECAP threshold. All stimulation levels were within normal and safe limits: charge densities remained below the Shannon limit and within the electrochemical “water window.”

      Loudness setting: In this study, CI stimuli were presented primarily at a single level—each electrode was stimulated at its ECAP threshold level for the tone-to-CI mapping experiments. We have added these details in the methods under the “Stimulus presentation for cortical sensory mapping in cochlear implanted rats” subsection.

      Recording state: All iEEG recordings reported in the manuscript were acute and performed under anesthesia. This is now stated explicitly at the start of the Methods section.

      (6) Plasticity and training effects warrant further consideration: although the manuscript reports no difference between naïve and trained rats, Figure 3 suggests greater across-trial variability for CI than NH that is not evident in the trained subset; examining relationships among behavioural performance, decoder performance, across-trial variability, and training duration would strengthen interpretation.

      We agree that plasticity and training effects are central questions for cochlear implant research and that iEEG is well suited to study how cortical representations evolve with CI use. However, the current dataset was collected mainly to compare cortical encoding of acoustic versus CI stimulation under matched, acute conditions (not necessarily after behavioral training with the implant, and we note that most studies of physiological responses to cochlear implant function in non-human species also do not incorporate aspects of training). All CI-evoked iEEG recordings were obtained immediately after implantation, before any CI-based behavioral training. As a result, any training effects reflected in the iEEG data can only arise from prior normal-hearing training, not from experience with CI stimuli themselves. Only a small subset of animals (N = 3 of 10) underwent behavioral training with cochlear implants, and their training histories (duration, performance levels, CI hardware status) are not uniform. This yields insufficient statistical power to meaningfully examine correlations among behavioral performance, decoder performance, across-trial variability, and training duration. While we note the reviewer’s observation that across-trial variability appears qualitatively different in the small, trained subset, we do not believe the current data justify strong conclusions about training-related plasticity.

      (7) Differentiating the CI rats stimulated directly or through the microphone of the speech processor -at least in the figures - would be useful to allow the reader to assess whether both stimulation strategies give rise to similar results.

      We agree that it is important to distinguish between rats stimulated directly via CI hardware and those stimulated acoustically through a speech processor. We now show in new Supplementary Figure 2, which animals received direct electrical stimulation and which were driven acoustically through the processor microphone. We also now plot tonotopic and cochleotopic maps for all CI animals in Supplementary Figure 3, with the stimulation mode indicated for each animal. As discussed in our response to comment #2 of Reviewer 3, we also provide validation that acoustic tones can be used to selectively drive individual electrodes via the speech processor. However, the sample sizes for the two stimulation strategies are small (N = 4 rats with direct CI stimulation, N = 3 rats with acoustic CI stimulation). For this reason, we have chosen not to draw strong statistical conclusions about differences between direct vs acoustic CI stimulation in the present manuscript.

      (8) Typographical error at the end of the introduction ("To this end we have designed and manufactured..."), and in the first paragraph of the Discussion ("...that both that...").”

      Thanks, we have updated the manuscript accordingly.

      (9) Inconsistent terminology: use a single form (e.g., "normal-hearing") throughout.

      Good suggestion, thanks. We have updated all main manuscript to only use normal-hearing. We found and changed two instances in which we used the acronym NH in lieu of normal-hearing, once early in the results section and once in the legend for Figure 3.

      (10) In Figure 3D (temporal), there appears to be an extra data point for the NH-trained group.

      Thank you for flagging this mis-labeling, which Reviewer 3 also pointed out. We have switched the appropriate data point in Figure 3D from ‘trained’ to ‘naïve’.

      (11) In Figure 4D, the yellow line is not defined; based on Figure 6D, it likely represents shuffled/chance performance and should be labeled accordingly (including beneath the chance line on the plots).

      We have updated Figure 6 to indicate that the yellow line does indeed reflect shuffled/chance.

      (12) Figure 8 would benefit from a control demonstrating that poor cross-modal decoding reflects train-test distribution differences rather than weak decoders (e.g., train on a subsample of NH and test on held-out NH), and from reporting decoding on raw ERP/HG features in addition to TCA-derived data.

      Good suggestion, thanks; we have now added this control. We agree that a positive control is necessary to show that poor tone→CI decoding reflects differences of underlying representations rather than a failure of the decoder or modeling approach. (Reviewer 2 raised the same point.)

      To validate our cross‑modal analysis pipeline, we re‑implemented the full procedure used in Figure 8, but instead of training on tone‑evoked responses and testing on CI‑evoked responses, we trained and tested on independent sets of tone‑evoked trials from the same animals (tone→tone). For each tone in each animal, we withheld 10 trials as a test set. Using the remaining trials, we fit the original TCA model to obtain spatial and temporal factors (Fig. 8A). We then fixed these factors and re‑optimized only the trial factors on the withheld tone‑evoked trials (Fig. 8B). The LDA decoder was trained on the trial factors from the original TCA fit and tested on the re‑optimized trial factors from the withheld trials, using the same classification pipeline as in the main analysis.

      As shown in the top panels of Figure 8C,D, this positive control yielded robust tone→tone generalization: predicted tone frequencies closely matched the actual tones, decoder performance was significantly above chance, and prediction errors were tightly clustered around the true stimulus, indicating that the decoder was tuned to tone frequency. In contrast, when we trained on tone‑evoked responses and tested on CI‑evoked responses, information transfer was markedly reduced (Fig. 8E-G).

      These results demonstrate that the TCA+decoder pipeline can reliably transfer information across independent tone‑evoked datasets, confirming that the method captures shared structure when it exists. The poor cross‑modal transfer between tone‑ and CI‑evoked activity therefore is unlikely to be due to a weak decoder or to a failure of the modeling pipeline, but instead reflects a genuine mismatch between CI and sound representations in auditory cortex. We have updated Figure 8 and the Results section to describe this positive control analysis and clarify the interpretation.

      (13) Perception and interpretation of signals are mentioned several times in the introduction, although perception is not explored in the manuscript (only neuronal processing). This might be confusing.

      We appreciate the need to distinguish between neuronal encoding and perception. We also feel we have been careful not to invoke relationships to perception when presenting analyses on iEEG measurements, but we did identify an opportunity to further clarify this distinction between neuronal processing and perception by adding text in the intro, as follows “for the auditory system to interpret patterns of evoked neural activity and inform downstream auditory areas.”

      (14) Figure 1C. Why is the performance of CI rats so much lower than what was previously published (Glennon et al., 2023)? Did the training duration change?

      The three animals that were behaviorally trained on the normal-hearing (pre-deafening) and cochlear implant task (post-deafening) are within the distribution of the full set of animals from Glennon et al. (2023). However, we note that for Glennon et al. (2023), as one of our behavioral criterion was days to d’ > 1, animals were trained daily until reaching that level and not included in the initial data set if they did not reach that level. However, as we were including animals in this study of iEEG responses that were not trained at all, we felt it appropriate to include this third animal as well, that was trained just for 3 days before recordings were made. The two other animals were trained for 9 and 13 days. We have now included this information in the methods.

      (15) The p-values = 0.5 should be given with an additional digit.

      We previously rounded to the nearest single decimal digit, for all p-values greater than 0.10. We have updated the figures and manuscript text to ensure precision at least to the second digit.

      Reviewer #2 (Recommendations for the authors):

      We thank the Reviewer for their thoughtful comments on our study.

      (1) Less noisy recording methods based on spike detection would provide stronger claims.

      We agree that spike recordings, particularly isolated single-unit activity, are powerful for testing hypotheses about sensory encoding in auditory cortex, and we plan to incorporate such approaches in future work. However, our decision to use iEEG arrays in the present study was deliberate and central to the scientific and translational goals of the project.

      First, iEEG and related population-level approaches such as scalp EEG (e.g., Lalor and Foxe, 2010; O’Sullivan et al., 2015) and fNIRS (e.g., Bortfeld et al., 2009; Peelle, 2017) are widely used in humans and have been highly successful in decoding sound- and speech-evoked responses, revealing fundamental principles of how sound and speech are encoded in the human brain. Because speech is uniquely human and cochlear implants are primarily designed to restore speech perception, aligning our recordings with clinically relevant, human-used modalities enhances the translational relevance of our work.

      Second, iEEG arrays provide distinct advantages over modern multi- and single-unit electrophysiology. Even with high-density probes, the spatial sampling of neuronal activity does not match the coverage of the 60-channel iEEG arrays used here, which span large extents of auditory cortex. One might instead consider optical methods such as calcium imaging to interrogate topographical encoding at single-neuron and mesoscale resolutions, as has been done in normal-hearing mice (Romero and Hight et al., 2019). However, calcium signals are intrinsically slow, limiting access to the temporal precision that is critical for CI encoding, and these tools are unlikely to be available in humans in the foreseeable future, substantially reducing their translational value.

      Using iEEG arrays, we show that CI-evoked responses are topographically organized, consistent with prior work (Klinke et al. 1999, Bierer and Middlebrooks 2002, Middlebrooks and Bierer 2002, including Adenis et al., 2024 now referenced in the manuscript). Our study extends these findings by exploiting simultaneous recordings across both spatial and temporal domains, which are essential for several key analyses (Figs. 3-8), including quantification of trial-by-trial variability, decoding of stimulus identity from single trials, and cross-modal comparisons between normal-hearing and CI-evoked iEEG responses.

      Thus, we believe that the strength of this study is due to, rather than in spite of, its use of iEEG arrays. This approach uniquely allows us to test hypotheses about CI encoding across cortical topography and time using a modality that is directly translatable to human research and clinical practice. In response to the reviewer’s concern, we have also (i) improved the statistical treatment of our data (by adopting linear mixed-effects models that incorporate both paired and unpaired observations), (ii) added additional positive controls (see response to comment #2), and (iii) collected new data that further validate our rodent CI model. Together, these additions strengthen the support for our conclusions while preserving the key advantages of the iEEG-based approach.

      (2) A positive control is necessary to claim the mismatch between CI and sound representations.

      We agree. We now have added a positive control specifically designed to validate our cross-modal analysis pipeline in our revised manuscript. As also suggested by Reviewer 1, the goal was to test whether our method can successfully transfer information when the training and test datasets are matched in modality (tone→tone), thereby ensuring that the observed failure of cross-modal transfer (tone→CI) is not an artifact of the analysis.

      To do this, we re-implemented the full pipeline used in Figure 8, but instead of training on tone-evoked responses and testing on CI-evoked responses, we trained and tested on independent sets of tone-evoked trials from the same animals. For each tone in each animal, we withheld 10 trials as a test set. Using the remaining trials, we fit the original TCA model to obtain spatial and temporal factors (Fig. 8A). We then fixed these factors and re-optimized only the trial factors on the withheld tone-evoked trials (Fig. 8B). The LDA decoder was trained on the trial factors from the original TCA fit and tested on the re-optimized trial factors from the withheld trials, using the same classification pipeline as elsewhere in the manuscript.

      As shown in the top panels of Figure 8C,D, this positive control yielded robust tone→tone generalization: predicted tone frequencies closely matched the actual tones, decoder performance was significantly above chance, and prediction errors were tightly clustered around the true stimulus, indicating that the decoder was tuned to tone frequency. In contrast, when we trained on tone-evoked responses and tested on CI-evoked responses, information transfer was markedly reduced and not different from shuffled controls (Fig. 8E-G).

      These results demonstrate that the TCA+decoder pipeline can reliably transfer information across independent tone-evoked datasets, confirming that the method captures shared structure when it exists. The poor cross-modal transfer between tone- and CI-evoked activity therefore cannot be attributed to a failure of the modeling pipeline but instead reflects a mismatch between CI and sound representations in auditory cortex. We have updated Figure 8, the methods, and the results section to include this new important analysis.

      Reviewer #3 (Recommendations for the authors):

      We thank reviewer 3’s appreciation for study design and the appropriateness of analyses taken. We also appreciate the recognition of noteworthiness, specifically that stimulus identity can be decoded on a single-trial basis and of the potential benefit of using central decoders in clinical settings.

      (1a) Animal heterogeneity: It is difficult to keep track of the animals used in this study, and some received a different protocol of stimulation (sounds through the speech processor vs. direct stimulation) and were also trained in a behavioral task using different target stimuli (4kHz vs. 22.6kHz, also no mention of the CI electrode used as a target).

      We have now clarified the animal cohorts and stimulation protocols in our revised manuscript. We added a new Supplementary Figure 2 that schematizes, for each animal if it underwent behavioral training with pure tones in the normal-hearing condition, if tone-evoked iEEG measurements were collected, if CI-evoked iEEG measurements were collected (and whether stimulation was direct or via the speech processor), and if it subsequently received CI-based behavioral training. Regarding the behavioral targets, we now specify in the Methods that for normal-hearing training, the target stimulus was a 22.6-kHz pure tone. For CI-trained animals, the target was either CI channel 3 (n = 2 rats) or CI channel 4 (n = 1 rat). Details about stimuli targets during behavior have been added to the methods section under “Behavioral training for tone and implant channel detection.”

      (1b) There is no comparison of the CI maps from rats tested with the speech processor and directly stimulated. How different were they? Was the frequency allocation of each electrode the same for each animal? Since data might already have intrinsic variability because of the grid placement, the mechanical deafening, and the cochlear implantation in each animal, such heterogeneity in the 'background' and stimulation protocol might blur the authors' results.

      Our study focuses on cortical encoding of single-channel CI stimulation, so it is indeed important to ensure that the stimuli are effectively delivered by a single electrode, regardless of whether they are driven acoustically via the speech processor or by direct electrical stimulation.

      Stimulation mode and frequency allocation: The project began with single-channel stimulation achieved by presenting pure tones to the speech processor (N=3 animals) and later transitioned to direct programmatic control of individual electrodes (N=4 animals) to simplify the experimental setup. In both cases, the goal was to activate only one CI channel at a time.

      For the programming speech-processor animals, the validation protocol described in Glennon et al. (2023) is as follows:

      - Set the number of active channels in the processor to 1 (the clinical default is 8) to avoid spectral spread across electrodes.

      - Disabled all additional signal-processing strategies (e.g., Scan, ASC, ADRO, SNR-NR, WNR).

      - Used customized frequency allocation tables that mapped narrow frequency bands to individual electrodes, as shown in Glennon et al., 2023, Extended Data Fig. 2.

      To confirm that a given tone drove only the intended electrode, we recorded tone-evoked electrodograms—measurements of the output at each electrode—and verified that only the targeted channel was active (Glennon et al., 2023, Extended Data Fig. 2). Thus, although the initial CI drive was acoustic, the effective stimulation at the array was restricted to a single electrode with a well-defined frequency allocation.

      For the direct-stimulation animals, we used the same underlying frequency allocations to choose which electrode to stimulate, but the pulses were delivered programmatically rather than via the speech processor. In both modes, the center frequency associated with each electrode was therefore defined consistently across animals, and stimulation was confined to one channel at a time.

      Comparison of maps across stimulation modes: We now explicitly indicate the stimulation mode (speech-processor vs direct) for each CI animal in Supplementary Figure 2 and plot the maps for all animals in Supplementary Figure 3. Qualitatively, the spatial organization of CI-evoked maps is similar across the two stimulation strategies; we do not observe systematic differences in map structure that would suggest large biases introduced by the stimulation mode. However, the sample sizes for each group are small (N = 3 speech-processor, N = 4 direct). For this reason, we have not performed formal between-mode statistics and instead treat stimulation mode as a source of minor heterogeneity, alongside inevitable variability from grid placement, mechanical deafening, and cochlear insertion. Given the electrodogram validation (Glennon et al., 2023, Extended Data Fig. 2) and consistent frequency allocation tables, we are confident that both approaches produce single-channel activation with comparable effective frequency assignments.

      (1c) The number of animals used is also confusing. The authors report 7 NH and 7 CI animals (14 total), 4 NH and 3 CI were trained before being implanted (so 3 naïve NH and 4 naïve CI remain). Figure 1C reports that only 3 trained NH performed with the CI (let us call them 3 NH->CI). But then Figure 1E reports only 1 trained NH->CI and only 1 trained NH and 3 naïve NH that got implanted later. On the other hand, Figure 1E reports only 1 true naïve CI animal, the 3 others being naïve NH that got implanted. For the sake of clarity, I would encourage the authors to provide a timeline of the procedures/stimulation protocols coupled with a schematic distribution of the animals.

      To address this, we have added a new Supplementary Figure 2 that provides, for each individual animal a chronological timeline (NH recordings, deafening, implantation, CI recordings); if it was behaviorally trained in the NH condition, the CI condition, or both; if CI stimulation was delivered via the speech processor or by direct electrical stimulation; and which stimulus conditions (tone-evoked iEEG, CI-evoked iEEG) were collected. This schematic makes it clear how the reported totals arise (7 NH and 7 CI for iEEG; 4 NH-trained and 3 CI-trained behaviorally) and shows which specific animals contribute to each panel in Figure 1 and to the later iEEG analyses. We now reference Supplementary Figure 2 in the Results when introducing the cohorts to guide readers through animal accounting.

      (2a) Methods and statistics: Deafening is only mechanical, with no direct or postmortem proof that deafening was complete. The authors cite previous studies, but that would have been a good control to have since mechanical deafening isn't as accepted as the chemical deafening, like Neomycin, especially when some of your animals were stimulated with pure tones through the speech processor.”

      We agree that rigorous verification of deafening is essential, particularly when some CI animals are driven acoustically through the speech processor. Ototoxic approaches (e.g., systemic or local neomycin) are one established method, but their effectiveness can be sensitive to dose and delivery, and they introduce systemic side-effects that can complicate long-term survival and recovery.

      Our laboratory has used the mechanical deafening procedure since it was first described in King et al. (2016) and more recently in Glennon et al. (2023). In King et al., mechanical and ototoxic methods were combined, and we found that ototoxic methods provided no more additional robustness in deafening compared to mechanical lesion. Instead, the additional time required for ototoxic drug application reduced survival times in what was already a very complex and long surgical procedure for bilateral deafening and unilateral cochlear implantation.

      In Glennon et al. (2023) we intentionally employed mechanical-only deafening to minimize side-effects while still achieving profound hearing loss in implanted animals. Glennon et al. (2023) provides an extensive validation of this mechanical-only protocol under the same surgical and experimental conditions as the present study. As we mentioned in our response to comment #3a of Referee 1, we assessed deafness through three measures:

      Histology: In N=4 deafened animals, inner hair cell loss was ~50% and outer hair cell loss was near complete at almost 100% in all animals.

      Physiology: For N=14 rats, acoustic ABRs were substantial before deafening but statistically similar to baseline noise after deafening.

      Behavior: For N=16 deafened rats, behavioral performance with implant on was d′: 1.7±0.1, but when implant was turned off in a subset of sessions, performance dropped to chance (d′: −0.05±0.1, P < 0.0001).

      This convergent anatomical, physiological, and behavioral evidence demonstrates that the mechanical procedure produces profound deafness, with no functionally relevant residual hearing at levels ≥90 dB SPL. Also as we mentioned in response to comment #3a of Referee 1, we believe that the behavioral criterion is most essential and also least common in the literature. Because the tones used to drive the speech processor in the current study were presented at 70 dB SPL, we have no reason to believe that residual acoustic hearing contributed to any of the CI-evoked responses we report.

      We now cite these validation data explicitly in the methods under the section “Bilateral sensorineural hearing loss” as follows “(mechanical only, as described and validated in Glennon et al. 2023)” to make clear why we consider the mechanical-only approach sufficient for ensuring deafness in the present experiments.

      (2b) What motivated the selection of 15 Principal Components for the PCA? That might need to be justified, maybe by scree plot or variance plot (Eigen Values or CEV), as if too many PCs are selected, you are at risk of losing information. Side comment for TCA: why is it important that the number of latent factors exceeds the number of tones or stimuli? Is there a way to justify this statement?

      We thank the reviewer for raising this point. Our choice of 15 components/latent factors was motivated by both theoretical and empirical considerations, which are now made explicit in the manuscript.

      For the PCA analyses, we selected 15 principal components for two reasons. First, because our decoder must discriminate between 10 tone conditions, we reasoned that providing at least as many dimensions as stimuli would be beneficial, while also allowing for the possibility that some components may carry little or no stimulus-selective information. We therefore chose a modest number of components that exceeded the number of tones (10) but avoided unnecessarily high dimensionality. Second, we empirically examined the variance explained as a function of the number of components. As shown in the new scree plots (Supplemental Fig. 4A), the cumulative variance explained enters a near-linear, low-slope regime beyond ~15 PCs, indicating diminishing returns for including additional components. Thus, 15 PCs capture a substantial fraction of the stimulus-related variance while minimizing the risk of overfitting and retaining a consistent dimensionality across animals.

      For the TCA analyses, we used 15 latent factors to match the dimensionality used in PCA and to ensure that the latent space was sufficiently flexible to represent the 10 tone conditions without being under-parameterized. In practice, increasing the number of TCA components reduces reconstruction error (Williams et al., 2018), but with diminishing improvement beyond a certain point. We therefore systematically evaluated model error as a function of the number of latent factors and found that error decreased rapidly up to ~15 components and then plateaued (Supplemental Fig. 4B). This pattern parallels the PCA scree plots and supports 15 as a reasonable trade-off between model flexibility and parsimony.

      We have updated the Results clarify these choices, as follows “The number of components (15) was chosen based on PCA scree plots (Supplemental Fig. 4A), which showed that explained variance entered a near‑linear, low‑slope regime beyond this point demonstrating a similar plateau in reconstruction error (Supplemental Fig. 4B).”

      (2c) Legend of Figure 2E, J states that a Student's paired t-test was used, meaning that only the 'linked' points of the graph were used (thus, comparing only animals that got tested NH then implanted). This is usually the same across the manuscript. Why not include all the points with an unpaired t-test? Otherwise, why are all the points plotted if they serve no purpose? This choice should be justified.

      We agree with this concern, which was also raised by Reviewer 1. We have revised our statistical approach accordingly in our revised manuscript. In the original submission, we used paired t-tests when animals contributed both normal-hearing (NH) and CI data, which meant that animals with only NH or only CI measurements were excluded from those comparisons even though they were shown in the plots.

      To address this, we have re-analyzed all normal-hearing vs. CI comparisons using linear mixed-effects models that include both paired and unpaired data within a single framework. This approach ensures that every plotted data point contributes to the statistical tests, properly accounts for within-animal dependence when both conditions are present, and avoids the loss of power that would arise from either paired-only or purely unpaired tests.

      The mixed-effects results are consistent with our original interpretations, with two comparisons becoming significant in the updated analysis: Fig. 2E (p = 0.048) and Fig. 6F (p = 0.027). We have updated the Results and figure legends to describe the use of mixed-effects models and to report these revised p-values. Together with the new tonotopy and cochleotopy analyses described above, these changes strengthen the statistical support for our conclusions without altering the overall interpretation of the data.

      (2d) Side comment: There are inconsistencies on the bar plots of Figure 6C (Missing a purple point) and Figure 3D (Temporal has 3 purple points).

      Thank you for flagging this mis-labeling (which Reviewer 1 also noticed). We have correctly updated the appropriate data point from trained to naive for Fig. 3D and from naive to trained for Fig. 6C.

      (3a) Pure tones and CI-evoked responses maps: It is the reviewer's understanding that Figure 2 is an averaged representation for all animals. Why is the tonotopic shift so dim for ERPs? The averaged maps aren't very convincing. How were the gradients on an animal-to-animal basis since Figure 2D is only an example animal? Also, everything has been evaluated at 70dB, where selectivity might not be best. It would have been easier to follow the tonotopic gradient at the CFs where contrasts are higher.

      We agree that the strength and interpretation of tonotopy/cochleotopy in our iEEG data needed to be presented more clearly. Reviewer 1 raised closely related concerns, and we have substantially expanded the analyses and explanations in response. Here we highlight the points that address your specific questions.

      Single-animal vs. averaged maps: We included both exemplar maps and population summaries in Figure 2. The panels analogous to Figure 2D show single-animal best-frequency (BF) or best-channel maps; these were chosen because they exhibit clear, interpretable gradients. In the exemplar shown, there is a local high-frequency (HF) region along the medial edge of the array that transitions to lower frequencies toward the rostral edge. For CI-evoked best-channel maps in the same animal, we observe a parallel pattern in which basal electrodes (e.g., electrode 8, representing higher frequencies) occupy the HF region and apical electrodes (e.g., electrode 1, lower frequencies) occupy the LF region.

      Averaged ERP maps, by contrast, necessarily blur some of this structure because iEEG is a summed field potential and animal-to-animal differences in array placement, cochlear insertion depth, and anatomy introduce variability. We have softened the language in the text to reflect that ERP-based tonotopy is coarse and weaker at the population level, while emphasizing that robust gradients are evident in single animals and in HG-based measures.

      Quantitative assessment across animals: To move beyond visual impressions, we added quantitative analyses that mirror those used in Romero and Hight et al. (2020) for calcium imaging data (Romero and Hight et al. 2020 and Fig. 2). For each map we computed local tonotopic gradient vectors at every pixel and summarized their magnitude/direction on a unit circle, then compared the mean vector strength to shuffled maps. Applied to our BF and best-channel maps, this analysis shows that both are significantly more ordered than shuffled controls (p < 10<sup>-10</sup>), indicating that the maps are tonotopic/cochleotopic rather than random, despite the apparent dimness of the gradients in some averaged ERP plots. These new results are described in the revised manuscript and shown in Romero and Hight et al. 2020 and Fig. 2.

      Effect of intensity (70 dB SPL) and “dim” gradients: We agree that stimulus level influences the apparent sharpness of tonotopy. Higher intensities tend to broaden tuning and compress the dynamic range of BF maps. As we now discuss in more detail (adapted from our response to Reviewer 1), tones were presented at 70 dB SPL, so we expect maps to emphasize mid-frequency regions (around 8 kHz) and to show somewhat broader tuning than maps derived at threshold. For CI stimulation, we used ECAP thresholds to set intensity, which is effective in our preparation because animals can robustly discriminate individual electrodes and these electrodes evoke clear cortical activity (King et al., 2015; Glennon et al., 2023).

      In summary, we clarified which panels in Figure 2 show single-animal exemplars vs population summaries, added quantitative analyses demonstrating spatial correlations are greater for adjacent stimuli compared to far-apart stimuli, and expanded the discussion of how recording modality and stimulus level influence the visibility of tonotopic gradients. These changes are intended to make the evidence for tonotopy/cochleotopy in our iEEG data (and its limitations) more transparent.

      (3b) Since new experiments might not be available, it is the reviewer's suggestion to add a supplementary figure showing a couple of animal examples following the format of Figures 2A and 2C that have more contrasted gradients to strengthen the group data. In the case of the CI-evoked responses map, this might also provide another argument to dismiss the potential monopolar smearing.

      Good suggestion, thanks. We now include a new Supplementary Figure 3 that shows additional single-animal examples for both tone-evoked and CI-evoked maps, following the same format as Figure 2C.

      Regarding monopolar stimulation, we agree that monopolar configurations are expected to be less spatially specific than bipolar or multipolar modes because current returns to an extracochlear reference electrode, potentially broadening the spread of excitation. We nevertheless chose monopolar stimulation because it is the predominant clinical configuration in human CI users and therefore most relevant for translational purposes. We acquired ECAP measurements of peripheral (spatial and temporal) tuning via a forward masking paradigm and demonstrate that monopolar is effectively tuned (Supplemental Fig. 2). Together with additional single-animal maps in Supplementary Figure 3, together with our vector-strength analysis (Romero and Hight et al. 2020 and Fig. 2), demonstrate that even under acute monopolar stimulation we observe structured cochleotopic organization in cortex, rather than the fully smeared patterns one might expect if monopolar spread completely dominated.

      We also note that all CI-evoked iEEG measurements were made acutely, immediately after implantation and before any CI-based behavioral experience. It is possible that with longer-term use and plasticity, cortical cochleotopy could become sharper than what we observe here under acute conditions. In this sense, our data provide a conservative baseline showing that even at the earliest stages of CI use, monopolar stimulation already engages tonotopically selective regions of auditory cortex. A longitudinal comparison of acute versus chronic maps would be an interesting direction for future work but is beyond the scope of the current study.

      (3c) Side comments: The legends of Figures 2D and 2I should mention that this is an animal example and not group data, as the rest of the figures are group data.

      Thank you for this suggestion to improve figure clarity. We have updated all of our figures, where appropriate, to indicate whether data are single or groups of animals.

      (3d) In general, some of the legends should be revised because they are sometimes too "strong". As an example, Figure 3B, D legend states: "Variability of iEEG measurements across trials (root mean square, rms) was consistently higher for cochlear implant-evoked compared to tone-evoked activity", despite three of the statistical tests being non-significant. The manuscript is correct, on the other hand.

      Good point. We revised the legend for Figure 3 to be consistent with the figure and the manuscript.

      (3e) The example spatial map given in Figure 3A for CI might not be the best choice since it is showing a pretty reliable trial-by-trial response, while your group data proves the opposite.

      We understand the reviewer’s concern and agree that the exemplar CI map in Figure 3A appears relatively reliable on a trial-by-trial basis. This example was chosen deliberately from an animal in which we had both NH- and CI-evoked iEEG recordings, so that the reader could visually compare the two conditions within the same preparation. In this animal, as in the group data, the differences between NH and CI trial-by-trial responses are subtle rather than dramatic.

      Our group-level analysis shows that the RMS error across trials is consistently higher for CI-evoked than for NH-evoked responses, but the absolute differences are small (< 0.1) and relatively uniform across animals. The spatial maps plotted in Figure 3A are representative of this pattern: both conditions show reasonably robust evoked responses, with CI responses nonetheless showing slightly greater variability. To avoid implying a stronger qualitative difference than is supported by the data, we have revised the text to emphasize that (i) CI-evoked responses remain clearly detectable on single trials, and (ii) the key effect is a small but consistent increase in variability across animals, as captured by the RMS error metrics, “We noted that the differences were qualitatively subtle (Fig. 3A, right panel), they were consistent across animals (Fig. 3B).”

      (4a) Decoders for CI stimulation Regarding CI stimulation, Pearson's correlations were truncated at a spacing of 5 electrodes. Likewise, none of the LDA classifiers show prediction for channels past CI-6. Again, that choice should be justified, or the missing channels should be presented.

      We truncated the correlation between electrodes at 5 because beyond that, the estimated means are significantly noisy. These estimated means are noisy because the number of data are significantly reduced, also significantly increasing the standard error. For example, for the maximum stimulus spacing, the number of pairwise correlations is at maximum the number of animals tested (i.e., N=7). We believe it’s important to be transparent, so we have included the non-truncated version of the figure here in this public review (Author response image 3). We leave the figures in the manuscript untouched but have updated the Figure 2 legend justify this selection of data.

      Author response image 3.

      Expanded figures for spatial correlations and LDA performance. A) The same data from manuscript Figure 2 are re-plotted but with expanded x-axes to include up to 4.5 octaves and 7 channels. Due to the smaller numbers of data at these points, the estimates for the mean spatial correlations are noisier. In all cases, the mean correlations are significantly higher for the first data point compared to the last 3 (NH, ERP p<0.001; NH, HG p<0.001; CI, ERP p=0.005; and CI, HG p=0.39, linear mixed effects models). B) The same data from manuscript figure 4 are re-plotted but with expanded x-axes to include up to ±3.5 octaves and ±6 channels.

      (4b) Finally, retrained PCA-LDA on spatial-only and temporal-only for CI are absent in Figure 3D. Since the authors were pretty consistent in showing both NH and CI alongside in the rest of the paper, it would be coherent to add the CI counterpart to Figure 3D, or maybe with a supplementary figure.

      We agree that consistency can be improved by including classifiers for CI-evoked measurements, though presumably for Fig. 6C and not Fig. 3D. Figure 6 has been updated accordingly.

    1. Author response:

      The following is the authors’ response to the original reviews.

      (1) We bioinformatically examined the repeat compositions of MLSs (Figure 3B), which clearly indicated that all MLSs are composed of repetitive sequences to a much greater extent than the rest of the genome.

      (2) We confirmed the blockage of chromosome breakage by the 4R-CBS mutations using a telomere-anchored PCR assay (Figure 5C-E).

      (3) We examined the effect of the 4R-CBS mutations on the expression of genes encoded in 4R-MDS by RNA-seq (Figure 9). This analysis unexpectedly revealed that gene expression from 4R-MDS is not significantly affected in the mutants, allowing us to extend our discussion.

      (4) We added two authors, Alix Lemoine and Tomoko Noto, who performed the experiments for these revisions.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Nagao and Mochizuki examine the fate of germline chromosome ends during somatic genome differentiation in the ciliate Tetrahymena thermophila. During sexual reproduction, a new somatic genome is created from a zygotic, germline-derived genome by extensive programmed DNA elimination events. It has been known for some time that the termini of the germline chromosomes are eliminated, but the exact process and kinetics of the elimination events have not been thoroughly investigated. The authors first use germline-specific telomere probes to show that the loss of these chromosome ends occurs with similar timing as other DNA elimination events. By comparative analysis of the assembled germline and somatic genomes, the authors find that the ends of each of the germline chromosomes are composed of a few hundred kilobases of micronuclear limited sequences (MLS) that are removed starting around 14 hours after the start of conjugation, which initiates sexual development. They then develop an in situ hybridization assay to track the fate of one end of chromosome 4 while simultaneously following the adjacent macronuclear destined sequence (MDS) retained in the new somatic genome. This allows the authors to more clearly show that these adjacent chromosomal segments are initially amplified in the developing genome before the terminal MLS is eliminated. Finally, they mutate the chromosome breakage sequence (CBS) that normally separates the MLS terminus from the adjacent MDS region, to show that strains that develop with only one mutant chromosome can produce viable sexual progeny, but it appears that both the MLS and the MDS from the mutant chromosome are lost. If both chromosome copies have the CBS mutation, the cells arrest during development and do not eliminate many germline-limited sequences and fail to produce viable progeny. Overall, this study provides many new insights into the fate of germline chromosome ends during somatic genome remodeling and suggests extensive coordination of different DNA elimination events in Tetrahymena.

      Strengths:

      Overall, the experiments were well executed with appropriate controls. The findings are generally robust. Importantly, the study provides several novel findings. First, the authors provide a fairly comprehensive characterization of the size of the MLS at the end of each germline chromosome. I'm not sure whether this has been published elsewhere. Second, the authors develop a novel method to study the fate of chromosome termini during development and use it to conclusively track the elimination of these termini. Third, the authors show that the elimination of these termini appears to occur concurrently with most other DNA elimination events during somatic genome differentiation. And fourth, the authors show that failure to separate these eliminated sequences from the normally retained chromosome alters the fate of these adjacent MDS and the loss of the cells' ability to produce viable progeny.

      Weaknesses:

      It appears the authors did extensive analysis of the MLS chromosome ends, but did not provide too much information related to their composition. If this has not been published elsewhere, it would be useful to describe the proportion of unique and repetitive sequences and provide more information about the general composition of the chromosome ends. Such information would help the reader understand the nature of these MLS and how they may or may not differ from other eliminated sequences.

      We now calculated the proportions of unique and repetitive sequences for each MLS, and these data are included in Figure 3B and described in the main text of the revised manuscript. A more comprehensive analysis of chromosome-end composition, including detailed characterization in the context of the complete MIC genome assembly, is beyond the scope of the current study and will be presented in a future publication.

      Although the development of the novel FISH probes for large chromosome ends allowed for these novel discoveries, the signal in several images was visible, but often quite faint. I'm not sure there is anything the authors could do to improve the signal-to-noise ratio, but one needs to stare at the images carefully to understand the findings.

      We have submitted higher-resolution images for the revised manuscript, which we believe much improve the visibility of faint signals.

      One main weakness in the opinion of this reviewer is that the authors did very little to understand why, when a terminal MLS and the adjacent MDS fail to get separated because of failure in chromosome breakage, both segments are eliminated. The authors propose that possibly essential genes in the MDS get silenced, and the resulting lack of gene expression is the issue, but this and other possibilities were not tested. The study would provide more mechanistic insight if they had tried to assess whether the MDS on the CBS mutant chromosome becomes enriched in silencing modifications (e.g., H3K9me3). Alternatively, the authors could have examined changes in gene expression for some of the loci on the neighbouring MDS.

      The 4R-CBS mutation causes two distinct defects that should be considered separately: (1) co-elimination of 4R-MLS and the adjacent 4R-MDS during uniparental transmission of the 4R-CBS mutation; and (2) a global block of DNA elimination during biparental transmission of the 4R-CBS mutation.

      For the first defect, 4R-MLS and 4R-MDS may simply co-segregate into the nuclear compartment where DNA elimination occurs when the chromosome break that normally separates 4R-MLS from 4R-MDS is blocked. In this scenario, no additional process, such as spreading of scnRNA production, heterochromatin formation, or gene silencing, would be required to induce co-elimination. This point was not clearly stated in the previous manuscript, and we have now added a discussion of it to the revised manuscript.

      The possibility of gene silencing within 4R-MDS was raised as a potential explanation for the second defect. To test this possibility, we performed RNA-seq analysis of wild-type and 4R-CBS mutant cells to determine whether gene expression from 4R-MDS is affected by mutations at 4R-CBS. Contrary to our expectations, we found that genes in 4R-MDS are not significantly down-regulated in 4R-CBS mutant cells compared with other genes. This result suggests that the DNA elimination defect in these cells cannot be explained by silencing of genes located within 4R-MDS. We have added these RNA-seq data to Figure 9 and described them in the Results section. We have also revised the Discussion to propose alternative possibilities that may guide future investigations.

      The other main weakness is that since the authors only mutated the end of one germline chromosome, it is not clear whether the elimination of the MDS adjacent to the terminal MLS on chromosome 4 when the CBS is mutated is a general phenomenon, i.e., would happen at all chromosome ends, or is unique to the situation at Chromosome 4R. Knowing whether it is a general phenomenon or not would provide important insight into the authors' findings.

      As was described in the manuscript, the short (CBS = 15 nt) target within AT-rich and repetitive regions prevent designing gRNAs specifically targeting some of the chromosome end CBSs. We tried to mutate the CBS sequences of the left end of the chromosome 3 (3L) and the left end of the chromosome 5 (5L) by the strategy we used to mutate 4R-CBS but failed. Therefore, to systematically mutate other chromosome-end CBSs, we need to establish a different strategy, such as combining template-based repairing to CRISPR-induced DSB. We have explained this technical limitation and stated that “Our data support a critical role for 4R-CBS in separating 4R-MLS from 4R-MDS, but it remains unclear whether all MIC chromosome ends are strictly CBS-dependent for their elimination.” in Discussion (Page 12).

      Reviewer #2 (Public review):

      Summary:

      Nagao and Mochizuki investigated how the germline (MIC) telomere was removed during programmed genome rearrangement in the developing somatic nucleus (MAC). Using an optimized oligo-FISH procedure, the authors demonstrated that MIC telomeres were co-eliminated with a large region of MIC-limited sequences (MLS) demarcated on the opposite side by a sub-telomeric chromosome breakage site (CBS). This conclusion was corroborated by the latest assembly of the Tetrahymena MIC genome. They further employed CRISPR-Cas9 mutagenesis to disrupt a specific sub-telomeric CBS (4R-CBS). In uniparental progeny (mutant X WT), DNA elimination of the sub-telomeric MLS was not affected, but the adjacent MAC-destined sequence (MDS) may be co-eliminated. However, in biparental progeny (mutant X mutant), global DNA elimination was arrested, revealing previously unrecognized connections between chromosome breakage and DNA elimination. It also paves the way for future studies into the underlying molecular mechanisms. The work is rigorous, well-controlled, and offers important insights into how eukaryotic genomes demarcate genic regions (retained DNA) and regions derived from transposable elements (TE; eliminated DNA) during differentiation. The identification of chromosome breakage sequences as barriers preventing the spread of silencing (and ultimately, DNA elimination) from TE-derived regions into functional somatic genes is a key conceptual contribution.

      Strengths:

      New method development: Oligo-FISH in Tetrahymena. This allows high-resolution visualization of critical genome rearrangement events during MIC-to-MAC differentiation. This method will be a very powerful tool in this area of study.

      Integration of cytological and genomic data. The conclusion is strongly supported by both analyses.

      Rigorous genetic analysis of the role played by 4R-CBS in separating the fate of sub-telomeric MLS (elimination) and MDS (retention). DNA elimination in ciliates has long been regarded as an extreme form of gene silencing. Now, chromosome breakage sequences can be viewed as an extreme form of gene insulators.

      Weaknesses:

      The finding of global disruption of DNA elimination in 4R-CBS mutant progeny is highly intriguing, but it's mostly presented as a hypothesis in the Discussion. The authors propose that the failure to separate MLS from MDS allows aberrant heterochromatin spreading from the former into the latter, potentially silencing genes required for DNA elimination itself. While supported by prior literature on heterochromatin feedback loops, the specific targets silenced are not identified. While results from ChIP-seq and small RNA-seq can greatly strengthen the paper, the reviewer understands that direct molecular characterization may be beyond the scope of the current work.

      As mentioned in our reply to Reviewer #1’s comment above, we performed RNA-seq on wild-type and 4R-CBS mutant cells at 13.5 hpm and 15 hpm and found that genes in 4R-MDS are not significantly downregulated in 4R-CBS mutant cells (Figure 9), suggesting that the DNA elimination defect in these cells cannot be explained by aberrant heterochromatin spreading. Therefore, the link between the chromosome break at 4R-CBS and general DNA elimination remains elusive and will be a very interesting subject for our future research. We have added these results and revised the discussion in the manuscript.

      Reviewer #3 (Public review):

      Programmed DNA elimination (PDE) is a process that removes a substantial amount of genomic DNA during development. While it contradicts the genome constancy rule, an increasing number of organisms have been found to undergo PDE, indicating its potential biological function. Single-cell ciliates have been used as a prominent model system for studying PDE, providing important mechanistic insights into this process. Many of those studies have focused on the excision of internally eliminated sequences (IES) and the subsequent repair using non-homologous end joining (NHEJ). These studies have led to the identification of small RNAs that mark retained or eliminated regions and the transposons that generate double-strand breaks.

      In this manuscript, Nagao and Mochizuki examined the other type of breaks in ciliates that were healed with telomere addition. They specifically focused on the sequences at the ends of the germline (MIC) chromosomes, which have received relatively less attention due to the technical challenges associated with the highly repetitive nature of the sequences. The authors used the Tetrahymena model and developed a set of new tools. They used a novel FISH strategy that enables the distinction between germline and somatic telomeres, as well as the retained and eliminated DNA near the chromosome ends. This allows them to track these sequences at the cellular level throughout the development process, where PDE occurs. They also analyzed the more comprehensive germline and somatic genomes and determined at the sequence level the loss of subtelomeric and telomere sequences at all chromosome ends. Their result is reminiscent of the PDE observed in nematodes, where all germline chromosome ends are removed and remodeled. Thus, the finding connects two independent PDE systems, a protozoan and a metazoan, and suggests the convergent evolution of chromosome end removal and remodeling in PDE.

      The majority of sites (8/10) at the junctions of retained and eliminated DNA at the chromosome ends contain a chromosome breakage sequence (CBS). The authors created a set of mutants that modify the CBS at the ends of chromosome 4R. CBS regions are challenging for CRISPR due to their AT-rich sequences, making the creation of the 4R-CBS mutants a significant breakthrough. They used the FISH assay to determine if PDE still occurs in these mutant strains with compromised CBS. Surprisingly, they found that instead of blocking PDE, its adjacent retained DNA is now eliminated, suggesting a co-elimination event when the breakage is impaired. Furthermore, in biparental mutant crosses, no PDE occurred, and no viable progeny were produced, indicating that the removal of chromosome ends is crucial for proper PDE and sexual progeny development. Overall, the work demonstrates a critical role for 4R-CBS in separating retained and eliminated DNA.

      We appreciate Reviewer 3’s assessment.

      Recommendations for the authors:

      Reviewing Editor Comments:

      All reviewers agree that this study makes an important contribution to the field; however, they also offered several suggestions for how the manuscript could be improved. In particular, we draw your attention to the comments from Reviewer #1, who suggests that the manuscript could benefit from additional information on the general composition of germline chromosome ends, where available.

      As noted in our response to Reviewer #1 in the Public Reviews above, we have included an analysis of the fraction of repetitive sequences for each MLS as Figure 3B in the revised manuscript, highlighting the highly repetitive nature of MLSs compared with the rest of the genome.

      Reviewer #1 (Recommendations for the authors):

      As mentioned in the weaknesses section, the authors could provide more information regarding the nature of the sequences that make up the terminal MLS. There have been reports that these are highly repetitive; is that the case? Also, did the authors identify common repeats that are not internal to mic chromosomes that could be used to track all terminal segments of the five chromosomes? This would complement their mic-telomere probe.

      As noted in our response to Reviewer #1’s Public Review above, we have added an analysis of the fraction of repetitive sequences for each MLS as Figure 3B in the revised manuscript, which confirms that MLSs are highly repetitive.

      Apart from the moderately conserved Telomere Associated Sequence (TAS), described by Kirk and Blackburn (1995) and of unknown function, we were unable to identify any obvious shared repeats unique to MLSs that could support the development of pan-MLS-specific probes.

      One major weakness is that the authors did little to determine the cause of the elimination of the adjacent MDS along the 4R-MLS when the CBS was mutated. It would really improve the study if the authors could show that:

      (1) Gene expression of genes on the MDS is reduced in 4r-CBS mutant progeny.

      (2) Heterochromatin modifications are unexpectedly acquired on the MDS in mutants relative to wild-type chromosomes.

      (3) Do scnRNA specific to the MDS region appear in the mutant progeny during development, but not in wild-type crosses?

      Any data that would help support the authors' hypothesis regarding how the MDS region is eliminated when the CBS is mutant would definitely strengthen the conclusions of the study.

      As noted in our response to Reviewer #1’s Public Review above, we performed RNA-seq on wild-type and 4R-CBS mutant cells at 13.5 hpm and 15 hpm. Our analysis showed that genes within the 4R-MDS are not significantly downregulated in 4R-CBS mutant cells (Figure 9), suggesting that the DNA elimination defect in these cells cannot be attributed to aberrant heterochromatin spreading. Therefore, the connection between the chromosome break at 4R-CBS and general DNA elimination remains unclear and represents an important avenue for future investigation. We have incorporated these results and revised the discussion accordingly in the updated manuscript.

      The other main weakness is that by mutating the CBS of only one chromosome arm, one can't know whether the loss of the MDS with the MLS in the mutants is generalizable for all chromosome arms or is unique to 4R. The authors noted that they were unable to make any other mutated CBSs. Another way to try to get to this question is to try to rescue the mutant by inserting a new CBS into the 4R arm such that some MLS remains linked to the 4R-MDS and see whether removing the mic telomere is the issue, or would a block of MLS attached to the 4R-MDS be sufficient to cause its elimination. I'm not sure where to exactly put the new CBS, but worth thinking about.

      To introduce a new CBS into 4R-MLS, we would need to insert a CBS-containing construct into the MIC by homologous recombination during conjugation and then select engineered transformants using a drug resistance marker expressed from the derived MAC. However, because 4R-MLS is still eliminated in the progeny of 4R-CBS mutants, the introduced marker would be lost from the MAC even if homologous recombination were successful. Therefore, although the strategy suggested by this reviewer is very interesting, several technical innovations are required to make such experiments feasible, leaving this approach for a future project.

      It seems somewhat curious that the mutation of the CBS completely blocks nuclear development. In Paramecium, the failure to complete internal DNA elimination events can lead to alternative telomere addition. The caveat being that, in Paramecium, telomere addition appears more promiscuous than in Tetrahymena. It would be helpful to know how absolute the failure to produce progeny is in these mutants. Is it zero progeny in 10<sup>6</sup>, 10<sup>7</sup>, 10<sup>8</sup> ..... mated cells? Can the authors provide a possible lowest possible frequency?

      The viability tests were performed using bulk mating of 2.5 × 10<sup>4</sup> cells for each cross. Because ~70-80% of mating pairs complete the conjugation process and produce exconjugants under our standard culture conditions, and because we did not detect any 6-mp-resistant progeny from MUT x MUT crosses, we estimate that the probability of obtaining viable progeny in these crosses was less than 1 progeny per ~2 × 10<sup>4</sup> mating pairs. The number of cells used for the viability assay is described in the “Viability Test of Sexual Progeny” section of Materials and Methods and the estimated frequency of progeny production from the mutants has been mentioned in Results section in the revised manuscript.

      The one implication of the study is that chromosome breakage and DNA elimination, two different events, are coupled. In most mutants that block scnRNA-directed DNA elimination, both IES excision and chromosome breakage occur. In the study by McDaniel, SL. et al (2016). DRH1, a p68-related RNA helicase, is required for chromosome breakage in Tetrahymena. Biology Open pii: bio.021576. doi: 10.1242/bio.021576, germline knockouts of DRH1 could complete IES excision, but not chromosome breakage, indicating that the processes can be uncoupled. It may be useful for the authors to discuss this previous work in relation to their finding that failure in chromosome breakage can lead to DNA elimination of neighboring sequences.

      So far, DRH1 is the only gene reported to be required for chromosome breakage without affecting DNA elimination in Tetrahymena. However, McDaniel SL et al. (2016) examined chromosome breakage at only two CBSs (distinct from 4R-CBS), and thus it remains unclear how broadly chromosome breakage, including that at 4R-CBS, is affected in the absence of DRH1. In addition, McDaniel SL et al. (2016) assessed DNA elimination at three different IESs using PCR, whereas our study examined elimination of the repetitive Tlr1 transposon using FISH. Therefore, without further analysis of the similarities and differences in chromosome breakage and DNA elimination phenotypes between DRH1 knockout cells and 4R-CBS mutants, it is difficult to draw meaningful conclusions. Accordingly, we have limited ourselves to stating the following in the Discussion of the revised manuscript: “Moreover, chromosome breakage can be inhibited without disrupting DNA elimination, as shown in cells lacking zygotic expression of the p68-like RNA helicase Drh1 (McDaniel et al., 2016).”

      Minor corrections:

      Page 7, line 3: the text "......inducing chromosome break" should either be "......inducing chromosome breaks" or "......inducing a chromosome break".

      Corrected as “inducing a chromosome break”.

      Page 13, line 13: "......large block...." should be "......large blocks......".

      Corrected as suggested.

      Reviewer #2 (Recommendations for the authors):

      The authors can experimentally validate that chromosome breakage at 4R-CBS is indeed disrupted by the mutations. A PCR-based assay testing de novo telomere addition is a standard tool. In addition, MLS-linked telomere should only appear transiently during conjugation in WT cells.

      Because it was previously unknown whether de novo telomere addition occurs at the ends of MLSs upon chromosome breakage, we tested this using a PCR-based assay. We detected telomere-added chromosome ends of 4R-MLS and 3L-MLS, which were undetectable until 10.5 hpm, appeared at 12 hpm, and gradually decreased by 18 hpm in wild-type cells (WT × WT cross). Importantly, the appearance of the telomere-added 4R-MLS end, but not the 3L-MLS end, was blocked in 4R-CBS mutants (Mut x Mut crosses), strongly supporting that the 4R-CBS mutations specifically disrupt chromosome breakage at 4R-CBS. These new data are shown in Figure 5C–E and described in the Results section.

      The high FISH background during conjugation may be caused by the abundant presence of dsRNA, which is resistant to RNase A treatment but may be degraded by RNase III.

      The high FISH background was observed in the parental MAC at 9 and 12 hpm (Figure 2, 4, and S2) where dsRNA accumulation was not detected in the previous studies (Woo et al. 2016; Shehzada et al. 2024). In contrast, the MIC at 3 hpm and the new MAC at 9 and 12 hpm, where strong dsRNA accumulation was detected, showed much weaker background FISH signals (Figure 2, 4, and S2). Therefore, we believe that dsRNA is not the main cause of the high FISH background.

      It is likely that the long MIC telomere is treated as IES and targeted for DNA elimination. Indeed, telomere-specific scnRNA is abundantly produced during conjugation (http://www.ncbi.nlm.nih.gov/pubmed/19460867).

      We have cited the suggested literature and the following description has been added in Discussion to relate the reported telomere-derived scnRNAs to the abundant scnRNAs produced from MIC chromosomal ends: “In addition, telomere-complementary scnRNAs were reported to be produced specifically during conjugation (Cao et al. 2009).”

      Global disruption of DNA elimination may be a direct effect (DNA excision machinery affected) or indirect (unrepaired DSB and checkpoint activation).

      It has been reported that unrepaired DSBs caused by loss of Ku80 (Tku80) do not block DNA elimination in Tetrahymena (Lin et al. 2012). Therefore, checkpoint activation by unrepaired DSBs, if it occurs, is unlikely to explain the DNA elimination defect observed in the progeny of 4R-CBS mutants. Nonetheless, this direct-versus-indirect issue would be relevant when considering whether disruption of specific 4R-MDS-encoded genes in 4R-CBS mutants could cause the DNA elimination defect. Our new RNA-seq analysis, however, suggests that this possibility is unlikely. Therefore, we did not add further discussion of this direct-versus-indirect issue.

      Minor points:

      The zoom-in boxes in most images are barely visible.

      We have modified the zoom-in boxes to make them clearer.

      Page 13: scnRNA precursors (Cai et al., 2025) (Cai et al., in press). Is it one paper or two?

      They are two papers and the latter was published reacently. We have updated the citation.

      Reviewer #3 (Recommendations for the authors):

      The manuscript is well-written, with clear data, thoughtful discussion, and concise presentation. I have only a few minor comments below.

      For Figure 4 and others, the right panel shows the stats and percentages, with positive and negative labels. It's a bit confusing at first glance. I think it can be clarified what positive and negative mean in the legend.

      The legends of Figure 4, Figure 6 and Supplementary Figure S2, have been modified as “The presence (Positive) or absence (Negative) of the 4R-MLS FISH signal in new MAC (An) in 50 cells per time point was examined.”

      The quality of the FISH images is low at their current resolution. It is difficult to get a clear view.

      In the initial version, some images were in low resolution when we combined them into a single pdf file for review. In the revised manuscript, the images have been replaced with high-resolution images.

      The co-elimination of neighboring 4R-MDS when 4R-CBS is mutated, can this be viewed as a fail-safe mechanism to ensure the elimination of the chromosome ends? Regardless, the result begs the question of the significance of end removal and remodeling of PDE. Some speculations in the discussion might be helpful.

      Because the neighboring 4R-MDS contains approximately 100 predicted genes, its co-elimination would likely be too risky to evolve as a fail-safe mechanism for ensuring chromosome-end elimination in every generation. Instead, we interpret this as an erroneous process that can still be compensated for through endoreplication of the remaining, normally processed 4R-MDS from the non-mutated copy.

      We further speculate that the connection between chromosome breakage at 4R-CBS and the essential PDE process may serve as an evolutionary pressure to preserve the 4R-CBS locus in a chromosome breakage-competent state. We have added the following discussion to the revised manuscript (Page 15): “The observed link between chromosome breakage at 4R-CBS and the essential DNA elimination process may reflect the biological significance of MLSs and the importance of their removal from the MAC. Coupling these processes may have evolved as a mechanism to ensure that only functional chromosome-end CBS loci are preferentially transmitted to future generations.”

      Figure 1, legend, line 3, "the sexual reproduction process", do you mean "the sexual reproduction proceeds or initiates"?

      We meant “conjugation” = “the sexual reproduction process”. To make this clearer, we have revised the legend as “conjugation, which is the sexual reproduction process of Tetrahymena”.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors conclude that mPFC is not required for avoidance, based on the minimal behavioral effects of optogenetic inhibition. While this interpretation is supported by the data, the choice of viral constructs could lead to an underestimation of the mPFC's role for other reasons. First, the choice of viral constructs could lead to an underestimation of the mPFC's role for several reasons. Specifically, the efficacy of eArch3.0 inhibition was not verified beyond histology, and its non-cell-type-specific nature could lead to disinhibition or compensatory activity in downstream regions. Although the authors' use of visual cortex (VI) inhibition as a control suggests that broad cortical inhibition does not impair avoidance, subcortical compensation cannot be ruled out. Additionally, Vgat-ChR2 targets only GABAergic neurons, potentially missing glutamatergic contributions. Addressing these limitations in the Discussion section would strengthen the manuscript.

      We thank the reviewer for these points. First, although we did not perform direct electrophysiological verification of eArch3.0 efficacy in mPFC in the present study, this construct has been extensively validated in prior work and is widely used to produce robust neuronal inhibition. In our experiments, the lack of behavioral effect with eArch3.0 inhibition converged with the results obtained using the independent Vgat-ChR2 approach, which we directly validated, supporting the conclusion that mPFC inhibition does not impair avoidance under these conditions. Our results are also consistent with previous studies showing that mPFC lesions do not impair avoidance behavior.

      Second, we agree that manipulating mPFC activity will necessarily influence downstream circuits, including subcortical regions, given the interconnected nature of these networks. Our goal was to test whether inhibiting mPFC activity alters avoidance behavior, not to isolate it from its targets. In this context, the absence of behavioral effects indicates that avoidance behavior can be supported without mPFC activity. While compensation is always a possibility, this usually reveals some impairment while compensation occurs, but we did not observe those effects. Our results are consistent with the idea that subcortical circuits normally mediate these behaviors.

      Finally, regarding Vgat-ChR2, activating GABAergic neurons is a well-established approach to suppress cortical activity, as these interneurons provide strong inhibition onto local glutamatergic neurons. Thus, this manipulation is expected to broadly reduce excitatory output in cortex. Indeed, the robust suppression of cortical activity we observed with GABAergic activation makes it unlikely that major glutamatergic contributions were missed.

      These points are in the paper, including the Discussion.

      Reviewer #2 (Public review):

      (1) There are few details on the linear mixed models in the methods. This section could be improved by including a mathematical description. More importantly, the reader never learns how accurately the models capture the data. Given that most conclusions rely on the models, it seems central to address this point carefully. For example, what is the explained variance, marginal, and conditional? Were the nested models compared to non-nested ones (e.g., AIC), what are the specific outputs of the likelihood ratio tests briefly mentioned in the methods?

      Model structure was defined a priori by the experimental design and hypotheses rather than selected through model comparison, but we verified the contribution of key model components (e.g., covariates, interactions, and random effects) using likelihood ratio tests comparing models. Regarding model performance, we now report for each model the marginal and conditional R<sup>2</sup> values (Nakagawa), which quantify variance explained by fixed effects alone and by the full mixed model including random effects. In addition, likelihood ratio test results for all fixed effects and interactions (χ<sup>2</sup> statistics) were already reported in the manuscript.

      (2) For several figures, there is a disconnect with the main text, in the sense that it is difficult to understand how statements in the main text connect with specific figure panels or bars in their graphs. This is particularly the case for the most complex figures, e.g., Figures 3, 4, and their supplements. It would be beneficial to introduce subfigure labels (A1, etc) and state explicitly in the main text what figure panel is described (in parentheses). Alternatively, breakdown the figures into multiple ones, decreasing ambiguity. This is important because it will help the reader better assess the strength of the results.

      We have significantly revised the manuscript to reduce ambiguity and thank the reviewer for each of their (28) requests, which we have implemented in full. We also added additional figure references to the Results to assist with readability. This has significantly improved clarity and readability.

      (3) It does not appear that the code and data used to produce the figures are made available. That would be very beneficial, given the complexity of the analysis and dataset collection procedures. It would also help readers better understand the results and probe their validity.

      As usual, we will share the full dataset in the VOR at Dryad after the revision is completed.

      Reviewer #3 (Public review):

      The main weakness, in my view, lies in the Results section. In the figures, the authors do not present any raw data, and the plots are shown as mean {plus minus} SEM without displaying the distribution of individual data points.

      We thank the reviewer for the recommendations. Individual data points are shown where appropriate (e.g., Fig. 1). However, most of our analyses involve repeated-measures, hierarchical data with multiple levels (cells and sessions nested within animals), where simple point overlays can be misleading or difficult to interpret without explicit linking across levels. We therefore use mean ± SEM visualizations for clarity in these summary figures, while preserving the full hierarchical structure in the statistical analysis through mixed-effects models. All data will be made available in the VOR to allow full inspection of the underlying distributions.

      It is both a strength and a weakness that the authors do not attempt to guide the reader through the Results section and instead present the findings with very little emphasis on the key outcomes of the GLM. While this approach is arguably the most transparent way to report results, it also makes the section quite difficult to follow and may discourage readers.

      I would recommend rewriting the Results section to make it more accessible to a broader audience. A similar issue applies to the figures: presenting all plots reflects a commendable commitment to transparency, but it would greatly benefit from a clearer narrative. As it stands, it is difficult to grasp the message of each figure by simply browsing through them.

      The full description (complexity) of the models is entirely in the legends and supplemental figures. This was done to make the results easier to follow. We have made all the changes noted above to facilitate readability while assuring there is enough transparency to assess the data. We think readability has significantly improved.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Below are a few specific suggestions related to the main weaknesses mentioned above.

      (1) P4 L9: The sentence starting with "However, most ..." sounds more like a statement than a contrast with the previous sentence. Therefore, please delete "However" and please add references to justify the statement.

      Done.

      (2) P8: Definition of movement peaks. It would be great to have three videos illustrating the mouse behavior in the three different movement peaks. This would allow the reader to better understand the differences between no peaks 3 sec prior, more than 5 seconds, and one example that does not fit these two categories. In addition, what percentage of all peaks to the no peaks 3 sec prior and more than 5 sec represent?

      We added the percentages. The “3 sec prior” represent ~23% and the “5 sec” represent ~31%. However, we do not think adding a single video of one movement per these 3 cases would be useful as the dataset is composed of thousands of these movements.

      (3) P8: Last paragraph. When you state that you performed a linear fit between DF/F and movement, do you mean speed? In addition, the statement "integrating both signals over a 200 ms window" is incomplete. How is the window selected? Is the window 200 ms around movement onset or movement peak speed?

      Yes, the movement variable used in the linear fit corresponds to speed. Regarding the 200 ms window, this analysis does not focus on specific behavioral events such as movement onset or peak speed. Instead, both ΔF/F and speed signals were segmented into consecutive 200 ms windows across the entire recording session, and the linear relationship was computed across these paired segments. Thus, the analysis captures the overall relationship between neural activity and ongoing movement, rather than eventaligned dynamics. We have revised the text to clarify both the use of speed and the implementation of the 200 ms window.

      (4) P14: Discussion of AA19 and AA39 tasks: It would be helpful to clearly specify what percentage of actions you would expect given no learning, is it the 23% action dashed line indicated in the top panel of Figure 2B?

      The expected percentage of actions under no learning is not fixed, as it depends on the rate of spontaneous (non–cue-driven) crossings. In these tasks, we estimate this baseline using behavior during the noUS condition, where the action rate is ~23% (Fig. 2B). In the AA19 and especially AA39 tasks, this baseline decreases because spontaneous inter-trial crossings (ITCs) are progressively reduced, leading to lower expected action rates under no-learning conditions. Thus, the 23% baseline derived from noUS is lower in the AA19/39 tasks. In other studies, we explicitly included NoCS (no-cue) trials to estimate chance performance; however, in the present design we rely on the noUS baseline and the observed changes in ITC rate. We have clarified this point in the text.

      (5) P15 L2: "Considering tone intensity (Fig. 2B), CS1 avoids latencies increased at medium and high intensities but not a low intensity." This is confusing. Are you referring to the AA39 triangles under CS1 in the middle panel, left? They are all above the dashed reference line. So the plot seems to contradict the statement. If you are referring to AA19, the red dots also seem to show the opposite of the statement.

      The dashed reference line reflects latency during the noUS condition and is included for visual reference; however, these values are not directly comparable to those in the AA tasks, as noUS latencies are largely unconstrained and reflect baseline behavior rather than learned responding. The statement in the text refers specifically to changes across AA conditions, consistent with our analysis approach throughout the manuscript, where values are compared to the immediately preceding condition. In this case, we are referring to AA39 (triangles) relative to AA19 (circles). Under this comparison, CS1 avoidance latencies increase at medium and high intensities, but not at low intensity, consistent with the statistical contrasts. We have revised the text to clarify the points.

      (6) P17: "Movement and neural measures subtract the baseline from the other three windows at a trial level." Do you mean to say that for each measure, the baseline was subtracted? How is baseline defined (over which time window)?

      The baseline is defined in that same paragraph as the −0.5 to 0 s pre-CS window. To improve clarity, we have revised the text to explicitly restate this definition in the sentence describing baseline subtraction.

      (7) P17: "Fig. 2-Supplement 2A,B shows model-derived marginal means of movement averaged across tone intensities." Some explanation needs to be provided, since the previous figures show a dependence of behavior on tone intensity. Are you doing this based on Fig. 2-S1?

      Yes, these results are derived from the same model of the full data shown in Fig. 2–S1. In this particular analysis, tone intensity was included in the model but not retained when computing marginal means and contrasts, effectively averaging across intensity levels. The rationale for this approach is that tone intensity was primarily used to increase behavioral variability, particularly error rates, which are otherwise low in this task. Averaging across intensity therefore improves statistical power and allows us to more clearly isolate the effects of the primary factors of interest. We have clarified this point in the text.

      (8) P18: "Orienting magnitude was strongly dependent on tone intensity...". However, in Figure 2-S2, there is no information about tone intensity. So how is the reader supposed to see this? Same issue on P19 when discussing the action window. Generally, the description of Figure 2-S1 and S2 is difficult to follow and should be improved. It is not clear that all panels are referred to in the text.

      We have revised the start of the Movement section to clarify how tone intensity is treated across analyses and figures. Specifically, tone intensity is included as a factor in all statistical models; however, for clarity of presentation, it is sometimes collapsed in figures to reduce dimensionality and to emphasize other task-related factors. This manipulation was introduced primarily to increase behavioral variability (particularly error rates), thereby improving sensitivity for estimating the effects of the other task variables.

      We have also clarified when we reference Fig. 2–S2 legend that, although intensity is not displayed in the figure for visualization purposes, it is included in the underlying model and its effects are reported in the supplement.

      (9) P22, 23: Windows are mentioned, but not defined or indicated in figures.

      We have clarified in the text that the same time windows defined for movement analyses (baseline, orienting, action, and from-action) were also used for the neural analyses.

      (10) P22: "Covariates were standardized within each window so that estimated marginal means reflected ΔF/F at average covariate values." It is unclear what was done exactly. What do you mean by "standardized"? Maybe give an example here and elaborate in the methods.

      By “standardized within each window,” we mean that covariates were z-scored within each analysis window (i.e., each covariate was transformed to have a mean of 0 and a standard deviation of 1 within that window). This ensures that estimated marginal means correspond to ΔF/F evaluated at the average covariate values within each window. We have clarified this in the Methods and Results.

      (11) P24-25: Indicating spurious action on Figure 3-S2 (and in Figure 3) would help the reader follow the argument in the main text.

      We clarified this in the legends by indicating that actions not classified as AA, PA, Escape, or PA Error are spurious actions.

      (12) P25: "After controlling for ..., but this includes the effects of aversive stimulation." The second part of this sentence was not clear.

      We have clarified this sentence to indicate that avoidance errors are followed by aversive stimulation (i.e., errors are punished).

      (13) P34L3: "Classs" -> "Class".

      Fixed.

      (14) P42 top paragraph: There are two references to Figure 5-S1 panel D, but there is no panel D on the figure.

      Fixed.

      (15) P57: The sentence starting with "Random effects were specified ..." is very difficult to follow.

      We have revised this sentence to improve clarity by separating the description of the random-effects structure from the model syntax.

      (16) P57: The windows analyzed are finally defined at the bottom of this page. The information also needs to be included early in the results to improve comprehension.

      This is now included in the main text when windows are first used in the movement section.

      (17) P58: Several R packages are mentioned by name, but without specifying that they are R packages, which would facilitate reading.

      We added R.

      (18) P58 top paragraph: "Tuckey's correction", do you mean "Tukey's HSD test"?

      We thank the reviewer for noting this. We used Holm-adjusted p-values for multiple comparisons (as implemented in emmeans) and have revised the text.

      (19) P63: "features extracted from F/F" do you mean "DF/F"?

      Yes, fixed.

      (20) Figure 1B speed plots: it is not possible to visualize the lines at the movement peak because they overlap completely. You can either add an inset on the left of the peak (for each panel), magnifying that region, or play with the transparency of the traces to improve visibility. There is a similar issue in Figure 5A, B. (Alternatively, if it is not possible to solve the issue graphically, explicitly state that traces overlap.)

      We have fixed this by making some traces dashed in Figure1 and 1-S1, which reveals the underlying traces. We also stated that the peak speed completely overlaps. In Figure 5, we stated that traces overlap as expected; transparency or dashing does not work well with the colors used in Figure 5 and in fact the overlap emphasizes the similarity of the movements.

      (21) Legend 1A: abbreviation CCF not defined. Is it anterior to the left? Abbreviation WM not defined. The right panels are unclear. The legend states that they show a schematic of the location of the optical fibers, but that was not clear. Do the dots indicate the location of the fibers? Is the green region indicative of V1? Same for dark gray in the mPFC panel. What are the lighter grey regions and the blue region? Does 'lateral' mean 'lateral from midline'? Please clarify these points.

      CCF is defined in Methods, and the typesetting process will adjust abbreviations as needed per the journal. We have defined MW and clarified all the other points in the legend.

      (22) 1B: "peaks taken at a fixed interval > 5 s", this is a bit confusing. If the interval is fixed, the exact time interval should be given. If it is > 5 s, then this suggests that it is not fixed. Do you mean "at intervals > 5 s"?

      Yes, fixed.

      (23) Figure 1-S1C: is the area the integral of the z-scored DF/F above zero DF/F? If so, it should have units of seconds (integral over dt of a dimensionless variable). Similarly, the Peak is a z-score value? In addition, is the time to peak in seconds? What is zero? Peak time of movement?

      We thank the reviewer for raising these points. We have clarified the terminology in the text and figure. Specifically, “area” was inaccurately labeled and refers to the mean z-scored ΔF/F within each analysis window (not a time integral). Peak values correspond to the maximum z-scored ΔF/F within the window, and time to peak is reported in seconds relative to the alignment point. We have also clarified the definition of time zero and included these definitions in Methods.

      (24) Figure 2-S1: It is not clear if this figure is obtained by averaging across all animals. Please explain in the legend.

      We clarified that values represent averages across mice.

      (25) Figure 2-S2: Are the speeds in A and B in units of cm/s (vertical axis)? This needs to be indicated.

      We have clarified in the figure legend that movement speed is expressed in cm/s.

      (26) Figure 5A, scale bar: It looks like a Delta is missing in front of F because the label reads 0.5 F/F instead of 0.5 DF/F. I am unclear why there are three colored traces for the speed panels. If the colors denote neuron classes, does this mean they were recorded in different sessions, allowing the authors to distinguish activation speed for each class separately?

      We fixed the scale bar typo. The speed traces in the bottom panels are shown to illustrate that movement is highly similar across activation types within each avoidance mode, indicating that the observed large differences in neural activity cannot be attributed to differences in movement. Minor differences in the speed traces arise because activation types are composed of neurons that can be recorded in the same or different sessions, and each activation type may not be present in every session. We added several sentences to this section that should fully clarify the issue.

      (27) Figure 4-S1 legend B: Please indicate why the two panels are missing for the PA case (for the confused reader).

      We have clarified in the legend that panels are not shown for correct CS2 passive avoids because these trials do not involve an action, and therefore from-action alignment cannot be defined.

      (28) Figure 5-S A, B: Units missing for speed.

      Fixed.

      Reviewer #3 (Recommendations for the authors):

      I cannot assess the scientific validity of the study design as it is too far away from my direct field of expertise. But I found the authors' arguments convincing, and the results sound pretty consistent with the little I know of the field. The recording methods are good and the statistical analysis robust. So my only recommendation for the authors would be to work on the figures to improve clarity.

      Thank you. We have introduced various changes that we hope will facilitate readability for a wider audience while preserving the necessary details.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study identifies mutations in alpha-tubulin that suppress Tau-induced neurodegeneration using the C. elegans model of Tauopathy, suggesting a potentially interesting role for microtubule properties in modulating Tau toxicity. These missense mutations cluster in the C-terminal Tau-interacting helix 12 region of alpha-tubulin genes (tba-1, tba-2, and mec-12). Further analysis, particularly using the strongest suppressor tba-2, shows that it rescues Tau-induced behavioral deficits and neuronal loss without significantly altering bulk tau-phosphorylation, aggregation, or binding to soluble tubulin. The authors suggest that altered microtubule properties underlie the neuroprotective effects, and manipulating microtubule properties may have therapeutic potential.

      Strengths:

      The study is conceptually interesting as it shows that Tau-induced neurotoxicity can, in this model, be partially uncoupled from canonical pathological hallmarks such as Tau-hyperphosphorylation and aggregation. The identification of multiple independent mutations in the same structural region of three alpha-tubulin genes provides support for the functional relevance of helix 12 in modulating Tau-induced toxicity. The authors demonstrate significant rescue of behavioral deficits (using motility and manual thrashing assays) and neuronal loss in both WT-tau and FTLD-associated TauV337M in combination with mutant alpha-tubulins, suggesting a general mechanism for tubulin-regulated modulation of Tau-toxicity. Moreover, the correlation between mutant tubulin expression levels and the extent of rescue supports a causal relationship.

      Weaknesses:

      One of the major claims of this manuscript is that altered microtubule properties suppress Tau toxicity. The only supporting evidence in this context provided by the authors is reduced taxol-stabilized microtubule mass, which does not fully explain neuronal loss or the rescue of behavioral deficits. What remains unclear is whether these mutations alter microtubule dynamics, catastrophe, lattice stability, or axonal transport.

      We agree with Reviewer #1’s critique that the evidence presented does not fully explain neuronal loss and requires further investigation. This first manuscript characterized the mutations discovered through forward genetic screening techniques and provided data to support the positive correlation mutant expression and level of suppression. We believe the studies and data presented here help to formulated the next testable hypotheses, and guide the next lines of experimentation. We are encouraged by Reviewer #1’s assessment that exploration of microtubule dynamics, catastrophe, lattice stability and axonal transport will be critical to testing the hypothesis that mutant tubulin drives suppression of tau toxicity through changes to microtubule properties. These suggestions are highly relevant and align with our priorities as we recently submitted an application for a 5-year research award to support these key questions.

      To address this specifically, the reviewer recommended “The microtubule-dependent axonal transport should be examined in tubulin mutants and compared with mutant tubulin + Tau conditions. Imaging of mitochondrial or synaptic vesicle markers, along with appropriate quantifications (velocity or run length), may provide a functional readout linking microtubule changes to neuronal survival.”

      We agree with the reviewer that these experiments will be highly valuable to further understand the mechanisms underlying suppression, and we have planned to complete these experiments upon receipt of funding that would directly support the completion of these experiments.

      The authors show that mutant tba-2 reduces total tau levels by ~45%. This level of reduction is likely significant but underexplored in the manuscript. Why are the Tau levels reduced? How is Tau getting cleared- is there enhanced autophagy or ubiquitin-proteasome pathway getting upregulated in tba-2 + Tau animals? Or one or more of the Tau species not detectable by the antibodies used in this study? The observation that the mec-12 mutant rescues Tau-induced phenotypes without altering Tau levels suggests that suppression can occur through Tau-independent mechanisms. This raises an important unresolved question regarding the extent to which suppression is Tau-dependent vs Tau-independent across different mutant alpha-tubulin genes, complicating the interpretation of the rescue phenotypes.

      We think the reviewer has addressed an important point that there may be both tau-dependent and tau-independent mechanisms at work here, and we will add greater nuance to this in our discussion. Additionally, we agree these two potential mechanistic pathways merit further exploration. To address this, we have planned to conduct experiments using reporter C. elegans lines crossed with our mutant tubulin/tau-transgenic lines to detect potential upregulation of these pathways as mechanisms for tau clearance.

      Given that Tau primarily associates with the microtubule lattice in vivo, measuring interactions with soluble tubulin may not fully capture biologically relevant binding dynamics and therefore does not exclude the possibility that these mutations alter tau-microtubule interactions at the lattice level or may affect the binding of other MAPs/regulators, thereby altering stability or trafficking.

      In the discussion we acknowledge the limitation of only examining the binding affinity between soluble tubulin and tau and intend to complete further studies with polymerized microtubules containing mutant α-tubulin. We will expand discussion of this in the text. Similar to reviewer 1, we have also concluded that the next line of experimentation will focus on mutant alpha-tubulin effects on the microtubule polymer such as changes to MAP interactions, stability and trafficking. We have applied for and hope to receive funding to address these questions in the near future.

      To address this concern specifically, we plan to conduct these experiments using C. elegans extracts to polymerize microtubules and subsequently test the binding of recombinant human tau. These co-sedimentation experiments are expected to be included in the revised manuscript.

      A large body of conclusions is drawn from behavioral rescue and biochemical assays. This limits the understanding of how molecular changes in tubulin might affect cellular mechanisms of neuroprotection. Are there changes in the neuronal microtubule organization, Tau localization, or its redistribution in the mutant alpha-tubulin background? Are there differences in soluble vs oligomeric vs insoluble Tau in mutant tba-2 and mec-12 animals?

      The reviewer raises relevant questions regarding elucidation of the mechanisms underlying mutant tubulin-mediated suppression at the cellular level. To address this concern we will analyze the cellular distribution of tau in neurons from mutant and non-mutant C. elegans.

      Ultimately, our goals are to identify and connect the underlying biochemical mechanisms with the observed prevention of cell death as Reviewer 1 has identified. Their suggestion to explore cellular-level changes such as mutant tubulin effects on tau distribution is highly relevant. We therefore plan to test this directly by imaging neurons in C. elegans strains expressing fluorescently labeled tau and/or immunohistochemical techniques to stain for tau in C. elegans neurons.

      The suppression of behavior in the co-pathology model is interesting but mechanistically insufficient, mainly because the underlying basis of suppression is not examined in these models. Moreover, it remains unclear whether tubulin-Tau genetically interacts with Aβ or TDP-43, and what cellular mechanisms account for the partial rescue observed in these co-pathology models.

      In agreement with Reviewer #1’s assessment, we have concluded these data, while interesting, do not substantially expand our understanding apart from the existing data. Without additional information regarding the underlying mechanisms, they do not provide substantial novel insights and we have therefore chosen to remove the co-pathology data sets from the revised version of the manuscript to refine the scope of the data and hypotheses discussed in this work.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Benbow et al. identifies, through a genetic screen, key tubulin mutants that, with high confidence, rescue tau-mediated ND phenotypes. This manuscript is well written, and the experimental results strongly support the authors' claims that these tubulin mutants can rescue ND-linked phenotypes in C. elegans while having little to no direct effect on Tau aggregation.

      Strengths:

      Benbow et al. use a relatively unbiased forward genetic screen to identify mutations associated with phenotypes that suppress tauopathy-related defects. The authors then logically focus on the various α-tubulin missense mutations identified in H12, which are known to localize to the external face of microtubules. The authors also carefully compare their established tauopathy-associated phenotypes in the WT TauH model, with and without specific α-tubulin mutations, using appropriate controls throughout. Lastly, the authors provide partial mechanistic insight into the α-tubulin mutant-mediated rescue, showing that these effects are independent of tau aggregation and tau phosphorylation, and instead suggest that the α-tubulin mutations may confer altered microtubule assembly properties based on the sedimentation assays.

      Weaknesses:

      While the claims are largely supported by the experimental outcomes, the authors at times do not provide enough detail in the text for readers to interpret the data sets independently. In addition, some claims appear to be slightly overstated relative to the data or the degree of error associated with those data.

      We appreciate the feedback regarding the need for additional clarity for independent analysis of the datasets. We will revise the figures and text to increase clarity for the readers. We will review statements and edit language in accordance with their degrees of error as appropriate.

      The authors measure tau binding affinities using soluble tubulin but do not assess tau binding to assembled microtubules. This is an important limitation, as the physiologically relevant interaction involves α/β-tubulin heterodimers, either free or incorporated into the microtubule lattice. Furthermore, the binding analysis appears to focus only on the D429N α-tubulin mutant, which further limits physiological relevance, as β-tubulin, which is also required for normal tau binding, is not explicitly considered.

      We acknowledge that the limited conclusions may be drawn from soluble tubulin interactions with tau and additional analysis with polymerized microtubules will be useful in understanding tau-microtubule binding affinity. The analysis was completed with isolated pools of tubulin from C. elegans, not recombinant mutant tubulin, so this is a heterogenous mixture of tubulin composed of α/β heterodimer subunits, and a mixture of the mutant isotype within the larger pool of wild type isotypes. While this further complicating the analysis, and is the likely source of variability, it incorporates the normal heterodimer subunit biochemistry.

      Given that tau prominently binds the microtubule lattice we agree with the reviewers that the assessment that experiments with polymerized microtubules containing mutant tubulin would offer a greater understanding of the effects of mutant alpha-tubulin on microtubule properties and potential mechanisms of toxic tau suppression. To test this directly we intend to complete co-sedimentation experiments using C. elegans extracts from wild type and mutant tubulin expressing C. elegans incubated with recombinant human tau.

      In conclusion, the thoughtful commentary and suggestions from reviewers will help improve the manuscript. We plan to complete the following experiments to address their concerns.

      (1) Assess tau localization in mutant tba-2 and mec-12 C. elegans as compared to tau-transgenic C. elegans without tubulin mutations. We plan to use immunohistochemical techniques and/or imaging of Dendra2-labeled tau to assess the sub-compartmental distribution of tau in C. elegans neurons. This addresses Reviewer #1’s question of whether the mutant tubulin changes tau localization in neurons.

      (2) Assess changes mutant-tubulin driven changes to tau affinity for polymerized microtubules. To address both reviewers concerns regarding the limitations of biding experiments with tau and soluble tubulin, We plan to use C. elegans extracts to tests whether microtubule polymers containing mutant alpha-tubulin alter tau-microtubule co-sedimentation.

      (3) Using C. elegans reporter lines we plan to assess whether tau clearance occurs in tba-2 mutant tubulin C. elegans through the upregulation of autophagy or ubiquitin degradation pathways.

      (4) Evaluate the neuroprotective effects of mutant alpha-tubulin in cholinergic neurons using a C. elegans strain expressing a fluorescent label specifically in cholinergic neurons.

      We plan to make textual revisions to increase clarity, aid in independent analysis of the presented datasets, and better address the possibility of both tau-dependent and tau-independent mechanisms. We appreciate the Reviewers attentive reading and thoughtful feedback for the improvement of this manuscript.

    1. Reviewer #3 (Public review):

      Summary

      In this paper, the authors have 5 human subjects learn to play Super Mario Bros while undergoing fMRI for 15 hrs each. They compare a reinforcement learning (RL) model (PPO), an imitation learning (IL) model, and a vision model (ResNet) in their ability to play the game, match human behavior, and, critically, explain human brain activity.

      The key findings can be summarized as follows:

      (1) RL, IL, and vision models explain similar amounts of variance in the BOLD signal (Fig 2a), with a significant but small trend of RL > IL > ResNet (Tab 1).

      (2) Untrained models with the same architecture explain a smaller but very similar amount of variance (Figure 2a, Table 1).

      (3) The brain maps across all models (and layers) are strikingly similar, with the strongest effects in visual, parietal, and motor regions (Figures 2b, 2d; Supplementary Material II).

      (4) Behavioral and neural performance are correlated across model checkpoints (but not levels), such that later checkpoints in training have better behavioral and neural encoding performance (Figures 3 & 4), although the neural effect plateaus pretty quickly.

      (5) Out-of-distribution performance is quite poor, both behaviorally (Figure 5a) and neurally (Figure 5b).

      I believe this work will be of interest to neuroscientists, cognitive scientists, and AI researchers alike. There has been a growing trend in neuroscience to adopt AI models as cognitive models of complex perception and action, while at the same time, AI researchers are increasingly looking at the brain for inspiration. The key finding of this paper -- that these models fail to generalize to out-of-distribution levels -- questions the core assumptions of this whole enterprise.

      Strengths:

      Unlike previous studies applying machine learning to naturalistic game-play, the authors take great care to make sure their models are evaluated on an equal footing, using equivalent or similar architectures/number of parameters and training data.

      While the number of subjects (5) is relatively small, the amount of data per subject (15 hours) is impressive, which is important for fitting the imitation learning & ResNet models and for obtaining reliable encoding performance for each individual subject. The authors employed a train/val/test split and held out sets, the gold standard in the literature.

      Overall, the paper was well-written and easy to follow. The figures clearly illustrate the main findings.

      Weaknesses:

      (1) Missing statistical tests

      I think the main weakness of the paper is that many of the claims are qualitative in nature and lack appropriate statistical tests, for example:

      - "The conv3 layer has the highest brain encoding score";<br /> - "Robust association between task performance and brain encoding" ;<br /> - "Level patterns strongly predict brain encoding";<br /> - "Brain encoding performance was severely degraded";<br /> - "Effect of training on brain encoding was apparent".

      While these effects are indeed qualitatively visible in the figures, it is unclear which of these differences are significant (with the notable exception of Table 1). I believe the paper would benefit substantially if these effects were quantified and every claim were supported by the appropriate statistical tests. As an example, with the exception of Table 1 and the corresponding paragraph, I could not find any p-values in the results section.

      (2) Missing model performance and human-likeness

      Also absent from the results is an assessment of model performance on the task and similarity to human performance/behavior. From Figures 3 and 4, we can see that the game score of PPO is around 500-1000 - how does that compare to the humans? We can also see that the imitation scores for IL are around 0.4-0.7, but what does that mean? Such results would be crucial to assess if the models have indeed learned to play the games and/or imitate the humans, and therefore, whether they would be good candidates as cognitive models (before even looking at brain activity). At minimum, plotting the human versus model game scores (see e.g. Tomov et al. 2023 Neuron, Figure 2) would be helpful; or, if you'd like to dig deeper, showing that human actions are more valuable or more likely under those models (see e.g. Cross et al. 2022 Neuron, Figure 2). It might also be helpful to look at imitation scores for the RL model and game performance of the imitation model -- I suspect they will both be bad, but they can at least serve as informative baselines for their counterparts.

      (3) Possible undertraining

      Relatedly, one possible explanation for why the Untrained model does so well is that all the models may be effectively undertrained. For example, while there are no training curves in the paper, it seems from the spacing of the checkpoint game scores (x-axis on Figure 3c) that the RL model may not have converged yet (it would be helpful if those were somehow colored by training epoch). Showing training curves would be helpful (i.e., something similar to Figure 3a, except with performance on the y-axis).

      Additionally, it would be great to provide more details regarding the PPO training protocol. How many episodes? How many steps per episode? How many steps for all of the training? Similarly, for the imitation learning model: batch size, number of epochs, optimizer, scheduler, etc.

      (4) Mysterious poor encoding performance of Untrained and ResNet models on the held-out set

      Critically, and related to that, I'm a little confused about the Untrained model results on the held-out set (Figure 5b, top row on the right). Why should those be any different from the test set results with the Untrained model (Figure 2a, right, fourth row from the top)? It makes sense why the other models are worse on the held-out set -- they have never been trained on any frames from those levels. However, the untrained model has not been trained on *any* frames from *any* levels, including the test set and the held-out set.

      The same is true for the ResNet model, which is pre-trained on a completely separate data set and yet similarly shows worse performance on the held-out set compared to the test set.

      This cannot be explained by the ridge regression, which has no parameters or hyperparameters fitted on either the test set or the held-out set.

      The big discrepancy in the untrained model & ResNet results between the test and the held-out set makes think that there is something substantially different about the levels in that held-out set; that they are truly out of distribution compared to the other 20 levels (e.g., maybe they're the last 2 hardest levels and look completely differently? e.g. ResNet proxy in Fig 5c shows worse performance than the mean, which is indicative of an anti-correlation). Alternatively, it may be some issue with the analysis pipeline. The poor generalization results are central to the claims of the paper, so I believe this should be clarified.

      (4) Brittleness conclusion rationale

      I'm not quite on board with the author's rationale that "[poor model performance on the out-of-distribution levels] demonstrates that the models we tested are limited in scope and may not provide a valid inference of brain-like processing, as human behavior remains robust and generalizable across levels".

      For one, unlike the models, humans were actually trained on those levels, so it would not be surprising if they perform just as well on them as on the other levels (but do they? Again, it would be great to see some behavioral data from the humans and the models).

      Second, as the authors themselves show, task performance and human-likeness do not really correlate with neural encoding across levels (Fig 4a & b, respectively), so even if model performance remained "robust and generalizable" on the held-out levels, that will not necessarily translate to good neural encoding.

      Thirdly, and perhaps most importantly, unless the test set and held-out set were sampled exclusively from the practice phase when the subjects have mastered all the levels (that doesn't seem to be the case, but the authors should clarify), then the humans are continuously learning, which means that their own internal representations of the game are evolving. That's not the case for the models, which I assume are in "inference mode" when their representations are extracted for neural encoding. That is, their weights are frozen. So there's a fundamental mismatch between the mode in which humans are operating (continuously learning and executing) and the mode in which the models are operating (just executing). While this is true for all the levels, it may partially account for the discrepancy in the held-out set specifically.

    1. Author response:

      The following is the authors’ response to the current reviews.

      Public Review:

      Reviewer #1 (Public review):

      Suggestions to clarify the study:

      In the revised version, the authors carefully consider these suggestions and provide further details, clarifications and even some new results. Regarding the question of how infection of a cell with one virus could lead to lower probability for a secondary infection, I think that it is possible that infected cells activate antiviral programs that lead, for example, to lower expression of surface receptors. This has been considered at least in hepatitis C virus infection. However, this is a minor point.

      Yes, the possibility that infection of a cell by a virion would reduce chance of infection by another virion was allowed in our model. However, such as a process will not result in apparent cooperativity (n>1) in our model, and thus, is irrelevant to the issue of apparent cooperativity we identified.

      Reviewer #2 (Public review):

      In their article, Peterson et al. wanted to show to what extent the classical "single hit" model of virion infection, where always the same quantity of virion is required to infect a cell, does not match with empirical observations based on human cytomegalovirus in vitro infection model, and how this would have practical impacts in experimental protocols.

      Strengths:

      The use of a very simple and robust experimental assay, where they infected cells with serially diluted virions and measured the proportion of infected cells with flow cytometry. This convincingly showed how the proportion of infected cells differed from a "single hit" model which they simulated using a simple mathematical model ("power-law model"), and better fitted a model where virions need to cooperate to infect cells.

      The use of different cell types and virus strains, which allows to draw some generalizations.

      The exploration of the mechanisms that could explain this apparent cooperation, using biologically plausible simulations.

      The practical consequences that this phenomenon has for lab virologists as well as modelers.

      Thank you.

      Weaknesses:

      The impossibility to discriminate between biological mechanisms is an important limitation of this study and calls for developing experimental designs able to further understand this question.

      The outcome of the virion clumping remains highly sensitive to the choice of the clumps size distribution, which is itself very complicated to estimate, especially at high dilution.

      The impossibility to directly fit the mathematical models to the data limit them to a qualitative discussion.

      Overall, this work is very valuable as it raises the general question of how the estimate of infectivity can be biased if extrapolated from a single virus titer assay. The observation that HCMV virions often cooperate and that this cooperation varies between context seems robust. The putative biological explanations would require further exploration.

      This topic is very well known in the case of segmented viruses and the semi-infectious particles, leading to the idea of studying "sociovirology", but to my knowledge this is the first time that it was explored for a non-segmented virus, and in the context of MOI estimation.

      Thank you. We would note, however, that inability to discriminate between alternative models is not a weakness per se. It shows that our work goes beyond a somewhat typical approach in mathematical modeling to offer a single explanation for a phenomenon in question (rather than focusing on discriminating between alternatives that is often hard to do).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I now understand better the graphical abstract. I think my eye was too much attracted by the increase in specific infectivity that you see for more than 1 genome/cell, which is not the point of your paper. I am wondering if you should not guide even more the reader, by pointing out that the fact that the initial decline in specific infectivity represents apparent cooperativity.

      Let’s hope that the readers are smart enough to understand what to focus their eyes on. At the end, this is a graphical abstract that is not supposed to have too much text explaining where to look.

      (2) For your one-inflated geometric distribution, I agree that the estimations would remain very hypothetical because you would have to make many assumptions, however I think a hurdle model where you would fit the P(clump size = 1)=f1 and P(clump size = (i) following a one-truncated geometric distribution would be more appropriate because it would lead to a distribution closer to your PDF from figure S11C.

      The issue is that our data are not in clump sizes but in diameter of the clump D. This is why we opted for using a mixture of continuous distributions, not a mixture of discrete distributions. We are sharing the DLS data, so others are welcome to do another try of fitting other types of distribution to the data.

      (3) For the DLS data, I understand your choice to include all the datapoints, however I find the interpretation confusing: if I understand correctly, you consider that f1, the fraction of the smaller distribution, represents clumps of one virion. However, its median size is 10 times smaller than a virion. So, the number of clumps with one virion would be overestimated. I think it would be helpful for the reader to clarify this aspect, either in the results around lines 503-512, or in the discussion. Could it be that at higher dilution, what is represented by this smaller distribution would almost only be debris because the virions are so rare?

      When fitting a mixture of two log-normal distributions f<sub>1</sub> represents the proportion of clumps of larger size (as was described in the materials and methods). The actual estimated value of f<sub>1</sub> is not highly relevant in calculating change in PDF of the distribution only for D>=d (230nm) as shown in Suppl Fig S11C. But we now realize that this variable f<sub>1</sub> may be confused with a variable f<sub>1</sub> used to denote the fraction of clumps with virion size=1 (in Fig 5C). We now mention that in the caption of Supp Fig S10.

      (4) For the dashed diagonal lines of fig 2, what I don't understand is the choice of the intercept that seems a bit random. I was wondering if it would not be more helpful to make it so that the dashed line intersects the observation for 1 genome/cell, which could then be interpreted as a deviation from the "single hit" model extrapolated outside of 1 genome/cell?

      The diagonal lines in Fig 2 are exactly the same in ALL panels, as are the x/y axes ranges; the slope of the line (equals to 1) allows visually to see when the regression (shown by think black lines) deviates from slope=1, i.e., indicates apparent cooperativity. We will keep the lines are they are. Thank you for the suggestion, though.


      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors conduct both experiments and modeling of human cytomegalovirus (HCMV) infection in vitro to study how the infectivity of the virus (measured by cell infection) scales with the viral concentration in the inoculum. A naïve thought would be that this is linear in the sense that doubling the virus concentration (and thus the total virus) in the inoculum would lead to doubling the fraction of infected cells. However, the authors show convincingly that this is not the case for HCMV, using multiple strains, two different target cells, and repeated experiments. In fact, they find that for some regimens (inoculum concentration), infected cells increase faster than the concentration of the inoculum, which they term "apparent cooperativity". The authors then provided possible explanations for this phenomenon and constructed mathematical models and simulations to implement these explanations. They show that these ideas do help explain the cooperativity, but they can't be conclusive as to what the correct explanation is. In any case, this advances our knowledge of the system, and it is very important when quantitative experiments involving MOI are performed.

      Strengths:

      Careful experiments using state-of-the-art methodologies and advancing multiple competing models to explain the data.

      Weaknesses:

      There are minor weaknesses in explaining the implementation of the model. However, some specific assumptions, which to this reviewer were unclear, could have a substantial impact on the results. For example, whether cell infection is independent or not. This is expanded below.

      Suggestions to clarify the study:

      (1) Mathematically, it is clear what "increase linearly" or "increase faster than linearly" (e.g., line 94) means. However, it may be confusing for some readers to then look at plots such as in Figure 2, which appear linear (but on the log-log scale) and about which the authors also say (line 326) "data best matching the linear relationship on a log-log scale".

      This is a good point. We included a clarification to indicate that linear on the log-log scale relationship does not imply linear relationship on the linear-linear scale. We wrote:

      “Because most data did not exhibit a linear relationship between virion concentration and infection probability we fitted the models to subsets of data best matching a linear relationship on a log-log scale. Note that linear relationship on log-log scale may still be nonlinear (on linear-linear scale) when n!=1.”

      (2) One of the main issues that is unclear to me is whether the authors assume that cell infection is independent of other cells. This could be a very important issue affecting their results, both when analyzing the experimental data and running the simulations. One possible outcome of infection could be the generation of innate mediators that could protect (alter the resistance) of nearby cells. I can imagine two opposite results of this: i) one possibility is that resistance would lead to lower infection frequencies and this would result in apparent sub-linear infection (contrary to the observations); or ii) inoculums with more virus lead to faster infection, which doesn't allow enough time for the "resistance" (innate effect) to spread (potentially leading to results similar to the observations, supra-linear infection).

      In our models we assumed cells to be independent of each other (see also responses to other similar points). Because we measure infection in individual cells, assuming cells are independent is a reasonable first approximation. However, the reviewer makes an excellent point that there may be some between-cell signaling happening in the culture that “alerts” or “conditions” cells to change their “resistance”. It is also possible that at higher genome/cell numbers, exposure of cells to virions or virion debris may change the state of cells in the culture, and more cells become “susceptible” to infection. This is a good point that we now list in Limitations subsection of Discussion; it is a good hypothesis to test in our future experiments. We write:

      “Accrued damage model is also consistent with the idea that at higher genome/cell values, the inoculum itself (including cell and/or virion debris) may impact overall susceptibility of all cells in the well, for example, making them more susceptible to infection. It may be expected, though, that exposing cells to debris would increase cell resistance to infection; this would result in n < 1 that we did not observe at small genomes/cell values.”

      (3) Another unclear aspect of cell infection is whether each cell only has one chance to be infected or multiple chances, i.e., do the authors run the simulation once over all the cells or more times?

      Each cell has only one chance to be infected. Algorithm 1 clearly states that; we will add an extra sentence in “Agent-based simulations” to indicate this point.

      (4) On the other hand, the authors address the complementary issue of the virus acting independently or not, with their clumping model (which includes nice experimental measurements). However, it was unclear to me what the assumption of the simulation is in this case. In the case of infection by a clump of virus or "viral compensation", when infection is successful (the cell becomes infected), how many viruses "disappear" and what happens to the rest? For example, one of the viruses of the clump is removed by infection, but the others are free to participate in another clump, or they also disappear. The only thing I found about this is the caption of Figure S10, and it seems to indicate that only the infected virus is removed. However, a typical assumption, I think, is that viruses aggregate to improve infection, but then the whole aggregate participates in infection of a single cell, and those viruses in the clump can't participate in other infections. Viral cooperativity with higher inocula in this case would be, perhaps, the result of larger numbers of clumps for higher inocula. This seems in agreement with Figure S8, but was a little unclear in the interpretation provided.

      This is a good point. We did not remove the clump if one of the virions in the clump manages to infect a cell, and indeed, this could be the reason why in some simulations we observe apparent cooperativity when modeling viral clumping. We have explored this in the revision and found that it does not really impact how infection rate scales with the genomes/cell (e.g., see Suppl Fig S8).

      (5) In algorithm 1, how does P_i, as defined, relate to equation 1?

      These are unrelated because eqn.(1) is a phenomenological model that links infection per cell to genomes per cell. P_i in algorithm 1 is “physics-inspired” potential barrier.

      (6) In line 228, and several other places (e.g., caption of Table S2), the authors refer to the probability of a single genome infecting a cell p(1)=exp(-lambda), but shouldn't it be p(1)=1-exp(-lambda) according to equation 1?

      Indeed, it was a typo, p(1)=1-exp(-lambda) per eqn 1. Thank you, it has been corrected in the revised paper.

      (7) In line 304, the accrued damage hypothesis is defined, but it is stated as a triggering of an antiviral response; one would assume that exposure to a virion should increase the resistance to infection. Otherwise, the authors are saying that evolution has come up with intracellular viral resistance mechanisms that are detrimental to the cell. As I mentioned above, this could also be a mechanism for non-independent cell infection. For example, infected cells signal to neighboring cells to "become resistance" to infection. This would also provide a mechanism for saturation at high levels.

      We do not know how exposure of a cell to one virion would change its “antiviral state”, i.e., to become more or less resistant to the next infection. If a cell becomes more resistant, there is no possibility to observe apparent cooperativity in infection of cells, so this hypothesis cannot explain our observations with n>1. Whether this mechanism plays a role in saturation of cell infection rate at lower than 1 value when genome/cell is large is unclear but is a possibility. We added this point to Discussion in revision (see our text above that includes this point).

      (8) In Figure 3, and likely other places, t-tests are used for comparisons, but with only an n=5 (experiments). Many would prefer a non-parametric test.

      We repeated the analyses in Fig 3 with Mann-Whitney test, results were the same, so we would like to keep results from the t-test in the paper.

      Reviewer #1 (Recommendations for the authors):

      (1) The strains of HCMV used have a fluorescent reporter "in place of the US11 gene". Can you provide a brief comment on whether and how this gene deletion affects HCMV replication?

      US11 is a resident ER protein that is considered an "immune evasion factor". It promotes ERAD of MHC I and has no observable effect on replication of HCMV in cultured cells (Berger 2000 JVI, Wiertz 1996 Cell). We now add this information in Materials and methods section of the paper. We write:

      “All BAC clones were modified to express green fluorescent protein (GFP) or the monomeric red fluorescent protein mCherry (mCherry) with En passant recombineering by replacing US11 with the eGFP or mCherry gene, respectively. US11 is a resident ER protein that is considered an “immune evasion factor”. It promotes ERAD of MHC I and has no observable effect on replication of HCMV in cultured cells [27, 28]. Infectious HCMV was recovered by electroporation of BAC-DNA into MRC5 cells which were then co-cultured with either HFFCs (TB and TR) or HFF-tet cells (ME).”

      (2) I didn't understand what the section "Virus titer assays" refers to. When was this used? How or why is this different from the "Virus stock dilution and dose-response assay"? Also in this section, you refer to NHDF cells - can you provide more information about these? And how does a different type of cell affect the titer assay (here measured as infected cells), since this is one of the main points of your paper?

      Apologies for the confusion. In Ryckman lab we routinely generate viral stock and titrate it using a specific cell type, Normal (or neonatal) Human Dermal Fibroblasts (NHDF). This way, the titer of the stock is consistent between experiments by different researchers in the lab. We then use standard 10-fold dilutions to define the number of infectious units per mL of the stock. We now name this subsection as “Quantification of viral stock infectivity using standard 10-fold dilutions”. After the stock was quantified, we then used that stock in our actual experiments with very small dilution factor df that allowed us to detect deviations of the rate of infection from single hit model.

      (3) In many places, "powerlaw" is written. This is usually written as two words, "power law".

      Because powerlaw comes together with “model”, we decided to use “power-law model”.

      (4) Line 75: "have" instead of "has"?

      (5) Line 84: "with" repeated.

      Corrected, thank you.

      (6) Line 116: This section "Cell lines" seems to describe three cell lines, "HFF cells and MRC5 cells" and then "EC" cells.

      HFF cells are fibroblasts used in our main experiments and MRC5 cells are another type of fibroblasts. We used MRC5 cells in the first step of recovering infection HCMV from BAC DNA (electroporation). We clarified this in Materials and methods. We write:

      “Cell lines. Human foreskin fibroblast cells (HFFCs or fibroblasts) and MRC5 cells (also fibroblasts) were cultured in Dulbecco’s modified Eagle’s medium (DMEM, Sigma) supplemented with 5% heat-inactivated fetal bovine serum (FBS, Rocky Mountain Biologicals, Missoula, MT, USA) and 5%Fetalgro® (Rocky Mountain Biologicals, Missoula, MT, USA). We used MRC5 cells in the first step of recovering infection HCMV from BAC DNA (electroporation). For main experiments we used HFFCs as fibroblasts. Human retinal pigment epithelial cells (ECs or ARPE-19, American Type Culture Collection, Manassas, VA, USA) were cultured in a 1:1 mixture of DMEM and Ham’s F-12 medium (DMEM:F-12, Gibco) and supplemented with 10% FBS.”

      (7) Line 188: Because the virus is double-stranded, do you have to divide the qPCR result by 2 to get genomes?

      This is typically accounted for in our calculations of genome/cell.

      (8) Line 200: Typically, one would write "500g" and not "500xg".

      Corrected.

      (9) Line 248: It would be clearer to write "cell type C different from cell type C2".

      Here C and C_2 refer to actual numbers of cell in the titration/growth experiments, so it is comparing numbers, not cell types. We kept the relationship as it is.

      (10) Definition of cell class: what is n in p_n, the total number of cells, or are these divided into n classes of resistance?

      This part was incorrectly copied from an earlier version, both cell resistance and virion infectivity was sampled from normal distributions with different mean and variances (see Table 1). We corrected the text to reflect this.

      (11) Line 272 to 273: Something seems to be missing, as the change of line doesn't make sense.

      Thank you. Edited to improve readability. Now it reads

      “Clumping hypothesis. In the basic model the number of virions a given cell is exposed to follows a Poisson distribution. However, it is well recognized that as virions are produced by infected cells, they may form clumps/aggregates; the number of virions per clump/aggregate may deviate from, for example, the Poisson distribution [33].”

      (12) Line 283: How lambda is chosen is not indicated here, only later (line 424), but at this point, one can confuse it with lambda in equation 1. Is it the same? It also doesn't seem to be indicated in your Table 1.

      The mean of the Poisson distribution in clump simulations lambda is not the same as lambda in eqn 1; we re-named the mean of Poisson distribution as lambda_c which is estimated by fitting a Poisson distribution to clump size distribution estimated from DLS experiments. Because it was dependent on the virus stock dilution, it is not listed in Table 1. However, we did perform additional simulations assuming lambda_c=2 (Suppl Fig S10).

      (13) Equation 6: I understand that you mostly used kappa=0, but in equation 6, would it be positive or negative (if not zero)?

      We probably expect kappa to be negative but we did not fully explore this extension of the model.

      (14) Line 350: Instead of "infection rates" would "infection frequencies" be better?

      We agree. Changed (also changed in the sentence above that line).

      (15) Line 366: I found this sentence a bit awkward.

      We edited it to the best of our ability to improve it.

      “Importantly, for most HCMV strain-target cell combinations we estimated n>1 (Figure 2 and Supplemental Table S2). With n>1 increase in virion concentration (i.e., higher genomes/cell values) results in a higher than linear increase in the probability of a cell to be infected (eqn. (1)) indicating cooperation between virions at infecting cells. We call this phenomenon “apparent cooperativity”.

      (16) Figure 2, panel L: I wonder if it would be better to include the panel with the name of the experiment, but no data. Currently, it takes a while to find what you are talking about in panel L (or at the very least, indicate the panel in the caption).

      Changed

      (17) Figure 2: When you say that experiments were done at least twice, are you referring to the GFP and mCherry versions of the experiment, or replicates within each of those fluorescent labels?

      Replicates with each of those labels.

      (18) Figure 3: What is the number on top of the black bars? I think it is the average of the paired fold change. Is this right? Why, in panel E, is it 1.32 when only one goes up?

      Yes, fold change. Indeed, 1.32 was a typo, it is 0.70, thank you for noting.

      (19) Line 408: delete the word "there".

      Done. Thank you.

      (20) Line 412: Instead of "The", it should be "Then".

      Done. Thank you.

      Reviewer #2 (Public review):

      In their article, Peterson et al. wanted to show to what extent the classical "single hit" model of virion infection, where one virion is required to infect a cell, does not match empirical observations based on human cytomegalovirus in vitro infection model, and how this would have practical impacts in experimental protocols.

      They first used a very simple experimental assay, where they infected cells with serially diluted virions and measured the proportion of infected cells with flow cytometry. From this, they could elegantly show how the proportion of infected cells differed from a "single hit" model, which they simulated using a simple mathematical model ("powerlaw model"), and better fit a model where virions need to cooperate to infect cells. They then explore which mechanism could explain this apparent cooperation:

      (1) Stochasticity alone cannot explain the results, although I am unsure how generalizable the results are, because the mathematical model chosen cannot, by design, explain such observations only by stochasticity.

      Our null model simulations are not just about stochasticity; they also include variability in virion infectivity and cell resistance to infection. We agree that simulations cannot truly prove that such variability cannot result in apparent cooperativity; however, we also provide a mathematical proof that increase in frequency of infected cells should be linear with virion concentration at small genome/cell numbers.

      (2) Virion clumping seemed not to be enough either to generally explain such a pattern. For that, they first use a mathematical model showing that the apparent cooperation would be small. However, I am unsure how extreme the scenario of simulated virion clumping is. They then used dynamic light scattering to measure the distribution of the sizes of clumps. From these estimates, they show that virion clumps cannot reproduce the observed virion cooperation in serial dilution assays. However, the authors remain unprecise on how the uncertainty of these clumps' size distribution would impact the results, as most clumps have a size smaller than a single virion, leaving therefore a limited number of clumps truly containing virions.

      As we stated in the paper, clumping may explain apparent cooperativity in simulations depending on how stock dilution impacts distribution of virions/clump. This could be explored further, however, better experimental measurements of virions/clump would be highly informative (but we do not have resources to do these experiments at present). Our point is that the degree of apparent cooperativity is dependent on the target cell used (n is smaller on epithelial cells than on fibroblasts) that is difficult to explain by clumping which is a virion property. Per comment by reviewer 1, we have done more analyses of the clumping model to investigate importance of clump removal per successful infection on the detected degree of apparent cooperativity. We found that it was not critical to our conclusions (Suppl Fig S8).

      The two models remain unidentifiable from each other but could explain the apparent virion cooperativity: either due to an increase in susceptibility of the cell each time a virion tries to infect it, or due to viral compensation, where lesser fit viruses are able to infect cells in co-infection with a better fit virion. Unfortunately, the authors here do not attempt to fit their mathematical model to the experimental data but only show that theoretical models and experimental data generate similar patterns regarding virion apparent cooperation.

      In the revision we now provide examples of our earlier simulations that “match” experimental data with a relatively high degree of apparent cooperativity (Supp Fig S9).

      Finally, the authors show that this virions cooperation could make the relationship between the estimated multiplicity of infection and viruses/cell deviate from the 1:1 relationship. Consequently, the dilution of a virion stock would lead to an even stronger decrease in infectivity, as more diluted virions can cooperate less for infection.

      Overall, this work is very valuable as it raises the general question of how the estimate of infectivity can be biased if extrapolated from a single virus titer assay. The observation that HCMV virions often cooperate and that this cooperation varies between contexts seems robust. The putative biological explanations would require further exploration.

      This topic is very well known in the case of segmented viruses and the semi-infectious particles, leading to the idea of studying "sociovirology", but to my knowledge, this is the first time that it was explored for a nonsegmented virus, and in the context of MOI estimation.

      Thank you.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      Two aspects of the work would benefit from further thought:

      (1) The simulation of virion clumps: in both cases (Poisson distribution or one-inflated geometric distribution), the proportion of clumps containing more than one virion will be small. For the Poisson distribution, as you fit the powerlaw model on the range of genomes/cell < ~ 3 genomes/cell (Figure 4B). I wonder to what extent this explains the sudden rise in infections/cells you observe above that limit. It would be interesting to plot the (cumulative) distribution of the clump sizes at different dilution levels to have a better idea.

      The reviewer has a good eye, indeed, the relationship between infection frequency and genomes/cell is linear up to a point, and we believe the inflection point reflects the genomes/cell values when clumps contain more than 1 virion. Here is the results of simulations with distribution of virions/clump plotted:

      Similarly, for the one-inflated geometric distribution, the proportion of clumps of size 1 is the sum of two events: f1, plus 1-f1 times the probability that the geometric distribution is zero, if I follow the methods on lines 287-294. I wonder if this is appropriate regarding the estimates made with the DLC. In particular, Figure 5C shows that the proportion of clumps of size 1 is more than ~ half of all the clumps, and does not seem to be the same distribution as the estimates made on Figure S9C. Maybe a hurdle model would be more appropriate?

      This is a fair point. In our analyses we found that modeling clump size distribution is tricky and required various assumptions. The issue with the DLS data is that we do not really know the distribution of intact virions per clump so how to relate the size of the clump to the number of virions in a clump is wide-open; we explored several possibilities and found that the answer (whether clumping results in apparent cooperativity) depends on assumptions of how clumps are modelled (e.g., compare Fig 4B and Suppl. Fig S11). Hurdle model is not appropriate for clumps because by our definition of a clump, it must have at least 1 virion. Our key observation, however, is that the degree of apparent cooperativity depends on the target cell type – and thus should be independent of virion clumping (unless there is viral cooperativity in the clumps). Overall, we decided that exploring more clumping models would take extra effort, but it is unclear if it brings any benefits to our conclusions.

      The analysis of the clump size distribution using dynamic light scattering, in Figure S8. If I interpret correctly, events with size < 230 nm should be excluded as they do not represent clumps of virions but rather media impurities or cell debris. Therefore, I don't understand the choice of fitting the whole set with a combination of two normal distributions, as even the larger normal distribution covers clumps < 230 nm. If the f1 indicated here is the one used in the methods line 287-294, this is then wrong because it does not represent the fraction of clumps of size 1, but rather debris.

      We used two normal (on log-scale) distributions when quantifying clump distribution data (Supp Fig S10) to avoid sub-selection of the data; in this way, two distribution fit the whole dataset with excellent quality. An alternative approach would be to sub-select data with size >=230nm and fit a normal (or similar) distribution of the clumps; such an approach may generate biases and/or unreliable estimates at high dilutions due to small number of clumps with large size (e.g., see Supp Fig S10S-X). In our simulations to model clump distribution and infection (Fig 5) we attempted to simulate the estimated clump size distribution (Suppl Fig S11C) only approximately. Again, because in our measurements we don’t really know the number of virions per clump, efforts to model exactly clump size distribution, we believe, are not going to give full answers.

      (2) Figure 4 and results lines 419-465: Why didn't you try to fit the different models to the data, instead of qualitatively comparing the estimate of n in the simulations with arbitrary parameters to the one for empirical data? Your models match the expectation of virion cooperation by design, so they are not more convincing for a virologist than logical non-quantitative reasoning. They would be of stronger evidence in my opinion if you could show how well they fit the data. You could then directly compare the different models' fits using goodness-of-fit metrics and decide whether one is better than another or if they all explain equally well the observations.

      Well, we have 11 different relationships between infection rate and genome/cell, finding parameter combinations that would match all the data with at least 2 alternative models seems excessive at present but it is a good direction as we get extra funding to continue this work. It is also difficult to extensively search for the parameter values that would result in a perfect fit of the stochastic simulations to data since the methods of fitting agent-based models to data are not fully developed. However, following this suggestion we now show results of simulations for the two alternative models (accrued damage and viral compensation) that we believe do match experimental data somewhat (see new Suppl Fig S9).

      Minor comments:

      (1) Graphical abstract: This requires more context as it is too rough here to help me understand the general idea of the paper. Plus, why does specific infectivity first decrease with genome/cell?

      We added few elements to the graphical abstract including the strain and target cell used. The decrease in specific infectivity at lower genome/cell is due to apparent cooperativity.

      (2) Equation (7): It would be beneficial for the reader if the reasoning behind the likelihood computation were further described.

      This is a relatively standard approach to model/estimate parameters of a binary outcome, e.g., see Wikipedia: https://en.wikipedia.org/wiki/Logistic_regression

      (3) Line 352-357: could the drop in infectivity also be enhanced/explained by increased cell mortality? Did you gate on cell viability during FCM?

      The infection rate was measured in live cells only, so increased cell mortality may be an explanation.

      (4) Figure 2: I don't understand the dashed diagonal lines: what do they represent exactly? Especially, wouldn't the single-hit model depend on p(1), in which case it should vary by cell x virus?

      As the caption to Figure 2 clearly states, diagonal dashed lines show the slope =1 (i.e, single hit model), so one would be able compare how far the data and/or model fit line deviate from 1. The note for p(1) in panel A is to illustrate how p(1) is calculated; obviously it varies by the strain-cell combination as is indicated in Suppl. Tab S2).

      (5) Fig3G: Is it not surprising to find a positive relationship between p(1) and n? I would have intuitively expected that the stricter the environment is, the more cooperation you observe. But maybe these viruses did not evolve in this context, and therefore, this relationship is different from what you expect from an evolutionary optimum.

      Well, we simply don’t know. The relationship simply suggests that there is connection between infectivity of a single virion and the degree of apparent cooperativity. We are not certain what is the context in which these viruses have evolved.

      (6) Flow cytometry assay: could it be possible that cells infected by more virions generate more fluorescent proteins and are therefore less likely to be false negatives? Maybe you could compare the fluorescence intensity distribution among infected cells in the context of low MOI vs high MOI?

      This is an interesting point. From presented flow cytometry plots (e.g., Suppl Fig S3), the MFI for infected cells does not seem to depend on the dilution (or genome/cell).

      (7) Figure S9B: I did not understand this figure. Are the axes labels correct? How is it possible to have less than 1 virion/well?

      The y axis shows a scaled number calculated from integrating estimated clump size distribution, we assume 1 “scaled” virion/well at highest virion/cell values. With scaling, yes, it is possible to have less than 1 virion/well.

      Reviewer #3 (Public review):

      Summary:

      The authors dilute fluorescent HCMV stocks in small steps (df ≈ 1.3-1.5) across 23 points, quantify infections by flow cytometry at 3 dpi, and fit a power-law model to estimate a cooperativity parameter n (n > 1 indicates apparent cooperativity). They compare fibroblasts vs epithelial cells and multiple strains/reporters, and explore alternative mechanisms (clumping, accrued damage, viral compensation) via analytical modeling and stochastic simulations. They discuss implications for titer/MOI estimation and suggest a method for detecting "apparent cooperativity," noting that for viruses showing this behavior, MOI estimation may be biased.

      Strengths:

      (1) High-resolution titration & rigor: The small-step dilution design (23 serial dilutions; tailored df) improves dose-response resolution beyond conventional 10× series.

      (2) Clear quantitative signal: Multiple strain-cell pairs show n > 1, with appropriate model fitting and visualization of the linear regime on log-log axes.

      (3) Mechanistic exploration: Side-by-side modeling of clumping vs accrued damage vs compensation frames testable hypotheses for cooperativity.

      Thank you.

      Weaknesses:

      (1) Secondary infection control: The authors argue that 3 dpi largely avoids progeny-mediated secondary infection; this claim should be strengthened (e.g., entry inhibitors/control infections) or add sensitivity checks showing results are robust to a small secondary-infection contribution.

      This is an important point. We do believe that the current knowledge about HCMV virion production time – it takes 3-4 days to make virions per multiple papers (see Fig 7 in Vonka and Benyesh-Melnick JB 1966; Fig 3B in Stanton et al JCI 2010; and Fig 1A in Li et al. PNAS 2015) – is sufficient to justify our experimental design but we do agree that an additional control to block novel infections with would be useful. We had previously performed experiments with a HCMV TB-gL-KO that cannot make infectious virions (but the stock virions can be made from complemented target cells). We will investigate if our titration experiments with this virus strain have sufficient resolution to detect apparent cooperativity. However, at present we do not have the resources to perform novel experiments.

      (2) Discriminating mechanisms: At present, simulations cannot distinguish between accrued damage and viral compensation. The authors should propose or add a decisive experiment (e.g., dual-color coinfection to quantify true coinfection rates versus "priming" without coinfection; timed sequential inocula) and outline expected signatures for each mechanism.

      Excellent suggestion. Because infection of a cell is a result of the joint viral infectivity and cell resistance, it may be hard to discriminate between these alternatives unless we specify them as particular molecular mechanisms. But we tried our and listed potential future experiments in the revised version of the paper. Specifically, we write:

      “Second, while we have proposed alternative mechanisms that may result in apparent cooperativity, at present we could not discriminate between these alternatives, in part, because the models lacked specifics – e.g., if virions interacting with a cell reduce its resistance to infection, what does it mean exactly [12]? If virions in a collection augment their infectivity (which may be expected for segmented viruses), how does that viral compensation actually work? Designing experiments that would discriminate between these alternatives would require focusing on a specific mechanism. For example, it may be that that the initiation of gene expression is difficult but is more efficient when there are more virions bringing in more tegument transactivators like pp72/ppUL35 [59]. Alternatively, it may be that there is a bona fide resistance mechanism at play here (e.g. “interferon”) that is antagonized by a viral tegument protein (like TRS1/IRS1 that acts against PKR and 2’5’OAS) [60]. Accrued damage model is also consistent with the idea that at higher genome/cell values, the inoculum itself (including cell and/or virion debris) may impact overall susceptibility of all cells in the well, for example, making them more susceptible to infection. It may be expected, though, that exposing cells to debris would increase cell resistance to infection; this would result in n < 1 that we did not observe at small genomes/cell values. Addressing these hypotheses is an area of future research that will require funding.”

      (3) Decline at high genomes/cell: Several datasets show a downturn at high input. Hypotheses should be provided (cytotoxicity, receptor depletion, and measurement ceiling) and any supportive controls.

      Another good point. We do not have a good explanation, but we do not believe this is because of saturation of available target cells. It seemed to only happen (or was most pronounced) with the ME stocks, which are typically lower in titer and so the higher MOI were nearly undiluted stock. It may be the effect of the conditioned medium. Or perhaps there are non-infectious particles like dense bodies (enveloped particles that lack a capsid and genome) and non-infectious, enveloped particles (NIEPs) that compete for receptors or otherwise damage cells and these don’t get diluted out at the higher doses. We included the point about cell death in Discussion of the revised version of the paper. Specifically, we write:

      “We also do not have a clear explanation of why infection frequency declines at high genomes/cell values for some strain-cell combinations (e.g., Figure 2A, C, D, I, J). Because we measured cell infection in live cells, increase in cell death at higher genomes/cell values may result in the decrease in the number of viable cells.”

      (4) Include experimental data: In Figure 6, please include the experimentally measured titers (IU/mL), if available.

      This is a model-simulated scenario, and as such, there is no measured titers.

      (5) MOI guidance: The practical guidance is important; please add a short "best-practice box" (how to determine titer at multiple genomes/cell and cell densities; when single-hit assumptions fail) for end-users.

      Good suggestion. We now include best-practice box using guidelines developed in Ryckman lab over the years in the revised version of the paper. This is how it reads:

      “Match viral titration methods to the experiment as far as possible. This includes using the same dilution of the viral stock, the cell type, duration of inoculation, and readout of infection.

      When possible, determine the degree of apparent cooperativity (“n”-value, eqn. (1)) for each virus strain/cell type pair being studied.

      If n= 1 (no cooperativity), it is reasonable to calculate experimental MOI based on stock infectivity value determined from a convenient stock dilution.

      If n > 1 or unknown, then stock infectivity should be determined at a dilution resulting in an MOI as close as possible to the desired experimental MOI. Alternatively, the inoculum size can be empirically determined to yield the desired number of infected cells. In these ways different virus/cell type pairs can be compared more fairly.

      Box 1: Recommendations on titrating viral stocks and on performing experiments when comparing different viral strains.”

      Reviewer #3 (Recommendations for the authors):

      FROM PUBLIC REVIEWS (2) Discriminating mechanisms: At present, simulations cannot distinguish between accrued damage and viral compensation. The authors should propose or add a decisive experiment (e.g., dual-color coinfection to quantify true coinfection rates versus "priming" without coinfection; timed sequential inocula) and outline expected signatures for each mechanism.

      This is a good point but to propose a good experiment we need to narrow down the “generic” mechanism to specific processes/genes. We put forward some ideas but clearly more work is needed here:

      “Second, while we have proposed alternative mechanisms that may result in apparent cooperativity, at present we could not discriminate between these alternatives, in part, because the models lacked specifics – e.g., if virions interacting with a cell reduce its resistance to infection, what does it mean exactly [12]? If virions in a collection augment their infectivity (which may be expected for segmented viruses), how does that viral compensation actually work? Designing experiments that would discriminate between these alternatives would require focusing on a specific mechanism. For example, it may be that that the initiation of gene expression is just difficult but is more efficient when there are more virions bringing in more tegument transactivators like pp72/ppUL35 [59]. Alternatively, it may be that there is a bona fide resistance mechanism at play here (e.g. “interferon”) that is antagonized by a viral tegument protein (like TRS1/IRS1 that acts against PKR and 2’5’OAS) [60]. Accrued damage model is also consistent with the idea that at higher genome/cell, the inoculum itself (including cell and/or virion debris) may impact overall susceptibility of all cells in culture, for example, making them more susceptible to infection. It may be expected, though, that exposing cells to debris would increase cell resistance to infection; this would result in n < 1 that we did not observe at small genomes/cell values. Addressing these hypotheses is an area of future research that will require funding.”

      (1) Methods transparency: Include raw spreadsheets or tables of dilution factors and per-well genome estimates used for Figure 1A; this will help reproducibility of the df = 1.3-1.5 pipeline.

      Provided as supplemental xlsx file.

      (2) Epithelial vs fibroblast contrast: Since n is lower on epithelial cells, expand on cell-intrinsic barriers that could dampen apparent cooperativity, and if this argues against simple clumping.

      Indeed, this is our point that we raised in Discussion. Since ECs show lower n than fibroblasts, this observation argues against clumps. Going forward the contrast between cell types will be an approach to understand mechanism. One difference is entry pathways, the ECs involve endocytosis and endosome acidification whereas the fibroblasts do not. There are clearly different receptors involved also, although they are not clearly characterized. One recent report that might be relevant is Ohman 2024 PNAS that shows the gH/gL/UL128-131 complex (aka, "pentamer") is not just dispensable for entry into fibroblasts, but inhibitory. They suggest that the pentamer might bind to a receptor on fibroblasts that activates a pathways that acts against viral IE expression, It could be that in this situation, more virions are really helpful to overcome that block, whatever it is. We now update this point in Discussion.

      (3) Visualization: In Figure 2, consider showing confidence bands for the fitted slope (n) within the colored fit window and reporting n {plus minus} SE in the panels.

      Because we used custom scripts to fit models to data, showing bands of model predictions was a bit complex and would interfere with data points. But we now show 95% Cis for the estimated value n (that are listed in Suppl. Tab S2).

      (4) Symbols: Define all symbols (e.g., V₀, n) on first use in the main text, not only in Methods.

      Done.

      (5) Plot axes check: Explain non-uniform axis labeling ("genomes/cell," "infections/cell").

      This comment was unclear – which labels were not “uniform”? Genomes/cell indicate the expected number of genomes (or virions) that a cell is on average exposed to, infections/cell indicates the probability that a cell actually gets infected.

      (6) Confidence interval for estimated parameters: Figure 3 A-C, please report estimated parameter intervals.

      These are listed in Suppl. Tab S2. Putting Cis for all estimates would clutter the figure making it hard to tell which CIs are for which estimate. But we put the Cis for estimated parameter n in Figure 2.

    1. Author response:

      Reviewer #1 (Public Review):

      This study by Charendoff et al provides interesting observations related to global histone hypermethylation in host cells, during Chlamydia trachomatis infections. The core observation they report is that the host histones are highly hypermethylated during infection, and this appears to be an amplifying effect due to continuous inhibition of demethylases, in part due to a metabolic shift in the host where succinate amounts (which inhibit demethylases) increases. The authors claim specifically due to the bacteria, since antibiotic treatment prevents histone hypermethylation (but leaves you wondering about cause/consequence correlations).

      The core observation of hyper methylation is very interesting, and well documented. There are a number of points to consider though in order to fully substantiate the findings, and close out loose ends. My comments are broad - and built around the interpretations (vs the data presented).

      (1) Related to observations coming Fig 1C etc, and connecting to Fig 3 - the hyper methylation appears to be across different protein arg/lys residues - and is not histone specific. So, is it just a consequence of high SAM pools and flux in infected cells? i.e. the bacterial infection increases SAM pools in cells, and provides an increase in substrate pools for the methyltransferases, leading to protein hyper methylation. The approach used here only measures steady-state SAM amounts (and not SAM flux or utilisation).

      For example, reduced SAM amounts in nuclei could be due to increased utilisation of SAM. The experiments done with the demethylase does not actually answer this question - if you decrease demethylase activity, you will get an increase in net methylation. The authors see an increase in net methylation in the infected cells - this would suggest that in addition (or perhaps primarily) to reduced demethylase activity, there could be much higher SAM utilisation/flux. Again, the over expression of JMJ proteins does not resolve this problem.

      This is an important point. Indeed, one limitation of the initial version of the paper was that we had measured SAM concentration only at one time point (40 hpi) and on the whole population. During revision we used a ratiometric sensor to measure SAM concentration in cells (PMID 34937909). We observed cell-to-cell heterogeneity in SAM levels in HeLa cells, as previously reported in other cell lines. Chlamydia inclusions develop asynchronously, which allows to observe, 40 hpi, a continuum of early (low bacterial load) to late (high bacterial load) stages of infection. We observed no correlation between bacterial load and SAM level, and SAM levels were globally similar when comparing infected and non-infected cells. This experiment strongly supports the hypothesis that protein hypermethylation is not due to an increase in SAM during infection. The data were added in the New Fig. 3. Note that the former Fig. 3 is now split into New Fig. 3 and New Fig. 4.

      (2) Adding to this - what happens to SAM pools in the cells treated with the inhibitors? This actually may not look like the slightly reduced SAM pool observed in infected cell nuclei. Also, what is the SAM/SAH ratio (a very useful indicator of methylation activity).

      Based on the high cell-to-cell heterogeneity of SAM levels observed with the ratiometric probe, we reasoned that measuring SAM/SAH ratio without single cell resolution would not bring crucial information. Also, the discrepancy between data displayed in new Fig. 3A (nuclear extracts) and 3C (live cell imaging) indicate that SAM might be less stable in cellular extracts from infected cells compared to non-infected ones, which would complicate the interpretation of the data. Therefore, we did not implement LC-MS/MS on nuclear extracts to measure SAM/SAH ratio.  

      (3) There is a correlation/implication issue here in Fig 2 - cells with C. trachoma's infection show hyper methylation. But these are the only cells with high C. trachomatis. So it is a bit ingenious to say that histone hyper methylation correlates with bacterial proliferation. The cells without bacteria don't have hyper methylation - and that does not have anything to do with the bacterial proliferation.

      In Fig. 2B, we compared the methylation signal within the population of infected cells only (excluding the uninfected cells). We edited the text to clarify this point. “We observed that, within the population of infected cells, the sum intensity of the mCherry signal was higher in cells that displayed hypermethylation of H3K9me3 than in cells with low level of H3K9me3, indicating that histone hypermethylation correlated with bacterial load (Fig. 2B).”

      (4) The claim that demethylase activity is down in infected cells again comes primarily from the increased succinate (2-fold) amounts in infected nuclei - and then correlated with experiments where succinate, (permeable) a-KG are supplemented in excess. While I personally like the hypothesis that the hypermethylation might be a result of an imbalance in cofactors (succinate vs a-KG) in infected cells, the data presented is very premature to make that conclusion. Again, steady state measurements of only succinate cannot provide a clear answer to that question. For example, is there a clear allocation/flux difference (between a-KG, and leading out to glutamate/glutamine, vs flux through the TCA and increased succinate accumulation? Is there a bottleneck/build-up of succinate in cells that might lead to the increase in nuclei? This also opens another direction of possible regulation - increased histone succinylation. When you see a large increase in succinate in the nucleus, before looking at demethylase activity - it becomes obvious if succinate itself increases histone succinylation (through HATs).

      Our work confirms the accumulation of succinate in cells infected by C. trachomatis, previously reported in Rother et al 2018. The reason for this accumulation remains to be investigated in detail. We have previously shown that OxPhos is relatively stable in infected cells (PMID 35931114), indicating that the flux through the TCA of the eukaryotic host proceeds normally. As mentioned in our discussion, the TCA of the bacteria is disrupted with several enzymes missing, although not in the step immediately downstream of succinate/fumarate production. Still, synthesis of succinate and fumarate (fumarate accumulation was observed in the Rother 2018 study) by bacterial enzymes might contribute to their accumulation in infected cells. The approach we chose to measure methylation at the proteome level is not suitable to look for histone succinylation, because of the diversity of post translational modifications on histones, which occur in combinations. However, following on this reviewer’s comment, we reanalysed the proteomic data to compare protein succinylation levels in infected and non-infected samples. We detected 41 succinylated peptides in the infected samples, against 23 in the uninfected samples. For many of these, we did not have quantitative data in all condition and only one protein, transportin 1 (TNPO1), reached statistical significance, with a 4-fold increase in succinylation in infected samples. Thus, while essentially qualitative, this analysis fully supports the hypothesis that succinate accumulates in infected cells. These data were added to Table S1 and to the result section.

      (5) What might the authors hypothesise about why this hyper methylation happens? It appears in some ways that hyper methylation happens - potentially due to a metabolic bottleneck that the bacteria triggers (and there is a build-up of SAM and/or succinate, and altered flux out of a-kg). The methylation is just a visible outcome - but may not be central to pathogenesis or viability.

      We discussed this question in the penultimate paragraph of the discussion by giving some elements of answer to the question: “Does it benefit the host or the bacteria? ». In our study, we showed that protein hypermethylation affected the transcriptional response of the host. We did not investigate whether the activity of some of the host proteins engaged in the response to infection were affected. It might be the case, considering that methylation is a common PTM regulating protein’s activity. Still, we agree with this reviewer that hypermethylation might not be central to pathogenesis or viability. Addressing this question would require a complex model in which protein methylation levels could be controlled experimentally.  

      Reviewer #2 (Public Review):

      Strengths:

      (1) Because the study compares genuinely infected cells with uninfected cells within the same infected cell population, it enables a clearer and more rigorous comparison.

      (2) By using multiple Chlamydia species and cells from multiple host species (human and mouse), and obtaining consistent findings across these systems, the study demonstrates the generality of bacterium-induced epigenomic alterations.

      (3) The study shows that the epigenomic changes are caused by reduced activity of JMJC domain-containing lysine demethylases, demonstrating through multiple complementary approaches-including the use of a demethylase inhibitor, overexpression of target-specific demethylases, and analysis from the perspective of cofactors required for JMJC domain-containing demethylases-that decreased lysine demethylase activity constitutes the molecular mechanism underlying the increased H3 methylation levels induced by Chlamydia infection.

      (4) By performing ChIP-seq analyses of H3K4me3 and H3K9me3, the study clearly delineates, on a genome-wide scale, how infection leads to increased levels of these epigenomic marks.

      Weakness:

      (1) Reduction of cofactors such as Fe2+ or a-KG decreases the activity of JMJC-domaincontaining lysine demethylases (thereby directly affecting histone H3 lysine methylation). However, these cofactors are also involved in the activities of other epigenetic regulators, such as TET enzymes that contribute to DNA demethylation and SIRT family proteins that mediate histone deacetylation. Therefore, it cannot be excluded that modulation of these factors indirectly leads to the changes in H3 lysine methylation dynamics targeted in this study.

      Indeed, reduction of the concentration of Fe2+ and aKG is expected to have other consequences in addition to the inhibition of JMJC-domain containing lysine demethylases on which we focus in this study. As a matter of fact, we reported a decrease in the methylation level of host DNA in infected cells, and we brought some elements that might explain the discrepancy between DNA and histone methylation status in the discussion (e.g., infected cells display enhanced expression of GADD45, which recruit TET enzymes and thus facilitate DNA demethylation). This example illustrates the complexity of host/pathogen interplay, which affect many parameters simultaneously. Indeed, we cannot rule out that modulation of enzymatic activities other than JMJC-domain containing lysine demethylase contribute significantly to the hypermethylation phenotype.

      (2) Related to point 1, although overexpression of JMJC-type demethylases has been shown to reduce the Chlamydia infection-induced increase in H3 lysine methylation, it is well known that over production of these enzymes, while target-specific, also leads to a genome-wide reduction of lysine methylation. Thus, a decrease in lysine methylation upon expression of these demethylases does not necessarily demonstrate that the infection-induced increase in H3 lysine methylation is caused by impaired JMJC-type demethylase activity.

      We fully agree. We included this experiment to show that increasing the expression of one demethylase only restored demethylation of its cognate target. This support the hypothesis that if the hypermethylation is due to poor demethylase activity, it is likely that several demethylases show impaired activity (as opposed to a scenario in which failure of activity of a single demethylase would indirectly affect all other methylation marks).  

      Reviewer #3 (Public Review):

      In this manuscript, the authors explore a molecular basis for hypermethylation of histones in epithelial cells infected with the obligate intracellular bacterial pathogen Chlamydia trachomatis. This is of particular interest given that Chlamydia is known to drastically alter host cell gene transcription, and histone hypermethylation would suggest a new way by which Chlamydia interferes with gene expression of its host. Histone methylation was previously implicated in the introduction of dsDNA breaks in infected cells, and the chlamydial effector NUE was reported to methylate histones, but the role of this modification in dictating host cell gene transcription has been unexplored. The authors use a suite of tools to approach this question, including various -omics techniques, genetic approaches, and biochemical assays. Overall, the manuscript provides many interesting pieces of data, though some of them are difficult to reconcile, which may reflect methodological hurdles that are not fully addressed in the current version of the manuscript. My major concerns regard the rationale/interpretation for various mechanistic experiments and that the heterogeneity of the histone hypermethylation phenotype is not addressed which I believe may explain some apparent inconsistencies in the results.

      We thank this reviewer for insightful comments. We address these two major concerns during revision and bring some elements in our responses below.

      Using an immunofluorescent approach, the authors show that a subpopulation of the nuclei in Chlamydia-infected cells (~10-20%) exhibit high amounts of methylated histone species. This occurs during the late stages of infection, near the time when Chlamydia would lyse the host cell and positively correlates with bacterial burden.

      Accordingly, halting chlamydial growth blocks the onset of histone hypermethylation. Exogenously supplying cofactors for histone demethylases, the low activity of which is implicated in the histone hypermethylation phenotype, reduces histone hypermethylation. In general, these data are compelling and raise interesting questions about the role of histone methylation in governing chlamydial egress from infected cells. Interestingly, these behaviors seem to arise independently of NUE, the secreted chlamydial histone methyltransferase, supporting the notion that a metabolic reprogramming may underlie the hypermethylation phenomenon.

      As noted above, the authors propose that hypermethylation arises due to decreased demethylase activity in infected cells. However, the data do not conclusively support this interpretation. For example, the approaches used to probe demethylase activity rely on (i) a direct biochemical measure of demethylase activity, (ii), pharmacological inhibition of demethylase, and (iii) heterologous expression of a specific demethylase. With the exception of (i), these approaches would be expected to alter histone methylation regardless of the source. That is, inhibition of demethylases should increase histone methylation regardless of whether the source of methylation is increased methylase or decreased demethylase activity. Similarly, overexpression of a demethylase would be expected to reduce cognate histone methylation arising either from increased methylase or decreased demethylase activity.

      We agree with the reviewer’s comments. The experiment using pharmacological inhibitors (ii) show that infected cells are sensitized to these inhibitors but doesn’t provide direct mechanistic insight. The experiment using heterologous expression of demethylases (iii) was included to show that increasing the expression of one demethylase only restored demethylation of its cognate target. This supports the hypothesis that several demethylases show impaired activity (as opposed to a scenario in which failure of activity of a single demethylase would indirectly affect all other methylation marks).  

      The most direct evidence for impaired demethylase activity come from the direct measure of demethylation of H3K4me3 in nuclear extract (i). It is strengthened by indirect evidence that metabolite concentrations hinder demethylase activities late in infection: 1/ iron and DMKG supply diminish hypermethylation of histone lysine residues 2/ succinate levels (a competitor of aKG) are two-fold higher in nuclei isolated from infected cells. This latter finding was confirmed during revision as we identified more succinylated proteins in infected samples compared to non-infected ones.

      We also considered the possibility that infected cells displayed increased histone methyl transferase (HMT) activity. This would be compatible with decrease KDM activity and could contribute to the histone hypermethylation. Unfortunately, this hypothesis cannot be tested directly (as we did for the measure of H3K4me3 demethylation activity). Indeed, SAM is notoriously labile and in vitro assays to measure HMT require to add exogenous SAM to cell extracts to detect any HMT activity, which would not allow us to test activity based on endogenous SAM levels.

      Instead, we used a ratiometric sensor to measure SAM concentration in cells (PMID 34937909). Chlamydia inclusions develop asynchronously, which allows to observe, 40 hpi, a continuum of early (low bacterial load) to late (high bacterial load) stages of infection. There was no correlation between bacterial load and SAM level, and this level was globally similar when comparing infected and non-infected cells. This experiment supports our hypothesis that protein hypermethylation is not due to an increase in SAM during infection.

      This experiment was also very interesting because it revealed a high cell-to-cell heterogeneity in SAM levels in HeLa cells. Thus, in some cells, SAM might be limiting, which could explain why only a fraction of cells display histone hypermethylation.

      Still, we cannot fully rule out the possibility that increase in SAM availability late in the infectious cycle in some cells, and is immediately consumed through protein methylation, resulting in no net [SAM] increase. The discussion was expanded to take these comments into consideration.

      Altogether, we think that the evidence of decrease KDM activities in infected cells late in infection are strong. Our data do not rule out the possibility that additional mechanisms may contribute.

      Moreover, the authors report that the effect of the demethylase inhibitor on histone hypermethylation is significantly potentiated by infection, suggesting that infected cells have greater methylase activity than uninfected cells, because the latter barely respond to the presence of demethylase inhibitor. In other words, a dramatic increase in histone methylation in the presence of demethylase inhibitor is most parsimoniously explained by increased methylation (no longer being removed by demethylase), not decreased demethylation (which would be analogous to treatment with demethylase inhibitor). The authors do not directly assay methylase activity. These concerns extend to the rationale used to justify experiments with infected mice, which the authors treat with the demethylase inhibitor.

      The observation that the same concentration of JIB-04 leads to an increase of histone methylation in infected cells and not in non-infected cells, is coherent with the data showing that aKG or iron supply diminish histone hypermethylation in infected cells. Indeed, the inhibitor is taken up similarly by infected and uninfected cells but the potency of the inhibitor will depend partly on levels of iron, aKG and succinate found in the cellular milieu so same concentration of inhibitor may inhibit demethylase activity in cells with higher succinate and/or low aKG and low iron but fail to inhibit demethylase activity in cells with higher iron or aKG or lower succinate. In other words, high iron, high aKG or low succinate will “buffer” JIB-04 and make it less potent since JIB-04 partly acts by competing with the iron (competitively) and the aKG (mixed competitive inhibition) PMID 23792809. The same phenomenon is expected for SD70 and TACH101 that share aspects of the mode of action of JIB-04 regarding partly competing for aKG and/or iron in the catalytic site.

      The authors perform experiments to characterize the consequence of hypermethylation genome-wide. Because the authors do not enrich for those cells which exhibit histone hypermethylation, the results reflect the mixed population, and therefore presumably dilute out important signal related to the phenomena under investigation. For example, the proteomic analysis of post-translational modifications identifies only one methylated histone species, whereas the immunofluorescent approach shows consistent effects across five different methylated histone species. Moreover, the chromatin immunoprecipitation analysis indicates that there is unexpectedly a lower density of methylated histones at regions which are also enriched in uninfected cells. The authors argue that this suggests increased methylation is happening "outside" of these histone-dense regions, but direct evidence in support of this claim is lacking.

      The caveat of bulk analyses as opposed to single cell resolution is indeed important to consider when analysing the chIP-seq data and we emphasized this point in the revised manuscript. We could have sorted the cells with high bacterial burden; this would probably have given stronger differences between the two samples. Still, the change in distribution of H3K4me3 in infected samples was very clear and statistically significant. A change in H3K9me3 distribution would be more difficult to catch, as the mark is more widespread.

      In sum, this paper provides compelling evidence in support of the notion that histones are hypermethylated at various residues late in chlamydial infection, that this process is modulated by known cofactors of demethylases, and is the result of high levels of bacterial replication in the cell. That histone hypermethylation governs host gene transcription during chlamydial infection suggests a relatively novel mechanism by which Chlamydia subverts the host cell to establish a replicative niche or egress to infect a new cell. The information obtained regarding the methylation status of host proteins and host gene transcription controlled by a metabolic cofactor during infection will be a useful resource for other researchers. However, in the current version of the manuscript, the mechanistic basis for these behaviors is relatively unclear.

      We thank this reviewer for constructive feedback. We believe that the mechanistic conclusions of our report have been strengthened during revision with additional experiments and text clarification.

    1. Note: This response was posted by the corresponding author to Review Commons. The content has not been altered except for formatting.

      Learn more at Review Commons


      Reply to the reviewers

      General Statements

      Thank you for providing an assessment of our manuscript. Below, we outline our revision plan. The revisions address four main areas: the relationship between the identified molecular signatures and fibrosis severity or disease etiology; the criteria used to identify disease-associated fibroblasts; the interpretation of the genes and biological processes highlighted by our analyses; and the broader biological insights supported by the study.

      As part of the revisions implemented, we have:

      Associated organ-specific fibrotic molecular signatures and fibrosis severity scores available in the clinical metadata, helping to relate the identified transcriptional patterns to biologically meaningful aspects of fibrosis. Extended supplementary figures that more clearly present the decision-making process used to identify fibroblast subpopulations associated with fibrosis. Revised the methods, figures, legends, and captions in response to the reviewers' suggestions to improve clarity. Expanded the discussion of the results by incorporating the literature suggested by the reviewers, thereby providing additional context for the identified fibrotic signatures. Extended our spatial analysis using a more robust identification of fibrotic regions.

      We plan to:

      Extend our cell-cell communication and spatial analysis using deconvolution methods Provide comparisons between our unsupervised multicellular factor analysis of multiple studies with our supervised fibrotic signatures to ensure coherence between analyses. Perform additional comparisons between specific pairs of organs and additional cell types, instead of focusing solely on the comparison of all organs simultaneously. Expand the results and discussion to clarify the relevance and limitations of our study. We believe these revisions will strengthen our resource manuscript and will help us to provide a robust and reliable description of fibrotic processes across organs.

      Description of the planned revisions

      Reviewer #1

      Reviewer #1, major comment 1: The group has been developing cutting edge bioinformatic tools for the community. The authors also provided scripts and the processed data for reproducibility. I have no doubt in their implementation of the methodology. I also understand the reasons of the objective tone throughout the manuscript. However, the authors made very little claims with biological significance. The conclusion of the study is vague with almost nothing mentioned in the abstract. What are the cross-organ effects in fibrosis identified in this study? I believe some additional claims would facilitate the reader with less technical knowledge to grasp the study better.

      We understand the concern of the reviewer regarding the lack of an explicit discussion of the biological significance in the abstract and other parts of the manuscript, as most of the manuscript is focused on the comparison of studies at different levels. Our study defines which fibrosis-associated transcriptional patterns are reproducibly detectable across the currently available public single-cell datasets, while also identifying where cross-organ interpretation remains limited. We observed that some disease-associated transcriptional patterns recur across organs and studies, particularly in mesenchymal and endothelial compartments. In contrast, other compartments, including myeloid cells, showed weaker cross-organ agreement, which may reflect either greater tissue-context dependence or stronger sensitivity to differences in disease stage, sampling, and annotation. Finally, we observed a convergence of fibrotic signals in a subset of mesenchymal cells and show which genes are specifically expressed in actively scarring regions across organs, with TIMP1 being consistently identified as highly expressed in fibrotic regions by disease associated fibroblasts across tissues and modalities.

      Our results should be interpreted as robust and reproducible cross-dataset fibrosis signatures rather than definitive evidence for a specific pathophysiological mechanism. Therefore, we believe that the primary contribution of this study lies not in assigning causal roles to individual genes or pathways, but in providing a systematic framework for identifying fibrosis-associated programs that are reproducibly observed across studies, organs, and disease etiologies. As our analysis is entirely computational, we intentionally avoid making strong mechanistic claims without experimental validation. Instead, we envision this resource as a means to prioritize candidates and generate hypotheses for future functional studies.

      To address the reviewer's concern, we will make more explicit claims of our observations within our abstract and throughout the text to make our intentions and conclusions clearer. We will be more explicit about what information we are providing with our resource and how it can best be leveraged. We will further include our conclusions about cross-organ agreements described above, as well as specific observations from our analyses that help the reader to get a better grasp of the study.

      Reviewer #1, major comment 3: The authors performed multicellular factor modeling in each organ and identified factors that are distinct in fibrotic and reference tissue in Fig. 2B, e.g., factors 1 and 2 in heart. Are these factors driven by specific biological pathways? Could these factors also be used to identify common biological functions in fibrotic tissue across organs?

      We agree that, in principle, the latent factors identified by the multicellular factor models could be interrogated for their biological interpretation. Each factor is associated with a gene-weight vector per cell type, which can be analyzed similarly to a differential expression signature to identify enriched pathways and biological processes.

      However, we chose not to pursue a systematic factor-level interpretation for three reasons. First, as shown in Suppl. Figure 3, the contribution of individual factors to the separation between fibrotic and reference samples varies substantially across organs. In some organs, the distinction is largely captured by a single factor, whereas in others it is distributed across multiple factors. Second, because the models were trained independently for each organ, there is no direct correspondence between factor identities across organs, making cross-organ comparisons of individual factors difficult to interpret. Finally, we were not able to capture a fibrosis-related transcriptomic program from all organs.

      We therefore used the multicellular factor analysis primarily as an unsupervised approach to assess whether common fibrosis-associated variation could be detected across datasets. The observation that fibrotic and reference samples consistently separated along latent factors suggested the presence of shared disease-associated signals. For the subsequent biological interpretation, however, we opted for a supervised analysis framework based on differential expression and downstream functional enrichment, which allowed more direct and robust comparisons across organs and disease contexts.

      We will revise the manuscript to make this rationale more transparent to the reader. In addition, we will include an analysis demonstrating that the gene weights associated with the disease-relevant latent factors closely resemble the corresponding organ effect sizes in heart, kidney, and lung, illustrating that biological interpretation at the factor level yields conclusions that are highly consistent with those obtained from the supervised differential expression analysis. This further supports our decision to base the downstream functional analyses on the organ effect sizes, which provide a more straightforward framework for cross-organ comparison.

      Reviewer #1, major comment 4: Although strong organ-specific effects, the author detected similar transcriptional changes in endothelial and mesenchymal cells in heart and lung at Fig. 3B. The analysis on disease-associated fibroblasts also showed much higher overlapped between heart and lung compared to, e.g., liver and kidney in Fig. 4C. Are there additional shared fibrosis features or functions in mesenchymal cells or disease-associated fibroblasts in heart and lung?

      Reviewer #1, major comment 5: There seems to be certain degree of similarities among the epithelial cells in kidney and lung in Fig. 4B.

      Shared response for comments 4 and 5:

      Given the high number of combinations of comparisons, we decided to focus on the most shared signals (mesenchymal and endothelial) in our manuscript. However, as the reviewer notes, there are other comparisons, such as the one between epithelial cells from kidney and lung, or in endothelial cells between heart and lung, that may be important to report. We plan to revise the text in section "Fibrotic disease programs within tissues" to explicitly discuss the observed similarity and plan to additionally show the shared genes driving these similarities in a supplementary Figure in the manuscript.

      Reviewer #1, major comment 8: TNC appears in the lower bottom of the list in Fig. 6C. It is unclear why TNC was chosen as a board therapeutic target in the end.

      We agree that the original wording may have implied that TNC was selected because it was the top-ranked candidate in Figure 6C. This was not our intention. Rather, we chose TNC as an illustrative example because it emerged from our analysis without prior manual prioritization, has already been linked to fibrosis in specific disease contexts, and has been explored experimentally as a therapeutic target. At the same time, its role has not been investigated broadly across fibrotic diseases, making it a useful example of how the presented framework can identify candidates that may have relevance beyond the settings in which they were originally studied.

      We will revise the text to clarify that TNC is presented as one representative example from the set of prioritized candidates rather than as the single most highly ranked therapeutic target.

      Reviewer #1, minor comment 1: Is there additional measure that account for the datasets with lower RNA counts shown in Fig. S1?

      We thank the reviewer for highlighting this potential source of technical variation. We did not apply an additional correction specifically to account for datasets with lower RNA counts. Instead, to minimize the impact of differences in sequencing depth and cell-level sparsity across datasets, the majority of our analyses were performed on pseudobulk profiles rather than individual cells. Pseudobulk aggregation substantially reduces the influence of variation in RNA counts between cells and datasets, providing more robust estimates of gene expression. We therefore believe that differences in RNA counts had a limited impact on the main conclusions of the study. To illustrate this point, we plan on showing additional quality control summary plots for our pseudobulked data.

      Reviewer #2

      Reviewer #2, major comment 3: Fig. 4B-C: the full list of organ-specific and overlapping genes should be given in a supplemental table.

      We thank the reviewer for this suggestion. We agree that providing the complete lists of organ-specific and overlapping genes improves the transparency and utility of the analysis. We will provide the full gene lists underlying Figures 4B-C as supplementary tables in the revised manuscript. These tables will provide the complete set of genes used for the reported overlap analyses and allow readers to further explore the identified organ-specific and shared fibrotic programs.

      Reviewer #2, major comment 6: Cell-cell communications analysis: It would be informative to add a circosplot highlighting the best cell-cell communication candidates in each organ. The authors should also provide the full list of predicted interactions in a supplementary table, including scores for each organ for each interaction. Additionally, it would be important to focus specifically on ligand-receptor pairs associated with growth factors and cytokines. While incorporating Visium data is very interesting and challenging, it may reduce sensitivity due to its relatively poor capture efficiency. This could particularly overemphasize the importance of collagens and other ECM-related factors, which are highly expressed.

      We agree that additional visualization and data availability would improve the presentation of the cell-cell communication analysis. Therefore, we will add additional organ-specific visualizations highlighting the highest-confidence cell-cell communication candidates within each organ, providing a more intuitive overview of the predicted interactions. Second, we plan to include the complete list of predicted ligand-receptor interactions as supplementary tables, including the corresponding scores for each organ and gene annotations (i.e. cytokine, growth factor, etc.), allowing readers to explore the full set of predictions underlying the analyses.

      We also agree that highly expressed extracellular matrix components, such as collagens and proteoglycans, can dominate CCC analyses, especially when investigating fibrotic diseases. Indeed, this consideration motivated our final therapeutic target prioritization strategy (Figure 6). In this analysis, we specifically excluded collagens and proteoglycans, thereby enriching for extracellular signaling molecules that are more likely to represent biologically informative and therapeutically actionable cell-cell communication events. We will modify the results section to clarify our rationale for this analysis.

      Reviewer #2, major comment 8: Visium Dataset Analysis: It would be interesting to compare fibrotic areas across different organs by performing niche or topic analyses using supervised deconvolution approaches (such as RCTD). This would allow for a better estimation of cell composition and functional annotations of fibrotic and inflammatory areas.

      We agree that a cell type deconvolution would provide an informative framework for characterizing the cellular composition of fibrotic niches and its association with the fibrotic signatures we derived from single-cell data. We plan to address the reviewer's suggestion by running a cell type deconvolution analysis of the Visium datasets to estimate the enrichment of major cell populations within scar regions and compare them across organs. We hope that these additional analyses will provide complementary information on the cellular composition of these areas.

      Reviewer #2, minor comment 1: p11: the authors conclude that "cell proportions differed not only between patients and organs, but also that there was no uniform abundance change in disease". This result may reflect technical variability, particularly due to dissociation biases from very different organs or the use of different platforms. This limitation should be discussed.

      We agree that differences in cell type proportions may not only reflect biological variation but can also be influenced by technical factors, including organ-specific dissociation biases, differences in tissue processing, and the use of distinct sequencing platforms. We will expand the text to explicitly acknowledge these potential confounding factors and to emphasize that the observed differences in cell abundances should be interpreted with appropriate caution.

      Reviewer #2, minor comment 3: Panel E in Fig. 5 is difficult to read and needs to be improved.

      To improve the readability of the figure, will include fewer ligand-receptor pairs and additionally add grey boxes in the background to help the reader to better distinguish the ligand-receptor pairs from each other.

      Reviewer #3

      Reviewer #3, minor comment 1: P5: Some context regarding expected differences between single cell and single nuclei datasets here would be good (especially if some differences are potentially important).

      We agree that adding context regarding the expected differences between single cell and single nuclei datasets would add value to the manuscript. These differences have been investigated in the past and were shown to have an impact on the RNA-sequencing results and their interpretations (Van Melkebeke et al. 2024; Lake et al. 2023; Feng et al. 2026; Denisenko et al. 2020; Litviňuková et al. 2020; Koenitzer et al. 2020). We therefore plan to include more background information, including the distinct capture biases and transcriptomic characteristics, to highlight that these differences should be considered when comparing datasets generated using different protocols.

      Reviewer #3, minor comment 6: *P12: Please clarify whether the multicellular factor model is fit jointly across all datasets within an organ, or separately per dataset followed by comparison. If fit jointly, how are batch/study effects handled? If fit separately, how are factors aligned across invocations? *

      Is it possible to say how much of this consistency across datasets is due to non-fibrotic or non-disease state regulation? Are the disease-associated factors driven by coordinated changes across multiple cell types, or primarily by one dominant cell type? And if the latter, is this related to expression magnitude, or cell type abundance?

      We agree that the description of the multicellular factor model in the original manuscript did not provide sufficient methodological detail.

      The multicellular factor model was fitted jointly across all datasets within each organ, resulting in one model per organ (four models in total). Following the strategy proposed in the MOFA+ framework (Argelaguet et al. 2020), individual studies were treated as groups within the model, allowing the integration of multiple datasets while accounting for study-specific effects. Because the model uses cell type-specific pseudobulk profiles as separate views, the inferred factors reflect coordinated transcriptional changes across cell types rather than differences in single-cell abundance. Pseudobulk aggregation substantially reduces the influence of cell number variation, and we applied quality control thresholds to ensure that only samples with sufficient counts for each cell type were included.

      To further clarify the relationship between latent factors and fibrosis, we plan to add an additional analysis showing the proportion of variance explained (R²) by each factor across studies and cell types. The R² can be used as a proxy of the importance of a cell-type in defining the latent factor. Whereas many latent factors capture sources of biological or technical variation unrelated to disease, only a subset consistently separates fibrotic from reference samples. These disease-associated factors therefore represent fibrosis-specific variation rather than general transcriptional structure and are the factors we highlighted in the manuscript text to support that different studies had a consistent disease signal.

      We will incorporate these clarifications into the manuscript to make the modeling framework and its interpretation more transparent and add additional analyses showing the variance explained as extra insights into the models.

      Reviewer #3, minor comment 12: *P23: What conclusions should be drawn from the broad cell-type communication comparisons between organs in Fig. 5A? The text reports which broad cell-type pairs account for many upregulated ligand-receptor interactions, but it is not clear whether these comparisons identify fibrosis-specific communication or mainly reflect broad tissue architecture, cell-type abundance, etc. *

      If the broad categories were chosen because finer cell-state annotations are not consistently available across studies, it would be helpful to state this limitation explicitly.

      We agree that the rationale and interpretation of the broad cell-cell communication analysis should be described more clearly in the manuscript.

      The analysis shown in Figure 5A is based on the organ-specific mixed-effects differential expression models and therefore reflects disease-associated changes in ligand and receptor expression between fibrotic and reference samples, rather than absolute expression levels. Therefore, Figure 5A shows which cell type pairs increase their communication in fibrosis, based on the amount of ligand-receptor pairs that are differentially expressed above a threshold. As the mixed-effects models run per cell type separately, it is unlikely that an increase in cell type proportion causes more upregulated communication events to another cell type with this type of analysis. Overall, we do not see a correlation between increase in cell type proportion in the tissue (Figure 2A) and number of upregulated genes with the mixed effect models (Figure 4A). Therefore, we do not think that cell type proportions have a high effect on this particular analysis.

      We also agree that the use of broad cell type categories warrants clarification. These categories were chosen because they can be robustly harmonized across the diverse datasets included in this meta-analysis, whereas finer cell-state annotations are not consistently available or comparable across studies and organs. We plan to revise the manuscript to clarify both the interpretation of Figure 5A and the rationale for using broad cell type categories in this analysis.

      Reviewer #3, minor comment 14: P31: The therapeutic suggestions should come with some discussion that this is association rather than causation, as it's not established that these are causal drivers. MOXD1 seems compelling, especially if this has been observed to have a potential therapeutic effect in other fibrotic diseases, and this is an excellent outcome that justifies the meta-analysis approach. TNC is somewhat more speculative in this regard, so if there is any mechanistic or other motivations, it would be good to include them here.

      We agree that the therapeutic implications of our findings should be interpreted with appropriate caution, as our analyses identify associations rather than causal drivers of fibrosis.

      These candidates were selected based on the combination of our computational prioritization results and the existing literature, rather than a causal role that has been established by our analysis. Our intention was to provide representative examples of how the presented framework can recover biologically plausible candidates with existing experimental support while simultaneously suggesting their potential relevance across a broader range of fibrotic diseases. We plan to revise the discussion to more clearly emphasize that the proposed therapeutic candidates represent hypothesis-generating observations that require experimental validation.

      Reviewer #3, minor comment 16: P31: It would be nice to have what you think the issues are with the lack of patient metadata, and how these issues might manifest in the analyses (this links with the previous comment regarding disease stage).

      The lack of detailed clinical and histological metadata substantially limits the range of biological and clinical questions that can be addressed, thereby reducing the value that can be extracted from the considerable effort and cost associated with large-scale tissue sequencing studies. In the current study, we are mostly restricted to comparing fibrotic and reference samples because information such as disease stage, fibrosis severity, time since diagnosis, medication, treatment history, tissue sampling location, and other clinical covariates is largely unavailable or inconsistently reported across studies. If these metadata were available, they could be explicitly incorporated into the statistical models, allowing analyses that relate transcriptional changes to clinically relevant variables such as fibrosis severity or disease progression rather than simply disease status.

      Furthermore, additional patient metadata would allow potential confounding factors to be accounted for or controlled in the analysis. For example, treatment effects or other clinical characteristics could be modeled directly or specific patient groups could be excluded where appropriate, leading to a clearer separation of disease-associated biology from technical or clinical confounders.

      We will expand the Discussion to more explicitly describe these limitations and their potential impact on the interpretation of our results.

      Description of the revisions that have already been incorporated in the transferred manuscript

      To facilitate review of the revised manuscript, we have grouped our responses into two categories. First, we address comments that resulted in substantial new analyses, figures, or modifications to the interpretation of the results. Second, we address minor and editorial comments, which have already been directly incorporated into the revised manuscript.

      3.1 Comments requiring additional analyses or substantial revisions

      Reviewer #1

      Reviewer #1, major comment 2: The authors have pooled the data from at least five different disease per organ to identify the pan-fibrosis signature across diseases. Some of the diseases, e.g., pneumonitis, ICM, MI, MCD, ALD) may present more acute remodeling compared to the rest, which might exhibit distinct features that mask the analysis. The extent of fibrosis also varies very significantly. A correlation with histological data is required.

      We agree that fibrotic diseases differ substantially with respect to disease etiology, disease stage, extent of remodeling, and the degree of fibrosis present in the tissue. We had highlighted this as a key limitation of the study in the discussion:

      "Second, the limited availability of patient metadata leaves many aspects unresolved, including the exact diagnosis, disease severity, tissue sampling location, and the extent of fibrosis. If these aspects were better documented, they could be accounted for in the analysis and could allow a clearer distinction of physiological from pathophysiological fibrotic processes. Third, we treated all disease etiologies collectively under the term "fibrosis". However, the degree of fibrotic remodeling likely varies between conditions, and the dataset remains imbalanced in terms of sample representation across organs."

      While comprehensive histological and disease severity information was not consistently available across the published datasets included in our meta-analysis, we were able to further investigate this question in the subset of studies for which fibrosis-related metadata were available. Specifically, we derived organ-specific fibrosis signatures, scored these signatures across patients, and performed a per-study normalization. In these datasets, our derived organ fibrosis scores correlated with available fibrosis severity measurements, supporting the biological relevance of the identified programs (Figure S5A-D).

      In addition, these analyses indicate that fibrosis signature scores vary across disease etiologies, consistent with the reviewer's suggestion that different diseases may exhibit distinct degrees of fibrotic remodeling (Figure S5E). However, given that most of the etiologies are covered by a single study, it is not possible to disentangle these results from the type of controls used by each study and technical variability.

      Nevertheless, because detailed histological and clinical metadata are available only for a limited subset of studies, we believe that a comprehensive analysis of fibrosis severity, disease chronicity, and etiology-specific remodeling is not possible with the currently available data. Future studies with more uniformly annotated patient cohorts will be well-positioned to address these questions in greater depth. Our findings should therefore be interpreted as identifying molecular programs consistently associated with fibrotic disease across diverse conditions, rather than as a direct measure of fibrosis severity itself. We have included these observations in the results section "Identification of shared gene programs per tissue":

      "As multiple disease etiologies and disease stages were integrated in each organ, we asked whether the extracted organ-consensus genes were associated with fibrosis severity. However, fibrosis severity measurements were unavailable for the majority of studies, preventing a systematic assessment of severity across the integrated dataset. To nevertheless evaluate whether the identified programs captured biologically meaningful aspects of fibrosis, we derived organ-specific fibrosis signatures, scored these signatures across patients, and performed a per-study normalization. In datasets containing fibrosis severity measurements, our derived fibrosis signature scores correlated with fibrosis severity, supporting the biological relevance of the identified programs (Figure S5A-D). Furthermore, we observed differences in signature scores across disease etiologies (Figure S5E). However, because disease etiologies were unevenly distributed across studies, it remains difficult to distinguish true biological differences from study-specific technical effects. Overall, these results suggest that there is a part of the fibrotic program that appears to be shared within most tissues, primarily found in endothelial, mesenchymal, and epithelial cells. Furthermore, our findings indicate that the identified organ-consensus programs capture biologically meaningful aspects of fibrosis."

      To explain our methodology, we further added this section to our methods:

      "Fibrosis severity scoring

      To associate the organ-consensus gene signature with fibrosis severity, we first extracted an organ-consensus gene set per organ from the organ-specific gene ranking. Specifically, for each cell type and organ, genes were ranked based on the random-effects meta-analysis estimate obtained from differential expression analyses across studies. Only genes detected in at least three studies were considered for downstream analyses. Positively associated genes were required to have a non-negative upper confidence interval bound and were ranked by decreasing effect size, whereas negatively associated genes were required to have a non-positive upper confidence interval bound and were ranked by increasing effect size. The top 200 positively associated genes and the top 100 negatively associated genes were retained for each cell type-organ combination.

      To give each sample a fibrosis score, pseudobulk profiles were generated for each study by aggregating raw counts across all annotated cells per sample, excluding samples with fewer than three annotated cell types. Pseudobulk count matrices were normalized to 10,000 counts per sample, followed by log-transformation. Gene set activities were inferred per sample using decoupler's (124) (v1.9.0) univariate linear model (ULM) with curated organ-consensus gene sets, yielding enrichment scores for each sample.

      Finally, these enrichment scores were normalized per study: For each study, the mean and standard deviation of enrichment scores were calculated for all control samples. Sample-level scores were then centered against the corresponding study-specific control mean and additionally converted to standardized scores by dividing by the control standard deviation."

      Reviewer #1, major comment 7: The graphs in Fig. S6A do not clearly present how the disease-associated fibroblasts are identified. The true identities of disease should also be plotted in these UMAPs. The results indicating these cells expressed myofibroblast signature should also be shown confirming that these cells are not other mesenchymal cells, e.g., pericytes or smooth muscle cells.

      We agree that the original supplementary figures did not sufficiently illustrate how disease-associated fibroblast populations were identified and distinguished from other mesenchymal cell types. To improve transparency, we have substantially expanded the original Figures S6A-C with four organ-specific supplementary figures (Figures S6-S9). For each organ, we now provide:

      Cluster-level compositional analyses showing changes in abundance between healthy and fibrotic samples. (A) Percentage of mesenchymal cell labels as disease-associated fibroblast (blue) and "rest" per study. (B) Expression of canonical marker genes for myofibroblasts, pericytes, and smooth muscle cells across clusters. (C) The top marker genes for the cluster(s) selected as disease-associated fibroblasts. (C) UMAP visualizations colored by disease etiology and disease condition (fibrosis vs. control), the study, and the original author-provided cell state annotations, including myofibroblast/activated fibroblast annotations where available. (D - G) UMAP visualizations colored by the final annotations used in the subsequent analysis. (H) These additions make the selection procedure substantially more transparent and provide multiple independent lines of evidence supporting the identification of disease-associated fibroblast populations.

      The rationale for the selected clusters is now evident from the revised supplementary figures. In the lung, the selected cluster 3 exhibits a clear increase in abundance in fibrotic samples, expresses canonical myofibroblast markers, and corresponds closely to activated fibroblast/myofibroblast annotations provided in the original studies. In the heart, the selected cluster 1 was the only population showing a robust disease-associated expansion together with strong myofibroblast marker expression and agreement with published annotations. Although another small cluster (cluster 4) displayed partial myofibroblast characteristics, its very low abundance would have a negligible impact on our pseudobulk-based analyses. In the liver, the selected cluster showed consistent expansion across studies and expressed canonical myofibroblast markers, although author-provided annotations were not available for direct comparison. Finally, the kidney datasets presented the greatest integration challenges, likely due to differences between single-cell and single-nucleus protocols. Here, we selected two clusters (cluster 0 and cluster 4) that increased in fibrosis and expressed fibroblast-associated markers, while excluding another expanding cluster (cluster 2) that showed a pericyte-like expression profile. Overall, our final annotations were broadly consistent with the original study annotations wherever such information was available.

      Changes in the manuscript:

      "We integrated the mesenchymal cell population per organ and identified a disease-associated cluster by compositional analysis (Figure 4A, Figures S7-Figure S10)."

      Furthermore, we added the following section to our methods to clarify our methodology:

      "Candidate clusters were required to show consistent enrichment in fibrotic samples and a transcriptional profile characteristic of activated fibroblasts/myofibroblasts. In cases where multiple candidate populations were present, clusters with low abundance or expression profiles inconsistent with myofibroblast identity (e.g., pericyte-like populations) were excluded. Final cluster assignments were validated against the original study annotations whenever available."

      Reviewer #2

      Reviewer #2, major comment 1: Fig.4A: Fibroblast Population Analysis. The authors integrated the fibroblast populations per organ to identify a disease-associated cluster by compositional analysis. In some models, more than one pathological clusters are revealed by the analysis. Shouldn't they be included as pathological, or at least excluded, from the reference population used as a control for differential expression?

      We thank the reviewer for this important comment. We agree that, in some organs, more than one cluster shows features associated with disease and that the selection of disease-associated fibroblast populations should therefore be carefully justified. To improve transparency, we have substantially expanded the supplementary analyses and replaced the original Figures S6A-C with four organ-specific supplementary figures (Figures S7-S10), as described in our answer to Reviewer #1, major comment 7.

      Regarding the reviewer's suggestion to exclude additional potentially pathological clusters from the reference population, we chose not to do so. In many cases, the identity of these secondary clusters is less clear, and excluding them would introduce an additional layer of subjective decision-making that may not necessarily improve robustness. Instead, we used a conservative strategy in which only well-supported disease-associated fibroblast populations were explicitly selected. Furthermore, all downstream analyses of disease-associated fibroblasts were performed using pseudobulk profiles. Because pseudobulk aggregation emphasizes broad transcriptional trends, we expect the resulting signatures to be relatively robust to the inclusion or exclusion of small, ambiguously annotated subpopulations. For these reasons, we believe that retaining the remaining mesenchymal populations in the reference group provides the most objective and reproducible framework for the differential expression analysis.

      For changes in the manuscript associated to this comment, please see our answer to Reviewer #1, major comment 7.

      Reviewer #2, major comment 7: Scar-specific cell-cell communication: Using only COL1A1 as a marker may not be the best option, as this gene is also expressed in normal areas. Suggestion: Use a score combining the best fibrosis-associated genes across the four organs to define fibrotic areas more accurately?

      We thank the reviewer for this suggestion. We agree that COL1A1 is not exclusively expressed in fibrotic regions and can also be detected in normal tissue. To make the analysis more robust, we revised our approach and no longer rely on a single marker gene. Instead, we now compute an enrichment score based on a broader set of established extracellular matrix components, including all collagens and proteoglycans collected by Naba et al. (2012), thereby identifying regions characterized by active matrix deposition rather than expression of COL1A1 alone.

      We then assess the spatial colocalization of candidate ligands and receptors with these ECM-enriched regions across the entire tissue section and focus on the strongest colocalization signals. Importantly, this spatial analysis is subsequently integrated with the disease-associated fibroblast analysis, allowing us to prioritize genes that are both enriched in disease-associated fibroblasts and localized to ECM-rich regions.

      We acknowledge that ECM-rich regions are not necessarily equivalent to fibrotic scar tissue and that some physiologically matrix-producing regions may also be captured by this approach. However, because the analysis is performed across entire tissue sections and multiple independent samples, we expect such regions to contribute primarily as background signal for fibrotic slides. By focusing on the strongest and most consistently colocalizing ligands and receptors across samples, the analysis is designed to identify signals robustly associated with ECM-rich regions rather than being driven by isolated areas of physiological matrix expression.

      We considered the reviewer's suggestion of defining fibrotic regions using fibrosis-associated genes derived from our single-cell analyses. However, we chose not to pursue this strategy because it would introduce a degree of circularity into the analysis. Specifically, the same fibrosis-associated genes would first be used to define fibrotic regions and evaluate for spatial association with candidate ligands and receptors. They would naturally be used again in the gene expression ranking of disease-associated fibroblasts. However, we would like to compare those genes we have found in our meta-analysis with an independent data-modality. Therefore, by instead using an independent ECM-based definition of scar regions, we avoid this potential bias and maintain a clearer separation between the identification of fibrotic regions and the prioritization of disease-associated signaling molecules.

      We compared the results from before (COL1A1-to-gene colocalization) to our results now (ECM enrichment-to-gene colocalization) and found high correlation values between both results for each organ (Review Plan Figure 1). To further show that we expect the pathophysiological ECM signature to largely overshadow physiological ECM expression, we quantified their scores per slide (Figure 6B). We think that our new analysis method is more robust than before, as we now combine several genes into one score.

      We have updated Figure 6 and its text with these new results in our manuscript:

      "To refine these insights, we next focused on identifying ligands and receptors that are specifically expressed in actively scarring regions. We prioritized these molecules because, as extracellular signaling factors and cell-surface proteins, they are directly accessible to therapeutic intervention and therefore represent particularly attractive candidate targets. Structural extracellular matrix molecules were excluded as candidate genes in this analysis and were used instead for the identification of fibrotic scar regions.

      Accordingly, we calculated an ECM enrichment score for each spatial spot, based on a broad set of established structural extracellular matrix components, consisting of all collagens and proteoglycans collected by Naba et al.(Naba et al. 2012). We then computed the spatial colocalization of all remaining ligands and receptors with the identified scarring regions (see methods). Finally, we compared the scar-localization of each gene per organ to the organ-consensus scores of disease fibroblasts (Figure 6A). ECM enrichment scores were significantly elevated in fibrotic compared with control samples across all four organs (Wilcoxon rank-sum test: heart p = 0.005; lung p = 0.002; liver p = 0.014; kidney p = 4e-6, Figure 6B), indicating that pathological extracellular matrix production substantially exceeds physiological ECM turnover. We overall observed a low correlation between scar localization of ligands and receptors and organ effect size in each organ (R in heart = 0.32, liver = 0.38, lung = 0.09, kidney = 0.11), suggesting several cell types and states to be involved in scar-tissue gene expression or a fibrotic gene expression change that goes beyond the scar area (Figure 6C). When comparing the overlap between top ranked genes per organ (upper 20th percentile in gene regulation and colocalization), we observed 8 genes that were identified in 3 out of 4 organs (VIM, TIMP1, FSTL1, CCN2, ANXA2, FBN1, FN1, THBS2), and 2 genes (TIMP2, MRC2) that were identified in all four organs (Figure 6D)."

      Furthermore, we updated Supplementary Figure 12 to include ECM enrichment scores instead of COL1A1 expression.

      Finally, we updated the methods section:

      "To identify actively scarring regions, we performed an enrichment analysis of the geneset consisting of Collagens and Proteoglycans using decoupler's (124) (v1.9.0) univariate linear model (ULM). The spatial colocalization of scarring regions and targets of interest was estimated with the bivariate Moran's R metric implemented in LIANA+ (130) v1.5.0 per target and Visium slide."

      Reviewer #3

      Reviewer #3, minor comment 13: P30: The staging or severity of each of the diseases seems like quite a strong confounder, especially if there is a bias for sampling tissues that are late stage. It would be nice to see this addressed more explicitly in the results, perhaps with some comparisons between those that are identified as earlier and later stage in the respective fibrotic diseases (if these annotations exist).

      We thank the reviewer for raising this important point. We agree that disease stage and severity are potential confounding factors in any meta-analysis of fibrotic diseases and that a bias toward sampling late-stage disease could influence the molecular programs identified.

      Unfortunately, disease staging and fibrosis severity annotations were not consistently available across the published datasets included in our analysis. As a result, we were unable to systematically stratify samples into early- and late-stage disease groups across all organs and disease etiologies. We have therefore highlighted this limitation in the discussion:

      "Second, the limited availability of patient metadata leaves many aspects unresolved, including the exact diagnosis, disease severity, tissue sampling location, and the extent of fibrosis. If these aspects were better documented, they could be accounted for in the analysis and could allow a clearer distinction of physiological from pathophysiological fibrotic processes."

      Nevertheless, we sought to address this concern in the subset of studies for which fibrosis-related severity measurements were available. Specifically, we derived organ-specific fibrosis signatures, scored these signatures across patients, and performed per-study normalization. In these datasets, fibrosis signature scores correlated with available fibrosis severity measurements, supporting the biological relevance of the identified programs (Figure S5A-D). In addition, these analyses indicate that fibrosis signature scores vary across disease etiologies, consistent with the reviewer's suggestion that different diseases may exhibit distinct degrees of fibrotic remodeling (Figure S5E).

      Nevertheless, because detailed histological and clinical metadata are available only for a limited subset of studies, we believe that a comprehensive analysis of fibrosis severity, disease chronicity, and etiology-specific remodeling is beyond the scope of the currently available data and that the currently available metadata are insufficient to robustly compare early- and late-stage disease across the full collection of datasets. We agree that a systematic investigation of stage-specific fibrotic programs would be highly valuable and represents an important direction for future studies using more comprehensively annotated patient cohorts.

      For changes in the manuscript associated to this comment, please see our answer to Reviewer #1, major comment 2.

      3.2 Editorial corrections or clarity improvements

      Reviewer #1

      Reviewer #1, major comment 6: The authors focused on the common functions between mesenchymal and endothelial cells among organs in Fig. 3H and I. Are there cell type specific effects here but shared across organs?

      We thank the reviewer for this question. The results shown in Figures 3H and 3I already represent cell type-specific functional enrichments, as the analyses were performed independently for each cell type before identifying pathways that are consistently altered across organs. Thus, the reported enrichments correspond to cell type-specific effects that are shared across fibrotic diseases in different tissues.

      At the same time, we agree with the reviewer that an interesting observation emerging from these analyses is the overlap in the enriched biological processes identified across different cell types. This suggests that, despite clear cell type-specific transcriptional responses, multiple cell populations converge on a common set of fibrosis-associated pathways. To avoid potential confusion, we have revised the text to clarify that Figures 3H and 3I display cell type-specific enrichments and that the overlap between cell types reflects convergence on shared biological processes rather than identical gene-level responses. Furthermore, we pointed out one difference shown in the plots: the enrichment of neuronal development and axonogenesis pathways in mesenchymal cells.

      "This association with development was further supported by the functional characterization of upregulated genes per organ and cell type."

      [...] "In addition, enrichment of neuronal development and axonogenesis pathways points to activation of projection-related programs, which were not present in the endothelial cell population (Figure 3I). "

      Reviewer #1, major comment 9: It is unclear why only known ligands and receptors are included in the therapeutic target identification analysis in Fig. 6B.

      Our intention was to focus the therapeutic target identification analysis on known ligands and receptors, while excluding major extracellular matrix (ECM) components, because ligands and receptors are generally more amenable to therapeutic intervention and therefore represent particularly attractive candidate targets. To clarify this rationale, we have revised the manuscript text to explicitly describe the criteria used for target selection and the motivation for restricting the analysis to this subset of genes. The corresponding clarification has been added to the results section "Scar-specific cell-cell communication":

      "To refine these insights, we next focused on identifying ligands and receptors that are specifically expressed in actively scarring regions. We prioritized these molecules because, as extracellular signaling factors and cell-surface proteins, they are directly accessible to therapeutic intervention and therefore represent particularly attractive candidate targets. Structural extracellular matrix molecules were excluded as candidate genes in this analysis and were used instead for the identification of fibrotic scar regions."

      Reviewer #1, minor comment 2: The description or legend for the colors is missing in Fig. 3A

      We thank the reviewer for this comment. The color legend was included in the original version of Figure 3A; however, we agree that its placement did not make it sufficiently prominent and may have reduced its visibility. To improve clarity, we have revised the figure layout and repositioned the legend of Figure 3A above the plot so that the color annotation is more readily identifiable.

      Reviewer #1, minor comment 3: FAP appears to be the top gene with robust upregulation in fibrotic heart, lung, liver, and kidney in Fig. 3E, which is also a well-establish surrogate of fibroblast activity and tissue fibrosis in clinical settings (for instance, PMID: 38279381) but not mentioned anywhere in the text.

      We thank the reviewer for highlighting the upregulation of FAP across fibrotic organs. We agree that FAP is a well-established marker of activated fibroblasts and tissue fibrosis and therefore deserves explicit mention at this stage of the analysis. We have revised the text accompanying Figure 3E to highlight FAP as one of the most consistently upregulated genes across organs and to note its established relevance in fibrotic disease:

      "One of the most robustly upregulated genes across organs was prolyl endopeptidase FAP (FAP), a well-established marker gene of activated fibroblasts that has been shown to be functionally relevant in fibrotic diseases in several clinical settings."

      Reviewer #1, minor comment 4: Although it is clear that this study was performed at a much larger scale, the additional gain compared to the previous attempt on identification of shared feature in fibrotic heart, lung, liver, and kidney should be mentioned (PMID: 41752153).

      We thank the reviewer for pointing out this relevant study (PMID: 41752153). We agree that it represents an important previous effort to identify shared features across fibrotic diseases and should be discussed. We have therefore revised the Introduction to acknowledge this work and clarify how the present study extends beyond it. Specifically, while the previous study compared fibrotic heart, lung, liver, and kidney tissues, it was based on a limited number of studies and disease contexts per organ. In contrast, our analysis integrates a substantially larger collection of datasets spanning multiple disease etiologies within each organ, enabling a more systematic assessment of conserved and tissue-specific fibrotic programs across diverse fibrotic diseases.

      • "Recent studies have sought to define shared molecular features across fibrotic diseases affecting the heart, lung, liver, and kidney (15). However, these analyses were based on one study and limited disease contexts per organ, restricting their ability to systematically assess the robustness and generalizability of shared fibrotic programs across diverse disease etiologies."*

      Reviewer #2

      Reviewer #2, major comment 4: Fig.4D: Among this top list, DNM3OS has been indeed characterized as a regulator of the TGF-β pathway in lung fibrosis and should be cited (PMID: 30964696). Interestingly, this lncRNA encodes a cluster of miRNA, including miR-199a-5p, that has been found deregulated in various fibrotic models including lung, kidney and liver (PMID: 23459460).

      We thank the reviewer for highlighting the functional relevance of DNM3OS in fibrosis to improve the manuscript. We checked the literature and agree that its role as a regulator of TGF-β signaling and the involvement of its associated miRNA cluster, including miR-199a-5p, provide important context for interpreting our findings.

      We have therefore expanded the discussion of Fig. 4D and DNM3OS in the manuscript and added the suggested references. Specifically, we now note that DNM3OS was consistently upregulated across organs and that both DNM3OS and its associated miRNA miR-199a-5p have been implicated as downstream effectors of TGF-β signaling involved in myofibroblast activation in lung fibrosis, as well as in experimental models of liver and kidney fibrosis.

      "Furthermore, long noncoding RNA dynamin 3 opposite strand (DNM3OS) was consistently upregulated across organs. DNM3OS and its associated miRNA, miR-199a-5p, have been identified as downstream effectors of TGF-β signaling and implicated in myofibroblast activation in lung fibrosis (76), as well as in experimental mouse models of liver and kidney fibrosis (77)."

      Reviewer #2, major comment 5: Fig. 3F-I and Fig. 4E: the list of the predicted downstream genes for each TF should be provided in a supplemental table

      The transcription factor target gene sets used in these analyses were not generated as part of this study but were obtained from previously published and publicly available regulatory network resources. Because these target gene lists are extensive and already available through the original resource, we did not include them as supplementary tables. To improve transparency and reproducibility, we have revised the manuscript to clearly state the source of these regulatory networks and provide the corresponding reference(s) and access information, allowing readers to retrieve the complete target gene sets used in our analyses. Therefore, in the section "Common aspects of fibrosis across tissues in endothelial and mesenchymal cells", we now state that the collection is publicly available and refer to the methods section:

      "From organ effect sizes, we also inferred transcription factor (TF) activities per organ using CollectTRI (54), a curated publicly available collection of TF-targets, and identified the most commonly upregulated TFs based on the up- or downregulation of the genes they regulate across organs (see methods)."

      In addition, we specifically state in the methods section how the regulons can be accessed:

      "CollecTRI regulons are publicly accessible as described in the original publication (64), for instance at https://zenodo.org/records/8192729?preview_file=CollecTRI_regulons.csv."

      Reviewer #2, minor comment 2: Several panels (Fig.3F-I, Fig.4E-F) need to be improved, in particular the dot plots. with the same order for organs than for the other panels and another range for the size of the dots (-log10 pvalue) to reduce the max size of the dot as well as the enrichment score to expand the value of the z-score.

      We thank the reviewer for these suggestions regarding figure presentation. To improve the readability and consistency of the dot plots, we have made several changes to the figures. We believe these changes substantially improve the interpretability of the figures while preserving the underlying biological signal.

      First, we reordered the organs in Figures 3F-I and 4E-F (see above, in answer to Reviewer #1, minor comment 2 and below, respectively) to match the ordering used throughout the remainder of the manuscript. Second, we expanded the displayed enrichment score range from −2 to 2 to −4 to 4. While many values remain relatively homogeneous, this reflects the fact that these panels were specifically designed to highlight the most consistently shared and strongly regulated signals across organs. Third, we adjusted the dot size scaling for the adjusted p-values. To further improve the visualization of statistical significance, we now explicitly indicate significance using circle outlines: features with an adjusted p-value

      Reviewer #2, minor comment 4: The study is meticulously designed and clearly presented, employing a robust combination of computational approaches. To the reviewer's knowledge, this is the first systematic, cross-organ meta-analysis of fibrosis, offering a comprehensive characterization of both organ-specific and shared gene programs associated with fibrotic processes. A particularly commendable aspect of this work is the provision of a rich and accessible dataset through an interactive data browser, which will serve as a valuable resource for the scientific community at large. The impact of this study is broad and multidisciplinary, benefiting not only computational biologists but also experimental biologists and clinicians working in the field of fibrosis.

      We appreciate the positive assessment of our work and would like to thank the reviewer for recognizing the value of the systematic cross-organ analysis and the interactive data browser. We are pleased that the reviewer considers the study to be a useful resource for the fibrosis research community and appreciates its potential relevance to computational and experimental researchers, as well as clinicians.

      Reviewer #3

      Reviewer #3, minor comment 2: P6: 43 {plus minus} 9 % - this looks a little strange as a percentage, leaving it as a count would probably be clearer as its quite a small number. Please clarify here what 'feature count' here refers to.

      We agree that the notation "43 {plus minus} 9%" may be less intuitive. However, we chose to retain the percentage because it summarizes the proportion of female samples across datasets rather than the total number of samples, which varies substantially between studies. To improve clarity, we removed the variability term and now report only the percentage of samples in the section Data curation for a cross-organ comparison of fibrotic diseases (p.6):

      "In studies with available gender information (16/22 datasets), 43 % of samples were female on average (Figure 1D)."

      In addition, we clarified the meaning of "feature count" by replacing this term with "gene count" throughout the text and in Suppl. Figure 1B.

      Reviewer #3, minor comment 3: P8: Caption: Could you expand a bit upon this 'molecular change severity' in the text?

      We thank the reviewer for pointing this out. We agree that at this point in the manuscript, the concept of "molecular change severity" is not clear yet. It is described later in the manuscript at the beginning of the section "Fibrotic disease programs within tissues" and refers to our analysis with scDist.To make this clearer at its first mention, we have revised the caption to explicitly direct readers to the relevant section and figures. The caption now states:

      "Studies displayed in grey were excluded after an initial assessment of molecular change severity between patient groups, as discussed in the section 'Fibrotic disease programs within tissues' (Figure S2A-D & methods)."

      We believe this addition improves clarity while avoiding duplication of the more detailed explanation provided later in the manuscript.

      Reviewer #3, minor comment 4: Do the author annotated cell types correspond reasonably well with your cell type labels, in those datasets where its present?

      We would like to clarify that we did not perform de novo cell type annotation in the studies except for two. Instead, we used the cell type annotations provided by the original study authors and harmonized them into broader cell type categories based on their names to enable comparisons across studies and organs. The mapping between the original study annotations and these harmonized categories is already provided in Supplementary Table 1. To make this more explicit, the text now states:

      "To enable a comparison across tissues, we grouped cells into five broad categories based on the author's annotations: endothelial-, epithelial-, mesenchymal-, lymphoid-, and myeloid cells (mappings available in Suppl. Table 1)."

      Furthermore, the consistency of these annotations was assessed by examining the expression of cell type marker genes, as shown in Figure 1F, which supports the validity of the harmonized cell type labels used throughout the study.

      Reviewer #3, minor comment 5: P11: A little more information on scDist and what the distances are calculated based on would be good here.

      We thank the reviewer for this suggestion. We agree that the original description did not sufficiently explain how ScDist quantifies molecular differences between conditions. We have therefore expanded the text to clarify that ScDist is a mixed-effects modeling framework and that larger distances correspond to stronger disease-associated transcriptional perturbations:

      "To do so, we applied ScDist (36), a mixed-effects modeling framework that quantifies transcriptomic differences between conditions while accounting for donor-to-donor variability (see methods). For each cell type, ScDist estimates a distance in gene expression space between healthy and fibrotic cells, with larger values indicating stronger disease-associated transcriptional changes."

      Furthermore, we added to the methods:

      "To assess disease-associated transcriptional shifts within each cell type, we applied scDist (v1.1.2) (117) to estimate transcriptional distances between fibrotic and control samples. ScDist assesses disease-associated transcriptional shifts within each cell type by using a linear mixed-effects model that separates condition-associated transcriptional changes from inter-individual variability by including the disease condition as a fixed effect and donor-specific variation as a random effect."

      Reviewer #3, minor comment 7: P14: Are these genes known to be implicated in fibrotic diseases? I know that this is discussed further later, but a few words here would be good.

      We added some context to some of the mentioned genes into the text:

      ** "Notably, several of the highest-ranked genes by our analysis are well-established stress-response and fibrosis markers, such as POSTN38,39, SPP140, VCAN41,42, COL15A121, C343,44, FABP445, and VWF46,47, providing confidence that the identified signatures capture true disease processes instead of study-specific occurrences."

      Reviewer #3, minor comment 8: P17: Fig 3H: enrichment -> enrichment score? (same elsewhere)

      We thank the reviewer for noting this ambiguity. We agree that the term "enrichment" was imprecise in this context. To improve clarity and consistency, we have revised the figure legends of Fig 3 F-I and Fig 4 E-F to explicitly refer to the reported metric as the enrichment score rather than simply enrichment. The updated figure 3 can be found in our answer to Reviewer #1, minor comment 2, the updates to Figure 4 in our answer to Reviewer #2, minor comment 2.

      Reviewer #3, minor comment 9: P19: ULM is used a few times in the captions, but only ever defined in the methods.

      We agree that the abbreviation ULM was not sufficiently defined in the main text and figure legends. To improve readability, we now define the term ULM as univariate linear model at its first occurrence in the figure legends (Figure 3I, page 18).

      The figure caption now reads:

      "For F-I: Dots show the enrichment score (positive: upregulated in fibrosis), while sizes show the -log10 of the adjusted p-values of univariate linear model (ULM) enrichments."

      Reviewer #3, minor comment 10: P20: 'disease relevant cell states' - this might need rewording to better reflect the compositional analysis, and not imply that this identifies cell states rather than clusters of cells.

      We agree that compositional analysis formally identifies cell clusters enriched in disease rather than directly establishing biological cell states. We have revised the text to refer to disease-associated mesenchymal populations/clusters identified through compositional analysis rather than "disease-relevant cell states":

      "To identify disease-associated mesenchymal subpopulations in our datasets, we integrated the mesenchymal cell population per organ and identified a disease-associated cluster by compositional analysis"

      "We also explored disease-associated mesenchymal subpopulation-specific gene expression and the spatial localization of ligands and receptors."

      Reviewer #3, minor comment 11: P22: Fig 4D: This could do with more dynamic range on the colour axis, as most things are near or above the scale.

      We thank the reviewer for this suggestion & agree that the original color scale provided limited visual separation between highly concordant features. We note that this is, in part, a consequence of the panel's design, as Figure 4D specifically highlights genes that are consistently and strongly regulated across organs and therefore exhibit relatively similar effect sizes. Nevertheless, to improve visual discrimination, we have adjusted the color scale of Figure 4D (and similarly, Figure 3 D and E) to provide greater dynamic range and enhance the visibility of differences between genes while preserving the underlying data. We believe this modification improves the interpretability of the figures. The new figures 3 and 4 can be found in our answers to Reviewer #1, minor comment 2 and Reviewer #3, minor comment 8, respectively.

      Reviewer #3, minor comment 15: It would be nice to keep the gene naming schemes consistent (i.e., MOXD1 and TNC), especially within the same discussion.

      We thank the reviewer for this suggestion and agree that consistent gene nomenclature improves readability. We have therefore revised the discussion text to use a consistent naming.

      Reviewer #3, minor comment 17: 'some studies have highlighted the disease-relevance of specific cell states' -> please cite

      To support this statement, we have added the appropriate references describing disease-relevant cell states in fibrotic tissues:

      "Lastly, with exception to the mesenchymal cell population, our analysis primarily focused on broad cell type categories, even though some studies have highlighted the disease-relevance of specific cell states (22, 73,74,7,75,33)".

      Reviewer #3, minor comment 18: Code availability: I think the 'fi' digraph in the link for https://github.com/saezlab/organfibrosis breaks it, but after correcting it manually I can access the repository.

      We thank the reviewer for noting this issue. The hyperlink functions correctly in the submitted manuscript PDF, but we are not sure in which format the reviewer received the manuscript. We will work with the editorial team during the publishing process to ensure that the repository link will be displayed correctly and remains fully accessible in the published version.

      Description of analyses that authors prefer not to carry out

      Reviewer #1

      -

      Reviewer #2

      Reviewer #2, major comment 2: Myeloid Cell Analysis: given the importance of myeloid cells in fibrotic processes, particularly the origin of pathological cells (often monocyte-derived macrophages), it would be highly informative to adopt a similar approach to determine whether myeloid subpopulations differ depending on the affected organ.**

      We thank the reviewer for this suggestion and agree that myeloid cells play a critical role in fibrosis. A systematic comparison of disease-associated myeloid states across organs would therefore be highly valuable. In the present study, however, we chose to focus our state-level analysis on mesenchymal cells because they represent the principal effector population responsible for extracellular matrix deposition and scar formation across fibrotic diseases and because they showed a promising overlap between tissues at the broad cell type level. In contrast, our cross-organ analyses indicate weaker transcriptional conservation among myeloid cells (highest cross-organ disease score prediction AUROC mesenchymal: 0.88; myeloid: 0.72), suggesting that organ-specific immune responses may contribute more strongly than shared fibrosis-associated programs.

      Moreover, our integrated dataset combines both single-cell and single-nucleus sequencing studies, which are known to differ in transcript capture and cell type recovery, especially in immune cells (Feng et al. 2026; Van Melkebeke et al. 2024b; Denisenko et al. 2020). These technical differences already complicated the robust comparison of mesenchymal populations, and we expect they would present an even greater challenge for the identification and comparison of fine-grained myeloid cell states across studies and organs. We therefore chose to focus our detailed state-level analysis on mesenchymal populations, where the biological question was most directly aligned with the central objective of identifying conserved fibrogenic programs across organs.

      Therefore, extending the same analysis to myeloid populations would require a comprehensive integration, annotation, and validation effort that would substantially expand the scope of the current study. We therefore chose to focus our in-depth state-level analysis on the mesenchymal compartment, which is most directly aligned with the central objective of identifying conserved fibrogenic programs across organs.

      Reviewer #3

      -

      References

      Argelaguet, Ricard, Damien Arnol, Danila Bredikhin, et al. 2020. "MOFA+: A Statistical Framework for Comprehensive Integration of Multi-Modal Single-Cell Data." Genome Biology 21 (1): 111. https://doi.org/10.1186/s13059-020-02015-1.

      Denisenko, Elena, Belinda B. Guo, Matthew Jones, et al. 2020. "Systematic Assessment of Tissue Dissociation and Storage Biases in Single-Cell and Single-Nucleus RNA-Seq Workflows." Genome Biology 21 (1): 130. https://doi.org/10.1186/s13059-020-02048-6.

      Feng, Xue, Yu Feng, Sayed Haidar Abbas Raza, Yun Ma, and Hongyu Deng. 2026. "Single Cell and Single Nucleus RNA Sequencing in Liver Tissues: Applications and Prospects in Model and Non-Model Organisms." Frontiers in Genetics 17 (April): 1781941. https://doi.org/10.3389/fgene.2026.1781941.

      Koenitzer, Jeffrey R., Haojia Wu, Jeffrey J. Atkinson, Steven L. Brody, and Benjamin D. Humphreys. 2020. "Single-Nucleus RNA-Sequencing Profiling of Mouse Lung. Reduced Dissociation Bias and Improved Rare Cell-Type Detection Compared with Single-Cell RNA Sequencing." American Journal of Respiratory Cell and Molecular Biology 63 (6): 739-47. https://doi.org/10.1165/rcmb.2020-0095MA.

      Lake, Blue B., Rajasree Menon, Seth Winfree, et al. 2023. "An Atlas of Healthy and Injured Cell States and Niches in the Human Kidney." Nature 619 (7970): 585-94. https://doi.org/10.1038/s41586-023-05769-3.

      Litviňuková, Monika, Carlos Talavera-López, Henrike Maatz, et al. 2020. "Cells of the Adult Human Heart." Nature 588 (7838): 466-72. https://doi.org/10.1038/s41586-020-2797-4.

      Naba, Alexandra, Karl R. Clauser, Sebastian Hoersch, Hui Liu, Steven A. Carr, and Richard O. Hynes. 2012. "The Matrisome: In Silico Definition and In Vivo Characterization by Proteomics of Normal and Tumor Extracellular Matrices*." Molecular & Cellular Proteomics 11 (4): M111.014647. https://doi.org/10.1074/mcp.M111.014647.

      Van Melkebeke, Lukas, Jef Verbeek, Dora Bihary, et al. 2024a. "Comparison of the Single-Cell and Single-Nucleus Hepatic Myeloid Landscape within Decompensated Cirrhosis Patients." Frontiers in Immunology 15 (February). https://doi.org/10.3389/fimmu.2024.1346520.

      Van Melkebeke, Lukas, Jef Verbeek, Dora Bihary, et al. 2024b. "Comparison of the Single-Cell and Single-Nucleus Hepatic Myeloid Landscape within Decompensated Cirrhosis Patients." Frontiers in Immunology 15 (February). https://doi.org/10.3389/fimmu.2024.1346520.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the paper, the authors propose a new RNA velocity method, TSvelo, which predicts the transcription rate linearly based on the expression of RNA levels of transcription factors. This framework is an extension of its recent work TFvelo by including unspliced reads and designing a coherent neuralODE framework. Improved performance was demonstrated in six diverse datasets.

      Strengths:

      Overall, this method introduces innovative solutions to link cell differentiation and gene regulation, with a balance between model complexity (neuralODE) and interpretability (raw gene space).

      We thank the reviewer for the positive evaluation of our work and for recognizing the novelty of the proposed framework. We appreciate the reviewer’s summary highlighting that TSvelo extends our previous method TFvelo by incorporating unspliced reads and introducing a coherent neuralODE framework to model transcription dynamics.

      We are encouraged that the reviewer recognizes the potential of our approach to link cell differentiation with gene regulatory mechanisms, while maintaining a balance between model expressiveness and interpretability in the gene expression space. In the revised manuscript, we have further clarified several methodological details and strengthened the presentation to better highlight these aspects.

      Weaknesses:

      While it seems to provide convincing results, there are multiple technical concerns for the authors to clarify and double-check.

      (1) The authors should clarify and discuss the TF-target map: here, the TF-target genes map is predefined by the TF binding's ChIP-seq data. This annotation is largely incomplete and mostly compiled from a set of bulk tissues. Therefore, for a certain population, the TF-target relation may change. This requires clarification and discussion, possibly exploring how to address this in the model. In addition, a regulon database could be added, e.g., DoRothEA?

      We thank the reviewer for this important comment. The TF–target maps used in TSvelo (e.g., derived from ChIP-seq-based resources such as ENCODE) reflect aggregated TF binding evidence collected across diverse bulk cell types and experimental conditions. As such, they are inherently incomplete and do not capture fully context-specific regulatory activity in a given primary tissue. In TSvelo, we therefore do not treat these annotations as fixed or cell-type-specific ground truth regulatory relationships. Instead, they are used as a permissive prior that encodes a broad set of potential regulatory interactions.

      Within the TSvelo framework, the contribution of each TF–target interaction is learned from data through weight estimation, allowing the model to down-weight or effectively ignore prior edges that are inconsistent with the observed single-cell expression dynamics. This design enables TSvelo to remain robust even when the prior TF–target map is noisy, incomplete, or derived from heterogeneous bulk contexts.

      Following the reviewer’s suggestion, we additionally incorporated the DoRothEA regulon database as an alternative prior with confidence-level filtering. We further performed ablation studies on the pancreas dataset and the gastrulation erythroid dataset using different TF–target resources, including ChEA, ENCODE, and their combinations with DoRothEA.

      The results on the pancreas dataset and the gastrulation erythroid dataset are shown in Figure S13 and Figure S14 respectively, which come up with the same conclusion. We observed highly consistent results across most TF–target prior combinations, including ChEA, ENCODE, ChEA+ENCODE, ChEA+DoRothEA, ENCODE+DoRothEA, and ChEA+ENCODE+DoRothEA. Using the pancreas dataset as example, the mean velocity consistency ranged from 0.985 to 0.995, the mean in-cluster coherence ranged from 0.983 to 0.992, and the mean cross-boundary direction correctness ranged from 0.719 to 0.740 across all settings. These consistently high and tightly bounded metrics indicate that TSvelo is largely insensitive to the specific choice of TF–target prior.

      The only configuration showing reduced stability was the use of DoRothEA alone, particularly in terms of cross-boundary direction correctness. This is likely due to its comparatively limited coverage of TF–target interactions. For instance, in the pancreas dataset, only 81 out of 2000 highly variable genes (HVGs) could be associated with TFs based on DoRothEA, corresponding to 102 TF–target links in total, which may restrict downstream regulatory modeling. In contrast, ChEA covered 1793 genes with 13,976 TF–target links, and ENCODE covered 1854 genes with 33,076 links. These results further suggest that integrating multiple TF–target resources could improve performance, likely due to increased coverage and complementary regulatory information.

      We further acknowledge that regulatory interactions are inherently context-dependent, and that no static TF–target resource can fully capture tissue-specific regulatory programs. In the revised Discussion, we explicitly clarify this limitation and highlight that incorporating context-specific regulatory data (e.g., single-cell chromatin accessibility or perturbation-based regulatory maps) represents an important direction for future improvement.

      (2) The authors should clarify how example genes are selected. This is particularly unclear in Figure 2d.

      We thank the reviewer for raising this point. The example genes shown in Fig. 2d were selected to illustrate representative scenarios where our method provides advantages, particularly cases in which the unspliced–spliced 2D phase portrait exhibits mixed or overlapping patterns that are difficult to model using conventional RNA velocity approaches. These examples are therefore intended to demonstrate the types of transcriptional dynamics that TSvelo is designed to better capture.

      To avoid the impression of selective presentation, we note that our conclusions are based on systematic evaluation across all genes and datasets. Additional visualizations for a broader set of genes on this dataset are provided in Fig. S3. We have clarified the example gene selection criteria in the revised manuscript.

      (3) The authors should clarify confidence in the statement in lines 179-180, that ANXA4 should initially decrease. This is particularly concerning, as TSvelo didn't capture the cell cycle transitions well during the initial part.

      We thank the reviewer for raising this point. The statement that ANXA4 initially decreases is based on the observed expression pattern in the dataset rather than on cell-cycle–related dynamics inferred by the model. Specifically, ANXA4 shows higher expression in Ductal cells compared to Ngn3 EP cells, and Ductal represents an earlier stage in the developmental trajectory. Therefore, along the Ductal to Ngn3 EP transition, ANXA4 naturally exhibits an initial decrease in expression. We have clarified this point in the revised manuscript.

      (4) A support reference should be added for the statement in line 260 that "neuron migrations are inside-out manner". There is no reference supporting this, and this statement is critical for the model assessment.

      We thank the reviewer for this suggestion. This pattern has been reported in previous studies [1,2], which have been added into the revised manuscript.

      To Improve clarity, we have also revised the statement in the manuscript as follows:

      “During cortical development, neurons follow an inside-out layering pattern in which earlier-born neurons populate the deep cortical layers, whereas later-born neurons migrate past them to occupy more superficial layers.”

      (1) Nadarajah, B., Parnavelas, J. Modes of neuronal migration in the developing cerebral cortex. Nat Rev Neurosci 3, 423–432 (2002).

      (2) Li, C., Virgilio, M.C., Collins, K.L. et al. Multi-omic single-cell velocity models epigenome–transcriptome interactions and improves cell fate prediction. Nat Biotechnol 41, 387–398 (2023).

      (5) The comparison to scMultiomics data is particularly interesting, as MultiVelo uses ATAC data to predict the transcription rate. It would be very insightful to add a direct comparison of the estimated transcription rate between using ATAC and directly using TFs' RNA expressions.

      We thank the reviewer for suggesting this highly interesting comparison between ATAC-derived regulatory activity and TF RNA-based proxies for transcription rate estimation.

      We have conducted the requested analysis by computing gene-wise chrome accessibility rate used in MultiVelo and the learned transcription rate from TSvelo, and evaluated their correlation across genes. As shown in Figure S15, the two estimates exhibit almost no global correlation across genes, indicating that they capture substantially different aspects of regulatory information.

      This discrepancy is not unexpected and reflects the fundamental differences between these modalities. scATAC-seq measures chromatin accessibility, which provides a proxy for cis-regulatory potential of genomic regions. However, ATAC signals are inherently sparse and often exhibit a near-binary structure, limiting their ability to directly capture fine-grained temporal regulatory dynamics. In contrast, TF RNA expression reflects downstream transcriptional output, which is shaped by multiple regulatory layers, including post-transcriptional regulation, protein activity, temporal delays, and indirect regulation through intermediate transcriptional or signaling pathways. As a result, these two modalities are expected to capture complementary but not directly comparable aspects of gene regulation.

      Overall, this result suggests that ATAC-based and TF RNA-based signals capture distinct aspects of gene regulation. This further implies that integrating both modalities may be beneficial for future models that aim to more comprehensively characterize transcriptional regulation. We have added this discussion to the supplementary information.

      (6) In Figure 6g, it should be clarified how the lineage was determined. Did the authors use the LARRY barcodes, predicted cell fate, or any other methods? Here, the best way is probably using the LARRY barcodes for individual clones.

      We thank the reviewer for this suggestion. The lineage assignment used in Fig. 6g is described in the Methods section (“Lineage segmentation and pseudotime initialization”). Briefly, lineages are inferred from the transcriptomic structure of the data by performing Leiden clustering followed by PAGA-based connectivity analysis. Starting from an initial Leiden cluster, the filtered PAGA graph defines the shortest paths to other clusters, which are considered as the detected lineages, and diffusion pseudotime (DPT) is then used to initialize pseudotime along each lineage. Thus, in this analysis lineages are determined from the expression-derived trajectory structure. We have clarified this point in the revised manuscript and refer readers to the Methods section.

      Reviewer #2 (Public review):

      Summary:

      Li et al. propose TSvelo, a computational framework for RNA velocity inference that models transcriptional regulation and gene-specific splicing using a neural ODE approach. The method is intended to improve trajectory reconstruction and capture dynamic gene expression changes in scRNA-seq data. However, the manuscript in its current form falls short in several critical areas, including rigorous validation, quantitative benchmarking, clarity of definitions, proper use of prior knowledge, and interpretive caution. Many of the authors' claims are not fully supported by the evidence.

      We thank the reviewer for the careful evaluation of our manuscript and for the constructive comments. We appreciate the concerns regarding validation, benchmarking, methodological clarity, and interpretation. In the revised manuscript, we have carefully addressed these points by adding additional analyses, clarifying methodological details, and moderating several claims to ensure they are fully supported by the data. Detailed responses to each comment are provided below.

      Major comments:

      (1) Modeling comments

      (a) Lines 512-513: How does the U-to-S delay validate the accuracy of pseudotime? Using only a single gene as an example is not sufficient for "validation."

      We thank the reviewer for this important clarification. In the revised manuscript, we have rephrased this part to clarify that Fig. 1a serves only as an illustrative example showing the U-to-S delay for a single gene. Accordingly, we have corrected our statement to indicate that the U-to-S delay is used to infer trajectory orientation, rather than to validate the accuracy of pseudotime.

      In addition, we have expanded the description to explain that U-to-S delay signals are aggregated across all genes to provide a more robust and comprehensive assessment for this purpose. Additional analysis is provided in our response to the next comment.

      (b) Lines 512-518: The authors propose a strategy for selecting the initial state, but do not benchmark how accurate this selection procedure is, nor do they provide sufficient rationale. While some genes may indeed exhibit U-to-S delay during lineage differentiation, why does the highest U-to-S delay score indicate the correct initiation states? Please provide mathematical justification and demonstrate accuracy beyond using a single gene example. Maybe a simulation with ground truth could help here, too.

      We thank the reviewer for this insightful comment. In the revised manuscript, we have clarified both the intuition and justification of this approach. Briefly, along a correctly oriented trajectory, unspliced (U) expression is expected to precede spliced (S) expression due to transcriptional dynamics. Ideally, this U-to-S delay would be observable at the level of individual genes. However, due to the high noise inherent in scRNA-seq data, such delays are often not consistently detectable on a per-gene basis. To address this, we aggregate U-to-S delay signals across all genes and determine the lineage orientation by maximizing a global delay score. Under this criterion, the cluster from which all outgoing lineages exhibit the highest aggregated U-to-S delay is inferred to correspond to the initial state.

      We emphasize that this approach relies on genome-wide aggregation rather than any single gene. Moreover, the same strategy is applied uniformly across all six datasets using identical parameter settings, demonstrating its robustness and stability. To further address the reviewer’s concern, we additionally present the U-to-S delay scores for each Leiden cluster when treated as the initial state across all datasets (Author response image 1). The results on all datasets suggest that the highest U-to-S delay scores can be used to detect the initial cluster.

      Author response image 1.

      The U-to-S delay scores for each Leiden cluster when treated as the initial state across all datasets.

      Following your suggestions, we also add a simulation study. We generated synthetic single-cell RNA velocity datasets using a mechanistic transcriptional dynamics model with one or multiple developmental branches. The system included 200 genes, among which 30 were designated as transcription factors (TFs).

      For each branch, we independently sampled a TF–target regulatory matrix W ϵ R<sup>30×200</sup> from a standard normal distribution to simulate distinct GRN structures. Gene expression dynamics were modeled using a coupled ordinary differential equation (ODE) system describing unspliced and spliced RNA abundances:

      where u and s denote unspliced and spliced RNA levels, respectively. The transcription rate α was computed as a nonlinear function of TF expression, defined as a weighted sum of spliced TF abundance, followed by clipping to ensure bounded activation.

      Each branch is initialized from the same randomly sampled initial condition drawn from a gamma distribution, allowing controlled divergence of trajectories driven solely by branch-specific regulatory programs.

      To simulate observed sequencing counts, we introduced technical noise by scaling latent expression levels with cell-specific library sizes drawn from a log-normal distribution. The resulting expression counts were generated using a negative binomial sampling model:

      where θ controls over dispersion, with smaller values corresponding to higher noise levels. The final datasets consist of paired unspliced (U) and spliced (S) count matrices with realistic transcriptional stochasticity and branching gene regulatory dynamics. For each branch, cells were further divided into three developmental stages for downstream analysis.

      We evaluated TSvelo on multiple simulated datasets with varying numbers of branches and noise levels. There are two or three branches start from the same root cell groups in these datasets (Branch 1: stage 0 - stage 1 - stage 2. Branch 2: stage 0 - stage 3 - stage 4. Branch 3: stage 0 - stage 5 - stage 6). The results of initial state identification based on the unspliced-to-spliced (U-to-S) delay, along with the corresponding 2D velocity stream visualizations, are presented in Supplementary Figure S1. These results demonstrate that the U-to-S delay–based initialization is robust and consistently identifies cells corresponding to the earliest developmental stage (“stage 0”) across different simulation settings. All additional results have been included in the Supplementary Information.

      (c) Equation (8): The formulation looks to be incorrect. If $$W \in \mathbb{R}^{G\times G}$$ and $$W' - \Gamma' \in \mathbb{R}^{K\times K}$$, how can they be aligned within the same row? Please clarify.

      We thank the reviewer for pointing this out. This was a typographical error in the manuscript. In the third line of Equation (8), the term should be W’ instead of W. We have corrected this in the revised manuscript to ensure dimensional consistency.

      (d) The use of prior knowledge graphs from ENCODE or ChEA to constrain regulation raises concerns. Much of the regulatory information in these databases comes from cell lines. How can such cell-line-based regulation be reliably applied to primary tissues, as is done throughout the manuscript? Additional experiments are needed to test the robustness of TSvelo with respect to prior knowledge.

      We thank the reviewer for this important comment. In TSvelo, TF–target networks from resources such as ENCODE and ChEA are incorporated as priors that guide the model toward biologically plausible regulatory structures. Importantly, the contribution of each TF–target interaction is learned from the data, allowing the model to down-weight or override potentially inaccurate or context-mismatched regulatory links. By aggregating signals across a large number of genes, the model further reduces sensitivity to noise and incompleteness in any single prior network.

      To evaluate robustness with respect to prior knowledge, we incorporated the DoRothEA regulon resource as an alternative TF–target prior with confidence-level filtering. We further performed ablation studies on the pancreas dataset and the gastrulation erythroid dataset using different TF–target resources, including ChEA, ENCODE, and their combinations with DoRothEA.

      The results on the pancreas dataset and the gastrulation erythroid dataset are shown in Figure S13 and Figure S14 respectively, which come up with the same conclusion. We observed highly consistent results across most TF–target prior combinations, including ChEA, ENCODE, ChEA+ENCODE, ChEA+DoRothEA, ENCODE+DoRothEA, and ChEA+ENCODE+DoRothEA. Using the pancreas dataset as example, the mean velocity consistency ranged from 0.985 to 0.995, the mean in-cluster coherence ranged from 0.983 to 0.992, and the mean cross-boundary direction correctness ranged from 0.719 to 0.740 across all settings. These consistently high and tightly bounded metrics indicate that TSvelo is largely insensitive to the specific choice of TF–target prior. Notably, these results further suggest that even when the underlying regulatory resources differ in origin (e.g., cell-line-derived vs. curated or aggregated datasets), the inferred dynamics remain stable.

      The only configuration showing reduced stability was the use of DoRothEA alone, particularly for cross-boundary direction correctness. This is likely due to its comparatively limited coverage of TF–target interactions. For instance, in the pancreas dataset, only 81 out of 2000 highly variable genes (HVGs) could be associated with TFs based on DoRothEA, corresponding to 102 TF–target links in total, which may limit downstream regulatory modeling. In contrast, ChEA covered 1793 genes with 13,976 TF–target links, and ENCODE covered 1854 genes with 33,076 links. These results further suggest that integrating multiple TF–target resources can improve performance, likely due to increased coverage and complementary regulatory information.

      We agree that regulatory interactions derived from resources such as ENCODE and ChEA may not fully generalize to primary tissues due to their context-dependent nature. In the revised Discussion, we explicitly clarify this limitation, particularly their inability to capture tissue-specific regulatory programs. We further highlight that incorporating context-specific regulatory data, such as single-cell chromatin accessibility or perturbation-based regulatory maps, represents an important direction for future improvement.

      (e) Lines 579-580: How is the grid search performed? More methodological details are required. If an existing method was used, please provide a citation.

      The grid search for the time step means that the model evaluates the loss in equation (10) across all candidate values of t<sub>step</sub> in the set {0,1,2,...,999}. This strategy was originally adopted in scVelo for optimizing the time step parameter. We have now added the corresponding citation to scVelo in the revised manuscript.

      (2) Application on pancreatic endocrine datasets

      (a) Lines 140-141: What is the definition of the final pseudotime-fitted time t or velocity pseudotime?

      There is no distinction between “final pseudotime”, “fitted time t” and “velocity pseudotime”. All of them refer to the same quantity in our framework. To eliminate any potential ambiguity, we have standardized the terminology by replacing “final pseudotime” with “pseudotime”.

      (b) Lines 143-144: The use of the velocity consistency metric to benchmark methods in multi-lineage datasets is incorrect. In multi-lineage differentiation systems, cells (e.g., those in fate priming stages) may inherently show inconsistency in their velocity. Thus, it is difficult to distinguish inconsistency caused by estimation error from that arising from biological signals. Velocity consistency metrics are only appropriate in systems with unidirectional trajectories (e.g., cell cycling). The abnormally high consistency values here raise concerns about whether the estimated velocities meaningfully capture lineage differences.

      We thank the reviewer for raising this important point regarding the use of the velocity consistency metric in multi-lineage systems. Velocity consistency was initially introduced by scVelo [1] and implemented as scvelo.velocity_confidence() in its package. Velocity consistency provides one of the few widely adopted quantitative criteria for benchmarking RNA velocities [2]. We agree that it is especially suitable for single-lineage processes. For datasets with clear multi-lineage differentiation (Fig. 5 and Fig. 6), we do not use this metric, precisely to avoid the issue highlighted by the reviewer.

      However, the pancreatic endocrine dataset (Fig. 2) exhibits minimal branching, making velocity consistency be more appropriate. As introduced by veloVI study, RNA velocities are supposed to change smoothly over the phenotypic manifold [3]. Higher consistency indicates that neighboring cells show compatible velocity directions, reflecting stable and coherence of the inferred velocity field. Additionally, multiple previous studies used velocity consistency to evaluate model performance on this pancreas dataset [2,3,4], providing a standard point of comparison.

      To better address your concerns, we have replaced the corresponding panel in Fig. 2 of the main text with an evaluation of cell-type separability in both the traditional 2D (unspliced–spliced) phase portrait and the learned 3D (α–unspliced–spliced) phase portrait by TSvelo (Author response image 4 in our response to your subsequent question). We appreciate your suggestions, as the comparison more clearly highlights the novelty and contribution of TSvelo and helps explain its improved performance. Now, the velocity consistency panel has been moved to the Supplementary Information. In addition, we have added a clearer explanation of the cross-boundary correctness metric in the revised manuscript.

      (1) Bergen, V., Lange, M., Peidli, S., Wolf, F. A., & Theis, F. J. (2020). Generalizing RNA velocity to transient cell states through dynamical modeling. Nature Biotechnology, 38(12), 1408-1414.

      (2) Luo, Y., Ren, J., Yang, Q. ... & Li, Q. (2026). Benchmarking RNA velocity methods across 17 independent studies, Cell Reports Methods, 101367.

      (3) Gayoso, A., Weiler, P., Lotfollahi, M., Klein, D., Hong, J., Streets, A., ... & Yosef, N. (2024). Deep generative modeling of transcriptional dynamics for RNA velocity analysis in single cells. Nature Methods, 21(1), 50-59.

      (4) Li, J., Pan, X., Yuan, Y., & Shen, H. B. (2024). TFvelo: gene regulation inspired RNA velocity estimation. Nature Communications, 15(1), 1387.

      (c) The improvement of TSvelo over other methods in terms of cross-boundary direction correctness looks marginal; a statistical test would help to assess its significance.

      We thank the reviewer for this insightful comment. In the revised manuscript, we have added statistical tests for evaluated metrics, including velocity consistency, cross-boundary direction correctness, and in-cluster coherence.

      As shown in Author response image 2, TSvelo significantly outperforms all baseline methods in terms of velocity consistency across both datasets. For in-cluster coherence, TSvelo achieves significantly better performance on the gastrulation (erythroid) dataset, while on the pancreas dataset it performs comparably to the best-performing baselines (UniTVelo and TFvelo) and significantly outperforms several competing methods, including CellDancer, Dynamo, and scVelo.

      For cross-boundary direction correctness, TSvelo shows consistent improvements in mean performance on the pancreas dataset (Author response image 3), and significantly outperforms Dynamo and scVelo on the gastrulation dataset. Although not all pairwise comparisons on cross-boundary direction correctness reach statistical significance, this is likely influenced by the limited number of independent samples (n = 7 and n = 4 for the two datasets, respectively), which reduces statistical power for detecting differences. Importantly, TSvelo still achieves the best average performance among all methods, indicating a consistent overall trend in favor of TSvelo.

      We have added these results into the revised manuscript.

      Author response image 2.

      The quantitative comparison between TSvelo and baseline approaches on the pancreas dataset (panel a) and the gastrulation erythroid dataset (panel b). In each plot, methods are ranked in descending order of their mean values. Numbers at the bottom indicate the sample size for each metric. Significance is determined using a one-sided Mann–Whitney U test. *****, ***, ** and * represent p < 0.00001, 0.0001 ≤ p < 0.001, 0.001 ≤ p < 0.01, and 0.01 ≤ p < 0.05, respectively.

      Author response image 3.

      The comparison of mean cross-boundary direction correctness on the pancreas dataset.

      (d) Lines 177-178: Based on the figure, TSvelo does not appear to clearly distinguish cell types. A quantitative metric, such as Adjusted Rand Index (ARI), should be provided.

      We thank the reviewer for this helpful suggestion. To quantitatively assess whether TSvelo can distinguish cell types, we evaluated the separability of cell-type labels in both the 2D (unspliced–spliced) phase portrait adopted by previous RNA velocity approaches, and the 3D (α–unspliced–spliced, α denotes the transcriptional rate) phase portrait introduced by TSvelo.

      Specifically, we evaluated how well the embedding preserves cell-type information using a k-nearest neighbors (kNN) classification accuracy with 5-fold cross-validation. Given an embedding matrix in 2D or 3D space (X 𝛜 ℝ<sup>n*d</sup>, where n is the number of cells and d is 2 or 3) and corresponding cell-type labels (y 𝛜 {1, … ,C}, we partition the data into five folds. For each fold (k), a kNN classifier with K = 5, denoted asf<sup>(k)</sup>, is trained on the training subset and evaluated on the held-out test subset. The classification accuracy for the k-th fold is defined as ℝ

      where n<sub>k</sub> is the number of samples in the test set and 1(.)is the indicator function. The final score is obtained by averaging across all folds:

      This metric directly assesses whether cells of the same type are positioned close to each other in the embedding space, and is widely used to quantify representation quality.

      Using this evaluation, we observed that the 3D phase portrait consistently achieves significantly higher accuracy than the 2D phase portrait (Author response image 4). The improvement is highly statistically significant (one-sided Mann–Whitney U test, p-value = 4.37 × 10<sup>-10</sup>), demonstrating that the 3D representation provides substantially better separation of cell types.

      We have added these quantitative results to the revised manuscript to complement the visual evidence and to clarify that TSvelo effectively distinguishes cell types in the learned representation.

      Author response image 4.

      The evaluation of the separability of cell-type labels in both the 2D (unspliced–spliced) phase portrait and the 3D (α–unspliced–spliced) phase portrait for the pancreas dataset.

      (e) Lines 179-183: The claim that traditional methods cannot capture dynamics in the unspliced-spliced phase portrait is vague. What specific aspect is not captured-the fitted values or something else? Evidence is lacking. Please provide a detailed explanation and quantitative metrics to support this claim.

      We thank the reviewer for this important comment. We have revised the text to more clearly illustrate this point using representative example genes as follows: “For instance, ANXA4 shows higher expression in Ductal cells compared to Ngn3 low EP cells, which mean its expression pattern exhibits an initial decrease followed by an increase. Such dynamics are not easily captured in the conventional unspliced–spliced phase portrait used by previous approaches, as many baseline methods implicitly assume a decreasing–then–increasing expression pattern. By comparison, TSvelo can still fit such expression pattern by using additional information from the 3D phase portrait.”

      In addition, we also clarify that the 2D u–s representation has limited capacity to separate heterogeneous dynamic cell states, which can affect downstream velocity field estimation. In the conventional 2D u–s phase portrait, cells from different dynamic regimes may overlap in the same region of the embedding space. This overlap reduces the identifiability of underlying transcriptional states and makes the inferred local dynamics more ambiguous. In contrast, TSvelo introduces an additional latent variable α, forming a 3D (α, u, s) phase portrait, which helps disentangle these mixed trajectories and yields a more structured and separable representation of cell dynamics. We have provided quantitative evidence in the previous response (Author response image 4). Briefly, the proposed 3D representation achieves consistently higher kNN classification accuracy (5-fold cross-validation, k=5) for cell state identification compared to the 2D u–s embedding.

      (3) Application to gastrulation erythroid datasets

      (a) Lines 191-194: The observation that velocity genes are enriched for erythropoiesis-related pathways is trivial, since the analysis is restricted to highly variable genes (HVGs) from an erythropoiesis dataset. This enrichment is expected and therefore not informative.

      We thank the reviewer for this comment and agree that such enrichment is expected given the use of HVGs from an erythropoiesis dataset. This analysis was included only as a preliminary sanity check to support the plausibility of the inferred velocity genes, rather than as a main result. We have accordingly simplified the description and clarified that this analysis serves only as a preliminary check in the revised manuscript.

      (b) Lines 227-228: It remains unclear how TSvelo "accurately captures the dynamics." What is the definition of dynamics in this context? Figure 3g shows unspliced/spliced vs. fitted time plots and phase portraits, but without a quantitative definition or measure, the claim of superiority cannot be supported. Visualization of a single gene is insufficient; a systematic and quantitative analysis is needed.

      We thank the reviewer for this important comment. We have revised the text to more clearly illustrate this point using representative example genes as follows: “For HSP90AB1, which exhibits a counter-clockwise pattern in the unspliced–spliced phase portrait, in contrast to the clockwise dynamics typically assumed by most baseline approaches, it is difficult for previous methods to capture this behavior, whereas TSvelo can still faithfully model such patterns. For genes such as RPS26, which have critical roles in the development in blood progenitors to erythroid40, the unspliced-spliced data is so noisy that cells of different types overlap in phase portrait. TSvelo can still captures the gene dynamics and reveals differences in transcription rates across cell types.”

      In addition, we explicitly emphasize the role of the 3D (α, u, s) phase portrait, which provides a more structured and separable representation of transcriptional states compared to the conventional 2D u–s space. This improved representation is the key factor underlying the advantages of TSvelo in modeling transcriptional processes. In the conventional 2D u–s phase portrait, cells from different transcriptional states may overlap, leading to reduced separability. In contrast, introducing the latent variable α expands the representation to a 3D space, which helps disentangle these mixed states and yields a clearer phase structure. Similar to our previous response in Author response image 4, we provide quantitative evidence on this gastrulation erythroid dataset in Figure S7, showing that the 3D representation achieves consistently higher kNN classification accuracy for cell state separation compared to the 2D u–s embedding (one-sided Mann–Whitney U test, p-value = 0.002).

      (4) Application to the mouse brain and other datasets

      (a) Lines 280-281: The authors cannot claim that velocity streams are smoother in TSvelo than in Multivelo based solely on 2D visualization. Similarly, claiming that one model predicts the correct differentiation trajectory from a 2D projection is over-interpretation, as has been discussed in prior literature see PMID: 37885016.

      We thank the reviewer for this important comment. Consistent with other RNA velocity studies, TSvelo employs the 2D UMAP stream plot for visualizing the results. We agree that conclusions based solely on 2D visualizations may lead to over-interpretation. Our intention was to provide an intuitive visualization rather than a rigorous quantitative comparison. Accordingly, we have revised the text to avoid making definitive claims about smoothness or correctness of differentiation trajectories based solely on 2D projections.

      (b) Lines 304-306: Beyond transcriptional signal estimation, how is regulation inferred solely from scRNA-seq data validated, especially compared with scATAC-seq data? Are there cases where transcriptome-based regulatory inference is supported by epigenomic evidence, thereby demonstrating TSvelo's GRN inference accuracy?

      We thank the reviewer for this important question regarding the validation of regulatory inference derived from scRNA-seq data and its comparison to scATAC-seq-based evidence.

      We would like to first clarify the scope of TSvelo. Similar to existing RNA velocity methods, the primary goal of TSvelo is to model transcriptional dynamics and accurately infer cell state transitions and cell fate trajectories. In this context, gene regulatory information is not inferred de novo from data, but incorporated as prior knowledge from curated TF–target databases to guide and constrain the dynamics modeling process, as described in our Introduction.

      We have conducted the requested analysis by computing gene-wise chrome accessibility rate used in MultiVelo and the learned transcription rate from TSvelo, and evaluated their correlation across genes. As shown in Figure S15, the two estimates exhibit almost no global correlation across genes, indicating that they capture substantially different aspects of regulatory information.

      This discrepancy is not unexpected and reflects the fundamental differences between these modalities. scATAC-seq measures chromatin accessibility, which provides a proxy for cis-regulatory potential of genomic regions. In contrast, TF RNA expression reflects downstream transcriptional output, which is shaped by multiple regulatory layers, including post-transcriptional regulation, protein activity, temporal delays, and indirect regulation through intermediate transcriptional or signaling pathways. As a result, these two modalities are expected to capture complementary but not directly comparable aspects of gene regulation.

      We acknowledge that scATAC-seq provides valuable complementary information on chromatin accessibility and regulatory potential, and will consider incorporating matched multi-omics data in future work. In the revised manuscript, we further clarify that TSvelo is an RNA velocity method that incorporates prior knowledge from curated TF–target databases, and we have added a discussion on the potential use of scATAC-seq data for future extension of our framework.

      (c) The claim that TSvelo can model multi-lineage datasets hinges on its use of PAGA for lineage segmentation, followed by independent modeling of dynamics within each subset. However, the procedure for merging results across subsets remains unclear.

      We thank the reviewer for pointing out that the merging step was not sufficiently described. After modeling dynamics independently within each lineage-specific subset, TSvelo integrates the results via a weighted aggregation procedure at the cell level.

      For each cell and each inferred quantity (e.g., velocity or other dynamic variables), we collect the estimates obtained from different lineage-specific models and combine them using a weighted average. The weights are defined by the size of each lineage, reflecting its statistical support. We have clarified details about this merging procedure in the Methods section.

      This aggregation reconciles multiple lineage-specific estimates for the same cell into a single value and mitigates discontinuities that could arise from directly combining independent lineage analyses. The resulting values define a unified set of dynamics for each cell across lineages.

      Reviewer #3 (Public review):

      Despite the abundance of RNA velocity tools, there are still major limitations, and there is strong skepticism about the results these methods lead to. In this paper, the authors try to address some limitations of current RNA velocity approaches by proposing a unified framework to jointly infer transcriptional and splicing dynamics. The method is then benchmarked on 6 real datasets against the most popular RNA velocity tools.

      While the approach has the potential to be of interest for the field, and may present improvements compared to existing approaches, there are some major limitations that should be addressed, particularly concerning the benchmark (see major comment 1).

      Major comments:

      (1) My main criticism concerns the benchmarking: real data lack a ground truth, and are absolutely not ideal for comparing methods, because one can only speculate what results appear to be more plausible.

      A solid and extensive simulation study, which covers various scenarios and possibly distinct data-generating models, is needed for comparing approaches. The authors should check, for example, the simulation studies in the BayVel approach (Section 4, BayVel: A Bayesian Framework for RNA Velocity Estimation in Single-Cell Transcriptomics). Clearly, all methods should be included in the simulation.

      Following your recommendation, we have added the simulation analysis to compare TSvelo with existing RNA velocity approaches. We generated synthetic single-cell RNA velocity datasets using a mechanistic transcriptional dynamics model with one or multiple developmental branches. The system included 200 genes, among which 30 were designated as transcription factors (TFs).

      For each branch, we independently sampled a TF–target regulatory matrix W ϵ ℝ<sup>30×200</sup> from a standard normal distribution to simulate distinct GRN structures. Gene expression dynamics were modeled using a coupled ordinary differential equation (ODE) system describing unspliced and spliced RNA abundances:

      where u and s denote unspliced and spliced RNA levels, respectively. The transcription rate α was computed as a nonlinear function of TF expression, defined as a weighted sum of spliced TF abundance, followed by clipping to ensure bounded activation.

      Each branch is initialized from the same randomly sampled initial condition drawn from a gamma distribution, allowing controlled divergence of trajectories driven solely by branch-specific regulatory programs.

      To simulate observed sequencing counts, we introduced technical noise by scaling latent expression levels with cell-specific library sizes drawn from a log-normal distribution. The resulting expression counts were generated using a negative binomial sampling model:

      where θ controls over dispersion, with smaller values corresponding to higher noise levels. The final datasets consist of paired unspliced (U) and spliced (S) count matrices with realistic transcriptional stochasticity and branching gene regulatory dynamics. For each branch, cells were further divided into three developmental stages for downstream analysis.

      We evaluated TSvelo and those splicing-based RNA velocity approaches on multiple simulated datasets with varying numbers of branches and noise levels. There are one, two or three branches start from the same cell group in these datasets (Branch 1: stage 0 - stage 1 - stage 2. Branch 2: stage 0 - stage 3 - stage 4. Branch 3: stage 0 - stage 5 - stage 6). We primarily assessed performance using the cross-boundary direction correctness (CBDir) metric, as it directly evaluates inferred trajectories against ground-truth cell stage annotations, which have been widely adopted in RNA velocity studies such as VeloAE and UniTvelo. In detail, Cross-boundary direction correctness assesses the accuracy of transitions from a source cluster to a target cluster by examining the boundary cells, and requires ground truth annotations. We directly run the function unitvelo.evaluate() provided in UniTVelo to obtain the Cross-boundary direction correctness. In detail, the CBDir is calculated as follows:

      where θ controls over dispersion, with smaller values corresponding to higher noise levels. The final datasets consist of paired unspliced (U) and spliced (S) count matrices with realistic transcriptional stochasticity and branching gene regulatory dynamics. For each branch, cells were further divided into three developmental stages for downstream analysis.

      where C<sub>A</sub> denotes the set of cells in the target cluster A, and N(c) represents the neighboring cells of a given cell c v<sub>c</sub> and x<sub>c</sub> denote the low-dimensional velocity and state vectors of cell c, respectively, and x<sub>c’</sub> denotes the state vector of its neighboring cell.

      As shown in Figure S2, TSvelo consistently achieves the highest accuracy across all simulation settings, particularly in scenarios with complex branching structures, which pose significant challenges for baseline methods.

      (2) Related to the above: since a ground truth is missing, the real data analyses need to be interpreted with caution. I recommend avoiding strong statements, such as "successfully captures the correct gene dynamics", or "accurately infer", in favour of milder statements supported by the data, such as "... aligns with the biological processes described" (as in page 12), or "results are compatible with current biological knowledge", etc...

      We thank the reviewer for this helpful comment. We agree that analyses on real datasets should be interpreted with appropriate caution because definitive ground truth is typically unavailable. Following the reviewer’s suggestion, we have revised the wording throughout the manuscript to avoid overly strong claims. For example, statements such as “successfully captures the correct gene dynamics” and “accurately infer” have been replaced with more cautious descriptions such as “consistent with known biological processes”.

      (3) Many methods perform RNA velocity analyses. While there is a brief description, I think it'd be useful to have a schematic summary (e.g., via a Table) of the main conceptual, mathematical, and computational characteristics of each approach.

      We thank the reviewer for this insightful suggestion. We agree that a structured summary of existing RNA velocity methods would improve clarity and accessibility. We have added a new summary table (Table S1) that systematically compares representative RNA velocity approaches in the supplementary information.

      (4) Related to the above: I struggled to identify the main conceptual novelty of TSvelo, compared to existing approaches. I recommend explaining this aspect more extensively.

      We thank the reviewer for this insightful comment. We agree that the conceptual novelty of TSvelo can be more clearly articulated.

      In the revised manuscript, we have expanded the discussion at the beginning of the Results section to explicitly highlight the key distinctions between TSvelo and existing approaches. Specifically, we now clarify that most existing RNA velocity methods predominantly focus on splicing dynamics and typically operate in a gene-wise manner, without capturing coordinated dynamics across genes. In contrast, TSvelo models the full cascade of transcriptional regulation, transcription, and splicing within a unified framework, and estimates RNA velocity jointly across all genes, thereby capturing their coordinated dynamics at the system level.

      (5) A computational benchmark is missing; I'd appreciate seeing the runtime and memory cost of all methods in a couple of datasets.

      We thank the reviewer for this helpful suggestion regarding computational benchmarking. In the revised manuscript, we have added a systematic comparison of runtime and GPU memory usage across TSvelo and ba methods using simulated datasets of increasing scale (600, 1200, and 1800 cells) on our NVIDIA GeForce RTX 3090 device with 24 GB memory.

      Table S2 shows differences in computational efficiency and resource requirements among methods. Specifically, classical methods such as scVelo and Dynamo exhibit very fast runtimes (10–24 seconds) and do not rely on GPU acceleration, reflecting their relatively lightweight modeling strategies. In contrast, deep learning–based approaches, including UniTVelo, cellDancer, and TSvelo, have higher computational costs due to their increased model complexity.

      TSvelo exhibits a stable GPU memory footprint (~1.26 GB) across different dataset sizes, indicating that its memory usage is primarily determined by model architecture rather than the number of cells. This level of memory consumption is well within the capacity of modern GPUs and does not pose practical limitations. In terms of runtime, TSvelo scales approximately linearly with dataset size. The higher computational cost of TSvelo is mainly due to its EM-style optimization procedure, where each M-step also involves multiple optimization updates to infer gene regulatory effects in a global model. This design enables TSvelo to explicitly incorporate regulatory priors and jointly model gene interactions, which is not supported by these baseline methods.

      To further improve runtime efficiency, TSvelo allows flexible control of the number of EM iterations. As shown in Figure S16 and Table S3, we evaluated performance under different iteration settings on the simulation dataset. The early stopping strategy employed in the EM framework of TSvelo, which will stop modeling if the loss is not further reduced in the last 3 iterations. Results show that convergence is typically achieved within 3 iterations for this dataset, and increasing the maximum number of iterations beyond this does not further change the results. Notably, even a single iteration already yields competitive performance, likely benefiting from the strong initialization based on unspliced-to-spliced temporal delay.

      Overall, these results highlight a trade-off between computational efficiency and modeling expressiveness. While TSvelo is more computationally demanding than classical approaches, it provides a more flexible framework for incorporating regulatory information and capturing complex gene interactions, which we believe justifies the additional computational cost in scenarios requiring accurate dynamical inference.

      (6) I think BayVel (mentioned above) should be added to the list of competing methods (both in the text and in the benchmarks). The package can be found here: https://github.com/elenasabbioni/BayVel_pkgJulia.

      We thank the reviewer for suggesting BayVel and for providing the repository link. We carefully review the available resources, including both the BayVel_pkgJulia and the BayVel_notebooks, and we appreciate the authors’ efforts in making their code and data publicly available.

      We note that BayVel repositories primarily provide scripts and data for reproducing the figures and results reported in their manuscript. However, at present, the available resources do not yet provide a complete guideline or standardized pipeline for applying BayVel to new datasets. To ensure a fair and reproducible comparison, we therefore tend to use BayVel results officially provided by the authors. We are grateful that the BayVel results on the pancreas dataset is released at BayVel_notebooks page: https://github.com/elenasabbioni/BayVel_notebooks/tree/main/real%20data/Pancreas/moments/output.

      Based on these results, we conducted comparisons across all methods on the pancreas dataset, with quantitative evaluations shown in Author response image 55. In each plot, methods are ranked in descending order of their mean values. Numbers at the bottom indicate the sample size for each metric. Statistical significance is assessed using a one-sided Mann–Whitney U test, where *****, ***, **, and * denote p < 0.00001, 0.0001 ≤ p < 0.001, 0.001 ≤ p < 0.01, and 0.01 ≤ p < 0.05, respectively.

      BayVel has now been included in the Introduction, and corresponding comparisons have been added in the revised manuscript.

      Author response image 5.

      The quantitative comparison between TSvelo and baseline approaches on the pancreas dataset. In each plot, methods are ranked in descending order of their mean values. Numbers at the bottom indicate the sample size for each metric. Significance is determined using a one-sided Mann–Whitney U test. *****, ****,***, ** and * represent p < 0.00001, 0.00001 ≤ p < 0.0001, 0.0001 ≤ p < 0.001, 0.001 ≤ p < 0.01, and 0.01 ≤ p < 0.05, respectively.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please carefully proofread the text. Some typos:

      (1) Line 110: differentia -> differential.

      (2) Line 280: ".," to be corrected.

      (3) Line 566: optimize -> optimizes.

      We thank the reviewer for carefully proofreading the manuscript and for pointing out these typographical errors. We have corrected the identified typos in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Regarding Major Comment 1 in the Public Review, I contacted BayVel authors, who told me that they'll upload all their scripts here within a few days: https://github.com/elenasabbioni/BayVel_notebooks

      Thank you very much for reaching out to the BayVel authors. We sincerely appreciate the BayVel authors’ efforts to make their scripts and results publicly available through BayVel_notebooks. We believe this is a valuable contribution that will greatly benefit the community.

      We have followed the repository and have now included BayVel in the revised manuscript, with corresponding comparisons added to both the main text and the benchmarking results.

      (2) Page 9 mentions "consistency", "coherence", and "correctness". Instead of these qualitative (and potentially subjective) evaluations, I'd appreciate using quantitative metrics or visual descriptions when differences are visually clear.

      We thank the reviewer for this insightful comment. The terms “velocity consistency,” “in-cluster coherence,” and “cross-boundary correctness” used in our manuscript are not intended as subjective descriptions. They correspond to commonly used evaluation criteria in this field and have been adopted as quantitative metrics in previous studies, such as VeloAE[1] and UniTVelo[2]. We have incorporated the following updated definition into the Methods section.

      (1) Velocity consistency (VCon). We used the scvelo.velocity_confidence() function from scVelo to evaluate velocity consistency, interpreting the results as a measure of how consistent velocities are within neighboring cells. Velocity consistency is especially suitable for evaluating the RNA velocity modeling on single lineage. For each cell , the velocity consistency is calculated as follows:

      Where N (c) represents the neighboring cells of a given cell c v<sub>c</sub> v<sub>c’</sub> denote the low-dimensional velocity vectors of cell cand its neighboring cell c’.

      (2) Cross-boundary direction correctness (CBDir). Cross-boundary direction correctness assesses the accuracy of transitions from a source cluster to a target cluster by examining the boundary cells, and requires ground truth annotations. We directly run the function unitvelo.evaluate() provided in UniTVelo to obtain the Cross-boundary direction correctness. In detail, the CBDir is calculated as follows:

      Where C<sub>A</sub> denotes the set of cells in the target cluster A, and represents the neighboring cells of a given cell c v<sub>c</sub> v<sub>c’</sub> denote the low-dimensional velocity and state vectors of cell cand its neighboring cell c’.

      (3) Within-cluster velocity coherence (ICCoh). Within-cluster velocity coherence measures the coherence of velocities within a single cluster using a cosine similarity score between cell velocities. We applied the function unitvelo.evaluate() provided by UniTVelo to directly compute the within-cluster velocity coherence. Using the same notation as defined above, the CBDir is calculated as follows:

      (1) Qiao, C. & Huang, Y. Representation learning of RNA velocity reveals robust cell transitions. Proceedings of the National Academy of Sciences 118, e2105859118 (2021).

      (2) Gao, M., Qiao, C. & Huang, Y. UniTVelo: temporally unified RNA velocity reinforces single-cell trajectory inference. Nature Communications 13, 6586 (2022).

      (3) At page 3, some objects are not defined after formula (3):

      ReLU finction, and w_gi

      Additionally, parenthesis of ReLU function should be bigger.

      We thank the reviewer for pointing this out. In the revised manuscript, we have explicitly defined the ReLU activation function and clarified that w<sub>gi</sub> represents the regulatory weight of TF i on the target gene g. In addition, we have adjusted the formatting of Eq. (3) by enlarging the parentheses in the ReLU function to improve readability.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The objective of this study was to infer the population dynamics (rates of differentiation, division and loss) and lineage relationships of NK cell subsets during an acute immune response and under homeostatic conditions.

      Strengths:

      A rich dataset and a detailed analysis of a particular class of stochastic models.

      Weaknesses: (relating to initial submission)

      The stochastic models used are quite simple; each population is considered homogeneous with first-order rates of division, death, and differentiation. In Markov process models such as these there is no dependence of cellular behavior on its history of divisions. In recent years models of clonal expansion and diversification, in the settings of T and B cells, have progressed beyond this picture. So I was a little surprised that there was no mention of the literature exploring the role of replicative history in differentiation (e.g. Bresser Nat Imm 2022), nor of the notion of family 'division destinies' (either in division number, or the time spent proliferating, as described by the Cyton and Cyton2 models developed by Hodgkin and collaborators; e.g. Heinzel Nat Imm 2017). The emerging view is that variability in clone (family) size arises may arise predominantly from the signals delivered at activation, which dictate each precursor's subsequent degree of expansion, rather than from the fluctuations deriving from division and death modeled as Poisson processes.

      As you pointed out, the Gerlach and Buchholz Science papers showed evidence for highly skewed distributions of family sizes, and correlations between family size and phenotypic composition. Is it possible that your observed correlations could arise if the propensity for immature CD27+ cells to differentiate into mature CD27- cells increases with division number? The relative frequency of the two populations would then also be impacted by differences in the division rates of each subset - one would need to explore this. But depending on the dependence of the differentiation rate on division number, there may be parameter regimes (and timepoints) at which the more differentiated cells can predominate within large clones even if they divide more slowly than their immature precursors. One might not then be able to rule out the two-state model. I would like to see a discussion or rebuttal of these issues.

      Comments on revisions:

      (1) The authors have put in a lot of effort to address the reviews and have explored alternative models carefully.

      We appreciate the reviewers’ comments.

      (2) In the sections relating to homeostasis and the endogenous response, as far as I can tell you are estimating net growth rates (the k parameters) throughout - this is to be expected if you're working with just cell numbers and no information relating to proliferation. In these sections there are many places where you refer to proliferation rates and death rates when I think you just mean net positive or net negative growth rates. It's important to be precise about this even if the language can get a bit repetitive. (These net rates of growth or loss relate to clonal rather than cellular dynamics, which may be worth explaining). Later, you do use data relating to dead cells, which in principle can be used to get independent measures of death rates, but these data were not used in the fitting.

      We have modified the main text to address the comment.

      (3) There is so much evidence that T and B cell differentiation are often contingent on division that it would be very reasonable to consider it as a possibility for NK cells too. (Differentiation could be asymmetric, as you explored, or simply symmetric with some probability per division). These processes can be cast into simple ODE models but no longer allow you to aggregate division and death rates - so for parameter estimation you need to add measures of proliferation (Ki67 or similar) or death. This may be worth some discussion?

      We have modified the main text (lines 242-245) to address the comment.

      Reviewer #2 (Public review):

      Summary:

      Wethington et al. investigated the mechanistic principles underlying antigen-specific proliferation and memory formation in mouse natural killer (NK) cells following exposure to mouse cytomegalovirus (MCMV), a phenomenon predominantly associated with CD8+ T cells. Using a stochastic modeling approach, the authors aimed to develop a quantitative model of NK cell clonal dynamics during MCMV infection. Starting from a single immature Ly49+CD27+ NK cell, a two-state linear model (with a death variant) explained the negative correlation between clone size at 8 dpi and the CD27+ fraction, but failed to reproduce the first and second moments of CD27+ and CD27− NK cell populations at 8 dpi. To address this limitation, the authors added an intermediate maturation state, yielding a three-stage model (CD27+Ly6C− → CD27−Ly6C− → CD27−Ly6C+) that fits the first and second moments under two constraints: CD27+ NK cells proliferate faster than CD27− NK cells, and clone size is negatively correlated with the CD27+ fraction (upper bound of −0.2). The model predicts high proliferation in the intermediate state and high death in mature CD27−Ly6C+ cells, and it was validated using Adams et al. (2021) NK reporter mice tracking CD27+/− populations after tamoxifen, allowing discrimination between bone marrow-derived and pre-existing peripheral NK cells. To test the prediction that mature CD27− NK cells have a higher death rate, the authors measured Ly49H+ NK cell viability in the mouse spleen at different time points post-MCMV infection. Data confirmed lower viability of mature (CD27−) than immature (CD27+) cells during days 4-8 post-infection, and a model variant supported that higher CD27− death increases their proportion in the dead cell compartment. Altogether, the authors propose a three-stage quantitative model of antigen-specific expansion and maturation of naïve Ly49H+ NK cells with the trajectory CD27+Ly6C− (immature) → CD27−Ly6C− (mature I) → CD27−Ly6C+ (mature II), highlighting high proliferation in the mature I state and increased death in the mature II state.

      Strengths:

      Models explaining correlations and first and second moments, supported by analytical investigations, stochastic simulations, and model selection, identify key processes in antigen-specific NK expansion and maturation. The work distinguishes expansion, contraction, and memory in NK cells from CD8+ T cells and informs NK therapy development.

      Weaknesses (relating to initial submission):

      The conclusions of this paper are largely supported by the available data. However, a comparative analysis with more recent works in the field would be desirable. Clarifications:

      (1) Initial Conditions and Grassmann Data: The Grassmann data is used solely as a constraint, while the simulated values of CD27+/CD27− cells could have been directly fitted to the Grassmann data, which assumes a 1:1 ratio of CD27+/CD27− at t = 0. This would allow an alternative initial condition rather than starting from a single CD27+ cell.

      (2) Correlation Coefficients in the Three-State Model: Although the parameter scan of the three-stage model (Figure 2) demonstrates the potential for negative correlations between colony size and the fraction of CD27+ cells, the calculated correlation coefficients using the fitted parameter values are not shown. Including these would validate that the fitted parameters lie in the negative-correlation regime.

      (3) Viability Dynamics and Adaptive Response: The authors measured the time evolution of CD27+/− dynamics and viability over 30 days post-infection (Figure 4). It would be valuable to test whether the three-state model can reproduce the adaptive response of CD27− cells to MCMV infection, particularly the observed drop in CD27− viability at 5 dpi and its rebound at 8 dpi. Demonstrating this would test whether the model can simultaneously explain viability dynamics and moment dynamics, and would enable sensitivity analysis of CD27− viability with respect to model parameters.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor points:

      (1) line 175 - Here I think you have only ruled out the two state model with no death, and not the two state model in general?

      Edited the sentence to address the comment.

      (2) Figures 2 and 5 - the phenotypes (CD27+ Ly6C-, etc.) should be clearly labeled above each cell type. Fig 1 could be improved in the same way.

      Done.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Kashiwagi et al. undertook a population analysis of dendritic spine nanostructure applied to the objective grouping of 8 mouse models of neuropsychiatric disorders. They report that spine morphology in cultured hippocampal neurons shows a higher similarity among schizophrenia mouse models (compared with autism spectrum disorder (ASD) mouse models), and identify an effect of Ecrg4 (encoding small secretory peptides) on spine dynamics and shape in these models.

      Strengths:

      The study developed a method for objectively comparing spine properties in primary hippocampal neuron cultures from 8 mouse models of psychiatric disorders at the population level using high-resolution structured illumination microscopy (SIM) imaging. This novel technique identified two distinct groups of mouse models according to the population-level spine properties: those with ASD-related gene mutations and those with schizophreniarelated gene mutations. Functional studies, including gene knockdown and overexpression experiments, identified an effect of Ecrg4 on the spine phenotype of the schizophrenia model mice.

      We thank the reviewer for finding our strategy novel and useful for identifying molecules associated with the spine phenotype in schizophrenia-related mouse models.

      Weaknesses:

      The main weakness is that the study is wholly in vitro, using cultured hippocampal neurons. The authors present this as an advantage, however, arguing that spine morphology as measured in a reduced culture system can demonstrate direct effects of gene mutations on neuronal phenotypes in the absence of indirect influences from non-neuronal cells or specific environments.

      We appreciate this reviewer's concern about the limitation of cultured hippocampal neurons in extracting disease-related spine phenotypes. While we fully recognize this limitation, we consider that this in vitro system has several advantages that contribute to translational research on mental disorders.

      First, our culture system has been shown to support the development of spine morphology similar to that of the hippocampal CA1 excitatory synapse in vivo. High-resolution imaging techniques confirmed that the in vitro spine structure was highly preserved compared with in vivo preparations (Kashiwagi et al., Nature Communications, 2019). The present study used the same culture system and SIM imaging. Therefore, the difference we detected in samples derived from disease models is likely to reflect impairment of molecular mechanisms underlying native structural development in vivo.

      Second, super-resolution imaging of thousands of spines in tissue preparations under precisely controlled conditions cannot be practically applied using currently available techniques. The advantage of our imaging and analytical pipeline is its reproducibility, which enabled us to compare the spine population data from eight different mouse models without normalization.

      Third, a reduced culture system can demonstrate the direct effects of gene mutations on synapse phenotypes, independent of environmental influences. This property is highly advantageous for screening chemical compounds that rescue spine phenotypes. Neuronal firing patterns and receptor functions can also be easily controlled in a culture system. The difference in spine structure between ASD- and schizophrenia-related mouse models is valuable information to establish a drug screening system.

      Fourth, establishing an in vitro system for evaluating synapse phenotypes could reduce the need for animal experiments. Researchers should be aware of the 3Rs principles. In the future, combined with differentiation techniques for human iPS cells, our in vitro approach will enable the evaluation of disease-related spine phenotypes without the need for animal experiments. The effort to establish a reliable culture system should not be eliminated.

      We modified our text to have a balanced discussion on both advantages and disadvantages of the in vitro culture system in the study of mental disorder mouse models, as follows:

      "Finally, while the spine phenotype identified in the human postmortem brain undoubtedly resulted from complex interactions among genetic background, environmental influences, and regulation by non-neuronal cells, data from pure neuronal cultures are more likely to reflect the direct effects of schizophrenia-related gene mutations on synaptic functions. This property may be advantageous for identifying synaptic molecules that regulate synapse phenotypes in schizophrenia-related mouse models. However, the phenotype observed in the culture system requires confirmation using in vivo experiments of mouse models or human tissue samples. Efficient in vitro screening combined with reliable in vivo evaluation of synapses will facilitate translational research on mental disorders."

      Another weakness is that CaMKIIαK42R/K42R mutant mice are presented as a schizophrenia model, the authors justifying this by saying that "CaMKII-related signaling pathway disruption has been implicated in the working memory deficits found in schizophrenia patients". Since mutations in CAMK2A cause autosomal dominant intellectual developmental disorder-53 (OMIM 617798) and autosomal recessive intellectual developmental disorder-63 (OMIM 618095), and mice carrying the CAMK2A E183V mutation exhibit ASD-related synaptic and behavioral phenotypes (PMID: 28130356), I think it's stretching credibility to refer to the CaMKIIαK42R/K42R mice as a schizophrenia model.

      We agree with this reviewer that CAMK2A mutations in humans are linked to multiple mental disorders, including developmental disorders, ASD, and schizophrenia. Association of gene mutations with the categories of mental disorders is not straightforward, as the symptoms of these disorders also overlap with each other. For the CaMKIIα K42R/K42R mutant, we considered the following points in its characterization as a model of mental disorder. Analysis of CaMKIIα +/- mice in Dr. Tsuyoshi Miyakawa's lab has provided evidence for the reduced CaMKIIα in schizophrenia-related phenotypes (Yamasaki et al., Mol Brain 2008; Frankland et al., Mol Brain Editorial 2008). It is also known that the CaMKIIα R8H mutation in the kinase domain is linked to schizophrenia (Brown et al., 2021). Both CaMKIIα R8H and CaMKIIα K42R mutations are located in the N-terminal domain and eliminate kinase activity. On the other hand, the representative CaMKIIα E183V mutation identified in ASD patients exhibits unique characteristics, including reduced kinase activity, decreased protein stability and expression levels, and disrupted interactions with ASD-associated proteins such as Shank3 (Stephenson et al., 2017). Importantly, reduced dendritic spines in neurons expressing CaMKIIα E183V is a property opposite to that of the CaMKIIα K42R/K42R mutant, which showed increased spine density (Koeberle et al. 2017).

      References related to this discussion.

      (1) Yamasaki et al., Mol Brain. 2008 DOI: 10.1186/1756-6606-1-6

      (2) Frankland et al. Mol Brain. 2008 DOI: 10.1186/1756-6606-1-5

      (3) Stephenson et al., J Neurosci. 2017 DOI: 10.1523/JNEUROSCI.2068-16.2017

      (4) Koeberle et al. Sci Rep. 2017 DOI: 10.1038/s41598-017-13728-y

      (5) Brown et al., iScience. 2021 DOI: 10.1016/j.isci.2021.103184

      We fully agree with the reviewer that different CAMK2A mutations likely cause distinct phenotypes observed in the broad spectrum of mental disorders. In the revised manuscript, we include a discussion of the relevant literature to categorize this mouse model appropriately.

      "CaMKII-related signaling pathway disruption has been implicated in the working memory deficits found in schizophrenia patients [45,46]. CAMK2A mutations in humans are linked to multiple mental disorders, including developmental disorders, ASD, and schizophrenia [47]. The K42R mutation of CAMK2A does not correspond to any known human genetic variant, but the CAMK2A R8H mutation is linked to schizophrenia [48]. Both R8H and K42R mutations in the N-terminal domain of CaMKIIα eliminate kinase activity; these mutations may have a similar impact on human mental disorders."

      Although the manuscript is largely well written, there are some instances of ambiguous/unspecific language. This extends to the title (Decoding Spine Nanostructure in Mental Disorders Reveals a Schizophrenia-1 Linked Role for Ecrg4), which gives no indication that the work was in vitro on cultured neurons derived from mouse models.

      We appreciate the reviewer for pointing out the lack of information about the experimental system in the title of this manuscript. According to the suggestion of the reviewer, we modified the title as "Decoding spine nanostructure in cultured neurons derived from mouse models of mental disorder reveals a schizophrenia-linked role for Ecrg4".

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique that they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes, which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. However, I do have a number of major concerns.

      We thank the reviewer for evaluating our findings as potentially very interesting and useful.

      (1) The main finding of spine nanostructure changes is done by carrying out a PCA on various structural parameters, creating spine density plots across PC1 and PC2, and then subtracting the WT density plot from the mutant. Then, spines in the areas with obvious differences only are analyzed, from which they derive the finding that, for example, spine sizes are smaller. However, this seems a circular approach. It is like first identifying where there might be a difference in the data, then only analyzing that part of the data. I welcome input from a statistician, but to me, this is at best unconventional and potentially misleading. I assume the overall means are not different (although this should be included), but could they look at the distribution of sizes and see if these are shifted?

      We appreciate the reviewer's concern regarding our analysis of spine population data. The intention of pre-selecting the areas showing differences between wild-type and mutant was to make a direct comparison between two subareas (one is enriched with wild-type spines and the other is enriched with mutant spines) and clarify that the spines of schizophreniarelated mouse models were smaller than wild-type spines. Conventional methods of comparing the total spine population using simple size parameters are not useful for this purpose, as shown in Supplementary Figure 2.

      To clarify the reviewer's concern, we revised the analysis of the spine population data for both Figure 3 and Figure 8.

      Figure 3: We first divided the feature space projected onto PC1 and PC2 into four areas with distinct structural properties: (1) small and short, (2) small and long, (3) large and short, and (4) large and long. Next, we calculated the normalized spine counts in the four areas for both wild-type and mutant spines and obtained the relative ratio (mutant/wild-type) for each area. As we performed three independent SIM imaging experiments (in one, we imaged both wild type and mutant culture dishes prepared from the same pregnant mouse), there are three independent datasets from 8 mouse models.

      We found that the spine ratio (mutant/wild-type) only in area 2 (small and long spines) differed significantly between genotypes. This result is shown in Fig. 3 and explained in the text. The spine ratios in areas 1 and 3 did not show a clear relationship to the genotypes, while the ratio in area 4 showed the opposite trend to that in area 2. The opposite trend between areas 2 and 4 indicates enrichment of both small and long spines in schizophrenia-related mouse models, consistent with our previous analysis.

      Figure 8: In this analysis, we aimed to evaluate the rescue effect of Ecrg4 shRNA relative to that of control shRNA. If Ecrg4 shRNA is effective, the spine population enriched in the control shRNA condition should be reduced in the Ecrg4 shRNA condition. To confirm this point in the revised manuscript, we first defined areas in the projected PC1-PC2 plane showing either enrichment or depletion of spines in the control shRNA condition (spine numbers increasing or decreasing by more than 3 × SD). We next measured the difference in spine numbers between the control and Ecrg4 shRNA conditions in either enriched or depleted areas. The expectation is that Ecrg4 shRNA treatment reduces the extent of both enrichment and depletion. The effect was significant in both the 22qdel and Setd1a mouse models, as indicated by permutation tests. This analysis was explained in the revised manuscript.

      (2) Despite extracting 64 parameters describing spine structure, only 5 of these seemed to be used for the PCA. It should be possible to use all parameters and show the same results. More information on PC1 and PC2 would be helpful, given that the rest of the paper is based on these - what features are they related to?

      We thank the reviewer for the advice on providing the rationale for parameter selection in PCA. We divided spines into 160-nm segments along their long axis, and the spine segments were used to calculate the 64 parameters, which include volume of each spine segment (20 segments), convex hull volume of each spine segment (20 segments), and convex hull ratio of each spine segment (20 segments). As most spines are shorter than 0.16 × 20 =3.2 μm, these segment-related parameters contain a large fraction of zero values, which affect the proper calculation of principal components. Therefore, we selected two parameters that reflect the principal structural features (length and volume), together with three other parameters that were mutually independent and also independent from the first two parameters (pairwise correlation coefficients < 0.3). These selection criteria were described in the original manuscript. We also confirmed that PCA using all 64 parameters yields a cross correlation map similar to that shown in Fig. 2B.

      Author response image 1.

      We provided additional information in the Materials and Methods section of the revised manuscript.

      As described previously, the pattern of four areas with distinct spine structures (1. small and short, 2. small and long, 3. large and short, 4. large and long) supports the idea that the PC1PC2 plane reflects the relationship between spine volume and length (Fig. 3A and B).

      These specific features could then be analyzed in the full dataset, without doing the cherry picking above.

      We provided the dataset for the relative enrichment of spine counts across four areas of the PC1-PC2 plane in Fig. 3A and B. This analysis provides a comprehensive view of spine population properties related to spine volume and length, without relying on a pre-set region of interest.

      It would also be helpful to demonstrate whether PC1 and 2 differ across groups - for example, the authors could break their WT data into 2 subsets and repeat the analysis.

      We noticed differences in the pattern of spine distribution across the PC1-PC2 planes in each experiment. The subtraction of the distributional data between wild-type and mutant samples effectively cancels out such differences. In general, the difference between two wild-type samples is smaller than that between wild-type and mutant samples, as shown in Author response image 2.

      Author response image 2.

      We added a description of variation across groups to the revised manuscript.

      (3) Throughout the paper, the 'n' used for statistical analysis is often spine, which is not appropriate. At a minimum, cell should be used, but ideally a nested mixed model, which would take into account factors like cell, culture, and animal, would be preferable. Also, all of these factors should be listed, with sufficient independent cultures.

      We agree that nested mixed models are more appropriate for evaluating genotype effects in most of our datasets. We confirm that the results of statistical analysis using nested mixed models were consistent with our previous conclusions in most cases.

      Figure 3: We performed three independent primary cultures of embryonic hippocampal tissue with genotypes of both wild-type and mutant from the same pregnant mice for each mouse model. In our new Figure 3, each data point represents an independent culture experiment, and group comparisons were performed using one-way ANOVA followed by Tukey's post hoc test. In this analysis, statistical analysis using neurons as units of 'n' is not possible, as the number of spines measured from a single neuron is insufficient to generate the density map shown in Figure 3. The statistical analysis was described in the revised text. The details of experimental conditions related to Figure 3 are provided in Supplementary Table 1.

      Figure 5A-C: We analyzed spine turnover rate using a linear mixed-effects model with genotype as a fixed effect and plate, cell, and dendrite as nested random effects. In both 22q deletion model and Setd1a model, there were significant effects of genotype (F(1,25) = 5.79, p = 0.024 for 22q deletion model and F(1,22) = 7.33, p = 0.013 for Setd1a model). In contrast, Nlgn3 mutant neurons did not show a significant difference (F(1,14) = 1.35, p = 0.26). This analysis was described in the revised text.

      Figure 5D-F: Spine lifetime was analyzed using a linear mixed-effects model accounting for the hierarchical structure of the data (spines nested within dendrites, cells, and culture plates). The analysis revealed a significant effect of genotype in both 22q deletion mutant and Setd1a mutant (22qdel mutant; F(1,336) =5.33, p=0.022, Setd1a mutant; F(1,282)=6.38, p=0.012 ). The neurons of both mutants exhibited significantly longer spine lifetimes compared with wild-type neurons (22qdel mutant; ratio = 1.28, 95% CI 1.04–1.58, Setd1a mutant; ratio = 1.35, 95% CI 1.07–1.70). In contrast, Nlg3 mutation did not significantly alter spine lifetime (ratio = 0.86, 95% CI 0.61–1.22; F(1,220)=0.69, p=0.41). This analysis was described in the revised text.

      Figure 5G-I: Spine volume trajectories were analyzed using linear mixed-effects models incorporating nested random effects (spine/dendrite/cell/culture plate) to account for the hierarchical structure of the data. In the 22q deletion model, newly formed spines were significantly smaller than those in wild-type neurons (genotype effect: p < 0.001). The spines in Setd1a mutant neurons also displayed significantly smaller volume than those in wild-type neurons (p < 10<sup>-7</sup>). There were also differences in the temporal profiles of spine growth in these two mutants (p < 0.001). In contrast, newly formed spines in the Nlgn3 mutant neurons were significantly larger than those in wild-type neurons (p < 10<sup>-4</sup>) with preserved time-course of spine growth. This analysis was described in the revised text.

      Figure 5J-L: Similar analyses using linear mixed-effects models incorporating nested random effects (spine within dendrite within cell within culture plate) identified significantly smaller initial spine size in the 22q deletion model (p < 10<sup>⁻6</sup>), while no significant differences in the initial spine volume were found for Setd1a mutants. The temporal trajectories of spine shrinkage before their loss were also not significantly altered in both 22qdel and Setd1a mutants. The Nlg3 mutant showed a significantly different time-course of spine shrinkage (p < 0.05), while the initial spine size was not altered. This analysis was described in the revised text.

      Figure 7A overexpression dataset: We analyzed plate-averaged lifetime values using a linear mixed-effects model with treatment as a fixed effect. There exists a significant main effect of treatment (F(3,8) = 4.59, p = 0.038), with post hoc examination showing a significant increase in lifetime by Ecrg4 overexpression (β = 0.49 ± 0.16 SE, t(8) = 3.16, p = 0.013). Figure 7A shRNA dataset: We also applied a linear mixed-effects model for plate-averaged lifetime values with treatment as a fixed effect. The analysis revealed no significant effect of treatment (F(2,6) = 0.29, p = 0.76).

      The analyses of overexpression and shRNA datasets were described in the revised text.

      Figure 8: As in Figure 3, we performed three independent primary cultures of embryonic hippocampal tissue with genotypes of both wild-type and mutant from the same pregnant mice for each mouse model. The culture plates were transfected with either a control shRNA or an Ecrg4 shRNA construct. Each data point represents an independent culture experiment, and the effect of Ecrg4 shRNA relative to that of control shRNA was evaluated using a permutation test. The data analysis was described in the revised text. The details of experimental conditions related to Figure 8 are provided in Supplementary Table 1.

      (4) The authors should confirm that all mutants are also on the C57BL/6J background, and clarify whether control cultures are from littermates (this would be important). Also, are control versus mutant cultures done simultaneously? There can be significant batch effects with cultures.

      The mutant mice we used in this study are on C57BL/6J or C57BL/6N background. It is known that C57BL/6J or C57BL/6N mice exhibit distinct phenotypes across a range of physiological, biochemical, and behavioral systems. However, it is less likely that our analysis is affected by differences between C57BL/6J and C57BL/6N, as we compared wild-type and mutant littermates on the same genetic background. This experimental design can also reduce the batch effects with different culture preparations. This point was described in the revised text.

      (5) The spine analysis uses cultures from 18-22 DIV - this is quite a large range. It would be worth checking whether age is a confounder or correlated with any parameters / principal components.

      We described in the method sections that culture samples were processed for imaging at 18-22 DIV. However, all the SIM imaging experiments for eight mutant mouse models were performed on samples fixed at DIV 19. The wide range of imaging experiments (DIV 18-22) includes test samples we used to optimize imaging conditions. In the revised manuscript, we specified the timing of SIM imaging.

      (6) The computational modelling is interesting, but again, I am concerned about some circularity. Parameter optimization was used to identify the best fit model that replicated the spine turnover rates, so it is somewhat circular to say that this matched the observations when one of these is the turnover rate.

      We appreciate the reviewer's comment on some circularity of the argument. We agree that the turnover rate is already incorporated into the simulation model and is not an appropriate criterion for the evaluation. We modified the text accordingly.

      It is more convincing for spine density and size, but why not go back and test whether parameter differences are actually seen - for example, it would be possible to extract the probability of nascent spine loss, etc.

      We thank the reviewer for giving this important suggestion. The probability of nascent spine loss is an important parameter, and we initially attempted to estimate it from the original data set. However, the upper limit of our time-lapse imaging is 24 h, which is insufficient to distinguish stable and nascent spines clearly. The difficulty of extracting all the necessary parameters for spine remodeling is our motivation for starting this computational modelling.

      More compelling would be to repeat the experiments and see if the model still fits the data. In the interpretation (line 314-318) it is stated that '... reduced spine maturation rate can account for the three key properties of schizophrenia-related spines...', which is interesting if true, but it has just been stated that the probability of spine destabilization is also higher in mutants (line 303) - the authors should test whether if the latter is set to be the same as controls whether all the findings are replicated.

      As suggested by the reviewer, we set the probability of spine destabilization equal across wild-type and mutant models and repeated the simulations. The results indicate that this modification has small effects on spine density (0.61 vs 0.62), spine turnover rate (0.22 vs 0.21), fraction of small spines (0.21 vs 0.20), and mean spine size (0.37 vs 0.36). We described this point in the revised manuscript.

      (7) No validation for overexpression or knockdown is shown, although it is mentioned in the methods - please include.

      As suggested by the reviewer, we validated overexpression and knockdown. The results are summarized in Supplementary Figure 8.

      Supplementary Figure 8A-C shows the immunocytochemistry of anti-Ecrg4, anti-Cip4, and anti-NPAS4 for the confirmation of overexpression of these molecules.

      Supplementary Figure 8D-E shows the confirmation of the appropriate size of exogenously expressed Ecrg4, Cip4, and NPAS4 by immunoblotting. (previous Supplementary Figure 10F is now Supplementary Figure 8E).

      Supplementary Figure 8F-H indicates the efficient knockdown of exogenously expressed Met-GFP, ARHGAP15-GFP, and Ecrg4-HA by respective shRNA constructs in COS-7 cells. (previous Supplementary Figure 10G is now Supplementary Figure 8H)

      Also, for the knockdown, a scrambled shRNA control would be preferable.

      We used Stealth RNAi Negative Control Duplexes (Invitrogen) as the shRNA control in this study. To confirm that this RNAi sequence does not affect spine turnover, we performed timelapse imaging of neurons transfected with GFP alone or with GFP and the Stealth RNAi Negative Control. No detectable change in spine turnover was observed (Supplementary Figure 8I), indicating that this RNAi control sequence is suitable for our study.

      (8) The finding regarding ecgr4 is interesting, but showing that some ecgr4 is expressed at boutons and spines and some in DCVs is not enough evidence to suggest that actively involved in the regulation of synapse formation and maturation (line 356).

      To reveal the active roles of Ecrg4 in spine regulation, we exogenously applied a synthetic Ecrg4 peptide to wild-type neurons and monitored both spine density and turnover rate after Ecrg4 application. The Ecrg4 application increased the spine turnover rate, whereas samples treated with the scrambled peptide did not. This result supports the active role of Ecrg4 in regulating spine turnover. The data were added as Supplementary Figures 9F and G.

      (9) The same caveats that apply to the analysis also apply to the ecgr4 rescue. In addition, while for 22q the control shRNA mutant vs WT looks vaguely like Figure 2, setd1a looks completely different.

      We thank the reviewer for pointing out the apparent difference in the pattern of spine population data between Figure 2 and Figure 8. We performed SIM analysis using DiI-labeled neurons in Figure 2, whereas the data in Figure 8 are derived from GFP-expressing neurons. The images of cell-surface labeling and cytoplasmic labeling cannot be analyzed in the same way, as it is necessary to adjust parameters in SIM image processing and PCA-based dimensional reduction. Consequently, the distribution of the spine population projected onto the PC1-PC2 plane differs between DiI-labeled neurons and GFP-expressing neurons. To facilitate the comparison of PCA analysis applied to GFP-expressing neurons, we replaced the weight matrix for GFP-expressing neurons with that previously calculated for the DiIlabeled neurons. This adjustment increased the similarity of the data distributions shown in Figures 2 and 8. The explanation for the different patterns in the spine population map between Figure 2 and Figure 8 was added to the revised text. The related explanation for the data processing was described in the Materials and Methods.

      And if rescued, surely shRNA in the mutant should now resemble control in WT, so there shouldn't be big differences, but in fact, there are just as many differences as comparing mutant vs wild-type? Plus, for spine features, they only compare mutant rescue with mutant control, but this is not ideal - something more like a 2-way ANOVA is really needed. Maybe input from a statistician might be useful here?

      We appreciate the reviewer's important comment and agree that the analytical approach used in the original manuscript was not optimal. We therefore revised our analysis to examine whether the difference observed between wild-type and mutant neurons was reduced by suppression of Ecrg4 expression.

      To this end, we first identified two regions in the PC1–PC2 plane where mutant spines were either enriched or depleted relative to wild-type neurons (Areas A and B). We then counted the number of spines located in Areas A and B in control shRNA-treated mutant neurons (normalized spine counts XA and XB). Next, we quantified spine counts in the same areas using data from Ecrg4-suppressed mutant neurons (normalized spine counts YA and YB). If XA > YA and XB < YB, suppression of Ecrg4 would indicate a shift toward rescue of the phenotype observed in control shRNA-treated mutant neurons. Indeed, the datasets were consistent with this shift in relative spine counts.

      To determine whether these differences exceeded those expected from random variation in spine counts, we performed a permutation test. Specifically, spine identities were randomly shuffled between the two conditions while preserving the total number of spines in each dataset. The observed differences were then compared with the distribution obtained from the permuted datasets to assess statistical significance.

      We found that all three culture replicates showed statistical significance in both areas A and B for both the 22qdel and Setd1a mutations. This analysis is described in the Result section.

      (10) Although this is a study entirely focused on spine changes in mouse models for Sz, there is no discussion (or citation) of the various studies that have examined this in the literature. For example, for Setd1a, smaller spines or reduced spine densities have been described in various papers (Mukai et al, Neuron 2019; Chen et al, Sci Adv 2022; Nagahama et al, Cell Rep 2020).

      We appreciate the reviewer's suggestion to include a discussion of schizophrenia-related mouse models. We added more information related to the Setd1a mouse model to the Discussion section.

      "Population-level spine properties were more homogeneous in schizophrenia models (those with gene mutations implicated in schizophrenia) than in the other 4 models studied, in part due to a shared tendency for smaller spines. This observation is consistent with previous studies on Setd1a mutant mice, which showed reduced spine width, decreased mushroomtype spines, and lower spine density in the prefrontal cortex [43,56,57]. In contrast to these findings, several previous studies reported reduced numbers of small spines in the postmortem cortical tissues of schizophrenia patients [22,58]. "

      (11) There is a conceptual problem with the models if being used to differentiate autism risk from Sz risk genes. It is difficult to find good mouse models for Sz, so the choice of 22q11.2del and Setd1a haploinsufficiency is completely reasonable. However, these are both syndromic. 22qdel syndrome involves multiple issues, including hearing loss, delayed development, and learning disabilities, and is associated with autism (20% have autism, as compared to 25% with Sz). Similarly, Setd1a is also strongly associated with autism as well as Sz (and also involves global developmental delay and intellectual disability). While I think this is still the best we can do, and it is reasonable to say that these models show biased risk for these developmental disorders, it definitely can't be used as an explanation for the higher variability seen in the autism risk models.

      We appreciate the reviewer's suggestion for more careful consideration of the interpretation of phenotypes in mouse models, with regard to their relation to clinical phenotypes in human patients. According to the suggestion of the reviewer, we modified the relevant text as follows:

      "The nanoscale features of dendritic spines in ASD-associated mouse models were more variable than those in schizophrenia-associated mouse models. This difference may be related to the broader clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. The four ASD-associated mouse models examined in this study, Nlgn3<sup>R451C/(y or R451C) , Syngap1<sup>+/-</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, may represent subgroups with different levels of hippocampal dysfunction. Among the four ASD-associated mouse models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those of the schizophrenia models. To understand this similarity, further analysis of neural circuit changes in both ASD- and schizophrenia-associated mouse models will be necessary. Analysis of the relationships between rare genetic variants and synapse phenotypes in mouse models may contribute to their eventual categorization. This information should be useful to understand the underlying mechanisms of the broader clinical spectrum of ASD."

      (12) I am not convinced that using dissociated cultures is 'more likely to reflect the direct impact of schizophrenia-related gene mutations on synaptic properties' - first, cultures do have non-neuronal cells, although here glial proliferation was arrested at 2 days, glia will be present with the protocol used (or if not, this needs demonstrating).

      In our culture system, the density of non-neuronal cells is low, and most neurons are not in direct contact with non-neuronal cells. We reported this method in Nat. Neurosci. 1999, where we utilized this culture system to visualize GFP-tagged PSD-95 in neurons using recombinant adenovirus. Because recombinant adenovirus shows higher infection efficiency in glial cells, it was essential for us to establish a culture condition that isolates neurons from glial cells.

      Second, activity levels will affect spine size, and activity patterns are very abnormal in dissociated cultures, so it is very possible that spine changes may not translate into in vivo scenarios. Overall, it is a weakness that the dissociated culture system has been used, which is not to say that it is not useful, and from a technical and practical perspective, there are good justifications.

      We appreciate the reviewer's comment on the advantages and disadvantages of using an in vitro culture system. This comment aligns with the first reviewer's. We modified our text to have a balanced discussion on the role of the in vitro culture system in the study of mental disorder mouse models as follows:

      "Finally, while the spine phenotype identified in the human postmortem brain undoubtedly resulted from complex interactions among genetic background, environmental influences, and regulation by non-neuronal cells, data from pure neuronal cultures are more likely to reflect the direct effects of schizophrenia-related gene mutations on synaptic functions. This property may be advantageous for identifying synaptic molecules that regulate synapse phenotypes in schizophrenia-related mouse models. However, the phenotype observed in the culture system requires confirmation using in vivo experiments of mouse models or human tissue samples. Efficient in vitro screening combined with reliable in vivo evaluation of synapses will facilitate translational research on mental disorders."

      (13) As a minor comment, the spine time-lapse imaging is a strength of the paper. I wonder about the interpretation of Figure 5. For example, the results in Figure 5G and J look as if they may be more that the spines grow to a smaller size and start from a smaller size, rather than necessarily the rate of growth.

      We thank the reviewer for the insightful comment. In the revised manuscript, we analyze the time-lapse data using linear mixed-effects models incorporating nested random effects (spine/dendrite/cell/culture plate). This analysis suggested the difference in the initial size of spines. This point is described in the revised manuscript as follows:

      "Schizophrenia-associated mouse models showed higher similarity in spine morphology, driven by reduced size and growth of nascent spines."

      "We further compared the initial increase in spine volume between genotypes (Figure 5G-I). Linear mixed-effects models incorporating nested random effects revealed significantly smaller initial spine volumes in both 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> models (genotype effect: p < 0.001 for 22q11.2<sup>del/+</sup> and p < 10<sup>-7</sup> for Setd1a<sup>+/-</sup>). The spines in both mutants also displayed a significant reduction in spine volume increase (p < 0.001). In contrast, newly formed spines in the Nlgn3<sup>R451C/(y or R451C)</sup> neurons were significantly larger than those in wild-type neurons (p < 10<sup>-4</sup>) with preserved time-course of spine growth.”

      We tested whether the initial size difference in spines can be incorporated into the computational simulation. However, due to the large variability in the initial spine size, it was difficult to perform parameter optimization in the model with additional factors. Therefore, we did not further pursue this possibility in this revision. This point is described in the revised text.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The manuscript would be strengthened if the following issues were adequately addressed:

      (1) It would be helpful to know more about the in/ex vivo dendritic spine phenotype of the mouse models of neuropsychiatric disorders, to allow readers to judge whether and how the in vitro spine phenotype in hippocampal neuronal cultures overlaps with/replicates the spine phenotype within the mouse brain.

      We appreciate this comment, but our currently available data is insufficient to specify the difference between in vitro and in vivo spine phenotypes. Our previous study, published in Nature. Comm. (2019), provided data showing that the overall distribution of spine size is similar between in vivo and in vitro conditions in the mouse hippocampus.

      (2) Although the manuscript is largely well written, there are instances of ambiguous language, particularly when describing the spine phenotypes. For example, we are told that "ASD mouse models showed a tendency of decreasing spine subpopulation with small volumes." This description and other examples should be expressed more clearly.

      Following the reviewer's suggestions, we revised the text to improve clarity. We modified the sentence "ASD mouse models showed a tendency of decreasing spine subpopulation with small volumes" to "ASD-related mouse models showed an opposite spine phenotype."To avoid possible confusion for readers, we have revised several sentences in the text to clarify the intended meaning.

      Also, I question whether the word "decoding", meaning to convert (a coded message) into intelligible language, is the most appropriate for the title and abstract.

      The original meaning of the word "decoding" is the conversion of a coded message into an intelligible form; however, in this study, we use the term in a broader sense, referring to the extraction of latent population-level properties of dendritic spines from multidimensional structural parameters. We believe this usage is consistent with its common use in neuroscience and systems biology, where "decoding" often refers to inferring underlying biological states or information from complex datasets.

      (3) The authors should reconsider whether CaMKIIαK42R/K42R mice should be described as a schizophrenia model, when mutations in CAMK2A are known to cause autosomal dominant intellectual developmental disorder-53 (OMIM 617798) and autosomal recessive intellectual developmental disorder-63 (OMIM 618095), and mice carrying the CAMK2A E183V mutation exhibit ASD-related synaptic and behavioral phenotypes (PMID: 28130356).

      We provided a detailed answer to this question in the previous part of the rebuttal.

      (4) The title doesn't adequately summarise the contents of the manuscript. It should mention mice/mouse models and cultured neurons.

      We also responded to this request in the previous part of the rebuttal.

      Reviewer #2 (Recommendations for the authors):

      (1) Please provide a supplementary table with all DEGs. Also, DEGs are listed if present in 'more than 2' models - does this mean they had to be in 3 or more? Please clarify.

      According to the reviewer's suggestion, we added data on DEGs shared by >2 mouse models in Supplementary Figure 7. We also added Supplementary Tables 2 and 3 for all DEGs. The phrase "in more than 2 models" means "in 3 or 4 models".

      (2) There are several references to 'schizophrenia mouse models' - it is worth rephrasing this to make clear that these are not mice with schizophrenia.

      We replaced the expression "schizophrenia (or ASD) mouse models" with "schizophrenia (or ASD)-associated mouse models" or similar appropriate wording throughout the manuscript.

      (3) Line 66: 'a recent...' - 2014 is not really recent.

      We removed the word "recent" from the sentence.

      (4) Figure S1: The legend says A-D, but they are not on the figure. Also, make clear whether this data is only WT data - it seems to be from disorder models, with 4 colors for each model - please clarify.

      We changed the sentence from "shown as A to D" to "shown as A to C". The datasets in Supplementary Figure 1 are wild-type only. Each graph uses four colors to represent wildtype data from four imaging datasets obtained from different mouse models. Graphs A to C correspond to spine length, surface area, and volume, respectively.

      (5) Methods, line 680-4: More detail here would be helpful.

      We added more explanation for the generation of subtraction maps.

      (6) Line 193: Make it clear this is hippocampal in the main text.

      We added "cultures of embryonic hippocampi" to the text.

      (7) Figure 5, D-F: Make clear that these are transient spines (as per main text)

      We added "Lifetimes of transient spines" to both the main text and figure legend.

      (8) Figure 6B: More detail is needed; no idea what this is - no axis label. D - also not clear what numbers on the y-axis mean. E - color scale??

      We added details to the figure legend, the axis labels for Figures 6B and 6D, and the color scale for Figure 6E.

      (9) Supplementary Figure 9 - not clear what matrices are actually showing, nor what the scale refers to - is this the number of shared DEGs? If so, please make it clearer.

      The matrices show the shared DEG numbers, as shown in their titles. The scale indicates DEG numbers. We added the explanation of the color code to the figure legend.

      (10) Please make clear in the main text that ecgr4 affected the turnover rate. It would be good to measure other parameters as well.

      We added the phrase "a significant increase in spine turnover rate by Ecrg4 overexpression" to the main text.

      (11) Figure 7: Suggest to label C on images as well, so obvious which is GFP/anti-HA overlay (and respective colors) and which is anti-HA staining.

      We added the labels with respective colors to Figure 7.

      (12) Ecgr4 is a precursor protein that is cleaved to produce several hormone-like peptides. Where is the HA tag - so which cleavage products will it label? Any antibodies that work in immunocytochem?

      HA tag was attached to the C-terminal domain. We predict that anti-HA binds to four cleavage products (the full-length Ecrg4, Augurin, Argilin, and Δ16). Among several commercially available antibodies, only the SIGMA product could detect cells expressing Ecrg4-HA by immunocytochemistry.

      (13) Supplementary Figure 10: Synaptosome would be a good addition.

      We isolated the fraction of synaptosomes using Syn-PER™ Synaptic Protein Extraction Reagent in Supplementary Figure 9A. We added this explanation to the Materials and Methods section.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      Strengths of this paper include the important question addressed and the elegant and innovative combination of methods, which led to clear insights into the sensory biology of self-righting, and that will be useful for others in the field. This is a substantial contribution to understanding how animals correct their body position. The manuscript is very clearly written and couched in interesting biology.

      Limitations:

      (1.1) The interpretation of functional experiments is complicated by the proposed excitatory and inhibitory roles of dorsal and ventral sensory neuron activity, respectively. So, while silencing of an excitatory (dorsal) element might slow righting, silencing of inputs that inhibit righting could speed the behavior. Silencing them together, as is done here, could nullify or mask important D-V-specific roles. Selective manipulation of cells along the D-V axis could help address this caveat.

      We highly appreciate the thoughtful comments by Rev1 pointing out the relative simplicity of our current inferences regarding the role of dorsal vs. ventral substrate contact, and agree with the suggestion that cells along the DV axis could have diverse roles in their contribution to self-righting. In this context, we wish to point out two aspects, one theoretical and one practical. Regarding theory, our view is that this may not be a simple case of “excitation vs. inhibition”, but rather one in which the coordinated and dynamic activity of distributed sensory neurons promotes differential action selection in alignment with environmental conditions – a framework that could involve many different behaviours with a still uncertain level of granularity (e.g., is self-righting different if the larva is rotated to 160º instead of exactly 180º?). Regarding the practical aspect, while this area represents a fascinating point for future investigation, it is currently limited by technological development, particularly in the context of this study where a relatively low-cost implementation has been used to probe the AP axis. Investigation of the DV axis would require further technological development, since optogenetic light would need to be precisely delivered from the side rather than from underneath, with a greater degree of resolution compared to the AP axis given the much smaller width of the larva (~120-140µm) relative to its length (~550-600µm). Therefore, whilst we appreciate these comments and suggestion, we believe this line of experiments is ideal for a follow-up investigation, rather than being implemented in the current study.

      (1.2) Prior studies from the authors implicated daIV neurons in the righting response. One of the main advances of the current manuscript is the clever demonstration of region-specific roles of sensory input. However, this is only confirmed with a general md driver, 190(2)80, and not with the subsetspecific Gal4, so it is not clear if daIV sensory neurons are also acting in a regionally-specific manner along the A-P axis.

      To address this interesting and important comment by Rev1 we have carried out a new experiment using an alternative driver to 109(2)80-Gal4 and testing the impact of these manipulations on larval behaviour. The revised version of our MS includes a new figure Supp Fig S3 which shows self-righting times when using the ppk-Gal4 driver with the opto-axial technique. As observed with the 109(2)80-Gal4 driver, self-righting was delayed in anterior but not posterior inhibition conditions, suggesting the daIV neurons act in a region-specific manner to trigger postural control behaviour.

      We have also conducted a head casting analysis in the ppk domain; in another new figure, Supp Fig S7, we also show that head casting behaviour is also increased in the same manner as with the 109(2)80-Gal4 driver.

      These new panels and figures are cited within the sub sections entitled “Optogenetic inhibition of anterior but not posterior multidendritic neurons delays self-righting” and “Inhibition of anterior multidendritic neurons is associated with increased head casting during self-righting”, on pages 25 and 28, respectively. We are grateful to Rev1 for this suggestion, which we consider qualitatively improves our paper.

      (1.3) The manuscript is narrowly focused on sensory neurons that initiate righting, which limits the advance given the known roles for daIV neurons in righting. With the suite of innovative new tools, there is a missed opportunity to gain a more general understanding of how sensory neurons contribute to the righting response, including promoting and inhibiting righting in different regions of the larva, as well as aspects of proprioceptive sensing that could be necessary for righting and account for some of the observed effects of 109(2)80.

      Once again, we appreciate this interesting comment by Rev1. We feel our study provides novelty in understanding how sensory neurons in different body regions contribute to the induction of the behaviour. We developed new technology to show that the activity of anterior sensory neurons is essential for normal righting and inhibiting this activity leads to a switch to a different behavioural regime. We feel this represents a substantial advancement in our understanding of how this behaviour is initiated that has not been previously described. Whilst we also appreciate there is likely to be a substantial role of proprioception in self-righting behaviour, our work here focuses on the external stimuli that elicit self-righting, as a detailed understanding of proprioception would be out of scope and require the development of further techniques to manipulate and measure larval posture. As detailed in the above comment, we feel that the more targeted investigation of daIV neurons can also shed some light on the cell-type specificity and inputs to the self-righting induction process.

      (1.4) Although the authors observe an influence of Hox genes in righting, the possible mechanisms are not pursued, resulting in an unsatisfying conclusion that these genes are somehow involved in a certain region-specific behavior by their region-specific expression. Are the cells properly maintained upon knockdown? Are axon or dendrite morphologies of the cells disrupted upon knockdown?

      We agree with this comment in that further investigating the effects of Hox expression on localised aspects of the sensory system poses an interesting line of investigation. Indeed, we are currently conducting a full scale analysis of Hox gene effects across the sensory field. As things stands, it is not clear how Hox gene expression could affect local sensory processes, a mechanism which could involve morphological changes, changes in neuronal excitability (e.g. due to changes in channel expression), synapse formation and/or efficiency, cell development and identity, and/or combinations of these effects, amongst other possibilities. It is clear that a complete and satisfying investigation of this mechanism for each of the Hox genes would pose a substantial amount of work so, while we acknowledge the merit of Rev1’s comment, we consider that adding a cellular-mechanistic analysis of Hox effects is out of scope for the present study and shall constitute a central matter for a followup study emerging from current projects. We think that our data on Hox expression/function as reported here should serve to open up the analysis of genetic regulation of local sensory function, an area in which we are currently working very actively.

      (1.5) There could be many reasons for delays in righting behavior in the various manipulations, including ineffective sensory 'triggering', incoherent muscle contraction patterns, initiation of inappropriate behaviors that interfere with righting sequencing, and deficits in sensing body position. The authors show that delays in righting upon silencing of 109(2)80 are caused by a switch to head casting behavior. Is this also the case for silencing of daIV neurons, Hox RNAi experiments, and silencing of CO neurons? Does daIII silencing reduce head casting to lead to faster righting responses?

      This is an insightful comment. In the revised version of the manuscript, we do indeed show that anterior inhibition of daIV neurons leads to the same head casting behaviour as with the 109(2)80 domain, which we interpret as an inability of the larvae to sense the underlying substrate (see page 28). We hope the new data addresses this comment, at least to an extent. While we acknowledge it would also be insightful to run this behavioural analysis for other experimental conditions, such as the daIII inhibition and Hox RNAi lines, these experiments pose a specific technical difficulty: the behavioural analysis relies on a deep neural network (DNN) which was trained solely on recordings of the opto-axial technique, meaning it does not translate well to other experimental situations. This problem is further compounded by the use of L1 larvae, which means recording resolution is insufficient to accurately define the body landmarks used in the posture tracking at a smaller scale. Therefore, the recourse for identifying behavioural changes is manual observation, which we feel is too inconsistent to address a quantitative question like this.

      (1.6) 109(2)80 is expressed in a number of central neurons, so at least some of the righting phenotype with this line could be due to silenced neurons in the CNS. This should at least be acknowledged in the manuscript and controlled for, if possible, with other Gal4 lines.

      We thank the reviewer for making this interesting comment. We have added a phrase to the section “Conditional inhibition of multidendritic neurons delays self-righting” (p21) which acknowledges the presence of 109(2)80 expression in the CNS (as reported by Hughes and Thomas). We agree that ideally, a variety of sensory Gal4 lines would be used to check for consistency of the effects. However, it is also important to note that 109(2)80 is one of the only available Gal4 lines with near sole md neuron expression, as other Gal4s also drive expression strongly in external sensory cells for example. Thus, re-running experiments with these other lines – which would involve a substantial investment of time and resources – would not be an ideal strategy. We feel that the new observation of (very) similar axial results using the ppk-Gal4, which does express solely in the daIV neurons, better helps to confirm the specificity of the findings to multidendritic neurons.

      Other points:

      (1.7) Interpretation of roles of Hox gene expression and function in righting response should consider previous data on Hox expression and function in multidendritic neurons reported by Parrish et al. Genes and Development, 2007.

      We thank Rev1 for pointing out this study, which is definitively important to discuss given our results on Hox genes. To address this gap, we have added an additional paragraph in the Discussion (p37) to discuss the documented effects of Hox genes on da neuron dendritic morphology and how our results can be interpreted in light of this.

      (1.8) The daIII silencing phenotype could conceivably be explained if these neurons act as the ventral inhibitors. Do the authors have evidence for or against such roles?

      This is another interesting suggestion. If the daIII neurons were to fulfil this role, then in theory, their inhibition would result in self-righting behaviour under conditions of combined dorsal and ventral substrate contact. This is not an experiment we performed, so we are currently unable to confirm or rule out this possibility. However, we note from casual observation that daIII inhibition does not cause larvae to spontaneously self-right. As mentioned above, our view is not one in which the system has “dorsal/ventral stimulators/inhibitors” for a given behaviour, but that action selection proceeds according to a coordination of many (dynamic) contextual clues. Given the new results with the axial inhibition of daIV neurons (see above) it might be more parsimonious to suggest that these “tiling” neurons are primarily responsible for detecting substrate contact around the full circumference of the animal, rather than this involving different cell types according to the different sides of the body.

      Reviewer #2 (Public review):

      Strengths:

      The work of Roseby et al. does what it says on the tin. The experimental design is elegant, introducing innovative methods that will likely benefit the fly behavior community, and the results are robustly supported, without overstatement.

      Weaknesses:

      The manuscript is clearly written, flows smoothly, and features well-designed experiments. Nevertheless, there are areas that could be improved. Below is a list of suggestions and questions that, if addressed, would strengthen this work:

      (2.1) Figure 1A illustrates the sequence of self-righting behavior in a first instar larva, while the experiments in the same figure are performed on third instar larvae. It would be helpful to clarify whether the sequence of self-righting movements differs between larval stages. Later on in the manuscript, experiments are conducted on first instar larvae without explanation for the choice of stage. Providing the rationale for using different larval stages would improve clarity.

      This is a very interesting point raised by Rev2. Most of our previous work on self-righting (e.g. PicaoOsorio et al. 2015 Science; Picao-Osorio, Baldaia et al. 2017 Genetics; Klann et al. 2021 Journal of Neuroscience) was focused on the first instar larva (L1) because this early stage: (i) represents the simplest form of all larval stages, (ii) allows meaningful comparisons with late embryonic processes guiding the development and physiology of the nervous system, (iii) captures the system in a relatively naïve state, that had limited if any exposure to external stimuli. Although these attributes remain valid for the investigation of the sensory stimuli that trigger self-righting, the implementation of the necessary regional physical measurements and manipulations used in this study (surface contact, opto-axial technique, deep neural network analysis) would be impossible to implement in the early forms of the larva simply due to its reduced size. Due to this, we employed L3s, which due to their larger dimensions enabled the development and use of the sophisticated regional stimulation techniques reported here. Yet, as Rev2 rightly points out, we return to the late embryo and early L1 at the point of conducting gene expression analyses as these are optimised for those early stages. The selection of larval stage according to experiment relies on the fact that all forms of the larva display self-righting (Issa, Picao-Osorio, et al. 2019 Current Biology), that SR does not differ according to larval stage and that the characterisation of the structure of the nervous system across larval stages has shown a large level of similarity and consistent topographically arranged connectivity between identified neurons (Gerhard et al. 2017 eLife).

      (2.2) What was the genotype of the larvae used for the initial behavioral characterization (Figure 1)? It is assumed they were wild type or w1118, but this should be stated explicitly. This also raises the question of whether different wild-type strains exhibit this behavior consistently or if there is variability among them. Has this been tested?

      Thank you to the reviewer for pointing this out. The genotype for Figure 1 was w<sup>1118</sup>; this has now been added to the figure legend and the results section – thank you to Rev2 for pointing this out. Although in this study we did not explicitly compare self-righting (SR) performance in wild type/control genotypes (as we are internally consistent in using w<sup>1118</sup>) based on previous data collected in our lab we know that self-righting times are similar and very consistent amongst inbred control lines such as w<sup>1118</sup>, yw, and Oregon Red. Furthermore, we can also add that when comparing SR times between these inbred populations with a highly polymorphic outbred Drosophila population (Martins et al. 2013 PLoS Pathogens) we observed that their SR time (i.e. 6.14s ± 1.06) was not significantly different from the inbred lines (p<0.05, U test) (Picao-Osorio, J. 2014 Doctoral Thesis, Chapter 4, p112).

      (2.3) Could the observed slight leftward bias in movement angles of the tail (Figure 1I and S1) be related to the experimental setup, for example, the way water is added during the unlocking procedure? It would be helpful to include some speculation on whether the authors believe this preference to be endogenous or potentially a technical artifact.

      This is an interesting comment, and we recognise that lateral manipulation biases in self-righting could indeed reflect experimental limitations or biological tendencies. At this point we cannot interpret these results as formal evidence of chirality, given that they may reflect subtle aspects of the micromanipulation of specimens. We are currently developing a motorised platform to conduct self-righting tests, which when fully developed, should help addressing the chirality question.

      (2.4) The genotype of the larvae used for Figure 2 experiments is missing.

      Thank you for pointing this out. These were again w<sup>1118</sup> larvae; this detail has now been added to the figure legend and the main text.

      (2.5) The experiment shown in Figure 2E-G reports the proportion of larvae exhibiting self-righting behavior. Is the self-righting speed comparable to that measured using the setup in Figure 1?

      Thank you for pointing this out. We have now added average self-righting times to the figure legends of figures 1 and 2. The self-righting times across for the dorsal + ventral contact conditions was notably longer than dorsal-only cases, which were also slightly longer than the “standard” case. This is perhaps to be expected, as the larvae are encountering unusual and ambiguous situations. We suggest the extra time could reflect an additional decision-making step or action flip-flopping process, or simply physical constraints on the movement (for example, not being able to use some parts of the body).

      (2.6) Line 496 states: "However, the effect size was smaller than that for the entire multidendritic population, suggesting neurons other than the daIVs are important for self-righting". Although I agree that this is the more parsimonious hypothesis, an alternative interpretation of the observed phenomenon could be that the effect is not due to the involvement of other neuronal populations, but rather to stronger Gal4 expression in daIVs with the general driver compared to the specific one. Have the authors (or someone else) measured or compared the relative strengths of these two drivers?

      We agree with this suggestion and to address this concern, we have added as part of our new figure Supp. Fig. S3, a dedicated panel S3C showing fluorescence measurements from ddaC using the 109(2)80-Gal4 and ppk-Gal4 lines. We found no difference in tdTomato fluorescence intensity, suggesting equal expression strength across the two Gal4 drivers. Our new results for axial daIV inhibition are also consistent with this effect size difference, further suggesting that inhibition of all md neurons poses stronger challenges for self-righting compared to the daIV neurons alone.

      (2.7) Is there a way to quantify or semi-quantify the expression of the Hox genes shown in Figure 6A? Also, was this experiment performed more than once (are there any technical replicates?), or was the amount of RNA material insufficient to allow replication?

      Unfortunately, we only had limited amounts of mRNA extracted from FACS-sorted 109(2)80>GFP cells to feed our reverse transcriptase reactions and used much of these samples for the experiment reported. After Rev2 suggestion we went back to our freezers, recovered traces of the samples used in the original experiment, and attempted a new amplification; despite this effort, this new experiment was unsuccessful. We feel that the main point deduced from the original experiment is valid in that we obtained amplicons of the expected size for all the Hox transcripts analysed and that for those cases in which we observed biological effects – i.e. Antp and Abd-B – we corroborated protein expression in the 109(2)80 domain using immunohistochemistry. We are currently expanding this project examining the roles of all Hox genes across the entire sensory system and shall report the expression patterns of all Hox genes in each of the subcomponents of the sensory system the future.

      (2.8) Since RNAi constructs can sometimes produce off-target effects, it is generally advisable to use more than one RNAi line per gene, targeting different regions. Given that Hox genes have been extensively studied, the RNAis used in Figure 6B are likely already characterized. If this were the case, it would strengthen the data to mention it explicitly and provide references documenting the specificity and knockdown efficiency of the Hox gene RNAis employed. For example, does Antp RNAi expression in the 109(2)80 domain decrease Antp protein levels in multidendritic anterior neurons in immunofluorescence assays?

      We used the TRiP RNAi lines, specifically the Valium10 selection available from the Bloomington Stock Centre. Unfortunately, there is not much information on how specific the Hox RNAi lines areor whether their might have off-target effects.

      (2.9) In addition to increasing self-righting time, does Antp downregulation also affect head casting behavior or head movement speed? A more detailed behavioral characterization of this genetic manipulation could help clarify how closely it relates to the behavioral phenotypes described in the previous experiments.

      This would be interesting line of investigation. As described in a previous comment, this is currently unfeasible for us given some important differences between experiments including larval stage and recording conditions. We have added some speculative comments to the manuscript describing the larval behaviour under Hox RNAi.

      (2.10) Does down-regulation of Antp in the daIV domain also increase self-righting time?

      Given the new results with axial effects of daIV neurons, we also sought to address this point with a new series of experiments expressing Hox RNAi constructs in the ppk-Gal4 domain. The new data is shown in a new figure (Figure S8) displaying self-righting times for ppk-Gal4-Hox-RNAi. Interestingly, we found no effect of any RNAi expression on self-righting times, suggesting that md types other than daIVs are under Hox regulation that is important for self-righting.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers were enthusiastic about the value and quality of this study by Roseby and colleagues. There were two main issues that emerged from the reviews that we're highlighting for the authors to address, should they choose to:

      (1) A little more cell-type resolution of the anterior region

      The anterior region includes a lot of sensory neurons that may be contributing to the effect. Some sensory neurons (e.g., daIV) have been implicated in righting - are these the ones carrying the anterior signal? Are dorsal sensory neurons promoting righting and ventral ones stalling it?

      We are not suggesting a complete sensory-neuron mapping in the anterior region. Instead, we propose the authors conduct a focused check: repeat the axial inhibition with a daIV-specific driver (same photomask assay) to show the A-P effect within the implicated class, and, if possible, replicate one key result with an alternative broad md driver to address Gal4 strength/off-target expression.

      As mentioned above (see Rev1 comment) we have indeed carried out a new experiment using an alternative driver to 109(2)80-Gal4 and testing the impact of these manipulations on larval behaviour. The revised version of our MS includes a new figure Supp Fig S3 which shows self-righting times when using the ppk-Gal4 driver with the opto-axial technique. As with the 109(2)80-Gal4 driver, self-righting was delayed in anterior but not posterior inhibition conditions, suggesting the daIV neurons specifically act in a region-specific manner to trigger postural control behaviour.

      Furthermore, in another new figure, Supp Fig S7, we show that head casting behaviour is also increased in the same manner as with the 109(2)80-Gal4 driver. These new panels and figures are cited within the sub-sections entitled “Optogenetic inhibition of anterior but not posterior multidendritic neurons delays self-righting” and “Inhibition of anterior multidendritic neurons is associated with increased head casting during self-righting”, on pages 25 and 28, respectively. We are grateful to R1 for this suggestion, which we consider qualitatively improves the quality of our paper.

      (2) The Hox section to strengthen this section, we recommend:

      (a) Confirm specificity/efficacy of knockdown (e.g., Antp protein reduction in targeted md neurons and a second RNAi line if available).

      This is a reasonable comment. For our experiments, we selected a UAS-Antp<sup>RNAi</sup> line (Bloomington #27675) given that this construct has been: (i) utilised in several previous studies as the main and single line to interfere with Anpt expression (e.g. Baek et al. 2013 Development; Paul et al. 2021 Nature Comms) and (ii) shown to display a consistent reduction in Antp protein levels of approximately 50% (see Poliacikova et al. 2024 Science Adv.). Furthermore, previous work comparing #27675 with other UAS-Antp<sup>RNAi</sup> lines has demonstrated that all available lines lead to a similar level of reduction in protein expression, although the #27675 line exhibits the most consistent effects (lower variability) (Poliacikova et al. 2024 Science Adv.). Unfortunately, at this point in time, we do not have the capacity to conduct new experiments with other RNAi lines, but consider that the information and arguments mentioned above should be reassuring about our choice of a reasonable and previously validated method to interfere with Antp expression.

      (b) Perform one temporal control (GAL80^ts) or a simple rescue, to separate developmental vs acute roles.

      This is a good and interesting suggestion, but we consider that the discrimination between developmental and physiological effects falls outside the scope of this study. Indeed, experiments of this kind are currently being conducted in our lab as part of a wider examination of Hox gene roles in the sensory system.

      (c) Place the results clearly in the context of prior work (e.g., Parrish 2007), so the mechanism isn't left hanging.

      This is an important point, and we have now done this. Many thanks for pointing this out.

      Reviewer #1 (Recommendations for the authors):

      (1.1) A Gal4 line for the pannier dorsal specification gene shows expression in dorsal sensory neurons, as described in Galindo et al., Development, 2023, and could help tease apart dorsal v. ventral contributions.

      This is an interesting suggestion. However, we understand that the pannier (pnr) Gal4 line mentioned in Galindo et al. 2023 is an enhancer trap inserted in the pnr locus which drives expression in neural as well as non-neural tissues such as the embryonic dorsal ectoderm (see: Calleja et al. 1996 Development; Stronach et al. 2014 Genetics). Although, as Rev1 rightly indicates, this line also labels dorsal cluster sensory neurons, including ddaC (cIV) and ddaF (cIII) neurons the fact that the line displays expression in non-neural tissues makes its use in behavioural experiments difficult as non-neural effects might affect the behavioural patterns studied. A possible way to instrument the pnrGal4 tool into behavioural analyses might involve the creation of the necessary variants to implement a split-Gal4 approach, but this, we believe, unfortunately falls out of the scope of this study.

      (1.2) Potential roles for daII neurons and daI neurons are not examined. Drivers have been described for daII neurons, and there are drivers that will target a majority of proprioceptive md neurons, so these could be examined to complete the analysis started here.

      This is another interesting suggestion by Rev1, but we consider that the fine-grain mapping of effects mediated by sensory neuron sub-clases falls outside the scope of this study aimed at mapping sensory regional effects on self-righting. This does not take the merit of the suggestion away, and indeed, experiments of this kind are currently being conducted in our lab as part of a comprehensive examination of Hox gene roles in the sensory system.

      (1.3) To account for 109(2)80 off targets, the authors could consider other lines that silence most or all md neurons (clh201-Gal4; 5-40-Gal4; 21-7-Gal4) that could at least have different central offtargets. Some other lines are broad somatosensory system drivers but sensory-specific (pebbledGal4).

      This is an interesting comment, and so are the suggestions made. Although to include this kind of verification would be interesting, when carrying out our experiments, we did not observe any central expression at all. Also, to repeat all our experiments in which we use the established and validated 109(2) 80 line using instead these four Gal4 lines, is unfortunately out of scope for us at this point in time. We will nonetheless consider these comments by Rev1 in future extensions of our work.

      (1.4) There is a typo on line 481; it should be "other".

      We are grateful to R1 for pointing this out. This has now been amended

      Reviewer #2 (Recommendations for the authors):

      (2.1) Lines 91-92 cite references describing self-righting behavior across different animal groups, which is illustrated in Figure 1B. It would be helpful to indicate these references directly in the figure. For example, instead of using dots to denote their presence (which are, in a way, redundant since the behavior is reported in all groups), numbers or letters could be used to refer to the specific papers describing them.

      Thank you for this suggestion. We have now replaced the original dots by an abridged citation of a key paper providing evidence in that specific animal group, e.g. Smith, et al. 1997; Rogers et al. 2015

      (2.2) In Figure 1A, the diagrams illustrate the two large dorsal tracheae, which nicely indicate the larva's orientation. However, since they are drawn in a very light gray, they can be difficult to distinguish without zooming in. It might improve clarity if the tracheae were made slightly more prominent.

      Thank you for this suggestion. We have now implemented this change.

      (2.3) In Figure 1E, the dotted line and green bar mark the segment of the recording corresponding to self-righting, which is then quantified in Figure 1G. Was the same procedure applied when analyzing tail speed, or was it limited to head speed? Figure 1F does not show a dotted line or green bar, which is confusing; it would be helpful to clarify the reason for this discrepancy. Also, in Figure 1G, there is an inset showing photos of the movement sequence with the green bar and the caption 'Trimmed to SR sequence,' which implies to me that for tail speed, the 0.75-1 segment of the recording was also used for quantification. I suggest adding the dotted line and green bar to Figure 1F and removing this inset from Figure 1G, as it appears quite small and disrupts the layout of the figure. If it is retained, the figure legend should explicitly refer to the inset.

      Thank you for pointing this out. We have amended these figures as suggested.

      (2.4) In Figures 1 and 2, the box plots include the individual data points, whereas Figures 3 and S2 do not. For data transparency, it would be important to show the individual measurements here as well. I strongly recommend adding them to the figure, or alternatively providing a clear rationale in the text for not doing so.

      Thank you for mentioning this. The reason data points are not shown in Fig 3 or S2 is because the variance extends the scale and compresses the box making it illegible. To make this clear we now explain this in the figure legends.

      (2.5) In Figures 4 and 5, the distribution of self-righting times from the optogenetic inhibition experiments is shown using bar graphs rather than box plots, as in the previous figures. This choice obscures the data distribution, since all bars reach down to zero. Replacing the bar graphs in Figures 4 and 5 with box plots would more clearly convey the experimental results.

      We thak Rev2 for this comment, which gives us an opportunity to clarify the matter. Distributions of SR times are drawn with bars because we compare means +/- variance in the analysis, and not medians +/- IQR as is done in the other experiments. The choice of visualisation reflects the analysis, which is what is recommended by statisticians. Plus, we also show the individual observations, meaning the distribution can be observed. We hope that it is now clear that we are not obscuring any distributions.

      (2.6) Figure 6 would benefit from some reorganization. Panel A is very small and dense with information, making it difficult to interpret without significant zooming. In particular, the FACS graph is nearly impossible to read, as the axes remain unclear even when enlarged. It might be best to either remove this graph and replace it with a cartoon version of FACS-sorted populations, and reorganize the figure to ensure legibility. Additionally, the current layout progresses from the bottom up, which takes time to follow. Comprehension could be improved if the sequence began with the larva dissection placed in the top left area of the figure, where readers typically look first (I appreciate that this is mentioned in the figure legend; however, a different layout might present the information more effectively).

      We appreciate the constructive spirit of this comment and have indeed considered Rev2 suggestions including drafting new layouts of this figure. After all this experimentation, we remain of the view that the original presentation is probably the best trade-off between size and clarity, offering more space for the appreciation of confocal imaging and its interpretation.

      Minor corrections:

      (1) Throughout the text, the word Drosophila appears sometimes in italics and sometimes in regular font; please standardize its formatting for consistency.

      Amended

      (2) Line 179: the use of three hyphens in the sentence "minimum --- in all cases < 30 s --- to avoid larval desiccation" is unusual; exchanging them for commas or brackets is advised.

      Amended

      (3) Line 183: in w1118, the numbers are usually in superscript (not subscript), and the w should be italicized.

      Amended

      (4) In line 783, there is an incorrect space between "is" and the comma in "...repertoire, which is , in...".

      Amended

      (5) In Figure 2G, the left panel appears partially cut off, which makes the text at the edges difficult to read. It might help to adjust the panel so that all labels are fully visible.

      Done

      (6) In the current version of the manuscript, Figure 5 is presented before Figure 4, which is confusing.

      This has been amended.

      (7) Two videos are included in the supplementary material, but I could not find any reference to them in the main text of the manuscript.

      This has been amended.

    1. Author response:

      We thank the editors and reviewers for their thoughtful and constructive evaluation of our manuscript. We are pleased that the reviewers found the study valuable and the evidence supporting a role for Yme1 in MDC formation solid. As described below, we plan to modify the manuscript to clarify the lipid model, better explain the relationship between Ups-family proteins and MICOS, distinguish MDC formation from Atg32-dependent mitophagy, clarify metabolic conditions, add statistical analyses where missing, and strengthen Yme1 validation with immunoblotting.

      eLife Assessment

      This valuable study demonstrates that the inner membrane protease YME1 contributes to the formation of mitochondrial-derived compartments in yeast through the modulation of both the lipid transporter UPS2 and the MICOS complex. The evidence supporting this model is solid, although this manuscript could be improved by providing additional evidence supporting the independent roles for UPS2 and MICOS regulation in this process. This work will be of interest to cell biologists, biochemists, and geneticists interested in understanding the molecular basis of mitochondrial regulation and function.

      We appreciate this positive assessment and agree that the roles of Ups-family lipid transport and MICOS in MDC regulation could be expanded further. This will be an important topic for future studies, especially with regard to how MICOS contributes to MDC formation. In the current revision, we will add new genetic data focused on PA-linked lipid metabolism through the yeast Pah1/Lipin pathway, which we think will help strengthen and clarify the lipid arm of the model. Our current interpretation is that Yme1-regulated Ups-family lipid transport and MICOS may both influence a shared mitochondrial membrane state that permits MDC formation. This interpretation is consistent with our genetic data and with known connections between Ups proteins, MICOS, and mitochondrial membrane organization.

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Balasubramaniam and colleagues continue this group's efforts to understand mitochondrial-derived compartments (MDCs) that bud off from yeast mitochondria in response to metabolic stress. In a previous genetic screen, they identified Ups lipid transfer proteins and the AAA-protease Yme1 as components that modulate MDC formation. In this study, the authors link these observations by showing that Yme1 modulates levels of Ups1, Ups2, as well as MICOS complex members in the mitochondrial proteome. Using genetic approaches, they then show that Yme1's role on MDCs is dependent on its catalytic activity (via an inactive mutant) and that YME1 shows genetic interactions with UPS1/2 and MIC10/MIC60. The overall model is that Yme1 activity responds to metabolic cues and acts via proteolysis of these two distinct mitochondrial machineries to regulate MDC biogenesis.

      Strengths:

      The strengths of the study are its integration of mitochondrial proteomics with strong genetic approaches, as well as synergy with the authors' previous studies on the role of lipids in MD genesis. The work is overall well carried-out and experiments are thoughtfully discussed.

      Weaknesses:

      The major weaknesses are a lack of mechanistic resolution surrounding the model, e.g., proposed or tested mechanisms by which Yme1 activity is regulated by metabolic cues, or how Ups1/2 activity and the MICOS contribute to MDC generation. The authors acknowledge these as open questions, but addressing them would still enhance the significance of the study.

      We thank the reviewer for the positive assessment, and we agree that the upstream regulation of this response remains an important open question. Yme1-dependent MDC regulation could involve changes in Yme1 activity, substrate accessibility, or broader changes in mitochondrial lipid and protein organization. Fully resolving how metabolic state gates this response will require future work, likely outside the scope of the current study.

      We also agree that the manuscript would benefit from a more developed discussion of how lipid changes could contribute to MDC formation. Our prior work showed that reduced mitochondrial PE promotes MDC formation, whereas cardiolipin is required for MDC biogenesis (Xiao et al., 2024). We proposed that reduced PE changes the membrane environment of mitochondrial outer membrane proteins, potentially affecting their stability, abundance, insertion, or lateral organization within the membrane. Such changes could increase the pool of proteins available for sorting into MDCs or make the outer membrane more permissive for domain formation. In the revision, we will connect this model more directly to Yme1-dependent regulation of Ups-family lipid transport.

      We will also expand the model to incorporate PA-linked metabolism. We did not initially focus heavily on Ups1 because complete loss of UPS1, or loss of downstream cardiolipin synthesis through CRD1, blocks MDC formation because cardiolipin is required. Thus, complete disruption of Ups1-dependent lipid transport may obscure the effects of more moderate changes in PA flux. To address this, we will include additional lipid measurements and new genetic data targeting PA metabolism through the yeast Pah1/Lipin pathway. Because Pah1 converts PA to DAG, this provides a way to alter PA-linked metabolism without simply eliminating cardiolipin synthesis. Our new data suggest that PA accumulation or altered PA-linked lipid flux may also promote MDC formation. Together, these findings support a broader model in which reduced PE and increased PA alter both the organization of OMM proteins and the physical properties of the membrane, including curvature and domain formation, thereby creating a membrane state that is more permissive for MDC biogenesis.

      Reviewer #2 (Public review):

      In this manuscript, the authors report a novel regulation of the outer mitochondrial membrane remodeling domains called mitochondria-derived compartments, MDCs. The team has previously established the main principles behind this recently identified quality control pathway, but the mechanisms that control MDCs formation remain incompletely understood. Using the baker's yeast model, the authors identify the conserved mitochondrial protease Yme1 as a crucial factor that regulates MDC formation. Mechanistically, Yme1's proteolytic function controls the levels of Ups1 and Ups2 lipid transfer proteins and the components of the membrane organizing complex called MICOS, thus providing a plausible model as to how Yme1-dependent proteolysis permits MDC formation through the removal of lipid and MICOS-dependent constraints. Finally, the authors show that this Yme1-mediated activity is also defined by metabolic conditions. In principle, this study is interesting and novel, and holds potential to provide new insights into the regulation of the MDC pathway that emerged as a new fundamental mitochondrial quality control mechanism. However, the following points should be carefully addressed.

      Major points:

      (1) Yme1 has been previously shown to regulate mitochondria-specific autophagy through Atg32 processing. Given the high similarity of the MDC pathway to piecemeal autophagy and the fact that both pathways share some of the core components, the authors should address the involvement of Atg32 in their model. It would also be important to include a brief discussion addressing the differences between piecemeal autophagy and the MDC pathway.

      We agree that this is an important point. The reason we did not focus on Atg32 in the current manuscript is that we previously investigated the relationship between MDC formation and Atg32-dependent mitophagy and found that Atg32 is dispensable for MDC formation (Hughes et al., 2016). Based on that result, we do not anticipate that Atg32 is required for the Yme1-dependent MDC phenotypes described here. This is also consistent with the different growth conditions associated with these pathways: Atg32-dependent mitophagy is stimulated under respiratory or post-diauxic conditions, whereas MDCs do not form under the respiratory conditions that stimulate Atg32-dependent mitophagy (Hughes et al., 2016; Raghuram and Hughes, 2024).

      We will clarify this distinction in the revised manuscript. In addition, to be thorough, we plan to generate and test the Atg32-GFP variant previously shown to block Yme1-dependent Atg32 processing and mitophagy (Wang et al., 2013). This will allow us to test directly whether preventing Yme1-dependent Atg32 cleavage affects MDC formation. If successful and interpretable, we will include these data in the revised manuscript.

      (2) The Rpt3 (P215L) expression experiment is interesting, but appears to be somewhat superficial due to the unclear mechanism by which the mitochondrial network morphology is restored in these cells. Could this result be replicated in the dnm1∆ mgm1∆ double deletion mutant, which is a well-established model for mitochondrial network restoration?

      We agree that the Rpt3(P215L) experiment is best viewed as a morphology control. The purpose was to test whether abnormal mitochondrial morphology alone explains the MDC defect in yme1Δ cells. Because Rpt3(P215L) improved mitochondrial morphology but did not restore MDC formation, we interpret this as evidence that morphology alone is not sufficient.

      We attempted to generate the requested dnm1Δ mgm1Δ yme1Δ triple-mutant combination, but that strain combination has not been viable in our hands. However, we do have dnm1Δ data showing that altering mitochondrial structure can rescue some morphological features but does not restore MDC formation in yme1Δ cells. We will include these data where appropriate and clarify that this experiment is intended as a morphology control.

      (3) Figure 3E. The changes in PE levels appear to be minor. While statistically significant, the observed differences may not be physiologically relevant. More in-depth lipidomic analysis data should be presented to substantiate the authors' argument and better address the questions at hand. Related to that, could PE or PA supplementation stimulate MDC formation?

      We agree that additional lipid data would strengthen this part of the manuscript. We initially streamlined the lipid section because we had previously examined the lipid requirements for MDC formation in detail, showing that reduced mitochondrial PE can promote MDC formation, whereas cardiolipin is required (Xiao et al., 2024). However, the current study would benefit from a broader analysis of the lipid changes associated with Yme1-dependent regulation.

      In the revision, we will expand the lipid data to include additional lipid species and incorporate these results into the model. We will also add new genetic data targeting PA metabolism through the yeast Pah1/Lipin pathway. Together, these data suggest that PA accumulation or altered PA-linked lipid flux may also contribute to MDC formation. This supports a broader lipid-balance or lipid-shunting model in which reduced PE, increased PA, or altered lipid distribution between mitochondrial membranes could influence OMM remodeling through effects on membrane curvature, OMM protein organization, or mitochondrial membrane contacts.

      We agree that direct PE or PA supplementation would be a valuable experiment. We have attempted lipid supplementation but have not been able to deliver these lipids effectively to yeast cells in a way that produces interpretable results. We are therefore focusing on lipid profiling and genetic approaches that alter lipid metabolism inside the cell.

      (4) The connection between rapamycin treatment and Yme1-regulated MDC formation is unclear and puzzling and needs to be explained better.

      We agree that this connection is not fully clear. In this manuscript, rapamycin is used primarily as a robust MDC-inducing condition. Our data do not define the full pathway connecting TORC1 inhibition to Yme1-dependent mitochondrial remodeling.

      In the revision, we will either clarify this point or reduce the emphasis on rapamycin as a mechanistic entry point. Our current interpretation is that rapamycin creates a metabolic/mitochondrial state in which Yme1-dependent remodeling of lipid and membrane-organization pathways becomes important for MDC formation. Whether this involves direct regulation of Yme1, altered substrate availability, altered membrane composition, or a combination of these remains open.

      (5) The MICOS complex is clearly involved in the regulation of MDC, but the manuscript misses the mark on providing compelling evidence and a clear explanation as to how MICOS contributes to said regulation.

      We agree that the mechanism by which MICOS regulates MDC formation remains an important open question and will be a major focus of future work. Our current data show that MICOS perturbation can partially restore MDC formation in yme1Δ cells, supporting a role for MICOS in this pathway. This analysis was motivated in part by the incomplete genetic suppression achieved through the lipid pathway alone, which suggested that additional Yme1-regulated factors contribute to MDC formation.

      MICOS therefore represents a strong candidate for this additional regulatory input. However, defining whether MICOS acts through lipid distribution, OMM-IMM organization, membrane architecture, or another mechanism will require a deeper investigation than is possible within the scope of the current study. We will clarify this point in the revised manuscript and present the current findings as the beginning of a broader investigation into how MICOS contributes to MDC biogenesis.

      Minor points:

      (1) The authors should discuss potential reasons for the dramatically different rates of MDC formation in the S288C and W303 background cells. Does this have anything to do with generally more robust mitochondrial functions in the latter cells?

      We agree this is worth discussing. One likely explanation is that the difference reflects broader differences in mitochondrial activity and metabolic state between these strain backgrounds. We and others have shown that W303 cells have more robust respiratory mitochondrial function than BY/S288C-derived cells, and in our hands W303 also shows lower MDC formation. This fits our broader model that MDCs are favored in glucose-grown or metabolically perturbed cells and do not form under respiratory conditions (Raghuram and Hughes, 2024). We do not yet know the genetic basis for this difference, so we will present this as an interesting future direction.

      (2) Proper statistical analyses should be provided for all the graphs presented.

      We will add statistical analyses where missing.

      (3) The authors should include Yme1 immunoblots to confirm the identity of strains being studied and validate the presence or overexpression of Yme1 and its catalytic mutant in their experiments.

      We agree that direct validation of Yme1 protein levels will strengthen the manuscript. Our quantitative mitochondrial proteomics already confirms strong depletion of Yme1 in yme1Δ cells, and we will also include quantitative proteomics showing increased Yme1 abundance in the overexpression strain. In addition, we have now obtained a Yme1 antibody from a colleague and will include immunoblots validating Yme1 loss, re-expression, catalytic mutant expression, and overexpression where appropriate.

      Reviewer #3 (Public review):

      Summary:

      Since describing MDCs over a decade ago, the lab of the corresponding author, Hughes, has been at the forefront of further characterizing these structures. Here, they follow up on recent work (PMID: 38497895), where a screen identified Yme1 as a potential regulator of MDCs. After confirming that Yme1-ko prevents MDCs that are usually induced via various established treatments (Rapamycin, cycloheximide, Concanavalin A), the authors confirmed that the proteolytic activity of Yme1 is required. Next, using proteomics, they identified how loss of Yme1 impacts the mitochondrial proteome with and without Rapamycin treatment to induce MDCs. From this result and based on insight from other published data implicating lipids, the focused initially on the lipid transfer protein Usp2, a known target of Yme1. Here, they showed that loss of Usp2 could partially rescue MDC formation in Yme1-ko cells. To look for other Yme1 targets that might also be involved in MDC formation, next, they investigated the MICOS complex, which was also notable in their proteomics data. They then showed that inhibiting MICOS also partially restored MDC formation in Yme1-ko cells. They then tested the combined effects of Usp2 and MDC inhibition on MDCs, which was limited by the fact that the combination of full MICOS disruption, Usp2-KO, and Yme1-KO was not viable. To circumvent this limitation, they investigated the knockout of individual MICOS subunits in combination with Usp2 and/or Yme1. Finally, they showed that growth conditions also mediate MDC formation in the context of Yme1 overexpression. In rich media, Yme1 overexpression induces MDCs on its own. However, this induction is lost upon amino acid starvation, suggesting that there are still other as-yet-unidentified factors regulating the formation of MDCs.

      Strengths:

      The authors use unbiased approaches and genetic models to begin unraveling a novel regulatory role of Yme1 in the formation of MDCs.

      Weaknesses:

      (1) The authors find both Ups1 and Ups2 in their screens, but only focus on Ups2 in this paper. It would be good to know why they did not also investigate Ups1, and its other protease Atp23, which could potentially act similarly to Yme1, or even rescue the loss of Yme1.

      We agree that Ups1 and Atp23 are important to consider. We initially focused on Ups2 because its deletion partially restores MDC formation in yme1Δ cells and because of its connection to mitochondrial PE synthesis, which we had previously shown to regulate MDC formation (Xiao et al., 2024). Ups1 is more difficult to assess genetically because complete loss of UPS1, or of downstream cardiolipin synthesis through CRD1, blocks MDC formation due to the requirement for cardiolipin. Thus, an ups1Δ phenotype cannot readily reveal whether a more moderate reduction in Ups1 activity, and the resulting accumulation or redistribution of PA, might promote MDC formation.

      In the revision, we will explain this rationale and include new genetic data targeting PA metabolism through the yeast Pah1/Lipin pathway. This provides a way to test the contribution of PA accumulation without simultaneously eliminating cardiolipin synthesis, and our initial results support a role for PA-linked lipid remodeling in partially bypassing the requirement for Yme1. We will also discuss Atp23 as a potentially important regulator of Ups1 and PA metabolism. A full investigation of Atp23 will be an important direction for future work.

      (2) I'm not convinced that the data support the notion that Usp2 and MICOS have distinct effects on MDCs. In Figure S3C-D, there is no statistical analysis to indicate whether the small differences between the MICOS-ko and the double knockout are significant. If MICOS-ko and Ups2-ko were acting through different mechanisms, one would expect their combination to be additive; this does not appear to be the case, as both single deletions and the double deletion all cause similar levels of MDCs (~30-40%). Rather, this result is what you would expect if they were working through the same mechanism. There also does not appear to be an additive effect in Figure 4F-G, when using the mic60-ko rather than the complete MICOS-ko. In this regard, the authors note in their discussion that 'loss of MICOS may disrupt membrane associations or alter lipid distribution between mitochondrial subcompartments' (lines 390-392). The latter situation seems like it would be the same mechanism as Usp2 and would more accurately explain their findings.

      This is a very good point, and we agree with the reviewer’s interpretation. The lack of strong additivity is consistent with Ups2 and MICOS acting within the same pathway or converging on a shared mechanism, rather than representing two separate mechanisms of MDC regulation. We did not intend to imply that these must be independent pathways. In the revised manuscript, we will ensure that the text reflects this interpretation and will add statistical analyses to the relevant comparisons.

      (3) The manuscript is missing key data confirming the re-expression or overexpression of Yme1 protein (Figure 1 E/G and Figure 5A). It is important to know the relative levels of expression of the re-expressed proteins to each other and to endogenous Yme1.

      We agree that direct validation of Yme1 protein levels is important. Our quantitative mitochondrial proteomics already confirms strong depletion of Yme1 in yme1Δ cells, and we will also include quantitative proteomics showing increased Yme1 abundance in the overexpression strain. In addition, we have now obtained a Yme1 antibody from a colleague and will add immunoblots validating Yme1 loss, re-expression, catalytic mutant expression, and overexpression.

      (4) Some clarification of the details for metabolically restrictive conditions would be helpful.

      Thanks for this suggestion. We will clarify these conditions throughout the manuscript and figure legends and will define exactly what we mean by low-amino-acid, amino-acid-free, synthetic, and rich media conditions. More broadly, MDC formation is strongly influenced by media composition and mitochondrial metabolic state. MDCs form less efficiently in synthetic media and do not form under conditions that promote respiratory mitochondrial function (Raghuram and Hughes, 2024).

      (5) Beyond just the presence/absence of MDCs, does more detailed quantification of their size/shape reveal any subtle differences between conditions?

      This is an interesting question. In our hands, MDC size and shape are variable and appear strongly influenced by mitochondrial fission/fusion state. Conditions that favor more fused mitochondrial networks can produce larger MDC-like structures, whereas fragmented networks can produce smaller structures. So far, we have not found a simple size or shape metric that explains the Yme1/Ups2/MICOS phenotypes better than MDC frequency.

      We will clarify this point in the revised manuscript and avoid implying that MDC frequency captures every possible morphological difference. More detailed morphometric analysis of MDC size, topology, and maturation state will be an important future direction, especially as we connect lipid remodeling to membrane curvature and MDC biogenesis.

      References

      Hughes, A.L., Hughes, C.E., Henderson, K.A., Yazvenko, N., and Gottschling, D.E. 2016. Selective sorting and destruction of mitochondrial membrane proteins in aged yeast. eLife. 5. doi: 10.7554/eLife.13943.

      Raghuram, N., and Hughes, A.L. 2024. Amino acids trigger MDC-dependent mitochondrial remodeling by altering mitochondrial function. bioRxiv. 2024.07.09.602707. doi: 10.1101/2024.07.09.602707.

      Wang, K., Jin, M., Liu, X., and Klionsky, D.J. 2013. Proteolytic processing of Atg32 by the mitochondrial i-AAA protease Yme1 regulates mitophagy. Autophagy. 9(11):1828–1836. doi: 10.4161/auto.26281.

      Xiao, T., English, A.M., Wilson, Z.N., Maschek, J.A., Cox, J.E., and Hughes, A.L. 2024. The phospholipids cardiolipin and phosphatidylethanolamine differentially regulate MDC biogenesis. Journal of Cell Biology. 223(5). doi: 10.1083/jcb.202302069.