10,000 Matching Annotations
  1. Last 7 days
    1. married French women had easier accessto divorce by consent than women in the US and the UK, and unlikewomen in the US, married or pregnant French women were not easy toforce out of their jobs.

      Not comp politics theory of religion causes democracy

    2. and whenthere was some degree of ambiguity about women’s preferences, partiessubject to high levels of political competition become open to the challengeof fighting over the women in the middle.

      Uncertainty in ideology will actually help women because multiple parties think they can win the vote

    3. The uncertainty surrounding women’s future loyalties drove abias toward the status quo electoral rules that could only be overcomewhen competition was high or during a moment of political realignment.

      The progressive era

    4. why in many countries the longest standingresistance to women’s inclusion came from centrists

      because they had the most to lose with the introduction of a new distribution and the median voter shifts

    5. universal franchise bill (14 percent of today’scountries), as a result of external imposition (30 percent), gradually, aftersome men had already gained political voice (42 percent)

      different processes

    Annotators

    1. agonism

      A state of conflict, struggle, or competition between opposing parties. The political landscape was defined by an intense agonism between the two major partieS

    2. liberal humanism frames

      Liberal humanism promotes valuable universal goals, but Andreotti questions who defines “progress” and how the ahead/behind model can reproduce HEADS UP dynamics — particularly salvationism and paternalism.

    3. humanitarian interventions which generally define help as themoral responsibility of those who are ahead in terms of internationaldevelopment.

      ts underlying assumptions make it particularly vulnerable to reproducing those patterns unless they are critically examined. That's an important distinction. Take something like Education for All. Andreotti isn't saying: “Universal education is bad.”

      She's asking something more uncomfortable: “What do we mean by education? Who designed the model being universalised? What kinds of knowledge does it recognise? What happens if we assume that people without our form of schooling are simply ‘behind’?”

      Same with human rights. She doesn't necessarily have to reject rights to ask: “Who gets to define what universal rights mean in practice, and how are they imposed or negotiated across very different societies?”

    4. In this sense, obstacles to human progress becomethe focus of government agreed targets

      Humanity is progressing toward certain universal goods → international institutions can identify what those goods are → nation states should deliver them → countries that are further ahead have a responsibility to help those further behind.

      Andreotti wants you to notice that there are several assumptions packed into that seemingly benevolent story. Who decided what “human progress” means? Why are economic development, democracy, formal education, etc. arranged into this particular model of progress? Who gets described as “ahead” and “behind”? What knowledge from supposedly “behind” societies disappears when the relationship is framed that way?

    5. From this perspective,global/development education, often associated with ideas of ‘socialresponsibility’ involves the export of expertise from those heading the way interms of economic development to those lagging behind

      That’s why she’s arguing for critical literacy. She wants global education to move from: “How can we help them?” toward questions like: “Why do we understand the problem this way? Who gets to define development? What histories and power relations are involved? Are our proposed solutions reproducing those same relations?” So yes, she’s critiquing the field from within it, with the aim of making it better rather than rejecting it.

    6. From this perspective,global/development education, often associated with ideas of ‘socialresponsibility’ involves the export of expertise from those heading the way interms of economic development to those lagging behind

      Global/development education can become problematic when it assumes “we are developed, they are behind, and our job is to help them catch up.”

    7. human capital

      The skills, knowledge, and experience possessed by an individual or population, viewed in terms of their value or cost to an organization or country. The government is investing heavily in education to improve the nation's human capital.

    8. economic rationalisation

      The process of reorganizing an economy, industry, or organization to achieve greater efficiency by eliminating waste, reducing costs, and streamlining operations. The government's economic rationalisation program aimed to reduce the budget deficit by closing inefficient state-owned enterprises

    9. socialengineering

      The use of centralized planning in an attempt to manage social change and regulate the future development and behavior of a society.

    10. technicist instrumentalis

      For example, imagine a school has low attendance in a poor community. A technicist-instrumentalist response might be: “Let’s introduce an attendance app, rewards for students, and a new monitoring system.” Those might genuinely help. But the approach becomes technicist if nobody asks things like: Why are students missing school? Is housing insecurity involved? Transport? Racism? Poverty? Is the school itself alienating students? Who defined attendance as the main problem?

    11. Cartesian subject (who believes that he can know himselfand everything else objectively)

      Cartesian subject = the idea of an independent, rational observer who believes they can know themselves and the world objectively, without their own position shaping that knowledge.

    12. universal reason (the idea of onerationality

      Universal reason = the assumption that there is one universally correct way of thinking, knowing and judging what is rational. Critical/postcolonial approaches ask whose idea of ‘rational’ became universal, and what other ways of knowing were excluded.

    13. critical andpostcritical

      Relating to a stage or mode of thought that follows the critical or analytical phase, often characterized by a return to belief, synthesis, or holistic interpretation after the dismantling effects of deconstruction or skepticism.

    14. liberal humanis

      A person who subscribes to the philosophical and political tradition that emphasizes the autonomy, rationality, and inherent dignity of the individual, typically advocating for democratic institutions and social progress.

    15. technicist instrumentalis

      A person whose worldview prioritizes technical proficiency and the use of people or processes strictly as tools to achieve specific, often narrow, outcomes.

    16. Therefore, it is important to remember that maps are useful as long as theyare not taken to be the territory that they represent and are used critically as astarting point of discussion.

      Map ≠ reality. A framework simplifies and interprets reality; use it to begin discussion, not as a final truth.

    Annotators

    1. Tools to download data from the GISCO (Geographic Information System of the Commission) Eurostat database <https://ec.europa.eu/eurostat/web/gisco>. Global and European map data available. This package is in no way officially related to or endorsed by Eurostat.

      Co je tohle?

    2. Pokud ve Zprávě naleznete chybu, kontaktujte: vojtech.kuna@mmr.gov.cz

      ono to vypadá, že je to poslední, ale za tím se generuje ta literatura. Přemýšlím, jak to zviditelnit - co to dát třeba do rámečku? A ještě tam doplňme, jak to citovat...

    1. Struktura těchto výdajů je stabilně spíše poptávkově orientovaná – převažují sociální dávky a podpora stavebního spoření a hypotečních úvěrů, zatímco kapitálové investice do výstavby či obnovy bytového fondu tvoří jen menší část.

      Tohle by mělo být v závěru a shrnutí

    2. (v roce 2023 dosahovala přibližně 7,9 mld. Kč (Ministerstvo financí ČR, 2025))

      tohle je třeba nějak přeformulovat, citace tam vytváří další závorku. Dala bych to prostě jako novou větu a citaci v závorce za tím.

    1. CMS

      改进点在于。 1. stop the world,把和 GC Roots 直接相关的进行标记 2. 解除 stw,应用线程和并发的依赖标记线程一起做; 3. stop the world,把标记期间应用线程影响到的线程进行标记; 4. 解除 stw,和应用程序同时一起清除没被标记的对象。

    1. Por ello, el implícito esencial de la metáfora de la vida como navegación es la posibilidad del naufragio. También, que el mar engulle todas las huellas, ciega los caminos y borra los rastros, que duran apenas el instante de la estela. Tanto quienes alcanzan el puerto seguro como los náufragos, “dejan tras de sí la misma intacta superficie”. Cada existencia, por tanto, transita el mar de la vida por primera vez. Pero la navegación es sobre todo promesa de mundos nuevos, esperanza de alcanzar las tierras prometidas, sospecha de que existen otras maneras de vivir y de pensar a las que sólo se accede soltando amarras de las riberas familiares y de la tierra firme de la costumbre.

      Este fragmento interpela de un modo muy bello a todos aquellos que se sienten migrantes. Tanto si deciden serlo como no serlo.

    2. Blumenberg trabaja sobre lo que llama “metáforas absolutas”, es decir no derivadas, de procedencia inmemorial, reelaboradas continuamente por las generaciones en la medida en que dan cuenta de algún aspecto fundamental de la existencia. Una de ellas, de difícil traducción, es la que en latín ha encontrado una formulación concisa y exacta: navigatio vitae (“navegación de la vida”), es decir la vida como viaje o como navegación incierta, como deriva en lo imprevisible y en lo ignoto. La precariedad y el riesgo constituyen el corazón de esta metáfora del tiempo humano.

      Por un lado esto me hace pensar en que los animales somos dueños del espacio y las plantas del tiempo. Entonces para nosotros vivir es movimiento (y para ellas crecimiento) a través del espacio, que se puede relacionar fácilmente con el viaje. Por otro lado, pienso en la gran adicción a viajar que tenemos actualmente. ¿A lo mejor se debe a estar gran literalidad del pensamiento que tenemos ahora? No somos capaces de las metáforas. Entonces viajamos constantemente pensando que así nos vamos a sentir más vivos porque el viaje es la metáfora de la vida y de cómo superar la precariedad, pero no tenemos herramientas para hacerlo de forma metafórica. Entonces lo hacemos de forma absolutamente literal. Sin encontrar nada valioso en ello, solo distracción y huida.

    1. Suppose he owns a fraction αd of the firm on the receiving ("destination") side of a transferOverføring. A payment or shift of resources between parties that does not affect total surplus because gains and losses cancel out. and a smaller fraction αs of the firm on the paying ("source") side. Shifting value X from the source to the destination — by mispricing a contract between them — yields him a private gain of roughly B≈X(αd−αs). He loses αsX as a shareholderAksjonær. An owner of shares in a corporation who bears residual risk and has voting rights. of the source firm but gains αdX as a shareholderAksjonær. An owner of shares in a corporation who bears residual risk and has voting rights. of the destination firm; the net, X(αd−αs), is positive precisely b

      use real numbers rather than X aplha B etc. make the example concrete

    1. Dear authors, as part of a group activity in our lab, we discussed your very interesting manuscript with the goal of reviewing it as well as improving our reviewing skills. The review below reflects thoughts and comments raised during this exercise. We hope these comments are helpful for strengthening the manuscript. Summary: This manuscript identifies the ER-resident MSP-domain protein MOSPD2 as a regulator of late endosome/lysosome (LE/Lys) homeostasis and proposes that MOSPD2 forms a functionally non-redundant complex with the LE/Lys cholesterol transporter STARD3 at ER-LE/Lys membrane contact sites. Loss of MOSPD2 or STARD3 increases LE/Lys number, shifts their distribution toward the cell periphery, causes cholesterol enrichment at LE/Lys, and impairs endolysosomal fusion. Structure-function and rescue experiments indicate that the FFAT-binding MSP domain of MOSPD2, the phospho-FFAT motif of STARD3, and the cholesterol-binding START domain of STARD3 are required for this function. The authors further show that STARD3 interacts preferentially with MOSPD2 over VAP-A/VAP-B and that STARD3 overexpression can overcome MOSPD2 deficiency, supporting a model in which partner affinity contributes to the functional specialization of ER-organelle contact sites. The study addresses an interesting question in membrane contact-site biology and combines genetic perturbation, imaging, biochemical interaction assays, and live-cell approaches. The comments below focus mainly on the interpretation of cholesterol redistribution, quantitative support for several imaging conclusions and the mechanistic connection between the observed LE/Lys phenotypes. Major comments: 1. Figure 2: the authors conclude that MOSPD2 is required for the normal distribution of free cholesterol based on increased filipin or D4-probe fluorescence within LE/Lys. However, these measurements alone do not distinguish redistribution of an unchanged cellular cholesterol pool from an overall increase in cellular cholesterol caused, for example, by altered uptake, synthesis, or degradation/esterification. Total cellular cholesterol should therefore be quantified biochemically or by lipidomics. 2. Figure 5&6: The authors see several correlated phenotypes after loss of MOSPD2/STARD3—cholesterol enrichment, impaired fusion, increased LE/Lys number, and more peripheral positioning—but the causal relationships between these phenotypes remain unclear. In particular, it is not established whether cholesterol accumulation causes the fusion defect and expansion of the LE/Lys compartment, or whether altered LE/Lys dynamics secondarily cause cholesterol accumulation. A cholesterol-manipulation experiment would help address causality. For example, reducing lysosomal cholesterol accumulation in MOSPD2/STARD3-deficient cells and testing whether fusion and LE/Lys number are rescued would directly connect the lipid phenotype to organelle dynamics. Conversely, an established perturbation that induces lysosomal cholesterol accumulation could be tested for phenocopy of the fusion defect. 3. Figure 1: The conclusion that MOSPD2 is enriched at discrete foci associated with LE/Lys appears to rely largely on representative images and line scans from only a few cells. The apparent overlap in the presented images is modest. This should be quantified across biological replicates and a substantially larger number of cells, using an appropriate contact-site or colocalization metric. 4. Figure 6: Given that contact sites frequently control fission events, it was unexpected that the LE/Ly phenotype in Fig. 6 (CD) represented a fusion defect rather than a fission defect. LE/Lys fusion as compared to fission. Fig. 6D's dual-color dextran pulse-chase test reveals t0 colocalization values that differ considerably between MOSPD2/STARD3 KO and WT, indicating variations in dextran internalization. Therefore, it would be difficult to attribute the LE/lys phenotype to a fusion defect. 5. Figure 9: In Fig. 9G, GFP-MOSPD2 expression is mostly endosomal, with no ER-like reticular morphology, whereas mClover3-MOSPD2 expression in Fig. 1D exhibits ER-like reticular morphology. Does this result from the coexpression of mCherry-STARD3? 6. Figure 10: Given MOSPD2 exhibited a binding affinity to ORP1L comparable to that of STARD3, as shown in supplementary Fig.S4A, so could ORP1L overexpression also override the phenotype caused by MOSPD2 loss. 7. Figure. 10: In Fig. 10F, it is unclear why the authors chose to rescue the MOSPD2 KO phenotype by overexpressing another MSP family member, VAPA. It would be more interesting and compelling to delete VAPA in the MOSPD2 KO background and co-express STARD3 to check if another MSP family member like VAPA can compensate for the loss of MOSPD2 in mediating STARD3 binding and rescue the LE/Ly phenotype.

      Minor comments: 1. Figure 1: The electron microscopy provides potentially valuable ultrastructural information, but the apparent increase in multivesicular/endolysosomal structures is not quantified. Quantification of organelle number, size, and morphology would make better use of these data. 2. Figure 1: LE/Lys positioning analysis: The definition of 0–0.5 of the normalized nucleus-to-cell-edge distance as 'perinuclear' encompasses a large fraction of the cytoplasm. Please justify this threshold and ideally show the full continuous distribution of normalized organelle distances in addition to the binary perinuclear/peripheral classification. 3. Figure S2: Please provide the rationale for selecting MRC5 cells as the second cell type. Statistical comparisons should primarily be made between the siRNA-treated control and target siRNA conditions rather than between target siRNA and non-transfected cells. 4. Figure 2: Cholesterol-probe enrichment within LE/Lys would be more interpretable if normalized to lysosomal area or membrane area, particularly because MOSPD2 loss changes LE/Lys number and potentially morphology. Colocalization/contact between the cholesterol probes and LE/Lys markers should also be quantified. 5. Figure 2: Inclusion of a positive control known to produce lysosomal cholesterol accumulation, such as an NPC-pathway perturbation, would help benchmark the magnitude and appearance of the cholesterol phenotype. 6. Figure 3: Please show that the MOSPD2 deletion and point mutants used for rescue retain the expected subcellular localization, particularly ER localization. 7. Figure 5: The relationship between LAMP1/LAMP2 abundance on Western blot and the imaging-based LE/Lys-number phenotype differs between the two STARD3 KO clones. Quantification of the Western blots and discussion of this clone-to-clone variability would be helpful. 8. Figure 5: A direct rescue of STARD3 KO cells with WT STARD3 for LE/Lys number and positioning would strengthen the initial phenocopy experiments. Although later structure-function experiments provide rescue data for LE/Lys number and cholesterol, the positioning phenotype is not equivalently rescued. 14. Figure 5: The apparent redistribution of the D4 cholesterol probe from the plasma membrane toward intracellular/LysoTracker-positive structures is potentially important and should be quantified directly, for example by comparing plasma-membrane and LE/Lys-associated probe fractions.

      1. Figure 6: The dextran pulse–chase assay suggests impaired convergence/fusion of endocytic compartments in MOSPD2- and STARD3-deficient cells. However, a difference between WT and KO cells is already apparent at t = 0, before the chase period. Please clarify the origin of this initial difference and whether differences in dextran uptake, endocytic trafficking, or the pre-existing LE/Lys compartment could contribute to the subsequent differences in colocalization.
    1. přístup k bydlení je omezen hlavně pro zranitelné skupiny obyvatel

      Tato část věty působí divně a chtělo by to trochu upravit, to spojení přístup je omezen mi nějak nesedí. "...a mezi různými skupinami domácností". Náročnější podmínky v přístupu a udržení bydlení panují zejména u zranitelných skupiny obyvatel jako jsou samoživitelé, senioři...

    1. synaptic pruning

      Definition: natural biological process where the brain removes weak, unused, or excess neural connections (synapses) to make the remaining circuits faster and more efficient.

    1. A female is born with all her eggs already produced.

      Since females hit puberty quicker and have a set number of eggs, is this why it’s common for females over 30s to have less chances of getting pregnant and having more likely dna issues if they do have a kid in those ages?

    2. common, serious problem in modern society due to the prevalence of diets high in fat and lifestyles

      This issue is still a common thing and high percentage in the U.S. currently.

    3. distalproximal development

      Definition: a directional pattern of physical growth and motor skill acquisition where control and maturation proceed from the center of the body outward toward the extremities.

    1. because the public key can be used to verify the signature

      A private key is used to create (or sign) a signature, while the corresponding public key is used to verify that signature.

    1. Domácnosti v družstevních bytech vydávají na bydlení o něco větší část příjmu než vlastníci, ale výrazně méně než nájemníci. Podíl výdajů na bydlení na disponibilním příjmu dosáhl maxima kolem roku 2012 a následně dlouhodobě klesal. Při srovnání je však třeba zohlednit, že do vykazovaných výdajů nejsou zahrnuty splátky anuity. U družstevního bydlení mohou být skutečné peněžní výdaje domácnosti na bydlení vyšší než vykazované výdaje na bydlení, protože část domácností vedle běžných výdajů splácí také anuitu spojenou s úvěrem bytového družstva. U družstevního bydlení se podíl domácností nadměrně zatížených výdaji na bydlení dlouhodobě snižuje. Družstevní bydlení vykazuje výrazně nižší míru přetížení než nájemní bydlení, avšak mírně vyšší než bydlení vlastnické.

      opakuje se?

    2. U družstevního bydlení mohou být skutečné peněžní výdaje domácnosti na bydlení vyšší než vykazované výdaje na bydlení, protože část domácností vedle běžných výdajů splácí také úvěr použitý na pořízení družstevního podílu.

      to platí i u osobního vlastnictví a splácení hypotéky, celkově trochu nebezpečná formulace s ohledem na způsob sledování zatížení výdaji na bydlení 😄

    3. Bytové domy

      osobně dávám významnější kategorie ve skládaném sloupci spíše dolů a nahoře je ten zbytek jako nezjištěno, jiné a pod. mám tendenci sloupcové grafy číst spíše zespodu nahoru, ale nevím, jestli to je obecná tendence

    4. Počet družstevních bytů

      Proč je tady počet družstevních bytů a v grafu pod tím počet bytů v domech vlastněných bytovými družstvy, je v tom rozdíl?

    5. Vývoj právního

      Možná to taky nepoužívám úplně konzistentně, ale asi bychom měli postupně sjednotit názvosloví, myslím, že vlastnické bydlení používáme někde jako osobní vlastnictví a někde ve smyslu owner-occupied housing, tzn. včetně družstevního. V mezinárodních datech se české družstva většinou řadí do vlastnické formy bydlení.

    6. Z ekonomického hlediska se tak více podobá vlastnickému bydlení než klasickému pronájmu.

      Já vlastně asi nevidím, z jakého hlediska je to podobné nájmu 😄 Spíše je to vlastnictví s tím extra mezi subjektem družstva, což může mít výhody a nevýhody, ale s nájemním bydlením, kde platíš přímo jenom za tu službu, tam moc žádnou podobnost nevidím. Asi to není potřeba měnit, jenom tak uvažuju, což jsou nejlepší komentáře.

    7. Anuita

      Není anuita obecně nějaká fixní splátka za časové období? Tady to působí, jako by to měla být celková zbývající nesplacená částka.

    1. Arturo Béjar, a consultant on the wellbeing team of Instagram (which is also owned by Facebook), hit send on a long-gestating email to Zuckerberg and other senior Facebook executives on a similar theme. “I wanted to bring to your attention what I believe is a critical gap in how we as a company approach harm, and how the people we serve experience it,” Béjar began, before laying out some shocking statistics from his team’s survey. For example: 24.4% of 13-15-year-olds on Instagram said they had received unwanted advances, and 21.8% said they had been the target of bullying on the platform in the past seven days. Furthermore, 51% of Instagram users said they had had a bad or harmful experience in the past seven days. Only 1% of them reported it, and of those, 2% had the content taken down – so just 0.02%. This was a far greater degree of harm than the company’s existing measurements suggested. Zuckerberg never replied.

      quite schocking numbers 2021. 25% teenagers received unwanted approaches, 22% target of bullying in the past week. 51% report bad experience within last week. Only 1% reports, and only 2% of that results in takedown, .02% overall.

  2. social-media-ethics-automation.github.io social-media-ethics-automation.github.io
    1. Zack Sharf. ‘Star Wars: The Last Jedi’ Backlash: Academic Study Reveals 50% of Online Hate Caused by Russian Trolls or Non-Humans. October 2018. URL:

      Star Wars the last Jedi movie had a lot of hate received towards it was actually political in nature as Morten Bay, the primary researcher of the academic study would argue. He believes that the movie promoted messaging both the alt right and Russian federation would disagree with, and proves that there were Russian troll accounts pushing the hate on the movie. I find this to be pretty concerning that people would go so far for political gain as to bot or push negative reviews on a movie that loosely alludes to modern day politics.

    1. eLife Assessment

      This study provides important findings on the composition of glutamate receptor (GluR) subunits across diverse Drosophila neuromuscular junctions, spanning larval and adult stages and multiple muscle types. Most previous work has focused on larval body wall muscles, and this study demonstrates that this model does not capture the full diversity of GluR composition across development and tissues. The revised manuscript provides a convincing account of this diversity and will be of broad interest to researchers across neuroscience.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Sustar et al. takes a methodical approach to document the types of glutamate receptor subunits that reside in Drosophila muscles, examining developmental stages spanning from larvae to adults. Prior work thoroughly documented the subunits operating in Drosophila larval body wall muscles. Most subsequent research focused on the glutamate receptor heterotetramers found in the body wall, composed of GluRIIA/C/D/E or GluRIIB/C/D/E subunits, along with auxiliary subunits like isoforms of Neto.

      For the current work, the authors report that the larval muscle glutamate receptor composition is not universal for all Drosophila muscles. They examine the following muscle systems: larval body wall, adult abdomen, adult leg coxa, and adult indirect flight. They also briefly examine adult muscle structures associated with the proboscis, neck, and haltere. The authors find that the receptor subunits in the adult abdomen (mostly) match those in the larval body wall. This makes sense given that the adult abdominal muscles are derived from the larval body wall. Yet not much else matches the larval body wall. For example, all (or most) of the GluRII-type subunits are missing from the adult indirect flight muscles. Leg muscles have GluRII-type subunits, but they do not have all of them expressed prominently, and they are missing GluRIIB. Additionally, leg muscles express a glutamate-gated chloride channel, which could be a source of inhibitory glutamatergic transmission. Interestingly, when it comes to non-abdominal adult muscles, one general theme seems to be an active promoter (GAL4 driver) for the kainate-type glutamate receptor called Clumsy. The authors propose that Clumsy could be key to understanding how functional GluR complexes are assembled in adult insects.

      Strengths:

      (1) Documenting the types of glutamate receptors that operate in diverse insect muscle systems is important because it uncovers fundamental information.

      (2) Much of the prior research focus has been on how the body wall muscle tetramers assemble and operate. It is a strength to demonstrate the other receptor solutions used by adult NMJs.

      (3) The work uses GAL4 drivers and immunohistochemistry (when possible) in combination to draw conclusions.

      (4) The muscle anatomical analyses are high quality. This allows the research group to reach refined conclusions.

      (5) The confocal-level images of synaptic active zones and their apposed glutamate receptor clusters are high quality.

      (6) The dataset adds specific detail and complementary context to bulk transcript analysis collections like FlyAtlas.

      (7) The Discussion section is insightful. It is not simply a recap of the data in the paper. It is also an integration of that data with prior work, and it suggests direct experiments that the research team or others could attempt downstream.

      Weaknesses:

      One can draw expression-level conclusions from these data. But genetic tests (e.g., would clumsy losses of function impair leg muscles?) could help the authors and the field draw stronger conclusions about the roles of some of these glutamate receptor gene products.

      Overall Assessment and Discussion:

      The data in this study are high quality, and the results support the main conclusion: adult muscle glutamate receptor clusters do not recapitulate the "canonical" larval body wall clusters.

      Comments on revised version.

      The authors have done a serious revision of their manuscript, and they have directly addressed context/interpretation-level concerns that I had from the first round.

    3. Reviewer #3 (Public review):

      The Sustar et al. manuscript catalogs glutamate receptor composition across distinct Drosophila NMJs: larval and adult abdominal NMJs, as well as NMJs on adult leg and flight muscles. This work is important and probably overdue. The larval NMJ is the exemplar NMJ in this system, and the identity of "essential" and "alternative" subunits at this stage is assumed by many to hold across developmental stages and NMJ types. Here, the authors show that there is surprising diversification among NMJ types and that the notion of essential/alternative subunits only holds true at larval NMJs.

      The study will generate interest in the Clumsy GluR subunit, which has not been well-characterized at all, but is widely expressed at adult NMJs. They also find striking extrasynaptic expression of glutamate-gated chloride channel GluRClalpha in adult leg and flight muscles, raising questions about its role. The study is interesting, logical, and well-written. The figures are clear, and the discussion was particularly thoughtful.

      Comments on revised version.

      This is a nicely revised manuscript that went a long way toward addressing the concerns in the first round. The response to the reviewers is clear and thoughtful. I think the question is framed very well and the authors do a much better job throughout making clear where the heterologous approaches agree and differ. I also think the Limitations section in the Discussion does a great job putting the work in context.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This study provides important findings on the expression of glutamate receptor (GluR) subunits across developmental stages and muscle types in Drosophila. It shows that adult muscle differs in GluR composition from larval body wall muscles, which have been the focus of most past studies. The study, while convincing, could be strengthened by acknowledging that it relies on heterogeneous methods and the absence of positive signals to infer receptor loss, which limits confidence in some of its claims. The findings illuminate how Drosophila excites muscles in diverse tissue types at different life stages, and are of interest to researchers across neuroscience.

      A note on the eLife assessment

      We are grateful for the assessment and for the reviewers' engagement. The assessment notes that the study "could be strengthened by acknowledging that it relies on heterogeneous methods and the absence of positive signals to infer receptor loss." We have taken this seriously and have addressed it in three ways in the revised manuscript:

      (i) We have replaced claims of absence with "not detected" throughout;

      (ii) We have added a paragraph describing the discrepancies between methods and the specific false-negative risks of each; and

      (ii) We have made explicit the internal positive controls that constrain our negative results, including the reciprocal GluRIIB/GluRIIC labeling within a single femur and the detection of Brp, Neto-β, and GluClα in the same flight muscle preparations in which GluRII subunits were not detected.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Sustar et al. takes a methodical approach to document the types of glutamate receptor subunits that reside in Drosophila muscles, examining developmental stages spanning from larvae to adults. Prior work thoroughly documented the subunits operating in Drosophila larval body wall muscles. Most subsequent research focused on the glutamate receptor heterotetramers found in the body wall, composed of GluRIIA/C/D/E or GluRIIB/C/D/E subunits, along with auxiliary subunits like isoforms of Neto.

      For the current work, the authors report that the larval muscle glutamate receptor composition is not universal for all Drosophila muscles. They examine the following muscle systems: larval body wall, adult abdomen, adult leg coxa, and adult indirect flight. They also briefly examine adult muscle structures associated with the proboscis, neck, and haltere. The authors find that the receptor subunits in the adult abdomen (mostly) match those in the larval body wall. This makes sense given that the adult abdominal muscles are derived from the larval body wall. Yet not much else matches the larval body wall. For example, all (or most) of the GluRII-type subunits are missing from the adult indirect flight muscles. Leg muscles have GluRII-type subunits, but they do not have all of them expressed prominently, and they are missing GluRIIB. Additionally, leg muscles express a glutamate-gated chloride channel, which could be a source of inhibitory glutamatergic transmission. Interestingly, when it comes to non-abdominal adult muscles, one general theme seems to be an active promoter (GAL4 driver) for the kainate-type glutamate receptor called Clumsy. The authors propose that Clumsy could be key to understanding how functional GluR complexes are assembled in adult insects.

      Strengths:

      (1) Documenting the types of glutamate receptors that operate in diverse insect muscle systems is important because it uncovers fundamental information.

      (2) Much of the prior research focus has been on how the body wall muscle tetramers assemble and operate. It is a strength to demonstrate the other receptor solutions used by adult NMJs.

      (3) The work uses GAL4 drivers and immunohistochemistry (when possible) in combination to draw conclusions.

      (4) The muscle anatomical analyses are of high quality. This allows the research group to reach refined conclusions.

      (5) The confocal-level images of synaptic active zones and their apposed glutamate receptor clusters are of high quality.

      Weaknesses:

      (1) There is a strawman argument that is used repeatedly to highlight the significance of the work. The argument implies that the field broadly assumes (or “tacitly” assumes) that the larval body wall glutamate receptor composition extrapolates to all muscles of the fly, including the adult. This reviewer cannot find evidence that this assumption or argument has been explicitly promulgated by others. More likely, others have not examined these muscles directly, and thus, they have not speculated one way or the other.

      The reviewer is correct that we could not find a paper stating explicitly that "the adult NMJ is molecularly identical to the larval NMJ," and we have revised the text to avoid implying that such an explicit claim exists. However, we do argue that larval body-wall findings are routinely presented as properties of "the Drosophila NMJ," "the fly NMJ," or "Drosophila muscle". That unqualified generalization is what has led readers (ourselves included) to extrapolate the canonical larval architecture to the adult.

      Many concrete examples can be found in the literature. He and Dickman (2025, Curr Opin Neurobiol) state that "two subtypes of ionotropic glutamate receptors (GluRs) mediate postsynaptic currents in Drosophila muscle, GluRIIA- and GluRIIB-containing," and describe their regulation by Neto as a general feature of "the fly NMJ". The word "adult" does not appear in the review. Chou et al. (2020, Neural Development) similarly assert that "at the Drosophila NMJ, iGluRs are heterometric tetramers composed of three invariant subunits, GluRIII/GluRIIC, GluRIID, and GluRIIE, as well as one of either GluRIIA or GluRIIB". A highly cited review of these receptors is titled "Glutamate receptors at the Drosophila neuromuscular junction" (DiAntonio, 2006) yet describes only the larval body-wall system, and Harris and Littleton (2015, Genetics) introduce the larval NMJ as "a well-established model glutamatergic synapse" before discussing "the fly NMJ" throughout. Notably, we could not find a review paper about Drosophila neuromuscular control that discussed differences in receptor composition between larval and adult muscles.

      This assumption is related to the genesis of this project, which came about due to the failure of pharmacological tools used in the larva (e.g., philanthotoxin) to block synaptic transmission at adult fly muscles. We (the Tuthill Lab) contacted several principal investigators studying adult fly motor control and within the larval NMJ field (including now co-author Dion Dickman) and asked for advice about why these drugs were not effective in paralyzing leg muscles. Nobody suggested that this could be because adult muscles use different glutamate receptors than the larvae. Indeed, it was only years later when we noticed a high level of GluCl expression in the FlyCellAtlas data that we realized it could be due to differences in receptor expression.

      We agree with the reviewer that adult muscle has rarely been examined directly (Rivlin et al., 2004 being the principal exception, which we cite and now discuss at greater length). Our argument is that this gap has been filled by an untested extrapolation from larvae rather than by data. Of the ~50 papers citing Rivlin et al. (2004), we could not find any that compare glutamate‑receptor subunit expression between larval and adult muscle. Follow‑up work pursued NMJ morphology, remodeling, adhesion molecules, and transporters, leaving the receptor question open.

      Overall, we respectfully disagree with the reviewer that this argument is a strawman, because it reflects our lived experiences as active investigators of adult and larval fly motor control. However, we take the reviewer’s point. We have removed language implying that an explicit model was being overturned, and we now frame the assumption as implicit and not universally held. We note only that the assumption was operative in practice: this project began when pharmacological tools that reliably block transmission at the larval NMJ (e.g., philanthotoxin) failed to paralyze adult leg muscles, and when we consulted colleagues who work on the larval NMJ and adult fly motor control, differences in receptor composition were not among the explanations offered. We recognize this is an anecdote rather than evidence, and we have kept it out of the manuscript.

      (2) Related - to the extent that there has been any tacit assumption about GluRIIC/D/E-anchored receptors being ubiquitous among adult muscles, tacit doubt was raised by Rivilin et al., 2004 (cited by the authors but not as a source of doubt) and by RNAseq datasets like FlyAtlas from 2022 (replicated in Figures s11 and s12). To be clear, the current analysis is better than a bulk transcript analysis from adult tissues. But rather than “overturning” a field or being paradigm-shifting, the current data seem confirmatory of FlyAtlas - and confirmatory of Rivlin et al., 2004, which explicitly concluded that larval and adult NMJs were different

      Rivlin et al. did not examine GluRIIC, D and E, or most of the other glutamate receptor subunits. Nevertheless, we revised the beginning of our Discussion (pg 8) to position our work in the context of their important findings:

      In 2004, Rivlin and colleagues conducted a careful survey of NMJs in seven muscles in the adult Drosophila thorax. They established that adult NMJs are morphologically distinct from larval NMJs, and that GluRIIA and GluRIIB are expressed in muscle-specific subsets, with some muscles mysteriously lacking both. Importantly, they speculated that some adult muscles must be using different glutamate receptor subunits (which were unknown at the time) to build synapses. This problem has remained unexplored for the last two decades.

      As the reviewer requested, we have removed "overturning"/"paradigm-shifting" language throughout. But we respectfully disagree that our analysis is “confirmatory” of the Fly Cell Atlas. The Fly Cell Atlas (Li et al., 2022) does not mention glutamate receptors, or the NMJ, anywhere in its analysis. It is a general single-nucleus transcriptomic resource whose adult-muscle glutamate-receptor content was never extracted, curated, or interpreted. Extracting meaning from the dataset required posing the question, curating the relevant cells and genes, and analyzing them, which is itself a meaningful scientific contribution.

      (3) One can draw expression-level conclusions from these data. But genetic tests (e.g., would clumsy losses of function impair leg muscles?) could help the authors and the field draw stronger conclusions about the roles of some of these glutamate receptor gene products. The current dataset falls short of definitively establishing the function of alternate glutamate receptor modules.

      We intend to perform functional and behavioral experiments with clumsy loss-of-function flies. However, this work will require creation of new genetic reagents and behavioral experiments that make it beyond the scope of this manuscript. We intend to address this important topic in a future paper.

      (4) The confocal synaptic images are of high quality. They are good enough that one could analyze how well Brp directly apposes a specific glutamate receptor subunit for all the associated imaging data underlying Figures S1-S8. No such analysis is done, but understanding what components seem to directly oppose the site of release could lead to better conclusions.

      In all tissues where IIA-E are expressed at the NMJ, the protein localization is directly adjacent to Brp (see representative images in Figure 2B and s5-8). This localization is often clear in single z-slices. We attempted to perform more detailed analyses, like measure the distance between Brp and GluR puncta, but found it very challenging due to the 3D orientation of the NMJ structure. We added a sentence to page 4 of the manuscript describing what we see and the overall similarity of IIA-E localization.

      Overall Assessment and Discussion:

      The data in this study are of high quality, and the results support the main conclusion: adult muscle glutamate receptor clusters do not recapitulate the “canonical” larval body wall clusters. This is important, and the data stand on their own. That is the most important part. This reviewer does have suggestions on how to put the current work in proper context; the current draft appears to overstate the novelty of the findings. Additionally, some sentences need editing for accuracy. None of those concerns impeach the excellent foundational data.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents a broad survey of glutamate receptor composition at the neuromuscular junction in Drosophila across developmental stages and muscle types. The topic is clearly important, and the central observation-that adult muscles differ substantially from the canonical larval NMJ-is interesting and potentially impactful. The dataset is extensive and will likely be of value to the community. However, in my view, there are significant limitations in how the data are generated and interpreted, which at present reduce the strength of the conclusions.

      Strengths:

      The study addresses a relevant and timely question and provides a large and systematic dataset. The finding that adult muscles diverge from larval NMJ organization is compelling and challenges a widely held assumption in the field. The breadth of approaches, including genetic reporters, immunohistochemistry, endogenous tagging, and transcriptomic data, is, in principle, a strong aspect of the work and allows for a broad overview of receptor expression across tissues and developmental stages. Even in its current form, the manuscript provides useful descriptive information that will be of interest to the community.

      Weaknesses:

      A major concern is the reliance on a heterogeneous combination of detection methods (GAL4 reporters, antibody staining, endogenous tagging, and RNA), which are treated largely as equivalent lines of evidence. These approaches differ substantially in what they measure and in their sensitivity and specificity. While convergence across methods can in principle be convincing, here this convergence is often inferred from the shared absence of signal. This is problematic because all methods used are susceptible to false negatives for different reasons. As a result, the repeated conclusion that specific GluR subunits are “absent” from adult muscles, including those previously considered essential, is not fully justified by the data presented.

      This issue is not only theoretical. The manuscript itself seemingly contains examples where methods disagree, demonstrating that detection is incomplete and method-dependent. These discrepancies could be better integrated into the interpretation. Instead, negative results across methods are often taken as strong evidence for absence, which overstates the certainty of the findings.

      In addition, antibody validation appears to rely largely on prior work in larval tissue. Given the structural and biochemical differences in adult muscles, it is not clear that staining performance is equivalent, particularly in cases where the signal is weak or undetected. This further complicates the interpretation of negative results.

      More generally, the manuscript moves in several places from descriptive observations to functional or mechanistic implications that are not directly supported. The suggestion that adult muscles operate with fundamentally different receptor assemblies is intriguing, but remains speculative without functional validation. At a minimum, the distinction between observation and interpretation should be made more explicit.

      I thus think that the current conclusions need to be more carefully constrained. Ideally, the study would be strengthened by at least one functional experiment, such as electrophysiological recordings from adult NMJs or perturbation of candidate receptors like GluClα or Clumsy. This would help to anchor the expression data in synaptic function.

      We thank the reviewer for a careful and constructive critique. We agree that inference from the absence of signal requires caution, and we have revised both the wording and the framing of the manuscript substantially in response. We would also like to draw attention to a feature of our dataset that we did not previously make explicit, and which we believe speaks directly to the core concern: a number of our negative results are internally controlled by positive signal from the same reagent, in the same tissue, in the same preparation.

      (i) In the adult femur, anti-GluRIIB labels NMJs on tibia extensor fibers but not on accessory tibia flexor fibers within the same bisected femur, processed, stained, and imaged together (Figure 3C-D).

      (ii) Anti-GluRIIC shows the reciprocal pattern in the same preparations (Figure 3E-F). The failure to detect GluRIIC in the tibia extensor therefore cannot be attributed to a general failure of that antibody in adult tissue, to fixation, or to reagent penetration, because the same antibody labels NMJs tens of microns away in the same section.

      (iii) In indirect flight muscle, where we detected none of the five GluRII subunits, we reliably detected Brp, Neto-β::sfGFP, and GluClα:V5 at and around NMJs in the same preparations (Figures 2B, 2D, 4C). Reagent access, epitope preservation, and NMJ identification are therefore all demonstrated in the same tissue where our negative results are strongest.

      These controls do not exclude the possibility that receptors are present below our detection threshold, and we have accordingly replaced "absent" with "not detected" throughout the manuscript. They do, however, argue that the muscle-specific differences we report reflect genuine biological variation rather than method failure. We have added a paragraph to the Discussion making this argument explicitly, alongside a new paragraph describing the discrepancies between methods, and we have added a description of antibody validation to the Methods.

      In summary, this is an interesting and potentially important study, but the current manuscript somewhat overinterprets heterogeneous and partly indirect evidence. It will already be useful in its present form, but could be more convincing if the authors more rigorously account for methodological limitations and moderate their claims accordingly.

      Reviewer #3 (Public review):

      The Sustar et al. manuscript catalogs glutamate receptor composition across distinct Drosophila NMJs: larval and adult abdominal NMJs, as well as NMJs on adult leg and flight muscles. This work is important and probably overdue. The larval NMJ is the exemplar NMJ in this system, and the identity of “essential” and “alternative” subunits at this stage is assumed by many to hold across developmental stages and NMJ types. Here, the authors show that there is surprising diversification among NMJ types and that the notion of essential/alternative subunits only holds true at larval NMJs.

      The study will generate interest in the Clumsy GluR subunit, which has not been well-characterized at all, but is widely expressed at adult NMJs. They also find striking extrasynaptic expression of glutamate-gated chloride channel GluRClalpha in adult leg and flight muscles, raising questions about its role. The study is interesting, logical, and well-written. The figures are clear, and the discussion was particularly thoughtful. I have a couple of comments that the authors could consider.

      (1) They cite Rivlin et al., (2004) in the Introduction as the sole previous study to investigate the molecular composition of adult NMJs, but do not mention this work again. In the Discussion, it would be helpful to compare/contrast their finding with those of the earlier work.

      Thank you for the suggestion. We have revised the beginning of our Discussion (pg. 8) to position our work in the context of Rivlin et al.’s important findings.

      (2) Were these analyses done in adults of consistent ages? It seems possible that the GluR subunit composition could be different in very young adults or in aged flies. The age of the animals should be mentioned in the Methods.

      Thank you for noticing this gap. We added age of flies to the Methods section, including the important detail that in cases where we failed to detect any GAL4 signal (e.g., GluRIID-GAL4), we additionally imaged very young and very old flies (0-2 hours old and 20-day-old). It did not change the results.

      (3) The broad expression of GluCl:V5 in adult leg and flight muscles is surprisingly robust and appears to light up the edges of all muscle fibers. Would the authors comment on the controls that were done to ensure that this staining is real and specific to animals carrying that V5 endogenous tag?

      We added a sentence to the results section (pg 8) to highlight the V5 negative staining controls that are shown in Figure s15.

      (4) The snRNAseq data in Figure S12 differ a bit from the IHC/GAL4 data summarized in the table in Figure 2. In particular, the data suggests that Ukar and Grik are widely expressed in adult muscles. Is there a reason not to include an “snRNA seq” column in Figure 2 alongside the data from GAL4 lines and IHC? To my mind, it is about as reliable as GAL4 lines that often capture only a subset of the full expression pattern. In this case, the snRNAseq data suggest that Ukar/Grik are likely at adult flight muscle NMJs, which might be important since NMJ was negative for everything except Neto-beta by IHC.

      We liked your suggestion to put these data together in one figure, and we tried to incorporate this during our revisions. However, because of the difficulty of combining multiple types of data into one table, we ultimately decided to keep them separate to avoid any confusion. The GAL4 analysis is qualitative (no expression, weak expression, strong expression) in percent of muscles. The RNA-seq data is quantitative and measured in fraction of muscle cells. Figure s13 now shows RNA-seq results in a table format that is parallel to the table with GAL4 results in Figure 2.

      We also added a paragraph to the end of the discussion to mull several discrepancies in results (e.g. Grik and Ukar) when using different methods.

      We did observe some discrepancies across the different experimental approaches. For GAL4 reporter lines, we used T2A gene trap lines where possible (for 15 of the 16 GluR genes) to determine receptor expression. However, some of these reporters were extremely weak and may have resulted in false negatives. For example, we did not detect Grik-GAL4 or Ukar-GAL4 expression in any tissue (Figure 2C), despite RNA-seq evidence that these are expressed robustly in adult muscles (Figures s11, s12). We do not have antibodies or protein reporters for these genes to help resolve this. Another discrepancy was with GluRIIE. In this case, antibody staining and RNA-seq did not detect GluRIIE in flight muscles (Figures s8, s11), while the GAL4>GFP reporter showed extremely faint signal in the flight muscles. Because of the marginal signal and the discrepancy with the other methods, we believe this could likely be a false positive. The table in Figure 2 does not effectively capture the nuance of this otherwise confusing result. Overall, however, we found that results from different approaches generally agreed and revealed intriguing differences between different muscle types.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Most of these recommendations are for text edits. Where experiments are noted (#8 and #9 below), they could be helpful, but they are not essential because the study has value as is.

      (1) It would be appropriate to scrub the entire manuscript of statements that say that the work is overturning an incorrect model about larval GluRII receptor tetramers working the same way in adult muscles. First, it is not necessary. But more importantly, by this reviewer’s reading, there is not an explicit model promoted by others that the larval body wall findings extrapolate to other muscles. The better way to frame the work is to state that it clearly shows the field has a lot of work to do to properly characterize glutamate receptor composition and function in Drosophila adults. Places in the text that the authors should edit include:

      - “...challenging assumptions about the uniformity of neuromuscular function” (Abstract)

      We deleted this clause.

      - “...it is often tacitly assumed...This extrapolation overlooks...during metamorphosis.” (Introduction)

      New wording: … “…it has commonly been assumed…” (Abstract) and “…we and others assumed…” (Intro)

      - “(With one notable exception, Rivlin et al., 2004)” (Introduction). This reference deserves more than an obscure parenthetical because it explicitly concluded that larval and adult NMJs differ.

      As described above, we revised the beginning of our Discussion (pg 8) to position our work in the context of Rivlin et al.’s important findings.

      In 2004, Rivlin and colleagues conducted a careful survey of NMJs in seven muscles in the adult Drosophila thorax. They established that adult NMJs are morphologically distinct from larval NMJs, and that GluRIIA and GluRIIB are expressed in muscle-specific subsets, with some muscles mysteriously lacking both. Importantly, they speculated that some adult muscles must be using different glutamate receptor subunits (which were unknown at the time) to build synapses. This problem has remained unexplored for the last two decades.

      - “Collectively, these results overturn the assumption of NMJ uniformity...” (Introduction)

      New wording: Collectively, these results reveal a diversity in NMJ composition”

      - “Our findings challenge the assumption...” (Conclusions)

      We deleted this clause.

      (2) (Related) The authors quibble in a couple of places that the GluRIIC/D/E subunits have been “previously considered essential for viability and NMJ function.” This reviewer’s understanding of those subunits has always been that they are essential for life, and indeed, they are essential because null mutants arrest as embryos. The places where the text seems to combine viability and NMJ function (including the Abstract) as incorrectly assigned roles for these subunits, should be more carefully edited. This reviewer agrees with the “NMJ function” conclusion since the authors are considering so many NMJs. Viability still holds.

      We agree with the reviewer, and we thank them for catching an imprecision that recurred throughout the manuscript. GluRIIC, GluRIID, and GluRIIE are essential for viability, and nothing in our data contradicts that conclusion.

      We think the point we were trying to make is narrower and sharper when stated precisely. Null mutants for these subunits arrest as embryos — that is, they die before the adult leg and flight muscles are formed. Those experiments therefore established a requirement for viability and for larval NMJ function, but they could not, in principle, have tested whether the same subunits are required at NMJs that arise only during pupal development. Our data do not revise the viability conclusion. They show that the subunit requirement established at the larval body-wall NMJ does not generalize to every NMJ in the animal.

      We have decoupled "viability" from "NMJ function" throughout the manuscript (Abstract, Introduction, Discussion) and have added a sentence to the Discussion making the developmental logic explicit:

      Importantly, because these mutants die before metamorphosis, these experiments established a requirement for viability and for larval NMJ function, but could not test whether the same subunits are required at adult NMJs, which are formed during pupal development.

      (3) “...the physiology and molecular architecture of adult fly NMJs remain largely unexplored...” The authors cite the Rivilin paper that examined the prothorax. But physiology has certainly been examined at DLMs as well as proboscis NMJs.

      As the manuscript says on page 2, the topic is largely unexplored compared to the larva. The Rivlin paper is an excellent contribution but was conducted over 20 years ago with more limited genetic reagents. There were also no follow-up studies. In the same amount of time, over 1000 papers have been published regarding the Drosophila larval NMJ.

      The reviewer is correct that physiology has been performed in adult DLMs and proboscis muscles. These studies, however, have focused on other biological questions and did not investigate the makeup of the NMJ or manipulate glutamate receptors with pharmacology or genetic tools. Our study reveals the need for such experiments.

      (4) “Surprisingly, we found that none of the primary GluRII subunits was strongly expressed in adult indirect flight muscles...” This sentence was difficult to parse when simultaneously examining the large dot for GluRIIE-Gal4 in the IFMs (Figure 2C), until the realization that the blank IHC result was driving the conclusion (no expression). The authors may want to delineate Gal4 vs. IHC for each dataset to avoid confusion for readers.

      We agree with the reviewer that this result is confusing. GluRIIE-GAL4 expression was a case in which a dot in the table didn’t capture the subtlety of the result. The GluRIIE-GAL4 was expressed very faintly in the IFM, to a degree that we almost called it negative.

      We added a paragraph to the Discussion in which we reflect on some data discrepancies:

      We did observe some discrepancies across the different experimental approaches. For GAL4 reporter lines, we used T2A gene trap lines where possible (for 15 of the 16 GluR genes) to determine receptor expression. However, some of these reporters were extremely weak and may have resulted in false negatives. For example, we did not detect Grik-GAL4 or Ukar-GAL4 expression in any tissue (Figure 2C), despite RNA-seq evidence that these are expressed robustly in adult muscles (Figures s11, s12). We do not have antibodies or protein reporters for these genes to help resolve this. Another discrepancy was with GluRIIE. In this case, antibody staining and RNA-seq did not detect GluRIIE in flight muscles (Figures s8, s11), while the GAL4>GFP reporter showed extremely faint signal in the flight muscles. Because of the marginal signal and the discrepancy with the other methods, we believe this could likely be a false positive. The table in Figure 2 does not effectively capture the nuance of this otherwise confusing result. Overall, however, we found that results from different approaches generally agreed and revealed intriguing differences between different muscle types.

      (5) “GluClalpha in adult muscles establishes the molecular identity of inhibitory glutamate responses described decades ago in other insects...” This reviewer really likes this idea. But the current data fall far short of the conclusion in this sentence.

      You are correct. We tempered this statement to “The identification of extrasynaptic GluClα in adult muscles establishes a possible mechanism for inhibitory glutamate responses described decades ago […]”

      (6) “One past study (Han et al., 2015) explored many combinations....” Nothing in this sentence is untrue. This reviewer just wanted to point out that one interesting observation from the Han paper was that an A/E combination alone yielded significant function in the heterologous system. To the extent that A and B have been described as competitive for the same C/D/E core, it is not a stretch to imagine that a B/E combination would have significant function, just like A/E. Or B/D/E as well.

      We thank the reviewer for this observation, which we think materially strengthens the plausibility of our interpretation, and we have incorporated it into the Discussion. Han et al. (2015) found that GluRIIA and GluRIIE together were sufficient to reconstitute functional receptors in a heterologous system, without GluRIIC or GluRIID. Because GluRIIA and GluRIIB are thought to compete for the same position within the tetramer, there is no obvious reason a GluRIIB/IIE or GluRIIB/IID/IIE complex should not also be functional. This provides a concrete precedent for exactly the kind of non-canonical combinations our expression data imply — for example, GluRIIB in the absence of detectable GluRIIC in tibia extensor fibers. We have added the following to the Discussion (p. 9):

      Notably, Han et al. (2015) found that GluRIIA and GluRIIE alone were sufficient to reconstitute functional receptors, without GluRIIC or GluRIID. Because GluRIIA and GluRIIB are thought to occupy the same position within the tetramer, this raises the possibility that GluRIIB-containing complexes lacking GluRIIC — such as the GluRIIB/IID/IIE combination implied by our data in the tibia extensor — could likewise be functional. Neto-β, which is required for receptor gating and synaptic clustering (Han et al., 2024; Kim et al., 2012), is expressed at all adult NMJs we examined (Figures 2D, 3B), so the auxiliary subunit requirement for such complexes would be satisfied. Direct tests of these subunit combinations by heterologous reconstitution are an important next step.

      (7) GluClalpha roles: The idea of a negative feedback controller is a good one. An alternative idea can be gleaned from the NMJ anatomy of circuits like the Mauthner Cell Escape Circuit in zebrafish. In that circuit, the fish have excitatory connections on some muscles and simultaneous control of inhibitory connections wired for other muscles to enable fast changes of direction.

      We thank the reviewer for this interesting comparison. We think the Mauthner arrangement is unlikely to apply here, for two reasons. First, GluClα appears to be expressed broadly across the leg and flight muscles we examined, rather than in a subset complementary to the excitatory receptors. Second, and more importantly, GluClα is extrasynaptic: it is distributed along the periphery of the muscle fiber and is not apposed to Brp-positive active zones (Figure 4C-D). Target-specific inhibitory transmission of the kind used in the Mauthner circuit requires inhibitory receptors positioned at synapses, whereas the localization we observe is better suited to tonic or slowly varying modulation of muscle excitability. We have added a sentence to the Discussion making this distinction, since the reviewer's question is one that readers will also have. And add to the manuscript, at the end of the GluClα hypothesis paragraph in the Discussion:

      We note that this arrangement differs from circuits in which excitatory and inhibitory transmission are targeted to different muscles to produce rapid directional control, as in the Mauthner cell escape circuit of larval zebrafish. While most arthropods possess GABAergic motor neurons that directly inhibit leg muscles, holometabolous insects such as Drosophila do not (Witten and Truman, 1998). And because GluClα is extrasynaptic rather than apposed to active zones, it is unlikely to support fast, target-specific inhibitory transmission.

      Two optional experimental suggestions:

      (8) The authors’ findings with Clumsy expression along with Clumsy message expression in the FlyAtlas RNAseq datasets - and finally, the unpublished work on Clumsy - all collectively suggest that it is a great target for adult knockdown. It is not required, but the authors’ conclusions would be bolstered if they were to test RNAi against Clumsy in chosen adult NMJs and check behavior (e.g., walking?) for defects.

      We intend to perform functional and behavioral experiments with clumsy loss-of-function flies. However, this work will require creation of new genetic reagents and behavioral experiments that make it beyond the scope of this manuscript. We intend to address this in a future paper.

      (9) The confocal synaptic images are of high quality. They are good enough that one could analyze how well Brp directly apposes a specific glutamate receptor subunit for all the associated imaging data underlying Figures S1-S8. No such analysis is done, but understanding what components seem to directly oppose the site of release could lead to better conclusions.

      In all tissues where IIA-E are expressed at the NMJ, the protein localization is directly adjacent to Brp (see representative images in Figure 2B and s5-8). This localization is often clear in single z-slices. We attempted to perform more detailed analyses, like measure the distance between Brp and GluR puncta, but found it very challenging due to the 3D orientation of the NMJ structure. We added a sentence to page 4 of the manuscript describing what we see and the overall similarity of IIA-E localization.

      Reviewer #2 (Recommendations for the authors):

      The authors should revise the wording throughout to more clearly distinguish between “absence” and “not detected,” in particular for statements regarding GluRII subunits in adult muscles. In the current form, some claims appear stronger than supported by the methods.

      Thank you for this important suggestion. We replaced “absence” with “not detected” throughout the manuscript. We also note this distinction explicitly in the Discussion.

      It would also improve clarity if the different experimental approaches (GAL4, antibody staining, tagging, RNA-seq) were presented more explicitly separated, rather than combined into a single readout. This would allow a better assessment of the evidence.

      We have kept the Results organized by finding rather than by technique, because the majority of our results were concordant across approaches and a technique-by-technique presentation would fragment each biological conclusion across four sections. However, we agree that the evidence should be separable by the reader, and it now is: Figure 2C reports GAL4 and IHC in separate columns, Figure s13 presents the RNA-seq results in a parallel table format, and the new Discussion paragraph identifies each case in which the methods disagreed and states which result we consider likely to be a false positive or false negative. Where a single method underlies a conclusion, we now say so in the text.

      The discrepancies between methods (e.g., RNA-seq vs GAL4 expression) should be discussed more directly, and the possibility of false negatives, especially for antibody staining in adult tissue, should be acknowledged more clearly.

      Thank you. Other reviewers also noticed this weakness of our manuscript, and as noted above, we added a paragraph in the discussion to highlight several important discrepancies in methods and possibilities of false positives.

      A functional experiment is not strictly required, but would clearly strengthen the manuscript. For example, electrophysiology at adult NMJs or perturbation of candidate receptors (e.g., GluClα or Clumsy) would help to link expression to function.

      We intend to perform functional and behavioral experiments with clumsy loss-of-function flies. However, this work will require creation of new genetic reagents and behavioral experiments that make it beyond the scope of this manuscript. These experiments and the accompanying behavioral analysis will form a separate study, which will be the foundation of a future paper.

    1. eLife Assessment

      Overall, this is an important manuscript that delivers an invaluable community resource. The execution is relatively simple, but the evidence of this resource is solid. The effort invested in generating the knockout lines that for validation experiments is a clear strength of the study, but it doesn't substitute for a careful evaluation of the same reagents in the end-users' specific protocols before making any biological statement.

    2. Reviewer #1 (Public review):

      Summary:

      The authors address the lack of validated tools for detection and quantification of proteins associated with amyotrophic lateral sclerosis (ALS) through an extensive screening of 303 commercially available antibodies to 33 protein targets. Their ALS-Reproducible Antibody Platform (ALS-RAP) delivers a validated antibody toolbox for ALS research, which will provide an advantageous starting point for researchers in this field. Ayoubi R. et al. showcase the characterization workflow, presenting as an example the characterization of antibodies targeting Galectin-1, encoded by the LGALS1 gene. A selection of these antibodies was also used to profile protein levels across human induced pluripotent stem cell (iPSC)-derived and primary neurological cell types, and the fundings support that ALS disease mechanism involves both neuronal and glial cells.

      Strengths:

      The knockout (KO)-based approach is definitively the major strength of this study, providing a high level of confidence in the data collected in human induced pluripotent stem cell (iPSC)-derived and primary neurological cell types. Important is also the focus on renewable reagents (monoclonal and recombinant antibodies). The extensive characterization of this set of antibodies will benefit any scientist interested in any of the 33 targets proteins, even in fields other than neuroscience.

      The authors perform an interesting protein profiling study assessing 27 proteins, comparing RNA and protein expression data, and using two independent WB preparations of the same cell types. The conclusions that can be drawn from this first assessment might not be final, but data are compelling, because they have been collected with reliable and validated antibodies.

      Another strength of this work is the data dissemination strategy, which includes the Only Good Antibodies (OGA) platform, where YCharOS data are curated and presented in an easy and intuitive manner that facilitates antibody selection by the end user for WB, IP and IF applications.

      The authors mentioned the development of single-chain variable fragment (scFv) recombinant antibodies raised by the SGC against the six proteins (ANXA11, OPTN, MATR3, PFN1, UBQLN2 and VCP) that had limited renewable antibodies that are commercially available. The development was optimized to generate antibodies particularly suitable for IP, and the clone selection process was carried out using IP coupled to mass spectrometry. They provide details on the screening process for the SGC-generated scFv antibodies, plus share that scFv sequences.

      Weakness:

      The protein profiling study is limited to WB data, and the authors did not provide any explanation on why there was no integration with IP and IF data, not even for those targets that have validated antibodies. Also, not all the cell types have been screened by chemiluminescence-based detection and by fluorescence-based WB, and the authors do not elaborate on the reason for such choice. Technical and practical considerations for such decisions could have been briefly described.

    3. Reviewer #2 (Public review):

      Overall, this is a solid manuscript that delivers an important community resource. The execution is relatively simple, but the value is real, the work is rigorously performed, and the open dissemination through Zenodo, the F1000Research YCharOS Gateway and OGA is well executed. The effort invested in generating the knockout lines for validation experiments is a clear strength of the study.

      Comments on revised version.

      I thank the authors for their responses. A number of the smaller points I raised are now resolved. However, I am disappointed to see that the more substantial concerns, which relate to the cell-type expression claims running through the Results and the Discussion, remain unaddressed. My original point was that these claims are not quantitatively supported and that the identity of the iPSC-derived cell types used to make them has not been documented, and I do not think this revision changes that. Most of what I asked for was quantification of experiments the authors have already performed, or quality control on the cell types themselves, and this has been answered with scope arguments rather than with data. The other points I originally raised remain.

      I give further details on my two main remaining concerns below:

      (1) The cell-type expression claims across the manuscript are still supported by qualitative Western blot images only. These include statements such as "higher signal in glial populations," "enriched in neuronal lineages," "predominantly detected in microglial populations," and the "convergence on multicellular disease mechanisms" narrative in the Discussion. I understand the authors' decision to keep ALS-RAP as a tool and resource paper rather than expand the experimental work. However, if that is the framing, the biological interpretations of these expression data need to be removed. In their response, the authors themselves acknowledge that "quantification across multiple biological replicates would be required to establish definitive differences in protein abundance between neuronal and glial cell types." As long as this quantification is not provided, every claim of cell-type-specific expression should be removed from the Abstract, Results and Discussion, and the data presented as a descriptive survey.

      (2) Related to the previous point, I do not think that quality control of the non-microglia iPSC lineages is beyond the scope of the study, as argued in the response. For ALS-RAP to be adopted as a resource, other laboratories need to trust that the authors' iPSC-derived neurons, astrocytes and oligodendrocytes are really neurons, astrocytes and oligodendrocytes. The fact that the differentiation protocols have been previously published does not confirm that they have been correctly implemented here, and this is precisely the kind of quality control that should be documented in the paper. In their response, the authors list the single markers already shown in Figure S3 (ISL1, TH, MBP, GFAP, CD44) as the answer to my concern, but a single positive marker per lineage does not establish purity.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The authors mentioned the development of single-chain variable fragment (scFv) recombinant antibodies raised by the SGC against the six proteins (ANXA11, OPTN, MATR3, PFN1, UBQLN2 and VCP) that had limited renewable antibodies that are commercially available. The development was optimized to generate antibodies particularly suitable for IP, and the clone selection process was carried out using IP coupled to mass spectrometry. Even though the generation of these novel reagents is not the focus of this work, the authors do not provide any data on this aspect.

      The scFv reagents represent an important component of this resource and warranted additional description. In the revised manuscript, we expanded the section describing their origin, selection strategy, and intended use. We also added an updated reference describing the scFv discovery methodology, together with a brief summary of the workflow. In addition, the sequences of all six scFvs (ANXA11, OPTN, MATR3, PFN1, UBQLN2, and VCP) are now provided in Supplementary Table 2. During the revision process, we also identified that two scFv plasmids were missing from Add gene; these have now been deposited, making all six reagents and their sequences publicly available.

      (2) The protein profiling study is limited to WB data, and the authors did not provide any explanation on why there was no integration with IP and IF data, not even for those targets that have validated antibodies. Also, not all the cell types have been screened by chemiluminescence-based detection and by fluorescence-based WB, and the authors do not elaborate on the reason for such a choice.

      ALS-RAP was conceived as an open science resource for the ALS community. Given the breadth of the project (33 ALS-associated proteins), it is unrealistic for a single laboratory to investigate the biochemical and cellular biology of every target. We therefore generated and openly disseminated antibody characterization data across three complementary applications—western blot (WB), immunoprecipitation (IP), and immunofluorescence (IF)—to provide the community with validated tools that can support diverse downstream studies, including protein abundance, protein interactions, and subcellular localization.

      The primary objective of the protein profiling study presented in this manuscript, however, was to compare relative protein abundance across multiple neuronal and glial cell types. Western blotting was selected because it provides a robust and semi-quantitative measurement of protein abundance across biological samples, making it well suited for systematic comparisons between cell types. In contrast, IP and IF were primarily included as antibody validation assays and as community resources for future mechanistic studies rather than for quantitative expression profiling.

      Generating the large quantities of iPSC-derived neuronal and glial cells required to profile 33 ALS-associated proteins is technically demanding and resource-intensive. Consequently, a limited number of protein-cell type combinations are absent from one of the two detection modalities because insufficient lysate remained for repeat analyses. Importantly, every protein was evaluated across the complete panel of neurological cell types in at least one western blot detection modality. In most cases, the omitted samples corresponded to cell types in which the target protein was undetectable or expressed at very low levels in the initial experiment.

      Reviewer #2 (Public review):

      (1) The rationale for the selection of these 33 genes is insufficient. The authors lean on the Nijs & VanDamme classification and on PubMed entry counts, but the number of PubMed entries is not a meaningful criterion for what constitutes an important ALS protein - some of the most disease-relevant genes are precisely those with fewer publications, while heavily cited genes such as CAV1 carry weak ALS-specific evidence. The authors should provide a more transparent and biologically motivated rationale for inclusion and exclusion (ClinGen evidence tier, replicated GWAS signals, large meta-analyses, ALSoD) and explain why specific risk genes outside this list were not part of ALS-RAP.

      We thank the reviewer for this comment and apologize if our rationale for target selection was not sufficiently clear. We agree that the number of PubMed entries should not be interpreted as a criterion for inclusion. Rather, we included this metric as a descriptive measure of how extensively each gene has been studied in the context of ALS. Specifically, comparing the number of publications for Gene versus Gene + ALS provides an indication of the relative attention a target has received within the ALS field, but it was not used to prioritize genes for inclusion.

      The selection of the 33 ALS-RAP targets was instead based on established genetic evidence for disease association. Specifically, we prioritized ALS-associated genes supported by replicated human genetic studies, including rare variants with high, intermediate, or modest effect sizes, following the classification framework proposed by Nijs and Van Damme. As described in the manuscript, 25 of the 33 targets have also been independently curated by the Clinical Genome Resource (ClinGen) ALS Gene Curation Expert Panel and classified as having evidence supporting an association with ALS (Table 1). The remaining genes were selected based on published human genetic studies and emerging evidence supporting their relevance to ALS biology.

      (2) "107 of 231 (46%) demonstrated specific target staining in IF." The criteria used to define "specific target staining" at the IF level are not stated. From the Galectin-1 example, the mosaic WT/KO strategy provides a binary readout, but for proteins with low expression, weak punctate staining or unusual subcellular distributions, a single threshold is unlikely to capture specificity uniformly across 231 antibodies.

      We thank the reviewer for highlighting this omission. We agree that the criteria used to classify antibodies in each application should be explicitly stated. In the revised manuscript, we now define the classification criteria for Western blot (WB), immunoprecipitation (IP), and immunofluorescence (IF) in the Methods section. Briefly, WB antibodies were considered target-selective when the predominant immunoreactive band at the expected molecular weight was substantially reduced or absent in the corresponding knockout (KO) lysate. IP antibodies were considered successful when they specifically enriched the target protein from wild-type lysates, as assessed by immunoblotting of the immunoprecipitated material. For IF, antibodies were considered specific when the mean fluorescence intensity in wild-type cells was at least 1.5-fold higher than that measured in KO cells using the mosaic WT/KO assay. A minimum of 250 WT and 250 KO cells were quantified for each antibody. These criteria were applied uniformly across all antibodies to provide a standardized and objective benchmarking framework. We have also added references to the detailed YCharOS characterization workflows previously described in Ayoubi et al., Nature Protocols (2025).

      (3) Several claims in the manuscript depend on differential protein abundance across cell types. As presented, these claims are supported by qualitative Western blot images only. They should be substantiated by quantification across multiple biological replicates.

      The expression patterns observed were reproducible across two independent batches of iPSC-derived neuronal and glial cell types generated using the same differentiation protocols. In addition, several proteins enriched in iPSC-derived microglia also showed enrichment in primary human microglia and monocyte-derived macrophages, providing independent support for these observations.

      We agree that quantification across multiple biological replicates would be required to establish definitive differences in protein abundance between neuronal and glial cell types. However, this was not the objective of the present study. Rather, our goal was to perform an initial survey of ALS-associated protein expression using rigorously validated antibodies and to identify expression patterns that may warrant further investigation.

      ALS-RAP was conceived as an open science resource for the ALS community. Given the breadth of the project (33 ALS-associated proteins), we do not envision that a single laboratory can comprehensively define the expression, localization, and function of every target across all relevant cell types, tissues, and disease states. Instead, we provide validated reagents together with an initial characterization of protein expression in well-defined iPSC-derived neuronal and glial cell models to enable and accelerate future studies by the broader research community.

      (4) This manuscript represents a unique opportunity to address antibody recognition of splicing variants, which is something of considerable value to the community. For each target, the predicted isoforms in Ensembl could be cross-referenced against the observed bands, and the pattern of bands compared across cell types could be informative about which isoforms each antibody captures. This would convert ambiguous "extra bands" into useful biological information and would substantially increase the value of the resource. I strongly encourage the authors to include this analysis.

      We agree that defining which protein isoforms are recognized by individual antibodies would be valuable to the research community. However, we believe that such an analysis would be difficult to interpret and potentially misleading based on the data generated in the present study.

      All antibodies were initially characterized in a common cancer cell line, and the best-performing reagents were subsequently used to profile ALS-associated proteins in iPSC-derived neuronal and glial cell types. Additional bands observed in neurological cells, but not in the original characterization cell line, may reflect several nonmutually exclusive possibilities, including cross-reactivity with proteins specifically expressed in neurological cells, recognition of cell type-specific splice isoforms, or post-translationally modified forms of the target protein.

      Without knockout validation in each neurological cell type, these possibilities cannot be distinguished confidently.

      In addition, protein migration on SDS-PAGE frequently differs from the theoretical molecular weight predicted from amino acid sequence alone. Consequently, assigning individual immunoreactive bands to Ensembl-predicted isoforms based solely on apparent molecular weight would be speculative and could lead to incorrect conclusions. We therefore believe that a rigorous isoform analysis would require complementary experimental approaches beyond the scope of the present resource study.

      (5) The iPSC-derived microglia receive a comprehensive QC panel (IBA1/PU.1 IF, CD45/CD11b flow, qRT-PCR for nine canonical markers; Figure S4), which allows the reader to assess culture purity. The other iPSC-derived lineages - motor neurons, dopaminergic neurons, oligodendrocytes and astrocytes- are validated by a single marker each in WB (Figure S3) without purity quantification. Given that several conclusions of the manuscript rest on the cell-type-specific detection of ALS-associated proteins, equivalent quality control should be performed for the other lineages so that the reader can evaluate the purity of each preparation.

      We thank the reviewer for this important comment and agree that documenting the identity of iPSC-derived cell populations is essential given their inherent heterogeneity. The differentiation protocols used in this study follow well-established, published methods that have been extensively validated by their developers and the broader stem cell community. We confirmed the identity of each differentiated lineage using established lineage-specific markers, including ISL1 for motor neurons, TH for dopaminergic neurons, SATB2 and CTIP2 for cortical neurons, MBP for oligodendrocytes, GFAP and CD44 for astrocytes, and CD11b and Iba1 for microglia.

      We agree that comprehensive quantitative purity analyses for every differentiated lineage would further strengthen the study. However, we believe that these experiments fall beyond the scope of the present work, whose primary objective was to establish a validated antibody resource rather than to optimize or benchmark differentiation protocols. We also note that our conclusions are intentionally limited to broad protein expression patterns rather than definitive quantitative comparisons between cell types.

      Finally, the inclusion of primary human microglia, fetal astrocytes, and monocyte-derived macrophages provides an independent layer of biological validation. The concordance of expression patterns observed between these primary cells and the corresponding iPSC-derived populations supports the robustness of our principal observations.

      (6) The robustness of the resource would be substantially increased by validating at least a subset of the targets in a second iPSC background, in at least some of the cell types analysed.

      We agree that validating antibody performance and protein expression across multiple iPSC genetic backgrounds would further strengthen the resource. However, we believe that this objective falls beyond the scope of the present study, whose primary aim was to establish a rigorously validated antibody toolbox for ALS-associated proteins and to provide an initial survey of their expression in representative neuronal and glial cell types.

      We envision ALS-RAP as the foundation of a broader community effort. Rather than comprehensively evaluating all 33 proteins across multiple genetic backgrounds and tissues, our goal is to provide validated, renewable reagents that will enable such studies by the broader ALS research community.

      (7) The newly developed SGC scFv antibodies are arguably the most novel reagent contribution of this manuscript, yet they receive a single sentence in the body of the paper. A more thorough description is warranted.

      We agree that the scFv reagents represent an important component of this resource and warranted additional description. In the revised manuscript, we expanded the section describing their origin, selection strategy, and intended use. We also added an updated reference describing the scFv discovery methodology, together with a brief summary of the workflow. In addition, the sequences of all six scFvs (ANXA11, OPTN, MATR3, PFN1, UBQLN2, and VCP) are now provided in Supplementary Table 2. During the revision process, we also identified that two scFv plasmids were missing from Addgene; these have now been deposited, making all six reagents and their sequences publicly available.

      (8) Accessibility of the resource through Zenodo is not straightforward - the reader currently has to navigate to individual antibody characterization reports one by one to extract recommendations for a given target. While the use of an established public repository is important for permanence, a dedicated ALS-RAP website with an interactive, searchable interface - filterable by target, application, host species and clonality - would meaningfully improve uptake. The relationship between such a portal and the existing OGA platform should also be clarified.

      We agree that navigating individual antibody characterization reports on Zenodo can be cumbersome for end users. This challenge was one of the motivations behind the development of the Only Good Antibodies (OGA) platform, which curates YCharOS antibody characterization data, including the ALS-RAP dataset, into an interactive and searchable interface. OGA allows users to identify antibodies by target, application, host species, clonality, and other relevant attributes, while providing direct links to the complete characterization reports archived on Zenodo. In this way, Zenodo serves as the permanent open repository for the complete datasets, whereas OGA provides a user-friendly interface for data exploration and antibody selection.

    1. eLife Assessment

      This study used several approaches (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial recordings) to address a significant question in epilepsy research, with additional relevance to EEG studies more broadly: Are high frequency oscillations (or "fast ripples", defined as >200 Hz by the authors) distinct from randomly occurring clustering of spikes? The results suggest fast ripples can occur by chance and how this may occur. The significance was considered important and the strength of evidence convincing, with minor revisions suggested to temper conclusions related to stochastic vs. oscillatory fast ripples.

    2. Reviewer #1 (Public review):

      Summary:

      This is a study utilizing several types of analyses (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial multi-scale recordings) to address a highly relevant conceptual question: Are fast ripples (FRs) distinct pathological entities or largely emergent products of stochastic spike clustering? The results can potentially reshape current approaches to incorporating fast ripples into the epilepsy surgery evaluation.

      Strengths:

      The conceptualization of fast ripples as potentially arising by chance is highly novel and builds effectively on questions raised in prior studies that have never been satisfactorily resolved. Integration across biological scales and models provides a rigorous approach, now improved by addressing theoretical concerns regarding validity of the shuffling approach and state dependence. The discussion has been updated to provide a more nuanced interpretation of the study's findings.

      Weaknesses:

      The authors have satisfactorily and thoughtfully addressed the critiques provided in the first review. However, there remain two points that I would like authors to address prior to publication:

      (1) Synchronized burst firing is a key feature of an epileptic site generating interictal discharges, and one that could generate either oscillatory or stochastic FRs as documented in multiple prior publications cited in the manuscript and/or in the prior review. Paroxysmal depolarization, for example, has been very well described, and consists of strong, disorganized burst firing (resulting in summated postsynaptic potentials strong enough to generate high gamma signal) in a neuronal population coinciding with a large low-frequency deflection. I would like to see the results described in this context, and to avoid blanket dismissal of stochastic FRs without a clear oscillatory component.

      (2) It would be highly useful to add a conclusion paragraph that spells out implications of the study for use of FRs as epileptic biomarkers in clinical invasive EEG recordings.

      Comments on revised version.

      The author's additions to the discussion are appreciated.

    3. Reviewer #2 (Public review):

      Summary:

      This paper asks an important question that has not been discussed much in the extensive literature on the High Frequency Oscillations (HFOs) that have been extensively studied in patients with epilepsy and experimental models of epilepsy. The question is whether the Fast Ripples (FRs), the HFOs in the 250-500 Hz frequency band, represent a pathological phenomenon or represent a physiological phenomenon that occurs in the healthy brain but happens to be more frequent in epileptic tissue. It is an important question that has not been systematically addressed until now. The authors conclude, from very extensive simulations, from extensive experimental animal studies (the systemic kianate model of epilepsy in rats), and from a modest amount of human data, that FRs occur in healthy brains as a result of the chance occurrence of bursts of action potentials, and that in epileptic tissue, their frequency of occurrence is approximately 30% higher than what is expected by chance. They conclude that FRs are not a separate phenomenon of epileptic tissue. This finding is reinforced by the recent findings of FRs in experimental models of Alzheimer's disease.

      Strengths:

      This is a valuable study because it asks an important and original question and because it evaluates it from several angles (simulation, tissue culture, experimental animals, and human patients). The simulations and the analyses of real data are performed very carefully and with original and solidly documented approaches, using extensive simulations and extensive data sets in the cultured cell data and in the in vivo experiments. The paper is clearly written and well-illustrated.

      Comments on revised version.

      The authors have appropriately addressed the questions I raised in the first review

    4. Reviewer #3 (Public review):

      Summary:

      An outstanding question in the field of high frequency oscillations (HFOs) in the context of epilepsy is how these oscillations emerge, considering that they occur at such high frequencies i.e., 250Hz well above the firing ability of single neurons. One hypothesis that has been suggested in the past is that neurons that fire in an out of phase fashion or rather at random intervals may contribute to a spectrum of HFOs ranging from 250-500Hz that observed in epilepsy. However, how possible it is that random action potentials could aggregate to the extent that they could give rise to HFOs in the so-called fast ripple (FRs) frequency range (>200 according to the authors) remains unclear. To test this hypothesis, they used computational modeling to randomly insert action potentials in a signal, and they found that this approach is sufficient to generate FRs. Some of the predictors of whether FRs could occur were neuronal count, firing rate and synchronization. Besides computational modeling, they used different model systems to test whether that would be possible to be observed in neuronal cultures, in epileptic rats (intrahippocampal kainic acid model), and human data. Neuronal cultures treated with picrotoxin did not show evidence that FRs could be generated more than chance aggregation of action potentials. They then asked whether synchronization and firing rate could play a role in the emergence of FRs. They found that changes in neural firing and synchronization, such as those occurring during differences phase of the sleep-wake cycle could affect the number of FRs occurring by chance aggregation, with more FRs seen during periods of wakefulness, a result that they replicated in human data.

      The authors largely achieve their proposed aims of demonstrating that random neuronal firing can, in principle, generate FRs. Results from this study could influence current thinking around mechanisms generating FRs in epilepsy. The use of different computational approaches and model systems could offer new analytical methodologies for the study of FRs in the context of brain disease.

      Strengths:

      (1) The authors used a multi-level approach combining computational modeling with experimental datasets, including neuronal cultures, a rat model of temporal lobe epilepsy and human data.

      (2) Identification of key parameters such as neuronal count, firing rate, synchronization and brain state in observed incidence of FRs generated through random aggregation of neural firing.

      (3) Cross-species validation increases the likelihood of generalizability of the findings.

      Comments on revised version.

      The authors have addressed my comments and have no further suggestions or any changes in my assessment.

    5. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a study utilizing several types of analyses (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial multi-scale recordings) to address a highly relevant conceptual question: Are fast ripples (FRs) distinct pathological entities or largely emergent products of stochastic spike clustering? The results can potentially reshape current approaches to incorporating fast ripples into the epilepsy surgery evaluation.

      Strengths:

      The conceptualization of fast ripples as potentially arising by chance is highly novel and builds effectively on questions raised in prior studies that have never been satisfactorily resolved. Integration across biological scales and models provides a rigorous approach, now improved by addressing theoretical concerns regarding validity of the shuffling approach and state dependence. The discussion has been updated to provide a more nuanced interpretation of the study's findings.

      Weaknesses:

      The authors have satisfactorily and thoughtfully addressed the critiques provided in the first review. However, there remain two points that I would like authors to address:

      (1) Synchronized burst firing is a key feature of an epileptic site generating interictal discharges, and one that could generate either oscillatory or stochastic FRs as documented in multiple prior publications cited in the manuscript and/or in the prior review. Paroxysmal depolarization, for example, has been very well described, and consists of strong, disorganized burst firing (resulting in summated postsynaptic potentials strong enough to generate high gamma signal) in a neuronal population coinciding with a large low-frequency deflection. I would like to see the results described in this context, and to avoid blanket dismissal of stochastic FRs without a clear oscillatory component.

      To describe the results in this context, we have updated the Discussion. It reads:

      “Depolarizations that promote burst firing should also increase the likelihood of FRs. Paroxysmal depolarization shifts, which are the cellular events underlying epileptiform abnormalities, would therefore be expected to co-occur with FRs. Accordingly, epileptiform abnormalities should be associated with increased gamma/high-gamma (> 80 Hz) activity at their source as has been observed in epileptogenic parenchyma (Ren et al., 2015; Sheybani et al., 2021; Weiss et al., 2015).”

      (2) It would be highly useful to add a conclusion paragraph that spells out implications of the study for use of FRs as epileptic biomarkers in clinical invasive EEG recordings.

      We fully agree that the main message should be further clarified. We have added this paragraph to this end:

      “Clinically, our findings do not argue for or against the use of FRs as epilepsy biomarkers, as we did not assess their relationship with the epileptogenic zone. However, they do indicate that FRs are not highly specific to epileptic tissue. We would, however, recommend selecting FRs with particularly long duration, which may offer greater specificity and potentially better reliability as a biomarker. The clinical value of FRs as distinct entities or as emergent oscillations remains to be clarified.”

      Please address the above critiques in Discussion, or elsewhere as deemed necessary by the authors.

      Reviewer #2 (Public review):

      Summary:

      This paper asks an important question that has not been discussed much in the extensive literature on the High Frequency Oscillations (HFOs) that have been extensively studied in patients with epilepsy and experimental models of epilepsy. The question is whether the Fast Ripples (FRs), the HFOs in the 250-500 Hz frequency band, represent a pathological phenomenon or represent a physiological phenomenon that occurs in the healthy brain but happens to be more frequent in epileptic tissue. It is an important question that has not been systematically addressed until now. The authors conclude, from very extensive simulations, from extensive experimental animal studies (the systemic kianate model of epilepsy in rats), and from a modest amount of human data, that FRs occur in healthy brains as a result of the chance occurrence of bursts of action potentials, and that in epileptic tissue, their frequency of occurrence is approximately 30% higher than what is expected by chance. They conclude that FRs are not a separate phenomenon of epileptic tissue. This finding is reinforced by the recent findings of FRs in experimental models of Alzheimer's disease.

      Strengths:

      This is a valuable study because it asks an important and original question and because it evaluates it from several angles (simulation, tissue culture, experimental animals, and human patients). The simulations and the analyses of real data are performed very carefully and with original and solidly documented approaches, using extensive simulations and extensive data sets in the cultured cell data and in the in vivo experiments. The paper is clearly written and well-illustrated.

      Comments on revised version.

      The authors have appropriately addressed the questions I raised in the first review.

      Reviewer #3 (Public review):

      Summary:

      An outstanding question in the field of high-frequency oscillations (HFOs) in the context of epilepsy is how these oscillations emerge, considering that they occur at such high frequencies i.e., 250Hz well above the firing ability of single neurons. One hypothesis that has been suggested in the past is that neurons that fire in an out-of-phase fashion or rather at random intervals may contribute to a spectrum of HFOs ranging from 250-500Hz that observed in epilepsy. However, how possible it is that random action potentials could aggregate to the extent that they could give rise to HFOs in the so-called fast ripple (FRs) frequency range (>200 according to the authors) remains unclear. To test this hypothesis, they used computational modeling to randomly insert action potentials in a signal, and they found that this approach is sufficient to generate FRs. Some of the predictors of whether FRs could occur were neuronal count, firing rate and synchronization. Besides computational modeling, they used different model systems to test whether that would be possible to be observed in neuronal cultures, in epileptic rats (intrahippocampal kainic acid model), and human data. Neuronal cultures treated with picrotoxin did not show evidence that FRs could be generated more than chance aggregation of action potentials. They then asked whether synchronization and firing rate could play a role in the emergence of FRs. They found that changes in neural firing and synchronization, such as those occurring during different phases of the sleep-wake cycle could affect the number of FRs occurring by chance aggregation, with more FRs seen during periods of wakefulness, a result that they replicated in human data.

      The authors largely achieve their proposed aims of demonstrating that random neuronal firing can, in principle, generate FRs. Results from this study could influence current thinking around mechanisms generating FRs in epilepsy. The use of different computational approaches and model systems could offer new analytical methodologies for the study of FRs in the context of brain disease.

      Strengths:

      (1) The authors used a multi-level approach combining computational modeling with experimental datasets, including neuronal cultures, a rat model of temporal lobe epilepsy and human data.

      (2) Identification of key parameters such as neuronal count, firing rate, synchronization and brain state in observed incidence of FRs generated through random aggregation of neural firing.

      (3) Cross-species validation increases the likelihood of generalizability of the findings.

      Minor weakness:

      (1) The analyses conducted in human data lack direct comparison with sleep data due to no available data, but would encourage future investigations directly comparing HFOs during wakefulness and nocturnal sleep.

      Comments on revised version.

      The authors have addressed my comments and I have no further suggestions.

    1. On the other hand, some bots are made with the intention of harming, countering, or deceiving others. For example, people use bots to spam advertisements at people. You can use bots as a way of buying fake followers [c9], or making fake crowds that appear to support a cause

      I wonder how much verification processes actually help towards identifying fake accounts? I find it disconcerting how a lot of peoples opinions on the internet could be skewed via algorithms that could show mass produced bot content.

    1. Flash permettait une entrée en matière de la grammaire interactive par la métaphore de la page blanche, à l’instar de tout logiciel d’édition WYSIWYG. Ce modèle, né de l’Apple HyperCard (1987) et de Macromedia Director, influence d’ailleurs toujours les logiciels de visual coding contemporains, tel que Scratch (2006) , destiné à l’appréhension de l’interactivité par les enfants. En combinant une interface de création visuelle et un environnement de programmation dans un ensemble cohérent, Flash permit à toute une génération de designers d’aborder la création numérique de façon unifiée. Beaucoup d’écoles qui se consacraient au design numérique et interactif l’ont intégré dans leurs maquettes pédagogiques dédiées à l’interactivité du Web, car il permettait d’approcher différents types d’objets éditoriaux, jusqu’au jeu vidéo (notamment pris en charge depuis 2005 par Unity). Son abandon progressif a plongé designers et enseignants dans un certain désarroi et a forcé l’adoption de méthodologies moins visuelles et donc plus difficiles d’accès pour ce public qui manipule finalement un matériau qui n’est graphique que par destination.

      coeur de la problématique de Christian Porri dans le cadre de l'enseignement et l'apprentisage, une REGRESSION

    2. Dans le cas du jeu vidéo en ligne, Flash a conféré aux navigateurs Web le statut de plateformes de diffusion tout en bouleversant certaines pratiques, en contribuant notamment à l’essor de ce que l’on appelait le casual gaming. Beaucoup se réappropriaient des gameplays basiques hérités du point and clic et en raffinaient l’interactivité. Sans bousculer fondamentalement les genres, cette nouvelle plateforme a donné au jeu vidéo des usages propres au Web en tant que réseau : interconnexions avec des contenus externes ou partage de données vers d’autres plateformes, notamment sociales. Ces jeux ont également pu être les relais d’objets informationnels identifiés comme des sites promotionnels d’agences , des pages dédiées au lancement de produits ou des relais d’événements culturels. Certains jeux sont aussi des magazines, transposés depuis un support papier préexistant ou conçus pour l’écran, trouvant alors de nouveaux territoires de diffusion et d’expression par l’interactivité et le ludique.

      Apparition et renforcemetn de nouveau "objets" du web, et une vraie expension pas les atout de diffusion.

    3. Dès l’an 2000, les designers qui produisaient des animations Flash se sont emparés de ses fonctions de programmation. La structure de son langage ActionScript, héritée des langages de programmation orientée objet, se formalise parfaitement dans une interface graphique utilisateur en menus et palettes, facilitant l’accès à des fonctions d’interaction puissantes, à la gestion d’événements utilisateur, sans que le code ne soit une connaissance préalable indispensable

      ruée sur l'experimentation en tout genre grace à l'outil

    4. Le format-média SWF s’intègre au code HTML d’une page et la transforme en une zone ouverte à d’autres manipulations que les seuls liens hypertextes.

      éléments de definition de ce qu'est "Flash"

    5. Flash trouve son origine dans le logiciel d’animation SmartSketch, édité par FutureWave en 19931414 Voir : Richard C. Moss, « The Rise and Fall of Adobe Flash », ArsTechnica, 2020, http://b-o.fr/moss. Macromedia le racheta en 1996 et eut l’intelligence commerciale de négocier son installation par défaut sur tout système d’exploitation via le navigateur Microsoft Internet Explorer. Dès 1997,

      methode de présentation de l'histoire et critique de ces éléments

    6. Déjà en octobre 2000, les gourous de l’UX design Jakob Nielsen et Don Norman estimaient qu’Adobe Flash (1996) provoquait une sorte de « maladie de l’utilisabilité11 Jakob Nielsen, « Flash: 99 % Bad », Nielsen/Norman Group, 2000, http://b-o.fr/nielsen », un cancer de l’usage. En bons défenseurs d’une forme de fonctionnalisme, ils prônaient l’abandon des médias « enrichis » pour recentrer le Web sur le contenu « seul », sans autre forme d’invention. En 2010, on entendait des affirmations au ton tout aussi péremptoire lorsque Steve Jobs appellait à bannir Flash22 Steve Jobs, « Thoughts on Flash », avril 2010, http://b-o.fr/jobs. Plus personne ne tolérait cette technologie ni tout ce qu’elle avait contribué à créer pendant presque quinze ans : design Web immersif, animations sophistiquées, contenus multimédias, introductions, interactions non standard, etc.

      Contexte de l'article, fin des année 2000.

    1. Checkout / web shop 1 hour 5 min Lost sales mount fast (Ch. 2) Customer database 4 hours 1 hour Serious, but less instant

      The checkout in TB1 depends on the database in TB2, so if TB2 has an RTO of 4 hours, the 1-hour RTO requirement for TB1 cannot be met.

    1. dobrovolnou možnost přestěhování do vhodnějšího typu bydlení.

      nebo spíše - umožnit více možností pro případné řešení bytové situace ve formě komunitního či asistovaného bydlení?

    2. silné sociální a emocionální vazby

      Tohle se v celém textu opakuje, je to asi logické a předpokládáme to, ale bylo by dobré to podpořit nějakou studií či odbornou literaturou.

    3. jednočlenným uspořádáním domácnosti a ženským pohlavím

      možná spíše napsat, že významně více ohrožuje samostatně bydlící seniorky a spojit s následující větou. Dodala bych také, že se jich situace častěji týká z důvodu vyššího věku dožití.

    4. Nejvíce jsou jím ohroženy

      to nadměrné zatížení je poměr domácností, ale zatížení výdaji na bydlení ukazuje poměr příjmů k výdajům. Přijde mi trochu, že to zde v těchto dvou odstavcích není zcela jasné. Zároveň mám trochu pocit, že už jsem něco takového četla výše...

    5. rozdělit výdaje na vytápění, energie, služby, údržbu nebo nájem mezi více členů domácnosti

      to mi také přijde, že už bylo několikrát výše napsáno.

    6. Vyšší

      K textu - opakuje se to, co bylo výše - asi se odkázat a už neopakovat. Ke grafu - Chybí nadpis grafu - to platí i výše. A barvy mi přijdou, že jsou jiné než máme ve zbytku. A chybí legenda

    7. Šetření SILC

      velké Š - napsala bych jen ...Z dostupných dat vyplývá a do závorky (asi na konec věty) dala citaci EU-SILC těch individuálních dat

    8. priority řadí dostupnější a bezbariérové bydlení pro seniory

      Tohle je podle mě důležité pro koncepci bydlení -- je potřeba to tedy se strategií mpsv provázat

    9. Je spojeno s dlouhodobými sousedskými vztahy, rodinnou historií, vzpomínkami a pocitem bezpečí. Přestěhování proto může znamenat nejen změnu bytových podmínek, ale také oslabení sociálních vazeb a ztrátu známého prostřed

      Trochu mi přijde, že se toto stále opakuje

    10. K relativně vyššímu prostorovému standardu seniorských domácností přispívá také jejich častější bydlení v rodinných domech, které zpravidla disponují větší podlahovou plochou a vyšším počtem obytných místností.

      To myslím, že už bylo vícekrát řečeno výše, Zde bych už větu škrtla

    11. Pro seniory je typické častější

      máme nějaký zdroj nebo potvrzení, že je to pro ně typické? Z dat to nevidíme, tak třeba z nějaké odborné literatury?

    12. Ještě výraznější příjmové rozdíly se projevují u nadměrného

      Možná bych více zmínil ten velký výkyv v zatížení i nadměrném zatížení u nejchudších domácností, který ukazuje na náchylnost na ceny energií u této skupiny. důchody se nominálně nezmenšovaly a placené nájmy takhle nahoru a dolu taky nekolísají

    13. ně 30 % a více

      Upravil bych formulaci, ať je jasné, že se jedná o podíl nadměrně zatížených domácností. Vím, že to je definované nahoře, ale stejně.

    14. ženám

      Možná bychom mohli ještě dodat, že v důsledku delší doby dožití žen, jsou hlavně ohroženy chudobou samostatně žijící seniorky.

    15. Nadměrné zatížení výdaji

      Já to chápu, ale tady bych to asi rozepsal ... buď míra nadměrného zatížení nebo klidně i Podíl domácností, které jsou nadměrně zatížené výdaji na bydlení nebo tak

    16. Právní důvod

      Možná bych zmínil, že také více bydlí ve vlastním oproti neseniorským stejně velkým domácnostem a možná je zajímavý i rozdíl dvou vs jednočlenné seniorské, kde ty jednočlenné se více přesouvají do nájmu nebo k rodině, i když stále z většiny zůstávají ve vlastním bydlení. Taky to teda může být tím, že mladší domácnosti se nastěhujou do bytu původně rodičů, kteří tam stále bydlí, zejména pokud už nemají partnera, což je obecně asi projev nedostupnosti pro mladší domácnosti.

    17. znam, který přesahuje jeho ekonomickou hodnotu

      ekonomie obecně nehodnotí, čemu jednotliví lidé přisuzují hodnotu a jak velkou, asi myslíš spíše monetární nebo tržní hodnotu

    1. eLife Assessment

      This landmark study is part of an impressive large-scale effort to assess the consistency of published findings in the field of Drosophila immunity. In this article, the authors report an analysis of claims made in approximately 400 research papers from the Drosophila immunity literature. The authors perform validation experiments to test a subset of the claims that are most inconsistent with current understanding in the field. These experiments do not necessarily aim to reproduce the original experiments, but instead to test the repeatability of the result using updated tools and methods. In some cases, the original claim is validated and in others it is not. The conclusions on the consistency of the claims, as presented in this paper, are based on mostly convincing evidence.

    2. Reviewer #1 (Public Review):

      This work revisits a substantial part of the published literature in the field of Drosophila innate immunity from 1959 to 2011. The strategy has been to restrain the analysis to some 400 articles and then to extract a main claim, two to four major claims and up to four minor claims totaling some 2000 claims overall. The consistency of these claims with the current state-of-the-art has been evaluated and reported on a dedicated Web site known as Reprosci and also in the text as well as in the 28 Supplements that report experimental verification, direct or indirect, e.g., using novel null mutants unavailable at the time, of a selected set of claims made in several articles. Of note, this review is mostly limited to the manuscript and its associated supplements and does not integrally cover the Reprosci website.

      Strengths:<br /> One major strength of this article is that it tackles the issue of reproducibility/consistency on a large-scale. Indeed, while many investigators have some serious doubts on some results found in the literature, few have the courage, or the means and time, to seriously challenge studies, especially if published by leaders in the field. The Discussion adequately states the major limitations of the Reprosci approach, which should be kept in mind by the reader to form its own opinion.

      This study also allows investigators not familiar with the field to have a clearer understanding of the questions at stake and to derive a more coherent global picture that allows them to better frame their own scientific questions. Besides a thorough and up-to-date knowledge of the literature used to assess the consistency of the claims with our current knowledge, a merit of this study is the undertaking of independent experiments to address some puzzling findings and the evidence presented is often convincing, albeit one should keep in mind the inherent limitations as several parameters are difficult to control, especially in the field of infections, as underlined by the authors themselves. Importantly, some work of the lead author has also been re-evaluated (e.g., Supplements S2-S4). Thus, while utmost caution should be exerted, and often is, in challenging claims, even if the challenge eventually proves to be not grounded, it is valuable to point out potential controversial issues to the scientific community.

      While this is not a point of this review, it should be acknowledged that the possibility to post comments on the ReproSci website will allow further readjustments by the community in the appreciation of the literature and also of the Reprosci assessments themselves and of its complementary additional experiments. As science is evolving continuously, this site allows the current state of knowledge to be updated as needed, not necessarily by the authors themselves.

      Comments on revised version:

      The major weaknesses identified in the original version of the manuscript have been adequately addressed by the authors. In a few cases, additional KO mutants have been generated and analyzed thus further strengthening the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      The authors present an ambitious and large-scale reproducibility analysis of 400 articles on Drosophila immunity published before 2011. They extract major and minor claims from each article, assess their verifiability through literature comparison and, when possible, through targeted experimental re-testing, and synthesize their findings in an openly accessible online database. The goal is to provide clarity to the community regarding claims that have been contradicted, incompletely supported, or insufficiently followed up in the literature, and to foster broader community participation in evaluating historical findings. The manuscript summarizes the major insights emerging from this systematic effort.

      Strengths:

      (1) Novelty and community value: This work represents a rare example of a systematic, transparent, and community-facing reproducibility project in a specific research domain. The creation of a dedicated public platform for disseminating and discussing these assessments is particularly innovative.

      (2) Breadth and depth: The authors analyze an impressive number of publications spanning multiple decades, and they couple literature-based assessments with new experimental data where follow-up is missing.

      (3) Clarity of purpose: The manuscript carefully distinguishes between assessing evidential support for claims and judging the scientific merit of historical work. This helps frame the project as constructive rather than punitive.

      (4) Metascientific relevance: The analysis identifies methodological and contextual factors that commonly underlie irreproducible claims, providing a useful guide for future study design and interpretation.

      (5) Transparency: Supplementary datasets and the public website provide an exceptional degree of openness, which should facilitate community engagement and further refinement.

      Comments on revised version:

      In the revised manuscript the authors addressed in a satisfactory manner the weaknesses I previously highlighted in my initial review.

    4. Reviewer #3 (Public review):

      Summary:

      In this ambitious study the authors set out to analyse the validity of a number of claims, both minor and major from 400 published articles within the field of Drosophila immunity that were published before 2011. The authors were able to determine initially if claims were supported by comparing to other published literature in the field and if required by experimentally testing 'unchallenged' claims that had not been followed up in subsequent published literature. Using this approach the authors identified a number of claims that had contradictory evidence using new methods or taking into account developments within the field post initial publication. They publish their findings on a publicly available website designed to enable the research community to assess published work within the field with greater clarity.

      Strengths:

      The work presented is rigorous and methodical, the data presentation is high quality and importantly the data presented support the conclusions. The discussion is balanced and the study is written considerately and respectfully highlighting that the aim of the study is not to assign merit to individual scientists or publications but rather to improve clarity for scientists across the field. The approach carried out by the researchers focuses on testing the validity of the claims made in the original papers rather than testing whether the original experimental methods produced reproducible results. This is an important point since there are many reasons why the original interpretation of data may have understandably led to the claims made. These potential explanations for irreproducible data or conclusions are discussed in detail by the authors for each claim investigated.

      The authors have generated an accompanying website which provides a valuable tool for the Drosophila Immunity research community that can be used to fact check key claims and encourages community engagement. This will achieve one important goal of this study - to prevent time loss for scientists who base their research on claims that are irreproducible. The authors rightly point out that it is impossible (and indeed undesirable) to avoid publication of irreproducible results within a field since science is 'an exploratory process where progress is made by constant course correction'. This study is however an important piece of work that will make that course correction more efficient.

      Comments on revised version:

      I'm impressed with the revisions carried out by the authors. No further comments.

    5. Reviewer #4 (Public review):

      Westlake, Lemaitre and colleagues have revised their manuscript regarding the reproducibility of studies on the Drosophila immune system. As was emphasized by all reviewers at the original submission of this manuscript, this paper captures a tremendous effort and makes a valuable contribution to the community.

      One strong conclusion to be drawn from the overall effort is that the scientific literature is, for the most part, very good. The vast majority of claims made in the 400 papers assembled for this project were substantiated, either with previously published follow-up work or with new experimentation presented here. Although the authors have chosen not to quantify the overall validation rate of scientific claims in the present study, that information is available on the ReproSci website and in the companion paper.

      In the revised manuscript, the authors make the distinction between repeatability/direct reproducibility and what they term "conceptual replication" and "indirect reproducibility", emphasizing that this study falls in the latter category because in many cases the present authors are not attempting to reproduce or replicate the original study. Instead, they are often using new tools to reinvestigate original claims and determine the accuracy of those claims, and whether they stand up to additional data. The authors have replaced the word "irreproducible" with "irreplicable" in the revised manuscript. I continue to believe the present work can be considered "validation" of the original studies and that the present work is accurately considered a "validation project". The clarification of terminology has strengthened the revised manuscript.

      As with any study, the strength and breadth of conclusions that can be drawn from the present work are a function of the strength and resolution of the experiments performed. This may vary across the supplements so readers should interpret the presented results accordingly. In my opinion, all of the supplements in the revised article are sufficiently convincing to warrant inclusion in the final manuscript.

      There are several instances where the authors have declined to do additional experiments in response to review. I am comfortable with these decisions, recognizing that the extra work required would be prohibitively much. Instead, the authors have adequately revised their presentation of findings. In some instances where the authors have declined to perform additional experiments, they have provided a compelling logic for their decision in the response to reviews, which unfortunately does not appear in the main manuscript. For example, in response to a comment by Reviewer 1, the authors wrote "In response to the suggestion of scanning multiple blots for quantification, we note that no clear enhancement of ModSP autoactivation was observed across different incubation times. We also explored alternative incubation conditions, but none resulted in a detectable increase in activation. Thus, additional quantification would not change the overall conclusion." This is convincing, but no readers will see it unless it appears in the manuscript or the public response to review. At the same time, I recognize the practical limitations in available space in the main article and supplement. There are a handful of instances where reviewers have made suggestions of additional/alternative experiments that would be more rigorous, accurate, or convincing than what is currently presented but the authors have declined to take up those recommendations. In some instances, I strongly agree with the authors' rationale and in other cases it is less so. However, I did not find any instance where the absence of the suggested additional experiment substantially undermined the conclusions being put forward in this manuscript. Therefore, I believe these decisions can be left to the discretion of the authors in accordance with the eLife publication model and interpretations can be left to the discretion of the readers.

      Reviewer 1 and Reviewer 4 both commented on a lack of demonstration of RNAi efficacy in the knockdown experiments and suggested that negative conclusions drawn from RNAi experiments should include controls showing that the RNAi was effective. The authors have declined to include such controls, and I recognize they may be impossible to add now without redoing the entire associated experiment. Instead, the authors point to prior use of the RNA constructs by other authors in the literature. That should be recognized as less than ideal, although it is probably substantially adequate. The authors have added statements to the Discussion that RNAi knockdowns are not equivalent to null mutations and negative results may appear because of hypomorphic effects. This is a positive addition.

      The replacement of the sphinx RNAi with a deletion of Sphinx 1 and Sphinx 2 as part of a general strengthening of the section on sphinx and spheroid is an extremely strong addition to the paper. I could not expect that level of effort for each of the validations and I acknowledge and appreciate that they have made that effort here.

      Reviewer 1 and Reviewer 4 both commented that some figures in the original submission were ambiguous with respect to sample size and distribution of data points, including undefined error bars and inappropriate graphical representation of experiments with low sample size. These concerns have been largely addressed in the revised manuscript, including with addition of statistical analyses in several cases. The authors note that the figures were generated by multiple individuals across several research teams so it would be a large effort unify their style. I agree that this would be an unreasonable effort at this stage, for minimal return that is primarily aesthetic. The authors have now added illustration of all data points in most figures and have defined the error bars. One response to Reviewer 1 "Data points should be shown in the figure" with respect to a figure in S16 was "Unfortunately, we could not address this point to the difficulty to reach the collaborator who provided the data." That response certainly would not be acceptable in the publication of a typical scientific paper. However, I sense a different standard is being applied to the present work and the lack of these data points in this particular figure does not decrease my confidence in the authors' overall conclusion.

      Overall, this is a unique piece of research that is impressive in its scope. Scientific finding is never complete; the process is eternally iterative. Inclusion of the present study in the scientific literature is a testament to collective willingness to engage in that iteration. The authors of this paper are making an exemplar contribution to accuracy in scientific research and should be commended for their effort, dedication, and fairness in what can be a delicate topic.

    6. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This fundamental study is part of an impressive, large-scale effort to assess the reproducibility of published findings in the field of Drosophila immunity. In a companion article, the authors analyze 400 papers published between 1959 and 2011, and assess how many of the claims in these papers have been tested in subsequent publications. In this article, the authors report the results of validation experiments to assess a subset of the claims that, according to the literature, have not been corroborated. While the evidence reported for some of these validation studies is convincing, it remains incomplete for others.

      We thank the reviewers for their careful assessment of our manuscript. We are fully aware that this review process required a substantial investment of time and effort, and we are deeply grateful for their engagement. We also thank the eLife editors for enabling an open and constructive evaluation process of the highest standard.

      We fully agree with the eLife assessment, including the concluding statement highlighting the limitations of our analysis. We also concur with most of the reviewers’ comments, which we have addressed in detail in the revised manuscript. Addressing these comments required extensive additional work, which explains the duration of the revision. This period was also valuable in allowing colleagues from the Drosophila community to provide feedback on our study, including identifying potential issues—for instance regarding SR-C1—that we have carefully re-examined.

      Major revisions include the addition of new experimental data in the sections on Hemese, SR-C1, Sphinx1/2, and Caspar. These results largely confirm our initial conclusions, with the exception of SR-C1. We have therefore revised our interpretation of SR-C1, confirming earlier findings that it promotes bacterial binding, while maintaining that it does not function as a major phagocytic receptor. A table with new list of new reagents has been added at the end of the supplement (Table Supplement S4). We also add a comment on the fact that mutation in the serine protease MP1 (work from B.L.) does not affect wound melanization as initially proposed (Tang et al., 2006).

      One final point: our study assesses replicability using alternative experimental approaches, and we cannot exclude the possibility that our conclusions may be inaccurate. This limitation is inherent to all reproducibility studies. Nonetheless, our work should encourage the community to exercise caution when interpreting certain previous claims, some of which have become widely accepted.

      We thank the reviewers and editors for their constructive input, which has significantly improved the manuscript. At this stage, we are ready to proceed toward a version of record together with the eLife assessment (which may still evolve). We continue to believe that our reproducibility project is unique in its scope and depth. It not only provides quantitative insight into replicability (See companion article)—suggesting that replication rates may be higher in certain fields, consistent with the maturation of knowledge over recent decades—but also contributes to clarifying and refining a number of statements within the field. Although aware of the cost of assessing reproducibility (hard work for articles that are not quoted and statements that are not taking in account), I consider it as one of my most important and original works.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work revisits a substantial part of the published literature in the field of Drosophila innate immunity from 1959 to 2011. The strategy has been to restrain the analysis to some 400 articles and then to extract a main claim, two to four major claims and up to four minor claims totaling some 2000 claims overall. The consistency of these claims with the current state-of the-art has been evaluated and reported on a dedicated Web site known as ReproSci and also in the text as well as in the 28 Supplements that report experimental verification, direct or indirect, e.g., using novel null mutants unavailable at the time, of a selected set of claims made in several articles. Of note, this review is mostly limited to the manuscript and its associated supplements and does not integrally cover the ReproSci website.

      Strengths:

      One major strength of this article is that it tackles the issue of reproducibility/consistency on a large scale. Indeed, while many investigators have some serious doubts about some results found in the literature, few have the courage, or the means and time, to seriously challenge studies, especially if published by leaders in the field. The Discussion adequately states the major limitations of the ReproSci approach, which should be kept in mind by the reader to form their own opinion.

      This study also allows investigators not familiar with the field to have a clearer understanding of the questions at stake and to derive a more coherent global picture that allows them to better frame their own scientific questions. Besides a thorough and up-to-date knowledge of the literature used to assess the consistency of the claims with our current knowledge, a merit of this study is the undertaking of independent experiments to address some puzzling findings and the evidence presented is often convincing, albeit one should keep in mind the inherent limitations as several parameters are difficult to control, especially in the field of infections, as underlined by the authors themselves. Importantly, some work of the lead author has also been re-evaluated (Supplements S2-S4). Thus, while utmost caution should be exerted, and often is, in challenging claims, even if the challenge eventually proves to be not grounded, it is valuable to point out potential controversial issues to the scientific community.

      While this is not a point of this review, it should be acknowledged that the possibility to post comments on the ReproSci website will allow further readjustments by the community in the appreciation of the literature and also of the ReproSci assessments themselves and of its complementary additional experiments.

      We praise the reviewer for this assessment of our project

      Weaknesses:

      Challenging the results from articles is, by its very nature, a highly sensitive issue, and utmost care should be taken when challenging claims. While the authors generally acknowledge the limitations of their approach in the main text and Supplements, there are a few instances where their challenges remain questionable and should be reassessed. This is certainly the case for Supplement S18, for which the ReproSci authors make a claim for a point that was not made in the publication under scrutiny. The authors of that study (Ramet et al., Immunity, 2001) never claimed that scavenger receptor SR-CI is a phagocytosis receptor, but that it is required for optimal binding of S2 cells to bacteria. Westlake et al. here have tested for a role of this scavenger receptor in phagocytosis, which had not been tested by Ramet et al. Thus, even though the ReproSci study brings additional knowledge to our understanding of the function of SR-CI by directly testing its involvement in phagocytosis by larval hemocytes, it did not address the major point of the Ramet et al. study, SR-CI binding to bacteria, and thus inappropriately concludes in Supplement S18 that "Contrary to (Ramet et al., 2001, Saleh et al., 2006), we find that SR-CI is unlikely to be a major Drosophila phagocytic receptor for bacteria in vivo." It follows that the results of Ramet et al. cannot be challenged by ReproSci as it did not address this program. Of note, Saleh et al. (2006) also mistakenly stated that SR-CI impaired phagocytosis in S2 cells and could be used as a positive control to monitor phagocytosis in S2 cells. Their assay appears to have actually not monitored phagocytosis but the association of FITC-labeled bacteria to S2 cells by FACS, as they did not mention quenching the fluorescence of bacteria associated with the surface with Trypan blue.

      We have revised our evaluation of Sr-C1, as we acknowledge that we did not initially read Rämet et al. carefully. We incorrectly assumed that this study demonstrated a role for SR-C1 in phagocytosis, whereas it actually reports a defect in binding only. We have updated our assessment of SR-C1, validating the results from Rämet et al., 2001. We still mention our results on phagocytosis (See below for more information).

      The inference method to assess the consistency of results with current knowledge also has limitations that should be better acknowledged. At times, the argument is made that the gene under scrutiny may not be expressed at the right time according to large-scale data or that the gene product was not detected in the hemolymph by a mass-spectrometry approach. While being in theory strong arguments, some genes, for instance, those encoding proteases at the apex of proteolytic activation cascades, need not necessarily be strongly expressed and might be released by a few cells. In addition, we are often lacking relevant information on the expression of genes of interest upon specific immune challenges such as infections with such and such pathogens.

      The pattern of expression is always used as an element supporting our assessment but never as the central point of decision (See below for more information).

      As regards mass spectrometry, there is always the issue of sensitivity that limits the force of the argument. Our understanding of melanization remains currently limited, and methods are lacking to accurately measure the killing activity associated with the triggering of the proPO activation cascade. In this study, the authors monitor only the blackening reaction of the wound site based on a semi-quantitative measurement. They are not attempting to use other assays, such as monitoring the cleavage of proPOs into active POs or measuring PO enzymatic activity. These techniques are sometimes difficult to implement, and they suffer at times from variability. Thus, caution should be exerted when drawing conclusions from just monitoring the melanization of wounds.

      We employed approaches similar to those used in the original studies under evaluation. However, we acknowledge that our analysis primarily relied on cuticular melanization, which is more straightforward to monitor. We have revised the manuscript to clarify this limitation, particularly in the case of PGRP-LE.

      Likewise, the study of phagocytosis is limited by several factors. As most studies in the field focus on adults, the potential role of phagocytosis in controlling Gram-negative bacterial infections is often masked by the efficiency of the strong IMD-mediated systemic immune response mediated by AMPs (Hanson et al, eLife, 2019). This problem can be bypassed in rare instances of intestinal infections by Gram-negative bacteria such as Serratia marcescens (Nehme et al., PLoS Pathogens, 2007) or Pseudomonas aeruginosa (Limmer et al. PNAS, 2011), which escape from the digestive tract into the hemocoel without triggering, at least initially, the systemic immune response. It is technically feasible to monitor bacterial uptake in adults by injecting fluorescently labeled bacteria and subsequently quenching the signal from noningested bacteria. Nonetheless, many investigators prefer to resort to ex vivo assays starting from hemocytes collected from third-instar wandering larvae as they are easier to collect and then to analyze, e.g., by FACS. However, it should be pointed out that these hemocytes have been strongly exposed to a peak of ecdysone, which may alter their properties. Like for S2 cells, it is thus not clear whether third-instar larval hemocytes faithfully reproduce the situation in adults. The phagocytic assays are often performed with killed bacteria. Evidence with live microorganisms is better, especially with pathogens. Assays with live bacteria require however, an antibody used in a differential permeabilization protocol. Furthermore, the killing method alters the surface of the microorganisms, a key property for phagocytic uptake. Bacterial surface changes are minimal when microorganisms are killed by X-ray or UV light. These limitations should be kept in mind when proceeding to inference analysis of the consistency of claims. Eater illustrates this point well. Westlake et al. state that:" [...] subsequent studies showed that a null mutation of eater does not impact phagocytosis". The authors refer here to Bretscher et al., Biology Open, 2015, in which binding to heat-killed E. coli was assessed in an ex vivo assay in third instar larvae. In contrast, Chung and Kocks (JBC, 2011) tested whether the recombinant extracellular N-terminal ligand-binding domain was able to bind to bacteria. They found that this domain binds to live Gram-positive bacteria but not to live Gram-negative bacteria. For the latter, killing bacteria with ethanol or heating, but not by formaldehyde treatment, allowed binding. More importantly, Chung and Kocks documented a complex picture in which AMPs may be needed to permeabilize the Gram-negative bacterial cell wall that would then allow access of at least the recombinant secreted Eater extracellular domain to peptidoglycan or peptidoglycan-associated molecules. Thus, the systemic Imd-dependent immune response would be required in vivo to allow Eater-dependent uptake of Gram-negative bacteria by adult hemocytes. In ex vivo assays, any AMPs may be diluted too much to effectively attack the bacterial membrane. A prediction is then that there should be an altered phagocytosis of Gramnegative bacteria in IMD-pathway mutants, e.g., an imd null mutant but not the hypomorphic imd[1] allele. This could easily be tested by ReproSci using the adult phagocytosis assay used by Kocks et al, Cell, 2005. At the very least, the part on the role of Eater in phagocytosis should take the Chung &Kocks study into account, and the conclusions modulated.

      We agree that experimental conditions can always influence the outcome and may partly explain discrepancies in reproducibility. However, our approach is based on conceptual replication, with the aim of testing the generalizability of the original findings rather than reproducing them under identical conditions. We often used approaches quite similar to the initial statements.

      In the case of Eater, our results challenge the conclusion from Kocks et al. that phagocytosis of Gram-negative bacteria is impaired in the absence of Eater, as this was not confirmed when using null eater mutants. Nevertheless, our data support the idea that Eater contributes to the phagocytosis of Gram-negative bacteria, likely in cooperation with NimC1.

      Importantly, our findings are fully consistent with those of Chung and Knocks, who showed that the extracellular domain of Eater binds Gram-positive but not Gram-negative bacteria. Therefore, our work refines a specific aspect of the original conclusion, without questioning the overall significance of the initial Eater studies.

      Another point is that some mutant phenotypes may be highly sensitive to the genetic background, for instance, even after isogenization in two different backgrounds. In the framework of a Reproducibility project, there might be no other option for such cases than direct reproduction of the experiment as relying solely on inference may not be reliable enough.

      With respect to the experimental part, some minor weaknesses have been noted. The authors rely on survival to infection experiments, but often do not show any control experiments with mock-challenged or noninfected mutant fly lines. In some cases, monitoring the microbial burden would have strengthened the evidence. For long survival experiments, a check on the health status of the lines (viral microbiota, Wolbachia) would have been welcome. Also, the experimental validation of reagents, RNAi lines, or KO lines is not documented in all cases.

      We thank the reviewer for their careful and constructive assessment. We agree with most of the points raised and have revised the manuscript accordingly. In particular, we have softened the wording throughout to make our evaluation less abrupt. We have also added a dedicated section on limitations at the end of the manuscript, highlighting potential sources of bias and the inherent constraints of our reproducibility approach.

      In addition, we have expanded the description of the tools used and the validation of the RNAi experiments (see below). We consider it unlikely that factors such as Wolbachia or Nora virus significantly affect our conclusions, as our assays rely on short-term survival experiments, whereas these variables are more likely to influence ageing-related phenotypes. That said, we fully acknowledge that unrecognized parameters—such as environmental conditions (e.g., humidity), which we did not systematically control—could contribute to discrepancies. This limitation, however, applies broadly to most attempts at experimental reproduction.

      More generally, many of the claims we challenge have not been revisited for over a decade. Our inability to validate some of them suggests that the underlying biology may be more complex than initially proposed. We recognize that some of our conclusions may be incorrect, and it would be surprising if all were entirely accurate. At the same time, reproducibility studies often face an asymmetry in standards of evidence, whereby contradicting a published claim requires a higher burden of proof than establishing it. We believe it is important to document and share such discrepancies. We also explicitly acknowledge in the revised manuscript that the ReproSci framework itself may contain errors. Even if a significant proportion of our assessments are not correct, we believe this work contributes meaningfully to clarifying ongoing debates in the field for most of them.

      Finally, since the launch of the website, we have received three comments from the community—one addressing a major claim (Sr-C1) and two addressing more minor points— which have led us to refine our conclusions. The limited feedback we received suggest a broad agreement of the community with our assessment. It illustrates the value of maintaining an open, community-driven resource. The public and evolving nature of the project is, in our view, a key strength.

      Reviewer #2 (Public review):

      Summary:

      The authors present an ambitious and large-scale reproducibility analysis of 400 articles on Drosophila immunity published before 2011. They extract major and minor claims from each article, assess their verifiability through literature comparison and, when possible, through targeted experimental re-testing, and synthesize their findings in an openly accessible online database. The goal is to provide clarity to the community regarding claims that have been contradicted, incompletely supported, or insufficiently followed up in the literature, and to foster broader community participation in evaluating historical findings. The manuscript summarizes the major insights emerging from this systematic effort.

      Strengths:

      (1) Novelty and community value: This work represents a rare example of a systematic, transparent, and community-facing reproducibility project in a specific research domain. The creation of a dedicated public platform for disseminating and discussing these assessments is particularly innovative.

      (2) Breadth and depth: The authors analyze an impressive number of publications spanning multiple decades, and they couple literature-based assessments with new experimental data where follow-up is missing.

      (3) Clarity of purpose: The manuscript carefully distinguishes between assessing evidential support for claims and judging the scientific merit of historical work. This helps frame the project as constructive rather than punitive.

      (4) Metascientific relevance: The analysis identifies methodological and contextual factors that commonly underlie irreproducible claims, providing a useful guide for future study design and interpretation.

      (5) Transparency: Supplementary datasets and the public website provide an exceptional degree of openness, which should facilitate community engagement and further refinement.

      We praise the reviewer for this assessment of our project.

      Weaknesses:

      (1) Subjectivity in selection: Despite the authors' efforts, the choice of which papers and claims to highlight cannot be entirely objective. This is an inherent limitation of any retrospective curation effort, but it remains important to acknowledge explicitly.

      We have added a section at the end of the discussion to discuss the limitation of our study.

      (2) Emphasis on irreproducible claims: The manuscript focuses primarily on claims that are challenged or found to be weakly supported. While understandable from the perspective of novelty, this emphasis may risk overshadowing the value of claims that are well supported and reproducible.

      We have added two sentences at the end of the introduction and at the beginning of the discussion that most claims are reproducible.

      (3) Framing and language: Certain passages could benefit from more neutral phrasing and avoidance of binary terms such as "correct" or "incorrect," in keeping with the open-ended and iterative nature of scientific progress.

      We have carefully checked the phrasing of our article to be more nuanced and avoid black-and-white assessments.

      (4) Community interaction with the dataset: While the website is an excellent resource, the manuscript could further clarify how the community is expected to contribute, challenge, or refine the annotations, especially given the large volume of supplementary data.

      This point is already addressed in the revised sentence: “We hope that the community accessible website will encourage researchers to share their perspectives and contribute data from diverse sources, thereby improving objectivity.” In addition, we have proactively engaged with principal investigators to motivate them to contribute to this effort by sharing unpublished results and information. However, it remains unclear to what extent scientists are genuinely interested in reproducibility, as has been noted in other studies on the topic. For example, articles that challenge previous findings are often under-cited, and the presence of contradictory evidence does not necessarily prevent researchers from continuing to cite the original claims. This is clearly a challenge of reproducibility study. Comments, evidence and corrections will not be integrated in the ReproSci website. We have already updated this database when revising this article and will continue to do it.

      (5) Minor inconsistency: The manuscript states that papers from 1959-2011 were included, but the Methods section mentions a range beginning in 1940.

      This should be aligned for clarity.

      We have solved this discrepancy by correcting the method section. This should be 1959-2011. Thank you for spotting this mistake.

      Impact and significance:

      This contribution is likely to have a meaningful impact on both the Drosophila immunity community and the broader scientific ecosystem. It highlights methodological pitfalls, encourages transparent post-publication evaluation, and offers a reusable framework that other fields could adopt. The work also has pedagogical value for early-career researchers entering the field, who often struggle to navigate contradictory or outdated claims. By centralizing and contextualizing these discussions, the manuscript should help accelerate more robust and reproducible research.

      Reviewer #3 (Public review):

      Summary:

      In this ambitious study, the authors set out to analyse the validity of a number of claims, both minor and major, from 400 published articles within the field of Drosophila immunity that were published before 2011. The authors were able to determine initially if claims were supported by comparing them to other published literature in the field and, if required, by experimentally testing 'unchallenged' claims that had not been followed up in subsequent published literature. Using this approach, the authors identified a number of claims that had contradictory evidence using new methods or taking into account developments within the field post-initial publication. They put their findings on a publicly available website designed to enable the research community to assess published work within the field with greater clarity.

      Strengths:

      The work presented is rigorous and methodical, the data presentation is high quality, and importantly, the data presented support the conclusions. The discussion is balanced, and the study is written considerately and respectfully, highlighting that the aim of the study is not to assign merit to individual scientists or publications but rather to improve clarity for scientists across the field. The approach carried out by the researchers focuses on testing the validity of the claims made in the original papers rather than testing whether the original experimental methods produced reproducible results. This is an important point since there are many reasons why the original interpretation of data may have understandably led to the claims made. These potential explanations for irreproducible data or conclusions are discussed in detail by the authors for each claim investigated.

      The authors have generated an accompanying website, which provides a valuable tool for the Drosophila Immunity research community that can be used to fact-check key claims and encourages community engagement. This will achieve one important goal of this study - to prevent time loss for scientists who base their research on claims that are irreproducible. The authors rightly point out that it is impossible (and indeed undesirable) to avoid publication of irreproducible results within a field since science is 'an exploratory process where progress is made by constant course correction'. This study is, however, an important piece of work that will make that course correction more efficient.

      Weaknesses:

      I have little to recommend for the improvement of this manuscript. As outlined in my comments above, I am very supportive of this manuscript and think it is a bold and ambitious body of work that is important for the Drosophila immunity field and beyond.

      We thank the reviewer for this positive assessment. This was indeed a huge amount of work with the hope this would be useful.

      Reviewer #4 (Public review):

      This is an important paper that can do much to set an example for thoughtful and rigorous evaluation of a discipline-wide body of literature. The compiled website of publications in Drosophila immunity is by itself a valuable contribution to the field. There is much to praise in this work, especially including the extensive and careful evaluation of the published literature.

      However, there are also cautions.

      We praise the reviewer for this assessment of our project.

      One notable concern is that the validation experiments are generally done at low sample sizes and low replication rates, and often lack statistical analysis. This is slippery ground for declaring a published study to be untrue. Since the conclusions reported here are nearly all negative, it is essential that the experiments be performed with adequate power to detect the originally described effects. At a minimum, they should be performed with the same sample size and replication structure as the originally reported studies.

      Many of the original claims we are assessing were obtained with lower sample size than our replicability experiments. Furthermore, we are often complementing our data with alternative approaches that are more robust (survival with additional bacterial strains, AMP kinetics with several time points, etc.). Although we can never claim for a total absence of an effect, our experimental designs are largely sufficient to demonstrate that we cannot report the major effects reported in the initial publications.

      The first section of Results should be an overview of the general accuracy of the literature.

      Of all claims made in the 400 evaluated papers, what proportion fell into each category of "verified", "unchallenged", "challenged", "mixed", or "partially verified"? This summary overview would provide a valuable assessment of the field as a whole. A detailed dispute of individual highlighted claims could follow the summary overview.

      We have described the statistic of our analysis on claim validity in a companion article that will be published back-to-back with this article. We have preferred to divide assessment of claims and statistics on replicability in two articles. The present one is to provide an overview of the main claims that we considered to be fragile. Following the reviewer’s comment, we have added at the start of the manuscript a sentence stating that most of the claims were found to be validated, highlighting the solidity of the field with a link to the other article.

      Section headings are phrased as declarative statements, "Gene X is not involved in process Y", which is more definitive phrasing than we typically use in scientific research. It implies proving a negative, which is difficult and rare, and the evidence provided in the present manuscript generally does not reach that threshold. A more common phrasing would be "We find no evidence that gene X contributes to process Y". A good model for this more qualified phrasing is the "We conclude that while Caspar might affect the Imd pathway in certain tissue-specific contexts, it is unlikely to act as a generic negative regulator of the Imd pathway," concluding the section on the role of Caspar. I am sure the authors feel that the softer, more qualified phrasing would undermine their article's goal of cleansing the literature of inaccuracies, but the hard declarative 'never' statements are difficult to justify unless every validation experiment is done with a high degree of rigor under a variety of experimental conditions. This caveat is acknowledged in the 3rd paragraph of the Discussion, but it is not reflected in the writing of the Results. The caveat should also appear in the Introduction.

      We fully agree with the reviewer and we have softened the way we discuss all the claims in the result section.

      The article is clear that "Claims were assessed as verified, unchallenged, challenged, mixed, or partially verified," but the project is called "reproducibility project" in the 7th line of the abstract, and the website is "ReproSci". The fourth line of the abstract and the introduction call some published research "irreproducible". Most of the present manuscript does not describe reproduction or replication. It describes validation, or independent experimental tests for consistency. Published work is considered validated if subsequent studies using distinct approaches yielded consistent results. For work that the authors consider suspicious, or that has not been subsequently tested, the new experiments provided here do not necessarily recreate the published experiment. Instead, the published result is evaluated with experiments that use different tools or methods, again testing for consistency of results. This is an important form of validation, but it is not reproduction, and it should not be referred to as such. I strongly suggest that variations of the words "reproducible" or "replication" be removed from the manuscript and replaced with "validation". This will be more scientifically accurate and will have the additional benefit of reducing the emotional charge that can be associated with declaring published research to be irreproducible.

      We agree on this point. In response to the reviewer’s comment, we have clarified our terminology. Specifically, we now use the term “conceptual replicability” in the abstract and consistently refer to “replicability” and “irreplicability” throughout the manuscript. We hope this revised terminology improves clarity.

      The manuscript includes an explanatory passage in the Results section, "Our project focuses on assessing the strength of the claims themselves (inferential/indirect reproducibility) rather than testing whether the original methods produce repeatable results (results/direct reproducibility). Thus, our conclusions do not directly challenge the initial results leading to a claim, but rather the general applicability of the claim itself." Rather than first appearing in Results, this statement should appear prominently in the abstract and introduction because it is a core element of the premise of the study. This can be combined with the content of the present Disclaimer section into a single paragraph in the Introduction instead of appearing in two redundant passages. I would again encourage the authors to substitute the word validation for reproduction, which would eliminate the need for the invented distinction between indirect versus direct reproduction. It is notable that the authors have chosen to title the relevant Methods section "Experimental Validation" and not "Replication".

      The disclaimer has been moved at the end of the introduction and one sentence has been added in the abstract.

      Experimental data "from various laboratories" in the last paragraph of the Introduction and the first paragraph of the Results are ambiguous. Since these new experiments are part of the central core of the manuscript, the specific laboratories contributing them should be named in the two paragraphs. If experiments are being contributed by all authors on the manuscript, it would suffice to say "the authors' laboratories". The attribution to "various labs" appears to be contradicted by the Discussion paragraph 2, which states "the host laboratory has expertise in" antibacterial and antifungal defense, implying a single lab. The claim of expertise by the lead author's laboratory is unnecessary and can be deleted if the Lemaitre lab is the ultimate source of all validation experiments.

      The authors and laboratories that did the experiments are listed in the author list and are easy to identify for experts. At the same time, we did not want to highlight the specific authors to avoid a negative impact on their career. We were very happy that a small number of labs helped in our project. Since all the supplement were reviewed by Hannah Westlake (who has left the lab and the project) and myself (BL), our expertise was important as we did not want to test articles far from what we are doing in the lab. I rephrase the sentence to make it clear.

      The passage on the controversial role of Duox in the gut is balanced and scholarly, and stands out for its discussion of multiple alternative lines of evidence in the published literature and supplement. This passage may benefit from research by multiple groups following up on the original claims that are not available for other claims, but the tone of the Duox section can be a model for the other sections.

      Comments on other sections and supplements:

      I understand the desire to explain how original results may have been obtained when they are not substantiated by subsequent experiments. However, statements such as "The initial results may have been obtained due to residual impurities in preparations of recombinant GNBP1" and "Non-replicable results on the roles of Spirit, Sphinx and Spheroide in Toll pathway activation may be due to off-target effects common to first-generation RNAi tools" are speculation. No experimental data are presented to support these assertions, so these statements and others like them (currently at the end of most "insights" sections) should not appear in Results. I recognize that the authors are trying to soften their criticism of prior studies by providing explanations for how errors may have occurred innocently. If they wish to do so, the speculative hypotheses should appear in the Discussion.

      The statement in Results that "The initial claim concerning wntD may be explained by a genetic background effect independent of wntD" similarly appears to be a speculation based on the reading of the main text Results. However, the Discussion clarifies that "Here, we obtained the same results as the authors of the claim when using the same mutant lines, but the result does not stand when using an independent mutant of the same gene, indicating the result was likely due to genetic background." That additional explanation in the Discussion greatly increases reader confidence in the Result and should be explained with reference to S5 in the Results. Such complete explanations should be provided everywhere possible without requiring the reader to check the Supplement in each instance.

      In some cases, such as "The results of the initial papers are likely due to the use of ubiquitous overexpression of PGRP-LE, resulting in melanization due to overactivation of the Imd pathway and resulting tissue damage", the claim to explain the original finding would be easy to test. The authors should perform those tests where they can, if they wish to retain the statements in the manuscript. Similarly, the claim "The published data are most consistent with a scenario in which RNAi generated off-target knockdown of a protein related to retinophilin/undertaker, while Undertaker itself is unlikely to have a role in phagocytosis" would be stronger if the authors searched the Drosophila genome for a plausible homolog that might have been impacted by the RNAi construct, and then put forth an argument as to why the off-target gene is more likely to have generated the original phenotype than the nominally targeted gene. There is a brief mention in S19 that junctophilin is the authors' preferred off-target candidate, but no evidence or rationale is presented to support that assertion. If the original RNAi line is still available, it would be easy enough to test whether junctophilin is knocked down as an offtarget, and ideally then to use an independent knockdown of junctophilin to recapitulate the original phenotype. Otherwise, the off-target knockdown hypothesis is idle speculation.

      A good model is the passage on extracellular DNA, which states, "experiments performed for ReproSci using the original DNAse IIlo hypomorph show that elevated Diptericin expression in the hypomorph is eliminated by outcrossing of chromosome II, and does not occur in an independent DNAse II null mutant, indicating that this effect is due to genetic background (Supplementary S11)." In this case, the authors have performed a clear experiment that explains the original finding, and inclusion of that explanation is warranted. Similar background replacement experiments in other validations are equally compelling.

      We believe that offering possible explanations for discrepancies between results is valuable for readers. For this reason, we have generally retained these elements in the Results section, with the exception of PGRP-SD. In that specific case, we removed the sentence stating that we could not identify the expected mutation in the original PGRP-SD mutants (see below). Although these points could have been moved to the Discussion, doing so would have made it difficult to directly associate each explanation with the corresponding claim. We therefore consider that keeping them in the Results section provides greater clarity. Our hypotheses are grounded in well-established considerations. It is known, for example, that early generations of RNAi constructs often exhibited off-target effects, that commercially available LPS preparations (e.g., from Sigma) were sometimes contaminated, and that genetic background can significantly influence immunological outcomes. It is therefore essential to consider and discuss such confounding factors, particularly those that were not fully appreciated at the time the field emerged.

      Identifying motivated collaborators and coordinating experimental validation proved to be extremely time-consuming, making it impractical to systematically provide experimental verification for every discrepancy. As a result, most of our explanations remain speculative, except for the cases of WntD and DNase II, where the observed differences can be attributed to genetic background effects.

      The statement "Analysis of several fly stocks expected to carry the PGRP-SDdS3 mutation used in the initial study revealed the presence of a wild-type copy PGRP-SD, suggesting that either the stock used in this study did not carry the expected mutation, or that the mutation was lost by contamination prior to sharing the stock with other labs" provides a documentable explanation of a potential error in the original two manuscripts, but the subsequent "analysis of several fly stocks" needs citations to published literature or explanation in the supplement. It is unclear from this passage how the wildtype allele in the purportedly mutant stocks could have led to the misattribution of function to PGRP-SD, so that should be explained more clearly in the manuscript.

      We have removed this statement in the revised version from the result section

      The originally claimed anorexia of the Gr28b mutation is explained as having been "likely obtained due to comparison to a wild-type line with unusually high feeding rates". This claim would be stronger if the wildtype line in question were named and data showing a high rate of feeding were presented in the supplement or cited from published literature. Otherwise, this appears to be speculation.

      The wild-type fly stocks we used are named in the supplement.

      In the section "The Toll immune pathway is not negatively regulated by wntD", FlyAtlas is cited as evidence that wntD is not expressed in adult flies. However, the FlyAtlas data is not adequately sensitive to make this claim conclusively. If the present authors wish to state that wntD is not expressed in adults, they should do a thorough test themselves and report it in the Supplement.

      Alternatively, the statement "data from FlyAtlas show that wntD is only expressed at the embryonic stage and not at the adult stage at which the experiments were performed by (Gordon et al., 2005a)" could be rephrased to something like "data from FlyAtlas show strong expression of wntD in the embryo but not the adult" and it should be followed by a direct statement that adult expression was also found to be near-undetectable by qPCR in supplement S5. That data is currently "not shown" in the supplement, but it should be shown because this is a central result that is being used to refute the original claim. This manuscript passage should also describe the expression data described in Gordon et al. (2005), for contrast, which was an experimental demonstration of expression in the embryo and a claim "RT-PCR was used to confirm expression of endogenous wntD RNA in adults (data not shown)."

      We agree on this but the expression pattern provided from fly atlas and other genomic resource is an additional element that reinforce our conclusion. For instance, the observation that the expression pattern of WntD is limited to early embryo is consistent with a role in the regulation of Toll pathway in D/V patterning but not immunity. The use of the expression pattern is not the key element of our assessments that rely on the use of loss-of-function mutation. We have rephrased the text in the manuscript according to Reviewer’s suggestion.

      Inclusion of the section on croquemort is curious because it seems to be focused exclusively on clearance of apoptotic cells in the embryo, not on anything related to immunity. The subsection is titled "Croquemort is not a phagocytic engulfment receptor for apoptotic cells or bacteria", but the text passage contains no mention of phagocytosis of bacteria, and phagocytosis of bacteria is not tested in the S17 supplement. I would suggest deleting this passage entirely if there is not going to be any discussion of the immune-related phenotypes.

      The reviewer is correct and we have in the revised version removed ‘bacteria’ from the title. We have kept the section S17 because this statement ‘Crq recognize and uptake apoptotic cells’ (Science, Immunity)’ is still a source of confusion while the role of crq is more downstream. We agree that efferocytosis might not considered as part of innate immunity but we consider these articles were important in the Drosophila immunity community at that time.

      The claim "Toll is not activated by overexpression of GNBP3 or Grass: Experiments performed for ReproSci find that contrary to previous reports, overexpression of GNBP3 (Goear et al., 2006) or Grass (El Chamy et al., 2008) in the absence of immune challenge does not effectively activate Toll signaling (Supplementaries S6, S7)" is overly strongly stated unless the authors can directly repeat the original published studies with identical experimental conditions. In the absence of that, the claim in the present manuscript needs to be softened to "we find no evidence that..." or something similar. The definitive claim "does not" presumes that the current experiments are more accurate or correct than the published ones, but no explanation is provided as to why that should be the case. In the absence of a clear and compelling argument as to why the current experiment is more accurate, it appears that there is one study (the original) that obtained a certain result and a second study (the present one) that did not. This can be reported as an inconsistency, but the second experiment does not prove that the first was an error.

      We have softened our statement and revised the supplement document. The title of the section is now ‘Toll is not strongly activated by overexpression of GNBP3 or Grass’. However we want to keep these supplements. The comment of the reviewer applies to all our assessment as we could be wrong. We are using experimental validation and not replication, although in this present case, we are using very similar approach to the described in the original study. Of note, the fact that over-expression of GNBP3 does not consistently activate Drs was already stated in Buchon et al 2019 and we spend quite sometimes in attempts to repeat this experiment. Finally, the Ferrandon team that produce this claim can easily repeat the experiment. This does not affect the main conclusion of their landmark article on GNBP3.

      The same comment applies to the refutation of the roles for Edin and IRC. Even though the current experiments are done in the context of a broader validation study, this does not automatically make them more correct. The present work should adhere to the same standards of reporting that we expect in any other piece of science.

      We fully agree with the reviewer but the comments he/she raised apply to all the assessments made in our article and nearly all the reproducibility experiments. Our assessment does not state that we are right but that we fail to conceptually reproduce the claim. This is clearly mentioned in our article. We hope to see clarification in the future. It may be one day shown that i) IRC indeed regulates intestinal immunity, ii) PGRP-LE will be shown to have an extracellular form, iii) 18W could indeed be a PRR for LPS (maybe in a specific tissues), iv) PGRPSD will be shown to function upstream of Toll, v) that Crq indeed binds apoptotic cell to internalize them… Our study highlights statements that are source of debate, but this is far from being clear. The authors of these challenged claims have the opportunity to contradict us.

      The statement "Furthermore, evidence from multiple papers suggests that this result, and other instances where mutations have been found to specifically eliminate Defensin expression, is likely due to segregating polymorphisms within Defensin that disrupt primer binding in some genetic backgrounds and lead to a false negative result (Supplementary S20)" should include citations to the multiple papers being referenced. This passage would benefit from a brief summary of the logic presented in S20 regarding the various means of quantifying Defensin expression.

      We appreciate the value of citing some of the papers underlying this claim. Key papers in question are discussed at length in Supplement S20, and the original papers do not realize the issue we are reporting. We therefore modified the text as follows: "Furthermore, this result, and other instances where mutations have been found to specifically eliminate Defensin expression, is likely due to segregating polymorphisms within Defensin that disrupt primer binding in some genetic backgrounds, leading to false negative results (Brennan 2007,Neyen 2014 and see Supplementary S20)”. The hypothesis we raise, namely that defensin function may involve specific polymorphisms that are absent from the reference Drosophila genome, is an important consideration that should be communicated to the community.

      In S22 Results, the statement "For general characterization of the IrcMB11278 mutant, including developmental and motor defects and survival to septic injury, see additional information on the ReproSci website" is not acceptable. All necessary information associated with the paper needs to be included in the Supplement. There cannot be supporting data relegated to an independent website with no guaranteed stability or version control. The same comment applies to "Our results show that eiger flies do not have reduced feeding compared to appropriate controls (See ReproSci website)" in S25.

      The ReproSci website contains additional data that we could not display in the supplement but may be useful to scientists. We believe that it is worth to mention them. We have however changed the writing stating ‘Experiments reported in the ReproSci website’ rather than ‘See ReproSci website’ to make it clear that the data are not part of the supplement.

      For IRC: We have added a new section that addresses the general characterization of the IrcMB11278 mutant in the revised version (Supplement 22 Figure 2 and 3 with associated text). We do not refer to the ReproSci website anymore.

      Supplement S21 appears to show a difference between the wildtype and hemese mutants in parasitoid encapsulation, which would support the original finding. However, the validation experiment is performed at a small sample size and is not replicated, so there can be no statistical analysis. There is no reported quantification of lamellocytes or total hemocytes. The validation experiment does not support the conclusion that the original study should be refuted. The S21 evaluation of hemese must either be performed rigorously or removed from the Supplement and the main text.

      To address reviewers’ comment, we generated two new Hemese null mutants, Hemese<sup>JP187</sup> and Hemese<sup>JP828</sup>, carrying deletions in the central region of the gene. Using these lines, we show that wild-type and Hemese mutant larvae display comparable levels of lamellocytes and similar numbers of melanized capsules following wasp infestation. We further quantified lamellocyte differentiation and included appropriate statistical analyses. We have removed the previous results obtained with the Hemese<sup>SK2</sup> mutant, which yielded similar observations. By reproducing our previous findings with independent mutant lines and by providing quantitative analyses, we strengthen our conclusion that Hemese does not act as a negative regulator of lamellocyte differentiation upon wasp infestation.

      In S22, the second sentence of the passage "Due to the fact that IrcMB11278 flies always survived at least 24h prior to death after becoming stuck to the substrate by their wings, we do not attribute the increased mortality in Ecc15-fed IrcMB11278 flies primarily to pathogen ingestion, but rather to locomotor defects. The difference in survival between sucrose-fed and Ecc15-fed IrcMB11278 flies may be explained by the increased viscosity of the Ecc15-containing substrate compared to the sucrose-containing substrate" is quite strange. The first sentence is plausible and a reasonable interpretation of the observations. But to then conclude that the difference between the bacterial treatment versus the control is more plausibly due to substrate viscosity than direct action of the bacteria on the fly is surprising. If the authors wish to put forward that interpretation, they need to test substrate viscosity and demonstrate that fly mortality correlates with viscosity. Otherwise, they must conclude that the validation experiment is consistent with the original study.

      The viscosity of a vial can easily be assessed by direct observation. Here, we provide an information that may guide further research but this is by no way a strong point of our article. Thus, we did not do any change to address this comment. We agree that the function of IRC should be entirely re-visited. We have however changed the name viscosity for stickiness to be more neutral.

      In S27, the visualization of eiger expression using a GFP reporter is very non-standard as a quantitative assay. The correct assay is qPCR, as is performed in other validation experiments, and which can easily be done on dissected fat body for a tissue-specific analysis. S27 Figure 1 should be replaced with a proper experiment and quantitative analysis. In S27 Figure 2, the authors should add a panel showing that eiger is successfully knocked down with each driver>construct combination. This is important because the data being reported show no effect of knockdown; it is therefore imperative to show that the knockdown is actually occurring. The same comment applies everywhere there is an RNAi to demonstrate a lack of effect.

      Eiger expression: Previous RNA-seq (Flyseq, Troha et al., 2018) or Affymetrix array using poly (A)- RNA (De Gregorio et al., 2001) have never identified eiger as an immune-induced gene. We have used a previously characterized GFP reporter gene using the same approach done by the author to attempt to see an induction in the fat body without success. Our conclusions have been nuanced.

      Eiger RNAi KD: The two RNAi lines we are using have already been used in many studies assessing the function of eiger ((#108814, reference here, #45253 reference here)

      The Drosomycin expression data in S3 Figure 2A look extremely noisy and are presented without error bars or statistical analysis. The S4 claim that sphinx and spheroid are not regulators of the Toll pathway because quantitative expression levels of these genes do not correlate with Toll target expression levels is an extremely weak inference. The RNAi did not work in S4, so no conclusion should be inferred from those experiments. Although the original claims in dispute may be errors in both cases, the validation data used to refute the original claims must be rigorous and of an acceptable scientific standard.

      S3: The inducibility of antimicrobial peptide genes is high, albeit very variable. Importantly, we do not observe any inhibition of Drs expression at the two time points (24 and 48 h) following infection with two distinct bacteria, M. luteus and E. faecalis. Although each experiment was performed once, this corresponds to four independent tests challenging the claim from my lab that “Spheroide is required for activation of the Toll pathway.” Furthermore, if Spheroide played a significant role in Toll pathway activation, increased susceptibility to E. faecalis, B. subtilis, and B. bassiana would be expected. However, both survival data (three independent experiments) and Drs expression profiles (four conditions) consistently indicate that Spheroide is not required for Toll pathway activation. For clarity, we have removed the data related to S. aureus.

      S4: To address the reviewer’s comments and reinforce our conclusion, we generated a deletion removing both sphinx 1 and sphinx 2 and analyzed Toll pathway activity in the absence of these two serine proteases. We find that flies lacking sphinx 1 and sphinx 2 exhibit wild-type induction of Drosomycin following M. luteus infection, as well as normal survival upon E. faecalis challenge. These new data provide strong evidence supporting our claim that, in contrast to our previous report (Kambris et al., 2006), Sphinx 1 and Sphinx 2 do not regulate the Toll pathway. Accordingly, we have replaced the earlier RNAi-based results with data obtained from loss-of-function mutants.

      In S6 Figure 1, it is inappropriate to plot n=2 data points as a histogram with mean and standard errors. If there are fewer than four independent points, all points should be plotted as a dot plot. This comment applies to many qPCR figures throughout the supplement. In S7 Figure 1, "one representative experiment" out of two performed is shown. This strongly suggests that the two replicates are noisy, and a cynical reader might suspect that the authors are trying to hide the variance. This also applies to S5 Fig 3. Particularly in the context of a validation study, it is imperative to present all data clearly and objectively, especially when these are the specific data that are being used to refute the claim.

      S6: The figure 1 of S6 shows all the points. Collectively, this represents 6 replications using 3 drivers of the claim ‘over-expression of GNBP3 efficiently activate the Toll pathway’. The observation, we are now reporting in S6, was already mentioned in Buchon et al 2009 (PNAS). To address reviewer’s concern, we have removed the standard deviation. Finally, our conclusion is rather soft ‘Overexpression of GNBP3 either ubiquitously or in the fat body does not effectively activate the Toll pathway’. At most, we find similar to (El Chamy et al., 2008) that overexpression may produce ‘weak but detectable’ activation of Toll’.

      S7: We could not address this point. The point stands, there is no condition where overexpression of native Grass corresponds with higher Drosomycin expression

      Other comments:

      In S26, the authors suggest that much of the observed melanization arises from excessive tissue damage associated with abdominal injection contrasted to the lesser damage associated with thoracic injection. I believe there may be a methodological difference here. The Methods of S27 are not entirely clear, but it appears that the validation experiment was done with a pinprick, whereas the original Mabary and Schneider study was done with injection via a pulled capillary. My lab group (and I personally) have extensive experience with both techniques. In our hands, pinpricks to the abdomen do indeed cause substantial injury, and the physically less pliable thorax is more robust to pinpricks. However, capillary injections to the abdomen do virtually no tissue damage - very probably less than thoracic injections - and result in substantially higher survivals of infection even than thoracic injections. Thus, the present manuscript may infer substantial tissue damage in the original study because they are employing a different technique.

      The reviewer is correct that injection and pinprick injury do not induce the same type of injury and there are differences across labs. We still believe that our approach is valuable. The point is that we confirmed a role of eiger in melanization but not due to its expression in the fat body. It is important to underline that our ReproSci is a conceptual replication project and that we never directly challenge the original observation but rather the robustness of the claim. Injection with a pulled capillary was attempted — in our hands, this caused extensive trauma and killed over 80% of flies it was attempted on. The pinprick on the other hand caused very little trauma. The first sentence of the Results/discussion section states this point.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It is questionable whether all the studies relevant to the field of assessment have been identified, starting with the own publications from the lead authors. Indeed, Buchon et al, Genes & Development, 2009, Buchon et al., Cell Host&Microbe, 2009, Buchon et al., BMC Biology, 2010 have not been included in the 400 articles, even though one of them is cited when discussing the role of Duox in the intestinal host defense. Other missing articles that fulfill the criteria set by the authors include, but may not be limited to, Chung et al. JBC, 2011, Nehme et al., PLoS One, 2011, Cronin et al., Science, 2009, Limmer et al. 2011.

      In our selection, we did not choose articles dealing with intestinal homeostasis and stem cells as well as microbiota/symbionts (Buchon 2009a,b and 2010 Cronin 2009). We also did not analyze articles centered on virulence factor (Limmer 2011). I agree that we missed Chung et al. JBC, 2011 and Nehme et al., PLoS One, 2011 that were published in 2011. Table S1 provide all the information on the article we selected and the ones we removed from the first search.

      One area of the field that would have benefited from further insights of ReproSci is that of innate immunity memory or priming (Pham et al., PLoS Pathogens, 2007, has been included in the 400 articles). That being said, the scope of this article is already considerable, and this is just a suggestion for future scrutiny.

      The reviewer should recognize that this study (Pham et al.) is inherently complex to reproduce. The field of innate immune priming has also matured considerably since the first publications. Importantly, the set of articles containing “non-validated claims” does not fully overlap with the category of studies that might be considered as weak or problematic.

      A major issue in science is the exaggeration of claims, which our approach does not capture. As a result, some articles that are strongly promoted yet make relatively trivial claims— often a source of frustration within the community—may still be classified as valid under our framework. This represents an important limitation of our approach.

      For the Supplements, it would be better to show all the data points, especially with respect to RTqPCR data. The nature of the error bars (SD, SE) is never indicated except for one exception (Figure 5, S28). When introducing or using novel lines affecting gene expression, such as RNAi lines or KO mutants, there are often no indications as to how these lines have been validated.

      This is especially critical for negative results.

      Errors bars: The errors bars were standard deviation. Sorry for this lacks of clarity. This is now indicated in the supplement. Experiments were done in multiple laboratories requiring complex coordination. For this reason, changing all the graphs was too complicated and could not be done for all the figures in the revised version. However data points are shown in many figures of the original draft and all the new figures of the revised version.

      For RNAi validation: Concerning the RNAi validation:

      (a) Sphinx ½: We have extended our work on Sphinx using a newly generated double

      mutant that remove both Sphinx1 and 2. Thus our conclusion does not rely anymore only the use of RNAi.

      (b) S12: Hep RNAi : We are using standard Hep-RNAi and BskDN fly tools that have been validated in many studies.

      Articles validating the HepRNAi (BDSC35210) can be found on this link. This line has also been validated in Figure 1 of the revised version of Supplement S12.

      Articles validating the BskDN (BDSC6409) can be found on this link.

      Of note, the Bloomington number of Hep RNAi was miss-annotated and has been corrected in the revised version.

      (a) S16. Dscam: The Trip RNAi lines has been used in other articles (ex. Bui et al., “Adjacent neuronal fascicle guides motoneuron dendritic branching 2024 (eNeuro)).

      (b) S23. Duox. We are using two RNAi lines including the original RNAi construct used in the article we assessed (Ha et al., 2005).

      (c) Eiger: we are using two distinct UAS-RNAi constructs that have been validated (#108814, reference here, #45253 reference here).

      Loss-of-function mutations:

      For previously described mutants: we used some previously described mutants available in stock centers or from other labs. These include Spheroide<sup>D104</sup>, wntD mutants Irc<sup>MB11278</sup>, NOS<sup>Δall</sup> and Gr28b. See references in the supplement. We have always checked the genotypes of these mutants.

      For newly generated mutants

      spirit<sup>DR32</sup> see Figure 1 in S2

      edin<sup>KO</sup>: we have added a new figure describing the mutant in S14 D

      ΔSphinx1/2 we have added a figure describing the mutant in S4

      DNaseII<sup>sk4</sup>: the exact frameshift mutation is indicated in S11

      SR-CI<sup>SK6</sup> and Δ(Sr-CI,Sr-CIII): we have added a figure describing the mutant in S18

      Hemese<sup>JP187</sup> and HemeseJP<sup>828</sup> we have added a figure describing the mutant in S21

      Listericin<sup>Δ1</sup> and Listericin<sup>Δ2</sup> mutation are described in Figure 1 in S15

      We have now included in the revised supplement a table with all the new tools generated in the framework of the ReproSci project (Supplementary Table 4).

      The only uncharacterized mutant is PGRP-LE<sup>Δ53</sup>, a gift from François Rouyer (Gif-sur-Yvette). But we also did our experiments with an established mutant: PGRP-LE<sup>112</sup>.

      Is the repeated use of "in vitro" to refer to cell culture studies appropriate? In vitro usually refers to experiments performed in the absence of any living organism or cell, e.g., biochemical reactions in a test tube.

      Accordingly, we have replaced in vitro by in cell culture when appropriate

      (A) MAIN TEXT

      (1) Introduction:

      PGRP-LC: Ramet et al, Nature 2002 reference is missing.

      "[...]epithelial repair through stem cell proliferation" is a typical disease tolerance mechanism and therefore unlikely to "[...]contribute to DETERRING ingested pathogens". In addition, is "detterring" really the appropriate word?

      The changes have been made.

      (2) Results:

      GNBP1: "and find no evidence" would read better as "and found no evidence". In addition, for GNBP3, the authors may also want to cite Mishima et al. (JBC, 2009) and possibly one on silkworm ßGRP (Takehasi et al., PNAS, 2009).

      The changes have been made.

      Wnt D: "Dif" and not "dif"=diffuse irregular facets

      The changes have been made.

      cGLR-STING pathway (Cai et al., 2022): subsequent studies indicated that cGLR ligands are RNA molecules, not DNA. Thus, whether such receptors would be able to detect naked DNA remains a remote possibility.

      We agree and have amended the text accordingly

      DSCAM1: the statement "not amenable to genetic studies" might be modulated into "not easily amenable to genetic studies" as exemplified by Dong et al. PLoS Biol, 2025 (which the authors may now add to their reference list as this article was published after submission of the current work) and all the literature on DSCAM and the Drosophila nervous system.

      We agree and have amended the text accordingly

      Psidin: "multiple papers suggest...". References are not given.

      We have amended the text and reported to the supplement. Some references have been added.

      NOS: "this study reports" should be "this study reported"; otherwise, the reader falsely believes that the authors refer to the Reprosci article. NOS expression: what about data from subsequent RNAseq studies and data from FlyCell Atlas?

      We agree with the reviewer and used the past time. For simplicity, we did not discuss Nos expression pattern.

      Eiger: the Kodra et al. reference appears to be missing in the reference list.

      The Kodra reference has been added.

      (3) Discussion:

      Third source of irreproducible claims: besides off-target effects, a limitation of the RNAi approach is that one can never be sure that a null phenotype is obtained. Thus, when obtaining negative results, it may be caused by a hypomorphic effect.

      This additional point has been added with an example that illustrates the point “Besides off-target effects, a limitation of the RNAi approach is that one can never be sure that it mimics a full null phenotype. As an example, the partial silencing of the serine protease gene Grass may explain why its role in the antifungal response was not detected using in vivo RNAi, but was revealed through the use of a null mutation (El Chamy et al., 2008; Kambris et al., 2006b)”.

      (4) Reference list:

      Cuttell et al, Gordon et al, 2005, and Melcarne et al, 2019 are duplicated.

      Thank you for noticing it. Change has been made.

      (5) Supplementary materials:

      Table S2 and S3 appear to have been interchanged.

      This has been corrected.

      Ouyang et al: BioRXiv should be indicated.

      This has been corrected.

      (B) INDIVIDUAL SUPPLEMENTS:

      (1) S1: enzymatic activity of GNBP1.

      The number of independent experiments does not appear to be mentioned. With respect to the western blot shown in Figure 2A, it seems that the intensity of the band corresponding to activated ModSP is lower in lane 1 than in other lanes. The suggestion is to scan multiple blots to quantify the activation of this protein in different conditions. Do the authors know the nature of the high mobility band that appears solely when S. aureus peptidoglycan is added?

      Regarding the lower intensity of the activated ModSP band in lane 1, a similar phenomenon has been reported in An evolutionarily conserved serine protease network mediates melanization and Toll activation in Drosophila (Fig. S2), where ModSP autoactivation was observed following incubation with PGRP-SA, GNBP1, and peptidoglycans. In current experimental design, the primary objective was to determine whether preincubation of peptidoglycans with GNBP1 further enhances ModSP activation. Therefore, our analysis focused on comparisons across different incubation time points rather than direct comparison with lane 1. Within this context, the lower intensity of the activated ModSP band in lane 1 is expected and does not affect the interpretation of the results

      In response to the suggestion of scanning multiple blots for quantification, we note that no clear enhancement of ModSP autoactivation was observed across different incubation times. We also explored alternative incubation conditions, but none resulted in a detectable increase in activation. Thus, additional quantification would not change the overall conclusion.

      (2) S2: Spirit

      It would have been nice to display the spirit mRNA levels in the KO and KD mutants. Is the catalytic domain removed by the deletion and can the authors exclude a potential expression of a truncated Spirit from a cryptic translation initiation site upon leaky Gal4 expression of the transgene, hence the need for RTqPCR data using primers that map 3' to the transposon insertion?

      We are using a deletion that remove all the Spirit isoforms that all start with the same ATG which was removed in the mutant, so it's very unlikely. There is an HSP70T = Terminator after the Gal4, which should exclude an initiation of translation after Gal4. However, we have verified the expected mutation by PCR.

      Figure 2: Crystal cells appear to be present in spirit larvae. They, however, appear to be smaller, which may be difficult to measure. Quantification of the number of crystal cells is necessary.

      We have removed this part in the revised version (Figure 2C and associated text) because it does not concern the assessed claims in Kambris et al et., 2006 that Spirit regulates the Toll pathway.

      Figure 3: It is not clear if pooled data are presented for survival experiments.

      Yes survivals show pooled data. We have updated figure legend to make this clearer.

      (3) S4: Sphinx 1/2

      In the absence of testing null mutant flies, the conclusion may be too strong. Can the authors totally exclude the expression of these proteases in a few cells? If they function as apical proteases, a minimal expression may be sufficient to trigger a proteolytic cascade.

      To address the reviewer’s comments, we generated a deletion removing both sphinx 1 and sphinx 2 and analyzed Toll pathway activity in the absence of these two serine proteases. We find that flies lacking sphinx 1 and sphinx 2 exhibit wild-type induction of Drosomycin following M. luteus infection, as well as normal survival upon E. faecalis challenge. These new data provide strong evidence supporting our claim that, in contrast to our previous report (Kambris et al., 2006), Sphinx 1 and Sphinx 2 do not regulate the Toll pathway. Accordingly, we have replaced the earlier RNAi-based results with data obtained from loss-of-function mutants.

      (4) S5: wntD

      Figure 1, if any, appears to be missing.

      Thank you for noticing this. This is now corrected Figure 2=>Figure 1 and Figure 3=> Figure

      Fig. 3C-C': please, homogenize the color code with other panels (yw in blue and not suddenly in orange)

      We have homogenized the color code in the revised version.

      (5) S6-S7: Activation of Toll by overexpression of GNBP3 or protease genes The point is not framed appropriately: the overexpression studies were not intended to determine whether the overexpression of these genes activates the Toll pathway to levels of activation encountered during infection, but as tools for epistatic analysis, and in the case of S8, to identify novel protease genes that might be involved in Toll pathway activation. The results of the epistatic analysis are not disputed and, in some cases, have been validated by independent approaches (e.g., by hemolymph transfer experiments), and the hits of the protease screen have been validated by other laboratories, namely grass and SPE. In Figure S5B of Gottar et al, Cell, 2006, a level of only 30% of immunized was achieved, whereas it was about 70% in Figure 3A. Thus, the level of induction by overexpression of GNBP3 is variable but nevertheless sufficient for epistatic analysis both for Gottar et al. and for El Chamy et al., Nat. Immunol, 2008. Furthermore, levels of activation sufficient for epistatic analysis of Toll pathway activation or stimulation of the melanization proteolytic activation cascade upon the overexpression of either GNBP3 or GNBP1 together with PGRP-SA have been reached (Gobert et al., 2003; Gottar et al. 2006; El Chamy et al, 2008; Matskevich et al., Eur. J. Immunol., 2010). Thus, the overexpression of GNBP3 functions for genetic analysis purposes in two independent teams, that however, are working in the same research Unit. While one cannot exclude that enough peptidoglycan metabolites, e.g., TCT, are provided by the microbiota, it cannot be excluded that ß-(1-3)glucans might originate from the food or mycobiota, which may be distinct from that in the Reprosci laboratory. This illustrates the difficulty of controlling all parameters.

      We fully agree that the notion that GNBP3 functions upstream of serine protease is validated. The point is that we cannot reproduce the observation that overexpression of GNBP3 trigger a strong Toll activation sufficient for epistatic analysis as shown in Gottar et al., 2006. The point is that ReproSci does not challenged any of the major claims of Gottar et al. 2006 but only a minor claim. The magnitude of gene induction relative to the unchallenged condition is an important consideration in the field of Drosophila immunity and can be misleading. For example, a gene described as being induced fivefold may appear significant, whereas the same gene can be induced up to 1,000-fold following infection. Therefore, absolute expression level compared to challenged level, rather than fold changes alone, should be carefully considered when interpreting immune gene activation.

      We have done our studies with two different constructs. The observation that overexpression of a gene activate an immune pathway is by itself a result suggesting a concentration-dependent mode of activation. For instance, overexpression of ModSP is sufficient to activate the Toll pathway (Buchon 2009). This was not the case for GNBP3, even at low level. We cannot exclude that this could be the case in a specific background but at least not all background. Over-expression of GNBP3 is used in 3 panels of Gottar 2006 and this was a significant piece of data. On my side, my lab has lost significant amount of time to reproduce this without success.

      Now we could be wrong in our assessment but we let the author of this study to show it.

      To take in consideration Reviewer’s comment, we are softened our conclusions : ‘Overexpression of GNBP3 either ubiquitously or in the fat body does not effectively activate the Toll pathway’. At most, we find similar to (El Chamy et al., 2008) that overexpression may produce ‘weak but detectable’ activation of Toll’.

      References:

      Alphabetical order is not respected (Matskevich before Kim), and the Mishima reference is not optimal: it should be the JBC article in the same year that documents also the biochemistry of binding to ß-glucans.

      We have done the changes in the revised version

      (6) S8: PGRP-LE not in hemolymph

      Fink et al., Mucosal Immunol. 2016 used the NP1-Gal4, which is not expressed in enteroendocrine cells, as stated in the text, but in enterocytes (and possibly the nervous system).

      For controls, is the wild-type control the w[iso] Drosdel background?

      Please provide the supplier of the GFP antibody.

      We have addressed all the points in the revised version.

      (7) S9: PGRP-LE and melanization

      Figure 1D-E: A quantification would bolster the claim.

      We agree that quantification of Figure 1D-E would reinforce the claims but we believe that our data are sufficient to state that PGRP-LE does not block the cuticular melanization upon septic injury (which is also consistent with an intracellular role). We have softened our conclusion.

      (8) S10: PGRP-LE and bacillus

      Please provide a description of the Bacillus subtilis strain.

      This is a strain that derived from the Jean Lambert collection of Strasbourg and used in Lemaitre B PNAS 1997 (J.Millet and A. Klier, Pasteur Institute of Paris). We have checked the strain and confirm that it was indeed a B. subtilis strain that we name B subtillis FS.

      It is not clear from the Mat & Meth whether the 2-3 independent replicates per genotype represent biological replicates in one experiment or correspond to several independent experiments. Are the pooled data displayed?

      This information was on the methods. It has been copied to the figure caption.

      (9) S11: DNAse II and regulation of AMP genes

      The exact larval stage is not provided in S11, nor in the original study.

      The experiment was done on wandering third-instar larvae as indicated in the material and methods

      The number of flies for survival experiments is rather low, but this is not a major issue given the outcome of the experiment.

      We agree with the reviewer.

      What is the concentration of the injected DNA? As the solution was likely viscous, was the DNA sheared?

      We have used a concentration of DNA of: 200 ng/ul, which is now indicated in the revised version of the Supplement. DNA was no sheared.

      Figure 1A&D: A positive control of immunization would have been nice. 1D: At which time point was Diptericin expression measured by RTqPCR?

      Unchallenged flies were used in Figure 1A while in Figure 1D Diptericin was measured 6 h after challenged. We agree that a bacterial challenge control would have been nice, but clearly the level of Diptericin in 1A corresponds to levels observed in unchallenged flies. Of note, we have tried several different concentrations, of more or less long "sizes" of DNA (short or very long fragments), and also different DNA sources (E. coli or Drosophila) but we did not see any effect on the activation of the Imd pathway.

      DNAseII-deficient larvae have wild-type hemocyte numbers. His sub-title is not accurate as Figure 3 actually reports an INCREASED number of hemocytes in the DNase II mutant.

      The reviewer is correct and we have adjusted the text.

      (10) S12: JNK pathway and AMP gene expression

      Hep RNAi lines and possibly also the Bsk-DN line should be validated by monitoring puckered mRNA levels after a stress that activates the JNK pathway. The absence of this control weakens the Reprosci conclusions.

      As stated above, we are using standard Hep-RNAi and BskDN fly tools that have been validated in many studies.

      Articles validating the HepRNAi (BDSC35210) can be found on this link. We have also validated the RNAi in new Figure 1 of Supplement S12.

      Articles validating the BskDN (BDSC6409) can be found on this link

      The Bloomington number of Hep RNAi was miss-annotated and has been corrected in the revised version

      (11) S13: Caspar

      There is clearly a trend for improved survival of the Caspar mutant in Figure 1A. Please, show the pooled data and appropriate statistical analysis. The legend to 1A mentions one overexpression strain that is not displayed and has likely not been performed because loft experiments were performed at 25{degree sign}C, whereas OE experiments should be performed at 29{degree sign}C, hence in separate experiments.

      There was a mistake in the original figure legend when we mentioned the impact of overexpression of caspar on survival without showing the data (no effect was observed). The text has been removed in the revised version because it is not related to the claims we assessed from the Caspar article (Kim et al.,2006). To addressed reviewer’s comments, we now present the pooled data in Figure 1 of the revised Supplement S13. The conclusion is that we do not observe any increase host survival in caspar mutant related to wild-type consistent with our observation that the Imd pathway is not strikingly over-activated in absence of infection in caspar loss-of-function mutants.

      Why was the c729-Gal4 driver chosen? Has it been validated as a bona fide fat body driver?

      The use of c729-Gal4 was suggested by, and borrowed from, the laboratory of Prof. Gaiti Hasan, who had characterized and used it in adult fat body experiments (Subramanian M et al., 2013, Dis. Mod. & Mech). The BDSC number for the corresponding stock is BDSC_6983. Based on Flybase data, it is also expressed in somatic cells of male/female reproductive systems in the adult.

      The RTqPCR experiments are not overwhelmingly convincing for a lack of data points: pooled data should be shown, inasmuch as there is a significant difference upon Caspar OE 6h after Ecc15 challenge.

      We now show in the revised version of the Supplement S13 the pooled data from different experiments with individual points shown.

      In conclusion and although our results are variable, they do not reproduce the results produced in the original article by (Kim et al., 2006) According to our results, we observe:

      (1) no increase Dpt expression in absence of challenge both in caspar deficient larvae and adults.

      (2) A possible inhibition of Dpt upon over-expression caspar at 6h in adult.

      - A possible increased Dpt expression in caspar loss of function mutant larvae upon challenge.

      The later may suggest an inhibition role upon overexpression of caspar that required further analysis. The text of the supplement has been amended to take all those points in consideration.

      (12) S15: listericin

      Introduction: The LLO (also referred to as hly) L. monocytogenes mutant does not prevent the entry of Listeria inside cells. The pore-forming toxin listeriolysin is required for pathogen escape from the phagosome.

      There is no reference to LLO in the introduction of the S15 listericin section.

      Has the STAT-92E RNAi line been validated? It is difficult to draw a definitive conclusion on the absence of the role of the JAK-STAT pathway in adults using an RNAi transgene because it may only imperfectly silence the expression of the target gene. "Contrary to Goto et al.2010, in our hands reduction of JAK-STAT activity through RNAi [...] in adult flies": Goto et al. apparently tested the JAK-STAT pathway only in S2 cells. This statement should therefore be modulated accordingly.

      We have validated the STAT92 RNA in the revised version by confirming that Turandot gene expression is reduced in this mutant (Supplement 15 Figure 2 C-D). We have also validated the PGRP-LE<sup>112</sup> null mutant (Supplement 15 Figure 1C).

      We have adjusted the conclusion: “Contrary to (Goto et al., 2010), we find that Listericin expression is not specifically dependent on PGRP-LE, and is not dependent on JAK-STAT signaling in flies. Of note, this could be the case in S2 cells where the initial experiments were done.

      Antibacterial effects of listericin: It would be cautious not to rely solely on survival experiments, but also to monitor the bacterial titer to draw a definitive conclusion (assay used by Goto et al.). A control overexpressing GFP or a non-relevant peptide would have been welcome to control for potential metabolic effects that would decrease the effect of the overexpression. More importantly, Goto et al. used a Cg-Gal4 driver that is also expressed in hemocytes, which is likely not the case for the Lpp driver. Is it known if the c564 driver is also expressed in hemocytes and to the same level as the Cg-Gal4 driver?

      We have used two fat body drivers to overexpress Listerin, Lpp and C564, the latter being expressed in hemocytes indeed. We agree that bacterial counting would provide a sensitive assay but in absence of any protective effect of Listericin overexpression, we do not expect any major titer change in L. monocytogenes growth. To reinforce our conclusion, we have generated Listericin null mutations (Fig 4C) and we have shown that they are not more susceptible than the wild-type to L. monocytogenes.

      (13) S16: DSCAM1

      The authors use an RNAi construct to knock down DSCAM1 expression. Which part of the transcripts is targeted? Did the authors verify that silencing was effective and would affect all DSCAM transcripts?

      The DScam RNAI (BDSC #38945) that we are using has been used in other articles (link) and is validated in Kamiyama et al., Dev cell (Figure S4).

      Data points should be shown in the figure.

      Unfortunately, we could not address this point to the difficulty to reach the collaborator who provided the data.

      Their conclusions would be strengthened by directly testing opsonization using the assay developed by Haller et al., EMBO Reports, 2018. Also, to derive the phagocytic index, the authors must have counted the number of hemocytes retrieved from bled larvae? Was there a decrease when DSCAM1 was silenced in hemocytes? The use of a hml-Gal4-Gal80ts driver would have allowed to bypass the issue of developmental effects.

      We expect an opsonin to be abundant in the hemolymph but we did not find any secreted form of dScam in our proteomics analyses (Rommelaere 2024). 20 years after the publication of this landmark article in Science, the absence of any confirmation is puzzling and we are reluctant to do any additional experiment taking into consideration that phagocytosis assays are often variable.

      (14) S18: SR-CI

      The major issue with this supplement has been dealt with in the Public Review: no phagocytic receptor function was claimed in the Ramet et al, 2001 study.

      We fully agree with the reviewer and with Mika Rämet and Monty Krieger (personal communication), and we apologize for this misunderstanding. In fact, Rämet et al. (2001) did not present data demonstrating a role for Sr-CI in phagocytosis per se in his article, but rather showed that this receptor is involved in bacterial binding. Our confusion arose from a subtitle in the Results section of their article (“SR-CI Binds Both Gram-Positive and Gram-Negative Bacteria and Is Necessary for Optimal Phagocytosis by S2 Cells”) and the article title that claim that Sr-CI is a pattern recognition receptor, which does not reflect the content of the article.

      To clarify this point, we have revisited the original claim. Consistent with Rämet et al. (2001), we now observe a defect in hemocyte binding to bacteria in Sr-CI mutants. Thus, we validate the claim. Because Sr-CI is often assumed to function as a phagocytic receptor, we have also retained the experiments showing that it is not essential for the phagocytosis of E. coli or S. aureus. All experiments have been repeated during the revision process using a null single mutation and a double mutant removing both Sr-CI and Sr-CIII.

      Introduction: Is it really adequate to mention a phagocytic receptor for a ligand, dsRNA, that gets endocytosed?

      In the introduction, we summarized what is known on the Sr-CI, including its implication in the uptake of dsRNA. As suggested by the reviewer, we used the term endocytosis to refer to the ability to uptake dsRNA.

      Methods: Have the authors excluded the potential off-target effect of CRISPR-Cas9 in other scavenger receptor genes? Monitoring their expression levels would also allow the detection of any compensatory effect in the null mutant by overexpression of other scavenger receptors.

      We are now using two mutant lines: one carrying a frameshift mutation (SK6) that specifically affects Sr-CI, and another with a deletion removing both Sr-CI and Sr-CIII. We have carefully validated both mutations by genomic sequencing. As our results are now consistent with the original findings of Rämet et al. (2001), we believe that the analyses presented in the Supplementary Information are genetically well supported.

      Results: Figure 1 is really confusing with so much variability in the data, which raises doubts about the reliability of the method. The reduction of the phagocytic index in B is not "mild". Again, one would like to see the individual data points and to know what the error bars are. For 1G, has any statistical analysis been performed? Finally, Panels A to G lack homogeneity in their presentation. Also, why are the absolute values of the y-axis varying so much from panel to panel (A-B vs C-D; E vs. F)?

      We have repeated our experiments using a new set of mutants, both generated in the isogenic DrosDel background. In the revised version, we include additional experiments assessing the contribution of Sr-CI to hemocyte binding to bacteria. Consistent with the original report by Rämet et al. (2001), we observe similar results. We have revised the text accordingly and moderated our claims. In addition, we performed further biological replicates and now provide appropriate statistical analyses.

      (15) S19: Undertaker

      Undertaker has been identified through a deficiency screen that restricted the genomic zone of interest to eight genes. RNAi was then used to identify the candidate gene as being Undertaker. While an off-target effect of RNAi cannot be excluded, it cannot account for the results of the deficiency screen (junctophilin is on the second chromosome, whereas undertaker is on the third). The deficiency phenotype could be rescued by an Undertaker transgene. While overexpression can indeed provide a rescuing activity due to high levels of ectopic expression, why would it do so in the context of an 82F deficiency, unless this deficiency carries a second site mutation, which is a rather remote possibility.

      Given that the deficiency line deletes at least 14 genes in a cluster, several of which are involved in growth and development, we assume that another gene(s) within the span of this deficiency is responsible for the effect on efferocytosis, or that some other mutation in the uncontrolled genetic background is responsible. The impact of Df(3R)3-4 on efferocytosis remains to be investigated.

      The authors claim that Undertaker is not expressed at the right embryonic stage and hardly in S2 cells, which is correct according to large-scale data. However, Cuttell et al. successfully stained S2 cells with an antibody raised against retinophilin/undertaker by the investigators who studied the role of this gene in the eye (Mecklenburg et al., 2007). Should we trust more largescale datasets based on transcriptomics or a specially developed tool? (Both junctophilin and undertaker are at the limit of detection of the proteomic data available on FlyBase; equivalent signals in hemocytes were measured in FlyCellAtlas). Thus, the argument of the absence of expression is not definitive.

      We agree on this but the expression pattern is an additional element that reinforce our conclusion. This is not the key element that relies on the use of loss-of-function mutation. Another element is the absence of follow-up study in the last 18 years.

      The new phagocytosis data set generated by Reprosci shows an absence of impact of the undertaker loft mutation in larval hemocytes. Whether these data can be extrapolated to S2 cells, embryonic or adult hemocytes, remains debatable, as are the explanations put forth by Reprosci in their conclusion to explain by a possible off-target effect of RNAi on junctophilin. It would be interesting to determine the phenotype of junctophilin mutants in the ex vivo larval hemocyte assay.

      We have analyzed the role of Undertaker using third instar larvae in an ex vivo assay and did not find any role in efferocytosis. I think that all our data and analysis of the literature is consistent with our statement that’ that undertaker is unlikely to have a role in phagocytosis’. Now, we can never exclude that Undertake may play a subtle role in this process. The role of Draper has been observed in larval hemocyte that represent macrophage of flies. We are using a null mutation in undertaker preventing the existence of off-targets. We have nuanced our conclusion in the article and add a sentence in the conclusion of the supplement: ‘We cannot however exclude that Undertaker play a role at a specific stage of development that we did not assess.’

      (16) S20: Psidin

      This supplement is duplicated

      The duplicated supplement has been removed.

      Phagocytosis assay: the authors use E. coli-GFP. How can they discriminate between internalization and simple binding at the surface of hemocytes? It would have been more appropriate to use a phalloidin, stained to another color rather than green. The suggestion here is to first use FITC-labeled bacteria, followed by Trypan blue treatment to quench the fluorescence of noninternalized bacteria. Second, pHrodo bacteria could be used to monitor the acidification of the phagosome in wt and psidin hemocytes.

      We agree with the reviewer that this would be better. Because our observations are consistent with the data reported in this original article, we did not feel inclined to extend our analysis.

      (17) S21: Hemese

      The pie charts in Figure 1 would benefit from a statistical analysis. Why do the encapsulated eggs appear to be much larger in the hemese mutant? Could it be that more lamellocytes are recruited and make a thicker capsule? Should Figure 1C be imperatively quantified?

      To extent our analysis, we generated two new Hemese null mutants, Hemese<sup>JP187</sup> and Hemese<sup>JP828</sup>, carrying deletions in the central region of the gene. Using these lines, we show that wild-type and <sup>Hemese</sup> mutant larvae display comparable levels of lamellocytes and similar numbers of melanized capsules following wasp infestation. We further quantified lamellocyte differentiation and included appropriate statistical analyses. We have removed the previous results obtained with the Hemese<sup>SK2</sup> mutant, which yielded similar observations. By reproducing our previous findings with independent mutant lines and by providing quantitative analyses, we strengthen our conclusion that Hemese does not act as a negative regulator of lamellocyte differentiation upon wasp infestation.

      (18) S22 IRC

      IRC is actually not a catalase but a heme peroxidase. For the experiments on C. albicans, can the authors exclude that the IRC flies succumb because they are more sensitive to ethanol than wt or spz flies (sucrose is fermented by C. albicans): a control would be to use glycerol (nonfermentable) instead of sucrose solution.

      We thank the reviewer and Carolina Barillas (personal communication) to inform us that IRC is not a catalase but an heme peroxidase. This is now included in the article and the supplement. It is clear that the name IRC for immune regulated catalase is misleading. We cannot exclude that the IRC flies succumb because they are more sensitive to ethanol than wt or spz flies but we did not test this interesting idea during the revision.

      Do IRC mutants succumb to the ingestion of killed bacteria?

      No, they get stuck in fly food and die.

      (19) S23 DUOX

      Introduction: When citing Kumar et al., the authors might also want to cite Bai et al. on the Bactrocera dorsalis peritrophic matrix.

      This reference has been added

      Results:

      Figure 1: any statistical analysis? Data points are not shown for the lef panel, and error bars are not defined.

      We have added in the revised version ‘When using the Mann-Whitney U test, we still observed a statistically significant difference in CFU counts at 6 hpi in Figure 2B, as well as at both 6 hpi and 24 hpi in Figure 2F. However, the magnitude of these differences is relatively small’. Error bars have been defined in the general revision of the supplements (see above).

      Figure 2A shows only a limited reduction in Duox expression at the transcriptional level. "successfully" might be a bit of an overstatement.

      We have removed the term ‘successfully’.

      Figure 2E is too small.

      We have increased the size of Figure 2E.

      The authors might want to discuss the role of Duox in enterocytes destined to die as described by Amchelavsky et al., Cell Reports, 2020.

      Good point, we have discussed this reference in the revised version.

      The Sajjadian reference is incomplete.

      The reference has been completed

      (20) S24: NOS

      Maybe the work from the Royet laboratory could be cited along with that of Neyen et al. and Zaidmann-Remy et al. in the Introduction.

      Good points. We have added two references related to the work of the Royet lab (Bosco-Drayon et al., 2012; Charroux et al., 2017).

      It is not clear from Rabinovich et al. (2016) which line should be the relevant wild-type control. Why use w[1118] in Figure 1 and Ore-R in Figure 2? A more definitive conclusion could be reached upon isogenization or testing a second null mutant in homozygous and transheterozygous conditions?

      We agree that ideally, we should have used isogenic wild-type control but in absence of any effect of NOS mutants, we consider that the use of w[1118] and Oregon wild-type control is acceptable, considering that they behave similarly in the assay we have used.

      (21) S25 eiger survival

      Survival experiments: Are pooled data shown? Have the authors monitored the health status of their fly lines? In some of the survivals with late phenotypes, e.g., egr mutants and E. faecalis, this may be an issue, inasmuch as noninfected or mock-infected controls are not shown (have they been performed?).

      The survival analysis showed the pooled of 40 flies and were often done with two sexes. They also include many controls. We cannot never exclude minor effects at late time points, but the critical message is that we cannot reproduced the previously published results.

      The survival graphs are difficult to analyze: the authors should definitely avoid light colors such as yellow that do not show up well in a white background.

      As stated above, these experiments were done by many scientists in different lab and it was difficult to homogenize the data. Therefore, we did not modify our graphs although we agree with the reviewers. We believe that line patterns and contrast differences between groups guarantee that the readers can still assess the data.

      (22) S27: eiger humoral response

      Introduction: Have the authors also checked the Flysickseq data from the Buchon lab?

      There is no induction of eiger in Flysick. This information along with the reference Troha et al., have been added to S27.

      Methods: Have the RNAi lines been validated, and how? Is one more efficient at silencing than the other? This set of data would complement Figure 1 with RTqPCR experiments.

      The two RNAi lines we are using have already been used in many studies assessing the function of eiger ((#108814, reference here, #45253 reference here)

      Figure 3 colors in survival graphs: see S25.

      We did not modify our graphs although we agree with the reviewers. By increasing the magnification, the readers can still assess the data. As stated above, these experiments were done by many scientists in different labs and it was difficult to homogenize the data.

      Would it have been worth to also test the transheterozygous Gr28b mutants to avoid any issue of second-site mutations?

      We could have extended the genetics approach but, given that both us and another group (Sang) found no effect on feeding in an independent mutant, and extremely high feeding in the claimed 'anorexic' mutant, this would not be relevant to the claim. Future study may tell us if Gr28b has clearly a feeding phenotype.

      Reviewer #2 (Recommendations for the authors):

      (1) Adjust the discussion to better contextualize subjectivity: Because the choice of claims and papers inevitably involves subjective decisions, consider expanding the discussion to acknowledge this limitation more explicitly and explain how the companion metascience manuscript complements this work.

      We have added a section at the end of the discussion to discuss the limitation of our study and we establish more link to the companion article.

      (2) Refine language and avoid binary judgments: Terms such as "correct," "incorrect," or similar definitive formulations should be softened. Emphasizing uncertainty and openness to future revision will better reflect the iterative nature of scientific inquiry.

      We are softening our language in several of our assessment to avoid binary judgement.

      (3) Highlight reproducible findings as well: While the focus on irreproducible or challenged claims is understandable, it may be helpful to balance this by more clearly acknowledging wellsupported claims or areas of strong reproducibility to reinforce the positive impact of robust findings.

      We have added two sentences at the end of the introduction and at the beginning of the discussion that most claims are reproducible and that we have focused here on nonreproducible claim.

      (4) Clarify how the community can engage with the dataset: Because you provide extensive supplementary material and a large online resource, consider adding more explicit guidance on how researchers can contribute comments, evidence, or corrections, and how these contributions will be integrated.

      This point is already addressed in the revised sentence: “We hope that the community-accessible website will encourage researchers to share their perspectives and contribute data from diverse sources, thereby improving objectivity.” In addition, we have proactively engaged with principal investigators to motivate them to contribute to this effort by sharing unpublished results and information. However, it remains unclear to what extent scientists are genuinely interested in reproducibility, as has been noted in other studies on the topic. For example, articles that challenge previous findings are often under-cited, and the presence of contradictory evidence does not necessarily prevent researchers from continuing to cite the original claims. This is clearly a challenge of reproducibility study. Comments, evidence and corrections will be integrated in the ReproSci website. We have already updated this database when revising this article and will continue to do it.

      (5) Resolve the date-range inconsistency: The Introduction mentions analysis of papers from 1959-2011, whereas the Methods section cites 1940-2011. This discrepancy should be corrected for consistency.

      We have solved this discrepancy by correcting the method section. This should be 1959-2011. Thank you for spotting this mistake.

      Reviewer #4 (Recommendations for the authors):

      (1) The "disclaimer" subsection should be incorporated into the Introduction, as it is central to the premise of the study.

      The disclaimer has been moved at the end of the introduction and one sentence has been added in the abstract.

      (2) The inconsistency of graphical representation and statistical analysis of survival and gene expression data across different supplements is distracting and makes the project feel cobbled together. A consistent and unified framework for analysis and presentation would feel more cohesive and rigorous.

      The experiments were conducted across multiple laboratories and required substantial coordination. As a result, modifying the graphs at this stage would be technically complex and would require monumental effort for a tiny return. Therefore, we regret that we were unable to implement this change. The most important is that data and analyses are presented transparently, allowing readers to evaluate and interpret the results independently.

      (3) The purported roles of eiger and Gr28b in immunity have nothing to do with each other, aside from both having been published by David Schneider's group. I would suggest not combining them into a single section unless there is to be a 'miscellaneous' section that contains work by multiple groups.

      We have done two separate sections.

      (4) It would be easier to comment on specific passages in the manuscript if line numbers had been included.

      Sorry for this. We have included number of lines in the revised version.

      (5) The use of the word "elegant" to describe melanization reactions on page 3 is a bit odd and feels out of place.

      It has been removed

      (6) The sentence "We conclude that PGRP-LE is an exclusively intracellular sensor, with a prominent role regulating the gut immune response (Bosco-Drayon et al., 2012; Neyen et al., 2012)" should be edited to read "as reported in (citations)".

      We have edited the sentence.

      (7) The sentence "Experiments performed for ReproSci and review of the literature suggest that JNK is unlikely to have a strong direct or systematic role in the regulation of Drosophila AMPs in vivo, and is not required in the adult fat body for AMP expression in response to infection (Supplementary S12)" lacks detail. I realize that the experiments are described in S12, but I would suggest a brief summary in the main text of the article as well.

      We have provided more information in the article.

      (8) The phrasing of "While many of the challenged claims discussed in this paper are known to be controversial by experts, or have fallen quietly out of current research interests, some are still a source of confusion. This is particularly true for young scientists new to the field of Drosophila immunity or scientists with primary interests outside of Drosophila immunity that use the field as a reference." could be edited for clarity. A more direct phrasing could help, such as "Even in the absence of published refutation, many of the challenged claims were already considered controversial by experts in the field. Others have quietly fallen out of the general current research interest. These may nevertheless still be a source of confusion for young scientists new to the field of Drosophila immunity or scientists with primary interests outside of Drosophila immunity that use the field as a reference."

      We have changed the sentence accordingly.

      (9) On Discussion page 15, the comma should be removed from "In this article we have not discussed a number of claims, that had no clear follow-up studies and remain unchallenged."

      The coma has been removed.

      (10) In Methods, the authors "acknowledge that our verification experiments could also be erroneous". It would be appropriate to also include such a disclaimer in the Introduction and/or Discussion.

      The disclaimer has been moved at the end of the introduction and one sentence has been added in the abstract.

    1. eLife Assessment

      This important study presents an impressive large-scale effort to assess the reproducibility of published findings in the field of Drosophila immunity. The authors analyse 400 papers published between 1959 and 2011 and assess how many of the claims in these papers have been tested in subsequent publications. In a companion article they report the results of experiments to test a subset of the claims that, according to the literature, have not been tested. The evidence presented is convincing, and the limitations inherent to evaluating reproducibility primarily on the basis of the published literature - as opposed to repeating experiments - are acknowledged. The authors also explore if various factors related to authors, institutions and journals influence reproducibility in this field: in general, most correlations between these factors and reproducibility are weak, and some are dependent on analysis choices.

    2. Reviewer #1 (Public review):

      Summary:

      The authors set out on the ambitious task of establishing the reproducibility of claims from the Drosophila immunity literature. Starting out from a corpus of 400 articles from 1959 and 2011, the authors sought to determine whether their claims were confirmed or contradicted by previous or subsequent publications. Additionally, they actively sought to replicate a subset of the claims for which no previous replications were available (although this set was not necessarily representative of the whole sample, as the authors focused on suspicious and/or easily testable claims). The focus of the article is on inferential reproducibility; thus, methods don't necessarily map exactly to the original ones.

      The article is a large-scale analysis of the individual replication findings, which are presented in a companion article (Westlake et al. doi.org/10.1101/2025.07.07.663442). In their retrospective analysis, the authors find that 61% of the original claims were verified by the literature, 7.5% were partially verified, and only 6.8% was challenged, with 23.8% having no replication available. This is in stark contrast with the result of their prospective replications, in which only 16% of claims were successfully reproduced.

      The authors proceed to investigate correlates of replicability, with the most consistent finding being that findings stemming from higher-ranked universities were more likely to be challenged - a finding that is observed in a multivariable model and multiple sensitivity analyses. Other trends observed in the initial descriptive analysis, such as less replicability in findings from "trophy" journals and duration of engagement with the field, seem less consistent and fail to reach significance in the multivariable model, and should thus be considered tentative.

      Although this is a major contribution to the field, it is important to note that the high replication rates are mostly based on a retrospective survey of the literature, while prospective replication rates in the subset of articles directly replicated by the authors were much lower. Thus, the possibility of biases relating both to the field's decisions about what to replicate (and what to publish) and to the authors' criteria in evaluating claims should be considered when comparing these results with prospective replicability estimates obtained in other fields of science.

      Strengths:

      The work presents a large-scale, in-depth analysis of a particular field of science that includes authors with deep domain expertise of the field. This is a rare endeavour to establish the reproducibility of a particular subfield of science, and I'd argue that we need many more of these in different areas.

      The project was built on a collaborative basis (https://ReproSci.epfl.ch/), using an online database (https://ReproSci.epfl.ch/), which was used to organize the annotations and comments of the community about the claims. The website remains online and can be a valuable resource to the Drosophila immunity community. A companion article details the individual replication findings (Westlake et al. doi.org/10.1101/2025.07.07.663442).

      The revised version of the preprint includes a number of sensitivity analyses as supplementary material, which help to evaluate how the main findings change according to analytical decisions.

      Main concerns:

      (1) Although a number sensitivity analyses have now been conducted in order to address reviewer comments, these have been relegated to the supplementary material. Their existence is acknowledged in the multivariable model section, but the actual results of the main analysis are not mentioned at all in the main text and not taken into account in the discussion. Thus, the casual reader is not made aware of how particular results are dependent on analysis choices.<br /> Although the main finding of the correlation between replicability and university ranking holds over nearly all the analyses conducted, others are not as consistent. The association of irreplicability with trophy journals, for instance, is not statistically significant in the multivariate model (suggesting that it may have been partially due to confounding of journal tier with institutional ranking). It becomes even weaker when unchallenged claims are removed from the analysis (see point 2 below). Similarly, the association of irreplicability with "exploratory" career style is already weak in the univariate analysis and becomes even more marginal in the multivariable model. The effect of the lead author having been a first author in the field, on the other hand, is stronger than that of journal impact factor or that of the continuity/exploratory style (even though it fails to reach formal significance), but seems to receive less attention in the discussion.<br /> These differences in strength of evidence could be made more clear both in the abstract and in the discussion, which seem to prioritize some trends over others when discussing results. Once more, presenting a brief description of how results are affected by the sensitivity analyses in the results would be important to make these nuances clearer to the reader.

      (2) I still do not understand the authors' choice of including unchallenged claims along with verified ones in the multivariate model used to investigate predictors of replicability. In particular, the justification given for this choice in the supplementary material (i.e. the fact that a finding remaining unchallenged is not a random event and might be influenced by variables such as journal tier and institutional ranking) actually seems like an additional reason for not including unchallenged claims along with verified ones: if findings from low-tier journals or less prestigious institutions are more likely to remain unchallenged, this could lead to a spurious claim of "less irreplicability" in these journals and institutions, which could be purely due to receiving less attention. Predictably, both associations are reduced when unchallenged claims are removed (although the reduction is much more marked in the journal case) in the sensitivity analysis (whose results, again, are only mentioned in the supplementary material).<br /> Moreover, multiple other statements by the authors themselves concerning unchallenged claims seem to contradict their own decision. In multiple points in the text, they reiterate that remaining unchallenged is associated with less replicability (as suggested by the fact that replicability rates were lower in prospective replications than in the literature, even for non-suspicious claims). The decision to include unchallenged claims along with verified ones thus seems unjustified - if anything, I'd have expected the opposite approach after reading the discussion. That said, I'd argue that the best decision would be to exclude them from the analysis, as done in Figure S4, and presenting this model as the main one.

    3. Reviewer #3 (Public review):

      Summary:

      The authors of this paper were trying to identify how reproducible, or not, their subfield (Drosophila immunity) was since its inception over 50 years ago. This required identifying not only the papers, but the specific claims made in the paper, assessing if these claims were followed up in the literature, and if so whether they supported or refuted the original claim. In addition to this large manually curated effort, the authors further investigated some claims that were left unchallenged in the literature by conducting replications themselves. This provided a rich corpus of the subfield that could be investigated into what characteristics influence reproducibility.

      A major strength of this study is the focus on a subfield, the detailing of identifying the main, major, and minor claims - which is a very challenging manual task - and then cataloging not only their assessment of if these claims were followed up in the literature, but also what characteristics might be contributing to reproducibility, which also included more manual effort to supplement the data that they were able to extract from the published papers. While this provides a rich dataset for analysis, there is a major weakness with this approach, which is not unique to this study.

      The main weakness is relying heavily on the published literature as the source of whether a claim was determined to be verified or not. Nonetheless it is understandable why the authors took this approach - it is the only way to get at a breadth of the literature. However, there are many documented issues with this stemming from every field of research - such as publication bias, selective reporting, all the way to fraud. Importantly, this limitation is highlighted throughout the paper. At the same time, it is not reasonable to expect this study to have conducted independent experimental replications for all 400 papers identified as this would have been a laborious and costly effort. To overcome this weakness, the authors leveraged their own expertise, and crowdsourcing within the Drosophila immunity community, to assess the validity of the claims. While this helps mitigate some of this weakness, it does introduce new challenges in interpretation, which is acknowledged in the limitations section of the discussion. Overall, this situates this study more as an expert assessment of the replicability of the published literature opposed to an assessment of independent experimental replications.

      The authors should be applauded for the monumental effort they put into this project, which does a wonderful job of having experts within a subfield engage their community to understand the connectiveness of the literature and attempt to understand how reliable specific results are and what factors might contribute to them. This project provides a nice blueprint for others to build from as well as leverage the data generated from this subfield and thus should have an impact in the broader discussion on reproducibility and reliability of research evidence.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study presents an impressive large-scale effort to assess the reproducibility of published findings in the field of Drosophila immunity. The authors analyse 400 papers published between 1959 and 2011, and assess how many of the claims in these papers have been tested in subsequent publications. In a companion article they report the results of experiments to test a subset of the claims that, according to the literature, have not been tested. The present article also explores if various factors related to authors, institutions and journals influence reproducibility in this field. The evidence supporting the claims is solid, but there is considerable scope for strengthening and extending the analysis. The limitations inherent to evaluating reproducibility based on the published literature should also be acknowledged.

      We have included in the abstract the limitation of evaluating replicability based on the published literature

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors set out on the ambitious task of establishing the reproducibility of claims from the Drosophila immunity literature. Starting out from a corpus of 400 articles from 1959 and 2011, the authors sought to determine whether their claims were confirmed or contradicted by previous or subsequent publications. Additionally, they actively sought to replicate a subset of the claims for which no previous replications were available (although this set was not representative of the whole sample, as the authors focused on suspicious and/or easily testable claims). The focus of the article is on inferential reproducibility; thus, methods don't necessarily map exactly to the original ones.

      The authors present a large-scale analysis of the individual replication findings, which are presented in a companion article (Westlake et al., 2025. DOI 10.1101/2025.07.07.663442). In their retrospective analysis of reproducibility, the authors find that 61% of the original claims were verified by the literature, 7.5% were partialy verified, and only 6.8% were challenged, with 23.8% having no replication available. This is in stark contrast with the result of their prospective replications, in which only 16% of claims were successfully reproduced. The authors proceed to investigate correlates of replicability, with the most consistent finding being that findings stemming from higher-ranked universities (and possibly from very high impact journals) were more likely to be challenged.

      Strengths:

      (1) The work presents a large-scale, in-depth analysis of a particular field of science that includes authors with deep domain expertise of the field. This is a rare endeavour to establish the reproducibility of a particular subfield of science, and I'd argue that we need many more of these in different areas.

      (2) The project was built on a collaborative basis (https://ReproSci.epfl.ch/), using an online database (https://ReproSci.epfl.ch/), which was used to organize the annotations and comments of the community about the claims. The website remains online and can be a valuable resource to the Drosophila immunity community.

      (3) Data and code are shared in the authors' GitHub repository, with a Jupyter notebook available to reproduce the results.

      We thank the reviewer for their positive comments.

      Main concerns:

      (1) Although the authors claim that "Drosophila immunity claims are mostly replicable", this conclusion is strictly based on the retrospective analysis - in which around 84% of the claims for which a published verification attempt was found. This is in very stark contrast with the findings that the authors replicate prospectively, of which only 16% are verified. Although this large discrepancy may be explained by the fact that the authors focused on unchallenged and suspicious claims (which seems to be their preferred explanation), an alternative hypothesis is that there is a large amount of confirmation bias in the Drosophila immunity literature, either because attempts to replicate previous findings tend to reach similar results due to researcher bias, or because results that validate previous findings are more likely to be published.

      Both explanations are plausible (and, not being an expert in the field, I'd have a hard time estimating their relative probability), and in the absence of prospective replication of a systematic sample of claims - which could determine whether the replication rate for a random sample of claims is as high as that observed in the literature -, both should be considered in the manuscript.

      I still believe that most of ‘verified’ claims are indeed solid, consistent with the feeling of increased knowledge in the Drosophila immunity. Of note, when I started in the field in the early 90, nearly all the articles in Drosophila developmental biology, the dominant field at that time, were solid. I agree that this may represent a past period. Still, perceived irreproducibility is high because affecting claims published in high-impact articles. To address the reviewer’s comments, we have included the possibility of confirmation bias in the discussion of the article.

      (2) The fact that the analysis of factors correlating with reproducibility includes both prospective and retrospective replications also leads to the possibility of confusion bias in this analysis. If most of the challenged claims come from the authors' prospective replications, while most of the verified ones come from those that were replicated by the literature, it becomes unclear whether the identified factors are correlated with actual reproducibility of the claims or with the likelihood that a given claim will be tested by other authors and that this replication will be published.

      We agree that the ReproSci is a human biased endeavor and may be subjected to bias, notably the one mentioned. Some verified claims might be considered as challenged and some challenged claims as verified. But this is likely not affecting most of our results. Our papers have been posted on a public archive and that the community could scrutinize our data (most authors have checked their claims!). Based on community feedback, we have had to change the status of three claims only: one verified claim became challenged, one challenged claim became unchallenged and one unchallenged became verified. We therefore did not expect major changes in our conclusions. We have provided an estimation of irreproducibility ranging from 10-20% and we believe this is a rather safe interval. This is still significant. The fact that the ReproSci project is a human endeavor, and therefore subject to inherent limitations, is already emphasized in the last section of the discussion.

      (3) The methods are very brief for a project of this size, and many of the aspects in determining whether claims were conceptually replicated and how replications were set up are missing.

      Some of these - such as the PubMed search string for the publications and a better description of the annotation process - are described in the companion article, but this could be more explicitly stated. Others, however, remain obscure. Statements such as "Claims were cross-checked with evidence from previous, contemporary and subsequent publications and assigned a verification category" summarize a very complex process for which more detail should be given - in particular because what constitutes inferential reproducibility is not a self-evident concept. And although I appreciate that what constitutes a replication is ultimately a case-by-case decision, a general description of the guidelines used by the authors to determine this should be provided. As these processes were done by one author and reviewed by another, it would also be useful to know the agreement rates between them to have a general sense of how reproducible the annotation process might be.

      The same gap in methods descriptions holds for the prospective replications. How were labs selected, how were experimental protocols developed, and how was the validity of the experiments as a conceptual replication assessed? I understand that providing the methods for each individual replication is beyond the scope of the article, but a general description of how they were developed would be important.

      We have extended the first section of the Result to mention that an extended Material and Methods can be found in the complementary article Westlake et al., eLife 2026 and on the ReproSci website indicated the resources described in the companion articles. Disagreement between the two authors were limited and mostly dealt with textual interpretation (the definition of the major claims, which can affect if they are verified or not). As mentioned above, our articles are now public since more 6 months and we got little request to change the status of a claim (only 3 over more than 1000). Of note many claims from some ancient articles are part of the common knowledge, are verified routinely by many labs in the field, and therefore were easy to categorize.

      We also added a sentence to also underline that the scientific value of an article does not fully correlate with its status. A number of problematic articles filled with exaggeration or confused claims were still categorized as verified. Thus, we screen for replicability but they are many other publication issues we could not capture in our study. This point is now mentioned in the discussion.

      (4) As far as I could tell, the large-scale analysis of the replication results was not preregistered, and many decisions seem somewhat ad hoc. In particular, the categorization of journals (e.g. low impact, high impact, "trophy") and universities (e.g. top 50, 51-100, 101+) relies on arbitrary thresholds, and it is unclear how much the results are dependent on these decisions, as no sensitivity analyses are provided.

      Particularly, for analyses that correlate reproducibility with continuous variable (such as year of publication, impact factor or university ranking, I'd strongly favor using these variables as continuous variables in the analysis (e.g. using logistic regression) rather than performing pairwise comparisons between categories determined by arbitrary cutoffs. This would not only reduce the impact of arbitrary thresholds in the analysis, but would also increase statistical power in the univariate analyses (as the whole sample can be used in at once) and reduce the number of parameters in the multivariate model (as they will be included as a single variable rather than multiple dummy variables when there are more than two categories).

      We did not pre-registered our study, which is quite different from other reproducibility projects. We did not know at the start how the project would unfold. Of note, I believe that ‘cleaning’ an extensive field by experts was by itself an achievement and represent another way to assess replicability, with its own limitation.

      We agree that the cut-off can be seen as arbitrary. However, publication year has been modeled as a continuous variable through a spline (to allow for non-linearity) in the multivariate analysis. For journal and institutional ranking, we would like to defend and keep the categorical specifications in the main text for two reasons:

      - Institution ranking: For institutional ranking, the largest institutional category is “Not Ranked” which we cannot assign a rank and ranks institutions individually only through the top 100. Below that it publishes bands (101-150, 151-200, 201-300, 301400, 401-500) of variable width, which are not straightforward to use continuously without making strong assumptions (e.g midpoint) unsupported by the data.

      - Prestige: As mentioned by another reviewer there are career consequences attached to Nature, Science and Cell that are different from other journal, and the effect of moving from a journal from an IF 3 to 8 is different than from 30 to 35. We provide a sensitivity analysis in the supplementary information where we refit the multivariate model with log-transformed impact factor continuous predictor. We obtain very similar results than these in the main text, with the exception that the impact factor of the journal becomes significant. We describe and discuss this results in the supplementary information (section S5), while justifying why we keep the categorical model in the main text.

      We are grateful to the reviewer for prompting these additions, which we agree strengthen the paper.

      The text added in the SI states:

      “As a sensitivity analysis, we refitted the multivariable hierarchical logistic model by replacing journal-impact-factor categories with log₂-transformed continuous impact factor, while retaining all other covariates and random intercepts from the primary analysis. University ranking was retained as a categorical variable because the 2010 Shanghai Academic Ranking reports individual ranks only for institutions ranked 1– 100 and otherwise reports rank bands (101–150, 151–200, 201–300, 301–400, and 401–500); moreover, 441 of 1,006 claims (43.8%) were from unranked institutions, which constitute a distinct group.

      Each doubling of journal impact factor was associated with 1.51-fold higher odds that a claim was challenged (OR 1.51, 94% HDI 1.13–2.08). At the mean impact factors of the low-impact and trophy-journal tiers (4.4 and 61, respectively; approximately 3.8 doublings), the continuous model implies a trophy-versus-low-impact odds ratio of approximately 4.8. By contrast, the primary model directly estimated the trophy-versus-low-impact odds ratio as 1.75 (94% HDI 0.59–4.91). Although this wide interval includes 4.8, the point estimates differ between the two models. We interpret the larger implied trophy-journal effect in the continuous model cautiously because it assumes that the log-linear slope estimated predominantly from the many journals with lower impact factors (around 2–12) applies unchanged to the sparsely represented trophy journals (62 claims across Science, Nature, and Cell). We therefore retained the categorical journal analysis as the primary analysis because it directly estimates the trophy-journal comparison using claims published in those journals. Estimates for all other covariates were broadly similar to those in the primary model. In particular, the association with Top 50 university affiliation remained evident (OR 4.04, 94% HDI 1.57–10.08, compared with OR 3.68, 94% HDI 1.50–9.28 in the primary model)”.

      [This response is referred later as “Response to categorical variable encoding”]

      (5) The multivariate model used to investigate predictors of replicability includes unchallenged claims along with verified ones in the outcome, which seems like an odd decision. If the intention is to analyze which factors are correlated with reproducibility, it would make more sense to remove the unchallenged findings, as these are likely uninformative in this sense. In fact, based on the authors' own replications of unchallenged findings, they may be more likely to belong the "challenged" category than to the "unchallenged" one if they were to be verified.

      We thank the reviewer for raising this comment, we agree that the classification of the unchallenged claims as not challenged is difficult given their nature. We refitted the multivariable analysis excluding unchallenged claims, with a classification of Challenged and Unchallenged (Verified, Partially Verified, Mixed). The results are shown in the supplementary section S4 and are extremely similar to the main model.

      The text added in the SI states:

      “We fit the multivariate model from the main text while excluding unchallenged claims from the analysis. The methods are similar and the convergence and diagnosis checks are good and not reported here for brevity. The model simultaneously adjusts for author demographics, laboratory attributes and journal or institutional prestige and predict a binary outcome for each claim: challenged or not challenged (including Verified, Partially Verified, Mixed). The re-analysis included 659 claims, of which 60 (9.1%) were challenged. The non-challenged reference group comprised of Verified (n=525), Partially Verified (n=64), Mixed (n=10). The Adjusted odds ratios and 94% highest-density intervals are reported in Figure S4. The results are very similar to the analysis with the unchallenged claims in the main text, with only university ranking (Top 50) being a significant predictor (OR 3.34 (1.28–9.01) vs the main text OR 3.68 (94% HDI 1.50– 9.28). The other variable with the biggest changes shows only modest differences:

      - High-impact journal OR 1.12 (vs 1.31)

      - Trophy journal OR 1.47 (vs 1.75)

      - Post-doc OR 1.44 (vs 1.16) and their intervals still include 1. Note that the excluding unchallenged decision condition on follow-up, which is not random and might be influenced by our key variable (journal tier, institutional ranking) and bias results for this model.”

      However, we did not replace the main model of the article by this particular model because when you chose only claim that are either verified or challenged (so you exclude unchallenged), the condition that made these claims included in this analysis could be influenced by a cofactor, so our odds ratios might be wrong.

      [This response is referred later as “Response to unchallenged exclusion”]

      Reviewer #2 (Public review):

      Summary:

      Lemaitre et al. conducted an analysis of 400 publications in the Drosophila immunity field (1959-2011), performing both univariable and multivariable analyses to identify factors that correlate with or influence the irreproducibility of scientific claims. Some of the findings are unexpected, for instance, neither the career stage of the PI nor that of the first author appears to matter that much, while others, such as the influence of institutional prestige or publication in "trophy journals," are more predictable. The results provide valuable insight into patterns of irreproducibility in academia and may help inform policies to improve research reproducibility in the field.

      Strengths:

      This study is based on a large, manually curated dataset, complemented by a companion paper (Westlake et al., 2025. DOI 10.1101/2025.07.07.663442) that provides additional details on experimentally documented cases. The statistical methods are appropriate, and the findings are both important and informative. The results are clearly presented and supported by accessible documentation through the ReproSci project.

      We thank the reviewer for their assessment.

      Weaknesses:

      The analysis is limited to a specific field (immunity) and model system (Drosophila). Since biological context may influence reproducibility -- for example, depending on whether mechanisms are more hardwired or variable -- and the model system itself may contribute to these effects (as the authors note), it remains unclear to what extent these findings generalize to other fields or organisms. The authors could expand the discussion to address the potential scope and limitations of the study's generalizability.

      We have added a limitation section at the end of the manuscript to discuss these points. We wrote:

      “Finally, the findings we report may be specific to a particular field, model organism, and time period, and may be strongly influenced by factors idiosyncratic to their development. Because this analysis is limited to a single field (immunity) and model system (Drosophila), it remains unclear to what extent these findings generalize to other fields, organisms, or more recent studies.”

      However, we believe that our study adds on the discussion on reproducibility by providing another approach. Our study also aligns with the feeling that knowledge of Drosophila innate immunity has increased so much in the last decades, and that science is still quite robust in many fields.

      Reviewer #3 (Public review):

      Summary:

      The authors of this paper were trying to identify how reproducible, or not, their subfield (Drosophila immunity) was since its inception over 50 years ago. This required identifying not only the papers, but the specific claims made in the paper, assessing if these claims were followed up in the literature, and if so whether the subsequent papers supported or refuted the original claim. In addition to this large manually curated effort, the authors further investigated some claims that were left unchallenged in the literature by conducting replications themselves. This provided a rich corpus of the subfield that could be investigated into what characteristics influence reproducibility.

      Strengths:

      A major strength of this study is the focus on a subfield, the detailing of identifying the main, major, and minor claims - which is a very challenging manual task - and then cataloging not only their assessment of if these claims were followed up in the literature, but also what characteristics might be contributing to reproducibility, which also included more manual effort to supplement the data that they were able to extract from the published papers. While this provides a rich dataset for analysis, there is a major weakness with this approach, which is not unique to this study.

      Weaknesses:

      The main weakness is relying heavily on the published literature as the source for if a claim was determined to be verified or not. There are many documented issues with this stemming from every field of research - such as publication bias, selective reporting, all the way to fraud. It's understandable why the authors took this approach - it is the only way to get at a breadth of the literature - however the flaw with this approach is it takes the literature as a solid ground truth, which it is not. At the same time, it is not reasonable to expect the authors to have conducted independent replications for all of the 400 papers they identified.

      However, there is a big difference trying to assess the reproducibility of the literature by using the literature as the 'ground truth' vs doing this independently like other large-scale replication projects have attempted to do. This means the interpretation of the data is a bit challenging.

      We acknowledge that our analysis necessarily relies on the published literature; however, two important points should be considered. First, we assessed replicability more than a decade after the original studies, allowing us to evaluate whether their claims have withstood subsequent scrutiny. Second, the evaluation draws on the authors’ strong expertise in the field. Accordingly, the validity of published claims was not accepted uncritically but subjected to informed expert assessment. This is exemplified by the case of the proposed microbicidal role of Duox, which we classified as challenged despite the original article accumulating more than 600 citations.

      Overall, our replicability analysis, like similar efforts, relies not only on literature but also on expert judgment. Of course it has inherent limitations, which are now more clearly articulated in the discussion.

      Below are suggestions for the authors and readers to consider:

      (1) I understand why the authors prefer to mention claims as their primary means of reporting what they found, but it is nested within paper, and that makes it very hard to understand how to interpret these results at times. I also cannot understand at the high-level the relationship between claims and papers. The methods suggest there are 3-4 major claims per paper, but at 400 papers and 1,006 claims, this averages to ~2.5 claims per paper. Can the authors consider describing this relationship better (e.g., distribution of claims and papers) and/or considering presenting the data two ways (primary figures as claims and complimentary supplementary figures with papers as the unit). This will help the reader interpret the data both ways without confusion. I am also curious how the results look when presented both ways (e.g., does shifting to the paper as the unit of analysis shift the figures and interpretation?). This is especially true since the first and last author analysis shows there is varying distribution of papers and claims by authors (and thus the relationship between these is important for the reader).

      Articles can contain claims that we found replicable and other that we did not. Thus, we believe than an evaluation by claim is more precise. The extraction of the different claims relies on the reading of the article, but is a human enterprise. There are clearly some unavoidable biases at this step. Of note, the articles analyzed in this study are very different one from another (size, novelty, amount of data…).

      However, we agree that the analysis at the article level is interesting, and we followed the reviewer suggestion and provided in Supplementary Section S6 a sensitivity analysis of using article-level effects for the multivariate analysis. We note that because only 43 articles (12.5%) contained at least one challenged claim, this analysis is less robust than the claim analysis. The text added in the SI states:

      “Two claims in the same paper might not be independent, so we might want to incorporate an article random effect to the primary Bayesian formulation. However, due to the low number of claim per article (mostly one to four) and only 43 articles (12.5%) contained at least one challenged claim (and most articles contains only one to four claims), we could not manage to fit this model properly (we observe MCMC divergences). As a work-around sensitivity analysis, we repeated the analysis with the article as the unit. Each of 345 articles with complete covariate data (representing 869 major claims, 60 of the 69 challenged claims in the full dataset; the remaining 9 fell in articles excluded for missing covariates) contributed as outcome the proportion of its major claims classified as challenged. The model was a (frequentist) binomial-link fractional-response generalized linear model with HC3 standard errors (the intervals reported here are 94% confidence intervals rather than posterior credible intervals).

      Each article contributed equally regardless of its number of claims. The Top-50 institutional effect remained significant under both analysis (OR 5.15, 94% CI 2.18– 12.15; versus a claim-level OR 3.68, 94% HDI 1.50–9.28). In this new model Trophy journals also showed had higher odds of challenged claims than low-impact journals (OR 2.90, 94% CI 1.24–6.78 vs a claim-level OR 1.75, 94% HDI 0.59–4.91). This impact was not significant in the primary model but became significant at using article-level analysis. Other estimates are consistent in direction and magnitude with the primary claim-level model. Model diagnostics were satisfactory.”

      Moreover, we added Figure S6 which display visually the relationship between the claims and the article.

      [This response is referred later as “Response to claim vs article level analysis”]

      (2) As mentioned above, I think the biggest weakness is that the authors are taking the literature at face value when assigning if a claim was validated or challenged vs gathering new independent evidence. This means the paper leans more on papers, making it more like a citation analysis vs an independent effort like other large-scale replication projects. I highly recommend the authors state this in their limitations section.

      We have now explicitly acknowledged this limitation in a dedicated section at the end of the revised manuscript. However, as noted above, we do not take the literature at face value. Notably, some of the claims we identify as challenged originate from highly cited articles, which could lead a naïve reader to assume their validity.

      Ultimately, this replicabilty project remains a human, expert-driven effort aimed at determining which claims withstand scrutiny within a field. The broader community was invited to engage with and respond to these assessments and we did get some feedbacks. Importantly, claims were not evaluated in isolation but in the context of the entire body of knowledge, and all evaluations were made transparent and explicitly justified in the ReproSci database.

      On top of that, I have questions that I could not figure out (though I acknowledge I did not dig super deep into the data to try). The main comment I have is How was verified (and challenged) determined? It seems from the methods it was determined by "Claims were cross-checked with evidence from previous, contemporary and subsequent publications and assigned a verification category". If this is true, and all claims were done this way - are verified claims double counted then? (e.g., an original claim is found by a future claim to be verified - and thus that future claim is also considered to be verified because of the original claim).

      When two articles reported the same main claim and this claim was judged as “verified,” each instance was counted in the analysis. However, when validation came from studies published after 2011, the claim was counted only once. Thus, our study specifically evaluates the replicability of claims originating from a defined period (pre-2011).

      We note that, in principle, repeated publication of the same finding across many articles could artificially inflate estimates of replicability. However, such cases appeared to be rare in practice, and studies reporting similar claims typically provided complementary evidence—for example, elucidating gene function through either genetic or biochemical approaches.

      Related, did the authors look at the strength of validation or challenged claims? That is, if there is a relationship mapping the authors did for original claims and follow-up claims, I would imagine some claims have deeper (i.e., more) claims that followed up on them vs others. This might be interested to look at as well.

      We did not perform a systematic mapping between original claims and subsequent follow-up studies. While this is an important and interesting aspect, it was considered beyond the scope of the present work. The current study already represents a substantial investment of time and effort, and the resulting database provides a valuable resource for future analyses. In particular, it could support a dedicated bibliometric study at a later stage.

      (3) I recommend the authors add sample sizes when not present (e.g., Fig 4C).

      The cited panel appears to correspond to Figure 5C in the revised numbering, as the current Figure 4 contains only panels A and B. We systematically checked the denominators of all figures and revised the captions to report the numbers of claims, authors or laboratories represented, together with the reasons for exclusions. We also added the numbers of authors and claims directly within the author-level scatter plots, with the observational unit explicitly identified.

      I also find that the sample sizes are a bit confusing, and I recommend the authors check them and add more explanation when not complete, like they did for Fig 4A. For example, Fig 7B equals to 178 labs (how did more than 156 labs get determined here?), and yet the total number of claims is 996 (opposed to 1,006).

      There are 156 labs total, 23 labs contributed publications during both junior and senior stages, and 11 claims lacked a junior/senior classification. This has been explicited in the caption.

      Another example, is why does Fig 8B not have all 156 labs accounted for?

      The panel does not include all 156 laboratories because we could obtain information about first author training for only 146 PI and laboratories while prior first-author status was unavailable or indeterminate for ten PI. This has been added to the caption.

      (Related to Fig 8B, I caution on reporting a p value and drawing strong conclusions from this very small sample size - 22 authors).

      The 22 authors are not the total sample size; they constitute the subgroup of PIs who had previously published a first-author paper on Drosophila immunity in another laboratory. In the submitted panel, they were compared with 117 PIs without this training history, giving 139 classified laboratories and 774 claims.

      As a last example, Fig 8C has al 156 labs and 1,006 claims - is that expected? I guess it means authors who published before 1995 (as shown in Figure 8A continued to publish after 1995?) in that case, it's all authors? But the text says when they 'set up their lab' after 1995, but how can that be?

      Thank you for identifying this inconsistency. Figure 8C intentionally includes the full cohort of 156 PIs and all 1,006 claims. Panel C compares research styles and does not impose a restriction based on year of entry into the field or year of claim publication. Thus, PIs who entered the field before 1995 are included together with all claims attributed to them; the panel is not restricted to claims published after 1995.

      The statement that only PIs who established their laboratories after 1995 were included was an erroneous carryover from an earlier version of the analysis. We removed this restriction from the Results and revised the caption to state that exploratory PIs (n = 91) contributed 366 claims, whereas continuity PIs (n = 65) contributed 640 claims.

      We also replaced references to the date when PIs “established their laboratory” with “year of entry into the field,” defined as the year of their first Drosophila-immunity publication as either first or last author. During the same check, we revised Panel B to include all 146 PIs eligible for the pre-PI training comparison, contributing 793 claims.

      (4) Finally, I think it would help if the authors expanded on the limitations generally and potential alternative explanations and/or driving factors. For example, the line "though likely underestimated' is indicated in the discussion about the low rate of challenged claims, it might be useful to call out how publication bias is likely the driver here and thus it needs to be carefully considered in the interpretation of this. Related, I caution the authors on overinterpreting their suggestive evidence. The abstract for example, states claims of what was found in their analysis, when these are suggestive at best, which the authors acknowledge in the paper. But since most people start with the abstract, I worry this is indicating stronger evidence than what the authors actually have.

      We have added a dedicated section on limitations at the end of the Discussion to explicitly acknowledge key constraints, and we also address these points throughout the Discussion where relevant. In addition, these limitations are now in the abstract.

      The authors should be applauded for the monumental effort they put into this project, which does a wonderful job of having experts within a subfield engage their community to understand the connectiveness of the literature and attempt to understand how reliable specific results are and what factors might contribute to them. This project provides a nice blueprint for others to build from as well as leverage the data generated from this subfield, and thus should have an impact in the broader discussion on reproducibility and reliability of research evidence.

      Thank you for this positive assessment!

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Comments and recommendations are provided in the order that they appear in the manuscript.

      Introduction:

      - In the second paragraph of the introduction, I don't think all the references refer to "molecular life sciences" (e.g. one of them is from ecology and evolution, for example). I'd also argue that Macleod et al. 2014 refers to risk of bias rather than to reproducibility per se.

      We have corrected this in the revised version.

      - The author use both "reproducibility" and "replicability", apparently as synonyms, but as definitions for these words vary across sources (e.g. https://arxiv.org/abs/1802.03311, https://www.nationalacademies.org/our-work/reproducibility-and-replicability-in-science, https://osf.io/br9sp/) it may be worth stating their definitions upfront.

      Thank you for the references. We agree that the literature contains multiple, sometimes conflicting definitions of these terms, which can lead to confusion. In response to the reviewer’s comment, we have clarified our terminology. Specifically, we now use the term “conceptual replicability” in the abstract and consistently refer to “replicability” and “irreplicability” throughout the manuscript. We hope this revised terminology improves clarity.

      - Why are analysis described as "exploratory" and "multivariate"? If it has not been preregistered, shouldn't the multivariate analysis be considered exploratory as well?

      Thanks, we agree. The word “exploratory” is confusing as we use it for its structural meaning and not its epistemological meaning. We have replaced it throughout by “descriptive”, as to not imply that the multivariate analysis is preregistered by contrast.

      Results:

      Methodology and claim assessment

      - The methods for the selection, annotation and classification of studies are described very cursorily here, and the Methods section doesn't add that much, but I'll keep these observations to the methods. Still, it's worth mentioning that it's hard to know how reproducible the annotation and classification process is from the information provided.

      We have provided more information on the method section on the way to annotate and class claims. Of note, there is more information on the methods on the associated publication Westlake et al., 2025 and on the ReproSci website.

      - "Then, experimental work was performed in several laboratories...". Once again, a much better description of how replications were set up and performed is needed.

      We have provided information on the number of laboratories (nine) involved in the reproduction of 45 major claims. The laboratories involved in replication are listed in the author list. All the files associated with experimental replication can be found in the supplementary materials of Westlake et al 2025 or the ReproSci website.

      - "These claims subjected to experimental validation were selected according to criteria such as appearing suspicious due to the absence of direct follow up or being straightforward to test experimentally". This is rather vague, but if no explicit criteria were set upfront this may be inevitable. Still, it would be useful to know who made this decision.

      A first list of claims to be replicated was made by Hannah Westlake and Bruno Lemaitre. The claims that were selected to be replicated were tested by the labs and experimental workers that had the expertise required to test them. The choice of the claims to be tested was not random. Claims that were suspicious are claims that were already tested in the host laboratory without success, claims whose validity are debated at meeting, claims from highranked journal with no follow-up. This is now indicated in the method of the revised version.

      Drosophila immunity claims are mostly reproducible

      - The percentages given in the initial description of the results pool the results of the retrospective analysis of the literature and the prospective replications performed. As mentioned in the Public Review, I think these are rather distinct ways to assess reproducibility, and would recommend that these two analyses are described separately.

      We agree with the reviewer but we believe that we have already well separated the results before and after experimental work in the revised version. This is clearly shown in Table 1. Our text says ‘Importantly, 6.8% of claims (69 out of 1006) were challenged. Among these, 44.9% (31 out of 69) were contested by published articles including 7.2% by the same authors, and 55.1% (38 out of 69) were challenged experimentally as part of the ReproSci project.’

      - "These results may reflect the robust scientific standards and methodological rigor of Drosophila research...". Once again, as mentioned in the Public Review, I think publication bias is a plausible explanation, so this should be mentioned as an alternative.

      We have mentioned publication bias in the revised version in the section on limitation. But an important point to underline is that this analysis did not take publication results as face fact. This is a critical analysis by experts in the field. As already mentioned, some claims like ‘Duox produces microbicidal ROS” ‘NO is a signaling molecule…’that have been mentioned in multiple articles were considered as challenged.

      A second point is that it is likely that publications in some fields are more replicable than in other. Actually, we did not discuss ‘reproducibility issues’ when I started in the field 30 year ago because this was not perceived as an issue. It is normal that reproducibility rate varies according to fields. Nevertheless, we do find important claims with high visibility that are not replicable in our sample. Thus, our findings nuance the ‘reproducibility crisis narrative’ but we still believe that there is concern with reproducibility in life science. Last point, the criterium we have used in our study did not map all the problematic papers, notably some claims might be fundamentally verified but exaggerated in the original paper.

      - It would be useful to break down the "partially verified" category somewhere. Illustrative examples of what is meant by "insightful data were accompanied by incomplete interpretations" or "incomplete data were paired with insightful interpretations" would be useful as well.

      All the claims are listed in the public database with an explanation as to why they got affected to a category. The categories mixed and partially verified refer to claims that are more complex to categorize.

      Here are are two examples from the ReproSci website:

      Claim: Relish Rel mutants are very susceptible to bacterial and fungal infection. Partially verified

      Assessment: Relish mutants are primarily susceptible to Gram-negative bacterial infections, and do not typically have increased susceptibility to Gram-positive bacteria and fungi which are primarily rebuffed by the Toll pathway, although Imd signaling may provide a minor contribution to defense against these. The high susceptibility of Relish mutant to fungi observed in this article is likely due to the presence of the ebony marker (Lemaitre et al., 1996) [here only one part of the statement has been verified]

      Claim: PGRP-SC1a is required for Lys-type peptidoglycan recognition and Toll activation. Mixed

      Assessment: A possible role for PGRP-SCs in Toll activity has been strongly debated. Challenged by (Bischoff et al., 2006) who found no effect of PGRP-SC RNAi on activation of Toll (Drs expression) in adult flies (although note that their data show a minor reduction of Drs expression in flies in response to E. faecalis). Supported by (Costechareyre et al., 2016) who used single -SC1 and -SC2 mutants to show that Toll activation (Drs expression) was reduced in -SC1 but particularly -SC2 mutants (~50%) in response to E. faecalis (although survival was not affected). The supplementary data of (Paredes et al., 2011) similarly show that Drs expression was reduced (50%) in response to M. luteus septic infection in PGRP-SCdelta flies (PGRP-SC1A/B -SC2 triple mutant), but concluded that this was due to a secondary effect of Imd overactivation. Note that the phenotype found by (Garver et al., 2006) was much stronger (complete ablation of Drs expression), which is not consistent with any subsequent results and argues for a secondary mutation in this line, although this is not consistent with the successful rescue of Drs expression by transgenic replacement of PGRP-SC1a. Downregulation of Toll in these mutants could be explained if the Toll pathway requires amidase activity of PGRP-SC for immunogenicity (e.g. the sugar backbone without stem peptides is a stronger elicitor than whole peptidoglycan), whereas similar cleavage reduces recognition by the Imd pathway (consistent with the demonstrated requirement of stem peptides for full stimulation of the Imd pathway (Chang et al., 2005; Stenbak et al., 2004)). See annotation for (Mellroth et al., 2003). [here the claim is mostly challenged but there are observations that go in the same directions]

      As shown by these examples, category assessment was complex process and relied on human expertise. However, there was very little contestation from the community after the submission of our articles and the opening of the ReproSci website.

      A significant fraction of unchallenged claims is non-reproducible

      - Excluding the 45 unchallenged major claims that were experimentally tested, we categorized the remaining 240 unchallenged claims into three groups". No information is provided on that classification process (either here or in the methods). Who made these assessments (which seem quite subjective and dependent on field expertise), and how do you know if such judgments are reproducible across different evaluators?

      This categorization was done by Hannah Westlake and reviewed by Bruno Lemaitre with few disagreements. This is now indicated in the revised version. The category ‘unchallenged logically consistent” and ‘‘unchallenged logically consistent’ refer to claims that although not directly verified are corroborated or in with current literature. See ReproSci for justifications.

      Example from ReproSci website:

      Claim: The minimal structure needed to activate the Toll pathway is a muropeptide dimer.

      Assesment: Unchallenged logically consistent This is consistent with (Park et al., 2007) who show that on a linearized strand of peptidoglycan, the minimal motif is at least 3 dimers. But it is expected that a non-linearized dimer linked by the peptide bridge could serve as the minimal motif by clustering peptidoglycan.

      Claim: eater null flies are susceptible to oral infection with Serratia marcescens Unchallenged logically consistent Inconsistent with the observation that Eater mutants successfully phagocytose Serratia marcescens (Bretscher et al., 2015). This may indicate that another factor affected by eater mutation is required for defense against S. marcescens, such as formation of lamellopodia and filopodia or adhesion of hemocytes to the body wall, or that Eater contributes to binding of Gram-negative bacteria but does not trigger phagocytosis in response to them as it does for Gram-positive bacteria.

      Higher representation of challenged claims in trophy journals and from top universities

      - Both impact factor and the Shanghai university ranking are continuous variables, but the authors opt to use them as categorical variables in the analysis (i.e. "low-impact", "highimpact", "trophy"; "Top 50", "51-100", "101+"). While there may be legitimate reasons to do this if they feel that these categories are a better descriptor of the underlying reality (e.g. perhaps Cell, Science and Nature are indeed in a category by themselves), it is an unusual decision that leads the comparison to ignore the distinctions between journals/universities within a category. Moreover, it also opens up the opportunity to analysis bias as categories can be set up in many different ways. Thus, if the analysis was not preregistered, I'd recommend that analyses using impact factor and ranking as quantitative variables are added as sensitivity analyses, as these seem to me to be the most natural/less ad hoc way to look at the issue.

      See “Response to categorical variable encoding” above.

      - As mentioned in the Public Review, the analysis here pools the results from the retrospective analysis based on the literature and the prospective one based on the performed replications. Although this may be justified from a sample size perspective (as the number of challenged claims is not that high), it would be useful to perform this analysis separately on the two sets of results as well, as different trends may be noticed. There are many biases that might come into play here (e.g. results from top institutions being more or less likely to have challenged/unchallenged claims published) and looking at the results separately could help in tearing apart these hypotheses.

      The ReproSci project analyzes all the papers of a community during a period of time. As such, the total number of claims is limited and we believe that pooling the data was justified. We provide as sensitivity analysis in Supplement S7, a multivariate model with claims in their pre-experimental classification state:

      “As a sensitivity analysis, we refitted the multivariable hierarchical logistic model using only the claim classifications available before experimental validation by the ReproSci project, as the choice of which claims to further validate could have been influenced by covariate (impact factor, university status). The 45 claims tested prospectively were therefore restored to their original “Unchallenged” classification, while all covariates and author-level random intercepts were retained from the primary analysis. The complete-case analysis included 869 claims, of which 28 (3.2%) were challenged. No predictor had a 94% highest-density interval excluding 1, including publication in trophy journals (OR 1.11, 94% HDI 0.33–3.77) and affiliation with a Top50 institution (OR 1.44, 94% HDI 0.48–3.94). Thus, the associations observed in the pooled analysis were not evident when using classifications based exclusively on the retrospective literature review, although estimates were imprecise because of the small number of challenged claims”

      The irreproducibility rate has increased over time as the field has grown in popularity

      - Again, why not use year as a continuous variable rather than using 5-year windows (as this would effectively include more information)? I think the categorization here is less ad hoc than in the journal/university case, but it is still an unusual decision. More important than this, however, is the fact that pairwise comparisons between periods is probably not an appropriate strategy to look for a time trend, as it excludes all information not related to the pair in each analysis. I'd strongly suggest substituting this for a straightforward regression with publication year as a continuous variable.

      We thank the reviewer for this important suggestion. We agree that categorizing publication year into five-year periods may obscure temporal trends.

      We therefore modelled publication year continuously using a spline in the multivariable analysis, thereby retaining the full temporal information while allowing for non-linearity. A straightforward linear regression would impose a linear trend, which may not adequately capture changes over time. We retained the five-year groupings solely for visualization.

      - "The subsequent increase in unchallenged claims may reflect the rapid conceptual expansion of the field, which likely outpaced the growth in the number of researchers". Do we have data on the growth in the number of researchers in the field? If so, it might be interesting to cite this.

      We do not have the number of researchers to assess the growth of the field, however we can see in Figure 8A an increase in the number of laboratories (as identified by last author name) after 1995. This together with the increase in the number of articles (Figure 3A) clearly show the expansion on of the field.

      - Isn't a "rise in unchallenged claims" over time expected by chance, as older findings will have more time to have been verified by someone else? I understand that the 14-year window probably mitigates this effect, but it still probably exists in some degree and should be mentioned.

      The reviewer is correct that the % unchallenged claims should increase over time, but this does not explain the very low number of unchallenged claims between 1992-2001.

      First-author patterns of irreproducibility

      - With 69 challenged claims and 289 first authors, it would be impossible to achieve "perfect equality" (e.g. a Gini index of 0.88). To verify how far the observed coefficient deviates from chance, it would be useful to obtain (perhaps via simulations) the Gini coefficient expected by chance given the number of challenged claims/authors, and perhaps derive a p value as well (e.g. the proportion of random permutations in which the index exceeds the observed value).

      Very good observation, thank you, we provide this analysis for both first and last author, in the main text and in a supplementary table.

      Added first author text:

      “The distribution of challenged claims among first authors was highly unequal with a Gini coefficient of 0.881 (inequality index ranging from 0 -perfect equality- to 1 extreme inequality-, Figure 4B): the top 10% of first authors accounted for 73.9% of challenged claims, and the top 20% accounted for all challenged claims. This inequality exceeded that expected by chance (reassigning randomly the 69 challenged labels across individual claims, while preserving each author’s number of claims, p = 0.000001). It also remained greater than expected when claims from the same paper were kept together (p = 0.015; Supplementary Table Sx). The concentration of challenged claims among first authors cannot be explained solely by differences in the number of claims or by clustering within papers.”

      Added last author text:

      “Challenged claims were also unevenly distributed among leading authors (Gini coefficient = 0.856): the top 10% accounted for 71.0% of challenged claims and the top 20% accounted for 94.2%. The observed inequality exceeded that expected when challenged labels were reassigned across individual claims (p = 0.00081), but not when claims from the same paper were kept together (p = 0.127; Supplementary Table S9). The apparent concentration among leading authors could be explained by multiple challenged claims arising from the same papers.”

      Added method text:

      “We compared the observed Gini coefficients with two permutation analyses, each based on 1,000,000 permutations. First, we randomly reassigned the challenged labels across individual claims while keeping every claim attached to its original author. This preserved the number of claims contributed by each author. Second, we kept the challenged-claim pattern of each paper together and reassigned these patterns among papers containing the same number of claims. This additionally accounted for the possibility that claims from the same paper were challenged because of a shared problem. For each permutation, we recalculated the Gini coefficient across authors and calculated the p-value as the proportion of simulated coefficients that equalled or exceeded the observed coefficient. Full results are reported in Supplementary.”

      - Claims in a single paper may not be fully independent from each other (as both could be irreproducible due to the same error), so it could be worth adding the article as a random variable here.

      See answer above: [“Response to claim vs article level analysis”]

      - In Fig. 4B, what defines the order of the dots between 0 and 250 (i.e. authors that have no challenged claims)? Is this merely arbitrary? This should be stated more clearly.

      Thank you, the order was arbitrary, it has now been fixed by using a secondary sorting (verified claim proportion), and Authors tied on challenged-claim proportion are ordered by verified-claim proportion has been added the legend.

      - In Fig. 5C, there are clearly less dots than there are authors. This is likely due to superposition (i.e. there are probably multiple circles with 1 article and 0% verified claims). Nevertheless, it means that the graph conveys an erroneous message. As the x axis is a discrete variable, it may be worth turning it into five categories (e.g. 1 to 5) and adding some jitter to each of them in order to let the reader know how many dots are in each category/%.

      Thank you, this was indeed due to superposition. After trying to add a jitter, the graph was still confusing, so we changed the representation to show circle whose size depends on the number of authors. The graph now conveys the size of each block properly.

      Lead-author patterns of irreproducibility

      - The same comment concerning the Gini index made for the first authors also holds here.

      We added a text, see answer above.

      - In Fig. 6B, the same comment made for figure 4B also holds.

      Thank you, the order was arbitrary, it has now been fixed by using a secondary sorting (verified claim proportion), and Authors tied on challenged-claim proportion are ordered by verified-claim proportion has been added the legend.

      - The division between "senior" and "junior" PIs is rather ad hoc here. Once more, why not use "years from first last-author publication" as a continuous variable as a more neutral way to analyze this (as done in Fig. 8)? Also note that it is not clear whether "published a lastauthor article at least five years prior to the considered publication" means any article or one about Drosophila immunity (as in Fig. 8A), and which database was used to examine this (PubMed? Other?).

      Senior and junior researchers were classified based on whether they had published a last-author article five years or more using Pubmed (for seniors). We did not specify articles in Drosophila immunity when considering seniority but only having a last author article five year before the publication.

      - In Fig. 7C, the same comment made for Fig. 5C also holds, although the solution here is less obvious as there are more possible numbers for "number of articles".

      Thank you. As in Fig. 5C, superposition obscured multiple authors occupying the same coordinates. We therefore revised Fig. 7C so that authors with the same number of articles and proportion of challenged claims are represented by a single circle. Circle size and the number shown inside indicate how many authors are represented.

      - As far as I could tell, the data on Fig. 8A refers to when authors published their first first/last author papers on Drosophila immunity (which makes it different from Fig. 7, which refers to any article, but I could be mistaken). If this is the case, I'm not sure it's correct to talk about "the period when principal investigators established their laboratories" as mentioned in the text, as (a) they could have published a first author paper in somebody else's lab or (b) they could have started a lab and only later published a paper on Drosophila immunity. "Year of entry in the field" as in the figure legend seems more appropriate.

      We have changed for ‘year of entry in the field’ as suggested by the reviewer

      - In the same figure, the 1995 cutoff seems completely arbitrary. Why not analyze this as a regression with year of entry as a continuous variable?

      The 1995 boundary is historical as it marks the expansion of Drosophila immunity from a marginal subject into a popular field. The expansion of the field at this specific time was probably driven by a combination of factors: major advances in Drosophila genetics and genomic resources and growing knowledge of antimicrobial peptides in the early 1990, a strong interest on innate immunity that drew attention on the power of Drosophila to answer key questions in this new area. These developments attracted new researchers to insect immunity, and the later discoveries of the IMD and Toll pathways in 1995 and 1996 gave the field an additional (and probably even stronger) boost. Our corpus holds 29 articles for 1959– 1994 (0.8 per year) against 371 for 1995–2011 (21.8 per year), and only 13 of the 156 PIs entered before 1995. The comparison is therefore a cohort contrast more than a contrast or search for a cutting point. Moreover, a regression would impose a linear relationship between the different years, which we believe is too constrained on regard of the observed data.

      - "We hypothesized that these authors, having gained prior hands-on experience on Drosophila immunity, would be less prone to publishing irreproducible claims.". This is a possibility, but given that most of the sample is retrospective, it's also possible that the field is more prone to publicly challenging findings from newcomers.

      Since we do not observe major difference in replicability between senior and junior PI, we still believe that training as first author in a traditional immunity laboratory compared to no training is significant.

      Irreproducibility according to research styles

      - The objective definition of "continuity" and "exploratory" PIs should be stated here for this to be interpretable (note that this is not clear in the Methods either).

      We have better defined how we separate "continuity" and "exploratory" PIs in the methods and in the result section.

      Multivariable analysis of predictors of claim irreproducibility

      - As mentioned in the Public Review, does it make sense to include "unchallenged" along with "verified" in the outcome? If the objective is to reduce the analysis to a "verified/not verified" claim, wouldn't it make more sense to remove the unchallenged findings (as these are likely uninformative, and based on the authors' own replications may be more related to the "challenged" category than to the "unchallenged" one?

      See above, “Response to unchallenged exclusion”

      - As stated previously, why not include journal impact factor, university ranking, year of first paper and year of first Drosophila immunity paper, as continuous variables rather than categorical ones (which would effectively lead to a model with less parameters when there are more than 2 categories). Particularly, there is more information to be gained from adding a single variable than from performing individual comparisons between categories when these have a natural order.

      See above, “Response to categorical variable encoding”

      Discussion:

      - As "conceptual reproducibility" is somewhat of a vague concept, it seems important to discuss (both in the Methods and Discussion) how this was operationalized. Although it's obvious some decisions of what constitutes a direct replication will vary on a case-by-case basis, general guidelines on how these criteria were set are needed.

      We have defined in the methods used to categorize article claims as well as the link with the other companion article. In the ReproSci database, all the assessment are justified allowing to see how we classified claims.

      - "Contrary to the more dramatic narratives...". Here, it is important to emphasize the key differences between this and the cited references: (a) the fact that the majority of the replicability of the sample was found on the basis of a retrospective sample and (b) the fact that the study is dealing with conceptual rather than direct replications. Also, the possibility of confirmation bias should be mentioned as an alternative hypothesis in the last sentence of this paragraph.

      This is a good point and we have highlighted that methodological differences between our study and other studies could explain differences in replicability rates.

      - "findings that, despite being published, have never been independently tested". A more accurate description may be "have never been independently confirmed or refuted in the published literature", as many (and perhaps most - see Baker 2016) replication attempts may go unpublished, as the following sentences themselves indicate.

      We have changed the text according to the reviewer’s suggestion.

      - The description of how findings came to be regarded as "suspicious" here is interesting - and an important part of how the sample was determined. In this sense, I'd consider this as part of the methods. Even though defining what makes something "suspicious" may not be completely systematic, a general description of the method that led the researchers to arrive at this list (e.g. the process that seems to be hinted at in this paragraph) deserves a thorough description.

      The description of how findings came to be regarded as "suspicious" is now detailed in the methods.

      - "Our results suggest the status of a claim being "unchallenged" is not a reliable proxy for its validity". I agree, but this is statement is in direct contradiction with the authors' decision to include unchallenged claims along with validated ones in the binary outcome of the multivariate model.

      See above, “Response to unchallenged exclusion”

      - "Trophy journals are more likely to publish articles with challenged claims than high or lowimpact, although the difference was not significant." Again, a single analysis using a continuous measure of impact should provide more statistical evidence than the pairwise comparisons.

      See above, “Response to categorical variable encoding”

      - "However, this higher rate of follow-up work cannot fully explain by itself the higher proportion of challenged claims in trophy journals." Why not, exactly? I don't remember seeing an objective analysis of it.

      If we remove the unchallenged claims, we still observe higher rate of irreplicability in trophy journal: 15.85% in trophy journal versus 9% in high-impact and 8.1% in low-impact. So we believe that the higher rate of irreplicability in trophy journals cannot be explained by a lower level of unchallenged.

      - The very large paragraph on pages 21-22 could be split into two, one about journals and the other about universities.

      This has been done.

      - "Our data align with the broader narrative of increasing rates of non-reproducible science.". Does it? The time trend did not seem very clear to me (and I would argue that it was not analyzed properly).

      As indicated in the text, we observed an increase in the rate of irreplicable claims over time; however, this trend did not reach statistical significance. We are confident in the robustness of our analysis and therefore did not pursue this question further. While our findings align with concerns about non-replicable science, they also offer a more nuanced perspective, showing that in some fields, replicability remains significantly high.

      - "A similar but more acute pattern was observed during the SARS-CoV-2 pandemic." References should be provided here.

      We have added this reference to suggest that articles published on SARS-CoV-2 pandemic are overall less reliable than others: An alarming retraction rate for scientific publications on Coronavirus Disease 2019 (COVID-19) Nicole Shu Ling Yeo-The https://doi.org/10.1080/08989621.2020.1782203

      - Some of the trends discussed here (e.g. time, continuity vs. exploratory style are not supported (even as a trend) by the multivariate model, and this caveat should be mentioned. An odds ratio of 0.89 with a very wide confidence interval may be too weak in terms of evidence strength to merit a whole paragraph in the introduction discussing it.

      We did not mention the question of time horizon or the distinction between continuity and exploratory research styles in the Introduction; however, we believe these issues merit consideration in the Discussion. In particular, the contrast between continuity-driven and exploratory approaches is noteworthy. Although we found no significant difference in the rate of challenged claims between these approaches, there was a significant difference in the rate of unchallenged claims, which are subsequently more likely to be challenged. This finding is important in the current funding landscape, where many agencies prioritize short-term, exploratory projects. Such incentives may inadvertently contribute to issues of irreproducibility.

      - There are many other limitations beyond those mentioned, in particular the fact that much of the analysis is retrospective and based on a potentially biased literature. The fact that the analysis does not seem to be preregistered and is dependent on a lot of ad hoc decisions also merits discussion. This should definitely be explored in more detail in the Limitations section.

      These limitations are now discussed in the limitation section at the end of the discussion. We agree on the fact that our analysis was not preregistered and that some of our conclusions are raised after analyzing the data. This article is part of a more global project including the databases with all the information. However, full objectivity when analyzing literature is not unachievable, and these critics are inherent to all reproducibility projects.

      Methods:

      Selected articles, annotation and experimental validation:

      - "In brief, a list of 400 publications published before 2011 was generated using a curated search string on the publicly available PubMed database.". Please state that the search string is available in the companion article (e.g. https://doi.org/10.1101/2025.07.07.663442) or include it here.

      This is now stated.

      - "Selected primary articles were annotated by a single researcher".

      By "annotation", do the authors mean the extraction of claims? This is not self-evident. Also, what does the "review" process entails? Do the authors have any data on agreement?

      We do not have data on agreement but overall, there were few discrepancies. We agree that the way to section article in separate claims could affect the conclusion in a number of cases. All the data are public and we did not have any request from the community. The fact that this project depends from arbitrage from the two annotators is mentioned in the limitations.

      - "Claims were cross-checked with evidence from previous, contemporary and subsequent publications and assigned a verification category."

      How was this process performed? This seems to be quite complex and there's hardly any information about it. And once again, do the authors have any data on how reproducible this process would be when performed by different people?

      All the data, notably claim assessment and justifications, are available on the website. Findings cross-checking the claim could be identified by knowledge of the authors, checking articles that quotes the articles. This was an enormous amount of work with human arbitrage. We expected more feedback from the community. The reaction of the community will be described in a companion article later.

      - The authors mentioned that annotations and verifications were made available for comments on the community, but how were these comments incorporated if different opinions were voiced? Can the authors provide data on how frequent these comments were, and how often they were incorporated?

      All data is available on the website, where members of the community can publicly comment on our assessments. The authors can see these comments on the ReproSci website

      Table 1:

      - Again, the three "partially verified" categories in Table 1 seem to refer to very different situations and it would be interesting to break these down somewhere.

      We agree that categories ‘mixed’ and ‘partially verified’ are complex. They represent situations where a decision was not easy. We prefer not to break down those categories in multiple sub-categories.

      First and last author classification:

      - What does "status of first author" mean?

      Status means their position: technician, PhD student, Post-doc, PI….

      - "PIs were manually classified as..." - there seems to be something missing in this sentence (e.g. "as senior or junior on the basis of whether they have...")

      We have added PIs: PIs were manually classified as senior or junior PIs on the basis on having published an article in Pubmed as last author more than five year ago.

      - The distinction between continuity and exploratory is not explained in objective terms. Is there any objective definition of what it means to "continue to work in the field". Publishing a paper in the last X years? And what counts as a "transient" incursion?

      We have better explained in the methods this distinction.

      Statistical analysis:

      - Can the authors define exactly what they mean by "weakly informative priors"?

      Weakly informative prior are Bayesian prior distribution that are meant to be vague as to let the data dominate the results (vs the prior belief of the scientist). They can be seen as very close to flat priors, which are not used here because they cause numerical convergence issues.

      - The authors mention a lot of variables included in the multivariable analysis, but little information is provided on the categorization. What are "low, high and trophy journals"? What university rankings are used?

      We used Shanghai Ranking’s 2010 Academic Ranking of World Universities as our ranking of university. Our binning of ranking in categories make the analysis less susceptible to changes in ranking system. We selected the Shanghai ranking as it primarily evaluates research output and awards (where Times also evaluate teaching, and QS also employer reputation), which we believe represent better the variable than may affect irreplaceability. For the categorization, please see answer above: “Response to categorical variable encoding”

      - "Exact formulas (...) are available in the public repository". What repository do authors mean? The project website? The GitHub repository? Please specify and provide a direct link if possible.

      We added a link to the main text (it was only in the supplementary document). The Supplementary document contains information on how to reproduce the results.

      Reviewer #2 (Recommendations for the authors):

      Specific comments:

      Some methodological details related to the main conclusions of the paper are missing.

      We have extended the methodological section and also better link this article to the companion article.

      - Impact factor: It is unclear which year(s) of journal impact factors were used. Trophy journals are defined as those with an impact factor >50 (Science, Nature, and Cell), but according to Figure 2B, there are only three such journals. Are "trophy journals" limited to these three, or are others included that meet the >50 impact factor threshold?

      These are the 2022 impact factor, we made that clear in the main text. Moreover, only 3 journals in our datasets had an impact factor >50: Nature, Cell, and Science (so the Trophy Journal category is limited to these 3, but not by definition). This is explicit in Table S1: List of journals with impact factor and claim assessment.

      - University ranking: Similarly, please clarify which year(s) of university ranking data were used. Since rankings vary depending on the system (e.g., QS, Times, Shanghai), were the conclusions consistent across multiple ranking sources?

      We used Shanghai Ranking’s 2010 Academic Ranking of World Universities as our ranking of university. Our binning of ranking in categories make the analysis less susceptible to changes in ranking system. We selected the Shanghai ranking as it primarily evaluate research output and awards (where Times also evaluate teaching, and QS also employer reputation), which we believe represent better the variable than may affect irreplaceability.

      Reviewer #3 (Recommendations for the authors):

      Table S4 - for leading author there is a variable of 'historical lab after 1998 continuity' - should that be 1995?

      Correct, we changed this label to Trained in Historical laboratory (Comparison restricted to claims published after 1995) to make it more explicit.

    1. Undocumented security depends on specific people remembering specific things, and that breaks in three predictable ways: New staff cannot follow rules nobody wrote down; when a key person leaves, their knowledge leaves with them; and different people do the same task differently, which is another word for inconsistently.

      This is a long sentence. Please break it up into smaller pieces to make it easier to follow.

    1. Draft → Submit for review → Reviewer chooses "Approve and publish" or "Request changes" → Published version → Sections use a snapshot. Authors cannot publish directly by default.

      PXU + RnD

    1. How are people’s expectations different for a bot and a “normal” user?

      Bots, I feel are more of an annoyance for many including me. It definitely depends on what it is used for, however, in many cases online like posting comments or engagement farming, it is pretty much useless and takes up space on the platform.

    1. https://indy.peergos.me/%F0%9F%93%93/2026/10/%F0%9F%91%A4/indy/%E2%99%96%F0%9F%8C%90/2026/10/8/@helia_strings-npm.html

      http://localhost:7777/#%7B%22nonce%22:%22wK6LMCEEkfqwPH6AIumLGD2EwLFzOmWe%22%2c%22ciphertext%22:%22rJG+3NTsp/gNJm2tebhPxrm/ndAqVwRtveTzI1anuFmrz+fU8Q/AnHsf9vRYhqtfC5AVbRK348h8fRRAM2DpC850Q1uHQgQsLbEFkkFrrwLA1z5+UQdLLgpf0kvO+Gy/9JVN8VrJXU5Rg01fUxo2PU18WwkGH1s7sHTDBAcpZfoKX1BIrN52SDWsHjCDMt5mQ0ZwzmcIFvaWX65oc/CuoneupcfRbBUGK3KBHKlT+T13Y0RPJZF7DZseXAtzMoOOAfgWQcAxCd3Xg3+MbMjxYNAhklgSvp87eY/A0C4RJRqrQ/zmVgVJPyUPLAg+UvTjBExZD95k5anrlEYhzK5gsuxQJ+OnJU3npTB5DzCerbwR8QqaBmfVaT3Qpy2KMApm/8ejcF6QsQJ3279k+PcVJnXuZKWZx7g5q45dVg==%22%7D

    1. Document de Briefing : Intégration du Design de Service et de la Maîtrise d'Usage dans le Bâti Scolaire et la Commande Publique


      Synthèse Exécutive

      Ce document de briefing synthétise les enseignements issus des travaux et podcasts organisés par la Direction Interministérielle de la Transformation Publique (DITP) en collaboration avec le ministère de l'Éducation nationale, de la Jeunesse et des Sports (MENJ).

      L'analyse porte sur la transformation des établissements scolaires et des bâtiments publics par l'application des méthodologies du design, de l'assistance à la maîtrise d'usage (AMU) et de la permanence architecturale.

      Enjeux et Constats Clés

      • Complexité institutionnelle : La conception et la rénovation des bâtiments scolaires souffrent d'une répartition rigide des compétences : l'État (MENJ) définit la politique éducative et emploie les personnels, tandis que les collectivités territoriales (communes, départements, régions) gèrent la construction, la maintenance et les usages périscolaires/extrascolaires.

      • Nécessité de dépasser le cloisonnement : Les processus traditionnels de la commande publique découpent séquentiellement les projets (programmation, conception, réalisation, gestion).

      Ce découpage empêche de prendre en compte la réalité des usages et génère souvent des inadaptations spatiales ou fonctionnelles.

      • Le rôle pivot du design de service : Placer les usagers (élèves, enseignants, agents d'entretien, surveillants, riverains, parents) au cœur de la conception dès la phase amont permet de lever les craintes d'exigences extravagantes, d'optimiser les coûts d'investissement et de fonctionnement, et de favoriser le climat scolaire.

      • Pratiques émergentes dans la commande publique : La permanence architecturale, la résidence de designers, la programmation ouverte ("en actes") et le droit à l'expérimentation locale émergent comme des leviers majeurs pour adapter la commande publique aux réalités du terrain et réactiver la démocratie locale.


      1. Cadre Institutionnel et Recomposition des Compétences

      1.1 La Répartition Réglementaire et ses Limites

      Le secteur scolaire en France concerne 63 000 lieux, accueillant quotidiennement 12 millions d'élèves et 1 million de personnels.

      La répartition théorique des compétences s'établit comme suit :

      • L'État : En charge du programme pédagogique, du cadre de travail et de l'exercice d'enseignement de ses agents.

      • Les Collectivités Territoriales : En charge du bâti, mais développant également leurs propres politiques éducatives locales (périscolaire, extrascolaire, restauration, mise à disposition de personnels comme les ATSEM).

      Dans les faits, le bâtiment détermine directement les pratiques pédagogiques et constitue le cadre de travail quotidien des agents.

      Le cloisonnement strict des compétences apparaît insuffisant pour traiter la complexité de l'espace scolaire.

      1.2 Le Positionnement de la DITP et du Ministère de l'Éducation Nationale

      Face à cette complexité, l'État privilégie la concertation et la mise à disposition de ressources plutôt que l'ajout de contraintes normatives :

      • La DITP (Mission Innovation) : Prône la territorialisation de l'innovation publique et la valorisation des compétences de design de service.

      La DITP soutient l'expérimentation locale pour tester, échouer, adapter et essaimer les bonnes pratiques à l'échelle nationale.

      • Le MENJ (Cellule bâti scolaire) : À la suite d'une concertation publique ayant recueilli plus de 200 000 contributions, le ministère développe des référentiels et outils axés sur les usages pour nourrir le dialogue entre collectivités et équipes pédagogiques.

      2. Méthodologie du Design de Service et Maîtrise d'Usage (AMU)

      2.1 Définition et Posture du Designer en Commande Publique

      Le design de service et l'assistance à la maîtrise d'usage ne se limitent pas à rendre un bâtiment "pratique".

      Il s'agit d'une compétence d'écoute, de traduction et de synthèse des besoins du terrain :

      [Besoins bruts des usagers] ──> [Filtre & Traduction par le Designer/AMO] ──> [Programme & Projet Architectural]

      • Le designer comme filtre et traducteur : Il canalise les demandes du terrain, élimine les attentes irréalistes et traduit les besoins d'usage en solutions spatiales ou organisationnelles concrètes.

      • L'action "en creux" des référentiels : Le designer intervient dans les espaces laissés libres par les référentiels techniques et programmatiques pour apporter du bon sens et de la flexibilité.

      2.2 Modalités d'Intervention : Permanence et Temps Long

      L'intégration de l'usage nécessite d'inscrire le designer ou l'architecte dans un temps long sur le terrain :

      • La résidence / permanence : Présence continue du designer pendant la phase d'étude et les années de chantier (ex. résidence de 3 ans).

      • Le prototypage et la préfiguration : Tester les usages en amont (lignes au sol, mobilier temporaire, réorganisation de pièces) pour éprouver les concepts avant leur réalisation définitive en dur.

      • La réversibilité : Conserver une marge d'adaptation post-livraison pour permettre au bâtiment de se bonifier avec le temps plutôt que de subire une dégradation anticipée.


      3. Retours d'Expérience et Analyses de Cas Concrets

      3.1 Département du Val-d'Oise : Expérimentation sur les Sanitaires et Cours

      • Contexte : Gestion d'un patrimoine d'une centaine de collèges face à des problèmes récurrents de dégradations et d'usages dans les sanitaires.

      • Démarche : Consultation des collégiens, des agents d'entretien et des surveillants sur 5 collèges pilotes.

      • Résultats : Démystification des craintes d'exigences extravagantes ; émergence de demandes rationalisées ; apaisement des équipes de gestion patrimoniale et sécurisation des chefs d'établissement grâce au principe de réversibilité des aménagements.

      3.2 Saint-Herblain (Collège Ernest Renan, Loire-Atlantique)

      • Contexte : Reconstruction d'un collège innovant, avec une mission AMO (Agathe Chiron) intégrant une compétence obligatoire de design dans le cahier des clauses techniques particulières (CCTP), attribuée à la maîtrise d'œuvre K Architectures (Charlotte Cauwer en résidence).

      • Innovations nées de la maîtrise d'usage :

        • Cabinet de co-enseignement : Espace intermédiaire situé entre deux salles de classe, permettant le travail en petits groupes sans couper les élèves à besoins particuliers du groupe-classe.
      • Centre de Création et d'Exposition (CCE) : Lieu hybride et mutualisé remplaçant la salle multimédia obsolète, géré de manière flexible pour les expositions, le travail numérique ou l'autonomie des élèves.

      • Redéfinition du parvis et de l'accès : Repositionnement de l'entrée principale pour créer un espace d'accueil ouvert, pacifiant les relations avec le quartier et permettant la présence des familles.

      3.3 Bagneux : "Le Lycée avant le Lycée"

      • Contexte : Absence de lycée d'enseignement général à Bagneux depuis 1957.

      Inscription au Plan Pluriannuel d'Investissement (PPI) de la Région Île-de-France.

      • Démarche : Projet partenarial associant la Ville de Bagneux, la démarche de La Preuve par 7 (Patrick Bouchain) et l'association Le Plus Petit Cirque du Monde, financé en partie par le promoteur privé BNP Promotion Immobilière sur un ancien site industriel de 16 hectares.

      • Méthode : Une permanence sur site préfigure le futur établissement à travers des chantiers participatifs, des actions culturelles et la construction d'un bâtiment éphémère.

      Le projet cherche à associer le Rectorat et la Région en amont des phases réglementaires classiques.

      3.4 Thiers : Rénovation Scolaire et Maillage Social

      • Contexte : Ville industrielle en mutation, comportement scolaire impacté par une forte précarité (8 groupes scolaires dont 6 en REP, turnover d'élèves atteignant 20 à 33 %).

      • Démarche : Réflexion globale couplant la rénovation pédagogique des écoles avec le programme Cœur de Ville, le Centre Social et le dispositif Territoires Zéro Chômeur de Longue Durée.

      • Enjeu : Éviter que la démocratie participative ne soit préemptée par les seules classes moyennes ; aller chercher les populations invisibles et intégrer la dimension socio-psychiatrique et de santé mentale dans l'accompagnement des familles.

      3.5 L'Île-Saint-Denis : Internalisation de la Culture de l'Usage

      • Contexte : Petite commune insulaire (8 000 habitants, budget contraint) devant réhabiliter et agrandir son dernier groupe scolaire en REP.

      • Freins rencontrés : Rédaction complexe du cahier des charges de programmation ; dérive du marché public attribuant un poids prépondérant au critère prix (40 %).

      • Levier retenu : Recrutement interne d'un agent dédié à la participation citoyenne et à l'usage pour diffuser cette culture de manière transversale dans tous les services municipaux et s'affranchir de la dépendance exclusive aux prestataires externes.


      4. Commande Publique : Obstacles, Leviers et Modèles Économiques

      Le tableau ci-dessous récapitule les blocages identifiés dans les processus traditionnels et les leviers d'action préconisés par les praticiens de l'urbanisme et du design public :

      | Domaines | Freins et Obstacles Identifiés | Leviers et Solutions Méthodologiques | | --- | --- | --- | | Séquençage des projets | Découpage étanche entre programmation, conception, réalisation et gestion. | Implication des gestionnaires et des usagers dès la définition initiale du besoin. | | Comptabilité publique | Séparation stricte et étanchéité réglementaire entre budget de fonctionnement et d'investissement. | Utilisation de la permanence architecturale comme un outil d'investissement préfigurateur. | | Culture du risque | Crainte juridique des techniciens et des élus face aux procédures de concertation et d'expérimentation. | Recours au droit à l'expérimentation locale ; création de jurisprudence et de "culture des précédents". | | Participation citoyenne | Captation de la parole par les catégories sociales les plus insérées ou les interlocuteurs habituels. | Permanence de terrain, chantiers ouverts, présence hors des salles de réunion traditionnelles. | | Conduite du changement | Récapitulation rigide des postures d'enseignement ou d'administration après livraison du bâtiment. | Accompagnement continu à la transformation des pratiques et ouverture mutualisée des espaces (ex. cours ouvertes). |

      4.1 La Permanence Architecturale comme Tiers-Acteur

      La permanence architecturale s'affirme comme une méthode centrale portée par des collectifs (tels que La Preuve par 7, Ici, Barbara, Zerm).

      Elle agit comme un espace de recherche-action directement implanté sur le terrain :

      • Elle permet de tester la réversibilité et d'étudier la faisabilité "en actes" d'un projet avant le gel des crédits d'investissement (ex. projets à l'Hôtel Pasteur de Rennes ou au Couvent des Clarisses à Roubaix).

      • Elle permet de créer un partenariat "Public-Privé-Particulier" en réimpliquant la société civile dans la gestion des biens communs publics.


      5. Citations Clés et Synthèse des Intervenants

      Le tableau suivant présente les déclarations marquantes des intervenants, illustrant la diversité des perspectives territoriales et institutionnelles :

      | Intervenant(e) | Fonction / Structure | Citation Clé / Thématique | | --- | --- | --- | | Pauline Lavagne d'Ortigue | Cheffe de la mission innovation (DITP) | "Sans les usagers d'une part et sans les agents d'autre part, il y a peu de chances que les politiques publiques et les services publics touchent leur cible et soient réellement efficaces pour l'intérêt général." | | Sidi Soilmi | Directeur de projet, Cellule bâti scolaire (MENJ) | "Contrairement à ce que peut faire souvent l'État quand c'est complexe, où l'on va rajouter de la norme, l'approche que nous avons retenue est tout à fait différente \[...\] pour construire mieux, il faut accepter de décloisonner." | | Ariane Epstein | Responsable du pôle Design (DITP) | "Notre mission à la DITP c'est : comment on fait aujourd'hui pour trouver ensemble les moyens d'accompagner, de démocratiser, de généraliser ces démarches pour que ce qui est de l'innovation unique à un endroit devienne la norme ensuite pour tout le monde." | | Sophie Ricard | Architecte-urbaniste (La Preuve par 7) | "La permanence architecturale, c'est un alibi pour réactiver nos démocraties locales \[...\] se dire qu'on ne doit plus être le maillon en bout de chaîne qui répond à une commande, mais qu'on doit être là dès le départ à l'écriture de cette commande." | | Cécile Roussel | Directrice Patrimoine (Val-d'Oise) | "On s'est rendu compte que l'on n'a jamais eu de demandes extravagantes, et a contrario, on avait des demandes relativement raisonnables sur lesquelles on était en capacité de proposer des solutions." | | Sébastien Mandoux | Principal (Collège Ernest Renan, Saint-Herblain) | "Quand la première traduction sur plan est présentée en assemblée générale plénière, il y a des professeurs qui disent : 'Mais ça c'est notre idée !' Et du coup, ils ont encore plus envie de s'investir dans le projet." | | Charlotte Cauwer | Designer de service (K Architectures) | "La concertation vient surtout se placer dans les creux du référentiel \[...\] Ma méthode, c'est de tenir le crayon : tout comme un écrivain public va tenir la machine à écrire, je tiens un crayon et devant eux j'interprète ce qu'ils me disent." | | Stéphane Rodier | Maire de Thiers | "Ce que l'on veut, c'est rompre avec la société comme verdict, comme aurait dit Didier Eribon \[...\] Il faut donner à l'idée le réel dont elle a besoin." | | Marie Anquez | Maire adjointe (L'Île-Saint-Denis) | "Quand on met le doigt dans l'usage, on n'en sort plus. Et surtout ce qui m'a beaucoup intéressé, c'est l'effet sur l'équipe enseignante qui était manifeste, et l'effet sur les enfants." | | Martine Marchand-Prochasson | Cheffe de projet Lycée (Bagneux) | "Souvent le nouveau naît de la contrainte, et le projet est né d'une anomalie : Bagneux n'a pas de lycée d'enseignement général \[...\] la caractéristique du projet c'est d'être partenarial." |


      6. Recommandations Pédagogiques et Stratégiques

      • Activer le travail d'explicitation des besoins dès la phase amont : Réunir systématiquement l'ensemble des parties prenantes (usagers, agents, techniciens, élus) avant la rédaction définitive des cahiers des charges ou des programmes de consultation.

      • Introduire la compétence Design dans les marchés publics : Spécifier explicitement dans les règlements de consultation et les CCTP l'exigence d'une compétence en design de service ou en assistance à la maîtrise d'usage au sein des équipes de maîtrise d'œuvre.

      • Inscrire les démarches dans le temps long par la résidence : Privilégier la présence physique régulière de professionnels de l'usage sur le terrain pour favoriser le prototype, l'écoute active et la co-construction du récit architectural.

      • Favoriser la réversibilité et l'adaptabilité : Concevoir des aménagements légers et modifiables permettant d'ajuster les espaces aux évolutions pédagogiques et d'apaiser les craintes d'irréversibilité des équipes de direction.

      • Valoriser et documenter le récit des démarches : S'appuyer sur des plateformes de ressources, des réseaux spécialisés (ex. Réseau Designers Publics) et des chaires de recherche universitaire pour formaliser les modèles économiques et juridiques reproductibles.

    1. Author note (corresponding author, on behalf of all authors): We have identified several points in this version that will be corrected or clarified in a revised version, to be posted shortly:

      Damages and legal costs. Some totals combine damages and legal costs without saying so. For example, the 2024/25 figure of £760m is total paid (damages plus NHS and claimant legal costs), of which £626m is damages. All totals will be labelled consistently, and the basis of each calculation stated. HRG figures. These are National Schedule of NHS Costs reference costs, not Payment Scheme tariff prices. The terminology will be corrected, and the explanation of the year-on-year change in planned caesarean unit costs revised. Labour attribution. Our eleven-code attribution (55.9% of obstetric damages) counts only labour-exclusive cause codes. Because NHS Resolution records a single primary cause per claim, intrapartum harm recorded under generic codes is excluded, so the figure is a lower bound. This will be stated explicitly, and we will add a sensitivity analysis. Any less conservative attribution increases rather than reduces the cost difference between planned vaginal birth and planned caesarean.

    1. eLife Assessment

      This useful study shows that neural activity before a conversational response contains information about the duration and timing of the subsequent turn. The evidence that this information originates from speech-production planning is incomplete. Because the analyzed EEG is recorded during processing of the partner's speech, and because that speech causally determines the subsequent response, production planning and comprehension/context are confounded in the current analyses. It would strengthen the production interpretation considerably if the authors could show that the relationship with forthcoming duration survives richer controls for the preceding turn, or ideally after accounting for EEG variance driven by the incoming speech.

    2. Reviewer #1 (Public review):

      This manuscript examines EEG activity during face-to-face Diapix conversations and asks whether neural activity preceding a participant's next speaking turn predicts the latency and duration of that turn. The authors report sustained ERP and alpha/beta effects, particularly associated with upcoming response duration, together with temporal-generalization decoding beginning approximately 1-1.3 s before speech onset. They interpret these results as evidence for two stages of conversational speech planning: an early process that activates and maintains the forthcoming response, and a later process related to commitment and motor execution. The naturalistic dyadic EEG setting, relatively large number of conversational turns, trial-level analyses, and convergence across ERP, oscillatory, and multivariate approaches are important strengths.

      I nevertheless have substantial concerns about the interpretation. Most importantly, the EEG analyzed during this period is recorded while participants are processing their partner's speech. Therefore, neural activity necessarily includes auditory, linguistic, semantic, pragmatic, and turn-boundary processing, in addition to any potential response-planning activity. Because properties of the incoming utterance are themselves strongly related to what the listener subsequently says, predicting subsequent response duration or latency does not by itself isolate production planning. The authors acknowledge some of these limitations, but the current title, theoretical framing, Figure 8, and conclusions remain considerably stronger than the evidence allows. Indeed, the Discussion itself acknowledges that the signals cannot be determined to reflect linguistic content, general planning demands, or motor preparation.

      I therefore think the study could become valuable, but substantial additional analyses and conceptual reframing are required.

      Major concerns:

      (1) The central problem is dissociating speech planning from listening/comprehension

      This is my main concern with the manuscript and also with the claims in the Discussion. The authors repeatedly infer that because EEG during the partner's turn predicts the duration or latency of the participant's subsequent response, this activity reflects speech planning. For example, the manuscript states that examining the period before self-speech onset "isolat[es] preparatory neural activity before articulation." I do not think this inference is justified.

      During precisely this interval, the participant is also listening to and comprehending the partner. The incoming utterance is not independent of the forthcoming response; it is what determines the response. A complex statement, several questions, a request for clarification, or an information-rich description will produce different auditory/linguistic/comprehension activity and may also systematically change the length and timing of the subsequent response. Consequently, both partner speech → listening/comprehension EEG and partner speech → subsequent response properties can generate an EEG-response relationship without requiring that the measured EEG activity itself reflect production planning.

      The authors control partner-turn duration, which is useful, and they show that self-duration, self-latency, and partner duration are behaviorally related. However, partner duration is only a very coarse property of the input and cannot control for what participants are actually processing. The analysis should consider, at minimum, partner word count/speech rate, acoustic properties, information density or lexical/syntactic complexity, dialogue act, question versus statement, and preferably semantic/discourse content. For example, the number of questions or propositions in the partner's turn may directly determine both comprehension demands and how much the listener subsequently needs to say.

      Given that transcripts and word-level timing are already available, the authors should be able to substantially improve this analysis. One particularly useful approach would be to construct an encoding model of the EEG driven by acoustic and linguistic features of the partner's speech and then test whether the residual neural activity continues to predict subsequent response properties. At minimum, richer input-related covariates are necessary.

      Without such analyses, I think the main conclusion should be reframed from "neural signatures of speech planning" to something closer to "pre-response neural activity during conversational listening predicts properties of the subsequent turn." The latter is supported by the data; the former requires stronger dissociation.

      (2) It is not clear that the analyzed period is actually a "listening interval".

      Relatedly, the definition of the analyzed interval needs careful reconsideration. Epochs are aligned to the participant's own speech onset and extend back 2.2 s, whereas a turn transition is accepted whenever self-onset falls between −1 and +1 s relative to partner offset.

      Therefore, the interval from −2 to 0 s relative to self-onset is not necessarily an interval during which the partner is continuously speaking. For positive response latencies, part of this interval contains silence following partner offset. If the preceding partner IPU is relatively short, earlier portions of the epoch could potentially contain a different conversational state altogether. Yet the manuscript describes the whole pre-onset period as being "while they are listening to their partner."

      This matters particularly for the comparison between the early and late periods. Different response-latency conditions will necessarily place the partner offset at systematically different positions within an epoch aligned to self-onset. The authors already demonstrate this problem for the anterior latency ERP, where its inflection tracks the distribution of partner offsets. This is an important observation, but I do not think the problem is restricted to that single component.

      The authors should quantify, separately for TW1 and TW2, how much time is occupied by: active partner speech, silence after partner offset, overlapping speech, and potentially other conversational states.

      I would strongly recommend repeating the critical analyses for periods/trials during which the partner is actually speaking, and/or explicitly modeling partner-speech presence and the temporal distance to partner offset. Analyses aligned with both partner offset and self-onset would also help separate responses to the conversational boundary from activity specifically preceding self-production.

      (3) The Early-versus-Late Planning hypotheses are not sufficiently differentiated

      I was confused by the theoretical contrast (e.g., Lines 53-62). The Early Planning account is described as proposing that response preparation begins once sufficient information is available, whereas the Late Planning account explicitly allows "some aspects of conceptual preparation" during listening but proposes that full articulatory planning is delayed until near the turn boundary.

      Under these definitions, both hypotheses allow for early planning. They differ mainly in the level of production planning that occurs early, that is, conceptual/content preparation versus articulatory/motor preparation. This is quite different from a simple early-versus-late timing hypothesis.

      This is problematic because the current EEG measures do not determine whether the early activity reflects conceptual planning, lexical formulation, articulatory planning, comprehension, or general cognitive demands. Indeed, the authors explicitly acknowledge this later. Consequently, observing EEG-behavior relationships 1 s before speech cannot adjudicate between the two hypotheses as currently formulated.

      The theoretical framework should therefore distinguish explicitly among at least: conceptual/message planning → linguistic formulation → articulatory/motor preparation → initiation.

      If the "Late Planning" hypothesis specifically concerns late articulatory planning, then the authors need evidence that distinguishes articulatory planning from earlier conceptual processes. Otherwise, I would describe the findings more conservatively as showing early behavior-related neural activity followed by stronger onset-proximal sensorimotor activity, rather than claiming that Early and Late Planning accounts have been reconciled.

      (4) Response duration is not equivalent to the amount or content of a response that has already been planned

      The strongest empirical finding concerns upcoming speaking duration. I agree that this is interesting. However, throughout the Discussion, duration is gradually interpreted as the "extent," "scale," or even the amount of content already prepared. For example, the authors argue that longer responses involve maintaining a larger set of propositions, lexical items, or planned sequences. Figure 8 goes even further and explicitly contrasts "High Content" and "Low Content," although no measure of speech content has been analyzed.

      I do not think response duration alone supports this inference.

      A 2-s utterance is not necessarily twice as much planned content as a 1-s utterance. Duration can vary because of speech rate, hesitations, lexical retrieval, syntactic formulation, repetitions, interactive feedback, or decisions made after speech has already begun. An eventual IPU duration is therefore not necessarily specified before articulation.

      This issue is especially important because the manuscript explicitly acknowledges that the analyses cannot establish whether the signals encode lexical or semantic content. It is then inconsistent to conclude later that early activity supports "what and how much to say."

      Since transcripts are available, the authors could directly test several alternatives: word count, content-word count, number of propositions, speech rate, syntactic complexity, semantic similarity between partner and self turns, and possibly dialogue-act categories. If the early signal predicts eventual linguistic content/amount rather than merely duration, the claim of maintained response specification would become considerably stronger. If not, the interpretation should remain at the level of response duration.

      (5) The ROI/time-window linear mixed-model (LMM) analysis appears circular

      I have a methodological concern about the relationship between the cluster analyses and subsequent LMMs. The authors first identify electrodes and time ranges showing significant differences between long versus short responses or fast versus slow responses. They then use those same data-driven clusters to define anterior/posterior ROIs and TW1/TW2 and extract EEG values from the same trials, after which the extracted EEG measures are tested as predictors of the same behavioral variables.

      This creates a potentially serious double-dipping/selection bias problem. The continuous LMM outcome is not statistically independent of the median-split outcome used to select the neural features. FDR correction at the LMM stage does not correct for this feature-selection procedure.

      I think the confirmatory analyses need independent feature definition, for example, predefined electrodes/time windows from previous literature, split-half analysis, leave-one-participant-out feature selection, or another cross-validated procedure in which the data used to identify the ROI/window are independent from those used to estimate and test the EEG-behavior relationship.

      This is particularly important because the LMM results are used as evidence that the effect survives behavioral covariates. That conclusion should come from statistically independent tests.

      (6) Several statistical and validation issues need additional attention

      There are several related concerns here. First, the trial-level LMMs include only (1 | participant), despite the fact that the observations come from dyadic interactions, multiple 4-min conversations, and temporally adjacent conversational turns. Participants within the same dyad do not generate independent observations because the speech of one participant directly determines the context for the other participant. I would expect the hierarchy to include at least a dyad/conversation block, and random slopes for critical within-participant predictors should be considered. Robustness to conversational autocorrelation should also be examined.

      Second, the authors describe response duration as approximately log-normal, but the model description indicates only centering/scaling rather than transformation. The authors should report residual diagnostics and consider modeling log-duration or using an appropriate non-Gaussian mixed model.

      Third, the decoding uses stratified 10-fold cross-validation, but it is unclear how conversational dependence is handled. Randomly assigning turns from the same 4-min conversation to training and testing sets may permit classifiers to exploit slow neural states, discourse context, or block-specific characteristics shared by temporally adjacent turns. A more convincing analysis would use blocked cross-validation, such as leave-one-conversation/run-out. The authors should also clarify whether decoding is performed independently within participants and whether median thresholds are participant-specific or global.

      These analyses are particularly important because temporal generalisation establishes stable predictive information, but not the functional identity of that information. A stable partner-speech/context signal could also produce a broad temporal-generalisation pattern.

      (7) ERP baseline correction and potential speech/muscle contamination require clarification

      The ERP preprocessing raises another concern. ERPs were baseline corrected using the entire −2200 to +500 ms epoch. This interval includes the first 500 ms after self-speech onset, the period most vulnerable to speech-related muscular activity. Because the main grouping variable is upcoming response duration, post-onset activity may itself systematically differ between long and short responses. Subtracting a baseline that includes this activity could introduce condition-dependent shifts in the pre-speech ERP.

      I therefore think the critical ERP results should be demonstrated without using post-speech activity in baseline estimation, for example, by using an appropriate pre-speech/regression-based baseline strategy, or by analyses demonstrating that the result is insensitive to baseline choice.

      There is also an intriguing inconsistency concerning EMG. The participant section states that facial hair was an exclusion criterion to ensure electrode contact for EMG recordings, whereas the Limitations section suggests that combining EEG with EMG would be useful in future work. If EMG was in fact recorded, it should be described and, ideally, analyzed to establish when peripheral speech preparation begins and whether the onset-proximal alpha/beta results remain significant after controlling for muscle activity. This would be particularly valuable for the proposed late "motor commitment" interpretation.

    3. Reviewer #2 (Public review):

      This is an excellent paper, well researched, correctly referenced and very well presented. The main addition to the existing literature, apart from supporting earlier results (see also Roberts et al. 2015), is the ability to detect the length of the upcoming conversational contribution from the neural signal. This is interesting because it shows very early and sustained planning of response during listening to the incoming turn from the interlocutor. Why is that interesting? Because this simultaneous listening and response planning is utilising the very same mental machinery - a kind of dual-tasking that humans generally find very hard, given extreme capacity constraints on short-term and working memory. This then talks to the issue of whether speech production and comprehension are more independent than current theory holds. Clearly, human language is one of the elite skills of the species, but this dual-task ability in just this domain may be a specialization evolved over deep evolutionary time. I would like to see this background sketched more clearly in the introduction - it would make the paper of more interest to the general reader.

      The other main contribution is methodological, namely the ability to extract meaningful EEG out of unconstrained verbal interaction, using, e.g. machine learning to filter out movement artefacts. Earlier work was held back by scepticism that this was possible, but this paper points clearly to one way to do this. A further contribution is the suggestion that Beta modulation is associated with the duration of the turn under planning, while Alpha modulation was in addition associated with preparation for execution. This needs to be tested in further studies, but could be a valuable addition to the analytical armoury.

      In the final discussion, a conceptual model is outlined in Figure 8. This looks very similar to that given in Levinson & Torreira 2015 and Levinson 2016 - at least if there is a difference, it would be good to see it clarified. Note that in that model we hypothesized that preparation of the upcoming turn was continued right up to completion and then held if necessary for ending cues from the interlocutor. Note that quite late delivery preparations have been noted both in Torreira et al. 2015 (breathing signal) and in Bögels & Levinson 2023; in the latter, there is evidence of tongue preparation sometimes occurring not tied time-wise to delivery (suggesting inhibition till incoming turn is ending). So our model already had an early and late component - the news here is that the early component can be robustly discerned in naturalistic verbal interaction.

      Also, in the discussion, the Pickering & Garrod IAM model is favourably discussed, even though that model imagines the production machinery is used to predict the unfolding comprehension of the incoming turn: in that case, the production machinery would have to be doing two things at once - doing comprehension and planning one's own turn. In a way, the main lesson of the findings of the submission is precisely that this is implausible.

      In general, I have no hesitation recommending publication of an important contribution to the neuroscience of human performance in ecologically more valid settings.

      References:

      Bögels S, Levinson SC. Ultrasound measurements of interactive turn-taking in question-answer sequences: Articulatory preparation is delayed but not tied to the response. PLoS One. 2023 Jul 5;18(7):e0276470. doi: 10.1371/journal.pone.0276470. PMID: 37405982; PMCID: PMC10321606.

      Griffin, Z.M. and Bock, K. (2000) What the eyes say about speaking. Psychol. Sci. 4, 274-279.

      Levinson SC. Turn-taking in Human Communication--Origins and Implications for Language Processing. Trends Cogn Sci. 2016 Jan;20(1):6-14. doi: 10.1016/j.tics.2015.10.010. Epub 2015 Dec 1. PMID: 26651245.

      Levinson SC, Torreira F. Timing in turn-taking and its implications for processing models of language. Front Psychol. 2015 Jun 12;6:731. doi: 10.3389/fpsyg.2015.00731. PMID: 26124727; PMCID: PMC4464110.

      Roberts SG, Torreira F, Levinson SC. The effects of processing and sequence organization on the timing of turn taking: a corpus study. Front Psychol. 2015 May 13;6:509. doi: 10.3389/fpsyg.2015.00509. PMID: 26029125; PMCID: PMC4429583.

    1. “dead internet theory.” [c1]

      I have about this many times before and it really makes me wonder just how many interactions on the internet are just between bots. This makes a lot of things on the internet feel less credible and I feel can definitely lead to more distrust in things like the news or other things like announcements by companies.

    1. Hi there, very good work and I believe this work adds to the field. I have one question that is important. while processing the spike-ins bam files do you remove the duplicates or you treat them as the target? In Henikoff's suggested pipeline the duplicate reads are marked but never removed , due to the inherited Tag5 primers insertion. MNAse, the case for cutANDrun does work in a similar fashion and cuts in specific AT rich areas that are NOT random. Thank you for the preprint