2,118 Matching Annotations
  1. Last 7 days
    1. Know of high-impact research we should consider for evaluation — or other work this tool should be doing? Leave an email if you're open to follow-up discussion; if we later introduce compensation for useful contributions, earlier contributors will be grandfathered in.

      add a note here and elsewhere, along the lines of: We realize that you might have doubts about the value of adding content, ratings and comments on a largely AI-driven site. We'll try to personally respond to each substantial comment or engagement. And if you have doubts or want to flag us specifically, feel free to email contact@unjournal.org.

    1. Since the workshop: a company-reported production run at 22,000 litres and a new hydrolysate review bear directly on the scale and media questions. Dated updates and their limits.

      This comes across as strictly positive, but if you look at the evidence, I believe there was also some less positive news about finance.

    2. 9, 2026. These developments

      The developments reported below make it look like things have only gone in a positive direction, but I believe that on the finance side, things may have been scaled back? This might be adjusted a bit to not present the wrong impression.

    3. differed over how far particular substitutions could go,

      In a tooltip link direct quotes/sections - everything stated here needs some sort of attribution.

    1. Rate/discuss (Team) Rate/discuss (Public)

      Potentially, I could make these the same, just one box rather than two, but then people can fill in the "team options" within that interface.

    2. Unjournal commission an evaluation? (-- Strong No … ++ Strong Yes)For aggregation and sorting, the five quick ratings translate to percentile-equivalent scores: -- = 10, - = 30, ~ = 50, + = 70, and ++ = 90. These are deliberately spaced away from 0 and 100..

      make it clear that 1. these are translated to percentiles and 2. the "quick ratings" are less weighted/noted separately as 'quick ratings'

    1. visible CM_01 point estimates

      Link and tooltip explain what this question is, also give an in-text abbreviated meaningful name for this question

  2. Sep 2026
    1. The authors have made the atlas, methods, data, and replication code easy to inspect. We found some research indexing and light public discussion, but no major institution or policy report using the measures yet.

      this is what authors did. not quite the same -- also interesting but make that separate

    1. Several plausible routes could lower cultivated meat production costs.

      this first sentence seems a bit contentless. "Could" is vague and "lower" is not really quantified ... do we mean 'substantially lower' in some measured way? Bring closer to parity?

    2. Several plausible routes could lower cultivated meat production costs. The harder question is whether they can work together in a reliable process at commercial scale. Our current reading of the evidence favors testing complete, internally consistent production scenarios before treating lower ingredient prices or higher cell densities as a new cost forecast.

      This is an AI-generated (Astra, Extra-high effort) synthesis takeaway. Even in its initial statement, the claims should be sourced and linked. The "not this but that" aspect of this reads a bit like AI style. Quote: "Our current reading" -- it's not clear who "us" is. What would it mean to "test complete internally consistent production scenarios"? It's not clear what's being contrasted here - one wouldn't treat lower ingredient prices as a new cost forecast, for example, but they might enter into such a cost forecast.

      David Reinstien: To the extent this synthesis is making concrete statements. I'm not sure if I agree with them. At the very least, this should be stated in a much more contingent and tentative format.

    1. Framing note: Recent industry consolidation — including several company closures in 2024–2025 — is one set of data points, not a verdict on any particular TEA's methodology. Company outcomes depend on funding conditions, market timing, and management decisions that are largely independent of the underlying cost trajectory. The workshop aimed to evaluate the technical evidence on its merits.

      This should be a fold or a tooltip.

    2. Public working summary Workshop held May 8, 2026 Last reviewed July 20, 2026 S2 was off-record and is excluded. Its section below gives only the announced session scope.

      This isn't really a well-organized "public working summary." It goes into too much detail about the structure of the workshop rather than the overall narrative and a reasonable synthesis of the working results/takeaways.

    1. Informal Pre-Session Wed May 6, 2026 · 11:00am–12:00pm ET · Zoom · recorded Casual orientation for participants with limited May 8 availability — mostly preparation and broad discussion. Walkthrough of the beliefs form and cost model. David Manheim (Technion/ALTER) and Mirjam Capuder (University of Maribor) participated. It covered introductions, a walkthrough of the interactive cost model dashboard, and early framing questions about key modeling uncertainties — including Manheim's perspective on expert elicitation methodology and how to structure belief-updating around contested TEA assumptions. Substantive points from this session are incorporated where relevant in the S1 and S3 notes below.

      This is too much detail of the workshop's logistics here. Logistical stuff should be in folding boxes or tooltips.

    2. These map onto the workshop's structure: cell-line engineering → Crux 1; hydrolysate substitution → Crux 2; albumin and recombinant proteins → beliefs-form question E4; "what share of a hybrid product is cultivated?" → Crux 4.

      I don't see how this maps on at all - it seems like you're confusing things.

    3. Regulatory milestones are not directly comparable across jurisdictions. The FDA's February 2026 inventory lists five completed US cultured-cell food consultations, including three completed in 2025; market authorization, inspection, and product scope can still differ.

      I don't see that discussion in the evaluation 2-source it more directly and look for a quote.

    4. Evaluation 2 of this forecast

      That was not an evaluation of a specific forecast -- It was an evaluation of a paper involving an expert-driven forecasting exercise.

    5. rigorous

      "rigorous" is subjective. Think of a rephrasing here. Humbird was widely cited and influential. Pastika et al was published in Nature Food a peer reviewed journal that buy one metric ranks number 1 in it's subcategory and it has been heavily cited

    1. Tool updated: Aug 28, 2026Normal cadence: papers twice weekly; feedback hourlyPaper discovery, scoring, matching, and the full dashboard rebuild normally run on Monday and Thursday. New rating and suggestion submissions are checked up to hourly. The dated chips show the actual latest completed updates.Last paper scan: Aug 31, 2026Last AI scoring: Aug 31, 2026Crux/PQ match update: Jul 23, 2026Human feedback: 17 current ratings across 13 papersPublic anonymous aggregate refreshed Sep 1, 2026. Each rater's latest rating per paper counts once. Discussion and a display name are included only after explicit public-sharing consent; emails, rating keys, and submission ids are never included.Earlier API cost: at least ~$4.06 reconstructed

      less bold font here, this is taking up too much attention. It's good, but it's busy/distracting ... you can put all the updates dates in a more compact form, without so much bold. Actually, maybe fold these or put them in a tooltip ... just say "Last updated" and giv ehte date of the last paper scan, and then a tooltip with the other dates and the 'cadence' etc.

    1. A public economics column lists the review among its sources. This suggests some circulation outside the World Bank, although the article does not assess the review’s methods or health-policy conclusions in detail.

      it's given as a reference but I didn't see much discussion

    1. Mean absolute score movement?

      give a column for the overall MAD of scores to give a basis for cpmparison here ... to know if "6.3 is relatively large given how the scale is used' etc

    2. Agreement alpha?

      link a description and explanation of these measures in this context. What are normal and comparison values for this? What is the cardinal meaning of this measure ... 0.5 vs 0.25 eetc., and how to interpret it?

    3. rivate review found that one run treated prior public scrutiny as substantial while the other called it unclear, materially changing the estimated value of a new evaluation

      that's very interesting!

    4. Prior human-rated comparison papers were less stable on both dimensions: ordinary rerun movement averaged 8.5 points and the wording effect averaged 5.8 points. 2/6 (33%) had wording effects larger than ordinary rerun movement; wording-mean alpha was 0.69.

      this lower stability -- could it be somehow because of contamination from seeing human wordings, or because of greater variation in the nature of these papers?

    5. Recent AI-impact/governance papers: ordinary same-wording reruns differed by 4.1 points on average, while the repeat-averaged wording effect was 1.1 points. 0/6 (0%) had a wording effect larger than their ordinary rerun variation, and wording means produced no broad-tier changes. Prior human-rated comparison papers were less stable on both dimensions: ordinary rerun movement averaged 8.5 points and the wording effect averaged 5.8 points. 2/6 (33%) had wor

      worth discussing the pooled analysis too

    6. Recent AI-impact/governance papers: ordinary same-wording reruns differed by 4.1 points on average, while the repeat-averaged wording effect was 1.1 points. 0/6 (0%) had a wording effect larger than their ordinary rerun variation, and wording means produced no broad-tier changes.

      to me this suggests the wording variation was not particularly important -- is that reasonable? How does this compare to results in similar contexts?

    7. ordinary same-wording reruns differed by 4.1 points on average, while the repeat-averaged wording effect was 1.1 points.

      4.1 percentage points on the 0-100 scale? How does that compare to the typical rating variation between such papers -- drawing the reference group in a few reasonable ways

    8. Held-out check: a third baseline execution is kept outside the balanced headline analysis and used only as a robustness check.

      explain (tooltip?) why 'holdout' is important here. I'm not sure I understand what it's doing in this context

    9. What is being tested: every stability number below describes AI model scores. “Previously human-rated papers” names a paper cohort selected because earlier Unjournal ratings are available; those ratings do not enter the stability calculations. Human and AI scores are compared separately in the amber reference panel.

      make this a folding box (beyond the first line)

    10. paper-level human ratings.

      would be fine to do so -- why not? These are already public IIRC. I'm also not worried about sharing AI-generated ratings. Remember, we're rating for priorization/impact, not research quality

    11. Average the two runs within each reviewed, equivalent wording, then compare those wording means. This reduces ordinary rerun noise in the wording comparison.

      I got this upon re-read but was confused at first. Explain better. And is this the approach recommended in the apepr, or is there a better decomposition

    12. Run each wording twice with the paper, rubric, model, settings, and output schema fixed. Differences within a wording estimate ordinary run-to-run variation.

      Need some context here -- which models (Fable, Sol, etc) and effort levels are these being tested on?

    13. prompt variants

      i want to see the nature of these prompt variants early on, to get a sense of things. Show some examples concisely in a tooltip or fold

  3. Aug 2026
    1. A shared first pass: Browse research that may inform policy, funding, or further evaluation. Scores rank the value of attention or evaluation; they are not quality grades.

      add: metrics of mention on social media, citations in white papers, etc., ?altmetrics type metrics

    1. Living reviewFilter to papers surfaced by one or more expert-maintained living literature reviews. A paper keeps this provenance even if it originally entered through another academic source. All Reviews ▾ AI Accountability ReviewBridging BoundariesExistential CrunchLauren Policy: Migration Literature ReviewMonitoring Gene DrivesNew Things Under the SunRegression to the MeatReskilledScaling in Human SocietiesThe Care GapThe Patentist Review / UJ overlapFind review posts that cite papers The Unjournal has already evaluated, has in its pipeline, or previously considered. This is intended as an engagement and follow-up queue. All Any UJ overlap UJ evaluated In pipeline / considered Targeted intakeRequested, topic-focused additions outside the broad recurring scan. Use this to include all records, exclude deliberately oversampled batches, or inspect only one targeted run. Include allExclude targeted intakeOnly targeted intakeOnly: AI impacts on global health and development (29)Only: Animal-welfare focused intake (19)Only: Empirical conflict replication-game intake (14)Only: Large-N GCR and existential-risk quantitative evidence (6)Soil-invertebrate crux sweep (6 cruxes) Crux/PQFilter to papers that map to one or more Unjournal Pivotal Questions or community cruxes (from the cruxes explorer). "PQ match only" restricts to papers mapped to a Pivotal Question. All Has crux/PQ match PQ match only RecencyFilter papers by how recently they were released or last updated. Useful for focusing on the newest research that may benefit most from timely evaluation. Any time Last 30 days Last 3 months Last 6 months Last year Last 2 years Last 5 years ModelThe AI model used to score each paper. New papers receive a standard GPT-5.5 assessment; stronger candidates receive a second, more thorough GPT-5.5 assessment with higher reasoning effort. "UJ historical" papers were prioritized by the Unjournal team (not AI). Papers marked WAIT are temporarily excluded from AI ranking pending a genuine model reassessment. All Models ▾ UJ historicalawaiting genuine model assessmentclaude-sonnet-4-6gpt-5.4gpt-5.4-minigpt-5.5 (codex headless, high)gpt-5.5 (codex headless, medium)gpt-5.5 (codex headless, xhigh)manual_public_followupopus (headless) Also show small-model scored (mini/haiku) Also show UJ historical Also show legal scholarship Targeted intake catalog · 74 records · 5 search families

      this is too cluttered -- it should be a filter

    1. Source text (unverified type) source noteText supplied by the discovery source; its status as a formal abstract has not been verified.

      the paper has an abstract -- why isn't it being shown here?

    2. AI-assisted prioritization for The Unjournal evaluation

      Make it clear here very prominently (with some reminders in key places below) that these are not ratings about research quality -- they consider the potential direct or near-direct impact of the work from a global priorities perspective.

    3. he AI premium is concentrated in frontier and intensive AI use, including closed-source models, paying or seasoned users, and long prompts, rather than casual or open-weight use.

      How is this identified? Are they talking about 'firms whose returns correlate more with the aggregate use of these prompts in the AI consumption data"?

    4. interaction-and-communication skill content strongly associated with higher AI exposure, while analytical, scientific, and operations-control skills load more negatively.

      this seems particularly interesting -- does this havean implication for the labor market?

    5. predicts

      is 'predicts' the claim of interest? This is all backwards looking 'prediction' in a sense. is there an implied claim about future returns?

    6. The paper informs live decisions about AI labor-market adaptation, AI industrial policy, competition and market-concentration monitoring, financial-risk exposure to AI shocks, and prioritization of retraining or adjustment support across occupations. It i

      in what ways? Need more examples.

    7. This is a high-value Unjournal candidate: it is an NBER working paper using unusually granular proprietary data on realized AI consumption to estimate how AI demand is priced across firms, sectors, countries, and occupations. The findings could inform decisions

      "inform" in what ways? We want some examples

    1. r sufficiently strong, discussed public feedback—places the paper in a private SQLite-to-Coda queu

      tooltip the exact rule and try to keep it updated

    1. Quick-rate mode: for each paper, how strongly should The Unjournal commission an evaluation? (-- Strong No … ++ Strong Yes). Each click is saved to The Unjournal's server. Clicking another option for the same paper updates this browser's rating rather than adding another. Add your name or email once below; the detailed form will reuse it.

      One thing about the quick rate form - I don't want to encourage people to rate the paper simply based on its title. I'd like them to at least read the abstract and probably more. At the moment, we're bundling the detailed rating with seeing details on the paper - we should potentially adjust that or put some signposts so people don't get that impression.

    2. More filters and source controls DisciplineThe academic field or sub-discipline of the paper, as classified by the AI model. Used to filter by research methodology and domain expertise. All Fields ▾ AI & Data ScienceAnimal WelfareDevelopment & AgriculturalEconomicsEnvironmental & ClimateLabor, Education & HealthOtherPhilosophy & EthicsPolicy & GovernancePolitical Science & LawPsychology & BehavioralStats, Finance & Methods SourceHow the paper entered this dashboard. Source is a discovery route, not a publication venue, endorsement, or quality judgment.Unjournal database: imported from The Unjournal's existing evaluation and prioritization workflow.Targeted public-paper follow-up: an independently public paper or DOI found during a requested topic-focused search. It does not mean the author supplied or endorsed the score.Academic feeds: NBER, CEPR, and RePEc working-paper feeds; arXiv and SSRN preprints; OpenAlex and Semantic Scholar indexes.EA Forum: research linked from EA Forum posts.Research organizations: Anthropic, DeepMind, and selected AI governance or safety organizations.Legal sources: legal-scholarship searches covering OpenAlex Law, law reviews, and the Institute for Law & AI. All Sources ▾ AI Governance (arXiv)AI Safety OrgsAnthropic ResearchDeepMind ResearchEA ForumNBEROpenAlexRePEcSSRNSemantic ScholarTargeted OpenAlex intakeTargeted public-paper follow-upUnjournal databasearXiv Targeted intakeRequested, topic-focused additions outside the broad recurring scan. Use this to include all records, exclude deliberately oversampled batches, or inspect only one targeted run. Include allExclude targeted intakeOnly targeted intakeOnly: AI, global health, and development curation (3)Only: Animal-welfare focused intake (19)Only: Empirical conflict replication-game intake (14)Only: Large-N GCR and existential-risk quantitative evidence (6)Only: Transformative AI, global health, and wellbeing (16)Soil-invertebrate crux sweep (6 cruxes) Crux/PQFilter to papers that map to one or more Unjournal Pivotal Questions or community cruxes (from the cruxes explorer). "PQ match only" restricts to papers mapped to a Pivotal Question. All Has crux/PQ match PQ match only RecencyFilter papers by how recently they were released or last updated. Useful for focusing on the newest research that may benefit most from timely evaluation. Any time Last 30 days Last 3 months Last 6 months Last year Last 2 years Last 5 years ModelThe AI model used to score each paper. Papers are scored in tiers: a fast model (gpt-5.4-mini) screens all candidates, a stronger model (gpt-5.4) re-scores the top papers, and the most capable model (gpt-5.4-pro) provides detailed analysis of the highest-ranked papers. "UJ historical" papers were prioritized by the Unjournal team (not AI). All Models ▾ UJ historicalclaude-sonnet-4-6escalatedgpt-5.4gpt-5.4-minigpt-5.5gpt-5.5 (codex headless, high)gpt-5.5 (codex headless, medium)gpt-5.5 (codex headless, xhigh)manual_public_followupopus (headless) Also show small-model scored (mini/haiku) Also show UJ historical Also show legal scholarship Targeted intake and prioritization requests - 64 records across 5 paper search families and 1 crux search family This shared catalog records topic-focused searches outside the broad recurring scans. Initial and follow-up passes are combined below as one search family. A targeted label records how a record was found; it is not an endorsement, quality judgment, prevalence estimate, or completed Unjournal prioritization decision.Large-N GCR and existential-risk quantitative evidenceResearch papers (6 papers; 2026-07-30) Credible large-N, cross-country, historical, forecasting, survey, or other quantitative evidence directly bearing on global catastrophic or existential risks. Public OpenAlex search, quantitative-design triage, deduplication, and prioritization scoring. This is a separate follow-up to the empirical-conflict search. It asks whether global catastrophic and existential-risk questions have credible quantitative literatures with enough observations or repeated judgments to support evaluation or replication. It excludes primarily conceptual argument, technical engineering, and ordinary civil-conflict studies without a direct catastrophic-risk connection. Inclusion is not an endorsement or a completed Unjournal team decision. View 6 records →Transformative AI, global health, and wellbeingResearch papers (16 papers; 2026-07-24) Economic, social-science, and policy research relevant to global health and wellbeing under transformative AI, especially implications for lower-income countries. Search-pass details (2)Transformative AI, global health, and wellbeing follow-up: Second-wave public-paper search focused on first-run gaps, followed by deduplication and Codex subscription scoring. This second wave follows up the initial requested intake and deliberately targets gaps it left: public finance and tax-base resilience, distribution and social protection, AI market power, worker attitudes, and developing-economy structural transformation. It uses only independently public paper or institutional publication pages. Inclusion is not an endorsement or a completed Unjournal team decision.Transformative AI, global health, and wellbeing: Public OpenAlex search, theme-fit triage, deduplication, and Codex subscription scoring. The search follows the linked Coefficient Giving request for proposals but is an independent Unjournal intake exercise. It emphasizes non-catastrophic transformative-AI scenarios, labor markets, fiscal capacity, social protection, biomedical and health-system bottlenecks, development strategy, and aid allocation. Inclusion does not imply endorsement by Coefficient Giving or an Unjournal team decision. Coefficient Giving request for proposals. View 16 records →

      this fold should also mention the 'tartgeted intake' filter above

    1. A steelman against the Bay-Area assumption that very large AI-lab-adjacent philanthropy will quickly become usable funding for AI safety and EA cause areas.

      I don't think it still really should be considered a "steelman". It's more of a model and a dashboard/Calculator, although the initial prompts did emphasize this steel manning motivation

    1. The reasoning goes that if there is always a high level of background risk to humanity, then we should expect to go extinct soon anyway, which means the importance of avoiding any one particular risk is not as valuable as it may seem. For more details see the full report here.

      This seems rather intuitive to me, but it's asking a slightly different question than what the original phrasing might seem to imply.

      I think the initial intuition that more risk means more value of reducing risk, comes from the natural idea that effort spent reducing a particular risk will reduce that risk proportionally. So, spending effort on reducing risks from car crashes, malaria in Africa, or heart disease, all else equal, we yield more value than spending comparable effort on reducing the risks of bear attacks. I guess this is the "importance" part of the ITN paradigm.

      But of course, the benefit of reducing the risk of car crashes is lower if we are facing other impending doom. Let's say we see an asteroid coming toward the Earth, or the threat of incoming nuclear war is high.

  4. Jul 2026
    1. DisciplineThe academic field or sub-discipline of the paper, as classified by the AI model. Used to filter by research methodology and domain expertise. All Fields ▾ AI & Data ScienceAnimal WelfareDevelopment & AgriculturalEconomicsEnvironmental & ClimateLabor, Education & HealthOtherPhilosophy & EthicsPolicy & GovernancePolitical Science & LawPsychology & BehavioralStats, Finance & Methods SourceHow the paper entered this dashboard. Source is a discovery route, not a publication venue, endorsement, or quality judgment.Unjournal database: imported from The Unjournal's existing evaluation and prioritization workflow.Targeted public-paper follow-up: an independently public paper or DOI found during a requested topic-focused search. It does not mean the author supplied or endorsed the score.Academic feeds: NBER, CEPR, and RePEc working-paper feeds; arXiv and SSRN preprints; OpenAlex and Semantic Scholar indexes.EA Forum: research linked from EA Forum posts.Research organizations: Anthropic, DeepMind, and selected AI governance or safety organizations.Legal sources: legal-scholarship searches covering OpenAlex Law, law reviews, and the Institute for Law & AI. All Sources ▾ AI Governance (arXiv)AI Safety OrgsANIMAL_LAWANIMAL_LAW_REVIEWAnthropic ResearchDeepMind ResearchEA ForumLaw & AI InstituteLaw reviewNBERNEP_LAWOpenAlexOpenAlex LawRePEcSSRNSSRN_LAWSemantic ScholarTargeted OpenAlex intakeTargeted public-paper follow-upUJ seedUnjournal databasearXiv Targeted intakeRequested, topic-focused additions outside the broad recurring scan. Use this to include all records, exclude deliberately oversampled batches, or inspect only one targeted run. Include allExclude targeted intakeOnly targeted intakeOnly: AI, global health, and development curation (3)Only: Animal-welfare focused intake (19)Only: Empirical conflict replication-game intake (14)Only: Large-N GCR and existential-risk quantitative evidence (6)Only: Transformative AI, global health, and wellbeing (16)Soil-invertebrate crux sweep (6 cruxes) Crux/PQFilter to papers that map to one or more Unjournal Pivotal Questions or community cruxes (from the cruxes explorer). "PQ match only" restricts to papers mapped to a Pivotal Question. All Has crux/PQ match PQ match only RecencyFilter papers by how recently they were released or last updated. Useful for focusing on the newest research that may benefit most from timely evaluation. Any time Last 30 days Last 3 months Last 6 months Last year Last 2 years Last 5 years

      make it easier/more prominent to 'clear all filters'

    1. 5. Three problems that a better estimate of the exchange rate won't fix

      wait -- I don't think that's logically correct. If indeed we have a good estimate of the exchange rate, isn't it, by default, embodying that these problems are in some sense solved?

    2. 3. What this is based on

      second column in table below should be narrower, firrst one wider. Use the space more carefully (and make that a persistent pattern/skill)

    3. Preliminary — do not read Unreviewed working draft, 30 July 2026. Not for circulation, citation, or quotation.

      "do not read" is too strong. Make more caveats instead ... it's a current working space, being continually adjusted. And it has not yet been reviewed by workshop participants or research contractors (something we aim to do soon)

    4. 2. Bottom line

      we need more caveats on this. It has not yet been reviewed by workshop participants or research contractors (something we aim to do soon)

    1. Optimized messaging shifts choices towards plant-based foodsource titleExact title resolved from the linked paper's bibliographic metadata; discovery-page anchor text and model output are not used as titles.

      Where is the abstract here? I don't see it.

    2. Strong near-term candidate for animal-welfare evaluation because it tests behaviorally relevant messaging around plant-b...

      What's this text doing here? I don't see it at the top in the other entries.

    3. 2 ratings

      Is there a way that someone could change their rating? It seems like every time I come to the page and rate it, it records separately.

    4. abstract provenanceSource-supplied abstract wording; HTML entities and whitespace normalized. Not independently compared with the paper PDF.

      This tooltip seems rather process-oriented and not particularly helpful for users of this page. You certainly don't need to give it every time, and I'm not quite sure what it means or how to interpret

    5. Rate on a 0–100 percentile scale relative to all papers in our database (Strong Yes=top quintile, Strong No=bottom). Try to form your opinion before reading the AI discussion above. Discuss why you gave that rating.

      This percentile scale instruction is not consistent with the negative-neutral-plus scale shown below. Let them give a percentile first, and if they don't want to do that, you can let them give the priority ratings with the negative-positive scale, or just save that for the quick raters.

    6. 2 ratings

      Seems like it was double counted here. I rated it with the quick rate and then with the detailed rating. It doesn't seem quite right.

    1. displayed submittable slider defaults, so some responses may be anchored or untouched.

      Clarify this - the language here doesn't make it clear what you're talking about, and can't we just skip the ones where they did not move the slider?

    2. Direct validation of the linear WELLBY against better-specified alternatives — per Benjamin, not yet done.

      How would this be conducted - give at least a hint, perhaps in a tooltip?

    3. What would change this answer

      "The answer" given is diffuse, hedged, and multipart. Which part of the answer would change because of these things?

    4. Do if you have budget or control an instrument

      This discussion seems to ignore another point we have been covering in this conversation, namely the potential failure of cardinality of this scale. Shouldn't this also be mentioned?

    5. Flag explicitly when a conclusion depends on the location of the neutral point.

      This seems great, but it comes with not much explanation. This isn't well joined up.

    6. Use directly measured life satisfaction wherever a trial reports it, in preference to a converted mental-health measure.

      I'm missing a justification of this. Was this justified above, or is it a recommendation coming directly from one of the authors/participants?

    7. nd publish sensitivity bounds

      Give a little more hint as to how someone would publish sensitivity bounds or perhaps a link we're referencing explaining how to do this.

    8. 1. The decision this is about

      This needs more background. You give a brief explanation of the relevance of the question, although more links would be helpful, but you haven't explained the context in a way that makes this page work by itself. What is this individual page meaning to do? How is it based on evidence coming out of our Pivotal question our evaluation and our workshop?

    9. Ratios move even when coefficients do not

      That's a confusing overstatement. For the ratios to move, the coefficients also have to move. I believe it's just that they're much more sensitive.

    10. instruments

      What do you mean by "instruments" here? I guess you're talking about the set of questions used to measure well-being, etc., such as the Cantrell ladder or a set of depression and anxiety questions? But I really don't understand the implication here. Yes, I see they overlap and they are not completely nested, but what's the relevance of that? You need to complete the argument, even if it means using more space, which you can be judicious with, using tool tips and folds.

    11. SDs inherit the population you measured them in A standard deviation is a property of a sample, not of a person.

      I don't completely understand what the implication of "a property of a sample, not a person" is. What's the relevance of that?

    12. Mechanically, comparisons will tend to favour interventions run in more heterogeneous populations.

      I don't see this, so obviously. Will these interventions not be more noisy in such populations, in line with the greater heterogeneity? But also, I don't even understand directionally what you're talking about here. If the population is homogenous, then there will be less underlying variation. If some intervention has a sort of unit-level effect, it will be recorded as being more impactful in more homogenous populations, not more heterogeneous populations. Consider this more carefully. Explain it better, including possible tool tips and folds for making the explanation complete, and also source this to and link a page or paper that makes this point explicitly.

    13. 2. Bottom line

      You can, to the extent that this is indeed justified, keep this formulation without putting in all the references I mentioned below, as long as you give a major caveat and signpost that these are explained in sections below - which I wouldput hyperlinks to.

    14. income-doubling to lives-saved, lives-saved to DALYs — then backing out the remaining si

      I don't completely see how this relates to the SD-SD method?

    15. 3. What this is based on

      Good to know what your sources are for this, but it's not linked narrowly enough that I can understand what is attributable to what: - what's attributable to evidence - what's attributable to careful arguments - what's only one person's claim, etc.

    16. cheap fixes are known

      That's a bold claim. I'm not sure how "cheap" these would be in practice. I think you should hedge this claim and provide evidence for it.

    17. ny single global conversion factor therefore systematically mis-weights domains

      I don't see how the "therefore" follows here. This is underexplained.

    18. Disability weights top out around 0.3, while depression and anxiety show losses above a full life-satisfaction point.

      Same comments here - you need to give sources for everything, but also I don't quite understand what it is you're talking about here. Ok, it's kind of understandable, but it needs a little bit better explanation. What are the "conditions" we're talking about, etc.?

    19. The likely error is directional, not symmetric. Compression and ceiling effects both push toward overstating the wellbeing gains from mental-health interventions relative to physical ones. A conversion factor that is wrong in a known direction should not be treated as noise.

      What's the epistemic basis for this - is this your own thinking? Is this something that was explicitly claimed in the workshop, reflecting a peer-reviewed paper etc? Again, we need to reference these things.

    20. han a near-1:1 point estimate implies

      You're not providing context here. Where is this near one-to-one point estimate that you're talking about coming from? (I know, but you should reference it and link it, etc., perhaps with tool tips with embedded hyperlinks, and.or folding boxes

    21. The SD-to-SD step is defensible as a working assumption, and it is nobody's preferred method — including its originators'. Keep using it; report it as an assumption; run sensitivity analysis. Do not describe it as established.

      This is a recommendation without any evidence presented. It's not reasoned transparently. You need to justify this or explain that it will be justified below and link that.

    22. The question funders ask us is not "is subjective wellbeing valid." It is: can I keep using this step, and how wrong might it make my ranking?

      This is your AI. Not this, but that slop. Make it more like my own language and avoid using this bold all the time.

    23. Founders Pledge's moral weights use it. Happier Lives Institute's cost-effectiveness work generates it. Anyone reading either inherits it.

      Link citations/references to this, please.

    24. one standard deviation of improvement on a depression or anxiety instrument as one standard deviation of improvement in life satisfaction

      avoid this excess bolding

    1. Identify the fragile assumptionIn your preferred bottom-up TEA, which single assumption would break the headline result if it were wrong by one order of magnitude? Which interactions …

      where did this (and the other questions ) come from? Tooltip the derivation, of course w/o identifying anyonw who wanted anonymity

    1. The PDF output is a draft field packet, not a filled official IRS PDF.

      highish value if do-able ... but rather good if it jus tprovides instrutions on exactly what to imput

    2. QEF, mark-to-market, purging elections, foreign-tax-credit allocation, and notional accumulation-unit determinations are not implemented.

      let's implement those -- these seem high value

    1. Job-matching and recommendation systems affect labor-market frictions, unemployment duration, and match quality, all of which have major welfare and policy consequences.

      But isn't this mostly relevant to rich country job markets? Is it reflective of an AI-impacted job market? Are recommended systems responsible for a large share of the job-match and productivity value?

    1. Human ratings visible here:

      should we/can we reeasonably incoprporate Unjournal prioritization ratings for this haere? Need to check with the team to understand the terms under which those were given/shared. NB David Reinstein gives full permission to share/report any of his prioritization ratings, but would want to consider before releasing his discussion content in Coda.

    1. Adjust these sliders to create your own priority score, then sort by Custom weights. This changes only your browser view; it does not change The Unjournal scorer.

      are these reflecting global social value in the way we usually prefer?

    1. The useful question now is not whether the old paper should be revived unchanged. It is what, if anything, is worth testing under current technology and policy.

      another annoying AI dichotomy. Just list the thing that IS important

    2. gument in an Essex Economics discussion paper and an unfinished theory project

      That paper was related to the project, but it doesn't matter that it was "Essex economics", just say "discussion paper".

    3. The strongest warning comes from take-up: Ofcom reports 532,000 UK social-tariff customers in June 2025, only 8.6% of a conservative proxy for eligible households.

      what does this indicate? I don's see it as evidence of a lack of a welfare gain. Perhaps a friction (~transactions cost) or a stigma issue?

    4. so "free redistribution" is too strong. The more defensible claim is narrower: this could create additional purchasing power for some low-income consumers with less public expe

      this is AI-speak ... the not this buut that juxtaposition

    5. lower pric

      Maybe this missses the usual 'what's in it for the retailer'? and the answer is 'it helps them price discriminate by the (likely) single most indicative measure of willingness to pay -- the income (or adjusted income', helping them increase their profits

    1. Calibration anchors are a small set of real example papers that teach The Unjournal's AI prioritization scorer what “value of evaluation” looks like in each area — including the boundary cases where a good paper is not a good candidate. Each anchor's rating is the team's real prioritization rating; the proposed lesson is an AI inference about the calibration takeaway. Use this page to confirm or correct each rating and lesson, suggest new anchors, and track the discussion.

      This anchoring page should allow people to specify the different dimensions of quality, that is, the different ratings.

      And the form should be more quantitative rather than "too high, too low." -- or maybe I'm missing the point?

      To be honest, I'm not sure what was intended by the question "not by the answer", "not a good anchor". ... What is the " lesson" here? That's confusing.

    1. Research training

      We should label 'what kind of research training' here. I'm not sure the right term ... social science, quantitative modeling of cost/benefit forecasting, etc., economics, statistics, etc.

    2. s a possible consolidation of things I already do at a smaller scale: modeling workshops on contested quantitative questions, Fermi-estimation and parameter-elicitation sessions, and supervising early-career research

      The Fermi estimation session thing is a bit of an overstatement -- something we're considering running soon

    1. How to read each column: the conventional product's retail price; the cultivated cost the model delivers to retail (your biomass slider + scaffold for cuts + markup); their ratio R; and the estimated share of that product's own market cultivated would capture (not a share of all meat). Share-bar colour: green > 30%, amber 8–30%, red < 8%.

      give some column headers instead

    2. One caveat to this species-by-species framing (our own intuition, not in his model): especially early on, a cultivated “chicken” or “shrimp” product may be received as its own distinct food rather than competing head-to-head only with the conventional product it imitates, so cross-category substitution could matter more than a per-species contest implies.

      "my" not "our" -- and make this a tootip

    3. On agreements: we draw on the same source literature and read it similarly. Humbird’s 2021 pessimism was driven mostly by amino-acid/media cost, which is not a hard thermodynamic constraint and which Pasitka’s 2024 empirical work (hydrolysate at $0.63/L) pushed down sharply. Both of us read this as suggesting cultivated meat likely lands a few-fold above conventional meat — not at parity, but not orders of magnitude off either.

      I still don't want to suggest that I am 'reading this' in a way suggesting a particular conclusion!

    4. Share colour: green > 30%, amber 8–30%, red < 8%. Ordering across species is driven almost entirely by the price ratio R (cost fixed, conventional price varies), plus the per-tier authenticity offset — reproducing Pablo's inversion: cheap chicken and pork resist, expensive beef, seafood, and luxury foie gras are penetrable.

      label this better -- what actually are the columns/outcomes here? And maybe put in whole shrimp and shrimp paste too

    5. may capture little chicken or pork share

      my intuitions -- at least in early stages, even if they see it as 'real meat' the chicken, fish, etc. imitating products may still be seen as distinct and not compete only with the product theory imitate

    6. His is the piece we have explicitly not built.

      --> His work could be seen as a 'missing piece' for our analysis.

      ['the piece' suggests there is only one missing piece']

    7. Our model

      Explain more carefully ... we are not just a 'model' but also a calculator allowing you to provide different assumptions, and a template/example for future modelers

    8. Takes cost as exogenous. Cost → retail price ratio → discrete-choice (logit) market share by species, product tier, and geography → diffusion over time.

      make wider, use the space generated