512 Matching Annotations
  1. Last 7 days
    1. AI’s economic impacts and governance are priorities for The Unjournal. We’re looking for research where independent scrutiny could inform decisions about economic growth, jobs, development, power, and policy. This collection includes both economic evidence and governance proposals.

      This is a pretty bland explanation. We should make a better case for why this is a priority. You can consider our existing language and discussion in doing this. At the moment, you open the "why" box, and it doesn't really tell you much of why, just some general stuff.

    2. METR's time horizons (with its 2026 update) if capability measurement counts as governance evidence; otherwise the liability set (Weil, Trout, Gans on staged access). → METR time horizons, Overcoming Judgment-Proofness (Weil), When Does Regulation by Insurance Work? (Trout), Staged Access and Liability (Gans)

      you are proposing these 'paired evaluations' -- but I don't think we have a sstem or structure for that

    3. Worth asking for code Survey and elicitation data; FRI has released data from earlier tournaments, so a request is reasonable.

      bold yellow background is too bright

    4. You made 8 Keep/Maybe/Drop picks or notes here before they were folded into the rating scale. 2 of those papers have no rating yet; the rating panel shows your earlier pick and note so you can convert them.

      Does everyone see this, or is this just me-specific?

    5. October 2026. We are choosing a handful of these papers to commission for public evaluation; nothing is decided yet. Ratings, picks and comments here feed that choice.

      Mentioning that we're aiming to decide and move forward before the latter part of October

    6. Why this area. AI governance is one of The Unjournal's stated focus areas, but we have published only one evaluation in it so far. Policy is moving fast (pacing proposals, export controls, the EU AI Act, state liability law), and the research feeding these debates is influential yet rarely gets independent expert scrutiny. We focus on governance and the economics of AI policy, not technical AI safety.

      This is getting a bit text-heavy. Use folds and tooltips more. ... Remove the "published only one evaluation in it so far." That's a bit negative and not entirely accurate if we think broadly

    1. Why this page exists. We are considering whether an independent evaluation of this paper would be useful. The synthesis score uses provisional weights, and the human sample is small. It is one input to the decision.

      this takes up a bit too much space, and sounds a bit too AI in it's 'not this but that, this is only one input'. Improve it and put more into a tooltip

    1. Latest completed updates Paper scan: Oct 5, 2026 AI scoring: Oct 5, 2026 Crux/PQ matches: Oct 5, 2026 Public feedback aggregate: Oct 5, 2026 Normal cadence Paper discovery, scoring, matching, and dashboard rebuilds normally run Monday and Thursday. Feedback aggregates refresh separately from paper discovery; see the feedback timestamp above. Feedback 25 current ratings across 20 papers. Each person's latest rating per paper counts once; discussion and display names appear only with explicit public-sharing consent. Data context 57 evaluated papers and 129 in the Unjournal pipeline. Earlier reconstructed API cost: at least ~$4.06. Search:

      Make it easier to clear all the filters.

    1. Is the crux description wrong? A better question phrasing or PQ mapping? Tell us here — it saves directly to The Unjournal, no GitHub needed.

      Add box and adjust the interface a bit -- encourage people to provide feedback along with their rating of the crux, not just whether it was described incorrectly.

    2. Dark mode Reset filters Share URL Quick-rate: ON + Suggest a crux Download CSV Filters (1) Reset Showing 8 of 281 cruxes Quick-rate mode on. On any row, click ▼▼ (strong downvote) · ▼ · ▲ · ▲▲ (strong upvote) to rate how valuable / decision-relevant the crux is. Votes save instantly (anonymous, no account).

      Explain the "rate the croaks" bit - and make that a bit more prominent, including the "quick rate" part.

    3. Coverage by cause area — click to filter · amber = legacy AI cluster · green = Unjournal core & in-scope AI

      Put the table below. Make it clearer that you can click on each column to sort by that column. make this general across pages like this.

      "Date" is this the date the paper/object was last made public, or updated?

      If possible, use tooltips to explain what each column means and how it was generated. For example Op/.Clarity it's probably something about the clarity of operationalization, but most people won't understand this.

    4. needs review research paper targeted run: AI governance and the economics of AI policy: priority papers

      The yellow color of some of these boxes takes a bit too much attention, and it also overlaps with the color used by the hypothesis tool. Use a more muted color or just gray.

    5. Brookings

      Use the column space a bit better in this table. In particular, the second column seems to be a bit too wide given the size of the content. Also, I know that the label "targeted: AI governance blah blah blah" is very long, but that could be split across two lines.

  2. Oct 2026
    1. and it overlaps the FRI paper. Note the lab authorship.

      overlaps it in what way? I thought that was more about forecswters, and this seems more like a theory model?

    2. Where the paper sits in the working list: shortlist, getting attention now, alternates, found on the dashboard, or older papers kept because they are still in use. Attention

      make it easier to click these to sort by them?. And make 'attention' sort hte default.

    3. Candidate papers on AI governance and the economics of AI policy. Working notes, not Unjournal decisions or evaluations.

      Give a little more context here. Explain why we're focusing on this and what stage this is at. Make it clear how this relates to the prioritization database too (without cluttering -- the latter cound be a tooltip)

    1. by setting concrete safety requirements for continued development and deployment, setting a speed limit on capabilities growth, monitoring high-stakes AI R&D experiments and building the option to shut them down if needed, requiring that AI R&D take place in secure and isolated environments (e.g.,air-gapped networks), establishing international incident-sharing, clarifying how deterrence would apply around an intelligence explosion so as to avoid unexpected escalation, securing international agreements to pace progress if needed without fear of falling behind, and developing verification tools in advance for such agreement

      They're listing a variety of tools the government or the labs could take on. I say it should close quote. They're suggesting that these will have measurable impacts on reducing the catastrophic risks from AI . They give quite a laundry list here. Evaluators could consider whether these ideas indeed do have strong potential in terms of both feasibility and likely impact. However, given that it's such a laundry list, it makes it hard to do a the intensive expert evaluations were the thing that we considered

    2. Preliminary evidence suggests that this feedback loop could radically accelerate AI progress, overcoming frictions such as diminishing returns to research labor and hard-to-automate tasks.

      Here's another definitive claim that could be investigated, at least if the paper here is presenting or making claims about this evidence. However, the word "could" softens this in a sort of vague way.

    3. ome experts think that AI R&D could be fully automated within the next few years.

      "some experts think" - is that a pretty vague statement? Are the authors making a claim about their own beliefs or the evidence about this ?

    4. Anthropic reported AI systems completing 26% of internal AI R&D work with only high-level supervision, compared to just 1% five months earlier.

      Here's one substantive claim - is this paper, presumably, standing by it. That's something that could be brought under some scrutiny as to whether those numbers are actually meaningful AND CORRECT

    1. Jazz on the Green: Annika SkooghCrediton Arts Centre · CreditonDoors listed at 18:30. Jazz with Sri Lankan food; check the organiser for the set time.

      add a link so people can put it on their calendar

    2. Jazz, within reach.

      Give this a different name. Try to use my voice more. This looks kind of PR/AI … Just "Live jazz in the Exeter area" is fine. Mention me as the creator a little more prominently and link some of my stuff, including my YouTube and my band listing

  3. Sep 2026
    1. NoteLearn: Understanding Uncertainty & Distributions

      I think this is too much content for a fold. Turn it into a tooltip as well as a link to a different page that has this fleshed out in detail - maybe that page already exists and it's just a matter of linking it

    2. NoteNew to this model? Start with the Simplest Model → — a shorter version focusing on some key levers with line-of-sight explanations. You can carry your settings over to this Advanced Model when you’re ready.

      This one is short. It could probably just be a tooltip

    3. NoteModel version and September 17 update This page runs the current model (engine 2026-09-17.1). See the September 2026 review and correction log. For comparisons with beliefs recorded at the May 8 workshop, use the hosted workshop-era model or its tagged source. The archived model is preserved as it was deployed and includes errors corrected here. Two supplied research critiques prompted optional median-preserving priors, explicit growth-factor price ranges, and structural comparisons for media use, GF dosage, and maturity dependence. Baseline numerical assumptions are retained. New controls include provenance tooltips and links to the discussion responses, justification, and remaining questions. These additions are AI-implemented scenario tests, pending expert review. NoteNew to this model? Start with the Simplest Model → — a shorter version focusing on some key levers with line-of-sight explanations. You can carry your settings over to this Advanced Model when you’re ready. ►Audio overviews(AI-generated · note)How Cultured Meat is Made ~11 minDownload MP3View/annotate scriptThe Cost Model Explained ~10 minDownload MP3View/annotate script WarningImportant: Model Status & Limitations This model is largely AI-generated and, as of September 2026, has not been independently validated. It is provided to fix ideas, make assumptions inspectable, and support discussion and comparison. Do not treat the outputs as authoritative cost estimates. Several limitations are decision-relevant, not cosmetic: the single latent-maturity factor is ad hoc; source dollars are not yet normalized to one real-dollar year; wet-biomass water content is not standardized; and the model does not separate the probability of reaching commercial scale from cost conditional on reaching it. The sensitivity chart is a dollar-swing ranking, not a variance decomposition. We are addressing these incrementally. Read the full critique and our responses → Limits/Critique We especially welcome expert review of: growth factor quantities and prices, bioreactor CAPEX ranges, and the maturity correlation structure — see below and the Sources section. NoteFrom workshop evidence to model revision The May 8, 2026 workshop is complete. Its outputs now form a linked evidence-to-model workflow: the summary records the technical arguments, the beliefs analysis shows the forecast spread, the demand bridge connects factory-gate cost to an adoption model, and the proposed modeling hack turns the remaining disagreements into testable model changes. Workshop summary → · Public beliefs form → · Cost → demand bridge → · Modeling hack proposal → The named beliefs-analysis dashboard remains an internal review artifact and is intentionally not linked from this public site; a redacted public synthesis should be produced separately. NoteWe Want Your Feedback For substantive or longer-form discussion — please post on 💬 GitHub Discussions. That’s where the conversation can get involved, where others can reply and build on each other, and where everything stays threaded and organized. See the 📖 Discussion Map for where to post what, or jump to a hub: 🧠 Substantive hub — bio / econ / stats / engineering / welfare (the main event) 🎯 PQ framing · 💬 Workshop logistics · 🖥️ Platform & UX For quick inline notes on specific text or a parameter, use Hypothesis (click the < tab on the right edge). For anything beyond a brief highlight, prefer GitHub Discussions so the conversation stays organized and discoverable. 🎧 Listen: Technical Review (22 min MP3) — Audio walkthrough of model architecture and areas for review Other ways to reach us: Open a GitHub issue · Email contact@unjournal.org function toggleHypothesis() { const btn = document.getElementById('toggle-hypothesis') || document.getElementById('vc-hyp'); // Track state via body class (avoids checking Hypothesis's internal display/transform state) const hidden = document.body.classList.contains('hyp-force-hidden'); if (hidden) { document.body.classList.remove('hyp-force-hidden'); if (btn) btn.textContent = '◉ Annotations'; } else { document.body.classList.add('hyp-force-hidden'); if (btn) btn.textContent = '▷ Annotations'; } } function toggleFullWidth() { const body = document.body; const btn = document.getElementById('toggle-fullwidth') || document.getElementById('vc-wide'); // Quarto full-page layout uses column classes and sidebar const contentSelectors = [ 'main.content', 'main', '.page-columns', '#quarto-content', '.column-page', '.column-body', '.column-body-outset', '.panel-fill', '.panel-sidebar', '#quarto-sidebar', '.page-layout-full .content' ]; if (body.classList.contains('fullwidth-mode')) { body.classList.remove('fullwidth-mode'); // Remove injected style const injected = document.getElementById('fullwidth-style'); if (injected) injected.remove(); btn.textContent = 'Expand Content'; } else { body.classList.add('fullwidth-mode'); // Inject a style tag to override Quarto's layout constraints if (!document.getElementById('fullwidth-style')) { const style = document.createElement('style'); style.id = 'fullwidth-style'; style.textContent = ` body.fullwidth-mode #quarto-content, body.fullwidth-mode main.content, body.fullwidth-mode main, body.fullwidth-mode .page-columns, body.fullwidth-mode .column-body, body.fullwidth-mode .column-page { max-width: 100% !important; width: 100% !important; padding-left: 1rem !important; padding-right: 1rem !important; } body.fullwidth-mode #quarto-sidebar, body.fullwidth-mode #quarto-margin-sidebar, body.fullwidth-mode .sidebar { display: none !important; } body.fullwidth-mode .page-columns { grid-template-columns: 1fr !important; } `; document.head.appendChild(style); } btn.textContent = 'Normal Width'; // Also hide hypothesis when going fullwidth hideHypothesis(); } } function toggleParams() { const btn = document.getElementById('toggle-params'); const vcBtn = document.getElementById('vc-params'); const sidebar = document.querySelector('.panel-sidebar'); const fill = document.querySelector('.panel-fill'); if (!sidebar) return; const isHidden = window.getComputedStyle(sidebar).display === 'none'; if (isHidden) { sidebar.classList.remove('params-hidden'); sidebar.style.removeProperty('display'); if (fill) { fill.style.removeProperty('grid-column'); fill.style.removeProperty('width'); } if (btn) btn.textContent = '◀ Hide Parameters (expand charts)'; if (vcBtn) { vcBtn.textContent = '◀ Parameters'; vcBtn.style.background = '#e8f4e8'; vcBtn.style.borderColor = '#5a7a5a'; } } else { sidebar.classList.add('params-hidden'); if (fill) fill.style.gridColumn = '1 / -1'; if (btn) btn.textContent = '▶ Show Parameters'; if (vcBtn) { vcBtn.textContent = '▶ Parameters'; vcBtn.style.background = '#f8f9fa'; vcBtn.style.borderColor = '#ccc'; } } } // Wide sidebar: turn the parameters panel into a fixed-position drawer overlay. // Using position:fixed avoids all CSS grid conflicts — the sidebar lifts out of // the document flow and floats on top. A dimmed backdrop lets users click outside // to close. The OJS cells inside still work because the DOM nodes are the same. function toggleWideSidebar() { const body = document.body; const sidebar = document.querySelector('.panel-sidebar'); const backdrop = document.getElementById('params-backdrop'); const btnV = document.getElementById('vc-wide-params'); const btnI = document.getElementById('toggle-wide-params'); if (body.classList.contains('sidebar-wide')) { // Close drawer body.classList.remove('sidebar-wide'); const s = document.getElementById('sidebar-wide-style'); if (s) s.remove(); if (backdrop) backdrop.style.display = 'none'; if (btnV) { btnV.textContent = '⟺ Wide params'; btnV.style.background = '#f8f9fa'; btnV.style.borderColor = '#ccc'; } if (btnI) btnI.textContent = '⟺ Wide view'; } else { // Open drawer: make sidebar a fixed overlay panel body.classList.add('sidebar-wide'); if (!document.getElementById('sidebar-wide-style')) { const s = document.createElement('style'); s.id = 'sidebar-wide-style'; s.textContent = ` body.sidebar-wide .panel-sidebar { position: fixed !important; top: var(--quarto-navbar-height, 62px) !important; left: 0 !important; width: min(62vw, 820px) !important; max-width: 92vw !important; height: calc(100vh - var(--quarto-navbar-height, 62px)) !important; max-height: calc(100vh - var(--quarto-navbar-height, 62px)) !important; overflow-y: auto !important; overflow-x: hidden !important; z-index: 6000 !important; background: white !important; box-shadow: 8px 0 32px rgba(0,0,0,0.22) !important; padding: 0.5rem 2rem 5rem 1.5rem !important; border-right: 3px solid #3498db !important; scrollbar-width: thin !important; } `; document.head.appendChild(s); } if (backdrop) backdrop.style.display = 'block'; if (btnV) { btnV.textContent = '↘ Narrow params'; btnV.style.background = '#e8f0fa'; btnV.style.borderColor = '#3498db'; } if (btnI) btnI.textContent = '↘ Narrow view'; // Scroll drawer to top when opening if (sidebar) sidebar.scrollTop = 0; } } // Re-initialise Tippy tooltips for abbr[title] and span[title] elements added by OJS. // OJS renders after Quarto's DOMContentLoaded Tippy pass, so elements inside OJS // html`` templates are missed on first pass. Span elements (used for ⓘ inline notes // in the main content area) are also included — native browser title tooltips clip // long text and are unreliable across browsers. function initOjsTooltips() { if (typeof tippy === 'undefined') return; document.querySelectorAll( '.panel-sidebar abbr[title], .panel-fill abbr[title], .panel-sidebar span[title], .panel-fill span[title]' ).forEach(function(el) { const t = el.getAttribute('title'); if (!el._tippy && t) { el.removeAttribute('title'); // prevent duplicate native browser tooltip tippy(el, { content: t, placement: 'bottom', maxWidth: 500, allowHTML: false }); } }); } setTimeout(initOjsTooltips, 2500); // first pass after OJS settles // Re-run whenever sidebar OR main content changes (OJS re-renders on slider change) document.addEventListener('DOMContentLoaded', function() { var obs = new MutationObserver(function() { initOjsTooltips(); }); var sidebar = document.querySelector('.panel-sidebar'); var fill = document.querySelector('.panel-fill'); if (sidebar) obs.observe(sidebar, { childList: true, subtree: true }); if (fill) obs.observe(fill, { childList: true, subtree: true }); }); // Reset all parameters: clear URL state and reload (most reliable approach, // since OJS viewof inputs can't be reset reliably from plain JS) function resetAllDefaults() { window.location.href = window.location.pathname + window.location.hash; } var _expandAllActive = false; var _expandAllObserver = null; function collapseAllDetails() { _expandAllActive = false; if (_expandAllObserver) { _expandAllObserver.disconnect(); _expandAllObserver = null; } document.querySelectorAll('details[open]').forEach(function(d) { d.removeAttribute('open'); }); } function expandAllDetails() { _expandAllActive = true; document.querySelectorAll('details').forEach(function(d) { d.setAttribute('open', ''); }); // OJS re-renders reactive cells after every slider change, replacing <details> elements // with fresh closed ones. The MutationObserver re-applies 'open' to any new ones. if (!_expandAllObserver) { _expandAllObserver = new MutationObserver(function(mutations) { if (!_expandAllActive) return; mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType !== 1) return; if (node.tagName === 'DETAILS') node.setAttribute('open', ''); if (node.querySelectorAll) node.querySelectorAll('details').forEach(function(d) { d.setAttribute('open', ''); }); }); }); }); _expandAllObserver.observe(document.body, { childList: true, subtree: true }); } } function toggleViewControls() { const content = document.getElementById('view-controls-content'); const icon = document.getElementById('vc-minimize'); if (!content) return; const isHidden = content.style.display === 'none'; content.style.display = isHidden ? 'flex' : 'none'; if (icon) icon.textContent = isHidden ? '−' : '≡'; } function toggleToc() { const btn = document.getElementById('toggle-toc') || document.getElementById('vc-toc'); // Quarto TOC can be in several locations const selectors = [ '#quarto-margin-sidebar', '#TOC', '.sidebar.toc-left', 'nav.toc', '#quarto-sidebar' ]; let toc = null; for (const sel of selectors) { toc = document.querySelector(sel); if (toc) break; } if (!toc) { btn.textContent = 'TOC not found'; return; } if (toc.style.display === 'none') { toc.style.display = ''; btn.textContent = 'Hide Table of Contents'; } else { toc.style.display = 'none'; btn.textContent = 'Show Table of Contents'; } } function hideHypothesis() { document.body.classList.add('hyp-force-hidden'); const btn = document.getElementById('toggle-hypothesis') || document.getElementById('vc-hyp'); if (btn) btn.textContent = '▷ Annotations'; } View ≡ ◀ Parameters ⟺ Wide params ▷ Contents ▷ Annotations ↔ Expand ▲ Collapse all ▼ Expand all /* ── Fullwidth mode: hide parameter sidebar, let charts fill the page ── */ .fullwidth-mode .page-columns { grid-template-columns: 1fr !important; max-width: 100% !important; width: 100% !important; } .fullwidth-mode #quarto-content, .fullwidth-mode main.content, .fullwidth-mode main, .fullwidth-mode .column-page, .fullwidth-mode .column-body, .fullwidth-mode .panel-fill { max-width: 100% !important; width: 100% !important; padding-left: 1rem !important; padding-right: 1rem !important; } /* Hide parameter sidebar and TOC in fullwidth mode */ .fullwidth-mode .panel-sidebar, .fullwidth-mode #quarto-sidebar, .fullwidth-mode #quarto-margin-sidebar, .fullwidth-mode .sidebar { display: none !important; } /* Hide Hypothesis in fullwidth mode, and when manually toggled off */ .fullwidth-mode .annotator-frame, .fullwidth-mode .hypothesis-sidebar, .fullwidth-mode hypothesis-sidebar, .hyp-force-hidden .annotator-frame, .hyp-force-hidden .hypothesis-sidebar, .hyp-force-hidden hypothesis-sidebar, .hyp-force-hidden iframe[src*="hypothes.is"] { display: none !important; visibility: hidden !important; pointer-events: none !important; } /* Sidebar overflow fix — offset by Quarto navbar so top buttons aren't hidden */ .panel-sidebar { overflow-y: auto !important; overflow-x: hidden !important; position: sticky !important; top: var(--quarto-navbar-height, 62px) !important; max-height: calc(100vh - var(--quarto-navbar-height, 62px)) !important; } .panel-sidebar details[open] { max-width: 100%; overflow-wrap: break-word; } /* When params sidebar is hidden, let charts take full width */ .params-hidden { display: none !important; } NoteNew to Cultured Meat? Read our deep dive: How Cultured Chicken is Made — a detailed guide covering cell banking, bioreactors, media composition, growth factors, and why each step affects costs. Quick summary: Cultured chicken is produced by growing avian muscle cells in bioreactors. The main cost drivers are media (amino acids, nutrients), growth factors (signaling proteins), bioreactors (capital equipment), and operating costs. For a side-by-side comparison of how published TEAs differ in their assumptions and estimates, see our TEA Comparison page. This model stops at a factory-gate cost; for what a given cost implies for consumer adoption and market share, see our Cost → Demand bridge, which connects to Pablo AMC’s demand-side model. CELL BANK → SEED TRAIN → PRODUCTION → HARVEST → PRODUCT [O] [OOO] [OOOOOOO] [===] [≡≡≡]

      So many folds at the top. Better than having them unfolded but I feel like there must be a better way of displaying this, perhaps with some horizontal boxes, perhaps some of these should go to external links, etc

    1. The four initial $100 response slots are provisionally filled as of 31 July 2026. Earlier individual commitments still stand1If you were promised an incentive for responding by a particular date, we will hold to that promise. If you have any questions, email contact@unjournal.org. A further $50 is reserved for eligible respondents who return for the follow-up/update round. Please check with us before assuming that a new response is compensated.

      Actually make this compensation but a folded box -- not central

    2. Once we've gathered this round of independent forecasts, we'll circulate the combined picture and the main points of disagreement, and invite you to revisit and update your estimate. For now, please answer from your own knowledge and reasoning. Naturally, we encourage you to consult your own notes, do a background web search, run calculations, etc.

      make this latter part of the paragraph a tooltip

    3. Include worlds without actual large-scale production. The scale assumption does not imply successful commercialization or price parity.

      This isn't a direct quote -- no need for a quote offset. Instead "If you imagine no plants producing at this scale, consider the hypothetical cost at this scale". And then the tooltip on 'hypothetical cost at this scale' should say, "in worlds where there is no production at this scale, you should consider what the cost would be if a firm had chosen to produce at such a scale"

    4. in 2036, at a scale of at least 2,000 tonnes per plant per year? Definition Use actual comparable plant data where available; otherwise estimate a hypothetical industry using the technology and input-supply conditions you expect in 2036. The average is total capital plus operating cost divided by total biomass output, not the cheapest or a representative plant. If hypothetical, specify the imagined production mix in your reasoning. Granting this scale does not assume technology breakthroughs, a mature supply chain, or a particular learning history.

      in the tooltip -- how can they use 'actual plant data'? this is for a future world that doesn't exist yet

      Only the second sentence is helpful here, and it can be a tooltip on 'production-weighted average'

    5. If hypothetical, specify the imagined production mix in your reasoning. Granting this scale does not assume technology breakthroughs, a mature supply chain, or a particular learning history.

      YOu can leave off this last paragraph ... it's confusing and cumbersome

    6. If your cost estimate excludes technically infeasible outcomes, explain that below.

      why would an estimate include technically infeasible outcomes? That wouldn't be a good estimate, I think.

    1. A shorter route into the technical and empirical disagreements behind 2036 cost projections. Start with the five-card overview; open only the cruxes or questions you care about. This page uses public S1 and S3 material, consent-checked beliefs, public-cleared written contributions, and official regulatory sources. No S2 discussion is used.

      This could be a tooltip or fold

    2. Here, a crux is not just a general uncertainty. It is a question whose answer would materially change the cost forecast or the chance that the forecasted industry exists.

      "not this but that" -- AI tic -- let's improve this

    1. PUBLIC WORKING SYNTHESIS · Forms checked Tue/Fri; new responses require human privacy review · Manual review: 31 July 2026 · AI-assisted synthesisAI drafted and updated the synthesis prose and calculations. The latest human check-in was 31 July 2026, when privacy handling, email follow-ups, and the proposed synthesis update were reviewed and adjusted. Human review does not mean that every claim has been independently verified.

      this is too emphasized -- remove the yellow

    1. Better Together? The Impact of Cash Transfers and Couples' Financial Planning on Household Dynamics and Economic Outcomessource titleLiteral title supplied by NBER metadata; not model-generated. Cause: Global health & dev. (LMICs) AI shortlist |Sarika Gupta, Jessica Leight, Daniel W. Maggio |NBER gpt-5.5 (codex headless, xhigh) Human 90/100 · n=1Sorry: this weighting is ad hoc. Quick --, -, ~, +, and ++ ratings map to 10, 30, 50, 70, and 90 and have weight 1. A full rating has weight 2 without detailed discussion; with detailed discussion it has weight 3 for public raters and 4 for team members. “Detailed” currently means at least 120 characters or 20 words. The AI score has weight 3 in the synthesis, in quick-rating equivalents. Each person’s latest rating counts once. Effective human weight for this paper: 1 quick-rating equivalent. Source-provided abstract (not AI-generated): Households' most consequential economic decisions are usually made jointly, yet behavioral interventions are typically t...Households' most consequential economic decisions are usually made jointly, yet behavioral interventions are typically targeted to individuals and optimized to overcome intra-personal constraints. We use a randomized controlled trial to study whether behavioral interventions targeted to couples can overcome interpersonal constraints related to intra-household decision-making. In rural Liberia, we layer a facilitated joint financial planning exercise atop a large unconditional cash transfer, assessing effects relative to cash only and a pure control. Relative to cash only, adding planning increases the amount of the transfer spent on productive investment and consequently improves economic outcomes by 0.16 standard deviation units. However, for households with the highest ex-ante conflict risk, these gains come at the cost of additional intimate partner violence. ▾ Open details, rate or discuss Full abstract, scoring context, and feedback options Quick rate---

      to reduce visual clutter and overwhelm, when you "open details" for one paper, can other content be grayed out in some way?

  4. Aug 2026
    1. If possible, give a 0–100 percentile relative to other papers we might evaluate. Otherwise use the five-point scale. Numeric ratings are included in anonymous public aggregates. Team ratings require your name and team email and are checked against the internal roster before they receive the one-assessor workflow shortcut. For a first pass without prior scores or analyses, use the independent-rating view.

      public sharing disclaimer here too

    2. for each paper, how strongly should The Unjournal commission an evaluation?

      be more clear -- rating this for -- is it potentially impactful and is there evaluation value ... not '"it a good paper"

    1. Targeted intake and prioritization requests - 74 records across 4 paper search families and 1 crux search family This shared catalog records topic-focused searches outside the broad recurring scans. Initial and follow-up passes are combined below as one search family. Use the Targeted intake filter above to include all records, exclude these deliberately oversampled searches, or inspect one family. A targeted label records how a record was found; it is not an endorsement, quality judgment, prevalence estimate, or completed Unjournal prioritization decision.AI impacts on global health and developmentResearch papers (29 papers; 2026-07-24) Economic, social-science, and policy research relevant to global health and wellbeing under transformative AI, especially implications for lower-income countries. Search-pass details (4)AI impacts on global health and development: LMIC labor and preparedness: Targeted public-paper search guided by the Coefficient Giving application discussion, manual source and thematic-fit verification, deduplication, and Codex subscription scoring. This pass follows the application discussion's risk-to-response framing: exposure estimates are inputs, not outcomes, and should be assessed alongside actual task content, adoption, infrastructure, institutions, service-trade exposure, and feasible policy responses. It deliberately includes competing estimates and early evidence on BPO and export-linked work, youth and expertise pathways, firm adoption, and frontline health care. LMICs are not treated as one labor market, and inclusion is not endorsement or a completed Unjournal team decision.Transformative AI, global health, and wellbeing follow-up: Second-wave public-paper search focused on first-run gaps, followed by deduplication and Codex subscription scoring. This second wave follows up the initial requested intake and deliberately targets gaps it left: public finance and tax-base resilience, distribution and social protection, AI market power, worker attitudes, and developing-economy structural transformation. It uses only independently public paper or institutional publication pages. Inclusion is not an endorsement or a completed Unjournal team decision.AI impacts on global health and development: Public OpenAlex search, theme-fit triage, deduplication, and Codex subscription scoring. The search follows the linked Coefficient Giving request for proposals but is an independent Unjournal intake exercise. It emphasizes non-catastrophic transformative-AI scenarios, labor markets, fiscal capacity, social protection, biomedical and health-system bottlenecks, development strategy, and aid allocation. Inclusion does not imply endorsement by Coefficient Giving or an Unjournal team decision.AI, global health, and development curation: Manual curation of 25 candidate leads with public-link verification. This curation was preserved on a development branch but was not merged as a live intake batch. The labeled dashboard records were already present through other sources and are marked as surfaced rather than added. Candidate leads without a verified public paper page were not published. Coefficient Giving request for proposals. View 29 records →Large-N GCR and existential-risk quantitative evidenceResearch papers (6 papers; 2026-07-30) Credible large-N, cross-country, historical, forecasting, survey, or other quantitative evidence directly bearing on global catastrophic or existential risks. Public OpenAlex search, quantitative-design triage, deduplication, and prioritization scoring. This is a separate follow-up to the empirical-conflict search. It asks whether global catastrophic and existential-risk questions have credible quantitative literatures with enough observations or repeated judgments to support evaluation or replication. It excludes primarily conceptual argument, technical engineering, and ordinary civil-conflict studies without a direct catastrophic-risk connection. Inclusion is not an endorsement or a completed Unjournal team decision. View 6 records →Empirical conflict replication-game intakeResearch papers (14 papers; 2026-07-24) Recent empirical economics, political science, and related research with conflict or political violence as an outcome. Search-pass details (2)Empirical conflict replication-game intake: Public OpenAlex search, venue and empirical-design triage, deduplication, and Codex subscription scoring. This run supports consideration of a topic-focused replication game. It favors empirical papers in a coherent interdisciplinary conflict literature, especially work using geocoded or commonly shared conflict data and designs suitable for robustness or replication discussion. It excludes broad international-relations commentary and historical theory. Policy relevance and the quality of underlying conflict measures are treated as substantive uncertainties, not assumed strengths.Empirical conflict replication-game follow-up: Second-wave public-paper search focused on canonical geocoded designs, common conflict datasets, and measurement fragility, followed by Codex subscription scoring. This second wave follows up the initial conflict intake with canonical and recent empirical papers that make the proposed replication game's shared data and robustness questions more concrete. It includes one explicit measurement-methods anchor alongside outcome papers. Inclusion is not an endorsement, and policy relevance and conflict-data quality remain open questions. View 14 records →Soil-invertebrate crux sweepCruxes (6 cruxes; 2026-07-11) Decision-relevant cruxes involving soil invertebrates, soil animals, land use, and indirect animal-welfare effects. Targeted search of public forum posts followed by manual curation. Added after a specific request to look beyond the broad forum scan for soil-invertebrate cruxes. The topic was deliberately oversampled, so these entries should not be interpreted as representative of the overall corpus or as completed Unjournal prioritization decisions. View 6 records →Animal-welfare focused intakeResearch papers (19 papers; 2026-06-22) Public research on animal-welfare economics, food systems, consumer behavior, and production systems. Search-pass details (2)Animal-welfare economics public-paper follow-up: Manual follow-up restricted to independently public papers and DOI records. These papers were added through a requested, topic-focused follow-up rather than the broad recurring scan. Private or unpublished program leads were excluded. The topic was deliberately oversampled, so inclusion is not evidence of prevalence in the wider literature, author endorsement, or a completed Unjournal team decision.Animal-welfare focused intake: Focused OpenAlex search followed by local prioritization. This focused scan added three independently public OpenAlex papers. Unpublished conference-program topics are not named or exposed here. The topic was deliberately oversampled, so inclusion is not evidence of prevalence in the wider literature, author endorsement, or a completed Unjournal team decision. View 19 records → What is this score? (weights for the active lens) Active lens: Evaluation priority. The priority score and default sort are a weighted average of the six sub-scores below (each renormalized if a sub-score is missing; out-of-scope papers are floored).CriterionWeightWhat it capturesPotential decision relevance30%Relevance to important real-world / global-welfare decisions.Methodological potential10%Methodological rigor and how evaluable the analysis is.Real-world influence10%Whether decision-makers are already using or citing the work.Prominence0%Attention the work already has (venue, authors, citations).Timing value20%Early-stage work benefits most from timely feedback.Neglectedness / value of evaluation30%How much independent evaluation would add beyond existing review.Research relevance weights importance, rigor, and relevance for readers looking for useful work. Evaluation priority (default) adds timing and value-of-evaluation (neglectedness) for deciding what The Unjournal should evaluate next. Custom weighting & sorting

      these folds are taking up a bit too much vertical space

    1. Search: SortOrder papers by overall AI priority score, recency, decision relevance, or other criteria. "Deeper Model First" shows papers scored by the most capable AI models at the top. Priority (Human-AI synthesis) Most Recent First Deeper Model First Potential Decision Relevance Prominence Timing Value Neglectedness Crux/PQ relevance Custom weights Rating sourceChoose the score shown in the circle and used by the default priority sort. Sorry: the current weighting is ad hoc. Quick --, -, ~, +, and ++ ratings map to 10, 30, 50, 70, and 90 and have weight 1. A full rating has weight 2 without detailed discussion; with detailed discussion it has weight 3 for public raters and 4 for team members. “Detailed” currently means at least 120 characters or 20 words. The AI score has weight 3 in the synthesis, in quick-rating equivalents. Each person's latest rating counts once. Papers without human feedback retain their AI score in synthesis mode. Human-AI synthesis AI ratings only Human ratings only Evaluation-relevance weightingON (default): rank by Evaluation priority for The Unjournal — adds timing and value-of-evaluation (neglectedness) on top of relevance and rigor. OFF: rank by Research relevance for readers — importance, rigor, and relevance. Unfold "What is this score?" below to see the exact weights. Cause AreaThe primary global issue this research addresses, aligned with Unjournal's focus areas: global health, development, animal welfare, climate, AI governance, catastrophic risks, and more.

      how do I clear all filters? make it easier to do that

    1. We're continuing the discussion asynchronously and sharing key materials here. This site is evolving into a resource page. Read the workshop summary →  ·  Survey results & synthesis → ·

      Don't flag the "Survey Results and Synthesis" link you're giving here - that's a survey about whether people found the workshop useful. That's less interesting for outside participants - we can share that publicly, but it shouldn't be at the top. It shouldn't be prioritized. Instead, you can now link the belief elicitation results https://uj-wellbeing-workshop.netlify.app/beliefs-overview-w9k3m.html and the synthesis of this (is the latter up yet?)

    1. WELLBY measurement value, conversion, and forecasting · March 16 2026 workshop · The Unjournal · Internal draft, last corrected August 3 2026 — estimates visible, for internal review only

      Encourage people to look at the subquestions too. !!! I post that more prominently.

    2. WELLBY measurement value, conversion, and forecasting · March 16 2026 workshop · The Unjournal · Internal draft, last corrected August 3 2026 — estimates visible, for internal review only

      Link the summary/synthesis here - it's highly connected to this page and these results.

    3. Data-quality warning: The form opened with values already set on some sliders — PQ1A at 70 with its 80% interval pre-set to 40 and 90, PQ3A at 25%, PQ3C at 35% — and it recorded only the final value, not whether a respondent moved anything. These were not neutral midpoints: on a 0–100 scale the form presented what looked like a complete, plausible answer, which is a stronger anchor than a midpoint would have been. Exact-default values may therefore be deliberate answers, anchoring effects, or untouched sliders, and the data cannot distinguish them. Among the 7 responses, one matches all three PQ1A defaults, two match the PQ3A default, and two match the PQ3C default.

      Update - I think we've asked people to follow up on this... and they probably should be excluded by default.

    4. June 18 update: Added Miles Kimball's June 17 response. The most visible effects are a higher forecast for GiveWell-linked charity uptake (new max 50%), a higher forecast for agreement among surveyed economists and practitioners (60%), and a much wider DALY/WELLBY conversion range because Kimball gives 1 WELLBY/DALY with a 0.1–10 interval while stressing that his team has not studied that conversion directly.

      Make these "updates" folded sections.

    5. Asked again August 3 which of his values were intended; no answer yet — see above

      update -- he has now responded and made it pretty clear that where he gave the default answer, it was probably once he skipped .

    1. Representation and exit Q-001 · Representation / reference perspectiveWho the firm represents, or which perspective frames trade-offs. ↓ reveals Q-009 · DeadlockWhat happens when two members cannot agree.

      A bit more description of this dependency. You've made it clear that question one may reveal question nine, but what answers to question one reveal this, or is it just putting any answer that reveals it?

  5. Jul 2026
    1. The authors published a detailed response and reported revising the working paper to address evaluator comments. The public evaluation summary records changes to framing, methods, transparency, and the cost-effectiveness analysis while the headline result remained broadly similar.

      missing some content here .. add Givewll's response etc

    1. Empirical conflict replication-game follow-upResearch papers (8 papers; 2026-07-24)

      most of this is on civil wars etc., and it can be more robust because the data is richer .. But for GCR/Xrisk there was not much found. Is there credible 'large-N' quant analysis worth targeting? -- maybe do a nother target5ed run for that?

    1. Empirical conflict replication-game intakeResearch papers (6 papers; 2026-07-24) Recent empirical economics, political science, and related research with conflict or political violence as an outcome. Public OpenAlex search, venue and empirical-design triage, deduplication, and Codex subscription scoring. This run supports consideration of a topic-focused replication game. It favors empirical papers in a coherent interdisciplinary conflict literature, especially work using geocoded or commonly shared conflict data and designs suitable for robustness or replication discussion. It excludes broad international-relations commentary and historical theory. Policy relevance and the quality of underlying conflict measures are treated as substantive uncertainties, not assumed strengths. View 6 papers →Empirical conflict replication-game follow-upResearch papers (8 papers; 2026-07-24) Follow-up to Empirical conflict replication-game intake. Empirical conflict-outcome studies suitable for a coherent replication exercise, emphasizing spatial data, shared event datasets, design comparability, and measurement sensitivity. Second-wave public-paper search focused on canonical geocoded designs, common conflict datasets, and measurement fragility, followed by Codex subscription scoring. This second wave follows up the initial conflict intake with canonical and recent empirical papers that make the proposed replication game's shared data and robustness questions more concrete. It includes one explicit measurement-methods anchor alongside outcome papers. Inclusion is not an endorsement, and policy relevance and conflict-data quality remain open questions. View 8 papers →

      merge these in the epxlanation

    2. 24Evaluation priority (0-100): how strongly this paper is recommended for a Unjournal evaluation. Weighted for UJ prioritization — decision relevance, neglectedness, timing, methodology, and influence. NOT a rating of research quality. Evidence from scaling laws AI, technology & governance Lower Priority |EA Forum gpt-5.5 (codex headless, medium) crux: A deep critique of AI 2027's bad timelin... · 42% ▾ Details / Rate This is a highly influential OpenAI arXiv preprint on n

      You changed the title!!! use actual titles!!

    1. 0–5 min Updates and agenda choices Agree on where the meeting's attention is most useful today. 5–25 min Actual resear

      this is just an EXAMPLE schedule

    2. Wellbeing, psychology, and welfare economics

      this isn't really a field grouo/not the focus of the Psych/attitudes field group ... e.g., they previously considered focusing on misinformation. Moral circles (specific empirical questions) could also be relevant, and perhaps human response to AI tools

    3. Which empirical claims about AI scaling are decision-relevant and uncertain enough to merit independent technical auditing?

      this is meta -- I was looking for actual examples

    4. It would add a more durable layer connecting those activities to the larger research questions and decisions we care about.

      that's a bit too strong perhaps ... not all of those are connected

    5. This is a discussion proposal, not a decided policy or a finished framework. The examples are starting points meant to be criticized, rewritten, expanded, or rejected at the meeting.

      shorten this ... and it's a discussion beyond the meeting

    1. AI assistant; Opus model useful for deep analysis of research papers and evaluation assistance.

      needs updating -- Fable is now latest as of 13 Jul 2026

    1. The useful question now is not whether the old paper should be revived unchanged. It is what, if anything, is worth testing under current technology and policy.

      this is an annoying AI 'not this but that' again

    1. Show detailed scoring rubrics & methodology

      "Show" --> SHOW/HIDE ... and make it clearer that this is the 'toggle to open the content below' ... it's not shared or offset or boxed, so that's hard to see

    2. Scores are calibrated against 353 actual human prioritization decisions from the Unjournal team. The AI scores are systematically compared to human assessor ratings, and field-specific corrections are applied. Read more about UJ’s prioritization process.

      Link a page on 'how this was 'calibrated' here'

    3. For prominent work (NBER, CEPR, World Bank, top journals), decision-relevance dominates. For less-prominent work, methodology becomes the tie-breaker—our evaluation could boost neglected but rigorous research.

      clarify this -- see note abobe

    4. Methodological potential

      this needs more discussion and thinking. Even for prominent work the strength of the data/approach may be weighted somewhat. E.g., all else equal, real world measured outcome data is usually preferred to recall data, which might be preferred to hypothetical choice data. However, for prominent work we don't weight/rate the 'methodological correctness' of the work in rating its potential for impact. If it's prominent and ~influential already, if it has major flaws that makes it perhaps _ more_ important, to be publicly evaluated. In contrast, for less prominent work, clear methodological, logical, contextual, and communications flaws make it less likely to have potential for impact, and thus less valuable to commission for evaluation.

    5. The list can be scored and sorted through either of two transparent weightings of the six sub-scores — toggle Evaluation-relevance weighting in the controls, and unfold What is this score? to see the live weights:

      (Tooltip?) We may want to ask users to specifically rate papers by these multiple categories, in part to help refine the model

    6. We welcome both team and public feedback.

      Explain how we are also soliciting human ratings, and these wil be incorporateed into this tool, both in the direct prioritization, and informing the AI model.

    7. Quick-rate mode

      A tooldtip should explain 'quick rate mode'. And we should have a folding box (and or a link on 'get involved') explaining how we want direct ratings -- either quick +/- ratings, or detailed discussion input on prioritizing these research objects according to their potential for impact

    1. ation tool and prototype AI-assisted evaluation and summary workflows, keeping humans in the loop for judgment and ratings.

      and benchmark and investigate AI capabilities, alignment, priorities, and tast. Link this (unsuccessful) proposal https://llm-uj-research-eval.netlify.app/proposal ... but still a WIP project

    2. and do the grant-writing.

      I'll do most grantwriting work, but we're also interested in getting a bit of help from others especially if they are writting into the grant substnitally

    3. We're building our next grant applications around what you'd want to do

      that's a bit too strong -- it's a factor, but we're not necessarily going to implement exactly what people propose

    4. elevant questions and the decision-makers who'd use the answers.

      add a box or subbox for the proposed modeling hack: https://6a3a77696a52ea869c36e13d--uj-cm-workshop.netlify.app/modeling-hack.html

      and Beliefs synthesis (still awaiting further responses, see https://uj-cm-workshop.netlify.app/beliefs-overview-r7p2k.html)

    5. I-assisted prioritization prototype An early tool that suggests research to evaluate and gives preliminary indicators, using our principles and prior prioritization as context.

      this looks good but put another more promionent box xplaining our LLM evaluation/benchmarking project: https://llm-uj-research-eval.netlify.app/

    6. ost of these are lightweight, and several were built quickly — they're meant to show the shape of what's possible, not to be polished products.

      make this atooltip

    1. so an IO question doubles as a welfare question and as a target for antitrust or procurement policy

      But the animal welfare implications will go in the reverse to the usual antitrust direction -- monopoly/monopsony power limits quantity and thus the AW burden, right? (iirc I've seen some version of this before).

      Key question for us: what is the vector we can potentially affect here/whats the ToC? (NB that's kind of our usual question)

  6. Jun 2026
    1. The focal cost question is live on Metaculus: CM_01 — production cost per kg →. The expert-aggregation version is at CM_03 →. If you forecast on Metaculus, please share your username below so we can link your contributions.

      Add a paragraph in this fold: You may also be interested in the Animal Welfare Futures forecasting tournament we co-launched with Metaculus and Sentient futures. $3,400 in prizes and only about 25 forecasters as of June 26. It includes some cultivated-meat questions (but not the specific cost questions here), alongside other animal-welfare-relevant questions you may find interesting.

      And change the folding block title to "Public forecasting on Metaculus; Animal Futures Tournament"

    1. Same quantity, but as of end-2027 — verifiable within a couple of years, which helps calibrate the longer-horizon estimate. If you expect no large-scale plants by then, estimate for advanced pilots and say so.

      use same lang as prev form as ,much as possible for consistency

    1. If you're a forecaster, modeler, or skeptic we've invited

      Thinking about this again. We're going to try to get the skeptics to share their beliefs right away and without seeing the beliefs specified by the workshop participants. This doesn't really need to be on this page, although you could mention it perhaps in a folding box for these people. The forum for the non-cultured meat f orecasters skeptics should probably be on its own standalone page. We also should let them know that we're going to be asking them for follow-up after they've considered the full set of beliefs, with additional compensation. So perhaps $100 for their initial belief statement plus $50 for that follow-up?

    2. What we're asking: share your beliefs on the focal cost question (CM_01) and any subquestions you have a view on, with brief reasoning, via the beliefs elicitation form. Roughly 30–60 minutes; partial responses are welcome.

      Ask them to do at least a reasonable share of the sub-questions... This seems like they are way too optional.

    3. CM Workshop · Beliefs Analysis

      At the bottom of this, put a link to the "Next Steps" page as well as a link to the "Subquestions" page, to direct people's interest towards these, which are pretty high value

    4. One contributor (a commissioned PQ evaluator, included anonymously at their request) submitted detailed answers in a written document using the original Pivotal-Questions wording, which differs from the workshop form on several questions. Their CM_14 and CM_17 map directly. Their CM_01, CM_12, CM_13 and CM_16 are shown as provisional conversions onto the workshop axes, using explicit stated assumptions:

      Put this at the bottom. It's too prominent here.

    5. Technical sub-questions and expert distribution questions. Only respondents who answered each question are shown.

      We should have aggregation for these also. -- I only see the individual responses below.

    6. CM_01: What will be the average production cost ($/kg, wet weight at harvest) of undifferentiated cultured chicken cell biomass in 2036, assuming large-scale commercial production is achieved?

      Signpost this a little bit; I'm worried people are only going to see this question. Note that this was the focal question, but there's a set of other questions that are particularly interesting - link the "subquestions" tab.

    7. Fairness note: prior workshop contributors who already submitted a detailed response are equally welcome to the same honorarium for a substantive post-workshop update.)

      Make the fairness note just the words "fairness note" and then a tooltip explaining that the difference in compensation is necessary to elicit participation from people with less inherent interest in learning about cultured meat. And that wWe hope to offer this sort of modest compensation more broadly in the future.

    8. without ties to cultured meat

      Italics not bold. Note that this helps balance (in an ad-hoc way) as people without a direct links to cultured meat may have been less likely to attend our conference. And this is also very important to help us get at the cruxes and make real progress in the public understanding.

    9. CM_01: What will be the average production cost ($/kg, wet weight at harvest) of undifferentiated cultured chicken cell biomass in 2036, assuming large-scale commercial production is achieved?

      Make the question title a bit more prominent here.

    Annotators

    1. This is not a conversion gate like valuation, liquidity, allocation, or deployment. It is a reason to discount or stress-test local confidence when the local social graph may benefit from optimistic expectatio

      This was a comment you didn't need to actually add literally.

    2. The stronger skeptical case is about conversion, not existence.The strongest steelman is not "there will be no AI-wealth philanthropy." That claim now looks weak. Anthropic has confidentially filed a draft S-1 S1Anthropic confidential draft S-1"confidentially submitted a draft registration statement on Form S-1"Supports the claim that an IPO/liquidity event is live, while still conditional.Open source, and OpenAI's nonprofit Foundation controls a large equity stak

      Is this really potentially a straw man here? Who was actually saying there will be no AI wealth philanthropy? Not sure we need this particular "not this, but that" Construction here.

    3. What would update this memo?

      Claude: Michael Dickens: "What would update this memo?" – as with the gates, no sense of prioritization is given. IMO by far the biggest uncertainty, about which we will get more information in the future, is: What will Anthropic's valuation be when the lockup ends? "Concrete donor vehicles" is also important evidence, but we won't get that until probably 6-24 months later.

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    4. How to plan if the skeptical case is live

      Claude: Michael Dickens: The recommendations under "How to plan if the skeptical case is live" don't really make sense. AFAIK ~zero orgs are planning as if they're guaranteed to get a huge pile of donations 1–2 years from now. I believe nonprofits mainly plan based on the money they already have on their books + short-term (<1 year) fundraising expectations. "How to plan if the skeptical case is live" is just "business as usual".

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    5. Notional company equity value, not spendable cash.

      Claude: Michael Dickens: My downward adjustments to the model aren't even the pessimistic case. The pessimistic case* is that the AI field collapses (investor funding dries up or something) and Anthropic stock is worth $0. Base rate says there's like a 50% chance that that will happen. Even optimistically, you should expect at least a 10–20% chance that Anthropic stockholders get nothing.

      *this is pessimistic for donations but I would actually prefer that this happen because it would lengthen timelines. so in a way it's the optimistic outcome

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    6. The steelman could be too pessimistic if founders or employees treat liquidity as an urgent moral obligation

      Claude: Michael Dickens: "The steelman could be too pessimistic if founders or employees treat liquidity as an urgent moral obligation" – TBH the BOTEC as written seems to me like it's already pricing in that founders/employees will treat donations as urgent, e.g. it's implying that Anthropic money will be disbursed faster than FTX Foundation money, which itself was disbursed at historic speed. IMO most likely reason why the model will end up underestimating is that Anthropic market cap ends up being like 10x higher than predicted.

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    7. Field absorption ceiling

      Claude: Michael Dickens: "Field absorption ceiling" is structured more sensibly than "Grantmaker capacity multiplier", but these two seem redundant because they're closely related. If orgs have more capacity to expand, grantmakers can deploy money faster by giving to those orgs. If there are more grantmakers, they can create more and bigger RFPs. etc. I would include one variable or the other, but not both.

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    8. it scales raw potential funding by the available people, institutions, judgment, and legal plumbing needed to make good grants.

      Claude: Michael Dickens: Also this would make more sense as a dollar amount, not a multiplier. Like there's a fixed total amount that grantmakers can reasonably disburse. You could model it in a more complicated way but IMO a simple cap is the way to do it. Or maybe don't use this parameter at all. I think it's probably worth including, but keep it simple.

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    9. Grantmaker capacity multiplier

      Claude: Michael Dickens: "Grantmaker capacity multiplier" seems nonsensical as written. Shouldn't the capacity max out at 1x? If grantmakers are a complete non-bottleneck, then the other parameters will dictate the amount disbursed; if they're a bottleneck, then the amount disbursed will be less. There's no way for grantmaker capacity to have a multiplier >1x.

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    10. Founder deployment by end-2026

      Claude: Michael Dickens: IMO "Deployment by end-2026" should use a different date. IPO 3-6 months from now plus 6 months lockup means no money will be deployed in 2026, unless Anthropic does a fast IPO + early lockup release. Even by the end of 2027, you're talking about a 3-9 month turnaround time on lockup ending -> grants being disbursed. FTX Foundation donated $190 million (pre-clawbacks) in about 6 months, which was ridiculously fast compared to a typical foundation, and that was still a pretty small % of its long-term budget (or at least, what was believed to be its long-term budget before FTX collapsed).

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    11. Conversion gates

      Claude: Michael Dickens: This model is supposed to illustrate how a lot of people are being too optimistic, but even then, I think most of the point estimates in the model are too optimistic. Consider that e.g. the median self-reported earner-to-give only donates (IIRC) 3% of their income.

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    12. Gate 8: Bay social incentives make optimistic funding rumors self-reinforcing

      Claude: Michael Dickens: Gate 8 ("Bay social incentives") seems uninteresting since it's not a claim in the same category as the others. It's more like a meta-level reason why people might not think about the other 7 gates.

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    13. A late-2026 IPO could still imply mid-2027 or later insider liquidity because of lockups.

      Claude: Michael Dickens: Gate 1 says IPOs have lock-ups. That's true but I basically don't think that matters because lock-ups are very predictable: they will announce how long it is, and that's exactly how long it will be. There's no uncertainty. The main reason it's relevant is that a lockup gives more time for AI valuations to fluctuate or collapse, but the text doesn't even mention this.

      Source: https://forum.effectivealtruism.org/posts/fFDM9RNckMC6ndtYZ/david_reinstein-s-shortform?commentId=JS2dmcbgfvy3xtBwj

    14. Read these as sequential gates. Realization handles liquidity and legal availability; follow-through handles intent; allocation handles cause choice; deployment handles the timing of actual grants by the deadline.

      where do the defaults come from? Explain, reference, link (maybe as tooltips)

    15. This may leave legally awkward, politically controversial, or non-lab-compatible work underfunded.

      this needs fleshing out -- not sure what this is about

    16. People who could puncture the rumor may also be financially incentivized not to alienate future donors.

      I don't see how this would 'alienate future donors'?

    1. main Hi Quick hello Suggest 5 min Call Phone call Suggest 15 min Video FaceTime / video Suggest 30 min Drive Drive-time call Suggest 20 min Msg Async text Suggest 1 hour

      What does the room mean here

    1. 4. Are you or your coauthors especially exposed to a durable negative public signal?

      This framing feels too negative and definitive for my taste. As some of the modeling and discussion gets at, a single 'negative public signal' should not be so damning as people seem to think.

    1. Three explicit flags that would substantially revise the estimates: (1) better proxy validity research (reversal learning, parental care as welfare indicators); (2) new data on understudied invertebrates; (3) theoretical advances o

      can these be restated as questions?

    2. The load-bearing belief is that neuron counts are only a defensible proxy for moral weight if they reliably correlate with the welfare-relevant capacities organisms actually posse

      Better stated as a question ... something like "Do neuron counts reliably correlate with the welfare-relevant capacities organisms actually possess?" (of course 'reliably' probably would need operationalization, and this elides the possibility that it may "reliably correlate" but the correlation may be low)

    3. Three explicit cruxes structuring the entire cross-cause model: (1) animals' moral weights relative to humans; (2) expected value of the long-run future; (3) preference for making a difference vs. expected value. Cause rankings reve

      these are good, but can they be separated out and flexhed out and made more explicit?

    4. The load-bearing belief is that any significant moral weight for animals, combined with Rethink Priorities' finding that corporate animal welfare campaigns are ~1000x more cost-effective than top global health interventions (e.g. AMF), implies Open Phil should prioritize animal welfare in neartermism. The author's position would change if the moral weights / welfare-range estimates favoring animals were substantially lower, or if a defensible reason were given for valuing human welfare units far above animal welfare units purely on species membership. The author explicitly asks what would have to change in OP's views to NOT prioritize animal welfare.

      Not sure this one is an actual crux, rather an attempt to draw out Open Phil (now CG) on their moral weighting

    5. the moral-weighting framework for suffering: whether one prioritizes duration of welfare improvement or the intensity/severity of suffering averted.

      'prioritizes' is not clear here. Do they better operationalize this ... how would it be decided/measured?

    6. What belief changes would actually alter donations or work — and what are the poster's actual cruxes? Author foregrounds room for more funding and marginal value.

      meta -- not an actual crix

    7. The load-bearing belief is that no compelling cost-effectiveness case exists for the alternatives; a promising cost-effectiveness estimate for institutional meat reduction campaigns would make the author excited about (and supportive of reallocating funding toward) them. To me, to be excited about such campaigns I'd need to see a promising cost-effectiveness estimate. High Shortlist EA Forum Did corporate campaigns in the US have any counterfactual impact? A quantitative model verified No — published 2019, before 2024 window Karolina Sarek 2019-06-24 Key uncertainty

      clarify -- hard to read this

    8. The estimated counterfactual impact of US corporate cage-free campaigns (2.1-10%) is load-bearing on the price elasticity of egg demand, a parameter drawn from a literature review with only ~13 observations.

      this is a fairly well-defined operationalized crux. It's old, but I guess 'we still are not confident'!

    9. He would be persuaded toward global health if shown a defensible rationale for valuing one unit of human welfare so much more than animal welfare that it justifies Open Phil funding GH ~6x as much as AW —

      this latter bit is closer to being a specific 'crux'. remember we want these stated as operationalizable questions,if possible. And try to find the single question crux that seems most important or most correlated to the others. You can have another column listing other cruxes raised.

    10. y The author's recommendation hinges on whether furnished-cage advocacy is actually more cost-effective per unit welfare than cage-free advocacy. He remains skeptical despite his own estimate (2.84x), and his position would shift on whether advocacy costs scale linearly with producer costs, whether infrastructure lock-in blocks later cage-free transitions, and whether furnished-cage campaigns would undermine the cohesiveness of global laying-hen advocacy.

      This one is getting good. It seems pretty relevant, but I want you to state it as an explicit opertionalizable question here.

    11. The author's position that shrimp welfare rests on a weak evidentiary base

      State the actual crux. Your description here is more about the implications of the crux.

      You can state somewhat general grounds, but then give a specific instantiation if possible. E.g., something about whether one form of slaughter is more painful to shrimp than another, or whether analgesia is good evidence, etc.

    12. Would update on: timelines, public/model concern for animals, indirect normativity, moral-circle expansion, and simulated-animal welfare.

      this is too general/vague. I want explicit defined crixes. What is the specific question they are uncertain about, how would you measure it, and what decision would it change?