2,204 Matching Annotations
  1. Last 7 days
    1. prioritization dashboardThis is a topic-focused working collection within the shared research-prioritization project. Each brief links to its dashboard record for scores, source material, and ratings. The dashboard contains a broader AI collection; these briefs are candidate assessments, not completed Unjournal evaluations or commissioning decisions..

      this should also have a link to the dashboard in addition to the tooltip

  2. whatisit-game.netlify.app whatisit-game.netlify.app
    1. hat are they selling?

      This category is not particularly fun. Instead of just the name of the company, maybe show the logo with the name, or their ad slogan or jingle, like their catchphrase.

    2. Who was closest?

      I can't see people's guesses on this screen - you should be able to see what people guessed when we want to see who was closest.

    3. Cloud software that lets companies store, combine and analyse large amounts of data, charged by how much they use.

      This one was super boring. Let's have fewer questions having to do with designing websites and tech support and things.

    4. technology professionals help each other.It later moved to experts-exchange.com, with a hyphen. Read it again without one.See it for yourselfDraft: not yet re-checked online

      This one is not completely fair because there actually is a dash in that actual website

    5. Davidelsbeth.geldhof@gmail.coFunniest guess (bonus, optional)Davidelsbeth.geldhof@gmail.co

      For some reason, I'm not seeing the guesses here. Some data got lost there.

    6. Give a clueTime!

      The timer is not really working. 1. You shouldn't have a shared timer; each person should have their own timer. 2. There needs to be some penalization if you go over the time.

    7. Search it yourselfChecked 2026-10-06

      When I search it myself with this button, I get very different results. Make it clear which search it's using, and make it more likely that when they click this button, they get the same search. Also, does it matter if we put it in quotes?

    8. Wikipedia's article on puppy chow

      Does it usually come up with Wikipedia's article? Because I'm getting that a lot, and maybe it's a little bit boring.

    9. The inventor and his sister call it 'Tarzan' swinging, and a jungle yell is optional. A 2003 re-examination cancelled all four claims.Read the patentChecked 2026-10-06 · US 6,368,227

      For bonus points, maybe people have to guess whether the patent still exists or whether it was canceled.

      Also, make it clear: are these all patents that were approved, or were some of them denied?

    10. The inventor and his sister call it 'Tarzan' swinging, and a jungle yell is optional. A 2003 re-examination cancelled all four claims.

      Give more pictures of things like this, maybe.

    11. Clue 1: It's a cute reference site for people who build websites.

      Fewer ones having to do with creating websites and tech stuff. That's not so fun.

    1. Should each group maintain a short, evolving map of its highest-value pivotal questions and most promising impactful research directions? A proposal for discussion—at and beyond the meeting. The examples are starting points, not settled views.

      This is obviously AI-generated... You see the "not this but that" and the passive voice, etc. Make it less so. Try to use natural human language and my own voice more, and also clearly acknowledge how it was created and who's monitoring it, etc.

    1. follow that label to the internal prioritization hub (sign-in required) for an independent follow-up assessment. The label marks a workflow step, not an endorsement.

      Update/tooltip -- - Note our streamlined policy -- summarize and tooltip this:

      Oct 1 2026: For ~the next ten papers/projects (until we have more funding), a piece of research (which anyone can nominate) will be commissioned for evaluation once two management team members "strongly approve". If multiple field specialists strongly approve, we'll also take this very seriously. We'll announce these "conditionally approved" research objects in the relevant Slack channels (let us know if you want email notifications for this), and then everyone on the team, including field specialists, has one week to weigh in. If no one raises a strong doubt, we'll go ahead.

    2. ind research worth evaluating: choose a topic, read the evidence, and add your judgment.

      Edit to "Find impactful research for your own work, and help The Unjournal prioritize research with impact-potential for evaluation/peer review. Sort and filter by field, sources, and "Cause (topic)" of your choice.

      (USe tooltips on some of this, let's not be too repetitive either)

    3. Get involved — Suggest research for this list ↓ | See how this tool will work (planned workflow) ↓ | Would you sponsor an evaluation? (idea, feedback wanted

      make this it's own top level thing -- not under 'about this prioritization tool. Also make it less bold and distracting.

    4. They are provisional, not ratings of research quality.

      They are not ratings of research quality. These are also provisional, and AI-generated, unless noted below.

    5. Scoring methodology

      Let's have scoring methodology hosted/open on another page.

      In general anything that goes beyond about 1 paragraph of text, that isn't core to the papers/prioritization, should continue on a linked page instead of an expander

    1. hey sometimes disobey instructions, cheat on tasks,misrepresent their work, and are unable to com-plete some tasks at all, necessitating human inter-vention [23–25].

      pdf format is is horrible here .much more useful to have a hover link to tell you what the reference actually is, what it says, and how it relates to the claim being made

    2. Preliminary evidence suggests that a software-driven intelligence explosion is possible. If onedoes happen, it could be the most consequen-tial technological development in history: AI sys-tems could rapidly eclipse human experts acrossmost domains and radically accelerate technolog-ical progress.

      Being a bit snarky, this is compounding three "coulds" or "suggests that" statements.

    3. should urgently obtain more visibility into the automation of AI R&D,develop ways to steer and constrain an intelligence explosion, and prepare society to adapt to anintelligence explosion’s impacts

      Surely, it's hard to critique "should" statements, as they're sort of expressing an opinion. However, they could be recast as "This would be welfare-improving in the following sense" or "If you want to achieve this, this would be helpful" - which could then be critiqued more directly. Obviously, this is an abstract and thus needs to be compact.

    4. Preliminary evidence suggests that it coul

      "Could" is, at best, a sort of existence of the possibility and maybe a directional statement. It's not very precise.

    5. could AI progress radically accelerate in an “intelligence explosion,” where years of advances arecompressed into months or less?

      I guess they mean the equivalent of "years of advances at today's pace" or "years of advances if AI were not used in this process at all." ?

    1. Because it is already receiving media and policy attention, the timing value for independent evaluation is especially high.

      Give some examples of this "media and policy attention ". OK you give it below. WHen you do this, please say "see below" and link the later section if possible

    2. The main concern is that this may be more of a synthesis and agenda-setting policy paper than an evaluable quantitative social-science contribution,

      that's a big concern for me (David). Also hard to evaluate if it's mostly a 'here's what we believe and advocate' paper.

    3. Source record Publication-stage and date evidence should be checked in the linked paper record. TARGETED_CURATED

      You keep giving the same link to the paper itself as if it's linked to evidence of other things outside the paper.

    4. The welfare case is strongest under assumptions that future sentient welfare counts substantially, AI R&D automation could materially accelerate dangerous capabilities, and governance preparation can reduce risk without causing major counterproductive acceleration or panic.

      this VoI ToC seems logical to me. (Although the "future sentient welfare counts substantially" is a bit vague -- either it's obviously true, or there's some assumption buried here about the importance of the long term future)

    5. An evaluation would focus on how well the empirical premises (the share of AI R&D already automated, the feedback-loop claims) support the policy asks, and on what is claimed versus assumed about speed.

      These focuses seem reasonable and useful to me.

    6. AI R&D could compress years of progress into months and asking governments for visibility, rate limits and emergency plans.

      "Could compress" is a bit imprecise. What is the actual main claim or claims the paper makes?

    7. Alan Chan, Christoph Winter, Andrew Barto, Jakub Pachocki, Geoffrey Hinton, Eric Horvitz, Yoshua Bengio, Dawn Song, Jack Clark, Hilary Greaves, Anton Korinek, Sam Manning, et al. · Catastrophic risks, the long-term future, forecasting

      nice to show affiliations if possible -- maybe as tooltips

  3. Sep 2026
    1. Econometrics, statistics and data science notes Working notes with a micro, behavioral and experimental focus. Web book Researching and writing for economics students A practical guide to economics dissertations and research projects. Course notes Microeconomics (MSc) Notes for an MSc microeconomics module at Exeter.

      Nice, but also you should make it more clear in some way that these were built a while back entirely without AI. (NB-- I do intend to update these and help fill in the gaps now that we have a range of tools for lookup, coding, etc., leveraging AI. That is somewhat of a priority. ... Once we start that project, I want to clearly maintain the "pre-AI" version that people can see if they truly prefer to see what the source was before all this AI stuff started to happen, but the AI versions are probably going to be the main one.

    2. Increasing effective giving The puzzle of why people don't give more effectively, what the evidence says, and what we still need to learn. Web book Impact of impact information on giving Field experiments and synthesis on whether telling donors about effectiveness changes what they give. Reports EA Market Testing Public reports from a team running marketing and messaging trials with effective-giving organizations. Data analysis

      These are research projects, not really interactive tools. I wouldn't give them quite as much space on the page, and they probably don't illustrate much of what the page was meant to illustrate. Yes, keep them, but don't make them quite so prominent. Perhaps an outlink to a page laying them out futher

    3. Games and everyday tools

      These navigation things are helpful, but maybe also let people switch to a simple list and give them some sort of filter they can use to filter these by meaningful categories.

    4. AI Foundations for Research Judgment How language models work, for researchers who have to judge their output: tokens, measurement, evidence workflows, compute economics and safety.

      Not sure I'd highlight this. Maybe this should be a less forefront one? It's basically just an internal tool that I built for myself to learn that aggregates information from the internet.

    5. and site GTD Trio Guitar, tuba and drums jazz trio in Western Massachusetts. Band site K-House Jazz A flexible live jazz group in Exeter, Devon.

      Make these take up less space on the page (the band pages) and make them less prominent. They are not as interesting.

    6. Play and study 18 versions of the 12-bar blues in any key, from the basic three-chord form to Charlie Parker's 'Bird' changes.

      If you can do so without too much clutter, maybe with a hover thing or a quick note, put an update date on each of these and explain which models/tools they were built with, such as "built with Claude Opus 5.5." Some of the tools below were built without any AI tools, and that should be noted as well. But please, please avoid clutter.

    7. plus the bands I play tuba and trumpet with

      I play tuba, trombone, and flugelhorn, not really trumpet. . Importantly, the band pages are not really interactive tools. Don't make those quite so prominent. Maybe those should just be mini links.

    1. AI-generatedThis is a model-generated prioritization rationale, not source text or a human assessment. Check factual claims against the source abstract and linked paper. (

      It's good to say that things are AI-generated, but this standout purple text is a bit distracting - make it gray.

    2. Stored

      Why do you say "stored" evaluation rationale - what is stored about it? You don't need to give so much under the hood here - at least not next to every entry - you can explain how we came up with this at the top or in linked pages.

    1. visible CM_01 point estimates

      Link and tooltip explain what this question is, also give an in-text abbreviated meaningful name for this question

    1. Model and planning decision. Retain the numerical forecasts and commitment ledger. Compare scenarios with faster grantmaker expansion and different cause allocations; do not add this ecosystem total to the AI-linked ledger. For an organization considering expansion, the useful next evidence is an eligible funding program, a decision timetable, and a grant it can actually budget against.

      This also requires some more context and explanation, perhaps with detail in additional fields or tooltips

    2. Evaluation capacity deserves attention. The overview counts only nine publicly accessible evaluators with at least three full-time staff in evaluation or grantmaking. This supports investigating capacity constraints, but headcount alone cannot calibrate our dollar cap.

      You're really not giving me any context here.

    1. 10 July 2026 revision: OpenAI GPT-5.6 Sol reviewed the Fable branch following the authors’ email discussion. Its additions and edits remain AI-generated and unreviewed by the authors. The revision preserves the earlier source text in Git history, corrects several overstatements, and adds a comparison with Epperson, Diederich, and Goeschl’s unit-donation experiment.

      Things like this could be tooltips.

    2. This is the “Fable-adapted” branch of the ICRC paper: consolidated, reorganized, and substantially extended by Claude (Fable 5) in July 2026, under David Reinstein’s direction. Every section carries a provenance tag:

      Use folding boxes more. There's too much detail. It's too cluttered.

    3. Impact Information

      "Impact information" is not completely accurate: 1. Per unit cost, not ultimate impact - food parcels, not lives saved or something like that. 2. We vary the inclusion of this information, but we don't vary the actual impact of the donation because we're only using a single donation to a single charity.... - Maybe this updates people's beliefs about about the impact

      In addition to providing evidence on the cost impact or the cost per output, at least, we are also framing it in a way that suggests to them that they are specifically buying an output rather than just donating to a broad pool, which reflects some of Epperson.

      What we do is not 'purely clean' perhaps but it is field-relevant

    4. that a well-designed unit-donation scheme can increase giving, especially with a large unit size. Their full treatment does more than report a unit cost: it reframes the decision as choosing physical units and can restrict choices to a unit grid

      This needs expansion or clarification. Perhaps in a footnote or tooltip’s

    5. It does not show that private philanthropy can replace government aid, or identify how to build political support for official development assistance

      This not that ai language. Adjust

    1. “DONATE 50CHF TODAY: your donation can supply 3 food parcels (ca. 17CHF/parcel for one month) to a Syrian family”

      consider doing a comparison between cost and cost_sug_50 because ??? @jan schmitz

    2. “DONATE TODAY: your donation can supply food parcels (ca. 17CHF/parcel for one month) to a Syrian family”

      this treatment relative to the control does 2 things -- tells them about hte cost, and makes it concrete what they are 'buying' w a donation

    3. The mailing went out on April 22, 2021. An unrelated ICRC “door-drop” campaign followed around May 25, and a TV spot ran in German-speaking Switzerland in the same period. This triggered the preregistration’s Case I contingency: the narrow window (gifts through May 31) is the principal sample for hypothesis tests, with the broad three-month window as descriptive robustness. The narrow flag keeps all letters and simply does not count gifts arriving after May 31.

      AI -- please make this language less AI sounding, and explain it more fully. Use tooltips for details

    1. I. Impact of providing information about unit cost in a fundraising solicitation — primary questions.

      strictly, this is providing information about 'unit cost'

    2. 3  Main results: The impact of impact information

      this page should link to or prominently show the actual (or translated) text of the varyuing treatments in the letter

    3. anipulation strength and relevance: the cost-per-output treatments must have been meaningful, salient, and representative of what charities actually do and would

      This is important and a natural criticism. I think we can make a credible claim of naturalness. We get minor support from “the other dimension of treatment did have an effect here”. But the biggest limiting criticism I see is “did donors even notice this in a meaningful way?”

    4. Our closest arm-level comparison holds the CHF 150 ask fixed. The control text asks for CHF 150 and mention

      Recall and reconsider our thinking: cost info means something different when you have a suggested donation vs when it’s open ended?

    5. scheme with a large unit size substantially increased average giving.

      Did they find the predictable incidence va amount trade off and nonlinearity? How to compare the size of the ask / size of “large” units across these contexts…? I guess there’s were typical lab experiment type stakes?

    6. not yet reviewed by the authors.

      I am going through it right now to see whether it makes sense in a general way, and adding a few comments and suggestions and questions. You will also want to selectively and then fully check things in a manual and “fully human” way.

      Aside: I suspect hybrid human ai research becomes more common soon, but for now I think people expect full human oversight

    1. Table 4.2: Preregistered amount contrasts (SUG-HI-LOW-AMT, H10–H11), narrow sample. Outcome: donation amount in CHF including zeros (revenue per letter).

      provide base rates in tables like this

    1. Know of high-impact research we should consider for evaluation — or other work this tool should be doing? Leave an email if you're open to follow-up discussion; if we later introduce compensation for useful contributions, earlier contributors will be grandfathered in.

      add a note here and elsewhere, along the lines of: We realize that you might have doubts about the value of adding content, ratings and comments on a largely AI-driven site. We'll try to personally respond to each substantial comment or engagement. And if you have doubts or want to flag us specifically, feel free to email contact@unjournal.org.

    1. Since the workshop: a company-reported production run at 22,000 litres and a new hydrolysate review bear directly on the scale and media questions. Dated updates and their limits.

      This comes across as strictly positive, but if you look at the evidence, I believe there was also some less positive news about finance.

    2. 9, 2026. These developments

      The developments reported below make it look like things have only gone in a positive direction, but I believe that on the finance side, things may have been scaled back? This might be adjusted a bit to not present the wrong impression.

    3. differed over how far particular substitutions could go,

      In a tooltip link direct quotes/sections - everything stated here needs some sort of attribution.

    1. Rate/discuss (Team) Rate/discuss (Public)

      Potentially, I could make these the same, just one box rather than two, but then people can fill in the "team options" within that interface.

    2. Unjournal commission an evaluation? (-- Strong No … ++ Strong Yes)For aggregation and sorting, the five quick ratings translate to percentile-equivalent scores: -- = 10, - = 30, ~ = 50, + = 70, and ++ = 90. These are deliberately spaced away from 0 and 100..

      make it clear that 1. these are translated to percentiles and 2. the "quick ratings" are less weighted/noted separately as 'quick ratings'

    1. The authors have made the atlas, methods, data, and replication code easy to inspect. We found some research indexing and light public discussion, but no major institution or policy report using the measures yet.

      this is what authors did. not quite the same -- also interesting but make that separate

    1. Several plausible routes could lower cultivated meat production costs.

      this first sentence seems a bit contentless. "Could" is vague and "lower" is not really quantified ... do we mean 'substantially lower' in some measured way? Bring closer to parity?

    2. Several plausible routes could lower cultivated meat production costs. The harder question is whether they can work together in a reliable process at commercial scale. Our current reading of the evidence favors testing complete, internally consistent production scenarios before treating lower ingredient prices or higher cell densities as a new cost forecast.

      This is an AI-generated (Astra, Extra-high effort) synthesis takeaway. Even in its initial statement, the claims should be sourced and linked. The "not this but that" aspect of this reads a bit like AI style. Quote: "Our current reading" -- it's not clear who "us" is. What would it mean to "test complete internally consistent production scenarios"? It's not clear what's being contrasted here - one wouldn't treat lower ingredient prices as a new cost forecast, for example, but they might enter into such a cost forecast.

      David Reinstien: To the extent this synthesis is making concrete statements. I'm not sure if I agree with them. At the very least, this should be stated in a much more contingent and tentative format.

    1. Framing note: Recent industry consolidation — including several company closures in 2024–2025 — is one set of data points, not a verdict on any particular TEA's methodology. Company outcomes depend on funding conditions, market timing, and management decisions that are largely independent of the underlying cost trajectory. The workshop aimed to evaluate the technical evidence on its merits.

      This should be a fold or a tooltip.

    2. Public working summary Workshop held May 8, 2026 Last reviewed July 20, 2026 S2 was off-record and is excluded. Its section below gives only the announced session scope.

      This isn't really a well-organized "public working summary." It goes into too much detail about the structure of the workshop rather than the overall narrative and a reasonable synthesis of the working results/takeaways.

    1. Informal Pre-Session Wed May 6, 2026 · 11:00am–12:00pm ET · Zoom · recorded Casual orientation for participants with limited May 8 availability — mostly preparation and broad discussion. Walkthrough of the beliefs form and cost model. David Manheim (Technion/ALTER) and Mirjam Capuder (University of Maribor) participated. It covered introductions, a walkthrough of the interactive cost model dashboard, and early framing questions about key modeling uncertainties — including Manheim's perspective on expert elicitation methodology and how to structure belief-updating around contested TEA assumptions. Substantive points from this session are incorporated where relevant in the S1 and S3 notes below.

      This is too much detail of the workshop's logistics here. Logistical stuff should be in folding boxes or tooltips.

    2. These map onto the workshop's structure: cell-line engineering → Crux 1; hydrolysate substitution → Crux 2; albumin and recombinant proteins → beliefs-form question E4; "what share of a hybrid product is cultivated?" → Crux 4.

      I don't see how this maps on at all - it seems like you're confusing things.

    3. Regulatory milestones are not directly comparable across jurisdictions. The FDA's February 2026 inventory lists five completed US cultured-cell food consultations, including three completed in 2025; market authorization, inspection, and product scope can still differ.

      I don't see that discussion in the evaluation 2-source it more directly and look for a quote.

    4. Evaluation 2 of this forecast

      That was not an evaluation of a specific forecast -- It was an evaluation of a paper involving an expert-driven forecasting exercise.

    5. rigorous

      "rigorous" is subjective. Think of a rephrasing here. Humbird was widely cited and influential. Pastika et al was published in Nature Food a peer reviewed journal that buy one metric ranks number 1 in it's subcategory and it has been heavily cited

    1. Tool updated: Aug 28, 2026Normal cadence: papers twice weekly; feedback hourlyPaper discovery, scoring, matching, and the full dashboard rebuild normally run on Monday and Thursday. New rating and suggestion submissions are checked up to hourly. The dated chips show the actual latest completed updates.Last paper scan: Aug 31, 2026Last AI scoring: Aug 31, 2026Crux/PQ match update: Jul 23, 2026Human feedback: 17 current ratings across 13 papersPublic anonymous aggregate refreshed Sep 1, 2026. Each rater's latest rating per paper counts once. Discussion and a display name are included only after explicit public-sharing consent; emails, rating keys, and submission ids are never included.Earlier API cost: at least ~$4.06 reconstructed

      less bold font here, this is taking up too much attention. It's good, but it's busy/distracting ... you can put all the updates dates in a more compact form, without so much bold. Actually, maybe fold these or put them in a tooltip ... just say "Last updated" and giv ehte date of the last paper scan, and then a tooltip with the other dates and the 'cadence' etc.

    1. A public economics column lists the review among its sources. This suggests some circulation outside the World Bank, although the article does not assess the review’s methods or health-policy conclusions in detail.

      it's given as a reference but I didn't see much discussion

    1. Mean absolute score movement?

      give a column for the overall MAD of scores to give a basis for cpmparison here ... to know if "6.3 is relatively large given how the scale is used' etc

    2. Agreement alpha?

      link a description and explanation of these measures in this context. What are normal and comparison values for this? What is the cardinal meaning of this measure ... 0.5 vs 0.25 eetc., and how to interpret it?

    3. rivate review found that one run treated prior public scrutiny as substantial while the other called it unclear, materially changing the estimated value of a new evaluation

      that's very interesting!

    4. Prior human-rated comparison papers were less stable on both dimensions: ordinary rerun movement averaged 8.5 points and the wording effect averaged 5.8 points. 2/6 (33%) had wording effects larger than ordinary rerun movement; wording-mean alpha was 0.69.

      this lower stability -- could it be somehow because of contamination from seeing human wordings, or because of greater variation in the nature of these papers?

    5. Recent AI-impact/governance papers: ordinary same-wording reruns differed by 4.1 points on average, while the repeat-averaged wording effect was 1.1 points. 0/6 (0%) had a wording effect larger than their ordinary rerun variation, and wording means produced no broad-tier changes. Prior human-rated comparison papers were less stable on both dimensions: ordinary rerun movement averaged 8.5 points and the wording effect averaged 5.8 points. 2/6 (33%) had wor

      worth discussing the pooled analysis too

    6. Recent AI-impact/governance papers: ordinary same-wording reruns differed by 4.1 points on average, while the repeat-averaged wording effect was 1.1 points. 0/6 (0%) had a wording effect larger than their ordinary rerun variation, and wording means produced no broad-tier changes.

      to me this suggests the wording variation was not particularly important -- is that reasonable? How does this compare to results in similar contexts?

    7. ordinary same-wording reruns differed by 4.1 points on average, while the repeat-averaged wording effect was 1.1 points.

      4.1 percentage points on the 0-100 scale? How does that compare to the typical rating variation between such papers -- drawing the reference group in a few reasonable ways

    8. Held-out check: a third baseline execution is kept outside the balanced headline analysis and used only as a robustness check.

      explain (tooltip?) why 'holdout' is important here. I'm not sure I understand what it's doing in this context

    9. What is being tested: every stability number below describes AI model scores. “Previously human-rated papers” names a paper cohort selected because earlier Unjournal ratings are available; those ratings do not enter the stability calculations. Human and AI scores are compared separately in the amber reference panel.

      make this a folding box (beyond the first line)

    10. paper-level human ratings.

      would be fine to do so -- why not? These are already public IIRC. I'm also not worried about sharing AI-generated ratings. Remember, we're rating for priorization/impact, not research quality

    11. Average the two runs within each reviewed, equivalent wording, then compare those wording means. This reduces ordinary rerun noise in the wording comparison.

      I got this upon re-read but was confused at first. Explain better. And is this the approach recommended in the apepr, or is there a better decomposition

    12. Run each wording twice with the paper, rubric, model, settings, and output schema fixed. Differences within a wording estimate ordinary run-to-run variation.

      Need some context here -- which models (Fable, Sol, etc) and effort levels are these being tested on?

    13. prompt variants

      i want to see the nature of these prompt variants early on, to get a sense of things. Show some examples concisely in a tooltip or fold

  4. Aug 2026
    1. A shared first pass: Browse research that may inform policy, funding, or further evaluation. Scores rank the value of attention or evaluation; they are not quality grades.

      add: metrics of mention on social media, citations in white papers, etc., ?altmetrics type metrics

    1. Living reviewFilter to papers surfaced by one or more expert-maintained living literature reviews. A paper keeps this provenance even if it originally entered through another academic source. All Reviews ▾ AI Accountability ReviewBridging BoundariesExistential CrunchLauren Policy: Migration Literature ReviewMonitoring Gene DrivesNew Things Under the SunRegression to the MeatReskilledScaling in Human SocietiesThe Care GapThe Patentist Review / UJ overlapFind review posts that cite papers The Unjournal has already evaluated, has in its pipeline, or previously considered. This is intended as an engagement and follow-up queue. All Any UJ overlap UJ evaluated In pipeline / considered Targeted intakeRequested, topic-focused additions outside the broad recurring scan. Use this to include all records, exclude deliberately oversampled batches, or inspect only one targeted run. Include allExclude targeted intakeOnly targeted intakeOnly: AI impacts on global health and development (29)Only: Animal-welfare focused intake (19)Only: Empirical conflict replication-game intake (14)Only: Large-N GCR and existential-risk quantitative evidence (6)Soil-invertebrate crux sweep (6 cruxes) Crux/PQFilter to papers that map to one or more Unjournal Pivotal Questions or community cruxes (from the cruxes explorer). "PQ match only" restricts to papers mapped to a Pivotal Question. All Has crux/PQ match PQ match only RecencyFilter papers by how recently they were released or last updated. Useful for focusing on the newest research that may benefit most from timely evaluation. Any time Last 30 days Last 3 months Last 6 months Last year Last 2 years Last 5 years ModelThe AI model used to score each paper. New papers receive a standard GPT-5.5 assessment; stronger candidates receive a second, more thorough GPT-5.5 assessment with higher reasoning effort. "UJ historical" papers were prioritized by the Unjournal team (not AI). Papers marked WAIT are temporarily excluded from AI ranking pending a genuine model reassessment. All Models ▾ UJ historicalawaiting genuine model assessmentclaude-sonnet-4-6gpt-5.4gpt-5.4-minigpt-5.5 (codex headless, high)gpt-5.5 (codex headless, medium)gpt-5.5 (codex headless, xhigh)manual_public_followupopus (headless) Also show small-model scored (mini/haiku) Also show UJ historical Also show legal scholarship Targeted intake catalog · 74 records · 5 search families

      this is too cluttered -- it should be a filter

    1. Source text (unverified type) source noteText supplied by the discovery source; its status as a formal abstract has not been verified.

      the paper has an abstract -- why isn't it being shown here?

    2. AI-assisted prioritization for The Unjournal evaluation

      Make it clear here very prominently (with some reminders in key places below) that these are not ratings about research quality -- they consider the potential direct or near-direct impact of the work from a global priorities perspective.

    3. he AI premium is concentrated in frontier and intensive AI use, including closed-source models, paying or seasoned users, and long prompts, rather than casual or open-weight use.

      How is this identified? Are they talking about 'firms whose returns correlate more with the aggregate use of these prompts in the AI consumption data"?

    4. interaction-and-communication skill content strongly associated with higher AI exposure, while analytical, scientific, and operations-control skills load more negatively.

      this seems particularly interesting -- does this havean implication for the labor market?

    5. predicts

      is 'predicts' the claim of interest? This is all backwards looking 'prediction' in a sense. is there an implied claim about future returns?

    6. The paper informs live decisions about AI labor-market adaptation, AI industrial policy, competition and market-concentration monitoring, financial-risk exposure to AI shocks, and prioritization of retraining or adjustment support across occupations. It i

      in what ways? Need more examples.

    7. This is a high-value Unjournal candidate: it is an NBER working paper using unusually granular proprietary data on realized AI consumption to estimate how AI demand is priced across firms, sectors, countries, and occupations. The findings could inform decisions

      "inform" in what ways? We want some examples

    1. r sufficiently strong, discussed public feedback—places the paper in a private SQLite-to-Coda queu

      tooltip the exact rule and try to keep it updated

    1. Quick-rate mode: for each paper, how strongly should The Unjournal commission an evaluation? (-- Strong No … ++ Strong Yes). Each click is saved to The Unjournal's server. Clicking another option for the same paper updates this browser's rating rather than adding another. Add your name or email once below; the detailed form will reuse it.

      One thing about the quick rate form - I don't want to encourage people to rate the paper simply based on its title. I'd like them to at least read the abstract and probably more. At the moment, we're bundling the detailed rating with seeing details on the paper - we should potentially adjust that or put some signposts so people don't get that impression.

    2. More filters and source controls DisciplineThe academic field or sub-discipline of the paper, as classified by the AI model. Used to filter by research methodology and domain expertise. All Fields ▾ AI & Data ScienceAnimal WelfareDevelopment & AgriculturalEconomicsEnvironmental & ClimateLabor, Education & HealthOtherPhilosophy & EthicsPolicy & GovernancePolitical Science & LawPsychology & BehavioralStats, Finance & Methods SourceHow the paper entered this dashboard. Source is a discovery route, not a publication venue, endorsement, or quality judgment.Unjournal database: imported from The Unjournal's existing evaluation and prioritization workflow.Targeted public-paper follow-up: an independently public paper or DOI found during a requested topic-focused search. It does not mean the author supplied or endorsed the score.Academic feeds: NBER, CEPR, and RePEc working-paper feeds; arXiv and SSRN preprints; OpenAlex and Semantic Scholar indexes.EA Forum: research linked from EA Forum posts.Research organizations: Anthropic, DeepMind, and selected AI governance or safety organizations.Legal sources: legal-scholarship searches covering OpenAlex Law, law reviews, and the Institute for Law & AI. All Sources ▾ AI Governance (arXiv)AI Safety OrgsAnthropic ResearchDeepMind ResearchEA ForumNBEROpenAlexRePEcSSRNSemantic ScholarTargeted OpenAlex intakeTargeted public-paper follow-upUnjournal databasearXiv Targeted intakeRequested, topic-focused additions outside the broad recurring scan. Use this to include all records, exclude deliberately oversampled batches, or inspect only one targeted run. Include allExclude targeted intakeOnly targeted intakeOnly: AI, global health, and development curation (3)Only: Animal-welfare focused intake (19)Only: Empirical conflict replication-game intake (14)Only: Large-N GCR and existential-risk quantitative evidence (6)Only: Transformative AI, global health, and wellbeing (16)Soil-invertebrate crux sweep (6 cruxes) Crux/PQFilter to papers that map to one or more Unjournal Pivotal Questions or community cruxes (from the cruxes explorer). "PQ match only" restricts to papers mapped to a Pivotal Question. All Has crux/PQ match PQ match only RecencyFilter papers by how recently they were released or last updated. Useful for focusing on the newest research that may benefit most from timely evaluation. Any time Last 30 days Last 3 months Last 6 months Last year Last 2 years Last 5 years ModelThe AI model used to score each paper. Papers are scored in tiers: a fast model (gpt-5.4-mini) screens all candidates, a stronger model (gpt-5.4) re-scores the top papers, and the most capable model (gpt-5.4-pro) provides detailed analysis of the highest-ranked papers. "UJ historical" papers were prioritized by the Unjournal team (not AI). All Models ▾ UJ historicalclaude-sonnet-4-6escalatedgpt-5.4gpt-5.4-minigpt-5.5gpt-5.5 (codex headless, high)gpt-5.5 (codex headless, medium)gpt-5.5 (codex headless, xhigh)manual_public_followupopus (headless) Also show small-model scored (mini/haiku) Also show UJ historical Also show legal scholarship Targeted intake and prioritization requests - 64 records across 5 paper search families and 1 crux search family This shared catalog records topic-focused searches outside the broad recurring scans. Initial and follow-up passes are combined below as one search family. A targeted label records how a record was found; it is not an endorsement, quality judgment, prevalence estimate, or completed Unjournal prioritization decision.Large-N GCR and existential-risk quantitative evidenceResearch papers (6 papers; 2026-07-30) Credible large-N, cross-country, historical, forecasting, survey, or other quantitative evidence directly bearing on global catastrophic or existential risks. Public OpenAlex search, quantitative-design triage, deduplication, and prioritization scoring. This is a separate follow-up to the empirical-conflict search. It asks whether global catastrophic and existential-risk questions have credible quantitative literatures with enough observations or repeated judgments to support evaluation or replication. It excludes primarily conceptual argument, technical engineering, and ordinary civil-conflict studies without a direct catastrophic-risk connection. Inclusion is not an endorsement or a completed Unjournal team decision. View 6 records →Transformative AI, global health, and wellbeingResearch papers (16 papers; 2026-07-24) Economic, social-science, and policy research relevant to global health and wellbeing under transformative AI, especially implications for lower-income countries. Search-pass details (2)Transformative AI, global health, and wellbeing follow-up: Second-wave public-paper search focused on first-run gaps, followed by deduplication and Codex subscription scoring. This second wave follows up the initial requested intake and deliberately targets gaps it left: public finance and tax-base resilience, distribution and social protection, AI market power, worker attitudes, and developing-economy structural transformation. It uses only independently public paper or institutional publication pages. Inclusion is not an endorsement or a completed Unjournal team decision.Transformative AI, global health, and wellbeing: Public OpenAlex search, theme-fit triage, deduplication, and Codex subscription scoring. The search follows the linked Coefficient Giving request for proposals but is an independent Unjournal intake exercise. It emphasizes non-catastrophic transformative-AI scenarios, labor markets, fiscal capacity, social protection, biomedical and health-system bottlenecks, development strategy, and aid allocation. Inclusion does not imply endorsement by Coefficient Giving or an Unjournal team decision. Coefficient Giving request for proposals. View 16 records →

      this fold should also mention the 'tartgeted intake' filter above

    1. The reasoning goes that if there is always a high level of background risk to humanity, then we should expect to go extinct soon anyway, which means the importance of avoiding any one particular risk is not as valuable as it may seem. For more details see the full report here.

      This seems rather intuitive to me, but it's asking a slightly different question than what the original phrasing might seem to imply.

      I think the initial intuition that more risk means more value of reducing risk, comes from the natural idea that effort spent reducing a particular risk will reduce that risk proportionally. So, spending effort on reducing risks from car crashes, malaria in Africa, or heart disease, all else equal, we yield more value than spending comparable effort on reducing the risks of bear attacks. I guess this is the "importance" part of the ITN paradigm.

      But of course, the benefit of reducing the risk of car crashes is lower if we are facing other impending doom. Let's say we see an asteroid coming toward the Earth, or the threat of incoming nuclear war is high.

  5. Jul 2026
    1. DisciplineThe academic field or sub-discipline of the paper, as classified by the AI model. Used to filter by research methodology and domain expertise. All Fields ▾ AI & Data ScienceAnimal WelfareDevelopment & AgriculturalEconomicsEnvironmental & ClimateLabor, Education & HealthOtherPhilosophy & EthicsPolicy & GovernancePolitical Science & LawPsychology & BehavioralStats, Finance & Methods SourceHow the paper entered this dashboard. Source is a discovery route, not a publication venue, endorsement, or quality judgment.Unjournal database: imported from The Unjournal's existing evaluation and prioritization workflow.Targeted public-paper follow-up: an independently public paper or DOI found during a requested topic-focused search. It does not mean the author supplied or endorsed the score.Academic feeds: NBER, CEPR, and RePEc working-paper feeds; arXiv and SSRN preprints; OpenAlex and Semantic Scholar indexes.EA Forum: research linked from EA Forum posts.Research organizations: Anthropic, DeepMind, and selected AI governance or safety organizations.Legal sources: legal-scholarship searches covering OpenAlex Law, law reviews, and the Institute for Law & AI. All Sources ▾ AI Governance (arXiv)AI Safety OrgsANIMAL_LAWANIMAL_LAW_REVIEWAnthropic ResearchDeepMind ResearchEA ForumLaw & AI InstituteLaw reviewNBERNEP_LAWOpenAlexOpenAlex LawRePEcSSRNSSRN_LAWSemantic ScholarTargeted OpenAlex intakeTargeted public-paper follow-upUJ seedUnjournal databasearXiv Targeted intakeRequested, topic-focused additions outside the broad recurring scan. Use this to include all records, exclude deliberately oversampled batches, or inspect only one targeted run. Include allExclude targeted intakeOnly targeted intakeOnly: AI, global health, and development curation (3)Only: Animal-welfare focused intake (19)Only: Empirical conflict replication-game intake (14)Only: Large-N GCR and existential-risk quantitative evidence (6)Only: Transformative AI, global health, and wellbeing (16)Soil-invertebrate crux sweep (6 cruxes) Crux/PQFilter to papers that map to one or more Unjournal Pivotal Questions or community cruxes (from the cruxes explorer). "PQ match only" restricts to papers mapped to a Pivotal Question. All Has crux/PQ match PQ match only RecencyFilter papers by how recently they were released or last updated. Useful for focusing on the newest research that may benefit most from timely evaluation. Any time Last 30 days Last 3 months Last 6 months Last year Last 2 years Last 5 years

      make it easier/more prominent to 'clear all filters'

    1. 5. Three problems that a better estimate of the exchange rate won't fix

      wait -- I don't think that's logically correct. If indeed we have a good estimate of the exchange rate, isn't it, by default, embodying that these problems are in some sense solved?

    2. 3. What this is based on

      second column in table below should be narrower, firrst one wider. Use the space more carefully (and make that a persistent pattern/skill)

    3. Preliminary — do not read Unreviewed working draft, 30 July 2026. Not for circulation, citation, or quotation.

      "do not read" is too strong. Make more caveats instead ... it's a current working space, being continually adjusted. And it has not yet been reviewed by workshop participants or research contractors (something we aim to do soon)

    4. 2. Bottom line

      we need more caveats on this. It has not yet been reviewed by workshop participants or research contractors (something we aim to do soon)

    1. Optimized messaging shifts choices towards plant-based foodsource titleExact title resolved from the linked paper's bibliographic metadata; discovery-page anchor text and model output are not used as titles.

      Where is the abstract here? I don't see it.

    2. Strong near-term candidate for animal-welfare evaluation because it tests behaviorally relevant messaging around plant-b...

      What's this text doing here? I don't see it at the top in the other entries.

    3. 2 ratings

      Is there a way that someone could change their rating? It seems like every time I come to the page and rate it, it records separately.

    4. abstract provenanceSource-supplied abstract wording; HTML entities and whitespace normalized. Not independently compared with the paper PDF.

      This tooltip seems rather process-oriented and not particularly helpful for users of this page. You certainly don't need to give it every time, and I'm not quite sure what it means or how to interpret

    5. Rate on a 0–100 percentile scale relative to all papers in our database (Strong Yes=top quintile, Strong No=bottom). Try to form your opinion before reading the AI discussion above. Discuss why you gave that rating.

      This percentile scale instruction is not consistent with the negative-neutral-plus scale shown below. Let them give a percentile first, and if they don't want to do that, you can let them give the priority ratings with the negative-positive scale, or just save that for the quick raters.

    6. 2 ratings

      Seems like it was double counted here. I rated it with the quick rate and then with the detailed rating. It doesn't seem quite right.