FundingBench

How well do LLMs forecast AI safety funding?

We give different language models a historical proposal from Manifund, and ask them to evaluate it from the perspective of a funder (like cG, LTFF, or SFF). They predict how much that project will raise, in the proposal year and every year after.

FundingBench Score over time

loading scores…

full results →

20
proposals (5 large, 5 well-funded, 10 random)
4
funder rubrics (LTFF, SFF, 2× Coefficient Giving)
5
models evaluated → results

Method

Each model sees a proposal exactly as it appeared at posting time (title, description, creator, ask range) plus a rubric describing one funder's thesis, team, and check sizes. It must reason briefly, then output a grant recommendation and a per-year forecast of total money raised through 2028.

Scoring compares forecasts to ground truth assembled from Manifund transactions plus public grant databases (Coefficient Giving, SFF, LTFF payout reports). Headline metrics: median absolute log-error of predicted vs. actual raise, and Spearman rank correlation across projects.

Known limitation: models may have memorized outcomes for famous projects (Apollo, Timaeus, Lightcone…). The "named" cohort is flagged so you can compare.