How well do LLMs forecast AI safety funding?
We give different language models a historical proposal from Manifund, and ask them to evaluate it from the perspective of a funder (like cG, LTFF, or SFF). They predict how much that project will raise, in the proposal year and every year after.
FundingBench Score over time
loading scores…
Method
Each model sees a proposal exactly as it appeared at posting time (title, description, creator, ask range) plus a rubric describing one funder's thesis, team, and check sizes. It must reason briefly, then output a grant recommendation and a per-year forecast of total money raised through 2028.
Scoring compares forecasts to ground truth assembled from Manifund transactions plus public grant databases (Coefficient Giving, SFF, LTFF payout reports). Headline metrics: median absolute log-error of predicted vs. actual raise, and Spearman rank correlation across projects.
Known limitation: models may have memorized outcomes for famous projects (Apollo, Timaeus, Lightcone…). The "named" cohort is flagged so you can compare.