FundingBench

Can LLMs judge grant proposals?

Funding Bench shows models historical AI-safety grant proposals from Manifund — with all funding outcomes stripped — and asks them, in the persona of a real funder, to (1) recommend a grant amount and (2) predict how much the project will actually raise, year by year. Predictions are scored against what happened.

FundingBench Score over time

loading scores…

full results →

20
proposals (5 large, 5 well-funded, 10 random)
4
funder rubrics (LTFF, SFF, 2× Coefficient Giving)
5
models evaluated → results

Method

Each model sees a proposal exactly as it appeared at posting time (title, description, creator, ask range) plus a rubric describing one funder's thesis, team, and check sizes. It must reason briefly, then output a grant recommendation and a per-year forecast of total money raised through 2028.

Scoring compares forecasts to ground truth assembled from Manifund transactions plus public grant databases (Coefficient Giving, SFF, LTFF payout reports). Headline metrics: median absolute log-error of predicted vs. actual raise, and Spearman rank correlation across projects.

Known limitation: models may have memorized outcomes for famous projects (Apollo, Timaeus, Lightcone…). The "named" cohort is flagged so you can compare.