Can LLMs judge grant proposals?
Funding Bench shows models historical AI-safety grant proposals from Manifund — with all funding outcomes stripped — and asks them, in the persona of a real funder, to (1) recommend a grant amount and (2) predict how much the project will actually raise, year by year. Predictions are scored against what happened.
FundingBench Score over time
loading scores…
Method
Each model sees a proposal exactly as it appeared at posting time (title, description, creator, ask range) plus a rubric describing one funder's thesis, team, and check sizes. It must reason briefly, then output a grant recommendation and a per-year forecast of total money raised through 2028.
Scoring compares forecasts to ground truth assembled from Manifund transactions plus public grant databases (Coefficient Giving, SFF, LTFF payout reports). Headline metrics: median absolute log-error of predicted vs. actual raise, and Spearman rank correlation across projects.
Known limitation: models may have memorized outcomes for famous projects (Apollo, Timaeus, Lightcone…). The "named" cohort is flagged so you can compare.