Abstract
Existing agent benchmarks emphasize tool use, web navigation, and software editing, but rarely isolate repeated operating decisions under scarce resources and delayed consequences. We introduce FounderBench, a controlled benchmark for startup-like sequential decision making. Given a simulated company state and noisy market information, an agent must choose from 13 structured business actions over a fixed horizon; a deterministic simulator updates cash, demand, quality, and risk, and task-specific code maps final outcomes to a 0–100 score. Free-form rationale is logged for audit but never scored. The suite contains 50 synthetic public tasks balanced across 10 operating families. Headline gap: a task-aware calibration policy averages 80.90 (37/50 solved), while the best hosted single run reaches only 67.69 (32/50); no human-founder baseline is reported. Rankings reverse between average score and solved count, and family profiles expose distinct operating failures. We release the suite, simulator and scoring code, adapters, validated task-level results, and regeneration scripts at https://github.com/louiswang524/FounderBench. All hosted rows are single runs on visible public tasks; scores do not establish real-world startup competence.
Keywords
Illustration
Citation
@article{Wang2026FounderBench,
title={FounderBench: Evaluating LLM Agents on Sequential Startup Decisions},
author={Yufeng Wang},
year={2026},
url={https://cspaper.org/openprint/20260723.0002v1},
journal={OpenPrint:20260723.0002v1}
}Version History
| Version | Released Date | Submitter |
|---|---|---|
v1Current | Jul 23, 2026 | Louis Wang |
