Flex-RSI: Benchmarking Recursive Self-Improvement of Frontier Agents
Overview: A benchmark for how well frontier agents improve themselves: given a task, the agent iteratively refines its own solution, tools and prompts to raise its score within a fixed budget. Tasks span visual perception and reasoning, robot control, 3D reconstruction and more, and we are actively adding tasks, tracks and models.













