Researchers have developed CivBench, a new evaluation framework using the strategy game Civilization VI to test the long-term reasoning and decision-making capabilities of frontier artificial intelligence models.
Key Points
- CivBench utilizes 76 custom tools to allow AI agents to interact with the Civilization VI game engine through text-based commands.
- The project identifies the "sensorium effect," where AI agents fail to perceive threats they do not explicitly query, and a "knowing-doing gap" where models fail to execute planned strategies.
- Testing revealed that models often struggle with long-horizon planning, frequently ignoring rival victory conditions even when provided with the necessary data.
- The benchmark includes a diary system that acts as external memory, allowing agents to maintain strategic continuity across hundreds of game turns.
- The framework is open-source and designed to help researchers measure how AI systems handle complex, multi-variable decision-making under uncertainty.