Researchers have introduced CryptanalysisBench, a new evaluation framework designed to measure the ability of frontier large language models to discover mathematical vulnerabilities in cryptographic algorithms and security protocols.
Key Points
- CryptanalysisBench includes 191 tasks across six cryptographic primitive families, such as block ciphers and hash functions, sourced from NIST standardization competitions.
- Five frontier models, including Claude Opus and GPT-5.5, successfully broke up to 86% of Tier 1 schemes with known historical vulnerabilities.
- The models identified previously unknown security flaws, including a key-recovery attack on the SpoC AEAD and a proof error in the KINDI protocol.
- Anthropic utilized the benchmark to identify new vulnerabilities within the Hawk algorithm and reduced-round AES.
- The benchmark is intended to serve as a stress-testing tool for evaluating the security of cryptographic schemes before they are deployed in production environments.