AUTO-UPDATED

Measuring LLMs’ Ability to Perform Cryptanalysis

Researchers have introduced CryptanalysisBench, a new evaluation framework designed to measure the ability of frontier large language models to discover mathematical vulnerabilities in cryptographic algorithms and security protocols.

Key Points

  • CryptanalysisBench includes 191 tasks across six cryptographic primitive families, such as block ciphers and hash functions, sourced from NIST standardization competitions.
  • Five frontier models, including Claude Opus and GPT-5.5, successfully broke up to 86% of Tier 1 schemes with known historical vulnerabilities.
  • The models identified previously unknown security flaws, including a key-recovery attack on the SpoC AEAD and a proof error in the KINDI protocol.
  • Anthropic utilized the benchmark to identify new vulnerabilities within the Hawk algorithm and reduced-round AES.
  • The benchmark is intended to serve as a stress-testing tool for evaluating the security of cryptographic schemes before they are deployed in production environments.

Why it Matters

This development signals that artificial intelligence is rapidly approaching the capability to perform complex cryptanalysis that matches or exceeds current human-led research. As these models become more proficient, they will serve as critical tools for both identifying security flaws in digital infrastructure and potentially challenging the integrity of existing encryption standards.
Schneier.com Published by Bruce Schneier
Read original