How Cloudflare’s CryptoLabe uses AI to map a post‑quantum migration

How Cloudflare’s CryptoLabe uses AI to map a post‑quantum migration

Cloudflare is aiming for a full post‑quantum (PQ) rollout by 2029. To keep track of where cryptography lives in its monolithic codebase, the company built an internal AI tool called CryptoLabe. The effort matters because any missed algorithm could become a weakness once quantum computers become practical.

The problem of finding crypto at scale

Cloudflare’s code lives in a single source‑control platform, which sounds like a shortcut for discovery. In practice the code is split across hundreds of repositories, hidden behind libraries, configuration files, and even dead branches. Simple greps for strings such as “RSA” or “X25519” either over‑count (by catching unused code) or under‑count (by missing defaults and indirect calls). Moreover, the same algorithm can be used in very different protocols—TLS, JWT, SSH—each with its own migration path. The company needed a way to answer three questions:

  1. What cryptographic primitives are in use?
  2. How are they being used?
  3. Which parts of the stack lack a PQ alternative?

CryptoLabe’s two‑stage AI workflow

CryptoLabe runs a discovery stage that maps a repository and extracts raw observations from source files, manifests, lockfiles, scripts, tests, and internal docs. Those observations feed a second analysis stage where a large‑language model (LLM) re‑examines the code, follows references across repositories, and checks runtime contexts. The model then assigns a classification or flags the finding as “More evidence needed”, “External dependency”, or “Unknown”.

The process is orchestrated with Cloudflare Workers, Durable Objects, and Workers AI. A scan starts from a dashboard request, passes through a scanner Worker, and stores a snapshot of the repository in R2 (Cloudflare’s object storage). Each analysis step runs in an isolated sandbox, ensuring the model never writes to the live code.

Classifications and what they mean

Classification Typical use case Migration hint
Classical encryption ECDHE (X25519, P‑256, P‑384), RSA key agreement, HPKE Replace with hybrid or pure PQ key exchange
Classical signature RSA or ECDSA signatures in certificates, TLS handshakes, JWTs Move to ML‑DSA or other PQ signature scheme
Classical token RS256/ES256 JWTs Follow RFC 9964 for a PQ replacement
PQ‑ready hybrid key exchange X25519MLKEM768 in TLS 1.3 Already hybrid; monitor for pure PQ alternatives
PQ‑ready Stand‑alone PQ algorithms like ML‑DSA outside TLS Verify deployment and performance impact

The table shows that most of Cloudflare’s current PQ work lives in hybrid TLS key exchange, while signatures and tokens remain classical.

The hidden cost of scaling AI‑driven code scans

Running an LLM over thousands of repositories is cheap per request but expensive in aggregate. Cloudflare routes each model call through AI Gateway, which lets them swap models when cheaper options appear. The real bottleneck showed up as HTTP 429 rate‑limit responses when many scans fired simultaneously. The team solved this by introducing a global Durable Object that throttles every request, forcing all scans to share a common cooldown.

The trade‑off is clear: you gain visibility but you must build a thin orchestration layer to avoid throttling spikes. The extra plumbing adds operational overhead and creates a single point of failure—if the pacing object crashes, all scans stall. In practice this means you should budget time for reliability testing before relying on AI for production‑grade audits.

What we can learn for our own codebases

  1. Centralized repos simplify discovery but do not eliminate hidden dependencies. Even with a single platform, configuration files in separate repos can drive cryptographic choices.
  2. AI can bridge the gap between pattern matching and contextual understanding. The model’s ability to follow references across files reduces false positives.
  3. Classification must be conservative. CryptoLabe prefers “More evidence needed” over guessing, which keeps the report trustworthy.
  4. Rate‑limit handling is a first‑class concern. A shared back‑off mechanism prevents a cascade of retries that would otherwise overwhelm the AI service.

Quick experiment you can run today

If you have a modest codebase on GitHub, try the following steps:

  1. Clone the repository locally.
  2. Use an open‑source LLM (e.g., Llama 3) with a simple prompt that lists files and asks the model to identify any lines containing known crypto primitives.
  3. Feed the model the results of a grep -R "RSA\|ECDSA\|X25519" as context, then ask it to classify each hit as encryption, signature, or token.
  4. Record the classifications in a CSV and compare them to a manual review of a few random hits.

Even a rough pass will highlight whether AI adds value beyond plain grepping for your project.

Sources

Read next

We count page views without cookies — no identifier, nothing stored on your device. Accept to allow cookies for analytics.