Five researchers at UC San Diego and Inria Nancy ran a computation that took five calendar months and 1,380 CPU core-years, and whose active ingredient was asking a hardware security module to sign numbers. Four billion, two hundred and ninety-four million of them, give or take. At the end of it they could forge any signature they liked under that module’s 1,024-bit RSA key, for ever, offline, without the key ever leaving the box and without factoring the modulus. The paper went up on the IACR Cryptology ePrint Archive on 20 September as 2026/2131, “Forging 1024-bit RSA signatures in nearly SNFS time”.
Nothing malfunctioned. The module was tamper-resistant, the key stayed inside, and every one of those billions of answers was correct, all exactly as advertised. The attack is the volume of correct answers, and the number of questions a key will answer appears on no certificate, in no audit report, and in the marketing for no product in this industry, ours included.
TL;DR
- Laura Shea, Miro Haller, Adam Suhl, Nadia Heninger and Emmanuel Thomé implemented a 2007 algorithm of Joux, Naccache and Thomé (yes, the same Thomé) to forge 1,024-bit RSA signatures: 1,380 CPU core-years and 232 oracle queries, against the 500,000 to 1,000,000 core-years it would take to factor the same modulus.
- They ran the queries through a real Thales Luna HSM. In their words, this demonstrates “the ability to impersonate the HSM through black-box API interactions, without exfiltrating the key”.
- What gets repriced is the label, not the key: RSA with a signing oracle is 15 to 30 bits weaker than the factoring-based estimates everyone quotes, and “even 4096-bit RSA does not appear to meet a 128-bit security level in this attack model”.
- The paper’s own Privacy Pass arithmetic is the lesson: identical key, identical mathematics, and the attack costs 17 million years, 2.3 days or $13 trillion depending purely on how the interface meters requests.
- One of the few concrete reasons to enable the dangerous raw-RSA mode is to run a padding scheme the hardware lacks. RSA-FDH is on that list, and RSA-FDH is how RFC 9381 builds an RSA-based verifiable random function.
- A provably fair game is a query interface to a secret. Satoshie’s queries are transactions: priced, counted by the chain, impossible for us to hide or to serve for free.
What they actually did
The threat model has a name from the 1990s: a lunchtime attack. You get temporary access to a signing or decryption oracle for somebody’s key, you lose that access, and then you are challenged to produce a forgery. The paper’s finding is that temporary access converts into a permanent capability. As the authors put it, the attacker “can effectively use oracle queries to steal abilities equivalent to possessing the private key, without actually computing it”.
It runs in two stages. A precomputation depending only on the public modulus took about 1,200 core-years; the query phase needed roughly 232 raw, unpadded signatures. After that, each forgery is offline, costs about 180 core-years and is repeatable at will against any target. Set against the 500,000 to 1,000,000 core-years current estimates give for factoring a 1,024-bit modulus outright, the whole attack lands at well under half of one per cent of the cost of the thing everyone treats as the security bound.
The hardware should unsettle custodians. They ran a Luna K6 on a PCIe slot and a remote Luna S750, measuring 6,500 signatures per second at 40 threads, the thread count the vendor’s own administration guide recommends. At that rate 232 queries is roughly seven and a half days of continuous signing, which is my arithmetic rather than theirs, and it is the only loud thing in the entire attack. A week of nothing but raw unpadded requests is visible to anybody counting. Nobody counts.
The authors are careful, and any honest write-up has to carry their caveats. They disabled FIPS 140-2 mode on their own device to enable raw RSA and used a key they generated themselves; they call the attack “practical in an academic sense rather than in the script-kiddie sense”; and they say it is not cause for immediate alarm in anyone’s cryptographic inventory. Their conclusion is about parameters and assumptions, in a window where NIST already plans to deprecate RSA by 2030 and disallow it by 2035.
A security level is a claim about an attacker, not about a key
Here is what to take away even if you never touch RSA. 1,024-bit RSA is “80-bit security”; this attack runs in about 265. 2,048-bit RSA is “112-bit security”; scaled up, the attack is 290 with 243 queries. 4,096-bit RSA is sold as comfortably past 128 bits, and here it extrapolates to 2119 with 257 queries.
No key changed. No implementation was buggy. No operator picked a wrong number. What changed is the sentence that comes before the bit count, the one that says what the adversary is assumed to be able to do, and that sentence is almost never printed next to the number it governs. A key size is a fact about mathematics. A security level is a claim about an attacker, and every real deployment quietly edits that claim by handing the attacker an interface, then keeps quoting the old figure.
The PKCS#11 standard that nearly every HSM implements “provides an interface for raw RSA operations that exactly matches the oracle query functionality required by our attack”. Not a bug, not a bypass. The standard interface. FIPS 140-3, ISO/IEC 19790 and Common Criteria, meanwhile, specify tamper resistance and side-channel countermeasures: they specify the walls. The attack walked through the door, and the door is meant to be there.
The same key, three security levels
The clearest demonstration in the paper is not the HSM at all. It is blind RSA, which hands out a signing oracle on purpose, because that is the entire point of the primitive: a server signs something it cannot see. Privacy Pass, the protocol that proves you passed a CAPTCHA without revealing who you are, runs blind RSA in production at Cloudflare, Fastly, Persona and Apple.
So the authors priced the 243 queries a 2,048-bit attack needs against each deployment. Cloudflare’s original implementation issued 30 tokens per CAPTCHA, so you would need about 238 CAPTCHAs inside one key-rotation epoch, which is less absurd than it sounds against a company that says it handles more than seven trillion HTTP requests a day, roughly 243. Apple appears to rate-limit issuance to one token per minute per device: 17 million years on one device, 2.3 days across 2.3 billion of them. Persona charges $1.50 per API call, so the same queries cost $13 trillion. GNU Taler mints coins with blind RSA signatures per denomination, so forging one for ever means first minting about $88 billion of one-cent coins, a figure that checks out exactly when you multiply it through.
Same algorithm. Same key length. Same mathematics. Four wildly different security levels, and the thing that sets them is a rate limit, a price per call, a coin denomination and a device count. Those are cryptographic parameters. None of them appears in a cipher suite, a certificate, a key size or an audit scope.
The reason someone would switch the dangerous mode on
The obvious objection is that no sane operator runs raw unpadded RSA, so this is a laboratory result about a configuration nobody ships. The paper anticipates it, and its answer is more uncomfortable than the objection. You enable raw mode when you need a padding scheme the hardware does not implement natively: ISO 9796-2, still used in e-passports and in the digital tachographs fitted to professional vehicles; ANSI X9.31; and RSA-FDH, which the authors note is “used in GNU Taler and RFC 9381 for an RSA-based verifiable random function”.
RFC 9381 is the VRF standard. Section 4 of it is RSA-FDH-VRF, with ciphersuites for SHA-256, SHA-384 and SHA-512, sitting alongside the elliptic-curve ones. Read that chain once more, slowly. One of the small number of concrete, documented reasons an engineer would walk into the HSM settings and turn off FIPS mode is that they are building verifiable randomness on hardware that does not know how to produce it natively.
Be precise, because the honest version is narrower than the scary one. Full domain hash hashes the message first, so an attacker cannot simply ask for signatures on the smooth numbers the algorithm wants, and the paper is explicit that a padded oracle gives little asymptotic speedup over factoring. The VRF construction is not the vulnerability. The switch you flick to get the VRF is. The fairness feature and the attack surface arrive in the same commit.
Something this blog has argued twice, which this paper corrects
When Coldcard’s seed-generation flaw let attackers sweep 594 BTC, we argued that randomness is the only security primitive whose failure output is indistinguishable from success, on the grounds that signatures reject, decryption garbles, and only weak entropy quietly returns an ordinary-looking number. We had said much the same a few days earlier, when an AI system weakened a post-quantum signature candidate in about 60 hours: cryptographic systems fail loudly, server-side RNG fails silently. This paper puts a hole in half of that.
A forged signature verifies. Against the real public key, for the real modulus, with the real exponent, and nothing in the stack raises a flag, because there is nothing wrong with it. It leaves no artefact, and neither does a completed precomputation, which runs on the attacker’s hardware and asks your key for nothing. So here is the amended claim, which is more useful than the original: cryptography fails loudly when the attacker has to guess, and silently when the attacker was allowed to ask. Randomness is still the worst case, because a game outcome leaves no forensic record at all. But “signatures reject” holds only in a threat model where the adversary never had an account.
Every fairness scheme is a query interface to a secret
Strip a commit-and-reveal casino down to its mechanism and you get a server holding a secret seed, a player supplying a client seed and a nonce, and an HMAC per round. That is a signing oracle with better branding: the player chooses part of the input, the house holds the key, the output comes back on demand, and the number of outputs the house will produce under one secret appears nowhere on the fairness page.
Which produces a question this industry has never been asked. Not “is the scheme sound”, which is answerable and usually answered, but: how many outcomes will you generate under a single secret, and who is counting? Rotating the server seed is the only query budget the industry actually operates, and it is sold as a privacy nicety for suspicious players rather than as what it is, the one control bounding how much of the generator anyone can observe.
Then look at demo mode. Free play runs the same generator as the real game on most platforms, because building a second one is work and the point of a demo is that it behaves like the product. The paying player is metered by their own wallet. The demo player is metered by nothing at all. An operator who would never let you place ten million bets will happily serve you ten million free spins and call it marketing, and if the security argument depends on an attacker not seeing many outputs, that is an unmetered oracle to your own secret with a “play for fun” button on it.
What we can claim, narrowly
Satoshie’s randomness comes from Chainlink VRF, and the property that matters for today’s story is not cryptographic at all. It is that every request is a transaction. Chainlink’s documentation describes the mechanism plainly: “For each request, Chainlink VRF generates one or more random values and cryptographic proof of how those values were determined. The proof is published and verified onchain before any consuming applications can use it.” A request costs gas, settles in a block, and is counted by the chain whether we like it or not.
That gives us three things the Luna HSM did not have. The query count is public, so the number nobody could see in San Diego is one we could not hide if we wanted to. Every query is priced, so there is no free path to the generator, the same control Persona gets from charging $1.50 a call, except that we neither set it nor can waive it. And no demo mode points at the live generator, because a draw that did not escrow a stake is not a draw.
The rest is unchanged: odds a function of entry count in deployed code you can read before staking, a prize escrowed at entry rather than quoted on a page, resolution and payout in one transaction, no admin key over a draw in flight. On the mathematics the authors are blunt: “Our results do not carry over to other popular signature algorithms such as ECDSA or Ed25519.”
Three honest limits
First, our safety on this axis is inherited rather than earned. We are on the right side of this paper because we chose a provider whose scheme is elliptic-curve based, not because we analysed the query behaviour of competing VRF constructions. That is a procurement decision in a cryptography decision’s clothes, and as we said when quantum-safe Bitcoin got 79% cheaper in a week, immutability means we have no crypto agility. If an equivalent surprise lands on the curve, you read about it when we do, and we cannot patch.
Second, our query log is permanent and public. Every proof we have ever produced sits on Base for ever, readable by anyone, which is exactly the transparency we sell, and in any attack model that strengthens with the number of observed outputs it is also a corpus that only grows and that we could not withdraw. Publishing everything is a one-way decision, and we made it.
Third, we publish no usage budget either, and could not enforce one. We do not hold the VRF key, so we cannot rotate it, which was the authors’ own short-term advice to blind-RSA deployments, and we do not track or disclose a per-key output count. The argument of this post is that a secret without a counter is a secret being spent, and our answer is that ours is spent in public, not that it is spent carefully. Those are different claims and we would rather make the true one. It is the same shape of gap as CCIP 2.0’s optional second verifier: a number that decides a security level, living somewhere nobody thinks to read.
Three questions to ask anything you play on
- How many outcomes will your generator produce under one secret, and who counts them? If nobody can answer, the security argument has an unbounded variable in it.
- Does the free version use the same generator as the paid one? If it does, the house has published an unmetered oracle to its own secret and called it a demo.
- If an outcome had been forged and had verified correctly, what would look different? “Nothing” is an acceptable answer only if detection was never part of the design, and it usually was not.
The Luna HSM in that lab did everything it was certified to do. It protected the key from extraction, it resisted tampering, it never leaked a bit of secret material, and it answered every question correctly, four billion times, at 6,500 answers a second. Correctness was never the property under attack. The number of times it was willing to be correct was. The only defence against that, in a vault or in a casino, is a number somebody decides to write down before the questions start arriving, and it is also the number that caps how much business you can do. That is why nobody sets it.
📷 Photo by Alexandre Daoust on Unsplash


