Frozen Heart: How Incomplete Fiat-Shamir Transcripts Broke the Soundness of Production Zero-Knowledge Proof Systems

The Fiat-Shamir transform is the standard method for turning an interactive public-coin proof into a non-interactive one. It replaces the verifier’s random challenge with the output of a hash function evaluated over a transcript. The security of the resulting proof depends on what goes into that hash. If the transcript omits the statement being proved, the transform is called weak, and the proof system can lose soundness entirely. This is not a theoretical edge case. A 2023 survey of open-source implementations found 36 weak Fiat-Shamir implementations affecting 12 different proof systems, with novel knowledge-soundness attacks demonstrated against Bulletproofs, Plonk, Spartan, and Wesolowski’s VDF (Dao, Miller, Wright, Grubbs, ePrint 2023/691).

What the weak/strong distinction actually is

The distinction is stated precisely in Bernhard, Pereira, and Warinschi’s 2016 paper on the Fiat-Shamir heuristic:

Both variants start with the prover making a commitment. The strong variant then hashes both the commitment and the statement to be proved, whereas the weak variant hashes only the commitment.

That is the entire difference. One extra input to a hash function. The paper’s title — “How not to Prove Yourself” — is a fair summary of the consequence: in situations where malicious provers can select their statements adaptively, the weak Fiat-Shamir transformation yields unsound and unextractable proofs.

The mechanism is straightforward. In an interactive proof, the verifier’s challenge is chosen after the prover commits, and the verifier knows which statement is being proved. In the non-interactive transform, the challenge is derived from the transcript. If the statement is not in the transcript, a malicious prover can choose a commitment, observe the resulting challenge, and then select a statement for which that challenge happens to satisfy the verification equation. The prover is no longer bound to a statement at the time the challenge is fixed. Soundness, which requires that a prover cannot produce a valid proof for a false statement, collapses.

Why adaptive statement selection is the normal case, not an edge case

The 2016 paper makes a point that protocol reviewers should internalize:

Yet such settings naturally occur in systems when zero-knowledge proofs are used to enforce honest behavior.

This is the load-bearing observation. Zero-knowledge proofs in deployed systems are rarely used as academic demonstrations of knowledge. They are used to constrain what a party can do next. A proof gates a state transition: a withdrawal, a tally, a mint, a bridge release. The party producing the proof has an incentive to find a statement that passes verification while violating the intended constraint. That is exactly the adaptive-statement setting in which weak Fiat-Shamir fails.

The Helios voting system is the canonical historical case. The 2016 paper shows that using weak Fiat-Shamir in Helios leads to several possible security breaches: for some standard types of elections, under plausible circumstances, malicious parties can cause the tallying procedure to run indefinitely and even tamper with the result of the election. Helios is not a blockchain system, but the structural pattern — a proof used to enforce honest behavior by a party with something to gain — is identical to the pattern in proof-gated on-chain systems.

The 2023 survey: the bug class survived into modern proof systems

The 2016 paper dealt with classic protocols. The natural question was whether modern proof systems, built with different arithmetization and commitment schemes, had avoided the same mistake. The 2023 survey answered that question:

We perform a survey of open-source implementations and find 36 weak F-S implementations affecting 12 different proof systems. For four of these — Bulletproofs, Plonk, Spartan, and Wesolowski’s VDF — we develop novel knowledge soundness attacks accompanied by rigorous proofs of their efficacy.

Two things are worth separating here. First, the count: 36 implementations across 12 proof systems. Second, the fact that the authors developed new attacks, not just re-applied known ones. The weak Fiat-Shamir failure mode is not a single bug that was fixed once. It reappears in new codebases because the underlying specification practice — what exactly goes into the transcript hash — is often left implicit.

The survey also notes that prior work had shown weak Fiat-Shamir can break classic protocols like Schnorr’s discrete log proof. The 2023 contribution is showing that the same class of error persists in systems deployed today.

The blockchain stake is concrete. The 2023 paper states that a weak Fiat-Shamir vulnerability could have led to the creation of unlimited currency in a private blockchain protocol. That is a potential impact described by the authors, not a confirmed mainnet exploit. The distinction matters: the evidence establishes that the vulnerability class is present in deployed code and that the consequences can be severe, not that a specific public chain was drained.

What strong Fiat-Shamir actually buys

The 2016 paper does not stop at describing the failure. It defines a form of adaptive security for zero-knowledge proofs in the random oracle model — essentially simulation-sound extractability — and shows that strong Fiat-Shamir yields secure non-interactive proofs. In the Helios setting, the authors further show that strong proofs achieve non-malleable encryption and satisfy ballot privacy.

The scope of that guarantee should be stated carefully. Strong Fiat-Shamir is sufficient for the adaptive-security notion defined in the paper. It is not a blanket claim that hashing the statement makes any proof system secure. The transcript must still bind the right things, and the underlying proof system must still be sound in the interactive setting. Strong Fiat-Shamir removes one specific failure mode; it does not remove the need to review the rest of the protocol.

Domain separation is not statement binding

These two properties are often conflated. A domain separator — a protocol identifier or context string mixed into the hash — prevents a proof from one protocol or context being replayed in another. Statement binding prevents a prover from choosing a statement after seeing the challenge. A transcript can include a domain separator and still fail statement binding if the public inputs, circuit identifier, or other verifier-relevant values are omitted. The two properties address different attacks and require separate checks.

An audit heuristic for transcript binding

For each challenge in a proof system’s transcript, ask: what can a malicious prover change after seeing this challenge? If the answer includes any public input, circuit identifier, verification key, or domain separator that the verifier relies on, the transcript is under-bound.

A concrete way to apply this: enumerate every hash invocation in the prover and verifier, list the inputs to each, and compare that list against the statement the verifier believes it is checking. The verifier’s belief is the specification. The hash inputs are the implementation. If they differ, the gap is the attack surface.

As a hypothetical illustration of the mechanism — not a report of a specific exploit — consider a proof system where the verifier accepts a proof for public input x, and the challenge is derived only from the prover’s commitment. A malicious prover could, in principle, search for a commitment that yields a challenge satisfying the verification equation for a different, favorable statement. This is the adaptive-selection pattern the strong variant is designed to block. The search may be expensive, but the point of soundness is that it should be infeasible, not merely costly.

Why the bug class persists

The 2023 survey’s finding of 36 weak implementations across 12 proof systems suggests the problem is systemic rather than incidental. Independent codebases, written by different teams, arrived at the same omission. That pattern points to a specification and education gap: the transcript’s hash-input list is often treated as an implementation detail rather than part of the protocol specification.

The corrective is to treat the hash-input list as normative. A protocol specification should state, explicitly, every value that enters each challenge derivation, and reviewers should verify the implementation against that list. When the list is implicit, the weak variant is the default outcome, because hashing only the commitment is the simpler code path.

Frequently asked questions

Is weak Fiat-Shamir always exploitable?

No. The 2016 paper is specific: the weak variant yields unsound and unextractable proofs in situations where malicious provers can select their statements adaptively. If the statement is fixed before the challenge is derived by some other mechanism, the attack does not apply. The practical question is whether the deployed system actually enforces that fixing, which is often not the case when proofs gate state transitions.

Does hashing the statement guarantee soundness?

No. Strong Fiat-Shamir yields secure non-interactive proofs for the adaptive-security notion defined in the 2016 paper, but that result assumes the underlying interactive proof is sound and that the transcript binds the correct values. Hashing the statement is necessary for statement binding, not sufficient for the security of the whole system.

How would a reviewer detect a weak transcript in practice?

Enumerate the inputs to every challenge-derivation hash in both prover and verifier. Confirm that the statement — public inputs, circuit or relation identifier, verification key, and any context the verifier relies on — is among those inputs. Then ask whether a malicious prover could vary any verifier-relevant value after the challenge is fixed. If yes, the transcript is under-bound.

Is this a blockchain-specific problem?

No. The Helios case predates the modern blockchain deployment wave, and the 2023 survey covers proof systems used in multiple contexts. The blockchain relevance is that proof-gated state transitions are a natural fit for the adaptive-statement threat model, and the 2023 paper notes that a weak Fiat-Shamir vulnerability could have led to unlimited currency creation in a private blockchain protocol.

Sources

  • Bernhard, Pereira, Warinschi. “How not to Prove Yourself: Pitfalls of the Fiat-Shamir Heuristic and Applications to Helios.” Cryptology ePrint Archive, Paper 2016/771. https://eprint.iacr.org/2016/771
  • Dao, Miller, Wright, Grubbs. “Weak Fiat-Shamir Attacks on Modern Proof Systems.” Cryptology ePrint Archive, Paper 2023/691. https://eprint.iacr.org/2023/691
  • Gabizon, Williamson, Ciobotaru. “PLONK: Permutations over Lagrange-bases for Oecumenical Noninteractive arguments of Knowledge.” Cryptology ePrint Archive, Paper 2019/953. https://eprint.iacr.org/2019/953