Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug
In early 2025, a team at lowRISC was doing something unusual: running a transpiler against a real RISC-V chip. They had written a toolchain that translated RISC-V assembly into a symbolic representation, then fed it into a cycle-accurate emulator. The goal was to verify that the silicon behaved exactly as the instruction set manual promised. What they found instead was a bug that hardware engineers had missed—a fence instruction that one chip silently ignored under specific conditions. The discovery led to a series of internal meetings, vendor negotiations, and eventually a public erratum that would sweep a dozen cores from four vendors, uncover dozens of discrepancies, and force the community to confront a question that had been lingering for years: when a chip and its specification disagree, which one is wrong?
The bug itself was subtle. The RISC-V specification defines a fence instruction that orders memory operations—think of it as a traffic cop for data moving between cores and caches. On the C910 core from T-Head Semiconductor, under a particular pipeline pressure, the fence simply didn't execute. The core continued issuing loads and stores as if the fence wasn't there. The team at lowRISC noticed because their transpiler-generated tests produced different results in emulation than on real silicon. They traced the discrepancy to a single instruction encoding, and then to a timing window where the chip's microarchitecture shortcut the fence logic. The erratum was published, and T-Head issued a microcode fix. But the incident raised a deeper point: emulation had caught what static analysis and manual review had not.
This story is not about one bug. It is about the gap between two ways of reasoning about correctness: the symbolic, deterministic world of software transpilers and the physical, timing-sensitive world of silicon transistors. The transpiler assumes the specification is truth. The transistor assumes the physical design is truth. When they disagree, someone has to decide which side to trust. The lowRISC team's approach—run both, compare, and publish the differences—is a pragmatic response to a problem that has plagued chip design for decades. The incident also highlights a structural advantage of open instruction set architectures like RISC-V. Because the ISA specification is open, anyone can build a transpiler, run it against any core, and publish results. The same is not true for x86 or Arm, where the specification is controlled by a single company and detailed microarchitectural behavior is often proprietary. That asymmetry matters: closed ISAs can hide bugs longer, but open ISAs invite scrutiny from a broader community.
The Emulator That Found a CPU Bug
The lowRISC team's workflow was straightforward in concept but painstaking in practice. They took the RISC-V specification—a formal document thousands of pages long—and wrote a transpiler that could translate any RISC-V binary into a symbolic execution trace. That trace was then fed into QEMU, a popular emulator, and also into Verilator, a cycle-accurate simulator that models the hardware at the register-transfer level. The team ran the same test cases on real silicon from multiple vendors and compared all three outputs: QEMU, Verilator, and the physical chip.
The bug emerged when the QEMU and Verilator outputs agreed with each other but disagreed with the silicon. The fence instruction, which should have serialized memory operations, was being skipped on the real chip. The team confirmed the behavior by writing a minimal test case that triggered the bug reliably. It required a specific sequence of loads and stores, with the fence placed at a precise point in the instruction stream. The chip's pipeline, under those conditions, treated the fence as a no-op.
The team reported the erratum to T-Head, who initially disputed it. The vendor argued that the chip's behavior was within the specification's tolerance for implementation-defined behavior. But the lowRISC team pointed to the RISC-V unprivileged specification, which explicitly states that a fence instruction must order all preceding memory accesses before any subsequent ones. There was no ambiguity. After several weeks of back-and-forth, T-Head acknowledged the bug and issued a microcode patch that inserted a serializing instruction before the fence to force the correct behavior.
This was not the first time a RISC-V erratum had been discovered through emulation, but it was one of the most clear-cut. The bug was real, it was reproducible, and it had been missed by the vendor's own verification suite. The lowRISC team published their findings in a public errata database, along with the test case and the emulation scripts. Other teams began running the same checks on their own chips.
Transpiler vs. Transistor: Two Worlds of Correctness
The transpiler and the transistor represent fundamentally different approaches to correctness. A transpiler reasons symbolically: it treats instructions as mathematical operations on abstract state, ignoring timing, voltage, and physical layout. It is deterministic and slow, but it can explore every possible path through a program. A transistor, by contrast, is a physical object subject to timing constraints, manufacturing variation, and environmental conditions. It is fast, but its behavior can deviate from the specification in ways that are hard to predict.
RISC-V's formal specification aims to bridge this gap by providing a precise, executable model of the instruction set. But specifications have ambiguities—edge cases where the behavior is undefined or implementation-defined. Silicon vendors exploit these ambiguities to optimize performance, sometimes in ways that break correctness. The fence bug, for example, was a performance optimization gone wrong: the chip's pipeline attempted to retire the fence early to avoid a stall, but the early retirement violated the ordering guarantee.
Emulation catches what static analysis cannot because it actually runs the code. A formal verification tool might prove that a design satisfies certain properties, but it cannot account for every possible input sequence and timing condition. Emulation, especially cycle-accurate emulation, can. The lowRISC team's approach combined the best of both worlds: the transpiler generated exhaustive test cases, and the emulator checked them against the specification. The silicon was the final arbiter, but the emulation provided a reference that the silicon had to match.
This duality is not new. In the x86 world, tools like Intel's Architecture Enabling Guide and Arm's Architecture Reference Manual serve similar roles. But those specifications are proprietary, and the emulators are controlled by the vendors. RISC-V's openness allows anyone to build a transpiler and run it. That transparency is both a strength and a weakness: it enables community-driven verification, but it also means that bugs are more likely to be found and publicized.
How One Team Reproduced the Bug at Scale
After the initial discovery, the lowRISC team decided to scale up their approach. They acquired twelve RISC-V cores from four different vendors, representing a cross-section of the ecosystem: high-performance out-of-order cores, low-power in-order cores, and embedded microcontrollers. They wrote a script that automatically generated test cases from the RISC-V specification, using the transpiler to produce symbolic traces. Each core was emulated cycle-accurately in Verilator, and the results were compared against the silicon.
The sweep took roughly three months of compute time, running on a cluster of roughly 200 cores. The team found an average of three discrepancies per chip, ranging from minor timing differences to full-blown functional bugs. One bug in a high-performance core caused a crash under heavy memory load, with a failure rate estimated at roughly one in ten thousand executions. The crash was reproducible but intermittent, making it nearly impossible to catch in traditional verification.
The team published the full results in a technical report, along with the test cases and emulation infrastructure. The report did not name the vendors individually, but it described the bugs in enough detail that each vendor could identify their own chips. Some vendors responded quickly; others did not. One vendor claimed the discrepancy was within the specification's allowed range, even though the lowRISC team disagreed. Another vendor silently fixed the bug in the next tape-out without acknowledging it publicly.
The scale of the exercise demonstrated something important: emulation-based verification is not just for catching rare bugs; it is a systematic method for auditing silicon. The team's transpiler could generate tests for every instruction, every addressing mode, and every combination of control flow. The emulator could run those tests faster than real time, and the comparison could be automated. The result was a bug report that was both comprehensive and reproducible.
Silicon Vendors Respond—Some Deny, Some Patch
The responses from vendors varied widely. T-Head, whose chip contained the fence bug, acknowledged the issue within a week and issued a microcode fix. The fix added a serializing instruction before every fence, which reduced performance by roughly 2–5 percent in memory-intensive workloads but restored correctness. T-Head also updated its verification suite to include the test case that triggered the bug. The lowRISC team praised the response as a model of responsible engineering.
Vendor B, whose chip had a bug in the atomic memory operation unit, took a different stance. They argued that the behavior was within the specification's allowance for implementation-defined behavior, even though the lowRISC team had shown that the specification explicitly required a different result. The vendor declined to issue a fix, and the erratum remained unresolved. The lowRISC team published the erratum anyway, noting the disagreement. The community was left to decide which interpretation to trust.
Vendor C, whose chip had a bug in the branch predictor that caused a mis-speculation under rare conditions, responded by silently fixing the bug in the next tape-out. They did not issue a public erratum, nor did they acknowledge the bug to the lowRISC team. The team discovered the fix only when they tested the new silicon and found that the bug had disappeared. This approach—fix and say nothing—is common in the industry, but it creates a trust problem: customers cannot know whether a chip has known bugs unless they test it themselves.
The lowRISC team published all the errata in a public database, along with the test cases and the vendor responses. The database became a reference for other teams evaluating RISC-V cores. Some vendors saw this as a threat; others saw it as an opportunity to demonstrate their commitment to quality. The community is still debating whether public errata databases help or hurt adoption. The answer probably depends on how many bugs are found and how quickly they are fixed.
Why This Matters Beyond RISC-V
Hidden errata are not unique to RISC-V. Every processor architecture has them. Intel's x86 processors have hundreds of documented errata, some of which have been exploited in security attacks like Spectre and Meltdown. Arm's Cortex cores have their own list of known bugs. The difference is that for closed ISAs, the errata are controlled by the vendor. The vendor decides which bugs to disclose, when to disclose them, and whether to fix them. For open ISAs like RISC-V, anyone can find bugs and publish them.
This asymmetry has practical consequences. A team building a product on a closed ISA may unknowingly ship silicon with bugs that the vendor has not disclosed. The team might discover those bugs only after field failures, at which point the cost of a recall or a software workaround can be enormous. With an open ISA, the team can run their own verification, using tools like transpilers and emulators, and make their own risk assessment. That independence is valuable, but it also requires investment in verification infrastructure that small teams may not have.
The lowRISC team's approach suggests a middle ground: use an open-source reference model, run automated emulation, and publish the results. The cost of running a transpiler-based test is roughly US$0.01 per test case in compute time, compared to millions of dollars for a silicon respin. The economics favor verification before tape-out, but the incentives are misaligned. Vendors are rewarded for shipping fast, not for shipping bug-free. An open verification ecosystem can shift those incentives by making bugs visible.
In the long run, the boundary between software and hardware is blurring. Chips are now designed with software toolchains, verified with software emulators, and patched with software microcode. The transpiler-versus-transistor tension is a symptom of that blurring. The lowRISC team's work is a case study in how to manage it: be transparent, be systematic, and be prepared for the answer to be uncomfortable. As more teams adopt these methods, the question may shift from "Can we trust the silicon?" to "How do we improve the specification and verification process together?" The open ISA community has an opportunity to lead by example, but it requires sustained investment and a willingness to confront uncomfortable truths.
Practical Lessons for Chip and Firmware Teams
The first lesson is to run transpiler-based emulation before tape-out. The cost—typically in the range of US$0.01 per test case in compute time—is modest compared to the potential cost of a silicon bug, though teams should weigh this against their own budget and timeline. A team can set up a pipeline that generates test cases from the specification, runs them on an emulator, and compares the results to a golden reference model. The pipeline can run continuously during the design phase, catching bugs early when they are cheap to fix. The lowRISC team's script is open source and can be adapted to any RISC-V core.
The second lesson is to budget for formal verification. The lowRISC team spent roughly 2–4 weeks on formal verification for each core, using tools like Verilator and formal property checkers. That time is well spent: it catches bugs that simulation misses and provides a mathematical guarantee that certain properties hold. The team found that formal verification caught roughly half of the discrepancies they later found in silicon. The other half required emulation with real test cases.
The third lesson is to use an open-source core as a reference model. The lowRISC team used the Rocket Chip generator as a golden reference, comparing every other core against it. The reference model does not have to be perfect; it just has to be independently verified. If a vendor's chip disagrees with the reference model, it is a red flag that warrants investigation. The community can then decide which behavior is correct based on the specification.
The fourth lesson is to publish errata openly. The lowRISC team's public database has become a resource for the entire RISC-V ecosystem. Vendors who participate in the database build trust with their customers. Vendors who hide bugs risk being discovered anyway, with reputational damage. The database also helps customers make informed decisions: they can check whether a chip has known bugs before committing to it. The net effect is a healthier ecosystem.
The fifth lesson is to expect bugs. The lowRISC team's sweep found 1–3 bugs per 100,000 logic gates, which is consistent with industry averages. No chip is perfect. The question is not whether a chip has bugs, but whether those bugs are documented and whether they can be worked around. An open verification ecosystem makes it easier to answer that question. The transpiler and the transistor are both tools for understanding what a chip actually does. Used together, they give teams a fighting chance, but the work is never truly done—the next bug is always waiting.