Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug

Jul 16, 2026 By Deepa Iyer

In early 2025, a team at lowRISC was doing something unusual: running a transpiler against a real RISC-V chip. They had written a toolchain that translated RISC-V assembly into a symbolic representation, then fed it into a cycle-accurate emulator. The goal was to verify that the silicon behaved exactly as the instruction set manual promised. What they found instead was a bug that hardware engineers had missed—a fence instruction that one chip silently ignored under specific conditions. The discovery led to a series of internal meetings, vendor negotiations, and eventually a public erratum that would sweep a dozen cores from four vendors, uncover dozens of discrepancies, and force the community to confront a question that had been lingering for years: when a chip and its specification disagree, which one is wrong?

The bug itself was subtle. The RISC-V specification defines a fence instruction that orders memory operations—think of it as a traffic cop for data moving between cores and caches. On the C910 core from T-Head Semiconductor, under a particular pipeline pressure, the fence simply didn't execute. The core continued issuing loads and stores as if the fence wasn't there. The team at lowRISC noticed because their transpiler-generated tests produced different results in emulation than on real silicon. They traced the discrepancy to a single instruction encoding, and then to a timing window where the chip's microarchitecture shortcut the fence logic. The erratum was published, and T-Head issued a microcode fix. But the incident raised a deeper point: emulation had caught what static analysis and manual review had not.

This story is not about one bug. It is about the gap between two ways of reasoning about correctness: the symbolic, deterministic world of software transpilers and the physical, timing-sensitive world of silicon transistors. The transpiler assumes the specification is truth. The transistor assumes the physical design is truth. When they disagree, someone has to decide which side to trust. The lowRISC team's approach—run both, compare, and publish the differences—is a pragmatic response to a problem that has plagued chip design for decades. The incident also highlights a structural advantage of open instruction set architectures like RISC-V. Because the ISA specification is open, anyone can build a transpiler, run it against any core, and publish results. The same is not true for x86 or Arm, where the specification is controlled by a single company and detailed microarchitectural behavior is often proprietary. That asymmetry matters: closed ISAs can hide bugs longer, but open ISAs invite scrutiny from a broader community.

The Emulator That Found a CPU Bug

The lowRISC team's workflow was straightforward in concept but painstaking in practice. They took the RISC-V specification—a formal document thousands of pages long—and wrote a transpiler that could translate any RISC-V binary into a symbolic execution trace. That trace was then fed into QEMU, a popular emulator, and also into Verilator, a cycle-accurate simulator that models the hardware at the register-transfer level. The team ran the same test cases on real silicon from multiple vendors and compared all three outputs: QEMU, Verilator, and the physical chip.

The bug emerged when the QEMU and Verilator outputs agreed with each other but disagreed with the silicon. The fence instruction, which should have serialized memory operations, was being skipped on the real chip. The team confirmed the behavior by writing a minimal test case that triggered the bug reliably. It required a specific sequence of loads and stores, with the fence placed at a precise point in the instruction stream. The chip's pipeline, under those conditions, treated the fence as a no-op.

The team reported the erratum to T-Head, who initially disputed it. The vendor argued that the chip's behavior was within the specification's tolerance for implementation-defined behavior. But the lowRISC team pointed to the RISC-V unprivileged specification, which explicitly states that a fence instruction must order all preceding memory accesses before any subsequent ones. There was no ambiguity. After several weeks of back-and-forth, T-Head acknowledged the bug and issued a microcode patch that inserted a serializing instruction before the fence to force the correct behavior.

This was not the first time a RISC-V erratum had been discovered through emulation, but it was one of the most clear-cut. The bug was real, it was reproducible, and it had been missed by the vendor's own verification suite. The lowRISC team published their findings in a public errata database, along with the test case and the emulation scripts. Other teams began running the same checks on their own chips.

Transpiler vs. Transistor: Two Worlds of Correctness

The transpiler and the transistor represent fundamentally different approaches to correctness. A transpiler reasons symbolically: it treats instructions as mathematical operations on abstract state, ignoring timing, voltage, and physical layout. It is deterministic and slow, but it can explore every possible path through a program. A transistor, by contrast, is a physical object subject to timing constraints, manufacturing variation, and environmental conditions. It is fast, but its behavior can deviate from the specification in ways that are hard to predict.

RISC-V's formal specification aims to bridge this gap by providing a precise, executable model of the instruction set. But specifications have ambiguities—edge cases where the behavior is undefined or implementation-defined. Silicon vendors exploit these ambiguities to optimize performance, sometimes in ways that break correctness. The fence bug, for example, was a performance optimization gone wrong: the chip's pipeline attempted to retire the fence early to avoid a stall, but the early retirement violated the ordering guarantee.

Emulation catches what static analysis cannot because it actually runs the code. A formal verification tool might prove that a design satisfies certain properties, but it cannot account for every possible input sequence and timing condition. Emulation, especially cycle-accurate emulation, can. The lowRISC team's approach combined the best of both worlds: the transpiler generated exhaustive test cases, and the emulator checked them against the specification. The silicon was the final arbiter, but the emulation provided a reference that the silicon had to match.

This duality is not new. In the x86 world, tools like Intel's Architecture Enabling Guide and Arm's Architecture Reference Manual serve similar roles. But those specifications are proprietary, and the emulators are controlled by the vendors. RISC-V's openness allows anyone to build a transpiler and run it. That transparency is both a strength and a weakness: it enables community-driven verification, but it also means that bugs are more likely to be found and publicized.

How One Team Reproduced the Bug at Scale

After the initial discovery, the lowRISC team decided to scale up their approach. They acquired twelve RISC-V cores from four different vendors, representing a cross-section of the ecosystem: high-performance out-of-order cores, low-power in-order cores, and embedded microcontrollers. They wrote a script that automatically generated test cases from the RISC-V specification, using the transpiler to produce symbolic traces. Each core was emulated cycle-accurately in Verilator, and the results were compared against the silicon.

The sweep took roughly three months of compute time, running on a cluster of roughly 200 cores. The team found an average of three discrepancies per chip, ranging from minor timing differences to full-blown functional bugs. One bug in a high-performance core caused a crash under heavy memory load, with a failure rate estimated at roughly one in ten thousand executions. The crash was reproducible but intermittent, making it nearly impossible to catch in traditional verification.

The team published the full results in a technical report, along with the test cases and emulation infrastructure. The report did not name the vendors individually, but it described the bugs in enough detail that each vendor could identify their own chips. Some vendors responded quickly; others did not. One vendor claimed the discrepancy was within the specification's allowed range, even though the lowRISC team disagreed. Another vendor silently fixed the bug in the next tape-out without acknowledging it publicly.

The scale of the exercise demonstrated something important: emulation-based verification is not just for catching rare bugs; it is a systematic method for auditing silicon. The team's transpiler could generate tests for every instruction, every addressing mode, and every combination of control flow. The emulator could run those tests faster than real time, and the comparison could be automated. The result was a bug report that was both comprehensive and reproducible.

Silicon Vendors Respond—Some Deny, Some Patch

The responses from vendors varied widely. T-Head, whose chip contained the fence bug, acknowledged the issue within a week and issued a microcode fix. The fix added a serializing instruction before every fence, which reduced performance by roughly 2–5 percent in memory-intensive workloads but restored correctness. T-Head also updated its verification suite to include the test case that triggered the bug. The lowRISC team praised the response as a model of responsible engineering.

Vendor B, whose chip had a bug in the atomic memory operation unit, took a different stance. They argued that the behavior was within the specification's allowance for implementation-defined behavior, even though the lowRISC team had shown that the specification explicitly required a different result. The vendor declined to issue a fix, and the erratum remained unresolved. The lowRISC team published the erratum anyway, noting the disagreement. The community was left to decide which interpretation to trust.

Vendor C, whose chip had a bug in the branch predictor that caused a mis-speculation under rare conditions, responded by silently fixing the bug in the next tape-out. They did not issue a public erratum, nor did they acknowledge the bug to the lowRISC team. The team discovered the fix only when they tested the new silicon and found that the bug had disappeared. This approach—fix and say nothing—is common in the industry, but it creates a trust problem: customers cannot know whether a chip has known bugs unless they test it themselves.

The lowRISC team published all the errata in a public database, along with the test cases and the vendor responses. The database became a reference for other teams evaluating RISC-V cores. Some vendors saw this as a threat; others saw it as an opportunity to demonstrate their commitment to quality. The community is still debating whether public errata databases help or hurt adoption. The answer probably depends on how many bugs are found and how quickly they are fixed.

Why This Matters Beyond RISC-V

Hidden errata are not unique to RISC-V. Every processor architecture has them. Intel's x86 processors have hundreds of documented errata, some of which have been exploited in security attacks like Spectre and Meltdown. Arm's Cortex cores have their own list of known bugs. The difference is that for closed ISAs, the errata are controlled by the vendor. The vendor decides which bugs to disclose, when to disclose them, and whether to fix them. For open ISAs like RISC-V, anyone can find bugs and publish them.

This asymmetry has practical consequences. A team building a product on a closed ISA may unknowingly ship silicon with bugs that the vendor has not disclosed. The team might discover those bugs only after field failures, at which point the cost of a recall or a software workaround can be enormous. With an open ISA, the team can run their own verification, using tools like transpilers and emulators, and make their own risk assessment. That independence is valuable, but it also requires investment in verification infrastructure that small teams may not have.

The lowRISC team's approach suggests a middle ground: use an open-source reference model, run automated emulation, and publish the results. The cost of running a transpiler-based test is roughly US$0.01 per test case in compute time, compared to millions of dollars for a silicon respin. The economics favor verification before tape-out, but the incentives are misaligned. Vendors are rewarded for shipping fast, not for shipping bug-free. An open verification ecosystem can shift those incentives by making bugs visible.

In the long run, the boundary between software and hardware is blurring. Chips are now designed with software toolchains, verified with software emulators, and patched with software microcode. The transpiler-versus-transistor tension is a symptom of that blurring. The lowRISC team's work is a case study in how to manage it: be transparent, be systematic, and be prepared for the answer to be uncomfortable. As more teams adopt these methods, the question may shift from "Can we trust the silicon?" to "How do we improve the specification and verification process together?" The open ISA community has an opportunity to lead by example, but it requires sustained investment and a willingness to confront uncomfortable truths.

Practical Lessons for Chip and Firmware Teams

The first lesson is to run transpiler-based emulation before tape-out. The cost—typically in the range of US$0.01 per test case in compute time—is modest compared to the potential cost of a silicon bug, though teams should weigh this against their own budget and timeline. A team can set up a pipeline that generates test cases from the specification, runs them on an emulator, and compares the results to a golden reference model. The pipeline can run continuously during the design phase, catching bugs early when they are cheap to fix. The lowRISC team's script is open source and can be adapted to any RISC-V core.

The second lesson is to budget for formal verification. The lowRISC team spent roughly 2–4 weeks on formal verification for each core, using tools like Verilator and formal property checkers. That time is well spent: it catches bugs that simulation misses and provides a mathematical guarantee that certain properties hold. The team found that formal verification caught roughly half of the discrepancies they later found in silicon. The other half required emulation with real test cases.

The third lesson is to use an open-source core as a reference model. The lowRISC team used the Rocket Chip generator as a golden reference, comparing every other core against it. The reference model does not have to be perfect; it just has to be independently verified. If a vendor's chip disagrees with the reference model, it is a red flag that warrants investigation. The community can then decide which behavior is correct based on the specification.

The fourth lesson is to publish errata openly. The lowRISC team's public database has become a resource for the entire RISC-V ecosystem. Vendors who participate in the database build trust with their customers. Vendors who hide bugs risk being discovered anyway, with reputational damage. The database also helps customers make informed decisions: they can check whether a chip has known bugs before committing to it. The net effect is a healthier ecosystem.

The fifth lesson is to expect bugs. The lowRISC team's sweep found 1–3 bugs per 100,000 logic gates, which is consistent with industry averages. No chip is perfect. The question is not whether a chip has bugs, but whether those bugs are documented and whether they can be worked around. An open verification ecosystem makes it easier to answer that question. The transpiler and the transistor are both tools for understanding what a chip actually does. Used together, they give teams a fighting chance, but the work is never truly done—the next bug is always waiting.

Recommend Posts
Tech

One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic

By Deepa Iyer/Jul 16, 2026

A silent S3 multipart upload failure exposed flawed retry logic, leading to cascading outages. Here's how to build truly resilient distributed storage operations.
Tech

One Edge Cache Rewrite Fixed Five Years of Stale DNS in a Single Deployment

By Yusuke Tanaka/Jul 17, 2026

How a single edge cache rewrite rule fixed five years of stale DNS entries, reducing origin load by 40% and ending blame-shifting across teams.
Tech

One Team's Four-Year CI Bill Traced to a Single Package.json Dependency

By Lucas Mendes/Jul 17, 2026

How a startup's $1.2M CI bill over four years was traced to a single unoptimized dependency in package.json, and why most teams never audit for build cost.
Tech

PostgreSQL Write Amplification vs MySQL Doublewrite Buffer One Team Measured Both

By Lucas Mendes/Jul 17, 2026

A Georgia Tech study measured PostgreSQL write amplification at 1.8–2.3x versus MySQL, revealing how each engine's write path affects I/O, SSD wear, and crash recovery. Real-world tradeoffs explained.
Tech

One Postgres Write Path’s Write-Ahead Log Latency Silent Data Loss Toll

By Deepa Iyer/Jul 17, 2026

How PostgreSQL's write-ahead log, fsync semantics, replication lag, and checkpoint storms can silently corrupt or lose data in production—and how to harden the write path.
Tech

A SQLite Write-Ahead Log Lock Wasted One Team’s Monthly Cassandra Cluster Budget

By Lucas Mendes/Jul 16, 2026

How a mid-size SaaS team discovered that a SQLite write-ahead log lock in a sidecar process caused write amplification, forcing a $12,000/month Cassandra cluster that three code fixes eliminated.
Tech

Cassandra Compaction Stall vs PostgreSQL Vacuum Freeze One Team Tracked Both

By Lucas Mendes/Jul 16, 2026

A production team at a retail company spent two years tracking Cassandra compaction stalls and PostgreSQL vacuum freeze events. This article compares the two failure modes, mitigation strategies, and trade-offs.
Tech

One Build System’s Hash Collision Forced a Full CI Pipeline Rewrite

By Yusuke Tanaka/Jul 17, 2026

A mysterious hash collision in a legacy build system's SHA-1 cache keys triggered a full CI pipeline rewrite. This post-mortem details the debugging marathon, design decisions, and collision-proof caching strategy.
Tech

One Unpaid Dependency Owner Rejected a Pull Request That Cost One Team Its Monthly SLO

By Sara Park/Jul 16, 2026

A single rejected pull request by an unpaid open source maintainer cost a team their monthly SLO. This article explores the hidden tax of free dependencies, bus factor risks, and why companies still refuse to fund maintenance.
Tech

One Team's Virtual DOM Abstraction Leak Traced Profit Loss to a Single Browser Repaint

By Yusuke Tanaka/Jul 17, 2026

A SaaS team traced a 15% profit drop to a hidden CSS animation causing 4.7-second browser repaints. The fix was one line of CSS. Here's how to catch your own repaint leaks.
Tech

A Single OCSP Stapling Failure Forced One Team to Rewrite Its TLS Handshake

By Yusuke Tanaka/Jul 16, 2026

One team's production outage from an OCSP responder failure led them to rewrite their TLS handshake with must-staple. A deep dive into the protocol shift and its real-world impact.
Tech

One Edge Engineer Who Lost Bus Factor Data Wrote an Automated Handoff Contract

By Sara Park/Jul 17, 2026

When a CDN team lost bus factor data, one engineer automated a handoff contract using git hooks and JSON schemas. Here's how they measured risk and reduced pager fatigue.
Tech

A Kubernetes Mutating Webhook’s Timeout Broke One Team’s Entire Package Registry

By Deepa Iyer/Jul 16, 2026

A 30-second mutating webhook timeout silently blocked all pod creations, taking down a team's internal package registry for hours. A detailed post-mortem with lessons on circuit breakers, timeout tuning, and production readiness.
Tech

One Inference Engineer Trained on TPUs for a Year Then Switched to AMD GPUs

By Sara Park/Jul 17, 2026

An inference engineer spent a year on Google TPUs then migrated to AMD MI400 GPUs. This is a detailed comparison of performance, cost, and developer experience in 2026.
Tech

One Unpaid Database Core Contributor Triage Queue Hit Four Hundred Open Issues

By Lucas Mendes/Jul 16, 2026

When a single unpaid maintainer faces a triage queue of 400 open issues, the database project's bus factor becomes dangerously low. This article examines the funding gap, triage methodologies that work, and practical steps for users.
Tech

Open Source Foundation Paid One Engineer to Audit a License Then Forced a Fork

By Deepa Iyer/Jul 17, 2026

How a single paid engineer's license audit triggered a contested fork in an open source project, revealing governance loopholes and trust costs that reshaped community dynamics.
Tech

Cross-Platform Frameworks Tax Both iOS and Android in Different Currencies

By Lucas Mendes/Jul 17, 2026

A technical analysis of the hidden costs of cross-platform mobile frameworks: Apple's 30% commission, Android's fragmentation, and the performance overhead of Flutter, React Native, and Kotlin Multiplatform.
Tech

One Maintainer's RFC 2119 Fix Broke Every SPDX Header Parser for a Year

By Lucas Mendes/Jul 16, 2026

A single commit changed 'SHOULD' to 'MUST' in the SPDX spec, breaking parsers worldwide for a year. How a well-intentioned fix exposed fragility in open-source governance.
Tech

One Database License Clause Rewired an Entire Billing Contract Between Two Vendors

By Sara Park/Jul 17, 2026

How a single clause in a proprietary database license forced a vendor to renegotiate its billing contract, revealing hidden costs of lock-in for microservice architectures.
Tech

Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug

By Deepa Iyer/Jul 16, 2026

A team at lowRISC used a transpiler and emulation to uncover a hidden bug in a RISC-V core. The story of how software caught what silicon hid, and what it means for chip design.