SW/HW Co-Simulation: What It’s For and How We Do It
Software/hardware co-simulation links executable software to a hardware model so you can verify behavior before silicon or FPGA bring-up.
Teams usually use co-sim at two levels:
- Unit-level validation (for example, HLS IP correctness and I/O contracts)
- System-level validation (firmware exercising resets, interrupts, backpressure, and peripheral models)
Methodologies span a spectrum of control and productivity:
- File-based (C/C++ vectors + vendor TB): functional checks with minimal setup; timing/handshakes live in the RTL testbench
- Verilator (C++): cycle-accurate, per-cycle control (
eval()), fast loops for fuzzing, CI, and performance studies - cocotb + Verilator (Python): cycle-accurate with Python ergonomics (
async/await, quick randomization, easy analysis)
In this blog I will deep dive to the above three flows with my open-source project CoSim_Demo as demonstration so you can pick the right co-sim lane for your project.
CoSim Without Tears: Picking the Right Flow for Your Project
If you build chips (or software that talks to them), you eventually run into SW/HW co-simulation: pre-silicon bring-up, FPGA prototyping, and those moments where you need to confirm a driver actually pokes the right bits.
But the best co-sim flow depends on what you are trying to achieve:
- HLS / FPGA prototyping: treat co-sim like a unit test harness for IP blocks. You are proving math and I/O behavior early, fast, and often.
- SoC firmware work: think system-level. You are not only checking firmware logic; you are validating that modeled peripherals behave like the real ones (resets, interrupts, ready/valid, backpressure, and timing interactions).
The two dials: control and abstraction
Most flows fall along two axes:
- How much cycle-level control you need
- Which language you want to write tests in
1) File-based co-sim (C/C++ + vendor sim TB)
When you just want to know “does it work?” without micromanaging clocks, file-based co-sim is the lowest-friction path. Your C/C++ code writes vectors and checks results; a prewritten RTL testbench handles timing and handshakes. It is simple, portable, and great for golden-vector testing. Trade-off: backpressure and dynamic handshakes can feel scripted and clunky.
2) Verilator co-sim (pure C++)
If you need to drive the design every cycle, tick clocks yourself, and profile runtime, Verilator is the power tool. You link a Verilated model into C++, call eval() per cycle, and directly control ready/valid behavior. It is fast (no file I/O) and very CI-friendly for large regressions and fuzzing.
3) cocotb + Verilator (Python)
If you prefer authoring tests in Python but still want cycle accuracy, cocotb gives you Python ergonomics on top of a Verilator backend. You keep Verilator speed/tracing while leveraging Python tooling (hypothesis, numpy, quick data munging, etc.). There is overhead, but for unit/regression scope it is usually a non-issue.
Tradeoffs at a Glance
The three approaches differ mostly in integration boundary, control style, and test authoring ergonomics.
| Aspect | File-based CoSim | Verilator-based CoSim (C++) | cocotb + Verilator (Python) |
|---|---|---|---|
| Integration boundary | Files between C/C++ and Verilog TB | In-process C++ API to Verilated model | Python coroutines to Verilator via cocotb FFI/VPI |
| Authoring language | C/C++ vectors/checkers + Verilog TB | C++ testbench + Verilog DUT | Python testbench + Verilog DUT |
| Timing fidelity | Transaction/time-stepped; manual cycle sync | Cycle-accurate (eval() per cycle) | Cycle-accurate (per-cycle drives/awaits) |
| Backpressure / handshakes | Awkward (pre-scripted) | Natural (toggle signals each cycle) | Natural (await, randomized valid/ready) |
| Performance | Slower (disk I/O, process barriers) | Fast (no file I/O; Verilator optimized loops) | Fast enough for unit/regression; Python overhead acceptable |
| Tooling | Vendor simulators (Questa/ModelSim, etc.) | Verilator + C++ toolchain | Verilator + Python + cocotb |
| Waveforms | Vendor formats (WLF/VCD), viewer varies | FST + GTKWave | VCD/FST via Verilator tracing + GTKWave |
| Coverage | Vendor coverage tooling | --coverage + verilator_coverage | Verilator coverage (extra args) + same tooling |
| Best for | Portable golden checks | High-iteration CI/fuzzing / native C++ control | Fast authoring, Python-heavy test ecosystems |
The Demo Setup: Same DUT, Three Harnesses
To make the comparison concrete, the rest of the post uses the same tiny DUT across three harnesses:
- File-based C/C++ with a vendor-style RTL testbench
- Verilator with an in-process C++ testbench
- cocotb + Verilator in Python
The important point is that the DUT stays the same while the methodology changes. That makes it easy to compare:
- Timing control
- Backpressure handling
- Runtime and iteration speed
- Waveforms and debug workflow
- Coverage and artifact generation
The cosim_demo style setup is intentionally small and copy-pasteable: Make targets, sample vectors, wave dumps, and coverage outputs you can keep.
What This Repo Focuses On
This design distills three reusable patterns that cover most practical co-sim needs while staying compact:
- File-based CoSim
- Verilator-based CoSim (C++)
- Verilator + Python cocotb (Make-based)
All three wrap the same simple DUT and emphasize:
- Scoreboard checking
- Directed + randomized traffic
- Reusable artifacts (waveforms, coverage, logs)
The Adder DUT
The DUT is a simple ready/valid adder with a two-deep output queue (main buffer + spill). It is small enough to understand quickly, but realistic enough to exercise backpressure and handshake behavior.
Key behaviors:
- Sustains one result per cycle when
out_readyis high - Absorbs one cycle of backpressure without stalling inputs
- Computes
in_a + in_bin a single cycle - Prioritizes draining the spill buffer
- Asserts
in_readywhenever the spill slot is free - Uses active-low reset (
rst_n) to clear buffers and valid flags
The following flow diagram shows the internal adder architecture:

The Examples in This Repo
Example #1: file-based/simple_adder_rv
This example demonstrates file-based co-sim with a clean file contract:
- Software emits inputs
- Verilog testbench reads/drives the DUT
- Outputs are captured and compared against golden results
It is designed for deterministic golden checks and portability across simulators.
The C++ host testbench invokes QuestaSim via std::system(), using a fixed command string to run the simulation to completion (run -all; quit -f). Afterward, it validates results by comparing outputs.txt with the software-computed golden values. The main() exit code reports the outcome: 0 for pass, 5 for fail.
The example end-to-end CoSim flow is shown below:

One can build and run this example by running make run, the simulation shall output the following:

Example #2: verilator-based/rv_adder_example
This example shows a simple yet portable Verilator-based Co-simulation example that verifies the ready/valid adder (i.e. adder_rv_simple) using a C++ testbench. The test drives both directed vectors and random streaming with backpressure, checks results with a scoreboard, dumps an FST waveform, and writes coverage.
The testbench main() is implemented as follows:
- Runs a directed smoke test first with no backpressure (
top->out_ready = 1) - Flushes any leftover outputs to clear internal DUT buffers
- Switches to randomized streaming using
std::mt19937_64 rng(1) - Randomizes
a,b,in_valid, andout_readyto exercise backpressure - Drives stimulus on the negedge
- Calls
top->eval()on the posedge, then traces/checks outputs - Uses a
std::queueas the golden scoreboard - Uses mismatch count as the
main()return code
This gives you cycle-accurate control with a very fast feedback loop. The following diagram demonstrates the end-to-end simulation flow:

As shown above, the screenshot confirms the expected artifacts and shows the coverage summary (e.g., Total coverage ... 82.00%) and the annotation directory hint.
Example #3: verilator-based/cocotb_rv_adder
This is a Make-based test that drives the same ready/valid adder on the Verilator backend. It shows how to verify a SystemVerilog ready/valid adder (adder_rv_simple.sv) using a cocotb Python testbench on a Verilated C++ simulator. It also includes FST waveform dumping and Verilator coverage with post-run source annotation.
The cocotb test coroutine test_adder_rv_simple( ) mirrors Example #2—running a directed smoke test followed by randomized traffic—but differs in stimulus and clock handling. Specifically, inputs are applied right after negedge phase via await FallingEdge(dut.clk), and outputs are sampled in the simulator’s Observed stage using await ReadOnly(). This ensures cycle-accurate observation without race conditions. The following code snippet demonstrates such input setting and output sampling:
# drive only on FallingEdge
await FallingEdge(dut.clk)
dut.in_valid.value = 1
dut.in_a.value = a
dut.in_b.value = b
# observe only after RisingEdge
await RisingEdge(dut.clk)
await ReadOnly()
# sample output here; no writes are allowed here
dut_sum = dut.out_sum.value
For cocotb–Verilator communication, VPI is the glue. Cocotb registers its callbacks into the Verilated C++ model via VPI, which is the standard way for external code to hook into the simulator’s event schedule and design hierarchy. In practice, cocotb talks to Verilator through VPI.
Quick primer: simulation event regions (why ReadOnly() matters)
At a given simulation time t, the kernel does not execute everything at once. It drains a sequence of event regions (queues), potentially across multiple delta cycles, before advancing to t + Δ. That ordering is what lets RTL and testbench code interact without races.
The nine regions (in order):
- Preponed - Snapshot values before updates (for example, immediate assertion sampling)
- Active - Normal execution; blocking assignments update values; NBAs are scheduled
- Inactive - Handles
#0and deferred same-time-slot events - NBA - Commits scheduled non-blocking assignments
- Re-NBA - Second NBA commit pass (for NBAs scheduled later)
- Observed - Read-only; assertions/coverage/sample points see final values for this time
- Reactive - Testbench/program/clocking-block drives (avoids DUT races)
- Re-Active - Re-evaluation triggered by reactive drives
- Postponed - Final read-only stage (
$monitor, dumps, PLI/VPI callbacks)
The following diagram demonstrates the end-to-end simulation flow:

One can build and run this example by running make run, the simulation shall output the following:

Summary
SW/HW co-simulation meaningfully speeds pre-silicon work—both chip design and FPGA prototyping—by letting you verify at the abstraction and control level that best fits the task. cosim_demo isn’t a heavyweight framework; it’s a set of small, production-ready patterns you can drop into real codebases. Use file-based co-sim for portable, deterministic golden checks; choose Verilator + C++ when you need cycle-accurate control, high speed, and coverage; reach for cocotb + Verilator when Python’s ergonomics make test authoring faster. Same DUT, three viewpoints—pick what fits today, keep the others handy as your verification needs evolve.
Closing Thoughts
If you want, I can also turn this into a follow-up implementation guide with concrete repository layout, Make targets, and a minimal ready/valid scoreboard example for each pattern.
