Make RyuSim simulations faster: threads, optimization levels, 2-state mode, and how to measure.
A RyuSim run costs time in two places. ryusim compile turns the design into C++ and builds it with Clang. The simulation binary then runs the design. Measure the two separately, and check which build path RyuSim took before you compare two runs.
Read these before tuning anything:
ryusim profile is not implemented. It does not compile or run the design. It prints a report in which every time is 0 and exits with status 0. Use time and a sampling profiler instead (see below).--threads with a value above 1 turns off the default trigger schedule for the whole build, even when the partitioner refuses to split the design. A refused --threads build runs a different schedule from a build without the flag, and it can be slower.--threads refuses most designs. Its cost model counts a loop body a fixed four times, whatever the trip count.-O0 by default, whatever -O says. The cold pool holds fork branch bodies and concurrent assertion evaluation as well as one-time setup code.-O1 or -O0 to bound Clang's compile time. The build log still prints the pool's level for that file, and ryusim-route.json does not record the change. build.ninja does: look for opt = -O1 or opt = -O0 lines.--opt ignores a pass name it does not know, without a message.This design runs 100 clock cycles and prints the count:
module counter;
logic clk = 0;
logic [31:0] count = 0;
always #5 clk = ~clk;
always_ff @(posedge clk) count <= count + 1;
initial begin
#1000;
$display("count=%0d", count);
$finish;
end
endmodule
Time the compile. -q is a global option, so it goes before compile:
time ryusim -q compile --top counter counter.sv
Time the simulation on its own:
time ./obj_dir/build/counter_sim
Compile time has two parts. RyuSim parses, elaborates and writes C++ (steps 1 to 3 in the compile output); step 4 runs Clang through ninja. --codegen-only stops after step 3, so timing it gives RyuSim's share:
time ryusim -q compile --codegen-only --top counter -o obj_gen counter.sv
A second compile into the same output directory is incremental. RyuSim rewrites only the generated files whose text changed, and ninja rebuilds only those. For the full cost of a compile, time it into an empty directory.
Step 4 runs one Clang job per core, up to 8. Set the count with -j.
The simulation binary keeps its symbols. A sampling profiler attributes time to the generated functions, which are named after the module: ryusim::sim::Counter::eval() for combinational logic and ryusim::sim::Counter::eval_ff() for clocked logic. With Linux perf (it needs permission to open perf events; see kernel.perf_event_paranoid):
perf record -g ./obj_dir/build/counter_sim
perf report --sort symbol
The global -g option is accepted but does not change the generated build.
Every compile writes obj_dir/ryusim-route.json. It records the choices that change the generated code. Keep it with any timing you compare, since two builds of one design can take different paths.
grep -E '"(schedule|hot_opt|cold_opt|two_state)"' obj_dir/ryusim-route.json
"two_state": false,
"schedule": "trigger",
"hot_opt": "O2",
"cold_opt": "O0",
| Field | What it tells you |
|---|---|
schedule | trigger: the default. Processes are ordered at compile time and run when their inputs change. legacy-fixpoint: eval() is called again until nothing changes. The compile prints a Note: trigger schedule unavailable line giving the reason. |
hot_opt, cold_opt | The Clang optimization level of each build pool. |
two_state | true when the build used --2state. |
modules.<name>.emitter | typed or legacy, with the reason. One legacy module puts the whole design on the legacy-fixpoint schedule. |
modules.<name>.two_state_fraction | The share of the module's signals proven free of X and Z, which RyuSim stores as plain 0/1 values. |
modules.<name>.execution_model | Process counts, comb_cycles (combinational loops found), and nba_buffered with nba_kept_why (nonblocking assignments that still go through a next-value buffer). |
-O (long form --opt-level) takes 0 to 3; the default is 2. It is a global option and goes before compile: ryusim -O3 compile ā¦. Placed after compile, it is rejected.
-O sets the Clang level for the hot build pool only. RyuSim splits each module's generated C++ into two files. The hot file (<module>.cpp) holds eval(), eval_ff(), clocked processes and user functions and tasks. The cold file (<module>__slow.cpp) holds constructors, initialization, initial(), fork branch bodies and concurrent assertion evaluation. The cold pool builds at -O0. --cold-opt changes that; it is listed with the unstable options in the CLI reference.
-O0 shortens step 4 of the compile and slows the simulation. The build log prints each file's pool and level:
[3/7] CXX hot -O2 build/obj/ryu_schedule.cpp.o
[4/7] CXX hot -O2 build/obj/counter.cpp.o
[5/7] CXX cold -O0 build/obj/ryu_tu_slow_0.cpp.o
--opt switches RyuSim's own IR passes on or off. It takes a comma list; a leading - turns a pass off, for example --opt=-nba-elision. It repeats, and a later setting wins. Every pass is on by default, except as noted. Turning a pass off is a way to isolate a suspected RyuSim bug, not a tuning step.
| Pass | What it does |
|---|---|
coverage | Coverage instrumentation. Runs only with --coverage. |
flat-eval | Marks processes eligible for inlining into the flat evaluation. |
dce | Removes logic that nothing observable reads. |
two-state | Proves which signals can never hold X or Z, so they are stored as 0/1 values. |
cse, const-fold | Common-subexpression elimination and constant folding. |
fusion | Fuses continuous assignments. |
handle-elide | Drops reference counting for function-local class handles that do not escape the function. |
process-class, trigger-schedule | Classify processes and compute each process's triggers and a static evaluation order. |
nba-elision | Removes next-value buffers for nonblocking assignments where evaluation order makes them unnecessary. |
partition | The --threads partitioner. On only when --threads is above 1. |
preeval | Folds calls to pure functions with constant arguments at elaboration. |
--threads N is a compile option. It asks RyuSim to split the design's RTL into partitions at compile time. When the split is accepted, the simulation runs each evaluation phase on N threads: the main thread plus Nā1 workers, with a barrier between phases. Without the flag, or with 0 or 1, the build is single-threaded. The Clang build in step 4 is parallel either way; -j controls it.
The partitioner groups work by clock domain and by module instance, and gives each signal one owning partition. It keeps these on the main thread: processes that call system tasks or DPI, use hierarchical references, force/release, intra-assignment timing or anything that suspends; signals written from an interface; signals written by a reactive-region process. A coverage build does not hand whole instances to worker threads.
It then estimates the speedup from its cost model and the barrier cost. Below 1.2Ć it refuses. The compile continues with a single-thread schedule and prints a note on stderr:
ryusim -q compile --threads 2 --top counter -o obj_t2 counter.sv
Note: partition: refused ā predicted speedup 0.0123457x < 1.2x at N=2 (serial=2, parallel=162, ⦠phases=1; ā¦); emitting single-thread schedule
Note: trigger schedule unavailable ā --threads 2 (the partition schedule runs on the eval() surface); this build runs the legacy-fixpoint adapter (route manifest "schedule": "legacy-fixpoint")
The second note is the limit to watch. Any --threads value above 1 moves the build from the trigger schedule to the legacy-fixpoint schedule, whether or not the partition was accepted. The route file shows it:
./obj_t2/build/counter_sim && grep '"schedule"' obj_t2/ryusim-route.json
So --threads helps only when the partition is accepted and the parallel phases win back more than the schedule change costs. Time the threaded build against a build without the flag, on the same machine, before you keep it.
While a threaded simulation runs, idle workers spin for about 10 ms before they sleep. A build with --threads N keeps N cores busy, so leave cores free for the rest of the machine's work, cocotb's Python included.
--2state (alias --two-state) builds a model in which every signal is 0 or 1, and uninitialized variables start at zero unless you pass --init-reg or --init-mem with one or random. It departs from IEEE 1800 on purpose.
It is sound for a design and testbench that never depend on X or Z. It changes results when the design resolves buses through Z, reads undriven nets, divides by zero, checks for X with $isunknown or X-propagation assertions, or relies on X from uninitialized registers to catch reset bugs. Initialization and 2-state gives the exact rules.
Three limits:
--2state; Limits and errors lists them.ryusim-route.json lists each one under two_state_residue with the process, the reason and the source line.A build without --2state already stores the signals it proves X/Z-free as 0/1 values (two_state_fraction in the route file). --2state extends that to every signal.
Every generated file includes RyuSim's runtime headers. Clang builds them once into a precompiled header (PCH), and the builds share it through a cache directory: $XDG_CACHE_HOME/ryusim/pch, or $HOME/.cache/ryusim/pch when XDG_CACHE_HOME is unset. The first compile after an install or upgrade builds the PCH; later compiles reuse it. ryusim -v compile ⦠prints the cache directory and key.
An entry is keyed on the RyuSim version, the contents and modification times of its header tree, the Clang version, and the exact flags of the build pool. A new RyuSim version, another Clang, or another -O level creates new entries. RyuSim never deletes old ones, so the directory grows. Clear it when no compile is running:
rm -rf "${XDG_CACHE_HOME:-$HOME/.cache}/ryusim/pch"
--pch-cache DIR puts the cache somewhere else, for example a directory your CI caches between jobs. --pch-cache off builds the PCH inside each output directory, on every clean compile. If the cache directory cannot be created, the compile warns and falls back to off.
These follow from how RyuSim generates and schedules code. The route file shows most of them.
two_state_fraction means more of this work.--2state too."emitter": "legacy" moves the whole design to the legacy-fixpoint schedule.combinational cycle through <signals>, and the loop is evaluated again until it settles, up to --max-settle-iters times per time slot.q <= q + 1), several processes write it, or it belongs to an interface or an unpacked array. nba_kept_why counts the reasons.-O0. Fork branches and concurrent assertion evaluation are in the cold pool.--trace-vcd, --trace-fst and --coverage add work to the simulation. Turn them off when you measure.