Performance

Make RyuSim simulations faster: threads, optimization levels, 2-state mode, and how to measure.

A RyuSim run costs time in two places. ryusim compile turns the design into C++ and builds it with Clang. The simulation binary then runs the design. Measure the two separately, and check which build path RyuSim took before you compare two runs.

Current limits

Read these before tuning anything:

Measure compile and run separately

This design runs 100 clock cycles and prints the count:

module counter; logic clk = 0; logic [31:0] count = 0; always #5 clk = ~clk; always_ff @(posedge clk) count <= count + 1; initial begin #1000; $display("count=%0d", count); $finish; end endmodule

Time the compile. -q is a global option, so it goes before compile:

time ryusim -q compile --top counter counter.sv

Time the simulation on its own:

time ./obj_dir/build/counter_sim

Compile time has two parts. RyuSim parses, elaborates and writes C++ (steps 1 to 3 in the compile output); step 4 runs Clang through ninja. --codegen-only stops after step 3, so timing it gives RyuSim's share:

time ryusim -q compile --codegen-only --top counter -o obj_gen counter.sv

A second compile into the same output directory is incremental. RyuSim rewrites only the generated files whose text changed, and ninja rebuilds only those. For the full cost of a compile, time it into an empty directory.

Step 4 runs one Clang job per core, up to 8. Set the count with -j.

Find where simulation time goes

The simulation binary keeps its symbols. A sampling profiler attributes time to the generated functions, which are named after the module: ryusim::sim::Counter::eval() for combinational logic and ryusim::sim::Counter::eval_ff() for clocked logic. With Linux perf (it needs permission to open perf events; see kernel.perf_event_paranoid):

perf record -g ./obj_dir/build/counter_sim perf report --sort symbol

The global -g option is accepted but does not change the generated build.

Check which path the build took

Every compile writes obj_dir/ryusim-route.json. It records the choices that change the generated code. Keep it with any timing you compare, since two builds of one design can take different paths.

grep -E '"(schedule|hot_opt|cold_opt|two_state)"' obj_dir/ryusim-route.json
"two_state": false, "schedule": "trigger", "hot_opt": "O2", "cold_opt": "O0",
FieldWhat it tells you
scheduletrigger: the default. Processes are ordered at compile time and run when their inputs change. legacy-fixpoint: eval() is called again until nothing changes. The compile prints a Note: trigger schedule unavailable line giving the reason.
hot_opt, cold_optThe Clang optimization level of each build pool.
two_statetrue when the build used --2state.
modules.<name>.emittertyped or legacy, with the reason. One legacy module puts the whole design on the legacy-fixpoint schedule.
modules.<name>.two_state_fractionThe share of the module's signals proven free of X and Z, which RyuSim stores as plain 0/1 values.
modules.<name>.execution_modelProcess counts, comb_cycles (combinational loops found), and nba_buffered with nba_kept_why (nonblocking assignments that still go through a next-value buffer).

Optimization level: -O

-O (long form --opt-level) takes 0 to 3; the default is 2. It is a global option and goes before compile: ryusim -O3 compile …. Placed after compile, it is rejected.

-O sets the Clang level for the hot build pool only. RyuSim splits each module's generated C++ into two files. The hot file (<module>.cpp) holds eval(), eval_ff(), clocked processes and user functions and tasks. The cold file (<module>__slow.cpp) holds constructors, initialization, initial(), fork branch bodies and concurrent assertion evaluation. The cold pool builds at -O0. --cold-opt changes that; it is listed with the unstable options in the CLI reference.

-O0 shortens step 4 of the compile and slows the simulation. The build log prints each file's pool and level:

[3/7] CXX hot -O2 build/obj/ryu_schedule.cpp.o [4/7] CXX hot -O2 build/obj/counter.cpp.o [5/7] CXX cold -O0 build/obj/ryu_tu_slow_0.cpp.o

IR passes: --opt

--opt switches RyuSim's own IR passes on or off. It takes a comma list; a leading - turns a pass off, for example --opt=-nba-elision. It repeats, and a later setting wins. Every pass is on by default, except as noted. Turning a pass off is a way to isolate a suspected RyuSim bug, not a tuning step.

PassWhat it does
coverageCoverage instrumentation. Runs only with --coverage.
flat-evalMarks processes eligible for inlining into the flat evaluation.
dceRemoves logic that nothing observable reads.
two-stateProves which signals can never hold X or Z, so they are stored as 0/1 values.
cse, const-foldCommon-subexpression elimination and constant folding.
fusionFuses continuous assignments.
handle-elideDrops reference counting for function-local class handles that do not escape the function.
process-class, trigger-scheduleClassify processes and compute each process's triggers and a static evaluation order.
nba-elisionRemoves next-value buffers for nonblocking assignments where evaluation order makes them unnecessary.
partitionThe --threads partitioner. On only when --threads is above 1.
preevalFolds calls to pure functions with constant arguments at elaboration.

Threads: --threads

--threads N is a compile option. It asks RyuSim to split the design's RTL into partitions at compile time. When the split is accepted, the simulation runs each evaluation phase on N threads: the main thread plus Nāˆ’1 workers, with a barrier between phases. Without the flag, or with 0 or 1, the build is single-threaded. The Clang build in step 4 is parallel either way; -j controls it.

The partitioner groups work by clock domain and by module instance, and gives each signal one owning partition. It keeps these on the main thread: processes that call system tasks or DPI, use hierarchical references, force/release, intra-assignment timing or anything that suspends; signals written from an interface; signals written by a reactive-region process. A coverage build does not hand whole instances to worker threads.

It then estimates the speedup from its cost model and the barrier cost. Below 1.2Ɨ it refuses. The compile continues with a single-thread schedule and prints a note on stderr:

ryusim -q compile --threads 2 --top counter -o obj_t2 counter.sv
Note: partition: refused — predicted speedup 0.0123457x < 1.2x at N=2 (serial=2, parallel=162, … phases=1; …); emitting single-thread schedule Note: trigger schedule unavailable — --threads 2 (the partition schedule runs on the eval() surface); this build runs the legacy-fixpoint adapter (route manifest "schedule": "legacy-fixpoint")

The second note is the limit to watch. Any --threads value above 1 moves the build from the trigger schedule to the legacy-fixpoint schedule, whether or not the partition was accepted. The route file shows it:

./obj_t2/build/counter_sim && grep '"schedule"' obj_t2/ryusim-route.json

So --threads helps only when the partition is accepted and the parallel phases win back more than the schedule change costs. Time the threaded build against a build without the flag, on the same machine, before you keep it.

While a threaded simulation runs, idle workers spin for about 10 ms before they sleep. A build with --threads N keeps N cores busy, so leave cores free for the rest of the machine's work, cocotb's Python included.

Two-state mode: --2state

--2state (alias --two-state) builds a model in which every signal is 0 or 1, and uninitialized variables start at zero unless you pass --init-reg or --init-mem with one or random. It departs from IEEE 1800 on purpose.

It is sound for a design and testbench that never depend on X or Z. It changes results when the design resolves buses through Z, reads undriven nets, divides by zero, checks for X with $isunknown or X-propagation assertions, or relies on X from uninitialized registers to catch reset bugs. Initialization and 2-state gives the exact rules.

Three limits:

A build without --2state already stores the signals it proves X/Z-free as 0/1 values (two_state_fraction in the route file). --2state extends that to every signal.

Precompiled header cache

Every generated file includes RyuSim's runtime headers. Clang builds them once into a precompiled header (PCH), and the builds share it through a cache directory: $XDG_CACHE_HOME/ryusim/pch, or $HOME/.cache/ryusim/pch when XDG_CACHE_HOME is unset. The first compile after an install or upgrade builds the PCH; later compiles reuse it. ryusim -v compile … prints the cache directory and key.

An entry is keyed on the RyuSim version, the contents and modification times of its header tree, the Clang version, and the exact flags of the build pool. A new RyuSim version, another Clang, or another -O level creates new entries. RyuSim never deletes old ones, so the directory grows. Clear it when no compile is running:

rm -rf "${XDG_CACHE_HOME:-$HOME/.cache}/ryusim/pch"

--pch-cache DIR puts the cache somewhere else, for example a directory your CI caches between jobs. --pch-cache off builds the PCH inside each output directory, on every clean compile. If the cache directory cannot be created, the compile warns and falls back to off.

What makes a design slow

These follow from how RyuSim generates and schedules code. The route file shows most of them.