RV32IM in gno: 48 instructions that rustc, clang, TinyGo and Zig already
emit. A guest here is not written in a new language and is not translated by
anything of mine. It arrives compiled by the real compiler and optimized by the
real optimizer, and this package is a decode loop.
Second guest on vmkit, which is the only way to find out whether
the host ABI was shaped around the first one. It was not: riscv implements
vmkit.Machine unchanged and shares the fuel meter, the snapshot codec and the
instance store with bf.
1m,err:=riscv.NewMachine(image,0x1000)// image is a flat .text blob2iferr!=nil{returnerr}3used,status:=m.Step(host,50_000)// running, halted, trapped, out of fuel4snap:=m.Snapshot()// resume in a later transaction
Does it clear the bar
The design issue set a kill criterion before any code existed: above roughly
15,000 gas per RV32I instruction the track is a curiosity, not a substrate, on
the reasoning that a guest would then get about 200k instructions per block,
which is a token transfer and not a program anyone would write without thinking
about it.
Measured 2026-09-25 with gno test -v (--- GAS:), gno built from
gnolang/gno master. The figure is a slope from two loop sizes, 4,006 and
16,006 instructions, so the fixed per-test overhead cancels. The slope is
stable to a few gas between runs; the fixed term is not, which is why it is a
slope:
rung
what changed
gas / instruction
0
decode.gno's helpers called per instruction
33,992
1
field extraction and the fetch inlined into the dispatch loop
18,834
2
predecode once at load, dispatch on a frequency-ordered opcode
8,487
It clears the bar with a 1.8x margin, which works out to about 354,000
guest instructions per block. The ladder is worth 4.0x end to end, and all of
it came from moving work off the per-instruction path rather than from doing
anything clever with the instruction set.
Reproduce with gno test -v . and read TestBenchRV32I1000 and
TestBenchRV32I4000.
Be careful what that number is not. 354k instructions is a real program and a
small one: a sort, a parser, a state machine, a few thousand iterations of a
loop. It is not a signature verification, which is millions of instructions, and
the chain already has secp256k1 natively for that reason.
Conformance
The official riscv-tests
suite passes: all 42 rv32ui cases and all 8 rv32um, embedded in
conformance_test.gno and run on every gno test.
This is the only thing here a third party wrote. Everything in riscv_test.gno
checks that the emulator agrees with its author's reading of the spec; this
checks that it agrees with the people who wrote the spec, using their
assertions, compiled from their source.
One case is skipped and the skip is declared rather than omitted:
rv32ui-fence_i. FENCE.I exists to make writes to the instruction stream
visible, and this machine predecodes its text and refuses stores into it, so
there is nothing to make visible. Self-modifying code is not unimplemented here,
it is excluded.
The suite's own environment sets up mtvec, delegates exceptions and enters
through mret, none of which exists on a hart with no CSRs. The corpus was
therefore built against a replacement riscv_test.h that keeps the _start
symbol and TESTNUM in gp and drops the machine-mode ceremony. The exit
convention is the suite's own, untouched: it already spells pass and fail as
li a7, 93; ecall, which is this host's exit syscall. What was replaced is how
a case starts and hands back its verdict, not what it checks. Rebuild it with
tools/riscv-guests/conformance.
Two faults were injected to confirm the suite bites rather than merely passing:
SRAI as a logical shift takes out rv32ui-srai and rv32ui-lui at case 3, and
REMU by zero returning 0 takes out rv32um-remu at case 8. Both are recorded in
the test file, and the harness's own pass and fail decoding is pinned by two
tests so a change that made every case report success would not read as green.
Guests a compiler produced
The claim this package rests on is that a program arrives compiled by the real
compiler, so two of them are shipped compiled and nothing is trusted to stand in
for that. Both are rebuildable: source, linker script and exact command are in
tools/riscv-guests, and the committed words
were verified to match a fresh build.
language
words
what it does
riscv.GuestFNV()
C, clang 19.1.7
49
hashes the call input with FNV-1a
riscv.GuestLedger()
Rust, rustc 1.98.1, #![no_std]
1,959
mint, send and burn from a script
1m,_:=riscv.NewMachine(riscv.GuestLedger(),riscv.DefaultEntry)2m.Step(host,50_000)3// input: "mint alice 100\nmint bob 50\nsend alice bob 30\n"4// output: "alice 70\nbob 80\n"
Between them they exercise what a hand-written program does not: a real function
prologue spilling ra and s0 to a stack the host set up, .bss addressed far
above a text segment the host refuses stores into, LLVM materializing constants
through lui/addi pairs at the sign-extension boundary that bit this package
twice, and MUL in an inner loop.
The Rust one goes further, and that is why it is worth its size. It links
Rust's real core and compiler_builtins, which it never calls directly: a
bounds check reaches core::panicking, and the balances are 64-bit, so a divide
reaches __udivdi3 because RV32 has no 64-bit divide instruction. The
emulator is running the standard library's own software arithmetic, not only
what the compiler emitted for my source. TestRustLedgerDoes64BitArithmetic
pins a balance past 2^32 to keep that path live.
Both pause and resume correctly when sliced, the Rust one seventeen instructions
at a time, in the middle of a library routine it never asked to call. That is
the property a realm depends on and the one most likely to break on code nobody
wrote to be sliceable.
A refusal from the ledger is halted, not trapped: line 2: insufficient balance, exit code 1. The program said no and the machine did not break, and a
chain has to be able to tell a user which happened.
The memory layout is not incidental. W xor X means a guest whose writable data
shared a page with its code would trap on its first store, so the linker script
is part of the ABI and each guest pins its own: the addresses in it are baked
into the instruction stream as lui immediates, so a shared script edited for
one guest would silently invalidate every image built against the old one.
What the GnoVM charges for, measured
The rungs above are not classic interpreter optimizations. They are answers to
how this host executes code, and switchcost_test.gno measures it directly,
200,000 iterations per figure:
what
gas each
one elementary operation, from the loop control
~280
a function call
466
a call through a table indexed by opcode
727
switch, hitting the first case
292
switch, hitting the 48th case
13,358
each case label scanned on the way
278
make([]byte, n), for any n
655,224, flat
Two of these decided the design:
The GnoVM scans switch cases in source order. Case 47 costs 45x case 0.
So the internal opcodes in predecode.gno are ordered by how often compiled
code executes them, and that ordering is load-bearing: sorting that list
alphabetically would make the interpreter slower. It also means a flat
48-case switch is only a win if the common instructions are near the top,
and that a table of function values (727 gas, position independent) beats a
switch for anything past the third case.
make does not charge for size. A megabyte costs what four kilobytes
costs. The flat 1 MiB address space is therefore free at load, and the thing
to economize is per-word work, not bytes.
The cost on the other side: predecode
Predecoding is 19,623 gas per word, paid at load and again on every resume,
because the arrays are not in the snapshot: only the memory pages are.
That is 2.3x what executing an instruction costs, which is high enough to state
plainly rather than bury:
It pays for itself once a slice executes about 1.9 instructions per word of
its text (19,623 / (18,834 - 8,487)). Any loop clears that immediately.
Straight-line code run once does not, and for that shape rung 1 is the faster
interpreter.
A large binary pays it every transaction. A 64 KiB text segment is 16,384
words, so a resume costs about 321M gas, roughly 11% of a block, before the
guest executes anything.
The fix is known and not done here: decode a word the first time it is reached
rather than all of them at load, keyed on a sentinel opcode. That costs one
comparison per instruction (~280 gas, about 3%) and charges only for code the
guest actually runs. Worth doing when a real binary shows up; not worth
guessing at now.
Merging the immediate decoder into the opcode switch already took predecode from
31,432 to 19,623 gas per word. The two were separate functions, and the second
one had a case clause listing 22 opcodes, every label of which was scanned.
Semantics
W xor X. A store into the text segment traps. Predecoding is only sound if
the code cannot change under the arrays, and refusing is also what an ELF text
segment does, so self-modifying code is not a thing this VM has. FENCE.I is
therefore unnecessary and FENCE is a no-op, which is what it architecturally
is on a machine with one hart.
exit(1) halts, it does not trap. A program that returns non-zero said no;
it did not break. A chain that conflates "your code failed" with "the VM failed"
cannot tell a user which one happened.
An unknown syscall traps rather than returning zero, so a guest built
against a newer syscall table fails loudly on an older host instead of silently
reading a success it never got.
The M extension's edge cases return values, they do not trap. Division by
zero gives all ones, and INT_MIN / -1 gives INT_MIN. The spec defines both
precisely, which is exactly what a chain needs: there is no host arithmetic
exception left to differ between nodes.
No F, no D, no C, no A. No floats means no rounding mode and nothing that
can differ between nodes. No compressed instructions means every instruction is
4 bytes and (pc - base) >> 2 is the whole fetch. No atomics because there is
one hart.
Misaligned is a trap, out of range is a trap. RISC-V leaves an access
outside physical memory to the platform, and a chain has to pick the behaviour
that cannot differ between nodes. Refusing is deterministic; wrapping would let
a guest alias two addresses and get different answers from different memory
sizes.
Snapshots
A 1 MiB address space that serialized in full on every pause would make
continuations cost more than re-running, which is the kill criterion vmkit set
for them. Memory is snapshotted by dirty page: a guest that has written two
pages pauses in about 8 KiB regardless of how much memory it was given, and
TestSnapshotCarriesOnlyTouchedPages asserts the proportionality rather than a
magic number.
The program is not written twice. It is in the memory pages already, marked
dirty by WriteImage, so a snapshot carries the text extent and re-derives the
predecode from the restored pages.
The syscall table
Eight calls, in the RISC-V Linux register convention a compiler already knows:
a7 selects, a0..a3 carry arguments, a0 receives the result. A no_std
guest needs three lines of inline assembly and no gno-specific runtime.
a7
call
maps to
93
exit(status)
halts the machine
64
write(fd, buf, len)
Host.Output
63
read(fd, buf, len)
Host.Input
90
height()
block height
91
now()
block time, Unix seconds
92
caller(buf, len)
the calling address
94
log(buf, len)
Host.Log
95
emit(type, typelen, body, bodylen)
Host.Emit
Deliberately small, and every one of them is something vmkit.Host already
offers. The Host is the audited surface; a syscall that reached past it would be
a second, unreviewed one. One call's buffer is capped at 64 KiB, because a guest
asking to write a gigabyte has found the host's allocator, not something to say.
Tests
The assembler in riscv_test.gno is written from the spec independently of the
decoders in decode.gno, so a field in the wrong place fails unless both files
put it in the same wrong place. decode_test.gno round-trips all five immediate
formats at their sign boundaries, which is where this went wrong twice while it
was being written: a 12-bit immediate makes addi x8, x0, 4000 load -96,
and a 13-bit B-type immediate makes a hand-written -8 branch jump forward by
4,088.
Part of moul/gno-contracts — moul's versioned gno.land contracts. See the repository for the full catalog, build/test tooling, and usage.
Dependency graph:
🧪 Highly experimental — potentially vibe-coded. Not audited; may break, change, or be removed at any time. Do not use with anything of value. Full disclaimer: DISCLAIMER.
Overview
Package riscv is an RV32IM emulator for gno.land: 48 instructions that every real compiler already targets.
A guest here is not written in a new language. `rustc --target riscv32im-unknown-none-elf`, `clang -target riscv32`, TinyGo and Zig all emit it, so a program arrives compiled by the real compiler and optimized by the real optimizer. The interpreter is a decode loop rather than a semantic minefield: fetch four bytes, switch on seven bits, do arithmetic.
No F or D. No floats means no rounding-mode ambiguity and nothing that can differ between nodes. The M extension's divide-by-zero and overflow cases are defined by the spec to RETURN a value rather than trap, which is exactly what a deterministic chain wants.
It implements vmkit.Machine, so it shares the fuel meter, the snapshot codec and the instance store with p/moul/x/vm/bf(/p/moul/x/vm/bf/v0). That sharing is the point: the second guest is what tells you whether the ABI was shaped around the first one.
1const( 2// PageSize is the snapshot granularity, not an MMU page: there is no 3// address translation here, only bookkeeping about what changed. 4PageSize=4096 5 6// MemSize is the address space. 1 MiB is far more than a no_std guest 7// needs and small enough that the flat allocation is not the cost. 8MemSize=1<<20 910pageCount=MemSize/PageSize11)
Memory is the guest's flat address space, with a snapshot that carries only the pages it actually touched.
The shape is the one the bf ladder argued for. Running memory is a flat []byte because a load or a store happens on a large fraction of instructions and an avl lookup per access would dominate the interpreter the way vmkit.Meter.Charge did there. Snapshot cost is solved separately, by a dirty bitmap: a program that writes one page pauses in one page, not in a megabyte, which is the same trick that lets a bf hello world pause in 43 bytes.
That answers the open question on the design issue ("memory size versus snapshot cost") without paying for it on every instruction.
DefaultEntry is where a flat text image is conventionally loaded and where execution starts. Low enough to leave a scratch page below it, aligned, and nothing about the machine requires it: NewMachine takes the entry point.
GuestFNV returns the flat text image of the compiled FNV-1a guest, ready for NewMachine at DefaultEntry. A fresh slice per call, so a caller cannot mutate the copy everyone else gets.
Image serializes instruction words little-endian, which is what RV32 is. Exported because anything assembling a program has to do exactly this, and getting the byte order wrong produces a decoder error a long way from the mistake.
1typeMachinestruct{ 2reg[32]uint32 3pcuint32 4mem*Memory 5 6// code is the text segment with its fields already pulled apart. See 7// predecode.gno for why the dispatch loop reads arrays and not memory. 8code*code 910statusvmkit.Status11trapstring1213// out is what the guest wrote through the write syscall, mirrored to the14// host as it goes. Kept here too so a snapshot round-trips it.15outLenint16}
LoadData places initialized data in memory, outside the text segment.
The other half of a flat-image loader. A program with a .data section needs its initial bytes in memory and needs to write them back afterwards, which the text segment forbids: W xor X is what makes the predecode sound. So the linker puts data at its own address and it arrives here separately, rather than being concatenated onto the code and silently becoming read-only.
The pages are marked dirty, so a snapshot carries the data the same way it carries the program.
Restore loads a snapshot. The program image is not carried separately: it lives in the memory pages, which were marked dirty when it was written, so a restored hart has its code without the instance storing it twice.
Snapshot serializes the hart: registers, pc, status, and only the memory pages the guest has written.
A 1 MiB address space serialized on every pause would make continuations cost more than re-running, which is exactly the kill criterion vmkit set for them. The dirty bitmap is what avoids it: a guest that has touched two pages pauses in about 8 KiB regardless of how much memory it was given.
Step runs until the program halts, traps, or spends `fuel` instructions.
Written the way the bf ladder concluded, and for the same measured reasons: the hot state is in locals and written back once, and the fuel counter is two locals rather than a vmkit.Meter, because calling a method per instruction cost 86% there and would cost more here, where an instruction is cheaper.
Two things in this loop look like premature optimization and are not, both measured in switchcost_test.gno on this VM:
The case order is the internal opcode order from predecode.gno, which is by dynamic frequency, because the GnoVM scans switch cases in source order at 278 gas each. Reordering the arms below changes what the interpreter costs.
Everything the loop touches is a local. Reading cop[i] through m.code.op instead would pay for two field lookups on every instruction.