LibJS: Add hand-written assembly bytecode interpreter - #8299
Merged
Merged
Conversation
awesomekling
force-pushed
the
asmint
branch
3 times, most recently
from
March 7, 2026 00:37
e23a205 to
91380a2
Compare
Replace individual bool bitfields in Object (m_is_extensible, m_has_parameter_map, m_has_magical_length_property, etc.) with a single u8 m_flags field and Flag:: constants. This consolidates 8 scattered bitfields into one byte with explicit bit positions, making them easy to access from generated assembly code at a known offset. It also converts the virtual is_function() and is_ecmascript_function_object() methods to flag-based checks, avoiding virtual dispatch for these hot queries. ProxyObject now explicitly clears the IsFunction flag in its constructor when wrapping a non-callable target, instead of relying on a virtual is_function() override.
Move Interpreter::get() and set() from the .cpp file into the header as inline methods. Make handle_exception(), perform_call(), perform_call_impl(), and the HandleExceptionResponse enum public so they can be called by the upcoming assembly interpreter's C++ glue code. Also add set_running_execution_context() for the same reason.
Move the Bytecode.def parser, field type info, and layout computation out of Rust/build.rs into a standalone BytecodeDef crate. This allows both the Rust bytecode codegen (build.rs) and the upcoming AsmIntGen tool to share a single source of truth for instruction field offsets and sizes. The AsmIntGen directory is excluded from the workspace since it has its own Cargo.toml and is built separately by CMake.
awesomekling
force-pushed
the
asmint
branch
2 times, most recently
from
March 7, 2026 10:51
878a051 to
2e24666
Compare
AsmIntGen is a Rust tool that compiles a custom assembly DSL into native x86_64 or aarch64 assembly (.S files). It reads struct field offsets from a generated constants file and instruction layouts from Bytecode.def (via the BytecodeDef crate) to emit platform-specific code for each bytecode handler. The DSL provides a portable instruction set with register aliases, field access syntax, labels, conditionals, and calls. Each backend (codegen_x86_64.rs, codegen_aarch64.rs) translates this into the appropriate platform assembly with correct calling conventions (SysV AMD64, AAPCS64).
Add a new interpreter that executes bytecode via generated assembly, written in a custom DSL (asmint.asm) that AsmIntGen compiles to native x86_64 or aarch64 code. The interpreter keeps the bytecode program counter and register file pointer in machine registers for fast access, dispatching opcodes through a jump table. Hot paths (arithmetic, comparisons, property access on simple objects) are handled entirely in assembly, with cold/complex operations calling into C++ helper functions defined in AsmInterpreter.cpp. A small build-time tool (gen_asm_offsets) uses offsetof() to emit struct field offsets as constants consumed by the DSL, ensuring the assembly stays in sync with C++ struct layouts. The interpreter is enabled by default on platforms that support it. The C++ interpreter can be selected via LIBJS_USE_CPP_INTERPRETER=1. Currently supported platforms: - Linux/x86_64 - Linux/aarch64 - macOS/x86_64 - macOS/aarch64
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds a new bytecode interpreter for LibJS that executes via generated assembly instead of the C++ indirect-threaded computed goto interpreter. The bytecode PC and register file pointer live in machine registers, opcodes dispatch through a jump table, and hot paths (arithmetic, comparisons, property access on simple objects) stay entirely in assembly. Cold and complex operations call into C++ helpers.
The assembly is written in a custom DSL (asmint.asm) that a Rust tool (AsmIntGen) compiles into native x86_64 or aarch64 assembly. Instruction layouts come from Bytecode.def and struct field offsets from a small build-time C++ tool, so the assembly stays in sync with C++ struct layouts automatically.
Supported on Linux and macOS, both x86_64 and aarch64. Enabled by default where supported. Set LIBJS_USE_CPP_INTERPRETER=1 to fall back.
Benchmarks look great. There's more work to do on squeezing the generated code for performance ofc :)
macOS (M3 MacBook Pro):
Linux (AMD Ryzen 7 PRO 8700GE)