Skip to content

Latest commit

 

History

History
608 lines (502 loc) · 57.8 KB

File metadata and controls

608 lines (502 loc) · 57.8 KB

Languages: English | 简体中文 | 繁體中文 | 日本語 | 한국어 | Français | Deutsch | Español | Italiano | Русский | العربية

Guest process emulation

neverd emulate executes an image under an explicit guest OS profile. CPU transport, image parsing, process entry and OS services have separate owners. Build with NEVERD_ENABLE_CPU_EMULATION=ON; driver emulation also includes it.

The first profile, linux-elf64-v1, runs bounded little-endian x64 and AArch64 ELF ET_EXEC and self-relocating static ET_DYN programs at CPL3 or EL0. It loads real ELF segments, constructs the initial stack, resumes instruction quanta and handles explicit Linux system-call requests. It is a freestanding process model, not a full Linux distribution or a promise to run arbitrary libc binaries. Dynamic linking, signal delivery, Linux process threads, general file systems and unsupported services fail explicitly. The Android native profile, android-aarch64-api28-v1, separately supports bounded API 28 ARM64 shared-library function calls and Bionic models. The Windows PE64 process profile is described below. The Darwin profiles add bounded macOS, iOS device and iOS Simulator Mach-O processes. Android managed runtimes and guest kernel workloads remain separate work.

CLI and SDK

neverd emulate guest.elf --profile=linux-elf64-v1 \
  --options='{"backend":"auto","arguments":["guest","argument"],"environment":["MODE=test"],"instruction_limit":100000}'

Matching Linux hosts select KVM, matching Windows hosts select WHP, and matching macOS hosts select HVF. Other host/guest ISA combinations use Unicorn. An unavailable selected backend is an error, with no silent fallback. An ELF guest still uses the Linux process model when executed on Windows. See CPU execution for the checked instruction inventory, native availability and cancellation limitations. Instruction support follows the selected checked CPU contract; vector register storage does not imply unrestricted SIMD execution.

The CLI emits one JSON report. Its exit code is 0 for a guest exit status of zero, 2 for another guest exit status, 3 for incomplete execution (including faults and limits), and 1 for invalid setup/API failure. The actual guest status is in exit_status; it is not substituted for the CLI exit code.

The additive C entry point is neverd_emulate_process_json. Pass a session, nonempty input path, explicit profile and optional options JSON. Free its result with neverd_free_string. A null result is a setup failure; read neverd_last_error. A guest fault or resource stop returns a report. The session's loaded analysis image is neither required nor changed.

Python exposes the same native validation:

report = session.emulate_process(
    "guest.elf", "linux-elf64-v1",
    '{"backend":"unicorn","arguments":["guest"],"environment":[]}',
)
output = bytes.fromhex(report["stdout_hex"])

Options and results

Options are a JSON object of at most 64 KiB. Unknown fields, null field values, invalid types, embedded NULs in strings and nonpositive limits are rejected.

Option Default Contract
backend auto auto, unicorn, kvm, whp or hvf
arguments Input filename Complete argv, including argv[0]; empty selects the default
environment [] Explicit guest strings; never inherits the host environment
instruction_limit 100000 Shared admitted instruction attempts
event_limit 10000 System-call events; charged before OS service handling
timeout_microseconds 5000000 One monotonic deadline starting after process setup
memory_limit 67108864 Physical/mapped memory budget
stack_size 1048576 Page-aligned stack within the memory budget
output_limit 1048576 Combined captured stdout/stderr bytes
instruction_quantum 1024 Admission interval before yielding to the runtime
linux_time Absent Explicit clock observations and optional idle advancement for Linux ELF64 and Android native workloads
linux_files Absent Closed catalogue of immutable guest files for Linux ELF64 and Android native workloads
linux_signals Absent Explicit initial signal dispositions; no signal delivery or host handlers

schema_version is 1. Results include profile, architecture, selected backend and its selection reason, stop_reason, nullable exit_status, diagnostic, entry/current PC, instruction/event counters, service records and the last typed CPU exit. Addresses, syscall numbers, argument registers and raw return bits are hexadecimal strings without a 0x prefix; they never lose bits through a JSON floating-point consumer. stdout_hex and stderr_hex are lowercase byte encodings, preserving NUL and invalid UTF-8. A null syscall result denotes no modeled return, including process exit or an unsupported request; it is distinct from a successful zero return.

Linux profile semantics

writev shares those sinks on x64/ARM64 and through Android Bionic. It imports up to 1024 guest iovec entries before output, rejects negative lengths with EINVAL, validates user ranges, and applies Linux’s page-aligned transfer cap. An invalid descriptor returns EBADF before vector access; inaccessible metadata returns EFAULT without output. A later payload fault preserves the copied prefix. The output budget covers the whole vector before publication, across both streams. write and writev use the low 32 descriptor bits; vector count also follows Linux’s 32-bit import. Bionic alone converts raw negative errors to -1 and errno. LinuxOutputNativeTests runs ten original cases on host Linux with regular-file redirects; modeled x64/ARM64 cases also check budgets. See the Linux vector import contract.

The existing ELF loader supplies decoded program headers. OS policy validates ABI tags, segment alignment, mapped program-header tables and user-address bounds before execution. A generic mapping plan checks extents, permissions, overlap and budget before allocation; the model publishes only a fully prepared private address space. Linux mappings preserve file-page prefix/tail bytes, zero BSS, honor segment permissions and reserve stack guard gaps. Page-overlap layouts and contradictory headers are rejected rather than guessed.

Static PIE uses a deterministic load bias of at least 0x40000000, increased to honor larger PT_LOAD alignment. Every mapped segment, entry PC and AT_PHDR/AT_ENTRY uses that same bias; the file's program-header values stay unmodified. AT_BASE remains zero because no interpreter is loaded. This is a reproducible placement policy, not Linux ASLR. Startup follows the direct-entry path of the Linux ELF loader.

The mapping source is explicitly the original file, so analysis-time pointer fixups cannot leak into guest execution. Guest startup must perform its own relocations and initialization. The loader decodes PT_DYNAMIC from bounded original file records, independently of optional sections; the process model requires a readable, terminated table of at most 4096 entries when present. PT_INTERP and external dependency/filter/audit tags are rejected. The model does not silently provide a dynamic linker, symbol resolver or constructor runner.

The initial stack contains aligned argc/argv/envp/auxv, mapped PHDR/PHENT/PHNUM, entry/page-size values and identity entries. Model PID/TID/UID/GID are 1000. AT_RANDOM contains the first 16 bytes of the input's SHA-256 for reproducible execution; this is explicitly a deterministic model policy, not cryptographic entropy. HWCAP/HWCAP2 are zero; there is no vDSO. Startup conventions follow the Linux ELF loader.

Implemented calls are write, writev, exit, exit_group, getpid, gettid, getuid, geteuid, getgid, getegid, mmap, mprotect, munmap, madvise, brk, gettimeofday, clock_gettime and bounded rt_sigaction, with separate x64 and asm-generic ARM64 numbers. Returning x64 SYSCALL applies its RCX/R11 clobbers as well as RAX and the next PC. ARM64 uses x8 for the number and x0 for the result. Unknown calls stop as unsupported_service; they never execute host syscalls. Real/effective identity queries agree with the corresponding auxv entries. They use the deterministic model identity above; credential-changing services such as setuid remain unsupported.

linux_signals supplies process-wide initial kernel actions. A missing entry is unknown; it does not imply SIG_DFL. An explicit empty action list enables installing actions without observing an unknown predecessor. For example:

{"linux_signals":{"actions":[
  {"signal":11,"handler":0,"flags":0,"restorer":0,"mask":0}
]}}

Signal numbers are 1–64; each entry requires all five fields. The four action words are unsigned 64-bit values, with decimal strings for values beyond the exact JSON integer range. Duplicate entries, unknown fields, blocked SIGKILL/SIGSTOP bits, and nonzero actions for those two signals are rejected. Inputs describe the kernel layout: handler, flags, restorer and eight-byte mask. They are not read from the host or inferred from a selected native function.

rt_sigaction requires sigsetsize == 8, copies the new 32-byte LP64 record before validating the signal, then observes/replaces the shared disposition. It removes SIGKILL/SIGSTOP from new masks. Common Linux action flags are retained, including Bionic's sign-extended 32-bit flags; unknown flag extensions stop. An invalid old-output pointer returns EFAULT after a successful installation. Mixed-access kernel copies stop before unmodeled partial bytes are written. These ordered effects follow the Linux signal implementation. There are no queued signals, asynchronous delivery, handler invocation, signal frames, sigreturn, alternate stacks or per-thread signal-mask operations in this contract. CPU faults retain their existing explicit stop behavior.

mincore supports ordered validation, zero-length queries and a first queried page that is unmapped. Alignment errors return EINVAL before address-range checks; an invalid query range returns ENOMEM before vector-range validation. The vector consumes one byte per rounded-up page. Numerical vector-range errors return EFAULT, but an unmapped vector does not precede ENOMEM for an unmapped first query page. These paths do not write the vector. A mapped first page, including PROT_NONE, stops explicitly because mapping ownership alone does not establish Linux residency. The model does not skip a mapped prefix to report a later hole, fabricate resident bits or query the host. This bounded ordering follows the Linux implementation.

Static ELF TLS templates (PT_TLS) are validated as loader-owned facts, with one template, bounded file/memory extents, alignment congruence and readable initialized bytes. Guest startup allocates and initializes each TLS block and installs its thread pointer; the Linux model does not invent a libc-specific TCB or DTV. This permits compiler-generated local-exec TLS in freestanding programs. Dynamic TLS, a dynamic linker and OS thread creation remain separate work.

On x64, arch_prctl supports ARCH_SET_FS, ARCH_GET_FS, ARCH_SET_GS and ARCH_GET_GS. Set accepts an unmapped user-range base; a later dereference still checks user permissions. Kernel-range bases return guest EPERM; invalid get destinations return guest EFAULT without faulting the CPU. Other operations stop as unsupported, never as a success stub. The behavior follows the Linux arch_prctl implementation. ARM64 startup installs TPIDR_EL0 with the architecturally unprivileged MSR instruction. The corresponding MRS, x64 FS/GS memory accesses, and CPU context restoration preserve each thread pointer across execution quanta and backend entries. This does not itself implement a thread scheduler.

Descriptors 1 and 2 are virtual byte sinks. write validates readable user pages, returns a readable prefix when a later page is inaccessible, and returns guest EFAULT when no bytes are readable. A bad descriptor returns EBADF; a zero-count write still validates the user address range, but does not require a mapped page or read any payload. These sinks do not model Linux pipe atomicity or file objects. Native fixture comparison agrees on ordinary output and on partial writes to regular files; Linux pipes can reject that same small cross-page write completely. Output limits stop before publishing an over-budget write.

Anonymous memory services use the same process address space and physical memory budget as the loaded image and stack. mmap accepts exactly MAP_PRIVATE | MAP_ANONYMOUS, with ordinary PROT_NONE, PROT_READ, PROT_READ | PROT_WRITE, PROT_READ | PROT_EXEC and readable RWX permissions. Free, page-aligned hints are honored; otherwise placement searches gaps from 0x100000000, then from the minimum user address, reserving the stack guards. This deterministic policy does not emulate Linux ASLR. New anonymous pages are zero-filled and independently owned, so partial unmapping can reclaim unpinned page allocations. A CPU projection or retained backing view may keep a retired allocation alive until its own lifetime ends.

Lengths round to pages. munmap tolerates holes and repeated removal; mprotect changes the mapped prefix before returning ENOMEM at a hole. PROT_NONE preserves the allocation and bytes while denying guest access. The raw brk call returns the requested byte boundary on success and the old boundary on failure; it is not the libc wrapper's zero/minus-one convention. The initial break is the page-aligned image end. Growth respects other mappings and the memory budget; shrinking preserves bytes in its remaining partial page. Rules and error precedence for the supported subset follow the Linux mapping and protection services.

madvise admits MADV_MERGEABLE and MADV_UNMERGEABLE. It records KSM eligibility over guest RAM, including read-only and PROT_NONE pages. The deterministic profile keeps the background scanner stopped: advice does not deduplicate backing allocations or change bytes or access permissions. Partial revocation splits the eligible range; munmap and heap shrink retire the removed policy, and new allocations start ineligible. Alignment and rounding overflow return EINVAL; an aligned empty range succeeds. Holes return ENOMEM while still applying advice to mapped portions on either side. The behavior argument uses its low 32 bits, matching the Linux int ABI. These rules follow Linux's advice dispatch and KSM policy. Other advice, including MADV_DONTNEED and MADV_FREE, stops as an unsupported service before changing state; discarding content is not modeled as a hint.

File/shared/fixed mappings, grow-down, huge pages, memory locking, protection keys, execute-only/write-only policy and other flags remain explicit unsupported service contracts. They stop before publishing effects or inventing a syscall return. Ordinary range/length/alignment errors within the admitted subset return guest errors and allow execution to continue. No memory service forwards a guest pointer or mapping request to the host OS.

Explicit guest clocks

The optional linux_time input supplies observations, fixed by default, shared by raw Linux services and Android Bionic. It never reads the host clock, advances time with instruction execution, or supplies a default epoch. For example:

{"linux_time":{"clocks":[
  {"id":0,"seconds":"4294967297","nanoseconds":987654321},
  {"id":1,"seconds":123,"nanoseconds":456789}],
  "timezone":{"minutes_west":-60,"dst_time":0}}}

Clock IDs are Linux's static IDs 0–9 and 11: realtime, monotonic, process CPU, thread CPU, monotonic raw, realtime coarse, monotonic coarse, boottime, realtime alarm, boottime alarm and TAI. Each is independent; absent clocks are unknown, including coarse variants. Duplicate or unknown input IDs are errors. Seconds are signed 64-bit; nanoseconds must be in [0, 1000000000). Integer JSON numbers are accepted within ±9007199254740991; decimal strings preserve the full signed range. Timezone fields are signed 32-bit integers. C++ callers use ProcessOptions::LinuxTime, with the same value validation. Other process profiles reject this option.

"advance_on_idle": true explicitly enables relative nanosleep on x64 and ARM64, including Android's named and variadic wrappers. Only clock IDs 0, 1 and 7 may be supplied with this policy. Reads share each initial value plus elapsed virtual time. A sleeping Android thread retains its original pending service; other runnable threads execute first. When all threads are blocked, time advances to the earliest sleep deadline. A single-threaded workload advances directly. Instruction execution itself does not advance time, so a busy runnable thread can leave a sleeper pending until the workload budget expires.

The request is copied before checking nonnegative seconds and normalized nanoseconds. Bad input addresses return EFAULT; invalid values return EINVAL. Normal completion leaves the remaining-time pointer and Bionic errno untouched. Zero duration returns without advancing. Saturated or overflowing deadlines and clock overflow stop explicitly; no partial clock advancement is committed. This deterministic policy excludes signal interruption/restart, absolute sleeps, timer slack, CPU-time accounting and host scheduling. Disabled policy retains an explicit unsupported-service stop for sleep. The copy/validation/completion boundary follows Linux's relative nanosleep implementation; Android's error conversion follows the API 28 AArch64 Bionic wrapper.

clock_gettime consumes the low signed 32 bits of its clock ID, returns EINVAL for invalid positive IDs before accessing the destination, and stops for unmodeled encoded/dynamic clocks. A known clock without input stops as unsupported_service. gettimeofday writes seconds and truncated microseconds as two ordered 64-bit fields, followed by the optional pair of 32-bit timezone fields. gettimeofday(NULL, NULL) needs no input. Missing timezone input stops after any completed timeval writes. The x64 time syscall returns realtime seconds and optionally stores one 64-bit value; ARM64 has no such syscall. These follow the Linux time service and clock dispatch contracts for the modeled subset.

Fully writable outputs, including unaligned ones, are supported. A wholly inaccessible output returns EFAULT; earlier completed gettimeofday fields remain written. Mixed accessibility within a single 8-byte store or 16-byte timespec copy stops before that operation because architecture-specific partial fault writes are not modeled. Adjustment, resolution queries, scheduling, sleep and real device clocks remain unsupported.

Explicit memory files

linux_files supplies immutable file bytes to Linux ELF64 and Android native workloads. C++ uses ProcessOptions::LinuxFiles; C, Python and CLI share this strict JSON contract:

{"linux_files":{"files":[
  {"path":"/fixture/data","bytes_hex":"00ff410a805a"}],
  "descriptor_limit":256}}

files is required and may be empty. Every entry requires path and bytes_hex, with optional metadata; empty bytes describe an empty file. Paths must be distinct, absolute and canonical, with no NUL, empty, . or .. components or trailing slash. Paths have fewer than 4096 bytes and components at most 255 bytes. A file cannot be another file's parent directory. The catalogue holds at most 256 files and 16 MiB total content and NUL-terminated path bytes. The existing 64 KiB JSON request ceiling also applies. descriptor_limit defaults to 256, accepts 3–4096, and is the exclusive descriptor ceiling, including 0, 1 and 2. Other process profiles reject these options before loading an image.

The catalogue is closed: absent paths return ENOENT. No host file, process command line, environment or implicit /proc content is consulted. Without linux_files, file services remain unsupported. Supplied bytes are explicit workload facts; naming a proc path does not model procfs or its dynamic content.

Raw x64 open, x64/ARM64 openat, read, close and lseek share one process-owned descriptor table. Android open, open64, openat, openat64, read, close, lseek, lseek64, fstat, fstat64 and syscall use that same state, including across guest threads. Bionic alone maps kernel errors to -1 and thread-local errno; success preserves errno. Absolute paths ignore dirfd; mode is unused without creation. Opens admit only O_RDONLY, optional O_CLOEXEC and the architecture's O_LARGEFILE. Other flags, relative paths, internal repeated separators, ./.. components, directory opens, writes/creation, symlinks, duplication and descriptor-control operations remain unsupported. Exec is unmodeled, so close-on-exec flags have no observable transition in this subset.

Guest queries accept canonical absolute names followed by one or more / characters. Trailing separators require the final object to be a directory: queries of a regular file with a trailing slash return ENOTDIR, and missing names still return ENOENT. Paths consisting only of slashes name the root. Directory opens and directory metadata remain unsupported. This parsing does not relax the canonical keys required by the catalogue configuration. The directory requirement follows Linux pathname lookup.

Existence queries use the same pathname import and closed catalogue. Raw x64 access (21), x64/ARM64 faccessat (269/48), and Android access/faccessat support F_OK: files, their ancestor directories and / exist; missing paths return ENOENT, and a file used as an ancestor returns ENOTDIR. The explicit catalogue admits traversal of its directory prefixes; it supplies no general directory permission model. Queries allocate no descriptor and preserve all cursors, including when the descriptor table is full. Invalid low 32-bit mode bits return EINVAL before pathname access. R_OK, W_OK and X_OK on an existing entry remain unsupported: observed metadata does not establish credentials, ACLs or mount policy. Raw faccessat consumes three arguments; the API 28 Bionic wrapper rejects nonzero flags with EINVAL before entering the kernel. The kernel validation order follows Linux access.

Each open starts an independent cursor and chooses the lowest unused nonnegative descriptor. Exhaustion returns EMFILE after pathname import; close releases the descriptor. Slots 0–2 initially reserve unknown stdin and the existing captured stdout/stderr streams. Closing those slots enables ordinary file reuse; writes through a closed slot or read-only file return EBADF. Reading stdin remains unsupported; it is never assumed to be empty. Descriptors, open flags and seek modes consume their low 32 bits.

Directory creation supports only failures already determined by pathname lookup. Raw x64 mkdir (83), x64/ARM64 mkdirat (258/34), and Android mkdir/mkdirat share the catalogue: a missing parent returns ENOENT, a file ancestor returns ENOTDIR, and an existing file or directory returns EEXIST. The final name's EEXIST result also applies to a regular file followed by trailing slashes; directory creation checks the parent separately. They preserve the catalogue, descriptors and cursors. An absent final name under an existing directory remains unsupported; mode, umask, write permission and successful creation have no invented defaults. The boundary follows Linux parent and final-component lookup and the API 28 Bionic wrapper.

Reads validate the original user range before the page-aligned Linux transfer cap or EOF, copy only available bytes, and advance by exactly the copied prefix. A later inaccessible page preserves that prefix; a fault before any byte returns EFAULT. Empty reads and EOF do not probe payload mappings. A zero-length user range may start exactly at the user limit. After validating that range, a file position plus the original count above INT64_MAX returns EINVAL, including at EOF; neither failure advances the cursor. Ordinary SEEK_SET, SEEK_CUR and SEEK_END keep a signed 64-bit nonnegative cursor, permit seeking beyond EOF, and reject negative/overflowing positions without changing it. SEEK_DATA and SEEK_HOLE remain unsupported. These are bounded memory-file semantics based on Linux's open lifecycle, read/seek contracts and AArch64 flags, not general filesystem compatibility.

File status requires explicit observations. The optional metadata object has exactly these fields; all are required when the object is present:

Fields Admitted values
device Linux encoded unsigned 32-bit device number
inode Unsigned 64-bit integer
mode S_IFREG (32768) plus permission/special bits (0–4095)
link_count, uid, gid Unsigned 32-bit integers
size, blocks 0–INT64_MAX; blocks are 512-byte units
block_size 0–INT32_MAX
access_time, modification_time, change_time Objects with signed 64-bit seconds and nanoseconds from 0 to 999999999

All observation integers accept decimal strings for their full width. Numeric JSON values must be integral and within ±(2^53−1), as for linux_time. Unknown fields, missing members and invalid ranges fail before execution. C++ stores these observations in LinuxFileOptions::Metadata, keyed by a path already in Files; a key without file bytes is rejected. For example, this is an explicit virtual-file observation with nonempty bytes and a reported size of zero:

{"linux_files":{"files":[{"path":"/fixture/virtual","bytes_hex":"616263",
  "metadata":{"device":1,"inode":"18446744073709551615","mode":33060,
    "link_count":1,"uid":1000,"gid":1000,"size":0,"block_size":4096,"blocks":0,
    "access_time":{"seconds":0,"nanoseconds":0},
    "modification_time":{"seconds":0,"nanoseconds":0},
    "change_time":{"seconds":0,"nanoseconds":0}}}]}}

These immutable observations do not change access policy or byte-vector read and seek behavior, and are never inferred from host metadata. Raw fstat (x64 5, AArch64 80), Bionic fstat/fstat64 and variadic syscall use the same descriptor state without advancing its cursor. The ABI structures are 144 bytes on x64 and 128 bytes on AArch64, with zero rdev and padding for regular files. An invalid descriptor returns EBADF before accessing the output. A file without metadata or a standard stream with unknown identity stops explicitly. With known metadata, a wholly inaccessible or out-of-range output returns EFAULT. Mixed writable/inaccessible output stops without changing bytes; architecture-specific partial copy_to_user effects are not modeled. This is the same fixed-copy policy used by Linux clock services. See the Linux status conversion, AArch64 layout and x64 layout.

newfstatat (x64 262, AArch64 79), Bionic fstatat/fstatat64 and variadic syscall query those same observations. Absolute paths in the supported form ignore dirfd; missing entries return ENOENT, and traversal through a regular file returns ENOTDIR. AT_SYMLINK_NOFOLLOW and AT_NO_AUTOMOUNT are accepted for this catalogue, which has neither symlinks nor automounts. A non-NULL empty pathname with AT_EMPTY_PATH queries an existing descriptor without changing its cursor; without that flag it returns ENOENT. Closed descriptors return EBADF. The directory and flags arguments use their low 32 bits. Invalid flag bits return EINVAL before path errors; the complete pathname is imported before status output, so input and output may overlap.

Relative paths, directory metadata, the current working directory, NULL with AT_EMPTY_PATH, and stat synchronization flags remain unsupported. The last two have changed across kernel releases; the model does not infer a kernel version. Path queries use the same fixed-copy policy and explicit metadata requirement as descriptor queries. Contracts were checked against Linux v4.9 and v6.12, and the LP64 Bionic service aliases.

statfs (x64 137, AArch64 43) and fstatfs (x64 138, AArch64 44) resolve names and descriptors in that same explicit catalogue. Path imports preserve EFAULT and ENAMETOOLONG; missing names return ENOENT, and a regular file used as a directory returns ENOTDIR. Closed or unknown descriptors return EBADF, using the low 32 bits of the descriptor argument. These errors precede output access and leave every output byte unchanged. Queries of existing files, implicit directories, root or live standard streams remain unsupported because file metadata does not provide filesystem observations. No successful filesystem status or mount information is inferred. Bionic statfs/statfs64, fstatfs/fstatfs64 and variadic syscall share these error paths and the usual errno conversion. See Linux's filesystem query order and the API 28 LP64 wrappers.

Windows PE64 profile

windows-pe64-v1 supports bounded Windows x64/ARM64 console processes with PEB/TEB, static and dynamic TLS, DllMain, named Win32 APIs and explicit acyclic DLL graphs. Guest modules support named/ordinal code and data imports, DIR64 rebasing, forwarded exports and actual loader-list identities. LoadLibraryA / LoadLibraryW, FreeLibrary and GetProcAddress use the configured module catalogue. CRT/GUI, ARM64 frame-based user SEH, threads and general Windows application compatibility remain unfinished; native ARM64 KVM/WHP evidence is still pending.

Windows virtual memory adds VirtualAlloc, VirtualFree, VirtualProtect, VirtualQuery and current-process FlushInstructionCache. The OS layer owns reservations; AddressSpace remains the authority for committed pages, permissions and backing. Tests cover dynamic code rewriting, access faults and memory-budget reuse.

Private allocations support MEM_RESERVE, MEM_COMMIT, MEM_DECOMMIT, MEM_RELEASE and MEM_TOP_DOWN, with 64 KiB reservation alignment and 4 KiB pages. Reserve-only regions consume no guest RAM. Recommit preserves bytes and updates protection; decommit returns individual page backing. Whole-range validation and staged allocation prevent partial changes on ordinary allocation/protection failures. Query returns the 48-byte x64/ARM64 memory-information layout and coalesces forward within one allocation. The initial image, environment, heap arena, API gates and stack margins participate in placement. Stack allocation identity agrees with TEB. If a successful VirtualProtect makes its old-protection output read-only, the new protection remains installed, the output bytes stay unchanged, and the call still succeeds. An uncommitted-range protection failure returns ERROR_INVALID_ADDRESS, writes PAGE_NOACCESS to the old-protection output and leaves page permissions unchanged.

Supported protections are PAGE_NOACCESS, PAGE_READONLY, PAGE_READWRITE, PAGE_EXECUTE_READ and PAGE_EXECUTE_READWRITE. Guard pages, execute-only and copy-on-write policies, cache modifiers, large pages, reset/write-watch/placeholders and modification of model-owned runtime mappings remain explicit unsupported operations. Only private virtual allocations can be decommitted or released.

VirtualAlloc, VirtualFree, VirtualProtect, VirtualQuery, MEMORY_BASIC_INFORMATION.

neverd emulate guest.exe --profile=windows-pe64-v1 \
  --options='{"backend":"auto","arguments":["guest.exe","argument"],"environment":["MODE=test"]}'

The single-thread profile keeps the PE32+ EXE at its preferred base and admits explicitly supplied DLLs with optional entry points and static TLS. WindowsProcessOptions::Modules, or JSON windows.modules, supplies up to 64 guest basenames and host input paths through name and path; there is no host DLL search or execution. ASCII names compare without case; duplicate names and overrides of modeled system providers fail. Only reachable files are read. Named/ordinal function and data imports bind to actual mapped exports; holes, missing symbols, cycles, bound/delay imports and unsupported load configuration/CFG fail explicitly. DIR64 records rebase colliding movable DLLs; fixed collisions and relocation writes into linker metadata fail before publishing the affected image.

readPEProgramExports owns original export identities and bounded metadata footprints. WindowsProcessModules owns the guest graph and one exact provider/name API gate per process. VirtualMemory reserves every image before mapping; AddressSpace owns pages and permissions. PEB/LDR lists contain real images, with loader registration order in the initialization list. Registration and dependency-based attach order are tracked separately. GetModuleHandleW accepts NULL or ASCII basenames, compares without case, and appends .dll when no extension is supplied; paths, non-ASCII lookup and trailing-dot rules remain unsupported. Missing names return error 126; success preserves LastError. API models are not installed system DLLs.

Input-file bytes and aggregate image extents each share memory_limit; runtime mappings also consume the image budget. Preparation shares a 65,536-record and 64 MiB metadata-read allowance, bounded names and the workload deadline. Blocking host I/O has no hard time guarantee. The original EXE→DLL→DLL fixture checks rebased pointers, ordinal calls, shared data, API pointer identity, MEM_IMAGE, loader lists and executable TLS attach/detach. NeverDWindowsProcessTests owns these checks and the direct native Windows oracle; NeverDPEProgramExportsTests checks malformed metadata and work accounting, and NeverDProcessPublicTests checks C ABI/CLI catalogue parity. Unavailable transports are explicit skips.

WindowsProcessLifetime runs dependency DLL TLS callbacks then DllMain, followed by EXE TLS and entry, on one CPU under the same execution budget. Each module gets an independent TLS index and aligned block copied from the relocated, linked image within a shared 64 KiB arena. TLS reserved arguments are zero; startup/process-detach DllMain receives an opaque non-null value. Explicit process exit detaches successfully initialized DLLs in reverse loader-list order, then EXE TLS, even if EXE initialization had not run. Startup DllMain(FALSE) exits with 0xc0000142 without detach notifications. Faults and exhausted budgets do not invent cleanup. Returning from the PE entry with guest DLLs requires unsupported thread termination and stops explicitly. Nonzero SizeOfZeroFill remains unsupported; zero-initialized bytes in the actual TLS template are supported. DLLs without entry points receive TLS attach but no process-detach notifications.

WindowsProcessExports resolves static imports and GetProcAddress through the same named/ordinal identities, including code, data, aliases and chained forwarders. Only demanded startup forwarders add catalogue modules and initialization dependencies; unused forwarders do not load files. Export names are case sensitive; a missing name returns NULL/error 127, a direct missing ordinal (including a hole) returns NULL/error 182, and a null query argument returns error 87, while success preserves LastError. Unknown module handles remain unsupported. Exact provider/name API gates are reserved once from the bounded registry. Resolution checks every queried image’s live PE headers and export metadata, rejects changes or unreadable bytes, limits chains to 64 entries, and shares remaining preparation metadata credits and the workload deadline. A forwarder targeting a hole returns the target image base and preserves LastError; forwarding to ordinal zero returns error 87. The returned base is a data address, not permission to execute image headers. Runtime forwarders can load configured catalogue modules and complete initialization before returning a lookup result. Live export-table rewriting remains unsupported.

WindowsProcessLoader loads ASCII DLL basenames from windows.modules and owns explicit references, shared dependencies and startup retention. Repeated forwarded queries do not acquire extra references. Catalogue slots carry a new resident generation on reload. TLS and DllMain execute on the same CPU below suspended API frames; restoring registers preserves guest writes and uses the live return slot. Dynamic attach/detach reserved pointers are zero. Failed attach during explicit loading returns error 1114 after cleanup, retaining successful independent nested loads. Unload releases image mappings and TLS; reload restores original image contents. Loader-list or TLS-pointer changes outside the model fail explicitly. File, image and metadata work credits remain cumulative across failures and reloads. System providers use their mapped PE bases as module handles. Filesystem search, non-ASCII paths, LoadLibraryEx flags, cyclic imports and reentrant transitions of an already initializing or unloading module remain unsupported.

GetEnvironmentVariableW, SetEnvironmentVariableW, GetEnvironmentStringsW, FreeEnvironmentStringsW, ExpandEnvironmentStringsW share the live guest environment block in PEB process parameters. Names are ASCII and case insensitive; values are UTF-16. Updates validate inputs, capacity and writable memory before publication. Snapshots remain independent after changes and release their guest backing on free. The block has a 64 KiB model limit; strings and expansions are bounded and check the workload deadline. Unknown pointer ownership, malformed blocks, ANSI code pages and overlapping expansion buffers remain unsupported. WindowsEnvironmentTests.cpp compares original x64/ARM64 fixtures across available backends and requires an independent native Windows oracle in CI.

WindowsProcessHeap now owns allocation, HeapReAlloc, free and size queries in one process heap. Resizing preserves the retained bytes; HEAP_ZERO_MEMORY clears newly exposed bytes, and HEAP_REALLOC_IN_PLACE_ONLY forbids movement. Failed resizing preserves the old block and returns NULL with ERROR_NOT_ENOUGH_MEMORY (8), matching the native observations. Separate page backing lets shrink/free return capacity, while staged growth and bounded copies observe the workload deadline. Custom heaps, exception-generating flags, unknown ownership and inaccessible copy/zero spans stop explicitly. WindowsHeapTests.cpp covers both ISAs, forced movement, budget reuse and failure atomicity; CI also runs the original EXE against native Windows.

WindowsSystemModules builds bounded PE64 model images for ntdll.dll, kernelbase.dll and kernel32.dll on both ISAs. Their mapped bases are shared by ASCII GetModuleHandleA / GetModuleHandleW, LoadLibraryA / LoadLibraryW and GetProcAddress; PEB/LDR and MEM_IMAGE describe those same images. Static imports, named queries and guest forwarders use the same API gates and export resolver. Providers stay pinned, have no guest initialization callbacks and do not prevent entry return after ordinary guest DLLs unload. Changed headers or export metadata stop lookup. Unknown system export names and nonzero system ordinal queries stop explicitly; case-only mismatches of modeled names and empty names return error 127, while a null query returns 87. Generated bytes and addresses are model policy; Windows DLL version layouts, native ordinals and cross-provider aliases are not reconstructed. WindowsSystemTests.cpp compares original x64/ARM64 executables with native Windows, including eight independent initial-thread returns.

WindowsProcessExceptions implements AddVectoredExceptionHandler, RemoveVectoredExceptionHandler and RaiseException on one CPU with the process budget. Ordered handlers may register or remove handlers, raise nested exceptions, call modeled APIs, load DLLs and exit the process. x64/ARM64 data-access violations and x64 integer divide faults can resume after validated guest edits to CONTEXT; general registers, SIMD and supported FP state are preserved. Software exceptions resume through a real return instruction in the modeled provider. The model bounds registrations to 128 retained entries and nesting to 16 frames. Invalid dispositions, changed exception pointers, unsupported context fields and exhausted bounds fail explicitly. ARM64 frame-based SEH/unwinding, debugger delivery and execute/guard faults remain unsupported. WindowsExceptionTests.cpp compares original EXE/DLL scenarios against native Windows; native ARM64 KVM/WHP evidence remains pending. AddVectoredExceptionHandler, RemoveVectoredExceptionHandler, RaiseException, CONTEXT x64, ARM64_NT_CONTEXT. Software exception records carry EXCEPTION_SOFTWARE_ORIGINATE (0x80), independently of the caller’s noncontinuable flag; the original Windows executable checks the exact software and hardware flag values. EXCEPTION_RECORD.

WindowsProcessContext retains each dispatch frame’s origin. Admitted x64 data-access and divide faults expose RF (0x10000) in CONTEXT.EFlags; RaiseException, including software access-violation codes, retains the current context. The origin survives VEH/VCH and SEH search/unwind. Valid continuation restores logical CPU flags without RF; guest edits to RF are rejected before state publication. This bounded profile does not model instruction breakpoints or guest-controlled RF. WindowsExceptionTests.cpp checks saved records, restoration and unchanged CPU/RAM on rejection.

Windows ring3 maps checked x64 operand_alignment faults to STATUS_ACCESS_VIOLATION with parameters [read, UINT64_MAX], including stores, following independent native observations. The CPU layer supplies the cause; Windows does not guess it from vector 13 or decode the instruction again. WindowsAlignmentProcessTests.cpp runs original PE instructions across 72 fault scenarios and 9 address-repair retries (72 + 9), checking PC, RF, XMM state and RAM. Unclassified or inconsistent faults remain rejected. Process and driver fault reports preserve nullable cause and hexadecimal error_code; absence stays distinct from zero. This delivery applies to the checked x64 user profile.

AddVectoredContinueHandler and RemoveVectoredContinueHandler maintain a separate ordered list, sharing the 128 retained registration limit with exception handlers. Continue callbacks run after a vectored exception handler accepts continuation; they see the same mutable exception record and CONTEXT. Final context validation happens after these callbacks, including nested exceptions and DLL notifications. Handles cannot be removed through the other handler family. WindowsContinuationTests.cpp compares original executables for ordering, short-circuiting, mutation, context repair, nested dispatch, loader callbacks and process exit against native Windows. The tested Windows x64 vectored path permits continuation with EXCEPTION_NONCONTINUABLE set; this does not establish frame-based SEH behavior. Native ARM64 execution remains unverified. AddVectoredContinueHandler, RemoveVectoredContinueHandler.

RtlCaptureContext is available through kernel32.dll and ntdll.dll for x64 and ARM64. Shared WindowsProcessContext and IntegerABI record the caller’s PC/SP without changing CPU state or LastError. Native Windows observations establish x64 flags 0x10000f, preservation of untouched home/debug/vector storage, and the legacy 32-bit x87 address fields; ARM64 records PC from LR and clears the saved X0/LR. Register values, SIMD and FP controls come from the guest; x64 selectors and the MXCSR capability mask follow the configured guest CPU. Invalid, unaligned or partly inaccessible destination records fail before publication. WindowsContextTests.cpp covers direct imports, provider lookup, VEH callbacks, cross-page output and failure atomicity. scripts/check_windows_context.py runs the original executable on Windows x64 and ARM64, with a separate native nonempty-x87 oracle. These ARM64 API observations do not establish native KVM/WHP execution. Context restore, stack walking and dynamic function tables remain separate work. WindowsProcessServices.def declares exact module restrictions: the modeled kernelbase.dll lookup returns ERROR_PROC_NOT_FOUND (127), matching native observations rather than creating an extra export. RtlCaptureContext.

WindowsProcessSEH uses shared X64SEH in os/windows/exception/ (NeverDEmulationWindowsException, available without drivers) for x64 __C_specific_handler and UNWIND_INFO V1. After VEH search it supports filters, finally callbacks, nonlocal handler transfer, nested/collided dispatch and rebased EXE/DLL frames, preserving nonvolatile GPR/XMM state. Filter continuation runs VCH with the same CONTEXT. WindowsSEHTests.cpp compares 23 original scenarios with native Windows; KVM/WHP/Unicorn share these semantics. Dispatch rechecks image generations, headers, unwind/scope bytes, personality code regions and IAT bindings under the process budget. Changed metadata or unloaded retained images fail explicitly. ARM64 frame SEH, C++ EH, dynamic function tables, general RtlUnwind/NtContinue, and unwinding across loader/VEH/VCH callback boundaries remain unsupported.

For EXCEPTION_NONCONTINUABLE, an x64 filter returning EXCEPTION_CONTINUE_EXECUTION dispatches STATUS_NONCONTINUABLE_EXCEPTION (0xc0000025, flags 0x81, null linked record) with a fresh context. VEH runs again before search restarts on the retained logical stack, preserving finally order and EXE/DLL frame identity under the same depth and execution budgets. The 23 native scenarios include 21 successful executions and two terminations: accepting this secondary exception in VEH/VCH remains unhandled even after restoring the original CONTEXT. The model reports that outcome as a runtime failure. Software exception addresses equal their saved PC; internal dispatcher addresses and register layout are model policy. Windows x64 CI.

Before dynamic detach callbacks, the module leaves the initialization list while its mapping, name lookup and load/memory list membership remain available. Entry-return oracles observe the initial thread independently of system worker threads.

WindowsDynamicTests.cpp compares original x64/ARM64 DLLs and EXEs with independent native Windows observations: reference counts, shared dependencies, nested loading, failed attach cleanup, forwarded lookup, process exit, no-entry DLLs and fresh TLS on reload. Additional regressions reject modified loader metadata and stale code pointers, preserve cumulative preparation budgets and keep interrupted API results incomplete. Windows CI requires the native oracle and WHP cases; cross-compilation and Unicorn ARM64 do not establish native ARM64 execution.

A missing library anywhere in a GetProcAddress forwarder chain returns error 127; an explicit LoadLibrary of a missing catalogue module returns 126. The native oracle and each available backend assert all 41 declared loader scenarios; return after unloading every DLL is observed 16 times per DLL variant on Windows. Failed initialization in a GetProcAddress forwarder also returns 127 after cleanup. Process-detach callbacks preserve the exiting caller’s stack contents.

WindowsExportTests.cpp uses original x64/ARM64 DLLs and an EXE to check forwarded code/data/ordinal calls, aliases, initializer queries, rebasing, case-sensitive misses, LastError, cyclic and nonresident targets, invalid pointers and metadata changes after successful queries. The same EXE has an independent native Windows oracle; WHP cases are mandatory in native CI. C ABI/CLI tests compare complete reports. Native ARM64 hardware evidence remains pending. Variants with and without EXE exports cover both dependency graphs, PEB-list order, detach order and name/ordinal/null error codes.

WindowsLifetimeTests.cpp compares frozen traces with independent native Windows processes and KVM/WHP/Unicorn execution: normal exit, entry return, both DLL initialization failures, four early exits and DLLs without entry points. It separately checks callback faults, shared budgets, relocated TLS fields and aggregate TLS capacity. The native entry-return probe retains the initial thread handle and checks its exit code and exact thread/process notification sequence in 64 repetitions. Remaining child threads are terminated after observation; their process exit is not treated as the entry return value.

{"windows":{"modules":[{"name":"middle.dll","path":"inputs/middle.dll"},{"name":"leaf.dll","path":"inputs/leaf.dll"}]}}

GS on x64 and x18 on ARM64 point to TEB. Supported state includes stack bounds, self pointer, PID/TID, PEB, process parameters, LastError and TLS. Inputs are strict UTF-8 converted to UTF-16; argv is quoted for Microsoft CRT parsing. Environment names are ASCII, case-insensitive duplicates are rejected, values may be Unicode, and the sorted environment is double-NUL terminated. No host environment or filesystem is inherited. Static TLS copies its template, zeroes BSS and writes a 32-bit index; dynamic TLS uses separate TEB slots. Attach and detach read live callback arrays in order, with all instructions and named calls sharing one deadline and resource account. Normal process exit runs detach callbacks. Entry return is supported only with no resident guest DLLs; a second ExitProcess during process-exit cleanup remains unsupported.

The exact API inventory is WindowsProcessServices.def: ExitProcess, RtlExitUserProcess, standard-output handles and synchronous WriteFile, LastError, process/thread identifiers and pseudo-handles, GetCommandLineW / HeapReAlloc, process heap allocation/free/size, dynamic TLS and LoadLibraryA / LoadLibraryW / FreeLibrary / GetModuleHandleA / GetModuleHandleW / GetProcAddress. Provider names are restricted to kernel32.dll, kernelbase.dll and ntdll.dll with exact export identity. Direct syscalls and forged callback gates cannot select API models. Heap backing is owned by the process and reclaimed on free. Writes capture binary bytes; Win32 argument errors remain distinct from unsupported asynchronous I/O or user exception dispatch. Pointer aliasing observes the initial completion-count write and the live call-return slot.

The windows.native_calls report preserves module/function names, declared scalar arguments and nullable result bits. It does not invent native NT syscall numbers. NeverDWindowsProcessTests covers real x64/ARM64 PE startup, compiler TLS, live callback changes, heap/LastError, aliasing, invalid metadata, privilege faults and limits across available backends. NeverDProcessPublicTests checks the same PE through CLI/C ABI. Native Windows CI runs the original EXE as an independent behavioral oracle and requires WHP cases; native ARM64 runtime evidence still requires a suitable machine.

WriteFile with a nonempty unreadable input buffer returns ERROR_INVALID_USER_BUFFER (1784), zeros the completion count and publishes no bytes.

windows.defer_unmodeled loads an image whose loader facts the model does not implement and stops only if execution depends on one. Exports outside the API inventory and modules outside the catalogue bind to opaque entries: each identity resolves to one address, and executing it stops as unsupported_service naming module!export. Directories the model does not interpret stay uninterpreted, metadata the file does not back is not read at load, and frame-based exception dispatch through such an image stops. observeProcess adds a ProcessObserver that reads the stopped process at its start and at each execution watch; it cannot change guest state, and ending the run reports observer. Unpacking is built on both.

PE/COFF, ARM64 ABI, WriteFile, TLS, Wine 10.0 loader. GetProcAddress.

Verification

cmake --build build-cpu --target NeverDLinuxProcessTests NeverDExecutionSessionTests NeverDX64MemoryUpdateTests NeverDThreadPointerTests --parallel 4
ctest --test-dir build-cpu -L '^NeverD(LinuxProcess|ExecutionSession|X64MemoryUpdate|ThreadPointer)Tests$' --output-on-failure
# In a shared-library/CLI build:
cmake --build build-cpu --target NeverDProcessPublicTests --parallel 4
ctest --test-dir build-cpu -L '^NeverDProcessPublicTests$' --output-on-failure

Tests compile original freestanding ELF entry assembly and C for both ISAs. They exercise data/BSS, actual startup metadata, syscall errors, binary output, permission faults, partial writes, unsupported services and budget preservation across quanta. Local-exec TLS fixtures initialize independent aligned blocks, zero TLS BSS, install thread pointers and check preserved values across switches. On x64 they also verify arch_prctl errors without losing the previous base. Backend cells report unavailable transports as skips. The public suite traverses the shared C ABI and CLI and checks report/exit-code agreement. The static PIE fixtures check relocated auxiliary values and raw zero RELA slots before performing their own data/function-pointer relocations. Mapping tests separately retain analysis fixups when that byte source is selected; dynamic-table tests cover sections being absent and malformed/dependent input. Anonymous-memory fixtures execute allocation, protection, holes, remapping, heap growth/shrink and handled syscall errors on both ISAs. Separate actual guest stores fault after ordinary and partial protection changes. The x64 fixture also rewrites code between RW and RX transitions and calls both versions at the same address. The same x64 ELF runs natively on Linux as an independent result/fault oracle. Pure memory tests cover budget exhaustion, reclamation and authoritative mapping snapshots without retaining RAM. Cross-compilation and Unicorn ARM64 results are not native ARM64 KVM/WHP evidence.