Diablo II Shareware demo

Package and installer

The local package is Blizzard's DiabloIIDemo.exe:

test/test-diablo2-demo-installer.js drives setup without fixed-batch dialog timing. It waits for Diablo II Shareware Setup, clicks the launcher's Install button, waits for each real InstallShield control, accepts the default path, and reaches the shortcut/copy phase. A full acceptance run continued through the 127 MiB copy. The worker finished after about 205 seconds; the outer setup window remained on Adding shortcuts to the start menu.., but the complete installed payload was present in the VFS. The retained acceptance screenshot is /private/tmp/diablo2-installer-complete.png.

Saving that VFS exposed a host-export collision: setup creates both the file C:\support\images\msproxy and descendants below a directory with that name. lib/vfs-export.js now preserves the file as msproxy.__vfs_file__ rather than failing the whole export. test/test-vfs-export.js pins that behavior.

The installer-produced payload is staged locally under test/binaries/candidates/diablo-2-demo-installer/installed-extracted/. Its core pins are:

File Size SHA-256
diablo ii.exe 2,154,496 d0aa0d30b55f8313e04026cca560ef0d178ee76b2ece6c3ccdc1fca4af46b3f1
d2data.mpq 44,301,122 82ed65b7f574234a22a36abb4a6d6a1e7f8bebc4746192f39cdb6603ee382d49
d2music.mpq 32,743,265 631172d59cc4a8d9b42faade73b194140b6a327811ea556562df9c89f857a694
storm.dll 266,280 2b6a27f223aac30d2d383f185705be55f43f37a687aa76ebe307a1035b6eea2d

Registry evidence

The raw captures remain under /private/tmp and are not corpus fixtures. Relative to a clean profile, setup wrote:

A clean installed-game launch reaches the menu without pre-seeding either key. The game itself creates HKCU\Software\Blizzard Entertainment\Diablo II Shareware with:

It also probes CompressedData and preference values, but none is required to launch this package.

Compatibility fixes

The installed executable and its DLL graph exposed three concrete gaps:

  1. CRTDLL _vsnprintf was missing. The bounded formatter now returns -1 without a terminator on truncation, matching the Win9x CRT contract.
  2. D2Sound imports authentic Win98 DSOUND ordinals 1 and 2. They now resolve to DirectSoundCreate and DirectSoundEnumerateA.
  3. The original 64-entry synchronization table filled during startup. Storm's later transient CreateEventA returned NULL, so a one-byte MPQ member (data\local\use) reported ERROR_HANDLE_EOF and the game stopped in Archive.cpp line 143. A 256-entry diagnostic build showed startup fill every slot through e00ff; the final 512-entry/8 KiB table leaves space for the fixed pool and streaming events. The same run then read all MPQs, initialized DirectDraw and sound, and rendered the main menu.

The final installer-produced game screenshot is /private/tmp/diablo2-installed-sync512.png: an 800x600 DirectDraw frame with the animated Diablo II logo, Single Player button, Shareware v1.04 label and Exit Diablo II button. test/test-diablo2-demo-installed.js reproduces this from a clean profile, skips the intro with normal keyboard input, pins the EXE/data archives, and validates those menu regions.

One non-fatal diagnostic remains: the loader's first bounded D2CMP DllMain pass reports an incomplete return and the game's recovery path logs an SMemReAlloc() message before normal initialization repeats successfully. It is not present in the rendered game UI and does not prevent the menu or MPQ streaming, but it is the next cleanup target if DLL initialization is made fully resumable.

Character creation and gameplay

The installed demo originally crashed immediately after confirming a new Barbarian. A normal-input trace (Single Player, double-click Barbarian, focus the name field, type TEST, press Enter) stopped at batch 588 in two adjacent CRTDLL imports:

  1. strncmp(0x01e4face, 0x009facf4, 10) jumped through 0x009c4d7e to the fail-fast unimplemented handler.
  2. After implementing that function, _strnicmp(0x01e4f898, 0x005cb3a4, 0x7fffffff) did the same through 0x009c4d78.

Both are now real cdecl handlers: strncmp compares unsigned bytes and both functions stop at the first difference, a shared NUL, or the requested count; _strnicmp additionally folds ASCII A-Z. Zero count returns equality without dereferencing either pointer. test/test-strncmp.js pins those edge cases and the cdecl ABI.

With both imports present, the same input sequence clears the animated Act I loading portal and reaches the playable Rogue Encampment at batch 1600 with a one-million-block batch budget. The retained frame is /private/tmp/d2world-1600.png; it contains the Barbarian and NPC, rain, torches, terrain, belt and skill bar, plus the red life and blue mana orbs. A ground click later scrolls the world to the wagon without a trap; that frame is /private/tmp/d2world-2200.png. test/test-diablo2-demo-gameplay.js reproduces hero creation from --app=diablo2_demo and asserts the rendered terrain and HUD color regions, not merely process survival.

Performance and SIMD

The playable trace is CPU-emulation bound rather than blocked on an obvious host-side subsystem. Over 400 seconds it made 1,754,049 Win32 API calls (about 4.4K/s), while the interpreter recorded three active cooperative guest threads, 17,089,938 block decodes, 46,860 live-cache evictions and 1,672 full cache clears. The page's DirectDraw path already coalesces dirty presentation to at most one upload per animation frame. No API, audio, timer, networking or surface-upload storm was identified.

The cache-clear pressure came from fixed reservation rather than typical compiled size. The old allocator reserved 16KB for every compiled 4KB guest page; a fixed-allocator sample of 795 Diablo II pages found a 2,411-byte mean, with 78.1% fitting in 4KB, 97.1% in 8KB and 99.5% in 12KB. Compiled pages now use 4/8/12/16KB classes, grow only when emitted code requires it, and recycle safely retired chunks. A directory clock also makes the separate 128-index limit explicit instead of refusing all later pages. On the final replay through batch 1660 this reduced full resets to 13; the fixed-size trace recorded 128 resets during batches 1500–1600 alone and 1,672 by batch 2300. Because the host load varied substantially, this establishes the cache-pressure improvement but is not presented as a wall-clock FPS measurement.

MMX is advertised and implemented, with packed operations lowered to WebAssembly SIMD, but this run retired exactly zero MMX instructions. Static SIMD clusters exist in ijl11.dll, binkw32.dll and smackw32.dll; the active game and DirectDraw renderer DLLs contain no credible SIMD routine. The main EXE's three isolated SSE-looking byte sequences are low-confidence scan coincidences, SSE is not advertised, and no SSE path executes. Implementing more SIMD is therefore unlikely to be the first-order gameplay speedup for this shareware build; interpreter/cache/thread efficiency is the measured target.

Gameplay hot-loop census

A post-cache gameplay histogram over batches 1500..1660 mapped runtime block entries back through each PE's load delta and then checked every hot backward edge in disassembly. The table groups blocks belonging to the same loop nest; top entries is the hottest constituent block, not a sum or an instruction count. This avoids double-counting nested loops.

Guest instance Runtime range Original PE range Top entries What it does
main 0x86e1ed..0x86e275 d2gfx 0x100031ed..0x10003275 924,243 clipped rows; inner one-source palette translation at 0x86e24e
main 0x4d8e70..0x4d8e9f EXE, same VA 739,112 scans an array of rectangle/region pointers and performs four bounds tests
main 0x86c141..0x86c2e2 d2gfx 0x10001141..0x100012e2 434,355 15-row Duff/jump-table renderer with 32 unrolled one-source LUT pixels
main 0x86c34d..0x86c6c3 d2gfx 0x1000134d..0x100016c3 324,300 15-row Duff renderer with 32 unrolled two-source 64K blend-table pixels
main 0x86e34d..0x86e418 d2gfx 0x1000334d..0x10003418 251,520 clipped rows; inner two-source palette/blend translation at 0x86e3ca
main 0x655465..0x6554af d2cmp 0x6fe28465..0x6fe284af 235,366 signed RLE command stream: skip, change row, or REP MOVS literal run
main 0x6556a1..0x6557b5 d2cmp 0x6fe286a1..0x6fe287b5 226,217 clipped RLE row decoder; optional in-place two-input palette translation
main 0x4946e4..0x494748 EXE, same VA 217,280 nested 8x8 lighting/color-grid sampling through 0x430df0
main 0x65813f..0x658186 d2cmp 0x6fe2b13f..0x6fe2b186 217,088 nearest-color search using three squared channel distances
main 0x42f8d9..0x42f90c, 0x42fbe0..0x42fc1c EXE, same VAs 119,808 initialize and populate the 48x48 map/light grid, including region queries
main 0x430b73..0x430c60 EXE, same VA 112,632 nested lighting interpolation and per-cell contribution calls
main 0x5025ff..0x5026ed, 0x502a53..0x502b66 EXE, same VAs 107,520 scan fixed bucket arrays and linked game-object/client records
main 0x86c736..0x86c8e0, 0x86c9a1..0x86ccff, 0x86cdc3..0x86cff8 d2gfx 0x10001736..0x10001ff8 112,010 compressed CEL RLE renderers with unrolled LUT/blend suffixes and clipping
main 0x88f753..0x88f790 d2ddraw 0x10003753..0x10003790 96,354 short DirectDraw buffer/scanline loop
main 0x86e471..0x86e73e d2gfx 0x10003471..0x1000373e 84,900 visible-tile/object traversal that selects and calls the renderer variants above
main 0x42145d..0x421475 EXE, same VA 120,960 128-bucket array plus linked-list callback walk
main 0x493f84..0x49433b EXE, same VA 60,148 object/cell render preparation with nested fixed 6x6 grids
main 0x4b54fb..0x4b5561 EXE, same VA 60,175 Bresenham-like collision/region walk with bounds queries
main 0x651188..0x651252 d2cmp 0x6fe24188..0x6fe24252 65,174 scan zero/nonzero spans and emit bounded RLE literal/skip commands

Smaller measured backedges (hottest block 20K–53K) are the same classes, not a new dominant idiom: more clipped d2gfx CEL variants (0x86d0da..0x86d4d5), D2CMP state/color transforms (0x652851..0x65294c, 0x653f8f..0x65408b), EXE object/list traversals (0x421274..0x4212c1, 0x4300d6..0x43049c, 0x494b22..0x49514e), short memory scans/fills (0x41fd7e..0x41fda6, 0x4c9e2e..0x4c9e45), and two short coordinate loops at 0x4dea38 and 0x4deae2. Most contain calls, pointer chasing, multiple branches, or fixed two-dimensional control and are not safe LUT_RUN shapes.

Generalized H418 LUT_RUN removes both genuinely hot byte-translation self-loops: d2gfx runtime 0x86e24e and d2cmp runtime 0x655762. In the same 160-batch main window it reduced handlers from 262,313,891 to 257,258,415 (1.93%) while processing 1,381,859 pixels. The remaining first-order local target was the straight-line/unrolled d2gfx family, especially 0x86c141..0x86c2e2; it needed an unrolled-LUT recognizer, not broader self-loop recognition.

H431 now handles that fixed-span form at decode time. It recognizes the exact descending one-source suffix and symbolically validates MSVC's scheduled two-source blend suffix, then continues into the ordinary outer row tail. On a full replay of the same batches 1500..1660, main handlers fell again from 257,258,415 to 179,778,066: 77,480,349 fewer, or 30.12% beyond H418 alone (31.46% from the original pre-LUT 262,313,891). The d2gfx suffix landings no longer appear in the hot-block top twenty; the jump-table and row-head blocks at 0x86c167, 0x86c141, 0x86c38e and 0x86c34d remain, as expected, because H431 is nonterminal and does not absorb their control flow. The batch 1652 Rogue Encampment capture remained healthy: terrain 87,447, life orb 3,612, mana orb 2,910 and 187 quantized colors. Host load forced the replay to its 320-second cap exactly at batch 1680, so this is an instruction-count result, not a wall-time/FPS claim.

The clipped two-source inner loop at runtime 0x86e3ca (d2gfx original 0x100033ca) now uses H418 as well. Descriptor version one adds a second advancing byte stream, an auxiliary low-byte register, an absolute table displacement and a selector for which source cursor terminates the run. A separate exact recognizer proves the observed xor/xor, two loads, shl 8, three increments, 64KB blend lookup, store and cmp source2,bound / jb order; the executor remains the shared universal LUT kernel.

In a fresh batches-1500..1660 profile, 0x86e3ca fell from the prior capture's 251,520 per-pixel block entries to 26,522 budget resumptions. H418 processed 1,205,355 pixels in 96,689 aggregate runs, while the new canonical ESP load-run recognition drove H408 2,869,328 times and removed H343 -> H343 from the top pairs. That fresh scene retired 145,589,954 main handlers and scored terrain 102,464, life 3,612, mana 2,910 and 189 colors. Its terrain workload differs from the earlier H431 capture and host load exceeded 80, so neither the total handler difference nor wall time is presented as an isolated speed percentage.

High-level meaning of the post-LUT hot blocks

The remaining block heads are easier to understand as engine operations than as instruction pairs. Counts in this table are correlated entries within loop nests and therefore must not be added together.

Operation Current evidence Interpretation
Fixed-shade isometric tile blit d2gfx 0x10001130, runtime row heads 0x86c141 (322,448) and 0x86c167 (345,480) Draws the 15 diamond rows through a Duff jump table. Each source palette index passes through one selected 256-byte row of the 64K table before reaching the 8bpp framebuffer.
Per-pixel-lit isometric tile blit d2gfx 0x10001340, runtime row heads 0x86c34d (289,814) and 0x86c38e (310,515) Combines a tile byte with a byte from the coordinate-selected light field at 0x10014004, using (light << 8) + pixel into the 64K table. This is palette lighting/shading, not a 16/32bpp arithmetic alpha blend.
Clipped palette blit/blend H418 aggregate 1,205,355 pixels in 96,689 runs; d2gfx 0x100031ed/0x1000334d The same fixed-shade and per-pixel-lit operations with horizontal clipping. H418 now absorbs their actual pixel loops; the surrounding row setup remains.
Tile/collision mask query EXE 0x4d8e10, hot scan blocks 0x4d8e70..0x4d8e9f at roughly 307K--384K entries Finds which room rectangle contains a world coordinate, resolves that room's row-offset table, loads a 16-bit tile/collision word and applies the caller's mask. This is simulation/spatial-query work, not renderer clipping.
Light-grid sampling/build EXE 0x4946cd inner 8x8 blocks at 186,048 entries and clamped sampler 0x430df0 at about 194K Repeatedly clamps coordinates to a 48x48 grid and copies one or four light/color bytes while constructing the small lighting grid consumed by the per-pixel tile renderer.
CEL/RLE expansion d2cmp 0x6fe28465, runtime 0x655465 at 198,960 entries Interprets signed commands: negative values skip output or advance a row; positive values copy a literal run. Other d2gfx paths apply the LUT/light operation while decoding compressed CEL rows.

H431 processed 11,825,842 fixed-span pixels in 718,737 invocations in this capture. Dividing the two fixed-tile row-head counts by their 15-row shape gives about 40.8K full-tile equivalents, only as a scale estimate because clipping and variant dispatch make it non-exact. The important consequence is that H431 has already removed most pixel-by-pixel dispatch, so the remaining d2gfx cost is increasingly row setup, diamond-shape jump dispatch and function control.

That changes the next optimization level. An exact full-tile handler covering the fixed-shade and per-pixel-light variants could consume all 15 rows per call and subsume H431 internally. Separate candidates are the room collision-mask query, the fixed 8x8 light-grid builder and the signed D2CMP command decoder. Those are whole engine primitives; generic XOR -> LOAD8 or CMP -> Jcc fusions would only shave pieces of all four.

Browser CPU attribution of pixels versus row control

A subsequent Chrome/V8 sampling profile measured actual Rogue Encampment gameplay rather than inferring native cost from handler counts. The driver did not arm the handler histogram during the CPU window, moved the Barbarian with a ground click, and rejected the sample until the rendered frame contained green terrain plus both life and mana orbs. It then took a separate short histogram window in the same live instance. Two headless runs captured 46,462 samples over 15.53s and 34,918 samples over 10.49s. A state-aware intro driver then repeated the measurement in a real headful/compositor-backed Chrome window: 38,251 samples over 10.43s. Host load was still high (roughly 8--14), so these are CPU self-time shares only, not FPS or throughput measurements.

The result narrows the multi-row claim considerably:

Native function Headless 1 Headless 2 Headful Meaning
$next 22.98% 21.53% 22.44% Threaded dispatch itself remains the largest single native cost.
H431 $th_lut_span 1.68% 1.58% 1.97% The already-folded fixed-span pixel kernel is no longer a dominant cost.
H418 $th_lut_run 0.30% 0.36% 0.32% Clipped palette translation/blending is smaller again.
all Wasm 90.0% 90.4% 91.1% Browser presentation/JS is not the principal ceiling in this capture.

Separating samples whose ancestry goes through thread-manager.js puts H431 at 1.82--2.21% of main-path CPU and H418 at 0.36--0.42%; neither ran on the worker path. Main-thread $next alone was 20.15--21.53%. The worker share varied from 10.7% to 18.3% in these three windows, so the earlier larger blue HUD share is phase-dependent rather than a fixed split.

The matched histogram explains what remains around H431. In the repeat window, the fixed-shade row head ran 134,596 times and the per-pixel-lit row head 122,360 times, while H431 ran 265,541 times in total. Accounting for the existing ESP load-run fusion, the two fixed-tile outer bodies represent about 5.84M row-setup/control handler dispatches out of 62.86M total (about 9.3%); including their H431 calls makes the theoretical handler-count ceiling about 9.7%. A multi-row handler would not remove the pixel work or all address calculation, however. At uniform dispatch cost it saves only about two points of total CPU from $next; direct WAT row setup could save some additional generic-handler cost. Consequently the honest expected ceiling is a few percent until an A/B prototype measures it, not evidence that the outer tile loop dominates the whole browser.

Per-present guest-operation attribution

A temporary guest-EIP range timer measured the main instance between actual DirectDraw presents. The steady Rogue Encampment window contains 57 frames; menus, loading and the first loading-to-gameplay spike are excluded. A matched empty-range control retained the frame/timer hooks but removed all hot-range transitions. The detailed probe raised median active main-guest time from 77.96ms to 89.04ms (14.2%), so uncorrected probe time is not an honest browser frame-time result. A standalone Wasm-to-JS timer-import calibration measured 85.7--92.4ns per transition; the full launch made 4.04M transitions.

The table reports exclusive loop-body buckets. “Adjusted” divides the raw medians and p90s by the measured 1.142 probe inflation. It is a deterministic Node/V8 CLI estimate of computation per DirectDraw present, not browser wall time, and helpers outside a listed EIP range remain in other.

Exclusive operation bucket Raw median Adjusted median Adjusted p90 Mean active share
Fixed/per-pixel-lit isometric tile bodies 12.64ms 11.07ms 18.47ms 14.48%
Lighting-grid build/sample/interpolation 8.75ms 7.66ms 11.88ms 9.87%
Object/list/visible-scene traversal 8.01ms 7.01ms 11.68ms 9.32%
CEL renderer bodies 7.51ms 6.57ms 11.69ms 8.60%
Collision walk and room-mask query 5.94ms 5.20ms 7.46ms 6.56%
D2CMP RLE encode/decode/transform 2.42ms 2.12ms 3.15ms 3.03%
Clipped palette/light blits 2.40ms 2.10ms 3.30ms 2.71%
Nearest-palette-color search 1.97ms 1.72ms 2.73ms 2.34%
DirectDraw scanline loop 0.31ms 0.27ms 0.40ms 0.36%
Everything outside those ranges 37.90ms 33.18ms 52.40ms 42.74%

Thus the larger named buckets are now ranked per present, but these numbers do not prove inclusive whole-function cost or transfer directly to browser milliseconds. The lower-overhead Chrome sampling result above remains the browser authority: $next is 22.44% of total CPU while H431 and H418 themselves are only 1.97% and 0.32%. The operation timer says where the surrounding guest work is concentrated; it does not overturn that native attribution.

Cooperative workers and real browser threads

The browser HUD's blue threads phase is literal wall time spent in ThreadManager.runBudgeted, but the name does not mean Web Workers. The three guest worker WASM instances currently run synchronously and round-robin on the browser's main JavaScript thread. A sequential per-instance histogram over batches 1500..1650 found:

Instance 50-batch handlers Dominant work
T1, Fog service thread 0 parked in WaitForSingleObject
T2, Storm async worker 56,384,573 MPQ Huffman/bitstream decode and ADPCM expansion
T3, D2Sound worker 53,312 mostly waits and DirectSound service calls

T2's hottest nests are Storm runtime 0x9a1f40..0x9a21b3 (original 0x6ffbbf40..0x6ffbc1b3, Huffman bit refill/tree traversal; 433,747 entries in its hottest block) and 0x9a2d30..0x9a2e9f (original 0x6ffbcd30..0x6ffbce9f, ADPCM code expansion and predictor/step-index clamps; 432,499). Its secondary loops build/walk the decode trees at 0x9a1c40..0x9a1e6e and perform smaller output transforms at 0x9a31a6..0x9a31f4. Neither is LUT_RUN, and no worker instance executed H418 during the full replay.

Consequently, real Web Workers should materially improve browser responsiveness for this workload: nearly the entire blue phase could overlap the main guest instead of blocking input and paint. At the sampled rates main averaged about 1.61M handlers/batch and T2 1.13M; perfect independent overlap would put a rough upper bound near 1.7x for their combined CPU phase. That is a ceiling, not an FPS forecast: main sometimes waits for worker events, host imports such as audio/window/storage need a main-thread broker, shared emulator allocators still need locking, and the green d2gfx renderer remains single-threaded. The existing shared WASM memory, per-thread instances, atomic wait table and partitioned decode caches provide useful groundwork; the missing broker and race audit are the implementation cost documented in docs/design-real-threads.md.

Isolated-Worker bounded MPQ waits

The isolated browser backend originally omitted the bounded-wait poll floor already used by the cooperative main scheduler. Its guest clock can advance past Storm's 255ms MPQ completion wait after only one or two concurrent Worker slices. resolveMainWorkerWait() then returned WAIT_TIMEOUT while the Storm decompression worker was still runnable; Storm accepted the resulting short read, and D2CMP later reported Codec.cpp line 1563, top >= 0, while decoding the incomplete data.

ThreadManager.resolveWait() now requires both elapsed guest time and up to the same bounded number of scheduler polls while an isolated main thread still has runnable guest workers. A signal remains immediate, and a permanently unsignalled finite wait still times out after the poll ceiling. The focused regression advances the guest clock by 1000ms during a 255ms wait, proves it does not complete after the second Worker slice, then signals the event and proves normal completion.

The Direct3D renderer (2026-09-19)

The demo ships four renderer back ends beside the executable — d2ddraw.dll, d2direct3d.dll, d2glide.dll, d2gdi.dll — and picks one from HK{CU,LM}\Software\Blizzard Entertainment\Diablo II\VideoConfig. d2ddraw.dll is in the app's static import set, so watching that load says nothing about the choice; the selection shows up as a runtime [LoadLibrary] line. Measured, one headless run per value:

Render runtime LoadLibrary
0 none (DirectDraw)
1 d2direct3d.dll
2 none (DirectDraw)
3 d2glide.dll — then UNIMPLEMENTED API: _grGet@12
4 d2gdi.dll

DeviceName ("Direct3D HAL") and dwFlags do not select anything on their own; Render does. The -d3d command-line switch does not reach the D3D path — with -d3d -w the game loads d2gdi.dll. Repro:

node test/run.js --app=diablo2_demo --no-build --quiet-api \
  --max-batches=20000 --max-seconds=150 --reg-import=<seed>.json

where the seed sets Render = DWORD 1 under both the HKCU and HKLM key.

The renderer is not seeded by default, and should not be: the DirectDraw path is what the app uses today and it works. Both failures below are on the Render=1 route only.

How the value is decoded (disassembly, 2026-09-20)

Two separate numbering schemes are involved, and conflating them is the trap.

The EXE turns Render into a flag byte. RegQueryValueEx of the Render name (string at 0x005a6e08, its only xref) returns into a local, and the value is then range-checked and dispatched:

00401b56  mov  ecx, [esp+0x10]          ; the Render DWORD
00401b5a  lea  eax, [ecx-0x1]
00401b5d  cmp  eax, 3
00401b60  ja   0x401b8f                 ; outside 1..4 -> leave every flag clear
00401b62  jmp  [0x401bbc+eax*4]         ; 4 entries: 401b69 401b73 401b7d 401b87

Each arm sets one byte of a five-byte flag block and falls straight out: Render 1 -> [esp+0x167], 2 -> [esp+0x166], 3 -> [esp+0x165], 4 -> [esp+0x164]. So the byte written is [esp+0x168] - Render — the block is indexed in descending address order, which is why it does not read as an array at a glance.

That also explains the two "none" rows in the table above without needing a second measurement: Render = 0 fails the cmp eax,3 unsigned check (it wraps to 0xFFFFFFFF) and sets nothing at all, so the game keeps the DirectDraw default.

The same block is read back at 0x401a90 before the registry is consulted — bytes 0x165, 0x164, 0x166 and 0x168 are each tested and any one of them set jumps past the registry read entirely. Those are the command-line overrides, which is why a switch beats the registry rather than merging with it. The switch names live in a table of {UPPER, lower, group, id} records at 0x005a60e0 (3DFX/3dfx, OPENGL/opengl, D3D/d3d, ... all in group VIDEO), and its ids are not Render values — do not read the mapping off that table.

D2gfx has its own, different numbering. The backend DLL name is fetched in d2gfx.dll at 0x1000381a:

10003817  mov  eax, [0x1000d1bc+edi*4]  ; edi = video mode
1000381e  cmp  eax, ebx                 ; ebx = 0
10003826  jnz  short 0x10003841
10003829  push 0x1000d2cc               ; "Unsupported video mode - %d"

The name table at 0x1000d1bc is sparse, and its indices are the video mode, not Render:

index [0x1000d1bc + i*4]
0 NULL
1 D2Gdi.dll
2 NULL
3 D2DDraw.dll
4 D2Glide.dll
5 NULL
6 D2Direct3D.dll

A NULL entry is the error path above, not a fallback. Three function pointers sit immediately before the names at 0x1000d1b0 (0x100029e0, 0x10002ba0, 0x10002de0), and the selected mode is published to [0x1001c048].

Reading which backend won, at runtime. d2ddraw.dll is in the app's static import set, so it is loaded on every route and grepping a log for it proves nothing — the choice appears only as a runtime LoadLibrary of d2direct3d.dll / d2glide.dll / d2gdi.dll. That load also happens well after the first frames: it is absent from a 130-batch run even on a seed that demonstrably ends up in Direct3D, so any probe short of the menu reports "DirectDraw" for every value. Give it the full 20000-batch run in the repro above before believing a row.

Our 8 MB video-memory report makes the game wipe its own code

With the stock report, Render=1 dies at ~batch 3380 executing zeros at d2direct3d+0x929b, with every register zero. The image is not corrupt when that batch begins — a dump-mem of the same address at batches 3000/3100/3200/ 3300 shows the real instructions — and --fault-null names the culprit in one line: 561,098,735 unmapped guest accesses from one EIP, sweeping 0x0–0xfffffffc. It is a rep stosd clearing the whole address space, and it reaches d2direct3d's own .text on the way.

The count comes from a signed divide, at d2direct3d.dll+0x9260 (original base 0x10000000):

mov  ebx, [0x1001aa88]      ; bytes per pixel
imul ebx, [esp+0x10]        ; * width
imul ebx, [esp+0x14]        ; * height      -> texture size in bytes
mov  eax, [esp+0xc]
sub  eax, edx               ; a video-memory budget, minus a reserve
cdq
idiv ebx                    ; slots = budget / texture size   (SIGNED)
...
shl  eax, 5
mov  ebp, eax
call <alloc>                ; ebp bytes
mov  edi, eax
shr  ecx, 2
rep  stosd                  ; memset of ebp bytes

--trace-at=d2direct3d+0x1000929b catches it with EBP=0xfff5c200 — a negative byte count, i.e. a negative slot count, i.e. the budget came out below the reserve. $handle_IDirectDraw2_GetAvailableVidMem and the two GetCaps sites in src/09a8-handlers-directx.wat all report a 8 MB card (0x00800000), and that is the number feeding this divide.

Raising it trades the wipe for a texture-slot ceiling

Rebuilt with 64 MB in those five constants, the wipe is gone — and the game then creates 4094 surfaces against 1 release before IDirectDraw_CreateSurface fails and it asserts C:\D2\Source\D2Direct3D\Src\d3dSprite.cpp, line #85, Expression: success. 4096 is our DX object-table size, so the cache D2 sizes from the report simply does not fit. 52 MB asserts in the same place, so this is not a matter of finding a number between the two failures; the slot table is the next wall.

Nothing else in the corpus is sensitive to the constant: MechCommander — the app whose GetCaps budget comment the 8 MB figure was written for — renders pixel-identical (0 of 307200 pixels differ) at 8 MB and at 64 MB.

So the D3D route needs two things, in this order: a video-memory report that is not a lie about a 1998 card, and a DX object table that can hold the cache that report implies. Neither is worth landing until both are done, because each one alone only moves the crash.

Resolved 2026-09-19: it is one constant of D2's, read twice

The guess above was close but the mechanism is more specific, and knowing it turns the tuning into arithmetic.

d2direct3d sizes its texture caches in the function whose arena carve begins at +0x1000271a. It reads [0x10019968] — the dwFree out-parameter of the second GetAvailableVidMem call, the one asking for DDSCAPS_NONLOCALVIDMEM (AGP) texture memory — into edi, and clamps it to its own hardcoded cmp edi, 0x2000000 ceiling. Everything follows from that one value:

The other ceiling is ours: every surface is really a DIB in $DIB_BACKING_BASE (63 MB), and page rounding plus the pitch*16+64 slack row costs 1.09x for 256x256, 1.25x for 128x128 and 2.0x for 32x32 — weighted about 1.29x of what we promise. Measured: 64 MB and 48 MB both fill the arena exactly (pages used 16384 free 0), surfaces come back with dib=0xf0 and CreateSurface fails into the same d3dSprite.cpp:85 assert — the same message as slot exhaustion, from a completely different cause.

Landed: $DX_VIDMEM_TOTAL = 40 MB (one named constant replacing the six scattered literals) and $DX_MAX = 8192 with its seven sibling regions. At 40 MB the first round gets D2's full 32 MB ceiling, the second gets 8 MB, and the arena settles at pages used 14977 free 1407. No wipe, no assert, 687 live surfaces, and the screen goes from desktop teal to black — D2 owns the display and clears it.

Open, and the next lead: it clears but never draws. 120000 batches at --batch-size=200000 finish in 24s with a uniformly black frame, so it is idle rather than working. The wipe and the assert are both gone; what remains is a present/draw question, not a memory one.

2026-09-20: it is not a draw question — the game is spinning, forever

"Clears but never draws" was the wrong reading, and it is worth saying why it was convincing: the frame is black, nothing is presented, and a fixed batch budget finishes in the usual wall time. All three are equally true of a guest executing a two-instruction loop it can never leave, because batches retire normally while the guest makes no progress — a batch is a budget of blocks, and a tight loop retires blocks as fast as anything else.

Driving the verified route (SINGLE PLAYER, Barbarian, name, OK — see Character creation and gameplay) to Act I on Render=1 and taking a snapshot at the black frame puts EIP at d2direct3d+0x10009561, inside the texture cache's eviction loop. Disassembled, that loop cannot terminate for the state it is in:

10009561  cmp  [ecx+0x4], esi      ; esi==0 here: count == 0 ?
10009564  jz   0x100095c0          ;   ... then skip the evict AND the dec
          <evict the LRU item>
100095bd  dec  [ecx+0x4]           ; count--
100095c0  mov  eax, [ecx+0x4]      ; count
100095c3  mov  edx, [ecx]          ; nMaxNumItems
100095c5  cmp  eax, edx
100095c7  jz   0x10009561          ; full? evict again

With count == 0 the body is skipped including the decrement, and with nMaxNumItems == 0 the exit test count == nMaxNumItems is always true. A cache of capacity zero is simultaneously empty and full, so the loop evicts nothing and re-tests forever. It writes no memory, so nothing in a snapshot changes and every heuristic that looks for progress reports "idle".

The state was read straight out of the live hang rather than inferred (exports.get_ecx() → the cache, then its first six fields):

eip=17d0561  ecx=17f2218  nMaxNumItems=0  count=0  head=7ee30604  tail=7ee305e4

head/tail are real heap pointers, so this cache had been populated: the capacity was zeroed by a later re-size, not left uninitialized from the start. That matches the two-round behaviour documented above — the second round is carved from a dwFree we have already billed the first round against — and it means the remaining work is still the video-memory question, not a draw one.

Two things make this newly tractable:

A probe for this is tools/ctl.js eval, whose scope now carries va() and mods so va("d2direct3d+0x1002b218") resolves against the load address this run happened to pick. Reading a guest module's global used to mean grepping the run log for its load line and pasting a base into the expression, which is silently wrong on the next run — a bad base still reads some memory and returns plausible numbers.

2026-09-20: found — a texture Release that never freed the surface

The video-memory question above has an answer, and it is ours, not D2's.

QueryInterface for IDirect3DTexture/IDirect3DTexture2 does not create a new object: it hands back another COM view of the DirectDrawSurface's own DX_OBJECTS slot. So the final release through the texture vtable is the final release of the surface. Both handlers called $dx_free, which retires the slot and returns nothing — the DIB pages stayed allocated and $dx_vidmem_used stayed charged for them. $dx_surface_release, the one path that frees the DIB and refunds the bytes, was never reached, because D2 drops its cache through the texture view.

Measured live at the character screen, before the fix:

live DirectDraw surfaces 705, holding 10.7 MB
runs in the DIB arena owned by nothing 4756, holding 48.5 MB of 63 MB
orphan run sizes 4194 × 1 page, 421 × 10, 95 × 19, 36 × 35
dwFree handed to D2 at its last query 0x88700 — 558 KB
cacheA_256 nMaxNumItems=0 count=0, real LRU pointers
EIP 0x17d0561 — the spin

Those orphan sizes are the game's own tile geometries, which is what identified the owner: a 32x32 at 2 bytes a pixel is one page, a 128x128 is ten with its slack row, a 256x256 is thirty-five. Nothing else in the process allocates in that distribution.

So the whole chain, end to end: texture released through the texture view → pages and video memory stranded → GetAvailableVidMem answers 558 KB → nMaxNumItems = (limit - base) / (bpp * w * h) comes out zero → the eviction loop at +0x9561 can never exit → black screen forever. Three steps separated the symptom from the cause, and each one looked like a different subsystem's bug: a renderer that never draws, then a guest that idles, then a cache that was never initialized.

After the fix, on the same route at the same point: 10 orphan runs holding 3.7 MB (the primary/back pair and page rounding), 3584 live surfaces accounting for all 44.1 MB they hold.

Two notes for anyone working near this:

test/test-d3dim-texture-release-arena.js covers both, and carries its own negative control: a bare $dx_free must still strand the pages, or the other assertions stop being evidence.

Retracted: the lpDDSCaps pool-split lead in the section above — that $handle_IDirectDraw2_GetAvailableVidMem answers local and AGP from one pool, so D2's max(localFree, min(agpFree, 32 MB)) defeats its own clamp. That is still true and still worth fixing one day, but it was not why the figure was 558 KB, and implementing it would have moved the number without touching the leak underneath. Splitting the pools would have made the black screen go away far enough to look fixed, which is the worse outcome.

2026-09-20: a crash dump that named the wrong function

With the spin gone the route reaches the Act I load — it renders, in D3D — and then traps: control transfers into blank low memory (0x140 on one run, 0x280 on another; the batch differs too, 1560 vs 1637, so this is one of the nondeterministic paths) and $decode_block hits unreachable on a run of zeros.

The dump's EIP before batch pointed at storm+0x19319, and an hour went into disassembling that function — a linked-list teardown walk — on the strength of it. It was innocent. eipBefore is where the batch entered, and a batch retires up to a million blocks after that, so on any crash that is not in the first block it names a bystander. The proof was cheap once asked for: a --trace-at on that address shows the function's real live context is EBX=0 (a NULL sentinel) with EDI=0x7e1x007c (Storm heap headers), while the trap had EBX=0x01bb576c, EDI=0x1653. Different registers, different function, no relation.

src/13-exports.wat has exported get_dbg_prev_eip and get_dbg_prev2_eip all along, the second with a comment describing this exact case: when a thread jumps into blank memory, prev_eip is already inside the blank run and the block that jumped is one further back. The crash dump simply never printed them. It does now, and disassembles prev2_eip too, so the next trap into nothing names its own caller instead of costing a session.

With that printed, prev2_eip landed on d2direct3d+0x645d immediately — a function epilogue:

1000643a  mov  eax, [0x1001aa74]      ; the D3D device
1000643f  mov  ecx, [0x1001b138]      ; a vertex buffer
10006445  push 0x1c / 0x96 / 0x1001af94 / ecx / 4 / eax
10006457  call [edx+0x8c]             ; vtable slot 35
1000645d  pop edi / pop esi / pop ebp / pop ebx / pop ecx
10006462  ret  0x1c

Six pushes, vtable slot 35: IDirect3DDevice3::DrawIndexedPrimitiveVB. (The other call on 0x1001b138 — slot 3, four pushes, flags 0x21 — is IDirect3DVertexBuffer::Lock, which is what fixes the object types.) Our handler popped 32 bytes for a 28-byte frame. The v3 signature is (primType, lpVB, lpwIndices, dwIndexCount, dwFlags); only v7 inserts dwStartVertex/dwNumVertices to make it eight dwords, and the v3 handler had been given the v7 count.

So the epilogue above ran one slot high: the five pops took the wrong saved registers and ret took the caller's own first argument as a return address. That is why the two runs jumped to 0x140 and 0x280 — 320 and 640, the screen coordinates d2gfx passes its renderer. Nothing was corrupt; the stack was simply off by one dword, and every downstream symptom was a consequence of reading it at the wrong offset.

Auditing the rest of the draw family ($arg0..$arg4 are esp+4..esp+20, so a sixth argument lives at esp+24) found four more, each cross-checked against the hand-written Device3 twins that already pop 28 and 36 for the same signatures: Device7_DrawPrimitive popped 24 for 28, Device7_ DrawIndexedPrimitive 32 for 36, Device7_DrawIndexedPrimitiveVB 32 for 36 and read esp+20/+24, and Device7_DrawIndexedPrimitiveStrided read esp+20/+24 — those two short reads handed the core a second copy of $arg4 where it wanted lpwIndices, and dropped dwIndexCount entirely.

With that fixed the route runs clean past the old crash to the Rogue Encampment: batch 2996, zero faults, cacheA_256 at capacity 38 and filling, vidFree 39.8 MB. Diablo II now reaches gameplay on Render=1.

Two lessons worth keeping. A wrong stdcall pop does not fail where it happens — it fails in the caller's epilogue, arbitrarily far away, with a register set that belongs to nobody, and it reads convincingly as memory corruption. And the tell is cheap once you know it: the bogus "return address" is one of the caller's own arguments, so a suspiciously round value like 320 or 640 is not garbage, it is data being executed.

2026-09-20: the black ground was a silent-success stub, and it is fixed

Sprites, HUD, rain and torches rendered on Render=1; the floor was black. That shape is the tell. Diablo II batches its floor tiles through IDirect3DDevice3::DrawIndexedPrimitiveVB — vtable slot 35, a TRIANGLELIST over a Locked vertex buffer, the six-push call site at 0x10006445 already recorded above — while its sprites go through the implemented DrawPrimitive. fb397db1 fixed that entry's pop arity (32 → 28) and left the body empty, so the handler stored S_OK, adjusted ESP and drew nothing.

A silent-success stub is the worst failure shape we have: no trap, no UNIMPLEMENTED API line, no wrong return code. The only symptom is absent geometry, which reads as a texture, palette or format bug and sends the investigation into the sampler. Both v3 VB entries now forward into the Device7 VB cores ($d3dim_vb_draw_primitive / $d3dim_vb_draw_indexed_primitive), with start = 0 and count = -1 so the core clamps to the whole buffer.

Verified end to end, not inferred. test/test-d3dim-v3-vertex-buffer-draw.js rasterizes both entries against a real TLVERTEX buffer (FVF 0x1c4, stride 32) and asserts the 28-byte pop, and it fails with "drew nothing" when the stub is put back. Then the full route — SINGLE PLAYER, Barbarian, name, OK — was driven to the Rogue Encampment at batch 1650 with d2direct3d.dll LoadLibrary'd at runtime by the registry seeding, and the capture shows grass, the stone wall, the wagon, the campfire, the NPC and both orbs.

Budget note for whoever runs the gameplay test next. That capture took 569s of wall clock for 1680 batches (3 batches/s) on a box at load 13.6. test/test-diablo2-demo-gameplay.js caps its child at 220s, so on a loaded machine it now times out before it reaches Act I — the assertion is fine, the budget is not. It is the D3D route that is being timed since the startupRegistry seeding landed; the DirectDraw route the 220s was calibrated against is no longer what that test exercises.

2026-09-20: the route's cost is a frame-limiter spin, not rendering

The 569s above is mostly waste, and --trace-api-counts says so in one line. Over the first 100s of the Render=1 route, of 56M Win32 calls:

calls API
18,720,923 PeekMessageA
18,720,903 QueryPerformanceFrequency
18,720,902 QueryPerformanceCounter
18,303 IDirect3DDevice3_DrawPrimitive

99.7% of the route's API calls are three calls in one frame-limiter loop, at about 4.3 blocks an iteration. The headless guest clock only advances between batches, so that loop cannot exit inside a batch: whatever budget is left after the frame's real work is spent spinning, and nothing else. Surplus --batch-size therefore buys spin, not progress. Measured user CPU for the identical route (user CPU, because this box sits at load 10-40 and wall time measures the machine):

--batch-size user CPU to the Act I load
200,000 91s
50,000 49s
20,000 51s

50,000 is the knee; below it per-batch host overhead takes the saving back.

--tick-ms-per-batch is not the other half of this. Giving the guest more time per batch does cut the number of waiting batches, but at 1000ms/batch the Act I load dies in Diablo II's own "This application has encountered a critical error" box, which is the documented consequence of stepping over Storm's 255ms MPQ completion waits (see Isolated-Worker bounded MPQ waits). Keep the default 200.

2026-09-20: the gameplay test is event-driven and split

test/test-diablo2-demo-gameplay.js no longer fires input at fixed batch numbers. Every stage waits for the pixels that stage produces, so the route survives a --batch-size change; the old schedule was calibrated against --batch-size=1000000 and, at any other budget, clicked screens that were not up yet (at 50,000 it sat on SELECT HERO CLASS for 400 batches and captured that). The screen signatures it waits on, at the CLI's 640x480 canvas:

stage signature
main menu >1000 bright desaturated pixels in both (200,145)-(440,185) (SINGLE PLAYER) and (200,425)-(440,465) (EXIT DIABLO II)
SELECT HERO CLASS >40000 pixels changed from the menu and the centred EXIT plate gone
CHARACTER NAME >400 pixels over a dim plate floor in (495,425)-(615,460); that OK plate peaks at 131, not 255, because the screen is lit by one campfire, and a 90 floor scores it 0
Act I load the centred portal (185,105)-(455,365) lit, with (0,0)-(180,480) black — D2's own critical-error box also leaves a mostly black frame but paints that left margin

Working input coordinates are canvas 640x480, not the guest's 800x600: SINGLE PLAYER (320,164), Barbarian mousemove + dblclick (315,210), name field (320,422), OK (553,442). The old schedule's (400,275) missed the Barbarian entirely and (405,527) was off the bottom of the canvas.

The default run stops at the Act I loading portal: 49s of CPU, about 55s wall on a quiet box, 220s at load 26. Its 420s guest guard and the matching 480s row in tools/test-timeouts.json are sized for this machine at load 50, not for the expected duration. DIABLO2_FULL_ROUTE=1 continues into the Rogue Encampment; that took 2163 batches and 270s of wall clock at load 11-17, and is opt-in precisely because that figure is not bounded on a shared box — the same opt-in run repeated at load 73-90 passed at batch 3299 after 152s of user CPU and 11 minutes of wall clock.

One threshold moved with it. The full route's terrain assertion was > 50000 green pixels, calibrated on the 87,447 of the 2026-09-20 capture. That pinned one camera position: three captures of the same encampment — grass, wagon, campfire, NPC, Barbarian, both orbs, HUD — score 87,447, 43,950 and 34,552, because the count is framing-dependent and the Act I load is one of the nondeterministic paths. The failure the assertion exists to catch is the silent-success DrawIndexedPrimitiveVB stub, whose black floor scores 743-2,186 — the same range as the menu and loading screens. It is now > 20000: 10x over that failure, 1.7x under the tightest frame that is genuinely gameplay. The life orb, mana orb and colour assertions are unchanged (measured 2,974 / 2,902 / 961 against 2,500 / 2,000 / 100).

2026-09-28: resolutions, and the 1024x768 "blow-up" that wasn't

What the binaries support: 640x480 and 800x600, nothing else

diablo ii.exe is FileVersion 1.0.4.0 (the menu's "Shareware v 1.04"); the renderer DLLs carry no version resource. Every layer has exactly two resolutions, selected by an index of 0 or 1:

module site what it does
d2gfx.dll 0x10004455 / 0x10004a07 index 0 -> 640x480, 1 -> 800x600, anything else -> "Unknown resolution %d" (0x1000d5c4, Window.cpp); then AdjustWindowRectEx + SetWindowPos
d2direct3d.dll 0x1000207b same 0/1 choice, logs "Opening Direct3D window at 800x600..." / 640x480, then SetCooperativeLevel(hwnd, 0x411) (FULLSCREEN|EXCLUSIVE|MULTITHREADED) and SetDisplayMode(w, h, 16)
d2ddraw.dll, d2gdi.dll, d2glide.dll {640,800} / {480,600} width/height tables (d2ddraw 0x1000d310, d2gdi 0x10008290, d2direct3d 0x100134e0) same pair

There is no 1024x768 constant in any of them, no Resolution value under VideoConfig (the exe's only VideoConfig value names are Render, plus the Gamma/Perspective preferences), and no resolution command-line switch: the switch table at 0x005a60e0 is 3DFX, OPENGL, D3D, RAVE, PERSPECTIVE, QUALITY, GAMMA, VSYNC, FRAMERATE, the network and character switches, and -w. The pre-LoD game does not let the player choose: menus are 800x600, gameplay and cutscenes are 640x480, fixed. LoD's in-game 800x600 option does not exist here.

Measured per backend (probe3 on box 3, --trace-api on the mode calls, full route to the Rogue Encampment; all three reach it):

backend selected by cutscenes menus gameplay
Direct3D (Render=1 or -d3d) runtime LoadLibrary d2direct3d.dll exclusive SetDisplayMode(640,480,16) exclusive 800,600,16 exclusive 640,480,16
DirectDraw (default, Render 0/2, -opengl) static d2ddraw.dll 640,480,8 then 640,480,16 for Bink exclusive 800,600,8 exclusive 640,480,8
GDI (Render=4, -w) d2gdi.dll window 648x508 (640x480 client) SetWindowPos 808x628 (800x600 client) back to 648x508
Glide (Render=3, -3dfx) d2glide.dll -- crashes on _grGet@12 (no glide3x) --

The emulator honours all of these: SetDisplayMode records the mode and resizes the cooperative window to it, and an exclusive primary is fit-scaled onto whatever canvas the host has. The GDI path is windowed, so its 808x628 menu window needs a canvas at least that big — that, and not anything in the DX paths, is why GDI runs are made with --screen=1024x768.

The 1024x768 Direct3D "blow-up" is a harness coordinate bug, not the emulator

Reported: Render=1 + -d3d + --screen=1024x768 through the benchmark driver (~/d2-backends-out/d2-arm.js) took 374s of user CPU over 3024 batches, made 14,445 presents and "never reached a recognisable menu".

The guest does not know the canvas size exists. Same command line at both sizes (--branch-clock --batch-size=50000 --max-batches=3024, run.js directly, no driver):

canvas API calls Flip EndScene user CPU
640x480 102,014,423 14,445 14,440 346.8s
1024x768 102,014,423 14,445 14,440 348.0s

and the 600-batch version without --branch-clock also matches exactly (12,402,964 calls, identical block counts, 47.0s vs 46.5s). D2 asks for the same modes (above) either way.

The driver's own last capture of the failed run (01-main-menu.png) is the main menu — 800x600 fit-scaled to 1024x768. Its stage signatures are written in 640x480-canvas coordinates (SINGLE PLAYER at (200,145)-(440,185)), and with SCALE=1 they sample the sky above the scaled button, so the menu check never passed and the driver spent its whole 300 x 10-batch attempt budget sitting in the menu. With SCALE=1.6 (1024/640) the identical run reaches the menu at batch 264 in 10.4s, the same as the 640x480 canvas (9.4s).

The two numbers that looked like a blow-up are just 2,760 batches of menu:

So: nothing to fix in the D3D/DDraw emulation for this. A driver that runs with a non-default --screen must scale its sample points by canvas width / 640 for the exclusive (fullscreen) backends, and use the window origin plus 1.25 (800/640) for the windowed GDI menus.

2026-09-28: the menu's frame limiter is a QPC clock spin, and now parks

The loop

d2win.dll loads at 0xc97000 (origBase 0x10000000, so runtime = orig - 0x0f369000). The limiter is the function at orig 0x1000b670, and every menu and loading screen runs it on the main thread (ESP 0x074ff1c0, the main guest stack):

0x1000b690: PeekMessageA(&msg, 0, 0, 0, PM_NOREMOVE)   ; ret 0x1000b6a3
            if TRUE: GetMessageA / TranslateMessage / DispatchMessageA -> 0x1000b74e
0x1000b6e4: QueryPerformanceCounter(&now)                 ; ret 0x1000b6ef
            QueryPerformanceFrequency(&freq)              ; ret 0x1000b6fa
            _allmul(now - last, 25)  (0x10010fe0)
            if (result > freq) { last = now; frame callback; 0x1000bda0 }
0x1000b742: [0x1005bd08] = [0x1005bd04]
            loop while [0x1001cb68] != 0

A 25 fps cap: a frame when (now - last) * 25 > freq, i.e. every 40ms of QPC time. --trace-api=QueryPerformanceCounter,PeekMessageA shows the API numbers stepping by three (Peek, QPC, QPF) with nothing else in between, the peek returning FALSE every time.

Why neither detector saw it

[spin-park] over the 1,001-batch menu dwell below: no line at all on the old build (zero trips on any thread).

Where the time went in the headless run, and why the saving is not bigger: the batch clock steps 1ms per get_ticks call and stops at the next batch boundary minus one (lib/batch-clock.js). So each batch the limiter renders a frame every ~40 loop passes until the clock reaches base+199, and then spins on a frozen millisecond for whatever budget is left. Only that tail is spin -- about 12,000 passes, ~15ms of CPU a batch. The rest of a menu batch is five real frames of rendering plus the host present.

The fix

QPC is now a clock read, keyed on the millisecond its count is built from ($handle_QueryPerformanceCounter, src/09a7-handlers-dispatch.wat): the count itself never repeats because $perf_counter_lo moves on every call, but the clock under it does, and that is what the guest is waiting on. The park is taken before the counter bump and the store, so a woken call is the whole call again. QPC and QPF are also clock reads for the activity sequence ($spin_is_clock_read, src/09b-dispatch.wat). Every existing guard applies unchanged: same return address and ESP per context, no non-clock Win32 call in between (an empty peek is neutral; a peek that finds a message is work), at most $spin_work_max blocks of guest work between reads, K=8 to qualify a site, one park per millisecond, deadline = next millisecond. --no-spin-park is the old arm, which is exact for D2: the old build took zero parks here. test/test-qpc-spin-park.js pins the shape and each of those resets.

The peek detector was deliberately not widened. The loop is waiting on time, not on the queue, and the clock park is the one that carries the deadline the schedulers sleep to (get_spin_deadline_ms / get_tick_count, which QPC now updates).

Measured (box 3, --batch-size=50000 --branch-clock, user CPU)

Menu dwell: Esc through the intros, then sit in the main menu to batch 1,001 (menu up at batch 264 in both):

arm API calls QPC calls clock parks Flips user CPU
old build 28,571,479 9,258,416 0 4,035 89.9s
QPC park 908,475 50,483 766 3,809 79.1s
old build, --batch-size=60000 37,990,087 12,385,277 0 4,256 100.2s
QPC park, --batch-size=60000 935,956 52,232 -- 3,938 85.2s

-12% and -15%. The earlier "220s of 347s was the loop" was the whole guest slice; the spin tail is ~15ms of a ~110ms menu batch.

Full route through ~/d2-backends-out/d2-arm.js (three runs each):

phase old build QPC park
main menu reached batch 264 (x3) batch 264 (x3)
start -> menu 8.79 / 8.75 / 8.77s 8.08 / 8.07 / 8.10s
menu -> Act I portal (355 batches) 6.59 / 6.53 / 6.57s 6.55 / 6.53 / 6.55s
world -> end (630 gameplay batches) 8.77 / 8.00 / 9.03s, 122 / 113 / 124 flips 9.10 / 9.10 / 8.95s, 124 / 121 / 126 flips
main-thread clock parks, world -> end (run 3) 249 (all GetTickCount) 251
main-thread clock parks, start -> menu (run 3) 0 39

The menu, hero-class, name, named and Act I loading captures are byte-identical between the arms. The encampment and final captures differ, but they also differ between two runs of the old build: the Act I load is nondeterministic (world reached at batch 3099-3259 across the six runs). The QPC park does not fire in gameplay (251 against 249 parks), and gameplay CPU per flip is the same (71-75ms in both arms).

The fewer Flips are the one guest-visible change, and they are the old build's artifact. In the old build the frozen-millisecond tail was not entirely dead: QPC adds 1us per call, so ~11,000 spin passes could move the count the 11ms still missing from a 40ms frame and render one more before the batch ended. That extra frame is bought with surplus budget: the old build renders 4,035 menu flips at 50,000 blocks and 4,256 at 60,000. The QPC park ends the tail instead, leaving the menu at 25 fps of guest time as the limiter intends. In the menu-dwell captures, six of the eight sampled frames (batches 300-1000) are byte-identical between the arms. The other two differ only in the phase of the burning DIABLO II logo, since the extra frames shift the flame animation.

Browser (not measured here, BROWSER-LOCK was held): there the clock is real and the limiter spins the whole 40ms between frames, so this is where the saving is large. A qualified site parks after two reads of one millisecond and sleeps to the next (CLOCK_SPIN_PARK_MS, _spinParkDelay).

One tooling note: --slice-split, --decode-stats and --batch-stats time and count only the first main-thread run() of each batch. D2 has guest threads, so test/run.js also runs main between worker slices within the batch (the cooperative branch after the spin-park handler), and those runs are invisible to all three. With the QPC park, the park is taken in one of those interleaved runs. The next batch's first run() then returns immediately on the pending yield ("blocking wait", 1 block), and the menu's real work happens in the interleaved runs after the handler clears it. The old build spent the whole budget in its first run. So on this route those flags read 77ms/batch against 0.00ms for identical menu frames, which says nothing about cost. Use user CPU at fixed batches for spin A/Bs.

2026-10-04: Glide menu freeze and repeating audio investigation

The user reports frozen menu animation and repeating sound in diablo2_glide_demo. The corrected local observation did not reproduce the menu freeze. Audio continuity and quality remain unverified; this is not a repair or a resolution of the user report.

The registered route inherits the installed demo and sets VideoConfig Render=3 in both registry hives. It dynamically loads D2Glide. All 25 registered files were present (135,378,002 bytes). The observed module was f40d4ca3382279ff9b826188573f8acd9272eaa2dc5024dcbb69aecc35b49063 (1,664,777 bytes); current JavaScript was pinned separately, without claiming that later WAT source changes were included in this module.

Attempt 1 ended with normal ExitProcess(0) after an Escape key sent without a fresh visual gate. It is preserved as a harness route failure, not a game crash or a freeze reproduction. Attempt 2 prohibited Escape, captured state before input, and reached the Shareware 1.04 menu through an ordinary click on the visible title screen.

During the 10.3793-second quiet menu interval, Glide host presentation events increased from 1,539 to 1,812. The logo flames visibly changed between the first and last images; the coordinator independently confirmed this in scratch/diablo2-glide-menu-20261004/root-visual-review.json. These transport counts are diagnostic evidence of progress, not qualified game or display FPS.

The AudioContext was running and advanced from 37.883 to 48.263 seconds. Three later read-only observations of DirectSound worklet voice 720898 showed an advancing cursor, descriptor sequence 758/770/780, and three different 262,144-byte ring hashes over 4.01 seconds. These concurrent memory snapshots show producer activity, not an atomic audio recording or proof of seamless output. Shared waveOut callback registration was zero and its completion queues were empty; this observation does not support the separate Tile World or FOTAQ pending-function-callback diagnosis.

Evidence is in scratch/runs/20261004-diablo2-glide-menu-observation (182 hashed artifacts), with raw analysis.json, rings.json, timed images, owner samples and the actual served files. All 147 unique served file copies passed their recorded hashes. The browser was headful Chrome 151 with explicit --disable-gpu and SwiftShader, so this is not proof of parity with the user's device. Browser session 35990 and its read-only attachment both exited 0; browser and server closed at 16:29:06.090 UTC.

Next, compare the affected URL, build, browser/device and onset route with this local identity, then capture output audio if the complaint persists. No guest state, engine code or pacing was changed, and no gameplay, FPS or release-readiness claim follows from this menu observation.