ExecuteData across sparse pages
2026-09-22. Public buffer-access follow-up to ExecuteBuffer lifetime work.
SetExecuteData read its scalar fields with guest accessors but copied the 24-byte dsStatus tail through a single g2w pointer. GetExecuteData similarly cleared 48 bytes and copied its status tail through single translated pointers. Adjacent guest pages need not be adjacent in WASM memory, so these operations could leave the output partially untouched and modify unrelated backing storage while still returning S_OK.
The new regression first failed on an uncached GetExecuteData output starting two bytes before a guest page boundary: scalar stores succeeded but the remaining output retained 0xcc instead of being cleared. The handlers now use guest_memset/guest_memmove for the full clear and both status copies. The existing cache-header helper still returns a WASM address; convert that base once with w2g before adding the guest-relative status offset. No whole caller buffer is treated as affine merely because its first byte translates.
Coverage
node test/test-d3dim-execute-data-sparse.js exercises sixteen layouts:
cached/uncached objects, boundaries crossing dwSize or dsStatus, and independent
contiguous/sparse input and output choices. It checks all 48 output bytes,
input preservation, output canaries, both interleaved backing-page guards,
and stdcall cleanup/stack guards. Objects use public creation and Release;
public Unlock prepares the cached cases. Tests retain the existing missing
cache behavior (zero status) and fixed 48-byte output policy, not a newly
asserted native size/flag contract.
All sixteen cases pass after the fix, as do the eight cached/uncached lifetime cases and scoped fragment, ESP/epilogue, logical-operand, silent-stub, duplicate and test-tier checks. Quiet remains 243 manual + 22 metadata; duplicates remain 117 groups / 471 members.
Execution audit still open
This fixes the caller's ExecuteData buffers, not the whole Execute engine. The audit found these additional dependencies:
- Cache allocation/copy duplication and the mismatched-cache replacement leak are addressed by the shared-owner follow-up below. Source-base return values and downstream consumers still use raw WASM pointers.
- PROCESSVERTICES source/destination traversal is addressed in the bounded vertex follow-up below; opcode records still arrive as WASM pointers.
- Execute and Pick walk instruction records with WASM-pointer arithmetic, then hand those addresses to opcode handlers. Repairing the backing copy alone cannot repair those consumers.
- Cache status helpers now use guest headers (follow-up below). The remaining migration must keep persistent addresses guest-relative and acquire bounded spans at consumers requiring contiguous records, with explicit release and writeback. Span allocation/failure limits must be verified before using a whole-buffer span as a fallback.
No native Win98, real x86 indirect-call, browser/game rendering or full-build validation is claimed by this focused regression.
Shared cache owner follow-up
The new cache regression reproduced the lazy source path leaking one allocation on size mismatch (live heap 74 instead of 73). Both lazy source lookup and Unlock refresh now use d3dim_execbuf_cache_ensure. It owns guest-address header reads, allocation, initialization, page-aware snapshot copying and replacement. A matching lazy lookup preserves its snapshot; matching Unlock refreshes only the payload and preserves status. A successful replacement clears status, publishes the complete new cache, and frees the old allocation. Failed allocation leaves the old cache owned and unchanged; a subsequent retry or final Release can still retire it. No failure switch is added to production.
test/test-d3dim-execute-cache.js first verified 32 direct/sparse source
snapshots and replacements across both paths, every snapshot byte, unchanged
interleaved backing, lazy reuse versus Unlock refresh, status preservation
and final heap balance. The fixture varies stored size explicitly to exercise
replacement; it does not claim a public resize API exists. It borrows its
payload and detaches it before public Release, leaving cache ownership real.
The extended regression injects failure only at the shared cache allocation
site and checks eight initial/replacement failures plus retries: no publication
on initial failure, old pointer/bytes/count preserved on replacement failure,
and balanced final cleanup. The unchanged lazy allocation-failure fallback
returns the original buffer's translated base; that fallback is not yet safe
for downstream cross-page consumers.
All 32 snapshot cases and eight fault/retry cases pass, along with the sixteen ExecuteData layouts, eight lifetime cases and scoped static gates. Quiet and duplicate counts remain unchanged.
This consolidates the writer/owner, not the entire execution engine. The source-base API is migrated in the bounded vertex follow-up. Cache-header/status readers are migrated in the follow-up below. The owner test forces sparse source storage, not sparse allocator-returned cache storage or complete vertex/instruction execution.
Guest-address header/status contract
The expanded ExecuteData regression relocated a real cache into borrowed nonaffine test pages, retaining the original owned allocation for cleanup. It reproduced lost status when the first identity DWORD crossed a page: cache_header rejected a valid cache after reading unrelated backing bytes.
The helper is now explicitly named d3dim_execbuf_cache_header_guest, reads both identity fields with guest accessors and returns the persistent guest address. All three consumers use that contract: public Set/GetExecuteData no longer reverse-translate a WASM header, and the opcode status writer uses guest stores for its six status DWORDs and its extent updates. The opcode record itself retains its existing WASM-address input contract; this does not repair the instruction decoder's separate sparse-read assumptions.
The regression expands to 32 layouts: absent, original, header-crossing and status-crossing caches, crossed caller dwSize/status fields, and independent input/output placement. Each cached case also invokes the opcode status writer with a contiguous status record and verifies readback. All three interleaved backing pages and caller canaries are checked. Borrowed caches are detached and their original heap-owned pointers restored before public Release. The opcode record requests status only; native extent semantics and the device-state-dependent extent calculation are not certified by this fixture.
All 32 ExecuteData layouts, 32 cache snapshots/replacements with eight forced allocation failures/retries, eight lifetime cycles and scoped static gates pass. Quiet remains 243 manual + 22 metadata and duplicates 117 / 471.
Bounded vertex follow-up
The snapshot accessor is now d3dim_execbuf_source_guest: both cached and allocation-failure/no-owner paths return guest addresses. PROCESSVERTICES keeps source/destination indexing in that address space. COPY uses guest_memmove. The transform modes gather at most one 32-byte source vertex and one 32-byte destination vertex for the existing math helpers, then write back the destination and release the source in reverse acquisition order. No whole-buffer contiguous copy or extra persistent render surface is added.
test/test-d3dim-execute-vertices-sparse.js reproduced untouched destination
bytes in a crossing-source transform. Eighteen sparse/control comparisons now
pass: three modes, absent/heap-owned/borrowed-sparse source caches, and source
or destination page crossings. Every buffer byte matches the contiguous
control; neighboring backing pages, buffer canaries and span cursor/overflow
counters remain unchanged. Two vertices per record check advancing across
the boundary rather than only transforming the first vertex. Fixtures attach
borrowed storage directly and detach it for cleanup, so this proves the vertex
helper's addressing, not native creation policy or full Execute dispatch.
The cache regression's eight failure/retry cases now assert guest-address fallback results. Its 32 snapshots/replacements, the adjacent vertex-buffer ProcessVertices regression and scoped static gates also pass. Quiet and duplicate inventories are unchanged. Native transform/lighting semantics are not newly certified by comparing two layouts of the same implementation.
Remaining: Execute/Pick instruction walking and record inputs, other primitive vertex readers, range/overflow validation, and full application rendering. The generic span arena's unsafe exhaustion fallback is removed in the shared exhaustion follow-up: exhaustion now stops explicitly before copying. This vertex path uses at most 64 additional bytes and the test observes no overflow; graceful recovery from arena exhaustion remains separate work.
Execute instruction/record follow-up (2026-09-22)
Execute now retains a guest-address cursor for instruction headers, operand records, branch targets and trace offsets. All nine record helpers share that contract: points, lines, triangles, matrix load/multiply, state walking, PROCESSVERTICES, branch and status. Scalar reads use gl8/gl16/gl32; SETSTATUS copies its 24 bytes with guest_memmove. No whole-stream gather or record-group scratch allocation is needed. The two existing helper-level tests now pass guest record addresses instead of translating them first.
node test/test-d3dim-execute-instructions-sparse.js first passed its
contiguous control, then failed on the first sparse layout: render state
remained zero instead of becoming one. With the migration it passes:
- One control and 195 page-boundary placements across every byte of a multi-opcode stream, including headers, multirecord operands and a taken branch that skips a state write. Untaken branches are also exercised.
- Matrix load/multiply, transform/light/render state, COPY vertex output, public GetExecuteData status readback, guest-relative opcode traces, stdcall stack checks, unchanged instruction bytes and surrounding canaries.
- A 2,200-record render-state group spanning nonaffine pages, larger than the 16KiB span arena, without consuming scratch or changing overflow counters.
- 65 sparse/control pixel comparisons for point, line and triangle record streams on a real 32x32 surface. Two records per opcode exercise record advancement. Triangle coverage uses point fill; vertex bytes themselves deliberately remain within one page. Interleaved backing pages stay intact.
The fixture uses public Execute/CreateExecuteBuffer/SetExecuteData/Unlock/ GetExecuteData/Release dispatch, with an initialized synthetic device and borrowed sparse buffer mappings. It restores the owned buffer before Release. It does not exercise x86 indirect calls or certify native driver behavior.
Adjacent ExecuteData (32 layouts), PROCESSVERTICES (18 comparisons), cache (32 snapshots/replacements plus eight allocation faults/retries), interface metadata and scoped static gates pass. Quiet remains 243 manual + 22 metadata; duplicate census remains 117 groups / 471 members. No browser/full-build or performance result is claimed. The unrelated in-progress GL fog hunk in the same core file is excluded from this change.
Still open: primitive vertex-base translation; invalid record sizes, bounds, indices, arithmetic overflow and branch loops; actual branch status/mask semantics (the current helper still assumes status zero); the end-of-range header comparison (tests pad EXIT to preserve that separate policy); full application rendering. Pick walking and output crossings were handled in the separate Pick review.
Primitive vertex follow-up (2026-09-22)
Points and lines now index vertices using guest addresses and read positions and color through guest scalar accessors. Triangles acquire three bounded 32-byte vertex spans for the existing raster/math helpers. A single cleanup block releases them in reverse order after all normal draw/skip paths, including point fill, wireframe, fully hidden triangles and partially clipped triangles. It also removes four duplicated record-advance/loop-continue sequences. Directly contiguous vertices use the span helper's existing no-copy path; at most 96 bytes are gathered by this caller, not the entire execute buffer.
The Execute regression first failed with unrelated positions/colors for a point whose first vertex crossed a page. It now compares 2,470 sparse/control vertex layouts: 95 byte-boundary placements for points and lines, plus all three triangle fill modes and eight positive/negative-rhw combinations. Every draw verifies unchanged input vertices and balanced span cursor/overflow counters; neighboring backing pages retain their sentinels. The all-visible controls must actually draw pixels, rather than merely agree on a blank image. The existing 195 instruction-boundary cases, 65 record/pixel comparisons and 2,200-record group remain covered. This is layout-equivalence evidence for the existing rasterizer, not new native clipping/lighting/texture conformance.
The 18 PROCESSVERTICES sparse comparisons and existing v3 vertex-buffer draw regression (961 indexed pixels) also pass, alongside interface/dispatch and scoped static gates. Quiet and duplicate inventories remain unchanged. No full-build, browser gameplay or performance claim is made by this change.
Remaining: malformed record sizes and vertex indices, overflow/range checking, branch status semantics and full application/browser verification. A trap on scratch exhaustion is still an emulator safety stop, not graceful API recovery.
Branch comparison follow-up (2026-09-22)
Microsoft's SDK d3dtypes.h describes D3DBRANCH as masking driver status, comparing the result to dwValue, and optionally negating the comparison. Its D3DSTATUS definition identifies dwStatus separately from the flags and extent. This is the primary-source basis for the comparison change; Wine implementation behavior is not its basis.
The old branch helper ignored dwMask and assumed status zero. Execute now
passes the execute-buffer identity, and the helper reads the same guest-safe
retained status that SETSTATUS writes and GetExecuteData returns. It compares
(status & mask) == value, with any nonzero bNegate reversing the result.
Target arithmetic, zero-offset termination and first-record-only handling
are unchanged, not newly certified by this change.
The extended Execute regression failed before the fix for status/mask 0x80000000, value zero, negate zero: it incorrectly skipped a state write. Afterward all 1,152 combinations pass: four statuses, four masks, four values, three negate values (0/1/2), status seeded through either SetExecuteData or a SETSTATUS instruction, and three direct/sparse stream placements. Each case checks the observable render-state write, retained status readback and public stdcall ABI. Existing 195 instruction boundaries, 65 record/pixel comparisons, 2,470 vertex layouts and the large record group still pass, as do the 32 ExecuteData layouts and scoped static gates. Quiet243+22 and dup117/471 remain unchanged. No native/browser/full-build/performance claim is made.
Important remaining status work: storage is still tied to the source cache; the regression intentionally Unlocks before seeding status. SetExecuteData and SETSTATUS without that owner can still lose status, and cache-allocation failure is not repaired here. SETSTATUS flag-selective updates, multi-record branch dispatch, clip-generated status and device-vs-buffer lifetime across Execute calls still require review/native evidence. The SDK's device-status lifetime comment is not sufficient evidence to certify our per-buffer storage.
Multi-record branch/status traversal (2026-09-22)
The same Microsoft SDK header defines D3DINSTRUCTION.bSize as the size of each data unit and wCount as the number that follow. Execute previously ignored that count for BRANCHFORWARD and used only the first SETSTATUS record. Both now traverse the declared count/stride. The first taken branch transfers control immediately; no match falls through. Zero records consume no operands. SETSTATUS applies each record in order using the existing status writer.
The regression initially trapped on a zero-record branch: the old helper interpreted the following render-state instruction as a branch record and jumped into invalid instruction bytes. After the fix, 68 cases pass across zero/one/three records and direct/sparse placement. Branch cases cover first, middle, last and no matches, multiple simultaneously matching records with distinct destinations, and taken zero-offset termination. SETSTATUS cases verify that the final record is retained, or the existing status survives a zero-record instruction. All calls check the public Execute ABI.
The complete Execute sparse/render suite, 32 ExecuteData layouts and scoped static gates pass; quiet243+22 and dup117/471 are unchanged. The test uses standard record sizes, not malformed sizes or an assertion that arbitrary padded records are native-compatible. Branch offset origin/zero-offset policy are preserved and regression-tested, not newly native-verified. Status flags, cache-independent ownership, clip status, device lifetime, buffer/index bounds and browser/full-application verification remain open.
SetExecuteData status ownership transaction (2026-09-22)
SetExecuteData now prepares its existing shared cache/status owner before publishing vertex/instruction fields. Previously it silently discarded the 24-byte status when no Unlock/source snapshot had created that owner. It now reuses d3dim_execbuf_cache_ensure rather than adding another allocation path or status table. Matching storage is reused without refreshing the source snapshot; a subsequent Unlock retains its existing refresh behavior.
If preparation fails, SetExecuteData returns E_OUTOFMEMORY with the prior descriptor, cache pointer, cache bytes and heap ownership unchanged. A retry can publish the complete descriptor/status. This error choice is a resource failure policy, not a native Win98 memory-pressure/error-precedence measurement. The existing 1MiB cache-capacity and null-input policies are unchanged.
Two RED results were reproduced: the formerly uncached sparse-data case lost all status bytes, and the fault-injected setter returned success. The 32 ExecuteData layouts now require full status retention even before any Unlock. The shared cache regression adds eight transactional failure/retry/reuse cases: direct/sparse buffer storage, absent/mismatched owner, and direct/page-crossing input. It checks every object-record byte on failure, old cache contents and pointer, input preservation, heap balance, full public GetExecuteData readback, and stdcall cleanup. Matching-owner updates succeed even with allocation faults armed, proving reuse rather than another allocation.
All those cases, the prior 32 cache snapshots/replacements and eight owner allocation faults, and the complete public Execute sparse/branch/render suite pass. Scoped static gates also pass; quiet243+22 and dup117/471 are unchanged. No full-build/browser/performance/native-driver claim is made. Preparing an uncached descriptor now allocates the existing header-plus-source snapshot; its memory cost is intentional and not benchmarked here. Status flags/extents, native device-vs-buffer lifetime, invalid ranges/indices and browser application verification remain open. Direct internal SETSTATUS calls with no prepared owner still do nothing; normal public instruction setup now prepares the owner or reports failure instead of accepting a descriptor without status storage.
Real-app follow-up: PROCESSVERTICES clip status (2026-09-22)
The full shared-main build passed, as did Viewer Open/menu browser coverage in cooperative and threaded modes. That browser test does not assert scene pixels. Globe's CLI Render-menu regression failed: both point and wireframe captures had zero lit pixels. Synthetic branch correctness was insufficient coverage.
DX tracing and a guest instruction dump identified this sequence:
SETSTATUS status = 0x01fff000 (D3DSTATUS_DEFAULT)
PROCESSVERTICES transform/light 56 vertices
BRANCHFORWARD mask = 0x0003f000, value = 0, negate = 1, offset = 0
TRIANGLE ... never reached
The masked comparison was correct; PROCESSVERTICES never changed the seeded intersection bits. A temporary rsync-mirrored control changed only the branch comparison back to its former assumed-zero shortcut. Globe then passed with 222 / 2982 / 31713 lit pixels in point / wire / solid modes. That shortcut is not the fix and was not restored on main.
PROCESSVERTICES now accumulates the six standard clip-plane union/intersection bits from each transformed homogeneous vector. It reads the existing projection scratch before perspective division and before the near-zero-w clamp, and updates the ExecuteBuffer's retained guest status with sparse-safe accessors. The caller explicitly passes the buffer owner; no extra status table or guessed owner is introduced. COPY and zero-vertex operations leave status unchanged.
The basis is Microsoft's SDK d3dtypes.h for clip/status bit definitions and transformation pipeline for homogeneous clipping. Clip edges follow this renderer's inclusive plane convention; native Win98 exact-edge/rounding behavior has not been measured. User-plane bits and ZNOTVISIBLE are preserved, not falsely calculated by a position-only transform. Their semantics, SETSTATUS flag selection, actual extent accumulation, device-vs-buffer native lifetime and malformed stream validation remain open.
The new public regression was RED for a single inside vertex (DEFAULT retained instead of clearing standard intersection bits). It covers all six planes, inside/outside mixtures, grouped and separate records, zero counts, all three vertex modes, direct/sparse layouts, non-unit and negative w, and a tiny w that would be misclassified after divisor clamping. Assertions include status readback, downstream branching, ABI, span balance and neighboring-page guards.
Validation on the fixed shared-main worktree:
- Full build/gates PASS (normal WASM 1,504,139 bytes; unchanged region layout
hash
4f4410e063257228). Other agents' unrelated pending changes were present; this is not a pristine-commit build or performance measurement. - 324 new clip-status cases PASS, together with 195 instruction splits, 65 record / 2470 vertex pixel comparisons, 1152 masked branches, 68 multi-record cases and the 2200-record group.
- PROCESSVERTICES 18 sparse/control cases and ExecuteData 32 layouts PASS; interface-spec and scoped gates PASS. Quiet243+22 and dup117/471 unchanged.
- Globe Render-menu PASS: all 11 items survive; point / wire / solid produce 222 / 2982 / 31713 lit pixels, identical counts to the temporary control.
- Viewer rendered-mesh selection / Change Color CLI regression PASS.
- Viewer Open / Renderer menu browser regression PASS again after the fix, in cooperative and threaded modes (input/menu coverage, not a pixel oracle).
SETSTATUS field selection (2026-09-22)
Microsoft's SDK D3DSTATUS.dwFlags selects status and extents independently
(D3DSETSTATUS_STATUS=1, D3DSETSTATUS_EXTENTS=2); see the SDK header linked
above. The unconditional 24-byte copy violated both directions: an extents-only
record changed subsequent branch decisions, and a status-only record replaced
the prior rectangle. A zero-selection record also changed both payloads.
The writer now copies only selected payload fields with guest-safe accesses. It retains the existing flags-word readback policy. Public Execute tests cover all four selections with direct/page-crossing records and verify downstream masked branches plus preservation of unselected extents. The existing data fixture covers 384 field-selection combinations across 32 input/output/cache layouts and three record placements, including non-affine cache headers and records. Its state-less device fixture verifies the supplied selected payload; it is not evidence that device-backed extent reporting is correct. The first RED assertion was selection zero overwriting both status and rectangle.
The old helper fixture passed a null device handle. Extending it to EXTENTS exposed that the unchecked internal device accessor treats that as slot zero, which was an ExecuteBuffer, reading its vertex offset as a device-state pointer. The fixture now supplies an explicit synthetic device with no state rather than depending on accidental null-handle behavior. Public cases use a real initialized state and cover the selected-status branch semantics independently.
Still open: when EXTENTS is selected on a real device, the pre-existing viewport rectangle substitution remains. This change does not remove that approximation or claim actual extent accumulation. That requires following PROCESSVERTICES_UPDATEEXTENTS and primitive/raster effects, including empty, clipped and culled draws, rather than merely deleting the substitution and leaving the guest's inverted sentinel rectangle untouched. Native flags-word readback/error policy and device-vs-buffer status lifetime also remain unproven.
Verification: all 384 data-field selections and 16 public branch cases PASS,
as do the existing 324 clip-status cases and complete sparse Execute suite.
Full shared-main build/gates PASS (1,504,180-byte normal WASM, unchanged layout
hash 4f4410e063257228). Quiet243+22 and dup117/471 remain unchanged. This was
a correctness run on a heavily loaded shared machine, not a performance result.
Globe's real Render-menu test also PASSes after rebuilding: all 11 items survive,
with unchanged point / wire / solid counts of 222 / 2982 / 31713 lit pixels.
Original DX5 documentation recovered: extent design evidence (2026-09-22)
The original Microsoft SDK is available as
idx5sdk.exe, via the
archive item. The downloaded 33,018,416-byte
self-extractor has SHA-1 b14370372307360a9e8de2ebd8fcd13173fd3b4a, matching
the archive metadata. Its nested DX5SDK.EXE contains the original English
Word references, dated July 14, 1997:
| Archive path | SHA-256 |
|---|---|
/cdrom/docs/worddoc/d3dimref.doc |
f76cbc4e1b3f0311df4205371703b31e911088fcf1d6744e5493aeff166fa6ad |
/cdrom/docs/worddoc/d3dimovr.doc |
947b8ef0f779d07ee37d28c939295a2ee5ef542b64d2259c943516c524791538 |
Extract with 7-Zip, then use textutil -convert txt on macOS to search without
running an installer. Working copies for this investigation are under
/private/tmp/wa-dx5-docs.59DpEl/; these proprietary source documents and SDK
binaries are not committed. This is original Microsoft documentation and sample
code, not a compatibility implementation or a modern API analogy.
The D3DPROCESSVERTICES, D3DSTATUS, and D3DEXECUTEDATA sections establish:
- UPDATEEXTENTS includes transformed vertices in the returned rectangle.
- Status accumulates over executions; the rectangle expands, and SETSTATUS supplies its reset.
- ExecuteData exposes the screen extent of rendered geometry.
Microsoft's sdk/samples/uvis/uvis.cpp provides the other half of the evidence:
CreateFireObjects seeds an inverted rectangle, uses TRANSFORMLIGHT without
UPDATEEXTENTS, and issues triangles. RenderFire later passes GetExecuteData's
rectangle to the retained-mode viewport's ForceUpdate. The accompanying
misc/d3dmacs.h writes the supplied process flags verbatim. Thus this real sample
depends on rendering updating extents even without the transform flag.
This changes the implementation/test plan, not production behavior yet:
- Track both explicit transformed-vertex updates and rendered primitive bounds; a pixel-write-only tracker would omit the former, while a flag-only tracker would omit the sample's draw path.
- SETSTATUS must install the supplied rectangle, including an inverted seed, rather than substitute the viewport. Preserve accumulation across executions and across records until an explicit reset.
- Cover no-draw transforms with the flag on/off, clipped and culled rendering, sparse status storage, multiple buffers and repeated executions, as well as Globe/Uvis presentation. Do not infer exact native rounding or whether a rejected primitive contributes solely from the documented bounding-box term.
The earlier suggestion that every fully clipped operation must preserve the rectangle was too broad: a flagged transform can request extents independently of subsequent drawing. Native edge rounding, COPY+UPDATEEXTENTS and rejection details still need evidence. The available v86 reference profile is 4bpp and cannot currently supply a DirectDraw rendering oracle (see native probe limitation); that is not a reason to label these policies verified. No production fix or Win98-conformance claim is made by this documentation recovery.
Execute extent accumulation (2026-09-22)
The viewport substitution is now removed. SETSTATUS copies its selected rectangle verbatim, including inverted seeds, without consulting device state. The helper no longer needs a device argument; the synthetic state-less device workaround described above is consequently removed from its test fixture.
Two producers expand the existing per-buffer status rectangle:
- PROCESSVERTICES with UPDATEEXTENTS adds processed screen coordinates, even without drawing. The implementation includes COPY and uses the software renderer's rounded coordinates with exclusive upper bounds.
- Execute's raster path adds nonempty, clipped rectangles/spans from flat, alpha, depth and textured drawing. Collection is restricted to that Execute's render target. Two per-instance globals identify the active scope; the actual rectangle remains in the existing guest-owned cache header. Execute saves and restores the scope, restoring it before presentation. No new allocation or second persistent extent table is introduced.
The rectangle persists across executions and Unlock refreshes. SETSTATUS can reset it after prior work. Public SetExecuteData still explicitly installs its supplied status as before; this change does not establish native lifetime rules across devices or concurrent use of one ExecuteBuffer.
Regression evidence:
- RED before the fix: an opaque point draw retained
[2048,2048,0,0]instead of its rendered bounds[2,2,22,4]. - Existing 65 instruction-record and 2470 vertex-layout pixel comparisons now also compare returned extents against the nonzero opaque pixel bounding box. These cover points, lines and three triangle fill modes, including reciprocal-W visibility masks, with sparse instruction and vertex storage.
- 72 no-draw transform cases cover three vertex modes, flag on/off, empty and two-vertex ranges, direct/sparse storage and accumulation over two executions. Another 36 cases check explicit reset, restoration of a synthetic enclosing collector, and isolation from drawing after Execute returns. This is not a real nested host-callback test.
- 384 data field-selection cases, 16 public SETSTATUS branch cases, 324 clip cases and the complete prior Execute instruction/branch suite pass.
- Full build passes: normal WASM 1,504,413 bytes, compatibility 1,506,819 bytes,
unchanged layout hash
4f4410e063257228. Silent inventory remains 243+22. - Viewer CLI mesh-selection/Change Color and browser Open/Renderer menu tests pass; browser covers both cooperative and threaded scheduling.
- Globe Render-menu passes twice with all 11 items and point/wire/solid counts 224 / 3642 / 13535. These differ from the prior viewport-substitution counts; the test checks menu behavior and fill-mode separation, not image equality. A 150-batch capture still shows the sphere and the known texture-band defect.
An isolated control compiled the same current source closure with only 09aa
and 09ab restored to pre-change HEAD. Running the same main-tree test with
WINE_ASSEMBLY_WASM=/private/tmp/wa-extent-control.wasm reproduces the old
222 / 2982 / 31713 counts. Both arms render a sphere in 150-batch captures,
but at different apparent sizes/orientations. The original SDK globe.c
advances its camera animation by 0.08 per move callback, independently of the
callback delta. Different extent-driven guest work can therefore change the
scene reached at a fixed block budget; this is a plausible explanation, not
proof that the entire image difference is harmless. A frame-aligned comparison
remains open. Control artifacts and captures are under /private/tmp/wa-extent-*
and /private/tmp/wa-globe-extent*; main's artifact was not replaced.
The full-build result above predates the concurrent 3c7b023e MMIO region
addition. It certifies this change against the then-current shared tree, not
that later region-map/artifact pairs remain synchronized.
Limits: these are coverage bounds before per-fragment depth, alpha-test and color-key rejection, so they can over-report pixels actually written. Native rounding, COPY+UPDATEEXTENTS and rejection details are not verified by the SDK text or the available 4bpp v86 profile. Dedicated multi-buffer, XY-offscreen, culling and textured/depth/alpha extent-oracle tests remain useful follow-ups. Uvis source is available but its executable is absent from the local corpus; its ForceUpdate interaction is not runtime-verified. No performance claim is made on this loaded shared machine.
Public extent-owner isolation follow-up (2026-09-22)
The instruction regression now creates two separate ExecuteBuffers through the public API and uses their real Lock-returned allocations, rather than substituting borrowed fixture mappings. It interleaves point draws and Unlock refreshes, checks both retained rectangles after each relevant operation, resets one through SETSTATUS, releases it, and continues accumulating into the survivor. A fully XY-offscreen point leaves the prior rectangle unchanged; a point crossing the 32x32 target's bottom/right edge reports the clipped box. These cases pass without a runtime change, closing the basic sequential multi-buffer and point XY-clipping gaps listed above.
The entire instruction regression also passes against freshly compiled current sources after the concurrent MMIO region addition, along with test-tier and whitespace checks. This is not a new full shipping-artifact build. Concurrent ExecuteBuffer access, multi-device native status lifetime, polygon XY clipping, culling and textured/depth/alpha extent-oracle cases remain unverified; the new isolation test does not certify them.
Frame-aligned Globe A/B (2026-09-22)
The first aligned captures are pixel-identical. The fixed-budget discrepancy above is not reproduced when the guest reaches the same camera callback count and animation time. This narrows the uncertainty; it does not certify every Render-menu transition or prove exact Win98 extent semantics.
Method: original Microsoft SDK globe.c advances moveCamera's static time by
0.08 per callback. In the corpus globe.exe, the float constant at 0x40a040
has one xref, fadd at 0x401711; the callback starts at 0x4016ec and its
time lives at 0x40b104. Both arms run the same current CLI/JS with separate
temporary artifacts built from one source snapshot (before 9bebacdb). The
control replaces only 09aa/09ab with their c5f9758c^ versions. Main's shipping
artifact is not overwritten. SHA-256:
- Control:
7082d64d4dae9491fd3407f12596353ba6758ac5a7729cc1df004f53297ec243 - Candidate:
7b57afb0ea15733dfa994b636e986f4b5d4f07676ef399199d4124deb4e4cfbf
Launch each with --app=dx_globe --no-build --wasm=ARTIFACT --quiet-api --control-stdin --frozen --count=0x4016ec --batch-size=100000 --max-seconds=180 --png=OUTPUT. Through control eval, set exports.set_bp(0x4016ec), then step
one batch at a time until get_eip() equals that address and get_count(0)
equals the selected count. Read the time float through the guest translation,
assert both arms match, and quit to capture the primary surface. The breakpoint
precedes the next callback, so the count describes completed prior callbacks.
Do not use CLI --break with this stdin driver: its interactive debug
prompt consumes the control stream on the next step. The initial diagnostic
hit that prompt and was terminated; it supplied no rendering comparison.
| Completed callbacks | Animation time, both arms | Batch, both arms | Differing RGBA pixels |
|---|---|---|---|
| 10 | 0.7999998927116394 | 92 | 0 / 307200 |
| 50 | 3.9999983310699463 | 132 | 0 / 307200 |
| 100 | 7.99999475479126 | 182 | 0 / 307200 |
Comparison uses tools/png-diff.js with zero tolerance. Temporary driver:
/private/tmp/wa-globe-frame-compare.js; build driver:
/private/tmp/wa-extent-control.js; capture/log prefix:
/private/tmp/wa-globe-frame-. Both images at callback 10 visibly contain the
sphere, including the same known texture defect. These are scene checkpoints,
not blank-image equality or performance measurements (host load was above 20).
Indexed-texture far-plane regression triage (2026-09-22)
The reported test-d3dim-indexed-texture.js failure reproduced at its
behind-eye far-plane assertion. It was a test-state error, not evidence that
the new Execute extent collector changed clipping. That fixture uses the
direct primitive helper, outside Execute collection.
Its projected inputs are (1,1,.5,1), (6,1,.5,1) and (3,5,-.5,-1)
(screen x/y, z/w, reciprocal w). Reconstructing homogeneous coordinates puts
the last vertex beyond the far plane. Each crossing has interpolation fraction
1/4 and w=1/2. Projecting the generated vertices gives (0,-1) and (7.5,-1),
with z/w=1 and reciprocal w=2. The resulting polygon is counterclockwise in
screen space, so default CCW culling correctly discards it. The original DX5
reference's D3DCULL section explicitly distinguishes NONE, CW and CCW; the
overview describes rasterization's screen-winding test. No Wine source used.
The regression now checks both generated intersections numerically, requires the strip to render with NONE and CW culling, and requires an untouched target with CCW culling. Subsequent sampler assertions reuse that same CCW geometry, so they explicitly select NONE instead of implicitly relying on the former unset-means-no-cull behavior. This preserves the renderer's correct default and adds coverage instead of weakening culling to satisfy a stale test.
The full indexed-texture suite passes after this correction, covering later UV-set, filter/address, color-key, blend and reversed-depth assertions that the early failure previously prevented from running. No production source change is needed. The separately reported blank Boids frame remains untriaged.