mirror of
https://github.com/IfcOpenShell/IfcOpenShell.git
synced 2026-09-22 15:38:03 +00:00
bonsai-0.9.0-alpha2608301841
18 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4a761b51f5 |
ifcviewer: pack the cull-hot instance fields and skip unchanged culls
On the single-threaded web build the CPU cull WAS the frame: 52-60 ms of a 60 ms frame at 640k instances (desktop hides the same cost across cores via std::async, which web cannot use without COOP/COEP+pthreads). Two changes, both also helping desktop: - ModelGpuData::CullInstance packs the six AABB floats and three ids the cull reads into 40 contiguous bytes. InstanceInfo is 232 bytes with the AABB 200 bytes away from the ids, so the walk paid two or three cache lines per instance. Rebuilt by rebuildCullInstances at model apply and inside uploadInstanceRecords, which every recompose, transform and colour-override change already funnels through. Measured on web: 94 ns/instance -> 36 ns/instance during a continuous orbit (~2.6x). - render() re-culls only when a cull input changed: the camera, a cull-relevant setting (contribution px, LOD px, x-ray, HiZ on/off), a fresh HiZ pyramid, or scene_epoch_ — bumped by chunk residency, visibility, colours, transforms, model add/remove/hide/unload. A frame requested for an overlay redraw, pick feedback, or a streaming tick where nothing landed draws from the buffers the last cull uploaded and skips the walk entirely. Benchmarks are exempt so bench numbers keep measuring the real cull. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
091b4d4113 |
ifcviewer: bound the CPU triangle shadow and the cull scratch to residency
The wasm heap grew past 2 GB on a 66-model session (surfacing first as the setBindGroup 2 GB TypeError, fixed separately) because CPU memory attached to loaded geometry never shrank while the GPU pool did: - mesh_triangles_cache — the dequantised positions + LOD0 indices the surface raycasts and measurement tools read — was filled once per mesh on first residency (gated on mesh_local_volumes == 0) and never released, converging over a session to the whole federation's geometry on the heap: 12 B/vertex + 4 B/index, 400 MB - 1 GB at this scale. And on web nothing reads it at all (no measurement tools yet). - Every chunk's cull scratch was reserved at model load (20 B/instance scene-wide) and the scratch + uploaded mirrors survived eviction. - Cull ran the HiZ test and emitted VisibleDrawGpu entries — then uploaded them — for non-resident chunks render() cannot draw. Now the shadow follows GPU residency: a per-mesh resident-chunk refcount (the spatial planner may duplicate a mesh into several chunks) is counted up in applyStreamedChunk and down in unloadChunk, releasing the mesh's entry at zero and refilling from the chunk bytes on the next residency. mesh_local_volumes (8 B/mesh) is kept across eviction so the Volume tool still covers evicted meshes. Hosts opt in via ViewportHost::wantsCpuMeshTriangles(): Qt yes, web no until the tools are ported — so on web the shadow costs nothing. Cull stops at the streaming counters for non-resident chunks, the eager scratch reserve is gone, and unloadChunk releases the scratch and uploaded mirrors. Clearing the mirrors also fixes a real staleness bug in unload/load: the model's cull buffers are recreated on load, and a stale mirror would make the memcmp dirty-check skip the first upload into the fresh (garbage) buffer. The heartbeat log reports the shadow (cpuTris). Measured on a 3-model / 990 MB scene: shadow tracks residency (493 MB at a 530 MB resident set, flat over minutes of streaming churn; previously monotonic), unload drops it to zero, reload refills it (verified via readbackMeshTriangles round trip). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8074541057 |
Surface VRAM shortfall to the user and let them unload models
When the geometry in view needs more GPU memory than the cache can hold, the viewer keeps the largest on-screen chunks resident and streams the rest as the camera moves. That is the right degradation, but it was invisible: nothing told the user the scene did not fit, and the only lever was removing or hiding models, neither of which is "keep it in the federation but stop spending GPU memory on it". Viewer core: - ModelGpuData::unloaded, with drawable() = !hidden && !unloaded now the test every cull / draw / pick / streaming pass uses. unloadModel evicts every chunk and releases the model's own buffers; loadModel recreates them from the CPU mirrors (no disk read) and lets chunks stream back. Recompose keeps the CPU instances current while a model is unloaded so a reload sees up-to-date transforms. The MeshGpu/InstanceGpu record builders are factored out so load and reload share them. - FrameStats reports the camera's working set: chunks wanted, how many of those are not resident, and their bytes. - modelVramBytes / isModelUnloaded accessors, forwarded by ViewportWindow. BonsaiViewer: - Models tree gains a memory column (name | MB | eye) refreshed once a second and on load-state changes; unloaded models read "unloaded" in italics. The viewport stays the single authority for the state; SessionState only carries the modelLoadStateChanged notification. - Context menu: "Unload Model" / "Load Model", distinct from hide and remove, reporting the MB freed in the status bar. - Status bar notice, independent of the perf-stats toggle, once the shortfall has persisted for 3 s (a moment of missing chunks after any camera move is normal): "GPU memory full: N of M visible chunks (X MB) not loaded", with a tooltip pointing at Unload. The perf label also shows "N/M chunks waiting". Verified on the GPU: unloading a 497 MB model frees it immediately with the others still rendering; reloading streams all 180 chunks back. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ab99024307 |
ifcviewer: budget the geometry cache and make required allocations fallible
Loading enough models drove the chunk pool to the driver's refusal point, after which the first click aborted: the pick attachments are allocated lazily, wgpu-native reported their OOM as a validation error nobody observed, and the invalid views reached wgpuQueueSubmit, which panics across the FFI boundary. Two policy defects compounding: the cache was allowed to take the last byte, and nothing but the pool's own growth was treated as fallible. GPU memory is now two tiers. Required allocations (per-pixel attachments, a model's metadata buffers, readback staging) are eager, deterministic and fallible; the chunk pool is an elastic cache that grows only to a budget and yields whenever a required allocation fails. - GpuBudget (pure, unit-tested): desktop derives the budget from the driver's free-memory report minus a reserve for the attachments at 4K; web keeps the wasm-heap cap; either lowers it on pressure. The budget's source differs per platform, the mechanism does not. - GpuAllocScope: the OOM/Validation error-scope dance in one place, synchronous on wgpu-native, provisional on Dawn-web. BufferPool's inline copy now uses it. - BufferPool::shrinkToCapacity releases whole sub-buffers newest-first after the owner empties them; growth clamps to the budget instead of overshooting. - ViewportCore::allocateRequired runs any required creation under a scope and, on failure, lowers the budget, evicts and releases cache sub-buffers, waits for the device to reclaim them, and retries until it fits or the cache is at its floor. Pick attachments are created with the other attachments in configureSurface; render() skips a frame rather than submit invalid views; a model whose buffers cannot fit is not loaded instead of aborting. Verified on a 4 GB GeForce: the pool clamps itself at the derived budget (256+256+67 MB for a 579 MB budget) and, in a standalone check against the real device, a pool grown to the driver's refusal point observes a failed required allocation, releases 320 MB and succeeds on retry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
2c1d445d5b |
ifcviewer-web: mint session model ids when a load is requested
A federated pick could be attributed to the wrong file. The model slot a host sees — ElementRef::model_index, modelProgress's index — is a rank in session_model_id order, and on web that id was minted at the END of the sidecar read chain, after three network round trips. So the ranking was the order the models' reads happened to finish in, not the order the host added them. With ~40 similarly-sized models over HTTP, adjacent models swapped and a click reported its neighbour's file; the host page then asked for a GUID the file does not contain. Mint the id at the top of loadSidecarMetadataWeb instead, which runs synchronously from load_sidecar_from_source_c and therefore in the order the host asked for its models. A load that fails partway just abandons its id, and the ranks compact over the surviving models as before. Positions are still positions, though: if one model fails to load, every later index shifts down one and a host mapping index into its own list silently drifts again. So also carry the source id — the handle the host minted itself when it registered the file — through ElementRef into the pick payload and getObjects rows, and document it as the way to attribute an object to a file. ModelGpuData::web_source_id defaults to -1 now, since 0 is a real source id and cannot double as "none". The test server grows a ?delay=<ms> knob so a test can force the losing interleaving: georef-a is added first and served slowly, and its objects must still come back as model 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
935562142e |
Apply a model's coordinate operation when its sidecar loads
.ifcview has carried the model's CoordinateOperation since v11 and the streaming reader has always parsed it, but applyCachedModel ignored it. The matrix only ever reached the scene because BonsaiViewer pushes it after every load via setModelCoordinateOperation. Nothing does that on web, so every model rendered in its local coordinates and two federated models with differing map conversions came out misaligned. Seed the matrix and the unit scales from the sidecar, and recompose the model afterwards. Seeding alone is not enough: the instance transforms in a sidecar are baked with identity federation matrices, and applyCachedModel uploads them as-is. The recompose also fixes a second case that had nothing to do with georeferencing — a model loaded while a federated false origin was already in force kept its unshifted transforms. ModelGpuData gains the unit scales because composeModelTransformation needs them to lift a transform's anchor point into metres, and on a sidecar-only load there is no IFC to read them back from. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
75c9da5098 |
ifcviewer: overhaul model/object ID tracking
Rename the two overloaded model identifiers and make object_id assignment single-authority, fixing a pick -> properties mismatch. Identifiers: - Per-model UUID fed_id -> model_id; the uint32 runtime handle model_id -> session_model_id (SessionState accessors + mirror hashes renamed to match). "fed_id" was a misnomer -- the federation is the whole collection, not one model. object_id assignment (fixes wrong class on click): - Producers (GeometryStreamer, .ifcview sidecar) now stamp model-LOCAL object_ids; ViewportCore::applyCachedModel is the sole authority that assigns the session-global id (base + local). Removed SceneLoader::next_object_id_, GeometryStreamer::lastObjectId(), and the streamer's start_object_id parameter. - The element table is stamped by the same base on both load paths (applySidecarData and onStreamerFinished), so registry ids match the ids pick returns. Previously the sidecar path double-rebased instances vs the registry (click IfcSite -> showed IfcDoor); the live-stream path had the same latent mismatch. Both closed. Naming / cleanup: - SceneLoader::addFiles -> queueModels; startStreamLoadFor -> loadFromGeometryStreamer; readSidecarMetadataOnly -> readSidecarMetadata. - Federation::addModel takes an explicit display_name (no QFileInfo fallback); callers pass QFileInfo(path).fileName(). - Disambiguate cryptic short locals (d->sidecar, m->model, c->chunk, ...) in SceneLoader, Federation, ViewportWindow, AreaMeasurement, SectionGizmoRenderer, and the SidecarData/SidecarReadPlan spots in ViewportCore. Tests: 125/125 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
66d558ec2d |
ifcviewer: rename sidecar transfer/record types; drop unused element hierarchy (sidecar v17)
Rename the streamer/sidecar transfer and record types to describe what they are rather than how they move: MeshChunk -> StreamedMesh InstanceChunk -> StreamedInstance InstanceCpu -> InstanceInfo PackedElementInfo -> ElementTableRecord uploadMeshChunk -> uploadStreamedMesh uploadInstanceChunk -> uploadStreamedInstance buildMeshChunk -> buildStreamedMesh and the two post-index sidecar metadata blocks: "critical" metadata -> "geometry" metadata (meshes/instances/georef/TOC) "deferred" metadata -> "element" metadata (elements + string table) parseSidecarCritical -> parseSidecarGeometryMetadata parseSidecarDeferred -> parseSidecarElementMetadata The one behavioural change: the element hierarchy (parent_id) was carried through ElementInfo, ElementTableRecord, and the sidecar element table but never consumed, so drop it and bump SIDECAR_VERSION 16 -> 17. No back-compat: regenerate sidecars. sample.ifcview is regenerated at v17. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0b8c787ac0 |
ifcviewer: v16 zstd-compressed sidecars (~10x smaller over the wire)
The .ifcview data is hugely redundant (repeated double instance matrices, patterned indices) — measured 12x zstd whole-file. Server Content-Encoding can't be used (it breaks HTTP Range), so compress PER-CHUNK into the format. Format (v16): geometry becomes per-chunk zstd(vertices)+zstd(indices) frames — each independently Range-fetchable, so streaming is intact — and the critical + deferred metadata blocks are single zstd frames. SidecarChunk carries the compressed blob offsets/sizes; applyStreamedChunk (render/upload) is UNCHANGED — decompression slots into the fetch. Full readSidecar (test/tooling) reconstructs by decompress+scatter. zstd: desktop links libzstd (also compresses at bake); the web build (Emscripten has no zstd port) FetchContent's the pinned zstd source and compiles its decompress-only subset for wasm — no vendored blob, same version as desktop. New SidecarCompress wraps it (compress guarded off under Emscripten). Both stream paths — desktop StreamingThread worker + sync fallback (readChunkGeometryCompressed) and web beginWebChunkLoad — decompress; readSidecarMetadataOnly / the web bootstrap / loadDeferredMetadataWeb decompress the metadata blocks. streamingByteProgress reports COMPRESSED bytes. MEASURED: a 752 MB v15 federation → 75 MB v16 (10x; per-file 6.7-15.3x); PP-PLP 118→15 MB, loads 13/13 chunks on web, 0 errors. Three fixes found while testing big federations on a real server: - Web-streamed race: streaming_from_web was set in the deferred-header callback (a round-trip after the model+chunks exist), so driveStreamingLoads could take the sync fopen path meanwhile → "failed to read/decompress chunk 0". Now set immediately after applyCachedModel. - OOM abort on 18 models: the pool grew unbounded until an alloc failed, but on web that's an uncatchable bad_alloc abort. Cap total pool capacity (setMaxTotalCapacity, 3 GB) so it stops before the heap ceiling, and raise MAXIMUM_MEMORY 2→4 GB (wasm32 max) for headroom. - Web never evicted (grow-or-block only). At the hard budget, fall through to the LRU/priority evictor so a big federation stays navigable (highest-contribution chunks win) instead of freezing with holes. 113/113 desktop + 6/6 web smoke pass. No back-compat: regenerate sidecars (desktop bakes v16; scratch conv tool migrates v15→v16). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5299f6c13c |
ifcviewer-web: contribution-cull streaming + combined loaded/needed/total bar
Streaming picked candidates on raw frustum visibility, so viewAll over a big
federation (all models in-frustum) fetched every chunk — even fine chunks that
project to sub-pixel and the renderer never draws. Gate streaming on the SAME
contribution decision the render path already makes.
- New per-chunk contribution_visible_count: instances that passed frustum AND
the contribution cull (projected radius >= min_radius_px), counted BEFORE HiZ
— so it's stable while the camera is still (unlike the HiZ-post counters,
which flip frame-to-frame and would thrash the loader) and shifts only on
navigation, when the working set should. The candidate gate skips chunks with
count 0: the network pulls only what's resolvable now; the rest stream in as
you approach. Verified: fit view needs the geometry, zoomed-out needs 0.
- Loading UI reworked from per-model segments to a combined bar over the whole
federation: dark track = not needed for this view, dim = needed-but-unloaded,
bright = loaded. "loaded / needed" = how done THIS view is; "needed / total" =
how much of the model the view requires — "Loading 45 / 90 MB for this view ·
12% of 718 MB total". Driven by ViewportCore::streamingByteProgress + ifcv_bytes_*.
Two fixes from testing: (1) show a distinct overhead phase while total==0
("Loading model data — X MB · Y/N models ready") so the metadata download
isn't a dead-looking bar; (2) re-assert display:block every active frame so the
bar REAPPEARS when navigation reveals new chunks (was only set in
beginLoadProgress → stayed hidden after the first catch-up). Per-model exports
kept for a future detailed view.
111/111 unit + 6/6 web smoke pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
4c38e52741 |
ifcviewer-web: multi-file loading (federation) via a per-model byte-source
The scene core is already multi-model — models_gpu_ is a map, applyCachedModel APPENDS, and per-model model_id / object_id rebasing / georef+transformation are how the desktop federates today. The only web-specific gap was the byte source: web had ONE global source (__ifcvFile/__ifcvUrl) and reset the scene on every load, so it could show one file at a time. Desktop meanwhile carries a per-model source (streaming_file_path). Mirror that on web: give each model its own web_source_id into a JS source registry (Module.__ifcvSources[id] = a picked File or a sized remote URL). beginWebChunkLoad, the metadata bootstrap, and the on-demand deferred fetch all read from the owning model's source, so several files stream concurrently into one federated scene — reusing all the shared machinery (viewAll, picking, the GUID fetch) untouched. - webReadRangesAsync / ifcvReadRangeInto / ifcvSourceSize take a source id. - loadSidecarMetadataWeb(source_id, …) appends (no resetScene); main_web exposes load_sidecar_from_source_c(id) + clear_scene_c(). - URL size resolution moved to JS (shell.html registers + sizes sources via HEAD/Range), retiring the C-side ifcvBeginUrlSource / ifcv_source_ready dance. - shell.html: source registry + "Open" (replace) / "Add" (append) buttons, multi-file selection; ?model= registers a URL source then loads. Verified: two sidecars from two sources stream into one scene, both fully resident, zero GPU errors. 111/111 unit + 6/6 web smoke pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
681de6f817 |
ifcviewer-web: pick logs the object's IFC GUID via the on-demand deferred fetch
First real consumer of the v15 deferred property block, and an end-to-end demonstration that on-demand property loading works. On a left-click pick, logSelectedObjectGuidWeb ensures the owning model's deferred block is loaded (loadDeferredMetadataWeb — a network fetch the FIRST time, cached after) and logs the picked object's GUID to the console. Fix uncovered while wiring it: applyCachedModel rebases instance object_ids to a per-model global base (object_id_base) to keep them unique across models, but the deferred elements carry the sidecar's original local ids — so a lookup by the picked (global) id missed. Store object_id_base on the model and rebase the elements by it when the deferred block loads. Verified: with a streamed model, the deferred block is fetched ONLY after the first pick (not at load), and the pick logs a valid 22-char IFC GUID. 111/111 unit + 6/6 web smoke pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
41a85a70ba |
ifcviewer: v15 — defer property metadata off the first-paint path
First-paint over a network is metadata-bound: the whole post-index metadata (~10 MB on a 118 MB model) had to download before any geometry. But ~25% of it — elements + string_table, the IFC element tree (names/GUIDs/hierarchy) — is used only for UI/picking, never for rendering (ViewportCore never touches it). v15 splits the post-index metadata into a render-CRITICAL block (meshes, instances, georef, chunk TOC) preceded by its byte length, then a DEFERRED block (elements + string_table). The web loader reads only the critical block before painting; the deferred block sits at a known, self-describing offset ([critical end, EOF)) and is fetched on demand. Desktop reads both (local). Web on-demand path is wired and complete (not yet called — no UI consumer): loadDeferredMetadataWeb(model_id) range-fetches + parses the deferred block into ModelGpuData.elements/string_table, at most once; the first consumer will be "show the selected object's name" on pick. No background prefetch — view-only sessions never download the property data (saves 2.64 MB on this model). parseSidecarTail split into parseSidecarCritical + parseSidecarDeferred (pure, unit-tested); StreamingSidecar gains the critical-block locator. Measured (118 MB model): critical metadata 10.35 -> 7.71 MB, deferred 2.64 MB off the path; first paint 10.5 -> 9.6 s @ 24 Mbps. (Instances still dominate the critical block — the next metadata lever.) Format -> v15, no back-compat; regenerate sidecars. 111/111 unit + 6/6 web smoke pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e1be2f208c |
ifcviewer: v14 chunk-contiguous sidecar + progressive network streaming
Makes large-model streaming over a network actually good — fixing read
amplification, then first-paint latency — building on the byte-range work.
v14 layout + TOC (SidecarLayout, pure + unit-tested)
The loader chunks meshes by spatial Morton order, but the sidecar stored
geometry in mesh-id order, so a chunk's meshes were scattered through the
file: streaming one chunk meant either hundreds of tiny range requests or
reading (and discarding) everything between them — a 113 MB model fetched
~340 MB, a 531 MB model 2.25 GB (4.2x). Fix: at bake, reorder meshes into
the loader's chunk order and rebuild vertex/index(LOD0+LOD1)/instance
sections so each chunk is one CONTIGUOUS byte range, and bake a chunk TOC
({first_mesh, mesh_count}). The loader builds chunks straight from the TOC
rather than re-deriving the plan — the float Morton quantisation isn't
bit-identical across toolchains (x86 baker vs wasm loader), so a re-derived
plan scatters the chunks. Format bumped to v14 (regenerate sidecars). The
reorder buckets instances by per-instance mesh_id (the baker never sets
MeshInfo.first_instance — trusting it scrambled every transform → geometry
at the origin). Multiset-verified on a 28,900-instance model: every
instance's placement + geometry preserved. Result: 531 MB fetches 531 MB
(1.0x) in 72 requests (was 2036).
Progressive streaming (concurrency cap + small chunks)
Even at 1x, geometry appeared only after ~the whole model arrived: the
browser multiplexes every in-flight Range request over one HTTP/2 conn, so
unbounded concurrency (9 in flight) split the bandwidth and nothing finished
until the end (measured: first paint after 113 of 118 MB / 35 s @ 24 Mbps).
Cap concurrent chunk loads (kMaxWebInflightChunks=2): the priority-sorted
top chunks finish and paint first, then the next → first paint 9 s. Chunk
size dropped 16->4 MB (cheap now that each chunk is one read; matches Cesium
3D Tiles / xeokit / SVF2) for smoother progression. First-paint is now
metadata-bound (~10 MB tail) — the next lever.
111/111 unit (new test_sidecar_layout: geometry preserved, contiguous layout,
Morton-identity) + 6/6 web smoke pass; desktop bake (SceneLoader) reorders
before writeSidecar; embedded web sample regenerated to v14.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
23dc5dac48 |
ifcviewer-web: stream remote sidecars over HTTP Range (?model=URL)
Adds a network byte-source alongside the local Blob one. The async-chunk
infra is source-agnostic — only the two JS primitives knew it was a Blob —
so this generalises them and reuses everything else:
- ifcvReadRangeInto: local → Blob.slice; remote → fetch() with a Range
header (206). If a server ignores Range and returns 200, the requested
window is sliced out so it still works (without the bandwidth saving).
- ifcvFileSize: Blob size, or the URL's total length resolved up front.
- ifcvBeginUrlSource: resolves total size (HEAD Content-Length, else a
0-0 ranged GET's Content-Range) then fires _ifcv_source_ready.
- The metadata bootstrap is extracted into a source-agnostic
loadSidecarMetadataWeb(label); loadSidecarFromBlobWeb / FromUrlWeb are
thin entries. streaming_from_blob → streaming_from_web (now covers both).
main_web exports load_sidecar_from_url_c(url); shell.html reads a
?model=URL query param and ccalls it once the app is live (same-origin
needs no CORS; cross-origin hosts must send CORS + Accept-Ranges).
Test: serve.mjs now answers HEAD + Range (206) and falls back to the
ifcviewer-web source dir for sample.ifcview (embedded in the wasm, not in
build-web). New smoke case loads ?model=/sample.ifcview and asserts it
renders via the Range path. 6/6 web smoke + 107/107 unit pass; desktop
unaffected (web-guarded; only the shared field rename touches it).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
584504dcdc |
ifcviewer-web: stream user sidecars via Blob.slice byte ranges (#88)
Picked files are no longer copied whole into the wasm heap. The browser
File object stays in JS (Module.__ifcvFile) and is read lazily through
Blob.slice byte ranges, so a 200-500 MB sidecar never enters wasm linear
memory — only chunk-sized slices do.
Mechanism (web-only, #if __EMSCRIPTEN__):
- JS glue (EM_JS): ifcvFileSize + ifcvReadRangeInto — slice [off,off+n)
of the File and copy it into a caller-provided heap pointer, then call
back _ifcv_on_range_done. No malloc across the boundary; C pre-sizes
the destination from the read plan.
- webReadRangesAsync: reuses planSidecarReadRanges to coalesce a range
set into Blob.slice reads (1 MB gap — each slice is an async hop),
scatters them into a destination laid out in input order, and fires a
continuation when the whole set lands. An in-flight map keyed by id
survives unordered_map rehash (scratch buffers are heap-owned).
- loadSidecarFromBlobWeb: async metadata load — head (16 B) -> index
count -> tail-to-EOF -> parseSidecarHead/Tail -> applyCachedModel, then
tags the model streaming_from_blob and frames it.
- driveStreamingLoads: blob-sourced models route to beginWebChunkLoad
(async vertex+index range reads -> applyStreamedChunk in the callback),
holding is_loading until the bytes arrive. The embedded MEMFS sample
keeps the synchronous fopen path.
shell.html stashes the File and calls _load_sidecar_from_blob_c instead of
FS.writeFile'ing the whole thing; EXPORTED_RUNTIME_METHODS=['FS'] dropped.
Desktop is untouched (the new members + driveStreamingLoads branch are all
emscripten-guarded). Web links clean; desktop rebuilds; 107/107 unit tests
pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
b123ee69d6 |
viewer: two-pass alpha transparency + Alt+X global x-ray cap
## The bug
FZK-Haus windows rendered fully opaque despite every piece of the
data path carrying alpha correctly: vertex format is RGBA u8x4,
InstanceCpu/InstanceGpu carry color_override_rgba8 with its alpha
byte, fs_main returns vec4(rgb, in.color.a). Cause: the main render
pipeline's color target had `blend = nullptr`, which in wgpu disables
the blend stage entirely — fragment RGBA overwrites the back buffer
unmodified, alpha discarded.
## Why "just enable blend" isn't enough
Two failure modes that don't go away with a one-liner:
1. `depthWriteEnabled = True` on the main pipeline would make a
transparent window-frame pane occlude geometry behind it in
depth, so the wall behind the window then fails the depth test
and never draws — you'd see the silhouette of the window with
whatever colour was in the back buffer before, not the wall.
2. Order-dependent blending across transparent surfaces in arbitrary
cull order — overlapping transparent surfaces would shift colours
as the camera moves.
Standard fix for a BIM viewer is two-pass opaque-then-transparent.
## What this commit adds
### Per-mesh "has any alpha < 255" classifier
* `ModelGpuData::mesh_has_alpha` (uint8_t vector, parallel to meshes).
* Sized in `applyCachedModel`.
* Populated in `applyStreamedChunk` by scanning each in-chunk mesh's
vertex bytes for a vertex's alpha byte < 255 (offset 11 within
the 12-byte vertex record — the 4th byte of the third u32, which
the shader reads as `w2 >> 24`). Single chunk-arrival site covers
both sidecar streaming and the worker-result drain. First-load
IFC-without-sidecar geometry still routes opaque until the sidecar
bake completes; A-path scan is deferred.
### Per-chunk opaque/transparent partition during cull
* `Chunk::opaque_visible_vertices` / `opaque_visible_draws`
(per-frame counts).
* Transient `visible_draws_scratch_transparent` +
`transparent_per_draw_vertex_counts` filled alongside the existing
opaque half during the cull walk. Post-walk concat appends
transparent entries onto the opaque half and continues the
cumulative prefix-sum sequence — single buffer, single bind
group, no doubling.
* Classifier inside the cull lambda:
`xray_active ? always_transparent
: override_active ? (override.alpha < 255)
: mesh_has_alpha[mesh_id]`
### Per-chunk uniform layout extension
From `[total_draws, total_verts, 0, 0]` to
`[total_draws, total_verts, opaque_verts, opaque_draws]`. The third
slot is what `render()` passes as `firstVertex` to the transparent-
pass draw call so the shader's vid lands in the transparent range of
the same visible_draws_scratch buffer.
### `main_pipeline_transparent_`
Copy of `main_pipeline_` with `color_target.blend = SrcAlpha /
OneMinusSrcAlpha`. depthWriteEnabled stays True (see below).
### Two-pass `render()`
Opaque pass (`main_pipeline_`, firstVertex=0,
vertexCount=opaque_visible_vertices) then transparent pass
(`main_pipeline_transparent_`, firstVertex=opaque_visible_vertices,
vertexCount=total - opaque). Each loop skips empty halves so an
opaque-only chunk costs one draw call, transparent-only one draw,
mixed chunks two.
### depth_transparent.depthWriteEnabled = True (NOT off)
Initially set False (standard "let further-back geometry paint
through transparent front faces" trick) but that broke the edge-
detect pass: edge detection reads the depth buffer to find
silhouette discontinuities, and windows-without-depth meant the
glass had no silhouette at all (panes looked like framed holes) and
the edges of opaque geometry behind the glass painted through at
full intensity. Keeping the write avoids that — trade-off is depth-
test occlusion between transparent surfaces (closer occludes
farther), which for BIM panes that don't overlap in screen space
is invisible. Real fix for the overlap case is OIT or sort-back-
to-front, not depth-write toggling.
## Alt+X global X-ray (drops in basically free)
* `xray_alpha_cap` field on FrameUniforms + WGSL counterpart, default
1.0 (no effect). fs_main clamps `out.a = min(in.color.a, cap)`.
* `ViewportWindow::xray_alpha_cap_` member, default 1.0. Alt+X
toggles between 1.0 and 0.3.
* Cull classifier sees `xray_alpha_cap_ < 1.0` and forces every
instance into the transparent pass so the blend stage actually
fires (an opaque-pass fragment with capped alpha would still
overwrite the back buffer).
* No per-instance state mutation needed — toggle is a single float
in a uniform plus a re-cull. Excluding objects from x-ray later
would mean tagging them so the classifier skips the force-
transparent branch for them, also small.
Stress-tested on FZK-Haus: window glass visibly translucent with
correct silhouette edges; Alt+X turns the whole scene to a tinted
ghost of itself and back without artefact.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
8ab5c31e75 |
refactor: merge ifcviewer-wgpu into ifcviewer, drop Wgpu prefix
The GL backend is gone (task #53). The wgpu/non-wgpu folder split and the Wgpu* class prefix were both disambiguation artefacts from the overlap period — now pure dead weight. ## Folder + library merge * `src/ifcviewer-wgpu/` → folded into `src/ifcviewer/` (git mv tracks every file as a rename so blame/log history survives). * `src/ifcviewer-wgpu-minimal/` → `src/ifcviewer-minimal/` (the exe was already named `IfcViewerMinimal`; this just brings the folder + CMake target name into line). * `src/ifcviewer-wgpu/tests/test_wgpu_{selection,visibility}.cpp` → `src/ifcviewer/tests/test_{selection,visibility}.cpp`, folded into the existing `add_ifcviewer_unit_test(...)` helper. * The `IfcViewerWgpu` static library is dissolved — its sources become part of the unified `IfcViewer` static library, which now bundles scene/loader + renderer in one target. The pre-merge circular dependency (IfcViewer linking IfcViewerWgpu just to get the ViewportWindow.h include path that SceneLoader.h needs) goes away. * The wgpu-native FetchContent block, the Cocoa/QuartzCore link on Apple, the OBJCXX-enabled `.mm` source, and the wgpu-native runtime install all move into `src/ifcviewer/CMakeLists.txt` unchanged. ## Type renames (Wgpu prefix dropped from every Wgpu* identifier) WgpuAreaMeasurement → AreaMeasurement WgpuBufferPool → BufferPool WgpuLengthMeasurement → LengthMeasurement WgpuMetalSurface → MetalSurface WgpuModelGpuData → ModelGpuData WgpuOverlayFrame → OverlayFrame WgpuOverlayRenderer → OverlayRenderer WgpuSectionPlane → SectionPlane WgpuSelectionState → SelectionState WgpuStreamingLoader → StreamingLoader WgpuStreamingThread → StreamingThread WgpuViewportWindow → ViewportWindow WgpuVisibilityState → VisibilityState CMake target IfcViewerWgpuMinimal → IfcViewerMinimal (exe name was already this since wgpu shipped as default). Deliberately kept: `onWgpuLog` (wgpu-native log callback — names a binding to an external API, not one of *our* types), and the WGPU* enum/struct prefixes from wgpu-native's own headers. `WgpuMemProbe` lives in the separate `src/wgpu-mem-probe/` standalone diagnostic project and isn't touched. ## Include-path updates Every `#include "../ifcviewer-wgpu/Wgpu<X>.h"` → `"../ifcviewer/<X>.h"`, every in-directory `#include "Wgpu<X>.h"` → `"<X>.h"`. Includes from sibling subdirectories (modules/, etc.) are updated to point at `../../../ifcviewer/` instead of `../../../ifcviewer-wgpu/`. ## cmake/CMakeLists.txt simplification The redundant `add_subdirectory(ifcviewer-wgpu)` blocks (one inside the BUILD_BONSAIVIEWER fan-in, one in the BONSAIVIEWER-less standalone block) collapse into a single unconditional `add_subdirectory(../src/ifcviewer ifcviewer)`. The standalone block keeps only `wgpu-mem-probe` (the diagnostic tool, unrelated to the viewer lib). ## Verification * Full build green: `IfcViewer` static lib, `IfcViewerMinimal` exe, `BonsaiViewer` exe, all four pre-existing ifcviewer unit tests, and the two new-location tests (`test_selection`, `test_visibility`). * No stray `Wgpu<X>` identifier remains across `src/ifcviewer/`, `src/bonsaiviewer/`, `src/ifcviewer-minimal/` (verified by grep). * Renames tracked by git as `R` entries — `git log --follow` on ViewportWindow.cpp etc. continues to show history through the move. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |