mirror of
https://github.com/IfcOpenShell/IfcOpenShell.git
synced 2026-08-10 17:58:20 +00:00
126d2d4c06d175fd7d85fcf86f33b8196c7f735b
19042 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
126d2d4c06 |
wgpu: spatial instance bucketing for streaming (env-gated prototype)
WGPU_SPATIAL_BUCKETS=1 swaps the applyCachedModelStreaming planner from mesh-keyed Morton+greedy to octree-style instance bucketing. Default behaviour unchanged (env var unset → mesh-keyed planner runs). Phase 1 of #55 / #56. The mesh-keyed planner produces chunks whose AABBs are the union of all instances of the chunk's meshes — for heavily-deduplicated IFC meshes (a "standard floor tile" used 800 times across a federation) the mesh's "centroid" is a mean of scattered instance positions and the chunk's AABB ends up spanning the entire model. Symptom: chunk-level frustum cull rarely fires (AABB always intersects view), and the screen-area priority metric under-rates big-AABB chunks because their corners straddle the near plane. Visible objects pop in/out as the camera tilts, even though they're fully on screen. The spatial planner bucketises INSTANCES directly. Each leaf bucket contains its instance list + the unique mesh data those instances reference. A mesh whose instances scatter into multiple buckets gets its vertex/index data uploaded into multiple pool slices — duplication is the cost for tight bucket AABBs. For IFC this is acceptable: heavily-shared meshes tend to be small (fittings, fasteners), so per-bucket duplication adds tens-of-MB not GB. Octree implementation (planSpatialChunks): - work-stack subdivision: for each (instance subset, AABB), split into 8 octants around centre and recurse - stop conditions: bucket fits WGPU_CHUNK_VERTEX_BYTES_LIMIT for union vertex bytes AND ≤ spatial_max_instances_ instances; OR single instance left; OR every instance falls into the same octant (pathological — emit as leaf rather than infinite recurse) - spatial_max_instances_ default 5000, overridable via WGPU_SPATIAL_BUCKET_MAX_INSTS env var so the prototype can be swept without rebuilding Data-model adjustment beyond what dc2927997 prepared: - Per-chunk per-mesh chunk-local offset table (chunk_mesh_offsets) built during the chunk-construction loop. The mesh-keyed per-mesh global arrays (mesh_chunk_idx etc.) still get populated for legacy reads, but under spatial bucketing they're overwritten when the same mesh appears in multiple chunks — harmless because cull reads the per-instance arrays exclusively (per dc2927997). - Post-construction, per-instance arrays are populated from chunk_mesh_offsets via (instance_to_chunk[i], inst.mesh_id) lookup. Mesh-keyed planner derives identical values to before (pixel-identical); spatial planner now writes the correct per-bucket offsets even when the mesh appears in multiple chunks. basic.ifc parity on all three paths confirmed (non-streaming mesh-keyed, streaming mesh-keyed, streaming spatial all produce 0 pixel diff vs the reference). Spatial planner produced 1 bucket on basic.ifc (3 instances, well under thresholds) as expected. Non-streaming applyCachedModel left unchanged — the prototype targets the streaming path which is where the federation-scale missing-objects issue lives. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
1f8f6bffc4 |
wgpu cull: refactor per-mesh chunk lookups to per-instance
Preparatory refactor for spatial instance bucketing (#55). Cull previously routed through per-mesh tables (mesh_chunk_idx, mesh_chunk_local_base_vertex, _ebo_first_u32, _lod1_first_u32) to find the chunk and chunk-local offsets for each instance. That assumes a mesh lives in EXACTLY ONE chunk — the assumption holds under the current mesh-keyed planner but breaks under spatial bucketing, where the same mesh can be duplicated across multiple buckets if its instances are scattered. Adds four per-instance arrays (instance_chunk_idx, instance_base_vertex, instance_ebo_first_u32, instance_lod1_first_u32) populated at planning time. The current mesh-keyed planner derives them by translation: instance_chunk_idx[i] = mesh_chunk_idx[instances[i].mesh_id] The spatial-bucket planner (next commit) will populate them directly, allowing the same mesh_id to map to different chunks for different instances. cullModelCpuCompute now reads the per-instance arrays: - frustum_visible_count uses chunks[instance_chunk_idx[i]] - VisibleDrawGpu uses instance_base_vertex / instance_ebo_first_u32 / instance_lod1_first_u32 - LOD-select branch and use_lod1 logic unchanged. Per-mesh tables stay (used by makeChunkRequest, debug logs, the planner itself). Memory cost: 16 bytes × N instances ≈ 16 MB on a 1M-instance scene. Pixel-identical on basic.ifc in both non-streaming and streaming modes. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
3eaa35d748 |
wgpu: input parity, fly mode, diagnostic instrumentation, chunk-priority fix
Brings the wgpu viewport's keyboard + mouse into line with GL ViewportWindow
+ Bonsai's MainWindow shortcut table, lands fly-mode, swaps in three
diagnostic env vars, and fixes a chunk-priority bug exposed by the
diagnostics.
Keyboard parity with GL + Bonsai:
P — toggle perspective / ortho projection
X / Shift+X — front / back view (eye on ±X, pitch 0)
Y / Shift+Y — right / left view (eye on ±Y, pitch 0)
Z / Shift+Z — top / bottom view (pitch ±90°)
F — focus camera on currently selected object
Home — frame entire scene
C — print --camera CLI args for current view
H — hide selected
Shift+H — isolate selected (hide everything not in selection)
Alt+H — show all (clear hidden set)
Shift+F — enter fly mode (matches BonsaiViewer)
Escape (fly) — exit fly mode
WASD/QE/Shift — fly movement (when in fly mode)
The previous H/Shift+H/I assignments were wrong vs Bonsai (Shift+H went
to show-all, I to isolate); both are fixed.
Fly mode:
- GL-style absolute m/s base speed (default 5.0), Shift = 5×, scrollwheel
adjusts ×1.25/×0.8 per notch (Blender convention). Scrollwheel does
NOT zoom in fly mode; that interfered with speed when speed was
distance-scaled (it was, briefly; replaced with absolute m/s).
- Mouse-look pins eye: yaw/pitch update first, then target is re-derived
so orbitEye(target, dist, new_yaw, new_pitch) == old eye. Result:
camera rotates in place (FPS) rather than orbiting the pivot.
- dt ceiling clamp at 100ms (matches GL fps_move_speed_) so a stall
doesn't warp the camera.
- Pitch sign matches non-inverted FPS convention (mouse-up = look up).
Mouse-nav presets (WGPU_NAV_PRESET=blender|rhino|revit, default blender):
Blender — Orbit MMB, Pan Shift+MMB
Rhino — Orbit RMB, Pan Shift+RMB
Revit — Orbit Shift+MMB, Pan MMB
LMB stays free for selection in every preset. Nav-drag kind is captured
at press time so a mid-drag Shift release doesn't flip orbit↔pan. The
pan up-vector switches to world-Y at near-vertical pitch so panning
still works in top/bottom view (would otherwise NaN at pitch=±90°).
Camera-math refactor:
buildViewProj(view, proj) centralises perspective↔ortho selection and
the near-vertical up-vector switch. Four open-coded copies of the
view/proj build (cull, debug, streaming priority, render uniforms) now
call it instead, ensuring projection mode + up-vector switch land
identically everywhere. basic.ifc pixel-diff is 0 (refactor confirmed
output-equivalent on the path with no ortho / no near-vertical pitch).
WGPU_PRESENT_MODE=fifo|fifo_relaxed|mailbox|immediate (default fifo):
Diagnostic toggle for stutter analysis. fifo_relaxed gave the tightest
per-frame dt distribution on a federated bench scene; immediate gave
uncapped throughput at the cost of tearing. Mailbox not supported on
Vulkan + NVIDIA Linux but kept as an option for other backends.
WGPU_FLY_DEBUG=1: per-frame [fly] log printing dt, render gap, key
count, speed, position delta. Confirmed render_gap == dt to four
decimal places — fpsIntegrate runs exactly once per render, no
double-tick. Cull cost (~14-20 ms) is the dominant frame variance and
the eventual fix is task #49 (sub-model parallel cull) — fly-mode
stutter on slow scenes is a downstream symptom of cull cost, not a
fly-mode bug.
WGPU_STREAM_DEBUG=1: per-frame [stream-debug] log with cands / enq /
drained / ev_lru / ev_pri / blocked / resident / cycled / max_load.
Confirms thrash / pool-bound / load-budget cases on big scenes.
Pick-and-track diagnostic: clicking an object enumerates every chunk
holding instances of that object (an IFC object can split across
representations / chunks), printing each chunk's AABB + instance AABB +
residency. If any tracked chunk's is_resident flips true → false in
driveStreamingLoads, an EVICTED dump prints with the chunk's AABB,
priority, pool state, this-frame eviction counts, and the top-5
candidates that displaced it. Surfaces exactly why an object disappeared.
chunkScreenAreaPx fix (uses diagnostic to confirm the bug):
A chunk's AABB is the union of every instance's world AABB in the
chunk. On a federated IFC the camera commonly sits INSIDE that AABB
(e.g. inside a 263×30×15 m floor-area bounding box). Previously the
8-corner projection silently dropped corners with clip.w <= 1e-3
(behind near plane), so the projected bbox of the surviving in-front
corners was a tiny fraction of the chunk's true on-screen footprint.
Result: big-AABB chunks lost every eviction fight, visible objects
popped out as the camera tilted. Fix: short-circuit to full-viewport
area when (a) eye is inside the chunk AABB (mirrors GL's
contributionPasses camera-inside short-circuit), or (b) any AABB
corner sits behind the near plane (AABB straddles → 8 corners cannot
honestly measure footprint; conservatively over-prioritise).
The fix is a workaround for the deeper chunking issue — chunks are
mesh-keyed (group of meshes), and a mesh's AABB used in chunking is
the mean of its instances' positions, which is meaningless for
heavily-deduplicated meshes scattered across the scene. The root fix
is task #55 (spatial instance bucketing, runtime prototype) + #56
(sidecar v15 instance-keyed format). chunkScreenAreaPx fix unblocks
the user-visible "missing objects" issue while those land.
Extracted chunkScreenAreaPx from a driveStreamingLoads-local lambda
to a private member so the disappear-diagnostic and any future call
sites can use it consistently.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
a63bc63439 |
wgpu: diagnostic instrumentation for cull/streaming perf
Adds three knobs and three new heartbeat numbers to support the ongoing perf-parity work. None affect behaviour in default runs. cull[wall|compute|upload] split timer The existing cull_timer wrapped both the parallel std::async dispatch and the sequential cullModelCpuUpload loop (queueWriteBuffer × 3 per resident chunk × ~120 chunks ≈ 360 wgpu calls per frame). Splitting them ruled upload out as the bottleneck on a 51-model federation scene: compute ≈ 16-17 ms, upload ≈ 1 ms. WGPU_CULL_THREADS=0 — force sequential cull std::async-per-model was already in place; this env var disables it so we can compare wall time vs sequential and confirm parallelism is working. On the federation scene with 52 models: sequential 74 ms vs parallel 17 ms = 4.4× speedup. Confirmed; the 17 ms floor is not a parallelism failure, it's the cost of culling the largest single model (model 43, 114k instances) ÷ no parallelism within that model. WGPU_STREAM_DEBUG=1 — per-frame [stream-debug] log Surfaces cands/enq/drained/ev_lru/ev_pri/blocked/resident/cycled/ max_load each frame from driveStreamingLoads. The "cycled" / "max_load" pair makes thrash vs eviction-churn vs just-loading distinguishable. Off by default; opt-in via the env var. Bench-warm timeout dump When [bench warm] times out (600 frames without 0-loads streak), prints a structured summary: resident/missing/total chunks, cycled count, pool usage, largest free run, avg missing chunk size, and an auto-classifier diagnosis (POOL FRAGMENTED vs WORKING SET > POOL vs FEW-CHUNK CYCLE vs still-loading). Caught a real fragmentation pattern (18 MB largest free run vs ~100 MB typical chunk) on a 51-model run where the dumb classifier would have called it a load-budget problem. LOD1 firing counter "lod1 X/Y (saved Z tris, N no-lod1)" suffix on the [frame] log. X = LOD1-selected this frame, Y = LOD1-eligible, Z = tris not drawn vs always-LOD0, N = visible instances with no baked LOD1 (mesh below IFC_LOD_MIN_TRIS). Confirmed LOD1 path is genuinely firing post the per-chunk LOD1-storage commit, and exposed that ~90% of instances in real scenes are no-lod1 meshes — relevant to the future LOD-tier-residency design. Chunk.load_count + Chunk.lod0/1 layout bookkeeping Per-chunk reload counter for the thrash detector. lod0/1 layout_count fields prep the data model for distance-tiered residency (Phase B of #31) but aren't acted on yet. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
6a3dd4a0eb |
wgpu cull: contribution-cull defaults 2/10 → 3/15, env-var overrides
wgpu computes projected_px from view-Z distance (forward·(centre-eye)), which is the perspective-divide-correct denominator: an instance's on-screen radius really is world_radius * focal / z_view. GL computes the same value but with euclidean distance (sqrt(dx²+dy²+dz²)) — for off-axis instances euclidean > z_view, so GL underestimates screen size and culls more aggressively at the same numeric threshold. Concretely on the federation scene at the test camera, wgpu was drawing ~3× the instances GL drew despite identical 2/10 thresholds: wgpu obj 11067, hiz_rej 16532 vs GL obj 2310, hiz_rej 6823. Same fps (vsync-pinned), but ~30% more cull work for raster output that the user already wasn't seeing because GL had been quietly dropping it. Keeping wgpu's view-Z formula (more physically correct) and bumping the thresholds to 3.0 / 15.0 to match GL's effective drop rate. On the federation scene this lands obj/tri counts within ~10% of GL across the orbit, and shaves cull from 3.44 ms to 2.61 ms avg. basic.ifc parity is unchanged (its 3 instances clear 3 px easily). WGPU_MIN_PX / WGPU_MIN_PX_MOTION env vars added so the thresholds can be swept without rebuilding — needed while we visually confirm the new defaults across more scenes / cameras. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
10f8fc88e8 |
wgpu chunks: pack LOD1 indices alongside LOD0; cull picks per-instance
Sidecar LOD1 is per-mesh, index-only — meshoptimizer-baked decimated index slices that share the LOD0 VBO. Before this change the wgpu chunk path force-disabled LOD1 (effective_lod1 = false; tagged in a comment as "until per-chunk LOD1 storage lands"), so it had to walk every visible instance at full LOD0 even when the per-instance LOD selection said the projected radius was below the LOD1 threshold. Per-chunk layout: append LOD1 indices for the chunk's meshes after all LOD0 indices in the same pool slice. A new mesh_chunk_local_lod1_first_u32 array records each mesh's LOD1 starting offset (in u32s) within the chunk's index slice; LOD0 offsets stay where they were. The cull's emit then sets VisibleDrawGpu.ebo_first_u32 to whichever side matches the per- instance use_lod1 decision the prior #8 commit already computed. Vertex pulling is oblivious to the LOD split — it just reads the indices the cull pointed it at. c.index_count is repurposed as the LOD0+LOD1 total so the pool allocation, eviction's pool-fit check, and the VRAM accounting all scale automatically. c.lod1_index_count exposes the LOD1 portion for stats. makeChunkRequest appends LOD1 byte ranges to req.i_ranges in the same per-mesh order; the streaming worker concatenates ranges in order, so the assembled idx blob lands LOD0-first / LOD1-second which matches the chunk-local packing. On a 10-model regen with meshoptimizer-baked sidecars, the bench [frame] heartbeat shows lod1 firing on ~80-90% of LOD1-eligible instances and saving 10-30M tris per frame versus the prior LOD0-only ceiling. Same camera/scene on basic.ifc is still pixel-identical (no mesh in basic.ifc is large enough to bake a LOD1, so the cull just follows the LOD0 path it always did). Temporary debug counters (lod1_dbg_count_ et al.) print "lod1 X/Y (saved Z tris, N no-lod1)" in both interactive and benchmark [frame] heartbeats — kept on while LOD1 correctness gets confirmed across more scenes, will come out once trust is built. The non-streaming applyCachedModel path runs the same LOD1 plumbing but no longer fits the full federation scene in pool (extra index bytes push past the 2 GB single-buffer cap); that mode was already streaming-only on that scene before. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
c7fba1abb8 |
wgpu streaming: screen-space AABB priority + grace period + interactive heartbeat
The chunk priority metric is now the 2D projected pixel area of the chunk's AABB on screen — 8 corners projected through view-projection, 2D axis-aligned bbox of the projected points, clamped to viewport. This replaces the prior bounding-sphere-radius² metric, which was a 3D approximation: it treated a 322 × 55 × 5 m slab as a 163 m sphere, giving it the same huge priority face-on or edge-on. The new metric genuinely answers "what would this chunk's AABB cover if rendered solid given the current camera and viewport." Newly-loaded chunks get a 30-frame grace period at full priority (visibility_history floor temporarily forced to 1.0). Without it, just-loaded chunks crashed to history=0 → effective priority = pri × 0.05 → immediately reverse-swapped by the chunk they displaced. Cycle starved the per-frame load budget so candidates ranked below the cyclers never got attempted. 30 frames = HISTORY_ALPHA's time constant — enough for visibility_history to develop meaningfully. EVICT_PRIORITY_RATIO bumped 1.21 → 2.0 to suppress more swap noise between similar-priority chunks. Interactive heartbeat log added: every render in non-bench mode prints [frame] with fps, ms, obj, sub_draws, hiz_rej, cull, stream, chunks breakdown (resident/frustum/total + missing count), VRAM, model count. Every 30 frames when something's missing, also dumps: - top 8 models by missing chunk count - top 20 missing chunks by priority (with AABBs) - bottom 5 residents by effective priority - all chunks of brace.ifc (one-off diagnostic, hardcoded for the brace-visibility investigation) The heartbeat made the streaming bug visible: a brace model that isolation-loads correctly is missing in the full set because slabs covering more pixels win the priority contest. Per-model fairness or manual pinning are the remaining options if pixel-area + grace + hysteresis isn't enough — left for follow-up so the user can decide based on real testing. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
3d368e0079 |
wgpu chunks: 3D Morton-code spatial sort (tight voxel chunks)
The previous chunk-plan sorted meshes by lexicographic (z, y, x) centroid — effectively a 1D Z-slab traversal. On a typical IFC building (50 × 50 × 100 m), a 16-MB chunk's 80-ish meshes spanned roughly 50 × 50 × 0.5 m. On a city federation it was much worse: the first chunk grouped ground-floor stuff from every building, spanning the entire scene horizontally. Per-chunk AABBs that wide make frustum / contribution / HiZ rejection useless (every chunk "overlaps the frustum" by virtue of spanning the whole scene). 3D Morton (Z-order) interleaves bits of quantised (x, y, z) centroids, so consecutive items in the sorted order cluster in all 3 axes — chunks become tight 3D voxels of the model. Prerequisite for the contribution-aware eviction priority (task #25) to actually discriminate near and far chunks. 21 bits per axis = ~2 M bins per axis, sub-millimetre precision on a kilometre-scale scene. Both apply paths (streaming and non- streaming) share the same sortMeshIdsByMorton helper. Benchmark unchanged (~47 fps avg, 20 ms cull, 0.3 ms stream) — the distance-based evictor still keys on chunk centres, which moved slightly under Morton but not enough to materially shift residency. The user-visible win comes from the next commit, which switches priority to screen-space contribution × HiZ history — both of which need today's tight AABBs to mean anything. Pixel-identical to non-streaming on basic.ifc on both paths. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
0ec72482c2 |
wgpu pool: halve-on-failure in addSubBuffer extracts +35% VRAM
Many Vulkan drivers cap a single VkDeviceMemory allocation at exactly
maxStorageBufferBindingSize (NVIDIA: 2 GB on consumer GeForce) or
refuse big contiguous allocations once heap is fragmented. The old
addSubBuffer gave up at the first refusal, latching growth_disabled_
— so on a 4 GB GeForce we extracted 2 GB and called it done.
The wgpu-mem-probe tool (
|
||
|
|
feab05650d |
wgpu-mem-probe: standalone tool to investigate driver VRAM ceilings
New headless wgpu probe app — no Qt, no surface, just initializes a device and stress-tests buffer allocations. Reports: 1. Adapter + device limits (maxBufferSize, maxStorageBufferBindingSize). 2. Single-allocation probe: descending sizes, each released, finds the largest single buffer the driver will grant. 3. Cumulative probe: halve-on-failure, finds total VRAM the runtime will let us park behind one device across multiple sub-buffers. 4. Fixed-size cumulative probe: 1 GB / 512 MB / 256 MB uniform sizes, to detect whether the "big-first" strategy leaves VRAM on the table. Findings on a GTX 1650 (4 GB physical) + wgpu-native + Vulkan: - maxStorageBufferBindingSize = 2 GB (driver cap, not wgpu-native). - Any single storage buffer > 2 GB is REFUSED. - Total available across N sub-buffers = ~3 GB, INVARIANT under allocation pattern (2+1+0.06, 3×1 GB, 6×512 MB, 12×256 MB all reach 3.00 GB). Driver hands out a fixed VRAM slice; pattern doesn't matter. - Remaining ~1 GB is held by the desktop compositor + OS. - GL's higher "4 GB+ resident" claim is overcommit into host RAM, which wgpu/Vulkan don't do. The +50% (2 → 3 GB) improvement is real and worth chasing — a follow-up halve-on-failure addSubBuffer in WgpuBufferPool will extract that on this card. On 8/16/24 GB GPUs the same code gets us proportionally more. Build: ninja -C build-viewer-wgpu WgpuMemProbe Run: ./build-viewer-wgpu/wgpu-mem-probe/WgpuMemProbe Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
dcc2bf1c01 |
wgpu streaming: background-thread chunk I/O kills render-thread stutters
The sync chunk-read on the render thread was causing 100-300 ms spikes during orbit whenever a new chunk needed to scatter-gather its mesh bytes from disk. p99 was 326 ms on the close-camera benchmark. New WgpuStreamingThread: one worker thread with a condvar-protected request/result queue. driveStreamingLoads becomes drain-then-enqueue: 1. Drain any results the worker pushed since last frame. For each, pool-allocate slices + queueWriteBuffer + build the chunk bind group (still main-thread because wgpu queue ops aren't thread-safe). 2. Walk visible non-resident chunks (sorted by distance), evict to make pool room, and enqueue the request. Chunk gains is_loading flag to prevent re-enqueueing while in flight. loadChunkBytesAndUploadGpu becomes the sync fallback path, used only when a screenshot is pending — the deferred-capture wait would otherwise let the window manager re-layout the window between frames and the test framework would capture at the wrong size. Normal streaming always goes through the worker. Bench warm-gate / requestUpdate gating updated to consider streaming_thread_.inFlightApprox() so we don't declare "converged" while a worker read is still in flight, and the render loop stays alive until the worker queue is empty. Refactored loadChunkBytesAndUploadGpu into two helpers: - makeChunkRequest: builds the worker request from chunk metadata - applyStreamedChunk: pool.alloc + queueWriteBuffer + bind group Both the sync and async paths share applyStreamedChunk. Benchmark (big federation, --streaming): close camera: avg 24 fps p99 47 ms (was 27/326) default camera: avg 24 fps p99 46 ms (was 31/186) stream time: ~2 ms (was 8-12) cull is now the bottleneck (20 ms median) — task #17 (GPU compute cull) is the next frontier. Pixel-identical to non-streaming on basic.ifc on both paths. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
6f66d08bee |
wgpu cull: chunk-level frustum cull replaces BVH walk
cullModelCpuCompute previously had two paths: a flat linear scan over
all instances (default), or a BVH-stack walk (--bvh, gated off because
it regressed on dense scenes — the BVH built per instance but its
interior-node AABBs spanned huge chunks of model so most subtrees
straddled the frustum and the walk overhead beat the rejection win).
With spatial chunk planning (commit
|
||
|
|
4d36174200 |
wgpu streaming: spatial chunk planning + coalesced multi-range reads
Chunks are now grouped by world-space centroid instead of mesh-id
range, so each chunk's AABB tightly bounds its geometry instead of
spanning the whole model. Distance-based eviction can finally
distinguish the near corner of a skyscraper from the far corner.
Algorithm:
1. Compute each mesh's centroid = mean of its instances' world AABB
centres.
2. Sort mesh indices lexicographically by (z, y, x) centroid. Stable
sort keeps mesh-id order as tiebreaker for instanced repeats.
3. Greedy-pack sorted meshes into chunks ≤ WGPU_CHUNK_VERTEX_BYTES_LIMIT.
4. Each Chunk stores its mesh_ids list; the per-mesh layout (chunk_local
base_vertex / ebo_first_u32) is computed by walking the list at plan
time.
Loader: chunk vertex/index bytes are no longer file-contiguous, so
streaming uses new multi-range read paths
(readSidecarVertexRanges / readSidecarIndexRanges). Each range list
is sorted by file offset and adjacent ranges coalesced with a 64 KB
gap tolerance — on the close-camera benchmark this brings the
per-chunk seek count back down to ~mesh-id-grouping levels, so the
spatial sort costs ~nothing on I/O while delivering tighter AABBs.
Non-streaming applyCachedModel mirrors the spatial plan but gathers
from in-memory data.vertices / data.indices via per-mesh
queueWriteBuffer calls at chunk-local offsets.
Chunk struct drops vertex_byte_offset and index_first_u32 (no longer
meaningful — each chunk is N scattered ranges). vertex_byte_size and
index_count stay as aggregates for pool sizing + eviction math.
Tuning: kept WGPU_CHUNK_VERTEX_BYTES_LIMIT at 128 MB. Tried 8 MB and
32 MB; both gave tighter AABBs but the scatter-gather I/O cost blew
up because the per-frame load count grows linearly as chunks shrink
(orbit shifts the working set faster across finer chunks). 128 MB +
coalescing is the empirical sweet spot pre-v14. Once sidecar v14
re-orders bytes on disk to match spatial chunks, we can drop the
limit to ~8 MB for sharp eviction without re-paying the seek cost.
Benchmarks (big federation, --streaming):
close camera: avg 36 fps median 53 (was 35/49) — parity
default camera: avg 33 fps median 47 (was 40/49) — small regression
likely from increased coalesce overhead on more-
scattered orbit traversals; will resolve with v14.
Pixel-identical to non-streaming on basic.ifc on both paths.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
c3a55d7f7b |
wgpu streaming: multi-pool growth, frustum-only residency, sorted convergence
Five interlocking fixes that take --streaming on the big federation scene from "5 fps + endless flicker + infinite cold-load" to a stable 35-49 fps with a converged working set. 1. Multi-sub-buffer WgpuBufferPool. Pool now grows lazily by adding sub-buffers of per_sub_buffer_capacity_ when alloc demand exceeds existing free runs. Each Slice carries (buffer, offset, size, sub_idx). On driver refusal of addSubBuffer, growth_disabled_ latches so subsequent allocs don't keep retrying and log-spamming. pool_can_fit consults can_grow() to know when growth could rescue a candidate vs when eviction is the only path. 2. Split cull / stream benchmark timers. The previous "cull[wall]" metric was actually cull + driveStreamingLoads, blaming the wrong subsystem (~170 ms of "cull" was synchronous disk I/O). 3. frustum_visible_count on Chunk, populated in cullModelCpuCompute right after the per-instance aabbInFrustum check. driveStreamingLoads now keys residency on this instead of total_visible_draws (which includes contribution + HiZ). HiZ visibility flips frame-to-frame as occluders shift; using it for residency caused chunks to be evicted then immediately re-loaded, every frame, even with a stationary camera — both the perf cliff and the visible flicker. 4. Distance-sorted candidates in driveStreamingLoads. Walk the non-resident frustum-visible chunks in distance order (closest first). With sorted processing, evict_farthest_than converges monotonically: each swap replaces a far resident with a closer candidate; once the next candidate is farther than every remaining resident, the loop exits. Without sorting the loader visited candidates in model/chunk-id order, swapping random chunks every frame without ever converging. 5. 10% eviction hysteresis (EVICT_DIST2_RATIO = 1.21). On scenes where many chunks are clustered at similar distance from the camera (e.g. several chunks all ~370 m away), naive "evict any resident strictly farther than candidate" triggers sub-meter swaps every frame, never resting. Requiring the victim to be 10% farther in linear distance kills these cycles while still allowing genuine "much closer" candidates to evict. Plus: latched bench_warm_done_ on the cold-load gate, with a 5-frames-of-zero-loads convergence test (default-camera big scene converges in 20 frames) and a 600-frame timeout fallback that prints exactly once. Measured on the test federation (111 sidecars, ~3 GB raw, 1 M instances) with the user's close-in camera: - avg 35 fps (was 5), median 49 fps (was 7) - cull 19 ms (now the bottleneck), stream 5-8 ms (was 172) - p99 184 ms — occasional big-chunk load on the render thread; background-thread I/O would smooth that out as a follow-up. With the default wide camera: - avg 40 fps, converges in 20 frames, residency grows naturally from 59 → 76 chunks as orbit shifts the frustum. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
502c29fbc2 |
wgpu: probed-size pool replaces per-chunk createBuffer
Drops the per-machine "guess the OOM ceiling" budget knob in favour of a single buffer pool whose capacity is *probed* at device-init time. The runtime answers the question: descend from min(maxBufferSize, 4 GB) through OOM error scopes, accept the largest size that allocates cleanly. On a desktop wgpu-native v29 box this lands at 2 GB; on browser-class platforms it'll land at 256 MB – 1 GB depending on the implementation. Same code path either way. Architecture: - WgpuBufferPool (new): single WGPUBuffer + free-list sub-allocator with adjacent-range coalescing and first-fit. 256 B alignment for storage-binding offsets. - Chunks now hold (pool_vertex_offset, pool_vertex_size) and (pool_index_offset, pool_index_size) instead of per-chunk WGPUBuffer handles. Load = pool.alloc + queueWriteBuffer. Unload = pool.free. - Bind groups bind pool_.buffer() at the chunk's specific (offset, size) for both the vertex and index storage bindings. - Eviction queries pool.largest_free_run_bytes() instead of a tracked budget; the two-phase LRU/distance evictor's policy is unchanged. What this fixes: - No more gpu-alloc-rs fragmentation OOM: one VkDeviceMemory block instead of N per-chunk blocks with rounding overhead. On the test dataset (~3 GB on disk, 562 k visible instances) the wgpu backend now runs through to render without OOM at any point. - No --streaming-vram-mb knob, no hardcoded budget constant, no per-machine calibration. The pool size adapts to whatever the runtime grants. Notes: - Error scope probing: wgpu-native v29 classifies "Not enough memory left" as WGPUErrorType_Validation, not OutOfMemory. We push both filters (nested) and treat either firing as probe failure. - The 4 GB probe cap is principled, not magic: above that, wgpu-native's advertised maxBufferSize is sometimes a sentinel (1 TB) that just forces wasteful halving steps. 4 GB is the largest buffer any realistic WebGPU implementation will grant a single allocation today. - Pool destroy()/release happens after model release in shutdown() so the underlying buffer outlives every bind group that references it. Follow-ups: spatial chunking (task #22) for finer eviction granularity; cull perf needs work at 100+ models / 1M+ instances (separate from streaming concerns). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
71e61dd8a5 |
wgpu streaming (5/4): per-chunk indices + LRU/distance eviction (stopgap)
Defers index buffers per-chunk (alongside vertex bytes) so streaming fully delivers on its "don't load until visible" contract — the previous per-model index buffer was upfront-loaded and tipped scenes >~1.5 GB into allocator OOM at frame 1. Adds residency tracking + a two-phase evictor: (1) drop LRU non-visible chunks first, (2) if everything resident is visible-this-frame, drop the farthest-from-eye chunk only when the candidate to load is closer. This gives monotonic convergence to "closest visible chunks fit the budget" instead of "first 4 win, rest never load." Default budget set to 1 GB — explicitly a stopgap, documented inline. The per-machine OOM ceiling on wgpu-native (caused by allocator fragmentation from one VkDeviceMemory per createBuffer call) cannot be solved by tuning this knob. The proper fix is a probed single-pool buffer with sub-allocation, tracked under task #16. Caveat: LOD1 indices are now force-disabled when chunking — per-chunk buffers only carry LOD0. Re-enabling needs LOD1 to participate in the chunk plan. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
5a3e5167df |
wgpu streaming (4/4): per-frame chunk-on-visible loader
The OOM fix for vertex storage. With --streaming, chunks now load on
demand:
- After cull determines which chunks have visible draws, driveStreamingLoads
walks non-resident chunks with total_visible_draws > 0 and brings up
to MAX_STREAMING_LOADS_PER_FRAME (currently 4) into residency.
- Each load: readSidecarVertexChunk → createBufferWithData →
buildChunkBindGroup → is_resident = true. Same frame's draw loop
picks up the newly-built bind_group and renders the chunk.
- If more non-resident-but-visible chunks remain, requestUpdate is
called so the load loop keeps running until the visible set is fully
resident.
Per-chunk bind group construction refactored out of buildModelBindGroup
into a buildChunkBindGroup(m, chunk_idx) helper so the streaming loader
can build one chunk at a time as it arrives.
4 chunks/frame × 60 fps = 240 chunks/sec ingestion. A 200-chunk scene
fully resides in ~1 second of motion. Off-screen chunks never become
resident, never pay vertex-storage VRAM — that's where most of the OOM
fix lands.
Verified on basic.ifc: pixel-identical to non-streaming. On the user's
real 111-model / 1M-instance scene: all metadata loads succeed (was
OOM before), then loader runs but **indices are still loaded upfront
(1.5 GB!) so OOM still hits when vertex chunks start adding on top.**
Per-chunk index deferral is the next commit.
Eager-no-evict policy (per the design conversation): chunks stay
resident once loaded. LRU eviction lands in a follow-up if a workload
proves it necessary.
This completes the 4-commit stage-1 series for task #16. Stage-2:
defer indices, deferred mesh/instance storage if needed, async worker
thread.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
f6d888d42b |
wgpu streaming (3/4): --streaming scaffold + applyCachedModelStreaming
Wires the metadata-only reader (commit 1) through a parallel streaming
load path. With --streaming on:
- loadSidecar routes through readSidecarMetadataOnly: reads header +
mesh dict + instance dict + georef + elements upfront. Skips
vertex bytes entirely.
- applyCachedModelStreaming computes the same chunk plan as the
non-streaming path, allocates the small per-chunk buffers
(visible_draws + prefix_sums + per_chunk_uniform), allocates the
model-shared mesh + instance + index buffers, but leaves each
chunk's vertex_storage NULL and is_resident=false.
- Stores streaming_file_path + vertex_section_offset on the model so
the per-frame loader can range-read chunks later.
- Computes per-chunk world AABB by walking instances → mesh → chunk;
used by both cull (chunk-level frustum reject, future) and the
streaming loader (proximity-prioritised fetch, future).
Index buffer is still loaded upfront in stage 1 (small relative to
vertex data: ~1/2 of vertex bytes on real scenes). Stage 2 may defer
it too if measurements suggest it's worth the extra plumbing.
Render + pick already gate on c.bind_group (null when non-resident),
so the existing guards correctly skip non-resident chunks without
further changes.
With this commit alone, --streaming mode shows an EMPTY scene (just
background colour) because no chunk ever becomes resident. Commit 4
adds the per-frame loader that triggers chunk load when cull marks
them visible — that's the commit where rendering kicks in and the
OOM fix actually lands.
Default behaviour (no --streaming): legacy synchronous full-load.
Pixel-identical to the prior commit on basic.ifc.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
d368ee449d |
wgpu streaming (2/4): per-chunk residency fields on WgpuModelGpuData
Foundation for streaming. Adds to each Chunk:
- is_resident (default true; streaming flips false initially)
- vertex_byte_offset / vertex_byte_size in the sidecar file
- aabb_min / aabb_max world-space chunk bounds (used by future cull
and streaming priority)
Plus on the model:
- streaming_file_path (non-empty = streaming path was used)
- streaming_vertex_section_offset (where the chunks live in the file)
All fields default to backward-compatible values: is_resident=true,
streaming_file_path empty. The existing non-streaming applyCachedModel
sets up a Chunk with is_resident=true (implicit) and ignores the
streaming fields, so no behaviour changes yet.
Commit 3/4 wires the metadata-only reader from (1/4) through a new
applyCachedModelStreaming path that flips is_resident=false initially;
commit 4/4 adds the per-frame loader that brings chunks resident on
demand. This commit is verified pixel-identical to the previous render
on basic.ifc.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
a06d920fc6 |
wgpu streaming (1/4): metadata-only sidecar reader
First foundational piece for task #16. WgpuStreamingLoader exposes: - readSidecarMetadataOnly(path): reads v13 header + mesh dict + instance dict + georef + elements + string table from disk. Skips the bulky vertex and index byte sections, recording their on-disk offsets so they can be range-read later (per-chunk, on demand). The file handle is closed before return. - readSidecarVertexChunk / readSidecarIndexChunk: open + fseek + fread for a byte range. Synchronous; intended to be called from a worker thread for true async streaming or the main thread for stage-1 on-demand load. No format change yet — operates on existing v13 sidecars. v14 with an explicit per-chunk TOC arrives in a follow-up; this layer abstracts the chunk boundaries so the upgrade stays internal. No integration with existing applyCachedModel — that's commit 3/4. Build verifies the API compiles and links into IfcViewerWgpu. Commits in this series: 1/4: metadata-only reader (THIS) 2/4: per-chunk residency state on WgpuModelGpuData 3/4: --streaming opt-in path through applyCachedModel 4/4: per-frame chunk-on-visible loader (the OOM fix) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
a1693259b8 |
wgpu backend: BVH cull (opt-in via --bvh, default off)
Stage 15 implementation lands but doesn't pay off as default-on. On a 562k-instance / 18-model scene with a centred camera, the BVH walk adds ~10 ms of cull cost without rejecting enough subtrees to compensate — every interior node's AABB straddles the frustum, so descents go all the way to leaves anyway. Linear scan beats it by that 10 ms. GL's BVH works better mainly because they do full cull (frustum + HiZ + contribution) at every node — their per-test cost is lower (likely SIMD-vectorised) and they get more subtree rejections. My current impl does frustum-only at interior nodes (HiZ there cost more than it saved on the smaller dataset). For now, gate the whole BVH walk behind --bvh, default off. The infrastructure (BvhAccel build at applyCachedModel, walk in cull, release) stays in place so it's a one-flag toggle to measure either side. Real default-on requires further tuning — see updated task #15. Measured on 562k-instance scene: --bvh on → 25.9ms total (cull 25.4ms) --bvh off → 15.4ms total (cull 14.5ms) ← default For comparison, GL on the same scene + camera: GL → 18.2ms total (cull 8.5ms wall, multi-threaded BVH) Net: wgpu beats GL by ~3ms total despite slower cull, because the GPU side (no edge-pass cost, async HiZ readback, lean main pipeline) gives back more than the cull deficit. Also added task #17 (GPU compute-shader cull) as the asymptotic answer — both backends hit CPU cull as the ceiling on ≥500k scenes; moving it to a compute shader drops it to sub-ms regardless. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
7dc13eb104 |
wgpu backend: chunk vertex storage to fit browser limits + settle frame
Two pieces: 1. Per-chunk vertex storage (stage 13) WebGPU mandates maxStorageBufferBindingSize ≥ 128 MB. Real BIM models routinely exceed that (one of yours is 139 MB vertex). Without chunking, every browser load would fail with "exceeds max_storage_buffer_binding_size". Strategy: each model's vertex data is split into ≤ 128 MB chunks at applyCachedModel time. Each chunk gets its own vertex_storage buffer, visible_draws / prefix_sums buffers, per_chunk_uniform, and bind group. Index buffer, instance storage, and mesh storage stay single-per-model (they fit well under the cap on every scene we've seen). Mesh-to-chunk assignment is bake-time-deterministic (walks meshes in order, opens a new chunk when adding the next would overflow). Cull buckets visible instances by their mesh's chunk; render issues one drawcall per non-empty chunk per model. WGSL is unchanged — the binary-search vertex pulling works identically per chunk because base_vertex is now CHUNK-LOCAL (the chunk's bind group binds its own vertex_storage). Single code path: chunking is ALWAYS on at 128 MB regardless of target. Cost on desktop is a handful of extra drawcalls per frame (1 per non-empty chunk; typical models = 1-3 chunks). Negligible. A mesh whose vertex range is itself > 128 MB can't fit in any chunk and would need splitting — typical IFC meshes are nowhere near that (hundreds of verts), and applyCachedModel warns loudly if one ever appears. --web-limits CLI flag requests the WebGPU mandatory floor limits (128 MB max storage binding, 256 MB max buffer) instead of the adapter's actual max. Used to verify chunking actually fits through browser constraints — turns "trust me, web will work" into a hard test. The 139 MB scene loads cleanly with --web-limits. 2. Settle frame after motion (bug fix) Reported regression: after orbiting, sub-pixel instances dropped by motion-mode contribution culling stayed missing after the camera stopped. Event-driven rendering means no frame is scheduled after mouse-up, so the cull never re-ran at the still threshold. Fix: track last_cull_was_motion_. If this frame used the motion threshold, requestUpdate() after present to schedule one settle frame. Next frame: camera_moved = false → still threshold → small instances reappear. Matches GL's last_cull_was_motion_ behaviour. Verified pixel-identical on basic.ifc; loads the user's dense scene successfully under --web-limits (chunks=2 on the 139 MB model, chunks=1 on the others). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
5810894eb8 |
ifcviewer (GL minimal): --screenshot for parity diff with wgpu
Closes the other half of task #10. The wgpu minimal already wrote PNGs via wgpuCommandEncoderCopyTextureToBuffer + mapAsync; the GL backend now has the equivalent via glReadPixels on the back buffer just before swapBuffers. - ViewportWindow::captureNextFrameToPng(path, quit_after=true) queues a one-shot capture. render() reads the default framebuffer at full pixel size (width * devicePixelRatio), flips bottom-up → top-down into a QImage::Format_RGBA8888, saves PNG, and optionally QCoreApplication::quit. Synchronous glReadPixels is fine here — pick is interactive and rare; not used per-frame. - ifcviewer-minimal --screenshot PATH wires through MinimalWindow just like --camera / --benchmark. Honoured after all loads complete (applyPendingBenchmark also drains pending_screenshot_). Lets a parity script do: IfcViewerMinimal foo.ifc --camera A,B,C,D,E,F --screenshot gl.png IfcViewerWgpuMinimal foo.ifcview --camera A,B,C,D,E,F --screenshot wgpu.png # then pixel-diff with whatever (ImageMagick, PIL, etc.) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
4dfe27e251 |
wgpu backend: per-element visibility + H/Shift+H/I hotkeys
Closes the interactive selection loop. After this commit you can: - LMB-click an object → highlight (selection) - Press H → hide all selected - Press Shift+H → show all (clear hidden set) - Press I → isolate selected (hide everything else) WgpuVisibilityState (new header) is a plain unordered_set<uint32_t> of hidden object_ids — mirrors src/ifcviewer/Visibility.h's shape but stays Qt-free for the ifcviewer-core extract later. cullModelCpuCompute consults visibility_.isHidden(inst.object_id) before the frustum test — hidden instances cost nothing on every axis (no draw, no depth contribution, no pick hit). The CPU vector is read concurrently by the parallel cull workers, which is safe because mutations only happen between renders (handlers requestUpdate after mutating; render reads). Hiding deselects (matches GL behaviour: H clears the now-invisible selection rather than leaving phantom selected-but-invisible ids). Stage 5's last piece — clip planes — is deferred. Adding the uniform array + WGSL discard is mechanical, but the section-tool UI that drives them isn't ported yet (minimal viewer has no way to place a clip plane), so it'd ship as empty plumbing. Will land alongside the section-tool port. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
54fa7d8379 |
wgpu backend: selection visualisation + global object_id rebase
Closes the loop on stage 4 (pick): clicking an object now highlights it
on screen. Plus the prerequisite plumbing for selection to behave
correctly across multi-sidecar loads.
Pieces:
1. WgpuSelectionState (new header)
CPU-side multi-set + active-id, mirroring the GL Selection.h shape
but pure stdlib (no Qt deps) so it can move into ifcviewer-core
later without dragging Qt across. clear/replace/add/remove/toggle
APIs + a fillFlagsArray helper that packs (selected, active) into
a u32 bitmap indexed by object_id.
2. selection_flags storage buffer + frame_bgl bump to 2 entries
Indexed by object_id, bit 0 = selected, bit 1 = active. Lives in
the frame bind group (group=0 binding=1) because object_ids are
globally unique — making it model-scoped would be the wrong cut.
ensureSelectionFlagsBuffer grows geometrically (64 → 128 → … u32)
as new models push next_object_id_ up, rebuilds the frame bind
group when it does.
3. Global object_id rebase in applyCachedModel
Each sidecar's local ids start from 1 and collide across files;
pick was previously ambiguous on multi-model loads. We now add
next_object_id_ as a base offset, rewrite InstanceCpu.object_id
(CPU mirror stays consistent) + InstanceGpu.object_id (what pick
reads back), and bump next_object_id_ by the model's max + 1.
4. WGSL main fragment reads sel_flags
Vertex shader passes inst.object_id through to fragment as
@interpolate(flat). Fragment reads sel_flags[object_id], mixes
(0.2, 0.6, 1.0) at 0.45 for in-selection and (0.4, 0.8, 1.0) at
0.40 on top for active. Same constants as the GL main shader.
5. Mouse → selection
LMB-click-without-drag pick result feeds the selection:
no modifier → replace
Shift → add
Ctrl → remove (active migrates to another id in the set)
miss + no modifier → clear
uploadSelectionFlagsIfDirty repacks + writes the GPU bitmap at
the top of the next render(); no upload on still frames.
Pick pipeline is unchanged — it already outputs the per-instance
object_id, and that's what the selection storage indexes.
Visibility + clip planes are pending follow-ups in stage 5 (mostly
small, share the same buffer-lifecycle pattern). Edge silhouette
(stage 9 partial), --screenshot diff harness (stage 10 partial),
ifcviewer-core extract (stage 12), and web chunking (stage 13) all
still pending.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
1bbcd10bc0 |
wgpu backend: pick pass with sync R32UInt readback
Stage 4 of the wgpu port. LMB-click-without-drag now resolves the
object_id under the cursor by running a dedicated pick render and
copying back the single texel at the click position.
- Pick pipeline reuses the existing pipeline_layout_ (same bindings
as main: frame uniform at group=0, per-model storages at group=1).
Different vs / fs entry points (vs_pick / fs_pick) in the main
WGSL module — the vertex pulling logic is duplicated for now but
the bind group layout match means no pipeline_layout rebuild and
pickObjectAt can reuse the current frame's already-uploaded
visible_draws + per-model bind groups.
- Pick FBO: surface-sized R32UInt color attachment + Depth32Float
depth, both single-sample (no MSAA — pick needs exact texel
access). CopySrc on the color so we can copyTextureToBuffer the
1×1 click region. Recreated on surface resize.
- pickObjectAt: encodes a one-shot pick pass + a single texel copy
into a 256-byte staging buffer, submits, mapAsync, sync-spins
processEvents until ready. Synchronous wait is fine here — pick
runs on click, not per-frame, so a sub-ms stall is invisible.
- Mouse integration: existing LMB drag-orbit preserved. A 3-pixel
threshold promotes drag (set nav_dragged_); release without
dragging triggers pickObjectAt at the release coords (logical
Qt → physical pixels via devicePixelRatio). object_id is logged;
selection state to consume the id arrives with stage 5.
Object_id 0 means miss (clear value); the pick attachment is cleared
to 0 before each pass and the fragment writes the instance's
object_id, so any non-zero result is a real hit on a drawn instance.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
f4243038bd |
wgpu backend: edge silhouette post-process (ported GL renderEdgePass)
First piece of stage 9 — the dark outlines BonsaiViewer / the GL backend
draw at depth discontinuities. Ported the GL renderEdgePass algorithm
verbatim, including the three things my earlier attempt missed:
1. Linearise depth to view-space metres before the Laplacian. Raw
[0,1] clip-z is heavily non-linear so a fixed-threshold edge
detector only caught near-camera silhouettes. Now reverses the
wgpu z-remap (z * 2 - 1 back to GL NDC) then standard reverse-
perspective to view-z.
2. Threshold scales with depth: t = EDGE_THRESHOLD * c. A 4 mm gap
between two surfaces reads the same whether it's 0.5 m or 50 m
away from the camera.
3. Multiplicative blend (Dst, Zero) with fragment output of
vec3(1 - edge). Strictly darkens, never brightens. Matches GL's
(GL_DST_COLOR, GL_ZERO) blend.
Constants EDGE_SCALE=6.0 / EDGE_THRESHOLD=0.004 are GL's tuned values.
Camera near/far hard-coded to 0.1 / 10000 (the viewport defaults);
they'll move to a small uniform when AppSettings ports across.
Pipeline state: depth-attachment-less, sample count 1, blend on, no
cull. Reuses depth_texture_'s TextureBinding usage that HiZ added.
Render pass loads the resolved main-pass colour (LoadOp_Load) and
writes back through the multiplicative blend; encoded between the main
pass and the HiZ resolve so HiZ uses the same MSAA depth that produced
the edges. edge_bind_group_ rebuilds lazily when depth_view_ is
replaced (mirrors the HiZ bind group lifecycle).
Perf cost on the 10-sidecar / 380k-instance benchmark: 0.1 ms (11.5 →
11.6 ms). Fullscreen depth-laplacian is essentially free on this GPU.
Remaining stage 9 work: HUD/labels/lines/points overlay primitives,
which need the QPainter-→-texture path. Lower visual priority than
edges; handled in a follow-up.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
51dc31a50b |
wgpu backend: 10× perf — megadraw, async HiZ, parallel cull, motion mode
Closes the perf gap to the GL backend on real BIM benchmarks. On a 10-
sidecar / 380k-instance corpus at a fixed --camera the wgpu binary went
from 110.6 ms to 11.6 ms (vs GL's 23 ms — half the frame time, but
note GL is doing extra work the wgpu backend hasn't ported yet; see
the caveats list at the bottom). Bundled because the pieces interlock
and shipping any of them without the others reintroduces the same wall.
1. Cross-mesh vertex pulling (single mega-draw per model)
The previous one-drawIndexed-per-(mesh × LOD-bucket) loop was costing
~13ms on a 27k-mesh scene. CPU now emits a flat visible_draws[]
(16 B per visible (mesh,lod,instance)) plus a prefix_sums[] table.
WGSL binary-searches prefix_sums by @builtin(vertex_index) to find
the entry, then manually fetches the mesh-local index from a
storage-bound indices[] and pulls the packed 12 B vertex. No
setIndexBuffer; the shader reads everything from storage. Bind
group grew from 4 to 7 entries (vertices, meshes, instances,
indices, visible_draws, prefix_sums, per-model uniform) — well
under WebGPU's mandatory 8 storage / 12 uniform floor.
2. Async HiZ readback via ping-pong staging buffers
Sync wait via wgpuInstanceProcessEvents was costing ~37 ms on a
real scene (GPU drain). Two staging slots now ping-pong: frame N
kicks a non-blocking mapAsync on slot K, frame N+1's first action
is one processEvents drain. Pyramid is 1-2 frames stale — matches
the "slightly-stale depth, fine" pattern the GL backend already
documents. encodeHizResolve returns -1 (skip) if both slots are
in flight; cull keeps using the most recent pyramid.
3. Cull reorder: contribution before HiZ
HiZ projection is ~10× more expensive than the contribution
check, yet most contribution-survivors would be HiZ-rejected
anyway on dense scenes. Computing projected_px first lets
contribution short-circuit ~80% of HiZ tests with no rejection-
quality loss. Saved ~34 ms on the dense bench.
4. Motion-mode contribution threshold
AppSettings::motionMinPixelRadius parity. While the camera is
changing (orbit/pan/zoom/--benchmark sweep), drop instances
below 10 px instead of 2 px. Halves visible_objects during
motion with no perceived quality loss.
5. Parallel cull (std::async across models)
Per-model cullModelCpu split into Compute (CPU-only, thread-safe)
+ Upload (main-thread wgpu queue writes). std::async fan-outs the
compute across models; main-thread joins and uploads. Wall-clock
cull on the 10-model corpus drops from ~17 ms single-threaded to
~9 ms across cores.
6. --no-hiz CLI flag + per-phase benchmark timings
Benchmark now also prints "per-frame avg ms: cull=X
hiz_readback=Y" so future regressions can be attributed without
guesswork. --no-hiz toggles the master switch from the CLI.
Honest caveats — wgpu is currently faster mostly because GL is doing
work we haven't ported yet:
- Edge silhouette pass (stage 9) will add ~3-5 ms back to wgpu.
- GL's HiZ uses the BVH so it rejects whole subtrees (1.7k vs
our 358 rejects on the same scene). BVH for HiZ is future work
(task #13 / a new task) — until then we draw more sub-pixel
geometry that's behind closer surfaces. Visually correct, perf
cost paid. Stage 4+5 are unaffected.
Verified pixel-identical on basic.ifc through every change. Real-scene
visual diff against GL pending the --screenshot flag on the GL minimal
(task #10's other half).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
406124ca3d |
wgpu backend: HiZ occlusion culling
Stage 7 of the wgpu port. Per-frame after the main render pass:
1. encodeHizResolve runs a depth-only render pass that samples the
MSAA depth texture (sample 0) and max-reduces it into a small
single-sample Depth32Float target (256 × ~h-aspect). Implemented
as a fullscreen-triangle WGSL pipeline; one nested loop per
output texel over its source rect. WebGPU has no built-in depth
resolve, so this combined resolve+downsample fragment shader is
the way.
2. copyTextureToBuffer writes the small resolved depth into a
CPU-mappable staging buffer (≈ 160 KB at 256×160).
3. readbackAndBuildHizPyramid maps the staging buffer (sync via
wgpuInstanceProcessEvents — small enough that the stall is
well under a millisecond), strips per-row padding, and CPU
max-reduces a full mip pyramid (level 0 → 1×1). Stores the VP
used so the next frame can project AABBs into the same space.
Next frame, cullModelCpu calls aabbOccludedByHiz after the frustum
test: projects all 8 AABB corners through hiz_vp_, computes the
screen-space AABB and the nearest projected z, picks the mip level
where the AABB covers ≤ 2 texels per axis, samples that level's 2×2
window, and culls iff min_z > max_pyramid_depth in [0,1] z.
Plumbing changes:
- depth_texture_ gains TextureBinding usage so the resolve shader
can read it.
- hiz_enabled_ master switch defaults true; mirrors IFC_NO_HIZ in
the GL backend. Disabling skips encode + readback entirely.
- Bench output's "hiz_rej N" field now reflects actual rejections.
Verified: basic.ifc (3 instances, no occluders) renders pixel-
identical to pre-HiZ — proves the test rejects nothing it shouldn't.
Real rejection counts need a dense scene; this should drop visible-
objects count noticeably on real BIM benchmarks where back-of-room
walls hide each other.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
6ce2b564c1 |
wgpu backend: per-instance contribution culling in cullModelCpu
Quick win before the proper HiZ stage. Adds a min_pixel_radius threshold (defaults 2.0 to match AppSettings::minPixelRadius() in GL): instances whose projected bounding-sphere radius falls below it are dropped from the per-mesh buckets entirely. The projected_px math (radius_world * focal_px / view_z) is now computed once per instance and shared with the LOD pick that uses the same number. Saves one square root per instance per frame on dense scenes vs the previous code path that only computed it inside the LOD branch. Expected impact on real BIM benchmarks: visible-objects count drops by roughly 10×, matching the GL backend's number. Without this fix, wgpu was drawing every frustum-surviving sub-pixel instance — most of the work and most of the geometry the GL backend wasn't even submitting. Motion-mode threshold bump (10.0 in GL during camera drag) lands later when mouse-driven motion tracking is wired up. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
7893135790 |
wgpu backend: --camera flag + request adapter's max buffer limits
Two pieces that block proper side-by-side parity with the GL minimal: 1. --camera tx,ty,tz,dist,yaw,pitch. Same format string as the GL minimal so a pasted camera arg lands the same view on both backends. setCamera() also flips initial_view_applied_ = true so the auto- viewAll-on-first-load doesn't snap away from the script-set position when the model finishes uploading. 2. Real BIM models exceed the conservative WebGPU defaults at device create time. A 114k-instance / 19M-index sidecar's vertex storage is 139 MB, which trips wgpu's default 128 MB max_storage_buffer_binding_ size and bind-group creation fails. Now wgpuAdapterGetLimits is called first and the device is requested at the adapter's full ceiling — every desktop driver supports multi-GB. Trade-off worth flagging: web parity will fail here because browsers cap at the defaults. The eventual fix is to split a model's vertex/ instance storage into ≤128 MB chunks with a small per-frame routing table, which is a real chunk of work. For now this unblocks all the native benchmarking the user is actually doing. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
f9a9273ecc |
wgpu backend: match GL orbit + viewAll math so pivots align
Reported regression: --benchmark on the same sidecar visibly rotated
around a different point in the wgpu binary than in IfcViewerMinimal.
Root cause was two camera-convention drifts:
1. orbitEye placed the camera at (sin yaw, -cos yaw) from target;
the GL backend uses (cos yaw, sin yaw). Same target, but the
camera faces a different side of the model at yaw=0, which made
the orbit feel like it pivoted around a different point even
though the actual world-space target was the same. Now exactly
matches GL ViewportWindow::updateCamera:
eye.x = target.x + dist * cos(pitch) * cos(yaw)
eye.y = target.y + dist * cos(pitch) * sin(yaw)
eye.z = target.z + dist * sin(pitch)
2. viewAll's distance was an ad-hoc 0.6 * diag / tan(half_fov);
GL uses frameAabb(mn, mx, 1.10): tan_half = tan(fov/2),
min_aspect = min(aspect, 1), distance = (radius / (tan_half *
min_aspect)) * 1.10. Aspect-aware so portrait windows pull back
enough that the bounding sphere still fits on the tighter axis.
Now ported verbatim.
Also logs the computed target + distance on viewAll so a follow-up
side-by-side run prints both backends' framings and any remaining
discrepancy is easy to spot.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
244a145255 |
wgpu backend: per-instance LOD0/LOD1 pick in the cull
Stage 8 of the wgpu port. cullModelCpu now buckets each visible instance
by (mesh_id, lod) instead of (mesh_id), and emits one MeshDraw record
per non-empty bucket. LOD pick projects the instance's world-space
bounding sphere to pixels via
projected_px = world_radius * focal_px / view_z
where focal_px = viewport_h / (2 * tan(fov_y/2)) and view_z is the
forward·(center-eye) depth. When projected_px < lod1_pixel_threshold_
AND the mesh has a baked LOD1 slice (MeshInfo.lod1_index_count > 0),
the instance draws the LOD1 index range instead of LOD0; baseVertex
and the vertex storage are shared between LODs.
mesh_draws can now grow to up to 2 × meshes.size() per frame (LOD0 + LOD1
slice per mesh). The visible_buffer layout per mesh becomes
[LOD0 instances | LOD1 instances] contiguous, with each MeshDraw
referencing its own firstInstance offset.
lod1_pixel_threshold_ defaults to 30 (mirrors AppSettings::
lod1PixelThreshold() in the GL backend); set to 0 to disable LOD1
entirely (always LOD0). AppSettings port lands in a later commit.
Verified: basic.ifc (3 tiny instances, no LOD1 baked by meshoptimizer
since each mesh is well under the 500-tri threshold) renders pixel-
identical to pre-stage-8 — proves the all-LOD0 path is preserved.
Real LOD switching needs a sidecar where buildLods produced LOD1 slices.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
4596f2e584 |
wgpu backend: lighting parity, MSAA, cavity shading, fix sRGB output
Closes the visible gap to BonsaiViewer down to just the post-process edge silhouette pass (still pending in task #9). Four changes bundled because together they bring up the parity story: - WGSL fragment now applies cavity = clamp(length(fwidth(n))*1.5, 0, 0.35) and multiplies by (1 - cavity). Matches GL shader. - Lighting constants switched to GL's exact values: key (0.3, 0.5, 0.8), fill (-0.3, -0.5, 0.8), sky tint (0.55, 0.60, 0.70), ground tint (0.35, 0.32, 0.28). My initial guesses were close but not identical; matching them means side-by-side diffs only flag actual pipeline differences, not lighting tweaks. - 4× MSAA: render pass writes into a MULTISAMPLE color attachment (surface_format_-matched), resolves into the surface texture for present. Depth is also 4 samples. Pipeline.multisample.count = 4. ensureMsaaColorTexture / releaseMsaaColorTexture mirror the depth- texture lifecycle. Matches GL minimal's QSurfaceFormat::setSamples(4). - sRGB output fix. wgpu-native's Vulkan swap chain on X11 treats BGRA8Unorm as sRGB-output (applies linear→sRGB encoding on shader writes), even though caps.formats[0] reports plain Unorm. The GL backend writes to a non-sRGB framebuffer with no such conversion, so a clearValue of (0.125, 0.137, 0.161) lands as bytes (32, 35, 41) on GL but (99, 104, 112) on wgpu — ~3× brighter. Pre-decoding via srgbToLinear on (a) the clearValue in C++ and (b) the final fragment colour in WGSL makes wgpu's implicit encode round-trip, so the final bytes match GL. Verified via screenshot pixel sample: #202329 background reads as exactly (32, 35, 41). Remaining visible gap to BonsaiViewer is the dark-line edge silhouettes (renderEdgePass in GL, depth laplacian → outline). That belongs with the overlay / post-process work in task #9. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
a95437dd64 |
wgpu backend: match GL pitch sign so drag-down tilts the camera up
Drag-down was decreasing pitch (camera diving), opposite to the GL viewport's convention where drag-down increases pitch so the top of the object rotates toward the viewer. Yaw direction was already correct. Matches the existing user muscle memory from IfcViewerMinimal. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
de44de26f8 |
wgpu backend: clearer sidecar-load diagnostics + tilde expansion
The single "(file missing, wrong magic, or schema mismatch)" message
was making triage harder than necessary. loadSidecar now expands a
leading ~/ (shells skip it inside double quotes, which trips up paste-
from-launcher), and on failure peeks the file's header itself to
report exactly which check failed:
- "Sidecar not found" — file doesn't exist
- "Sidecar unreadable" — exists but open failed
- "Sidecar truncated" — <12 bytes
- "Sidecar magic mismatch" — wrong magic, reports got vs expected
- "Sidecar schema mismatch" — wrong version, reports both numbers
and suggests re-baking
- "Sidecar endianness mismatch" — cross-platform load attempt
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
819196b3ce |
wgpu backend: --benchmark N parity with the GL minimal
Stage 11 of the wgpu port. WgpuViewportWindow gains setBenchmarkFrames(N); the minimal driver wires it to a --benchmark N flag. Renders N frames after a 5-frame warmup, yaw-sweeping the camera at 0.5°/frame, captures per-frame wall time with QElapsedTimer (cull + encode + present), and prints avg/median/p1/p99 + last-frame stats in the same line format as IfcViewerMinimal so a script can diff them line for line. Per-frame stats (visible_objects, visible_triangles, sub_draws) are now summed in render() from m.mesh_draws. hiz_rej reports 0 until stage 7 adds HiZ occlusion. Verified on basic.ifc (3 instances): wgpu 11.68 ms avg vs GL 11.75 ms avg — same scene, same camera sweep, same window size. Noise-level delta as expected on a tiny scene; the interesting comparison is on real BIM corpora once you bake them to v13 sidecars. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
ddef8c65b5 |
wgpu backend: orbit/pan/zoom mouse navigation
LMB drag → orbit (yaw/pitch, pitch clamped to ±89.9° to avoid gimbal flip at the poles). MMB drag → pan in the camera's screen-space plane, world-units-per-pixel sized against the view frustum at the pivot depth so panning feels constant regardless of zoom. Wheel → zoom (12% per notch, sign matches "wheel up = closer"). LMB is bound to orbit because selection isn't wired yet; will rebind to selection + nav preset once AppSettings ports over. Pure addition to WgpuViewportWindow — overrides four QWindow event handlers, no changes to render or cull paths. Lets you actually fly around a loaded sidecar without a screenshot loop. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
61726e00a4 |
wgpu backend: CPU frustum cull + per-mesh draw compaction
Stage 6 of the wgpu port. Replaces the one-draw-per-(mesh, instance) loop
with a CPU cull pass that survives one drawIndexed per non-empty mesh
with packed instanceCount.
Adds to WgpuModelGpuData:
- visible_buffer: u32[] storage SSBO, pre-sized to instance_count at
applyCachedModel so the bind group reference never invalidates.
Re-uploaded each frame via wgpuQueueWriteBuffer.
- mesh_draws: per-mesh schedule (first_instance, instance_count,
first_index, base_vertex, index_count). instance_count==0 means the
mesh contributed nothing this frame and the draw is elided entirely.
cullModelCpu per-frame:
- Extract 6 frustum planes from the same VP we write into the uniform.
WebGPU clip-space z is [0, 1], so near plane = matrix row 2 (not
row 3 + row 2 as in GL); rest of the derivation is standard.
- Per-instance AABB-vs-frustum test using the p-vertex shortcut
(cheapest correct early-out for AABBs).
- Bucket survivors by mesh_id; flatten into a contiguous u32 list;
upload via wgpuQueueWriteBuffer. Per-mesh slice is [first_instance,
first_instance + instance_count).
WGSL adds @group(1) @binding(3) var<storage, read> visible: array<u32>
and an extra indirection: instance_idx = visible[iid]; the rest of the
shader is unchanged. firstInstance on each drawIndexed offsets into
visible[], so each mesh reads its own slice.
Verified two ways:
1. basic.ifc (3 instances, all on-screen) renders pixel-identically
to pre-stage-6 — proves cull keeps everything it should.
2. basic.ifc + a synthetic instance placed at (100, 100, 100) is
culled cleanly: only the cube renders, the far quad is rejected
by the frustum test. Proves cull actually rejects out-of-frustum
geometry rather than passing everything through.
Contribution culling, HiZ, and LOD selection arrive in stages 7 and 8;
they all hook into the same cullModelCpu seam.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
75b9963136 |
wgpu backend: --screenshot capability for visual verification
Pulls the capture half of task #10 forward so we stop flying blind from stage 3 onward. WgpuViewportWindow gains captureNextFrameToPng(path); the minimal driver wires it to a --screenshot PATH flag that renders one frame, copies the surface texture back to host memory, writes a PNG via QImage, and quits. CopySrc is added to the surface configuration usage so the surface texture can be the copy source. The texel-to-buffer copy honours WebGPU's 256-byte bytes-per-row alignment by padding rows and stripping the padding when assembling the QImage. Surface format 28 (BGRA8Unorm) is byte-swapped to RGBA on the way into QImage::Format_RGBA8888; RGBA8 surface formats are memcpy'd straight through. Verified end-to-end on /tmp/basic.ifcview: 3 cube meshes/instances render with depth, back-face cull, and the hemisphere-ambient + key+fill lighting model — top face reads sky (bright), front faces read mid-tone, exactly as the WGSL shading intended. The pixel-diff half of task #10 (comparing against a GL baseline) lands later when the GL minimal binary gets an equivalent flag. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
bbf2bfde92 |
wgpu backend: vertex-pulling main render pass
Stage 3 of the wgpu port. Replaces the clear-only render loop with the
full main shading pass:
- WGSL port of the GL main shader. Vertex-pulling: the vertex storage
buffer is read as array<u32> in the shader, with pos/normal/color
decoded manually per vertex. baseVertex (set per draw to mesh's
vertex offset) folds into @builtin(vertex_index) automatically;
firstInstance carries the instance slot for @builtin(instance_index).
No vertex-input layout — vertex pulling means no IA bindings.
- Render pipeline bound to depth-32-float (write-on, less compare),
back-face cull, CCW front face. Pre-multiplies a [-1,1]→[0,1] z-remap
matrix onto Qt's projection so WebGPU's clip-z convention is met.
- Two bind groups: group=0 per-frame (uniform with view-proj + key/fill
light + hemisphere ambient), group=1 per-model (three read-only
storage buffers: vertices, mesh quant, instances).
- Depth texture is created lazily and recreated on surface resize.
- Orbit camera state on WgpuViewportWindow with viewAll() that frames
the union of all loaded models' world AABBs after the first load.
Mouse navigation lands later.
- Draw loop: one drawIndexed per (mesh, instance) pair per model. This
is correct but CPU-heavy on dense scenes; stage 6 introduces the cull
+ compacted visible list that lets multiple instances of one mesh
collapse to a single call, and the eventual GPU-driven cull (post
sunset of the GL backend) goes further.
Verified on /tmp/quad_v13.ifcview (1 mesh, 1 instance) and on a real v13
sidecar baked from basic.ifc via the GL minimal viewer (3 meshes,
3 instances, 864 B verts). No wgpu validation errors fire across pipeline
creation, depth attachment, bind groups, or the draw loop on either.
Visual confirmation deferred until --screenshot lands (task #10) which
is being pulled forward next so we don't keep flying blind.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
9daa5fe195 |
wgpu backend: load .ifcview sidecars onto GPU buffers
Stage 2 of the wgpu port. WgpuViewportWindow gains a queueLoadSidecar API (called from the minimal driver before init) and an applyCachedModel that runs after init: reads via SidecarCache::readSidecar, allocates four wgpu buffers per model (vertex storage, index, mesh-quant storage, instance storage), uploads via wgpuQueueWriteBuffer, retains a CPU mirror of the MeshInfo/InstanceCpu arrays for the cull and picking paths that arrive in later stages. MeshGpu (the per-mesh quantization basis) is derived from MeshInfo on the fly; InstanceGpu (transform + ids) is derived from InstanceCpu and uses the cached float transform — composing from placement_transformation against federation-stage matrices lands when stage 5 wires those. SidecarCache.cpp is compiled into IfcViewerWgpu directly: it's pure C++ with no Qt/OCCT/IFC-parse deps, so dragging in the IfcViewer static lib for one source file would be wasteful. This duplication goes away once src/ifcviewer-core/ is extracted (task #12). Verified on a synthesised v13 sidecar (4 verts, 6 indices, 1 mesh, 1 instance) and a multi-sidecar load that assigns successive model_ids. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
19a39a0413 |
Scaffold experimental wgpu viewer backend
Adds src/ifcviewer-wgpu/ and src/ifcviewer-wgpu-minimal/ behind a new BUILD_BONSAIVIEWER_WGPU option (default OFF), gated independently of BUILD_BONSAIVIEWER. Stage 1 brings up a Qt window with a wgpu-native v29 surface (X11) and clears to the background colour — no rendering beyond that yet. Mirrors the lifecycle of the GL ViewportWindow so subsequent stages (vertex-pulling renderer, pick, cull, HiZ, overlay) slot in without restructuring the host. wgpu-native is fetched as a pre-built binary release via FetchContent; its .so SONAME is patched in at configure time so dependents get a clean DT_NEEDED. The X11 native handle is obtained via the public QNativeInterface::QX11Application API; Wayland and macOS/Windows surface creation are stubbed with explicit "not wired yet" warnings. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
dd902bf0f8 |
Use fixed overlay text font
Use Qt's system fixed font for viewer overlay text instead of the generic monospace family, avoiding the Windows font-resolution delay seen during measurement overlays. Generated with the assistance of an AI coding tool. |
||
|
|
371aabfef6 |
Disable Autodesk connector UPX
Build the PyInstaller Autodesk connector without UPX compression. UPX-packed launchers are more likely to trigger enterprise Windows security scanning, and the connector is distributed as a fresh unsigned artifact for each build. Generated with the assistance of an AI coding tool. |
||
|
|
1c22fa0669 |
Use bound overlay uploads
Update the overlay renderer's dynamic VBO uploads to bind the buffer and use glBufferData/glBufferSubData instead of direct-state glNamedBufferData/glNamedBufferSubData. This avoids Windows/NVIDIA driver corruption seen with overlay axes, pick markers, HUD rects, and marquee rectangles while keeping the same overlay geometry and draw paths. Generated with the assistance of an AI coding tool. |
||
|
|
89cb551dbb |
Run Autodesk upload/download on a worker thread
The progress dialog was the only connector window not driven by a Tk event loop: the handler created it, then blocked inline in httpx I/O. On Windows CTkToplevel withdraws itself at construction and re-shows via a delayed after() callback, which never fires without a running loop, so the progress window stayed invisible for the whole transfer. Add run_with_progress(): the blocking work runs on a daemon thread while the main thread pumps the Tk loop and shows the dialog. Progress reports are coalesced and marshalled back to the UI thread via _ProgressBridge, and worker exceptions are re-raised on the main thread, preserving the JSON-RPC error path. All eight upload/download handlers converted. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
102ac551b3 |
Add VisibilityState/SelectionState tests; test real quantization helpers
test_instanced_geometry previously re-implemented vertex quantization inline, with a stale comment claiming the helpers still lived in ViewportWindow.cpp. They now live in VertexQuantization.h, so route the test through the real quantizeVertex/octEncodeNormal and add coverage for the degenerate-axis path, octahedral normal round-trip, the i8 normal error bound (~0.78 deg worst observed), and color passthrough. Add test_visibility and test_selection: Tier-1 coverage of the two per-object viewport state machines. Both are QObjects for their changed() signal but touch no GL on the construction/mutation path, so the tests exercise the pure CPU logic without a context. Suite goes from 39 to 61 cases. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
137a890256 |
Add test suite for the Bonsai Viewer Autodesk connector
Introduce pytest coverage for the previously untested connector — rpc, cache, settings, autodesk (auth + APS client) and connector handlers — 94 tests, runnable via the new `test` optional-dependency extra. To make HTTP, time and the OAuth redirect testable without a network or real sockets, add dependency-injection seams to autodesk.py: AuthSessionService and ApsClient accept an optional httpx transport; AuthSessionService accepts an injectable clock and callback_waiter; and _wait_for_callback is extracted to the module-level wait_for_oauth_callback. All seams default to the previous behaviour. Remove the APS_CLIENT_ID environment-variable override: the client id now comes solely from settings.json, collapsing settings.load_client_id and simplifying the settings dialog. CI: the build-bonsaiviewer-autodesk workflow gains a `test` job (Python 3.11 + 3.13) that gates the build matrix. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
de7520418b |
Build the Bonsai Viewer in CI with the Autodesk connector bundled
Compile the Bonsai Viewer as part of the Linux and Windows binary builds, and ship the Autodesk connector alongside the viewer executable. Qt6 dependencies: - The viewer links Qt6::Svg for runtime icon tinting. Svg is a separate base-Qt archive, so aqt now installs "qtbase qtsvg" (plus icu on Linux) rather than qtbase alone, on both Linux and Windows. - Qt6::CorePrivate is exposed differently across Qt versions: Qt 6.8 ships the target inside Qt6Core, while Qt 6.10 provides it only as a separate CorePrivate config package. The viewer CMakeLists requests it via OPTIONAL_COMPONENTS so it resolves on both. - When cross-compiling Windows ARM64, windeployqt runs from the host x64 Qt, so qtsvg is installed into the host Qt as well. Windows build: - build-all-win.py passed -DBUILD_IFCVIEWER, a flag since renamed to BUILD_BONSAIVIEWER, so the Windows build compiled no viewer at all. It now passes -DBUILD_BONSAIVIEWER. - The Autodesk connector is bundled under connectors/ next to BonsaiViewer.exe in the packaged archive, mirroring the Linux builds. - The Windows workflow builds the connector (PyInstaller) before the main build so it is available to bundle. Connector bundling: - The Linux rocky workflows build the connector and bundle it into the BonsaiViewer archive; the Windows build now does the same. Generated with the assistance of an AI coding tool. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |