Closes the visible gap to BonsaiViewer down to just the post-process
edge silhouette pass (still pending in task #9). Four changes bundled
because together they bring up the parity story:
- WGSL fragment now applies cavity = clamp(length(fwidth(n))*1.5,
0, 0.35) and multiplies by (1 - cavity). Matches GL shader.
- Lighting constants switched to GL's exact values: key (0.3, 0.5,
0.8), fill (-0.3, -0.5, 0.8), sky tint (0.55, 0.60, 0.70), ground
tint (0.35, 0.32, 0.28). My initial guesses were close but not
identical; matching them means side-by-side diffs only flag actual
pipeline differences, not lighting tweaks.
- 4× MSAA: render pass writes into a MULTISAMPLE color attachment
(surface_format_-matched), resolves into the surface texture for
present. Depth is also 4 samples. Pipeline.multisample.count = 4.
ensureMsaaColorTexture / releaseMsaaColorTexture mirror the depth-
texture lifecycle. Matches GL minimal's QSurfaceFormat::setSamples(4).
- sRGB output fix. wgpu-native's Vulkan swap chain on X11 treats
BGRA8Unorm as sRGB-output (applies linear→sRGB encoding on shader
writes), even though caps.formats[0] reports plain Unorm. The GL
backend writes to a non-sRGB framebuffer with no such conversion,
so a clearValue of (0.125, 0.137, 0.161) lands as bytes (32, 35,
41) on GL but (99, 104, 112) on wgpu — ~3× brighter. Pre-decoding
via srgbToLinear on (a) the clearValue in C++ and (b) the final
fragment colour in WGSL makes wgpu's implicit encode round-trip,
so the final bytes match GL. Verified via screenshot pixel sample:
#202329 background reads as exactly (32, 35, 41).
Remaining visible gap to BonsaiViewer is the dark-line edge silhouettes
(renderEdgePass in GL, depth laplacian → outline). That belongs with
the overlay / post-process work in task #9.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Drag-down was decreasing pitch (camera diving), opposite to the GL
viewport's convention where drag-down increases pitch so the top of
the object rotates toward the viewer. Yaw direction was already
correct. Matches the existing user muscle memory from IfcViewerMinimal.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The single "(file missing, wrong magic, or schema mismatch)" message
was making triage harder than necessary. loadSidecar now expands a
leading ~/ (shells skip it inside double quotes, which trips up paste-
from-launcher), and on failure peeks the file's header itself to
report exactly which check failed:
- "Sidecar not found" — file doesn't exist
- "Sidecar unreadable" — exists but open failed
- "Sidecar truncated" — <12 bytes
- "Sidecar magic mismatch" — wrong magic, reports got vs expected
- "Sidecar schema mismatch" — wrong version, reports both numbers
and suggests re-baking
- "Sidecar endianness mismatch" — cross-platform load attempt
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Stage 11 of the wgpu port. WgpuViewportWindow gains setBenchmarkFrames(N);
the minimal driver wires it to a --benchmark N flag. Renders N frames
after a 5-frame warmup, yaw-sweeping the camera at 0.5°/frame, captures
per-frame wall time with QElapsedTimer (cull + encode + present), and
prints avg/median/p1/p99 + last-frame stats in the same line format as
IfcViewerMinimal so a script can diff them line for line.
Per-frame stats (visible_objects, visible_triangles, sub_draws) are now
summed in render() from m.mesh_draws. hiz_rej reports 0 until stage 7
adds HiZ occlusion.
Verified on basic.ifc (3 instances): wgpu 11.68 ms avg vs GL 11.75 ms
avg — same scene, same camera sweep, same window size. Noise-level
delta as expected on a tiny scene; the interesting comparison is on
real BIM corpora once you bake them to v13 sidecars.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
LMB drag → orbit (yaw/pitch, pitch clamped to ±89.9° to avoid gimbal
flip at the poles). MMB drag → pan in the camera's screen-space plane,
world-units-per-pixel sized against the view frustum at the pivot depth
so panning feels constant regardless of zoom. Wheel → zoom (12% per
notch, sign matches "wheel up = closer"). LMB is bound to orbit because
selection isn't wired yet; will rebind to selection + nav preset once
AppSettings ports over.
Pure addition to WgpuViewportWindow — overrides four QWindow event
handlers, no changes to render or cull paths. Lets you actually fly
around a loaded sidecar without a screenshot loop.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Stage 6 of the wgpu port. Replaces the one-draw-per-(mesh, instance) loop
with a CPU cull pass that survives one drawIndexed per non-empty mesh
with packed instanceCount.
Adds to WgpuModelGpuData:
- visible_buffer: u32[] storage SSBO, pre-sized to instance_count at
applyCachedModel so the bind group reference never invalidates.
Re-uploaded each frame via wgpuQueueWriteBuffer.
- mesh_draws: per-mesh schedule (first_instance, instance_count,
first_index, base_vertex, index_count). instance_count==0 means the
mesh contributed nothing this frame and the draw is elided entirely.
cullModelCpu per-frame:
- Extract 6 frustum planes from the same VP we write into the uniform.
WebGPU clip-space z is [0, 1], so near plane = matrix row 2 (not
row 3 + row 2 as in GL); rest of the derivation is standard.
- Per-instance AABB-vs-frustum test using the p-vertex shortcut
(cheapest correct early-out for AABBs).
- Bucket survivors by mesh_id; flatten into a contiguous u32 list;
upload via wgpuQueueWriteBuffer. Per-mesh slice is [first_instance,
first_instance + instance_count).
WGSL adds @group(1) @binding(3) var<storage, read> visible: array<u32>
and an extra indirection: instance_idx = visible[iid]; the rest of the
shader is unchanged. firstInstance on each drawIndexed offsets into
visible[], so each mesh reads its own slice.
Verified two ways:
1. basic.ifc (3 instances, all on-screen) renders pixel-identically
to pre-stage-6 — proves cull keeps everything it should.
2. basic.ifc + a synthetic instance placed at (100, 100, 100) is
culled cleanly: only the cube renders, the far quad is rejected
by the frustum test. Proves cull actually rejects out-of-frustum
geometry rather than passing everything through.
Contribution culling, HiZ, and LOD selection arrive in stages 7 and 8;
they all hook into the same cullModelCpu seam.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Pulls the capture half of task #10 forward so we stop flying blind from
stage 3 onward. WgpuViewportWindow gains captureNextFrameToPng(path);
the minimal driver wires it to a --screenshot PATH flag that renders
one frame, copies the surface texture back to host memory, writes a
PNG via QImage, and quits.
CopySrc is added to the surface configuration usage so the surface
texture can be the copy source. The texel-to-buffer copy honours
WebGPU's 256-byte bytes-per-row alignment by padding rows and stripping
the padding when assembling the QImage. Surface format 28 (BGRA8Unorm)
is byte-swapped to RGBA on the way into QImage::Format_RGBA8888;
RGBA8 surface formats are memcpy'd straight through.
Verified end-to-end on /tmp/basic.ifcview: 3 cube meshes/instances
render with depth, back-face cull, and the hemisphere-ambient + key+fill
lighting model — top face reads sky (bright), front faces read mid-tone,
exactly as the WGSL shading intended. The pixel-diff half of task #10
(comparing against a GL baseline) lands later when the GL minimal binary
gets an equivalent flag.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Stage 3 of the wgpu port. Replaces the clear-only render loop with the
full main shading pass:
- WGSL port of the GL main shader. Vertex-pulling: the vertex storage
buffer is read as array<u32> in the shader, with pos/normal/color
decoded manually per vertex. baseVertex (set per draw to mesh's
vertex offset) folds into @builtin(vertex_index) automatically;
firstInstance carries the instance slot for @builtin(instance_index).
No vertex-input layout — vertex pulling means no IA bindings.
- Render pipeline bound to depth-32-float (write-on, less compare),
back-face cull, CCW front face. Pre-multiplies a [-1,1]→[0,1] z-remap
matrix onto Qt's projection so WebGPU's clip-z convention is met.
- Two bind groups: group=0 per-frame (uniform with view-proj + key/fill
light + hemisphere ambient), group=1 per-model (three read-only
storage buffers: vertices, mesh quant, instances).
- Depth texture is created lazily and recreated on surface resize.
- Orbit camera state on WgpuViewportWindow with viewAll() that frames
the union of all loaded models' world AABBs after the first load.
Mouse navigation lands later.
- Draw loop: one drawIndexed per (mesh, instance) pair per model. This
is correct but CPU-heavy on dense scenes; stage 6 introduces the cull
+ compacted visible list that lets multiple instances of one mesh
collapse to a single call, and the eventual GPU-driven cull (post
sunset of the GL backend) goes further.
Verified on /tmp/quad_v13.ifcview (1 mesh, 1 instance) and on a real v13
sidecar baked from basic.ifc via the GL minimal viewer (3 meshes,
3 instances, 864 B verts). No wgpu validation errors fire across pipeline
creation, depth attachment, bind groups, or the draw loop on either.
Visual confirmation deferred until --screenshot lands (task #10) which
is being pulled forward next so we don't keep flying blind.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Stage 2 of the wgpu port. WgpuViewportWindow gains a queueLoadSidecar
API (called from the minimal driver before init) and an applyCachedModel
that runs after init: reads via SidecarCache::readSidecar, allocates
four wgpu buffers per model (vertex storage, index, mesh-quant storage,
instance storage), uploads via wgpuQueueWriteBuffer, retains a CPU
mirror of the MeshInfo/InstanceCpu arrays for the cull and picking
paths that arrive in later stages.
MeshGpu (the per-mesh quantization basis) is derived from MeshInfo on
the fly; InstanceGpu (transform + ids) is derived from InstanceCpu and
uses the cached float transform — composing from placement_transformation
against federation-stage matrices lands when stage 5 wires those.
SidecarCache.cpp is compiled into IfcViewerWgpu directly: it's pure
C++ with no Qt/OCCT/IFC-parse deps, so dragging in the IfcViewer
static lib for one source file would be wasteful. This duplication
goes away once src/ifcviewer-core/ is extracted (task #12).
Verified on a synthesised v13 sidecar (4 verts, 6 indices, 1 mesh,
1 instance) and a multi-sidecar load that assigns successive model_ids.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds src/ifcviewer-wgpu/ and src/ifcviewer-wgpu-minimal/ behind a new
BUILD_BONSAIVIEWER_WGPU option (default OFF), gated independently of
BUILD_BONSAIVIEWER. Stage 1 brings up a Qt window with a wgpu-native
v29 surface (X11) and clears to the background colour — no rendering
beyond that yet. Mirrors the lifecycle of the GL ViewportWindow so
subsequent stages (vertex-pulling renderer, pick, cull, HiZ, overlay)
slot in without restructuring the host.
wgpu-native is fetched as a pre-built binary release via FetchContent;
its .so SONAME is patched in at configure time so dependents get a
clean DT_NEEDED. The X11 native handle is obtained via the public
QNativeInterface::QX11Application API; Wayland and macOS/Windows
surface creation are stubbed with explicit "not wired yet" warnings.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>