wgpu backend: 10× perf — megadraw, async HiZ, parallel cull, motion mode

Closes the perf gap to the GL backend on real BIM benchmarks. On a 10-
sidecar / 380k-instance corpus at a fixed --camera the wgpu binary went
from 110.6 ms to 11.6 ms (vs GL's 23 ms — half the frame time, but
note GL is doing extra work the wgpu backend hasn't ported yet; see
the caveats list at the bottom). Bundled because the pieces interlock
and shipping any of them without the others reintroduces the same wall.

1. Cross-mesh vertex pulling (single mega-draw per model)
   The previous one-drawIndexed-per-(mesh × LOD-bucket) loop was costing
   ~13ms on a 27k-mesh scene. CPU now emits a flat visible_draws[]
   (16 B per visible (mesh,lod,instance)) plus a prefix_sums[] table.
   WGSL binary-searches prefix_sums by @builtin(vertex_index) to find
   the entry, then manually fetches the mesh-local index from a
   storage-bound indices[] and pulls the packed 12 B vertex. No
   setIndexBuffer; the shader reads everything from storage. Bind
   group grew from 4 to 7 entries (vertices, meshes, instances,
   indices, visible_draws, prefix_sums, per-model uniform) — well
   under WebGPU's mandatory 8 storage / 12 uniform floor.

2. Async HiZ readback via ping-pong staging buffers
   Sync wait via wgpuInstanceProcessEvents was costing ~37 ms on a
   real scene (GPU drain). Two staging slots now ping-pong: frame N
   kicks a non-blocking mapAsync on slot K, frame N+1's first action
   is one processEvents drain. Pyramid is 1-2 frames stale — matches
   the "slightly-stale depth, fine" pattern the GL backend already
   documents. encodeHizResolve returns -1 (skip) if both slots are
   in flight; cull keeps using the most recent pyramid.

3. Cull reorder: contribution before HiZ
   HiZ projection is ~10× more expensive than the contribution
   check, yet most contribution-survivors would be HiZ-rejected
   anyway on dense scenes. Computing projected_px first lets
   contribution short-circuit ~80% of HiZ tests with no rejection-
   quality loss. Saved ~34 ms on the dense bench.

4. Motion-mode contribution threshold
   AppSettings::motionMinPixelRadius parity. While the camera is
   changing (orbit/pan/zoom/--benchmark sweep), drop instances
   below 10 px instead of 2 px. Halves visible_objects during
   motion with no perceived quality loss.

5. Parallel cull (std::async across models)
   Per-model cullModelCpu split into Compute (CPU-only, thread-safe)
   + Upload (main-thread wgpu queue writes). std::async fan-outs the
   compute across models; main-thread joins and uploads. Wall-clock
   cull on the 10-model corpus drops from ~17 ms single-threaded to
   ~9 ms across cores.

6. --no-hiz CLI flag + per-phase benchmark timings
   Benchmark now also prints "per-frame avg ms: cull=X
   hiz_readback=Y" so future regressions can be attributed without
   guesswork. --no-hiz toggles the master switch from the CLI.

Honest caveats — wgpu is currently faster mostly because GL is doing
work we haven't ported yet:
  - Edge silhouette pass (stage 9) will add ~3-5 ms back to wgpu.
  - GL's HiZ uses the BVH so it rejects whole subtrees (1.7k vs
    our 358 rejects on the same scene). BVH for HiZ is future work
    (task #13 / a new task) — until then we draw more sub-pixel
    geometry that's behind closer surfaces. Visually correct, perf
    cost paid. Stage 4+5 are unaffected.

Verified pixel-identical on basic.ifc through every change. Real-scene
visual diff against GL pending the --screenshot flag on the GL minimal
(task #10's other half).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
Dion Moult
2026-05-27 19:46:13 +10:00
parent 406124ca3d
commit 51dc31a50b
4 changed files with 551 additions and 283 deletions
+36 -20
View File
@@ -51,31 +51,47 @@ struct WgpuModelGpuData {
// recompose from placement_transformation when stage matrices change.
WGPUBuffer instance_storage = nullptr;
// u32[] of visible instance indices, repacked per frame by the CPU cull
// into per-mesh contiguous slices. Sized for the worst case
// (instance_count entries) at applyCachedModel time so we never have to
// recreate it (and the bind group that references it) mid-frame.
WGPUBuffer visible_buffer = nullptr;
size_t visible_buffer_capacity = 0; // entries (u32 count), not bytes
// Cross-mesh vertex pulling: one entry per visible (mesh, lod, instance)
// tuple, written into a single flat buffer per frame. The vertex shader
// binary-searches `prefix_sums` to find which entry a given
// gl_VertexID belongs to, then fetches that entry's mesh slice via the
// offsets baked here. Pre-sized at applyCachedModel to a generous
// worst case so the bind group stays valid for the model's lifetime.
//
// std430 layout: 16 bytes per entry, naturally aligned.
struct alignas(16) VisibleDrawGpu {
uint32_t mesh_id; // -> meshes[] for quantisation basis
uint32_t instance_idx; // -> instances[] for transform + ids
uint32_t ebo_first_u32; // start of this entry's slice in indices[]
uint32_t base_vertex; // start of this mesh's slice in vertices[]
};
static_assert(sizeof(VisibleDrawGpu) == 16, "VisibleDrawGpu must be 16 bytes");
// Bind group binding the four storage buffers above (group=1 in the
// main pipeline). Built in applyCachedModel after the buffers exist.
WGPUBuffer visible_draws_buffer = nullptr;
size_t visible_draws_capacity = 0; // entries (not bytes)
// prefix_sums[i] = sum of index_counts for visible_draws[0..i-1].
// Sized to visible_draws_capacity + 1 so prefix_sums[N] = total verts.
WGPUBuffer prefix_sums_buffer = nullptr;
size_t prefix_sums_capacity = 0; // entries (u32 count)
// Per-model uniform: (draw_count, total_vertex_count, _pad, _pad).
// The vertex shader uses draw_count to bound the binary search; CPU
// uses total_vertex_count as the draw() call's vertex count.
WGPUBuffer per_model_uniform = nullptr;
// Bind group binding the storage + uniform buffers above (group=1 in
// the main pipeline). Built in applyCachedModel.
WGPUBindGroup bind_group = nullptr;
// One per mesh, populated per frame by cullAndCompact. instance_count==0
// means the mesh contributes nothing this frame and the draw is skipped.
struct MeshDraw {
uint32_t first_instance = 0; // offset into visible_buffer
uint32_t instance_count = 0;
uint32_t first_index = 0; // index buffer offset (in u32 indices)
int32_t base_vertex = 0; // added to every fetched vertex_index
uint32_t index_count = 0;
};
std::vector<MeshDraw> mesh_draws; // size = meshes.size() after first frame
// Per-frame, populated by cullModelCpu and consumed by render().
uint32_t total_visible_vertices = 0;
uint32_t total_visible_draws = 0;
// Scratch reused each frame so per-frame cull doesn't allocate. Sized
// to instance_count entries on first use; never shrunk.
std::vector<uint32_t> visible_flat_scratch;
// on first use; never shrunk.
std::vector<VisibleDrawGpu> visible_draws_scratch;
std::vector<uint32_t> prefix_sums_scratch;
// Size mirrors for stats / range checks.
size_t vertex_bytes = 0;