wgpu: probed-size pool replaces per-chunk createBuffer

Drops the per-machine "guess the OOM ceiling" budget knob in favour of
a single buffer pool whose capacity is *probed* at device-init time.
The runtime answers the question: descend from min(maxBufferSize, 4 GB)
through OOM error scopes, accept the largest size that allocates
cleanly. On a desktop wgpu-native v29 box this lands at 2 GB; on
browser-class platforms it'll land at 256 MB – 1 GB depending on the
implementation. Same code path either way.

Architecture:
- WgpuBufferPool (new): single WGPUBuffer + free-list sub-allocator
  with adjacent-range coalescing and first-fit. 256 B alignment for
  storage-binding offsets.
- Chunks now hold (pool_vertex_offset, pool_vertex_size) and
  (pool_index_offset, pool_index_size) instead of per-chunk WGPUBuffer
  handles. Load = pool.alloc + queueWriteBuffer. Unload = pool.free.
- Bind groups bind pool_.buffer() at the chunk's specific (offset, size)
  for both the vertex and index storage bindings.
- Eviction queries pool.largest_free_run_bytes() instead of a tracked
  budget; the two-phase LRU/distance evictor's policy is unchanged.

What this fixes:
- No more gpu-alloc-rs fragmentation OOM: one VkDeviceMemory block
  instead of N per-chunk blocks with rounding overhead. On the test
  dataset (~3 GB on disk, 562 k visible instances) the wgpu backend
  now runs through to render without OOM at any point.
- No --streaming-vram-mb knob, no hardcoded budget constant, no
  per-machine calibration. The pool size adapts to whatever the
  runtime grants.

Notes:
- Error scope probing: wgpu-native v29 classifies "Not enough memory
  left" as WGPUErrorType_Validation, not OutOfMemory. We push both
  filters (nested) and treat either firing as probe failure.
- The 4 GB probe cap is principled, not magic: above that, wgpu-native's
  advertised maxBufferSize is sometimes a sentinel (1 TB) that just
  forces wasteful halving steps. 4 GB is the largest buffer any
  realistic WebGPU implementation will grant a single allocation today.
- Pool destroy()/release happens after model release in shutdown() so
  the underlying buffer outlives every bind group that references it.

Follow-ups: spatial chunking (task #22) for finer eviction granularity;
cull perf needs work at 100+ models / 1M+ instances (separate from
streaming concerns).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
Dion Moult
2026-05-28 12:21:52 +10:00
parent 71e61dd8a5
commit 502c29fbc2
6 changed files with 503 additions and 148 deletions
+27 -17
View File
@@ -63,24 +63,31 @@ struct WgpuModelGpuData {
};
static_assert(sizeof(VisibleDrawGpu) == 16, "VisibleDrawGpu must be 16 bytes");
// Per-chunk state. Each chunk owns its vertex_storage (sized ≤ 128 MB)
// plus a small set of per-frame buffers (visible_draws, prefix_sums,
// uniform) and a bind group that pulls in the chunk's vertex_storage
// alongside the model-shared index/mesh/instance buffers. Rendering
// issues one drawcall per non-empty chunk.
// Per-chunk state. Each chunk references a vertex range and an
// index range inside WgpuViewportWindow::pool_, plus a small set of
// per-frame buffers (visible_draws, prefix_sums, uniform) and a bind
// group that binds the pool ranges alongside the model-shared
// mesh/instance storage. Rendering issues one drawcall per non-empty
// chunk.
//
// Streaming (task #16): a chunk may be marked is_resident=false; its
// vertex_storage + bind_group are then null until the streaming loader
// brings it in. Other per-chunk buffers (visible_draws etc.) stay
// allocated regardless because cull still needs them. Non-streaming
// path always sets is_resident=true so existing code is unchanged.
// pool ranges (pool_*_size == 0) and bind_group are then unclaimed
// until the streaming loader brings it in. Other per-chunk buffers
// (visible_draws etc.) stay allocated regardless because cull still
// needs them. Non-streaming path always sets is_resident=true and
// populates pool ranges at applyCachedModel time.
struct Chunk {
WGPUBuffer vertex_storage = nullptr;
// Per-chunk index buffer (was previously per-model). Each chunk
// covers a contiguous range of mesh ids, so its indices form a
// contiguous slice of the model's overall index data. Storing
// per-chunk lets streaming defer index loading alongside vertices.
WGPUBuffer index_buffer = nullptr;
// Pool-allocated vertex bytes. When resident, pool_vertex_size > 0
// and the range [pool_vertex_offset, pool_vertex_offset + pool_vertex_size)
// in WgpuViewportWindow::pool_ holds this chunk's vertex_storage.
// When non-resident, both are 0.
uint64_t pool_vertex_offset = 0;
uint64_t pool_vertex_size = 0;
// Pool-allocated index bytes. Same lifetime as the vertex range —
// either both resident or both freed.
uint64_t pool_index_offset = 0;
uint64_t pool_index_size = 0;
WGPUBuffer visible_draws_buffer = nullptr;
WGPUBuffer prefix_sums_buffer = nullptr;
WGPUBuffer per_chunk_uniform = nullptr;
@@ -183,8 +190,11 @@ struct WgpuModelGpuData {
bool hidden = false;
};
// Release every wgpu handle in `m` and clear its size mirrors. Safe to call
class WgpuBufferPool;
// Release every wgpu handle in `m` (including per-chunk and per-model pool
// ranges via `pool.free()`) and clear its size mirrors. Safe to call
// repeatedly; idempotent on already-released entries.
void releaseWgpuModelGpuData(WgpuModelGpuData& m);
void releaseWgpuModelGpuData(WgpuModelGpuData& m, WgpuBufferPool& pool);
#endif // WGPUMODELGPUDATA_H