mirror of
https://github.com/IfcOpenShell/IfcOpenShell.git
synced 2026-09-24 18:20:06 +00:00
Rewrite README for instancing pipeline and refocus Phase 3
The previous README described a pre-instancing world (32-byte world-
coord vertices with per-vertex object_id, ObjectDrawInfo structs, EBO
reordering after BVH build, and a Phase 3 plan built around moving
draw submission to the GPU). Most of that is either gone or already
solved:
- Vertices are now 28 B local-coord; per-instance transforms live
in an SSBO read through a visible-index SSBO and gl_BaseInstanceARB.
- ObjectDrawInfo is replaced by MeshInfo + InstanceCpu + InstanceGpu.
- No EBO reorder on BVH build — the BVH is over instance AABBs and
the mesh/EBO layout is orthogonal.
- Draw-call submission is already one glMultiDrawElementsIndirect
per model; the old Phase 3 goal is met.
New content worth keeping:
- GPU instancing section documents the mesh/instance/visible/indirect
buffer contract the whole renderer hangs off of.
- Reflection-aware two-pass draw is documented (det<0 placements,
forward/reverse slice split, glFrontFace toggle).
- reorient-shells and backface culling are called out as correctness
+ perf levers with their tradeoffs.
- Phase 3 is rewritten around the actual bottleneck surfaced by
profiling: per-frame glNamedBufferSubData stalls on the visible
and indirect buffers. Includes the diagnostic methodology (empty-
screen jump to 60 fps, window/MSAA invariance, upload-comment-out
experiment) so future-me remembers why this is the next step.
- 3A (persistent mapped ring buffers, near-term) and 3B (GPU-side
compute cull, longer-term) split out with scope estimates.
- Roadmap updated: instancing / MDI / reflections / reorient-shells
/ backface cull all ticked; 3A surfaced as the next open item.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
+314
-427
@@ -1,75 +1,115 @@
|
|||||||
# IfcViewer
|
# IfcViewer
|
||||||
|
|
||||||
A high-performance native IFC viewer built on IfcOpenShell's C++ geometry engine with a Qt6 interface and OpenGL 4.5 rendering.
|
A high-performance native IFC viewer built on IfcOpenShell's C++ geometry
|
||||||
|
engine with a Qt6 interface and OpenGL 4.5 rendering.
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
|
|
||||||
```
|
```
|
||||||
+-------------------------------------------+
|
+---------------------------------------------------+
|
||||||
| Qt6 Application (MainWindow) |
|
| Qt6 Application (MainWindow) |
|
||||||
| +----------+ +--------------------------+|
|
| +----------+ +----------------------------------+|
|
||||||
| | Element | | 3D Viewport ||
|
| | Element | | 3D Viewport ||
|
||||||
| | Tree | | (QWindow + OpenGL 4.5) ||
|
| | Tree | | (QWindow + OpenGL 4.5 Core) ||
|
||||||
| | (per- | | ||
|
| | (per- | | ||
|
||||||
| | model) | | Per-model VAO/VBO/EBO ||
|
| | model) | | Per-model: VAO/VBO/EBO ||
|
||||||
| +----------+ | glMultiDrawElements ||
|
| +----------+ | instance SSBO ||
|
||||||
| | Property | | BVH frustum culling ||
|
| | Property | | visible SSBO ||
|
||||||
| | Table | | GPU pick pass ||
|
| | Table | | indirect buffer ||
|
||||||
| +----------+ +--------------------------+|
|
| +----------+ | glMultiDrawElementsIndirect ||
|
||||||
| | Status / Progress / Stats |
|
| | Status / Progress / Stats |
|
||||||
+-------------------------------------------+
|
+---------------------------------------------------+
|
||||||
^ ^
|
^ ^
|
||||||
| |
|
| |
|
||||||
element metadata UploadChunks / Sidecar
|
element metadata MeshChunk / InstanceChunk / Sidecar
|
||||||
| |
|
| |
|
||||||
+-------------------------------------------+
|
+---------------------------------------------------+
|
||||||
| GeometryStreamer (one per loaded model) |
|
| GeometryStreamer (one per loaded model) |
|
||||||
| IfcGeom::Iterator with N threads |
|
| IfcGeom::Iterator with N threads |
|
||||||
| (models loaded sequentially) |
|
| Dedups representations -> MeshChunk |
|
||||||
+-------------------------------------------+
|
| Emits one InstanceChunk per placement |
|
||||||
|
+---------------------------------------------------+
|
||||||
```
|
```
|
||||||
|
|
||||||
### Key design decisions
|
### Key design decisions
|
||||||
|
|
||||||
- **QWindow viewport** embedded via `QWidget::createWindowContainer()`. This gives us a raw native surface for OpenGL, bypassing `QOpenGLWidget`'s compositor overhead.
|
- **QWindow viewport** embedded via `QWidget::createWindowContainer()`. Gives
|
||||||
- **Per-model GPU buffers**: each loaded model gets its own VAO/VBO/EBO. No shared buffer, no cross-model copies on growth. Removing a model frees its GPU memory immediately.
|
us a raw native surface for OpenGL, bypassing `QOpenGLWidget`'s compositor
|
||||||
- **Interleaved vertex format**: position (3 floats) + normal (3 floats) + object ID (1 float, bitcast uint32) + color (RGBA8 packed into 1 float) = 32 bytes per vertex.
|
overhead.
|
||||||
- **Progressive GPU upload**: bulk sidecar loads allocate empty GPU buffers, then stream data in 48 MB chunks per frame. VBO uploads first (no objects visible), then EBO (objects appear progressively as their index range lands). The viewport stays interactive throughout — you can orbit already-loaded models while new ones stream in.
|
- **GPU instancing as the central pillar.** IFC models are dominated by
|
||||||
- **Non-blocking sidecar loading**: sidecar files are read on a background thread. The heavy disk I/O (potentially gigabytes) never blocks the render loop. Only the final GPU upload and tree population happen on the main thread.
|
repeated geometry — identical doors, windows, studs, pipes placed at
|
||||||
- **BVH frustum culling**: per-model BVH trees cull entire subtrees of objects in one frustum test, reducing per-frame cost from O(N) to O(log N). Falls back to linear scan during progressive upload; BVH activates once the model is fully loaded.
|
different transforms. IfcOpenShell's iterator surfaces representation
|
||||||
- **GPU object picking**: a second render pass writes object IDs to an R32UI framebuffer. Click reads back one pixel. No CPU-side raycasting.
|
identity, so we upload each unique mesh exactly once and keep per-placement
|
||||||
- **Multi-model support**: multiple IFC files can be loaded simultaneously. Each model gets its own `GeometryStreamer` (owning the `ifcopenshell::file` for property lookup). Models are loaded sequentially. Per-model visibility toggle and removal are supported.
|
data (transform, object id, optional colour override) in a separate SSBO.
|
||||||
- **Multi-threaded tessellation**: `IfcGeom::Iterator` runs on a background thread and internally parallelizes geometry conversion across all CPU cores.
|
For real projects this collapses tens of millions of triangles of duplicate
|
||||||
- **Non-blocking streaming**: the iterator emits `UploadChunk` signals via Qt's queued connection. The main thread uploads to the GPU without blocking iteration.
|
vertex data into a few hundred MB of unique meshes.
|
||||||
- **World coordinates**: geometry is emitted in world space (`use-world-coords=true`) so no per-object transform matrices are needed on the GPU.
|
- **Per-model GPU buffers**: each loaded model gets its own
|
||||||
|
VAO/VBO/EBO/instance-SSBO/visible-SSBO/indirect-buffer. No cross-model
|
||||||
|
growth copies. Removing a model frees its GPU memory immediately.
|
||||||
|
- **Local-coordinate vertex format (28 B):** position (3 floats) + normal
|
||||||
|
(3 floats) + packed RGBA8 colour (1 uint). The per-instance transform is
|
||||||
|
applied in the vertex shader via an SSBO lookup. No world-baked vertex data.
|
||||||
|
- **Multi-draw indirect:** every frame the CPU builds a flat list of visible
|
||||||
|
instance indices and one `DrawElementsIndirectCommand` per non-empty mesh,
|
||||||
|
then issues a single `glMultiDrawElementsIndirect` per model. 50k visible
|
||||||
|
instances across 8k unique meshes collapse to one driver-side command
|
||||||
|
submission per model.
|
||||||
|
- **BVH frustum culling over instances**: per-model BVH trees cull whole
|
||||||
|
subtrees of placements with one frustum test. Falls back to a linear scan
|
||||||
|
during progressive upload and for very small models (< 32 instances).
|
||||||
|
- **Reflection-aware two-pass draw:** IFC placements can have negative-
|
||||||
|
determinant transforms (mirrored families). These flip the screen-space
|
||||||
|
winding of their triangles, which would make them vanish under
|
||||||
|
`GL_CULL_FACE`. The cull pass buckets visible instances into forward
|
||||||
|
(det ≥ 0) and reverse (det < 0) slices and the renderer issues two MDI
|
||||||
|
calls per model with `glFrontFace` toggled between them.
|
||||||
|
- **`reorient-shells` enabled in the iterator:** makes face winding
|
||||||
|
consistent within a shell at geometry-gen time — the only place this can
|
||||||
|
actually be fixed. Without it, files with inside-out faces produce dark
|
||||||
|
patches and swiss-cheese under backface culling. Costs iterator time but
|
||||||
|
is cached in the sidecar.
|
||||||
|
- **Progressive rendering during streaming:** the viewport is drawable
|
||||||
|
before `finalizeModel()`. Instances are pushed to the SSBO one at a time
|
||||||
|
via `glNamedBufferSubData` as they arrive, and the linear-scan cull path
|
||||||
|
handles them until the BVH is built. Orbit and pan remain interactive
|
||||||
|
through load.
|
||||||
|
- **Non-blocking sidecar loading**: sidecars are read on a background
|
||||||
|
thread; only the final GPU upload touches the main thread.
|
||||||
|
- **GPU object picking**: a second render pass writes object IDs into an
|
||||||
|
R32UI framebuffer. Click reads back one pixel. No CPU-side raycasting.
|
||||||
|
- **Multi-model support**: multiple IFCs can be loaded simultaneously.
|
||||||
|
Each gets its own `GeometryStreamer` (which owns the `ifcopenshell::file`
|
||||||
|
for property lookup). Models load sequentially. Per-model
|
||||||
|
hide/show/remove.
|
||||||
|
|
||||||
### Files
|
### Files
|
||||||
|
|
||||||
| File | Purpose |
|
| File | Purpose |
|
||||||
|------|---------|
|
|------|---------|
|
||||||
| `main.cpp` | Application entry point, GL 4.5 surface format, CLI argument parsing |
|
| `main.cpp` | Application entry, GL 4.5 surface format, CLI argument parsing |
|
||||||
| `MainWindow.h/cpp` | Qt main window: multi-model project management, element tree, property table, status bar |
|
| `MainWindow.h/cpp` | Qt main window: multi-model project, element tree, properties, status |
|
||||||
| `ViewportWindow.h/cpp` | OpenGL 4.5 Core renderer: shaders, buffer management, camera, frustum culling, BVH traversal, picking |
|
| `ViewportWindow.h/cpp` | OpenGL 4.5 Core renderer: shaders, buffers, camera, culling, MDI draw, picking |
|
||||||
| `GeometryStreamer.h/cpp` | Background geometry processing: loads IFC, runs iterator, emits chunks (one per model) |
|
| `GeometryStreamer.h/cpp` | Background iterator runner; emits `MeshChunk` + `InstanceChunk` |
|
||||||
| `BvhAccel.h/cpp` | BVH construction (median-split), per-model trees, EBO reordering |
|
| `InstancedGeometry.h` | Shared structs: `MeshInfo`, `InstanceCpu`, `InstanceGpu`, chunk records |
|
||||||
| `SidecarCache.h/cpp` | Raw binary `.ifcview` sidecar read/write |
|
| `BvhAccel.h/cpp` | Median-split BVH builder; operates on instance world-AABBs |
|
||||||
| `AppSettings.h/cpp` | Persisted application preferences (geometry library, show stats) |
|
| `SidecarCache.h/cpp` | Raw binary `.ifcview` (v4) sidecar read/write |
|
||||||
| `SettingsWindow.h/cpp` | Settings dialog UI |
|
| `AppSettings.h/cpp` | Persisted preferences (geometry library, stats overlay, backface culling) |
|
||||||
|
| `SettingsWindow.h/cpp` | Settings dialog |
|
||||||
| `CMakeLists.txt` | Build configuration |
|
| `CMakeLists.txt` | Build configuration |
|
||||||
|
|
||||||
## Dependencies
|
## Dependencies
|
||||||
|
|
||||||
- **Qt6** (Core, Gui, Widgets)
|
- **Qt6** (Core, Gui, Widgets, OpenGL)
|
||||||
- **OpenGL 4.5** (GL_ARB_direct_state_access) - available on Windows and Linux; macOS will need a Vulkan/MoltenVK backend (not yet implemented)
|
- **OpenGL 4.5** with `GL_ARB_direct_state_access` and
|
||||||
- **IfcOpenShell C++ libraries** (IfcParse, IfcGeom, and their dependencies: Open CASCADE, Boost, Eigen3, optionally CGAL)
|
`GL_ARB_shader_draw_parameters` — available on Windows and Linux. macOS
|
||||||
|
will need a Vulkan/MoltenVK backend (not yet implemented; macOS caps out
|
||||||
|
at GL 4.1).
|
||||||
|
- **IfcOpenShell C++ libraries** (IfcParse, IfcGeom, and their
|
||||||
|
dependencies: Open CASCADE, Boost, Eigen3, optionally CGAL).
|
||||||
|
|
||||||
## Building
|
## Building
|
||||||
|
|
||||||
IfcViewer is built as part of the IfcOpenShell CMake project. You do not need to build everything - disable the targets you don't need.
|
IfcViewer is part of the IfcOpenShell CMake project. From the repo root:
|
||||||
|
|
||||||
### Minimal build (IfcViewer only)
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
mkdir build && cd build
|
mkdir build && cd build
|
||||||
@@ -89,25 +129,13 @@ cmake ../cmake \
|
|||||||
make -j$(nproc) IfcViewer
|
make -j$(nproc) IfcViewer
|
||||||
```
|
```
|
||||||
|
|
||||||
This builds only IfcParse, IfcGeom (with geometry kernels), and IfcViewer itself. All other targets (IfcConvert, Python bindings, serializers, etc.) are skipped.
|
|
||||||
|
|
||||||
If Qt6 is not in a standard location, pass `-DQT_DIR=/path/to/qt6`.
|
If Qt6 is not in a standard location, pass `-DQT_DIR=/path/to/qt6`.
|
||||||
|
|
||||||
### Full build with IfcViewer enabled
|
|
||||||
|
|
||||||
```sh
|
|
||||||
cmake ../cmake -DBUILD_IFCVIEWER=ON
|
|
||||||
make -j$(nproc)
|
|
||||||
```
|
|
||||||
|
|
||||||
## Usage
|
## Usage
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
# Open one or more files from the command line
|
|
||||||
./IfcViewer arch.ifc struct.ifc mep.ifc
|
./IfcViewer arch.ifc struct.ifc mep.ifc
|
||||||
|
./IfcViewer # then File -> Add Files
|
||||||
# Or use File -> Add Files from the menu (supports multiselect)
|
|
||||||
./IfcViewer
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Controls
|
### Controls
|
||||||
@@ -117,449 +145,308 @@ make -j$(nproc)
|
|||||||
| Middle mouse drag | Orbit camera |
|
| Middle mouse drag | Orbit camera |
|
||||||
| Shift + middle mouse drag | Pan camera |
|
| Shift + middle mouse drag | Pan camera |
|
||||||
| Scroll wheel | Zoom |
|
| Scroll wheel | Zoom |
|
||||||
| Left click | Select object (highlights in viewport and tree) |
|
| Left click | Select object |
|
||||||
|
|
||||||
### Keyboard shortcuts
|
### Keyboard
|
||||||
|
|
||||||
| Key | Action |
|
| Key | Action |
|
||||||
|-----|--------|
|
|-----|--------|
|
||||||
| Ctrl+O | Add files |
|
| Ctrl+O | Add files |
|
||||||
| Ctrl+Q | Quit |
|
| Ctrl+Q | Quit |
|
||||||
|
|
||||||
|
### Settings
|
||||||
|
|
||||||
|
- **Geometry Library** — kernel string passed to IfcOpenShell (default
|
||||||
|
`hybrid-cgal-simple-opencascade`).
|
||||||
|
- **Show Performance Stats** — overlay FPS / object / triangle / draw
|
||||||
|
counts in the status bar.
|
||||||
|
- **Backface Culling** — `GL_CULL_FACE` on closed solids. Default on.
|
||||||
|
Disable if a model uses open shells and you see missing faces.
|
||||||
|
|
||||||
## Performance Strategy
|
## Performance Strategy
|
||||||
|
|
||||||
The viewer targets smooth orbiting at 60 fps on models up to 1 million IFC objects.
|
The viewer targets smooth orbiting at 60 fps on real-world multi-discipline
|
||||||
Rendering performance is addressed in three phases. Each phase builds on the
|
BIM projects (a "real job" being ~50 models, several million placements,
|
||||||
previous one, and the system is designed so that smaller models never pay for
|
hundreds of millions of rasterised triangles when everything is in view).
|
||||||
optimizations they don't need.
|
|
||||||
|
|
||||||
### Phase 1: Per-Object Frustum Culling (CPU)
|
Rendering performance has evolved in phases. Each builds on the previous,
|
||||||
|
and smaller models never pay for optimisations they don't need.
|
||||||
|
|
||||||
**Status:** Implemented.
|
### Phase 1 — Per-object Frustum Culling
|
||||||
|
|
||||||
The simplest win: don't draw what's off screen.
|
**Status:** implemented (and still the fallback for small models / during
|
||||||
|
streaming).
|
||||||
|
|
||||||
#### Data model
|
Six view-frustum planes are extracted from the view-projection matrix each
|
||||||
|
frame. Each instance's world AABB is tested with the p-vertex / n-vertex
|
||||||
|
method (one dot product + one compare per plane, 6 planes).
|
||||||
|
|
||||||
During `uploadChunk()`, the viewport records a small metadata struct for every
|
Surviving instance indices are written into a per-mesh bucket, then
|
||||||
object that enters the GPU buffers:
|
flattened into a single `uint[]` (the "visible SSBO", binding = 1) and
|
||||||
|
accompanied by one `DrawElementsIndirectCommand` per non-empty mesh.
|
||||||
|
One `glMultiDrawElementsIndirect` call per model draws everything.
|
||||||
|
|
||||||
```cpp
|
Cost: ~6 dot products per instance per frame. Fine up to ~100 k instances
|
||||||
struct ObjectDrawInfo {
|
per frame; above that the linear scan shows up in profiles, motivating
|
||||||
uint32_t index_offset; // byte offset into the model's EBO
|
Phase 2.
|
||||||
uint32_t index_count; // number of indices (triangles * 3)
|
|
||||||
uint32_t model_id; // which model this object belongs to
|
|
||||||
float aabb_min[3]; // world-space axis-aligned bounding box
|
|
||||||
float aabb_max[3]; // (computed from vertex positions at upload time)
|
|
||||||
};
|
|
||||||
```
|
|
||||||
|
|
||||||
This costs 32 bytes per object. For 1M objects that's ~32 MB of CPU-side
|
### Phase 2 — BVH Acceleration + Sidecar Cache
|
||||||
metadata — negligible next to the vertex data.
|
|
||||||
|
|
||||||
#### Frustum extraction
|
**Status:** implemented.
|
||||||
|
|
||||||
Each frame, before drawing, six clip planes are extracted from the
|
For models exceeding ~32 instances, a bounding volume hierarchy groups
|
||||||
view-projection matrix (`VP = proj * view`). The standard Griess-Hartmann
|
nearby placements into a binary tree and culls entire subtrees with a
|
||||||
method pulls them directly from the matrix rows:
|
single frustum test. This reduces per-frame work from O(N) to O(log N) in
|
||||||
|
the best case (camera zoomed to a corner) and remains well under 1 ms for
|
||||||
|
100 k instances in the worst case (everything on screen).
|
||||||
|
|
||||||
```
|
A BVH was chosen over an octree because BIM data is spatially non-uniform
|
||||||
left = VP[3] + VP[0]
|
— dense MEP risers in one zone, sparse open atria in another. An octree
|
||||||
right = VP[3] - VP[0]
|
subdivides space uniformly, wasting nodes on empty regions and creating
|
||||||
bottom = VP[3] + VP[1]
|
deep chains in dense ones. A BVH adapts its splits to the actual
|
||||||
top = VP[3] - VP[1]
|
placement distribution.
|
||||||
near = VP[3] + VP[2]
|
|
||||||
far = VP[3] - VP[2]
|
|
||||||
```
|
|
||||||
|
|
||||||
Each plane is stored as (a, b, c, d) and normalized so that
|
#### Activation
|
||||||
`a*x + b*y + c*z + d` gives the signed distance from the plane.
|
|
||||||
|
|
||||||
#### AABB-frustum test
|
The BVH is optional and non-disruptive. Until it is built, the Phase 1
|
||||||
|
linear scan handles culling. The renderer checks for a BVH per model and
|
||||||
|
falls back to the scan for any model that doesn't have one.
|
||||||
|
|
||||||
For each object, the AABB is tested against all six planes using the
|
It activates in one of two ways:
|
||||||
"p-vertex / n-vertex" method:
|
|
||||||
|
|
||||||
- For each plane, find the AABB corner most in the direction of the plane
|
1. **Sidecar hit** — the `.ifcview` file next to the `.ifc` is found and
|
||||||
normal (the p-vertex).
|
valid; its instance data is uploaded and the BVH rebuilt on the fly
|
||||||
- If the p-vertex is on the negative side of the plane, the entire AABB is
|
from the restored AABBs (cheap — `< 100 ms` for 100 k placements).
|
||||||
outside the frustum → cull.
|
2. **After streaming** — `finalizeModel()` builds the BVH synchronously
|
||||||
- If any plane culls the object, skip it.
|
once all chunks are in (instances already live on the GPU, so there's
|
||||||
|
no EBO re-sort to do). The sidecar is written afterwards.
|
||||||
|
|
||||||
This test is conservative: it never culls a visible object, but may
|
Models under 32 instances skip the BVH.
|
||||||
occasionally keep an invisible one (when the AABB straddles a frustum corner).
|
|
||||||
That's fine — false positives just cost a few extra triangles.
|
|
||||||
|
|
||||||
#### Drawing visible objects
|
#### BVH node layout (32 B, two per cache line)
|
||||||
|
|
||||||
The surviving objects' `(index_count, index_offset)` pairs are passed to
|
|
||||||
`glMultiDrawElements()` in a single call. This replaces the previous single
|
|
||||||
`glDrawElements()` that drew everything. The GPU processes only the index
|
|
||||||
ranges that survived the frustum test.
|
|
||||||
|
|
||||||
Alternatively, for the pick pass (which runs less frequently), the same
|
|
||||||
visibility list is reused — objects culled from the main pass are also culled
|
|
||||||
from picking.
|
|
||||||
|
|
||||||
#### Performance characteristics
|
|
||||||
|
|
||||||
| Metric | Value |
|
|
||||||
|--------|-------|
|
|
||||||
| Per-object cost | ~6 dot products + 6 comparisons per frame |
|
|
||||||
| 50k objects | ~0.3 ms on a modern CPU core |
|
|
||||||
| 500k objects | ~3 ms (starts to matter at 60 fps) |
|
|
||||||
| 1M objects | ~6 ms (too expensive — need phase 3) |
|
|
||||||
| Memory overhead | 32 bytes/object |
|
|
||||||
| Load-time overhead | Near zero (AABB computed during existing upload) |
|
|
||||||
|
|
||||||
Phase 1 is sufficient for models up to ~100k objects. Beyond that, the CPU-side
|
|
||||||
frustum test becomes a measurable fraction of the frame budget, motivating
|
|
||||||
phase 3.
|
|
||||||
|
|
||||||
### Phase 2: BVH Acceleration (optional, for large models)
|
|
||||||
|
|
||||||
**Status:** Implemented.
|
|
||||||
|
|
||||||
For models exceeding ~100 objects, a bounding volume hierarchy (BVH) groups
|
|
||||||
nearby objects into a binary tree and culls entire subtrees in one frustum
|
|
||||||
test. This reduces the number of AABB-frustum tests from O(N_objects) to
|
|
||||||
O(log N) in the best case (camera zoomed into a corner) and gives a constant
|
|
||||||
overhead for the common case where most of the model is on screen.
|
|
||||||
|
|
||||||
A BVH was chosen over an octree because BIM data is spatially non-uniform —
|
|
||||||
dense MEP risers in one zone, sparse open atriums in another. An octree
|
|
||||||
subdivides space uniformly, wasting nodes on empty regions and creating deep
|
|
||||||
chains in dense ones. A BVH adapts its splits to the actual object
|
|
||||||
distribution, producing balanced trees regardless of density variation.
|
|
||||||
|
|
||||||
#### When the BVH activates
|
|
||||||
|
|
||||||
The BVH is **optional and non-disruptive**. Until it is built, phase 1's
|
|
||||||
linear scan handles all culling. The rendering loop checks for an active BVH
|
|
||||||
and falls back to the linear scan for any model that doesn't have one.
|
|
||||||
|
|
||||||
The BVH activates in one of two ways:
|
|
||||||
|
|
||||||
1. **Sidecar cache exists**: If a `.ifcview` file is found next to the `.ifc`
|
|
||||||
file, the BVH is loaded from it instantly (raw memory read, no parsing).
|
|
||||||
The model uses BVH culling from the first frame after loading.
|
|
||||||
2. **Automatic build**: After streaming finishes, a background thread builds
|
|
||||||
the BVH from the per-object AABBs already computed in phase 1. Until it
|
|
||||||
completes, phase 1 culling handles visibility. On completion, the render
|
|
||||||
thread picks up the BVH on the next frame. The sidecar is written for
|
|
||||||
future loads.
|
|
||||||
|
|
||||||
Models with fewer than 32 objects skip the BVH entirely — the overhead of tree
|
|
||||||
traversal is worse than a linear scan at that scale.
|
|
||||||
|
|
||||||
#### BVH node layout
|
|
||||||
|
|
||||||
Each node is 32 bytes, so two nodes fit in one 64-byte cache line:
|
|
||||||
|
|
||||||
```cpp
|
```cpp
|
||||||
struct BvhNode {
|
struct BvhNode {
|
||||||
float aabb_min[3]; // world-space bounding box (12 bytes)
|
float aabb_min[3]; // 12 B
|
||||||
float aabb_max[3]; // (12 bytes)
|
float aabb_max[3]; // 12 B
|
||||||
uint32_t right_or_first; // interior: right child index; leaf: first object index (4 bytes)
|
uint32_t right_or_first; // interior: right child index; leaf: first item index
|
||||||
uint16_t count; // 0 = interior node; >0 = leaf with this many objects (2 bytes)
|
uint16_t count; // 0 = interior, >0 = leaf
|
||||||
uint16_t axis; // split axis for interior (0=x, 1=y, 2=z); unused for leaf (2 bytes)
|
uint16_t axis; // 0/1/2 for interior; unused for leaf
|
||||||
};
|
};
|
||||||
```
|
```
|
||||||
|
|
||||||
Interior nodes store the right child index; the left child is always the
|
Left child is always the next node (pre-order DFS). Leaf items are
|
||||||
immediately next node in the array (implicit in pre-order DFS layout, no
|
indices into the per-model `instances` array; the parallel `bvh_items[]`
|
||||||
pointer needed). Leaf nodes reference a contiguous range in a sorted
|
array carries the world AABBs.
|
||||||
object-index array.
|
|
||||||
|
|
||||||
The BVH is stored as a flat `std::vector<BvhNode>` in pre-order DFS layout.
|
#### Build: object-median split
|
||||||
This means a depth-first traversal (which is what frustum culling does) reads
|
|
||||||
memory sequentially, maximizing prefetch and cache-line utilization.
|
|
||||||
|
|
||||||
#### Build algorithm: object-median split
|
1. Compute centroid of each item's AABB.
|
||||||
|
2. Pick the longest axis of the node's AABB.
|
||||||
|
3. `std::nth_element` partitions at the median on that axis — O(n).
|
||||||
|
4. Recurse until a leaf holds ≤ 8 items.
|
||||||
|
|
||||||
1. Compute the centroid of each object's AABB.
|
O(n log n) total. No SAH — for frustum culling (6-plane tests, early
|
||||||
2. Find the longest axis of the current node's bounding box.
|
subtree reject) the quality difference vs median is negligible.
|
||||||
3. Use `std::nth_element` to partition objects at the median centroid on that
|
|
||||||
axis. This is O(n) — no full sort needed.
|
|
||||||
4. Recurse on each half. Terminate when the node contains ≤ 8 objects (leaf).
|
|
||||||
5. Write nodes into the flat array in pre-order DFS.
|
|
||||||
|
|
||||||
Total build time is O(n log n). For 100k objects this is well under 100 ms on
|
#### Traversal: stack-based, no recursion
|
||||||
a single core.
|
|
||||||
|
|
||||||
SAH (Surface Area Heuristic) is the gold standard for ray-tracing BVHs, but
|
|
||||||
for frustum culling — where we test 6 planes and early-out entire subtrees —
|
|
||||||
the quality difference vs. median split is negligible. Median split is simpler
|
|
||||||
and produces reliably balanced trees.
|
|
||||||
|
|
||||||
#### Frustum traversal
|
|
||||||
|
|
||||||
The traversal uses an explicit stack on the C++ stack (no heap allocation,
|
|
||||||
no recursion):
|
|
||||||
|
|
||||||
```
|
```
|
||||||
stack[64] = {0} // start at root; depth 64 handles billions of objects
|
stack[64] = { 0 } // root
|
||||||
while stack not empty:
|
while stack not empty:
|
||||||
node = nodes[stack.pop()]
|
node = nodes[stack.pop()]
|
||||||
if node AABB outside frustum: continue // cull entire subtree
|
if node.aabb outside frustum: continue
|
||||||
if leaf:
|
if leaf:
|
||||||
for each object in node:
|
for each item in node:
|
||||||
if object AABB in frustum: emit to visible list
|
if item.aabb in frustum: emit to visible list
|
||||||
else:
|
else:
|
||||||
push right child, push left child // left processed first (DFS)
|
push right child, push left child // left processed first (DFS)
|
||||||
```
|
```
|
||||||
|
|
||||||
When the camera is zoomed into a corner of the model, the traversal skips
|
Depth 64 is enough for billions of items on any balanced tree. The stack
|
||||||
large portions of the tree after testing only a handful of interior nodes.
|
is on the C++ stack, zero per-frame allocation.
|
||||||
When zoomed out to see everything, the traversal visits all leaves but the
|
|
||||||
overhead of the interior-node tests is small relative to the leaf work.
|
|
||||||
|
|
||||||
#### Per-model BVH
|
#### Sidecar format (`.ifcview`, v4)
|
||||||
|
|
||||||
Each loaded model gets its own BVH. During frustum culling, the outer loop
|
Raw memory dump, Blender-`.blend`-style — no serialisation, no parsing.
|
||||||
iterates over models (skipping hidden/removed ones); the inner loop traverses
|
Stores everything needed to skip the `IfcGeom::Iterator` pass:
|
||||||
that model's BVH. This means hiding or removing a model is free — just skip
|
|
||||||
its BVH, no tree modification needed.
|
|
||||||
|
|
||||||
```cpp
|
```
|
||||||
struct ModelBvh {
|
SidecarHeader (magic "IFVW", version, endian, ...)
|
||||||
uint32_t model_id;
|
uint64_t source_file_size
|
||||||
std::vector<BvhNode> nodes; // flat BVH node array
|
uint32_t + float[] vertex data (7 floats × N_verts, local coords)
|
||||||
std::vector<uint32_t> object_indices; // indices into object_draw_info_
|
uint32_t + uint32_t[] index data (mesh-local)
|
||||||
|
uint32_t + MeshInfo[] per-unique-mesh metadata (48 B each)
|
||||||
|
uint32_t + InstanceCpu[] per-placement records (transform + AABB + ids)
|
||||||
|
uint32_t + PackedElementInfo[] element tree records
|
||||||
|
uint32_t + char[] string table
|
||||||
|
```
|
||||||
|
|
||||||
|
Staleness check: `source_file_size` vs actual file size. Mismatched →
|
||||||
|
reject and rebuild. Endianness marker rejects cross-arch caches.
|
||||||
|
|
||||||
|
### GPU Instancing pipeline (the central pillar)
|
||||||
|
|
||||||
|
Everything above plugs into a single data-flow, worth documenting on its
|
||||||
|
own because it's what makes the whole thing fast.
|
||||||
|
|
||||||
|
Per-model state on the GPU:
|
||||||
|
|
||||||
|
| Buffer | Contents | Lifetime |
|
||||||
|
|--------|----------|----------|
|
||||||
|
| `VBO` | Interleaved local-coord vertex data (28 B/vert). One range per unique representation. | Grow-on-demand during streaming; static after finalize. |
|
||||||
|
| `EBO` | Mesh-local uint32 indices. One range per unique representation. | Same. |
|
||||||
|
| `SSBO` (binding 0) | `InstanceGpu[]` (80 B each: mat4 transform, object_id, color_override, pad). | Appended during streaming, static after finalize. |
|
||||||
|
| `visible SSBO` (binding 1) | `uint32[]` — flat list of visible instance indices, ordered by mesh, uploaded each frame. | Rewritten every frame. |
|
||||||
|
| Draw-indirect buffer | `DrawElementsIndirectCommand[]` — one per non-empty mesh, uploaded each frame. | Rewritten every frame. |
|
||||||
|
|
||||||
|
Draw command:
|
||||||
|
|
||||||
|
```c
|
||||||
|
struct DrawElementsIndirectCommand {
|
||||||
|
uint32_t count; // mesh.index_count
|
||||||
|
uint32_t instanceCount; // visible-list length for this mesh
|
||||||
|
uint32_t firstIndex; // mesh.ebo_byte_offset / 4
|
||||||
|
uint32_t baseVertex; // mesh.vbo_byte_offset / 28
|
||||||
|
uint32_t baseInstance; // offset into the flat visible-index array
|
||||||
};
|
};
|
||||||
```
|
```
|
||||||
|
|
||||||
#### EBO re-sorting
|
The vertex shader reads `visible[gl_BaseInstanceARB + gl_InstanceID]` to
|
||||||
|
get the real instance id, then indexes into the instance SSBO:
|
||||||
|
|
||||||
For BVH culling to maximise GPU cache performance, the EBO is re-sorted so
|
```glsl
|
||||||
that objects in the same BVH leaf are contiguous. This happens via **deferred
|
uint slot = uint(gl_BaseInstanceARB) + uint(gl_InstanceID);
|
||||||
compaction**:
|
uint iid = visible[slot];
|
||||||
|
InstanceRecord inst = instances[iid];
|
||||||
1. During initial load, geometry uploads in iterator order (fast first frame,
|
gl_Position = u_view_projection * inst.transform * vec4(a_position, 1.0);
|
||||||
phase 1 culling active).
|
|
||||||
2. After the BVH build completes on the background thread:
|
|
||||||
a. Walk the BVH leaves in DFS order.
|
|
||||||
b. For each object in each leaf, copy its index data to a new EBO buffer,
|
|
||||||
updating `ObjectDrawInfo::index_offset` accordingly.
|
|
||||||
c. Package the reordered EBO + updated draw info as a `BvhBuildResult`.
|
|
||||||
3. The render thread picks up the result on the next frame: one
|
|
||||||
`glNamedBufferSubData` call to re-upload the EBO, then swap in the new
|
|
||||||
draw info and activate the BVH. One frame of stutter, bounded by EBO
|
|
||||||
upload time (~5 ms for 32 MB).
|
|
||||||
|
|
||||||
#### Async build and render-thread handoff
|
|
||||||
|
|
||||||
The BVH build must not stall the render loop:
|
|
||||||
|
|
||||||
1. `buildBvhAsync()` snapshots `object_draw_info_` under the upload mutex,
|
|
||||||
then launches a `std::thread`.
|
|
||||||
2. The thread builds the BVH and reordered EBO, then stores the result in a
|
|
||||||
`pending_bvh_result_` pointer under a separate mutex.
|
|
||||||
3. At the top of each `render()` call, `applyBvhResult()` checks for a
|
|
||||||
pending result. If found, it re-uploads the EBO (requires GL context),
|
|
||||||
swaps the draw info, and activates the BVH.
|
|
||||||
4. Until the BVH is ready, phase 1's linear scan runs every frame as before.
|
|
||||||
|
|
||||||
#### Preprocessed sidecar format (`.ifcview`)
|
|
||||||
|
|
||||||
The sidecar is a raw memory dump (Blender `.blend`-style) — no serialization
|
|
||||||
format, no parsing. It stores everything needed to display the model without
|
|
||||||
re-tessellating: vertex data, index data, per-object metadata, element tree
|
|
||||||
info, and the BVH. Loading is just `fread` into vectors → GPU upload →
|
|
||||||
render. The expensive `IfcGeom::Iterator` tessellation is skipped entirely.
|
|
||||||
|
|
||||||
The IFC file is still parsed on demand (in background) for detailed property
|
|
||||||
lookup; the sidecar provides the basic properties (name, type, GUID)
|
|
||||||
immediately.
|
|
||||||
|
|
||||||
```
|
|
||||||
SidecarHeader (16 bytes: magic, version, endian, reserved)
|
|
||||||
uint64_t source_file_size
|
|
||||||
|
|
||||||
uint32_t + float[] vertex data (interleaved, 8 floats/vertex)
|
|
||||||
uint32_t + uint32_t[] index data (global indices, ready for EBO)
|
|
||||||
uint32_t + ObjectDrawInfo[] per-object draw metadata
|
|
||||||
uint32_t + PackedElementInfo[] element tree records (fixed-size)
|
|
||||||
uint32_t + char[] string table (concatenated UTF-8: guid, name, type)
|
|
||||||
|
|
||||||
uint32_t num_bvh_models
|
|
||||||
per model:
|
|
||||||
uint32_t model_id
|
|
||||||
uint32_t + BvhNode[] BVH node array
|
|
||||||
uint32_t + uint32_t[] object indices
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Staleness check: `source_file_size` is compared against the actual IFC file
|
`gl_BaseInstanceARB` requires `GL_ARB_shader_draw_parameters`, which is
|
||||||
size. If mismatched, the sidecar is stale and is rebuilt. This is cheap and
|
available on all GL-4.6-capable drivers.
|
||||||
sufficient for a local cache (no hash computation on multi-GB files).
|
|
||||||
|
|
||||||
Endianness: if the marker reads back as `0x01020304`, the file was written on
|
Reflection handling: at upload time we store a parallel
|
||||||
the same architecture — just `fread` the structs directly. Otherwise, reject
|
`instance_reflected[]` byte array (1 if the transform's upper-3×3 has
|
||||||
the sidecar and rebuild.
|
det < 0). The cull pass produces two flat visible-list slices — fwd
|
||||||
|
(non-reflected) first, rev (reflected) after — concatenated into one
|
||||||
|
buffer. The renderer issues MDI twice: fwd with `glFrontFace(GL_CCW)`,
|
||||||
|
rev with `glFrontFace(GL_CW)`. `GL_CULL_FACE` stays on and does the
|
||||||
|
right thing in both passes.
|
||||||
|
|
||||||
#### Performance characteristics
|
### Current bottleneck — Phase 3 as designed is already obsolete
|
||||||
|
|
||||||
|
The original README's Phase 3 ("GPU-driven indirect draw") described
|
||||||
|
moving draw submission to the GPU via compute. In the meantime, GPU
|
||||||
|
instancing and MDI made the CPU-side draw cost essentially free (10
|
||||||
|
`glMultiDrawElementsIndirect` calls per frame for 10 models). **That
|
||||||
|
goal is met.** The real Phase 3 problem is different.
|
||||||
|
|
||||||
|
#### Diagnosed on a 10-model / 379 k-instance / 128 M-triangle scene
|
||||||
|
|
||||||
|
Observed numbers (everything in view, no movement):
|
||||||
|
|
||||||
| Metric | Value |
|
| Metric | Value |
|
||||||
|--------|-------|
|
|--------|-------|
|
||||||
| BVH build time (100k objects) | < 100 ms (single-threaded, background) |
|
| FPS | 10 |
|
||||||
| Per-frame traversal (100k objects, 50% visible) | ~0.1 ms |
|
| Frame time | ~100 ms |
|
||||||
| Per-frame traversal (100k objects, 5% visible) | ~0.02 ms |
|
| gl_draws | 10 |
|
||||||
| Memory overhead | 32 bytes/node + 4 bytes/object index (~1.5× object count) |
|
| Sub-draws packed in indirect buffers | 67 037 |
|
||||||
| EBO reorder (one-time) | 1–5 ms upload for 32 MB EBO |
|
|
||||||
| Sidecar file size | ~same as geometry data (vertices + indices + metadata) |
|
|
||||||
| Sidecar read time | bounded by disk I/O (~500 ms for 640 MB, ~2 s for 2.8 GB from NVMe) |
|
|
||||||
| GPU upload time | progressive: ~48 MB/frame (~1 s for 2.8 GB at 60 fps, non-blocking) |
|
|
||||||
|
|
||||||
#### Spatial coherence bonus
|
Elimination experiments:
|
||||||
|
|
||||||
Beyond culling, BVH-leaf-sorted EBOs improve GPU cache performance. When the
|
| Probe | Result | Interpretation |
|
||||||
GPU rasterizes a leaf's triangles, the vertices are close together in the VBO,
|
|-------|--------|----------------|
|
||||||
so the post-transform vertex cache hits more often. This can yield 10–20%
|
| Camera off-screen (nothing visible) | → 60 fps | GPU is idle; CPU path is cheap |
|
||||||
rasterization speedup even when nothing is culled (e.g. zoomed out to see the
|
| Resize window to 1/4 area | no change | Not fragment/raster bound |
|
||||||
whole model).
|
| `setSamples(4)` → `setSamples(1)` | no change | Not MSAA/resolve bound |
|
||||||
|
| Comment out the two `glNamedBufferSubData` in `cullAndUploadVisible` | → 60 fps (screen blank) | **The per-frame uploads are the bottleneck.** |
|
||||||
|
|
||||||
### Phase 3: GPU-Driven Indirect Draw
|
So the bottleneck is two `glNamedBufferSubData` calls per model per
|
||||||
|
frame uploading ~1.5 MB (visible list) + ~1.3 MB (indirect buffer).
|
||||||
|
3 MB/frame / 60 fps = 180 MB/s — trivial for the bus, but `glNamedBufferSubData`
|
||||||
|
against a buffer the GPU is still reading forces the driver to stall
|
||||||
|
the CPU or orphan/reallocate the backing store, and we're hitting that
|
||||||
|
on 20 buffers per frame.
|
||||||
|
|
||||||
For models with 500k+ objects, even tile-level CPU culling is fast, but the
|
### Phase 3 (proposed) — Eliminate per-frame upload stalls
|
||||||
real bottleneck shifts to draw call submission. Phase 3 moves all per-frame
|
|
||||||
visibility decisions to the GPU via compute shaders and indirect draw commands.
|
|
||||||
|
|
||||||
#### How it works
|
Two ways to attack it, in ascending order of effort:
|
||||||
|
|
||||||
Phase 3 builds on the BVH from phase 2. It does not replace the BVH — it
|
#### 3A. Persistent mapped ring buffers (near-term)
|
||||||
moves the per-frame traversal to the GPU.
|
|
||||||
|
|
||||||
1. **Upload phase** (once, at load time):
|
Allocate each of the per-frame-written buffers with
|
||||||
- Per-leaf AABBs from the BVH are uploaded to a GPU SSBO (`leaf_aabbs`).
|
`glBufferStorage(GL_MAP_PERSISTENT_BIT | GL_MAP_COHERENT_BIT | GL_MAP_WRITE_BIT)`
|
||||||
- One `DrawElementsIndirectCommand` per BVH leaf is written to an indirect
|
at 3× the needed size. Keep one `void*` from `glMapBufferRange` forever.
|
||||||
draw buffer:
|
Each frame, write the CPU-side data into slice `frame % 3` and bind
|
||||||
```c
|
that slice via `glBindBufferRange`. The GPU reads slice N−1 while the
|
||||||
struct DrawElementsIndirectCommand {
|
CPU writes slice N — no driver sync, no orphan, no stall.
|
||||||
uint count; // leaf's total index count
|
|
||||||
uint instanceCount; // 1
|
|
||||||
uint firstIndex; // offset into EBO (from BVH leaf order)
|
|
||||||
uint baseVertex; // 0 (indices are global)
|
|
||||||
uint baseInstance; // leaf_id (available in shader via gl_DrawID)
|
|
||||||
};
|
|
||||||
```
|
|
||||||
- A "template" copy of the indirect buffer is kept so the compute shader
|
|
||||||
can reset culled commands each frame without re-uploading from CPU.
|
|
||||||
|
|
||||||
2. **Cull phase** (every frame, on the GPU):
|
Scope: ~80 lines across `ModelGpuData` + `cullAndUploadVisible` +
|
||||||
- The CPU uploads 6 frustum plane vec4s as a uniform or small UBO.
|
binding in `render()` / `renderPickPass()`. No algorithmic change, no
|
||||||
- A compute shader dispatches `ceil(N_leaves / 64)` workgroups:
|
shader change. Expected result on the stats scene: 10 fps → ~60 fps
|
||||||
```glsl
|
(the measured ceiling once uploads are removed).
|
||||||
layout(local_size_x = 64) in;
|
|
||||||
|
|
||||||
void main() {
|
#### 3B. GPU-side culling (longer-term)
|
||||||
uint leaf_id = gl_GlobalInvocationID.x;
|
|
||||||
if (leaf_id >= leaf_count) return;
|
|
||||||
|
|
||||||
// Copy from template (resets any previously zeroed commands)
|
Push culling itself to the GPU. A compute shader reads the
|
||||||
commands[leaf_id] = template_commands[leaf_id];
|
`InstanceCpu`-equivalent SSBO + frustum planes, builds the visible list
|
||||||
|
and indirect commands in-place via atomics. Zero CPU→GPU per-frame
|
||||||
|
bytes. Also lays the foundation for occlusion and contribution culling
|
||||||
|
(both want to run on the GPU anyway, with access to the depth buffer
|
||||||
|
or screen-space projection).
|
||||||
|
|
||||||
// Frustum test
|
Scope: compute shader + atomic counter + BVH-traversal-on-GPU (or a
|
||||||
if (!aabb_vs_frustum(leaf_aabbs[leaf_id], frustum_planes)) {
|
linear compute scan — simpler and still gains most of the win since
|
||||||
commands[leaf_id].count = 0; // culled: GPU skips zero-count draws
|
traversal isn't the bottleneck once upload is gone). Bigger change;
|
||||||
}
|
worth doing after 3A is measured, because 3A may be enough for a long
|
||||||
}
|
while.
|
||||||
```
|
|
||||||
- A memory barrier ensures the indirect buffer is visible to the draw stage.
|
|
||||||
|
|
||||||
3. **Draw phase** (every frame):
|
### Planned follow-ups (post-Phase-3)
|
||||||
- One call: `glMultiDrawElementsIndirect(GL_TRIANGLES, GL_UNSIGNED_INT,
|
|
||||||
nullptr, N_leaves, 0)`.
|
|
||||||
- The GPU reads the indirect buffer, skips tiles with `count == 0`, and
|
|
||||||
draws the rest. Zero CPU-side per-object or per-tile work.
|
|
||||||
|
|
||||||
#### What the CPU does per frame
|
- **Screen-space contribution cull.** Reject instances whose projected
|
||||||
|
screen-space AABB is below a pixel threshold. Cheap CPU-side filter
|
||||||
|
that eliminates distant MEP detail. Big win on unfiltered plant-room
|
||||||
|
scenes.
|
||||||
|
- **Hierarchical-Z occlusion culling.** Render large occluders, build a
|
||||||
|
depth pyramid, test BVH / instance AABBs against it. In dense BIM,
|
||||||
|
most geometry is behind other geometry from any given viewpoint; this
|
||||||
|
is historically a 3–10× reduction in drawn instances.
|
||||||
|
- **Distance / contribution LOD.** Unique meshes pre-simplified at load
|
||||||
|
time; compute shader selects an LOD per instance per frame based on
|
||||||
|
screen-space size. Same visible-SSBO plumbing, different `firstIndex`.
|
||||||
|
- **Mesh shaders / meshlets.** Ceiling-raising but overkill until the
|
||||||
|
above are exhausted.
|
||||||
|
|
||||||
1. Upload 6 vec4 frustum planes (96 bytes).
|
## Summary table
|
||||||
2. Dispatch one compute shader.
|
|
||||||
3. Issue one `glMultiDrawElementsIndirect`.
|
|
||||||
4. Swap buffers.
|
|
||||||
|
|
||||||
That's it. The CPU frame time is essentially constant regardless of model size.
|
|
||||||
|
|
||||||
#### Future extensions (enabled by this architecture)
|
|
||||||
|
|
||||||
Once the compute-based cull pass exists, it's straightforward to add:
|
|
||||||
|
|
||||||
- **Hierarchical-Z occlusion culling**: render a coarse depth buffer from the
|
|
||||||
previous frame, then test BVH leaf AABBs against it in the compute shader.
|
|
||||||
Leaves fully behind closer geometry get culled. This handles interior-heavy
|
|
||||||
BIM models well (most rooms are occluded from any given viewpoint).
|
|
||||||
- **Distance-based LOD**: the compute shader can select different index ranges
|
|
||||||
(coarse vs. fine tessellation) per leaf based on distance to camera.
|
|
||||||
- **Contribution culling**: leaves whose screen-space projection is below a
|
|
||||||
pixel threshold get `count = 0`. Removes distant small objects.
|
|
||||||
|
|
||||||
#### Performance characteristics
|
|
||||||
|
|
||||||
| Metric | Value |
|
|
||||||
|--------|-------|
|
|
||||||
| CPU per-frame work | ~0.01 ms (constant, independent of model size) |
|
|
||||||
| GPU compute dispatch | ~0.02 ms for 2k leaves |
|
|
||||||
| Draw call overhead | 1 indirect multi-draw call |
|
|
||||||
| GPU memory overhead | ~48 bytes/leaf (AABB SSBO) + 20 bytes/leaf (indirect commands) × 2 (template + live) |
|
|
||||||
| Total for 2k leaves | ~176 KB GPU memory |
|
|
||||||
| Implementation complexity | High (compute shaders, SSBOs, memory barriers, indirect draw) |
|
|
||||||
|
|
||||||
#### When to use
|
|
||||||
|
|
||||||
Phase 3 is worthwhile when:
|
|
||||||
|
|
||||||
- The model has 500k+ objects (CPU frustum testing > 3 ms).
|
|
||||||
- Smooth 60 fps orbiting is required during interaction.
|
|
||||||
- The GPU has compute shader support (OpenGL 4.3+, which is guaranteed since
|
|
||||||
the viewer requires 4.5).
|
|
||||||
|
|
||||||
For models under 100k objects, phase 1 alone is sufficient. For 100k–500k,
|
|
||||||
phase 2 (BVH) keeps CPU culling well under 1 ms. Phase 3 is the final step
|
|
||||||
that makes the CPU frame time constant.
|
|
||||||
|
|
||||||
### Summary
|
|
||||||
|
|
||||||
```
|
```
|
||||||
Model size Active phases CPU cull cost Draw calls
|
Scene size Bottleneck Fix
|
||||||
───────────── ────────────── ────────────── ──────────
|
----------- ---------- ---
|
||||||
< 10k objects Phase 1 ~0.06 ms 1 multi-draw
|
< 100k instances CPU cull scan Phase 1 only (current)
|
||||||
10k–100k Phase 1 ~0.6 ms 1 multi-draw
|
100k–500k CPU cull scan BVH (Phase 2) — done
|
||||||
100k–500k Phase 1 + 2 ~0.01 ms 1 multi-draw
|
500k+ across many models visible/indirect Phase 3A mapped rings
|
||||||
500k–1M+ Phase 1 + 2 + 3 ~0 (GPU) 1 indirect multi-draw
|
buffer uploads (next)
|
||||||
```
|
--- --- ---
|
||||||
|
multi-million + occlusion-heavy fragment / overdraw HiZ occlusion + LOD
|
||||||
The load path:
|
|
||||||
|
|
||||||
```
|
|
||||||
open(model.ifc):
|
|
||||||
├─ sidecar exists (.ifcview)?
|
|
||||||
│ ├─ yes: background thread reads sidecar file (non-blocking I/O)
|
|
||||||
│ │ → allocate per-model VAO/VBO/EBO (empty, exact size)
|
|
||||||
│ │ → progressive GPU upload: 48 MB/frame VBO, then EBO
|
|
||||||
│ │ → objects appear as EBO chunks land
|
|
||||||
│ │ → BVH activates once fully loaded
|
|
||||||
│ │ → viewport interactive throughout
|
|
||||||
│ └─ no: stream from IFC via GeometryStreamer
|
|
||||||
│ → uploadChunk() appends to per-model buffers (immediately drawable)
|
|
||||||
│ → phase 1 linear-scan culling active from first chunk
|
|
||||||
│ → on completion: background BVH build, re-sort EBO, save .ifcview
|
|
||||||
└─ rendering (per model, per frame):
|
|
||||||
├─ phase 3 available? → compute cull + indirect multi-draw
|
|
||||||
├─ BVH available? → BVH traversal + glMultiDrawElements
|
|
||||||
└─ else / progressive → linear scan of active objects + glMultiDrawElements
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Roadmap
|
## Roadmap
|
||||||
|
|
||||||
- [x] Material color support (per-vertex RGBA8)
|
- [x] Material colour support (per-vertex RGBA8)
|
||||||
- [x] Per-model GPU buffers (VAO/VBO/EBO per model, no cross-model copies)
|
- [x] Per-model GPU buffers (VAO/VBO/EBO per model, no cross-model copies)
|
||||||
- [x] Per-object frustum culling (phase 1)
|
- [x] Per-object frustum culling (Phase 1)
|
||||||
- [x] BVH acceleration with per-model trees (phase 2)
|
- [x] BVH acceleration with per-model trees (Phase 2)
|
||||||
- [x] Raw binary `.ifcview` sidecar cache (full geometry + BVH, Blender-style)
|
- [x] Raw binary `.ifcview` sidecar cache
|
||||||
- [x] Non-blocking sidecar loading (background thread I/O)
|
- [x] Non-blocking sidecar loading (background thread I/O)
|
||||||
- [x] Progressive GPU upload (48 MB/frame chunked VBO/EBO transfer)
|
- [x] Progressive GPU upload (VBO/EBO growth + streaming-time instance appends)
|
||||||
- [ ] GPU-driven indirect draw (phase 3)
|
- [x] GPU instancing (unique meshes + per-placement SSBO)
|
||||||
|
- [x] `glMultiDrawElementsIndirect` draw path
|
||||||
|
- [x] Reflection-aware two-pass draw for mirrored placements
|
||||||
|
- [x] Backface culling (user-toggleable, default on)
|
||||||
|
- [x] `reorient-shells` enabled in iterator
|
||||||
|
- [ ] **Phase 3A — persistent-mapped ring buffers for visible + indirect** (next)
|
||||||
|
- [ ] Phase 3B — GPU-side compute-shader culling
|
||||||
|
- [ ] Screen-space contribution culling
|
||||||
- [ ] Hierarchical-Z occlusion culling
|
- [ ] Hierarchical-Z occlusion culling
|
||||||
- [ ] Distance-based LOD selection
|
- [ ] Distance-based LOD selection
|
||||||
- [ ] Vulkan/MoltenVK backend for macOS
|
- [ ] Vulkan/MoltenVK backend for macOS
|
||||||
|
|||||||
Reference in New Issue
Block a user