- cantilever_cpu_comparison.jl research demo explained
- Warning: NOT user-facing, low-level performance research
- Documents element vs nodal assembly comparison
- Results: nodal 4.7× faster, 2× fewer CG iterations
- Points users to proper examples (linear_static.jl)
- Direct use of ElementAssemblyData and NodeToElementsMap
New 125-line README documenting:
- Two main demonstrations (GPU+MPI and Krylov solver)
- Requirements (Julia 1.9+, MPI, optional CUDA)
- Key insights: type stability required for GPU/MPI/Krylov
- Nodal assembly pattern explanation
- Architecture validation (v0.5.1 vs v1.0 comparison)
- References to benchmarks and design docs
- Contributing guidelines for new demos