Comprehensive benchmark comparing three nonlinear solver strategies:
1. Traditional Newton (full Jacobian assembly + direct solve)
2. Matrix-free Newton-Krylov (GMRES, no Jacobian matrix)
3. Matrix-free with Anderson acceleration (accelerated convergence)
Problem: 3D nonlinear elasticity with cubic nonlinearity
- r(u) = K·u + β·(K·u)³ - f
- Jacobian-vector product via finite differences: J·v ≈ [r(u+ε·v) - r(u)]/ε
Key findings validated:
- Matrix-free eliminates Jacobian assembly cost
- Anderson acceleration reduces iteration count
- GPU acceleration for large problems (memory bandwidth bound)
- GMRES with adaptive tolerance (Eisenstat-Walker formula)
Includes both CPU and GPU implementations with performance comparison
showing memory usage, iteration counts, and wall-clock times for systems
ranging from 1K to 1M DOFs (835 lines, full implementation).
Complete execution output from material_models_benchmark.jl validation:
Performance results:
- Linear Elastic: 5.1× speedup (Tensors.jl vs Voigt/Dict)
- Neo-Hookean Manual: 2.0× speedup over old approach
- Perfect Plasticity: 21.0× speedup (zero allocations vs Dict)
- Average speedup: 9.4× (validates 5-50× claim range)
Key validation:
- All new implementations: ZERO allocations (confirmed)
- Manual derivatives: 21.1× faster than automatic differentiation
- Type stability: All @code_warntype checks pass (no red flags)
- AbstractMaterialState hierarchy: State handling identical for all materials
Demonstrates Newton iteration state handling for both stateless (LinearElastic,
NoState) and stateful (PerfectPlasticity, PlasticityState) materials.
Compares 5 strategies for accessing integration points during assembly:
- OLD: Runtime dispatch + mutable struct with Dict (type-unstable)
- Option A: Compile-time function returning tuples
- Option B: Store in element as NTuple (current approach)
- Option C: Compile-time with Vec{D} from Tensors.jl
- Option D: Pre-computed global constants
- Option E: Function returning pre-computed constant
Validates that compile-time generation (Options C-E) matches the golden
standard architecture from nodal assembly demos. Includes realistic FEM
assembly comparison showing performance difference between old runtime
dispatch and new compile-time approach.
Recommendation: Option C (compile-time with Vec) or D/E (pre-computed)
for zero allocation and full inlining, matching eval_basis! pattern.
Compares two state update strategies for GPU optimization:
- Strategy 1: Array of Structs (AoS) - immutable elements with embedded state
- Strategy 2: Structure of Arrays (SoA) - separate geometry and mutable state
Validates that SoA achieves 5-10× better memory bandwidth due to coalesced
access patterns. Benchmarks both CPU and GPU implementations with detailed
performance metrics including bandwidth utilization.
Key findings:
- SoA enables coalesced memory access (consecutive threads → consecutive memory)
- AoS suffers from pointer chasing and non-coalesced access
- SoA has zero allocations (in-place updates vs element reconstruction)
- Tests with 10k, 100k, 1M elements showing scalability
- Validates zero-allocation claim for compute_deformation_gradient()
- Allocation analysis with @allocated macro
- Performance benchmarking with BenchmarkTools
- LLVM IR analysis for optimization verification
- Tests finite strain formulation F = I + ∇u
- Hex8 element with 10% stretch in x-direction
- Jacobian computation J = ∑ X_i ⊗ dN_i/dξ
- Measures median time in nanoseconds/microseconds
- Confirms type stability and inlineability
- 240 lines analyzing deformation gradient computation performance
- Tests 1 to 5000 fields to find crossover point
- Confirms stack copying is O(n) at 0.16 ns/field
- Confirms Dict mutation is O(1) at 7 ns constant
- Crossover at 100 fields (800 bytes) for updates
- Typical FEM elements (20-60 fields) well below crossover
- Immutable wins for access and iteration at ALL sizes
- Generates 5 publication-quality plots
- Exports JSON + CSV with system specs
- System: Intel Xeon Gold 6326, 32 cores, 503 GB RAM
Created comprehensive documentation and benchmark demonstrating why immutable
elements with type-stable fields are 40-130x faster than mutable Dict-based
elements.
benchmarks/element_immutability_benchmark.jl:
- Compares mutable (Dict) vs immutable (NamedTuple) implementations
- Measures field access, updates, assembly loops, large-scale meshes
- Results: 40x faster field access, 130x faster assembly, zero allocations
docs/design/IMMUTABILITY.md:
- Explains counterintuitive API change: element = update(element, ...)
- Benchmarks show 40-130x speedup despite 'copying' elements
- Key insight: Type stability >> mutation, compiler optimizes away copies
- Migration guide: old mutable API → new immutable API
- GPU/HPC rationale: Only bits types work on GPU (no pointers)
Key Results:
- Field access: 1ns vs 45ns (40x faster)
- Assembly: 9ns vs 1124ns per element (130x faster)
- Large mesh: 0.01ms vs 1.2ms for 1000 elements (120x faster)
- Memory: 0 allocations vs 70,000 allocations
- GPU: Compatible (bits types) vs Incompatible (pointers)
This documents a fundamental architectural decision for JuliaFEM 1.0.
RESEARCH QUESTION: Should JuliaFEM use hand-calculated derivatives or AD?
Created comprehensive benchmark comparing:
- Manual: Hand-calculated derivatives (traditional FEM)
- AD: Tensors.jl gradient() (automatic differentiation)
RESULTS (AMD Ryzen 9, Julia 1.12.1):
- Manual: 8.7 ns, 0 allocations
- AD: 268.1 ns, 0 allocations
- AD is 30× SLOWER than manual
KEY FINDINGS:
✅ Both achieve zero allocations (Tensors.jl is well-optimized)
❌ AD has 30× compute overhead from dual number arithmetic
⚠️ In assembly loops: millions of calls = 10+ seconds extra per solve
RECOMMENDATION:
- Keep manual derivatives for common elements (Tet10, Hex8, Quad4, etc.)
- Use AD for prototyping and rare elements
- Unit test manual vs AD to catch errors
- Future: Generate derivatives symbolically (Symbolics.jl)
WHY NOT AD EVERYWHERE?
Assembly is hottest path in FEM. 30× overhead = unacceptable for
production code. Users will notice the performance difference.
WHY NOT ABANDON AD?
- Excellent for prototyping
- Required for exotic bases (NURBS)
- Perfect for unit testing manual derivatives
- Zero allocations impressive
Files:
- benchmarks/tet10_derivatives_benchmark.jl (runnable benchmark)
- docs/benchmarks/shape_function_derivatives_ad_vs_manual.md (analysis)
Dependencies added: BenchmarkTools
This answers the research question definitively with data.