chore: remove benchmarks directory

Remove benchmarks directory as it has been moved to another
repository

- Remove benchmarks/ directory and all contents
- 37 files removed including benchmark scripts and results
This commit is contained in:
Jukka Aho
2025-12-12 21:41:48 +02:00
parent 1cbc236f89
commit c4745e8cd1
37 changed files with 0 additions and 9727 deletions
-244
View File
@@ -1,244 +0,0 @@
# GPU Benchmark Suite
Comprehensive benchmarks demonstrating GPU-friendly FEM architecture strategies.
## Prerequisites
```bash
# Add required packages
julia --project=. -e 'using Pkg; Pkg.add(["CUDA", "Tensors", "BenchmarkTools", "IterativeSolvers"])'
```
## Benchmarks
### 1. State Management Strategy Comparison
**File:** `gpu_state_management_benchmark.jl`
**What it tests:**
- **Strategy 1:** Immutable elements (Array of Structs - AoS)
- **Strategy 2:** Separate mutable state (Structure of Arrays - SoA)
**Metrics:**
- Memory bandwidth (GB/s)
- Execution time
- Allocation counts
**Run:**
```bash
julia --project=. benchmarks/gpu_state_management_benchmark.jl
```
**Expected Results:**
- CPU: Strategy 2 is 3-5× faster (better cache utilization)
- GPU: Strategy 2 is 5-10× faster (coalesced memory access)
- Strategy 2 bandwidth: 500-900 GB/s (depending on GPU)
- Strategy 1 bandwidth: 50-150 GB/s (non-coalesced)
### 2. Matrix-Free Newton-Krylov
**File:** `matrix_free_gpu_benchmark.jl`
**What it tests:**
- **Traditional Newton:** Full Jacobian assembly + direct solve
- **Matrix-Free NK:** Jacobian-free with GMRES
- **Anderson-Accelerated:** Matrix-Free + Anderson acceleration
**Metrics:**
- Total time
- Iterations to convergence
- Time per iteration
**Run:**
```bash
julia --project=. benchmarks/matrix_free_gpu_benchmark.jl
```
**Expected Results:**
- Matrix-Free: 2-4× faster than traditional (no assembly)
- Anderson: 2-3× fewer iterations (superlinear convergence)
- GPU: Additional 3-10× speedup for large problems (>10K DOFs)
- **Total:** 5-10× speedup with all optimizations
## Understanding the Results
### State Management Benchmark Output
```text
GPU Benchmark: 1000000 elements
====================================================================
📊 Strategy 1 (Immutable Elements - AoS on GPU):
Time: 12.456 ms
Bandwidth: 87.3 GB/s
📊 Strategy 2 (Separate State - SoA on GPU):
Time: 1.234 ms
Bandwidth: 743.2 GB/s
✅ GPU Speedup (Strategy 2 / Strategy 1): 10.09×
✅ Bandwidth Improvement: 8.51×
Strategy 1: 87.3 GB/s (non-coalesced)
Strategy 2: 743.2 GB/s (coalesced)
```
**Interpretation:**
- Strategy 2 achieves 8-10× higher bandwidth
- Approaching GPU memory bandwidth limit (~1000 GB/s for high-end GPUs)
- Coalesced memory access is critical for GPU performance
### Matrix-Free Benchmark Output
```text
CPU Benchmark: 10000 DOFs
====================================================================
📊 Traditional Newton (Full Jacobian):
Time: 456.78 ms
Iterations: 8
Time/iter: 57.10 ms
📊 Matrix-Free Newton-Krylov:
Time: 123.45 ms
Iterations: 10
Time/iter: 12.35 ms
📊 Anderson-Accelerated Matrix-Free:
Time: 67.89 ms
Iterations: 5
Time/iter: 13.58 ms
✅ CPU Speedups:
Matrix-Free vs Traditional: 3.70×
Anderson vs Traditional: 6.73×
Anderson vs Matrix-Free: 1.82×
```
**Interpretation:**
- Matrix-Free eliminates expensive Jacobian assembly
- Anderson reduces Newton iterations (superlinear convergence)
- Combined: 6-10× speedup on CPU, more on GPU
## Hardware Requirements
### Minimum
- CPU: x86_64 with AVX2
- RAM: 8 GB
- Julia: 1.9+
### Recommended for GPU Benchmarks
- GPU: NVIDIA with CUDA Compute Capability 7.0+ (RTX 20XX or newer)
- VRAM: 4 GB+
- CUDA: 11.0+
### Tested On
- NVIDIA RTX 4090 (24 GB VRAM)
- NVIDIA RTX 3090 (24 GB VRAM)
- NVIDIA RTX 3080 (10 GB VRAM)
## Troubleshooting
### "CUDA not available"
If you see this warning, the benchmark will run CPU-only comparison:
```julia
⚠️ CUDA not available! Running CPU-only comparison.
```
**Solution:**
1. Check GPU is recognized: `nvidia-smi`
2. Verify CUDA installation: `julia -e 'using CUDA; CUDA.versioninfo()'`
3. Rebuild CUDA.jl: `julia --project=. -e 'using Pkg; Pkg.build("CUDA")'`
### Out of Memory Errors
If benchmarks crash with OOM:
```julia
ERROR: Out of memory
```
**Solution:**
- Reduce problem sizes in benchmark scripts
- Edit `sizes = [10_000, 100_000, 1_000_000]` → smaller values
- Close other GPU applications
### Slow CPU Benchmarks
Matrix-Free benchmark may be slow on CPU for large problems (>100K DOFs).
**Solution:**
- Use smaller test sizes for CPU
- Focus on GPU results for large problems
- Enable BLAS threading: `export JULIA_NUM_THREADS=8`
## Performance Expectations
### State Management (1M elements)
| Strategy | CPU Time | GPU Time | GPU Bandwidth |
|----------|----------|----------|---------------|
| Strategy 1 (AoS) | 18 ms | 12 ms | 80-150 GB/s |
| Strategy 2 (SoA) | 5 ms | 1.2 ms | 500-900 GB/s |
| **Speedup** | **3.6×** | **10×** | **6-10×** |
### Matrix-Free Newton-Krylov (100K DOFs)
| Method | CPU Time | GPU Time | Iterations |
|--------|----------|----------|------------|
| Traditional Newton | 8200 ms | N/A | 20 |
| Matrix-Free NK | 2100 ms | 450 ms | 20 |
| MF-NK + Anderson | 1050 ms | 180 ms | 8 |
| **Speedup** | **7.8×** | **45×** | **2.5× fewer** |
## Validation
Both benchmarks validate correctness:
1. **State Management:**
- Verifies final state matches between strategies
- Checks zero allocations for Strategy 2
2. **Matrix-Free:**
- Compares solution accuracy (||u_traditional - u_matrixfree|| < 1e-6)
- Validates convergence to same residual norm
## Citation
If you use these benchmarks in research, please cite:
```bibtex
@software{juliafem2025,
title = {JuliaFEM: GPU-Accelerated Finite Element Method},
author = {Aho, Jukka},
year = {2025},
url = {https://github.com/JuliaFEM/JuliaFEM.jl}
}
```
## References
1. **Knoll & Keyes (2004):** "Jacobian-free NewtonKrylov methods"
2. **Walker & Ni (2011):** "Anderson acceleration for fixed-point iterations"
3. **CUDA Programming Guide:** <https://docs.nvidia.com/cuda/>
## Contact
Questions or issues? Open an issue at: <https://github.com/JuliaFEM/JuliaFEM.jl/issues>
-85
View File
@@ -1,85 +0,0 @@
# Benchmark Validation Results
**Date:** November 9, 2025
**Platform:** Julia 1.12.1
**Document:** `docs/book/zero_allocation_fields.md`
**Benchmark:** `benchmarks/field_storage_comparison.jl`
## Summary
**All performance claims validated**
The zero-allocation field storage design achieves **9-92× speedup** over `Dict{String,Any}` with **zero allocations in hot paths**.
## Measured Results
| Test | OLD (Dict) | NEW (Typed) | Speedup |
|------|------------|-------------|---------|
| Constant field access | 19.2ns, 0 allocs | 2.1ns, 0 allocs | **9×** |
| Nodal field access | 262ns, 3 allocs | 6.5ns, 0 allocs | **40×** |
| Interpolation (uncached) | 2.6μs, 50 allocs | 44ns, 2 allocs | **59×** |
| Interpolation (cached) | 2.6μs, 50 allocs | 53ns, **0 allocs** ✅ | **49×** |
| Assembly (1000 elem) | 109μs, 4000 allocs | 1.2μs, **0 allocs** ✅ | **92×** |
## Key Achievements
1.**Zero allocations** in cached interpolation (53ns)
2.**Zero allocations** in assembly loop (1.2μs vs 109μs)
3.**Type stability** eliminates runtime dispatch
4.**9-92× speedup** across all operations
5.**Simple implementation** (~200 LOC for field types)
## Design Validated
The `NamedTuple` + typed field structs approach is proven effective:
```julia
# Simple field types
struct ConstantField{T}
value::T
end
struct NodalField{T}
values::Matrix{T}
end
# Type-stable container
fields = (
youngs_modulus = ConstantField(210e3),
displacement = NodalField(zeros(3, 1000)),
)
# Fast access (zero allocations)
E = fields.youngs_modulus.value # 2.1ns, 0 allocs
u = @view fields.displacement.values[:, nodes] # 6.5ns, 0 allocs
```
## Claims Verification
| Claim | Measured | Status |
|-------|----------|--------|
| 50× faster | 9-92× across operations | ✅ VALIDATED |
| 0 allocations | 0 allocs in hot paths | ✅ VALIDATED |
| Type stability | No runtime dispatch | ✅ VALIDATED |
| Simple implementation | ~200 LOC field types | ✅ VALIDATED |
## Reproduction
```bash
cd /home/juajukka/dev/JuliaFEM.jl
julia --project=. benchmarks/field_storage_comparison.jl
```
## Next Steps
1. ✅ Document written and validated
2. ⏭️ Implement field types in `src/fields/types.jl`
3. ⏭️ Update `Element` struct for `ElementSet` pattern
4. ⏭️ Add CI benchmarks to prevent regression
5. ⏭️ Migrate examples to new field system
## Conclusion
The zero-allocation field storage design is **ready for v1.0 implementation**. Measured performance exceeds targets with 9-92× speedup and zero allocations in hot paths.
**Design Decision:** Use `NamedTuple` of typed field structs for JuliaFEM v1.0
@@ -1,412 +0,0 @@
# Benchmark: Basis Function Access Patterns for Tet10 (Realistic 3D Case)
#
# Focus: 10-node quadratic tetrahedron (Tet10) - the workhorse for 3D simulations
#
# Goal: Find the fastest way to access basis functions and their DERIVATIVES with:
# 1. Return all 10 basis functions as tuple (zero allocation)
# 2. Return single basis function by index (must be inlineable)
# 3. Return all 10 derivatives as tuple of Vec{3} (zero allocation)
# 4. Return single derivative by index (must be inlineable)
# 5. Pass topology separately (separation of concerns)
# 6. Must be type-stable and superfast
#
# Use Case:
# - Stiffness matrix assembly: Need derivatives (B matrix construction)
# - Mass matrix assembly: Need basis functions (M matrix construction)
# - Nodal assembly: Need single basis function/derivative at a time
#
# Usage: julia --project=. benchmarks/basis_function_access_patterns.jl
using BenchmarkTools
using Tensors
# ============================================================================
# Define minimal types for testing
# ============================================================================
abstract type AbstractTopology end
struct Tetrahedron <: AbstractTopology end
abstract type AbstractBasis end
struct Lagrange{P} <: AbstractBasis end
# Vec type comes from Tensors.jl (already available in JuliaFEM)
# Vec{3,Float64} for 3D gradients
# ============================================================================
# Strategy 1: Return tuple, index with getindex
# ============================================================================
# Advantages: Natural Julia syntax, type-stable
# Disadvantages: Might not inline getindex?
"""
Get all basis functions for Triangle, P1 Lagrange (3 functions).
Returns tuple of 3 Float64 values.
"""
@inline function get_basis_functions_v1(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
u, v = xi
# P1 triangle: N1 = 1-u-v, N2 = u, N3 = v
return (1 - u - v, u, v)
end
"""
Get single basis function by index (1-based).
"""
@inline function get_basis_function_v1(topology::Triangle, basis::Lagrange{1}, xi::Vec{2,T}, i::Int) where T
N_all = get_basis_functions_v1(topology, basis, xi)
return N_all[i] # Tuple indexing
end
# ============================================================================
# Strategy 2: Generated function for single access
# ============================================================================
# Advantages: Compiler can specialize for each index
# Disadvantages: More complex code
@inline function get_basis_functions_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
u, v = xi
return (1 - u - v, u, v)
end
"""
Use @generated to create specialized code for each index at compile time.
"""
@generated function get_basis_function_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}, ::Val{I}) where {T,I}
if I == 1
return :(1 - xi[1] - xi[2])
elseif I == 2
return :(xi[1])
elseif I == 3
return :(xi[2])
else
return :(error("Invalid basis function index: $I"))
end
end
# ============================================================================
# Strategy 3: Manual dispatch on Val (type-stable index)
# ============================================================================
# Advantages: Explicit, clear what's happening
# Disadvantages: Verbose, need to write each case
@inline function get_basis_functions_v3(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
u, v = xi
return (1 - u - v, u, v)
end
@inline get_basis_function_v3(t::Triangle, b::Lagrange{1}, xi::Vec{2,T}, ::Val{1}) where T = 1 - xi[1] - xi[2]
@inline get_basis_function_v3(t::Triangle, b::Lagrange{1}, xi::Vec{2,T}, ::Val{2}) where T = xi[1]
@inline get_basis_function_v3(t::Triangle, b::Lagrange{1}, xi::Vec{2,T}, ::Val{3}) where T = xi[2]
# ============================================================================
# Strategy 4: Struct with getindex (most Julian)
# ============================================================================
# Advantages: Can use N[i] syntax naturally
# Disadvantages: Extra struct allocation?
struct BasisFunctions{N,T}
data::NTuple{N,T}
end
Base.@propagate_inbounds Base.getindex(bf::BasisFunctions, i::Int) = bf.data[i]
Base.length(::BasisFunctions{N}) where N = N
@inline function get_basis_functions_v4(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
u, v = xi
return BasisFunctions((1 - u - v, u, v))
end
# Can use natural indexing
@inline function get_basis_function_v4(topology::Triangle, basis::Lagrange{1}, xi::Vec{2,T}, i::Int) where T
N = get_basis_functions_v4(topology, basis, xi)
return N[i]
end
# ============================================================================
# Strategy 5: Separate implementation per basis function (extreme specialization)
# ============================================================================
# Advantages: Maximum performance, no tuple allocation at all
# Disadvantages: Lots of code duplication
@inline function get_basis_function_1_v5(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
return 1 - xi[1] - xi[2]
end
@inline function get_basis_function_2_v5(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
return xi[1]
end
@inline function get_basis_function_3_v5(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
return xi[2]
end
@inline function get_basis_functions_v5(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
u, v = xi
return (1 - u - v, u, v)
end
# ============================================================================
# Benchmark: Access all basis functions (typical in assembly loop)
# ============================================================================
function benchmark_all_access()
println("\n" * "="^80)
println("BENCHMARK: Access ALL basis functions")
println("="^80)
topology = Triangle()
basis = Lagrange{1}()
xi = Vec(0.25, 0.25)
println("\nStrategy 1: Tuple return + getindex")
@btime get_basis_functions_v1($topology, $basis, $xi)
println("\nStrategy 2: Generated function")
@btime get_basis_functions_v2($topology, $basis, $xi)
println("\nStrategy 3: Val dispatch")
@btime get_basis_functions_v3($topology, $basis, $xi)
println("\nStrategy 4: BasisFunctions struct")
@btime get_basis_functions_v4($topology, $basis, $xi)
println("\nStrategy 5: Separate functions")
@btime get_basis_functions_v5($topology, $basis, $xi)
# Verify all return same values
r1 = get_basis_functions_v1(topology, basis, xi)
r2 = get_basis_functions_v2(topology, basis, xi)
r3 = get_basis_functions_v3(topology, basis, xi)
r4 = get_basis_functions_v4(topology, basis, xi).data
r5 = get_basis_functions_v5(topology, basis, xi)
@assert r1 == r2 == r3 == r4 == r5 "Results don't match!"
println("\n✓ All strategies return identical values: $r1")
end
# ============================================================================
# Benchmark: Access SINGLE basis function (for nodal assembly)
# ============================================================================
function benchmark_single_access()
println("\n" * "="^80)
println("BENCHMARK: Access SINGLE basis function (nodal assembly)")
println("="^80)
topology = Triangle()
basis = Lagrange{1}()
xi = Vec(0.25, 0.25)
println("\nStrategy 1: Tuple + runtime index")
@btime get_basis_function_v1($topology, $basis, $xi, 2)
println("\nStrategy 2: Generated function with Val{2}")
@btime get_basis_function_v2($topology, $basis, $xi, Val(2))
println("\nStrategy 3: Val dispatch")
@btime get_basis_function_v3($topology, $basis, $xi, Val(2))
println("\nStrategy 4: BasisFunctions struct + index")
@btime get_basis_function_v4($topology, $basis, $xi, 2)
println("\nStrategy 5: Direct function call")
@btime get_basis_function_2_v5($topology, $basis, $xi)
# Verify all return same value
r1 = get_basis_function_v1(topology, basis, xi, 2)
r2 = get_basis_function_v2(topology, basis, xi, Val(2))
r3 = get_basis_function_v3(topology, basis, xi, Val(2))
r4 = get_basis_function_v4(topology, basis, xi, 2)
r5 = get_basis_function_2_v5(topology, basis, xi)
@assert r1 == r2 == r3 == r4 == r5 "Results don't match!"
println("\n✓ All strategies return identical value: $r1")
end
# ============================================================================
# Benchmark: Typical assembly loop pattern
# ============================================================================
function benchmark_assembly_loop()
println("\n" * "="^80)
println("BENCHMARK: Typical assembly loop (iterate over all basis functions)")
println("="^80)
topology = Triangle()
basis = Lagrange{1}()
xi = Vec(0.25, 0.25)
# Pattern 1: Get all, iterate over tuple
println("\nPattern 1: Get all as tuple, iterate")
function assemble_v1()
N_all = get_basis_functions_v1(topology, basis, xi)
s = 0.0
for N_i in N_all
s += N_i * N_i # Dummy computation
end
return s
end
@btime assemble_v1()
# Pattern 2: Get all, index in loop
println("\nPattern 2: Get all, index with i")
function assemble_v2()
N_all = get_basis_functions_v1(topology, basis, xi)
s = 0.0
for i in 1:3
s += N_all[i] * N_all[i]
end
return s
end
@btime assemble_v2()
# Pattern 3: Get one at a time (nodal assembly style)
println("\nPattern 3: Get one at a time with Val")
function assemble_v3()
s = 0.0
# Unrolled loop (what compiler would do with Val)
N1 = get_basis_function_v3(topology, basis, xi, Val(1))
s += N1 * N1
N2 = get_basis_function_v3(topology, basis, xi, Val(2))
s += N2 * N2
N3 = get_basis_function_v3(topology, basis, xi, Val(3))
s += N3 * N3
return s
end
@btime assemble_v3()
# Pattern 4: Direct function calls (strategy 5)
println("\nPattern 4: Direct function calls (extreme specialization)")
function assemble_v4()
s = 0.0
N1 = get_basis_function_1_v5(topology, basis, xi)
s += N1 * N1
N2 = get_basis_function_2_v5(topology, basis, xi)
s += N2 * N2
N3 = get_basis_function_3_v5(topology, basis, xi)
s += N3 * N3
return s
end
@btime assemble_v4()
# Verify all compute same result
r1 = assemble_v1()
r2 = assemble_v2()
r3 = assemble_v3()
r4 = assemble_v4()
@assert r1 == r2 == r3 == r4 "Assembly results don't match!"
println("\n✓ All patterns compute same result: $r1")
end
# ============================================================================
# Benchmark: Basis derivatives (return Vec)
# ============================================================================
function benchmark_derivatives()
println("\n" * "="^80)
println("BENCHMARK: Basis function DERIVATIVES (return Vec)")
println("="^80)
topology = Triangle()
basis = Lagrange{1}()
xi = Vec(0.25, 0.25)
# Triangle P1 derivatives (constant):
# dN1/d(u,v) = (-1, -1)
# dN2/d(u,v) = (1, 0)
# dN3/d(u,v) = (0, 1)
println("\nStrategy 1: Return tuple of Vecs")
@inline function get_basis_derivatives_v1(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
return (Vec(-1.0, -1.0), Vec(1.0, 0.0), Vec(0.0, 1.0))
end
@btime get_basis_derivatives_v1($topology, $basis, $xi)
println("\nStrategy 2: Return single Vec with Val indexing")
@inline get_basis_derivative_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}, ::Val{1}) where T = Vec(-1.0, -1.0)
@inline get_basis_derivative_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}, ::Val{2}) where T = Vec(1.0, 0.0)
@inline get_basis_derivative_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}, ::Val{3}) where T = Vec(0.0, 1.0)
@btime get_basis_derivative_v2($topology, $basis, $xi, Val(2))
# Verify
all_derivs = get_basis_derivatives_v1(topology, basis, xi)
single_deriv = get_basis_derivative_v2(topology, basis, xi, Val(2))
@assert all_derivs[2] == single_deriv
println("\n✓ Derivatives match: $single_deriv")
end
# ============================================================================
# Main execution
# ============================================================================
function main()
println("\n")
println("" * "="^78 * "")
println("" * " "^78 * "")
println("" * " "^20 * "BASIS FUNCTION ACCESS PATTERNS BENCHMARK" * " "^18 * "")
println("" * " "^78 * "")
println("" * "="^78 * "")
println("\nGoal: Find fastest way to access basis functions for nodal assembly")
println("Requirements:")
println(" - Zero allocation")
println(" - Type stable")
println(" - Inlineable")
println(" - Support both 'all at once' and 'one at a time' access")
benchmark_all_access()
benchmark_single_access()
benchmark_assembly_loop()
benchmark_derivatives()
println("\n" * "="^80)
println("SUMMARY & RECOMMENDATIONS")
println("="^80)
println("""
For TRADITIONAL ASSEMBLY (get all basis functions at integration point):
→ Use Strategy 1 or 4: Simple tuple return
→ Should be 0-5 ns, zero allocation
For NODAL ASSEMBLY (get single basis function):
→ Use Strategy 3: Val dispatch for compile-time index
→ Should be 0-2 ns, zero allocation, fully inlined
→ Usage: get_basis_function(Triangle(), Lagrange{1}(), xi, Val(i))
For DERIVATIVES:
→ Return tuple of Vec for all derivatives
→ Use Val indexing for single derivative
→ Same performance as basis functions
RECOMMENDED API:
```julia
# Get all basis functions (returns tuple)
N_all = get_basis_functions(Triangle(), Lagrange{1}(), xi)
# Get single basis function (Val for compile-time specialization)
N_i = get_basis_function(Triangle(), Lagrange{1}(), xi, Val(i))
# Get all derivatives (returns tuple of Vec)
dN_all = get_basis_derivatives(Triangle(), Lagrange{1}(), xi)
# Get single derivative (returns Vec)
dN_i = get_basis_derivative(Triangle(), Lagrange{1}(), xi, Val(i))
```
WHY Val?
- Compiler knows index at compile time
- Can generate optimal code for each basis function
- Zero runtime overhead
- Type stable
NOTE: For runtime indexing (i not known at compile time), tuple indexing
is still very fast (typically 1-2 ns overhead).
""")
println("\n" * "="^80)
end
# Run benchmarks
if abspath(PROGRAM_FILE) == @__FILE__
main()
end
-504
View File
@@ -1,504 +0,0 @@
# Benchmark: Basis Function Access Patterns for Tet10 (Realistic 3D Case)
#
# Focus: 10-node quadratic tetrahedron (Tet10) - the workhorse for 3D simulations
#
# Goal: Find the fastest way to access basis functions and their DERIVATIVES with:
# 1. Return all 10 basis functions as tuple (zero allocation)
# 2. Return single basis function by index (must be inlineable)
# 3. Return all 10 derivatives as tuple of Vec{3} (zero allocation)
# 4. Return single derivative by index (must be inlineable)
# 5. Pass topology separately (separation of concerns: Element(Tetrahedron, Lagrange{2}, ...))
# 6. Must be type-stable and superfast
#
# Use Case:
# - Stiffness matrix assembly: Need derivatives (B matrix construction)
# - Mass matrix assembly: Need basis functions (M matrix construction)
# - Nodal assembly: Need single basis function/derivative at a time
#
# Usage: julia --project=. benchmarks/basis_function_access_tet10.jl
using BenchmarkTools
using Tensors
# ============================================================================
# Define minimal types for testing
# ============================================================================
abstract type AbstractTopology end
struct Tetrahedron <: AbstractTopology end
abstract type AbstractBasis end
struct Lagrange{P} <: AbstractBasis end
# ============================================================================
# Tet10 Basis Functions (Quadratic Tetrahedron, 10 nodes)
# ============================================================================
# Node numbering:
# 1-4: vertices
# 5-10: edge midpoints (5: 1-2, 6: 2-3, 7: 3-1, 8: 1-4, 9: 2-4, 10: 3-4)
#
# Parametric coordinates: (u, v, w) where u+v+w ≤ 1
# Reference element: vertices at (0,0,0), (1,0,0), (0,1,0), (0,0,1)
"""
Get all 10 basis functions for Tet10 at parametric point (u,v,w).
Returns NTuple{10, Float64}.
The basis functions are:
- N1 = (1-u-v-w)(1-2u-2v-2w) [vertex 1]
- N2 = u(2u-1) [vertex 2]
- N3 = v(2v-1) [vertex 3]
- N4 = w(2w-1) [vertex 4]
- N5 = 4u(1-u-v-w) [edge 1-2]
- N6 = 4uv [edge 2-3]
- N7 = 4v(1-u-v-w) [edge 3-1]
- N8 = 4w(1-u-v-w) [edge 1-4]
- N9 = 4uw [edge 2-4]
- N10= 4vw [edge 3-4]
"""
@inline function get_basis_functions(::Tetrahedron, ::Lagrange{2}, xi::Vec{3,T}) where T
u, v, w = xi
λ = 1 - u - v - w # barycentric coordinate for vertex 1
# Vertex nodes (1-4)
N1 = λ * (2λ - 1)
N2 = u * (2u - 1)
N3 = v * (2v - 1)
N4 = w * (2w - 1)
# Edge midpoint nodes (5-10)
N5 = 4 * u * λ
N6 = 4 * u * v
N7 = 4 * v * λ
N8 = 4 * w * λ
N9 = 4 * u * w
N10 = 4 * v * w
return (N1, N2, N3, N4, N5, N6, N7, N8, N9, N10)
end
"""
Get all 10 basis function derivatives for Tet10.
Returns NTuple{10, Vec{3,Float64}}.
Each derivative is ∇N_i = (∂N_i/∂u, ∂N_i/∂v, ∂N_i/∂w)
"""
@inline function get_basis_derivatives(::Tetrahedron, ::Lagrange{2}, xi::Vec{3,T}) where T
u, v, w = xi
λ = 1 - u - v - w
# Derivatives of vertex nodes
dN1 = Vec(-3 + 4u + 4v + 4w, -3 + 4u + 4v + 4w, -3 + 4u + 4v + 4w)
dN2 = Vec(4u - 1, 0.0, 0.0)
dN3 = Vec(0.0, 4v - 1, 0.0)
dN4 = Vec(0.0, 0.0, 4w - 1)
# Derivatives of edge midpoint nodes
dN5 = Vec(4λ - 4u, -4u, -4u)
dN6 = Vec(4v, 4u, 0.0)
dN7 = Vec(-4v, 4λ - 4v, -4v)
dN8 = Vec(-4w, -4w, 4λ - 4w)
dN9 = Vec(4w, 0.0, 4u)
dN10 = Vec(0.0, 4w, 4v)
return (dN1, dN2, dN3, dN4, dN5, dN6, dN7, dN8, dN9, dN10)
end
# ============================================================================
# Strategy 1: Tuple indexing (runtime index)
# ============================================================================
@inline function get_basis_function_v1(topology::Tetrahedron, basis::Lagrange{2},
xi::Vec{3,T}, i::Int) where T
N_all = get_basis_functions(topology, basis, xi)
return N_all[i]
end
@inline function get_basis_derivative_v1(topology::Tetrahedron, basis::Lagrange{2},
xi::Vec{3,T}, i::Int) where T
dN_all = get_basis_derivatives(topology, basis, xi)
return dN_all[i]
end
# ============================================================================
# Strategy 2: Val dispatch (compile-time index)
# ============================================================================
@inline function get_basis_function_v2(t::Tetrahedron, b::Lagrange{2},
xi::Vec{3,T}, ::Val{I}) where {T,I}
N_all = get_basis_functions(t, b, xi)
return N_all[I]
end
@inline function get_basis_derivative_v2(t::Tetrahedron, b::Lagrange{2},
xi::Vec{3,T}, ::Val{I}) where {T,I}
dN_all = get_basis_derivatives(t, b, xi)
return dN_all[I]
end
# ============================================================================
# Strategy 3: Generated function (compute only requested basis function)
# ============================================================================
@generated function get_basis_function_v3(::Tetrahedron, ::Lagrange{2},
xi::Vec{3,T}, ::Val{I}) where {T,I}
# Generate specialized code for each index
if I == 1
return quote
u, v, w = xi
λ = 1 - u - v - w
return λ * (2λ - 1)
end
elseif I == 2
return quote
u = xi[1]
return u * (2u - 1)
end
elseif I == 3
return quote
v = xi[2]
return v * (2v - 1)
end
elseif I == 4
return quote
w = xi[3]
return w * (2w - 1)
end
elseif I == 5
return quote
u, v, w = xi
λ = 1 - u - v - w
return 4 * u * λ
end
elseif I == 6
return quote
u, v = xi[1], xi[2]
return 4 * u * v
end
elseif I == 7
return quote
u, v, w = xi
λ = 1 - u - v - w
return 4 * v * λ
end
elseif I == 8
return quote
u, v, w = xi
λ = 1 - u - v - w
return 4 * w * λ
end
elseif I == 9
return quote
u, w = xi[1], xi[3]
return 4 * u * w
end
elseif I == 10
return quote
v, w = xi[2], xi[3]
return 4 * v * w
end
else
return :(error("Invalid basis function index: $I for Tet10"))
end
end
@generated function get_basis_derivative_v3(::Tetrahedron, ::Lagrange{2},
xi::Vec{3,T}, ::Val{I}) where {T,I}
if I == 1
return quote
u, v, w = xi
return Vec(-3 + 4u + 4v + 4w, -3 + 4u + 4v + 4w, -3 + 4u + 4v + 4w)
end
elseif I == 2
return quote
u = xi[1]
return Vec(4u - 1, 0.0, 0.0)
end
elseif I == 3
return quote
v = xi[2]
return Vec(0.0, 4v - 1, 0.0)
end
elseif I == 4
return quote
w = xi[3]
return Vec(0.0, 0.0, 4w - 1)
end
elseif I == 5
return quote
u, v, w = xi
λ = 1 - u - v - w
return Vec(4λ - 4u, -4u, -4u)
end
elseif I == 6
return quote
u, v = xi[1], xi[2]
return Vec(4v, 4u, 0.0)
end
elseif I == 7
return quote
u, v, w = xi
λ = 1 - u - v - w
return Vec(-4v, 4λ - 4v, -4v)
end
elseif I == 8
return quote
u, v, w = xi
λ = 1 - u - v - w
return Vec(-4w, -4w, 4λ - 4w)
end
elseif I == 9
return quote
u, w = xi[1], xi[3]
return Vec(4w, 0.0, 4u)
end
elseif I == 10
return quote
v, w = xi[2], xi[3]
return Vec(0.0, 4w, 4v)
end
else
return :(error("Invalid basis function index: $I for Tet10"))
end
end
# ============================================================================
# Benchmark Functions
# ============================================================================
function benchmark_all_basis_functions()
println("\n" * "="^80)
println("BENCHMARK 1: Get ALL 10 basis functions")
println("="^80)
println("Use case: Mass matrix assembly, need all N_i at integration point")
topology = Tetrahedron()
basis = Lagrange{2}()
xi = Vec(0.25, 0.25, 0.2) # Typical integration point
println("\nAccess all 10 basis functions:")
@btime get_basis_functions($topology, $basis, $xi)
result = get_basis_functions(topology, basis, xi)
println("\n✓ Result (10 values): ", result)
println("✓ Sum of basis functions (partition of unity): ", sum(result))
@assert abs(sum(result) - 1.0) < 1e-10 "Partition of unity violated!"
end
function benchmark_single_basis_function()
println("\n" * "="^80)
println("BENCHMARK 2: Get SINGLE basis function (nodal assembly)")
println("="^80)
println("Use case: Nodal assembly, need N_i for specific node")
topology = Tetrahedron()
basis = Lagrange{2}()
xi = Vec(0.25, 0.25, 0.2)
node_idx = 5 # Edge midpoint node
println("\nStrategy 1: Tuple + runtime index")
@btime get_basis_function_v1($topology, $basis, $xi, $node_idx)
println("\nStrategy 2: Val dispatch (compile-time index)")
@btime get_basis_function_v2($topology, $basis, $xi, Val($node_idx))
println("\nStrategy 3: @generated function (minimal computation)")
@btime get_basis_function_v3($topology, $basis, $xi, Val($node_idx))
# Verify all return same value
r1 = get_basis_function_v1(topology, basis, xi, node_idx)
r2 = get_basis_function_v2(topology, basis, xi, Val(node_idx))
r3 = get_basis_function_v3(topology, basis, xi, Val(node_idx))
@assert r1 r2 r3 "Strategies return different values!"
println("\n✓ All strategies return: N_$node_idx = $r1")
end
function benchmark_all_derivatives()
println("\n" * "="^80)
println("BENCHMARK 3: Get ALL 10 basis function derivatives")
println("="^80)
println("Use case: Stiffness matrix assembly (B matrix construction)")
println("Most important benchmark for 3D simulations!")
topology = Tetrahedron()
basis = Lagrange{2}()
xi = Vec(0.25, 0.25, 0.2)
println("\nAccess all 10 derivatives (each is Vec{3}):")
@btime get_basis_derivatives($topology, $basis, $xi)
result = get_basis_derivatives(topology, basis, xi)
println("\n✓ Result (10 Vec{3} gradients):")
for (i, dN) in enumerate(result)
println(" ∇N_$i = $dN")
end
end
function benchmark_single_derivative()
println("\n" * "="^80)
println("BENCHMARK 4: Get SINGLE basis function derivative")
println("="^80)
println("Use case: Nodal assembly for stiffness matrix")
topology = Tetrahedron()
basis = Lagrange{2}()
xi = Vec(0.25, 0.25, 0.2)
node_idx = 5
println("\nStrategy 1: Tuple + runtime index")
@btime get_basis_derivative_v1($topology, $basis, $xi, $node_idx)
println("\nStrategy 2: Val dispatch")
@btime get_basis_derivative_v2($topology, $basis, $xi, Val($node_idx))
println("\nStrategy 3: @generated function")
@btime get_basis_derivative_v3($topology, $basis, $xi, Val($node_idx))
# Verify
r1 = get_basis_derivative_v1(topology, basis, xi, node_idx)
r2 = get_basis_derivative_v2(topology, basis, xi, Val(node_idx))
r3 = get_basis_derivative_v3(topology, basis, xi, Val(node_idx))
@assert r1 r2 r3 "Strategies return different values!"
println("\n✓ All strategies return: ∇N_$node_idx = $r1")
end
function benchmark_stiffness_assembly_pattern()
println("\n" * "="^80)
println("BENCHMARK 5: Realistic stiffness matrix assembly loop")
println("="^80)
println("Use case: Compute element stiffness K_e (typical FEM inner loop)")
println("Pattern: B^T D B where B = strain-displacement matrix")
topology = Tetrahedron()
basis = Lagrange{2}()
xi = Vec(0.25, 0.25, 0.2) # Integration point
# Typical pattern: Get all derivatives, compute B matrix terms
println("\nPattern 1: Get all derivatives at once (traditional)")
stiffness_loop_v1 = let topology = topology, basis = basis, xi = xi
() -> begin
dN_all = get_basis_derivatives(topology, basis, xi)
s = 0.0
# Simplified: compute trace of K (just for benchmarking)
for i in 1:10
for j in 1:10
dNi = dN_all[i]
dNj = dN_all[j]
s += dot(dNi, dNj) # Simplified K_ij computation
end
end
return s
end
end
@btime $stiffness_loop_v1()
println("\nPattern 2: Get derivatives with @generated (manual unroll)")
stiffness_loop_v3 = let topology = topology, basis = basis, xi = xi
() -> begin
s = 0.0
# Unrolled loop (compiler would do this with Val)
dN1 = get_basis_derivative_v3(topology, basis, xi, Val(1))
dN2 = get_basis_derivative_v3(topology, basis, xi, Val(2))
dN3 = get_basis_derivative_v3(topology, basis, xi, Val(3))
dN4 = get_basis_derivative_v3(topology, basis, xi, Val(4))
dN5 = get_basis_derivative_v3(topology, basis, xi, Val(5))
dN6 = get_basis_derivative_v3(topology, basis, xi, Val(6))
dN7 = get_basis_derivative_v3(topology, basis, xi, Val(7))
dN8 = get_basis_derivative_v3(topology, basis, xi, Val(8))
dN9 = get_basis_derivative_v3(topology, basis, xi, Val(9))
dN10 = get_basis_derivative_v3(topology, basis, xi, Val(10))
# Compute all pairs (100 dot products)
for dNi in (dN1, dN2, dN3, dN4, dN5, dN6, dN7, dN8, dN9, dN10)
for dNj in (dN1, dN2, dN3, dN4, dN5, dN6, dN7, dN8, dN9, dN10)
s += dot(dNi, dNj)
end
end
return s
end
end
@btime $stiffness_loop_v3()
# Verify all compute same result
r1 = stiffness_loop_v1()
r3 = stiffness_loop_v3()
@assert abs(r1 - r3) < 1e-10 "Assembly patterns give different results!"
println("\n✓ Assembly result: $r1")
end
# ============================================================================
# Main execution
# ============================================================================
function main()
println("\n")
println("" * "="^78 * "")
println("" * " "^78 * "")
println("" * " "^15 * "TET10 BASIS FUNCTION ACCESS BENCHMARK" * " "^25 * "")
println("" * " "^20 * "(10-node Quadratic Tetrahedron)" * " "^26 * "")
println("" * " "^78 * "")
println("" * "="^78 * "")
println("\nElement: Tet10 (10-node quadratic tetrahedron)")
println("Nodes: 4 vertices + 6 edge midpoints")
println("Polynomial degree: P2 (quadratic)")
println("Dimension: 3D")
println("\nSeparation of concerns API:")
println(" Element(Tetrahedron, Lagrange{2}, connectivity)")
println(" get_basis_functions(Tetrahedron(), Lagrange{2}(), xi)")
println(" get_basis_derivatives(Tetrahedron(), Lagrange{2}(), xi)")
benchmark_all_basis_functions()
benchmark_single_basis_function()
benchmark_all_derivatives()
benchmark_single_derivative()
benchmark_stiffness_assembly_pattern()
println("\n" * "="^80)
println("SUMMARY & RECOMMENDATIONS FOR 3D SIMULATIONS")
println("="^80)
println("""
FOR STIFFNESS MATRIX ASSEMBLY (derivatives):
→ Use get_basis_derivatives() - returns all 10 gradients as tuple
→ Expected: ~10-30 ns, zero allocation
→ This is the HOT PATH for 3D FEM!
FOR MASS MATRIX ASSEMBLY (basis functions):
→ Use get_basis_functions() - returns all 10 values as tuple
→ Expected: ~5-15 ns, zero allocation
FOR NODAL ASSEMBLY (single node operations):
→ Use Val dispatch: get_basis_derivative(t, b, xi, Val(i))
→ @generated gives minimal computation (only compute requested function)
→ Expected: ~5-10 ns per node
API DESIGN DECISION:
```julia
# Separation of concerns (RECOMMENDED):
Element(Tetrahedron, Lagrange{2}, connectivity)
# Functions take topology explicitly:
dN_all = get_basis_derivatives(Tetrahedron(), Lagrange{2}(), xi)
dN_i = get_basis_derivative(Tetrahedron(), Lagrange{2}(), xi, Val(i))
```
WHY THIS API?
- Clear separation: topology is geometry, basis is interpolation
- Topology passed to basis evaluation (no redundancy in type parameters)
- Type-stable, zero-allocation, fully inlined
- Works with any basis type (Lagrange, Hierarchical, Nedelec, etc.)
PERFORMANCE TARGET:
- 10-node Tet10 derivatives: < 30 ns (achieved!)
- 100× faster than old Dict-based approach
- Ready for million-element meshes
""")
println("\n" * "="^80)
end
# Run benchmarks
if abspath(PROGRAM_FILE) == @__FILE__
main()
end
-240
View File
@@ -1,240 +0,0 @@
# Performance Analysis for Deformation Gradient Implementation
# This script analyzes the machine code generated and validates zero-allocation claims
using JuliaFEM
using Tensors
using BenchmarkTools
using InteractiveUtils
# Load the deformation gradient code
include("../src/physics/deformation_gradient.jl")
println("="^80)
println("DEFORMATION GRADIENT PERFORMANCE ANALYSIS")
println("="^80)
println()
# Setup test data
X_nodes = (
Vec(0.0, 0.0, 0.0),
Vec(1.0, 0.0, 0.0),
Vec(1.0, 1.0, 0.0),
Vec(0.0, 1.0, 0.0),
Vec(0.0, 0.0, 1.0),
Vec(1.0, 0.0, 1.0),
Vec(1.0, 1.0, 1.0),
Vec(0.0, 1.0, 1.0)
)
u_nodes = (
Vec(0.0, 0.0, 0.0),
Vec(0.1, 0.0, 0.0),
Vec(0.1, 0.0, 0.0),
Vec(0.0, 0.0, 0.0),
Vec(0.0, 0.0, 0.0),
Vec(0.1, 0.0, 0.0),
Vec(0.1, 0.0, 0.0),
Vec(0.0, 0.0, 0.0)
)
ξ = Vec(0.0, 0.0, 0.0)
dN_dξ = get_basis_derivatives(Hexahedron(), Lagrange{Hexahedron,1}(), ξ)
# Compute Jacobian
function compute_jacobian(X_nodes, dN_dξ)
J = zero(Tensor{2,3,Float64,9})
for i in 1:8
J += X_nodes[i] dN_dξ[i]
end
return J
end
J = compute_jacobian(X_nodes, dN_dξ)
# ============================================================================
# 1. ALLOCATION ANALYSIS
# ============================================================================
println("1. ALLOCATION ANALYSIS")
println("-"^80)
# Warm up (compile)
F = compute_deformation_gradient(X_nodes, u_nodes, dN_dξ, J, FiniteStrain())
# Measure allocations
allocs = @allocated compute_deformation_gradient(X_nodes, u_nodes, dN_dξ, J, FiniteStrain())
println("Allocations: $allocs bytes")
if allocs == 0
println("✅ ZERO ALLOCATIONS CONFIRMED!")
else
println("❌ WARNING: Found $allocs bytes allocated!")
end
println()
# ============================================================================
# 2. PERFORMANCE BENCHMARKING
# ============================================================================
println("2. PERFORMANCE BENCHMARKING")
println("-"^80)
println("Benchmarking compute_deformation_gradient...")
result = @benchmark compute_deformation_gradient($X_nodes, $u_nodes, $dN_dξ, $J, FiniteStrain())
println(result)
println()
median_time_ns = median(result.times)
println("Median time: $(median_time_ns) ns = $(median_time_ns/1000) μs")
println()
# ============================================================================
# 3. LLVM IR ANALYSIS
# ============================================================================
println("3. LLVM IR ANALYSIS")
println("-"^80)
println("Examining LLVM IR for signs of optimization...")
println()
io_llvm = IOBuffer()
code_llvm(io_llvm, compute_deformation_gradient,
typeof((X_nodes, u_nodes, dN_dξ, J, FiniteStrain())))
llvm_code = String(take!(io_llvm))
# Count key indicators
n_allocations = count(r"@julia.gc_alloc_obj", llvm_code)
n_stores = count(r"store", llvm_code)
n_loads = count(r"load", llvm_code)
n_vector_ops = count(r"<\d+ x ", llvm_code) # SIMD vector operations
println("LLVM IR Statistics:")
println(" - GC allocations: $n_allocations")
println(" - Store operations: $n_stores")
println(" - Load operations: $n_loads")
println(" - Vector operations (SIMD): $n_vector_ops")
println()
if n_allocations == 0
println("✅ No GC allocations in LLVM IR!")
else
println("❌ WARNING: Found $n_allocations GC allocation calls!")
end
if n_vector_ops > 0
println("✅ SIMD vectorization detected!")
end
println()
# Print full LLVM (first 100 lines)
println("Full LLVM IR (first 100 lines):")
println("-"^80)
llvm_lines = split(llvm_code, '\n')
for (i, line) in enumerate(llvm_lines[1:min(100, length(llvm_lines))])
println(line)
end
println()
# ============================================================================
# 4. NATIVE ASSEMBLY ANALYSIS
# ============================================================================
println("4. NATIVE ASSEMBLY ANALYSIS")
println("-"^80)
println("Examining native assembly...")
println()
io_native = IOBuffer()
code_native(io_native, compute_deformation_gradient,
typeof((X_nodes, u_nodes, dN_dξ, J, FiniteStrain())))
native_code = String(take!(io_native))
# Count key assembly features
n_movsd = count(r"movsd", native_code) # Scalar moves
n_movapd = count(r"movapd", native_code) # Aligned packed moves
n_movupd = count(r"movupd", native_code) # Unaligned packed moves
n_mulpd = count(r"mulpd", native_code) # Packed multiply
n_addpd = count(r"addpd", native_code) # Packed add
n_call = count(r"call", native_code) # Function calls
println("Native Assembly Statistics:")
println(" - Scalar moves (movsd): $n_movsd")
println(" - Aligned packed moves (movapd): $n_movapd")
println(" - Unaligned packed moves (movupd): $n_movupd")
println(" - Packed multiplies (mulpd): $n_mulpd")
println(" - Packed adds (addpd): $n_addpd")
println(" - Function calls: $n_call")
println()
if n_mulpd > 0 || n_addpd > 0
println("✅ SSE/AVX SIMD instructions detected!")
end
if n_call == 0
println("✅ Fully inlined - no function calls!")
else
println("⚠️ Note: $n_call function calls detected (may include math library)")
end
println()
# Print full assembly (first 100 lines)
println("Full Native Assembly (first 100 lines):")
println("-"^80)
native_lines = split(native_code, '\n')
for (i, line) in enumerate(native_lines[1:min(100, length(native_lines))])
println(line)
end
println()
# ============================================================================
# 5. TYPE STABILITY ANALYSIS
# ============================================================================
println("5. TYPE STABILITY ANALYSIS")
println("-"^80)
println("Checking type stability with @code_warntype...")
println()
io_warntype = IOBuffer()
code_warntype(io_warntype, compute_deformation_gradient,
typeof((X_nodes, u_nodes, dN_dξ, J, FiniteStrain())))
warntype_output = String(take!(io_native))
# Check for type instabilities
has_any = contains(warntype_output, "Any")
has_union = contains(warntype_output, "Union{")
if has_any
println("⚠️ WARNING: 'Any' types detected (type instability)")
else
println("✅ No 'Any' types detected!")
end
if has_union
println("⚠️ Note: Union types detected (may be intentional)")
else
println("✅ No Union types detected!")
end
println()
# Print warntype output (first 50 lines)
println("@code_warntype output (first 50 lines):")
println("-"^80)
warntype_lines = split(warntype_output, '\n')
for (i, line) in enumerate(warntype_lines[1:min(50, length(warntype_lines))])
println(line)
end
println()
# ============================================================================
# 6. SUMMARY
# ============================================================================
println("="^80)
println("PERFORMANCE SUMMARY")
println("="^80)
println()
println("✅ Implementation validated as:")
println(" - Zero allocation (confirmed)")
println(" - Type stable")
println(" - SIMD optimized ($(n_vector_ops) vector ops in LLVM)")
println(" - Median execution time: $(round(median_time_ns, digits=2)) ns")
println()
println("This implementation achieves the best possible performance for")
println("deformation gradient computation in Julia.")
println()
println("="^80)
@@ -1,396 +0,0 @@
# ==============================================================================
# ELEMENT IMMUTABILITY BENCHMARK
# ==============================================================================
#
# Purpose: Demonstrate why immutable elements with type-stable fields are faster
# than mutable elements with Dict-based fields, despite seeming
# counterintuitive.
#
# Hypothesis: Immutable + type-stable >> Mutable + Dict
#
# What we measure:
# 1. Field access time (reading)
# 2. Field update time (writing)
# 3. Memory allocations
# 4. Assembly loop performance (realistic FEM workload)
#
# Expected results:
# - Dict lookup: O(1) amortized, but ~100ns overhead per access
# - Type-stable access: O(1), but ~1ns (inlined, no overhead)
# - Immutable update: Allocates new struct, but compiler optimizes away
# - Dict update: Mutates in-place, but loses type stability
#
# Conclusion: For FEM assembly (tight loops, millions of field accesses),
# type stability dominates. Immutability enables GPU/HPC.
#
# ==============================================================================
using BenchmarkTools
using Statistics
println("="^80)
println("ELEMENT IMMUTABILITY BENCHMARK")
println("="^80)
println()
println("Comparing two implementations of P2 Lagrange Tetrahedron (Tet10):")
println(" 1. Mutable element with Dict-based fields (OLD API)")
println(" 2. Immutable element with NamedTuple fields (NEW API)")
println()
println("Measuring: field access, field update, assembly loop")
println("="^80)
println()
# ==============================================================================
# IMPLEMENTATION 1: Mutable Element with Dict-based Fields (OLD)
# ==============================================================================
"""
Mutable element: fields stored in Dict{Symbol,Any}
- Pro: Can add/remove fields dynamically
- Con: Type-unstable, Dict lookup overhead, no GPU support
"""
mutable struct MutableElement
id::UInt
connectivity::Vector{UInt}
fields::Dict{Symbol,Any} # Type-unstable!
end
function MutableElement(connectivity::Vector{UInt})
return MutableElement(UInt(0), connectivity, Dict{Symbol,Any}())
end
# Old-style update: mutate in-place
function update_field!(elem::MutableElement, field_name::Symbol, value)
elem.fields[field_name] = value
return nothing
end
# Old-style access: Dict lookup
function get_field(elem::MutableElement, field_name::Symbol)
return elem.fields[field_name]
end
# ==============================================================================
# IMPLEMENTATION 2: Immutable Element with NamedTuple Fields (NEW)
# ==============================================================================
"""
Immutable element: fields stored in NamedTuple
- Pro: Type-stable, zero overhead access, GPU-compatible
- Con: Cannot mutate, must create new element (but compiler optimizes!)
"""
struct ImmutableElement{F}
id::UInt
connectivity::NTuple{10,UInt} # Fixed size, stack-allocated
fields::F # Type-stable! (NamedTuple)
end
function ImmutableElement(connectivity::NTuple{10,UInt}, fields::NamedTuple)
return ImmutableElement{typeof(fields)}(UInt(0), connectivity, fields)
end
# New-style update: return new element (immutable)
function update_field(elem::ImmutableElement, updates::NamedTuple)
new_fields = merge(elem.fields, updates)
return ImmutableElement(elem.connectivity, new_fields)
end
# New-style access: direct field access (inlined!)
function get_field(elem::ImmutableElement, field_name::Symbol)
return getfield(elem.fields, field_name)
end
# ==============================================================================
# BENCHMARK 1: Field Access (Read Performance)
# ==============================================================================
println("BENCHMARK 1: Field Access (Reading E, ν, ρ in tight loop)")
println("-"^80)
# Setup test elements
connectivity_vec = UInt.(1:10)
connectivity_tuple = ntuple(i -> UInt(i), 10)
mutable_elem = MutableElement(connectivity_vec)
update_field!(mutable_elem, :E, 210e9)
update_field!(mutable_elem, :ν, 0.3)
update_field!(mutable_elem, :ρ, 7850.0)
immutable_elem = ImmutableElement(connectivity_tuple, (E=210e9, ν=0.3, ρ=7850.0))
# Benchmark: Read fields 1000 times (simulating assembly loop)
function read_fields_mutable(elem, n)
sum_val = 0.0
for _ in 1:n
E = get_field(elem, :E)
ν = get_field(elem, :ν)
ρ = get_field(elem, :ρ)
sum_val += E + ν + ρ
end
return sum_val
end
function read_fields_immutable(elem, n)
sum_val = 0.0
for _ in 1:n
E = get_field(elem, :E)
ν = get_field(elem, :ν)
ρ = get_field(elem, :ρ)
sum_val += E + ν + ρ
end
return sum_val
end
n_reads = 1000
println("Reading fields $n_reads times:")
println()
t_mutable = @benchmark read_fields_mutable($mutable_elem, $n_reads)
println("Mutable (Dict): ", minimum(t_mutable.times) / n_reads, " ns/read")
println(" Median: ", median(t_mutable.times) / n_reads, " ns/read")
println(" Allocs: ", t_mutable.allocs)
t_immutable = @benchmark read_fields_immutable($immutable_elem, $n_reads)
println("Immutable (Tuple): ", minimum(t_immutable.times) / n_reads, " ns/read")
println(" Median: ", median(t_immutable.times) / n_reads, " ns/read")
println(" Allocs: ", t_immutable.allocs)
speedup_read = minimum(t_mutable.times) / minimum(t_immutable.times)
println()
println("Speedup: ", round(speedup_read, digits=1), "x faster")
println()
# ==============================================================================
# BENCHMARK 2: Field Update (Write Performance)
# ==============================================================================
println("BENCHMARK 2: Field Update (Updating temperature field)")
println("-"^80)
# Benchmark: Update temperature field 100 times
function update_temperature_mutable(elem, n)
for i in 1:n
update_field!(elem, :temperature, Float64(i) * 293.15)
end
return get_field(elem, :temperature)
end
function update_temperature_immutable(elem, n)
current = elem
for i in 1:n
current = update_field(current, (temperature=Float64(i) * 293.15,))
end
return get_field(current, :temperature)
end
n_updates = 100
println("Updating temperature field $n_updates times:")
println()
# Reset elements
mutable_elem2 = MutableElement(connectivity_vec)
update_field!(mutable_elem2, :E, 210e9)
update_field!(mutable_elem2, :ν, 0.3)
immutable_elem2 = ImmutableElement(connectivity_tuple, (E=210e9, ν=0.3))
t_mutable_update = @benchmark update_temperature_mutable($mutable_elem2, $n_updates)
println("Mutable (mutate): ", minimum(t_mutable_update.times) / n_updates, " ns/update")
println(" Median: ", median(t_mutable_update.times) / n_updates, " ns/update")
println(" Allocs: ", t_mutable_update.allocs)
println(" Memory: ", t_mutable_update.memory, " bytes")
t_immutable_update = @benchmark update_temperature_immutable($immutable_elem2, $n_updates)
println("Immutable (copy): ", minimum(t_immutable_update.times) / n_updates, " ns/update")
println(" Median: ", median(t_immutable_update.times) / n_updates, " ns/update")
println(" Allocs: ", t_immutable_update.allocs)
println(" Memory: ", t_immutable_update.memory, " bytes")
println()
println("Note: Immutable creates new structs, but compiler optimizes stack allocation")
println()
# ==============================================================================
# BENCHMARK 3: Realistic Assembly Loop (FEM Workload)
# ==============================================================================
println("BENCHMARK 3: Realistic FEM Assembly Loop")
println("-"^80)
println("Simulating element stiffness matrix assembly:")
println(" - Read E, ν from element fields")
println(" - Compute 10 Gauss integration points")
println(" - Each point: read fields, compute B matrix, add to K")
println()
# Simplified assembly kernel
function assemble_stiffness_mutable(elem)
E = get_field(elem, :E)
ν = get_field(elem, :ν)
# Compute material matrix (simplified)
λ = E * ν / ((1 + ν) * (1 - 2ν))
μ = E / (2 * (1 + ν))
K = 0.0
# Simulate 10 integration points
for ip in 1:10
# Simulate field reads at integration point
E_ip = get_field(elem, :E)
ν_ip = get_field(elem, :ν)
# Simplified stiffness contribution
detJ = 1.0 + 0.1 * ip # Fake Jacobian
weight = 0.1
K += (λ + 2μ) * detJ * weight
end
return K
end
function assemble_stiffness_immutable(elem)
E = get_field(elem, :E)
ν = get_field(elem, :ν)
# Compute material matrix (simplified)
λ = E * ν / ((1 + ν) * (1 - 2ν))
μ = E / (2 * (1 + ν))
K = 0.0
# Simulate 10 integration points
for ip in 1:10
# Simulate field reads at integration point
E_ip = get_field(elem, :E)
ν_ip = get_field(elem, :ν)
# Simplified stiffness contribution
detJ = 1.0 + 0.1 * ip # Fake Jacobian
weight = 0.1
K += (λ + 2μ) * detJ * weight
end
return K
end
t_assembly_mutable = @benchmark assemble_stiffness_mutable($mutable_elem)
t_assembly_immutable = @benchmark assemble_stiffness_immutable($immutable_elem)
println("Assembly time per element:")
println()
println("Mutable (Dict): ", minimum(t_assembly_mutable.times), " ns")
println(" Median: ", median(t_assembly_mutable.times), " ns")
println(" Allocs: ", t_assembly_mutable.allocs)
println("Immutable (Tuple): ", minimum(t_assembly_immutable.times), " ns")
println(" Median: ", median(t_assembly_immutable.times), " ns")
println(" Allocs: ", t_assembly_immutable.allocs)
speedup_assembly = minimum(t_assembly_mutable.times) / minimum(t_assembly_immutable.times)
println()
println("Speedup: ", round(speedup_assembly, digits=1), "x faster")
println()
# ==============================================================================
# BENCHMARK 4: Large-Scale Mesh (1000 elements)
# ==============================================================================
println("BENCHMARK 4: Large-Scale Assembly (1000 elements)")
println("-"^80)
n_elements = 1000
# Create mesh
mutable_mesh = [
begin
elem = MutableElement(UInt.(1:10) .+ UInt(i * 10))
update_field!(elem, :E, 210e9)
update_field!(elem, :ν, 0.3)
elem
end for i in 1:n_elements
]
immutable_mesh = [
begin
conn = ntuple(j -> UInt(j + i * 10), 10)
ImmutableElement(conn, (E=210e9, ν=0.3))
end for i in 1:n_elements
]
function assemble_mesh_mutable(mesh)
K_total = 0.0
for elem in mesh
K_total += assemble_stiffness_mutable(elem)
end
return K_total
end
function assemble_mesh_immutable(mesh)
K_total = 0.0
for elem in mesh
K_total += assemble_stiffness_immutable(elem)
end
return K_total
end
println("Assembling $n_elements elements:")
println()
t_mesh_mutable = @benchmark assemble_mesh_mutable($mutable_mesh)
println("Mutable (Dict): ", minimum(t_mesh_mutable.times) / 1e6, " ms")
println(" Median: ", median(t_mesh_mutable.times) / 1e6, " ms")
println(" Allocs: ", t_mesh_mutable.allocs)
println(" Memory: ", t_mesh_mutable.memory / 1024, " KB")
t_mesh_immutable = @benchmark assemble_mesh_immutable($immutable_mesh)
println("Immutable (Tuple): ", minimum(t_mesh_immutable.times) / 1e6, " ms")
println(" Median: ", median(t_mesh_immutable.times) / 1e6, " ms")
println(" Allocs: ", t_mesh_immutable.allocs)
println(" Memory: ", t_mesh_immutable.memory / 1024, " KB")
speedup_mesh = minimum(t_mesh_mutable.times) / minimum(t_mesh_immutable.times)
println()
println("Speedup: ", round(speedup_mesh, digits=1), "x faster")
println()
# ==============================================================================
# SUMMARY
# ==============================================================================
println("="^80)
println("SUMMARY")
println("="^80)
println()
println("Key findings:")
println()
println("1. Field Access:")
println(" - Type-stable (immutable) is ", round(speedup_read, digits=1), "x faster")
println(" - Dict lookup: ~100-200ns overhead per access")
println(" - NamedTuple: ~1ns (inlined, zero overhead)")
println()
println("2. Assembly Performance:")
println(" - Single element: ", round(speedup_assembly, digits=1), "x faster")
println(" - Large mesh: ", round(speedup_mesh, digits=1), "x faster")
println()
println("3. Memory:")
println(" - Immutable elements: no allocations in hot path")
println(" - Mutable elements: Dict overhead + dynamic dispatch")
println()
println("4. GPU/HPC Compatibility:")
println(" - Immutable: ✓ All bits types, can transfer to GPU")
println(" - Mutable: ✗ Pointers, heap allocations, no GPU support")
println()
println("CONCLUSION:")
println("-"^80)
println("Despite appearing counterintuitive, IMMUTABLE elements with type-stable")
println("fields are SIGNIFICANTLY FASTER for FEM assembly. The key insight:")
println()
println(" • Dict lookup cost dominates in tight loops (millions of accesses)")
println(" • Type stability enables compiler optimizations (inlining, SIMD)")
println(" • Immutability enables GPU/HPC parallelization (no race conditions)")
println(" • Modern compilers optimize away struct copies on stack")
println()
println("For FEM with millions of field accesses per assembly, type stability")
println("is the critical factor. Immutability is a small price for 10-100x speedup.")
println()
println("="^80)
-334
View File
@@ -1,334 +0,0 @@
#!/usr/bin/env julia
#
# Benchmark: Dict{String,Any} vs Type-Stable Field Storage
#
# This benchmark validates the performance claims in:
# docs/book/zero_allocation_fields.md
#
# Expected results:
# - Constant field access: 50× faster, 0 allocations
# - Nodal field access: 50× faster, 0 allocations
# - Interpolation: 16× faster with zero allocations (cached)
#
using BenchmarkTools
using LinearAlgebra
# ============================================================================
# Field Type Definitions (from proposal)
# ============================================================================
abstract type AbstractField{T} end
struct ConstantField{T} <: AbstractField{T}
value::T
end
struct NodalField{T} <: AbstractField{T}
values::Matrix{T} # N_components × N_nodes
end
# Accessors
@inline value(f::ConstantField) = f.value
@inline function value(f::NodalField, node_ids::AbstractVector{Int})
return @view f.values[:, node_ids]
end
# ============================================================================
# Mock Element (simplified for benchmarking)
# ============================================================================
struct MockElement
connectivity::Vector{Int}
end
function mock_eval_basis(x::Vector{Float64})
# Mock basis function values for 8-node element
return [0.1, 0.15, 0.05, 0.1, 0.2, 0.15, 0.15, 0.1]
end
# ============================================================================
# Setup: OLD (Dict-based) vs NEW (Typed)
# ============================================================================
const OLD_FIELDS = Dict{String,Any}(
"youngs_modulus" => 210e3,
"poissons_ratio" => 0.3,
"displacement" => zeros(3, 8),
)
const NEW_FIELDS = (
youngs_modulus=ConstantField(210e3),
poissons_ratio=ConstantField(0.3),
displacement=NodalField(zeros(3, 8)),
)
# ============================================================================
# Benchmark 1: Constant Field Access
# ============================================================================
println("="^70)
println("Benchmark 1: Constant Field Access")
println("="^70)
println("\nOLD (Dict{String,Any}):")
old_constant = @benchmark $OLD_FIELDS["youngs_modulus"]
display(old_constant)
println("\nNEW (ConstantField):")
new_constant = @benchmark value($NEW_FIELDS.youngs_modulus)
display(new_constant)
old_time_1 = median(old_constant).time
new_time_1 = median(new_constant).time
speedup_1 = old_time_1 / new_time_1
allocs_old_1 = median(old_constant).allocs
allocs_new_1 = median(new_constant).allocs
println("\n📊 Results:")
println(" OLD: $(round(old_time_1, digits=1)) ns, $(allocs_old_1) allocations")
println(" NEW: $(round(new_time_1, digits=1)) ns, $(allocs_new_1) allocations")
println(" Speedup: $(round(speedup_1, digits=1))×")
println(" Allocation reduction: $(allocs_old_1 - allocs_new_1)")
# ============================================================================
# Benchmark 2: Nodal Field Access
# ============================================================================
println("\n" * "="^70)
println("Benchmark 2: Nodal Field Access (4 nodes)")
println("="^70)
node_ids = [1, 2, 3, 4]
println("\nOLD (Dict with Array slicing):")
old_nodal = @benchmark $OLD_FIELDS["displacement"][:, $node_ids]
display(old_nodal)
println("\nNEW (NodalField with @view):")
new_nodal = @benchmark value($NEW_FIELDS.displacement, $node_ids)
display(new_nodal)
old_time_2 = median(old_nodal).time
new_time_2 = median(new_nodal).time
speedup_2 = old_time_2 / new_time_2
allocs_old_2 = median(old_nodal).allocs
allocs_new_2 = median(new_nodal).allocs
println("\n📊 Results:")
println(" OLD: $(round(old_time_2, digits=1)) ns, $(allocs_old_2) allocations")
println(" NEW: $(round(new_time_2, digits=1)) ns, $(allocs_new_2) allocations")
println(" Speedup: $(round(speedup_2, digits=1))×")
println(" Allocation reduction: $(allocs_old_2 - allocs_new_2)")
# ============================================================================
# Benchmark 3: Interpolation (Without Cache)
# ============================================================================
println("\n" * "="^70)
println("Benchmark 3: Spatial Interpolation (No Cache)")
println("="^70)
element = MockElement([1, 2, 3, 4, 5, 6, 7, 8])
x = [0.1, 0.2, 0.3]
N = mock_eval_basis(x)
function interpolate_old(element, N, fields_dict)
u = fields_dict["displacement"] # Type: Any
result = zeros(3)
for i in 1:length(N)
result .+= N[i] .* u[:, element.connectivity[i]]
end
return result
end
function interpolate_new(element, N, fields)
u_nodal = value(fields.displacement, element.connectivity)
result = zeros(3)
for i in 1:length(N)
result .+= N[i] .* @view u_nodal[:, i]
end
return result
end
println("\nOLD (Dict-based):")
old_interp = @benchmark interpolate_old($element, $N, $OLD_FIELDS)
display(old_interp)
println("\nNEW (Typed fields):")
new_interp = @benchmark interpolate_new($element, $N, $NEW_FIELDS)
display(new_interp)
old_time_3 = median(old_interp).time
new_time_3 = median(new_interp).time
speedup_3 = old_time_3 / new_time_3
allocs_old_3 = median(old_interp).allocs
allocs_new_3 = median(new_interp).allocs
println("\n📊 Results:")
println(" OLD: $(round(old_time_3/1000, digits=1)) μs, $(allocs_old_3) allocations")
println(" NEW: $(round(new_time_3, digits=1)) ns, $(allocs_new_3) allocations")
println(" Speedup: $(round(speedup_3, digits=1))×")
println(" Allocation reduction: $(allocs_old_3 - allocs_new_3)")
# ============================================================================
# Benchmark 4: Interpolation (With Cache - Zero Allocation Target)
# ============================================================================
println("\n" * "="^70)
println("Benchmark 4: Spatial Interpolation (WITH Cache)")
println("="^70)
struct InterpolationCache
result::Vector{Float64}
end
function interpolate_cached!(cache, element, N, fields)
u_nodal = value(fields.displacement, element.connectivity)
fill!(cache.result, 0.0)
for i in eachindex(N)
cache.result .+= N[i] .* @view u_nodal[:, i]
end
return cache.result
end
cache = InterpolationCache(zeros(3))
println("\nNEW (Cached - Zero Allocation Target):")
cached_interp = @benchmark interpolate_cached!($cache, $element, $N, $NEW_FIELDS)
display(cached_interp)
cached_time = median(cached_interp).time
cached_allocs = median(cached_interp).allocs
speedup_4 = old_time_3 / cached_time
println("\n📊 Results:")
println(" Cached: $(round(cached_time, digits=1)) ns, $(cached_allocs) allocations")
println(" Speedup vs OLD: $(round(speedup_4, digits=1))×")
println(" Zero allocation target: $(cached_allocs == 0 ? "✅ MET" : "❌ FAILED")")
# ============================================================================
# Benchmark 5: Assembly Loop (1000 elements)
# ============================================================================
println("\n" * "="^70)
println("Benchmark 5: Assembly Loop (1000 elements)")
println("="^70)
n_elements = 1000
elements = [MockElement(collect(1:8)) for _ in 1:n_elements]
function assemble_old_style(elements, fields_dict)
total = 0.0
for element in elements
E = fields_dict["youngs_modulus"] # Type-unstable access
ν = fields_dict["poissons_ratio"]
# Mock stiffness computation
K_local = E * (1 - ν^2)
total += K_local
end
return total
end
function assemble_new_style(elements, fields)
E = value(fields.youngs_modulus) # Type-stable access (once)
ν = value(fields.poissons_ratio)
total = 0.0
for element in elements
# Mock stiffness computation
K_local = E * (1 - ν^2)
total += K_local
end
return total
end
println("\nOLD (Dict access in loop):")
old_assembly = @benchmark assemble_old_style($elements, $OLD_FIELDS)
display(old_assembly)
println("\nNEW (Typed fields, hoisted access):")
new_assembly = @benchmark assemble_new_style($elements, $NEW_FIELDS)
display(new_assembly)
old_time_5 = median(old_assembly).time
new_time_5 = median(new_assembly).time
speedup_5 = old_time_5 / new_time_5
allocs_old_5 = median(old_assembly).allocs
allocs_new_5 = median(new_assembly).allocs
println("\n📊 Results:")
println(" OLD: $(round(old_time_5/1000, digits=1)) μs, $(allocs_old_5) allocations")
println(" NEW: $(round(new_time_5/1000, digits=1)) μs, $(allocs_new_5) allocations")
println(" Speedup: $(round(speedup_5, digits=1))×")
println(" Allocation reduction: $(allocs_old_5 - allocs_new_5)")
# ============================================================================
# Summary and Validation
# ============================================================================
println("\n" * "="^70)
println("SUMMARY - Validation Against Claims")
println("="^70)
validation_passed = true
# Claim 1: Constant field access should be ~50× faster, 0 allocations
println("\n1. Constant Field Access:")
println(" Claimed: ~50× faster, 0 allocations")
println(" Actual: $(round(speedup_1, digits=1))× faster, $(allocs_new_1) allocations")
if speedup_1 >= 10 && allocs_new_1 == 0
println(" Status: ✅ VALIDATED ($(round(speedup_1, digits=1))× > 10× threshold)")
else
println(" Status: ⚠️ PARTIAL (speedup or allocation target not met)")
validation_passed = false
end
# Claim 2: Nodal field access should be ~50× faster, 0 allocations
println("\n2. Nodal Field Access:")
println(" Claimed: ~50× faster, 0 allocations")
println(" Actual: $(round(speedup_2, digits=1))× faster, $(allocs_new_2) allocations")
if speedup_2 >= 10 && allocs_new_2 == 0
println(" Status: ✅ VALIDATED ($(round(speedup_2, digits=1))× > 10× threshold)")
else
println(" Status: ⚠️ PARTIAL (speedup or allocation target not met)")
validation_passed = false
end
# Claim 3: Cached interpolation should be ~16× faster, 0 allocations
println("\n3. Interpolation (Cached):")
println(" Claimed: ~16× faster, 0 allocations")
println(" Actual: $(round(speedup_4, digits=1))× faster, $(cached_allocs) allocations")
if speedup_4 >= 10 && cached_allocs == 0
println(" Status: ✅ VALIDATED ($(round(speedup_4, digits=1))× > 10× threshold)")
else
println(" Status: ⚠️ PARTIAL (speedup or allocation target not met)")
validation_passed = false
end
# Claim 4: Assembly should be 10-100× faster
println("\n4. Assembly Loop:")
println(" Claimed: 10-100× faster")
println(" Actual: $(round(speedup_5, digits=1))× faster")
if speedup_5 >= 10
println(" Status: ✅ VALIDATED ($(round(speedup_5, digits=1))× > 10× threshold)")
else
println(" Status: ⚠️ PARTIAL (speedup target not met)")
validation_passed = false
end
println("\n" * "="^70)
if validation_passed
println("✅ ALL PERFORMANCE CLAIMS VALIDATED")
else
println("⚠️ SOME CLAIMS NOT FULLY VALIDATED (but likely still significant improvement)")
end
println("="^70)
println("\nKey Insights:")
println(" • Type stability (NamedTuple) eliminates runtime dispatch")
println(" • Zero allocations achieved with @view and pre-allocated caches")
println(" • Hoisting invariant access out of loops provides massive speedup")
println(" • The combination gives 10-100× speedup in realistic scenarios")
println("\n✅ This validates the NamedTuple + typed fields design for v1.0")
@@ -1,468 +0,0 @@
"""
GPU State Management Benchmark
Demonstrates Strategy 1 (immutable elements) vs Strategy 2 (separate mutable state)
and validates memory coalescing patterns on actual GPU hardware.
Run with:
julia --project=. benchmarks/gpu_state_management_benchmark.jl
"""
using CUDA
using Tensors
using BenchmarkTools
using Printf
# Check GPU availability
if !CUDA.functional()
error("CUDA not available! This benchmark requires a CUDA-capable GPU.")
end
println("GPU Device: $(CUDA.device())")
println("GPU Memory: $(CUDA.name(CUDA.device())) - $(round(CUDA.total_memory()/1e9, digits=1)) GB")
println()
# ============================================================================
# Strategy 1: Immutable Elements (Array of Structs - AoS)
# ============================================================================
"""
Strategy 1: Element contains its own state (immutable).
Update creates new element (allocation + copy).
"""
struct Element_Strategy1{T}
connectivity::NTuple{8,Int32}
material_id::Int32
# State (plastic strain, hardening)
ε_p::SymmetricTensor{2,3,T,6}
α::T
end
"""
Update state for Strategy 1 (returns new element - allocation!).
"""
function update_element_strategy1(elem::Element_Strategy1{T}, Δε_p, Δα) where T
return Element_Strategy1(
elem.connectivity,
elem.material_id,
elem.ε_p + Δε_p,
elem.α + Δα
)
end
"""
CPU kernel: Update all elements (Strategy 1).
"""
function update_elements_strategy1_cpu!(
elements::Vector{Element_Strategy1{T}},
strain_increments::Vector{SymmetricTensor{2,3,T,6}},
hardening_increments::Vector{T}
) where T
n = length(elements)
for i in 1:n
elements[i] = update_element_strategy1(
elements[i],
strain_increments[i],
hardening_increments[i]
)
end
end
"""
GPU kernel: Update all elements (Strategy 1).
Problem: Each thread accesses scattered memory (pointer chasing).
"""
function update_elements_strategy1_kernel!(
elements::CuDeviceVector{Element_Strategy1{T}},
strain_increments::CuDeviceVector{SymmetricTensor{2,3,T,6}},
hardening_increments::CuDeviceVector{T}
) where T
i = (blockIdx().x - 1) * blockDim().x + threadIdx().x
if i <= length(elements)
elem = elements[i] # Non-coalesced read!
Δε_p = strain_increments[i]
Δα = hardening_increments[i]
# Update (creates new element - allocation on GPU!)
new_elem = Element_Strategy1(
elem.connectivity,
elem.material_id,
elem.ε_p + Δε_p,
elem.α + Δα
)
elements[i] = new_elem # Non-coalesced write!
end
return nothing
end
function update_elements_strategy1_gpu!(
elements::CuVector{Element_Strategy1{T}},
strain_increments::CuVector{SymmetricTensor{2,3,T,6}},
hardening_increments::CuVector{T}
) where T
n = length(elements)
threads = 256
blocks = cld(n, threads)
@cuda threads = threads blocks = blocks update_elements_strategy1_kernel!(
elements, strain_increments, hardening_increments
)
CUDA.synchronize()
end
# ============================================================================
# Strategy 2: Separate Mutable State (Structure of Arrays - SoA)
# ============================================================================
"""
Strategy 2: Geometry is immutable, state is separate and mutable.
"""
struct ElementGeometry
connectivity::NTuple{8,Int32}
material_id::Int32
end
"""
Mutable state storage (flat arrays for GPU coalescing).
"""
mutable struct AssemblyState{T,VecT}
# Plastic strain (Voigt notation: 6 components per state)
ε_p_flat::VecT # [N_states × 6]
# Hardening variable (1 component per state)
α_flat::VecT # [N_states]
n_states::Int
end
function AssemblyState{T}(n_states::Int) where T
return AssemblyState{T,Vector{T}}(
zeros(T, n_states * 6),
zeros(T, n_states),
n_states
)
end
"""
CPU kernel: Update state (Strategy 2 - in-place!).
"""
function update_state_strategy2_cpu!(
state::AssemblyState{T,Vector{T}},
strain_increments::Vector{SymmetricTensor{2,3,T,6}},
hardening_increments::Vector{T}
) where T
n = state.n_states
for i in 1:n
# Flat indexing (cache-friendly!)
offset = (i - 1) * 6
Δε_p = strain_increments[i]
# Update in-place (no allocation!)
state.ε_p_flat[offset+1] += Δε_p[1, 1]
state.ε_p_flat[offset+2] += Δε_p[2, 2]
state.ε_p_flat[offset+3] += Δε_p[3, 3]
state.ε_p_flat[offset+4] += Δε_p[1, 2]
state.ε_p_flat[offset+5] += Δε_p[1, 3]
state.ε_p_flat[offset+6] += Δε_p[2, 3]
state.α_flat[i] += hardening_increments[i]
end
end
"""
GPU kernel: Update state (Strategy 2).
Advantage: Coalesced memory access!
- Thread 0 accesses state.ε_p_flat[0:5]
- Thread 1 accesses state.ε_p_flat[6:11]
- Thread 2 accesses state.ε_p_flat[12:17]
All consecutive in memory!
"""
function update_state_strategy2_kernel!(
ε_p_flat::CuDeviceVector{T},
α_flat::CuDeviceVector{T},
strain_increments_flat::CuDeviceVector{T},
hardening_increments::CuDeviceVector{T},
n_states::Int
) where T
i = (blockIdx().x - 1) * blockDim().x + threadIdx().x
if i <= n_states
# Flat indexing (coalesced access!)
offset = (i - 1) * 6
strain_offset = (i - 1) * 6
# Update plastic strain (6 consecutive reads/writes)
ε_p_flat[offset+1] += strain_increments_flat[strain_offset+1]
ε_p_flat[offset+2] += strain_increments_flat[strain_offset+2]
ε_p_flat[offset+3] += strain_increments_flat[strain_offset+3]
ε_p_flat[offset+4] += strain_increments_flat[strain_offset+4]
ε_p_flat[offset+5] += strain_increments_flat[strain_offset+5]
ε_p_flat[offset+6] += strain_increments_flat[strain_offset+6]
# Update hardening (1 read/write)
α_flat[i] += hardening_increments[i]
end
return nothing
end
function update_state_strategy2_gpu!(
state_gpu::AssemblyState{T,<:CuVector{T}},
strain_increments_flat::CuVector{T},
hardening_increments::CuVector{T}
) where T
n = state_gpu.n_states
threads = 256
blocks = cld(n, threads)
@cuda threads = threads blocks = blocks update_state_strategy2_kernel!(
state_gpu.ε_p_flat,
state_gpu.α_flat,
strain_increments_flat,
hardening_increments,
n
)
CUDA.synchronize()
end
# ============================================================================
# Benchmark Setup
# ============================================================================
function setup_benchmark(n_elements::Int)
T = Float64
# Create random strain increments
Δε_p_tensors = [SymmetricTensor{2,3}((
rand(T) * 1e-5,
rand(T) * 1e-5,
rand(T) * 1e-5,
rand(T) * 1e-6,
rand(T) * 1e-6,
rand(T) * 1e-6
)) for _ in 1:n_elements]
Δα = rand(T, n_elements) .* 1e-5
# Strategy 1: Array of immutable elements
elements_s1 = [Element_Strategy1(
ntuple(j -> Int32(j), 8),
Int32(1),
zero(SymmetricTensor{2,3,T}),
zero(T)
) for _ in 1:n_elements]
# Strategy 2: Separate geometry and state
geometry_s2 = [ElementGeometry(
ntuple(j -> Int32(j), 8),
Int32(1)
) for _ in 1:n_elements]
state_s2 = AssemblyState{T}(n_elements)
return Δε_p_tensors, Δα, elements_s1, geometry_s2, state_s2
end
# ============================================================================
# CPU Benchmarks
# ============================================================================
function benchmark_cpu(n_elements::Int)
println("="^70)
println("CPU Benchmark: $n_elements elements")
println("="^70)
Δε_p, Δα, elements_s1, geometry_s2, state_s2 = setup_benchmark(n_elements)
# Strategy 1: Update immutable elements
println("\n📊 Strategy 1 (Immutable Elements - AoS):")
elements_s1_copy = copy(elements_s1)
t1 = @belapsed update_elements_strategy1_cpu!(
$elements_s1_copy, $Δε_p, $Δα
) samples = 10
println(" Time: $(round(t1 * 1000, digits=3)) ms")
println(" Bandwidth: N/A (CPU cache)")
# Check allocations
allocs = @allocated update_elements_strategy1_cpu!(elements_s1_copy, Δε_p, Δα)
println(" Allocations: $(allocs) bytes ($(allocs ÷ n_elements) bytes/element)")
# Strategy 2: Update mutable state
println("\n📊 Strategy 2 (Separate State - SoA):")
state_s2_copy = deepcopy(state_s2)
t2 = @belapsed update_state_strategy2_cpu!(
$state_s2_copy, $Δε_p, $Δα
) samples = 10
println(" Time: $(round(t2 * 1000, digits=3)) ms")
println(" Bandwidth: N/A (CPU cache)")
# Check allocations
allocs2 = @allocated update_state_strategy2_cpu!(state_s2_copy, Δε_p, Δα)
println(" Allocations: $(allocs2) bytes")
# Speedup
speedup = t1 / t2
println("\n✅ CPU Speedup (Strategy 2 / Strategy 1): $(round(speedup, digits=2))×")
println()
end
# ============================================================================
# GPU Benchmarks
# ============================================================================
function benchmark_gpu(n_elements::Int)
println("="^70)
println("GPU Benchmark: $n_elements elements")
println("="^70)
T = Float64
Δε_p, Δα, elements_s1, geometry_s2, state_s2 = setup_benchmark(n_elements)
# ========================================================================
# Strategy 1: GPU
# ========================================================================
println("\n📊 Strategy 1 (Immutable Elements - AoS on GPU):")
# Transfer to GPU
elements_s1_gpu = CuArray(elements_s1)
Δε_p_gpu = CuArray(Δε_p)
Δα_gpu = CuArray(Δα)
# Warmup
update_elements_strategy1_gpu!(elements_s1_gpu, Δε_p_gpu, Δα_gpu)
# Benchmark
t1_gpu = CUDA.@elapsed begin
update_elements_strategy1_gpu!(elements_s1_gpu, Δε_p_gpu, Δα_gpu)
end
println(" Time: $(round(t1_gpu * 1000, digits=3)) ms")
# Estimate bandwidth (reading + writing entire element)
bytes_per_elem = sizeof(Element_Strategy1{T})
total_bytes = bytes_per_elem * n_elements * 2 # Read + write
bandwidth_s1 = total_bytes / t1_gpu / 1e9
println(" Bandwidth: $(round(bandwidth_s1, digits=1)) GB/s")
# ========================================================================
# Strategy 2: GPU
# ========================================================================
println("\n📊 Strategy 2 (Separate State - SoA on GPU):")
# Transfer to GPU (flat arrays!)
state_s2_gpu = AssemblyState{T,CuVector{T}}(
CuArray(state_s2.ε_p_flat),
CuArray(state_s2.α_flat),
state_s2.n_states
)
# Flatten strain increments for GPU
Δε_p_flat = zeros(T, n_elements * 6)
for i in 1:n_elements
offset = (i - 1) * 6
ε = Δε_p[i]
Δε_p_flat[offset+1] = ε[1, 1]
Δε_p_flat[offset+2] = ε[2, 2]
Δε_p_flat[offset+3] = ε[3, 3]
Δε_p_flat[offset+4] = ε[1, 2]
Δε_p_flat[offset+5] = ε[1, 3]
Δε_p_flat[offset+6] = ε[2, 3]
end
Δε_p_flat_gpu = CuArray(Δε_p_flat)
Δα_flat_gpu = CuArray(Δα)
# Warmup
update_state_strategy2_gpu!(state_s2_gpu, Δε_p_flat_gpu, Δα_flat_gpu)
# Benchmark
t2_gpu = CUDA.@elapsed begin
update_state_strategy2_gpu!(state_s2_gpu, Δε_p_flat_gpu, Δα_flat_gpu)
end
println(" Time: $(round(t2_gpu * 1000, digits=3)) ms")
# Estimate bandwidth (only state data, not geometry!)
bytes_per_state = 6 * sizeof(T) + sizeof(T) # 6 strain + 1 hardening
total_bytes_s2 = bytes_per_state * n_elements * 2 # Read + write
bandwidth_s2 = total_bytes_s2 / t2_gpu / 1e9
println(" Bandwidth: $(round(bandwidth_s2, digits=1)) GB/s")
# ========================================================================
# Comparison
# ========================================================================
speedup = t1_gpu / t2_gpu
bandwidth_ratio = bandwidth_s2 / bandwidth_s1
println("\n✅ GPU Speedup (Strategy 2 / Strategy 1): $(round(speedup, digits=2))×")
println("✅ Bandwidth Improvement: $(round(bandwidth_ratio, digits=2))×")
println(" Strategy 1: $(round(bandwidth_s1, digits=1)) GB/s (non-coalesced)")
println(" Strategy 2: $(round(bandwidth_s2, digits=1)) GB/s (coalesced)")
# Theoretical peak (example: RTX 4090 = ~1000 GB/s)
gpu_name = CUDA.name(CUDA.device())
println("\n💡 GPU Memory Bandwidth:")
println(" Achieved: $(round(bandwidth_s2, digits=1)) GB/s")
println(" Device: $gpu_name")
println()
# Cleanup
CUDA.unsafe_free!(elements_s1_gpu)
CUDA.unsafe_free!(Δε_p_gpu)
CUDA.unsafe_free!(Δα_gpu)
CUDA.unsafe_free!(state_s2_gpu.ε_p_flat)
CUDA.unsafe_free!(state_s2_gpu.α_flat)
CUDA.unsafe_free!(Δε_p_flat_gpu)
CUDA.unsafe_free!(Δα_flat_gpu)
end
# ============================================================================
# Main Benchmark
# ============================================================================
function main()
println("\n" * "=" * 70)
println("GPU State Management Strategy Benchmark")
println("=" * 70)
println()
# Test sizes
sizes = [10_000, 100_000, 1_000_000]
for n in sizes
# CPU benchmark
benchmark_cpu(n)
# GPU benchmark
benchmark_gpu(n)
println()
end
println("="^70)
println("Benchmark Complete!")
println("="^70)
println()
println("Key Findings:")
println(" - Strategy 1 (AoS): Non-coalesced memory access on GPU")
println(" - Strategy 2 (SoA): Coalesced memory access on GPU")
println(" - Strategy 2 achieves 5-10× higher memory bandwidth")
println(" - Strategy 2 has zero allocations (in-place update)")
println()
end
# Run benchmark
if abspath(PROGRAM_FILE) == @__FILE__
main()
end
-301
View File
@@ -1,301 +0,0 @@
# Benchmark: Integration Point Access Patterns
# ============================================
#
# This benchmark compares different approaches to storing and accessing
# integration points in finite element assembly loops.
#
# Key Question: What's the fastest way to get integration points?
#
# Approaches tested:
# 1. OLD: Runtime dispatch + mutable struct with Dict (type-unstable)
# 2. NEW Option A: Compile-time function (like eval_basis!)
# 3. NEW Option B: Store in element as NTuple
# 4. NEW Option C: Hybrid (compile-time generation + caching)
using BenchmarkTools
using Tensors
using StaticArrays
# ============================================================================
# OLD APPROACH: Runtime dispatch with mutable struct
# ============================================================================
struct OldIP
id::UInt
weight::Float64
coords::Tuple{Vararg{Float64}}
fields::Dict{String,Any} # Type-unstable!
end
function get_integration_points_old(::Type{Val{:Tri3}})
# Simulate runtime dispatch to get IPs
return [
OldIP(UInt(1), 0.5, (1 / 3, 1 / 3), Dict{String,Any}()),
]
end
function assembly_loop_old()
sum_val = 0.0
for _ in 1:1000 # Simulate 1000 elements
ips = get_integration_points_old(Val{:Tri3})
for ip in ips
w = ip.weight
xi, eta = ip.coords
# Simulate some computation
sum_val += w * (xi + eta)
end
end
return sum_val
end
# ============================================================================
# NEW OPTION A: Compile-time function (zero allocation)
# ============================================================================
"""
Return integration points as tuple at compile time.
Similar to eval_basis! - zero allocation, fully inlined.
"""
@inline function get_integration_points!(::Type{Val{:Tri3_Gauss1}})
# Return as tuple of (weight, coords) pairs
return (
(0.5, (1 / 3, 1 / 3)),
)
end
@inline function get_integration_points!(::Type{Val{:Tet4_Gauss1}})
return (
(1 / 24, (0.25, 0.25, 0.25)),
)
end
function assembly_loop_option_a()
sum_val = 0.0
for _ in 1:1000
ips = get_integration_points!(Val{:Tri3_Gauss1})
for (w, (xi, eta)) in ips
sum_val += w * (xi + eta)
end
end
return sum_val
end
# ============================================================================
# NEW OPTION B: Store in element as NTuple (what we have now)
# ============================================================================
struct IntegrationPoint{D}
ξ::NTuple{D,Float64}
weight::Float64
end
struct MockElement{NIP}
ips::NTuple{NIP,IntegrationPoint{2}}
end
function create_element_b()
ips = (
IntegrationPoint((1 / 3, 1 / 3), 0.5),
)
return MockElement(ips)
end
function assembly_loop_option_b()
elements = [create_element_b() for _ in 1:1000]
sum_val = 0.0
for element in elements
for ip in element.ips
w = ip.weight
xi, eta = ip.ξ
sum_val += w * (xi + eta)
end
end
return sum_val
end
# ============================================================================
# NEW OPTION C: Compile-time with Tensors.jl Vec (recommended for FEM)
# ============================================================================
"""
Return integration points with Vec{D} coordinates (Tensors.jl).
This matches the golden standard from nodal assembly demos.
"""
@inline function get_integration_points_vec!(::Type{Val{:Tri3_Gauss1}})
return (
(0.5, Vec{2}((1 / 3, 1 / 3))),
)
end
@inline function get_integration_points_vec!(::Type{Val{:Tet4_Gauss1}})
return (
(1 / 24, Vec{3}((0.25, 0.25, 0.25))),
)
end
function assembly_loop_option_c()
sum_val = 0.0
for _ in 1:1000
ips = get_integration_points_vec!(Val{:Tri3_Gauss1})
for (w, xi) in ips
# Vec arithmetic is optimized by Tensors.jl
sum_val += w * sum(xi)
end
end
return sum_val
end
# ============================================================================
# NEW OPTION D: Pre-generated global constants (ultimate zero-cost)
# ============================================================================
const TRI3_GAUSS1_IPS = (
(0.5, Vec{2}((1 / 3, 1 / 3))),
)
const TET4_GAUSS1_IPS = (
(1 / 24, Vec{3}((0.25, 0.25, 0.25))),
)
function assembly_loop_option_d()
sum_val = 0.0
for _ in 1:1000
for (w, xi) in TRI3_GAUSS1_IPS
sum_val += w * sum(xi)
end
end
return sum_val
end
# ============================================================================
# OPTION E: Hybrid - Function returns pre-computed constant
# ============================================================================
@inline get_ips_tri3_gauss1() = TRI3_GAUSS1_IPS
@inline get_ips_tet4_gauss1() = TET4_GAUSS1_IPS
function assembly_loop_option_e()
sum_val = 0.0
for _ in 1:1000
for (w, xi) in get_ips_tri3_gauss1()
sum_val += w * sum(xi)
end
end
return sum_val
end
# ============================================================================
# Run Benchmarks
# ============================================================================
println("="^80)
println("Integration Point Access Pattern Benchmark")
println("="^80)
println()
println("OLD APPROACH: Runtime dispatch + mutable struct with Dict")
println("-"^80)
@btime assembly_loop_old()
println()
println("OPTION A: Compile-time function returning tuples")
println("-"^80)
@btime assembly_loop_option_a()
println()
println("OPTION B: Store in element as NTuple (current approach)")
println("-"^80)
@btime assembly_loop_option_b()
println()
println("OPTION C: Compile-time function with Vec{D} (Tensors.jl)")
println("-"^80)
@btime assembly_loop_option_c()
println()
println("OPTION D: Pre-generated global constants")
println("-"^80)
@btime assembly_loop_option_d()
println()
println("OPTION E: Function returning pre-computed constant")
println("-"^80)
@btime assembly_loop_option_e()
println()
# ============================================================================
# Realistic FEM Assembly Benchmark
# ============================================================================
println()
println("="^80)
println("REALISTIC FEM ASSEMBLY COMPARISON")
println("="^80)
println()
# Simulate realistic element stiffness computation
function compute_element_stiffness_old(element_type::Type{Val{:Tri3}})
K_elem = zeros(6, 6)
ips = get_integration_points_old(element_type)
for ip in ips
w = ip.weight
xi, eta = ip.coords
# Simulate shape function evaluation and stiffness computation
N1 = 1 - xi - eta
N2 = xi
N3 = eta
# Accumulate (simplified)
K_elem[1, 1] += w * (N1^2)
end
return K_elem[1, 1]
end
function compute_element_stiffness_new()
K_elem = 0.0
for (w, xi) in get_integration_points_vec!(Val{:Tri3_Gauss1})
# Vec arithmetic
xi_val = xi[1]
eta_val = xi[2]
N1 = 1 - xi_val - eta_val
K_elem += w * (N1^2)
end
return K_elem
end
println("OLD: Realistic element stiffness assembly")
@btime begin
sum_val = 0.0
for _ in 1:1000
sum_val += compute_element_stiffness_old(Val{:Tri3})
end
sum_val
end
println()
println("NEW: Realistic element stiffness assembly")
@btime begin
sum_val = 0.0
for _ in 1:1000
sum_val += compute_element_stiffness_new()
end
sum_val
end
println()
println("="^80)
println("SUMMARY")
println("="^80)
println()
println("Expected ranking (fastest to slowest):")
println("1. Option D/E: Pre-computed constants (ultimate zero-cost)")
println("2. Option C: Compile-time with Vec{D} (recommended for FEM)")
println("3. Option A: Compile-time with plain tuples")
println("4. Option B: Stored in element NTuple (slight overhead)")
println("5. OLD: Runtime dispatch + Dict (type-unstable)")
println()
println("RECOMMENDATION:")
println(" Use Option C or D/E for integration points:")
println(" - Compile-time generation like eval_basis!")
println(" - Return as Tuple of (weight, Vec{D}) pairs")
println(" - Zero allocation, fully inlined")
println(" - Matches golden standard architecture")
-272
View File
@@ -1,272 +0,0 @@
"""
Performance analysis for LinearElastic material model.
Analyzes:
1. Execution time (@btime)
2. Memory allocations (@allocated)
3. Type stability (@code_warntype)
4. LLVM IR optimization (code_llvm)
5. Native assembly (code_native)
"""
using BenchmarkTools
using Tensors
using InteractiveUtils
# Load implementation
include("../src/materials/linear_elastic.jl")
println("="^80)
println("LINEAR ELASTIC MATERIAL - PERFORMANCE ANALYSIS")
println("="^80)
println()
# Test material (steel)
steel = LinearElastic(E=200e9, ν=0.3)
# Test strain (uniaxial extension)
ε = SymmetricTensor{2,3}((0.001, 0.0, 0.0, 0.0, 0.0, 0.0))
println("Material: Steel (E = 200 GPa, ν = 0.3)")
println("Strain: Uniaxial extension (ε₁₁ = 0.001)")
println()
# ============================================================================
# BENCHMARK 1: Execution Time
# ============================================================================
println("BENCHMARK 1: Execution Time")
println("-"^80)
# Warmup
compute_stress(steel, ε, nothing, 0.0)
# Benchmark
println("Running @btime compute_stress(steel, ε, nothing, 0.0)...")
t = @benchmark compute_stress($steel, $ε, nothing, 0.0)
println()
display(t)
println()
println()
# ============================================================================
# BENCHMARK 2: Memory Allocations
# ============================================================================
println("BENCHMARK 2: Memory Allocations")
println("-"^80)
# First call to compile
compute_stress(steel, ε, nothing, 0.0)
# Check allocations
allocs = @allocated compute_stress(steel, ε, nothing, 0.0)
println("Allocations: $allocs bytes")
if allocs == 0
println("✅ ZERO ALLOCATIONS (stack-only computation)")
else
println("⚠️ WARNING: Non-zero allocations detected!")
end
println()
println()
# ============================================================================
# BENCHMARK 3: Type Stability
# ============================================================================
println("BENCHMARK 3: Type Stability")
println("-"^80)
println("Running @code_warntype compute_stress(steel, ε, nothing, 0.0)...")
println()
@code_warntype compute_stress(steel, ε, nothing, 0.0)
println()
println()
# ============================================================================
# BENCHMARK 4: LLVM IR Analysis
# ============================================================================
println("BENCHMARK 4: LLVM IR Analysis")
println("-"^80)
println("Running @code_llvm compute_stress(steel, ε, nothing, 0.0)...")
println()
@code_llvm compute_stress(steel, ε, nothing, 0.0)
println()
println()
# ============================================================================
# BENCHMARK 5: Native Assembly
# ============================================================================
println("BENCHMARK 5: Native Assembly")
println("-"^80)
println("Running @code_native compute_stress(steel, ε, nothing, 0.0)...")
println()
@code_native compute_stress(steel, ε, nothing, 0.0)
println()
println()
# ============================================================================
# LLVM IR INSPECTION (Detailed Analysis)
# ============================================================================
println("LLVM IR INSPECTION")
println("-"^80)
# Get LLVM IR as string
llvm_ir = sprint(io -> code_llvm(io, compute_stress, typeof.((steel, ε, nothing, 0.0))))
# Count key operations
n_fadd = count(r"fadd", llvm_ir)
n_fmul = count(r"fmul", llvm_ir)
n_load = count(r"load", llvm_ir)
n_store = count(r"store", llvm_ir)
n_call = count(r"call", llvm_ir)
n_alloca = count(r"alloca", llvm_ir)
# Count vector operations (SIMD)
n_vector_ops = count(r"<\d+ x ", llvm_ir)
n_shufflevector = count(r"shufflevector", llvm_ir)
n_insertelement = count(r"insertelement", llvm_ir)
n_extractelement = count(r"extractelement", llvm_ir)
println("LLVM Operations Count:")
println(" Floating-point additions: $n_fadd")
println(" Floating-point multiplications: $n_fmul")
println(" Memory loads: $n_load")
println(" Memory stores: $n_store")
println(" Function calls: $n_call")
println(" Stack allocations (alloca): $n_alloca")
println()
println("SIMD Vectorization:")
println(" Vector operations: $n_vector_ops")
println(" Shuffle operations: $n_shufflevector")
println(" Insert element operations: $n_insertelement")
println(" Extract element operations: $n_extractelement")
println()
if n_call == 0
println("✅ No function calls (fully inlined)")
else
println("⚠️ Contains $n_call function calls (may not be fully inlined)")
end
if n_alloca == 0
println("✅ No stack allocations (register-only computation)")
else
println("️ Contains $n_alloca stack allocations")
end
println()
println()
# ============================================================================
# NATIVE ASSEMBLY INSPECTION
# ============================================================================
println("NATIVE ASSEMBLY INSPECTION")
println("-"^80)
# Get native assembly as string
native_asm = sprint(io -> code_native(io, compute_stress, typeof.((steel, ε, nothing, 0.0))))
# Count SIMD instructions (AVX/SSE)
n_vmul = count(r"vmul", native_asm)
n_vadd = count(r"vadd", native_asm)
n_vsub = count(r"vsub", native_asm)
n_vfma = count(r"vfma", native_asm)
n_vmov = count(r"vmov", native_asm)
n_vbroadcast = count(r"vbroadcast", native_asm)
# Count total vector instructions
n_total_simd = n_vmul + n_vadd + n_vsub + n_vfma + n_vmov + n_vbroadcast
println("x86-64 Assembly SIMD Instructions:")
println(" vmulpd/vmulsd: $n_vmul")
println(" vaddpd/vaddsd: $n_vadd")
println(" vsubpd/vsubsd: $n_vsub")
println(" vfmadd/vfmsub: $n_vfma (fused multiply-add)")
println(" vmovapd/vmovsd: $n_vmov")
println(" vbroadcast: $n_vbroadcast")
println(" Total SIMD ops: $n_total_simd")
println()
if n_vfma > 0
println("✅ FMA (Fused Multiply-Add) instructions detected (optimal)")
end
if n_total_simd > 0
println("✅ SIMD vectorization active (AVX/AVX2)")
else
println("⚠️ No SIMD instructions detected")
end
println()
println()
# ============================================================================
# PERFORMANCE SUMMARY
# ============================================================================
println("="^80)
println("PERFORMANCE SUMMARY")
println("="^80)
println()
# Extract median time from benchmark
median_time = median(t.times)
median_ns = median_time # Already in nanoseconds
println("Execution Time:")
println(" Median: $(round(median_ns, digits=2)) ns")
println(" Mean: $(round(mean(t.times), digits=2)) ns")
println(" Minimum: $(round(minimum(t.times), digits=2)) ns")
println()
println("Memory:")
println(" Allocations: $allocs bytes")
if allocs == 0
println(" ✅ Zero allocation (confirmed)")
end
println()
println("Code Quality:")
if n_call == 0
println(" ✅ Fully inlined (no function calls)")
end
if n_alloca == 0
println(" ✅ Register-only computation (no stack usage)")
end
if n_total_simd > 0
println(" ✅ SIMD optimized ($n_total_simd vector instructions)")
end
if n_vfma > 0
println(" ✅ FMA instructions ($n_vfma fused multiply-adds)")
end
println()
println("Expected Operations:")
println(" Hooke's law: σ = λ·tr(ε)·I + 2μ·ε")
println(" - 1 trace computation: 3 additions")
println(" - 1 scalar multiplication: 1 multiply")
println(" - 6 scalar multiplications for diagonal")
println(" - 6 additions for final stress")
println(" Tangent: 𝔻 = λ·I⊗I + 2μ·𝕀ˢʸᵐ")
println(" - Constant tensor construction (may be compile-time)")
println()
# Theoretical lower bound
theoretical_flops = 3 + 1 + 6 + 6 # From expected operations
println("Theoretical minimum FLOPs: ~$theoretical_flops")
println("LLVM FLOPs: $(n_fadd + n_fmul)")
println()
# Throughput calculation
elements_per_second = 1e9 / median_ns
println("Throughput:")
println(" ~$(round(elements_per_second / 1e6, digits=1)) million stress evaluations/second/core")
println()
println("✅ Implementation validated as:")
println(" - Zero allocation (confirmed)")
println(" - Type stable")
if n_total_simd > 0
println(" - SIMD optimized ($n_total_simd vector ops)")
end
println(" - Median execution time: $(round(median_ns, digits=2)) ns")
println()
println("="^80)
-865
View File
@@ -1,865 +0,0 @@
"""
Material Models Performance Benchmark (Extended Version)
Validates performance claims from docs/book/material_modeling.md:
- Zero allocation claims
- 5-50× speedup over Voigt/Dict approach
- Type stability analysis (especially 'nothing' return for stateless materials)
- Manual vs automatic differentiation for Neo-Hookean
- Material state handling for Newton iterations
Compares:
1. New approach: Tensors.jl with SymmetricTensor
2. Old approach: Voigt notation with arrays/Dict
3. Neo-Hookean: Manual derivatives vs automatic differentiation
Materials tested:
- Linear Elastic (Hookean) - Stateless
- Neo-Hookean Hyperelasticity - Stateless (AD and manual versions)
- Perfect Plasticity (von Mises) - Stateful
Type hierarchy:
- AbstractMaterial - Base type for all materials
- AbstractMaterialState - Base type for material internal state
- NoState - For stateless materials
- PlasticityState - For plasticity with history
"""
using Tensors
using BenchmarkTools
using LinearAlgebra
using InteractiveUtils # For @code_warntype
println("="^80)
println("Material Models Performance Benchmark (Extended)")
println("="^80)
println()
#=============================================================================
TYPE HIERARCHY
=============================================================================#
"""
Abstract base type for all materials.
All concrete materials must implement:
- `compute_stress(material, ε, state_old, Δt) -> (σ, 𝔻, state_new)`
- `initial_state(material) -> AbstractMaterialState`
"""
abstract type AbstractMaterial end
"""
Abstract base type for material internal state.
Used to track history-dependent variables during Newton iterations:
- Old state (beginning of time step)
- Trial state (current Newton iteration)
- New state (converged solution)
"""
abstract type AbstractMaterialState end
"""
State for stateless materials (no history dependence).
Using singleton type instead of `nothing` for type hierarchy consistency.
Performance identical to `nothing` (zero-sized type).
"""
struct NoState <: AbstractMaterialState end
"""
Initial state for stateless materials.
"""
initial_state(::AbstractMaterial) = NoState()
#=============================================================================
NEW APPROACH: Tensors.jl Implementation
=============================================================================#
# ---------------------------------------------------------------------------
# 1. Linear Elastic (Hookean)
# ---------------------------------------------------------------------------
"""Linear elastic material with Tensors.jl"""
struct LinearElastic <: AbstractMaterial
E::Float64 # Young's modulus [Pa]
ν::Float64 # Poisson's ratio [-]
end
LinearElastic(; E, ν) = LinearElastic(E, ν)
λ(mat::LinearElastic) = mat.E * mat.ν / ((1 + mat.ν) * (1 - 2mat.ν))
μ(mat::LinearElastic) = mat.E / (2(1 + mat.ν))
"""Compute stress for linear elastic material."""
function compute_stress(
material::LinearElastic,
ε::SymmetricTensor{2,3,T},
state_old::NoState,
Δt::Float64
) where T
# Lamé parameters
λ_val = λ(material)
μ_val = μ(material)
# Identity tensor
I = one(ε)
# Hooke's law: σ = λ·tr(ε)·I + 2μ·ε
σ = λ_val * tr(ε) * I + 2μ_val * ε
# Tangent modulus: 𝔻 = λ I⊗I + 2μ 𝕀ˢʸᵐ
𝕀ˢʸᵐ = one(SymmetricTensor{4,3,T}) # Symmetric 4th order identity
𝔻 = λ_val * I I + 2μ_val * 𝕀ˢʸᵐ
return σ, 𝔻, NoState() # No state change (stateless)
end# ---------------------------------------------------------------------------
# 2. Neo-Hookean Hyperelasticity (Automatic Differentiation)
# ---------------------------------------------------------------------------
"""Neo-Hookean hyperelastic material (using automatic differentiation)."""
struct NeoHookeanAD <: AbstractMaterial
μ::Float64 # Shear modulus [Pa]
λ::Float64 # Lamé parameter [Pa]
end
function NeoHookeanAD(; E, ν)
μ = E / (2(1 + ν))
λ = E * ν / ((1 + ν) * (1 - 2ν))
return NeoHookeanAD(μ, λ)
end
"""Strain energy density for Neo-Hookean model."""
function strain_energy(material::NeoHookeanAD, C::SymmetricTensor{2,3})
μ, λ = material.μ, material.λ
# Invariants
I₁ = tr(C)
J = (det(C))
# Strain energy: ψ = μ/2(I₁ - 3) - μln(J) + λ/2·ln²(J)
ψ = μ / 2 * (I₁ - 3) - μ * log(J) + λ / 2 * log(J)^2
return ψ
end
"""Compute stress for Neo-Hookean material using automatic differentiation."""
function compute_stress(
material::NeoHookeanAD,
E::SymmetricTensor{2,3,T}, # Green-Lagrange strain
state_old::NoState,
Δt::Float64
) where T
# Right Cauchy-Green tensor: C = 2E + I
I = one(E)
C = 2E + I
# Strain energy function (closure capturing material)
ψ(C_) = strain_energy(material, C_)
# Automatic differentiation!
𝔻, S = hessian(ψ, C, :all) # Returns both hessian and gradient!
# Note: We want S = 2·∂ψ/∂C, 𝔻 = 4·∂²ψ/∂C²
S = 2 * S
𝔻 = 4 * 𝔻
return S, 𝔻, NoState() # No state change (stateless)
end
# ---------------------------------------------------------------------------
# 3. Neo-Hookean Hyperelasticity (Manual Derivatives)
# ---------------------------------------------------------------------------
"""
Neo-Hookean hyperelastic material (hand-coded derivatives).
Strain energy: ψ(C) = μ/2(I₁ - 3) - μln(J) + λ/2·ln²(J)
Where:
- I₁ = tr(C) - First invariant
- J = √det(C) - Jacobian determinant
Derivatives (computed by hand):
- S = 2∂ψ/∂C = μ(I - C⁻¹) + λln(J)C⁻¹
- 𝔻 = 4∂²ψ/∂C² = λ(C⁻¹⊗C⁻¹) + 2(μ - λln(J))∂C⁻¹/∂C
The second derivative uses the identity:
∂C⁻¹/∂C : X = -C⁻¹:(X:C⁻¹) for any symmetric X
"""
struct NeoHookeanManual <: AbstractMaterial
μ::Float64 # Shear modulus [Pa]
λ::Float64 # Lamé parameter [Pa]
end
function NeoHookeanManual(; E, ν)
μ = E / (2(1 + ν))
λ = E * ν / ((1 + ν) * (1 - 2ν))
return NeoHookeanManual(μ, λ)
end
"""Compute stress for Neo-Hookean material with manual derivatives."""
function compute_stress(
material::NeoHookeanManual,
E::SymmetricTensor{2,3,T}, # Green-Lagrange strain
state_old::NoState,
Δt::Float64
) where T
μ, λ = material.μ, material.λ
# Right Cauchy-Green tensor: C = 2E + I
I = one(E)
C = 2E + I
# Invariants
J = (det(C))
C_inv = inv(C)
# Second Piola-Kirchhoff stress: S = μ(I - C⁻¹) + λln(J)C⁻¹
S = μ * (I - C_inv) + λ * log(J) * C_inv
# Material tangent: 𝔻 = 4∂²ψ/∂C²
# Term 1: λ(C⁻¹⊗C⁻¹)
𝔻₁ = λ * (C_inv C_inv)
# Term 2: 2(μ - λln(J))∂C⁻¹/∂C
# The derivative ∂C⁻¹/∂C can be computed as:
# (∂C⁻¹/∂C)ᵢⱼₖₗ = -1/2(C⁻¹ᵢₖC⁻¹ⱼₗ + C⁻¹ᵢₗC⁻¹ⱼₖ)
#
# For SymmetricTensor, we build this fourth-order tensor
# by exploiting the symmetry structure
# Build the symmetric fourth-order tensor manually
# This is the most expensive part of the computation
𝕀ˢʸᵐ = one(SymmetricTensor{4,3,T})
# For compressible Neo-Hookean, the full tangent is:
# 𝔻 = λ(C⁻¹⊗C⁻¹) - 2(μ - λln(J))(C⁻¹⊙C⁻¹)
# where ⊙ is the symmetric dyadic product for fourth-order tensors
# Construct C⁻¹⊗C⁻¹ part (already have 𝔻₁)
# Construct symmetric part: use Tensors.jl identity operations
# The fourth-order identity for symmetric tensors handles this
coeff = 2(μ - λ * log(J))
# For the symmetric outer product of C⁻¹ with itself,
# we can use the following approach:
# Build component-wise using Voigt ordering
# Simplified: Use the property that for small strains,
# this reduces to a simpler form. For full nonlinear case:
𝔻₂ = -coeff * inv_symmetric_outer(C_inv)
𝔻 = 𝔻₁ + 𝔻₂
return S, 𝔻, NoState()
end
"""
Compute symmetric fourth-order tensor from inverse: ∂C⁻¹/∂C
For symmetric second-order tensor C⁻¹, compute the fourth-order tensor:
(∂C⁻¹/∂C)ᵢⱼₖₗ = -1/2(C⁻¹ᵢₖC⁻¹ⱼₗ + C⁻¹ᵢₗC⁻¹ⱼₖ)
This appears in the material tangent of hyperelastic materials.
"""
function inv_symmetric_outer(C_inv::SymmetricTensor{2,3,T}) where T
# Extract components (Voigt notation: 11, 22, 33, 12, 23, 13)
c = [C_inv[1, 1], C_inv[2, 2], C_inv[3, 3],
C_inv[1, 2], C_inv[2, 3], C_inv[1, 3]]
# Build fourth-order tensor in Voigt notation (6x6 matrix representation)
# Then convert to SymmetricTensor{4,3}
#
# This is the -1/2(CᵢₖCⱼₗ + CᵢₗCⱼₖ) tensor
# For now, use a simpler approximation that works for Neo-Hookean
# Full implementation would build all 36 components
# Use outer product and symmetrize
result = C_inv C_inv
# Add symmetric component
# (This is a simplified version - full implementation needs more care)
return result
end
# ---------------------------------------------------------------------------
# 4. Perfect Plasticity (von Mises)
# ---------------------------------------------------------------------------
"""Perfect plasticity with von Mises yield criterion."""
struct PerfectPlasticity <: AbstractMaterial
E::Float64 # Young's modulus [Pa]
ν::Float64 # Poisson's ratio [-]
σ_y::Float64 # Yield stress [Pa]
end
PerfectPlasticity(; E, ν, σ_y) = PerfectPlasticity(E, ν, σ_y)
λ(mat::PerfectPlasticity) = mat.E * mat.ν / ((1 + mat.ν) * (1 - 2mat.ν))
μ(mat::PerfectPlasticity) = mat.E / (2(1 + mat.ν))
"""
Internal state for plasticity (history-dependent variables).
This struct is passed through Newton iterations:
- state_old: State at beginning of time step (t_n)
- state_trial: Trial state during iteration (may not converge)
- state_new: Updated state for next iteration (t_n+1)
"""
struct PlasticityState{T} <: AbstractMaterialState
ε_p::SymmetricTensor{2,3,T} # Plastic strain
α::T # Equivalent plastic strain
end
"""Initial state for plasticity (zero plastic strain)."""
initial_state(::PerfectPlasticity) = PlasticityState(zero(SymmetricTensor{2,3}), 0.0)
"""Von Mises equivalent stress."""
function von_mises_stress(σ::SymmetricTensor{2,3})
s = dev(σ) # Deviatoric stress
return (3 / 2 * s s)
end
"""Compute stress for perfectly plastic material with radial return."""
function compute_stress(
material::PerfectPlasticity,
ε::SymmetricTensor{2,3,T},
state_old::PlasticityState{T},
Δt::Float64
) where T
# Material parameters
λ_val = λ(material)
μ_val = μ(material)
σ_y = material.σ_y
# Elastic constitutive tensor
I = one(ε)
𝕀ˢʸᵐ = one(SymmetricTensor{4,3,T})
𝔻ᵉ = λ_val * I I + 2μ_val * 𝕀ˢʸᵐ
# Elastic predictor
ε_e = ε - state_old.ε_p
σ_trial = λ_val * tr(ε_e) * I + 2μ_val * ε_e
σ_eq_trial = von_mises_stress(σ_trial)
# Yield function
f = σ_eq_trial - σ_y
if f 0.0
# Elastic step
σ = σ_trial
𝔻 = 𝔻ᵉ
state_new = state_old
else
# Plastic step: Radial return
s_trial = dev(σ_trial)
p = tr(σ_trial) / 3
# Return to yield surface
σ = p * I + (σ_y / σ_eq_trial) * s_trial
# Plastic multiplier
Δγ = f / (3μ_val)
# Flow direction
n = (3 / 2) * s_trial / σ_eq_trial
# Update plastic strain
ε_p_new = state_old.ε_p + Δγ * n
α_new = state_old.α + Δγ
state_new = PlasticityState(ε_p_new, α_new)
# Algorithmic tangent (simplified)
θ = 1 - σ_y / σ_eq_trial
β = 6μ_val^2 / (3μ_val + θ * 3μ_val)
𝔻 = 𝔻ᵉ - β * (n n)
end
return σ, 𝔻, state_new
end
#=============================================================================
OLD APPROACH: Voigt Notation + Array Implementation
=============================================================================#
"""Old-style linear elastic with Voigt notation."""
struct LinearElasticOld
E::Float64
ν::Float64
end
"""Compute 6×6 constitutive matrix (Voigt notation)."""
function constitutive_matrix(mat::LinearElasticOld)
E, ν = mat.E, mat.ν
λ = E * ν / ((1 + ν) * (1 - 2ν))
μ = E / (2(1 + ν))
D = zeros(6, 6)
D[1:3, 1:3] .= λ
D[1, 1] = D[2, 2] = D[3, 3] = λ + 2μ
D[4, 4] = D[5, 5] = D[6, 6] = μ
return D
end
"""Compute stress (old approach with arrays)."""
function compute_stress_old(
material::LinearElasticOld,
ε_vec::Vector{Float64}, # [ε11, ε22, ε33, 2ε12, 2ε23, 2ε13]
state_old::Dict{String,Any},
Δt::Float64
)
D = constitutive_matrix(material)
σ_vec = D * ε_vec
return σ_vec, D, state_old
end
"""Old-style Neo-Hookean (manual derivatives)."""
struct NeoHookeanOld
μ::Float64
λ::Float64
end
"""Compute stress manually (simplified, no actual derivatives for brevity)."""
function compute_stress_old(
material::NeoHookeanOld,
E_vec::Vector{Float64},
state_old::Dict{String,Any},
Δt::Float64
)
# This would normally have 50+ lines of manual derivative calculations
# For benchmark purposes, just do some array operations
D = zeros(6, 6)
for i in 1:6
D[i, i] = material.μ + material.λ / 3
end
σ_vec = D * E_vec
return σ_vec, D, state_old
end
"""Old-style plasticity with Dict storage."""
struct PerfectPlasticityOld
E::Float64
ν::Float64
σ_y::Float64
end
"""Compute stress with Dict field storage."""
function compute_stress_old(
material::PerfectPlasticityOld,
ε_vec::Vector{Float64},
state_old::Dict{String,Any},
Δt::Float64
)
# Get plastic strain from Dict (type instability!)
if haskey(state_old, "epsilon_plastic")
ε_p_vec = state_old["epsilon_plastic"]
else
ε_p_vec = zeros(6)
end
# Elastic trial
D = constitutive_matrix(LinearElasticOld(material.E, material.ν))
ε_e_vec = ε_vec - ε_p_vec
σ_trial_vec = D * ε_e_vec
# Von Mises check (manual calculation with arrays)
s11, s22, s33 = σ_trial_vec[1:3]
s12, s23, s13 = σ_trial_vec[4:6]
p = (s11 + s22 + s33) / 3
dev_vec = [s11 - p, s22 - p, s33 - p, s12, s23, s13]
σ_eq = (3 / 2 * (dev_vec[1]^2 + dev_vec[2]^2 + dev_vec[3]^2 +
2 * (dev_vec[4]^2 + dev_vec[5]^2 + dev_vec[6]^2)))
f = σ_eq - material.σ_y
state_new = copy(state_old)
if f > 0.0
# Plastic correction
factor = material.σ_y / σ_eq
σ_vec = [p, p, p, 0.0, 0.0, 0.0] + factor * dev_vec
# Update state in Dict
Δγ = f / (3 * material.E / (2(1 + material.ν)))
n_vec = (3 / 2) * dev_vec / σ_eq
state_new["epsilon_plastic"] = ε_p_vec + Δγ * n_vec
else
σ_vec = σ_trial_vec
end
return σ_vec, D, state_new
end
#=============================================================================
MATERIAL STATE HANDLING FOR NEWTON ITERATIONS
=============================================================================#
"""
Example: How to handle material state during Newton-Raphson iterations.
In FEM nonlinear analysis, each time step requires iterative solution:
1. **Beginning of time step (t_n):**
- state_old = converged state from previous time step
2. **During Newton iterations (t_n → t_n+1):**
- For each iteration k = 1, 2, ...
- Compute: σ, 𝔻, state_trial = compute_stress(material, ε_k, state_old, Δt)
- state_trial is NOT committed yet (iteration may not converge)
3. **After convergence:**
- state_new = state_trial from final iteration
- Commit: state_old ← state_new for next time step
This ensures:
- Failed iterations don't corrupt material history
- Material state is consistent with converged solution
- Internal variables (plastic strain, damage, etc.) evolve correctly
"""
"""
Simulate Newton-Raphson iteration with material state handling.
Returns:
- converged: Whether iterations converged
- n_iter: Number of iterations
- state_converged: Final material state (only valid if converged)
"""
function newton_with_material_state(
material::AbstractMaterial,
ε_target::SymmetricTensor{2,3},
state_old::AbstractMaterialState,
Δt::Float64;
max_iter=10,
tol=1e-8
)
println(" Newton iteration with material state tracking:")
println(" " * "="^60)
# Initial guess
ε_k = zero(ε_target)
for k in 1:max_iter
# Compute stress and tangent (state_trial is NOT committed yet!)
σ_k, 𝔻_k, state_trial = compute_stress(material, ε_k, state_old, Δt)
println(" Iteration $k:")
println(" strain: $(norm(ε_k))")
println(" stress: $(norm(σ_k))")
println(" state: $(state_trial)")
# Residual (simplified: just strain error)
r = norm(ε_k - ε_target)
if r < tol
println(" → Converged!")
println(" Final state committed: $(state_trial)")
return true, k, state_trial
end
# Newton update (simplified)
ε_k = ε_k + 0.5 * (ε_target - ε_k)
end
println(" → Failed to converge!")
println(" State NOT committed (keeping state_old)")
return false, max_iter, state_old # Keep old state on failure!
end
println()
println("="^80)
println("NEWTON ITERATION STATE HANDLING EXAMPLE")
println("="^80)
println()
# Example 1: Stateless material (LinearElastic)
println("Example 1: Stateless Material (LinearElastic)")
println("-"^80)
steel_example = LinearElastic(E=200e9, ν=0.3)
state_stateless = initial_state(steel_example)
ε_test = SymmetricTensor{2,3}((0.001, 0.0, 0.0, 0.0, 0.0, 0.0))
converged, n_iter, state_final = newton_with_material_state(
steel_example, ε_test, state_stateless, 1.0, max_iter=3
)
println("Result: state_final = $state_final (NoState, always)")
println()
# Example 2: Stateful material (PerfectPlasticity)
println("Example 2: Stateful Material (PerfectPlasticity)")
println("-"^80)
plastic_example = PerfectPlasticity(E=200e9, ν=0.3, σ_y=250e6)
state_stateful = initial_state(plastic_example)
ε_test_plastic = SymmetricTensor{2,3}((0.002, 0.0, 0.0, 0.0, 0.0, 0.0)) # Large strain → plastic
converged, n_iter, state_final = newton_with_material_state(
plastic_example, ε_test_plastic, state_stateful, 1.0, max_iter=3
)
println("Result: state_final = $state_final (plastic strain accumulated)")
println()
println("Key insight: State handling is IDENTICAL for all materials due to")
println("AbstractMaterialState type hierarchy. Assembly code doesn't need")
println("to know whether material is stateless or stateful!")
println()
#=============================================================================
BENCHMARK SETUP
=============================================================================#
println("Setting up materials and test cases...")
println()
# Materials (realistic steel properties)
steel_new = LinearElastic(E=200e9, ν=0.3)
steel_old = LinearElasticOld(200e9, 0.3)
rubber_ad = NeoHookeanAD(E=10e6, ν=0.45)
rubber_manual = NeoHookeanManual(E=10e6, ν=0.45)
rubber_old = NeoHookeanOld(10e6 / (2 * 1.45), 10e6 * 0.45 / (1.45 * 0.1))
plastic_new = PerfectPlasticity(E=200e9, ν=0.3, σ_y=250e6)
plastic_old = PerfectPlasticityOld(200e9, 0.3, 250e6)
# Test strain (small elastic deformation)
ε11, ε22, ε33 = 0.001, -0.0003, -0.0003 # Uniaxial tension with Poisson effect
ε12, ε23, ε13 = 0.0, 0.0, 0.0
# New approach: SymmetricTensor
ε_tensor = SymmetricTensor{2,3}((ε11, ε12, ε13, ε22, ε23, ε33))
E_tensor = ε_tensor # For Neo-Hookean (Green-Lagrange ≈ small strain here)
# Old approach: Voigt vector (note factor of 2 for shear!)
ε_voigt = [ε11, ε22, ε33, 2 * ε12, 2 * ε23, 2 * ε13]
# States (using proper type hierarchy)
state_nostate = NoState()
state_dict_empty = Dict{String,Any}()
state_plastic_new = initial_state(plastic_new)
state_plastic_old = Dict{String,Any}("epsilon_plastic" => zeros(6))
println("Materials configured:")
println(" - Linear Elastic: E = 200 GPa, ν = 0.3")
println(" - Neo-Hookean (AD): μ ≈ 3.4 MPa, λ ≈ 45 MPa (automatic differentiation)")
println(" - Neo-Hookean (Manual): μ ≈ 3.4 MPa, λ ≈ 45 MPa (hand-coded derivatives)")
println(" - Perfect Plasticity: E = 200 GPa, σ_y = 250 MPa")
println()
println("Test strain: ε11 = 0.001 (uniaxial tension)")
println()
#=============================================================================
TYPE STABILITY CHECK
=============================================================================#
println("="^80)
println("TYPE STABILITY ANALYSIS")
println("="^80)
println()
println("Checking for type instabilities...")
println()
# Check LinearElastic
println("1. Linear Elastic (Tensors.jl):")
@code_warntype compute_stress(steel_new, ε_tensor, state_nostate, 0.0)
println()
println("2. Linear Elastic (Old Voigt/Dict):")
@code_warntype compute_stress_old(steel_old, ε_voigt, state_dict_empty, 0.0)
println()
println("3. Neo-Hookean AD (Tensors.jl with automatic differentiation):")
@code_warntype compute_stress(rubber_ad, E_tensor, state_nostate, 0.0)
println()
println("4. Neo-Hookean Manual (Tensors.jl with hand-coded derivatives):")
@code_warntype compute_stress(rubber_manual, E_tensor, state_nostate, 0.0)
println()
println("5. Perfect Plasticity (Tensors.jl):")
@code_warntype compute_stress(plastic_new, ε_tensor, state_plastic_new, 0.0)
println()
println("6. Perfect Plasticity (Old Dict):")
@code_warntype compute_stress_old(plastic_old, ε_voigt, state_plastic_old, 0.0)
println()
#=============================================================================
ALLOCATION TESTS
=============================================================================#
println("="^80)
println("ALLOCATION TESTS")
println("="^80)
println()
println("Testing for allocations (should be 0 for new approach)...")
println()
# Linear Elastic
println("1. Linear Elastic")
println(" NEW (Tensors.jl):")
allocs_le_new = @allocated compute_stress(steel_new, ε_tensor, state_nostate, 0.0)
println(" Allocations: $allocs_le_new bytes")
println(" OLD (Voigt/Dict):")
allocs_le_old = @allocated compute_stress_old(steel_old, ε_voigt, state_dict_empty, 0.0)
println(" Allocations: $allocs_le_old bytes")
println()
# Neo-Hookean
println("2. Neo-Hookean")
println(" NEW (Tensors.jl + AD):")
allocs_nh_ad = @allocated compute_stress(rubber_ad, E_tensor, state_nostate, 0.0)
println(" Allocations: $allocs_nh_ad bytes")
println(" NEW (Tensors.jl + Manual):")
allocs_nh_manual = @allocated compute_stress(rubber_manual, E_tensor, state_nostate, 0.0)
println(" Allocations: $allocs_nh_manual bytes")
println(" OLD (Array):")
allocs_nh_old = @allocated compute_stress_old(rubber_old, ε_voigt, state_dict_empty, 0.0)
println(" Allocations: $allocs_nh_old bytes")
println()
# Perfect Plasticity
println("3. Perfect Plasticity (elastic branch)")
println(" NEW (Tensors.jl):")
allocs_pp_new = @allocated compute_stress(plastic_new, ε_tensor, state_plastic_new, 0.0)
println(" Allocations: $allocs_pp_new bytes")
println(" OLD (Dict):")
allocs_pp_old = @allocated compute_stress_old(plastic_old, ε_voigt, state_plastic_old, 0.0)
println(" Allocations: $allocs_pp_old bytes")
println()
#=============================================================================
PERFORMANCE BENCHMARKS
=============================================================================#
println("="^80)
println("PERFORMANCE BENCHMARKS")
println("="^80)
println()
println("Running detailed benchmarks (this may take a minute)...")
println()
# Linear Elastic
println("1. LINEAR ELASTIC")
println("-"^40)
println("NEW (Tensors.jl):")
bench_le_new = @benchmark compute_stress($steel_new, $ε_tensor, $state_nostate, 0.0)
display(bench_le_new)
println()
println("OLD (Voigt/Dict):")
bench_le_old = @benchmark compute_stress_old($steel_old, $ε_voigt, $state_dict_empty, 0.0)
display(bench_le_old)
println()
speedup_le = median(bench_le_old.times) / median(bench_le_new.times)
println("SPEEDUP: $(round(speedup_le, digits=1))×")
println()
# Neo-Hookean
println("2. NEO-HOOKEAN")
println("-"^40)
println("NEW (Tensors.jl + Automatic Differentiation):")
bench_nh_ad = @benchmark compute_stress($rubber_ad, $E_tensor, $state_nostate, 0.0)
display(bench_nh_ad)
println()
println("NEW (Tensors.jl + Manual Derivatives):")
bench_nh_manual = @benchmark compute_stress($rubber_manual, $E_tensor, $state_nostate, 0.0)
display(bench_nh_manual)
println()
println("OLD (Array):")
bench_nh_old = @benchmark compute_stress_old($rubber_old, $ε_voigt, $state_dict_empty, 0.0)
display(bench_nh_old)
println()
speedup_nh_ad = median(bench_nh_old.times) / median(bench_nh_ad.times)
speedup_nh_manual = median(bench_nh_old.times) / median(bench_nh_manual.times)
ad_overhead = median(bench_nh_ad.times) / median(bench_nh_manual.times)
println("SPEEDUP (AD): $(round(speedup_nh_ad, digits=1))×")
println("SPEEDUP (Manual): $(round(speedup_nh_manual, digits=1))×")
println("AD OVERHEAD: $(round(ad_overhead, digits=1))× (AD / Manual)")
println()
# Perfect Plasticity
println("3. PERFECT PLASTICITY (elastic branch)")
println("-"^40)
println("NEW (Tensors.jl):")
bench_pp_new = @benchmark compute_stress($plastic_new, $ε_tensor, $state_plastic_new, 0.0)
display(bench_pp_new)
println()
println("OLD (Dict):")
bench_pp_old = @benchmark compute_stress_old($plastic_old, $ε_voigt, $state_plastic_old, 0.0)
display(bench_pp_old)
println()
speedup_pp = median(bench_pp_old.times) / median(bench_pp_new.times)
println("SPEEDUP: $(round(speedup_pp, digits=1))×")
println()
#=============================================================================
SUMMARY
=============================================================================#
println("="^80)
println("SUMMARY")
println("="^80)
println()
println("ALLOCATIONS:")
println(" LinearElastic: NEW = $allocs_le_new bytes, OLD = $allocs_le_old bytes")
println(" NeoHookean (AD): NEW = $allocs_nh_ad bytes, OLD = $allocs_nh_old bytes")
println(" NeoHookean (Manual): NEW = $allocs_nh_manual bytes")
println(" PerfectPlasticity: NEW = $allocs_pp_new bytes, OLD = $allocs_pp_old bytes")
println()
println("MEDIAN TIMING:")
println(" LinearElastic: NEW = $(median(bench_le_new.times)) ns, OLD = $(median(bench_le_old.times)) ns")
println(" NeoHookean (AD): NEW = $(median(bench_nh_ad.times)) ns, OLD = $(median(bench_nh_old.times)) ns")
println(" NeoHookean (Manual): NEW = $(median(bench_nh_manual.times)) ns")
println(" PerfectPlasticity: NEW = $(median(bench_pp_new.times)) ns, OLD = $(median(bench_pp_old.times)) ns")
println()
println("SPEEDUP (OLD / NEW):")
println(" LinearElastic: $(round(speedup_le, digits=1))×")
println(" NeoHookean (AD): $(round(speedup_nh_ad, digits=1))×")
println(" NeoHookean (Manual): $(round(speedup_nh_manual, digits=1))×")
println(" PerfectPlasticity: $(round(speedup_pp, digits=1))×")
println()
println("AD OVERHEAD:")
println(" NeoHookean: AD is $(round(ad_overhead, digits=1))× slower than manual derivatives")
println()
avg_speedup = (speedup_le + speedup_nh_manual + speedup_pp) / 3
println("AVERAGE SPEEDUP: $(round(avg_speedup, digits=1))× (using manual Neo-Hookean)")
println()
# Validate claims
println("VALIDATION OF CLAIMS:")
println(" - Zero allocations for new approach: ",
allocs_le_new == 0 && allocs_nh_ad == 0 && allocs_nh_manual == 0 && allocs_pp_new == 0 ? "✓ PASS" : "✗ FAIL")
println(" - Manual derivatives outperform AD: ",
median(bench_nh_manual.times) < median(bench_nh_ad.times) ? "✓ PASS" : "✗ FAIL")
println(" - Type stability with NoState return: Check @code_warntype output above")
println()
println("="^80)
println("Benchmark complete! Results saved to: material_models_benchmark_results.txt")
println("="^80)
@@ -1,922 +0,0 @@
================================================================================
Material Models Performance Benchmark (Extended)
================================================================================
================================================================================
NEWTON ITERATION STATE HANDLING EXAMPLE
================================================================================
Example 1: Stateless Material (LinearElastic)
--------------------------------------------------------------------------------
Newton iteration with material state tracking:
============================================================
Iteration 1:
strain: 0.0
stress: 0.0
state: NoState()
Iteration 2:
strain: 0.0005
stress: 1.5741063022831637e8
state: NoState()
Iteration 3:
strain: 0.00075
stress: 2.3611594534247452e8
state: NoState()
→ Failed to converge!
State NOT committed (keeping state_old)
Result: state_final = NoState() (NoState, always)
Example 2: Stateful Material (PerfectPlasticity)
--------------------------------------------------------------------------------
Newton iteration with material state tracking:
============================================================
Iteration 1:
strain: 0.0
stress: 0.0
state: PlasticityState{Float64}([0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0], 0.0)
Iteration 2:
strain: 0.001
stress: 3.1482126045663273e8
state: PlasticityState{Float64}([0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0], 0.0)
Iteration 3:
strain: 0.0015
stress: 4.7223189068494904e8
state: PlasticityState{Float64}([0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0], 0.0)
→ Failed to converge!
State NOT committed (keeping state_old)
Result: state_final = PlasticityState{Float64}([0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0], 0.0) (plastic strain accumulated)
Key insight: State handling is IDENTICAL for all materials due to
AbstractMaterialState type hierarchy. Assembly code doesn't need
to know whether material is stateless or stateful!
Setting up materials and test cases...
Materials configured:
- Linear Elastic: E = 200 GPa, ν = 0.3
- Neo-Hookean (AD): μ ≈ 3.4 MPa, λ ≈ 45 MPa (automatic differentiation)
- Neo-Hookean (Manual): μ ≈ 3.4 MPa, λ ≈ 45 MPa (hand-coded derivatives)
- Perfect Plasticity: E = 200 GPa, σ_y = 250 MPa
Test strain: ε11 = 0.001 (uniaxial tension)
================================================================================
TYPE STABILITY ANALYSIS
================================================================================
Checking for type instabilities...
1. Linear Elastic (Tensors.jl):
MethodInstance for compute_stress(::LinearElastic, ::SymmetricTensor{2, 3, Float64, 6}, ::NoState, ::Float64)
from compute_stress(material::LinearElastic, ε::SymmetricTensor{2, 3, T}, state_old::NoState, Δt::Float64) where T @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:94
Static Parameters
T = Float64
Arguments
#self#::Core.Const(Main.compute_stress)
material::LinearElastic
ε::SymmetricTensor{2, 3, Float64, 6}
state_old::Core.Const(NoState())
Δt::Float64
Locals
𝔻::SymmetricTensor{4, 3, Float64, 36}
𝕀ˢʸᵐ::SymmetricTensor{4, 3, Float64, 36}
σ::SymmetricTensor{2, 3, Float64, 6}
I::SymmetricTensor{2, 3, Float64, 6}
μ_val::Float64
λ_val::Float64
Body::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
1 ─ %1 = Main.λ::Core.Const(Main.λ)
│ (λ_val = (%1)(material))
│ %3 = Main.μ::Core.Const(Main.μ)
│ (μ_val = (%3)(material))
│ %5 = Main.one::Core.Const(one)
│ (I = (%5)(ε))
│ %7 = Main.:+::Core.Const(+)
│ %8 = Main.:*::Core.Const(*)
│ %9 = λ_val::Float64
│ %10 = Main.tr::Core.Const(LinearAlgebra.tr)
│ %11 = (%10)(ε)::Float64
│ %12 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ %13 = (%8)(%9, %11, %12)::SymmetricTensor{2, 3, Float64, 6}
│ %14 = Main.:*::Core.Const(*)
│ %15 = Main.:*::Core.Const(*)
│ %16 = μ_val::Float64
│ %17 = (%15)(2, %16)::Float64
│ %18 = (%14)(%17, ε)::SymmetricTensor{2, 3, Float64, 6}
│ (σ = (%7)(%13, %18))
│ %20 = Main.one::Core.Const(one)
│ %21 = Main.SymmetricTensor::Core.Const(SymmetricTensor)
│ %22 = $(Expr(:static_parameter, 1))::Core.Const(Float64)
│ %23 = Core.apply_type(%21, 4, 3, %22)::Core.Const(SymmetricTensor{4, 3, Float64})
│ (𝕀ˢʸᵐ = (%20)(%23))
│ %25 = Main.:+::Core.Const(+)
│ %26 = Main.:⊗::Core.Const(Tensors.otimes)
│ %27 = Main.:*::Core.Const(*)
│ %28 = λ_val::Float64
│ %29 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ %30 = (%27)(%28, %29)::SymmetricTensor{2, 3, Float64, 6}
│ %31 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ %32 = (%26)(%30, %31)::SymmetricTensor{4, 3, Float64, 36}
│ %33 = Main.:*::Core.Const(*)
│ %34 = Main.:*::Core.Const(*)
│ %35 = μ_val::Float64
│ %36 = (%34)(2, %35)::Float64
│ %37 = 𝕀ˢʸᵐ::Core.Const([1.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.5 0.0; 0.5 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.5; 0.0 0.0 0.0; 0.5 0.0 0.0;;;; 0.0 0.5 0.0; 0.5 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.5; 0.0 0.5 0.0;;;; 0.0 0.0 0.5; 0.0 0.0 0.0; 0.5 0.0 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.5; 0.0 0.5 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 1.0])
│ %38 = (%33)(%36, %37)::SymmetricTensor{4, 3, Float64, 36}
│ (𝔻 = (%25)(%32, %38))
│ %40 = σ::SymmetricTensor{2, 3, Float64, 6}
│ %41 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
│ %42 = Main.NoState::Core.Const(NoState)
│ %43 = (%42)()::Core.Const(NoState())
│ %44 = Core.tuple(%40, %41, %43)::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
└── return %44
2. Linear Elastic (Old Voigt/Dict):
MethodInstance for compute_stress_old(::LinearElasticOld, ::Vector{Float64}, ::Dict{String, Any}, ::Float64)
from compute_stress_old(material::LinearElasticOld, ε_vec::Vector{Float64}, state_old::Dict{String, Any}, Δt::Float64) @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:413
Arguments
#self#::Core.Const(Main.compute_stress_old)
material::LinearElasticOld
ε_vec::Vector{Float64}
state_old::Dict{String, Any}
Δt::Float64
Locals
σ_vec::Vector{Float64}
D::Matrix{Float64}
Body::Tuple{Vector{Float64}, Matrix{Float64}, Dict{String, Any}}
1 ─ %1 = Main.constitutive_matrix::Core.Const(Main.constitutive_matrix)
│ (D = (%1)(material))
│ %3 = Main.:*::Core.Const(*)
│ %4 = D::Matrix{Float64}
│ (σ_vec = (%3)(%4, ε_vec))
│ %6 = σ_vec::Vector{Float64}
│ %7 = D::Matrix{Float64}
│ %8 = Core.tuple(%6, %7, state_old)::Tuple{Vector{Float64}, Matrix{Float64}, Dict{String, Any}}
└── return %8
3. Neo-Hookean AD (Tensors.jl with automatic differentiation):
MethodInstance for compute_stress(::NeoHookeanAD, ::SymmetricTensor{2, 3, Float64, 6}, ::NoState, ::Float64)
from compute_stress(material::NeoHookeanAD, E::SymmetricTensor{2, 3, T}, state_old::NoState, Δt::Float64) where T @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:147
Static Parameters
T = Float64
Arguments
#self#::Core.Const(Main.compute_stress)
material::NeoHookeanAD
E::SymmetricTensor{2, 3, Float64, 6}
state_old::Core.Const(NoState())
Δt::Float64
Locals
@_6::Int64
S::SymmetricTensor{2, 3, Float64, 6}
𝔻::SymmetricTensor{4, 3, Float64, 36}
ψ::var"#ψ#compute_stress##0"{NeoHookeanAD}
C::SymmetricTensor{2, 3, Float64, 6}
I::SymmetricTensor{2, 3, Float64, 6}
Body::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
1 ─ %1 = Main.one::Core.Const(one)
│ (I = (%1)(E))
│ %3 = Main.:+::Core.Const(+)
│ %4 = Main.:*::Core.Const(*)
│ %5 = (%4)(2, E)::SymmetricTensor{2, 3, Float64, 6}
│ %6 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ (C = (%3)(%5, %6))
│ %8 = Main.:(var"#ψ#compute_stress##0")::Core.Const(var"#ψ#compute_stress##0")
│ %9 = Core._typeof_captured_variable(material)::Core.Const(NeoHookeanAD)
│ %10 = Core.apply_type(%8, %9)::Core.Const(var"#ψ#compute_stress##0"{NeoHookeanAD})
│ (ψ = %new(%10, material))
│ %12 = Main.hessian::Core.Const(Tensors.hessian)
│ %13 = ψ::var"#ψ#compute_stress##0"{NeoHookeanAD}
│ %14 = C::SymmetricTensor{2, 3, Float64, 6}
│ %15 = (%12)(%13, %14, :all)::Tuple{SymmetricTensor{4, 3, Float64, 36}, SymmetricTensor{2, 3, Float64, 6}, Float64}
│ %16 = Base.indexed_iterate(%15, 1)::Core.PartialStruct(Tuple{SymmetricTensor{4, 3, Float64, 36}, Int64}, Any[SymmetricTensor{4, 3, Float64, 36}, Core.Const(2)])
│ (𝔻 = Core.getfield(%16, 1))
│ (@_6 = Core.getfield(%16, 2))
│ %19 = @_6::Core.Const(2)
│ %20 = Base.indexed_iterate(%15, 2, %19)::Core.PartialStruct(Tuple{SymmetricTensor{2, 3, Float64, 6}, Int64}, Any[SymmetricTensor{2, 3, Float64, 6}, Core.Const(3)])
│ (S = Core.getfield(%20, 1))
│ %22 = Main.:*::Core.Const(*)
│ %23 = S::SymmetricTensor{2, 3, Float64, 6}
│ (S = (%22)(2, %23))
│ %25 = Main.:*::Core.Const(*)
│ %26 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
│ (𝔻 = (%25)(4, %26))
│ %28 = S::SymmetricTensor{2, 3, Float64, 6}
│ %29 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
│ %30 = Main.NoState::Core.Const(NoState)
│ %31 = (%30)()::Core.Const(NoState())
│ %32 = Core.tuple(%28, %29, %31)::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
└── return %32
4. Neo-Hookean Manual (Tensors.jl with hand-coded derivatives):
MethodInstance for compute_stress(::NeoHookeanManual, ::SymmetricTensor{2, 3, Float64, 6}, ::NoState, ::Float64)
from compute_stress(material::NeoHookeanManual, E::SymmetricTensor{2, 3, T}, state_old::NoState, Δt::Float64) where T @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:203
Static Parameters
T = Float64
Arguments
#self#::Core.Const(Main.compute_stress)
material::NeoHookeanManual
E::SymmetricTensor{2, 3, Float64, 6}
state_old::Core.Const(NoState())
Δt::Float64
Locals
𝔻::SymmetricTensor{4, 3, Float64, 36}
𝔻₂::SymmetricTensor{4, 3, Float64, 36}
coeff::Float64
𝕀ˢʸᵐ::SymmetricTensor{4, 3, Float64, 36}
𝔻₁::SymmetricTensor{4, 3, Float64, 36}
S::SymmetricTensor{2, 3, Float64, 6}
C_inv::SymmetricTensor{2, 3, Float64, 6}
J::Float64
C::SymmetricTensor{2, 3, Float64, 6}
I::SymmetricTensor{2, 3, Float64, 6}
λ::Float64
μ::Float64
Body::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
1 ─ %1 = Base.getproperty(material, :μ)::Float64
│ %2 = Base.getproperty(material, :λ)::Float64
│ (μ = %1)
│ (λ = %2)
│ %5 = Main.one::Core.Const(one)
│ (I = (%5)(E))
│ %7 = Main.:+::Core.Const(+)
│ %8 = Main.:*::Core.Const(*)
│ %9 = (%8)(2, E)::SymmetricTensor{2, 3, Float64, 6}
│ %10 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ (C = (%7)(%9, %10))
│ %12 = Main.:√::Core.Const(sqrt)
│ %13 = Main.det::Core.Const(LinearAlgebra.det)
│ %14 = C::SymmetricTensor{2, 3, Float64, 6}
│ %15 = (%13)(%14)::Float64
│ (J = (%12)(%15))
│ %17 = Main.inv::Core.Const(inv)
│ %18 = C::SymmetricTensor{2, 3, Float64, 6}
│ (C_inv = (%17)(%18))
│ %20 = Main.:+::Core.Const(+)
│ %21 = Main.:*::Core.Const(*)
│ %22 = μ::Float64
│ %23 = Main.:-::Core.Const(-)
│ %24 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ %25 = C_inv::SymmetricTensor{2, 3, Float64, 6}
│ %26 = (%23)(%24, %25)::SymmetricTensor{2, 3, Float64, 6}
│ %27 = (%21)(%22, %26)::SymmetricTensor{2, 3, Float64, 6}
│ %28 = Main.:*::Core.Const(*)
│ %29 = λ::Float64
│ %30 = Main.log::Core.Const(log)
│ %31 = J::Float64
│ %32 = (%30)(%31)::Float64
│ %33 = C_inv::SymmetricTensor{2, 3, Float64, 6}
│ %34 = (%28)(%29, %32, %33)::SymmetricTensor{2, 3, Float64, 6}
│ (S = (%20)(%27, %34))
│ %36 = Main.:*::Core.Const(*)
│ %37 = λ::Float64
│ %38 = Main.:⊗::Core.Const(Tensors.otimes)
│ %39 = C_inv::SymmetricTensor{2, 3, Float64, 6}
│ %40 = C_inv::SymmetricTensor{2, 3, Float64, 6}
│ %41 = (%38)(%39, %40)::SymmetricTensor{4, 3, Float64, 36}
│ (𝔻₁ = (%36)(%37, %41))
│ %43 = Main.one::Core.Const(one)
│ %44 = Main.SymmetricTensor::Core.Const(SymmetricTensor)
│ %45 = $(Expr(:static_parameter, 1))::Core.Const(Float64)
│ %46 = Core.apply_type(%44, 4, 3, %45)::Core.Const(SymmetricTensor{4, 3, Float64})
│ (𝕀ˢʸᵐ = (%43)(%46))
│ %48 = Main.:*::Core.Const(*)
│ %49 = Main.:-::Core.Const(-)
│ %50 = μ::Float64
│ %51 = Main.:*::Core.Const(*)
│ %52 = λ::Float64
│ %53 = Main.log::Core.Const(log)
│ %54 = J::Float64
│ %55 = (%53)(%54)::Float64
│ %56 = (%51)(%52, %55)::Float64
│ %57 = (%49)(%50, %56)::Float64
│ (coeff = (%48)(2, %57))
│ %59 = Main.:*::Core.Const(*)
│ %60 = Main.:-::Core.Const(-)
│ %61 = coeff::Float64
│ %62 = (%60)(%61)::Float64
│ %63 = Main.inv_symmetric_outer::Core.Const(Main.inv_symmetric_outer)
│ %64 = C_inv::SymmetricTensor{2, 3, Float64, 6}
│ %65 = (%63)(%64)::SymmetricTensor{4, 3, Float64, 36}
│ (𝔻₂ = (%59)(%62, %65))
│ %67 = Main.:+::Core.Const(+)
│ %68 = 𝔻₁::SymmetricTensor{4, 3, Float64, 36}
│ %69 = 𝔻₂::SymmetricTensor{4, 3, Float64, 36}
│ (𝔻 = (%67)(%68, %69))
│ %71 = S::SymmetricTensor{2, 3, Float64, 6}
│ %72 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
│ %73 = Main.NoState::Core.Const(NoState)
│ %74 = (%73)()::Core.Const(NoState())
│ %75 = Core.tuple(%71, %72, %74)::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
└── return %75
5. Perfect Plasticity (Tensors.jl):
MethodInstance for compute_stress(::PerfectPlasticity, ::SymmetricTensor{2, 3, Float64, 6}, ::PlasticityState{Float64}, ::Float64)
from compute_stress(material::PerfectPlasticity, ε::SymmetricTensor{2, 3, T}, state_old::PlasticityState{T}, Δt::Float64) where T @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:328
Static Parameters
T = Float64
Arguments
#self#::Core.Const(Main.compute_stress)
material::PerfectPlasticity
ε::SymmetricTensor{2, 3, Float64, 6}
state_old::PlasticityState{Float64}
Δt::Float64
Locals
𝔻::SymmetricTensor{4, 3, Float64, 36}
β::Float64
θ::Float64
state_new::PlasticityState{Float64}
α_new::Float64
ε_p_new::SymmetricTensor{2, 3, Float64, 6}
n::SymmetricTensor{2, 3, Float64, 6}
Δγ::Float64
σ::SymmetricTensor{2, 3, Float64, 6}
p::Float64
s_trial::SymmetricTensor{2, 3, Float64, 6}
f::Float64
σ_eq_trial::Float64
σ_trial::SymmetricTensor{2, 3, Float64, 6}
ε_e::SymmetricTensor{2, 3, Float64, 6}
𝔻ᵉ::SymmetricTensor{4, 3, Float64, 36}
𝕀ˢʸᵐ::SymmetricTensor{4, 3, Float64, 36}
I::SymmetricTensor{2, 3, Float64, 6}
σ_y::Float64
μ_val::Float64
λ_val::Float64
Body::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, PlasticityState{Float64}}
1 ─ Core.NewvarNode(:(𝔻))
│ Core.NewvarNode(:(β))
│ Core.NewvarNode(:(θ))
│ Core.NewvarNode(:(state_new))
│ Core.NewvarNode(:(α_new))
│ Core.NewvarNode(:(ε_p_new))
│ Core.NewvarNode(:(n))
│ Core.NewvarNode(:(Δγ))
│ Core.NewvarNode(:(σ))
│ Core.NewvarNode(:(p))
│ Core.NewvarNode(:(s_trial))
│ %12 = Main.λ::Core.Const(Main.λ)
│ (λ_val = (%12)(material))
│ %14 = Main.μ::Core.Const(Main.μ)
│ (μ_val = (%14)(material))
│ (σ_y = Base.getproperty(material, :σ_y))
│ %17 = Main.one::Core.Const(one)
│ (I = (%17)(ε))
│ %19 = Main.one::Core.Const(one)
│ %20 = Main.SymmetricTensor::Core.Const(SymmetricTensor)
│ %21 = $(Expr(:static_parameter, 1))::Core.Const(Float64)
│ %22 = Core.apply_type(%20, 4, 3, %21)::Core.Const(SymmetricTensor{4, 3, Float64})
│ (𝕀ˢʸᵐ = (%19)(%22))
│ %24 = Main.:+::Core.Const(+)
│ %25 = Main.:⊗::Core.Const(Tensors.otimes)
│ %26 = Main.:*::Core.Const(*)
│ %27 = λ_val::Float64
│ %28 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ %29 = (%26)(%27, %28)::SymmetricTensor{2, 3, Float64, 6}
│ %30 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ %31 = (%25)(%29, %30)::SymmetricTensor{4, 3, Float64, 36}
│ %32 = Main.:*::Core.Const(*)
│ %33 = Main.:*::Core.Const(*)
│ %34 = μ_val::Float64
│ %35 = (%33)(2, %34)::Float64
│ %36 = 𝕀ˢʸᵐ::Core.Const([1.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.5 0.0; 0.5 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.5; 0.0 0.0 0.0; 0.5 0.0 0.0;;;; 0.0 0.5 0.0; 0.5 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.5; 0.0 0.5 0.0;;;; 0.0 0.0 0.5; 0.0 0.0 0.0; 0.5 0.0 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.5; 0.0 0.5 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 1.0])
│ %37 = (%32)(%35, %36)::SymmetricTensor{4, 3, Float64, 36}
│ (𝔻ᵉ = (%24)(%31, %37))
│ %39 = Main.:-::Core.Const(-)
│ %40 = Base.getproperty(state_old, :ε_p)::SYMMETRICTENSOR{2, 3, FLOAT64}
│ (ε_e = (%39)(ε, %40))
│ %42 = Main.:+::Core.Const(+)
│ %43 = Main.:*::Core.Const(*)
│ %44 = λ_val::Float64
│ %45 = Main.tr::Core.Const(LinearAlgebra.tr)
│ %46 = ε_e::SymmetricTensor{2, 3, Float64, 6}
│ %47 = (%45)(%46)::Float64
│ %48 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ %49 = (%43)(%44, %47, %48)::SymmetricTensor{2, 3, Float64, 6}
│ %50 = Main.:*::Core.Const(*)
│ %51 = Main.:*::Core.Const(*)
│ %52 = μ_val::Float64
│ %53 = (%51)(2, %52)::Float64
│ %54 = ε_e::SymmetricTensor{2, 3, Float64, 6}
│ %55 = (%50)(%53, %54)::SymmetricTensor{2, 3, Float64, 6}
│ (σ_trial = (%42)(%49, %55))
│ %57 = Main.von_mises_stress::Core.Const(Main.von_mises_stress)
│ %58 = σ_trial::SymmetricTensor{2, 3, Float64, 6}
│ (σ_eq_trial = (%57)(%58))
│ %60 = Main.:-::Core.Const(-)
│ %61 = σ_eq_trial::Float64
│ %62 = σ_y::Float64
│ (f = (%60)(%61, %62))
│ %64 = Main.:≤::Core.Const(<=)
│ %65 = f::Float64
│ %66 = (%64)(%65, 0.0)::Bool
└── goto #3 if not %66
2 ─ %68 = σ_trial::SymmetricTensor{2, 3, Float64, 6}
│ (σ = %68)
│ %70 = 𝔻ᵉ::SymmetricTensor{4, 3, Float64, 36}
│ (𝔻 = %70)
│ %72 = state_old::PlasticityState{Float64}
│ (state_new = %72)
└── goto #4
3 ─ %75 = Main.dev::Core.Const(Tensors.dev)
│ %76 = σ_trial::SymmetricTensor{2, 3, Float64, 6}
│ (s_trial = (%75)(%76))
│ %78 = Main.:/::Core.Const(/)
│ %79 = Main.tr::Core.Const(LinearAlgebra.tr)
│ %80 = σ_trial::SymmetricTensor{2, 3, Float64, 6}
│ %81 = (%79)(%80)::Float64
│ (p = (%78)(%81, 3))
│ %83 = Main.:+::Core.Const(+)
│ %84 = Main.:*::Core.Const(*)
│ %85 = p::Float64
│ %86 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
│ %87 = (%84)(%85, %86)::SymmetricTensor{2, 3, Float64, 6}
│ %88 = Main.:*::Core.Const(*)
│ %89 = Main.:/::Core.Const(/)
│ %90 = σ_y::Float64
│ %91 = σ_eq_trial::Float64
│ %92 = (%89)(%90, %91)::Float64
│ %93 = s_trial::SymmetricTensor{2, 3, Float64, 6}
│ %94 = (%88)(%92, %93)::SymmetricTensor{2, 3, Float64, 6}
│ (σ = (%83)(%87, %94))
│ %96 = Main.:/::Core.Const(/)
│ %97 = f::Float64
│ %98 = Main.:*::Core.Const(*)
│ %99 = μ_val::Float64
│ %100 = (%98)(3, %99)::Float64
│ (Δγ = (%96)(%97, %100))
│ %102 = Main.:/::Core.Const(/)
│ %103 = Main.:*::Core.Const(*)
│ %104 = Main.:√::Core.Const(sqrt)
│ %105 = Main.:/::Core.Const(/)
│ %106 = (%105)(3, 2)::Core.Const(1.5)
│ %107 = (%104)(%106)::Core.Const(1.224744871391589)
│ %108 = s_trial::SymmetricTensor{2, 3, Float64, 6}
│ %109 = (%103)(%107, %108)::SymmetricTensor{2, 3, Float64, 6}
│ %110 = σ_eq_trial::Float64
│ (n = (%102)(%109, %110))
│ %112 = Main.:+::Core.Const(+)
│ %113 = Base.getproperty(state_old, :ε_p)::SYMMETRICTENSOR{2, 3, FLOAT64}
│ %114 = Main.:*::Core.Const(*)
│ %115 = Δγ::Float64
│ %116 = n::SymmetricTensor{2, 3, Float64, 6}
│ %117 = (%114)(%115, %116)::SymmetricTensor{2, 3, Float64, 6}
│ (ε_p_new = (%112)(%113, %117))
│ %119 = Main.:+::Core.Const(+)
│ %120 = Base.getproperty(state_old, :α)::Float64
│ %121 = Δγ::Float64
│ (α_new = (%119)(%120, %121))
│ %123 = Main.PlasticityState::Core.Const(PlasticityState)
│ %124 = ε_p_new::SymmetricTensor{2, 3, Float64, 6}
│ %125 = α_new::Float64
│ (state_new = (%123)(%124, %125))
│ %127 = Main.:-::Core.Const(-)
│ %128 = Main.:/::Core.Const(/)
│ %129 = σ_y::Float64
│ %130 = σ_eq_trial::Float64
│ %131 = (%128)(%129, %130)::Float64
│ (θ = (%127)(1, %131))
│ %133 = Main.:/::Core.Const(/)
│ %134 = Main.:*::Core.Const(*)
│ %135 = Main.:^::Core.Const(^)
│ %136 = μ_val::Float64
│ %137 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
│ %138 = (%137)()::Core.Const(Val{2}())
│ %139 = Base.literal_pow(%135, %136, %138)::Float64
│ %140 = (%134)(6, %139)::Float64
│ %141 = Main.:+::Core.Const(+)
│ %142 = Main.:*::Core.Const(*)
│ %143 = μ_val::Float64
│ %144 = (%142)(3, %143)::Float64
│ %145 = Main.:*::Core.Const(*)
│ %146 = θ::Float64
│ %147 = Main.:*::Core.Const(*)
│ %148 = μ_val::Float64
│ %149 = (%147)(3, %148)::Float64
│ %150 = (%145)(%146, %149)::Float64
│ %151 = (%141)(%144, %150)::Float64
│ (β = (%133)(%140, %151))
│ %153 = Main.:-::Core.Const(-)
│ %154 = 𝔻ᵉ::SymmetricTensor{4, 3, Float64, 36}
│ %155 = Main.:*::Core.Const(*)
│ %156 = β::Float64
│ %157 = Main.:⊗::Core.Const(Tensors.otimes)
│ %158 = n::SymmetricTensor{2, 3, Float64, 6}
│ %159 = n::SymmetricTensor{2, 3, Float64, 6}
│ %160 = (%157)(%158, %159)::SymmetricTensor{4, 3, Float64, 36}
│ %161 = (%155)(%156, %160)::SymmetricTensor{4, 3, Float64, 36}
└── (𝔻 = (%153)(%154, %161))
4 ┄ %163 = σ::SymmetricTensor{2, 3, Float64, 6}
│ %164 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
│ %165 = state_new::PlasticityState{Float64}
│ %166 = Core.tuple(%163, %164, %165)::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, PlasticityState{Float64}}
└── return %166
6. Perfect Plasticity (Old Dict):
MethodInstance for compute_stress_old(::PerfectPlasticityOld, ::Vector{Float64}, ::Dict{String, Any}, ::Float64)
from compute_stress_old(material::PerfectPlasticityOld, ε_vec::Vector{Float64}, state_old::Dict{String, Any}, Δt::Float64) @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:455
Arguments
#self#::Core.Const(Main.compute_stress_old)
material::PerfectPlasticityOld
ε_vec::Vector{Float64}
state_old::Dict{String, Any}
Δt::Float64
Locals
@_6::ANY
@_7::ANY
σ_vec::ANY
n_vec::ANY
Δγ::ANY
factor::ANY
state_new::Dict{String, Any}
f::ANY
σ_eq::ANY
dev_vec::ANY
p::ANY
s13::ANY
s23::ANY
s12::ANY
s33::ANY
s22::ANY
s11::ANY
σ_trial_vec::ANY
ε_e_vec::ANY
D::Matrix{Float64}
ε_p_vec::ANY
Body::TUPLE{ANY, MATRIX{FLOAT64}, DICT{STRING, ANY}}
1 ─ Core.NewvarNode(:(@_6))
│ Core.NewvarNode(:(@_7))
│ Core.NewvarNode(:(σ_vec))
│ Core.NewvarNode(:(n_vec))
│ Core.NewvarNode(:(Δγ))
│ Core.NewvarNode(:(factor))
│ Core.NewvarNode(:(state_new))
│ Core.NewvarNode(:(f))
│ Core.NewvarNode(:(σ_eq))
│ Core.NewvarNode(:(dev_vec))
│ Core.NewvarNode(:(p))
│ Core.NewvarNode(:(s13))
│ Core.NewvarNode(:(s23))
│ Core.NewvarNode(:(s12))
│ Core.NewvarNode(:(s33))
│ Core.NewvarNode(:(s22))
│ Core.NewvarNode(:(s11))
│ Core.NewvarNode(:(σ_trial_vec))
│ Core.NewvarNode(:(ε_e_vec))
│ Core.NewvarNode(:(D))
│ Core.NewvarNode(:(ε_p_vec))
│ %22 = Main.haskey::Core.Const(haskey)
│ %23 = (%22)(state_old, "epsilon_plastic")::Bool
└── goto #3 if not %23
2 ─ (ε_p_vec = Base.getindex(state_old, "epsilon_plastic"))
└── goto #4
3 ─ %27 = Main.zeros::Core.Const(zeros)
└── (ε_p_vec = (%27)(6))
4 ┄ %29 = Main.constitutive_matrix::Core.Const(Main.constitutive_matrix)
│ %30 = Main.LinearElasticOld::Core.Const(LinearElasticOld)
│ %31 = Base.getproperty(material, :E)::Float64
│ %32 = Base.getproperty(material, :ν)::Float64
│ %33 = (%30)(%31, %32)::LinearElasticOld
│ (D = (%29)(%33))
│ %35 = Main.:-::Core.Const(-)
│ %36 = ε_p_vec::ANY
│ (ε_e_vec = (%35)(ε_vec, %36))
│ %38 = Main.:*::Core.Const(*)
│ %39 = D::Matrix{Float64}
│ %40 = ε_e_vec::ANY
│ (σ_trial_vec = (%38)(%39, %40))
│ %42 = σ_trial_vec::ANY
│ %43 = Main.:(:)::Core.Const(Colon())
│ %44 = (%43)(1, 3)::Core.Const(1:3)
│ %45 = Base.getindex(%42, %44)::ANY
│ %46 = Base.indexed_iterate(%45, 1)::ANY
│ (s11 = Core.getfield(%46, 1))
│ (@_7 = Core.getfield(%46, 2))
│ %49 = @_7::ANY
│ %50 = Base.indexed_iterate(%45, 2, %49)::ANY
│ (s22 = Core.getfield(%50, 1))
│ (@_7 = Core.getfield(%50, 2))
│ %53 = @_7::ANY
│ %54 = Base.indexed_iterate(%45, 3, %53)::ANY
│ (s33 = Core.getfield(%54, 1))
│ %56 = σ_trial_vec::ANY
│ %57 = Main.:(:)::Core.Const(Colon())
│ %58 = (%57)(4, 6)::Core.Const(4:6)
│ %59 = Base.getindex(%56, %58)::ANY
│ %60 = Base.indexed_iterate(%59, 1)::ANY
│ (s12 = Core.getfield(%60, 1))
│ (@_6 = Core.getfield(%60, 2))
│ %63 = @_6::ANY
│ %64 = Base.indexed_iterate(%59, 2, %63)::ANY
│ (s23 = Core.getfield(%64, 1))
│ (@_6 = Core.getfield(%64, 2))
│ %67 = @_6::ANY
│ %68 = Base.indexed_iterate(%59, 3, %67)::ANY
│ (s13 = Core.getfield(%68, 1))
│ %70 = Main.:/::Core.Const(/)
│ %71 = Main.:+::Core.Const(+)
│ %72 = s11::ANY
│ %73 = s22::ANY
│ %74 = s33::ANY
│ %75 = (%71)(%72, %73, %74)::ANY
│ (p = (%70)(%75, 3))
│ %77 = Main.:-::Core.Const(-)
│ %78 = s11::ANY
│ %79 = p::ANY
│ %80 = (%77)(%78, %79)::ANY
│ %81 = Main.:-::Core.Const(-)
│ %82 = s22::ANY
│ %83 = p::ANY
│ %84 = (%81)(%82, %83)::ANY
│ %85 = Main.:-::Core.Const(-)
│ %86 = s33::ANY
│ %87 = p::ANY
│ %88 = (%85)(%86, %87)::ANY
│ %89 = s12::ANY
│ %90 = s23::ANY
│ %91 = s13::ANY
│ (dev_vec = Base.vect(%80, %84, %88, %89, %90, %91))
│ %93 = Main.:√::Core.Const(sqrt)
│ %94 = Main.:*::Core.Const(*)
│ %95 = Main.:/::Core.Const(/)
│ %96 = (%95)(3, 2)::Core.Const(1.5)
│ %97 = Main.:+::Core.Const(+)
│ %98 = Main.:^::Core.Const(^)
│ %99 = dev_vec::ANY
│ %100 = Base.getindex(%99, 1)::ANY
│ %101 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
│ %102 = (%101)()::Core.Const(Val{2}())
│ %103 = Base.literal_pow(%98, %100, %102)::ANY
│ %104 = Main.:^::Core.Const(^)
│ %105 = dev_vec::ANY
│ %106 = Base.getindex(%105, 2)::ANY
│ %107 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
│ %108 = (%107)()::Core.Const(Val{2}())
│ %109 = Base.literal_pow(%104, %106, %108)::ANY
│ %110 = Main.:^::Core.Const(^)
│ %111 = dev_vec::ANY
│ %112 = Base.getindex(%111, 3)::ANY
│ %113 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
│ %114 = (%113)()::Core.Const(Val{2}())
│ %115 = Base.literal_pow(%110, %112, %114)::ANY
│ %116 = Main.:*::Core.Const(*)
│ %117 = Main.:+::Core.Const(+)
│ %118 = Main.:^::Core.Const(^)
│ %119 = dev_vec::ANY
│ %120 = Base.getindex(%119, 4)::ANY
│ %121 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
│ %122 = (%121)()::Core.Const(Val{2}())
│ %123 = Base.literal_pow(%118, %120, %122)::ANY
│ %124 = Main.:^::Core.Const(^)
│ %125 = dev_vec::ANY
│ %126 = Base.getindex(%125, 5)::ANY
│ %127 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
│ %128 = (%127)()::Core.Const(Val{2}())
│ %129 = Base.literal_pow(%124, %126, %128)::ANY
│ %130 = Main.:^::Core.Const(^)
│ %131 = dev_vec::ANY
│ %132 = Base.getindex(%131, 6)::ANY
│ %133 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
│ %134 = (%133)()::Core.Const(Val{2}())
│ %135 = Base.literal_pow(%130, %132, %134)::ANY
│ %136 = (%117)(%123, %129, %135)::ANY
│ %137 = (%116)(2, %136)::ANY
│ %138 = (%97)(%103, %109, %115, %137)::ANY
│ %139 = (%94)(%96, %138)::ANY
│ (σ_eq = (%93)(%139))
│ %141 = Main.:-::Core.Const(-)
│ %142 = σ_eq::ANY
│ %143 = Base.getproperty(material, :σ_y)::Float64
│ (f = (%141)(%142, %143))
│ %145 = Main.copy::Core.Const(copy)
│ (state_new = (%145)(state_old))
│ %147 = Main.:>::Core.Const(>)
│ %148 = f::ANY
│ %149 = (%147)(%148, 0.0)::ANY
└── goto #6 if not %149
5 ─ %151 = Main.:/::Core.Const(/)
│ %152 = Base.getproperty(material, :σ_y)::Float64
│ %153 = σ_eq::ANY
│ (factor = (%151)(%152, %153))
│ %155 = Main.:+::Core.Const(+)
│ %156 = p::ANY
│ %157 = p::ANY
│ %158 = p::ANY
│ %159 = Base.vect(%156, %157, %158, 0.0, 0.0, 0.0)::ANY
│ %160 = Main.:*::Core.Const(*)
│ %161 = factor::ANY
│ %162 = dev_vec::ANY
│ %163 = (%160)(%161, %162)::ANY
│ (σ_vec = (%155)(%159, %163))
│ %165 = Main.:/::Core.Const(/)
│ %166 = f::ANY
│ %167 = Main.:/::Core.Const(/)
│ %168 = Main.:*::Core.Const(*)
│ %169 = Base.getproperty(material, :E)::Float64
│ %170 = (%168)(3, %169)::Float64
│ %171 = Main.:*::Core.Const(*)
│ %172 = Main.:+::Core.Const(+)
│ %173 = Base.getproperty(material, :ν)::Float64
│ %174 = (%172)(1, %173)::Float64
│ %175 = (%171)(2, %174)::Float64
│ %176 = (%167)(%170, %175)::Float64
│ (Δγ = (%165)(%166, %176))
│ %178 = Main.:/::Core.Const(/)
│ %179 = Main.:*::Core.Const(*)
│ %180 = Main.:√::Core.Const(sqrt)
│ %181 = Main.:/::Core.Const(/)
│ %182 = (%181)(3, 2)::Core.Const(1.5)
│ %183 = (%180)(%182)::Core.Const(1.224744871391589)
│ %184 = dev_vec::ANY
│ %185 = (%179)(%183, %184)::ANY
│ %186 = σ_eq::ANY
│ (n_vec = (%178)(%185, %186))
│ %188 = Main.:+::Core.Const(+)
│ %189 = ε_p_vec::ANY
│ %190 = Main.:*::Core.Const(*)
│ %191 = Δγ::ANY
│ %192 = n_vec::ANY
│ %193 = (%190)(%191, %192)::ANY
│ %194 = (%188)(%189, %193)::ANY
│ %195 = state_new::Dict{String, Any}
│ Base.setindex!(%195, %194, "epsilon_plastic")
└── goto #7
6 ─ %198 = σ_trial_vec::ANY
└── (σ_vec = %198)
7 ┄ %200 = σ_vec::ANY
│ %201 = D::Matrix{Float64}
│ %202 = state_new::Dict{String, Any}
│ %203 = Core.tuple(%200, %201, %202)::TUPLE{ANY, MATRIX{FLOAT64}, DICT{STRING, ANY}}
└── return %203
================================================================================
ALLOCATION TESTS
================================================================================
Testing for allocations (should be 0 for new approach)...
1. Linear Elastic
NEW (Tensors.jl):
Allocations: 0 bytes
OLD (Voigt/Dict):
Allocations: 496 bytes
2. Neo-Hookean
NEW (Tensors.jl + AD):
Allocations: 0 bytes
NEW (Tensors.jl + Manual):
Allocations: 0 bytes
OLD (Array):
Allocations: 496 bytes
3. Perfect Plasticity (elastic branch)
NEW (Tensors.jl):
Allocations: 0 bytes
OLD (Dict):
Allocations: 8828848 bytes
================================================================================
PERFORMANCE BENCHMARKS
================================================================================
Running detailed benchmarks (this may take a minute)...
1. LINEAR ELASTIC
----------------------------------------
NEW (Tensors.jl):
BenchmarkTools.Trial: 10000 samples with 997 evaluations per sample.
Range (min … max): 19.464 ns … 45.831 ns ┊ GC (min … max): 0.00% … 0.00%
Time (median): 19.577 ns ┊ GC (median): 0.00%
Time (mean ± σ): 19.670 ns ± 0.675 ns ┊ GC (mean ± σ): 0.00% ± 0.00%
▁█▄
███▇▄▃▂▂▂▂▂▁▁▂▂▂▂▂▂▂▂▂▂▂▂▁▁▁▂▁▁▁▂▁▁▁▂▂▁▁▁▁▂▁▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂ ▂
19.5 ns Histogram: frequency by time 22.8 ns <
Memory estimate: 0 bytes, allocs estimate: 0.
OLD (Voigt/Dict):
BenchmarkTools.Trial: 10000 samples with 950 evaluations per sample.
Range (min … max): 93.356 ns … 8.107 μs ┊ GC (min … max): 0.00% … 97.34%
Time (median): 100.107 ns ┊ GC (median): 0.00%
Time (mean ± σ): 139.733 ns ± 249.840 ns ┊ GC (mean ± σ): 25.28% ± 13.73%
█▂ ▁ ▁
██▄▄██▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▆▇▇▄▃▄▅▄▄▅▅▅▅▅▃▆ █
93.4 ns Histogram: log(frequency) by time 1.74 μs <
Memory estimate: 496 bytes, allocs estimate: 4.
SPEEDUP: 5.1×
2. NEO-HOOKEAN
----------------------------------------
NEW (Tensors.jl + Automatic Differentiation):
BenchmarkTools.Trial: 10000 samples with 23 evaluations per sample.
Range (min … max): 1.050 μs … 2.802 μs ┊ GC (min … max): 0.00% … 0.00%
Time (median): 1.051 μs ┊ GC (median): 0.00%
Time (mean ± σ): 1.055 μs ± 31.780 ns ┊ GC (mean ± σ): 0.00% ± 0.00%
█▄▂▂▁▂▁▂▂▁▁▁▁▁▁▁▁▁▂▁▁▂▁▁▁▁▁▁▁▁▁▁▁▂▁▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂ ▂
1.05 μs Histogram: frequency by time 1.19 μs <
Memory estimate: 0 bytes, allocs estimate: 0.
NEW (Tensors.jl + Manual Derivatives):
BenchmarkTools.Trial: 10000 samples with 987 evaluations per sample.
Range (min … max): 49.806 ns … 1.787 μs ┊ GC (min … max): 0.00% … 0.00%
Time (median): 49.922 ns ┊ GC (median): 0.00%
Time (mean ± σ): 50.262 ns ± 17.399 ns ┊ GC (mean ± σ): 0.00% ± 0.00%
▅██▄▁ ▂
█████▆▄▁▄▄▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▁▃▄▄▅▅▆▆▆▆▆▇▇█▇▇▇▇▇██▇▇▆▇▇ █
49.8 ns Histogram: log(frequency) by time 53.3 ns <
Memory estimate: 0 bytes, allocs estimate: 0.
OLD (Array):
BenchmarkTools.Trial: 10000 samples with 955 evaluations per sample.
Range (min … max): 91.182 ns … 9.723 μs ┊ GC (min … max): 0.00% … 97.56%
Time (median): 99.922 ns ┊ GC (median): 0.00%
Time (mean ± σ): 142.795 ns ± 307.090 ns ┊ GC (mean ± σ): 22.77% ± 11.92%
█▃ ▄▁ ▁
██▆▄██▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▄▇█ █
91.2 ns Histogram: log(frequency) by time 1.77 μs <
Memory estimate: 496 bytes, allocs estimate: 4.
SPEEDUP (AD): 0.1×
SPEEDUP (Manual): 2.0×
AD OVERHEAD: 21.1× (AD / Manual)
3. PERFECT PLASTICITY (elastic branch)
----------------------------------------
NEW (Tensors.jl):
BenchmarkTools.Trial: 10000 samples with 976 evaluations per sample.
Range (min … max): 69.677 ns … 151.814 ns ┊ GC (min … max): 0.00% … 0.00%
Time (median): 70.389 ns ┊ GC (median): 0.00%
Time (mean ± σ): 70.566 ns ± 1.326 ns ┊ GC (mean ± σ): 0.00% ± 0.00%
▂▆▇▆▇▆▇▇▇█▇▆▇▆▅▃ ▁▁▁▁▁▁▂▁▁▁▁▁ ▃
▆███████████████████▇▇▇▆▇▆▆▄▃▁▁▁▁▁▆▅▅▆▇▆▇▇██████████████████ █
69.7 ns Histogram: log(frequency) by time 73.9 ns <
Memory estimate: 0 bytes, allocs estimate: 0.
OLD (Dict):
BenchmarkTools.Trial: 10000 samples with 10 evaluations per sample.
Range (min … max): 1.371 μs … 998.345 μs ┊ GC (min … max): 0.00% … 99.45%
Time (median): 1.480 μs ┊ GC (median): 0.00%
Time (mean ± σ): 1.701 μs ± 9.970 μs ┊ GC (mean ± σ): 5.84% ± 0.99%
▅█▆▂
▁▂▃▆████▆▅▄▃▃▃▃▂▂▂▂▂▁▁▂▁▁▁▁▁▂▂▂▂▂▂▂▂▂▂▃▃▃▃▂▂▂▂▂▂▂▂▂▂▁▁▁▁▁▁▁ ▂
1.37 μs Histogram: frequency by time 2.16 μs <
Memory estimate: 1.98 KiB, allocs estimate: 53.
SPEEDUP: 21.0×
================================================================================
SUMMARY
================================================================================
ALLOCATIONS:
LinearElastic: NEW = 0 bytes, OLD = 496 bytes
NeoHookean (AD): NEW = 0 bytes, OLD = 496 bytes
NeoHookean (Manual): NEW = 0 bytes
PerfectPlasticity: NEW = 0 bytes, OLD = 8828848 bytes
MEDIAN TIMING:
LinearElastic: NEW = 19.576730190571716 ns, OLD = 100.10684210526315 ns
NeoHookean (AD): NEW = 1051.304347826087 ns, OLD = 99.92198952879582 ns
NeoHookean (Manual): NEW = 49.92198581560284 ns
PerfectPlasticity: NEW = 70.38934426229508 ns, OLD = 1479.55 ns
SPEEDUP (OLD / NEW):
LinearElastic: 5.1×
NeoHookean (AD): 0.1×
NeoHookean (Manual): 2.0×
PerfectPlasticity: 21.0×
AD OVERHEAD:
NeoHookean: AD is 21.1× slower than manual derivatives
AVERAGE SPEEDUP: 9.4× (using manual Neo-Hookean)
VALIDATION OF CLAIMS:
- Zero allocations for new approach: ✓ PASS
- Manual derivatives outperform AD: ✓ PASS
- Type stability with NoState return: Check @code_warntype output above
================================================================================
Benchmark complete! Results saved to: material_models_benchmark_results.txt
================================================================================
-835
View File
@@ -1,835 +0,0 @@
"""
Matrix-Free Newton-Krylov GPU Benchmark
Demonstrates traditional Newton vs Matrix-Free Newton-Krylov with Anderson
acceleration, running on GPU.
Run with:
julia --project=. benchmarks/matrix_free_gpu_benchmark.jl
"""
using CUDA
using LinearAlgebra
using IterativeSolvers
using Printf
# Check GPU availability
if !CUDA.functional()
@warn "CUDA not available! Running CPU-only comparison."
USE_GPU = false
else
println("GPU Device: $(CUDA.name(CUDA.device()))")
println("GPU Memory: $(CUDA.total_memory() / 1e9) GB")
println()
USE_GPU = true
end
# ============================================================================
# Problem Setup: 3D Nonlinear Elasticity
# ============================================================================
"""
Residual for 3D nonlinear elasticity with cubic nonlinearity.
r(u) = K·u + β·(K·u).^3 - f
where K is stiffness matrix, β is nonlinearity parameter.
"""
struct NonlinearProblem{T,MatT,VecT}
K::MatT # Stiffness matrix (sparse or LinearMap)
f::VecT # Force vector
β::T # Nonlinearity parameter
n::Int # DOF count
end
"""
Compute residual: r(u) = K·u + β·(K·u).^3 - f
"""
function compute_residual!(r::AbstractVector, prob::NonlinearProblem, u::AbstractVector)
# Linear part
mul!(r, prob.K, u) # r = K·u
# Nonlinear part: r += β·(K·u).^3
if prob.β != 0
# Reuse r (which contains K·u)
@. r = r + prob.β * r^3
end
# Apply forcing
@. r = r - prob.f
return r
end
"""
Jacobian-vector product: J·v [r(u+ε·v) - r(u)] / ε (finite difference)
"""
function jacobian_vector_product!(
Jv::AbstractVector,
prob::NonlinearProblem,
u::AbstractVector,
v::AbstractVector,
r_u::AbstractVector, # Pre-computed r(u)
temp::AbstractVector # Workspace
)
ε = 1e-7
# temp = u + ε·v
@. temp = u + ε * v
# Jv = r(u + ε·v)
compute_residual!(Jv, prob, temp)
# Jv = [r(u + ε·v) - r(u)] / ε
@. Jv = (Jv - r_u) / ε
return Jv
end
# ============================================================================
# Traditional Newton Solver
# ============================================================================
"""
Traditional Newton with full Jacobian assembly.
u_{k+1} = u_k - J(u_k)^{-1} · r(u_k)
Expensive: Assembles full Jacobian matrix at each iteration.
"""
function newton_traditional!(
u::AbstractVector{T},
prob::NonlinearProblem{T},
r::AbstractVector{T},
du::AbstractVector{T};
tol=1e-8,
max_iter=20,
verbose=true
) where T
n = length(u)
# Build Jacobian matrix (expensive!)
# J ≈ K + 3β·diag((K·u).^2)·K
Ku = prob.K * u
J = copy(prob.K)
for iter in 1:max_iter
# Compute residual
compute_residual!(r, prob, u)
norm_r = norm(r)
if verbose
@printf(" Iter %2d: ||r|| = %.6e\n", iter, norm_r)
end
if norm_r < tol
if verbose
println(" ✅ Converged!")
end
return iter
end
# Update Jacobian (expensive!)
Ku .= prob.K * u
for i in 1:n
J[i, i] = prob.K[i, i] + 3 * prob.β * Ku[i]^2 * prob.K[i, i]
end
# Solve linear system (expensive!)
du .= -(J \ r)
# Update
u .+= du
end
if verbose
println(" ⚠️ Did not converge in $max_iter iterations")
end
return max_iter
end
# ============================================================================
# Helper: Matrix-free operator wrapper for GMRES
# ============================================================================
"""
Wrapper to make a function look like a matrix for GMRES.
"""
struct MatrixFreeOperator{F}
matvec!::F
n::Int
end
Base.size(A::MatrixFreeOperator) = (A.n, A.n)
Base.size(A::MatrixFreeOperator, d::Int) = d <= 2 ? A.n : 1
Base.eltype(::MatrixFreeOperator{F}) where F = Float64
function LinearAlgebra.mul!(y, A::MatrixFreeOperator, x)
A.matvec!(y, x)
return y
end
# ============================================================================
# Matrix-Free Newton-Krylov
# ============================================================================
"""
Matrix-Free Newton-Krylov with GMRES.
J·v [r(u+ε·v) - r(u)] / ε (no matrix!)
du = gmres(Jv_op, -r)
u_{k+1} = u_k + du
Cheap: Only residual evaluations, no Jacobian assembly.
"""
function newton_matrix_free!(
u::AbstractVector{T},
prob::NonlinearProblem{T},
r::AbstractVector{T},
du::AbstractVector{T},
temp::AbstractVector{T},
Jv::AbstractVector{T};
tol=1e-8,
max_iter=20,
gmres_tol=1e-6,
verbose=true
) where T
for iter in 1:max_iter
# Compute residual
compute_residual!(r, prob, u)
norm_r = norm(r)
if verbose
@printf(" Iter %2d: ||r|| = %.6e", iter, norm_r)
end
if norm_r < tol
if verbose
println(" ✅ Converged!")
end
return iter
end
# Matrix-free operator: J·v
function Jv_matvec!(out, v)
jacobian_vector_product!(out, prob, u, v, r, temp)
return out
end
Jv_op = MatrixFreeOperator(Jv_matvec!, length(u))
# Solve J·du = -r using GMRES (matrix-free!)
du .= 0
gmres!(du, Jv_op, -r;
abstol=gmres_tol, reltol=0, maxiter=50, verbose=false)
gmres_iters = 50 # Would need to extract from gmres! return
if verbose
@printf(" [GMRES: ~%d iters]\n", gmres_iters)
end
# Update
u .+= du
end
if verbose
println(" ⚠️ Did not converge in $max_iter iterations")
end
return max_iter
end
# ============================================================================
# Anderson-Accelerated Newton-Krylov
# ============================================================================
"""
Anderson acceleration for Newton-Krylov.
Combines m previous iterates via least-squares:
u_new = αᵢ·uᵢ where argmin || αᵢ·rᵢ||² s.t. αᵢ = 1
Transforms linear convergence superlinear convergence.
"""
function anderson_newton_matrix_free!(
u::AbstractVector{T},
prob::NonlinearProblem{T},
r::AbstractVector{T},
du::AbstractVector{T},
temp::AbstractVector{T},
Jv::AbstractVector{T};
m=5, # Anderson history
tol=1e-8,
max_iter=20,
gmres_tol=1e-6,
verbose=true
) where T
n = length(u)
# Anderson history
U_history = [zeros(T, n) for _ in 1:m]
R_history = [zeros(T, n) for _ in 1:m]
history_count = 0
for iter in 1:max_iter
# Compute residual
compute_residual!(r, prob, u)
norm_r = norm(r)
if verbose
@printf(" Iter %2d: ||r|| = %.6e", iter, norm_r)
end
if norm_r < tol
if verbose
println(" ✅ Converged!")
end
return iter
end
# Matrix-free operator
function Jv_matvec!(out, v)
jacobian_vector_product!(out, prob, u, v, r, temp)
return out
end
Jv_op = MatrixFreeOperator(Jv_matvec!, n)
# Solve J·du = -r using GMRES
du .= 0
gmres!(du, Jv_op, -r;
abstol=gmres_tol, reltol=0, maxiter=50, verbose=false)
# Store in history (circular buffer)
idx = mod1(history_count + 1, m)
U_history[idx] .= u
R_history[idx] .= r
history_count = min(history_count + 1, m)
if verbose
@printf(" [GMRES: ~50 iters, history: %d]", history_count)
end
# Anderson acceleration (if enough history)
if history_count >= 2
# Build residual difference matrix
k = history_count
R_diff = zeros(T, n, k)
for i in 1:k
R_diff[:, i] .= R_history[i] .- r
end
# Check condition number before QR
# If matrix is ill-conditioned, skip Anderson this iteration
R_norm = norm(R_diff)
if R_norm < 1e-10
# Matrix too small, use standard update
u .+= du
if verbose
println(" [Anderson: skipped (residuals too small)]")
end
else
# Least-squares: min ||R_diff·α||² s.t. sum(α) = 1
# Use QR factorization with regularization
try
Q, Rt = qr(R_diff)
# Add small regularization to diagonal if needed
Rt_diag = diag(Rt)
if any(abs.(Rt_diag) .< 1e-12)
# Add Tikhonov regularization
λ = 1e-8
Rt_reg = Rt + λ * I
α = Rt_reg \ (Q' * r)
else
α = Rt \ (Q' * r)
end
α ./= sum(α) # Normalize
# Combine previous iterates
u_combined = zeros(T, n)
for i in 1:k
u_combined .+= α[i] .* U_history[i]
end
# Update with combination
u .= u_combined .+ du
if verbose
println(" [Anderson: α=$(round.(α, digits=3))]")
end
catch e
# If QR fails, fall back to standard Newton
u .+= du
if verbose
println(" [Anderson: failed ($e), using standard update]")
end
end
end
else
# Standard Newton update
u .+= du
if verbose
println()
end
end
end
if verbose
println(" ⚠️ Did not converge in $max_iter iterations")
end
return max_iter
end
# ============================================================================
# GPU Implementations
# ============================================================================
"""
GPU version of residual computation.
"""
function compute_residual_gpu!(
r::CuVector{T},
K::CuMatrix{T},
u::CuVector{T},
f::CuVector{T},
β::T
) where T
# r = K·u
mul!(r, K, u)
# r = r + β·r³ - f
if β != 0
r .= r .+ β .* r .^ 3 .- f
else
r .= r .- f
end
return r
end
"""
GPU version of Matrix-Free Newton-Krylov.
"""
function newton_matrix_free_gpu!(
u::CuVector{T},
K::CuMatrix{T},
f::CuVector{T},
β::T;
tol=1e-8,
max_iter=20,
gmres_tol=1e-6,
verbose=true
) where T
n = length(u)
r = CUDA.zeros(T, n)
du = CUDA.zeros(T, n)
temp = CUDA.zeros(T, n)
Jv = CUDA.zeros(T, n)
for iter in 1:max_iter
# Compute residual on GPU
compute_residual_gpu!(r, K, u, f, β)
norm_r = norm(Array(r)) # Transfer to CPU for norm
if verbose
@printf(" Iter %2d: ||r|| = %.6e\n", iter, norm_r)
end
if norm_r < tol
if verbose
println(" ✅ Converged!")
end
return iter
end
# Jacobian-vector product (on GPU)
ε = T(1e-7)
function Jv_matvec_gpu!(out_cpu, v_cpu)
v = CuArray(v_cpu)
# temp = u + ε·v
temp .= u .+ ε .* v
# Jv = r(u + ε·v)
compute_residual_gpu!(Jv, K, temp, f, β)
# Jv = [r(u + ε·v) - r(u)] / ε
Jv .= (Jv .- r) ./ ε
out_cpu .= Array(Jv)
return out_cpu
end
Jv_op_gpu = MatrixFreeOperator(Jv_matvec_gpu!, n)
# Solve on CPU (GMRES doesn't have GPU version in IterativeSolvers.jl)
r_cpu = Array(r)
du_cpu = zeros(T, n)
gmres!(du_cpu, Jv_op_gpu, -r_cpu;
abstol=gmres_tol, reltol=0, maxiter=50, verbose=false)
# Update on GPU
du .= CuArray(du_cpu)
u .+= du
end
if verbose
println(" ⚠️ Did not converge in $max_iter iterations")
end
return max_iter
end
# GPU Anderson-Accelerated Newton-Krylov
function anderson_newton_matrix_free_gpu!(
u::CuVector{T},
K::CuMatrix{T},
f::CuVector{T},
β::T;
tol=1e-8,
max_iter=20,
gmres_tol=1e-6,
history_size=5,
verbose=true
) where T
n = length(u)
r = CUDA.zeros(T, n)
du = CUDA.zeros(T, n)
temp = CUDA.zeros(T, n)
Jv = CUDA.zeros(T, n)
# Anderson acceleration storage (CPU)
R_history = Vector{Vector{T}}()
U_history = Vector{Vector{T}}()
history_count = 0
for iter in 1:max_iter
# Compute residual on GPU
compute_residual_gpu!(r, K, u, f, β)
norm_r = norm(Array(r)) # Transfer to CPU for norm
if verbose
@printf(" Iter %2d: ||r|| = %.6e", iter, norm_r)
end
if norm_r < tol
if verbose
println("\n ✅ Converged!")
end
return iter
end
# Jacobian-vector product (on GPU)
ε = T(1e-7)
function Jv_matvec_gpu!(out_cpu, v_cpu)
v = CuArray(v_cpu)
# temp = u + ε·v
temp .= u .+ ε .* v
# Jv = r(u + ε·v)
compute_residual_gpu!(Jv, K, temp, f, β)
# Jv = [r(u + ε·v) - r(u)] / ε
Jv .= (Jv .- r) ./ ε
out_cpu .= Array(Jv)
return out_cpu
end
Jv_op_gpu = MatrixFreeOperator(Jv_matvec_gpu!, n)
# Solve on CPU (GMRES doesn't have GPU version)
r_cpu = Array(r)
u_cpu = Array(u)
du_cpu = zeros(T, n)
gmres!(du_cpu, Jv_op_gpu, -r_cpu;
abstol=gmres_tol, reltol=0, maxiter=50, verbose=false)
# Anderson acceleration (on CPU)
if history_count >= 2
# Build residual difference matrix
k = history_count
R_diff = zeros(T, n, k)
for i in 1:k
R_diff[:, i] .= R_history[i] .- r_cpu
end
# Check condition number before QR
R_norm = norm(R_diff)
if R_norm < 1e-10
# Matrix too small, use standard update
u .+= CuArray(du_cpu)
if verbose
println(" [Anderson: skipped (residuals too small)]")
end
else
# Least-squares with regularization
try
Q, Rt = qr(R_diff)
# Add small regularization to diagonal if needed
Rt_diag = diag(Rt)
if any(abs.(Rt_diag) .< 1e-12)
# Add Tikhonov regularization
λ = 1e-8
Rt_reg = Rt + λ * I
α = Rt_reg \ (Q' * r_cpu)
else
α = Rt \ (Q' * r_cpu)
end
α ./= sum(α) # Normalize
# Combine previous iterates
u_combined = zeros(T, n)
for i in 1:k
u_combined .+= α[i] .* U_history[i]
end
# Update with combination (transfer to GPU)
u .= CuArray(u_combined .+ du_cpu)
if verbose
println(" [Anderson: α=$(round.(α, digits=3))]")
end
catch e
# If QR fails, fall back to standard Newton
u .+= CuArray(du_cpu)
if verbose
println(" [Anderson: failed ($e), using standard update]")
end
end
end
else
# Standard Newton update (transfer du to GPU)
u .+= CuArray(du_cpu)
if verbose
println()
end
end
# Store history (on CPU to avoid GPU memory overhead)
push!(R_history, copy(r_cpu))
push!(U_history, copy(u_cpu))
history_count += 1
# Maintain history size
if history_count > history_size
popfirst!(R_history)
popfirst!(U_history)
history_count = history_size
end
end
if verbose
println(" ⚠️ Did not converge in $max_iter iterations")
end
return max_iter
end
# ============================================================================
# Benchmark Runners
# ============================================================================
function benchmark_cpu(n::Int)
println("="^70)
println("CPU Benchmark: $n DOFs")
println("="^70)
# Setup problem
T = Float64
K = Matrix(Tridiagonal(
-ones(T, n - 1),
2ones(T, n),
-ones(T, n - 1)
))
f = ones(T, n) * 0.1
β = T(1e-3) # Nonlinearity
prob = NonlinearProblem(K, f, β, n)
# Initial guess
u0 = zeros(T, n)
# Allocate workspace
r = zeros(T, n)
du = zeros(T, n)
temp = zeros(T, n)
Jv = zeros(T, n)
# Benchmark Traditional Newton
println("\n📊 Traditional Newton (Full Jacobian):")
u_trad = copy(u0)
t_trad = @elapsed iters_trad = newton_traditional!(u_trad, prob, r, du; verbose=false)
println(" Time: $(round(t_trad * 1000, digits=2)) ms")
println(" Iterations: $iters_trad")
println(" Time/iter: $(round(t_trad / iters_trad * 1000, digits=2)) ms")
# Benchmark Matrix-Free
println("\n📊 Matrix-Free Newton-Krylov:")
u_mf = copy(u0)
t_mf = @elapsed iters_mf = newton_matrix_free!(
u_mf, prob, r, du, temp, Jv; verbose=false
)
println(" Time: $(round(t_mf * 1000, digits=2)) ms")
println(" Iterations: $iters_mf")
println(" Time/iter: $(round(t_mf / iters_mf * 1000, digits=2)) ms")
# Benchmark Anderson-Accelerated
println("\n📊 Anderson-Accelerated Matrix-Free:")
u_anderson = copy(u0)
t_anderson = @elapsed iters_anderson = anderson_newton_matrix_free!(
u_anderson, prob, r, du, temp, Jv; m=5, verbose=false
)
println(" Time: $(round(t_anderson * 1000, digits=2)) ms")
println(" Iterations: $iters_anderson")
println(" Time/iter: $(round(t_anderson / iters_anderson * 1000, digits=2)) ms")
# Speedups
println("\n✅ CPU Speedups:")
println(" Matrix-Free vs Traditional: $(round(t_trad / t_mf, digits=2))×")
println(" Anderson vs Traditional: $(round(t_trad / t_anderson, digits=2))×")
println(" Anderson vs Matrix-Free: $(round(t_mf / t_anderson, digits=2))×")
println()
end
function benchmark_gpu(n::Int)
if !USE_GPU
println("⚠️ GPU not available, skipping GPU benchmark\n")
return
end
println("="^70)
println("GPU Benchmark: $n DOFs")
println("="^70)
# Setup problem
T = Float64
K_cpu = Matrix(Tridiagonal(
-ones(T, n - 1),
2ones(T, n),
-ones(T, n - 1)
))
f_cpu = ones(T, n) * 0.1
β = T(1e-3)
# Transfer to GPU
K_gpu = CuArray(K_cpu)
f_gpu = CuArray(f_cpu)
u0_gpu = CUDA.zeros(T, n)
# Benchmark Matrix-Free on GPU
println("\n📊 Matrix-Free Newton-Krylov (GPU):")
u_gpu = copy(u0_gpu)
# Warmup
newton_matrix_free_gpu!(u_gpu, K_gpu, f_gpu, β; max_iter=2, verbose=false)
# Benchmark
CUDA.synchronize()
t_gpu = CUDA.@elapsed begin
iters_gpu = newton_matrix_free_gpu!(u_gpu, K_gpu, f_gpu, β; verbose=false)
CUDA.synchronize()
end
println(" Time: $(round(t_gpu * 1000, digits=2)) ms")
println(" Iterations: $iters_gpu")
println(" Time/iter: $(round(t_gpu / iters_gpu * 1000, digits=2)) ms")
# Compare with CPU
prob_cpu = NonlinearProblem(K_cpu, f_cpu, β, n)
u_cpu = zeros(T, n)
r = zeros(T, n)
du = zeros(T, n)
temp = zeros(T, n)
Jv = zeros(T, n)
t_cpu = @elapsed iters_cpu = newton_matrix_free!(
u_cpu, prob_cpu, r, du, temp, Jv; verbose=false
)
println("\n✅ GPU vs CPU Speedup: $(round(t_cpu / t_gpu, digits=2))×")
println(" CPU: $(round(t_cpu * 1000, digits=2)) ms")
println(" GPU: $(round(t_gpu * 1000, digits=2)) ms")
# Benchmark Anderson-Accelerated on GPU
println("\n📊 Anderson-Accelerated Newton-Krylov (GPU):")
u_gpu_anderson = copy(u0_gpu)
# Warmup
anderson_newton_matrix_free_gpu!(u_gpu_anderson, K_gpu, f_gpu, β; max_iter=2, verbose=false)
# Benchmark
CUDA.synchronize()
t_gpu_anderson = CUDA.@elapsed begin
iters_gpu_anderson = anderson_newton_matrix_free_gpu!(u_gpu_anderson, K_gpu, f_gpu, β; verbose=false)
CUDA.synchronize()
end
println(" Time: $(round(t_gpu_anderson * 1000, digits=2)) ms")
println(" Iterations: $iters_gpu_anderson")
println(" Time/iter: $(round(t_gpu_anderson / iters_gpu_anderson * 1000, digits=2)) ms")
# Compare with CPU Anderson
u_cpu_anderson = zeros(T, n)
t_cpu_anderson = @elapsed iters_cpu_anderson = anderson_newton_matrix_free!(
u_cpu_anderson, prob_cpu, r, du, temp, Jv; verbose=false
)
println("\n✅ GPU vs CPU Speedup (Anderson): $(round(t_cpu_anderson / t_gpu_anderson, digits=2))×")
println(" CPU: $(round(t_cpu_anderson * 1000, digits=2)) ms")
println(" GPU: $(round(t_gpu_anderson * 1000, digits=2)) ms")
# Overall comparison
println("\n📊 Summary:")
println(" Matrix-Free GPU speedup: $(round(t_cpu / t_gpu, digits=2))×")
println(" Anderson GPU speedup: $(round(t_cpu_anderson / t_gpu_anderson, digits=2))×")
println()
end
# ============================================================================
# Main
# ============================================================================
function main()
println("\n" * "="^70)
println("Matrix-Free Newton-Krylov GPU Benchmark")
println("="^70)
println()
# Test sizes (reasonable for demonstration)
sizes = [1000, 5_000, 10_000]
for n in sizes
# CPU comparison
benchmark_cpu(n)
# GPU benchmark
if USE_GPU
benchmark_gpu(n)
end
end
println("="^70)
println("Benchmark Complete!")
println("="^70)
println()
println("Key Findings:")
println(" - Matrix-Free eliminates Jacobian assembly cost")
println(" - Anderson acceleration reduces Newton iterations")
println(" - GPU provides additional speedup for large problems")
println(" - Total speedup: 5-10× depending on problem size")
println()
end
if abspath(PROGRAM_FILE) == @__FILE__
main()
end
-571
View File
@@ -1,571 +0,0 @@
#!/usr/bin/env julia
#
# Multi-GPU Nodal Assembly Benchmark with MPI + CUDA
#
# Usage:
# mpirun -np 2 julia --project=. benchmarks/multigpu_mpi_benchmark.jl
# mpirun -np 4 julia --project=. benchmarks/multigpu_mpi_benchmark.jl
#
# Each MPI rank gets one GPU
using MPI
using CUDA
using LinearAlgebra
using Printf
MPI.Init()
const comm = MPI.COMM_WORLD
const rank = MPI.Comm_rank(comm)
const nranks = MPI.Comm_size(comm)
# Set GPU device based on rank
if CUDA.functional()
CUDA.device!(rank % CUDA.ndevices())
if rank == 0
println("="^70)
println("Multi-GPU Nodal Assembly Benchmark (MPI + CUDA)")
println("="^70)
println("MPI ranks: $nranks")
println("CUDA devices: $(CUDA.ndevices())")
println("CUDA functional: $(CUDA.functional())")
println("="^70)
println()
end
else
if rank == 0
println("ERROR: CUDA not functional!")
println("Install CUDA.jl: using Pkg; Pkg.add(\"CUDA\")")
end
MPI.Finalize()
exit(1)
end
# ============================================================================
# Data Structures
# ============================================================================
struct Node
id::Int32
x::Float32
y::Float32
z::Float32
end
struct Element
id::Int32
connectivity::NTuple{8,Int32} # Hex8
end
struct Partition
rank::Int
owned_nodes::UnitRange{Int}
ghost_nodes::Vector{Int}
local_elements::Vector{Int}
node_to_elements::Vector{Vector{Int}}
interface_neighbors::Vector{Int} # Neighbor ranks
interface_send::Dict{Int,Vector{Int}} # rank → local DOF indices to send
interface_recv::Dict{Int,Vector{Int}} # rank → local DOF indices to receive
end
# ============================================================================
# Mesh Generation
# ============================================================================
function create_hex_mesh(nx, ny, nz)
"""Create structured hexahedral mesh"""
n_nodes = nx * ny * nz
n_elements = (nx - 1) * (ny - 1) * (nz - 1)
nodes = Node[]
for k in 1:nz, j in 1:ny, i in 1:nx
node_id = Int32((k - 1) * nx * ny + (j - 1) * nx + i)
push!(nodes, Node(node_id, Float32(i), Float32(j), Float32(k)))
end
elements = Element[]
for k in 1:(nz-1), j in 1:(ny-1), i in 1:(nx-1)
n1 = Int32((k - 1) * nx * ny + (j - 1) * nx + i)
n2 = n1 + 1
n3 = n2 + nx
n4 = n1 + nx
n5 = n1 + nx * ny
n6 = n2 + nx * ny
n7 = n3 + nx * ny
n8 = n4 + nx * ny
elem_id = Int32(length(elements) + 1)
push!(elements, Element(elem_id, (n1, n2, n3, n4, n5, n6, n7, n8)))
end
return nodes, elements
end
function build_node_to_elements(nodes, elements)
node_to_elems = [Int[] for _ in 1:length(nodes)]
for (elem_id, element) in enumerate(elements)
for node_id in element.connectivity
push!(node_to_elems[node_id], elem_id)
end
end
return node_to_elems
end
# ============================================================================
# Partitioning
# ============================================================================
function partition_mesh_for_rank(nodes, elements, my_rank, n_ranks)
"""Create partition for this MPI rank"""
n_nodes = length(nodes)
nodes_per_rank = ceil(Int, n_nodes / n_ranks)
# Owned nodes
start_node = my_rank * nodes_per_rank + 1
end_node = min((my_rank + 1) * nodes_per_rank, n_nodes)
owned_nodes = start_node:end_node
node_to_elems = build_node_to_elements(nodes, elements)
# Find local elements (touching owned nodes)
local_elements = Int[]
ghost_nodes = Set{Int}()
for (elem_id, element) in enumerate(elements)
if any(Int(nid) in owned_nodes for nid in element.connectivity)
push!(local_elements, elem_id)
for nid in element.connectivity
if !(Int(nid) in owned_nodes)
push!(ghost_nodes, Int(nid))
end
end
end
end
# Build local node_to_elements
local_node_to_elems = [
filter(eid -> eid in local_elements, node_to_elems[nid])
for nid in owned_nodes
]
# Find interface nodes with each neighbor
interface_send = Dict{Int,Vector{Int}}()
interface_recv = Dict{Int,Vector{Int}}()
for neighbor_rank in 0:(n_ranks-1)
if neighbor_rank == my_rank
continue
end
neighbor_start = neighbor_rank * nodes_per_rank + 1
neighbor_end = min((neighbor_rank + 1) * nodes_per_rank, n_nodes)
neighbor_owned = neighbor_start:neighbor_end
# Nodes I own that neighbor needs (I send)
send_nodes = Int[]
for elem_id in local_elements
element = elements[elem_id]
has_neighbor = any(Int(nid) in neighbor_owned for nid in element.connectivity)
if has_neighbor
for nid in element.connectivity
if Int(nid) in owned_nodes && !(Int(nid) in send_nodes)
push!(send_nodes, Int(nid))
end
end
end
end
# Nodes neighbor owns that I need (I receive)
recv_nodes = Int[]
for nid in ghost_nodes
if Int(nid) in neighbor_owned
push!(recv_nodes, Int(nid))
end
end
if !isempty(send_nodes) || !isempty(recv_nodes)
# Convert to local DOF indices
send_dofs = Int[]
for nid in send_nodes
local_idx = nid - start_node + 1
for d in 0:2
push!(send_dofs, (local_idx - 1) * 3 + d + 1)
end
end
recv_dofs = Int[]
for nid in recv_nodes
ghost_idx = findfirst(==(nid), sort(collect(ghost_nodes)))
for d in 0:2
# Ghost DOFs come after owned DOFs
push!(recv_dofs, length(owned_nodes) * 3 + (ghost_idx - 1) * 3 + d + 1)
end
end
if !isempty(send_dofs)
interface_send[neighbor_rank] = send_dofs
end
if !isempty(recv_dofs)
interface_recv[neighbor_rank] = recv_dofs
end
end
end
interface_neighbors = sort(collect(keys(interface_send) keys(interface_recv)))
return Partition(
my_rank,
owned_nodes,
sort(collect(ghost_nodes)),
local_elements,
local_node_to_elems,
interface_neighbors,
interface_send,
interface_recv
)
end
# ============================================================================
# GPU Kernel: Nodal Assembly
# ============================================================================
function gpu_matvec_kernel!(
y::CuDeviceArray{Float32,1},
x::CuDeviceArray{Float32,1},
nodes::CuDeviceArray{Node,1},
elements::CuDeviceArray{Element,1},
node_to_elems_offsets::CuDeviceArray{Int32,1},
node_to_elems_data::CuDeviceArray{Int32,1},
n_owned_nodes::Int32,
)
idx = (blockIdx().x - 1) * blockDim().x + threadIdx().x
if idx > n_owned_nodes
return
end
# This thread processes owned node idx
node = nodes[idx]
dof_start = (idx - 1) * 3 + 1
# Initialize nodal contribution
y1 = Float32(0.0)
y2 = Float32(0.0)
y3 = Float32(0.0)
# Get connected elements using CSR-like format (1-based indexing)
if idx + Int32(1) > length(node_to_elems_offsets)
return
end
elem_start = node_to_elems_offsets[idx] + Int32(1)
elem_end = node_to_elems_offsets[idx+Int32(1)]
for i in elem_start:elem_end
if i > length(node_to_elems_data)
return
end
elem_id = node_to_elems_data[i]
if elem_id > length(elements)
return
end
element = elements[elem_id]
# Add contribution from all nodes in this element
for j in 1:8
nid = element.connectivity[j]
x_dof_start = (nid - 1) * 3 + 1
# Mock stiffness contribution
y1 += Float32(0.1) * x[x_dof_start]
y2 += Float32(0.1) * x[x_dof_start+1]
y3 += Float32(0.1) * x[x_dof_start+2]
end
end
# Write to output
y[dof_start] = y1
y[dof_start+1] = y2
y[dof_start+2] = y3
return nothing
end
# ============================================================================
# Multi-GPU Communication
# ============================================================================
function exchange_ghost_values!(
x_local::CuArray{Float32,1},
partition::Partition,
comm::MPI.Comm
)
"""Exchange interface DOF values between MPI ranks"""
# Prepare send/recv buffers on CPU
send_bufs = Dict{Int,Vector{Float32}}()
recv_bufs = Dict{Int,Vector{Float32}}()
# Copy data from GPU to CPU for sending
x_cpu = Array(x_local)
for neighbor in partition.interface_neighbors
if haskey(partition.interface_send, neighbor)
send_dofs = partition.interface_send[neighbor]
send_bufs[neighbor] = x_cpu[send_dofs]
end
if haskey(partition.interface_recv, neighbor)
recv_dofs = partition.interface_recv[neighbor]
recv_bufs[neighbor] = zeros(Float32, length(recv_dofs))
end
end
# MPI communication
requests = MPI.Request[]
# Post receives
for neighbor in partition.interface_neighbors
if haskey(recv_bufs, neighbor)
req = MPI.Irecv!(recv_bufs[neighbor], comm; source=neighbor, tag=neighbor)
push!(requests, req)
end
end
# Post sends
for neighbor in partition.interface_neighbors
if haskey(send_bufs, neighbor)
req = MPI.Isend(send_bufs[neighbor], comm; dest=neighbor, tag=partition.rank)
push!(requests, req)
end
end
# Wait for all communications
MPI.Waitall(requests)
# Copy received data back to GPU
for neighbor in partition.interface_neighbors
if haskey(partition.interface_recv, neighbor)
recv_dofs = partition.interface_recv[neighbor]
x_cpu[recv_dofs] .= recv_bufs[neighbor]
end
end
# Update GPU array
copyto!(x_local, x_cpu)
end
# ============================================================================
# Benchmark
# ============================================================================
function run_multigpu_benchmark(nx, ny, nz, n_warmup=5, n_runs=10)
if rank == 0
println("\n" * "="^70)
println("Multi-GPU Benchmark: $nx × $ny × $nz mesh")
println("="^70)
end
# Create full mesh on all ranks
nodes, elements = create_hex_mesh(nx, ny, nz)
if rank == 0
println(" Total nodes: ", length(nodes))
println(" Total elements: ", length(elements))
println(" Total DOFs: ", 3 * length(nodes))
end
# Partition for this rank
partition = partition_mesh_for_rank(nodes, elements, rank, nranks)
n_owned = length(partition.owned_nodes)
n_ghost = length(partition.ghost_nodes)
n_local_dofs = 3 * (n_owned + n_ghost)
println("Rank $rank: $n_owned owned nodes, $n_ghost ghost nodes, " *
"$(length(partition.local_elements)) elements")
# Prepare GPU data
local_nodes = [nodes[i] for i in vcat(collect(partition.owned_nodes), partition.ghost_nodes)]
# Create mapping from global node ID to local index
global_to_local_node = Dict{Int,Int32}()
for (local_idx, global_nid) in enumerate(vcat(collect(partition.owned_nodes), partition.ghost_nodes))
global_to_local_node[global_nid] = Int32(local_idx)
end
# Remap element connectivity to local node indices
local_elements = Element[]
for global_eid in partition.local_elements
element = elements[global_eid]
# Convert global node IDs to local indices
local_conn = ntuple(8) do i
global_nid = Int(element.connectivity[i])
global_to_local_node[global_nid]
end
push!(local_elements, Element(element.id, local_conn))
end
# Create mapping from global element ID to local index (for CSR data)
global_to_local_elem = Dict{Int,Int}()
for (local_idx, global_id) in enumerate(partition.local_elements)
global_to_local_elem[global_id] = local_idx
end
# Convert node_to_elements to GPU-friendly flat format
# Format: offsets array + flat data array (CSR-like)
# IMPORTANT: Convert global element IDs to local indices
node_to_elems_offsets = Int32[0]
node_to_elems_data = Int32[]
for arr in partition.node_to_elements
# Map global element IDs to local indices
local_indices = [global_to_local_elem[global_id] for global_id in arr]
append!(node_to_elems_data, Int32.(local_indices))
push!(node_to_elems_offsets, length(node_to_elems_data))
end
# Debug: check element ID range
if rank == 0 && length(node_to_elems_data) > 0
min_elem_id = minimum(node_to_elems_data)
max_elem_id = maximum(node_to_elems_data)
println("\nCSR data element ID range: $min_elem_id to $max_elem_id")
println("Local elements array size: $(length(local_elements))")
if max_elem_id > length(local_elements)
println("❌ WARNING: Element ID $max_elem_id > array size $(length(local_elements))")
end
end
# Transfer to GPU
nodes_gpu = CuArray(local_nodes)
elements_gpu = CuArray(local_elements)
node_to_elems_offsets_gpu = CuArray(node_to_elems_offsets)
node_to_elems_data_gpu = CuArray(node_to_elems_data)
# Debug: print array sizes
if rank == 0
println("\nArray sizes on GPU:")
println(" nodes: $(length(nodes_gpu))")
println(" elements: $(length(elements_gpu))")
println(" node_to_elems_offsets: $(length(node_to_elems_offsets_gpu))")
println(" node_to_elems_data: $(length(node_to_elems_data_gpu))")
println(" Expected offsets length: $(n_owned + 1)")
end
# Test vectors
x_local = CUDA.rand(Float32, n_local_dofs)
y_local = CUDA.zeros(Float32, n_local_dofs)
# Kernel launch parameters
threads_per_block = 256
n_blocks = cld(n_owned, threads_per_block)
if rank == 0
println("\nGPU configuration:")
println(" Threads per block: $threads_per_block")
println(" Blocks per rank: $n_blocks")
end
# Warmup
for _ in 1:n_warmup
exchange_ghost_values!(x_local, partition, comm)
CUDA.@sync @cuda threads = threads_per_block blocks = n_blocks gpu_matvec_kernel!(
y_local, x_local, nodes_gpu, elements_gpu,
node_to_elems_offsets_gpu, node_to_elems_data_gpu, Int32(n_owned)
)
end
MPI.Barrier(comm)
# Benchmark
times = Float64[]
comm_times = Float64[]
compute_times = Float64[]
for _ in 1:n_runs
t_start = time_ns()
# Communication
t_comm_start = time_ns()
exchange_ghost_values!(x_local, partition, comm)
MPI.Barrier(comm)
t_comm_end = time_ns()
# Computation
t_compute_start = time_ns()
CUDA.@sync @cuda threads = threads_per_block blocks = n_blocks gpu_matvec_kernel!(
y_local, x_local, nodes_gpu, elements_gpu,
node_to_elems_offsets_gpu, node_to_elems_data_gpu, Int32(n_owned)
)
MPI.Barrier(comm)
t_compute_end = time_ns()
t_end = time_ns()
push!(times, (t_end - t_start) / 1e9)
push!(comm_times, (t_comm_end - t_comm_start) / 1e9)
push!(compute_times, (t_compute_end - t_compute_start) / 1e9)
end
# Gather results
local_time = minimum(times)
local_comm = minimum(comm_times)
local_compute = minimum(compute_times)
all_times = MPI.Gather(local_time, 0, comm)
all_comm = MPI.Gather(local_comm, 0, comm)
all_compute = MPI.Gather(local_compute, 0, comm)
if rank == 0
println("\nResults:")
println(" Rank | Owned Nodes | Total Time | Comm Time | Compute Time | Comm %")
println(" " * "-"^70)
for r in 0:(nranks-1)
nodes_str = lpad(string(length(partition.owned_nodes)), 11)
total_str = @sprintf("%.3f ms", all_times[r+1] * 1000)
comm_str = @sprintf("%.3f ms", all_comm[r+1] * 1000)
compute_str = @sprintf("%.3f ms", all_compute[r+1] * 1000)
comm_pct = @sprintf("%.1f%%", all_comm[r+1] / all_times[r+1] * 100)
println(" $r | $nodes_str | $(lpad(total_str, 10)) | " *
"$(lpad(comm_str, 9)) | $(lpad(compute_str, 12)) | $(lpad(comm_pct, 6))")
end
max_time = maximum(all_times)
avg_compute = sum(all_compute) / length(all_compute)
avg_comm = sum(all_comm) / length(all_comm)
println("\n Maximum time: ", @sprintf("%.3f ms", max_time * 1000))
println(" Average compute: ", @sprintf("%.3f ms", avg_compute * 1000))
println(" Average communication: ", @sprintf("%.3f ms", avg_comm * 1000))
println(" Communication overhead: ", @sprintf("%.1f%%", avg_comm / max_time * 100))
throughput = length(nodes) / max_time / 1e6
println(" Throughput: ", @sprintf("%.2f Mnodes/s", throughput))
end
end
# ============================================================================
# Main
# ============================================================================
if rank == 0
println("Starting benchmarks...")
println()
end
# Run benchmarks with increasing mesh sizes
run_multigpu_benchmark(30, 30, 30)
run_multigpu_benchmark(50, 50, 50)
run_multigpu_benchmark(70, 70, 70)
if rank == 0
println("\n" * "="^70)
println("✓ Multi-GPU Benchmark Complete")
println("="^70)
end
MPI.Finalize()
-329
View File
@@ -1,329 +0,0 @@
# Multi-GPU Nodal Assembly Results
**Date:** November 9, 2025
**Hardware:** 2× MPI ranks, 1× NVIDIA RTX A2000 12GB (shared between ranks)
**Software:** Julia 1.12.1, CUDA.jl, MPI.jl
---
## Executive Summary
**✅ Multi-GPU implementation WORKS!** Successfully ran multi-GPU nodal assembly with MPI + CUDA.
**Key Achievement:** Implemented nodal assembly on GPU with:
- CSR-format node-to-elements connectivity (zero allocation)
- Proper global-to-local index remapping for elements and nodes
- MPI ghost value exchange for domain interfaces
- Validated correctness (all ranks complete successfully)
**Performance:**
- **Throughput:** 115-302 Mnodes/s (scales with mesh size)
- **Communication overhead:** 29-61% (MPI transfers dominate for small/medium meshes)
- **Compute performance:** GPU kernel is fast (0.16-0.56 ms), communication is bottleneck
---
## Detailed Results
### Benchmark Configuration
- **MPI ranks:** 2
- **GPU per rank:** 1 (shared device for both ranks in this test)
- **Partitioning:** Slab decomposition (nodes split evenly)
- **Element type:** Hex8 (8-node hexahedron)
- **Kernel:** Mock stiffness (simplified matvec for validation)
- **Warmup:** 10 iterations
- **Measurement:** 100 iterations (timed)
### Performance Table
| Mesh Size | Total Nodes | Total DOFs | Owned/Rank | Ghost/Rank | Throughput | Comm % |
|-----------|-------------|------------|------------|------------|------------|--------|
| 30³ | 27,000 | 81,000 | 13,500 | 900 | 114.84 Mnodes/s | 29.1% |
| 50³ | 125,000 | 375,000 | 62,500 | 2,500 | 130.64 Mnodes/s | 60.7% |
| 70³ | 343,000 | 1,029,000 | 171,500 | 4,900 | 301.83 Mnodes/s | 50.9% |
### Detailed Timing Breakdown
**30×30×30 mesh:**
```
Rank 0: Total 0.235 ms = Compute 0.159 ms + Comm 0.069 ms (29.4%)
Rank 1: Total 0.227 ms = Compute 0.159 ms + Comm 0.068 ms (29.9%)
Throughput: 114.84 Mnodes/s
```
**50×50×50 mesh:**
```
Rank 0: Total 0.957 ms = Compute 0.373 ms + Comm 0.581 ms (60.7%)
Rank 1: Total 0.957 ms = Compute 0.375 ms + Comm 0.581 ms (60.7%)
Throughput: 130.64 Mnodes/s
```
**70×70×70 mesh:**
```
Rank 0: Total 1.136 ms = Compute 0.557 ms + Comm 0.579 ms (50.9%)
Rank 1: Total 1.136 ms = Compute 0.557 ms + Comm 0.579 ms (50.9%)
Throughput: 301.83 Mnodes/s
```
---
## Comparison with CPU Baseline
**From `nodal_assembly_scalability.jl` (validated Nov 9, 2025):**
### CPU Multi-Threading (8 threads, single node)
| Mesh Size | Nodes | Single-Thread | 8 Threads | Speedup | Efficiency |
|-----------|---------|---------------|-----------|---------|------------|
| 20³ | 8,000 | 3.4 Mnodes/s | 49.6 Mnodes/s | 14.6× | 182% |
| 40³ | 64,000 | 3.6 Mnodes/s | 54.8 Mnodes/s | 15.2× | 189% |
| 60³ | 216,000 | 3.6 Mnodes/s | 47.9 Mnodes/s | 13.3× | 166% |
### CPU Partitioned (4 partitions, sequential)
| Mesh Size | Nodes | Throughput | Speedup vs Single-Thread | Interface Overhead |
|-----------|---------|------------|--------------------------|-------------------|
| 20³ | 8,000 | 29.7 Mnodes/s | 8.5× | 40.1% |
| 40³ | 64,000 | 28.8 Mnodes/s | 8.0× | 22.4% |
| 60³ | 216,000 | 27.3 Mnodes/s | 7.6× | 10.3% |
### GPU vs CPU Comparison
**Throughput Comparison (approximate mesh sizes):**
| Mesh | CPU Single-Thread | CPU 8-Thread | CPU 4-Partition | GPU 2-Rank (MPI) | GPU Speedup vs 8-Thread |
|------|-------------------|--------------|-----------------|------------------|-------------------------|
| ~30³ | 3.5 Mnodes/s | ~50 Mnodes/s | ~29 Mnodes/s | 114.84 Mnodes/s | **2.3×** |
| ~60³ | 3.6 Mnodes/s | 47.9 Mnodes/s | 27.3 Mnodes/s | ~200 Mnodes/s (interpolated) | **4.2×** |
| 70³ | 3.6 Mnodes/s | ~48 Mnodes/s (est) | ~28 Mnodes/s (est) | 301.83 Mnodes/s | **6.3×** |
**Key Observations:**
1. ✅ GPU is **2-6× faster** than CPU 8-thread for same mesh size
2. ✅ GPU throughput scales better with mesh size (114 → 302 Mnodes/s)
3. ⚠️ GPU communication overhead (29-61%) higher than CPU partitioned (10-40%)
4. 🎯 GPU shines on larger meshes (70³: 6.3× faster than CPU)
---
## Analysis & Insights
### What Worked Well ✅
1. **Nodal assembly pattern on GPU:**
- Each thread processes one node (no atomics!)
- Gathers contributions from connected elements
- Direct write to owned DOFs (no race conditions)
2. **CSR-format node-to-elements:**
- `offsets[node_id]` → start of element list
- `data[offsets[i]:offsets[i+1]]` → element IDs
- Zero allocation, type-stable, GPU-friendly
3. **Global-to-local index remapping:**
- Element IDs: Global mesh → Local partition indices
- Node IDs in connectivity: Global mesh → Local partition indices
- Critical for correctness with sliced arrays
4. **MPI ghost exchange:**
- Interface nodes identified correctly
- Ghost values exchanged between ranks
- Enables domain decomposition
### Performance Bottlenecks ⚠️
1. **Communication overhead dominates small/medium meshes:**
- 30³ mesh: 29% communication
- 50³ mesh: **61% communication** (worst case!)
- 70³ mesh: 51% communication
- **Root cause:** MPI transfers CPU ↔ GPU for every iteration
2. **Single GPU shared between 2 MPI ranks:**
- Both ranks compete for same GPU
- No true parallelism in this test configuration
- Need multiple GPUs for real multi-GPU scaling
3. **Mock kernel (simplified stiffness):**
- Real FEM kernel would be more compute-intensive
- Would reduce communication percentage
- Current kernel is memory-bound
### Opportunities for Improvement 🎯
1. **CUDA-aware MPI:**
- Direct GPU-to-GPU transfers (no CPU staging)
- Can reduce communication time by 50-80%
- Requires recompilation of MPI with CUDA support
2. **Multiple physical GPUs:**
- Current test uses 1 GPU for 2 ranks (shared)
- True multi-GPU: Each rank gets own GPU
- Would enable concurrent execution
3. **Larger elements (higher-order):**
- Tet10, Hex20, Hex27 have more work per element
- More compute per node → reduces communication %
- Better compute/communication ratio
4. **Full element stiffness:**
- Real FEM: Integration loops, material models, plasticity
- 10-100× more work per element
- Communication becomes negligible (<5%)
5. **Batched assembly:**
- Assemble multiple timesteps before MPI sync
- Amortize communication cost
- Useful for explicit dynamics
---
## Technical Details
### Data Structures
**Node (immutable, 32 bytes):**
```julia
struct Node
id::Int32
x::Float32
y::Float32
z::Float32
end
```
**Element (immutable, 36 bytes):**
```julia
struct Element
id::Int32
connectivity::NTuple{8, Int32} # Hex8
end
```
**Partition:**
- `owned_nodes`: Nodes owned by this rank
- `ghost_nodes`: Nodes owned by neighbors (interface)
- `local_elements`: Elements touching owned nodes
- `node_to_elements`: Inverse connectivity (node → elements)
- `interface_send/recv`: MPI communication patterns
### GPU Kernel (Simplified)
```julia
function gpu_matvec_kernel!(y, x, nodes, elements, offsets, data, n_owned)
idx = (blockIdx().x - 1) * blockDim().x + threadIdx().x
if idx > n_owned
return # Thread beyond owned nodes
end
# Initialize accumulator
fx = fy = fz = 0.0f0
# Loop over connected elements (CSR access)
elem_start = offsets[idx] + 1
elem_end = offsets[idx + 1]
for i in elem_start:elem_end
elem_id = data[i]
element = elements[elem_id]
# Gather from element nodes
for j in 1:8
nid = element.connectivity[j]
dof_base = (nid - 1) * 3 + 1
fx += 0.1f0 * x[dof_base]
fy += 0.1f0 * x[dof_base + 1]
fz += 0.1f0 * x[dof_base + 2]
end
end
# Write result (owned DOFs only)
dof_base = (idx - 1) * 3 + 1
y[dof_base] = fx
y[dof_base + 1] = fy
y[dof_base + 2] = fz
end
```
### Key Implementation Challenges & Solutions
**Challenge 1:** CuArray{CuArray} not supported
**Solution:** Flatten to CSR format (offsets + data arrays)
**Challenge 2:** Global element IDs in CSR data, but local array
**Solution:** Create `global_to_local_elem` mapping, remap before GPU transfer
**Challenge 3:** Global node IDs in element connectivity
**Solution:** Create `global_to_local_node` mapping, rebuild elements with local indices
**Challenge 4:** BoundsError during kernel execution
**Solution:** All three index spaces must be consistent (nodes, elements, DOFs)
---
## Conclusions
### Claims We Can Now Make ✅
1. ✅ **Nodal assembly works on GPU** - Validated with working implementation
2. ✅ **2-6× faster than CPU multi-threading** - Real measurements on same mesh
3. ✅ **Scales to 343K nodes / 1M DOFs** - Successfully ran 70³ mesh
4. ✅ **Communication overhead acceptable** - 29-61% (will improve with CUDA-aware MPI)
5. ✅ **CSR format enables zero-allocation** - No dynamic memory in kernel
### Claims We CANNOT Yet Make ⚠️
1. ⚠️ **Multi-GPU strong scaling** - Only tested 1 GPU with 2 ranks (not true multi-GPU)
2. ⚠️ **Production-ready performance** - Mock kernel, needs real FEM stiffness
3. ⚠️ **Weak scaling to N GPUs** - Need cluster with multiple GPUs
4. ⚠️ **Better than Gridap/Ferrite** - Haven't compared with other libraries
5. ⚠️ **Contact mechanics on GPU** - Not yet implemented
### Next Steps 🎯
**Immediate (validate architecture):**
1. Test with multiple physical GPUs (2-4 GPUs on cluster)
2. Implement real element stiffness (not mock)
3. Measure CUDA-aware MPI improvement
4. Add higher-order elements (Tet10, Hex20)
**Short-term (production features):**
1. Material state updates on GPU (plasticity, damage)
2. Contact detection and assembly on GPU
3. Preconditioned GMRES on GPU (full solver)
4. Integration with JuliaFEM element library
**Long-term (scale-up):**
1. Weak scaling study (1-64 GPUs)
2. Strong scaling study (fixed problem, varying GPUs)
3. Comparison with Gridap.jl + PETSc
4. Real-world contact mechanics problem (1M+ DOFs)
---
## Files & Artifacts
**Benchmark code:**
- `benchmarks/multigpu_mpi_benchmark.jl` (555 lines, working)
**CPU baseline (validated):**
- `benchmarks/nodal_assembly_scalability.jl` (500 lines)
**Documentation:**
- `docs/book/multigpu_nodal_assembly.md` (design, needs update with real data)
- `docs/book/nodal_assembly_gpu_pattern.md` (architecture)
- `demos/gpu_nodal_assembly_demo.jl` (educational demo)
**This report:**
- `benchmarks/multigpu_results_2025-11-09.md`
---
## Acknowledgments
**User (Jukka):** Demanded real measurements, not designs. Caught AI making unvalidated claims. Insisted on "Just run" - forcing validation before documentation.
**Key Insight:** "Did you actually run that code?" - Best engineering feedback possible. No more design documents without validation!
---
**Status:** ✅ Multi-GPU architecture VALIDATED
**Verdict:** Nodal assembly on GPU is **feasible and fast**. Communication overhead acceptable. Ready for production implementation.
-260
View File
@@ -1,260 +0,0 @@
"""
Performance Analysis: NeoHookean Hyperelastic Material
Benchmarks automatic differentiation overhead in hyperelastic stress computation.
Key Questions:
1. What is the cost of AD compared to LinearElastic?
2. Is the implementation allocation-free?
3. How does performance scale with problem size?
4. What is the breakdown of strain energy vs stress vs tangent?
Results inform whether AD is suitable for production FEM assembly loops.
"""
using BenchmarkTools
using Tensors
using LinearAlgebra
using Printf
# Load implementations
include("../src/materials/abstract_material.jl")
include("../src/materials/linear_elastic.jl")
include("../src/materials/neo_hookean.jl")
println("="^80)
println("NeoHookean Hyperelastic Material - Performance Analysis")
println("="^80)
println()
# =============================================================================
# Test 1: Single Stress Evaluation (typical FEM use case)
# =============================================================================
println("Test 1: Single Stress Evaluation")
println("-"^80)
# Create materials
neo = NeoHookean(E_mod=200e9, nu=0.3) # Steel-like properties
linear = LinearElastic(E=200e9, ν=0.3)
# Test strain (moderate deformation)
E_strain = SymmetricTensor{2,3}((0.01, 0.005, 0.003, -0.002, 0.004, 0.006))
# Compile first
compute_stress(neo, E_strain, nothing, 0.0)
compute_stress(linear, E_strain, nothing, 0.0)
# Benchmark
println("\nLinearElastic (manual derivatives):")
t_linear = @benchmark compute_stress($linear, $E_strain, nothing, 0.0)
display(t_linear)
println()
println("\nNeoHookean (automatic differentiation):")
t_neo = @benchmark compute_stress($neo, $E_strain, nothing, 0.0)
display(t_neo)
println()
# Compute overhead
overhead = median(t_neo).time / median(t_linear).time
println("\nAD Overhead: $(round(overhead, digits=1))x")
println("Absolute difference: $(round((median(t_neo).time - median(t_linear).time)/1e3, digits=1)) μs")
# =============================================================================
# Test 2: Allocation Check
# =============================================================================
println("\n" * "="^80)
println("Test 2: Allocation Analysis")
println("-"^80)
allocs_neo = @allocated compute_stress(neo, E_strain, nothing, 0.0)
allocs_linear = @allocated compute_stress(linear, E_strain, nothing, 0.0)
println("LinearElastic allocations: $allocs_linear bytes")
println("NeoHookean allocations: $allocs_neo bytes")
if allocs_neo == 0
println("✅ Zero-allocation achieved!")
else
println("⚠️ Allocations detected - investigate")
end
# =============================================================================
# Test 3: Component Breakdown (where does time go?)
# =============================================================================
println("\n" * "="^80)
println("Test 3: Component Breakdown")
println("-"^80)
# Test strain energy only
C = 2E_strain + one(E_strain)
println("\nStrain energy computation:")
t_energy = @benchmark strain_energy($neo, $C)
display(t_energy)
println()
# Stress only (includes strain energy + gradient)
println("\nComplete stress computation (energy + gradient + hessian):")
display(t_neo)
println()
energy_fraction = median(t_energy).time / median(t_neo).time
println("Strain energy is $(round(energy_fraction*100, digits=1))% of total time")
println("AD overhead (gradient + hessian) is $(round((1-energy_fraction)*100, digits=1))% of total time")
# =============================================================================
# Test 4: Deformation Magnitude Sensitivity
# =============================================================================
println("\n" * "="^80)
println("Test 4: Performance vs Deformation Magnitude")
println("-"^80)
strain_levels = [1e-6, 1e-4, 1e-2, 0.1, 0.5]
times = Float64[]
for ε_mag in strain_levels
E_test = SymmetricTensor{2,3}((ε_mag, 0.0, 0.0, 0.0, 0.0, 0.0))
compute_stress(neo, E_test, nothing, 0.0) # Compile
t = @benchmark compute_stress($neo, $E_test, nothing, 0.0) samples = 1000
push!(times, median(t).time)
@printf("Strain magnitude: %.1e → Time: %.1f ns\n", ε_mag, median(t).time)
end
time_variation = (maximum(times) - minimum(times)) / minimum(times) * 100
println("\nTime variation across strain levels: $(round(time_variation, digits=1))%")
if time_variation < 10
println("✅ Performance is strain-independent (good for Newton solvers)")
else
println("⚠️ Performance varies with strain (may impact Newton convergence)")
end
# =============================================================================
# Test 5: Tangent Accuracy vs Finite Difference
# =============================================================================
println("\n" * "="^80)
println("Test 5: Tangent Accuracy (AD vs Finite Difference)")
println("-"^80)
E_base = SymmetricTensor{2,3}((0.01, 0.005, 0.003, -0.002, 0.004, 0.006))
S_ad, 𝔻_ad, _ = compute_stress(neo, E_base, nothing, 0.0)
# Finite difference tangent
ε_fd = 1e-8
errors = Float64[]
for i in 1:6
E_pert_data = collect(E_base.data)
E_pert_data[i] += ε_fd
E_pert = SymmetricTensor{2,3}(tuple(E_pert_data...))
S_pert, _, _ = compute_stress(neo, E_pert, nothing, 0.0)
∂S∂E_fd = (S_pert - S_ad) / ε_fd
error = norm(∂S∂E_fd) # Simplified error metric
push!(errors, error)
end
println("Tangent norm comparison:")
println(" AD tangent: $(round(norm(𝔻_ad), sigdigits=6))")
println(" FD derivative: $(round(mean(errors), sigdigits=6))")
println(" Relative diff: $(round((norm(𝔻_ad) - mean(errors))/norm(𝔻_ad)*100, digits=2))%")
println("\n✅ AD provides exact derivatives (limited only by machine precision)")
# =============================================================================
# Test 6: Memory Footprint
# =============================================================================
println("\n" * "="^80)
println("Test 6: Memory Footprint")
println("-"^80)
println("Struct sizes:")
println(" NeoHookean: $(sizeof(neo)) bytes")
println(" LinearElastic: $(sizeof(linear)) bytes")
println("\nReturn value sizes:")
println(" SymmetricTensor{2,3}: $(sizeof(S_ad)) bytes")
println(" SymmetricTensor{4,3}: $(sizeof(𝔻_ad)) bytes")
println(" Total per evaluation: $(sizeof(S_ad) + sizeof(𝔻_ad)) bytes")
# =============================================================================
# Test 7: Scaling with Multiple Evaluations
# =============================================================================
println("\n" * "="^80)
println("Test 7: Assembly Loop Simulation (1000 evaluations)")
println("-"^80)
n_evals = 1000
strain_samples = [SymmetricTensor{2,3}((rand(), rand(), rand(), rand(), rand(), rand())) * 0.01
for _ in 1:n_evals]
# Compile
for E in strain_samples[1:10]
compute_stress(neo, E, nothing, 0.0)
compute_stress(linear, E, nothing, 0.0)
end
println("\nLinearElastic ($n_evals evaluations):")
t_linear_loop = @benchmark begin
for E in $strain_samples
compute_stress($linear, E, nothing, 0.0)
end
end
display(t_linear_loop)
println()
println("\nNeoHookean ($n_evals evaluations):")
t_neo_loop = @benchmark begin
for E in $strain_samples
compute_stress($neo, E, nothing, 0.0)
end
end
display(t_neo_loop)
println()
overhead_loop = median(t_neo_loop).time / median(t_linear_loop).time
per_eval_neo = median(t_neo_loop).time / n_evals
per_eval_linear = median(t_linear_loop).time / n_evals
println("Per-evaluation time:")
println(" LinearElastic: $(round(per_eval_linear, digits=1)) ns")
println(" NeoHookean: $(round(per_eval_neo, digits=1)) ns")
println(" Overhead: $(round(overhead_loop, digits=1))x")
# =============================================================================
# Summary and Recommendations
# =============================================================================
println("\n" * "="^80)
println("SUMMARY AND RECOMMENDATIONS")
println("="^80)
total_overhead = median(t_neo).time / median(t_linear).time
println("\n📊 Performance Metrics:")
println(" Single evaluation: $(round(median(t_neo).time, digits=1)) ns")
println(" AD overhead: $(round(total_overhead, digits=1))x")
println(" Allocations: $(allocs_neo) bytes")
println(" Strain-independent: $(time_variation < 10 ? "Yes ✅" : "No ⚠️")")
println("\n🎯 Recommendations:")
if total_overhead < 5
println(" ✅ EXCELLENT: AD overhead < 5x, suitable for production FEM")
elseif total_overhead < 10
println(" ✅ GOOD: AD overhead < 10x, acceptable for most applications")
elseif total_overhead < 20
println(" ⚠️ MODERATE: AD overhead < 20x, consider for prototyping only")
else
println(" ❌ HIGH: AD overhead > 20x, manual derivatives recommended for production")
end
if allocs_neo == 0
println(" ✅ Zero allocations achieved - suitable for tight loops")
else
println(" ⚠️ Allocations detected - profile and optimize")
end
println("\n💡 Use Cases:")
println(" • Research code: Strongly recommended (correctness > speed)")
println(" • Prototyping: Excellent (rapid implementation)")
println(" • Production: $(total_overhead < 10 ? "Acceptable" : "Profile first") ($(round(total_overhead, digits=1))x overhead)")
println(" • Contact: Excellent (unsymmetric tangent, complex derivatives)")
println("\n" * "="^80)
-438
View File
@@ -1,438 +0,0 @@
#!/usr/bin/env julia
#
# Nodal Assembly Scalability Benchmark
#
# Tests:
# 1. Single-threaded baseline
# 2. Multi-threaded scaling (2, 4, 8 threads)
# 3. Actual speedup measurements
# 4. Interface communication overhead
#
# Run with: julia --project=. -t 8 benchmarks/nodal_assembly_scalability.jl
using LinearAlgebra
using Printf
using Base.Threads
println("="^70)
println("Nodal Assembly Scalability Benchmark")
println("="^70)
println()
println("Julia threads available: ", nthreads())
println()
# ============================================================================
# Data Structures
# ============================================================================
struct Node
id::Int
x::Float64
y::Float64
z::Float64
end
struct Element
id::Int
connectivity::NTuple{8,Int} # Hex8
end
struct Partition
rank::Int
owned_nodes::UnitRange{Int}
ghost_nodes::Vector{Int}
local_elements::Vector{Int}
node_to_elements::Vector{Vector{Int}}
interface_nodes::Dict{Int,Vector{Int}}
end
# ============================================================================
# Mesh Generation
# ============================================================================
function create_hex_mesh(nx, ny, nz)
"""Create structured hexahedral mesh"""
n_nodes = nx * ny * nz
n_elements = (nx - 1) * (ny - 1) * (nz - 1)
# Create nodes
nodes = Node[]
for k in 1:nz, j in 1:ny, i in 1:nx
node_id = (k - 1) * nx * ny + (j - 1) * nx + i
push!(nodes, Node(node_id, Float64(i), Float64(j), Float64(k)))
end
# Create elements (Hex8)
elements = Element[]
for k in 1:(nz-1), j in 1:(ny-1), i in 1:(nx-1)
n1 = (k - 1) * nx * ny + (j - 1) * nx + i
n2 = n1 + 1
n3 = n2 + nx
n4 = n1 + nx
n5 = n1 + nx * ny
n6 = n2 + nx * ny
n7 = n3 + nx * ny
n8 = n4 + nx * ny
elem_id = length(elements) + 1
push!(elements, Element(elem_id, (n1, n2, n3, n4, n5, n6, n7, n8)))
end
return nodes, elements
end
function build_node_to_elements(nodes, elements)
"""Build inverse connectivity"""
node_to_elems = [Int[] for _ in 1:length(nodes)]
for (elem_id, element) in enumerate(elements)
for node_id in element.connectivity
push!(node_to_elems[node_id], elem_id)
end
end
return node_to_elems
end
# ============================================================================
# Partitioning
# ============================================================================
function partition_mesh(nodes, elements, n_partitions)
"""Partition mesh by nodes"""
n_nodes = length(nodes)
nodes_per_partition = ceil(Int, n_nodes / n_partitions)
node_to_elems = build_node_to_elements(nodes, elements)
partitions = Partition[]
for rank in 0:(n_partitions-1)
# Owned nodes
start_node = rank * nodes_per_partition + 1
end_node = min((rank + 1) * nodes_per_partition, n_nodes)
owned_nodes = start_node:end_node
# Find elements touching owned nodes
local_elements = Int[]
ghost_nodes = Set{Int}()
for (elem_id, element) in enumerate(elements)
if any(nid in owned_nodes for nid in element.connectivity)
push!(local_elements, elem_id)
# Mark ghost nodes
for nid in element.connectivity
if !(nid in owned_nodes)
push!(ghost_nodes, nid)
end
end
end
end
# Build local node_to_elements (owned nodes only)
local_node_to_elems = [node_to_elems[nid] for nid in owned_nodes]
# Find interface nodes (owned nodes that couple to other partitions)
interface = Dict{Int,Vector{Int}}()
for neighbor_rank in 0:(n_partitions-1)
if neighbor_rank == rank
continue
end
neighbor_start = neighbor_rank * nodes_per_partition + 1
neighbor_end = min((neighbor_rank + 1) * nodes_per_partition, n_nodes)
neighbor_owned = neighbor_start:neighbor_end
interface_with_neighbor = Int[]
for elem_id in local_elements
element = elements[elem_id]
has_owned = any(nid in owned_nodes for nid in element.connectivity)
has_neighbor = any(nid in neighbor_owned for nid in element.connectivity)
if has_owned && has_neighbor
for nid in element.connectivity
if nid in owned_nodes && !(nid in interface_with_neighbor)
push!(interface_with_neighbor, nid)
end
end
end
end
if !isempty(interface_with_neighbor)
interface[neighbor_rank] = interface_with_neighbor
end
end
partition = Partition(
rank,
owned_nodes,
collect(ghost_nodes),
local_elements,
local_node_to_elems,
interface
)
push!(partitions, partition)
end
return partitions
end
# ============================================================================
# Nodal Assembly (Matrix-Vector Product)
# ============================================================================
function matvec_single_threaded!(y, x, nodes, elements, node_to_elements)
"""Single-threaded nodal assembly"""
fill!(y, 0.0)
for node_id in 1:length(nodes)
node = nodes[node_id]
# Get node DOFs
dof_start = (node_id - 1) * 3 + 1
y_nodal = zeros(3)
# Gather from connected elements
for elem_id in node_to_elements[node_id]
element = elements[elem_id]
# Mock element contribution (just for timing)
for nid in element.connectivity
x_dof_start = (nid - 1) * 3 + 1
for d in 1:3
y_nodal[d] += 0.1 * x[x_dof_start+d-1]
end
end
end
# Write to global
for d in 1:3
y[dof_start+d-1] = y_nodal[d]
end
end
end
function matvec_multi_threaded!(y, x, nodes, elements, node_to_elements)
"""Multi-threaded nodal assembly (direct, no partitioning)"""
fill!(y, 0.0)
@threads for node_id in 1:length(nodes)
node = nodes[node_id]
dof_start = (node_id - 1) * 3 + 1
y_nodal = zeros(3)
for elem_id in node_to_elements[node_id]
element = elements[elem_id]
for nid in element.connectivity
x_dof_start = (nid - 1) * 3 + 1
for d in 1:3
y_nodal[d] += 0.1 * x[x_dof_start+d-1]
end
end
end
for d in 1:3
y[dof_start+d-1] = y_nodal[d]
end
end
end
function matvec_partitioned!(y, x, nodes, elements, partitions)
"""Multi-threaded with explicit partitioning (simulates multi-GPU)"""
fill!(y, 0.0)
# Each partition processed by one thread
@threads for partition in partitions
# Process owned nodes
for (local_idx, node_id) in enumerate(partition.owned_nodes)
node = nodes[node_id]
dof_start = (node_id - 1) * 3 + 1
y_nodal = zeros(3)
for elem_id in partition.node_to_elements[local_idx]
element = elements[elem_id]
for nid in element.connectivity
x_dof_start = (nid - 1) * 3 + 1
for d in 1:3
y_nodal[d] += 0.1 * x[x_dof_start+d-1]
end
end
end
for d in 1:3
y[dof_start+d-1] = y_nodal[d]
end
end
end
end
# ============================================================================
# Benchmarks
# ============================================================================
function run_benchmark(name, nx, ny, nz, n_warmup=2, n_runs=10)
println("\n" * "="^70)
println("Benchmark: $name")
println(" Mesh: $nx × $ny × $nz = $(nx*ny*nz) nodes, $((nx-1)*(ny-1)*(nz-1)) elements")
println("="^70)
# Create mesh
print("Creating mesh... ")
nodes, elements = create_hex_mesh(nx, ny, nz)
node_to_elements = build_node_to_elements(nodes, elements)
n_dofs = 3 * length(nodes)
println("")
println(" Nodes: ", length(nodes))
println(" Elements: ", length(elements))
println(" DOFs: ", n_dofs)
println(" Avg elements/node: ", sum(length.(node_to_elements)) / length(nodes))
# Test vectors
x = randn(n_dofs)
y_ref = zeros(n_dofs)
y_test = zeros(n_dofs)
# ========================================================================
# 1. Single-threaded baseline
# ========================================================================
println("\n1. Single-threaded baseline:")
# Warmup
for _ in 1:n_warmup
matvec_single_threaded!(y_ref, x, nodes, elements, node_to_elements)
end
# Benchmark
times = Float64[]
for _ in 1:n_runs
t_start = time_ns()
matvec_single_threaded!(y_ref, x, nodes, elements, node_to_elements)
t_end = time_ns()
push!(times, (t_end - t_start) / 1e9)
end
t_single = minimum(times)
println(" Time: ", @sprintf("%.3f ms", t_single * 1000))
println(" Throughput: ", @sprintf("%.2f Mnodes/s", length(nodes) / t_single / 1e6))
# ========================================================================
# 2. Multi-threaded (all available threads)
# ========================================================================
if nthreads() > 1
println("\n2. Multi-threaded ($(nthreads()) threads):")
# Warmup
for _ in 1:n_warmup
matvec_multi_threaded!(y_test, x, nodes, elements, node_to_elements)
end
# Verify correctness
error = norm(y_test - y_ref) / norm(y_ref)
println(" Verification: ", error < 1e-10 ? "✓ PASS" : "✗ FAIL (error=$error)")
# Benchmark
times = Float64[]
for _ in 1:n_runs
t_start = time_ns()
matvec_multi_threaded!(y_test, x, nodes, elements, node_to_elements)
t_end = time_ns()
push!(times, (t_end - t_start) / 1e9)
end
t_multi = minimum(times)
speedup = t_single / t_multi
efficiency = speedup / nthreads() * 100
println(" Time: ", @sprintf("%.3f ms", t_multi * 1000))
println(" Speedup: ", @sprintf("%.2fx", speedup))
println(" Efficiency: ", @sprintf("%.1f%%", efficiency))
println(" Throughput: ", @sprintf("%.2f Mnodes/s", length(nodes) / t_multi / 1e6))
end
# ========================================================================
# 3. Partitioned (simulates multi-GPU)
# ========================================================================
if nthreads() >= 4
n_partitions = 4
println("\n3. Partitioned ($n_partitions partitions, 4 threads):")
# Create partitions
print(" Creating partitions... ")
partitions = partition_mesh(nodes, elements, n_partitions)
println("")
# Print partition info
for partition in partitions
n_owned = length(partition.owned_nodes)
n_ghost = length(partition.ghost_nodes)
n_interface = sum(length.(values(partition.interface_nodes)))
interface_pct = n_interface / n_owned * 100
println(" Partition $(partition.rank): $n_owned owned, $n_ghost ghost, " *
"$n_interface interface ($(round(interface_pct, digits=1))%)")
end
# Warmup
for _ in 1:n_warmup
matvec_partitioned!(y_test, x, nodes, elements, partitions)
end
# Verify correctness
error = norm(y_test - y_ref) / norm(y_ref)
println(" Verification: ", error < 1e-10 ? "✓ PASS" : "✗ FAIL (error=$error)")
# Benchmark
times = Float64[]
for _ in 1:n_runs
t_start = time_ns()
matvec_partitioned!(y_test, x, nodes, elements, partitions)
t_end = time_ns()
push!(times, (t_end - t_start) / 1e9)
end
t_partitioned = minimum(times)
speedup = t_single / t_partitioned
efficiency = speedup / n_partitions * 100
println(" Time: ", @sprintf("%.3f ms", t_partitioned * 1000))
println(" Speedup: ", @sprintf("%.2fx", speedup))
println(" Efficiency: ", @sprintf("%.1f%%", efficiency))
println(" Throughput: ", @sprintf("%.2f Mnodes/s", length(nodes) / t_partitioned / 1e6))
end
println()
end
# ============================================================================
# Run Benchmarks
# ============================================================================
# Small mesh
run_benchmark("Small Mesh", 20, 20, 20)
# Medium mesh
run_benchmark("Medium Mesh", 40, 40, 40)
# Large mesh (if enough threads)
if nthreads() >= 4
run_benchmark("Large Mesh", 60, 60, 60)
end
println("="^70)
println("✓ Benchmark Complete")
println("="^70)
println()
println("Notes:")
println(" - Single-threaded: Baseline performance")
println(" - Multi-threaded: All available threads, direct parallelization")
println(" - Partitioned: Simulates multi-GPU with explicit partitions")
println(" - Efficiency = Speedup / N_threads × 100%")
println(" - Near 100% efficiency = perfect scaling")
println()
-320
View File
@@ -1,320 +0,0 @@
"""
Performance Analysis: PerfectPlasticity Material Model
Comprehensive benchmarking of J2 plasticity implementation with radial return mapping.
Tests:
1. Single evaluation performance (elastic vs plastic)
2. Zero-allocation verification
3. Type stability verification
4. State overhead measurement
5. Hardening parameter sensitivity
6. Assembly loop simulation
7. Comparison to LinearElastic and NeoHookean
8. Strain level scalability
Run with:
julia --project=. benchmarks/perfect_plasticity_analysis.jl
"""
using BenchmarkTools
using Tensors
using Statistics
using Printf
using Dates
# Load implementations
include("../src/materials/abstract_material.jl")
include("../src/materials/linear_elastic.jl")
include("../src/materials/neo_hookean.jl")
include("../src/materials/perfect_plasticity.jl")
println("="^80)
println("PERFECT PLASTICITY MATERIAL - PERFORMANCE ANALYSIS")
println("="^80)
println()
# ==============================================================================
# TEST 1: Single Evaluation - Elastic Path
# ==============================================================================
println("TEST 1: Single Evaluation - Elastic Path")
println("-"^80)
steel = PerfectPlasticity(E=200e9, ν=0.3, σ_y=250e6, H=1e9)
ε_elastic = SymmetricTensor{2,3}((1e-5, 0.0, 0.0, 0.0, 0.0, 0.0)) # Below yield
state = PlasticityState()
# Benchmark elastic path
bench_elastic = @benchmark compute_stress($steel, $ε_elastic, $state, 0.0)
t_elastic = median(bench_elastic.times)
allocs_elastic = bench_elastic.allocs
println("Elastic path (no yielding):")
println(" Time: ", @sprintf("%.2f ns", t_elastic))
println(" Allocations: ", allocs_elastic)
println(" Memory: ", bench_elastic.memory, " bytes")
println()
# ==============================================================================
# TEST 2: Single Evaluation - Plastic Path
# ==============================================================================
println("TEST 2: Single Evaluation - Plastic Path")
println("-"^80)
ε_plastic = SymmetricTensor{2,3}((0.003, 0.0, 0.0, 0.0, 0.0, 0.0)) # Beyond yield
# Benchmark plastic path
bench_plastic = @benchmark compute_stress($steel, $ε_plastic, $state, 0.0)
t_plastic = median(bench_plastic.times)
allocs_plastic = bench_plastic.allocs
println("Plastic path (radial return):")
println(" Time: ", @sprintf("%.2f ns", t_plastic))
println(" Allocations: ", allocs_plastic)
println(" Memory: ", bench_plastic.memory, " bytes")
println()
println("Plastic overhead:")
println(" Ratio: ", @sprintf("%.2fx", t_plastic / t_elastic))
println()
# ==============================================================================
# TEST 3: State Management Overhead
# ==============================================================================
println("TEST 3: State Management Overhead")
println("-"^80)
# Compare with and without state history
σ1, 𝔻1, state1 = compute_stress(steel, ε_plastic, nothing, 0.0) # Fresh state
σ2, 𝔻2, state2 = compute_stress(steel, ε_plastic, state1, 0.0) # With history
bench_fresh = @benchmark compute_stress($steel, $ε_plastic, nothing, 0.0)
bench_history = @benchmark compute_stress($steel, $ε_plastic, $state1, 0.0)
println("Fresh state (ε_p = 0, α = 0):")
println(" Time: ", @sprintf("%.2f ns", median(bench_fresh.times)))
println()
println("With history (ε_p ≠ 0, α ≠ 0):")
println(" Time: ", @sprintf("%.2f ns", median(bench_history.times)))
println()
println("State overhead: ", @sprintf("%.1f%%",
(median(bench_history.times) - median(bench_fresh.times)) / median(bench_fresh.times) * 100))
println()
# ==============================================================================
# TEST 4: Hardening Parameter Sensitivity
# ==============================================================================
println("TEST 4: Hardening Parameter Sensitivity")
println("-"^80)
hardening_values = [0.0, 1e8, 1e9, 10e9, 100e9] # Perfect to strong hardening
times_H = Float64[]
for H in hardening_values
mat = PerfectPlasticity(E=200e9, ν=0.3, σ_y=250e6, H=H)
bench = @benchmark compute_stress($mat, $ε_plastic, $state, 0.0)
push!(times_H, median(bench.times))
end
println("H (Pa) Time (ns) Overhead")
println(repeat("-", 45))
for (H, t) in zip(hardening_values, times_H)
overhead = (t - times_H[1]) / times_H[1] * 100
println(@sprintf("%-15.1e %8.2f %+6.1f%%", H, t, overhead))
end
println()
# ==============================================================================
# TEST 5: Comparison to Other Materials
# ==============================================================================
println("TEST 5: Comparison to Other Materials")
println("-"^80)
# LinearElastic
linear = LinearElastic(E=200e9, ν=0.3)
bench_linear = @benchmark compute_stress($linear, $ε_plastic)
t_linear = median(bench_linear.times)
# NeoHookean (uses Green-Lagrange strain for small deformation)
μ_neo = 200e9 / (2 * (1 + 0.3))
λ_neo = 200e9 * 0.3 / ((1 + 0.3) * (1 - 2 * 0.3))
neo = NeoHookean(μ=μ_neo, λ=λ_neo)
E_gl = ε_plastic # For small strains, E_GL ≈ ε
bench_neo = @benchmark compute_stress($neo, $E_gl)
t_neo = median(bench_neo.times)
println("Material Time (ns) Ratio vs Linear")
println(repeat("-", 50))
println(@sprintf("LinearElastic %8.2f 1.00x (baseline)", t_linear))
println(@sprintf("PerfectPlasticity %8.2f %.2fx", t_plastic, t_plastic / t_linear))
println(@sprintf("NeoHookean %8.2f %.2fx", t_neo, t_neo / t_linear))
println()
println("Performance ranking:")
println(" 1. LinearElastic (fastest, no state, manual derivatives)")
println(" 2. PerfectPlasticity (", @sprintf("%.1fx", t_plastic / t_linear),
" - state management + radial return)")
println(" 3. NeoHookean (", @sprintf("%.1fx", t_neo / t_linear),
" - automatic differentiation overhead)")
println()
# ==============================================================================
# TEST 6: Assembly Loop Simulation
# ==============================================================================
println("TEST 6: Assembly Loop Simulation (1000 Gauss points)")
println("-"^80)
n_gauss = 1000
strains = [SymmetricTensor{2,3}((0.001 + 0.002 * rand(), 0.0, 0.0, 0.0, 0.0, 0.0))
for _ in 1:n_gauss]
# Elastic assembly
function assembly_elastic(material, strains)
total = zero(SymmetricTensor{2,3})
for ε in strains
σ, _, _ = compute_stress(material, ε)
total += σ
end
return total
end
# Plastic assembly (stateful)
function assembly_plastic(material, strains, state)
total = zero(SymmetricTensor{2,3})
for ε in strains
σ, _, state = compute_stress(material, ε, state, 0.0)
total += σ
end
return total, state
end
bench_asm_linear = @benchmark assembly_elastic($linear, $strains)
bench_asm_plastic = @benchmark assembly_plastic($steel, $strains, $state)
t_asm_linear = median(bench_asm_linear.times) / 1e6 # Convert to ms
t_asm_plastic = median(bench_asm_plastic.times) / 1e6
println("LinearElastic assembly: ", @sprintf("%.3f ms", t_asm_linear))
println("PerfectPlasticity assembly: ", @sprintf("%.3f ms", t_asm_plastic))
println("Overhead: ", @sprintf("%.2fx", t_asm_plastic / t_asm_linear))
println()
# ==============================================================================
# TEST 7: Strain Level Scalability
# ==============================================================================
println("TEST 7: Strain Level Scalability")
println("-"^80)
strain_magnitudes = [0.0005, 0.001, 0.002, 0.005, 0.01, 0.02]
times_strain = Float64[]
yields = Bool[]
for ε_mag in strain_magnitudes
ε_test = SymmetricTensor{2,3}((ε_mag, 0.0, 0.0, 0.0, 0.0, 0.0))
σ_test, _, state_test = compute_stress(steel, ε_test, nothing, 0.0)
bench = @benchmark compute_stress($steel, $ε_test, nothing, 0.0)
push!(times_strain, median(bench.times))
push!(yields, state_test.κ > 0.0)
end
println("ε_magnitude Time (ns) Yielded?")
println(repeat("-", 40))
for (ε_mag, t, y) in zip(strain_magnitudes, times_strain, yields)
status = y ? "YES" : "no"
println(@sprintf("%.4f %8.2f %s", ε_mag, t, status))
end
println()
# ==============================================================================
# TEST 8: Cyclic Loading Performance
# ==============================================================================
println("TEST 8: Cyclic Loading (Bauschinger Effect)")
println("-"^80)
# Simulate cyclic loading path
ε_cycle = [
SymmetricTensor{2,3}((0.003, 0.0, 0.0, 0.0, 0.0, 0.0)), # Tension
SymmetricTensor{2,3}((0.0, 0.0, 0.0, 0.0, 0.0, 0.0)), # Unload
SymmetricTensor{2,3}((-0.002, 0.0, 0.0, 0.0, 0.0, 0.0)), # Compression
SymmetricTensor{2,3}((0.0, 0.0, 0.0, 0.0, 0.0, 0.0)), # Unload
]
function cyclic_loading(material, strains, state)
for ε in strains
σ, 𝔻, state = compute_stress(material, ε, state, 0.0)
end
return state
end
bench_cyclic = @benchmark cyclic_loading($steel, $ε_cycle, $state)
t_cyclic = median(bench_cyclic.times)
println("Cyclic loading (4 load steps):")
println(" Total time: ", @sprintf("%.2f ns", t_cyclic))
println(" Per load step: ", @sprintf("%.2f ns", t_cyclic / 4))
println()
# ==============================================================================
# TEST 9: Type Stability Verification
# ==============================================================================
println("TEST 9: Type Stability")
println("-"^80)
using InteractiveUtils
println("Return type inference:")
result_type = @code_typed compute_stress(steel, ε_plastic, state, 0.0)
println(" ✓ Type stable: ", result_type[2])
println()
# ==============================================================================
# SUMMARY
# ==============================================================================
println("="^80)
println("SUMMARY")
println("="^80)
println()
println("Performance Characteristics:")
println(" • Elastic path: ", @sprintf("%.0f ns", t_elastic), " (no allocations)")
println(" • Plastic path: ", @sprintf("%.0f ns", t_plastic), " (~128 bytes for state)")
println(" • Plastic overhead:", @sprintf("%.2fx", t_plastic / t_elastic))
println()
println("Comparison to other materials:")
println("", @sprintf("%.2fx", t_plastic / t_linear), " slower than LinearElastic (baseline)")
println("", @sprintf("%.2fx", t_neo / t_plastic), " faster than NeoHookean (AD)")
println()
println("Key findings:")
println(" ✓ Zero allocations on elastic path")
println(" ✓ Minimal allocations on plastic path (state struct only)")
println(" ✓ Type stable")
println(" ✓ Hardening parameter has negligible performance impact")
println(" ✓ Performance independent of strain level")
println(" ✓ Suitable for production FEM with ~",
@sprintf("%.0f", 1e9 / t_plastic), " evaluations/second")
println()
println("Recommendations:")
if t_plastic < 500
println(" ✓ Excellent performance - suitable for all applications")
elseif t_plastic < 1000
println(" ✓ Good performance - suitable for most applications")
println(" • Consider caching for problems with >10M DOF")
else
println(" ⚠ Acceptable performance - profile before using with >1M DOF")
println(" • Consider precomputation for repeated analyses")
end
println()
println("Expected performance in FEM assembly:")
println(" • Small problems (<10K DOF): Negligible overhead")
println(" • Medium problems (10K-1M DOF): ", @sprintf("<%.1f seconds", 1e6 * t_plastic / 1e9))
println(" • Large problems (>1M DOF): ", @sprintf("<%.1f seconds", 1e7 * t_plastic / 1e9))
println()
println("="^80)
println("Analysis complete: ", now())
println("="^80)
-348
View File
@@ -1,348 +0,0 @@
Precompiling packages...
669.3 ms ✓ EpollShim_jll
725.5 ms ✓ libfdk_aac_jll
760.9 ms ✓ Graphite2_jll
752.3 ms ✓ fzf_jll
756.4 ms ✓ LERC_jll
822.1 ms ✓ Xorg_libICE_jll
815.2 ms ✓ LAME_jll
797.9 ms ✓ Ogg_jll
788.9 ms ✓ x265_jll
829.1 ms ✓ libaom_jll
849.2 ms ✓ mtdev_jll
843.9 ms ✓ MbedTLS_jll
840.0 ms ✓ x264_jll
860.8 ms ✓ XZ_jll
674.8 ms ✓ libevdev_jll
769.9 ms ✓ Opus_jll
743.6 ms ✓ eudev_jll
738.4 ms ✓ FriBidi_jll
707.7 ms ✓ Dbus_jll
730.8 ms ✓ Xorg_libxkbfile_jll
779.8 ms ✓ Xorg_xcb_util_jll
770.5 ms ✓ Xorg_libXi_jll
772.1 ms ✓ Xorg_libXrandr_jll
793.7 ms ✓ Xorg_libXcursor_jll
892.3 ms ✓ Wayland_jll
2016.1 ms ✓ ColorVectorSpace
1008.1 ms ✓ HarfBuzz_jll
711.6 ms ✓ JLFzf
1304.3 ms ✓ Ghostscript_jll
737.4 ms ✓ Xorg_libSM_jll
1455.9 ms ✓ RecipesBase
803.4 ms ✓ libvorbis_jll
731.4 ms ✓ libinput_jll
785.6 ms ✓ Libtiff_jll
742.5 ms ✓ Xorg_xkbcomp_jll
731.2 ms ✓ Xorg_xcb_util_image_jll
733.0 ms ✓ Xorg_xcb_util_keysyms_jll
741.3 ms ✓ Xorg_xcb_util_renderutil_jll
2530.6 ms ✓ StatsBase
764.4 ms ✓ Xorg_xcb_util_wm_jll
1120.5 ms ✓ MbedTLS
879.4 ms ✓ libass_jll
653.5 ms ✓ Xorg_xkeyboard_config_jll
962.5 ms ✓ Pango_jll
711.2 ms ✓ Xorg_xcb_util_cursor_jll
758.4 ms ✓ xkbcommon_jll
1332.1 ms ✓ FFMPEG_jll
999.6 ms ✓ Vulkan_Loader_jll
1149.2 ms ✓ libdecor_jll
839.5 ms ✓ FFMPEG
2960.5 ms ✓ Latexify
974.9 ms ✓ GLFW_jll
4035.5 ms ✓ ColorSchemes
852.9 ms ✓ Latexify → SparseArraysExt
1490.1 ms ✓ Qt6Base_jll
918.9 ms ✓ Qt6ShaderTools_jll
1397.3 ms ✓ GR_jll
2814.7 ms ✓ Qt6Declarative_jll
1317.4 ms ✓ Qt6Wayland_jll
7309.2 ms ✓ PlotUtils
12361.6 ms ✓ HTTP
2980.3 ms ✓ PlotThemes
3588.3 ms ✓ RecipesPipeline
5614.7 ms ✓ GR
64840.3 ms ✓ Plots
65 dependencies successfully precompiled in 87 seconds. 112 already precompiled.
================================================================================
SYSTEM INFORMATION
================================================================================
CPU Model: Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz
CPU Cores: 32 threads (32 physical cores)
CPU Speed: 3300 MHz
Julia Version: 1.12.1
OS: Linux x86_64-linux-gnu
Word Size: 64 bits
Approximate CPU Cache Sizes:
L1 Cache: ~32-64 KB per core (typical)
L2 Cache: ~256-512 KB per core (typical)
L3 Cache: ~8-32 MB shared (typical)
Note: Testing up to 8KB structs to exceed L1 cache
================================================================================
SYSTEM INFORMATION
================================================================================
Julia Version: 1.12.1
CPU Model: Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz
CPU Cores: 32
Total Memory: 503.35 GB
L1 Cache: 48K
L2 Cache: 1280K
L3 Cache: 24576K
================================================================================
STRUCT SIZE SCALING BENCHMARK
================================================================================
Testing hypothesis: Immutable slows down with struct size, mutable stays constant
Testing struct with 1 Float64 fields (8 bytes)...
Access: Mut=4.62ns Imm=2.02ns Speedup=2.3x
Update: Mut=7.2ns Imm=2.32ns Speedup=3.1x
Iterate: Mut=11.77ns Imm=2.02ns Speedup=5.8x
Copy: 2.02ns
Testing struct with 2 Float64 fields (16 bytes)...
Access: Mut=4.9ns Imm=2.02ns Speedup=2.4x
Update: Mut=7.21ns Imm=2.31ns Speedup=3.1x
Iterate: Mut=14.49ns Imm=2.02ns Speedup=7.2x
Copy: 2.03ns
Testing struct with 5 Float64 fields (40 bytes)...
Access: Mut=4.9ns Imm=2.37ns Speedup=2.1x
Update: Mut=7.2ns Imm=2.6ns Speedup=2.8x
Iterate: Mut=16.71ns Imm=2.31ns Speedup=7.2x
Copy: 2.38ns
Testing struct with 10 Float64 fields (80 bytes)...
Access: Mut=4.62ns Imm=2.03ns Speedup=2.3x
Update: Mut=7.2ns Imm=2.6ns Speedup=2.8x
Iterate: Mut=23.13ns Imm=2.6ns Speedup=8.9x
Copy: 3.31ns
Testing struct with 20 Float64 fields (160 bytes)...
Access: Mut=4.62ns Imm=2.02ns Speedup=2.3x
Update: Mut=7.2ns Imm=3.16ns Speedup=2.3x
Iterate: Mut=63.13ns Imm=5.52ns Speedup=11.4x
Copy: 3.18ns
Testing struct with 50 Float64 fields (400 bytes)...
Access: Mut=4.62ns Imm=2.03ns Speedup=2.3x
Update: Mut=7.2ns Imm=6.02ns Speedup=1.2x
Iterate: Mut=196.85ns Imm=24.78ns Speedup=7.9x
Copy: 6.9ns
Testing struct with 100 Float64 fields (800 bytes)...
Access: Mut=4.62ns Imm=2.02ns Speedup=2.3x
Update: Mut=7.2ns Imm=11.75ns Speedup=0.6x
Iterate: Mut=260.56ns Imm=68.5ns Speedup=3.8x
Copy: 9.48ns
Testing struct with 200 Float64 fields (1600 bytes)...
Access: Mut=4.62ns Imm=2.37ns Speedup=1.9x
Update: Mut=7.41ns Imm=24.63ns Speedup=0.3x
Iterate: Mut=777.91ns Imm=155.16ns Speedup=5.0x
Copy: 20.95ns
Testing struct with 500 Float64 fields (4000 bytes)...
Access: Mut=4.62ns Imm=2.03ns Speedup=2.3x
Update: Mut=7.2ns Imm=80.84ns Speedup=0.1x
Iterate: Mut=1156.3ns Imm=499.36ns Speedup=2.3x
Copy: 58.68ns
Testing struct with 1000 Float64 fields (8000 bytes)...
Access: Mut=4.67ns Imm=2.08ns Speedup=2.2x
Update: Mut=7.2ns Imm=193.26ns Speedup=0.0x
Iterate: Mut=3465.75ns Imm=1100.2ns Speedup=3.2x
Copy: 45.7ns
Testing struct with 2000 Float64 fields (16000 bytes)...
Access: Mut=4.62ns Imm=2.02ns Speedup=2.3x
Update: Mut=7.2ns Imm=410.85ns Speedup=0.0x
Iterate: Mut=4811.71ns Imm=2226.22ns Speedup=2.2x
Copy: 82.95ns
Testing struct with 5000 Float64 fields (40000 bytes)...
Access: Mut=4.62ns Imm=2.08ns Speedup=2.2x
Update: Mut=7.2ns Imm=1936.1ns Speedup=0.0x
Iterate: Mut=33969.0ns Imm=5667.17ns Speedup=6.0x
Copy: 867.04ns
================================================================================
RESULTS SUMMARY
================================================================================
Field Access Performance:
Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup
----------------------------------------------------------------------
1 | 8 | 4.62 | 2.02 | 2.3x
2 | 16 | 4.90 | 2.02 | 2.4x
5 | 40 | 4.90 | 2.37 | 2.1x
10 | 80 | 4.62 | 2.03 | 2.3x
20 | 160 | 4.62 | 2.02 | 2.3x
50 | 400 | 4.62 | 2.03 | 2.3x
100 | 800 | 4.62 | 2.02 | 2.3x
200 | 1600 | 4.62 | 2.37 | 1.9x
500 | 4000 | 4.62 | 2.03 | 2.3x
1000 | 8000 | 4.67 | 2.08 | 2.2x
2000 | 16000 | 4.62 | 2.02 | 2.3x
5000 | 40000 | 4.62 | 2.08 | 2.2x
Field Update Performance:
Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup
----------------------------------------------------------------------
1 | 8 | 7.20 | 2.31 | 3.1x
2 | 16 | 7.21 | 2.31 | 3.1x
5 | 40 | 7.20 | 2.60 | 2.8x
10 | 80 | 7.20 | 2.60 | 2.8x
20 | 160 | 7.20 | 3.16 | 2.3x
50 | 400 | 7.20 | 6.02 | 1.2x
100 | 800 | 7.20 | 11.75 | 0.6x
200 | 1600 | 7.41 | 24.63 | 0.3x
500 | 4000 | 7.20 | 80.84 | 0.1x
1000 | 8000 | 7.20 | 193.26 | 0.0x
2000 | 16000 | 7.20 | 410.85 | 0.0x
5000 | 40000 | 7.20 | 1936.10 | 0.0x
Iteration Performance:
Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup
----------------------------------------------------------------------
1 | 8 | 11.77 | 2.02 | 5.8x
2 | 16 | 14.49 | 2.02 | 7.2x
5 | 40 | 16.71 | 2.31 | 7.2x
10 | 80 | 23.13 | 2.60 | 8.9x
20 | 160 | 63.13 | 5.52 | 11.4x
50 | 400 | 196.85 | 24.78 | 7.9x
100 | 800 | 260.56 | 68.50 | 3.8x
200 | 1600 | 777.91 | 155.16 | 5.0x
500 | 4000 | 1156.30 | 499.36 | 2.3x
1000 | 8000 | 3465.75 | 1100.20 | 3.2x
2000 | 16000 | 4811.71 | 2226.22 | 2.2x
5000 | 40000 | 33969.00 | 5667.17 | 6.0x
Immutable Copy Cost (ns):
Size (fields) | Bytes | Copy Time (ns)
----------------------------------------
1 | 8 | 2.02
2 | 16 | 2.03
5 | 40 | 2.38
10 | 80 | 3.31
20 | 160 | 3.18
50 | 400 | 6.90
100 | 800 | 9.48
200 | 1600 | 20.95
500 | 4000 | 58.68
1000 | 8000 | 45.70
2000 | 16000 | 82.95
5000 | 40000 | 867.04
================================================================================
ANALYSIS
================================================================================
✓ Immutable ALWAYS faster for field access (even at 1000 fields = 8KB)
Minimum speedup: 1.9x at 5000 fields
⚠ Mutable wins for field update at 100 fields
✓ Immutable ALWAYS faster for iteration (even at 1000 fields = 8KB)
Minimum speedup: 2.2x at 5000 fields
Scaling Analysis:
Copy time scaling:
Linear fit: time(ns) = -25.47 + 0.1587 * nfields
Per-field cost: 0.1587 ns/field
Base overhead: -25.47 ns
Is copy time linear? (checking R²)
R² = 0.9006
⚠ Copy time not perfectly linear (compiler optimizations?)
KEY INSIGHT:
--------------------------------------------------------------------------------
Even at 1000 fields (8KB struct), immutable is STILL faster because:
1. Dict lookup cost (~40-50ns) >> copy cost per field (~0.1587ns)
2. Type stability enables compiler optimizations (inlining, SIMD)
3. Stack allocation has better cache locality than heap pointers
Theoretical crossover point (if it exists):
Would occur at ~413 fields (3KB)
================================================================================
CONCLUSION
================================================================================
Your intuition about O(n) scaling is CORRECT, BUT:
• Dict lookup base cost is SO high (~40-50ns)
• Copy cost per field is SO low (~0.1587ns)
• Compiler optimizations are SO good (inlining, SIMD, escape analysis)
That immutable wins even for unrealistically large structs (8KB+)!
For typical FEM elements:
• Material properties: 3-10 fields (24-80 bytes)
• State variables: 10-50 fields (80-400 bytes)
• Even with 100 fields (800 bytes), immutable is >10x faster
Type stability > Everything else.
================================================================================
SAVING DATA
================================================================================
✓ Data saved to: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/struct_size_scaling.json
✓ CSV saved to: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/struct_size_scaling.csv
================================================================================
GENERATING PLOTS
================================================================================
]1337;ReportCellSizeP+q544e\GKS: cannot open display - headless operation mode active
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/field_access_scaling.png
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/field_update_scaling.png
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/iteration_scaling.png
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/copy_cost_linear.png
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/speedup_ratios.png
All plots saved to: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results
================================================================================
SAVING DATA
================================================================================
✓ Data saved to: benchmarks/results/struct_scaling_20251109_201256.json
✓ CSV saved to: benchmarks/results/struct_scaling_20251109_201256.csv
================================================================================
GENERATING PLOTS
================================================================================
┌ Warning: Assignment to `p1` in soft scope is ambiguous because a global variable by the same name exists: `p1` will be treated as a new local. Disambiguate by using `local p1` to suppress this warning or `global p1` to assign to the existing global variable.
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:633
┌ Warning: Assignment to `p2` in soft scope is ambiguous because a global variable by the same name exists: `p2` will be treated as a new local. Disambiguate by using `local p2` to suppress this warning or `global p2` to assign to the existing global variable.
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:649
┌ Warning: Assignment to `p3` in soft scope is ambiguous because a global variable by the same name exists: `p3` will be treated as a new local. Disambiguate by using `local p3` to suppress this warning or `global p3` to assign to the existing global variable.
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:664
┌ Warning: Assignment to `p4` in soft scope is ambiguous because a global variable by the same name exists: `p4` will be treated as a new local. Disambiguate by using `local p4` to suppress this warning or `global p4` to assign to the existing global variable.
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:679
┌ Warning: Assignment to `p5` in soft scope is ambiguous because a global variable by the same name exists: `p5` will be treated as a new local. Disambiguate by using `local p5` to suppress this warning or `global p5` to assign to the existing global variable.
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:697
✓ Saved: field_access_20251109_201256.png
✓ Saved: field_update_20251109_201256.png
✓ Saved: iteration_20251109_201256.png
✓ Saved: speedup_factors_20251109_201256.png
✓ Saved: copy_cost_20251109_201256.png
✓ Saved: combined_20251109_201256.png
All plots saved successfully!
================================================================================
Binary file not shown.

Before

Width:  |  Height:  |  Size: 153 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 36 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 53 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 34 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 46 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 31 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 56 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 36 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 44 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 54 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 54 KiB

@@ -1,13 +0,0 @@
nfields,bytes,mut_access_ns,imm_access_ns,speedup_access,mut_update_ns,imm_update_ns,speedup_update,mut_iter_ns,imm_iter_ns,speedup_iter,imm_copy_ns
1,8,4.615,2.023,2.2812654473554126,7.1991991991991995,2.315,3.1098052696324836,11.773773773773774,2.024,5.817081904038426,2.023
2,16,4.904,2.023,2.4241225902125554,7.207207207207207,2.314,3.11460985618289,14.48997995991984,2.025,7.1555456592196744,2.029
5,40,4.897,2.373,2.063632532659081,7.201201201201201,2.604,2.765438249309217,16.70941883767535,2.312,7.227257282731553,2.383
10,80,4.621,2.029,2.2774765894529327,7.1991991991991995,2.602,2.7667944654877785,23.13152610441767,2.602,8.889902422912249,3.311
20,160,4.618,2.023,2.282748393475037,7.201201201201201,3.1633266533066133,2.2764646179280326,63.131632653061224,5.523,11.43067764857165,3.177
50,400,4.617,2.029,2.275505174963036,7.197197197197197,6.022,1.1951506471599465,196.84902597402598,24.783132530120483,7.942862982909166,6.895895895895896
100,800,4.618,2.025,2.280493827160494,7.198198198198198,11.745745745745745,0.612834498039884,260.55786350148367,68.50307377049181,3.803593753682347,9.476476476476476
200,1600,4.621,2.373,1.9473240623683101,7.41041041041041,24.633534136546185,0.30082611651798524,777.9142857142857,155.1606475716065,5.013605562294905,20.948897795591183
500,4000,4.624,2.027,2.281203749383325,7.198198198198198,80.83854166666667,0.08904413723690831,1156.3,499.35567010309273,2.3155840000000003,58.68463886063072
1000,8000,4.671,2.084,2.241362763915547,7.2002002002002,193.26257861635222,0.03725604952469045,3465.75,1100.2,3.1501090710779853,45.70171890798787
2000,16000,4.625,2.024,2.2850790513833994,7.197197197197197,410.8542713567839,0.0175176399491468,4811.714285714285,2226.222222222222,2.1613809428742545,82.94813278008299
5000,40000,4.623,2.083,2.2193951032165145,7.2042042042042045,1936.1,0.0037209876577677828,33969.0,5667.166666666667,5.994000529365056,867.0408163265306
1 nfields bytes mut_access_ns imm_access_ns speedup_access mut_update_ns imm_update_ns speedup_update mut_iter_ns imm_iter_ns speedup_iter imm_copy_ns
2 1 8 4.615 2.023 2.2812654473554126 7.1991991991991995 2.315 3.1098052696324836 11.773773773773774 2.024 5.817081904038426 2.023
3 2 16 4.904 2.023 2.4241225902125554 7.207207207207207 2.314 3.11460985618289 14.48997995991984 2.025 7.1555456592196744 2.029
4 5 40 4.897 2.373 2.063632532659081 7.201201201201201 2.604 2.765438249309217 16.70941883767535 2.312 7.227257282731553 2.383
5 10 80 4.621 2.029 2.2774765894529327 7.1991991991991995 2.602 2.7667944654877785 23.13152610441767 2.602 8.889902422912249 3.311
6 20 160 4.618 2.023 2.282748393475037 7.201201201201201 3.1633266533066133 2.2764646179280326 63.131632653061224 5.523 11.43067764857165 3.177
7 50 400 4.617 2.029 2.275505174963036 7.197197197197197 6.022 1.1951506471599465 196.84902597402598 24.783132530120483 7.942862982909166 6.895895895895896
8 100 800 4.618 2.025 2.280493827160494 7.198198198198198 11.745745745745745 0.612834498039884 260.55786350148367 68.50307377049181 3.803593753682347 9.476476476476476
9 200 1600 4.621 2.373 1.9473240623683101 7.41041041041041 24.633534136546185 0.30082611651798524 777.9142857142857 155.1606475716065 5.013605562294905 20.948897795591183
10 500 4000 4.624 2.027 2.281203749383325 7.198198198198198 80.83854166666667 0.08904413723690831 1156.3 499.35567010309273 2.3155840000000003 58.68463886063072
11 1000 8000 4.671 2.084 2.241362763915547 7.2002002002002 193.26257861635222 0.03725604952469045 3465.75 1100.2 3.1501090710779853 45.70171890798787
12 2000 16000 4.625 2.024 2.2850790513833994 7.197197197197197 410.8542713567839 0.0175176399491468 4811.714285714285 2226.222222222222 2.1613809428742545 82.94813278008299
13 5000 40000 4.623 2.083 2.2193951032165145 7.2042042042042045 1936.1 0.0037209876577677828 33969.0 5667.166666666667 5.994000529365056 867.0408163265306
@@ -1,187 +0,0 @@
{
"julia_version": "1.12.1",
"analysis": {
"copy_intercept_ns": -25.467962218157183,
"copy_slope_ns_per_field": 0.1586673181436861,
"r_squared": 0.9005527188987935
},
"results": [
{
"mutable_update_ns": 7.1991991991991995,
"immutable_copy_ns": 2.023,
"mutable_access_ns": 4.615,
"bytes": 8,
"speedup_update": 3.1098052696324836,
"speedup_iter": 5.817081904038426,
"immutable_access_ns": 2.023,
"speedup_access": 2.2812654473554126,
"nfields": 1,
"immutable_update_ns": 2.315,
"mutable_iter_ns": 11.773773773773774,
"immutable_iter_ns": 2.024
},
{
"mutable_update_ns": 7.207207207207207,
"immutable_copy_ns": 2.029,
"mutable_access_ns": 4.904,
"bytes": 16,
"speedup_update": 3.11460985618289,
"speedup_iter": 7.1555456592196744,
"immutable_access_ns": 2.023,
"speedup_access": 2.4241225902125554,
"nfields": 2,
"immutable_update_ns": 2.314,
"mutable_iter_ns": 14.48997995991984,
"immutable_iter_ns": 2.025
},
{
"mutable_update_ns": 7.201201201201201,
"immutable_copy_ns": 2.383,
"mutable_access_ns": 4.897,
"bytes": 40,
"speedup_update": 2.765438249309217,
"speedup_iter": 7.227257282731553,
"immutable_access_ns": 2.373,
"speedup_access": 2.063632532659081,
"nfields": 5,
"immutable_update_ns": 2.604,
"mutable_iter_ns": 16.70941883767535,
"immutable_iter_ns": 2.312
},
{
"mutable_update_ns": 7.1991991991991995,
"immutable_copy_ns": 3.311,
"mutable_access_ns": 4.621,
"bytes": 80,
"speedup_update": 2.7667944654877785,
"speedup_iter": 8.889902422912249,
"immutable_access_ns": 2.029,
"speedup_access": 2.2774765894529327,
"nfields": 10,
"immutable_update_ns": 2.602,
"mutable_iter_ns": 23.13152610441767,
"immutable_iter_ns": 2.602
},
{
"mutable_update_ns": 7.201201201201201,
"immutable_copy_ns": 3.177,
"mutable_access_ns": 4.618,
"bytes": 160,
"speedup_update": 2.2764646179280326,
"speedup_iter": 11.43067764857165,
"immutable_access_ns": 2.023,
"speedup_access": 2.282748393475037,
"nfields": 20,
"immutable_update_ns": 3.1633266533066133,
"mutable_iter_ns": 63.131632653061224,
"immutable_iter_ns": 5.523
},
{
"mutable_update_ns": 7.197197197197197,
"immutable_copy_ns": 6.895895895895896,
"mutable_access_ns": 4.617,
"bytes": 400,
"speedup_update": 1.1951506471599465,
"speedup_iter": 7.942862982909166,
"immutable_access_ns": 2.029,
"speedup_access": 2.275505174963036,
"nfields": 50,
"immutable_update_ns": 6.022,
"mutable_iter_ns": 196.84902597402598,
"immutable_iter_ns": 24.783132530120483
},
{
"mutable_update_ns": 7.198198198198198,
"immutable_copy_ns": 9.476476476476476,
"mutable_access_ns": 4.618,
"bytes": 800,
"speedup_update": 0.612834498039884,
"speedup_iter": 3.803593753682347,
"immutable_access_ns": 2.025,
"speedup_access": 2.280493827160494,
"nfields": 100,
"immutable_update_ns": 11.745745745745745,
"mutable_iter_ns": 260.55786350148367,
"immutable_iter_ns": 68.50307377049181
},
{
"mutable_update_ns": 7.41041041041041,
"immutable_copy_ns": 20.948897795591183,
"mutable_access_ns": 4.621,
"bytes": 1600,
"speedup_update": 0.30082611651798524,
"speedup_iter": 5.013605562294905,
"immutable_access_ns": 2.373,
"speedup_access": 1.9473240623683101,
"nfields": 200,
"immutable_update_ns": 24.633534136546185,
"mutable_iter_ns": 777.9142857142857,
"immutable_iter_ns": 155.1606475716065
},
{
"mutable_update_ns": 7.198198198198198,
"immutable_copy_ns": 58.68463886063072,
"mutable_access_ns": 4.624,
"bytes": 4000,
"speedup_update": 0.08904413723690831,
"speedup_iter": 2.3155840000000003,
"immutable_access_ns": 2.027,
"speedup_access": 2.281203749383325,
"nfields": 500,
"immutable_update_ns": 80.83854166666667,
"mutable_iter_ns": 1156.3,
"immutable_iter_ns": 499.35567010309273
},
{
"mutable_update_ns": 7.2002002002002,
"immutable_copy_ns": 45.70171890798787,
"mutable_access_ns": 4.671,
"bytes": 8000,
"speedup_update": 0.03725604952469045,
"speedup_iter": 3.1501090710779853,
"immutable_access_ns": 2.084,
"speedup_access": 2.241362763915547,
"nfields": 1000,
"immutable_update_ns": 193.26257861635222,
"mutable_iter_ns": 3465.75,
"immutable_iter_ns": 1100.2
},
{
"mutable_update_ns": 7.197197197197197,
"immutable_copy_ns": 82.94813278008299,
"mutable_access_ns": 4.625,
"bytes": 16000,
"speedup_update": 0.0175176399491468,
"speedup_iter": 2.1613809428742545,
"immutable_access_ns": 2.024,
"speedup_access": 2.2850790513833994,
"nfields": 2000,
"immutable_update_ns": 410.8542713567839,
"mutable_iter_ns": 4811.714285714285,
"immutable_iter_ns": 2226.222222222222
},
{
"mutable_update_ns": 7.2042042042042045,
"immutable_copy_ns": 867.0408163265306,
"mutable_access_ns": 4.623,
"bytes": 40000,
"speedup_update": 0.0037209876577677828,
"speedup_iter": 5.994000529365056,
"immutable_access_ns": 2.083,
"speedup_access": 2.2193951032165145,
"nfields": 5000,
"immutable_update_ns": 1936.1,
"mutable_iter_ns": 33969.0,
"immutable_iter_ns": 5667.166666666667
}
],
"timestamp": "20251109_201256",
"system": {
"cpu_model": "Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz",
"cpu_speed_mhz": 3300,
"machine": "x86_64-linux-gnu",
"word_size": 64,
"cpu_cores": 32,
"os": "Linux"
}
}
@@ -1,13 +0,0 @@
nfields,bytes,mut_access_ns,imm_access_ns,speedup_access,mut_update_ns,imm_update_ns,speedup_update,imm_copy_ns,mut_iter_ns,imm_iter_ns,speedup_iter
1,8,4.615,2.023,2.2812654473554126,7.1991991991991995,2.315,3.1098052696324836,2.023,11.773773773773774,2.024,5.817081904038426
2,16,4.904,2.023,2.4241225902125554,7.207207207207207,2.314,3.11460985618289,2.029,14.48997995991984,2.025,7.1555456592196744
5,40,4.897,2.373,2.063632532659081,7.201201201201201,2.604,2.765438249309217,2.383,16.70941883767535,2.312,7.227257282731553
10,80,4.621,2.029,2.2774765894529327,7.1991991991991995,2.602,2.7667944654877785,3.311,23.13152610441767,2.602,8.889902422912249
20,160,4.618,2.023,2.282748393475037,7.201201201201201,3.1633266533066133,2.2764646179280326,3.177,63.131632653061224,5.523,11.43067764857165
50,400,4.617,2.029,2.275505174963036,7.197197197197197,6.022,1.1951506471599465,6.895895895895896,196.84902597402598,24.783132530120483,7.942862982909166
100,800,4.618,2.025,2.280493827160494,7.198198198198198,11.745745745745745,0.612834498039884,9.476476476476476,260.55786350148367,68.50307377049181,3.803593753682347
200,1600,4.621,2.373,1.9473240623683101,7.41041041041041,24.633534136546185,0.30082611651798524,20.948897795591183,777.9142857142857,155.1606475716065,5.013605562294905
500,4000,4.624,2.027,2.281203749383325,7.198198198198198,80.83854166666667,0.08904413723690831,58.68463886063072,1156.3,499.35567010309273,2.3155840000000003
1000,8000,4.671,2.084,2.241362763915547,7.2002002002002,193.26257861635222,0.03725604952469045,45.70171890798787,3465.75,1100.2,3.1501090710779853
2000,16000,4.625,2.024,2.2850790513833994,7.197197197197197,410.8542713567839,0.0175176399491468,82.94813278008299,4811.714285714285,2226.222222222222,2.1613809428742545
5000,40000,4.623,2.083,2.2193951032165145,7.2042042042042045,1936.1,0.0037209876577677828,867.0408163265306,33969.0,5667.166666666667,5.994000529365056
1 nfields bytes mut_access_ns imm_access_ns speedup_access mut_update_ns imm_update_ns speedup_update imm_copy_ns mut_iter_ns imm_iter_ns speedup_iter
2 1 8 4.615 2.023 2.2812654473554126 7.1991991991991995 2.315 3.1098052696324836 2.023 11.773773773773774 2.024 5.817081904038426
3 2 16 4.904 2.023 2.4241225902125554 7.207207207207207 2.314 3.11460985618289 2.029 14.48997995991984 2.025 7.1555456592196744
4 5 40 4.897 2.373 2.063632532659081 7.201201201201201 2.604 2.765438249309217 2.383 16.70941883767535 2.312 7.227257282731553
5 10 80 4.621 2.029 2.2774765894529327 7.1991991991991995 2.602 2.7667944654877785 3.311 23.13152610441767 2.602 8.889902422912249
6 20 160 4.618 2.023 2.282748393475037 7.201201201201201 3.1633266533066133 2.2764646179280326 3.177 63.131632653061224 5.523 11.43067764857165
7 50 400 4.617 2.029 2.275505174963036 7.197197197197197 6.022 1.1951506471599465 6.895895895895896 196.84902597402598 24.783132530120483 7.942862982909166
8 100 800 4.618 2.025 2.280493827160494 7.198198198198198 11.745745745745745 0.612834498039884 9.476476476476476 260.55786350148367 68.50307377049181 3.803593753682347
9 200 1600 4.621 2.373 1.9473240623683101 7.41041041041041 24.633534136546185 0.30082611651798524 20.948897795591183 777.9142857142857 155.1606475716065 5.013605562294905
10 500 4000 4.624 2.027 2.281203749383325 7.198198198198198 80.83854166666667 0.08904413723690831 58.68463886063072 1156.3 499.35567010309273 2.3155840000000003
11 1000 8000 4.671 2.084 2.241362763915547 7.2002002002002 193.26257861635222 0.03725604952469045 45.70171890798787 3465.75 1100.2 3.1501090710779853
12 2000 16000 4.625 2.024 2.2850790513833994 7.197197197197197 410.8542713567839 0.0175176399491468 82.94813278008299 4811.714285714285 2226.222222222222 2.1613809428742545
13 5000 40000 4.623 2.083 2.2193951032165145 7.2042042042042045 1936.1 0.0037209876577677828 867.0408163265306 33969.0 5667.166666666667 5.994000529365056
-196
View File
@@ -1,196 +0,0 @@
{
"struct_sizes": [
1,
2,
5,
10,
20,
50,
100,
200,
500,
1000,
2000,
5000
],
"system_info": {
"cpu_model": "Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz",
"julia_version": "1.12.1",
"total_memory_gb": 503.35,
"l1_cache": "48K",
"l2_cache": "1280K",
"l3_cache": "24576K",
"cpu_cores": 32
},
"results": [
{
"mutable_update_ns": 7.1991991991991995,
"immutable_copy_ns": 2.023,
"mutable_access_ns": 4.615,
"bytes": 8,
"speedup_update": 3.1098052696324836,
"speedup_iter": 5.817081904038426,
"immutable_access_ns": 2.023,
"speedup_access": 2.2812654473554126,
"nfields": 1,
"immutable_update_ns": 2.315,
"mutable_iter_ns": 11.773773773773774,
"immutable_iter_ns": 2.024
},
{
"mutable_update_ns": 7.207207207207207,
"immutable_copy_ns": 2.029,
"mutable_access_ns": 4.904,
"bytes": 16,
"speedup_update": 3.11460985618289,
"speedup_iter": 7.1555456592196744,
"immutable_access_ns": 2.023,
"speedup_access": 2.4241225902125554,
"nfields": 2,
"immutable_update_ns": 2.314,
"mutable_iter_ns": 14.48997995991984,
"immutable_iter_ns": 2.025
},
{
"mutable_update_ns": 7.201201201201201,
"immutable_copy_ns": 2.383,
"mutable_access_ns": 4.897,
"bytes": 40,
"speedup_update": 2.765438249309217,
"speedup_iter": 7.227257282731553,
"immutable_access_ns": 2.373,
"speedup_access": 2.063632532659081,
"nfields": 5,
"immutable_update_ns": 2.604,
"mutable_iter_ns": 16.70941883767535,
"immutable_iter_ns": 2.312
},
{
"mutable_update_ns": 7.1991991991991995,
"immutable_copy_ns": 3.311,
"mutable_access_ns": 4.621,
"bytes": 80,
"speedup_update": 2.7667944654877785,
"speedup_iter": 8.889902422912249,
"immutable_access_ns": 2.029,
"speedup_access": 2.2774765894529327,
"nfields": 10,
"immutable_update_ns": 2.602,
"mutable_iter_ns": 23.13152610441767,
"immutable_iter_ns": 2.602
},
{
"mutable_update_ns": 7.201201201201201,
"immutable_copy_ns": 3.177,
"mutable_access_ns": 4.618,
"bytes": 160,
"speedup_update": 2.2764646179280326,
"speedup_iter": 11.43067764857165,
"immutable_access_ns": 2.023,
"speedup_access": 2.282748393475037,
"nfields": 20,
"immutable_update_ns": 3.1633266533066133,
"mutable_iter_ns": 63.131632653061224,
"immutable_iter_ns": 5.523
},
{
"mutable_update_ns": 7.197197197197197,
"immutable_copy_ns": 6.895895895895896,
"mutable_access_ns": 4.617,
"bytes": 400,
"speedup_update": 1.1951506471599465,
"speedup_iter": 7.942862982909166,
"immutable_access_ns": 2.029,
"speedup_access": 2.275505174963036,
"nfields": 50,
"immutable_update_ns": 6.022,
"mutable_iter_ns": 196.84902597402598,
"immutable_iter_ns": 24.783132530120483
},
{
"mutable_update_ns": 7.198198198198198,
"immutable_copy_ns": 9.476476476476476,
"mutable_access_ns": 4.618,
"bytes": 800,
"speedup_update": 0.612834498039884,
"speedup_iter": 3.803593753682347,
"immutable_access_ns": 2.025,
"speedup_access": 2.280493827160494,
"nfields": 100,
"immutable_update_ns": 11.745745745745745,
"mutable_iter_ns": 260.55786350148367,
"immutable_iter_ns": 68.50307377049181
},
{
"mutable_update_ns": 7.41041041041041,
"immutable_copy_ns": 20.948897795591183,
"mutable_access_ns": 4.621,
"bytes": 1600,
"speedup_update": 0.30082611651798524,
"speedup_iter": 5.013605562294905,
"immutable_access_ns": 2.373,
"speedup_access": 1.9473240623683101,
"nfields": 200,
"immutable_update_ns": 24.633534136546185,
"mutable_iter_ns": 777.9142857142857,
"immutable_iter_ns": 155.1606475716065
},
{
"mutable_update_ns": 7.198198198198198,
"immutable_copy_ns": 58.68463886063072,
"mutable_access_ns": 4.624,
"bytes": 4000,
"speedup_update": 0.08904413723690831,
"speedup_iter": 2.3155840000000003,
"immutable_access_ns": 2.027,
"speedup_access": 2.281203749383325,
"nfields": 500,
"immutable_update_ns": 80.83854166666667,
"mutable_iter_ns": 1156.3,
"immutable_iter_ns": 499.35567010309273
},
{
"mutable_update_ns": 7.2002002002002,
"immutable_copy_ns": 45.70171890798787,
"mutable_access_ns": 4.671,
"bytes": 8000,
"speedup_update": 0.03725604952469045,
"speedup_iter": 3.1501090710779853,
"immutable_access_ns": 2.084,
"speedup_access": 2.241362763915547,
"nfields": 1000,
"immutable_update_ns": 193.26257861635222,
"mutable_iter_ns": 3465.75,
"immutable_iter_ns": 1100.2
},
{
"mutable_update_ns": 7.197197197197197,
"immutable_copy_ns": 82.94813278008299,
"mutable_access_ns": 4.625,
"bytes": 16000,
"speedup_update": 0.0175176399491468,
"speedup_iter": 2.1613809428742545,
"immutable_access_ns": 2.024,
"speedup_access": 2.2850790513833994,
"nfields": 2000,
"immutable_update_ns": 410.8542713567839,
"mutable_iter_ns": 4811.714285714285,
"immutable_iter_ns": 2226.222222222222
},
{
"mutable_update_ns": 7.2042042042042045,
"immutable_copy_ns": 867.0408163265306,
"mutable_access_ns": 4.623,
"bytes": 40000,
"speedup_update": 0.0037209876577677828,
"speedup_iter": 5.994000529365056,
"immutable_access_ns": 2.083,
"speedup_access": 2.2193951032165145,
"nfields": 5000,
"immutable_update_ns": 1936.1,
"mutable_iter_ns": 33969.0,
"immutable_iter_ns": 5667.166666666667
}
],
"timestamp": "2025-11-09T20:12:52.503"
}
-731
View File
@@ -1,731 +0,0 @@
# Benchmark: Does immutable performance degrade with struct size?
# Theory: Stack copying is O(n), heap pointers are O(1)
# Question: At what size does mutable win?
# NOTE: Using packages from global environment (not project)
using BenchmarkTools
using Printf
using JSON
using Plots
using JSON
using Dates
# Get system information
println("="^80)
println("SYSTEM INFORMATION")
println("="^80)
println()
# CPU info
cpu_info = Sys.cpu_info()
println("CPU Model: ", cpu_info[1].model)
println("CPU Cores: ", Sys.CPU_THREADS, " threads (", length(cpu_info), " physical cores)")
println("CPU Speed: ", cpu_info[1].speed, " MHz")
println()
# Julia and system info
println("Julia Version: ", VERSION)
println("OS: ", Sys.KERNEL, " ", Sys.MACHINE)
println("Word Size: ", Sys.WORD_SIZE, " bits")
println()
# Memory and cache info (approximate)
println("Approximate CPU Cache Sizes:")
println(" L1 Cache: ~32-64 KB per core (typical)")
println(" L2 Cache: ~256-512 KB per core (typical)")
println(" L3 Cache: ~8-32 MB shared (typical)")
println()
println("Note: Testing up to 8KB structs to exceed L1 cache")
println()
using BenchmarkTools
using Printf
using JSON
using Plots
# Collect system information
function get_system_info()
info = Dict{String,Any}()
info["julia_version"] = string(VERSION)
info["cpu_model"] = Sys.cpu_info()[1].model
info["cpu_cores"] = Sys.CPU_THREADS
info["total_memory_gb"] = round(Sys.total_memory() / 1024^3, digits=2)
# Try to get CPU cache info (Linux)
try
if Sys.islinux()
l1_cache = read("/sys/devices/system/cpu/cpu0/cache/index0/size", String) |> strip
l2_cache = read("/sys/devices/system/cpu/cpu0/cache/index2/size", String) |> strip
l3_cache = read("/sys/devices/system/cpu/cpu0/cache/index3/size", String) |> strip
info["l1_cache"] = l1_cache
info["l2_cache"] = l2_cache
info["l3_cache"] = l3_cache
end
catch
info["cache_info"] = "Not available"
end
return info
end
system_info = get_system_info()
println("="^80)
println("SYSTEM INFORMATION")
println("="^80)
println("Julia Version: $(system_info["julia_version"])")
println("CPU Model: $(system_info["cpu_model"])")
println("CPU Cores: $(system_info["cpu_cores"])")
println("Total Memory: $(system_info["total_memory_gb"]) GB")
if haskey(system_info, "l1_cache")
println("L1 Cache: $(system_info["l1_cache"])")
println("L2 Cache: $(system_info["l2_cache"])")
println("L3 Cache: $(system_info["l3_cache"])")
end
println()
println("="^80)
println("STRUCT SIZE SCALING BENCHMARK")
println("="^80)
println()
println("Testing hypothesis: Immutable slows down with struct size, mutable stays constant")
println()
# Test different struct sizes (number of Float64 fields)
# Extended range to go well beyond register file and L1 cache
STRUCT_SIZES = [1, 2, 5, 10, 20, 50, 100, 200, 500, 1000, 2000, 5000]
results = []
for nfields in STRUCT_SIZES
println("Testing struct with $nfields Float64 fields ($(nfields * 8) bytes)...")
# Generate mutable version
field_names_mut = [Symbol("field$i") for i in 1:nfields]
# Mutable: Dict-based
mutable_data = Dict{Symbol,Float64}()
for fname in field_names_mut
mutable_data[fname] = rand()
end
# Immutable: NamedTuple-based
immutable_data = NamedTuple{Tuple(field_names_mut)}(Tuple(rand() for _ in 1:nfields))
# Benchmark 1: Field Access (read first field)
first_field = field_names_mut[1]
time_mut_access = @belapsed $mutable_data[$first_field]
time_imm_access = @belapsed $immutable_data.$first_field
# Benchmark 2: Field Update (change first field)
time_mut_update = @belapsed begin
$mutable_data[$first_field] = 42.0
end
time_imm_update = @belapsed begin
$immutable_data = (; $immutable_data..., $first_field=42.0)
end
# Benchmark 3: Struct Copy (merge with empty to force copy)
time_imm_copy = @belapsed merge($immutable_data, NamedTuple())
# Benchmark 4: Iteration over all fields
time_mut_iter = @belapsed begin
sum = 0.0
for (k, v) in $mutable_data
sum += v
end
sum
end
time_imm_iter = @belapsed begin
sum = 0.0
for v in $immutable_data
sum += v
end
sum
end
speedup_access = time_mut_access / time_imm_access
speedup_update = time_mut_update / time_imm_update
speedup_iter = time_mut_iter / time_imm_iter
push!(results, (
nfields=nfields,
bytes=nfields * 8,
# Access times
mut_access=time_mut_access,
imm_access=time_imm_access,
speedup_access=speedup_access,
# Update times
mut_update=time_mut_update,
imm_update=time_imm_update,
speedup_update=speedup_update,
# Copy time
imm_copy=time_imm_copy,
# Iteration times
mut_iter=time_mut_iter,
imm_iter=time_imm_iter,
speedup_iter=speedup_iter
))
println(" Access: Mut=$(round(time_mut_access*1e9, digits=2))ns Imm=$(round(time_imm_access*1e9, digits=2))ns Speedup=$(round(speedup_access, digits=1))x")
println(" Update: Mut=$(round(time_mut_update*1e9, digits=2))ns Imm=$(round(time_imm_update*1e9, digits=2))ns Speedup=$(round(speedup_update, digits=1))x")
println(" Iterate: Mut=$(round(time_mut_iter*1e9, digits=2))ns Imm=$(round(time_imm_iter*1e9, digits=2))ns Speedup=$(round(speedup_iter, digits=1))x")
println(" Copy: $(round(time_imm_copy*1e9, digits=2))ns")
println()
end
println("="^80)
println("RESULTS SUMMARY")
println("="^80)
println()
println("Field Access Performance:")
println("Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup")
println("-"^70)
for r in results
@printf("%13d | %5d | %12.2f | %14.2f | %6.1fx\n",
r.nfields, r.bytes, r.mut_access * 1e9, r.imm_access * 1e9, r.speedup_access)
end
println()
println("Field Update Performance:")
println("Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup")
println("-"^70)
for r in results
@printf("%13d | %5d | %12.2f | %14.2f | %6.1fx\n",
r.nfields, r.bytes, r.mut_update * 1e9, r.imm_update * 1e9, r.speedup_update)
end
println()
println("Iteration Performance:")
println("Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup")
println("-"^70)
for r in results
@printf("%13d | %5d | %12.2f | %14.2f | %6.1fx\n",
r.nfields, r.bytes, r.mut_iter * 1e9, r.imm_iter * 1e9, r.speedup_iter)
end
println()
println("Immutable Copy Cost (ns):")
println("Size (fields) | Bytes | Copy Time (ns)")
println("-"^40)
for r in results
@printf("%13d | %5d | %13.2f\n", r.nfields, r.bytes, r.imm_copy * 1e9)
end
println()
# Analysis
println("="^80)
println("ANALYSIS")
println("="^80)
println()
# Check if mutable ever wins
access_wins = [r for r in results if r.speedup_access < 1.0]
update_wins = [r for r in results if r.speedup_update < 1.0]
iter_wins = [r for r in results if r.speedup_iter < 1.0]
if isempty(access_wins)
println("✓ Immutable ALWAYS faster for field access (even at 1000 fields = 8KB)")
min_speedup = minimum(r.speedup_access for r in results)
println(" Minimum speedup: $(round(min_speedup, digits=1))x at $(results[end].nfields) fields")
else
println("⚠ Mutable wins for field access at $(access_wins[1].nfields) fields")
end
println()
if isempty(update_wins)
println("✓ Immutable ALWAYS faster for field update (even at 1000 fields = 8KB)")
min_speedup = minimum(r.speedup_update for r in results)
println(" Minimum speedup: $(round(min_speedup, digits=1))x at $(results[end].nfields) fields")
else
println("⚠ Mutable wins for field update at $(update_wins[1].nfields) fields")
end
println()
if isempty(iter_wins)
println("✓ Immutable ALWAYS faster for iteration (even at 1000 fields = 8KB)")
min_speedup = minimum(r.speedup_iter for r in results)
println(" Minimum speedup: $(round(min_speedup, digits=1))x at $(results[end].nfields) fields")
else
println("⚠ Mutable wins for iteration at $(iter_wins[1].nfields) fields")
end
println()
# Check scaling behavior
println("Scaling Analysis:")
println()
# Linear regression on copy time vs size
sizes = [r.nfields for r in results]
copy_times = [r.imm_copy * 1e9 for r in results] # Convert to ns
# Simple linear fit: time = a + b*size
n = length(sizes)
mean_size = sum(sizes) / n
mean_time = sum(copy_times) / n
cov = sum((sizes[i] - mean_size) * (copy_times[i] - mean_time) for i in 1:n) / n
var_size = sum((s - mean_size)^2 for s in sizes) / n
slope = cov / var_size
intercept = mean_time - slope * mean_size
println("Copy time scaling:")
println(" Linear fit: time(ns) = $(round(intercept, digits=2)) + $(round(slope, digits=4)) * nfields")
println(" Per-field cost: $(round(slope, digits=4)) ns/field")
println(" Base overhead: $(round(intercept, digits=2)) ns")
println()
# Check if copy time grows linearly
println("Is copy time linear? (checking R²)")
ss_tot = sum((t - mean_time)^2 for t in copy_times)
ss_res = sum((copy_times[i] - (intercept + slope * sizes[i]))^2 for i in 1:n)
r_squared = 1 - ss_res / ss_tot
println(" R² = $(round(r_squared, digits=4))")
if r_squared > 0.95
println(" ✓ Copy time is linear in struct size (as expected)")
else
println(" ⚠ Copy time not perfectly linear (compiler optimizations?)")
end
println()
# Key insight
println("KEY INSIGHT:")
println("-"^80)
println()
println("Even at 1000 fields (8KB struct), immutable is STILL faster because:")
println(" 1. Dict lookup cost (~40-50ns) >> copy cost per field (~$(round(slope, digits=4))ns)")
println(" 2. Type stability enables compiler optimizations (inlining, SIMD)")
println(" 3. Stack allocation has better cache locality than heap pointers")
println()
println("Theoretical crossover point (if it exists):")
crossover_fields = (40.0 - intercept) / slope # When copy cost = Dict lookup
println(" Would occur at ~$(round(Int, crossover_fields)) fields ($(round(Int, crossover_fields*8/1024))KB)")
if crossover_fields > 1000
println(" But this is beyond any realistic FEM element!")
end
println()
println("="^80)
println("CONCLUSION")
println("="^80)
println()
println("Your intuition about O(n) scaling is CORRECT, BUT:")
println()
println(" • Dict lookup base cost is SO high (~40-50ns)")
println(" • Copy cost per field is SO low (~$(round(slope, digits=4))ns)")
println(" • Compiler optimizations are SO good (inlining, SIMD, escape analysis)")
println()
println("That immutable wins even for unrealistically large structs (8KB+)!")
println()
println("For typical FEM elements:")
println(" • Material properties: 3-10 fields (24-80 bytes)")
println(" • State variables: 10-50 fields (80-400 bytes)")
println(" • Even with 100 fields (800 bytes), immutable is >10x faster")
println()
println("Type stability > Everything else.")
println()
# ============================================================================
# SAVE DATA TO DISK
# ============================================================================
println("="^80)
println("SAVING DATA")
println("="^80)
println()
# Create results directory
results_dir = joinpath(@__DIR__, "results")
mkpath(results_dir)
# Prepare data for JSON
data_to_save = Dict(
"system_info" => system_info,
"timestamp" => string(now()),
"struct_sizes" => STRUCT_SIZES,
"results" => [
Dict(
"nfields" => r.nfields,
"bytes" => r.bytes,
"mutable_access_ns" => r.mut_access * 1e9,
"immutable_access_ns" => r.imm_access * 1e9,
"speedup_access" => r.speedup_access,
"mutable_update_ns" => r.mut_update * 1e9,
"immutable_update_ns" => r.imm_update * 1e9,
"speedup_update" => r.speedup_update,
"immutable_copy_ns" => r.imm_copy * 1e9,
"mutable_iter_ns" => r.mut_iter * 1e9,
"immutable_iter_ns" => r.imm_iter * 1e9,
"speedup_iter" => r.speedup_iter
)
for r in results
]
)
# Save as JSON
json_file = joinpath(results_dir, "struct_size_scaling.json")
open(json_file, "w") do f
JSON.print(f, data_to_save, 2)
end
println("✓ Data saved to: $json_file")
# Save as CSV for easy plotting in other tools
csv_file = joinpath(results_dir, "struct_size_scaling.csv")
open(csv_file, "w") do f
println(f, "nfields,bytes,mut_access_ns,imm_access_ns,speedup_access,mut_update_ns,imm_update_ns,speedup_update,imm_copy_ns,mut_iter_ns,imm_iter_ns,speedup_iter")
for r in results
println(f, "$(r.nfields),$(r.bytes),$(r.mut_access*1e9),$(r.imm_access*1e9),$(r.speedup_access),$(r.mut_update*1e9),$(r.imm_update*1e9),$(r.speedup_update),$(r.imm_copy*1e9),$(r.mut_iter*1e9),$(r.imm_iter*1e9),$(r.speedup_iter)")
end
end
println("✓ CSV saved to: $csv_file")
println()
# ============================================================================
# GENERATE PLOTS
# ============================================================================
println("="^80)
println("GENERATING PLOTS")
println("="^80)
println()
# Extract data for plotting
bytes_vals = [r.bytes for r in results]
mut_access = [r.mut_access * 1e9 for r in results]
imm_access = [r.imm_access * 1e9 for r in results]
mut_update = [r.mut_update * 1e9 for r in results]
imm_update = [r.imm_update * 1e9 for r in results]
mut_iter = [r.mut_iter * 1e9 for r in results]
imm_iter = [r.imm_iter * 1e9 for r in results]
imm_copy = [r.imm_copy * 1e9 for r in results]
# Typical FEM element sizes
fem_small = 40 # 5 fields (E, ν, ρ, etc.)
fem_medium = 160 # 20 fields (material + state)
fem_large = 400 # 50 fields (complex plasticity)
# Plot 1: Field Access Performance
p1 = plot(bytes_vals, mut_access,
label="Mutable (Dict)",
xlabel="Struct Size (bytes)",
ylabel="Time (nanoseconds)",
title="Field Access Performance vs Struct Size",
linewidth=2,
marker=:circle,
legend=:topleft,
size=(800, 600))
plot!(p1, bytes_vals, imm_access,
label="Immutable (NamedTuple)",
linewidth=2,
marker=:square)
vline!(p1, [fem_small, fem_medium, fem_large],
label="Typical FEM sizes",
linestyle=:dash,
linecolor=:gray,
linewidth=1)
annotate!(p1, fem_small, maximum(mut_access) * 0.9, text("Small\n(5 fields)", 8, :left))
annotate!(p1, fem_medium, maximum(mut_access) * 0.8, text("Medium\n(20 fields)", 8, :left))
annotate!(p1, fem_large, maximum(mut_access) * 0.7, text("Large\n(50 fields)", 8, :left))
plot_file1 = joinpath(results_dir, "field_access_scaling.png")
savefig(p1, plot_file1)
println("✓ Plot saved: $plot_file1")
# Plot 2: Field Update Performance (showing crossover)
p2 = plot(bytes_vals, mut_update,
label="Mutable (Dict)",
xlabel="Struct Size (bytes)",
ylabel="Time (nanoseconds)",
title="Field Update Performance vs Struct Size (Crossover at ~800 bytes)",
linewidth=2,
marker=:circle,
legend=:topleft,
size=(800, 600))
plot!(p2, bytes_vals, imm_update,
label="Immutable (NamedTuple)",
linewidth=2,
marker=:square)
vline!(p2, [fem_small, fem_medium, fem_large, 800],
label=["", "", "", "Crossover (~100 fields)"],
linestyle=[:dash, :dash, :dash, :dot],
linecolor=[:gray, :gray, :gray, :red],
linewidth=[1, 1, 1, 2])
annotate!(p2, fem_small, maximum(imm_update) * 0.2, text("Small", 8, :left))
annotate!(p2, fem_medium, maximum(imm_update) * 0.3, text("Medium", 8, :left))
annotate!(p2, fem_large, maximum(imm_update) * 0.4, text("Large", 8, :left))
plot_file2 = joinpath(results_dir, "field_update_scaling.png")
savefig(p2, plot_file2)
println("✓ Plot saved: $plot_file2")
# Plot 3: Iteration Performance
p3 = plot(bytes_vals, mut_iter,
label="Mutable (Dict)",
xlabel="Struct Size (bytes)",
ylabel="Time (nanoseconds)",
title="Iteration Performance vs Struct Size",
linewidth=2,
marker=:circle,
legend=:topleft,
size=(800, 600),
yscale=:log10)
plot!(p3, bytes_vals, imm_iter,
label="Immutable (NamedTuple)",
linewidth=2,
marker=:square)
vline!(p3, [fem_small, fem_medium, fem_large],
label="Typical FEM sizes",
linestyle=:dash,
linecolor=:gray,
linewidth=1)
plot_file3 = joinpath(results_dir, "iteration_scaling.png")
savefig(p3, plot_file3)
println("✓ Plot saved: $plot_file3")
# Plot 4: Copy Cost (linear scaling)
p4 = plot(bytes_vals, imm_copy,
label="Measured",
xlabel="Struct Size (bytes)",
ylabel="Copy Time (nanoseconds)",
title="Immutable Struct Copy Cost (Linear Scaling)",
linewidth=2,
marker=:circle,
legend=:topright,
size=(800, 600))
# Add linear fit line
plot!(p4, bytes_vals, [intercept + slope * (b / 8) for b in bytes_vals],
label="Linear fit: $(round(intercept, digits=1)) + $(round(slope, digits=3)) × nfields",
linestyle=:dash,
linewidth=2)
vline!(p4, [fem_small, fem_medium, fem_large],
label="Typical FEM sizes",
linestyle=:dash,
linecolor=:gray,
linewidth=1)
plot_file4 = joinpath(results_dir, "copy_cost_linear.png")
savefig(p4, plot_file4)
println("✓ Plot saved: $plot_file4")
# Plot 5: Speedup ratios (showing where immutable wins)
p5 = plot(bytes_vals, [r.speedup_access for r in results],
label="Field Access",
xlabel="Struct Size (bytes)",
ylabel="Speedup (Immutable / Mutable)",
title="Performance Speedup: Immutable vs Mutable",
linewidth=2,
marker=:circle,
legend=:right,
size=(800, 600))
plot!(p5, bytes_vals, [r.speedup_update for r in results],
label="Field Update",
linewidth=2,
marker=:square)
plot!(p5, bytes_vals, [r.speedup_iter for r in results],
label="Iteration",
linewidth=2,
marker=:diamond)
hline!(p5, [1.0],
label="Break-even",
linestyle=:dot,
linecolor=:black,
linewidth=2)
vline!(p5, [fem_small, fem_medium, fem_large],
label="",
linestyle=:dash,
linecolor=:gray,
linewidth=1)
annotate!(p5, fem_large, 0.5, text("Typical FEM range →", 8, :left))
plot_file5 = joinpath(results_dir, "speedup_ratios.png")
savefig(p5, plot_file5)
println("✓ Plot saved: $plot_file5")
println()
println("All plots saved to: $results_dir")
println()
# Save results to JSON
println("="^80)
println("SAVING DATA")
println("="^80)
println()
timestamp = Dates.format(now(), "yyyymmdd_HHMMSS")
output_dir = "benchmarks/results"
mkpath(output_dir)
# Prepare data for saving
benchmark_data = Dict(
"timestamp" => timestamp,
"julia_version" => string(VERSION),
"system" => Dict(
"cpu_model" => cpu_info[1].model,
"cpu_cores" => Sys.CPU_THREADS,
"cpu_speed_mhz" => cpu_info[1].speed,
"os" => string(Sys.KERNEL),
"machine" => string(Sys.MACHINE),
"word_size" => Sys.WORD_SIZE
),
"results" => [
Dict(
"nfields" => r.nfields,
"bytes" => r.bytes,
"mutable_access_ns" => r.mut_access * 1e9,
"immutable_access_ns" => r.imm_access * 1e9,
"speedup_access" => r.speedup_access,
"mutable_update_ns" => r.mut_update * 1e9,
"immutable_update_ns" => r.imm_update * 1e9,
"speedup_update" => r.speedup_update,
"mutable_iter_ns" => r.mut_iter * 1e9,
"immutable_iter_ns" => r.imm_iter * 1e9,
"speedup_iter" => r.speedup_iter,
"immutable_copy_ns" => r.imm_copy * 1e9
) for r in results
],
"analysis" => Dict(
"copy_slope_ns_per_field" => slope,
"copy_intercept_ns" => intercept,
"r_squared" => r_squared
)
)
json_file = joinpath(output_dir, "struct_scaling_$(timestamp).json")
open(json_file, "w") do f
JSON.print(f, benchmark_data, 2)
end
println("✓ Data saved to: $json_file")
println()
# Also save as CSV for easy plotting
csv_file = joinpath(output_dir, "struct_scaling_$(timestamp).csv")
open(csv_file, "w") do f
println(f, "nfields,bytes,mut_access_ns,imm_access_ns,speedup_access,mut_update_ns,imm_update_ns,speedup_update,mut_iter_ns,imm_iter_ns,speedup_iter,imm_copy_ns")
for r in results
println(f, "$(r.nfields),$(r.bytes),$(r.mut_access*1e9),$(r.imm_access*1e9),$(r.speedup_access),$(r.mut_update*1e9),$(r.imm_update*1e9),$(r.speedup_update),$(r.mut_iter*1e9),$(r.imm_iter*1e9),$(r.speedup_iter),$(r.imm_copy*1e9)")
end
end
println("✓ CSV saved to: $csv_file")
println()
println("="^80)
println("GENERATING PLOTS")
println("="^80)
println()
# Note: Using Plots from global environment
try
# Import from global environment
pushfirst!(LOAD_PATH, "@stdlib")
import Plots
# Set backend
Plots.gr()
# Extract data for plotting
bytes_sizes = [r.bytes for r in results]
# Plot 1: Field Access Performance
p1 = Plots.plot(bytes_sizes, [r.mut_access * 1e9 for r in results],
label="Mutable (Dict)", linewidth=2, marker=:circle,
xlabel="Struct Size (bytes)", ylabel="Time (nanoseconds)",
title="Field Access Performance",
legend=:topleft, xscale=:log10, grid=true)
Plots.plot!(p1, bytes_sizes, [r.imm_access * 1e9 for r in results],
label="Immutable (NamedTuple)", linewidth=2, marker=:square)
# Add typical FEM element size markers
Plots.vline!(p1, [40, 400], label="Typical FEM (5-50 fields)",
linestyle=:dash, linewidth=1, color=:gray)
Plots.savefig(p1, joinpath(output_dir, "field_access_$(timestamp).png"))
println("✓ Saved: field_access_$(timestamp).png")
# Plot 2: Field Update Performance
p2 = Plots.plot(bytes_sizes, [r.mut_update * 1e9 for r in results],
label="Mutable (Dict)", linewidth=2, marker=:circle,
xlabel="Struct Size (bytes)", ylabel="Time (nanoseconds)",
title="Field Update Performance",
legend=:topleft, xscale=:log10, grid=true)
Plots.plot!(p2, bytes_sizes, [r.imm_update * 1e9 for r in results],
label="Immutable (NamedTuple)", linewidth=2, marker=:square)
Plots.vline!(p2, [40, 400], label="Typical FEM (5-50 fields)",
linestyle=:dash, linewidth=1, color=:gray)
Plots.savefig(p2, joinpath(output_dir, "field_update_$(timestamp).png"))
println("✓ Saved: field_update_$(timestamp).png")
# Plot 3: Iteration Performance
p3 = Plots.plot(bytes_sizes, [r.mut_iter * 1e9 for r in results],
label="Mutable (Dict)", linewidth=2, marker=:circle,
xlabel="Struct Size (bytes)", ylabel="Time (nanoseconds)",
title="Field Iteration Performance",
legend=:topleft, xscale=:log10, yscale=:log10, grid=true)
Plots.plot!(p3, bytes_sizes, [r.imm_iter * 1e9 for r in results],
label="Immutable (NamedTuple)", linewidth=2, marker=:square)
Plots.vline!(p3, [40, 400], label="Typical FEM (5-50 fields)",
linestyle=:dash, linewidth=1, color=:gray)
Plots.savefig(p3, joinpath(output_dir, "iteration_$(timestamp).png"))
println("✓ Saved: iteration_$(timestamp).png")
# Plot 4: Speedup Factors
p4 = Plots.plot(bytes_sizes, [r.speedup_access for r in results],
label="Access Speedup", linewidth=2, marker=:circle,
xlabel="Struct Size (bytes)", ylabel="Speedup Factor (Immutable/Mutable)",
title="Performance Advantage of Immutable Elements",
legend=:right, xscale=:log10, grid=true)
Plots.plot!(p4, bytes_sizes, [r.speedup_update for r in results],
label="Update Speedup", linewidth=2, marker=:square)
Plots.plot!(p4, bytes_sizes, [r.speedup_iter for r in results],
label="Iteration Speedup", linewidth=2, marker=:diamond)
Plots.hline!(p4, [1.0], label="Break-even", linestyle=:dash, color=:black, linewidth=1)
Plots.vline!(p4, [40, 400], label="Typical FEM",
linestyle=:dash, linewidth=1, color=:gray)
Plots.savefig(p4, joinpath(output_dir, "speedup_factors_$(timestamp).png"))
println("✓ Saved: speedup_factors_$(timestamp).png")
# Plot 5: Copy Cost Scaling
p5 = Plots.plot(bytes_sizes, [r.imm_copy * 1e9 for r in results],
label="Measured", linewidth=2, marker=:circle,
xlabel="Struct Size (bytes)", ylabel="Copy Time (nanoseconds)",
title="Immutable Struct Copy Cost",
legend=:topleft, xscale=:log10, grid=true)
# Add linear fit
fitted = [intercept + slope * r.nfields for r in results]
Plots.plot!(p5, bytes_sizes, fitted,
label="Linear Fit ($(round(slope, digits=4)) ns/field)",
linewidth=2, linestyle=:dash)
Plots.vline!(p5, [40, 400], label="Typical FEM",
linestyle=:dash, linewidth=1, color=:gray)
Plots.savefig(p5, joinpath(output_dir, "copy_cost_$(timestamp).png"))
println("✓ Saved: copy_cost_$(timestamp).png")
# Combined plot
layout = Plots.@layout [a b; c d]
p_combined = Plots.plot(p1, p2, p3, p4, layout=layout, size=(1200, 900))
Plots.savefig(p_combined, joinpath(output_dir, "combined_$(timestamp).png"))
println("✓ Saved: combined_$(timestamp).png")
println()
println("All plots saved successfully!")
catch e
println("⚠ Could not generate plots (Plots.jl not available in global environment)")
println(" Error: $e")
println(" Install with: julia -e 'using Pkg; Pkg.add(\"Plots\")'")
end
println()
println("="^80)
-159
View File
@@ -1,159 +0,0 @@
#!/bin/bash
# Quick test script for GPU benchmarks
echo "======================================================================"
echo "GPU Benchmark Quick Test"
echo "======================================================================"
echo ""
# Check Julia
echo "Checking Julia installation..."
if ! command -v julia &> /dev/null; then
echo "❌ Julia not found! Please install Julia 1.9+"
exit 1
fi
julia_version=$(julia --version)
echo "✅ Found: $julia_version"
echo ""
# Check GPU
echo "Checking GPU availability..."
if command -v nvidia-smi &> /dev/null; then
echo "✅ NVIDIA GPU detected:"
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader
echo ""
else
echo "⚠️ No NVIDIA GPU detected. Benchmarks will run CPU-only."
echo ""
fi
# Check CUDA.jl
echo "Checking CUDA.jl..."
julia --project=. -e '
using Pkg
try
using CUDA
if CUDA.functional()
println("✅ CUDA.jl functional: ", CUDA.name(CUDA.device()))
else
println("⚠️ CUDA.jl installed but GPU not functional")
end
catch
println("⚠️ CUDA.jl not installed. Run: Pkg.add(\"CUDA\")")
end
' 2>/dev/null
echo ""
# Run quick state management test (small size)
echo "======================================================================"
echo "Test 1: State Management (10K elements)"
echo "======================================================================"
julia --project=. -e '
n = 10_000
println("Running state management benchmark with $n elements...")
include("benchmarks/gpu_state_management_benchmark.jl")
# Override main() to run smaller test
Δε_p, Δα, elements_s1, geometry_s2, state_s2 = setup_benchmark(n)
# CPU test
println("\n📊 CPU Test:")
state_s2_copy = deepcopy(state_s2)
t = @elapsed update_state_strategy2_cpu!(state_s2_copy, Δε_p, Δα)
println("Time: $(round(t * 1000, digits=2)) ms")
println("✅ CPU benchmark works!")
# GPU test (if available)
if USE_GPU
println("\n📊 GPU Test:")
try
T = Float64
state_gpu = AssemblyState{T}(
CUDA.zeros(T, n * 6),
CUDA.zeros(T, n),
n
)
Δε_p_flat = zeros(T, n * 6)
Δα_gpu = CuArray(Δα)
for i in 1:n
offset = (i - 1) * 6
ε = Δε_p[i]
Δε_p_flat[offset + 1] = ε[1, 1]
Δε_p_flat[offset + 2] = ε[2, 2]
Δε_p_flat[offset + 3] = ε[3, 3]
Δε_p_flat[offset + 4] = ε[1, 2]
Δε_p_flat[offset + 5] = ε[1, 3]
Δε_p_flat[offset + 6] = ε[2, 3]
end
Δε_p_flat_gpu = CuArray(Δε_p_flat)
update_state_strategy2_gpu!(state_gpu, Δε_p_flat_gpu, Δα_gpu)
println("✅ GPU benchmark works!")
catch e
println("⚠️ GPU test failed: $e")
end
end
'
echo ""
# Run quick matrix-free test (small size)
echo "======================================================================"
echo "Test 2: Matrix-Free Newton-Krylov (1K DOFs)"
echo "======================================================================"
julia --project=. -e '
n = 1000
println("Running matrix-free benchmark with $n DOFs...")
include("benchmarks/matrix_free_gpu_benchmark.jl")
# Override to run small test
T = Float64
K = Matrix(Tridiagonal(-ones(T, n-1), 2ones(T, n), -ones(T, n-1)))
f = ones(T, n) * 0.1
β = T(1e-3)
prob = NonlinearProblem(K, f, β, n)
println("\n📊 CPU Test:")
u = zeros(T, n)
r = zeros(T, n)
du = zeros(T, n)
temp = zeros(T, n)
Jv = zeros(T, n)
t = @elapsed iters = newton_matrix_free!(u, prob, r, du, temp, Jv;
max_iter=10, verbose=false)
println("Time: $(round(t * 1000, digits=2)) ms")
println("Iterations: $iters")
println("✅ CPU benchmark works!")
if USE_GPU
println("\n📊 GPU Test:")
try
K_gpu = CuArray(K)
f_gpu = CuArray(f)
u_gpu = CUDA.zeros(T, n)
t_gpu = CUDA.@elapsed begin
iters_gpu = newton_matrix_free_gpu!(u_gpu, K_gpu, f_gpu, β;
max_iter=10, verbose=false)
CUDA.synchronize()
end
println("Time: $(round(t_gpu * 1000, digits=2)) ms")
println("Iterations: $iters_gpu")
println("✅ GPU benchmark works!")
catch e
println("⚠️ GPU test failed: $e")
end
end
'
echo ""
echo "======================================================================"
echo "Quick Test Complete!"
echo "======================================================================"
echo ""
echo "To run full benchmarks:"
echo " julia --project=. benchmarks/gpu_state_management_benchmark.jl"
echo " julia --project=. benchmarks/matrix_free_gpu_benchmark.jl"
echo ""
-284
View File
@@ -1,284 +0,0 @@
# Benchmark: Tet10 Shape Function Derivatives - Manual vs AD
#
# Compares two approaches:
# 1. Manual: Hand-calculated derivatives (traditional FEM)
# 2. AD: Automatic differentiation using Tensors.jl gradient()
#
# Run with: julia --project=. benchmarks/tet10_derivatives_benchmark.jl
using BenchmarkTools
using Tensors
using Printf
println("="^70)
println("Tet10 Shape Function Derivatives: Manual vs AD Benchmark")
println("="^70)
println()
# ============================================================================
# METHOD 1: MANUAL (Hand-Calculated Derivatives)
# ============================================================================
"""
Tet10 shape functions (manual implementation).
Reference element: ξ [0,1], η [0,1], ζ [0,1], ξ+η+ζ 1
"""
module ManualTet10
using Tensors
# Shape functions
@inline N1(ξ, η, ζ) = (1 - ξ - η - ζ) * (2 * (1 - ξ - η - ζ) - 1)
@inline N2(ξ, η, ζ) = ξ * (2 * ξ - 1)
@inline N3(ξ, η, ζ) = η * (2 * η - 1)
@inline N4(ξ, η, ζ) = ζ * (2 * ζ - 1)
@inline N5(ξ, η, ζ) = 4 * ξ * (1 - ξ - η - ζ)
@inline N6(ξ, η, ζ) = 4 * ξ * η
@inline N7(ξ, η, ζ) = 4 * η * (1 - ξ - η - ζ)
@inline N8(ξ, η, ζ) = 4 * ζ * (1 - ξ - η - ζ)
@inline N9(ξ, η, ζ) = 4 * ξ * ζ
@inline N10(ξ, η, ζ) = 4 * η * ζ
# Derivatives (calculated by hand - error-prone!)
@inline dN1_dξ(ξ, η, ζ) = 4 * ξ + 4 * η + 4 * ζ - 3
@inline dN1_dη(ξ, η, ζ) = 4 * ξ + 4 * η + 4 * ζ - 3
@inline dN1_dζ(ξ, η, ζ) = 4 * ξ + 4 * η + 4 * ζ - 3
@inline dN2_dξ(ξ, η, ζ) = 4 * ξ - 1
@inline dN2_dη(ξ, η, ζ) = 0.0
@inline dN2_dζ(ξ, η, ζ) = 0.0
@inline dN3_dξ(ξ, η, ζ) = 0.0
@inline dN3_dη(ξ, η, ζ) = 4 * η - 1
@inline dN3_dζ(ξ, η, ζ) = 0.0
@inline dN4_dξ(ξ, η, ζ) = 0.0
@inline dN4_dη(ξ, η, ζ) = 0.0
@inline dN4_dζ(ξ, η, ζ) = 4 * ζ - 1
@inline dN5_dξ(ξ, η, ζ) = 4 * (1 - 2 * ξ - η - ζ)
@inline dN5_dη(ξ, η, ζ) = -4 * ξ
@inline dN5_dζ(ξ, η, ζ) = -4 * ξ
@inline dN6_dξ(ξ, η, ζ) = 4 * η
@inline dN6_dη(ξ, η, ζ) = 4 * ξ
@inline dN6_dζ(ξ, η, ζ) = 0.0
@inline dN7_dξ(ξ, η, ζ) = -4 * η
@inline dN7_dη(ξ, η, ζ) = 4 * (1 - ξ - 2 * η - ζ)
@inline dN7_dζ(ξ, η, ζ) = -4 * η
@inline dN8_dξ(ξ, η, ζ) = -4 * ζ
@inline dN8_dη(ξ, η, ζ) = -4 * ζ
@inline dN8_dζ(ξ, η, ζ) = 4 * (1 - ξ - η - 2 * ζ)
@inline dN9_dξ(ξ, η, ζ) = 4 * ζ
@inline dN9_dη(ξ, η, ζ) = 0.0
@inline dN9_dζ(ξ, η, ζ) = 4 * ξ
@inline dN10_dξ(ξ, η, ζ) = 0.0
@inline dN10_dη(ξ, η, ζ) = 4 * ζ
@inline dN10_dζ(ξ, η, ζ) = 4 * η
# Evaluation function (returns tuple - zero allocation)
@inline function eval_basis_and_grad(xi::Vec{3})
ξ, η, ζ = xi[1], xi[2], xi[3]
N = (N1(ξ, η, ζ), N2(ξ, η, ζ), N3(ξ, η, ζ), N4(ξ, η, ζ), N5(ξ, η, ζ),
N6(ξ, η, ζ), N7(ξ, η, ζ), N8(ξ, η, ζ), N9(ξ, η, ζ), N10(ξ, η, ζ))
dN = (Vec(dN1_dξ(ξ, η, ζ), dN1_dη(ξ, η, ζ), dN1_dζ(ξ, η, ζ)),
Vec(dN2_dξ(ξ, η, ζ), dN2_dη(ξ, η, ζ), dN2_dζ(ξ, η, ζ)),
Vec(dN3_dξ(ξ, η, ζ), dN3_dη(ξ, η, ζ), dN3_dζ(ξ, η, ζ)),
Vec(dN4_dξ(ξ, η, ζ), dN4_dη(ξ, η, ζ), dN4_dζ(ξ, η, ζ)),
Vec(dN5_dξ(ξ, η, ζ), dN5_dη(ξ, η, ζ), dN5_dζ(ξ, η, ζ)),
Vec(dN6_dξ(ξ, η, ζ), dN6_dη(ξ, η, ζ), dN6_dζ(ξ, η, ζ)),
Vec(dN7_dξ(ξ, η, ζ), dN7_dη(ξ, η, ζ), dN7_dζ(ξ, η, ζ)),
Vec(dN8_dξ(ξ, η, ζ), dN8_dη(ξ, η, ζ), dN8_dζ(ξ, η, ζ)),
Vec(dN9_dξ(ξ, η, ζ), dN9_dη(ξ, η, ζ), dN9_dζ(ξ, η, ζ)),
Vec(dN10_dξ(ξ, η, ζ), dN10_dη(ξ, η, ζ), dN10_dζ(ξ, η, ζ)))
return N, dN
end
end
# ============================================================================
# ============================================================================
# METHOD 2: AD (Tensors.jl gradient)
# ============================================================================
module ADTet10
using Tensors
# Just shape functions (no manual derivatives!)
@inline N1(xi) = (1 - xi[1] - xi[2] - xi[3]) * (2 * (1 - xi[1] - xi[2] - xi[3]) - 1)
@inline N2(xi) = xi[1] * (2 * xi[1] - 1)
@inline N3(xi) = xi[2] * (2 * xi[2] - 1)
@inline N4(xi) = xi[3] * (2 * xi[3] - 1)
@inline N5(xi) = 4 * xi[1] * (1 - xi[1] - xi[2] - xi[3])
@inline N6(xi) = 4 * xi[1] * xi[2]
@inline N7(xi) = 4 * xi[2] * (1 - xi[1] - xi[2] - xi[3])
@inline N8(xi) = 4 * xi[3] * (1 - xi[1] - xi[2] - xi[3])
@inline N9(xi) = 4 * xi[1] * xi[3]
@inline N10(xi) = 4 * xi[2] * xi[3]
const shape_fns = (N1, N2, N3, N4, N5, N6, N7, N8, N9, N10)
@inline function eval_basis_and_grad(xi::Vec{3})
# Evaluate basis functions
N = ntuple(i -> shape_fns[i](xi), 10)
# Compute gradients with Tensors.jl gradient()
dN = ntuple(i -> gradient(shape_fns[i], xi), 10)
return N, dN
end
end
# ============================================================================
# BENCHMARKING
# ============================================================================
println("Setting up benchmark...")
println()
# Test point (typical integration point)
const ξ_test = Vec(0.25, 0.25, 0.25)
# Verification: Both methods should give same results
println("Verifying correctness...")
N_manual, dN_manual = ManualTet10.eval_basis_and_grad(ξ_test)
N_ad, dN_ad = ADTet10.eval_basis_and_grad(ξ_test)
println(" Manual basis: ", N_manual)
println(" AD basis: ", N_ad)
println()
# Check agreement
rtol = 1e-10
if !all(isapprox.(N_manual, N_ad, rtol=rtol))
@warn "Manual and AD basis functions disagree!"
end
# Check derivatives
for i in 1:10
if !isapprox(dN_manual[i], dN_ad[i], rtol=rtol)
@warn "Manual and AD derivative $i disagree!" dN_manual[i] dN_ad[i]
end
end
println("✓ Both methods agree (within tolerance)")
println()
# ============================================================================
# RUN BENCHMARKS
# ============================================================================
println("Running benchmarks (this may take a minute)...")
println()
# Warm-up
for _ in 1:1000
ManualTet10.eval_basis_and_grad(ξ_test)
ADTet10.eval_basis_and_grad(ξ_test)
end
# Benchmark each method
b_manual = @benchmark ManualTet10.eval_basis_and_grad($ξ_test)
b_ad = @benchmark ADTet10.eval_basis_and_grad($ξ_test)
# ============================================================================
# RESULTS
# ============================================================================
# ============================================================================
# RESULTS
# ============================================================================
println("="^70)
println("RESULTS")
println("="^70)
println()
# Extract median times
t_manual = median(b_manual.times)
t_ad = median(b_ad.times)
# Extract allocations
alloc_manual = b_manual.allocs
alloc_ad = b_ad.allocs
# Calculate relative speed
rel_ad = t_ad / t_manual
println("Method | Time (ns) | Allocations | Relative Speed")
println("----------------|-----------|-------------|----------------")
@printf "Manual | %9.1f | %11d | %.2f× (baseline)\n" t_manual alloc_manual 1.0
@printf "AD (Tensors.jl) | %9.1f | %11d | %.2f×\n" t_ad alloc_ad rel_ad
println()
# Detailed stats
println("Detailed Statistics:")
println()
println("Manual (Hand-Calculated):")
display(b_manual)
println()
println()
println("AD (Tensors.jl gradient):")
display(b_ad)
println()
println()
# ============================================================================
# ANALYSIS
# ============================================================================
println("="^70)
println("ANALYSIS")
println("="^70)
println()
if rel_ad < 2.0
println("🎉 RECOMMENDATION: Use AD everywhere!")
println()
println("Tensors.jl AD is within 2× of manual, providing:")
println(" ✓ Zero maintenance burden")
println(" ✓ No manual derivative errors")
println(" ✓ Easy to add new elements")
println(" ✓ Supports any basis type")
println()
println("Small performance cost is acceptable for these benefits.")
elseif rel_ad < 5.0
println("⚠️ RECOMMENDATION: Hybrid approach")
println()
println("AD is 2-5× slower than manual. Consider:")
println(" • Common elements (Tet10, Hex8, Quad4): Manual")
println(" • Rare elements: AD-generated")
println(" • Research/prototype elements: Always AD")
println()
println("This balances performance and maintainability.")
else
println("❌ RECOMMENDATION: Manual derivatives (with symbolic generation)")
println()
println("AD is >5× slower than manual. For performance-critical code:")
println(" • Generate derivatives with SymPy/Symbolics.jl")
println(" • Unit test against AD to verify correctness")
println(" • Accept the maintenance burden")
println()
println("Consider AD only for prototyping.")
end
println()
println("Memory analysis:")
if alloc_manual == 0 && alloc_ad == 0
println(" ✓ Both methods achieve zero allocations (excellent!)")
elseif alloc_manual == 0 && alloc_ad > 0
println(" ⚠️ AD allocates (", alloc_ad, " allocs)")
println(" This will hurt performance in tight loops.")
else
println(" ⚠️ Unexpected allocation pattern - investigate!")
end
println()
println("="^70)
println("Benchmark complete! Results saved to console.")
println("="^70)