chore: remove benchmarks directory
Remove benchmarks directory as it has been moved to another repository - Remove benchmarks/ directory and all contents - 37 files removed including benchmark scripts and results
@@ -1,244 +0,0 @@
|
||||
# GPU Benchmark Suite
|
||||
|
||||
Comprehensive benchmarks demonstrating GPU-friendly FEM architecture strategies.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
```bash
|
||||
# Add required packages
|
||||
julia --project=. -e 'using Pkg; Pkg.add(["CUDA", "Tensors", "BenchmarkTools", "IterativeSolvers"])'
|
||||
```
|
||||
|
||||
## Benchmarks
|
||||
|
||||
### 1. State Management Strategy Comparison
|
||||
|
||||
**File:** `gpu_state_management_benchmark.jl`
|
||||
|
||||
**What it tests:**
|
||||
|
||||
- **Strategy 1:** Immutable elements (Array of Structs - AoS)
|
||||
- **Strategy 2:** Separate mutable state (Structure of Arrays - SoA)
|
||||
|
||||
**Metrics:**
|
||||
|
||||
- Memory bandwidth (GB/s)
|
||||
- Execution time
|
||||
- Allocation counts
|
||||
|
||||
**Run:**
|
||||
|
||||
```bash
|
||||
julia --project=. benchmarks/gpu_state_management_benchmark.jl
|
||||
```
|
||||
|
||||
**Expected Results:**
|
||||
|
||||
- CPU: Strategy 2 is 3-5× faster (better cache utilization)
|
||||
- GPU: Strategy 2 is 5-10× faster (coalesced memory access)
|
||||
- Strategy 2 bandwidth: 500-900 GB/s (depending on GPU)
|
||||
- Strategy 1 bandwidth: 50-150 GB/s (non-coalesced)
|
||||
|
||||
### 2. Matrix-Free Newton-Krylov
|
||||
|
||||
**File:** `matrix_free_gpu_benchmark.jl`
|
||||
|
||||
**What it tests:**
|
||||
|
||||
- **Traditional Newton:** Full Jacobian assembly + direct solve
|
||||
- **Matrix-Free NK:** Jacobian-free with GMRES
|
||||
- **Anderson-Accelerated:** Matrix-Free + Anderson acceleration
|
||||
|
||||
**Metrics:**
|
||||
|
||||
- Total time
|
||||
- Iterations to convergence
|
||||
- Time per iteration
|
||||
|
||||
**Run:**
|
||||
|
||||
```bash
|
||||
julia --project=. benchmarks/matrix_free_gpu_benchmark.jl
|
||||
```
|
||||
|
||||
**Expected Results:**
|
||||
|
||||
- Matrix-Free: 2-4× faster than traditional (no assembly)
|
||||
- Anderson: 2-3× fewer iterations (superlinear convergence)
|
||||
- GPU: Additional 3-10× speedup for large problems (>10K DOFs)
|
||||
- **Total:** 5-10× speedup with all optimizations
|
||||
|
||||
## Understanding the Results
|
||||
|
||||
### State Management Benchmark Output
|
||||
|
||||
```text
|
||||
GPU Benchmark: 1000000 elements
|
||||
====================================================================
|
||||
|
||||
📊 Strategy 1 (Immutable Elements - AoS on GPU):
|
||||
Time: 12.456 ms
|
||||
Bandwidth: 87.3 GB/s
|
||||
|
||||
📊 Strategy 2 (Separate State - SoA on GPU):
|
||||
Time: 1.234 ms
|
||||
Bandwidth: 743.2 GB/s
|
||||
|
||||
✅ GPU Speedup (Strategy 2 / Strategy 1): 10.09×
|
||||
✅ Bandwidth Improvement: 8.51×
|
||||
Strategy 1: 87.3 GB/s (non-coalesced)
|
||||
Strategy 2: 743.2 GB/s (coalesced)
|
||||
```
|
||||
|
||||
**Interpretation:**
|
||||
|
||||
- Strategy 2 achieves 8-10× higher bandwidth
|
||||
- Approaching GPU memory bandwidth limit (~1000 GB/s for high-end GPUs)
|
||||
- Coalesced memory access is critical for GPU performance
|
||||
|
||||
### Matrix-Free Benchmark Output
|
||||
|
||||
```text
|
||||
CPU Benchmark: 10000 DOFs
|
||||
====================================================================
|
||||
|
||||
📊 Traditional Newton (Full Jacobian):
|
||||
Time: 456.78 ms
|
||||
Iterations: 8
|
||||
Time/iter: 57.10 ms
|
||||
|
||||
📊 Matrix-Free Newton-Krylov:
|
||||
Time: 123.45 ms
|
||||
Iterations: 10
|
||||
Time/iter: 12.35 ms
|
||||
|
||||
📊 Anderson-Accelerated Matrix-Free:
|
||||
Time: 67.89 ms
|
||||
Iterations: 5
|
||||
Time/iter: 13.58 ms
|
||||
|
||||
✅ CPU Speedups:
|
||||
Matrix-Free vs Traditional: 3.70×
|
||||
Anderson vs Traditional: 6.73×
|
||||
Anderson vs Matrix-Free: 1.82×
|
||||
```
|
||||
|
||||
**Interpretation:**
|
||||
|
||||
- Matrix-Free eliminates expensive Jacobian assembly
|
||||
- Anderson reduces Newton iterations (superlinear convergence)
|
||||
- Combined: 6-10× speedup on CPU, more on GPU
|
||||
|
||||
## Hardware Requirements
|
||||
|
||||
### Minimum
|
||||
|
||||
- CPU: x86_64 with AVX2
|
||||
- RAM: 8 GB
|
||||
- Julia: 1.9+
|
||||
|
||||
### Recommended for GPU Benchmarks
|
||||
|
||||
- GPU: NVIDIA with CUDA Compute Capability 7.0+ (RTX 20XX or newer)
|
||||
- VRAM: 4 GB+
|
||||
- CUDA: 11.0+
|
||||
|
||||
### Tested On
|
||||
|
||||
- NVIDIA RTX 4090 (24 GB VRAM)
|
||||
- NVIDIA RTX 3090 (24 GB VRAM)
|
||||
- NVIDIA RTX 3080 (10 GB VRAM)
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### "CUDA not available"
|
||||
|
||||
If you see this warning, the benchmark will run CPU-only comparison:
|
||||
|
||||
```julia
|
||||
⚠️ CUDA not available! Running CPU-only comparison.
|
||||
```
|
||||
|
||||
**Solution:**
|
||||
|
||||
1. Check GPU is recognized: `nvidia-smi`
|
||||
2. Verify CUDA installation: `julia -e 'using CUDA; CUDA.versioninfo()'`
|
||||
3. Rebuild CUDA.jl: `julia --project=. -e 'using Pkg; Pkg.build("CUDA")'`
|
||||
|
||||
### Out of Memory Errors
|
||||
|
||||
If benchmarks crash with OOM:
|
||||
|
||||
```julia
|
||||
ERROR: Out of memory
|
||||
```
|
||||
|
||||
**Solution:**
|
||||
|
||||
- Reduce problem sizes in benchmark scripts
|
||||
- Edit `sizes = [10_000, 100_000, 1_000_000]` → smaller values
|
||||
- Close other GPU applications
|
||||
|
||||
### Slow CPU Benchmarks
|
||||
|
||||
Matrix-Free benchmark may be slow on CPU for large problems (>100K DOFs).
|
||||
|
||||
**Solution:**
|
||||
|
||||
- Use smaller test sizes for CPU
|
||||
- Focus on GPU results for large problems
|
||||
- Enable BLAS threading: `export JULIA_NUM_THREADS=8`
|
||||
|
||||
## Performance Expectations
|
||||
|
||||
### State Management (1M elements)
|
||||
|
||||
| Strategy | CPU Time | GPU Time | GPU Bandwidth |
|
||||
|----------|----------|----------|---------------|
|
||||
| Strategy 1 (AoS) | 18 ms | 12 ms | 80-150 GB/s |
|
||||
| Strategy 2 (SoA) | 5 ms | 1.2 ms | 500-900 GB/s |
|
||||
| **Speedup** | **3.6×** | **10×** | **6-10×** |
|
||||
|
||||
### Matrix-Free Newton-Krylov (100K DOFs)
|
||||
|
||||
| Method | CPU Time | GPU Time | Iterations |
|
||||
|--------|----------|----------|------------|
|
||||
| Traditional Newton | 8200 ms | N/A | 20 |
|
||||
| Matrix-Free NK | 2100 ms | 450 ms | 20 |
|
||||
| MF-NK + Anderson | 1050 ms | 180 ms | 8 |
|
||||
| **Speedup** | **7.8×** | **45×** | **2.5× fewer** |
|
||||
|
||||
## Validation
|
||||
|
||||
Both benchmarks validate correctness:
|
||||
|
||||
1. **State Management:**
|
||||
- Verifies final state matches between strategies
|
||||
- Checks zero allocations for Strategy 2
|
||||
|
||||
2. **Matrix-Free:**
|
||||
- Compares solution accuracy (||u_traditional - u_matrixfree|| < 1e-6)
|
||||
- Validates convergence to same residual norm
|
||||
|
||||
## Citation
|
||||
|
||||
If you use these benchmarks in research, please cite:
|
||||
|
||||
```bibtex
|
||||
@software{juliafem2025,
|
||||
title = {JuliaFEM: GPU-Accelerated Finite Element Method},
|
||||
author = {Aho, Jukka},
|
||||
year = {2025},
|
||||
url = {https://github.com/JuliaFEM/JuliaFEM.jl}
|
||||
}
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
1. **Knoll & Keyes (2004):** "Jacobian-free Newton–Krylov methods"
|
||||
2. **Walker & Ni (2011):** "Anderson acceleration for fixed-point iterations"
|
||||
3. **CUDA Programming Guide:** <https://docs.nvidia.com/cuda/>
|
||||
|
||||
## Contact
|
||||
|
||||
Questions or issues? Open an issue at: <https://github.com/JuliaFEM/JuliaFEM.jl/issues>
|
||||
@@ -1,85 +0,0 @@
|
||||
# Benchmark Validation Results
|
||||
|
||||
**Date:** November 9, 2025
|
||||
**Platform:** Julia 1.12.1
|
||||
**Document:** `docs/book/zero_allocation_fields.md`
|
||||
**Benchmark:** `benchmarks/field_storage_comparison.jl`
|
||||
|
||||
## Summary
|
||||
|
||||
✅ **All performance claims validated**
|
||||
|
||||
The zero-allocation field storage design achieves **9-92× speedup** over `Dict{String,Any}` with **zero allocations in hot paths**.
|
||||
|
||||
## Measured Results
|
||||
|
||||
| Test | OLD (Dict) | NEW (Typed) | Speedup |
|
||||
|------|------------|-------------|---------|
|
||||
| Constant field access | 19.2ns, 0 allocs | 2.1ns, 0 allocs | **9×** |
|
||||
| Nodal field access | 262ns, 3 allocs | 6.5ns, 0 allocs | **40×** |
|
||||
| Interpolation (uncached) | 2.6μs, 50 allocs | 44ns, 2 allocs | **59×** |
|
||||
| Interpolation (cached) | 2.6μs, 50 allocs | 53ns, **0 allocs** ✅ | **49×** |
|
||||
| Assembly (1000 elem) | 109μs, 4000 allocs | 1.2μs, **0 allocs** ✅ | **92×** |
|
||||
|
||||
## Key Achievements
|
||||
|
||||
1. ✅ **Zero allocations** in cached interpolation (53ns)
|
||||
2. ✅ **Zero allocations** in assembly loop (1.2μs vs 109μs)
|
||||
3. ✅ **Type stability** eliminates runtime dispatch
|
||||
4. ✅ **9-92× speedup** across all operations
|
||||
5. ✅ **Simple implementation** (~200 LOC for field types)
|
||||
|
||||
## Design Validated
|
||||
|
||||
The `NamedTuple` + typed field structs approach is proven effective:
|
||||
|
||||
```julia
|
||||
# Simple field types
|
||||
struct ConstantField{T}
|
||||
value::T
|
||||
end
|
||||
|
||||
struct NodalField{T}
|
||||
values::Matrix{T}
|
||||
end
|
||||
|
||||
# Type-stable container
|
||||
fields = (
|
||||
youngs_modulus = ConstantField(210e3),
|
||||
displacement = NodalField(zeros(3, 1000)),
|
||||
)
|
||||
|
||||
# Fast access (zero allocations)
|
||||
E = fields.youngs_modulus.value # 2.1ns, 0 allocs
|
||||
u = @view fields.displacement.values[:, nodes] # 6.5ns, 0 allocs
|
||||
```
|
||||
|
||||
## Claims Verification
|
||||
|
||||
| Claim | Measured | Status |
|
||||
|-------|----------|--------|
|
||||
| 50× faster | 9-92× across operations | ✅ VALIDATED |
|
||||
| 0 allocations | 0 allocs in hot paths | ✅ VALIDATED |
|
||||
| Type stability | No runtime dispatch | ✅ VALIDATED |
|
||||
| Simple implementation | ~200 LOC field types | ✅ VALIDATED |
|
||||
|
||||
## Reproduction
|
||||
|
||||
```bash
|
||||
cd /home/juajukka/dev/JuliaFEM.jl
|
||||
julia --project=. benchmarks/field_storage_comparison.jl
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. ✅ Document written and validated
|
||||
2. ⏭️ Implement field types in `src/fields/types.jl`
|
||||
3. ⏭️ Update `Element` struct for `ElementSet` pattern
|
||||
4. ⏭️ Add CI benchmarks to prevent regression
|
||||
5. ⏭️ Migrate examples to new field system
|
||||
|
||||
## Conclusion
|
||||
|
||||
The zero-allocation field storage design is **ready for v1.0 implementation**. Measured performance exceeds targets with 9-92× speedup and zero allocations in hot paths.
|
||||
|
||||
**Design Decision:** Use `NamedTuple` of typed field structs for JuliaFEM v1.0
|
||||
@@ -1,412 +0,0 @@
|
||||
# Benchmark: Basis Function Access Patterns for Tet10 (Realistic 3D Case)
|
||||
#
|
||||
# Focus: 10-node quadratic tetrahedron (Tet10) - the workhorse for 3D simulations
|
||||
#
|
||||
# Goal: Find the fastest way to access basis functions and their DERIVATIVES with:
|
||||
# 1. Return all 10 basis functions as tuple (zero allocation)
|
||||
# 2. Return single basis function by index (must be inlineable)
|
||||
# 3. Return all 10 derivatives as tuple of Vec{3} (zero allocation)
|
||||
# 4. Return single derivative by index (must be inlineable)
|
||||
# 5. Pass topology separately (separation of concerns)
|
||||
# 6. Must be type-stable and superfast
|
||||
#
|
||||
# Use Case:
|
||||
# - Stiffness matrix assembly: Need derivatives (B matrix construction)
|
||||
# - Mass matrix assembly: Need basis functions (M matrix construction)
|
||||
# - Nodal assembly: Need single basis function/derivative at a time
|
||||
#
|
||||
# Usage: julia --project=. benchmarks/basis_function_access_patterns.jl
|
||||
|
||||
using BenchmarkTools
|
||||
using Tensors
|
||||
|
||||
# ============================================================================
|
||||
# Define minimal types for testing
|
||||
# ============================================================================
|
||||
|
||||
abstract type AbstractTopology end
|
||||
struct Tetrahedron <: AbstractTopology end
|
||||
|
||||
abstract type AbstractBasis end
|
||||
struct Lagrange{P} <: AbstractBasis end
|
||||
|
||||
# Vec type comes from Tensors.jl (already available in JuliaFEM)
|
||||
# Vec{3,Float64} for 3D gradients
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 1: Return tuple, index with getindex
|
||||
# ============================================================================
|
||||
# Advantages: Natural Julia syntax, type-stable
|
||||
# Disadvantages: Might not inline getindex?
|
||||
|
||||
"""
|
||||
Get all basis functions for Triangle, P1 Lagrange (3 functions).
|
||||
Returns tuple of 3 Float64 values.
|
||||
"""
|
||||
@inline function get_basis_functions_v1(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
u, v = xi
|
||||
# P1 triangle: N1 = 1-u-v, N2 = u, N3 = v
|
||||
return (1 - u - v, u, v)
|
||||
end
|
||||
|
||||
"""
|
||||
Get single basis function by index (1-based).
|
||||
"""
|
||||
@inline function get_basis_function_v1(topology::Triangle, basis::Lagrange{1}, xi::Vec{2,T}, i::Int) where T
|
||||
N_all = get_basis_functions_v1(topology, basis, xi)
|
||||
return N_all[i] # Tuple indexing
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 2: Generated function for single access
|
||||
# ============================================================================
|
||||
# Advantages: Compiler can specialize for each index
|
||||
# Disadvantages: More complex code
|
||||
|
||||
@inline function get_basis_functions_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
u, v = xi
|
||||
return (1 - u - v, u, v)
|
||||
end
|
||||
|
||||
"""
|
||||
Use @generated to create specialized code for each index at compile time.
|
||||
"""
|
||||
@generated function get_basis_function_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}, ::Val{I}) where {T,I}
|
||||
if I == 1
|
||||
return :(1 - xi[1] - xi[2])
|
||||
elseif I == 2
|
||||
return :(xi[1])
|
||||
elseif I == 3
|
||||
return :(xi[2])
|
||||
else
|
||||
return :(error("Invalid basis function index: $I"))
|
||||
end
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 3: Manual dispatch on Val (type-stable index)
|
||||
# ============================================================================
|
||||
# Advantages: Explicit, clear what's happening
|
||||
# Disadvantages: Verbose, need to write each case
|
||||
|
||||
@inline function get_basis_functions_v3(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
u, v = xi
|
||||
return (1 - u - v, u, v)
|
||||
end
|
||||
|
||||
@inline get_basis_function_v3(t::Triangle, b::Lagrange{1}, xi::Vec{2,T}, ::Val{1}) where T = 1 - xi[1] - xi[2]
|
||||
@inline get_basis_function_v3(t::Triangle, b::Lagrange{1}, xi::Vec{2,T}, ::Val{2}) where T = xi[1]
|
||||
@inline get_basis_function_v3(t::Triangle, b::Lagrange{1}, xi::Vec{2,T}, ::Val{3}) where T = xi[2]
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 4: Struct with getindex (most Julian)
|
||||
# ============================================================================
|
||||
# Advantages: Can use N[i] syntax naturally
|
||||
# Disadvantages: Extra struct allocation?
|
||||
|
||||
struct BasisFunctions{N,T}
|
||||
data::NTuple{N,T}
|
||||
end
|
||||
|
||||
Base.@propagate_inbounds Base.getindex(bf::BasisFunctions, i::Int) = bf.data[i]
|
||||
Base.length(::BasisFunctions{N}) where N = N
|
||||
|
||||
@inline function get_basis_functions_v4(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
u, v = xi
|
||||
return BasisFunctions((1 - u - v, u, v))
|
||||
end
|
||||
|
||||
# Can use natural indexing
|
||||
@inline function get_basis_function_v4(topology::Triangle, basis::Lagrange{1}, xi::Vec{2,T}, i::Int) where T
|
||||
N = get_basis_functions_v4(topology, basis, xi)
|
||||
return N[i]
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 5: Separate implementation per basis function (extreme specialization)
|
||||
# ============================================================================
|
||||
# Advantages: Maximum performance, no tuple allocation at all
|
||||
# Disadvantages: Lots of code duplication
|
||||
|
||||
@inline function get_basis_function_1_v5(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
return 1 - xi[1] - xi[2]
|
||||
end
|
||||
|
||||
@inline function get_basis_function_2_v5(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
return xi[1]
|
||||
end
|
||||
|
||||
@inline function get_basis_function_3_v5(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
return xi[2]
|
||||
end
|
||||
|
||||
@inline function get_basis_functions_v5(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
u, v = xi
|
||||
return (1 - u - v, u, v)
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark: Access all basis functions (typical in assembly loop)
|
||||
# ============================================================================
|
||||
|
||||
function benchmark_all_access()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK: Access ALL basis functions")
|
||||
println("="^80)
|
||||
|
||||
topology = Triangle()
|
||||
basis = Lagrange{1}()
|
||||
xi = Vec(0.25, 0.25)
|
||||
|
||||
println("\nStrategy 1: Tuple return + getindex")
|
||||
@btime get_basis_functions_v1($topology, $basis, $xi)
|
||||
|
||||
println("\nStrategy 2: Generated function")
|
||||
@btime get_basis_functions_v2($topology, $basis, $xi)
|
||||
|
||||
println("\nStrategy 3: Val dispatch")
|
||||
@btime get_basis_functions_v3($topology, $basis, $xi)
|
||||
|
||||
println("\nStrategy 4: BasisFunctions struct")
|
||||
@btime get_basis_functions_v4($topology, $basis, $xi)
|
||||
|
||||
println("\nStrategy 5: Separate functions")
|
||||
@btime get_basis_functions_v5($topology, $basis, $xi)
|
||||
|
||||
# Verify all return same values
|
||||
r1 = get_basis_functions_v1(topology, basis, xi)
|
||||
r2 = get_basis_functions_v2(topology, basis, xi)
|
||||
r3 = get_basis_functions_v3(topology, basis, xi)
|
||||
r4 = get_basis_functions_v4(topology, basis, xi).data
|
||||
r5 = get_basis_functions_v5(topology, basis, xi)
|
||||
|
||||
@assert r1 == r2 == r3 == r4 == r5 "Results don't match!"
|
||||
println("\n✓ All strategies return identical values: $r1")
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark: Access SINGLE basis function (for nodal assembly)
|
||||
# ============================================================================
|
||||
|
||||
function benchmark_single_access()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK: Access SINGLE basis function (nodal assembly)")
|
||||
println("="^80)
|
||||
|
||||
topology = Triangle()
|
||||
basis = Lagrange{1}()
|
||||
xi = Vec(0.25, 0.25)
|
||||
|
||||
println("\nStrategy 1: Tuple + runtime index")
|
||||
@btime get_basis_function_v1($topology, $basis, $xi, 2)
|
||||
|
||||
println("\nStrategy 2: Generated function with Val{2}")
|
||||
@btime get_basis_function_v2($topology, $basis, $xi, Val(2))
|
||||
|
||||
println("\nStrategy 3: Val dispatch")
|
||||
@btime get_basis_function_v3($topology, $basis, $xi, Val(2))
|
||||
|
||||
println("\nStrategy 4: BasisFunctions struct + index")
|
||||
@btime get_basis_function_v4($topology, $basis, $xi, 2)
|
||||
|
||||
println("\nStrategy 5: Direct function call")
|
||||
@btime get_basis_function_2_v5($topology, $basis, $xi)
|
||||
|
||||
# Verify all return same value
|
||||
r1 = get_basis_function_v1(topology, basis, xi, 2)
|
||||
r2 = get_basis_function_v2(topology, basis, xi, Val(2))
|
||||
r3 = get_basis_function_v3(topology, basis, xi, Val(2))
|
||||
r4 = get_basis_function_v4(topology, basis, xi, 2)
|
||||
r5 = get_basis_function_2_v5(topology, basis, xi)
|
||||
|
||||
@assert r1 == r2 == r3 == r4 == r5 "Results don't match!"
|
||||
println("\n✓ All strategies return identical value: $r1")
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark: Typical assembly loop pattern
|
||||
# ============================================================================
|
||||
|
||||
function benchmark_assembly_loop()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK: Typical assembly loop (iterate over all basis functions)")
|
||||
println("="^80)
|
||||
|
||||
topology = Triangle()
|
||||
basis = Lagrange{1}()
|
||||
xi = Vec(0.25, 0.25)
|
||||
|
||||
# Pattern 1: Get all, iterate over tuple
|
||||
println("\nPattern 1: Get all as tuple, iterate")
|
||||
function assemble_v1()
|
||||
N_all = get_basis_functions_v1(topology, basis, xi)
|
||||
s = 0.0
|
||||
for N_i in N_all
|
||||
s += N_i * N_i # Dummy computation
|
||||
end
|
||||
return s
|
||||
end
|
||||
@btime assemble_v1()
|
||||
|
||||
# Pattern 2: Get all, index in loop
|
||||
println("\nPattern 2: Get all, index with i")
|
||||
function assemble_v2()
|
||||
N_all = get_basis_functions_v1(topology, basis, xi)
|
||||
s = 0.0
|
||||
for i in 1:3
|
||||
s += N_all[i] * N_all[i]
|
||||
end
|
||||
return s
|
||||
end
|
||||
@btime assemble_v2()
|
||||
|
||||
# Pattern 3: Get one at a time (nodal assembly style)
|
||||
println("\nPattern 3: Get one at a time with Val")
|
||||
function assemble_v3()
|
||||
s = 0.0
|
||||
# Unrolled loop (what compiler would do with Val)
|
||||
N1 = get_basis_function_v3(topology, basis, xi, Val(1))
|
||||
s += N1 * N1
|
||||
N2 = get_basis_function_v3(topology, basis, xi, Val(2))
|
||||
s += N2 * N2
|
||||
N3 = get_basis_function_v3(topology, basis, xi, Val(3))
|
||||
s += N3 * N3
|
||||
return s
|
||||
end
|
||||
@btime assemble_v3()
|
||||
|
||||
# Pattern 4: Direct function calls (strategy 5)
|
||||
println("\nPattern 4: Direct function calls (extreme specialization)")
|
||||
function assemble_v4()
|
||||
s = 0.0
|
||||
N1 = get_basis_function_1_v5(topology, basis, xi)
|
||||
s += N1 * N1
|
||||
N2 = get_basis_function_2_v5(topology, basis, xi)
|
||||
s += N2 * N2
|
||||
N3 = get_basis_function_3_v5(topology, basis, xi)
|
||||
s += N3 * N3
|
||||
return s
|
||||
end
|
||||
@btime assemble_v4()
|
||||
|
||||
# Verify all compute same result
|
||||
r1 = assemble_v1()
|
||||
r2 = assemble_v2()
|
||||
r3 = assemble_v3()
|
||||
r4 = assemble_v4()
|
||||
@assert r1 == r2 == r3 == r4 "Assembly results don't match!"
|
||||
println("\n✓ All patterns compute same result: $r1")
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark: Basis derivatives (return Vec)
|
||||
# ============================================================================
|
||||
|
||||
function benchmark_derivatives()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK: Basis function DERIVATIVES (return Vec)")
|
||||
println("="^80)
|
||||
|
||||
topology = Triangle()
|
||||
basis = Lagrange{1}()
|
||||
xi = Vec(0.25, 0.25)
|
||||
|
||||
# Triangle P1 derivatives (constant):
|
||||
# dN1/d(u,v) = (-1, -1)
|
||||
# dN2/d(u,v) = (1, 0)
|
||||
# dN3/d(u,v) = (0, 1)
|
||||
|
||||
println("\nStrategy 1: Return tuple of Vecs")
|
||||
@inline function get_basis_derivatives_v1(::Triangle, ::Lagrange{1}, xi::Vec{2,T}) where T
|
||||
return (Vec(-1.0, -1.0), Vec(1.0, 0.0), Vec(0.0, 1.0))
|
||||
end
|
||||
@btime get_basis_derivatives_v1($topology, $basis, $xi)
|
||||
|
||||
println("\nStrategy 2: Return single Vec with Val indexing")
|
||||
@inline get_basis_derivative_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}, ::Val{1}) where T = Vec(-1.0, -1.0)
|
||||
@inline get_basis_derivative_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}, ::Val{2}) where T = Vec(1.0, 0.0)
|
||||
@inline get_basis_derivative_v2(::Triangle, ::Lagrange{1}, xi::Vec{2,T}, ::Val{3}) where T = Vec(0.0, 1.0)
|
||||
@btime get_basis_derivative_v2($topology, $basis, $xi, Val(2))
|
||||
|
||||
# Verify
|
||||
all_derivs = get_basis_derivatives_v1(topology, basis, xi)
|
||||
single_deriv = get_basis_derivative_v2(topology, basis, xi, Val(2))
|
||||
@assert all_derivs[2] == single_deriv
|
||||
println("\n✓ Derivatives match: $single_deriv")
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Main execution
|
||||
# ============================================================================
|
||||
|
||||
function main()
|
||||
println("\n")
|
||||
println("╔" * "="^78 * "╗")
|
||||
println("║" * " "^78 * "║")
|
||||
println("║" * " "^20 * "BASIS FUNCTION ACCESS PATTERNS BENCHMARK" * " "^18 * "║")
|
||||
println("║" * " "^78 * "║")
|
||||
println("╚" * "="^78 * "╝")
|
||||
|
||||
println("\nGoal: Find fastest way to access basis functions for nodal assembly")
|
||||
println("Requirements:")
|
||||
println(" - Zero allocation")
|
||||
println(" - Type stable")
|
||||
println(" - Inlineable")
|
||||
println(" - Support both 'all at once' and 'one at a time' access")
|
||||
|
||||
benchmark_all_access()
|
||||
benchmark_single_access()
|
||||
benchmark_assembly_loop()
|
||||
benchmark_derivatives()
|
||||
|
||||
println("\n" * "="^80)
|
||||
println("SUMMARY & RECOMMENDATIONS")
|
||||
println("="^80)
|
||||
println("""
|
||||
|
||||
For TRADITIONAL ASSEMBLY (get all basis functions at integration point):
|
||||
→ Use Strategy 1 or 4: Simple tuple return
|
||||
→ Should be 0-5 ns, zero allocation
|
||||
|
||||
For NODAL ASSEMBLY (get single basis function):
|
||||
→ Use Strategy 3: Val dispatch for compile-time index
|
||||
→ Should be 0-2 ns, zero allocation, fully inlined
|
||||
→ Usage: get_basis_function(Triangle(), Lagrange{1}(), xi, Val(i))
|
||||
|
||||
For DERIVATIVES:
|
||||
→ Return tuple of Vec for all derivatives
|
||||
→ Use Val indexing for single derivative
|
||||
→ Same performance as basis functions
|
||||
|
||||
RECOMMENDED API:
|
||||
```julia
|
||||
# Get all basis functions (returns tuple)
|
||||
N_all = get_basis_functions(Triangle(), Lagrange{1}(), xi)
|
||||
|
||||
# Get single basis function (Val for compile-time specialization)
|
||||
N_i = get_basis_function(Triangle(), Lagrange{1}(), xi, Val(i))
|
||||
|
||||
# Get all derivatives (returns tuple of Vec)
|
||||
dN_all = get_basis_derivatives(Triangle(), Lagrange{1}(), xi)
|
||||
|
||||
# Get single derivative (returns Vec)
|
||||
dN_i = get_basis_derivative(Triangle(), Lagrange{1}(), xi, Val(i))
|
||||
```
|
||||
|
||||
WHY Val?
|
||||
- Compiler knows index at compile time
|
||||
- Can generate optimal code for each basis function
|
||||
- Zero runtime overhead
|
||||
- Type stable
|
||||
|
||||
NOTE: For runtime indexing (i not known at compile time), tuple indexing
|
||||
is still very fast (typically 1-2 ns overhead).
|
||||
""")
|
||||
|
||||
println("\n" * "="^80)
|
||||
end
|
||||
|
||||
# Run benchmarks
|
||||
if abspath(PROGRAM_FILE) == @__FILE__
|
||||
main()
|
||||
end
|
||||
@@ -1,504 +0,0 @@
|
||||
# Benchmark: Basis Function Access Patterns for Tet10 (Realistic 3D Case)
|
||||
#
|
||||
# Focus: 10-node quadratic tetrahedron (Tet10) - the workhorse for 3D simulations
|
||||
#
|
||||
# Goal: Find the fastest way to access basis functions and their DERIVATIVES with:
|
||||
# 1. Return all 10 basis functions as tuple (zero allocation)
|
||||
# 2. Return single basis function by index (must be inlineable)
|
||||
# 3. Return all 10 derivatives as tuple of Vec{3} (zero allocation)
|
||||
# 4. Return single derivative by index (must be inlineable)
|
||||
# 5. Pass topology separately (separation of concerns: Element(Tetrahedron, Lagrange{2}, ...))
|
||||
# 6. Must be type-stable and superfast
|
||||
#
|
||||
# Use Case:
|
||||
# - Stiffness matrix assembly: Need derivatives (B matrix construction)
|
||||
# - Mass matrix assembly: Need basis functions (M matrix construction)
|
||||
# - Nodal assembly: Need single basis function/derivative at a time
|
||||
#
|
||||
# Usage: julia --project=. benchmarks/basis_function_access_tet10.jl
|
||||
|
||||
using BenchmarkTools
|
||||
using Tensors
|
||||
|
||||
# ============================================================================
|
||||
# Define minimal types for testing
|
||||
# ============================================================================
|
||||
|
||||
abstract type AbstractTopology end
|
||||
struct Tetrahedron <: AbstractTopology end
|
||||
|
||||
abstract type AbstractBasis end
|
||||
struct Lagrange{P} <: AbstractBasis end
|
||||
|
||||
# ============================================================================
|
||||
# Tet10 Basis Functions (Quadratic Tetrahedron, 10 nodes)
|
||||
# ============================================================================
|
||||
# Node numbering:
|
||||
# 1-4: vertices
|
||||
# 5-10: edge midpoints (5: 1-2, 6: 2-3, 7: 3-1, 8: 1-4, 9: 2-4, 10: 3-4)
|
||||
#
|
||||
# Parametric coordinates: (u, v, w) where u+v+w ≤ 1
|
||||
# Reference element: vertices at (0,0,0), (1,0,0), (0,1,0), (0,0,1)
|
||||
|
||||
"""
|
||||
Get all 10 basis functions for Tet10 at parametric point (u,v,w).
|
||||
Returns NTuple{10, Float64}.
|
||||
|
||||
The basis functions are:
|
||||
- N1 = (1-u-v-w)(1-2u-2v-2w) [vertex 1]
|
||||
- N2 = u(2u-1) [vertex 2]
|
||||
- N3 = v(2v-1) [vertex 3]
|
||||
- N4 = w(2w-1) [vertex 4]
|
||||
- N5 = 4u(1-u-v-w) [edge 1-2]
|
||||
- N6 = 4uv [edge 2-3]
|
||||
- N7 = 4v(1-u-v-w) [edge 3-1]
|
||||
- N8 = 4w(1-u-v-w) [edge 1-4]
|
||||
- N9 = 4uw [edge 2-4]
|
||||
- N10= 4vw [edge 3-4]
|
||||
"""
|
||||
@inline function get_basis_functions(::Tetrahedron, ::Lagrange{2}, xi::Vec{3,T}) where T
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w # barycentric coordinate for vertex 1
|
||||
|
||||
# Vertex nodes (1-4)
|
||||
N1 = λ * (2λ - 1)
|
||||
N2 = u * (2u - 1)
|
||||
N3 = v * (2v - 1)
|
||||
N4 = w * (2w - 1)
|
||||
|
||||
# Edge midpoint nodes (5-10)
|
||||
N5 = 4 * u * λ
|
||||
N6 = 4 * u * v
|
||||
N7 = 4 * v * λ
|
||||
N8 = 4 * w * λ
|
||||
N9 = 4 * u * w
|
||||
N10 = 4 * v * w
|
||||
|
||||
return (N1, N2, N3, N4, N5, N6, N7, N8, N9, N10)
|
||||
end
|
||||
|
||||
"""
|
||||
Get all 10 basis function derivatives for Tet10.
|
||||
Returns NTuple{10, Vec{3,Float64}}.
|
||||
|
||||
Each derivative is ∇N_i = (∂N_i/∂u, ∂N_i/∂v, ∂N_i/∂w)
|
||||
"""
|
||||
@inline function get_basis_derivatives(::Tetrahedron, ::Lagrange{2}, xi::Vec{3,T}) where T
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w
|
||||
|
||||
# Derivatives of vertex nodes
|
||||
dN1 = Vec(-3 + 4u + 4v + 4w, -3 + 4u + 4v + 4w, -3 + 4u + 4v + 4w)
|
||||
dN2 = Vec(4u - 1, 0.0, 0.0)
|
||||
dN3 = Vec(0.0, 4v - 1, 0.0)
|
||||
dN4 = Vec(0.0, 0.0, 4w - 1)
|
||||
|
||||
# Derivatives of edge midpoint nodes
|
||||
dN5 = Vec(4λ - 4u, -4u, -4u)
|
||||
dN6 = Vec(4v, 4u, 0.0)
|
||||
dN7 = Vec(-4v, 4λ - 4v, -4v)
|
||||
dN8 = Vec(-4w, -4w, 4λ - 4w)
|
||||
dN9 = Vec(4w, 0.0, 4u)
|
||||
dN10 = Vec(0.0, 4w, 4v)
|
||||
|
||||
return (dN1, dN2, dN3, dN4, dN5, dN6, dN7, dN8, dN9, dN10)
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 1: Tuple indexing (runtime index)
|
||||
# ============================================================================
|
||||
|
||||
@inline function get_basis_function_v1(topology::Tetrahedron, basis::Lagrange{2},
|
||||
xi::Vec{3,T}, i::Int) where T
|
||||
N_all = get_basis_functions(topology, basis, xi)
|
||||
return N_all[i]
|
||||
end
|
||||
|
||||
@inline function get_basis_derivative_v1(topology::Tetrahedron, basis::Lagrange{2},
|
||||
xi::Vec{3,T}, i::Int) where T
|
||||
dN_all = get_basis_derivatives(topology, basis, xi)
|
||||
return dN_all[i]
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 2: Val dispatch (compile-time index)
|
||||
# ============================================================================
|
||||
|
||||
@inline function get_basis_function_v2(t::Tetrahedron, b::Lagrange{2},
|
||||
xi::Vec{3,T}, ::Val{I}) where {T,I}
|
||||
N_all = get_basis_functions(t, b, xi)
|
||||
return N_all[I]
|
||||
end
|
||||
|
||||
@inline function get_basis_derivative_v2(t::Tetrahedron, b::Lagrange{2},
|
||||
xi::Vec{3,T}, ::Val{I}) where {T,I}
|
||||
dN_all = get_basis_derivatives(t, b, xi)
|
||||
return dN_all[I]
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 3: Generated function (compute only requested basis function)
|
||||
# ============================================================================
|
||||
|
||||
@generated function get_basis_function_v3(::Tetrahedron, ::Lagrange{2},
|
||||
xi::Vec{3,T}, ::Val{I}) where {T,I}
|
||||
# Generate specialized code for each index
|
||||
if I == 1
|
||||
return quote
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w
|
||||
return λ * (2λ - 1)
|
||||
end
|
||||
elseif I == 2
|
||||
return quote
|
||||
u = xi[1]
|
||||
return u * (2u - 1)
|
||||
end
|
||||
elseif I == 3
|
||||
return quote
|
||||
v = xi[2]
|
||||
return v * (2v - 1)
|
||||
end
|
||||
elseif I == 4
|
||||
return quote
|
||||
w = xi[3]
|
||||
return w * (2w - 1)
|
||||
end
|
||||
elseif I == 5
|
||||
return quote
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w
|
||||
return 4 * u * λ
|
||||
end
|
||||
elseif I == 6
|
||||
return quote
|
||||
u, v = xi[1], xi[2]
|
||||
return 4 * u * v
|
||||
end
|
||||
elseif I == 7
|
||||
return quote
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w
|
||||
return 4 * v * λ
|
||||
end
|
||||
elseif I == 8
|
||||
return quote
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w
|
||||
return 4 * w * λ
|
||||
end
|
||||
elseif I == 9
|
||||
return quote
|
||||
u, w = xi[1], xi[3]
|
||||
return 4 * u * w
|
||||
end
|
||||
elseif I == 10
|
||||
return quote
|
||||
v, w = xi[2], xi[3]
|
||||
return 4 * v * w
|
||||
end
|
||||
else
|
||||
return :(error("Invalid basis function index: $I for Tet10"))
|
||||
end
|
||||
end
|
||||
|
||||
@generated function get_basis_derivative_v3(::Tetrahedron, ::Lagrange{2},
|
||||
xi::Vec{3,T}, ::Val{I}) where {T,I}
|
||||
if I == 1
|
||||
return quote
|
||||
u, v, w = xi
|
||||
return Vec(-3 + 4u + 4v + 4w, -3 + 4u + 4v + 4w, -3 + 4u + 4v + 4w)
|
||||
end
|
||||
elseif I == 2
|
||||
return quote
|
||||
u = xi[1]
|
||||
return Vec(4u - 1, 0.0, 0.0)
|
||||
end
|
||||
elseif I == 3
|
||||
return quote
|
||||
v = xi[2]
|
||||
return Vec(0.0, 4v - 1, 0.0)
|
||||
end
|
||||
elseif I == 4
|
||||
return quote
|
||||
w = xi[3]
|
||||
return Vec(0.0, 0.0, 4w - 1)
|
||||
end
|
||||
elseif I == 5
|
||||
return quote
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w
|
||||
return Vec(4λ - 4u, -4u, -4u)
|
||||
end
|
||||
elseif I == 6
|
||||
return quote
|
||||
u, v = xi[1], xi[2]
|
||||
return Vec(4v, 4u, 0.0)
|
||||
end
|
||||
elseif I == 7
|
||||
return quote
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w
|
||||
return Vec(-4v, 4λ - 4v, -4v)
|
||||
end
|
||||
elseif I == 8
|
||||
return quote
|
||||
u, v, w = xi
|
||||
λ = 1 - u - v - w
|
||||
return Vec(-4w, -4w, 4λ - 4w)
|
||||
end
|
||||
elseif I == 9
|
||||
return quote
|
||||
u, w = xi[1], xi[3]
|
||||
return Vec(4w, 0.0, 4u)
|
||||
end
|
||||
elseif I == 10
|
||||
return quote
|
||||
v, w = xi[2], xi[3]
|
||||
return Vec(0.0, 4w, 4v)
|
||||
end
|
||||
else
|
||||
return :(error("Invalid basis function index: $I for Tet10"))
|
||||
end
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark Functions
|
||||
# ============================================================================
|
||||
|
||||
function benchmark_all_basis_functions()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK 1: Get ALL 10 basis functions")
|
||||
println("="^80)
|
||||
println("Use case: Mass matrix assembly, need all N_i at integration point")
|
||||
|
||||
topology = Tetrahedron()
|
||||
basis = Lagrange{2}()
|
||||
xi = Vec(0.25, 0.25, 0.2) # Typical integration point
|
||||
|
||||
println("\nAccess all 10 basis functions:")
|
||||
@btime get_basis_functions($topology, $basis, $xi)
|
||||
|
||||
result = get_basis_functions(topology, basis, xi)
|
||||
println("\n✓ Result (10 values): ", result)
|
||||
println("✓ Sum of basis functions (partition of unity): ", sum(result))
|
||||
@assert abs(sum(result) - 1.0) < 1e-10 "Partition of unity violated!"
|
||||
end
|
||||
|
||||
function benchmark_single_basis_function()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK 2: Get SINGLE basis function (nodal assembly)")
|
||||
println("="^80)
|
||||
println("Use case: Nodal assembly, need N_i for specific node")
|
||||
|
||||
topology = Tetrahedron()
|
||||
basis = Lagrange{2}()
|
||||
xi = Vec(0.25, 0.25, 0.2)
|
||||
node_idx = 5 # Edge midpoint node
|
||||
|
||||
println("\nStrategy 1: Tuple + runtime index")
|
||||
@btime get_basis_function_v1($topology, $basis, $xi, $node_idx)
|
||||
|
||||
println("\nStrategy 2: Val dispatch (compile-time index)")
|
||||
@btime get_basis_function_v2($topology, $basis, $xi, Val($node_idx))
|
||||
|
||||
println("\nStrategy 3: @generated function (minimal computation)")
|
||||
@btime get_basis_function_v3($topology, $basis, $xi, Val($node_idx))
|
||||
|
||||
# Verify all return same value
|
||||
r1 = get_basis_function_v1(topology, basis, xi, node_idx)
|
||||
r2 = get_basis_function_v2(topology, basis, xi, Val(node_idx))
|
||||
r3 = get_basis_function_v3(topology, basis, xi, Val(node_idx))
|
||||
@assert r1 ≈ r2 ≈ r3 "Strategies return different values!"
|
||||
println("\n✓ All strategies return: N_$node_idx = $r1")
|
||||
end
|
||||
|
||||
function benchmark_all_derivatives()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK 3: Get ALL 10 basis function derivatives")
|
||||
println("="^80)
|
||||
println("Use case: Stiffness matrix assembly (B matrix construction)")
|
||||
println("Most important benchmark for 3D simulations!")
|
||||
|
||||
topology = Tetrahedron()
|
||||
basis = Lagrange{2}()
|
||||
xi = Vec(0.25, 0.25, 0.2)
|
||||
|
||||
println("\nAccess all 10 derivatives (each is Vec{3}):")
|
||||
@btime get_basis_derivatives($topology, $basis, $xi)
|
||||
|
||||
result = get_basis_derivatives(topology, basis, xi)
|
||||
println("\n✓ Result (10 Vec{3} gradients):")
|
||||
for (i, dN) in enumerate(result)
|
||||
println(" ∇N_$i = $dN")
|
||||
end
|
||||
end
|
||||
|
||||
function benchmark_single_derivative()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK 4: Get SINGLE basis function derivative")
|
||||
println("="^80)
|
||||
println("Use case: Nodal assembly for stiffness matrix")
|
||||
|
||||
topology = Tetrahedron()
|
||||
basis = Lagrange{2}()
|
||||
xi = Vec(0.25, 0.25, 0.2)
|
||||
node_idx = 5
|
||||
|
||||
println("\nStrategy 1: Tuple + runtime index")
|
||||
@btime get_basis_derivative_v1($topology, $basis, $xi, $node_idx)
|
||||
|
||||
println("\nStrategy 2: Val dispatch")
|
||||
@btime get_basis_derivative_v2($topology, $basis, $xi, Val($node_idx))
|
||||
|
||||
println("\nStrategy 3: @generated function")
|
||||
@btime get_basis_derivative_v3($topology, $basis, $xi, Val($node_idx))
|
||||
|
||||
# Verify
|
||||
r1 = get_basis_derivative_v1(topology, basis, xi, node_idx)
|
||||
r2 = get_basis_derivative_v2(topology, basis, xi, Val(node_idx))
|
||||
r3 = get_basis_derivative_v3(topology, basis, xi, Val(node_idx))
|
||||
@assert r1 ≈ r2 ≈ r3 "Strategies return different values!"
|
||||
println("\n✓ All strategies return: ∇N_$node_idx = $r1")
|
||||
end
|
||||
|
||||
function benchmark_stiffness_assembly_pattern()
|
||||
println("\n" * "="^80)
|
||||
println("BENCHMARK 5: Realistic stiffness matrix assembly loop")
|
||||
println("="^80)
|
||||
println("Use case: Compute element stiffness K_e (typical FEM inner loop)")
|
||||
println("Pattern: B^T D B where B = strain-displacement matrix")
|
||||
|
||||
topology = Tetrahedron()
|
||||
basis = Lagrange{2}()
|
||||
xi = Vec(0.25, 0.25, 0.2) # Integration point
|
||||
|
||||
# Typical pattern: Get all derivatives, compute B matrix terms
|
||||
println("\nPattern 1: Get all derivatives at once (traditional)")
|
||||
stiffness_loop_v1 = let topology = topology, basis = basis, xi = xi
|
||||
() -> begin
|
||||
dN_all = get_basis_derivatives(topology, basis, xi)
|
||||
s = 0.0
|
||||
# Simplified: compute trace of K (just for benchmarking)
|
||||
for i in 1:10
|
||||
for j in 1:10
|
||||
dNi = dN_all[i]
|
||||
dNj = dN_all[j]
|
||||
s += dot(dNi, dNj) # Simplified K_ij computation
|
||||
end
|
||||
end
|
||||
return s
|
||||
end
|
||||
end
|
||||
@btime $stiffness_loop_v1()
|
||||
|
||||
println("\nPattern 2: Get derivatives with @generated (manual unroll)")
|
||||
stiffness_loop_v3 = let topology = topology, basis = basis, xi = xi
|
||||
() -> begin
|
||||
s = 0.0
|
||||
# Unrolled loop (compiler would do this with Val)
|
||||
dN1 = get_basis_derivative_v3(topology, basis, xi, Val(1))
|
||||
dN2 = get_basis_derivative_v3(topology, basis, xi, Val(2))
|
||||
dN3 = get_basis_derivative_v3(topology, basis, xi, Val(3))
|
||||
dN4 = get_basis_derivative_v3(topology, basis, xi, Val(4))
|
||||
dN5 = get_basis_derivative_v3(topology, basis, xi, Val(5))
|
||||
dN6 = get_basis_derivative_v3(topology, basis, xi, Val(6))
|
||||
dN7 = get_basis_derivative_v3(topology, basis, xi, Val(7))
|
||||
dN8 = get_basis_derivative_v3(topology, basis, xi, Val(8))
|
||||
dN9 = get_basis_derivative_v3(topology, basis, xi, Val(9))
|
||||
dN10 = get_basis_derivative_v3(topology, basis, xi, Val(10))
|
||||
|
||||
# Compute all pairs (100 dot products)
|
||||
for dNi in (dN1, dN2, dN3, dN4, dN5, dN6, dN7, dN8, dN9, dN10)
|
||||
for dNj in (dN1, dN2, dN3, dN4, dN5, dN6, dN7, dN8, dN9, dN10)
|
||||
s += dot(dNi, dNj)
|
||||
end
|
||||
end
|
||||
return s
|
||||
end
|
||||
end
|
||||
@btime $stiffness_loop_v3()
|
||||
|
||||
# Verify all compute same result
|
||||
r1 = stiffness_loop_v1()
|
||||
r3 = stiffness_loop_v3()
|
||||
@assert abs(r1 - r3) < 1e-10 "Assembly patterns give different results!"
|
||||
println("\n✓ Assembly result: $r1")
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Main execution
|
||||
# ============================================================================
|
||||
|
||||
function main()
|
||||
println("\n")
|
||||
println("╔" * "="^78 * "╗")
|
||||
println("║" * " "^78 * "║")
|
||||
println("║" * " "^15 * "TET10 BASIS FUNCTION ACCESS BENCHMARK" * " "^25 * "║")
|
||||
println("║" * " "^20 * "(10-node Quadratic Tetrahedron)" * " "^26 * "║")
|
||||
println("║" * " "^78 * "║")
|
||||
println("╚" * "="^78 * "╝")
|
||||
|
||||
println("\nElement: Tet10 (10-node quadratic tetrahedron)")
|
||||
println("Nodes: 4 vertices + 6 edge midpoints")
|
||||
println("Polynomial degree: P2 (quadratic)")
|
||||
println("Dimension: 3D")
|
||||
println("\nSeparation of concerns API:")
|
||||
println(" Element(Tetrahedron, Lagrange{2}, connectivity)")
|
||||
println(" get_basis_functions(Tetrahedron(), Lagrange{2}(), xi)")
|
||||
println(" get_basis_derivatives(Tetrahedron(), Lagrange{2}(), xi)")
|
||||
|
||||
benchmark_all_basis_functions()
|
||||
benchmark_single_basis_function()
|
||||
benchmark_all_derivatives()
|
||||
benchmark_single_derivative()
|
||||
benchmark_stiffness_assembly_pattern()
|
||||
|
||||
println("\n" * "="^80)
|
||||
println("SUMMARY & RECOMMENDATIONS FOR 3D SIMULATIONS")
|
||||
println("="^80)
|
||||
println("""
|
||||
|
||||
FOR STIFFNESS MATRIX ASSEMBLY (derivatives):
|
||||
→ Use get_basis_derivatives() - returns all 10 gradients as tuple
|
||||
→ Expected: ~10-30 ns, zero allocation
|
||||
→ This is the HOT PATH for 3D FEM!
|
||||
|
||||
FOR MASS MATRIX ASSEMBLY (basis functions):
|
||||
→ Use get_basis_functions() - returns all 10 values as tuple
|
||||
→ Expected: ~5-15 ns, zero allocation
|
||||
|
||||
FOR NODAL ASSEMBLY (single node operations):
|
||||
→ Use Val dispatch: get_basis_derivative(t, b, xi, Val(i))
|
||||
→ @generated gives minimal computation (only compute requested function)
|
||||
→ Expected: ~5-10 ns per node
|
||||
|
||||
API DESIGN DECISION:
|
||||
```julia
|
||||
# Separation of concerns (RECOMMENDED):
|
||||
Element(Tetrahedron, Lagrange{2}, connectivity)
|
||||
|
||||
# Functions take topology explicitly:
|
||||
dN_all = get_basis_derivatives(Tetrahedron(), Lagrange{2}(), xi)
|
||||
dN_i = get_basis_derivative(Tetrahedron(), Lagrange{2}(), xi, Val(i))
|
||||
```
|
||||
|
||||
WHY THIS API?
|
||||
- Clear separation: topology is geometry, basis is interpolation
|
||||
- Topology passed to basis evaluation (no redundancy in type parameters)
|
||||
- Type-stable, zero-allocation, fully inlined
|
||||
- Works with any basis type (Lagrange, Hierarchical, Nedelec, etc.)
|
||||
|
||||
PERFORMANCE TARGET:
|
||||
- 10-node Tet10 derivatives: < 30 ns (achieved!)
|
||||
- 100× faster than old Dict-based approach
|
||||
- Ready for million-element meshes
|
||||
""")
|
||||
|
||||
println("\n" * "="^80)
|
||||
end
|
||||
|
||||
# Run benchmarks
|
||||
if abspath(PROGRAM_FILE) == @__FILE__
|
||||
main()
|
||||
end
|
||||
@@ -1,240 +0,0 @@
|
||||
# Performance Analysis for Deformation Gradient Implementation
|
||||
# This script analyzes the machine code generated and validates zero-allocation claims
|
||||
|
||||
using JuliaFEM
|
||||
using Tensors
|
||||
using BenchmarkTools
|
||||
using InteractiveUtils
|
||||
|
||||
# Load the deformation gradient code
|
||||
include("../src/physics/deformation_gradient.jl")
|
||||
|
||||
println("="^80)
|
||||
println("DEFORMATION GRADIENT PERFORMANCE ANALYSIS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Setup test data
|
||||
X_nodes = (
|
||||
Vec(0.0, 0.0, 0.0),
|
||||
Vec(1.0, 0.0, 0.0),
|
||||
Vec(1.0, 1.0, 0.0),
|
||||
Vec(0.0, 1.0, 0.0),
|
||||
Vec(0.0, 0.0, 1.0),
|
||||
Vec(1.0, 0.0, 1.0),
|
||||
Vec(1.0, 1.0, 1.0),
|
||||
Vec(0.0, 1.0, 1.0)
|
||||
)
|
||||
|
||||
u_nodes = (
|
||||
Vec(0.0, 0.0, 0.0),
|
||||
Vec(0.1, 0.0, 0.0),
|
||||
Vec(0.1, 0.0, 0.0),
|
||||
Vec(0.0, 0.0, 0.0),
|
||||
Vec(0.0, 0.0, 0.0),
|
||||
Vec(0.1, 0.0, 0.0),
|
||||
Vec(0.1, 0.0, 0.0),
|
||||
Vec(0.0, 0.0, 0.0)
|
||||
)
|
||||
|
||||
ξ = Vec(0.0, 0.0, 0.0)
|
||||
dN_dξ = get_basis_derivatives(Hexahedron(), Lagrange{Hexahedron,1}(), ξ)
|
||||
|
||||
# Compute Jacobian
|
||||
function compute_jacobian(X_nodes, dN_dξ)
|
||||
J = zero(Tensor{2,3,Float64,9})
|
||||
for i in 1:8
|
||||
J += X_nodes[i] ⊗ dN_dξ[i]
|
||||
end
|
||||
return J
|
||||
end
|
||||
|
||||
J = compute_jacobian(X_nodes, dN_dξ)
|
||||
|
||||
# ============================================================================
|
||||
# 1. ALLOCATION ANALYSIS
|
||||
# ============================================================================
|
||||
println("1. ALLOCATION ANALYSIS")
|
||||
println("-"^80)
|
||||
|
||||
# Warm up (compile)
|
||||
F = compute_deformation_gradient(X_nodes, u_nodes, dN_dξ, J, FiniteStrain())
|
||||
|
||||
# Measure allocations
|
||||
allocs = @allocated compute_deformation_gradient(X_nodes, u_nodes, dN_dξ, J, FiniteStrain())
|
||||
println("Allocations: $allocs bytes")
|
||||
|
||||
if allocs == 0
|
||||
println("✅ ZERO ALLOCATIONS CONFIRMED!")
|
||||
else
|
||||
println("❌ WARNING: Found $allocs bytes allocated!")
|
||||
end
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# 2. PERFORMANCE BENCHMARKING
|
||||
# ============================================================================
|
||||
println("2. PERFORMANCE BENCHMARKING")
|
||||
println("-"^80)
|
||||
|
||||
println("Benchmarking compute_deformation_gradient...")
|
||||
result = @benchmark compute_deformation_gradient($X_nodes, $u_nodes, $dN_dξ, $J, FiniteStrain())
|
||||
println(result)
|
||||
println()
|
||||
|
||||
median_time_ns = median(result.times)
|
||||
println("Median time: $(median_time_ns) ns = $(median_time_ns/1000) μs")
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# 3. LLVM IR ANALYSIS
|
||||
# ============================================================================
|
||||
println("3. LLVM IR ANALYSIS")
|
||||
println("-"^80)
|
||||
println("Examining LLVM IR for signs of optimization...")
|
||||
println()
|
||||
|
||||
io_llvm = IOBuffer()
|
||||
code_llvm(io_llvm, compute_deformation_gradient,
|
||||
typeof((X_nodes, u_nodes, dN_dξ, J, FiniteStrain())))
|
||||
llvm_code = String(take!(io_llvm))
|
||||
|
||||
# Count key indicators
|
||||
n_allocations = count(r"@julia.gc_alloc_obj", llvm_code)
|
||||
n_stores = count(r"store", llvm_code)
|
||||
n_loads = count(r"load", llvm_code)
|
||||
n_vector_ops = count(r"<\d+ x ", llvm_code) # SIMD vector operations
|
||||
|
||||
println("LLVM IR Statistics:")
|
||||
println(" - GC allocations: $n_allocations")
|
||||
println(" - Store operations: $n_stores")
|
||||
println(" - Load operations: $n_loads")
|
||||
println(" - Vector operations (SIMD): $n_vector_ops")
|
||||
println()
|
||||
|
||||
if n_allocations == 0
|
||||
println("✅ No GC allocations in LLVM IR!")
|
||||
else
|
||||
println("❌ WARNING: Found $n_allocations GC allocation calls!")
|
||||
end
|
||||
|
||||
if n_vector_ops > 0
|
||||
println("✅ SIMD vectorization detected!")
|
||||
end
|
||||
println()
|
||||
|
||||
# Print full LLVM (first 100 lines)
|
||||
println("Full LLVM IR (first 100 lines):")
|
||||
println("-"^80)
|
||||
llvm_lines = split(llvm_code, '\n')
|
||||
for (i, line) in enumerate(llvm_lines[1:min(100, length(llvm_lines))])
|
||||
println(line)
|
||||
end
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# 4. NATIVE ASSEMBLY ANALYSIS
|
||||
# ============================================================================
|
||||
println("4. NATIVE ASSEMBLY ANALYSIS")
|
||||
println("-"^80)
|
||||
println("Examining native assembly...")
|
||||
println()
|
||||
|
||||
io_native = IOBuffer()
|
||||
code_native(io_native, compute_deformation_gradient,
|
||||
typeof((X_nodes, u_nodes, dN_dξ, J, FiniteStrain())))
|
||||
native_code = String(take!(io_native))
|
||||
|
||||
# Count key assembly features
|
||||
n_movsd = count(r"movsd", native_code) # Scalar moves
|
||||
n_movapd = count(r"movapd", native_code) # Aligned packed moves
|
||||
n_movupd = count(r"movupd", native_code) # Unaligned packed moves
|
||||
n_mulpd = count(r"mulpd", native_code) # Packed multiply
|
||||
n_addpd = count(r"addpd", native_code) # Packed add
|
||||
n_call = count(r"call", native_code) # Function calls
|
||||
|
||||
println("Native Assembly Statistics:")
|
||||
println(" - Scalar moves (movsd): $n_movsd")
|
||||
println(" - Aligned packed moves (movapd): $n_movapd")
|
||||
println(" - Unaligned packed moves (movupd): $n_movupd")
|
||||
println(" - Packed multiplies (mulpd): $n_mulpd")
|
||||
println(" - Packed adds (addpd): $n_addpd")
|
||||
println(" - Function calls: $n_call")
|
||||
println()
|
||||
|
||||
if n_mulpd > 0 || n_addpd > 0
|
||||
println("✅ SSE/AVX SIMD instructions detected!")
|
||||
end
|
||||
|
||||
if n_call == 0
|
||||
println("✅ Fully inlined - no function calls!")
|
||||
else
|
||||
println("⚠️ Note: $n_call function calls detected (may include math library)")
|
||||
end
|
||||
println()
|
||||
|
||||
# Print full assembly (first 100 lines)
|
||||
println("Full Native Assembly (first 100 lines):")
|
||||
println("-"^80)
|
||||
native_lines = split(native_code, '\n')
|
||||
for (i, line) in enumerate(native_lines[1:min(100, length(native_lines))])
|
||||
println(line)
|
||||
end
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# 5. TYPE STABILITY ANALYSIS
|
||||
# ============================================================================
|
||||
println("5. TYPE STABILITY ANALYSIS")
|
||||
println("-"^80)
|
||||
println("Checking type stability with @code_warntype...")
|
||||
println()
|
||||
|
||||
io_warntype = IOBuffer()
|
||||
code_warntype(io_warntype, compute_deformation_gradient,
|
||||
typeof((X_nodes, u_nodes, dN_dξ, J, FiniteStrain())))
|
||||
warntype_output = String(take!(io_native))
|
||||
|
||||
# Check for type instabilities
|
||||
has_any = contains(warntype_output, "Any")
|
||||
has_union = contains(warntype_output, "Union{")
|
||||
|
||||
if has_any
|
||||
println("⚠️ WARNING: 'Any' types detected (type instability)")
|
||||
else
|
||||
println("✅ No 'Any' types detected!")
|
||||
end
|
||||
|
||||
if has_union
|
||||
println("⚠️ Note: Union types detected (may be intentional)")
|
||||
else
|
||||
println("✅ No Union types detected!")
|
||||
end
|
||||
println()
|
||||
|
||||
# Print warntype output (first 50 lines)
|
||||
println("@code_warntype output (first 50 lines):")
|
||||
println("-"^80)
|
||||
warntype_lines = split(warntype_output, '\n')
|
||||
for (i, line) in enumerate(warntype_lines[1:min(50, length(warntype_lines))])
|
||||
println(line)
|
||||
end
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# 6. SUMMARY
|
||||
# ============================================================================
|
||||
println("="^80)
|
||||
println("PERFORMANCE SUMMARY")
|
||||
println("="^80)
|
||||
println()
|
||||
println("✅ Implementation validated as:")
|
||||
println(" - Zero allocation (confirmed)")
|
||||
println(" - Type stable")
|
||||
println(" - SIMD optimized ($(n_vector_ops) vector ops in LLVM)")
|
||||
println(" - Median execution time: $(round(median_time_ns, digits=2)) ns")
|
||||
println()
|
||||
println("This implementation achieves the best possible performance for")
|
||||
println("deformation gradient computation in Julia.")
|
||||
println()
|
||||
println("="^80)
|
||||
@@ -1,396 +0,0 @@
|
||||
# ==============================================================================
|
||||
# ELEMENT IMMUTABILITY BENCHMARK
|
||||
# ==============================================================================
|
||||
#
|
||||
# Purpose: Demonstrate why immutable elements with type-stable fields are faster
|
||||
# than mutable elements with Dict-based fields, despite seeming
|
||||
# counterintuitive.
|
||||
#
|
||||
# Hypothesis: Immutable + type-stable >> Mutable + Dict
|
||||
#
|
||||
# What we measure:
|
||||
# 1. Field access time (reading)
|
||||
# 2. Field update time (writing)
|
||||
# 3. Memory allocations
|
||||
# 4. Assembly loop performance (realistic FEM workload)
|
||||
#
|
||||
# Expected results:
|
||||
# - Dict lookup: O(1) amortized, but ~100ns overhead per access
|
||||
# - Type-stable access: O(1), but ~1ns (inlined, no overhead)
|
||||
# - Immutable update: Allocates new struct, but compiler optimizes away
|
||||
# - Dict update: Mutates in-place, but loses type stability
|
||||
#
|
||||
# Conclusion: For FEM assembly (tight loops, millions of field accesses),
|
||||
# type stability dominates. Immutability enables GPU/HPC.
|
||||
#
|
||||
# ==============================================================================
|
||||
|
||||
using BenchmarkTools
|
||||
using Statistics
|
||||
|
||||
println("="^80)
|
||||
println("ELEMENT IMMUTABILITY BENCHMARK")
|
||||
println("="^80)
|
||||
println()
|
||||
println("Comparing two implementations of P2 Lagrange Tetrahedron (Tet10):")
|
||||
println(" 1. Mutable element with Dict-based fields (OLD API)")
|
||||
println(" 2. Immutable element with NamedTuple fields (NEW API)")
|
||||
println()
|
||||
println("Measuring: field access, field update, assembly loop")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# IMPLEMENTATION 1: Mutable Element with Dict-based Fields (OLD)
|
||||
# ==============================================================================
|
||||
|
||||
"""
|
||||
Mutable element: fields stored in Dict{Symbol,Any}
|
||||
- Pro: Can add/remove fields dynamically
|
||||
- Con: Type-unstable, Dict lookup overhead, no GPU support
|
||||
"""
|
||||
mutable struct MutableElement
|
||||
id::UInt
|
||||
connectivity::Vector{UInt}
|
||||
fields::Dict{Symbol,Any} # Type-unstable!
|
||||
end
|
||||
|
||||
function MutableElement(connectivity::Vector{UInt})
|
||||
return MutableElement(UInt(0), connectivity, Dict{Symbol,Any}())
|
||||
end
|
||||
|
||||
# Old-style update: mutate in-place
|
||||
function update_field!(elem::MutableElement, field_name::Symbol, value)
|
||||
elem.fields[field_name] = value
|
||||
return nothing
|
||||
end
|
||||
|
||||
# Old-style access: Dict lookup
|
||||
function get_field(elem::MutableElement, field_name::Symbol)
|
||||
return elem.fields[field_name]
|
||||
end
|
||||
|
||||
# ==============================================================================
|
||||
# IMPLEMENTATION 2: Immutable Element with NamedTuple Fields (NEW)
|
||||
# ==============================================================================
|
||||
|
||||
"""
|
||||
Immutable element: fields stored in NamedTuple
|
||||
- Pro: Type-stable, zero overhead access, GPU-compatible
|
||||
- Con: Cannot mutate, must create new element (but compiler optimizes!)
|
||||
"""
|
||||
struct ImmutableElement{F}
|
||||
id::UInt
|
||||
connectivity::NTuple{10,UInt} # Fixed size, stack-allocated
|
||||
fields::F # Type-stable! (NamedTuple)
|
||||
end
|
||||
|
||||
function ImmutableElement(connectivity::NTuple{10,UInt}, fields::NamedTuple)
|
||||
return ImmutableElement{typeof(fields)}(UInt(0), connectivity, fields)
|
||||
end
|
||||
|
||||
# New-style update: return new element (immutable)
|
||||
function update_field(elem::ImmutableElement, updates::NamedTuple)
|
||||
new_fields = merge(elem.fields, updates)
|
||||
return ImmutableElement(elem.connectivity, new_fields)
|
||||
end
|
||||
|
||||
# New-style access: direct field access (inlined!)
|
||||
function get_field(elem::ImmutableElement, field_name::Symbol)
|
||||
return getfield(elem.fields, field_name)
|
||||
end
|
||||
|
||||
# ==============================================================================
|
||||
# BENCHMARK 1: Field Access (Read Performance)
|
||||
# ==============================================================================
|
||||
|
||||
println("BENCHMARK 1: Field Access (Reading E, ν, ρ in tight loop)")
|
||||
println("-"^80)
|
||||
|
||||
# Setup test elements
|
||||
connectivity_vec = UInt.(1:10)
|
||||
connectivity_tuple = ntuple(i -> UInt(i), 10)
|
||||
|
||||
mutable_elem = MutableElement(connectivity_vec)
|
||||
update_field!(mutable_elem, :E, 210e9)
|
||||
update_field!(mutable_elem, :ν, 0.3)
|
||||
update_field!(mutable_elem, :ρ, 7850.0)
|
||||
|
||||
immutable_elem = ImmutableElement(connectivity_tuple, (E=210e9, ν=0.3, ρ=7850.0))
|
||||
|
||||
# Benchmark: Read fields 1000 times (simulating assembly loop)
|
||||
function read_fields_mutable(elem, n)
|
||||
sum_val = 0.0
|
||||
for _ in 1:n
|
||||
E = get_field(elem, :E)
|
||||
ν = get_field(elem, :ν)
|
||||
ρ = get_field(elem, :ρ)
|
||||
sum_val += E + ν + ρ
|
||||
end
|
||||
return sum_val
|
||||
end
|
||||
|
||||
function read_fields_immutable(elem, n)
|
||||
sum_val = 0.0
|
||||
for _ in 1:n
|
||||
E = get_field(elem, :E)
|
||||
ν = get_field(elem, :ν)
|
||||
ρ = get_field(elem, :ρ)
|
||||
sum_val += E + ν + ρ
|
||||
end
|
||||
return sum_val
|
||||
end
|
||||
|
||||
n_reads = 1000
|
||||
|
||||
println("Reading fields $n_reads times:")
|
||||
println()
|
||||
|
||||
t_mutable = @benchmark read_fields_mutable($mutable_elem, $n_reads)
|
||||
println("Mutable (Dict): ", minimum(t_mutable.times) / n_reads, " ns/read")
|
||||
println(" Median: ", median(t_mutable.times) / n_reads, " ns/read")
|
||||
println(" Allocs: ", t_mutable.allocs)
|
||||
|
||||
t_immutable = @benchmark read_fields_immutable($immutable_elem, $n_reads)
|
||||
println("Immutable (Tuple): ", minimum(t_immutable.times) / n_reads, " ns/read")
|
||||
println(" Median: ", median(t_immutable.times) / n_reads, " ns/read")
|
||||
println(" Allocs: ", t_immutable.allocs)
|
||||
|
||||
speedup_read = minimum(t_mutable.times) / minimum(t_immutable.times)
|
||||
println()
|
||||
println("Speedup: ", round(speedup_read, digits=1), "x faster")
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# BENCHMARK 2: Field Update (Write Performance)
|
||||
# ==============================================================================
|
||||
|
||||
println("BENCHMARK 2: Field Update (Updating temperature field)")
|
||||
println("-"^80)
|
||||
|
||||
# Benchmark: Update temperature field 100 times
|
||||
function update_temperature_mutable(elem, n)
|
||||
for i in 1:n
|
||||
update_field!(elem, :temperature, Float64(i) * 293.15)
|
||||
end
|
||||
return get_field(elem, :temperature)
|
||||
end
|
||||
|
||||
function update_temperature_immutable(elem, n)
|
||||
current = elem
|
||||
for i in 1:n
|
||||
current = update_field(current, (temperature=Float64(i) * 293.15,))
|
||||
end
|
||||
return get_field(current, :temperature)
|
||||
end
|
||||
|
||||
n_updates = 100
|
||||
|
||||
println("Updating temperature field $n_updates times:")
|
||||
println()
|
||||
|
||||
# Reset elements
|
||||
mutable_elem2 = MutableElement(connectivity_vec)
|
||||
update_field!(mutable_elem2, :E, 210e9)
|
||||
update_field!(mutable_elem2, :ν, 0.3)
|
||||
|
||||
immutable_elem2 = ImmutableElement(connectivity_tuple, (E=210e9, ν=0.3))
|
||||
|
||||
t_mutable_update = @benchmark update_temperature_mutable($mutable_elem2, $n_updates)
|
||||
println("Mutable (mutate): ", minimum(t_mutable_update.times) / n_updates, " ns/update")
|
||||
println(" Median: ", median(t_mutable_update.times) / n_updates, " ns/update")
|
||||
println(" Allocs: ", t_mutable_update.allocs)
|
||||
println(" Memory: ", t_mutable_update.memory, " bytes")
|
||||
|
||||
t_immutable_update = @benchmark update_temperature_immutable($immutable_elem2, $n_updates)
|
||||
println("Immutable (copy): ", minimum(t_immutable_update.times) / n_updates, " ns/update")
|
||||
println(" Median: ", median(t_immutable_update.times) / n_updates, " ns/update")
|
||||
println(" Allocs: ", t_immutable_update.allocs)
|
||||
println(" Memory: ", t_immutable_update.memory, " bytes")
|
||||
|
||||
println()
|
||||
println("Note: Immutable creates new structs, but compiler optimizes stack allocation")
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# BENCHMARK 3: Realistic Assembly Loop (FEM Workload)
|
||||
# ==============================================================================
|
||||
|
||||
println("BENCHMARK 3: Realistic FEM Assembly Loop")
|
||||
println("-"^80)
|
||||
println("Simulating element stiffness matrix assembly:")
|
||||
println(" - Read E, ν from element fields")
|
||||
println(" - Compute 10 Gauss integration points")
|
||||
println(" - Each point: read fields, compute B matrix, add to K")
|
||||
println()
|
||||
|
||||
# Simplified assembly kernel
|
||||
function assemble_stiffness_mutable(elem)
|
||||
E = get_field(elem, :E)
|
||||
ν = get_field(elem, :ν)
|
||||
|
||||
# Compute material matrix (simplified)
|
||||
λ = E * ν / ((1 + ν) * (1 - 2ν))
|
||||
μ = E / (2 * (1 + ν))
|
||||
|
||||
K = 0.0
|
||||
# Simulate 10 integration points
|
||||
for ip in 1:10
|
||||
# Simulate field reads at integration point
|
||||
E_ip = get_field(elem, :E)
|
||||
ν_ip = get_field(elem, :ν)
|
||||
|
||||
# Simplified stiffness contribution
|
||||
detJ = 1.0 + 0.1 * ip # Fake Jacobian
|
||||
weight = 0.1
|
||||
K += (λ + 2μ) * detJ * weight
|
||||
end
|
||||
|
||||
return K
|
||||
end
|
||||
|
||||
function assemble_stiffness_immutable(elem)
|
||||
E = get_field(elem, :E)
|
||||
ν = get_field(elem, :ν)
|
||||
|
||||
# Compute material matrix (simplified)
|
||||
λ = E * ν / ((1 + ν) * (1 - 2ν))
|
||||
μ = E / (2 * (1 + ν))
|
||||
|
||||
K = 0.0
|
||||
# Simulate 10 integration points
|
||||
for ip in 1:10
|
||||
# Simulate field reads at integration point
|
||||
E_ip = get_field(elem, :E)
|
||||
ν_ip = get_field(elem, :ν)
|
||||
|
||||
# Simplified stiffness contribution
|
||||
detJ = 1.0 + 0.1 * ip # Fake Jacobian
|
||||
weight = 0.1
|
||||
K += (λ + 2μ) * detJ * weight
|
||||
end
|
||||
|
||||
return K
|
||||
end
|
||||
|
||||
t_assembly_mutable = @benchmark assemble_stiffness_mutable($mutable_elem)
|
||||
t_assembly_immutable = @benchmark assemble_stiffness_immutable($immutable_elem)
|
||||
|
||||
println("Assembly time per element:")
|
||||
println()
|
||||
println("Mutable (Dict): ", minimum(t_assembly_mutable.times), " ns")
|
||||
println(" Median: ", median(t_assembly_mutable.times), " ns")
|
||||
println(" Allocs: ", t_assembly_mutable.allocs)
|
||||
|
||||
println("Immutable (Tuple): ", minimum(t_assembly_immutable.times), " ns")
|
||||
println(" Median: ", median(t_assembly_immutable.times), " ns")
|
||||
println(" Allocs: ", t_assembly_immutable.allocs)
|
||||
|
||||
speedup_assembly = minimum(t_assembly_mutable.times) / minimum(t_assembly_immutable.times)
|
||||
println()
|
||||
println("Speedup: ", round(speedup_assembly, digits=1), "x faster")
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# BENCHMARK 4: Large-Scale Mesh (1000 elements)
|
||||
# ==============================================================================
|
||||
|
||||
println("BENCHMARK 4: Large-Scale Assembly (1000 elements)")
|
||||
println("-"^80)
|
||||
|
||||
n_elements = 1000
|
||||
|
||||
# Create mesh
|
||||
mutable_mesh = [
|
||||
begin
|
||||
elem = MutableElement(UInt.(1:10) .+ UInt(i * 10))
|
||||
update_field!(elem, :E, 210e9)
|
||||
update_field!(elem, :ν, 0.3)
|
||||
elem
|
||||
end for i in 1:n_elements
|
||||
]
|
||||
|
||||
immutable_mesh = [
|
||||
begin
|
||||
conn = ntuple(j -> UInt(j + i * 10), 10)
|
||||
ImmutableElement(conn, (E=210e9, ν=0.3))
|
||||
end for i in 1:n_elements
|
||||
]
|
||||
|
||||
function assemble_mesh_mutable(mesh)
|
||||
K_total = 0.0
|
||||
for elem in mesh
|
||||
K_total += assemble_stiffness_mutable(elem)
|
||||
end
|
||||
return K_total
|
||||
end
|
||||
|
||||
function assemble_mesh_immutable(mesh)
|
||||
K_total = 0.0
|
||||
for elem in mesh
|
||||
K_total += assemble_stiffness_immutable(elem)
|
||||
end
|
||||
return K_total
|
||||
end
|
||||
|
||||
println("Assembling $n_elements elements:")
|
||||
println()
|
||||
|
||||
t_mesh_mutable = @benchmark assemble_mesh_mutable($mutable_mesh)
|
||||
println("Mutable (Dict): ", minimum(t_mesh_mutable.times) / 1e6, " ms")
|
||||
println(" Median: ", median(t_mesh_mutable.times) / 1e6, " ms")
|
||||
println(" Allocs: ", t_mesh_mutable.allocs)
|
||||
println(" Memory: ", t_mesh_mutable.memory / 1024, " KB")
|
||||
|
||||
t_mesh_immutable = @benchmark assemble_mesh_immutable($immutable_mesh)
|
||||
println("Immutable (Tuple): ", minimum(t_mesh_immutable.times) / 1e6, " ms")
|
||||
println(" Median: ", median(t_mesh_immutable.times) / 1e6, " ms")
|
||||
println(" Allocs: ", t_mesh_immutable.allocs)
|
||||
println(" Memory: ", t_mesh_immutable.memory / 1024, " KB")
|
||||
|
||||
speedup_mesh = minimum(t_mesh_mutable.times) / minimum(t_mesh_immutable.times)
|
||||
println()
|
||||
println("Speedup: ", round(speedup_mesh, digits=1), "x faster")
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# SUMMARY
|
||||
# ==============================================================================
|
||||
|
||||
println("="^80)
|
||||
println("SUMMARY")
|
||||
println("="^80)
|
||||
println()
|
||||
println("Key findings:")
|
||||
println()
|
||||
println("1. Field Access:")
|
||||
println(" - Type-stable (immutable) is ", round(speedup_read, digits=1), "x faster")
|
||||
println(" - Dict lookup: ~100-200ns overhead per access")
|
||||
println(" - NamedTuple: ~1ns (inlined, zero overhead)")
|
||||
println()
|
||||
println("2. Assembly Performance:")
|
||||
println(" - Single element: ", round(speedup_assembly, digits=1), "x faster")
|
||||
println(" - Large mesh: ", round(speedup_mesh, digits=1), "x faster")
|
||||
println()
|
||||
println("3. Memory:")
|
||||
println(" - Immutable elements: no allocations in hot path")
|
||||
println(" - Mutable elements: Dict overhead + dynamic dispatch")
|
||||
println()
|
||||
println("4. GPU/HPC Compatibility:")
|
||||
println(" - Immutable: ✓ All bits types, can transfer to GPU")
|
||||
println(" - Mutable: ✗ Pointers, heap allocations, no GPU support")
|
||||
println()
|
||||
println("CONCLUSION:")
|
||||
println("-"^80)
|
||||
println("Despite appearing counterintuitive, IMMUTABLE elements with type-stable")
|
||||
println("fields are SIGNIFICANTLY FASTER for FEM assembly. The key insight:")
|
||||
println()
|
||||
println(" • Dict lookup cost dominates in tight loops (millions of accesses)")
|
||||
println(" • Type stability enables compiler optimizations (inlining, SIMD)")
|
||||
println(" • Immutability enables GPU/HPC parallelization (no race conditions)")
|
||||
println(" • Modern compilers optimize away struct copies on stack")
|
||||
println()
|
||||
println("For FEM with millions of field accesses per assembly, type stability")
|
||||
println("is the critical factor. Immutability is a small price for 10-100x speedup.")
|
||||
println()
|
||||
println("="^80)
|
||||
@@ -1,334 +0,0 @@
|
||||
#!/usr/bin/env julia
|
||||
#
|
||||
# Benchmark: Dict{String,Any} vs Type-Stable Field Storage
|
||||
#
|
||||
# This benchmark validates the performance claims in:
|
||||
# docs/book/zero_allocation_fields.md
|
||||
#
|
||||
# Expected results:
|
||||
# - Constant field access: 50× faster, 0 allocations
|
||||
# - Nodal field access: 50× faster, 0 allocations
|
||||
# - Interpolation: 16× faster with zero allocations (cached)
|
||||
#
|
||||
|
||||
using BenchmarkTools
|
||||
using LinearAlgebra
|
||||
|
||||
# ============================================================================
|
||||
# Field Type Definitions (from proposal)
|
||||
# ============================================================================
|
||||
|
||||
abstract type AbstractField{T} end
|
||||
|
||||
struct ConstantField{T} <: AbstractField{T}
|
||||
value::T
|
||||
end
|
||||
|
||||
struct NodalField{T} <: AbstractField{T}
|
||||
values::Matrix{T} # N_components × N_nodes
|
||||
end
|
||||
|
||||
# Accessors
|
||||
@inline value(f::ConstantField) = f.value
|
||||
@inline function value(f::NodalField, node_ids::AbstractVector{Int})
|
||||
return @view f.values[:, node_ids]
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Mock Element (simplified for benchmarking)
|
||||
# ============================================================================
|
||||
|
||||
struct MockElement
|
||||
connectivity::Vector{Int}
|
||||
end
|
||||
|
||||
function mock_eval_basis(x::Vector{Float64})
|
||||
# Mock basis function values for 8-node element
|
||||
return [0.1, 0.15, 0.05, 0.1, 0.2, 0.15, 0.15, 0.1]
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Setup: OLD (Dict-based) vs NEW (Typed)
|
||||
# ============================================================================
|
||||
|
||||
const OLD_FIELDS = Dict{String,Any}(
|
||||
"youngs_modulus" => 210e3,
|
||||
"poissons_ratio" => 0.3,
|
||||
"displacement" => zeros(3, 8),
|
||||
)
|
||||
|
||||
const NEW_FIELDS = (
|
||||
youngs_modulus=ConstantField(210e3),
|
||||
poissons_ratio=ConstantField(0.3),
|
||||
displacement=NodalField(zeros(3, 8)),
|
||||
)
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark 1: Constant Field Access
|
||||
# ============================================================================
|
||||
|
||||
println("="^70)
|
||||
println("Benchmark 1: Constant Field Access")
|
||||
println("="^70)
|
||||
|
||||
println("\nOLD (Dict{String,Any}):")
|
||||
old_constant = @benchmark $OLD_FIELDS["youngs_modulus"]
|
||||
display(old_constant)
|
||||
|
||||
println("\nNEW (ConstantField):")
|
||||
new_constant = @benchmark value($NEW_FIELDS.youngs_modulus)
|
||||
display(new_constant)
|
||||
|
||||
old_time_1 = median(old_constant).time
|
||||
new_time_1 = median(new_constant).time
|
||||
speedup_1 = old_time_1 / new_time_1
|
||||
allocs_old_1 = median(old_constant).allocs
|
||||
allocs_new_1 = median(new_constant).allocs
|
||||
|
||||
println("\n📊 Results:")
|
||||
println(" OLD: $(round(old_time_1, digits=1)) ns, $(allocs_old_1) allocations")
|
||||
println(" NEW: $(round(new_time_1, digits=1)) ns, $(allocs_new_1) allocations")
|
||||
println(" Speedup: $(round(speedup_1, digits=1))×")
|
||||
println(" Allocation reduction: $(allocs_old_1 - allocs_new_1)")
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark 2: Nodal Field Access
|
||||
# ============================================================================
|
||||
|
||||
println("\n" * "="^70)
|
||||
println("Benchmark 2: Nodal Field Access (4 nodes)")
|
||||
println("="^70)
|
||||
|
||||
node_ids = [1, 2, 3, 4]
|
||||
|
||||
println("\nOLD (Dict with Array slicing):")
|
||||
old_nodal = @benchmark $OLD_FIELDS["displacement"][:, $node_ids]
|
||||
display(old_nodal)
|
||||
|
||||
println("\nNEW (NodalField with @view):")
|
||||
new_nodal = @benchmark value($NEW_FIELDS.displacement, $node_ids)
|
||||
display(new_nodal)
|
||||
|
||||
old_time_2 = median(old_nodal).time
|
||||
new_time_2 = median(new_nodal).time
|
||||
speedup_2 = old_time_2 / new_time_2
|
||||
allocs_old_2 = median(old_nodal).allocs
|
||||
allocs_new_2 = median(new_nodal).allocs
|
||||
|
||||
println("\n📊 Results:")
|
||||
println(" OLD: $(round(old_time_2, digits=1)) ns, $(allocs_old_2) allocations")
|
||||
println(" NEW: $(round(new_time_2, digits=1)) ns, $(allocs_new_2) allocations")
|
||||
println(" Speedup: $(round(speedup_2, digits=1))×")
|
||||
println(" Allocation reduction: $(allocs_old_2 - allocs_new_2)")
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark 3: Interpolation (Without Cache)
|
||||
# ============================================================================
|
||||
|
||||
println("\n" * "="^70)
|
||||
println("Benchmark 3: Spatial Interpolation (No Cache)")
|
||||
println("="^70)
|
||||
|
||||
element = MockElement([1, 2, 3, 4, 5, 6, 7, 8])
|
||||
x = [0.1, 0.2, 0.3]
|
||||
N = mock_eval_basis(x)
|
||||
|
||||
function interpolate_old(element, N, fields_dict)
|
||||
u = fields_dict["displacement"] # Type: Any
|
||||
result = zeros(3)
|
||||
for i in 1:length(N)
|
||||
result .+= N[i] .* u[:, element.connectivity[i]]
|
||||
end
|
||||
return result
|
||||
end
|
||||
|
||||
function interpolate_new(element, N, fields)
|
||||
u_nodal = value(fields.displacement, element.connectivity)
|
||||
result = zeros(3)
|
||||
for i in 1:length(N)
|
||||
result .+= N[i] .* @view u_nodal[:, i]
|
||||
end
|
||||
return result
|
||||
end
|
||||
|
||||
println("\nOLD (Dict-based):")
|
||||
old_interp = @benchmark interpolate_old($element, $N, $OLD_FIELDS)
|
||||
display(old_interp)
|
||||
|
||||
println("\nNEW (Typed fields):")
|
||||
new_interp = @benchmark interpolate_new($element, $N, $NEW_FIELDS)
|
||||
display(new_interp)
|
||||
|
||||
old_time_3 = median(old_interp).time
|
||||
new_time_3 = median(new_interp).time
|
||||
speedup_3 = old_time_3 / new_time_3
|
||||
allocs_old_3 = median(old_interp).allocs
|
||||
allocs_new_3 = median(new_interp).allocs
|
||||
|
||||
println("\n📊 Results:")
|
||||
println(" OLD: $(round(old_time_3/1000, digits=1)) μs, $(allocs_old_3) allocations")
|
||||
println(" NEW: $(round(new_time_3, digits=1)) ns, $(allocs_new_3) allocations")
|
||||
println(" Speedup: $(round(speedup_3, digits=1))×")
|
||||
println(" Allocation reduction: $(allocs_old_3 - allocs_new_3)")
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark 4: Interpolation (With Cache - Zero Allocation Target)
|
||||
# ============================================================================
|
||||
|
||||
println("\n" * "="^70)
|
||||
println("Benchmark 4: Spatial Interpolation (WITH Cache)")
|
||||
println("="^70)
|
||||
|
||||
struct InterpolationCache
|
||||
result::Vector{Float64}
|
||||
end
|
||||
|
||||
function interpolate_cached!(cache, element, N, fields)
|
||||
u_nodal = value(fields.displacement, element.connectivity)
|
||||
fill!(cache.result, 0.0)
|
||||
for i in eachindex(N)
|
||||
cache.result .+= N[i] .* @view u_nodal[:, i]
|
||||
end
|
||||
return cache.result
|
||||
end
|
||||
|
||||
cache = InterpolationCache(zeros(3))
|
||||
|
||||
println("\nNEW (Cached - Zero Allocation Target):")
|
||||
cached_interp = @benchmark interpolate_cached!($cache, $element, $N, $NEW_FIELDS)
|
||||
display(cached_interp)
|
||||
|
||||
cached_time = median(cached_interp).time
|
||||
cached_allocs = median(cached_interp).allocs
|
||||
speedup_4 = old_time_3 / cached_time
|
||||
|
||||
println("\n📊 Results:")
|
||||
println(" Cached: $(round(cached_time, digits=1)) ns, $(cached_allocs) allocations")
|
||||
println(" Speedup vs OLD: $(round(speedup_4, digits=1))×")
|
||||
println(" Zero allocation target: $(cached_allocs == 0 ? "✅ MET" : "❌ FAILED")")
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark 5: Assembly Loop (1000 elements)
|
||||
# ============================================================================
|
||||
|
||||
println("\n" * "="^70)
|
||||
println("Benchmark 5: Assembly Loop (1000 elements)")
|
||||
println("="^70)
|
||||
|
||||
n_elements = 1000
|
||||
elements = [MockElement(collect(1:8)) for _ in 1:n_elements]
|
||||
|
||||
function assemble_old_style(elements, fields_dict)
|
||||
total = 0.0
|
||||
for element in elements
|
||||
E = fields_dict["youngs_modulus"] # Type-unstable access
|
||||
ν = fields_dict["poissons_ratio"]
|
||||
|
||||
# Mock stiffness computation
|
||||
K_local = E * (1 - ν^2)
|
||||
total += K_local
|
||||
end
|
||||
return total
|
||||
end
|
||||
|
||||
function assemble_new_style(elements, fields)
|
||||
E = value(fields.youngs_modulus) # Type-stable access (once)
|
||||
ν = value(fields.poissons_ratio)
|
||||
|
||||
total = 0.0
|
||||
for element in elements
|
||||
# Mock stiffness computation
|
||||
K_local = E * (1 - ν^2)
|
||||
total += K_local
|
||||
end
|
||||
return total
|
||||
end
|
||||
|
||||
println("\nOLD (Dict access in loop):")
|
||||
old_assembly = @benchmark assemble_old_style($elements, $OLD_FIELDS)
|
||||
display(old_assembly)
|
||||
|
||||
println("\nNEW (Typed fields, hoisted access):")
|
||||
new_assembly = @benchmark assemble_new_style($elements, $NEW_FIELDS)
|
||||
display(new_assembly)
|
||||
|
||||
old_time_5 = median(old_assembly).time
|
||||
new_time_5 = median(new_assembly).time
|
||||
speedup_5 = old_time_5 / new_time_5
|
||||
allocs_old_5 = median(old_assembly).allocs
|
||||
allocs_new_5 = median(new_assembly).allocs
|
||||
|
||||
println("\n📊 Results:")
|
||||
println(" OLD: $(round(old_time_5/1000, digits=1)) μs, $(allocs_old_5) allocations")
|
||||
println(" NEW: $(round(new_time_5/1000, digits=1)) μs, $(allocs_new_5) allocations")
|
||||
println(" Speedup: $(round(speedup_5, digits=1))×")
|
||||
println(" Allocation reduction: $(allocs_old_5 - allocs_new_5)")
|
||||
|
||||
# ============================================================================
|
||||
# Summary and Validation
|
||||
# ============================================================================
|
||||
|
||||
println("\n" * "="^70)
|
||||
println("SUMMARY - Validation Against Claims")
|
||||
println("="^70)
|
||||
|
||||
validation_passed = true
|
||||
|
||||
# Claim 1: Constant field access should be ~50× faster, 0 allocations
|
||||
println("\n1. Constant Field Access:")
|
||||
println(" Claimed: ~50× faster, 0 allocations")
|
||||
println(" Actual: $(round(speedup_1, digits=1))× faster, $(allocs_new_1) allocations")
|
||||
if speedup_1 >= 10 && allocs_new_1 == 0
|
||||
println(" Status: ✅ VALIDATED ($(round(speedup_1, digits=1))× > 10× threshold)")
|
||||
else
|
||||
println(" Status: ⚠️ PARTIAL (speedup or allocation target not met)")
|
||||
validation_passed = false
|
||||
end
|
||||
|
||||
# Claim 2: Nodal field access should be ~50× faster, 0 allocations
|
||||
println("\n2. Nodal Field Access:")
|
||||
println(" Claimed: ~50× faster, 0 allocations")
|
||||
println(" Actual: $(round(speedup_2, digits=1))× faster, $(allocs_new_2) allocations")
|
||||
if speedup_2 >= 10 && allocs_new_2 == 0
|
||||
println(" Status: ✅ VALIDATED ($(round(speedup_2, digits=1))× > 10× threshold)")
|
||||
else
|
||||
println(" Status: ⚠️ PARTIAL (speedup or allocation target not met)")
|
||||
validation_passed = false
|
||||
end
|
||||
|
||||
# Claim 3: Cached interpolation should be ~16× faster, 0 allocations
|
||||
println("\n3. Interpolation (Cached):")
|
||||
println(" Claimed: ~16× faster, 0 allocations")
|
||||
println(" Actual: $(round(speedup_4, digits=1))× faster, $(cached_allocs) allocations")
|
||||
if speedup_4 >= 10 && cached_allocs == 0
|
||||
println(" Status: ✅ VALIDATED ($(round(speedup_4, digits=1))× > 10× threshold)")
|
||||
else
|
||||
println(" Status: ⚠️ PARTIAL (speedup or allocation target not met)")
|
||||
validation_passed = false
|
||||
end
|
||||
|
||||
# Claim 4: Assembly should be 10-100× faster
|
||||
println("\n4. Assembly Loop:")
|
||||
println(" Claimed: 10-100× faster")
|
||||
println(" Actual: $(round(speedup_5, digits=1))× faster")
|
||||
if speedup_5 >= 10
|
||||
println(" Status: ✅ VALIDATED ($(round(speedup_5, digits=1))× > 10× threshold)")
|
||||
else
|
||||
println(" Status: ⚠️ PARTIAL (speedup target not met)")
|
||||
validation_passed = false
|
||||
end
|
||||
|
||||
println("\n" * "="^70)
|
||||
if validation_passed
|
||||
println("✅ ALL PERFORMANCE CLAIMS VALIDATED")
|
||||
else
|
||||
println("⚠️ SOME CLAIMS NOT FULLY VALIDATED (but likely still significant improvement)")
|
||||
end
|
||||
println("="^70)
|
||||
|
||||
println("\nKey Insights:")
|
||||
println(" • Type stability (NamedTuple) eliminates runtime dispatch")
|
||||
println(" • Zero allocations achieved with @view and pre-allocated caches")
|
||||
println(" • Hoisting invariant access out of loops provides massive speedup")
|
||||
println(" • The combination gives 10-100× speedup in realistic scenarios")
|
||||
println("\n✅ This validates the NamedTuple + typed fields design for v1.0")
|
||||
@@ -1,468 +0,0 @@
|
||||
"""
|
||||
GPU State Management Benchmark
|
||||
|
||||
Demonstrates Strategy 1 (immutable elements) vs Strategy 2 (separate mutable state)
|
||||
and validates memory coalescing patterns on actual GPU hardware.
|
||||
|
||||
Run with:
|
||||
julia --project=. benchmarks/gpu_state_management_benchmark.jl
|
||||
"""
|
||||
|
||||
using CUDA
|
||||
using Tensors
|
||||
using BenchmarkTools
|
||||
using Printf
|
||||
|
||||
# Check GPU availability
|
||||
if !CUDA.functional()
|
||||
error("CUDA not available! This benchmark requires a CUDA-capable GPU.")
|
||||
end
|
||||
|
||||
println("GPU Device: $(CUDA.device())")
|
||||
println("GPU Memory: $(CUDA.name(CUDA.device())) - $(round(CUDA.total_memory()/1e9, digits=1)) GB")
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 1: Immutable Elements (Array of Structs - AoS)
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Strategy 1: Element contains its own state (immutable).
|
||||
Update creates new element (allocation + copy).
|
||||
"""
|
||||
struct Element_Strategy1{T}
|
||||
connectivity::NTuple{8,Int32}
|
||||
material_id::Int32
|
||||
# State (plastic strain, hardening)
|
||||
ε_p::SymmetricTensor{2,3,T,6}
|
||||
α::T
|
||||
end
|
||||
|
||||
"""
|
||||
Update state for Strategy 1 (returns new element - allocation!).
|
||||
"""
|
||||
function update_element_strategy1(elem::Element_Strategy1{T}, Δε_p, Δα) where T
|
||||
return Element_Strategy1(
|
||||
elem.connectivity,
|
||||
elem.material_id,
|
||||
elem.ε_p + Δε_p,
|
||||
elem.α + Δα
|
||||
)
|
||||
end
|
||||
|
||||
"""
|
||||
CPU kernel: Update all elements (Strategy 1).
|
||||
"""
|
||||
function update_elements_strategy1_cpu!(
|
||||
elements::Vector{Element_Strategy1{T}},
|
||||
strain_increments::Vector{SymmetricTensor{2,3,T,6}},
|
||||
hardening_increments::Vector{T}
|
||||
) where T
|
||||
n = length(elements)
|
||||
for i in 1:n
|
||||
elements[i] = update_element_strategy1(
|
||||
elements[i],
|
||||
strain_increments[i],
|
||||
hardening_increments[i]
|
||||
)
|
||||
end
|
||||
end
|
||||
|
||||
"""
|
||||
GPU kernel: Update all elements (Strategy 1).
|
||||
|
||||
Problem: Each thread accesses scattered memory (pointer chasing).
|
||||
"""
|
||||
function update_elements_strategy1_kernel!(
|
||||
elements::CuDeviceVector{Element_Strategy1{T}},
|
||||
strain_increments::CuDeviceVector{SymmetricTensor{2,3,T,6}},
|
||||
hardening_increments::CuDeviceVector{T}
|
||||
) where T
|
||||
i = (blockIdx().x - 1) * blockDim().x + threadIdx().x
|
||||
|
||||
if i <= length(elements)
|
||||
elem = elements[i] # Non-coalesced read!
|
||||
Δε_p = strain_increments[i]
|
||||
Δα = hardening_increments[i]
|
||||
|
||||
# Update (creates new element - allocation on GPU!)
|
||||
new_elem = Element_Strategy1(
|
||||
elem.connectivity,
|
||||
elem.material_id,
|
||||
elem.ε_p + Δε_p,
|
||||
elem.α + Δα
|
||||
)
|
||||
|
||||
elements[i] = new_elem # Non-coalesced write!
|
||||
end
|
||||
|
||||
return nothing
|
||||
end
|
||||
|
||||
function update_elements_strategy1_gpu!(
|
||||
elements::CuVector{Element_Strategy1{T}},
|
||||
strain_increments::CuVector{SymmetricTensor{2,3,T,6}},
|
||||
hardening_increments::CuVector{T}
|
||||
) where T
|
||||
n = length(elements)
|
||||
threads = 256
|
||||
blocks = cld(n, threads)
|
||||
|
||||
@cuda threads = threads blocks = blocks update_elements_strategy1_kernel!(
|
||||
elements, strain_increments, hardening_increments
|
||||
)
|
||||
CUDA.synchronize()
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Strategy 2: Separate Mutable State (Structure of Arrays - SoA)
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Strategy 2: Geometry is immutable, state is separate and mutable.
|
||||
"""
|
||||
struct ElementGeometry
|
||||
connectivity::NTuple{8,Int32}
|
||||
material_id::Int32
|
||||
end
|
||||
|
||||
"""
|
||||
Mutable state storage (flat arrays for GPU coalescing).
|
||||
"""
|
||||
mutable struct AssemblyState{T,VecT}
|
||||
# Plastic strain (Voigt notation: 6 components per state)
|
||||
ε_p_flat::VecT # [N_states × 6]
|
||||
|
||||
# Hardening variable (1 component per state)
|
||||
α_flat::VecT # [N_states]
|
||||
|
||||
n_states::Int
|
||||
end
|
||||
|
||||
function AssemblyState{T}(n_states::Int) where T
|
||||
return AssemblyState{T,Vector{T}}(
|
||||
zeros(T, n_states * 6),
|
||||
zeros(T, n_states),
|
||||
n_states
|
||||
)
|
||||
end
|
||||
|
||||
"""
|
||||
CPU kernel: Update state (Strategy 2 - in-place!).
|
||||
"""
|
||||
function update_state_strategy2_cpu!(
|
||||
state::AssemblyState{T,Vector{T}},
|
||||
strain_increments::Vector{SymmetricTensor{2,3,T,6}},
|
||||
hardening_increments::Vector{T}
|
||||
) where T
|
||||
n = state.n_states
|
||||
|
||||
for i in 1:n
|
||||
# Flat indexing (cache-friendly!)
|
||||
offset = (i - 1) * 6
|
||||
|
||||
Δε_p = strain_increments[i]
|
||||
|
||||
# Update in-place (no allocation!)
|
||||
state.ε_p_flat[offset+1] += Δε_p[1, 1]
|
||||
state.ε_p_flat[offset+2] += Δε_p[2, 2]
|
||||
state.ε_p_flat[offset+3] += Δε_p[3, 3]
|
||||
state.ε_p_flat[offset+4] += Δε_p[1, 2]
|
||||
state.ε_p_flat[offset+5] += Δε_p[1, 3]
|
||||
state.ε_p_flat[offset+6] += Δε_p[2, 3]
|
||||
|
||||
state.α_flat[i] += hardening_increments[i]
|
||||
end
|
||||
end
|
||||
|
||||
"""
|
||||
GPU kernel: Update state (Strategy 2).
|
||||
|
||||
Advantage: Coalesced memory access!
|
||||
- Thread 0 accesses state.ε_p_flat[0:5]
|
||||
- Thread 1 accesses state.ε_p_flat[6:11]
|
||||
- Thread 2 accesses state.ε_p_flat[12:17]
|
||||
All consecutive in memory!
|
||||
"""
|
||||
function update_state_strategy2_kernel!(
|
||||
ε_p_flat::CuDeviceVector{T},
|
||||
α_flat::CuDeviceVector{T},
|
||||
strain_increments_flat::CuDeviceVector{T},
|
||||
hardening_increments::CuDeviceVector{T},
|
||||
n_states::Int
|
||||
) where T
|
||||
i = (blockIdx().x - 1) * blockDim().x + threadIdx().x
|
||||
|
||||
if i <= n_states
|
||||
# Flat indexing (coalesced access!)
|
||||
offset = (i - 1) * 6
|
||||
strain_offset = (i - 1) * 6
|
||||
|
||||
# Update plastic strain (6 consecutive reads/writes)
|
||||
ε_p_flat[offset+1] += strain_increments_flat[strain_offset+1]
|
||||
ε_p_flat[offset+2] += strain_increments_flat[strain_offset+2]
|
||||
ε_p_flat[offset+3] += strain_increments_flat[strain_offset+3]
|
||||
ε_p_flat[offset+4] += strain_increments_flat[strain_offset+4]
|
||||
ε_p_flat[offset+5] += strain_increments_flat[strain_offset+5]
|
||||
ε_p_flat[offset+6] += strain_increments_flat[strain_offset+6]
|
||||
|
||||
# Update hardening (1 read/write)
|
||||
α_flat[i] += hardening_increments[i]
|
||||
end
|
||||
|
||||
return nothing
|
||||
end
|
||||
|
||||
function update_state_strategy2_gpu!(
|
||||
state_gpu::AssemblyState{T,<:CuVector{T}},
|
||||
strain_increments_flat::CuVector{T},
|
||||
hardening_increments::CuVector{T}
|
||||
) where T
|
||||
n = state_gpu.n_states
|
||||
threads = 256
|
||||
blocks = cld(n, threads)
|
||||
|
||||
@cuda threads = threads blocks = blocks update_state_strategy2_kernel!(
|
||||
state_gpu.ε_p_flat,
|
||||
state_gpu.α_flat,
|
||||
strain_increments_flat,
|
||||
hardening_increments,
|
||||
n
|
||||
)
|
||||
CUDA.synchronize()
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark Setup
|
||||
# ============================================================================
|
||||
|
||||
function setup_benchmark(n_elements::Int)
|
||||
T = Float64
|
||||
|
||||
# Create random strain increments
|
||||
Δε_p_tensors = [SymmetricTensor{2,3}((
|
||||
rand(T) * 1e-5,
|
||||
rand(T) * 1e-5,
|
||||
rand(T) * 1e-5,
|
||||
rand(T) * 1e-6,
|
||||
rand(T) * 1e-6,
|
||||
rand(T) * 1e-6
|
||||
)) for _ in 1:n_elements]
|
||||
|
||||
Δα = rand(T, n_elements) .* 1e-5
|
||||
|
||||
# Strategy 1: Array of immutable elements
|
||||
elements_s1 = [Element_Strategy1(
|
||||
ntuple(j -> Int32(j), 8),
|
||||
Int32(1),
|
||||
zero(SymmetricTensor{2,3,T}),
|
||||
zero(T)
|
||||
) for _ in 1:n_elements]
|
||||
|
||||
# Strategy 2: Separate geometry and state
|
||||
geometry_s2 = [ElementGeometry(
|
||||
ntuple(j -> Int32(j), 8),
|
||||
Int32(1)
|
||||
) for _ in 1:n_elements]
|
||||
|
||||
state_s2 = AssemblyState{T}(n_elements)
|
||||
|
||||
return Δε_p_tensors, Δα, elements_s1, geometry_s2, state_s2
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# CPU Benchmarks
|
||||
# ============================================================================
|
||||
|
||||
function benchmark_cpu(n_elements::Int)
|
||||
println("="^70)
|
||||
println("CPU Benchmark: $n_elements elements")
|
||||
println("="^70)
|
||||
|
||||
Δε_p, Δα, elements_s1, geometry_s2, state_s2 = setup_benchmark(n_elements)
|
||||
|
||||
# Strategy 1: Update immutable elements
|
||||
println("\n📊 Strategy 1 (Immutable Elements - AoS):")
|
||||
elements_s1_copy = copy(elements_s1)
|
||||
t1 = @belapsed update_elements_strategy1_cpu!(
|
||||
$elements_s1_copy, $Δε_p, $Δα
|
||||
) samples = 10
|
||||
|
||||
println(" Time: $(round(t1 * 1000, digits=3)) ms")
|
||||
println(" Bandwidth: N/A (CPU cache)")
|
||||
|
||||
# Check allocations
|
||||
allocs = @allocated update_elements_strategy1_cpu!(elements_s1_copy, Δε_p, Δα)
|
||||
println(" Allocations: $(allocs) bytes ($(allocs ÷ n_elements) bytes/element)")
|
||||
|
||||
# Strategy 2: Update mutable state
|
||||
println("\n📊 Strategy 2 (Separate State - SoA):")
|
||||
state_s2_copy = deepcopy(state_s2)
|
||||
t2 = @belapsed update_state_strategy2_cpu!(
|
||||
$state_s2_copy, $Δε_p, $Δα
|
||||
) samples = 10
|
||||
|
||||
println(" Time: $(round(t2 * 1000, digits=3)) ms")
|
||||
println(" Bandwidth: N/A (CPU cache)")
|
||||
|
||||
# Check allocations
|
||||
allocs2 = @allocated update_state_strategy2_cpu!(state_s2_copy, Δε_p, Δα)
|
||||
println(" Allocations: $(allocs2) bytes")
|
||||
|
||||
# Speedup
|
||||
speedup = t1 / t2
|
||||
println("\n✅ CPU Speedup (Strategy 2 / Strategy 1): $(round(speedup, digits=2))×")
|
||||
|
||||
println()
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# GPU Benchmarks
|
||||
# ============================================================================
|
||||
|
||||
function benchmark_gpu(n_elements::Int)
|
||||
println("="^70)
|
||||
println("GPU Benchmark: $n_elements elements")
|
||||
println("="^70)
|
||||
|
||||
T = Float64
|
||||
Δε_p, Δα, elements_s1, geometry_s2, state_s2 = setup_benchmark(n_elements)
|
||||
|
||||
# ========================================================================
|
||||
# Strategy 1: GPU
|
||||
# ========================================================================
|
||||
println("\n📊 Strategy 1 (Immutable Elements - AoS on GPU):")
|
||||
|
||||
# Transfer to GPU
|
||||
elements_s1_gpu = CuArray(elements_s1)
|
||||
Δε_p_gpu = CuArray(Δε_p)
|
||||
Δα_gpu = CuArray(Δα)
|
||||
|
||||
# Warmup
|
||||
update_elements_strategy1_gpu!(elements_s1_gpu, Δε_p_gpu, Δα_gpu)
|
||||
|
||||
# Benchmark
|
||||
t1_gpu = CUDA.@elapsed begin
|
||||
update_elements_strategy1_gpu!(elements_s1_gpu, Δε_p_gpu, Δα_gpu)
|
||||
end
|
||||
|
||||
println(" Time: $(round(t1_gpu * 1000, digits=3)) ms")
|
||||
|
||||
# Estimate bandwidth (reading + writing entire element)
|
||||
bytes_per_elem = sizeof(Element_Strategy1{T})
|
||||
total_bytes = bytes_per_elem * n_elements * 2 # Read + write
|
||||
bandwidth_s1 = total_bytes / t1_gpu / 1e9
|
||||
println(" Bandwidth: $(round(bandwidth_s1, digits=1)) GB/s")
|
||||
|
||||
# ========================================================================
|
||||
# Strategy 2: GPU
|
||||
# ========================================================================
|
||||
println("\n📊 Strategy 2 (Separate State - SoA on GPU):")
|
||||
|
||||
# Transfer to GPU (flat arrays!)
|
||||
state_s2_gpu = AssemblyState{T,CuVector{T}}(
|
||||
CuArray(state_s2.ε_p_flat),
|
||||
CuArray(state_s2.α_flat),
|
||||
state_s2.n_states
|
||||
)
|
||||
|
||||
# Flatten strain increments for GPU
|
||||
Δε_p_flat = zeros(T, n_elements * 6)
|
||||
for i in 1:n_elements
|
||||
offset = (i - 1) * 6
|
||||
ε = Δε_p[i]
|
||||
Δε_p_flat[offset+1] = ε[1, 1]
|
||||
Δε_p_flat[offset+2] = ε[2, 2]
|
||||
Δε_p_flat[offset+3] = ε[3, 3]
|
||||
Δε_p_flat[offset+4] = ε[1, 2]
|
||||
Δε_p_flat[offset+5] = ε[1, 3]
|
||||
Δε_p_flat[offset+6] = ε[2, 3]
|
||||
end
|
||||
|
||||
Δε_p_flat_gpu = CuArray(Δε_p_flat)
|
||||
Δα_flat_gpu = CuArray(Δα)
|
||||
|
||||
# Warmup
|
||||
update_state_strategy2_gpu!(state_s2_gpu, Δε_p_flat_gpu, Δα_flat_gpu)
|
||||
|
||||
# Benchmark
|
||||
t2_gpu = CUDA.@elapsed begin
|
||||
update_state_strategy2_gpu!(state_s2_gpu, Δε_p_flat_gpu, Δα_flat_gpu)
|
||||
end
|
||||
|
||||
println(" Time: $(round(t2_gpu * 1000, digits=3)) ms")
|
||||
|
||||
# Estimate bandwidth (only state data, not geometry!)
|
||||
bytes_per_state = 6 * sizeof(T) + sizeof(T) # 6 strain + 1 hardening
|
||||
total_bytes_s2 = bytes_per_state * n_elements * 2 # Read + write
|
||||
bandwidth_s2 = total_bytes_s2 / t2_gpu / 1e9
|
||||
println(" Bandwidth: $(round(bandwidth_s2, digits=1)) GB/s")
|
||||
|
||||
# ========================================================================
|
||||
# Comparison
|
||||
# ========================================================================
|
||||
speedup = t1_gpu / t2_gpu
|
||||
bandwidth_ratio = bandwidth_s2 / bandwidth_s1
|
||||
|
||||
println("\n✅ GPU Speedup (Strategy 2 / Strategy 1): $(round(speedup, digits=2))×")
|
||||
println("✅ Bandwidth Improvement: $(round(bandwidth_ratio, digits=2))×")
|
||||
println(" Strategy 1: $(round(bandwidth_s1, digits=1)) GB/s (non-coalesced)")
|
||||
println(" Strategy 2: $(round(bandwidth_s2, digits=1)) GB/s (coalesced)")
|
||||
|
||||
# Theoretical peak (example: RTX 4090 = ~1000 GB/s)
|
||||
gpu_name = CUDA.name(CUDA.device())
|
||||
println("\n💡 GPU Memory Bandwidth:")
|
||||
println(" Achieved: $(round(bandwidth_s2, digits=1)) GB/s")
|
||||
println(" Device: $gpu_name")
|
||||
|
||||
println()
|
||||
|
||||
# Cleanup
|
||||
CUDA.unsafe_free!(elements_s1_gpu)
|
||||
CUDA.unsafe_free!(Δε_p_gpu)
|
||||
CUDA.unsafe_free!(Δα_gpu)
|
||||
CUDA.unsafe_free!(state_s2_gpu.ε_p_flat)
|
||||
CUDA.unsafe_free!(state_s2_gpu.α_flat)
|
||||
CUDA.unsafe_free!(Δε_p_flat_gpu)
|
||||
CUDA.unsafe_free!(Δα_flat_gpu)
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Main Benchmark
|
||||
# ============================================================================
|
||||
|
||||
function main()
|
||||
println("\n" * "=" * 70)
|
||||
println("GPU State Management Strategy Benchmark")
|
||||
println("=" * 70)
|
||||
println()
|
||||
|
||||
# Test sizes
|
||||
sizes = [10_000, 100_000, 1_000_000]
|
||||
|
||||
for n in sizes
|
||||
# CPU benchmark
|
||||
benchmark_cpu(n)
|
||||
|
||||
# GPU benchmark
|
||||
benchmark_gpu(n)
|
||||
|
||||
println()
|
||||
end
|
||||
|
||||
println("="^70)
|
||||
println("Benchmark Complete!")
|
||||
println("="^70)
|
||||
println()
|
||||
println("Key Findings:")
|
||||
println(" - Strategy 1 (AoS): Non-coalesced memory access on GPU")
|
||||
println(" - Strategy 2 (SoA): Coalesced memory access on GPU")
|
||||
println(" - Strategy 2 achieves 5-10× higher memory bandwidth")
|
||||
println(" - Strategy 2 has zero allocations (in-place update)")
|
||||
println()
|
||||
end
|
||||
|
||||
# Run benchmark
|
||||
if abspath(PROGRAM_FILE) == @__FILE__
|
||||
main()
|
||||
end
|
||||
@@ -1,301 +0,0 @@
|
||||
# Benchmark: Integration Point Access Patterns
|
||||
# ============================================
|
||||
#
|
||||
# This benchmark compares different approaches to storing and accessing
|
||||
# integration points in finite element assembly loops.
|
||||
#
|
||||
# Key Question: What's the fastest way to get integration points?
|
||||
#
|
||||
# Approaches tested:
|
||||
# 1. OLD: Runtime dispatch + mutable struct with Dict (type-unstable)
|
||||
# 2. NEW Option A: Compile-time function (like eval_basis!)
|
||||
# 3. NEW Option B: Store in element as NTuple
|
||||
# 4. NEW Option C: Hybrid (compile-time generation + caching)
|
||||
|
||||
using BenchmarkTools
|
||||
using Tensors
|
||||
using StaticArrays
|
||||
|
||||
# ============================================================================
|
||||
# OLD APPROACH: Runtime dispatch with mutable struct
|
||||
# ============================================================================
|
||||
|
||||
struct OldIP
|
||||
id::UInt
|
||||
weight::Float64
|
||||
coords::Tuple{Vararg{Float64}}
|
||||
fields::Dict{String,Any} # Type-unstable!
|
||||
end
|
||||
|
||||
function get_integration_points_old(::Type{Val{:Tri3}})
|
||||
# Simulate runtime dispatch to get IPs
|
||||
return [
|
||||
OldIP(UInt(1), 0.5, (1 / 3, 1 / 3), Dict{String,Any}()),
|
||||
]
|
||||
end
|
||||
|
||||
function assembly_loop_old()
|
||||
sum_val = 0.0
|
||||
for _ in 1:1000 # Simulate 1000 elements
|
||||
ips = get_integration_points_old(Val{:Tri3})
|
||||
for ip in ips
|
||||
w = ip.weight
|
||||
xi, eta = ip.coords
|
||||
# Simulate some computation
|
||||
sum_val += w * (xi + eta)
|
||||
end
|
||||
end
|
||||
return sum_val
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# NEW OPTION A: Compile-time function (zero allocation)
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Return integration points as tuple at compile time.
|
||||
Similar to eval_basis! - zero allocation, fully inlined.
|
||||
"""
|
||||
@inline function get_integration_points!(::Type{Val{:Tri3_Gauss1}})
|
||||
# Return as tuple of (weight, coords) pairs
|
||||
return (
|
||||
(0.5, (1 / 3, 1 / 3)),
|
||||
)
|
||||
end
|
||||
|
||||
@inline function get_integration_points!(::Type{Val{:Tet4_Gauss1}})
|
||||
return (
|
||||
(1 / 24, (0.25, 0.25, 0.25)),
|
||||
)
|
||||
end
|
||||
|
||||
function assembly_loop_option_a()
|
||||
sum_val = 0.0
|
||||
for _ in 1:1000
|
||||
ips = get_integration_points!(Val{:Tri3_Gauss1})
|
||||
for (w, (xi, eta)) in ips
|
||||
sum_val += w * (xi + eta)
|
||||
end
|
||||
end
|
||||
return sum_val
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# NEW OPTION B: Store in element as NTuple (what we have now)
|
||||
# ============================================================================
|
||||
|
||||
struct IntegrationPoint{D}
|
||||
ξ::NTuple{D,Float64}
|
||||
weight::Float64
|
||||
end
|
||||
|
||||
struct MockElement{NIP}
|
||||
ips::NTuple{NIP,IntegrationPoint{2}}
|
||||
end
|
||||
|
||||
function create_element_b()
|
||||
ips = (
|
||||
IntegrationPoint((1 / 3, 1 / 3), 0.5),
|
||||
)
|
||||
return MockElement(ips)
|
||||
end
|
||||
|
||||
function assembly_loop_option_b()
|
||||
elements = [create_element_b() for _ in 1:1000]
|
||||
sum_val = 0.0
|
||||
for element in elements
|
||||
for ip in element.ips
|
||||
w = ip.weight
|
||||
xi, eta = ip.ξ
|
||||
sum_val += w * (xi + eta)
|
||||
end
|
||||
end
|
||||
return sum_val
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# NEW OPTION C: Compile-time with Tensors.jl Vec (recommended for FEM)
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Return integration points with Vec{D} coordinates (Tensors.jl).
|
||||
This matches the golden standard from nodal assembly demos.
|
||||
"""
|
||||
@inline function get_integration_points_vec!(::Type{Val{:Tri3_Gauss1}})
|
||||
return (
|
||||
(0.5, Vec{2}((1 / 3, 1 / 3))),
|
||||
)
|
||||
end
|
||||
|
||||
@inline function get_integration_points_vec!(::Type{Val{:Tet4_Gauss1}})
|
||||
return (
|
||||
(1 / 24, Vec{3}((0.25, 0.25, 0.25))),
|
||||
)
|
||||
end
|
||||
|
||||
function assembly_loop_option_c()
|
||||
sum_val = 0.0
|
||||
for _ in 1:1000
|
||||
ips = get_integration_points_vec!(Val{:Tri3_Gauss1})
|
||||
for (w, xi) in ips
|
||||
# Vec arithmetic is optimized by Tensors.jl
|
||||
sum_val += w * sum(xi)
|
||||
end
|
||||
end
|
||||
return sum_val
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# NEW OPTION D: Pre-generated global constants (ultimate zero-cost)
|
||||
# ============================================================================
|
||||
|
||||
const TRI3_GAUSS1_IPS = (
|
||||
(0.5, Vec{2}((1 / 3, 1 / 3))),
|
||||
)
|
||||
|
||||
const TET4_GAUSS1_IPS = (
|
||||
(1 / 24, Vec{3}((0.25, 0.25, 0.25))),
|
||||
)
|
||||
|
||||
function assembly_loop_option_d()
|
||||
sum_val = 0.0
|
||||
for _ in 1:1000
|
||||
for (w, xi) in TRI3_GAUSS1_IPS
|
||||
sum_val += w * sum(xi)
|
||||
end
|
||||
end
|
||||
return sum_val
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# OPTION E: Hybrid - Function returns pre-computed constant
|
||||
# ============================================================================
|
||||
|
||||
@inline get_ips_tri3_gauss1() = TRI3_GAUSS1_IPS
|
||||
@inline get_ips_tet4_gauss1() = TET4_GAUSS1_IPS
|
||||
|
||||
function assembly_loop_option_e()
|
||||
sum_val = 0.0
|
||||
for _ in 1:1000
|
||||
for (w, xi) in get_ips_tri3_gauss1()
|
||||
sum_val += w * sum(xi)
|
||||
end
|
||||
end
|
||||
return sum_val
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Run Benchmarks
|
||||
# ============================================================================
|
||||
|
||||
println("="^80)
|
||||
println("Integration Point Access Pattern Benchmark")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
println("OLD APPROACH: Runtime dispatch + mutable struct with Dict")
|
||||
println("-"^80)
|
||||
@btime assembly_loop_old()
|
||||
println()
|
||||
|
||||
println("OPTION A: Compile-time function returning tuples")
|
||||
println("-"^80)
|
||||
@btime assembly_loop_option_a()
|
||||
println()
|
||||
|
||||
println("OPTION B: Store in element as NTuple (current approach)")
|
||||
println("-"^80)
|
||||
@btime assembly_loop_option_b()
|
||||
println()
|
||||
|
||||
println("OPTION C: Compile-time function with Vec{D} (Tensors.jl)")
|
||||
println("-"^80)
|
||||
@btime assembly_loop_option_c()
|
||||
println()
|
||||
|
||||
println("OPTION D: Pre-generated global constants")
|
||||
println("-"^80)
|
||||
@btime assembly_loop_option_d()
|
||||
println()
|
||||
|
||||
println("OPTION E: Function returning pre-computed constant")
|
||||
println("-"^80)
|
||||
@btime assembly_loop_option_e()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# Realistic FEM Assembly Benchmark
|
||||
# ============================================================================
|
||||
|
||||
println()
|
||||
println("="^80)
|
||||
println("REALISTIC FEM ASSEMBLY COMPARISON")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Simulate realistic element stiffness computation
|
||||
function compute_element_stiffness_old(element_type::Type{Val{:Tri3}})
|
||||
K_elem = zeros(6, 6)
|
||||
ips = get_integration_points_old(element_type)
|
||||
for ip in ips
|
||||
w = ip.weight
|
||||
xi, eta = ip.coords
|
||||
# Simulate shape function evaluation and stiffness computation
|
||||
N1 = 1 - xi - eta
|
||||
N2 = xi
|
||||
N3 = eta
|
||||
# Accumulate (simplified)
|
||||
K_elem[1, 1] += w * (N1^2)
|
||||
end
|
||||
return K_elem[1, 1]
|
||||
end
|
||||
|
||||
function compute_element_stiffness_new()
|
||||
K_elem = 0.0
|
||||
for (w, xi) in get_integration_points_vec!(Val{:Tri3_Gauss1})
|
||||
# Vec arithmetic
|
||||
xi_val = xi[1]
|
||||
eta_val = xi[2]
|
||||
N1 = 1 - xi_val - eta_val
|
||||
K_elem += w * (N1^2)
|
||||
end
|
||||
return K_elem
|
||||
end
|
||||
|
||||
println("OLD: Realistic element stiffness assembly")
|
||||
@btime begin
|
||||
sum_val = 0.0
|
||||
for _ in 1:1000
|
||||
sum_val += compute_element_stiffness_old(Val{:Tri3})
|
||||
end
|
||||
sum_val
|
||||
end
|
||||
|
||||
println()
|
||||
println("NEW: Realistic element stiffness assembly")
|
||||
@btime begin
|
||||
sum_val = 0.0
|
||||
for _ in 1:1000
|
||||
sum_val += compute_element_stiffness_new()
|
||||
end
|
||||
sum_val
|
||||
end
|
||||
|
||||
println()
|
||||
println("="^80)
|
||||
println("SUMMARY")
|
||||
println("="^80)
|
||||
println()
|
||||
println("Expected ranking (fastest to slowest):")
|
||||
println("1. Option D/E: Pre-computed constants (ultimate zero-cost)")
|
||||
println("2. Option C: Compile-time with Vec{D} (recommended for FEM)")
|
||||
println("3. Option A: Compile-time with plain tuples")
|
||||
println("4. Option B: Stored in element NTuple (slight overhead)")
|
||||
println("5. OLD: Runtime dispatch + Dict (type-unstable)")
|
||||
println()
|
||||
println("RECOMMENDATION:")
|
||||
println(" Use Option C or D/E for integration points:")
|
||||
println(" - Compile-time generation like eval_basis!")
|
||||
println(" - Return as Tuple of (weight, Vec{D}) pairs")
|
||||
println(" - Zero allocation, fully inlined")
|
||||
println(" - Matches golden standard architecture")
|
||||
@@ -1,272 +0,0 @@
|
||||
"""
|
||||
Performance analysis for LinearElastic material model.
|
||||
|
||||
Analyzes:
|
||||
1. Execution time (@btime)
|
||||
2. Memory allocations (@allocated)
|
||||
3. Type stability (@code_warntype)
|
||||
4. LLVM IR optimization (code_llvm)
|
||||
5. Native assembly (code_native)
|
||||
"""
|
||||
|
||||
using BenchmarkTools
|
||||
using Tensors
|
||||
using InteractiveUtils
|
||||
|
||||
# Load implementation
|
||||
include("../src/materials/linear_elastic.jl")
|
||||
|
||||
println("="^80)
|
||||
println("LINEAR ELASTIC MATERIAL - PERFORMANCE ANALYSIS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Test material (steel)
|
||||
steel = LinearElastic(E=200e9, ν=0.3)
|
||||
|
||||
# Test strain (uniaxial extension)
|
||||
ε = SymmetricTensor{2,3}((0.001, 0.0, 0.0, 0.0, 0.0, 0.0))
|
||||
|
||||
println("Material: Steel (E = 200 GPa, ν = 0.3)")
|
||||
println("Strain: Uniaxial extension (ε₁₁ = 0.001)")
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# BENCHMARK 1: Execution Time
|
||||
# ============================================================================
|
||||
println("BENCHMARK 1: Execution Time")
|
||||
println("-"^80)
|
||||
|
||||
# Warmup
|
||||
compute_stress(steel, ε, nothing, 0.0)
|
||||
|
||||
# Benchmark
|
||||
println("Running @btime compute_stress(steel, ε, nothing, 0.0)...")
|
||||
t = @benchmark compute_stress($steel, $ε, nothing, 0.0)
|
||||
println()
|
||||
display(t)
|
||||
println()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# BENCHMARK 2: Memory Allocations
|
||||
# ============================================================================
|
||||
println("BENCHMARK 2: Memory Allocations")
|
||||
println("-"^80)
|
||||
|
||||
# First call to compile
|
||||
compute_stress(steel, ε, nothing, 0.0)
|
||||
|
||||
# Check allocations
|
||||
allocs = @allocated compute_stress(steel, ε, nothing, 0.0)
|
||||
println("Allocations: $allocs bytes")
|
||||
|
||||
if allocs == 0
|
||||
println("✅ ZERO ALLOCATIONS (stack-only computation)")
|
||||
else
|
||||
println("⚠️ WARNING: Non-zero allocations detected!")
|
||||
end
|
||||
println()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# BENCHMARK 3: Type Stability
|
||||
# ============================================================================
|
||||
println("BENCHMARK 3: Type Stability")
|
||||
println("-"^80)
|
||||
|
||||
println("Running @code_warntype compute_stress(steel, ε, nothing, 0.0)...")
|
||||
println()
|
||||
@code_warntype compute_stress(steel, ε, nothing, 0.0)
|
||||
println()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# BENCHMARK 4: LLVM IR Analysis
|
||||
# ============================================================================
|
||||
println("BENCHMARK 4: LLVM IR Analysis")
|
||||
println("-"^80)
|
||||
|
||||
println("Running @code_llvm compute_stress(steel, ε, nothing, 0.0)...")
|
||||
println()
|
||||
@code_llvm compute_stress(steel, ε, nothing, 0.0)
|
||||
println()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# BENCHMARK 5: Native Assembly
|
||||
# ============================================================================
|
||||
println("BENCHMARK 5: Native Assembly")
|
||||
println("-"^80)
|
||||
|
||||
println("Running @code_native compute_stress(steel, ε, nothing, 0.0)...")
|
||||
println()
|
||||
@code_native compute_stress(steel, ε, nothing, 0.0)
|
||||
println()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# LLVM IR INSPECTION (Detailed Analysis)
|
||||
# ============================================================================
|
||||
println("LLVM IR INSPECTION")
|
||||
println("-"^80)
|
||||
|
||||
# Get LLVM IR as string
|
||||
llvm_ir = sprint(io -> code_llvm(io, compute_stress, typeof.((steel, ε, nothing, 0.0))))
|
||||
|
||||
# Count key operations
|
||||
n_fadd = count(r"fadd", llvm_ir)
|
||||
n_fmul = count(r"fmul", llvm_ir)
|
||||
n_load = count(r"load", llvm_ir)
|
||||
n_store = count(r"store", llvm_ir)
|
||||
n_call = count(r"call", llvm_ir)
|
||||
n_alloca = count(r"alloca", llvm_ir)
|
||||
|
||||
# Count vector operations (SIMD)
|
||||
n_vector_ops = count(r"<\d+ x ", llvm_ir)
|
||||
n_shufflevector = count(r"shufflevector", llvm_ir)
|
||||
n_insertelement = count(r"insertelement", llvm_ir)
|
||||
n_extractelement = count(r"extractelement", llvm_ir)
|
||||
|
||||
println("LLVM Operations Count:")
|
||||
println(" Floating-point additions: $n_fadd")
|
||||
println(" Floating-point multiplications: $n_fmul")
|
||||
println(" Memory loads: $n_load")
|
||||
println(" Memory stores: $n_store")
|
||||
println(" Function calls: $n_call")
|
||||
println(" Stack allocations (alloca): $n_alloca")
|
||||
println()
|
||||
println("SIMD Vectorization:")
|
||||
println(" Vector operations: $n_vector_ops")
|
||||
println(" Shuffle operations: $n_shufflevector")
|
||||
println(" Insert element operations: $n_insertelement")
|
||||
println(" Extract element operations: $n_extractelement")
|
||||
println()
|
||||
|
||||
if n_call == 0
|
||||
println("✅ No function calls (fully inlined)")
|
||||
else
|
||||
println("⚠️ Contains $n_call function calls (may not be fully inlined)")
|
||||
end
|
||||
|
||||
if n_alloca == 0
|
||||
println("✅ No stack allocations (register-only computation)")
|
||||
else
|
||||
println("ℹ️ Contains $n_alloca stack allocations")
|
||||
end
|
||||
println()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# NATIVE ASSEMBLY INSPECTION
|
||||
# ============================================================================
|
||||
println("NATIVE ASSEMBLY INSPECTION")
|
||||
println("-"^80)
|
||||
|
||||
# Get native assembly as string
|
||||
native_asm = sprint(io -> code_native(io, compute_stress, typeof.((steel, ε, nothing, 0.0))))
|
||||
|
||||
# Count SIMD instructions (AVX/SSE)
|
||||
n_vmul = count(r"vmul", native_asm)
|
||||
n_vadd = count(r"vadd", native_asm)
|
||||
n_vsub = count(r"vsub", native_asm)
|
||||
n_vfma = count(r"vfma", native_asm)
|
||||
n_vmov = count(r"vmov", native_asm)
|
||||
n_vbroadcast = count(r"vbroadcast", native_asm)
|
||||
|
||||
# Count total vector instructions
|
||||
n_total_simd = n_vmul + n_vadd + n_vsub + n_vfma + n_vmov + n_vbroadcast
|
||||
|
||||
println("x86-64 Assembly SIMD Instructions:")
|
||||
println(" vmulpd/vmulsd: $n_vmul")
|
||||
println(" vaddpd/vaddsd: $n_vadd")
|
||||
println(" vsubpd/vsubsd: $n_vsub")
|
||||
println(" vfmadd/vfmsub: $n_vfma (fused multiply-add)")
|
||||
println(" vmovapd/vmovsd: $n_vmov")
|
||||
println(" vbroadcast: $n_vbroadcast")
|
||||
println(" Total SIMD ops: $n_total_simd")
|
||||
println()
|
||||
|
||||
if n_vfma > 0
|
||||
println("✅ FMA (Fused Multiply-Add) instructions detected (optimal)")
|
||||
end
|
||||
|
||||
if n_total_simd > 0
|
||||
println("✅ SIMD vectorization active (AVX/AVX2)")
|
||||
else
|
||||
println("⚠️ No SIMD instructions detected")
|
||||
end
|
||||
println()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# PERFORMANCE SUMMARY
|
||||
# ============================================================================
|
||||
println("="^80)
|
||||
println("PERFORMANCE SUMMARY")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Extract median time from benchmark
|
||||
median_time = median(t.times)
|
||||
median_ns = median_time # Already in nanoseconds
|
||||
|
||||
println("Execution Time:")
|
||||
println(" Median: $(round(median_ns, digits=2)) ns")
|
||||
println(" Mean: $(round(mean(t.times), digits=2)) ns")
|
||||
println(" Minimum: $(round(minimum(t.times), digits=2)) ns")
|
||||
println()
|
||||
|
||||
println("Memory:")
|
||||
println(" Allocations: $allocs bytes")
|
||||
if allocs == 0
|
||||
println(" ✅ Zero allocation (confirmed)")
|
||||
end
|
||||
println()
|
||||
|
||||
println("Code Quality:")
|
||||
if n_call == 0
|
||||
println(" ✅ Fully inlined (no function calls)")
|
||||
end
|
||||
if n_alloca == 0
|
||||
println(" ✅ Register-only computation (no stack usage)")
|
||||
end
|
||||
if n_total_simd > 0
|
||||
println(" ✅ SIMD optimized ($n_total_simd vector instructions)")
|
||||
end
|
||||
if n_vfma > 0
|
||||
println(" ✅ FMA instructions ($n_vfma fused multiply-adds)")
|
||||
end
|
||||
println()
|
||||
|
||||
println("Expected Operations:")
|
||||
println(" Hooke's law: σ = λ·tr(ε)·I + 2μ·ε")
|
||||
println(" - 1 trace computation: 3 additions")
|
||||
println(" - 1 scalar multiplication: 1 multiply")
|
||||
println(" - 6 scalar multiplications for diagonal")
|
||||
println(" - 6 additions for final stress")
|
||||
println(" Tangent: 𝔻 = λ·I⊗I + 2μ·𝕀ˢʸᵐ")
|
||||
println(" - Constant tensor construction (may be compile-time)")
|
||||
println()
|
||||
|
||||
# Theoretical lower bound
|
||||
theoretical_flops = 3 + 1 + 6 + 6 # From expected operations
|
||||
println("Theoretical minimum FLOPs: ~$theoretical_flops")
|
||||
println("LLVM FLOPs: $(n_fadd + n_fmul)")
|
||||
println()
|
||||
|
||||
# Throughput calculation
|
||||
elements_per_second = 1e9 / median_ns
|
||||
println("Throughput:")
|
||||
println(" ~$(round(elements_per_second / 1e6, digits=1)) million stress evaluations/second/core")
|
||||
println()
|
||||
|
||||
println("✅ Implementation validated as:")
|
||||
println(" - Zero allocation (confirmed)")
|
||||
println(" - Type stable")
|
||||
if n_total_simd > 0
|
||||
println(" - SIMD optimized ($n_total_simd vector ops)")
|
||||
end
|
||||
println(" - Median execution time: $(round(median_ns, digits=2)) ns")
|
||||
println()
|
||||
println("="^80)
|
||||
@@ -1,865 +0,0 @@
|
||||
"""
|
||||
Material Models Performance Benchmark (Extended Version)
|
||||
|
||||
Validates performance claims from docs/book/material_modeling.md:
|
||||
- Zero allocation claims
|
||||
- 5-50× speedup over Voigt/Dict approach
|
||||
- Type stability analysis (especially 'nothing' return for stateless materials)
|
||||
- Manual vs automatic differentiation for Neo-Hookean
|
||||
- Material state handling for Newton iterations
|
||||
|
||||
Compares:
|
||||
1. New approach: Tensors.jl with SymmetricTensor
|
||||
2. Old approach: Voigt notation with arrays/Dict
|
||||
3. Neo-Hookean: Manual derivatives vs automatic differentiation
|
||||
|
||||
Materials tested:
|
||||
- Linear Elastic (Hookean) - Stateless
|
||||
- Neo-Hookean Hyperelasticity - Stateless (AD and manual versions)
|
||||
- Perfect Plasticity (von Mises) - Stateful
|
||||
|
||||
Type hierarchy:
|
||||
- AbstractMaterial - Base type for all materials
|
||||
- AbstractMaterialState - Base type for material internal state
|
||||
- NoState - For stateless materials
|
||||
- PlasticityState - For plasticity with history
|
||||
"""
|
||||
|
||||
using Tensors
|
||||
using BenchmarkTools
|
||||
using LinearAlgebra
|
||||
using InteractiveUtils # For @code_warntype
|
||||
|
||||
println("="^80)
|
||||
println("Material Models Performance Benchmark (Extended)")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
#=============================================================================
|
||||
TYPE HIERARCHY
|
||||
=============================================================================#
|
||||
|
||||
"""
|
||||
Abstract base type for all materials.
|
||||
|
||||
All concrete materials must implement:
|
||||
- `compute_stress(material, ε, state_old, Δt) -> (σ, 𝔻, state_new)`
|
||||
- `initial_state(material) -> AbstractMaterialState`
|
||||
"""
|
||||
abstract type AbstractMaterial end
|
||||
|
||||
"""
|
||||
Abstract base type for material internal state.
|
||||
|
||||
Used to track history-dependent variables during Newton iterations:
|
||||
- Old state (beginning of time step)
|
||||
- Trial state (current Newton iteration)
|
||||
- New state (converged solution)
|
||||
"""
|
||||
abstract type AbstractMaterialState end
|
||||
|
||||
"""
|
||||
State for stateless materials (no history dependence).
|
||||
|
||||
Using singleton type instead of `nothing` for type hierarchy consistency.
|
||||
Performance identical to `nothing` (zero-sized type).
|
||||
"""
|
||||
struct NoState <: AbstractMaterialState end
|
||||
|
||||
"""
|
||||
Initial state for stateless materials.
|
||||
"""
|
||||
initial_state(::AbstractMaterial) = NoState()
|
||||
|
||||
#=============================================================================
|
||||
NEW APPROACH: Tensors.jl Implementation
|
||||
=============================================================================#
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 1. Linear Elastic (Hookean)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
"""Linear elastic material with Tensors.jl"""
|
||||
struct LinearElastic <: AbstractMaterial
|
||||
E::Float64 # Young's modulus [Pa]
|
||||
ν::Float64 # Poisson's ratio [-]
|
||||
end
|
||||
|
||||
LinearElastic(; E, ν) = LinearElastic(E, ν)
|
||||
|
||||
λ(mat::LinearElastic) = mat.E * mat.ν / ((1 + mat.ν) * (1 - 2mat.ν))
|
||||
μ(mat::LinearElastic) = mat.E / (2(1 + mat.ν))
|
||||
|
||||
"""Compute stress for linear elastic material."""
|
||||
function compute_stress(
|
||||
material::LinearElastic,
|
||||
ε::SymmetricTensor{2,3,T},
|
||||
state_old::NoState,
|
||||
Δt::Float64
|
||||
) where T
|
||||
|
||||
# Lamé parameters
|
||||
λ_val = λ(material)
|
||||
μ_val = μ(material)
|
||||
|
||||
# Identity tensor
|
||||
I = one(ε)
|
||||
|
||||
# Hooke's law: σ = λ·tr(ε)·I + 2μ·ε
|
||||
σ = λ_val * tr(ε) * I + 2μ_val * ε
|
||||
|
||||
# Tangent modulus: 𝔻 = λ I⊗I + 2μ 𝕀ˢʸᵐ
|
||||
𝕀ˢʸᵐ = one(SymmetricTensor{4,3,T}) # Symmetric 4th order identity
|
||||
𝔻 = λ_val * I ⊗ I + 2μ_val * 𝕀ˢʸᵐ
|
||||
|
||||
return σ, 𝔻, NoState() # No state change (stateless)
|
||||
end# ---------------------------------------------------------------------------
|
||||
# 2. Neo-Hookean Hyperelasticity (Automatic Differentiation)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
"""Neo-Hookean hyperelastic material (using automatic differentiation)."""
|
||||
struct NeoHookeanAD <: AbstractMaterial
|
||||
μ::Float64 # Shear modulus [Pa]
|
||||
λ::Float64 # Lamé parameter [Pa]
|
||||
end
|
||||
|
||||
function NeoHookeanAD(; E, ν)
|
||||
μ = E / (2(1 + ν))
|
||||
λ = E * ν / ((1 + ν) * (1 - 2ν))
|
||||
return NeoHookeanAD(μ, λ)
|
||||
end
|
||||
|
||||
"""Strain energy density for Neo-Hookean model."""
|
||||
function strain_energy(material::NeoHookeanAD, C::SymmetricTensor{2,3})
|
||||
μ, λ = material.μ, material.λ
|
||||
|
||||
# Invariants
|
||||
I₁ = tr(C)
|
||||
J = √(det(C))
|
||||
|
||||
# Strain energy: ψ = μ/2(I₁ - 3) - μln(J) + λ/2·ln²(J)
|
||||
ψ = μ / 2 * (I₁ - 3) - μ * log(J) + λ / 2 * log(J)^2
|
||||
|
||||
return ψ
|
||||
end
|
||||
|
||||
"""Compute stress for Neo-Hookean material using automatic differentiation."""
|
||||
function compute_stress(
|
||||
material::NeoHookeanAD,
|
||||
E::SymmetricTensor{2,3,T}, # Green-Lagrange strain
|
||||
state_old::NoState,
|
||||
Δt::Float64
|
||||
) where T
|
||||
|
||||
# Right Cauchy-Green tensor: C = 2E + I
|
||||
I = one(E)
|
||||
C = 2E + I
|
||||
|
||||
# Strain energy function (closure capturing material)
|
||||
ψ(C_) = strain_energy(material, C_)
|
||||
|
||||
# Automatic differentiation!
|
||||
𝔻, S = hessian(ψ, C, :all) # Returns both hessian and gradient!
|
||||
|
||||
# Note: We want S = 2·∂ψ/∂C, 𝔻 = 4·∂²ψ/∂C²
|
||||
S = 2 * S
|
||||
𝔻 = 4 * 𝔻
|
||||
|
||||
return S, 𝔻, NoState() # No state change (stateless)
|
||||
end
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 3. Neo-Hookean Hyperelasticity (Manual Derivatives)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
"""
|
||||
Neo-Hookean hyperelastic material (hand-coded derivatives).
|
||||
|
||||
Strain energy: ψ(C) = μ/2(I₁ - 3) - μln(J) + λ/2·ln²(J)
|
||||
|
||||
Where:
|
||||
- I₁ = tr(C) - First invariant
|
||||
- J = √det(C) - Jacobian determinant
|
||||
|
||||
Derivatives (computed by hand):
|
||||
- S = 2∂ψ/∂C = μ(I - C⁻¹) + λln(J)C⁻¹
|
||||
- 𝔻 = 4∂²ψ/∂C² = λ(C⁻¹⊗C⁻¹) + 2(μ - λln(J))∂C⁻¹/∂C
|
||||
|
||||
The second derivative uses the identity:
|
||||
∂C⁻¹/∂C : X = -C⁻¹:(X:C⁻¹) for any symmetric X
|
||||
"""
|
||||
struct NeoHookeanManual <: AbstractMaterial
|
||||
μ::Float64 # Shear modulus [Pa]
|
||||
λ::Float64 # Lamé parameter [Pa]
|
||||
end
|
||||
|
||||
function NeoHookeanManual(; E, ν)
|
||||
μ = E / (2(1 + ν))
|
||||
λ = E * ν / ((1 + ν) * (1 - 2ν))
|
||||
return NeoHookeanManual(μ, λ)
|
||||
end
|
||||
|
||||
"""Compute stress for Neo-Hookean material with manual derivatives."""
|
||||
function compute_stress(
|
||||
material::NeoHookeanManual,
|
||||
E::SymmetricTensor{2,3,T}, # Green-Lagrange strain
|
||||
state_old::NoState,
|
||||
Δt::Float64
|
||||
) where T
|
||||
μ, λ = material.μ, material.λ
|
||||
|
||||
# Right Cauchy-Green tensor: C = 2E + I
|
||||
I = one(E)
|
||||
C = 2E + I
|
||||
|
||||
# Invariants
|
||||
J = √(det(C))
|
||||
C_inv = inv(C)
|
||||
|
||||
# Second Piola-Kirchhoff stress: S = μ(I - C⁻¹) + λln(J)C⁻¹
|
||||
S = μ * (I - C_inv) + λ * log(J) * C_inv
|
||||
|
||||
# Material tangent: 𝔻 = 4∂²ψ/∂C²
|
||||
# Term 1: λ(C⁻¹⊗C⁻¹)
|
||||
𝔻₁ = λ * (C_inv ⊗ C_inv)
|
||||
|
||||
# Term 2: 2(μ - λln(J))∂C⁻¹/∂C
|
||||
# The derivative ∂C⁻¹/∂C can be computed as:
|
||||
# (∂C⁻¹/∂C)ᵢⱼₖₗ = -1/2(C⁻¹ᵢₖC⁻¹ⱼₗ + C⁻¹ᵢₗC⁻¹ⱼₖ)
|
||||
#
|
||||
# For SymmetricTensor, we build this fourth-order tensor
|
||||
# by exploiting the symmetry structure
|
||||
|
||||
# Build the symmetric fourth-order tensor manually
|
||||
# This is the most expensive part of the computation
|
||||
𝕀ˢʸᵐ = one(SymmetricTensor{4,3,T})
|
||||
|
||||
# For compressible Neo-Hookean, the full tangent is:
|
||||
# 𝔻 = λ(C⁻¹⊗C⁻¹) - 2(μ - λln(J))(C⁻¹⊙C⁻¹)
|
||||
# where ⊙ is the symmetric dyadic product for fourth-order tensors
|
||||
|
||||
# Construct C⁻¹⊗C⁻¹ part (already have 𝔻₁)
|
||||
# Construct symmetric part: use Tensors.jl identity operations
|
||||
# The fourth-order identity for symmetric tensors handles this
|
||||
|
||||
coeff = 2(μ - λ * log(J))
|
||||
|
||||
# For the symmetric outer product of C⁻¹ with itself,
|
||||
# we can use the following approach:
|
||||
# Build component-wise using Voigt ordering
|
||||
|
||||
# Simplified: Use the property that for small strains,
|
||||
# this reduces to a simpler form. For full nonlinear case:
|
||||
𝔻₂ = -coeff * inv_symmetric_outer(C_inv)
|
||||
|
||||
𝔻 = 𝔻₁ + 𝔻₂
|
||||
|
||||
return S, 𝔻, NoState()
|
||||
end
|
||||
|
||||
"""
|
||||
Compute symmetric fourth-order tensor from inverse: ∂C⁻¹/∂C
|
||||
|
||||
For symmetric second-order tensor C⁻¹, compute the fourth-order tensor:
|
||||
(∂C⁻¹/∂C)ᵢⱼₖₗ = -1/2(C⁻¹ᵢₖC⁻¹ⱼₗ + C⁻¹ᵢₗC⁻¹ⱼₖ)
|
||||
|
||||
This appears in the material tangent of hyperelastic materials.
|
||||
"""
|
||||
function inv_symmetric_outer(C_inv::SymmetricTensor{2,3,T}) where T
|
||||
# Extract components (Voigt notation: 11, 22, 33, 12, 23, 13)
|
||||
c = [C_inv[1, 1], C_inv[2, 2], C_inv[3, 3],
|
||||
C_inv[1, 2], C_inv[2, 3], C_inv[1, 3]]
|
||||
|
||||
# Build fourth-order tensor in Voigt notation (6x6 matrix representation)
|
||||
# Then convert to SymmetricTensor{4,3}
|
||||
#
|
||||
# This is the -1/2(CᵢₖCⱼₗ + CᵢₗCⱼₖ) tensor
|
||||
|
||||
# For now, use a simpler approximation that works for Neo-Hookean
|
||||
# Full implementation would build all 36 components
|
||||
|
||||
# Use outer product and symmetrize
|
||||
result = C_inv ⊗ C_inv
|
||||
|
||||
# Add symmetric component
|
||||
# (This is a simplified version - full implementation needs more care)
|
||||
return result
|
||||
end
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 4. Perfect Plasticity (von Mises)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
"""Perfect plasticity with von Mises yield criterion."""
|
||||
struct PerfectPlasticity <: AbstractMaterial
|
||||
E::Float64 # Young's modulus [Pa]
|
||||
ν::Float64 # Poisson's ratio [-]
|
||||
σ_y::Float64 # Yield stress [Pa]
|
||||
end
|
||||
|
||||
PerfectPlasticity(; E, ν, σ_y) = PerfectPlasticity(E, ν, σ_y)
|
||||
|
||||
λ(mat::PerfectPlasticity) = mat.E * mat.ν / ((1 + mat.ν) * (1 - 2mat.ν))
|
||||
μ(mat::PerfectPlasticity) = mat.E / (2(1 + mat.ν))
|
||||
|
||||
"""
|
||||
Internal state for plasticity (history-dependent variables).
|
||||
|
||||
This struct is passed through Newton iterations:
|
||||
- state_old: State at beginning of time step (t_n)
|
||||
- state_trial: Trial state during iteration (may not converge)
|
||||
- state_new: Updated state for next iteration (t_n+1)
|
||||
"""
|
||||
struct PlasticityState{T} <: AbstractMaterialState
|
||||
ε_p::SymmetricTensor{2,3,T} # Plastic strain
|
||||
α::T # Equivalent plastic strain
|
||||
end
|
||||
|
||||
"""Initial state for plasticity (zero plastic strain)."""
|
||||
initial_state(::PerfectPlasticity) = PlasticityState(zero(SymmetricTensor{2,3}), 0.0)
|
||||
|
||||
"""Von Mises equivalent stress."""
|
||||
function von_mises_stress(σ::SymmetricTensor{2,3})
|
||||
s = dev(σ) # Deviatoric stress
|
||||
return √(3 / 2 * s ⊡ s)
|
||||
end
|
||||
|
||||
"""Compute stress for perfectly plastic material with radial return."""
|
||||
function compute_stress(
|
||||
material::PerfectPlasticity,
|
||||
ε::SymmetricTensor{2,3,T},
|
||||
state_old::PlasticityState{T},
|
||||
Δt::Float64
|
||||
) where T
|
||||
|
||||
# Material parameters
|
||||
λ_val = λ(material)
|
||||
μ_val = μ(material)
|
||||
σ_y = material.σ_y
|
||||
|
||||
# Elastic constitutive tensor
|
||||
I = one(ε)
|
||||
𝕀ˢʸᵐ = one(SymmetricTensor{4,3,T})
|
||||
𝔻ᵉ = λ_val * I ⊗ I + 2μ_val * 𝕀ˢʸᵐ
|
||||
|
||||
# Elastic predictor
|
||||
ε_e = ε - state_old.ε_p
|
||||
σ_trial = λ_val * tr(ε_e) * I + 2μ_val * ε_e
|
||||
σ_eq_trial = von_mises_stress(σ_trial)
|
||||
|
||||
# Yield function
|
||||
f = σ_eq_trial - σ_y
|
||||
|
||||
if f ≤ 0.0
|
||||
# Elastic step
|
||||
σ = σ_trial
|
||||
𝔻 = 𝔻ᵉ
|
||||
state_new = state_old
|
||||
else
|
||||
# Plastic step: Radial return
|
||||
s_trial = dev(σ_trial)
|
||||
p = tr(σ_trial) / 3
|
||||
|
||||
# Return to yield surface
|
||||
σ = p * I + (σ_y / σ_eq_trial) * s_trial
|
||||
|
||||
# Plastic multiplier
|
||||
Δγ = f / (3μ_val)
|
||||
|
||||
# Flow direction
|
||||
n = √(3 / 2) * s_trial / σ_eq_trial
|
||||
|
||||
# Update plastic strain
|
||||
ε_p_new = state_old.ε_p + Δγ * n
|
||||
α_new = state_old.α + Δγ
|
||||
|
||||
state_new = PlasticityState(ε_p_new, α_new)
|
||||
|
||||
# Algorithmic tangent (simplified)
|
||||
θ = 1 - σ_y / σ_eq_trial
|
||||
β = 6μ_val^2 / (3μ_val + θ * 3μ_val)
|
||||
|
||||
𝔻 = 𝔻ᵉ - β * (n ⊗ n)
|
||||
end
|
||||
|
||||
return σ, 𝔻, state_new
|
||||
end
|
||||
|
||||
#=============================================================================
|
||||
OLD APPROACH: Voigt Notation + Array Implementation
|
||||
=============================================================================#
|
||||
|
||||
"""Old-style linear elastic with Voigt notation."""
|
||||
struct LinearElasticOld
|
||||
E::Float64
|
||||
ν::Float64
|
||||
end
|
||||
|
||||
"""Compute 6×6 constitutive matrix (Voigt notation)."""
|
||||
function constitutive_matrix(mat::LinearElasticOld)
|
||||
E, ν = mat.E, mat.ν
|
||||
λ = E * ν / ((1 + ν) * (1 - 2ν))
|
||||
μ = E / (2(1 + ν))
|
||||
|
||||
D = zeros(6, 6)
|
||||
D[1:3, 1:3] .= λ
|
||||
D[1, 1] = D[2, 2] = D[3, 3] = λ + 2μ
|
||||
D[4, 4] = D[5, 5] = D[6, 6] = μ
|
||||
|
||||
return D
|
||||
end
|
||||
|
||||
"""Compute stress (old approach with arrays)."""
|
||||
function compute_stress_old(
|
||||
material::LinearElasticOld,
|
||||
ε_vec::Vector{Float64}, # [ε11, ε22, ε33, 2ε12, 2ε23, 2ε13]
|
||||
state_old::Dict{String,Any},
|
||||
Δt::Float64
|
||||
)
|
||||
D = constitutive_matrix(material)
|
||||
σ_vec = D * ε_vec
|
||||
return σ_vec, D, state_old
|
||||
end
|
||||
|
||||
"""Old-style Neo-Hookean (manual derivatives)."""
|
||||
struct NeoHookeanOld
|
||||
μ::Float64
|
||||
λ::Float64
|
||||
end
|
||||
|
||||
"""Compute stress manually (simplified, no actual derivatives for brevity)."""
|
||||
function compute_stress_old(
|
||||
material::NeoHookeanOld,
|
||||
E_vec::Vector{Float64},
|
||||
state_old::Dict{String,Any},
|
||||
Δt::Float64
|
||||
)
|
||||
# This would normally have 50+ lines of manual derivative calculations
|
||||
# For benchmark purposes, just do some array operations
|
||||
D = zeros(6, 6)
|
||||
for i in 1:6
|
||||
D[i, i] = material.μ + material.λ / 3
|
||||
end
|
||||
σ_vec = D * E_vec
|
||||
return σ_vec, D, state_old
|
||||
end
|
||||
|
||||
"""Old-style plasticity with Dict storage."""
|
||||
struct PerfectPlasticityOld
|
||||
E::Float64
|
||||
ν::Float64
|
||||
σ_y::Float64
|
||||
end
|
||||
|
||||
"""Compute stress with Dict field storage."""
|
||||
function compute_stress_old(
|
||||
material::PerfectPlasticityOld,
|
||||
ε_vec::Vector{Float64},
|
||||
state_old::Dict{String,Any},
|
||||
Δt::Float64
|
||||
)
|
||||
# Get plastic strain from Dict (type instability!)
|
||||
if haskey(state_old, "epsilon_plastic")
|
||||
ε_p_vec = state_old["epsilon_plastic"]
|
||||
else
|
||||
ε_p_vec = zeros(6)
|
||||
end
|
||||
|
||||
# Elastic trial
|
||||
D = constitutive_matrix(LinearElasticOld(material.E, material.ν))
|
||||
ε_e_vec = ε_vec - ε_p_vec
|
||||
σ_trial_vec = D * ε_e_vec
|
||||
|
||||
# Von Mises check (manual calculation with arrays)
|
||||
s11, s22, s33 = σ_trial_vec[1:3]
|
||||
s12, s23, s13 = σ_trial_vec[4:6]
|
||||
p = (s11 + s22 + s33) / 3
|
||||
dev_vec = [s11 - p, s22 - p, s33 - p, s12, s23, s13]
|
||||
σ_eq = √(3 / 2 * (dev_vec[1]^2 + dev_vec[2]^2 + dev_vec[3]^2 +
|
||||
2 * (dev_vec[4]^2 + dev_vec[5]^2 + dev_vec[6]^2)))
|
||||
|
||||
f = σ_eq - material.σ_y
|
||||
|
||||
state_new = copy(state_old)
|
||||
|
||||
if f > 0.0
|
||||
# Plastic correction
|
||||
factor = material.σ_y / σ_eq
|
||||
σ_vec = [p, p, p, 0.0, 0.0, 0.0] + factor * dev_vec
|
||||
|
||||
# Update state in Dict
|
||||
Δγ = f / (3 * material.E / (2(1 + material.ν)))
|
||||
n_vec = √(3 / 2) * dev_vec / σ_eq
|
||||
state_new["epsilon_plastic"] = ε_p_vec + Δγ * n_vec
|
||||
else
|
||||
σ_vec = σ_trial_vec
|
||||
end
|
||||
|
||||
return σ_vec, D, state_new
|
||||
end
|
||||
|
||||
#=============================================================================
|
||||
MATERIAL STATE HANDLING FOR NEWTON ITERATIONS
|
||||
=============================================================================#
|
||||
|
||||
"""
|
||||
Example: How to handle material state during Newton-Raphson iterations.
|
||||
|
||||
In FEM nonlinear analysis, each time step requires iterative solution:
|
||||
|
||||
1. **Beginning of time step (t_n):**
|
||||
- state_old = converged state from previous time step
|
||||
|
||||
2. **During Newton iterations (t_n → t_n+1):**
|
||||
- For each iteration k = 1, 2, ...
|
||||
- Compute: σ, 𝔻, state_trial = compute_stress(material, ε_k, state_old, Δt)
|
||||
- state_trial is NOT committed yet (iteration may not converge)
|
||||
|
||||
3. **After convergence:**
|
||||
- state_new = state_trial from final iteration
|
||||
- Commit: state_old ← state_new for next time step
|
||||
|
||||
This ensures:
|
||||
- Failed iterations don't corrupt material history
|
||||
- Material state is consistent with converged solution
|
||||
- Internal variables (plastic strain, damage, etc.) evolve correctly
|
||||
"""
|
||||
|
||||
"""
|
||||
Simulate Newton-Raphson iteration with material state handling.
|
||||
|
||||
Returns:
|
||||
- converged: Whether iterations converged
|
||||
- n_iter: Number of iterations
|
||||
- state_converged: Final material state (only valid if converged)
|
||||
"""
|
||||
function newton_with_material_state(
|
||||
material::AbstractMaterial,
|
||||
ε_target::SymmetricTensor{2,3},
|
||||
state_old::AbstractMaterialState,
|
||||
Δt::Float64;
|
||||
max_iter=10,
|
||||
tol=1e-8
|
||||
)
|
||||
println(" Newton iteration with material state tracking:")
|
||||
println(" " * "="^60)
|
||||
|
||||
# Initial guess
|
||||
ε_k = zero(ε_target)
|
||||
|
||||
for k in 1:max_iter
|
||||
# Compute stress and tangent (state_trial is NOT committed yet!)
|
||||
σ_k, 𝔻_k, state_trial = compute_stress(material, ε_k, state_old, Δt)
|
||||
|
||||
println(" Iteration $k:")
|
||||
println(" strain: $(norm(ε_k))")
|
||||
println(" stress: $(norm(σ_k))")
|
||||
println(" state: $(state_trial)")
|
||||
|
||||
# Residual (simplified: just strain error)
|
||||
r = norm(ε_k - ε_target)
|
||||
|
||||
if r < tol
|
||||
println(" → Converged!")
|
||||
println(" Final state committed: $(state_trial)")
|
||||
return true, k, state_trial
|
||||
end
|
||||
|
||||
# Newton update (simplified)
|
||||
ε_k = ε_k + 0.5 * (ε_target - ε_k)
|
||||
end
|
||||
|
||||
println(" → Failed to converge!")
|
||||
println(" State NOT committed (keeping state_old)")
|
||||
return false, max_iter, state_old # Keep old state on failure!
|
||||
end
|
||||
|
||||
println()
|
||||
println("="^80)
|
||||
println("NEWTON ITERATION STATE HANDLING EXAMPLE")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Example 1: Stateless material (LinearElastic)
|
||||
println("Example 1: Stateless Material (LinearElastic)")
|
||||
println("-"^80)
|
||||
steel_example = LinearElastic(E=200e9, ν=0.3)
|
||||
state_stateless = initial_state(steel_example)
|
||||
ε_test = SymmetricTensor{2,3}((0.001, 0.0, 0.0, 0.0, 0.0, 0.0))
|
||||
|
||||
converged, n_iter, state_final = newton_with_material_state(
|
||||
steel_example, ε_test, state_stateless, 1.0, max_iter=3
|
||||
)
|
||||
println("Result: state_final = $state_final (NoState, always)")
|
||||
println()
|
||||
|
||||
# Example 2: Stateful material (PerfectPlasticity)
|
||||
println("Example 2: Stateful Material (PerfectPlasticity)")
|
||||
println("-"^80)
|
||||
plastic_example = PerfectPlasticity(E=200e9, ν=0.3, σ_y=250e6)
|
||||
state_stateful = initial_state(plastic_example)
|
||||
ε_test_plastic = SymmetricTensor{2,3}((0.002, 0.0, 0.0, 0.0, 0.0, 0.0)) # Large strain → plastic
|
||||
|
||||
converged, n_iter, state_final = newton_with_material_state(
|
||||
plastic_example, ε_test_plastic, state_stateful, 1.0, max_iter=3
|
||||
)
|
||||
println("Result: state_final = $state_final (plastic strain accumulated)")
|
||||
println()
|
||||
|
||||
println("Key insight: State handling is IDENTICAL for all materials due to")
|
||||
println("AbstractMaterialState type hierarchy. Assembly code doesn't need")
|
||||
println("to know whether material is stateless or stateful!")
|
||||
println()
|
||||
|
||||
#=============================================================================
|
||||
BENCHMARK SETUP
|
||||
=============================================================================#
|
||||
|
||||
println("Setting up materials and test cases...")
|
||||
println()
|
||||
|
||||
# Materials (realistic steel properties)
|
||||
steel_new = LinearElastic(E=200e9, ν=0.3)
|
||||
steel_old = LinearElasticOld(200e9, 0.3)
|
||||
|
||||
rubber_ad = NeoHookeanAD(E=10e6, ν=0.45)
|
||||
rubber_manual = NeoHookeanManual(E=10e6, ν=0.45)
|
||||
rubber_old = NeoHookeanOld(10e6 / (2 * 1.45), 10e6 * 0.45 / (1.45 * 0.1))
|
||||
|
||||
plastic_new = PerfectPlasticity(E=200e9, ν=0.3, σ_y=250e6)
|
||||
plastic_old = PerfectPlasticityOld(200e9, 0.3, 250e6)
|
||||
|
||||
# Test strain (small elastic deformation)
|
||||
ε11, ε22, ε33 = 0.001, -0.0003, -0.0003 # Uniaxial tension with Poisson effect
|
||||
ε12, ε23, ε13 = 0.0, 0.0, 0.0
|
||||
|
||||
# New approach: SymmetricTensor
|
||||
ε_tensor = SymmetricTensor{2,3}((ε11, ε12, ε13, ε22, ε23, ε33))
|
||||
E_tensor = ε_tensor # For Neo-Hookean (Green-Lagrange ≈ small strain here)
|
||||
|
||||
# Old approach: Voigt vector (note factor of 2 for shear!)
|
||||
ε_voigt = [ε11, ε22, ε33, 2 * ε12, 2 * ε23, 2 * ε13]
|
||||
|
||||
# States (using proper type hierarchy)
|
||||
state_nostate = NoState()
|
||||
state_dict_empty = Dict{String,Any}()
|
||||
state_plastic_new = initial_state(plastic_new)
|
||||
state_plastic_old = Dict{String,Any}("epsilon_plastic" => zeros(6))
|
||||
|
||||
println("Materials configured:")
|
||||
println(" - Linear Elastic: E = 200 GPa, ν = 0.3")
|
||||
println(" - Neo-Hookean (AD): μ ≈ 3.4 MPa, λ ≈ 45 MPa (automatic differentiation)")
|
||||
println(" - Neo-Hookean (Manual): μ ≈ 3.4 MPa, λ ≈ 45 MPa (hand-coded derivatives)")
|
||||
println(" - Perfect Plasticity: E = 200 GPa, σ_y = 250 MPa")
|
||||
println()
|
||||
println("Test strain: ε11 = 0.001 (uniaxial tension)")
|
||||
println()
|
||||
|
||||
#=============================================================================
|
||||
TYPE STABILITY CHECK
|
||||
=============================================================================#
|
||||
|
||||
println("="^80)
|
||||
println("TYPE STABILITY ANALYSIS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
println("Checking for type instabilities...")
|
||||
println()
|
||||
|
||||
# Check LinearElastic
|
||||
println("1. Linear Elastic (Tensors.jl):")
|
||||
@code_warntype compute_stress(steel_new, ε_tensor, state_nostate, 0.0)
|
||||
println()
|
||||
|
||||
println("2. Linear Elastic (Old Voigt/Dict):")
|
||||
@code_warntype compute_stress_old(steel_old, ε_voigt, state_dict_empty, 0.0)
|
||||
println()
|
||||
|
||||
println("3. Neo-Hookean AD (Tensors.jl with automatic differentiation):")
|
||||
@code_warntype compute_stress(rubber_ad, E_tensor, state_nostate, 0.0)
|
||||
println()
|
||||
|
||||
println("4. Neo-Hookean Manual (Tensors.jl with hand-coded derivatives):")
|
||||
@code_warntype compute_stress(rubber_manual, E_tensor, state_nostate, 0.0)
|
||||
println()
|
||||
|
||||
println("5. Perfect Plasticity (Tensors.jl):")
|
||||
@code_warntype compute_stress(plastic_new, ε_tensor, state_plastic_new, 0.0)
|
||||
println()
|
||||
|
||||
println("6. Perfect Plasticity (Old Dict):")
|
||||
@code_warntype compute_stress_old(plastic_old, ε_voigt, state_plastic_old, 0.0)
|
||||
println()
|
||||
|
||||
#=============================================================================
|
||||
ALLOCATION TESTS
|
||||
=============================================================================#
|
||||
|
||||
println("="^80)
|
||||
println("ALLOCATION TESTS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
println("Testing for allocations (should be 0 for new approach)...")
|
||||
println()
|
||||
|
||||
# Linear Elastic
|
||||
println("1. Linear Elastic")
|
||||
println(" NEW (Tensors.jl):")
|
||||
allocs_le_new = @allocated compute_stress(steel_new, ε_tensor, state_nostate, 0.0)
|
||||
println(" Allocations: $allocs_le_new bytes")
|
||||
|
||||
println(" OLD (Voigt/Dict):")
|
||||
allocs_le_old = @allocated compute_stress_old(steel_old, ε_voigt, state_dict_empty, 0.0)
|
||||
println(" Allocations: $allocs_le_old bytes")
|
||||
println()
|
||||
|
||||
# Neo-Hookean
|
||||
println("2. Neo-Hookean")
|
||||
println(" NEW (Tensors.jl + AD):")
|
||||
allocs_nh_ad = @allocated compute_stress(rubber_ad, E_tensor, state_nostate, 0.0)
|
||||
println(" Allocations: $allocs_nh_ad bytes")
|
||||
|
||||
println(" NEW (Tensors.jl + Manual):")
|
||||
allocs_nh_manual = @allocated compute_stress(rubber_manual, E_tensor, state_nostate, 0.0)
|
||||
println(" Allocations: $allocs_nh_manual bytes")
|
||||
|
||||
println(" OLD (Array):")
|
||||
allocs_nh_old = @allocated compute_stress_old(rubber_old, ε_voigt, state_dict_empty, 0.0)
|
||||
println(" Allocations: $allocs_nh_old bytes")
|
||||
println()
|
||||
|
||||
# Perfect Plasticity
|
||||
println("3. Perfect Plasticity (elastic branch)")
|
||||
println(" NEW (Tensors.jl):")
|
||||
allocs_pp_new = @allocated compute_stress(plastic_new, ε_tensor, state_plastic_new, 0.0)
|
||||
println(" Allocations: $allocs_pp_new bytes")
|
||||
|
||||
println(" OLD (Dict):")
|
||||
allocs_pp_old = @allocated compute_stress_old(plastic_old, ε_voigt, state_plastic_old, 0.0)
|
||||
println(" Allocations: $allocs_pp_old bytes")
|
||||
println()
|
||||
|
||||
#=============================================================================
|
||||
PERFORMANCE BENCHMARKS
|
||||
=============================================================================#
|
||||
|
||||
println("="^80)
|
||||
println("PERFORMANCE BENCHMARKS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
println("Running detailed benchmarks (this may take a minute)...")
|
||||
println()
|
||||
|
||||
# Linear Elastic
|
||||
println("1. LINEAR ELASTIC")
|
||||
println("-"^40)
|
||||
println("NEW (Tensors.jl):")
|
||||
bench_le_new = @benchmark compute_stress($steel_new, $ε_tensor, $state_nostate, 0.0)
|
||||
display(bench_le_new)
|
||||
println()
|
||||
|
||||
println("OLD (Voigt/Dict):")
|
||||
bench_le_old = @benchmark compute_stress_old($steel_old, $ε_voigt, $state_dict_empty, 0.0)
|
||||
display(bench_le_old)
|
||||
println()
|
||||
|
||||
speedup_le = median(bench_le_old.times) / median(bench_le_new.times)
|
||||
println("SPEEDUP: $(round(speedup_le, digits=1))×")
|
||||
println()
|
||||
|
||||
# Neo-Hookean
|
||||
println("2. NEO-HOOKEAN")
|
||||
println("-"^40)
|
||||
println("NEW (Tensors.jl + Automatic Differentiation):")
|
||||
bench_nh_ad = @benchmark compute_stress($rubber_ad, $E_tensor, $state_nostate, 0.0)
|
||||
display(bench_nh_ad)
|
||||
println()
|
||||
|
||||
println("NEW (Tensors.jl + Manual Derivatives):")
|
||||
bench_nh_manual = @benchmark compute_stress($rubber_manual, $E_tensor, $state_nostate, 0.0)
|
||||
display(bench_nh_manual)
|
||||
println()
|
||||
|
||||
println("OLD (Array):")
|
||||
bench_nh_old = @benchmark compute_stress_old($rubber_old, $ε_voigt, $state_dict_empty, 0.0)
|
||||
display(bench_nh_old)
|
||||
println()
|
||||
|
||||
speedup_nh_ad = median(bench_nh_old.times) / median(bench_nh_ad.times)
|
||||
speedup_nh_manual = median(bench_nh_old.times) / median(bench_nh_manual.times)
|
||||
ad_overhead = median(bench_nh_ad.times) / median(bench_nh_manual.times)
|
||||
println("SPEEDUP (AD): $(round(speedup_nh_ad, digits=1))×")
|
||||
println("SPEEDUP (Manual): $(round(speedup_nh_manual, digits=1))×")
|
||||
println("AD OVERHEAD: $(round(ad_overhead, digits=1))× (AD / Manual)")
|
||||
println()
|
||||
|
||||
# Perfect Plasticity
|
||||
println("3. PERFECT PLASTICITY (elastic branch)")
|
||||
println("-"^40)
|
||||
println("NEW (Tensors.jl):")
|
||||
bench_pp_new = @benchmark compute_stress($plastic_new, $ε_tensor, $state_plastic_new, 0.0)
|
||||
display(bench_pp_new)
|
||||
println()
|
||||
|
||||
println("OLD (Dict):")
|
||||
bench_pp_old = @benchmark compute_stress_old($plastic_old, $ε_voigt, $state_plastic_old, 0.0)
|
||||
display(bench_pp_old)
|
||||
println()
|
||||
|
||||
speedup_pp = median(bench_pp_old.times) / median(bench_pp_new.times)
|
||||
println("SPEEDUP: $(round(speedup_pp, digits=1))×")
|
||||
println()
|
||||
|
||||
#=============================================================================
|
||||
SUMMARY
|
||||
=============================================================================#
|
||||
|
||||
println("="^80)
|
||||
println("SUMMARY")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
println("ALLOCATIONS:")
|
||||
println(" LinearElastic: NEW = $allocs_le_new bytes, OLD = $allocs_le_old bytes")
|
||||
println(" NeoHookean (AD): NEW = $allocs_nh_ad bytes, OLD = $allocs_nh_old bytes")
|
||||
println(" NeoHookean (Manual): NEW = $allocs_nh_manual bytes")
|
||||
println(" PerfectPlasticity: NEW = $allocs_pp_new bytes, OLD = $allocs_pp_old bytes")
|
||||
println()
|
||||
|
||||
println("MEDIAN TIMING:")
|
||||
println(" LinearElastic: NEW = $(median(bench_le_new.times)) ns, OLD = $(median(bench_le_old.times)) ns")
|
||||
println(" NeoHookean (AD): NEW = $(median(bench_nh_ad.times)) ns, OLD = $(median(bench_nh_old.times)) ns")
|
||||
println(" NeoHookean (Manual): NEW = $(median(bench_nh_manual.times)) ns")
|
||||
println(" PerfectPlasticity: NEW = $(median(bench_pp_new.times)) ns, OLD = $(median(bench_pp_old.times)) ns")
|
||||
println()
|
||||
|
||||
println("SPEEDUP (OLD / NEW):")
|
||||
println(" LinearElastic: $(round(speedup_le, digits=1))×")
|
||||
println(" NeoHookean (AD): $(round(speedup_nh_ad, digits=1))×")
|
||||
println(" NeoHookean (Manual): $(round(speedup_nh_manual, digits=1))×")
|
||||
println(" PerfectPlasticity: $(round(speedup_pp, digits=1))×")
|
||||
println()
|
||||
|
||||
println("AD OVERHEAD:")
|
||||
println(" NeoHookean: AD is $(round(ad_overhead, digits=1))× slower than manual derivatives")
|
||||
println()
|
||||
|
||||
avg_speedup = (speedup_le + speedup_nh_manual + speedup_pp) / 3
|
||||
println("AVERAGE SPEEDUP: $(round(avg_speedup, digits=1))× (using manual Neo-Hookean)")
|
||||
println()
|
||||
|
||||
# Validate claims
|
||||
println("VALIDATION OF CLAIMS:")
|
||||
println(" - Zero allocations for new approach: ",
|
||||
allocs_le_new == 0 && allocs_nh_ad == 0 && allocs_nh_manual == 0 && allocs_pp_new == 0 ? "✓ PASS" : "✗ FAIL")
|
||||
println(" - Manual derivatives outperform AD: ",
|
||||
median(bench_nh_manual.times) < median(bench_nh_ad.times) ? "✓ PASS" : "✗ FAIL")
|
||||
println(" - Type stability with NoState return: Check @code_warntype output above")
|
||||
println()
|
||||
|
||||
println("="^80)
|
||||
println("Benchmark complete! Results saved to: material_models_benchmark_results.txt")
|
||||
println("="^80)
|
||||
@@ -1,922 +0,0 @@
|
||||
================================================================================
|
||||
Material Models Performance Benchmark (Extended)
|
||||
================================================================================
|
||||
|
||||
|
||||
================================================================================
|
||||
NEWTON ITERATION STATE HANDLING EXAMPLE
|
||||
================================================================================
|
||||
|
||||
Example 1: Stateless Material (LinearElastic)
|
||||
--------------------------------------------------------------------------------
|
||||
Newton iteration with material state tracking:
|
||||
============================================================
|
||||
Iteration 1:
|
||||
strain: 0.0
|
||||
stress: 0.0
|
||||
state: NoState()
|
||||
Iteration 2:
|
||||
strain: 0.0005
|
||||
stress: 1.5741063022831637e8
|
||||
state: NoState()
|
||||
Iteration 3:
|
||||
strain: 0.00075
|
||||
stress: 2.3611594534247452e8
|
||||
state: NoState()
|
||||
→ Failed to converge!
|
||||
State NOT committed (keeping state_old)
|
||||
Result: state_final = NoState() (NoState, always)
|
||||
|
||||
Example 2: Stateful Material (PerfectPlasticity)
|
||||
--------------------------------------------------------------------------------
|
||||
Newton iteration with material state tracking:
|
||||
============================================================
|
||||
Iteration 1:
|
||||
strain: 0.0
|
||||
stress: 0.0
|
||||
state: PlasticityState{Float64}([0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0], 0.0)
|
||||
Iteration 2:
|
||||
strain: 0.001
|
||||
stress: 3.1482126045663273e8
|
||||
state: PlasticityState{Float64}([0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0], 0.0)
|
||||
Iteration 3:
|
||||
strain: 0.0015
|
||||
stress: 4.7223189068494904e8
|
||||
state: PlasticityState{Float64}([0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0], 0.0)
|
||||
→ Failed to converge!
|
||||
State NOT committed (keeping state_old)
|
||||
Result: state_final = PlasticityState{Float64}([0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0], 0.0) (plastic strain accumulated)
|
||||
|
||||
Key insight: State handling is IDENTICAL for all materials due to
|
||||
AbstractMaterialState type hierarchy. Assembly code doesn't need
|
||||
to know whether material is stateless or stateful!
|
||||
|
||||
Setting up materials and test cases...
|
||||
|
||||
Materials configured:
|
||||
- Linear Elastic: E = 200 GPa, ν = 0.3
|
||||
- Neo-Hookean (AD): μ ≈ 3.4 MPa, λ ≈ 45 MPa (automatic differentiation)
|
||||
- Neo-Hookean (Manual): μ ≈ 3.4 MPa, λ ≈ 45 MPa (hand-coded derivatives)
|
||||
- Perfect Plasticity: E = 200 GPa, σ_y = 250 MPa
|
||||
|
||||
Test strain: ε11 = 0.001 (uniaxial tension)
|
||||
|
||||
================================================================================
|
||||
TYPE STABILITY ANALYSIS
|
||||
================================================================================
|
||||
|
||||
Checking for type instabilities...
|
||||
|
||||
1. Linear Elastic (Tensors.jl):
|
||||
MethodInstance for compute_stress(::LinearElastic, ::SymmetricTensor{2, 3, Float64, 6}, ::NoState, ::Float64)
|
||||
from compute_stress(material::LinearElastic, ε::SymmetricTensor{2, 3, T}, state_old::NoState, Δt::Float64) where T @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:94
|
||||
Static Parameters
|
||||
T = Float64
|
||||
Arguments
|
||||
#self#::Core.Const(Main.compute_stress)
|
||||
material::LinearElastic
|
||||
ε::SymmetricTensor{2, 3, Float64, 6}
|
||||
state_old::Core.Const(NoState())
|
||||
Δt::Float64
|
||||
Locals
|
||||
𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
𝕀ˢʸᵐ::SymmetricTensor{4, 3, Float64, 36}
|
||||
σ::SymmetricTensor{2, 3, Float64, 6}
|
||||
I::SymmetricTensor{2, 3, Float64, 6}
|
||||
μ_val::Float64
|
||||
λ_val::Float64
|
||||
Body::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
|
||||
1 ─ %1 = Main.λ::Core.Const(Main.λ)
|
||||
│ (λ_val = (%1)(material))
|
||||
│ %3 = Main.μ::Core.Const(Main.μ)
|
||||
│ (μ_val = (%3)(material))
|
||||
│ %5 = Main.one::Core.Const(one)
|
||||
│ (I = (%5)(ε))
|
||||
│ %7 = Main.:+::Core.Const(+)
|
||||
│ %8 = Main.:*::Core.Const(*)
|
||||
│ %9 = λ_val::Float64
|
||||
│ %10 = Main.tr::Core.Const(LinearAlgebra.tr)
|
||||
│ %11 = (%10)(ε)::Float64
|
||||
│ %12 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ %13 = (%8)(%9, %11, %12)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %14 = Main.:*::Core.Const(*)
|
||||
│ %15 = Main.:*::Core.Const(*)
|
||||
│ %16 = μ_val::Float64
|
||||
│ %17 = (%15)(2, %16)::Float64
|
||||
│ %18 = (%14)(%17, ε)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (σ = (%7)(%13, %18))
|
||||
│ %20 = Main.one::Core.Const(one)
|
||||
│ %21 = Main.SymmetricTensor::Core.Const(SymmetricTensor)
|
||||
│ %22 = $(Expr(:static_parameter, 1))::Core.Const(Float64)
|
||||
│ %23 = Core.apply_type(%21, 4, 3, %22)::Core.Const(SymmetricTensor{4, 3, Float64})
|
||||
│ (𝕀ˢʸᵐ = (%20)(%23))
|
||||
│ %25 = Main.:+::Core.Const(+)
|
||||
│ %26 = Main.:⊗::Core.Const(Tensors.otimes)
|
||||
│ %27 = Main.:*::Core.Const(*)
|
||||
│ %28 = λ_val::Float64
|
||||
│ %29 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ %30 = (%27)(%28, %29)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %31 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ %32 = (%26)(%30, %31)::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %33 = Main.:*::Core.Const(*)
|
||||
│ %34 = Main.:*::Core.Const(*)
|
||||
│ %35 = μ_val::Float64
|
||||
│ %36 = (%34)(2, %35)::Float64
|
||||
│ %37 = 𝕀ˢʸᵐ::Core.Const([1.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.5 0.0; 0.5 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.5; 0.0 0.0 0.0; 0.5 0.0 0.0;;;; 0.0 0.5 0.0; 0.5 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.5; 0.0 0.5 0.0;;;; 0.0 0.0 0.5; 0.0 0.0 0.0; 0.5 0.0 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.5; 0.0 0.5 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 1.0])
|
||||
│ %38 = (%33)(%36, %37)::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ (𝔻 = (%25)(%32, %38))
|
||||
│ %40 = σ::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %41 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %42 = Main.NoState::Core.Const(NoState)
|
||||
│ %43 = (%42)()::Core.Const(NoState())
|
||||
│ %44 = Core.tuple(%40, %41, %43)::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
|
||||
└── return %44
|
||||
|
||||
|
||||
2. Linear Elastic (Old Voigt/Dict):
|
||||
MethodInstance for compute_stress_old(::LinearElasticOld, ::Vector{Float64}, ::Dict{String, Any}, ::Float64)
|
||||
from compute_stress_old(material::LinearElasticOld, ε_vec::Vector{Float64}, state_old::Dict{String, Any}, Δt::Float64) @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:413
|
||||
Arguments
|
||||
#self#::Core.Const(Main.compute_stress_old)
|
||||
material::LinearElasticOld
|
||||
ε_vec::Vector{Float64}
|
||||
state_old::Dict{String, Any}
|
||||
Δt::Float64
|
||||
Locals
|
||||
σ_vec::Vector{Float64}
|
||||
D::Matrix{Float64}
|
||||
Body::Tuple{Vector{Float64}, Matrix{Float64}, Dict{String, Any}}
|
||||
1 ─ %1 = Main.constitutive_matrix::Core.Const(Main.constitutive_matrix)
|
||||
│ (D = (%1)(material))
|
||||
│ %3 = Main.:*::Core.Const(*)
|
||||
│ %4 = D::Matrix{Float64}
|
||||
│ (σ_vec = (%3)(%4, ε_vec))
|
||||
│ %6 = σ_vec::Vector{Float64}
|
||||
│ %7 = D::Matrix{Float64}
|
||||
│ %8 = Core.tuple(%6, %7, state_old)::Tuple{Vector{Float64}, Matrix{Float64}, Dict{String, Any}}
|
||||
└── return %8
|
||||
|
||||
|
||||
3. Neo-Hookean AD (Tensors.jl with automatic differentiation):
|
||||
MethodInstance for compute_stress(::NeoHookeanAD, ::SymmetricTensor{2, 3, Float64, 6}, ::NoState, ::Float64)
|
||||
from compute_stress(material::NeoHookeanAD, E::SymmetricTensor{2, 3, T}, state_old::NoState, Δt::Float64) where T @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:147
|
||||
Static Parameters
|
||||
T = Float64
|
||||
Arguments
|
||||
#self#::Core.Const(Main.compute_stress)
|
||||
material::NeoHookeanAD
|
||||
E::SymmetricTensor{2, 3, Float64, 6}
|
||||
state_old::Core.Const(NoState())
|
||||
Δt::Float64
|
||||
Locals
|
||||
@_6::Int64
|
||||
S::SymmetricTensor{2, 3, Float64, 6}
|
||||
𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
ψ::var"#ψ#compute_stress##0"{NeoHookeanAD}
|
||||
C::SymmetricTensor{2, 3, Float64, 6}
|
||||
I::SymmetricTensor{2, 3, Float64, 6}
|
||||
Body::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
|
||||
1 ─ %1 = Main.one::Core.Const(one)
|
||||
│ (I = (%1)(E))
|
||||
│ %3 = Main.:+::Core.Const(+)
|
||||
│ %4 = Main.:*::Core.Const(*)
|
||||
│ %5 = (%4)(2, E)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %6 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ (C = (%3)(%5, %6))
|
||||
│ %8 = Main.:(var"#ψ#compute_stress##0")::Core.Const(var"#ψ#compute_stress##0")
|
||||
│ %9 = Core._typeof_captured_variable(material)::Core.Const(NeoHookeanAD)
|
||||
│ %10 = Core.apply_type(%8, %9)::Core.Const(var"#ψ#compute_stress##0"{NeoHookeanAD})
|
||||
│ (ψ = %new(%10, material))
|
||||
│ %12 = Main.hessian::Core.Const(Tensors.hessian)
|
||||
│ %13 = ψ::var"#ψ#compute_stress##0"{NeoHookeanAD}
|
||||
│ %14 = C::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %15 = (%12)(%13, %14, :all)::Tuple{SymmetricTensor{4, 3, Float64, 36}, SymmetricTensor{2, 3, Float64, 6}, Float64}
|
||||
│ %16 = Base.indexed_iterate(%15, 1)::Core.PartialStruct(Tuple{SymmetricTensor{4, 3, Float64, 36}, Int64}, Any[SymmetricTensor{4, 3, Float64, 36}, Core.Const(2)])
|
||||
│ (𝔻 = Core.getfield(%16, 1))
|
||||
│ (@_6 = Core.getfield(%16, 2))
|
||||
│ %19 = @_6::Core.Const(2)
|
||||
│ %20 = Base.indexed_iterate(%15, 2, %19)::Core.PartialStruct(Tuple{SymmetricTensor{2, 3, Float64, 6}, Int64}, Any[SymmetricTensor{2, 3, Float64, 6}, Core.Const(3)])
|
||||
│ (S = Core.getfield(%20, 1))
|
||||
│ %22 = Main.:*::Core.Const(*)
|
||||
│ %23 = S::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (S = (%22)(2, %23))
|
||||
│ %25 = Main.:*::Core.Const(*)
|
||||
│ %26 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ (𝔻 = (%25)(4, %26))
|
||||
│ %28 = S::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %29 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %30 = Main.NoState::Core.Const(NoState)
|
||||
│ %31 = (%30)()::Core.Const(NoState())
|
||||
│ %32 = Core.tuple(%28, %29, %31)::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
|
||||
└── return %32
|
||||
|
||||
|
||||
4. Neo-Hookean Manual (Tensors.jl with hand-coded derivatives):
|
||||
MethodInstance for compute_stress(::NeoHookeanManual, ::SymmetricTensor{2, 3, Float64, 6}, ::NoState, ::Float64)
|
||||
from compute_stress(material::NeoHookeanManual, E::SymmetricTensor{2, 3, T}, state_old::NoState, Δt::Float64) where T @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:203
|
||||
Static Parameters
|
||||
T = Float64
|
||||
Arguments
|
||||
#self#::Core.Const(Main.compute_stress)
|
||||
material::NeoHookeanManual
|
||||
E::SymmetricTensor{2, 3, Float64, 6}
|
||||
state_old::Core.Const(NoState())
|
||||
Δt::Float64
|
||||
Locals
|
||||
𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
𝔻₂::SymmetricTensor{4, 3, Float64, 36}
|
||||
coeff::Float64
|
||||
𝕀ˢʸᵐ::SymmetricTensor{4, 3, Float64, 36}
|
||||
𝔻₁::SymmetricTensor{4, 3, Float64, 36}
|
||||
S::SymmetricTensor{2, 3, Float64, 6}
|
||||
C_inv::SymmetricTensor{2, 3, Float64, 6}
|
||||
J::Float64
|
||||
C::SymmetricTensor{2, 3, Float64, 6}
|
||||
I::SymmetricTensor{2, 3, Float64, 6}
|
||||
λ::Float64
|
||||
μ::Float64
|
||||
Body::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
|
||||
1 ─ %1 = Base.getproperty(material, :μ)::Float64
|
||||
│ %2 = Base.getproperty(material, :λ)::Float64
|
||||
│ (μ = %1)
|
||||
│ (λ = %2)
|
||||
│ %5 = Main.one::Core.Const(one)
|
||||
│ (I = (%5)(E))
|
||||
│ %7 = Main.:+::Core.Const(+)
|
||||
│ %8 = Main.:*::Core.Const(*)
|
||||
│ %9 = (%8)(2, E)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %10 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ (C = (%7)(%9, %10))
|
||||
│ %12 = Main.:√::Core.Const(sqrt)
|
||||
│ %13 = Main.det::Core.Const(LinearAlgebra.det)
|
||||
│ %14 = C::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %15 = (%13)(%14)::Float64
|
||||
│ (J = (%12)(%15))
|
||||
│ %17 = Main.inv::Core.Const(inv)
|
||||
│ %18 = C::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (C_inv = (%17)(%18))
|
||||
│ %20 = Main.:+::Core.Const(+)
|
||||
│ %21 = Main.:*::Core.Const(*)
|
||||
│ %22 = μ::Float64
|
||||
│ %23 = Main.:-::Core.Const(-)
|
||||
│ %24 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ %25 = C_inv::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %26 = (%23)(%24, %25)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %27 = (%21)(%22, %26)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %28 = Main.:*::Core.Const(*)
|
||||
│ %29 = λ::Float64
|
||||
│ %30 = Main.log::Core.Const(log)
|
||||
│ %31 = J::Float64
|
||||
│ %32 = (%30)(%31)::Float64
|
||||
│ %33 = C_inv::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %34 = (%28)(%29, %32, %33)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (S = (%20)(%27, %34))
|
||||
│ %36 = Main.:*::Core.Const(*)
|
||||
│ %37 = λ::Float64
|
||||
│ %38 = Main.:⊗::Core.Const(Tensors.otimes)
|
||||
│ %39 = C_inv::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %40 = C_inv::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %41 = (%38)(%39, %40)::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ (𝔻₁ = (%36)(%37, %41))
|
||||
│ %43 = Main.one::Core.Const(one)
|
||||
│ %44 = Main.SymmetricTensor::Core.Const(SymmetricTensor)
|
||||
│ %45 = $(Expr(:static_parameter, 1))::Core.Const(Float64)
|
||||
│ %46 = Core.apply_type(%44, 4, 3, %45)::Core.Const(SymmetricTensor{4, 3, Float64})
|
||||
│ (𝕀ˢʸᵐ = (%43)(%46))
|
||||
│ %48 = Main.:*::Core.Const(*)
|
||||
│ %49 = Main.:-::Core.Const(-)
|
||||
│ %50 = μ::Float64
|
||||
│ %51 = Main.:*::Core.Const(*)
|
||||
│ %52 = λ::Float64
|
||||
│ %53 = Main.log::Core.Const(log)
|
||||
│ %54 = J::Float64
|
||||
│ %55 = (%53)(%54)::Float64
|
||||
│ %56 = (%51)(%52, %55)::Float64
|
||||
│ %57 = (%49)(%50, %56)::Float64
|
||||
│ (coeff = (%48)(2, %57))
|
||||
│ %59 = Main.:*::Core.Const(*)
|
||||
│ %60 = Main.:-::Core.Const(-)
|
||||
│ %61 = coeff::Float64
|
||||
│ %62 = (%60)(%61)::Float64
|
||||
│ %63 = Main.inv_symmetric_outer::Core.Const(Main.inv_symmetric_outer)
|
||||
│ %64 = C_inv::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %65 = (%63)(%64)::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ (𝔻₂ = (%59)(%62, %65))
|
||||
│ %67 = Main.:+::Core.Const(+)
|
||||
│ %68 = 𝔻₁::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %69 = 𝔻₂::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ (𝔻 = (%67)(%68, %69))
|
||||
│ %71 = S::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %72 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %73 = Main.NoState::Core.Const(NoState)
|
||||
│ %74 = (%73)()::Core.Const(NoState())
|
||||
│ %75 = Core.tuple(%71, %72, %74)::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, NoState}
|
||||
└── return %75
|
||||
|
||||
|
||||
5. Perfect Plasticity (Tensors.jl):
|
||||
MethodInstance for compute_stress(::PerfectPlasticity, ::SymmetricTensor{2, 3, Float64, 6}, ::PlasticityState{Float64}, ::Float64)
|
||||
from compute_stress(material::PerfectPlasticity, ε::SymmetricTensor{2, 3, T}, state_old::PlasticityState{T}, Δt::Float64) where T @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:328
|
||||
Static Parameters
|
||||
T = Float64
|
||||
Arguments
|
||||
#self#::Core.Const(Main.compute_stress)
|
||||
material::PerfectPlasticity
|
||||
ε::SymmetricTensor{2, 3, Float64, 6}
|
||||
state_old::PlasticityState{Float64}
|
||||
Δt::Float64
|
||||
Locals
|
||||
𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
β::Float64
|
||||
θ::Float64
|
||||
state_new::PlasticityState{Float64}
|
||||
α_new::Float64
|
||||
ε_p_new::SymmetricTensor{2, 3, Float64, 6}
|
||||
n::SymmetricTensor{2, 3, Float64, 6}
|
||||
Δγ::Float64
|
||||
σ::SymmetricTensor{2, 3, Float64, 6}
|
||||
p::Float64
|
||||
s_trial::SymmetricTensor{2, 3, Float64, 6}
|
||||
f::Float64
|
||||
σ_eq_trial::Float64
|
||||
σ_trial::SymmetricTensor{2, 3, Float64, 6}
|
||||
ε_e::SymmetricTensor{2, 3, Float64, 6}
|
||||
𝔻ᵉ::SymmetricTensor{4, 3, Float64, 36}
|
||||
𝕀ˢʸᵐ::SymmetricTensor{4, 3, Float64, 36}
|
||||
I::SymmetricTensor{2, 3, Float64, 6}
|
||||
σ_y::Float64
|
||||
μ_val::Float64
|
||||
λ_val::Float64
|
||||
Body::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, PlasticityState{Float64}}
|
||||
1 ─ Core.NewvarNode(:(𝔻))
|
||||
│ Core.NewvarNode(:(β))
|
||||
│ Core.NewvarNode(:(θ))
|
||||
│ Core.NewvarNode(:(state_new))
|
||||
│ Core.NewvarNode(:(α_new))
|
||||
│ Core.NewvarNode(:(ε_p_new))
|
||||
│ Core.NewvarNode(:(n))
|
||||
│ Core.NewvarNode(:(Δγ))
|
||||
│ Core.NewvarNode(:(σ))
|
||||
│ Core.NewvarNode(:(p))
|
||||
│ Core.NewvarNode(:(s_trial))
|
||||
│ %12 = Main.λ::Core.Const(Main.λ)
|
||||
│ (λ_val = (%12)(material))
|
||||
│ %14 = Main.μ::Core.Const(Main.μ)
|
||||
│ (μ_val = (%14)(material))
|
||||
│ (σ_y = Base.getproperty(material, :σ_y))
|
||||
│ %17 = Main.one::Core.Const(one)
|
||||
│ (I = (%17)(ε))
|
||||
│ %19 = Main.one::Core.Const(one)
|
||||
│ %20 = Main.SymmetricTensor::Core.Const(SymmetricTensor)
|
||||
│ %21 = $(Expr(:static_parameter, 1))::Core.Const(Float64)
|
||||
│ %22 = Core.apply_type(%20, 4, 3, %21)::Core.Const(SymmetricTensor{4, 3, Float64})
|
||||
│ (𝕀ˢʸᵐ = (%19)(%22))
|
||||
│ %24 = Main.:+::Core.Const(+)
|
||||
│ %25 = Main.:⊗::Core.Const(Tensors.otimes)
|
||||
│ %26 = Main.:*::Core.Const(*)
|
||||
│ %27 = λ_val::Float64
|
||||
│ %28 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ %29 = (%26)(%27, %28)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %30 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ %31 = (%25)(%29, %30)::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %32 = Main.:*::Core.Const(*)
|
||||
│ %33 = Main.:*::Core.Const(*)
|
||||
│ %34 = μ_val::Float64
|
||||
│ %35 = (%33)(2, %34)::Float64
|
||||
│ %36 = 𝕀ˢʸᵐ::Core.Const([1.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.5 0.0; 0.5 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.5; 0.0 0.0 0.0; 0.5 0.0 0.0;;;; 0.0 0.5 0.0; 0.5 0.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.5; 0.0 0.5 0.0;;;; 0.0 0.0 0.5; 0.0 0.0 0.0; 0.5 0.0 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.5; 0.0 0.5 0.0;;; 0.0 0.0 0.0; 0.0 0.0 0.0; 0.0 0.0 1.0])
|
||||
│ %37 = (%32)(%35, %36)::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ (𝔻ᵉ = (%24)(%31, %37))
|
||||
│ %39 = Main.:-::Core.Const(-)
|
||||
│ %40 = Base.getproperty(state_old, :ε_p)::SYMMETRICTENSOR{2, 3, FLOAT64}
|
||||
│ (ε_e = (%39)(ε, %40))
|
||||
│ %42 = Main.:+::Core.Const(+)
|
||||
│ %43 = Main.:*::Core.Const(*)
|
||||
│ %44 = λ_val::Float64
|
||||
│ %45 = Main.tr::Core.Const(LinearAlgebra.tr)
|
||||
│ %46 = ε_e::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %47 = (%45)(%46)::Float64
|
||||
│ %48 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ %49 = (%43)(%44, %47, %48)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %50 = Main.:*::Core.Const(*)
|
||||
│ %51 = Main.:*::Core.Const(*)
|
||||
│ %52 = μ_val::Float64
|
||||
│ %53 = (%51)(2, %52)::Float64
|
||||
│ %54 = ε_e::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %55 = (%50)(%53, %54)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (σ_trial = (%42)(%49, %55))
|
||||
│ %57 = Main.von_mises_stress::Core.Const(Main.von_mises_stress)
|
||||
│ %58 = σ_trial::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (σ_eq_trial = (%57)(%58))
|
||||
│ %60 = Main.:-::Core.Const(-)
|
||||
│ %61 = σ_eq_trial::Float64
|
||||
│ %62 = σ_y::Float64
|
||||
│ (f = (%60)(%61, %62))
|
||||
│ %64 = Main.:≤::Core.Const(<=)
|
||||
│ %65 = f::Float64
|
||||
│ %66 = (%64)(%65, 0.0)::Bool
|
||||
└── goto #3 if not %66
|
||||
2 ─ %68 = σ_trial::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (σ = %68)
|
||||
│ %70 = 𝔻ᵉ::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ (𝔻 = %70)
|
||||
│ %72 = state_old::PlasticityState{Float64}
|
||||
│ (state_new = %72)
|
||||
└── goto #4
|
||||
3 ─ %75 = Main.dev::Core.Const(Tensors.dev)
|
||||
│ %76 = σ_trial::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (s_trial = (%75)(%76))
|
||||
│ %78 = Main.:/::Core.Const(/)
|
||||
│ %79 = Main.tr::Core.Const(LinearAlgebra.tr)
|
||||
│ %80 = σ_trial::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %81 = (%79)(%80)::Float64
|
||||
│ (p = (%78)(%81, 3))
|
||||
│ %83 = Main.:+::Core.Const(+)
|
||||
│ %84 = Main.:*::Core.Const(*)
|
||||
│ %85 = p::Float64
|
||||
│ %86 = I::Core.Const([1.0 0.0 0.0; 0.0 1.0 0.0; 0.0 0.0 1.0])
|
||||
│ %87 = (%84)(%85, %86)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %88 = Main.:*::Core.Const(*)
|
||||
│ %89 = Main.:/::Core.Const(/)
|
||||
│ %90 = σ_y::Float64
|
||||
│ %91 = σ_eq_trial::Float64
|
||||
│ %92 = (%89)(%90, %91)::Float64
|
||||
│ %93 = s_trial::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %94 = (%88)(%92, %93)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (σ = (%83)(%87, %94))
|
||||
│ %96 = Main.:/::Core.Const(/)
|
||||
│ %97 = f::Float64
|
||||
│ %98 = Main.:*::Core.Const(*)
|
||||
│ %99 = μ_val::Float64
|
||||
│ %100 = (%98)(3, %99)::Float64
|
||||
│ (Δγ = (%96)(%97, %100))
|
||||
│ %102 = Main.:/::Core.Const(/)
|
||||
│ %103 = Main.:*::Core.Const(*)
|
||||
│ %104 = Main.:√::Core.Const(sqrt)
|
||||
│ %105 = Main.:/::Core.Const(/)
|
||||
│ %106 = (%105)(3, 2)::Core.Const(1.5)
|
||||
│ %107 = (%104)(%106)::Core.Const(1.224744871391589)
|
||||
│ %108 = s_trial::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %109 = (%103)(%107, %108)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %110 = σ_eq_trial::Float64
|
||||
│ (n = (%102)(%109, %110))
|
||||
│ %112 = Main.:+::Core.Const(+)
|
||||
│ %113 = Base.getproperty(state_old, :ε_p)::SYMMETRICTENSOR{2, 3, FLOAT64}
|
||||
│ %114 = Main.:*::Core.Const(*)
|
||||
│ %115 = Δγ::Float64
|
||||
│ %116 = n::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %117 = (%114)(%115, %116)::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ (ε_p_new = (%112)(%113, %117))
|
||||
│ %119 = Main.:+::Core.Const(+)
|
||||
│ %120 = Base.getproperty(state_old, :α)::Float64
|
||||
│ %121 = Δγ::Float64
|
||||
│ (α_new = (%119)(%120, %121))
|
||||
│ %123 = Main.PlasticityState::Core.Const(PlasticityState)
|
||||
│ %124 = ε_p_new::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %125 = α_new::Float64
|
||||
│ (state_new = (%123)(%124, %125))
|
||||
│ %127 = Main.:-::Core.Const(-)
|
||||
│ %128 = Main.:/::Core.Const(/)
|
||||
│ %129 = σ_y::Float64
|
||||
│ %130 = σ_eq_trial::Float64
|
||||
│ %131 = (%128)(%129, %130)::Float64
|
||||
│ (θ = (%127)(1, %131))
|
||||
│ %133 = Main.:/::Core.Const(/)
|
||||
│ %134 = Main.:*::Core.Const(*)
|
||||
│ %135 = Main.:^::Core.Const(^)
|
||||
│ %136 = μ_val::Float64
|
||||
│ %137 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
|
||||
│ %138 = (%137)()::Core.Const(Val{2}())
|
||||
│ %139 = Base.literal_pow(%135, %136, %138)::Float64
|
||||
│ %140 = (%134)(6, %139)::Float64
|
||||
│ %141 = Main.:+::Core.Const(+)
|
||||
│ %142 = Main.:*::Core.Const(*)
|
||||
│ %143 = μ_val::Float64
|
||||
│ %144 = (%142)(3, %143)::Float64
|
||||
│ %145 = Main.:*::Core.Const(*)
|
||||
│ %146 = θ::Float64
|
||||
│ %147 = Main.:*::Core.Const(*)
|
||||
│ %148 = μ_val::Float64
|
||||
│ %149 = (%147)(3, %148)::Float64
|
||||
│ %150 = (%145)(%146, %149)::Float64
|
||||
│ %151 = (%141)(%144, %150)::Float64
|
||||
│ (β = (%133)(%140, %151))
|
||||
│ %153 = Main.:-::Core.Const(-)
|
||||
│ %154 = 𝔻ᵉ::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %155 = Main.:*::Core.Const(*)
|
||||
│ %156 = β::Float64
|
||||
│ %157 = Main.:⊗::Core.Const(Tensors.otimes)
|
||||
│ %158 = n::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %159 = n::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %160 = (%157)(%158, %159)::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %161 = (%155)(%156, %160)::SymmetricTensor{4, 3, Float64, 36}
|
||||
└── (𝔻 = (%153)(%154, %161))
|
||||
4 ┄ %163 = σ::SymmetricTensor{2, 3, Float64, 6}
|
||||
│ %164 = 𝔻::SymmetricTensor{4, 3, Float64, 36}
|
||||
│ %165 = state_new::PlasticityState{Float64}
|
||||
│ %166 = Core.tuple(%163, %164, %165)::Tuple{SymmetricTensor{2, 3, Float64, 6}, SymmetricTensor{4, 3, Float64, 36}, PlasticityState{Float64}}
|
||||
└── return %166
|
||||
|
||||
|
||||
6. Perfect Plasticity (Old Dict):
|
||||
MethodInstance for compute_stress_old(::PerfectPlasticityOld, ::Vector{Float64}, ::Dict{String, Any}, ::Float64)
|
||||
from compute_stress_old(material::PerfectPlasticityOld, ε_vec::Vector{Float64}, state_old::Dict{String, Any}, Δt::Float64) @ Main ~/dev/JuliaFEM.jl/benchmarks/material_models_benchmark.jl:455
|
||||
Arguments
|
||||
#self#::Core.Const(Main.compute_stress_old)
|
||||
material::PerfectPlasticityOld
|
||||
ε_vec::Vector{Float64}
|
||||
state_old::Dict{String, Any}
|
||||
Δt::Float64
|
||||
Locals
|
||||
@_6::ANY
|
||||
@_7::ANY
|
||||
σ_vec::ANY
|
||||
n_vec::ANY
|
||||
Δγ::ANY
|
||||
factor::ANY
|
||||
state_new::Dict{String, Any}
|
||||
f::ANY
|
||||
σ_eq::ANY
|
||||
dev_vec::ANY
|
||||
p::ANY
|
||||
s13::ANY
|
||||
s23::ANY
|
||||
s12::ANY
|
||||
s33::ANY
|
||||
s22::ANY
|
||||
s11::ANY
|
||||
σ_trial_vec::ANY
|
||||
ε_e_vec::ANY
|
||||
D::Matrix{Float64}
|
||||
ε_p_vec::ANY
|
||||
Body::TUPLE{ANY, MATRIX{FLOAT64}, DICT{STRING, ANY}}
|
||||
1 ─ Core.NewvarNode(:(@_6))
|
||||
│ Core.NewvarNode(:(@_7))
|
||||
│ Core.NewvarNode(:(σ_vec))
|
||||
│ Core.NewvarNode(:(n_vec))
|
||||
│ Core.NewvarNode(:(Δγ))
|
||||
│ Core.NewvarNode(:(factor))
|
||||
│ Core.NewvarNode(:(state_new))
|
||||
│ Core.NewvarNode(:(f))
|
||||
│ Core.NewvarNode(:(σ_eq))
|
||||
│ Core.NewvarNode(:(dev_vec))
|
||||
│ Core.NewvarNode(:(p))
|
||||
│ Core.NewvarNode(:(s13))
|
||||
│ Core.NewvarNode(:(s23))
|
||||
│ Core.NewvarNode(:(s12))
|
||||
│ Core.NewvarNode(:(s33))
|
||||
│ Core.NewvarNode(:(s22))
|
||||
│ Core.NewvarNode(:(s11))
|
||||
│ Core.NewvarNode(:(σ_trial_vec))
|
||||
│ Core.NewvarNode(:(ε_e_vec))
|
||||
│ Core.NewvarNode(:(D))
|
||||
│ Core.NewvarNode(:(ε_p_vec))
|
||||
│ %22 = Main.haskey::Core.Const(haskey)
|
||||
│ %23 = (%22)(state_old, "epsilon_plastic")::Bool
|
||||
└── goto #3 if not %23
|
||||
2 ─ (ε_p_vec = Base.getindex(state_old, "epsilon_plastic"))
|
||||
└── goto #4
|
||||
3 ─ %27 = Main.zeros::Core.Const(zeros)
|
||||
└── (ε_p_vec = (%27)(6))
|
||||
4 ┄ %29 = Main.constitutive_matrix::Core.Const(Main.constitutive_matrix)
|
||||
│ %30 = Main.LinearElasticOld::Core.Const(LinearElasticOld)
|
||||
│ %31 = Base.getproperty(material, :E)::Float64
|
||||
│ %32 = Base.getproperty(material, :ν)::Float64
|
||||
│ %33 = (%30)(%31, %32)::LinearElasticOld
|
||||
│ (D = (%29)(%33))
|
||||
│ %35 = Main.:-::Core.Const(-)
|
||||
│ %36 = ε_p_vec::ANY
|
||||
│ (ε_e_vec = (%35)(ε_vec, %36))
|
||||
│ %38 = Main.:*::Core.Const(*)
|
||||
│ %39 = D::Matrix{Float64}
|
||||
│ %40 = ε_e_vec::ANY
|
||||
│ (σ_trial_vec = (%38)(%39, %40))
|
||||
│ %42 = σ_trial_vec::ANY
|
||||
│ %43 = Main.:(:)::Core.Const(Colon())
|
||||
│ %44 = (%43)(1, 3)::Core.Const(1:3)
|
||||
│ %45 = Base.getindex(%42, %44)::ANY
|
||||
│ %46 = Base.indexed_iterate(%45, 1)::ANY
|
||||
│ (s11 = Core.getfield(%46, 1))
|
||||
│ (@_7 = Core.getfield(%46, 2))
|
||||
│ %49 = @_7::ANY
|
||||
│ %50 = Base.indexed_iterate(%45, 2, %49)::ANY
|
||||
│ (s22 = Core.getfield(%50, 1))
|
||||
│ (@_7 = Core.getfield(%50, 2))
|
||||
│ %53 = @_7::ANY
|
||||
│ %54 = Base.indexed_iterate(%45, 3, %53)::ANY
|
||||
│ (s33 = Core.getfield(%54, 1))
|
||||
│ %56 = σ_trial_vec::ANY
|
||||
│ %57 = Main.:(:)::Core.Const(Colon())
|
||||
│ %58 = (%57)(4, 6)::Core.Const(4:6)
|
||||
│ %59 = Base.getindex(%56, %58)::ANY
|
||||
│ %60 = Base.indexed_iterate(%59, 1)::ANY
|
||||
│ (s12 = Core.getfield(%60, 1))
|
||||
│ (@_6 = Core.getfield(%60, 2))
|
||||
│ %63 = @_6::ANY
|
||||
│ %64 = Base.indexed_iterate(%59, 2, %63)::ANY
|
||||
│ (s23 = Core.getfield(%64, 1))
|
||||
│ (@_6 = Core.getfield(%64, 2))
|
||||
│ %67 = @_6::ANY
|
||||
│ %68 = Base.indexed_iterate(%59, 3, %67)::ANY
|
||||
│ (s13 = Core.getfield(%68, 1))
|
||||
│ %70 = Main.:/::Core.Const(/)
|
||||
│ %71 = Main.:+::Core.Const(+)
|
||||
│ %72 = s11::ANY
|
||||
│ %73 = s22::ANY
|
||||
│ %74 = s33::ANY
|
||||
│ %75 = (%71)(%72, %73, %74)::ANY
|
||||
│ (p = (%70)(%75, 3))
|
||||
│ %77 = Main.:-::Core.Const(-)
|
||||
│ %78 = s11::ANY
|
||||
│ %79 = p::ANY
|
||||
│ %80 = (%77)(%78, %79)::ANY
|
||||
│ %81 = Main.:-::Core.Const(-)
|
||||
│ %82 = s22::ANY
|
||||
│ %83 = p::ANY
|
||||
│ %84 = (%81)(%82, %83)::ANY
|
||||
│ %85 = Main.:-::Core.Const(-)
|
||||
│ %86 = s33::ANY
|
||||
│ %87 = p::ANY
|
||||
│ %88 = (%85)(%86, %87)::ANY
|
||||
│ %89 = s12::ANY
|
||||
│ %90 = s23::ANY
|
||||
│ %91 = s13::ANY
|
||||
│ (dev_vec = Base.vect(%80, %84, %88, %89, %90, %91))
|
||||
│ %93 = Main.:√::Core.Const(sqrt)
|
||||
│ %94 = Main.:*::Core.Const(*)
|
||||
│ %95 = Main.:/::Core.Const(/)
|
||||
│ %96 = (%95)(3, 2)::Core.Const(1.5)
|
||||
│ %97 = Main.:+::Core.Const(+)
|
||||
│ %98 = Main.:^::Core.Const(^)
|
||||
│ %99 = dev_vec::ANY
|
||||
│ %100 = Base.getindex(%99, 1)::ANY
|
||||
│ %101 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
|
||||
│ %102 = (%101)()::Core.Const(Val{2}())
|
||||
│ %103 = Base.literal_pow(%98, %100, %102)::ANY
|
||||
│ %104 = Main.:^::Core.Const(^)
|
||||
│ %105 = dev_vec::ANY
|
||||
│ %106 = Base.getindex(%105, 2)::ANY
|
||||
│ %107 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
|
||||
│ %108 = (%107)()::Core.Const(Val{2}())
|
||||
│ %109 = Base.literal_pow(%104, %106, %108)::ANY
|
||||
│ %110 = Main.:^::Core.Const(^)
|
||||
│ %111 = dev_vec::ANY
|
||||
│ %112 = Base.getindex(%111, 3)::ANY
|
||||
│ %113 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
|
||||
│ %114 = (%113)()::Core.Const(Val{2}())
|
||||
│ %115 = Base.literal_pow(%110, %112, %114)::ANY
|
||||
│ %116 = Main.:*::Core.Const(*)
|
||||
│ %117 = Main.:+::Core.Const(+)
|
||||
│ %118 = Main.:^::Core.Const(^)
|
||||
│ %119 = dev_vec::ANY
|
||||
│ %120 = Base.getindex(%119, 4)::ANY
|
||||
│ %121 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
|
||||
│ %122 = (%121)()::Core.Const(Val{2}())
|
||||
│ %123 = Base.literal_pow(%118, %120, %122)::ANY
|
||||
│ %124 = Main.:^::Core.Const(^)
|
||||
│ %125 = dev_vec::ANY
|
||||
│ %126 = Base.getindex(%125, 5)::ANY
|
||||
│ %127 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
|
||||
│ %128 = (%127)()::Core.Const(Val{2}())
|
||||
│ %129 = Base.literal_pow(%124, %126, %128)::ANY
|
||||
│ %130 = Main.:^::Core.Const(^)
|
||||
│ %131 = dev_vec::ANY
|
||||
│ %132 = Base.getindex(%131, 6)::ANY
|
||||
│ %133 = Core.apply_type(Base.Val, 2)::Core.Const(Val{2})
|
||||
│ %134 = (%133)()::Core.Const(Val{2}())
|
||||
│ %135 = Base.literal_pow(%130, %132, %134)::ANY
|
||||
│ %136 = (%117)(%123, %129, %135)::ANY
|
||||
│ %137 = (%116)(2, %136)::ANY
|
||||
│ %138 = (%97)(%103, %109, %115, %137)::ANY
|
||||
│ %139 = (%94)(%96, %138)::ANY
|
||||
│ (σ_eq = (%93)(%139))
|
||||
│ %141 = Main.:-::Core.Const(-)
|
||||
│ %142 = σ_eq::ANY
|
||||
│ %143 = Base.getproperty(material, :σ_y)::Float64
|
||||
│ (f = (%141)(%142, %143))
|
||||
│ %145 = Main.copy::Core.Const(copy)
|
||||
│ (state_new = (%145)(state_old))
|
||||
│ %147 = Main.:>::Core.Const(>)
|
||||
│ %148 = f::ANY
|
||||
│ %149 = (%147)(%148, 0.0)::ANY
|
||||
└── goto #6 if not %149
|
||||
5 ─ %151 = Main.:/::Core.Const(/)
|
||||
│ %152 = Base.getproperty(material, :σ_y)::Float64
|
||||
│ %153 = σ_eq::ANY
|
||||
│ (factor = (%151)(%152, %153))
|
||||
│ %155 = Main.:+::Core.Const(+)
|
||||
│ %156 = p::ANY
|
||||
│ %157 = p::ANY
|
||||
│ %158 = p::ANY
|
||||
│ %159 = Base.vect(%156, %157, %158, 0.0, 0.0, 0.0)::ANY
|
||||
│ %160 = Main.:*::Core.Const(*)
|
||||
│ %161 = factor::ANY
|
||||
│ %162 = dev_vec::ANY
|
||||
│ %163 = (%160)(%161, %162)::ANY
|
||||
│ (σ_vec = (%155)(%159, %163))
|
||||
│ %165 = Main.:/::Core.Const(/)
|
||||
│ %166 = f::ANY
|
||||
│ %167 = Main.:/::Core.Const(/)
|
||||
│ %168 = Main.:*::Core.Const(*)
|
||||
│ %169 = Base.getproperty(material, :E)::Float64
|
||||
│ %170 = (%168)(3, %169)::Float64
|
||||
│ %171 = Main.:*::Core.Const(*)
|
||||
│ %172 = Main.:+::Core.Const(+)
|
||||
│ %173 = Base.getproperty(material, :ν)::Float64
|
||||
│ %174 = (%172)(1, %173)::Float64
|
||||
│ %175 = (%171)(2, %174)::Float64
|
||||
│ %176 = (%167)(%170, %175)::Float64
|
||||
│ (Δγ = (%165)(%166, %176))
|
||||
│ %178 = Main.:/::Core.Const(/)
|
||||
│ %179 = Main.:*::Core.Const(*)
|
||||
│ %180 = Main.:√::Core.Const(sqrt)
|
||||
│ %181 = Main.:/::Core.Const(/)
|
||||
│ %182 = (%181)(3, 2)::Core.Const(1.5)
|
||||
│ %183 = (%180)(%182)::Core.Const(1.224744871391589)
|
||||
│ %184 = dev_vec::ANY
|
||||
│ %185 = (%179)(%183, %184)::ANY
|
||||
│ %186 = σ_eq::ANY
|
||||
│ (n_vec = (%178)(%185, %186))
|
||||
│ %188 = Main.:+::Core.Const(+)
|
||||
│ %189 = ε_p_vec::ANY
|
||||
│ %190 = Main.:*::Core.Const(*)
|
||||
│ %191 = Δγ::ANY
|
||||
│ %192 = n_vec::ANY
|
||||
│ %193 = (%190)(%191, %192)::ANY
|
||||
│ %194 = (%188)(%189, %193)::ANY
|
||||
│ %195 = state_new::Dict{String, Any}
|
||||
│ Base.setindex!(%195, %194, "epsilon_plastic")
|
||||
└── goto #7
|
||||
6 ─ %198 = σ_trial_vec::ANY
|
||||
└── (σ_vec = %198)
|
||||
7 ┄ %200 = σ_vec::ANY
|
||||
│ %201 = D::Matrix{Float64}
|
||||
│ %202 = state_new::Dict{String, Any}
|
||||
│ %203 = Core.tuple(%200, %201, %202)::TUPLE{ANY, MATRIX{FLOAT64}, DICT{STRING, ANY}}
|
||||
└── return %203
|
||||
|
||||
|
||||
================================================================================
|
||||
ALLOCATION TESTS
|
||||
================================================================================
|
||||
|
||||
Testing for allocations (should be 0 for new approach)...
|
||||
|
||||
1. Linear Elastic
|
||||
NEW (Tensors.jl):
|
||||
Allocations: 0 bytes
|
||||
OLD (Voigt/Dict):
|
||||
Allocations: 496 bytes
|
||||
|
||||
2. Neo-Hookean
|
||||
NEW (Tensors.jl + AD):
|
||||
Allocations: 0 bytes
|
||||
NEW (Tensors.jl + Manual):
|
||||
Allocations: 0 bytes
|
||||
OLD (Array):
|
||||
Allocations: 496 bytes
|
||||
|
||||
3. Perfect Plasticity (elastic branch)
|
||||
NEW (Tensors.jl):
|
||||
Allocations: 0 bytes
|
||||
OLD (Dict):
|
||||
Allocations: 8828848 bytes
|
||||
|
||||
================================================================================
|
||||
PERFORMANCE BENCHMARKS
|
||||
================================================================================
|
||||
|
||||
Running detailed benchmarks (this may take a minute)...
|
||||
|
||||
1. LINEAR ELASTIC
|
||||
----------------------------------------
|
||||
NEW (Tensors.jl):
|
||||
BenchmarkTools.Trial: 10000 samples with 997 evaluations per sample.
|
||||
Range (min … max): 19.464 ns … 45.831 ns ┊ GC (min … max): 0.00% … 0.00%
|
||||
Time (median): 19.577 ns ┊ GC (median): 0.00%
|
||||
Time (mean ± σ): 19.670 ns ± 0.675 ns ┊ GC (mean ± σ): 0.00% ± 0.00%
|
||||
|
||||
▁█▄
|
||||
███▇▄▃▂▂▂▂▂▁▁▂▂▂▂▂▂▂▂▂▂▂▂▁▁▁▂▁▁▁▂▁▁▁▂▂▁▁▁▁▂▁▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂ ▂
|
||||
19.5 ns Histogram: frequency by time 22.8 ns <
|
||||
|
||||
Memory estimate: 0 bytes, allocs estimate: 0.
|
||||
|
||||
OLD (Voigt/Dict):
|
||||
BenchmarkTools.Trial: 10000 samples with 950 evaluations per sample.
|
||||
Range (min … max): 93.356 ns … 8.107 μs ┊ GC (min … max): 0.00% … 97.34%
|
||||
Time (median): 100.107 ns ┊ GC (median): 0.00%
|
||||
Time (mean ± σ): 139.733 ns ± 249.840 ns ┊ GC (mean ± σ): 25.28% ± 13.73%
|
||||
|
||||
█▂ ▁ ▁
|
||||
██▄▄██▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▆▇▇▄▃▄▅▄▄▅▅▅▅▅▃▆ █
|
||||
93.4 ns Histogram: log(frequency) by time 1.74 μs <
|
||||
|
||||
Memory estimate: 496 bytes, allocs estimate: 4.
|
||||
|
||||
SPEEDUP: 5.1×
|
||||
|
||||
2. NEO-HOOKEAN
|
||||
----------------------------------------
|
||||
NEW (Tensors.jl + Automatic Differentiation):
|
||||
BenchmarkTools.Trial: 10000 samples with 23 evaluations per sample.
|
||||
Range (min … max): 1.050 μs … 2.802 μs ┊ GC (min … max): 0.00% … 0.00%
|
||||
Time (median): 1.051 μs ┊ GC (median): 0.00%
|
||||
Time (mean ± σ): 1.055 μs ± 31.780 ns ┊ GC (mean ± σ): 0.00% ± 0.00%
|
||||
|
||||
█
|
||||
█▄▂▂▁▂▁▂▂▁▁▁▁▁▁▁▁▁▂▁▁▂▁▁▁▁▁▁▁▁▁▁▁▂▁▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂ ▂
|
||||
1.05 μs Histogram: frequency by time 1.19 μs <
|
||||
|
||||
Memory estimate: 0 bytes, allocs estimate: 0.
|
||||
|
||||
NEW (Tensors.jl + Manual Derivatives):
|
||||
BenchmarkTools.Trial: 10000 samples with 987 evaluations per sample.
|
||||
Range (min … max): 49.806 ns … 1.787 μs ┊ GC (min … max): 0.00% … 0.00%
|
||||
Time (median): 49.922 ns ┊ GC (median): 0.00%
|
||||
Time (mean ± σ): 50.262 ns ± 17.399 ns ┊ GC (mean ± σ): 0.00% ± 0.00%
|
||||
|
||||
▅██▄▁ ▂
|
||||
█████▆▄▁▄▄▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▁▃▄▄▅▅▆▆▆▆▆▇▇█▇▇▇▇▇██▇▇▆▇▇ █
|
||||
49.8 ns Histogram: log(frequency) by time 53.3 ns <
|
||||
|
||||
Memory estimate: 0 bytes, allocs estimate: 0.
|
||||
|
||||
OLD (Array):
|
||||
BenchmarkTools.Trial: 10000 samples with 955 evaluations per sample.
|
||||
Range (min … max): 91.182 ns … 9.723 μs ┊ GC (min … max): 0.00% … 97.56%
|
||||
Time (median): 99.922 ns ┊ GC (median): 0.00%
|
||||
Time (mean ± σ): 142.795 ns ± 307.090 ns ┊ GC (mean ± σ): 22.77% ± 11.92%
|
||||
|
||||
█▃ ▄▁ ▁
|
||||
██▆▄██▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▄▇█ █
|
||||
91.2 ns Histogram: log(frequency) by time 1.77 μs <
|
||||
|
||||
Memory estimate: 496 bytes, allocs estimate: 4.
|
||||
|
||||
SPEEDUP (AD): 0.1×
|
||||
SPEEDUP (Manual): 2.0×
|
||||
AD OVERHEAD: 21.1× (AD / Manual)
|
||||
|
||||
3. PERFECT PLASTICITY (elastic branch)
|
||||
----------------------------------------
|
||||
NEW (Tensors.jl):
|
||||
BenchmarkTools.Trial: 10000 samples with 976 evaluations per sample.
|
||||
Range (min … max): 69.677 ns … 151.814 ns ┊ GC (min … max): 0.00% … 0.00%
|
||||
Time (median): 70.389 ns ┊ GC (median): 0.00%
|
||||
Time (mean ± σ): 70.566 ns ± 1.326 ns ┊ GC (mean ± σ): 0.00% ± 0.00%
|
||||
|
||||
▂▆▇▆▇▆▇▇▇█▇▆▇▆▅▃ ▁▁▁▁▁▁▂▁▁▁▁▁ ▃
|
||||
▆███████████████████▇▇▇▆▇▆▆▄▃▁▁▁▁▁▆▅▅▆▇▆▇▇██████████████████ █
|
||||
69.7 ns Histogram: log(frequency) by time 73.9 ns <
|
||||
|
||||
Memory estimate: 0 bytes, allocs estimate: 0.
|
||||
|
||||
OLD (Dict):
|
||||
BenchmarkTools.Trial: 10000 samples with 10 evaluations per sample.
|
||||
Range (min … max): 1.371 μs … 998.345 μs ┊ GC (min … max): 0.00% … 99.45%
|
||||
Time (median): 1.480 μs ┊ GC (median): 0.00%
|
||||
Time (mean ± σ): 1.701 μs ± 9.970 μs ┊ GC (mean ± σ): 5.84% ± 0.99%
|
||||
|
||||
▅█▆▂
|
||||
▁▂▃▆████▆▅▄▃▃▃▃▂▂▂▂▂▁▁▂▁▁▁▁▁▂▂▂▂▂▂▂▂▂▂▃▃▃▃▂▂▂▂▂▂▂▂▂▂▁▁▁▁▁▁▁ ▂
|
||||
1.37 μs Histogram: frequency by time 2.16 μs <
|
||||
|
||||
Memory estimate: 1.98 KiB, allocs estimate: 53.
|
||||
|
||||
SPEEDUP: 21.0×
|
||||
|
||||
================================================================================
|
||||
SUMMARY
|
||||
================================================================================
|
||||
|
||||
ALLOCATIONS:
|
||||
LinearElastic: NEW = 0 bytes, OLD = 496 bytes
|
||||
NeoHookean (AD): NEW = 0 bytes, OLD = 496 bytes
|
||||
NeoHookean (Manual): NEW = 0 bytes
|
||||
PerfectPlasticity: NEW = 0 bytes, OLD = 8828848 bytes
|
||||
|
||||
MEDIAN TIMING:
|
||||
LinearElastic: NEW = 19.576730190571716 ns, OLD = 100.10684210526315 ns
|
||||
NeoHookean (AD): NEW = 1051.304347826087 ns, OLD = 99.92198952879582 ns
|
||||
NeoHookean (Manual): NEW = 49.92198581560284 ns
|
||||
PerfectPlasticity: NEW = 70.38934426229508 ns, OLD = 1479.55 ns
|
||||
|
||||
SPEEDUP (OLD / NEW):
|
||||
LinearElastic: 5.1×
|
||||
NeoHookean (AD): 0.1×
|
||||
NeoHookean (Manual): 2.0×
|
||||
PerfectPlasticity: 21.0×
|
||||
|
||||
AD OVERHEAD:
|
||||
NeoHookean: AD is 21.1× slower than manual derivatives
|
||||
|
||||
AVERAGE SPEEDUP: 9.4× (using manual Neo-Hookean)
|
||||
|
||||
VALIDATION OF CLAIMS:
|
||||
- Zero allocations for new approach: ✓ PASS
|
||||
- Manual derivatives outperform AD: ✓ PASS
|
||||
- Type stability with NoState return: Check @code_warntype output above
|
||||
|
||||
================================================================================
|
||||
Benchmark complete! Results saved to: material_models_benchmark_results.txt
|
||||
================================================================================
|
||||
@@ -1,835 +0,0 @@
|
||||
"""
|
||||
Matrix-Free Newton-Krylov GPU Benchmark
|
||||
|
||||
Demonstrates traditional Newton vs Matrix-Free Newton-Krylov with Anderson
|
||||
acceleration, running on GPU.
|
||||
|
||||
Run with:
|
||||
julia --project=. benchmarks/matrix_free_gpu_benchmark.jl
|
||||
"""
|
||||
|
||||
using CUDA
|
||||
using LinearAlgebra
|
||||
using IterativeSolvers
|
||||
using Printf
|
||||
|
||||
# Check GPU availability
|
||||
if !CUDA.functional()
|
||||
@warn "CUDA not available! Running CPU-only comparison."
|
||||
USE_GPU = false
|
||||
else
|
||||
println("GPU Device: $(CUDA.name(CUDA.device()))")
|
||||
println("GPU Memory: $(CUDA.total_memory() / 1e9) GB")
|
||||
println()
|
||||
USE_GPU = true
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Problem Setup: 3D Nonlinear Elasticity
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Residual for 3D nonlinear elasticity with cubic nonlinearity.
|
||||
|
||||
r(u) = K·u + β·(K·u).^3 - f
|
||||
|
||||
where K is stiffness matrix, β is nonlinearity parameter.
|
||||
"""
|
||||
struct NonlinearProblem{T,MatT,VecT}
|
||||
K::MatT # Stiffness matrix (sparse or LinearMap)
|
||||
f::VecT # Force vector
|
||||
β::T # Nonlinearity parameter
|
||||
n::Int # DOF count
|
||||
end
|
||||
|
||||
"""
|
||||
Compute residual: r(u) = K·u + β·(K·u).^3 - f
|
||||
"""
|
||||
function compute_residual!(r::AbstractVector, prob::NonlinearProblem, u::AbstractVector)
|
||||
# Linear part
|
||||
mul!(r, prob.K, u) # r = K·u
|
||||
|
||||
# Nonlinear part: r += β·(K·u).^3
|
||||
if prob.β != 0
|
||||
# Reuse r (which contains K·u)
|
||||
@. r = r + prob.β * r^3
|
||||
end
|
||||
|
||||
# Apply forcing
|
||||
@. r = r - prob.f
|
||||
|
||||
return r
|
||||
end
|
||||
|
||||
"""
|
||||
Jacobian-vector product: J·v ≈ [r(u+ε·v) - r(u)] / ε (finite difference)
|
||||
"""
|
||||
function jacobian_vector_product!(
|
||||
Jv::AbstractVector,
|
||||
prob::NonlinearProblem,
|
||||
u::AbstractVector,
|
||||
v::AbstractVector,
|
||||
r_u::AbstractVector, # Pre-computed r(u)
|
||||
temp::AbstractVector # Workspace
|
||||
)
|
||||
ε = 1e-7
|
||||
|
||||
# temp = u + ε·v
|
||||
@. temp = u + ε * v
|
||||
|
||||
# Jv = r(u + ε·v)
|
||||
compute_residual!(Jv, prob, temp)
|
||||
|
||||
# Jv = [r(u + ε·v) - r(u)] / ε
|
||||
@. Jv = (Jv - r_u) / ε
|
||||
|
||||
return Jv
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Traditional Newton Solver
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Traditional Newton with full Jacobian assembly.
|
||||
|
||||
u_{k+1} = u_k - J(u_k)^{-1} · r(u_k)
|
||||
|
||||
Expensive: Assembles full Jacobian matrix at each iteration.
|
||||
"""
|
||||
function newton_traditional!(
|
||||
u::AbstractVector{T},
|
||||
prob::NonlinearProblem{T},
|
||||
r::AbstractVector{T},
|
||||
du::AbstractVector{T};
|
||||
tol=1e-8,
|
||||
max_iter=20,
|
||||
verbose=true
|
||||
) where T
|
||||
|
||||
n = length(u)
|
||||
|
||||
# Build Jacobian matrix (expensive!)
|
||||
# J ≈ K + 3β·diag((K·u).^2)·K
|
||||
Ku = prob.K * u
|
||||
J = copy(prob.K)
|
||||
|
||||
for iter in 1:max_iter
|
||||
# Compute residual
|
||||
compute_residual!(r, prob, u)
|
||||
|
||||
norm_r = norm(r)
|
||||
|
||||
if verbose
|
||||
@printf(" Iter %2d: ||r|| = %.6e\n", iter, norm_r)
|
||||
end
|
||||
|
||||
if norm_r < tol
|
||||
if verbose
|
||||
println(" ✅ Converged!")
|
||||
end
|
||||
return iter
|
||||
end
|
||||
|
||||
# Update Jacobian (expensive!)
|
||||
Ku .= prob.K * u
|
||||
for i in 1:n
|
||||
J[i, i] = prob.K[i, i] + 3 * prob.β * Ku[i]^2 * prob.K[i, i]
|
||||
end
|
||||
|
||||
# Solve linear system (expensive!)
|
||||
du .= -(J \ r)
|
||||
|
||||
# Update
|
||||
u .+= du
|
||||
end
|
||||
|
||||
if verbose
|
||||
println(" ⚠️ Did not converge in $max_iter iterations")
|
||||
end
|
||||
|
||||
return max_iter
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Helper: Matrix-free operator wrapper for GMRES
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Wrapper to make a function look like a matrix for GMRES.
|
||||
"""
|
||||
struct MatrixFreeOperator{F}
|
||||
matvec!::F
|
||||
n::Int
|
||||
end
|
||||
|
||||
Base.size(A::MatrixFreeOperator) = (A.n, A.n)
|
||||
Base.size(A::MatrixFreeOperator, d::Int) = d <= 2 ? A.n : 1
|
||||
Base.eltype(::MatrixFreeOperator{F}) where F = Float64
|
||||
|
||||
function LinearAlgebra.mul!(y, A::MatrixFreeOperator, x)
|
||||
A.matvec!(y, x)
|
||||
return y
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Matrix-Free Newton-Krylov
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Matrix-Free Newton-Krylov with GMRES.
|
||||
|
||||
J·v ≈ [r(u+ε·v) - r(u)] / ε (no matrix!)
|
||||
du = gmres(Jv_op, -r)
|
||||
u_{k+1} = u_k + du
|
||||
|
||||
Cheap: Only residual evaluations, no Jacobian assembly.
|
||||
"""
|
||||
function newton_matrix_free!(
|
||||
u::AbstractVector{T},
|
||||
prob::NonlinearProblem{T},
|
||||
r::AbstractVector{T},
|
||||
du::AbstractVector{T},
|
||||
temp::AbstractVector{T},
|
||||
Jv::AbstractVector{T};
|
||||
tol=1e-8,
|
||||
max_iter=20,
|
||||
gmres_tol=1e-6,
|
||||
verbose=true
|
||||
) where T
|
||||
|
||||
for iter in 1:max_iter
|
||||
# Compute residual
|
||||
compute_residual!(r, prob, u)
|
||||
|
||||
norm_r = norm(r)
|
||||
|
||||
if verbose
|
||||
@printf(" Iter %2d: ||r|| = %.6e", iter, norm_r)
|
||||
end
|
||||
|
||||
if norm_r < tol
|
||||
if verbose
|
||||
println(" ✅ Converged!")
|
||||
end
|
||||
return iter
|
||||
end
|
||||
|
||||
# Matrix-free operator: J·v
|
||||
function Jv_matvec!(out, v)
|
||||
jacobian_vector_product!(out, prob, u, v, r, temp)
|
||||
return out
|
||||
end
|
||||
Jv_op = MatrixFreeOperator(Jv_matvec!, length(u))
|
||||
|
||||
# Solve J·du = -r using GMRES (matrix-free!)
|
||||
du .= 0
|
||||
gmres!(du, Jv_op, -r;
|
||||
abstol=gmres_tol, reltol=0, maxiter=50, verbose=false)
|
||||
|
||||
gmres_iters = 50 # Would need to extract from gmres! return
|
||||
|
||||
if verbose
|
||||
@printf(" [GMRES: ~%d iters]\n", gmres_iters)
|
||||
end
|
||||
|
||||
# Update
|
||||
u .+= du
|
||||
end
|
||||
|
||||
if verbose
|
||||
println(" ⚠️ Did not converge in $max_iter iterations")
|
||||
end
|
||||
|
||||
return max_iter
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Anderson-Accelerated Newton-Krylov
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Anderson acceleration for Newton-Krylov.
|
||||
|
||||
Combines m previous iterates via least-squares:
|
||||
u_new = ∑ αᵢ·uᵢ where argmin ||∑ αᵢ·rᵢ||² s.t. ∑ αᵢ = 1
|
||||
|
||||
Transforms linear convergence → superlinear convergence.
|
||||
"""
|
||||
function anderson_newton_matrix_free!(
|
||||
u::AbstractVector{T},
|
||||
prob::NonlinearProblem{T},
|
||||
r::AbstractVector{T},
|
||||
du::AbstractVector{T},
|
||||
temp::AbstractVector{T},
|
||||
Jv::AbstractVector{T};
|
||||
m=5, # Anderson history
|
||||
tol=1e-8,
|
||||
max_iter=20,
|
||||
gmres_tol=1e-6,
|
||||
verbose=true
|
||||
) where T
|
||||
|
||||
n = length(u)
|
||||
|
||||
# Anderson history
|
||||
U_history = [zeros(T, n) for _ in 1:m]
|
||||
R_history = [zeros(T, n) for _ in 1:m]
|
||||
history_count = 0
|
||||
|
||||
for iter in 1:max_iter
|
||||
# Compute residual
|
||||
compute_residual!(r, prob, u)
|
||||
|
||||
norm_r = norm(r)
|
||||
|
||||
if verbose
|
||||
@printf(" Iter %2d: ||r|| = %.6e", iter, norm_r)
|
||||
end
|
||||
|
||||
if norm_r < tol
|
||||
if verbose
|
||||
println(" ✅ Converged!")
|
||||
end
|
||||
return iter
|
||||
end
|
||||
|
||||
# Matrix-free operator
|
||||
function Jv_matvec!(out, v)
|
||||
jacobian_vector_product!(out, prob, u, v, r, temp)
|
||||
return out
|
||||
end
|
||||
Jv_op = MatrixFreeOperator(Jv_matvec!, n)
|
||||
|
||||
# Solve J·du = -r using GMRES
|
||||
du .= 0
|
||||
gmres!(du, Jv_op, -r;
|
||||
abstol=gmres_tol, reltol=0, maxiter=50, verbose=false)
|
||||
|
||||
# Store in history (circular buffer)
|
||||
idx = mod1(history_count + 1, m)
|
||||
U_history[idx] .= u
|
||||
R_history[idx] .= r
|
||||
history_count = min(history_count + 1, m)
|
||||
|
||||
if verbose
|
||||
@printf(" [GMRES: ~50 iters, history: %d]", history_count)
|
||||
end
|
||||
|
||||
# Anderson acceleration (if enough history)
|
||||
if history_count >= 2
|
||||
# Build residual difference matrix
|
||||
k = history_count
|
||||
R_diff = zeros(T, n, k)
|
||||
for i in 1:k
|
||||
R_diff[:, i] .= R_history[i] .- r
|
||||
end
|
||||
|
||||
# Check condition number before QR
|
||||
# If matrix is ill-conditioned, skip Anderson this iteration
|
||||
R_norm = norm(R_diff)
|
||||
if R_norm < 1e-10
|
||||
# Matrix too small, use standard update
|
||||
u .+= du
|
||||
if verbose
|
||||
println(" [Anderson: skipped (residuals too small)]")
|
||||
end
|
||||
else
|
||||
# Least-squares: min ||R_diff·α||² s.t. sum(α) = 1
|
||||
# Use QR factorization with regularization
|
||||
try
|
||||
Q, Rt = qr(R_diff)
|
||||
|
||||
# Add small regularization to diagonal if needed
|
||||
Rt_diag = diag(Rt)
|
||||
if any(abs.(Rt_diag) .< 1e-12)
|
||||
# Add Tikhonov regularization
|
||||
λ = 1e-8
|
||||
Rt_reg = Rt + λ * I
|
||||
α = Rt_reg \ (Q' * r)
|
||||
else
|
||||
α = Rt \ (Q' * r)
|
||||
end
|
||||
|
||||
α ./= sum(α) # Normalize
|
||||
|
||||
# Combine previous iterates
|
||||
u_combined = zeros(T, n)
|
||||
for i in 1:k
|
||||
u_combined .+= α[i] .* U_history[i]
|
||||
end
|
||||
|
||||
# Update with combination
|
||||
u .= u_combined .+ du
|
||||
|
||||
if verbose
|
||||
println(" [Anderson: α=$(round.(α, digits=3))]")
|
||||
end
|
||||
catch e
|
||||
# If QR fails, fall back to standard Newton
|
||||
u .+= du
|
||||
if verbose
|
||||
println(" [Anderson: failed ($e), using standard update]")
|
||||
end
|
||||
end
|
||||
end
|
||||
else
|
||||
# Standard Newton update
|
||||
u .+= du
|
||||
|
||||
if verbose
|
||||
println()
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
if verbose
|
||||
println(" ⚠️ Did not converge in $max_iter iterations")
|
||||
end
|
||||
|
||||
return max_iter
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# GPU Implementations
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
GPU version of residual computation.
|
||||
"""
|
||||
function compute_residual_gpu!(
|
||||
r::CuVector{T},
|
||||
K::CuMatrix{T},
|
||||
u::CuVector{T},
|
||||
f::CuVector{T},
|
||||
β::T
|
||||
) where T
|
||||
# r = K·u
|
||||
mul!(r, K, u)
|
||||
|
||||
# r = r + β·r³ - f
|
||||
if β != 0
|
||||
r .= r .+ β .* r .^ 3 .- f
|
||||
else
|
||||
r .= r .- f
|
||||
end
|
||||
|
||||
return r
|
||||
end
|
||||
|
||||
"""
|
||||
GPU version of Matrix-Free Newton-Krylov.
|
||||
"""
|
||||
function newton_matrix_free_gpu!(
|
||||
u::CuVector{T},
|
||||
K::CuMatrix{T},
|
||||
f::CuVector{T},
|
||||
β::T;
|
||||
tol=1e-8,
|
||||
max_iter=20,
|
||||
gmres_tol=1e-6,
|
||||
verbose=true
|
||||
) where T
|
||||
|
||||
n = length(u)
|
||||
r = CUDA.zeros(T, n)
|
||||
du = CUDA.zeros(T, n)
|
||||
temp = CUDA.zeros(T, n)
|
||||
Jv = CUDA.zeros(T, n)
|
||||
|
||||
for iter in 1:max_iter
|
||||
# Compute residual on GPU
|
||||
compute_residual_gpu!(r, K, u, f, β)
|
||||
|
||||
norm_r = norm(Array(r)) # Transfer to CPU for norm
|
||||
|
||||
if verbose
|
||||
@printf(" Iter %2d: ||r|| = %.6e\n", iter, norm_r)
|
||||
end
|
||||
|
||||
if norm_r < tol
|
||||
if verbose
|
||||
println(" ✅ Converged!")
|
||||
end
|
||||
return iter
|
||||
end
|
||||
|
||||
# Jacobian-vector product (on GPU)
|
||||
ε = T(1e-7)
|
||||
function Jv_matvec_gpu!(out_cpu, v_cpu)
|
||||
v = CuArray(v_cpu)
|
||||
# temp = u + ε·v
|
||||
temp .= u .+ ε .* v
|
||||
# Jv = r(u + ε·v)
|
||||
compute_residual_gpu!(Jv, K, temp, f, β)
|
||||
# Jv = [r(u + ε·v) - r(u)] / ε
|
||||
Jv .= (Jv .- r) ./ ε
|
||||
out_cpu .= Array(Jv)
|
||||
return out_cpu
|
||||
end
|
||||
Jv_op_gpu = MatrixFreeOperator(Jv_matvec_gpu!, n)
|
||||
|
||||
# Solve on CPU (GMRES doesn't have GPU version in IterativeSolvers.jl)
|
||||
r_cpu = Array(r)
|
||||
du_cpu = zeros(T, n)
|
||||
gmres!(du_cpu, Jv_op_gpu, -r_cpu;
|
||||
abstol=gmres_tol, reltol=0, maxiter=50, verbose=false)
|
||||
|
||||
# Update on GPU
|
||||
du .= CuArray(du_cpu)
|
||||
u .+= du
|
||||
end
|
||||
|
||||
if verbose
|
||||
println(" ⚠️ Did not converge in $max_iter iterations")
|
||||
end
|
||||
|
||||
return max_iter
|
||||
end
|
||||
|
||||
# GPU Anderson-Accelerated Newton-Krylov
|
||||
function anderson_newton_matrix_free_gpu!(
|
||||
u::CuVector{T},
|
||||
K::CuMatrix{T},
|
||||
f::CuVector{T},
|
||||
β::T;
|
||||
tol=1e-8,
|
||||
max_iter=20,
|
||||
gmres_tol=1e-6,
|
||||
history_size=5,
|
||||
verbose=true
|
||||
) where T
|
||||
|
||||
n = length(u)
|
||||
r = CUDA.zeros(T, n)
|
||||
du = CUDA.zeros(T, n)
|
||||
temp = CUDA.zeros(T, n)
|
||||
Jv = CUDA.zeros(T, n)
|
||||
|
||||
# Anderson acceleration storage (CPU)
|
||||
R_history = Vector{Vector{T}}()
|
||||
U_history = Vector{Vector{T}}()
|
||||
history_count = 0
|
||||
|
||||
for iter in 1:max_iter
|
||||
# Compute residual on GPU
|
||||
compute_residual_gpu!(r, K, u, f, β)
|
||||
|
||||
norm_r = norm(Array(r)) # Transfer to CPU for norm
|
||||
|
||||
if verbose
|
||||
@printf(" Iter %2d: ||r|| = %.6e", iter, norm_r)
|
||||
end
|
||||
|
||||
if norm_r < tol
|
||||
if verbose
|
||||
println("\n ✅ Converged!")
|
||||
end
|
||||
return iter
|
||||
end
|
||||
|
||||
# Jacobian-vector product (on GPU)
|
||||
ε = T(1e-7)
|
||||
function Jv_matvec_gpu!(out_cpu, v_cpu)
|
||||
v = CuArray(v_cpu)
|
||||
# temp = u + ε·v
|
||||
temp .= u .+ ε .* v
|
||||
# Jv = r(u + ε·v)
|
||||
compute_residual_gpu!(Jv, K, temp, f, β)
|
||||
# Jv = [r(u + ε·v) - r(u)] / ε
|
||||
Jv .= (Jv .- r) ./ ε
|
||||
out_cpu .= Array(Jv)
|
||||
return out_cpu
|
||||
end
|
||||
Jv_op_gpu = MatrixFreeOperator(Jv_matvec_gpu!, n)
|
||||
|
||||
# Solve on CPU (GMRES doesn't have GPU version)
|
||||
r_cpu = Array(r)
|
||||
u_cpu = Array(u)
|
||||
du_cpu = zeros(T, n)
|
||||
gmres!(du_cpu, Jv_op_gpu, -r_cpu;
|
||||
abstol=gmres_tol, reltol=0, maxiter=50, verbose=false)
|
||||
|
||||
# Anderson acceleration (on CPU)
|
||||
if history_count >= 2
|
||||
# Build residual difference matrix
|
||||
k = history_count
|
||||
R_diff = zeros(T, n, k)
|
||||
for i in 1:k
|
||||
R_diff[:, i] .= R_history[i] .- r_cpu
|
||||
end
|
||||
|
||||
# Check condition number before QR
|
||||
R_norm = norm(R_diff)
|
||||
if R_norm < 1e-10
|
||||
# Matrix too small, use standard update
|
||||
u .+= CuArray(du_cpu)
|
||||
if verbose
|
||||
println(" [Anderson: skipped (residuals too small)]")
|
||||
end
|
||||
else
|
||||
# Least-squares with regularization
|
||||
try
|
||||
Q, Rt = qr(R_diff)
|
||||
|
||||
# Add small regularization to diagonal if needed
|
||||
Rt_diag = diag(Rt)
|
||||
if any(abs.(Rt_diag) .< 1e-12)
|
||||
# Add Tikhonov regularization
|
||||
λ = 1e-8
|
||||
Rt_reg = Rt + λ * I
|
||||
α = Rt_reg \ (Q' * r_cpu)
|
||||
else
|
||||
α = Rt \ (Q' * r_cpu)
|
||||
end
|
||||
|
||||
α ./= sum(α) # Normalize
|
||||
|
||||
# Combine previous iterates
|
||||
u_combined = zeros(T, n)
|
||||
for i in 1:k
|
||||
u_combined .+= α[i] .* U_history[i]
|
||||
end
|
||||
|
||||
# Update with combination (transfer to GPU)
|
||||
u .= CuArray(u_combined .+ du_cpu)
|
||||
|
||||
if verbose
|
||||
println(" [Anderson: α=$(round.(α, digits=3))]")
|
||||
end
|
||||
catch e
|
||||
# If QR fails, fall back to standard Newton
|
||||
u .+= CuArray(du_cpu)
|
||||
if verbose
|
||||
println(" [Anderson: failed ($e), using standard update]")
|
||||
end
|
||||
end
|
||||
end
|
||||
else
|
||||
# Standard Newton update (transfer du to GPU)
|
||||
u .+= CuArray(du_cpu)
|
||||
if verbose
|
||||
println()
|
||||
end
|
||||
end
|
||||
|
||||
# Store history (on CPU to avoid GPU memory overhead)
|
||||
push!(R_history, copy(r_cpu))
|
||||
push!(U_history, copy(u_cpu))
|
||||
history_count += 1
|
||||
|
||||
# Maintain history size
|
||||
if history_count > history_size
|
||||
popfirst!(R_history)
|
||||
popfirst!(U_history)
|
||||
history_count = history_size
|
||||
end
|
||||
end
|
||||
|
||||
if verbose
|
||||
println(" ⚠️ Did not converge in $max_iter iterations")
|
||||
end
|
||||
|
||||
return max_iter
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark Runners
|
||||
# ============================================================================
|
||||
|
||||
function benchmark_cpu(n::Int)
|
||||
println("="^70)
|
||||
println("CPU Benchmark: $n DOFs")
|
||||
println("="^70)
|
||||
|
||||
# Setup problem
|
||||
T = Float64
|
||||
K = Matrix(Tridiagonal(
|
||||
-ones(T, n - 1),
|
||||
2ones(T, n),
|
||||
-ones(T, n - 1)
|
||||
))
|
||||
f = ones(T, n) * 0.1
|
||||
β = T(1e-3) # Nonlinearity
|
||||
|
||||
prob = NonlinearProblem(K, f, β, n)
|
||||
|
||||
# Initial guess
|
||||
u0 = zeros(T, n)
|
||||
|
||||
# Allocate workspace
|
||||
r = zeros(T, n)
|
||||
du = zeros(T, n)
|
||||
temp = zeros(T, n)
|
||||
Jv = zeros(T, n)
|
||||
|
||||
# Benchmark Traditional Newton
|
||||
println("\n📊 Traditional Newton (Full Jacobian):")
|
||||
u_trad = copy(u0)
|
||||
t_trad = @elapsed iters_trad = newton_traditional!(u_trad, prob, r, du; verbose=false)
|
||||
println(" Time: $(round(t_trad * 1000, digits=2)) ms")
|
||||
println(" Iterations: $iters_trad")
|
||||
println(" Time/iter: $(round(t_trad / iters_trad * 1000, digits=2)) ms")
|
||||
|
||||
# Benchmark Matrix-Free
|
||||
println("\n📊 Matrix-Free Newton-Krylov:")
|
||||
u_mf = copy(u0)
|
||||
t_mf = @elapsed iters_mf = newton_matrix_free!(
|
||||
u_mf, prob, r, du, temp, Jv; verbose=false
|
||||
)
|
||||
println(" Time: $(round(t_mf * 1000, digits=2)) ms")
|
||||
println(" Iterations: $iters_mf")
|
||||
println(" Time/iter: $(round(t_mf / iters_mf * 1000, digits=2)) ms")
|
||||
|
||||
# Benchmark Anderson-Accelerated
|
||||
println("\n📊 Anderson-Accelerated Matrix-Free:")
|
||||
u_anderson = copy(u0)
|
||||
t_anderson = @elapsed iters_anderson = anderson_newton_matrix_free!(
|
||||
u_anderson, prob, r, du, temp, Jv; m=5, verbose=false
|
||||
)
|
||||
println(" Time: $(round(t_anderson * 1000, digits=2)) ms")
|
||||
println(" Iterations: $iters_anderson")
|
||||
println(" Time/iter: $(round(t_anderson / iters_anderson * 1000, digits=2)) ms")
|
||||
|
||||
# Speedups
|
||||
println("\n✅ CPU Speedups:")
|
||||
println(" Matrix-Free vs Traditional: $(round(t_trad / t_mf, digits=2))×")
|
||||
println(" Anderson vs Traditional: $(round(t_trad / t_anderson, digits=2))×")
|
||||
println(" Anderson vs Matrix-Free: $(round(t_mf / t_anderson, digits=2))×")
|
||||
|
||||
println()
|
||||
end
|
||||
|
||||
function benchmark_gpu(n::Int)
|
||||
if !USE_GPU
|
||||
println("⚠️ GPU not available, skipping GPU benchmark\n")
|
||||
return
|
||||
end
|
||||
|
||||
println("="^70)
|
||||
println("GPU Benchmark: $n DOFs")
|
||||
println("="^70)
|
||||
|
||||
# Setup problem
|
||||
T = Float64
|
||||
K_cpu = Matrix(Tridiagonal(
|
||||
-ones(T, n - 1),
|
||||
2ones(T, n),
|
||||
-ones(T, n - 1)
|
||||
))
|
||||
f_cpu = ones(T, n) * 0.1
|
||||
β = T(1e-3)
|
||||
|
||||
# Transfer to GPU
|
||||
K_gpu = CuArray(K_cpu)
|
||||
f_gpu = CuArray(f_cpu)
|
||||
u0_gpu = CUDA.zeros(T, n)
|
||||
|
||||
# Benchmark Matrix-Free on GPU
|
||||
println("\n📊 Matrix-Free Newton-Krylov (GPU):")
|
||||
u_gpu = copy(u0_gpu)
|
||||
|
||||
# Warmup
|
||||
newton_matrix_free_gpu!(u_gpu, K_gpu, f_gpu, β; max_iter=2, verbose=false)
|
||||
|
||||
# Benchmark
|
||||
CUDA.synchronize()
|
||||
t_gpu = CUDA.@elapsed begin
|
||||
iters_gpu = newton_matrix_free_gpu!(u_gpu, K_gpu, f_gpu, β; verbose=false)
|
||||
CUDA.synchronize()
|
||||
end
|
||||
|
||||
println(" Time: $(round(t_gpu * 1000, digits=2)) ms")
|
||||
println(" Iterations: $iters_gpu")
|
||||
println(" Time/iter: $(round(t_gpu / iters_gpu * 1000, digits=2)) ms")
|
||||
|
||||
# Compare with CPU
|
||||
prob_cpu = NonlinearProblem(K_cpu, f_cpu, β, n)
|
||||
u_cpu = zeros(T, n)
|
||||
r = zeros(T, n)
|
||||
du = zeros(T, n)
|
||||
temp = zeros(T, n)
|
||||
Jv = zeros(T, n)
|
||||
|
||||
t_cpu = @elapsed iters_cpu = newton_matrix_free!(
|
||||
u_cpu, prob_cpu, r, du, temp, Jv; verbose=false
|
||||
)
|
||||
|
||||
println("\n✅ GPU vs CPU Speedup: $(round(t_cpu / t_gpu, digits=2))×")
|
||||
println(" CPU: $(round(t_cpu * 1000, digits=2)) ms")
|
||||
println(" GPU: $(round(t_gpu * 1000, digits=2)) ms")
|
||||
|
||||
# Benchmark Anderson-Accelerated on GPU
|
||||
println("\n📊 Anderson-Accelerated Newton-Krylov (GPU):")
|
||||
u_gpu_anderson = copy(u0_gpu)
|
||||
|
||||
# Warmup
|
||||
anderson_newton_matrix_free_gpu!(u_gpu_anderson, K_gpu, f_gpu, β; max_iter=2, verbose=false)
|
||||
|
||||
# Benchmark
|
||||
CUDA.synchronize()
|
||||
t_gpu_anderson = CUDA.@elapsed begin
|
||||
iters_gpu_anderson = anderson_newton_matrix_free_gpu!(u_gpu_anderson, K_gpu, f_gpu, β; verbose=false)
|
||||
CUDA.synchronize()
|
||||
end
|
||||
|
||||
println(" Time: $(round(t_gpu_anderson * 1000, digits=2)) ms")
|
||||
println(" Iterations: $iters_gpu_anderson")
|
||||
println(" Time/iter: $(round(t_gpu_anderson / iters_gpu_anderson * 1000, digits=2)) ms")
|
||||
|
||||
# Compare with CPU Anderson
|
||||
u_cpu_anderson = zeros(T, n)
|
||||
t_cpu_anderson = @elapsed iters_cpu_anderson = anderson_newton_matrix_free!(
|
||||
u_cpu_anderson, prob_cpu, r, du, temp, Jv; verbose=false
|
||||
)
|
||||
|
||||
println("\n✅ GPU vs CPU Speedup (Anderson): $(round(t_cpu_anderson / t_gpu_anderson, digits=2))×")
|
||||
println(" CPU: $(round(t_cpu_anderson * 1000, digits=2)) ms")
|
||||
println(" GPU: $(round(t_gpu_anderson * 1000, digits=2)) ms")
|
||||
|
||||
# Overall comparison
|
||||
println("\n📊 Summary:")
|
||||
println(" Matrix-Free GPU speedup: $(round(t_cpu / t_gpu, digits=2))×")
|
||||
println(" Anderson GPU speedup: $(round(t_cpu_anderson / t_gpu_anderson, digits=2))×")
|
||||
|
||||
println()
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Main
|
||||
# ============================================================================
|
||||
|
||||
function main()
|
||||
println("\n" * "="^70)
|
||||
println("Matrix-Free Newton-Krylov GPU Benchmark")
|
||||
println("="^70)
|
||||
println()
|
||||
|
||||
# Test sizes (reasonable for demonstration)
|
||||
sizes = [1000, 5_000, 10_000]
|
||||
|
||||
for n in sizes
|
||||
# CPU comparison
|
||||
benchmark_cpu(n)
|
||||
|
||||
# GPU benchmark
|
||||
if USE_GPU
|
||||
benchmark_gpu(n)
|
||||
end
|
||||
end
|
||||
|
||||
println("="^70)
|
||||
println("Benchmark Complete!")
|
||||
println("="^70)
|
||||
println()
|
||||
println("Key Findings:")
|
||||
println(" - Matrix-Free eliminates Jacobian assembly cost")
|
||||
println(" - Anderson acceleration reduces Newton iterations")
|
||||
println(" - GPU provides additional speedup for large problems")
|
||||
println(" - Total speedup: 5-10× depending on problem size")
|
||||
println()
|
||||
end
|
||||
|
||||
if abspath(PROGRAM_FILE) == @__FILE__
|
||||
main()
|
||||
end
|
||||
@@ -1,571 +0,0 @@
|
||||
#!/usr/bin/env julia
|
||||
#
|
||||
# Multi-GPU Nodal Assembly Benchmark with MPI + CUDA
|
||||
#
|
||||
# Usage:
|
||||
# mpirun -np 2 julia --project=. benchmarks/multigpu_mpi_benchmark.jl
|
||||
# mpirun -np 4 julia --project=. benchmarks/multigpu_mpi_benchmark.jl
|
||||
#
|
||||
# Each MPI rank gets one GPU
|
||||
|
||||
using MPI
|
||||
using CUDA
|
||||
using LinearAlgebra
|
||||
using Printf
|
||||
|
||||
MPI.Init()
|
||||
|
||||
const comm = MPI.COMM_WORLD
|
||||
const rank = MPI.Comm_rank(comm)
|
||||
const nranks = MPI.Comm_size(comm)
|
||||
|
||||
# Set GPU device based on rank
|
||||
if CUDA.functional()
|
||||
CUDA.device!(rank % CUDA.ndevices())
|
||||
if rank == 0
|
||||
println("="^70)
|
||||
println("Multi-GPU Nodal Assembly Benchmark (MPI + CUDA)")
|
||||
println("="^70)
|
||||
println("MPI ranks: $nranks")
|
||||
println("CUDA devices: $(CUDA.ndevices())")
|
||||
println("CUDA functional: $(CUDA.functional())")
|
||||
println("="^70)
|
||||
println()
|
||||
end
|
||||
else
|
||||
if rank == 0
|
||||
println("ERROR: CUDA not functional!")
|
||||
println("Install CUDA.jl: using Pkg; Pkg.add(\"CUDA\")")
|
||||
end
|
||||
MPI.Finalize()
|
||||
exit(1)
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Data Structures
|
||||
# ============================================================================
|
||||
|
||||
struct Node
|
||||
id::Int32
|
||||
x::Float32
|
||||
y::Float32
|
||||
z::Float32
|
||||
end
|
||||
|
||||
struct Element
|
||||
id::Int32
|
||||
connectivity::NTuple{8,Int32} # Hex8
|
||||
end
|
||||
|
||||
struct Partition
|
||||
rank::Int
|
||||
owned_nodes::UnitRange{Int}
|
||||
ghost_nodes::Vector{Int}
|
||||
local_elements::Vector{Int}
|
||||
node_to_elements::Vector{Vector{Int}}
|
||||
interface_neighbors::Vector{Int} # Neighbor ranks
|
||||
interface_send::Dict{Int,Vector{Int}} # rank → local DOF indices to send
|
||||
interface_recv::Dict{Int,Vector{Int}} # rank → local DOF indices to receive
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Mesh Generation
|
||||
# ============================================================================
|
||||
|
||||
function create_hex_mesh(nx, ny, nz)
|
||||
"""Create structured hexahedral mesh"""
|
||||
n_nodes = nx * ny * nz
|
||||
n_elements = (nx - 1) * (ny - 1) * (nz - 1)
|
||||
|
||||
nodes = Node[]
|
||||
for k in 1:nz, j in 1:ny, i in 1:nx
|
||||
node_id = Int32((k - 1) * nx * ny + (j - 1) * nx + i)
|
||||
push!(nodes, Node(node_id, Float32(i), Float32(j), Float32(k)))
|
||||
end
|
||||
|
||||
elements = Element[]
|
||||
for k in 1:(nz-1), j in 1:(ny-1), i in 1:(nx-1)
|
||||
n1 = Int32((k - 1) * nx * ny + (j - 1) * nx + i)
|
||||
n2 = n1 + 1
|
||||
n3 = n2 + nx
|
||||
n4 = n1 + nx
|
||||
n5 = n1 + nx * ny
|
||||
n6 = n2 + nx * ny
|
||||
n7 = n3 + nx * ny
|
||||
n8 = n4 + nx * ny
|
||||
|
||||
elem_id = Int32(length(elements) + 1)
|
||||
push!(elements, Element(elem_id, (n1, n2, n3, n4, n5, n6, n7, n8)))
|
||||
end
|
||||
|
||||
return nodes, elements
|
||||
end
|
||||
|
||||
function build_node_to_elements(nodes, elements)
|
||||
node_to_elems = [Int[] for _ in 1:length(nodes)]
|
||||
|
||||
for (elem_id, element) in enumerate(elements)
|
||||
for node_id in element.connectivity
|
||||
push!(node_to_elems[node_id], elem_id)
|
||||
end
|
||||
end
|
||||
|
||||
return node_to_elems
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Partitioning
|
||||
# ============================================================================
|
||||
|
||||
function partition_mesh_for_rank(nodes, elements, my_rank, n_ranks)
|
||||
"""Create partition for this MPI rank"""
|
||||
n_nodes = length(nodes)
|
||||
nodes_per_rank = ceil(Int, n_nodes / n_ranks)
|
||||
|
||||
# Owned nodes
|
||||
start_node = my_rank * nodes_per_rank + 1
|
||||
end_node = min((my_rank + 1) * nodes_per_rank, n_nodes)
|
||||
owned_nodes = start_node:end_node
|
||||
|
||||
node_to_elems = build_node_to_elements(nodes, elements)
|
||||
|
||||
# Find local elements (touching owned nodes)
|
||||
local_elements = Int[]
|
||||
ghost_nodes = Set{Int}()
|
||||
|
||||
for (elem_id, element) in enumerate(elements)
|
||||
if any(Int(nid) in owned_nodes for nid in element.connectivity)
|
||||
push!(local_elements, elem_id)
|
||||
|
||||
for nid in element.connectivity
|
||||
if !(Int(nid) in owned_nodes)
|
||||
push!(ghost_nodes, Int(nid))
|
||||
end
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
# Build local node_to_elements
|
||||
local_node_to_elems = [
|
||||
filter(eid -> eid in local_elements, node_to_elems[nid])
|
||||
for nid in owned_nodes
|
||||
]
|
||||
|
||||
# Find interface nodes with each neighbor
|
||||
interface_send = Dict{Int,Vector{Int}}()
|
||||
interface_recv = Dict{Int,Vector{Int}}()
|
||||
|
||||
for neighbor_rank in 0:(n_ranks-1)
|
||||
if neighbor_rank == my_rank
|
||||
continue
|
||||
end
|
||||
|
||||
neighbor_start = neighbor_rank * nodes_per_rank + 1
|
||||
neighbor_end = min((neighbor_rank + 1) * nodes_per_rank, n_nodes)
|
||||
neighbor_owned = neighbor_start:neighbor_end
|
||||
|
||||
# Nodes I own that neighbor needs (I send)
|
||||
send_nodes = Int[]
|
||||
for elem_id in local_elements
|
||||
element = elements[elem_id]
|
||||
has_neighbor = any(Int(nid) in neighbor_owned for nid in element.connectivity)
|
||||
if has_neighbor
|
||||
for nid in element.connectivity
|
||||
if Int(nid) in owned_nodes && !(Int(nid) in send_nodes)
|
||||
push!(send_nodes, Int(nid))
|
||||
end
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
# Nodes neighbor owns that I need (I receive)
|
||||
recv_nodes = Int[]
|
||||
for nid in ghost_nodes
|
||||
if Int(nid) in neighbor_owned
|
||||
push!(recv_nodes, Int(nid))
|
||||
end
|
||||
end
|
||||
|
||||
if !isempty(send_nodes) || !isempty(recv_nodes)
|
||||
# Convert to local DOF indices
|
||||
send_dofs = Int[]
|
||||
for nid in send_nodes
|
||||
local_idx = nid - start_node + 1
|
||||
for d in 0:2
|
||||
push!(send_dofs, (local_idx - 1) * 3 + d + 1)
|
||||
end
|
||||
end
|
||||
|
||||
recv_dofs = Int[]
|
||||
for nid in recv_nodes
|
||||
ghost_idx = findfirst(==(nid), sort(collect(ghost_nodes)))
|
||||
for d in 0:2
|
||||
# Ghost DOFs come after owned DOFs
|
||||
push!(recv_dofs, length(owned_nodes) * 3 + (ghost_idx - 1) * 3 + d + 1)
|
||||
end
|
||||
end
|
||||
|
||||
if !isempty(send_dofs)
|
||||
interface_send[neighbor_rank] = send_dofs
|
||||
end
|
||||
if !isempty(recv_dofs)
|
||||
interface_recv[neighbor_rank] = recv_dofs
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
interface_neighbors = sort(collect(keys(interface_send) ∪ keys(interface_recv)))
|
||||
|
||||
return Partition(
|
||||
my_rank,
|
||||
owned_nodes,
|
||||
sort(collect(ghost_nodes)),
|
||||
local_elements,
|
||||
local_node_to_elems,
|
||||
interface_neighbors,
|
||||
interface_send,
|
||||
interface_recv
|
||||
)
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# GPU Kernel: Nodal Assembly
|
||||
# ============================================================================
|
||||
|
||||
function gpu_matvec_kernel!(
|
||||
y::CuDeviceArray{Float32,1},
|
||||
x::CuDeviceArray{Float32,1},
|
||||
nodes::CuDeviceArray{Node,1},
|
||||
elements::CuDeviceArray{Element,1},
|
||||
node_to_elems_offsets::CuDeviceArray{Int32,1},
|
||||
node_to_elems_data::CuDeviceArray{Int32,1},
|
||||
n_owned_nodes::Int32,
|
||||
)
|
||||
idx = (blockIdx().x - 1) * blockDim().x + threadIdx().x
|
||||
|
||||
if idx > n_owned_nodes
|
||||
return
|
||||
end
|
||||
|
||||
# This thread processes owned node idx
|
||||
node = nodes[idx]
|
||||
|
||||
dof_start = (idx - 1) * 3 + 1
|
||||
|
||||
# Initialize nodal contribution
|
||||
y1 = Float32(0.0)
|
||||
y2 = Float32(0.0)
|
||||
y3 = Float32(0.0)
|
||||
|
||||
# Get connected elements using CSR-like format (1-based indexing)
|
||||
if idx + Int32(1) > length(node_to_elems_offsets)
|
||||
return
|
||||
end
|
||||
|
||||
elem_start = node_to_elems_offsets[idx] + Int32(1)
|
||||
elem_end = node_to_elems_offsets[idx+Int32(1)]
|
||||
|
||||
for i in elem_start:elem_end
|
||||
if i > length(node_to_elems_data)
|
||||
return
|
||||
end
|
||||
elem_id = node_to_elems_data[i]
|
||||
if elem_id > length(elements)
|
||||
return
|
||||
end
|
||||
element = elements[elem_id]
|
||||
|
||||
# Add contribution from all nodes in this element
|
||||
for j in 1:8
|
||||
nid = element.connectivity[j]
|
||||
x_dof_start = (nid - 1) * 3 + 1
|
||||
|
||||
# Mock stiffness contribution
|
||||
y1 += Float32(0.1) * x[x_dof_start]
|
||||
y2 += Float32(0.1) * x[x_dof_start+1]
|
||||
y3 += Float32(0.1) * x[x_dof_start+2]
|
||||
end
|
||||
end
|
||||
|
||||
# Write to output
|
||||
y[dof_start] = y1
|
||||
y[dof_start+1] = y2
|
||||
y[dof_start+2] = y3
|
||||
|
||||
return nothing
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Multi-GPU Communication
|
||||
# ============================================================================
|
||||
|
||||
function exchange_ghost_values!(
|
||||
x_local::CuArray{Float32,1},
|
||||
partition::Partition,
|
||||
comm::MPI.Comm
|
||||
)
|
||||
"""Exchange interface DOF values between MPI ranks"""
|
||||
|
||||
# Prepare send/recv buffers on CPU
|
||||
send_bufs = Dict{Int,Vector{Float32}}()
|
||||
recv_bufs = Dict{Int,Vector{Float32}}()
|
||||
|
||||
# Copy data from GPU to CPU for sending
|
||||
x_cpu = Array(x_local)
|
||||
|
||||
for neighbor in partition.interface_neighbors
|
||||
if haskey(partition.interface_send, neighbor)
|
||||
send_dofs = partition.interface_send[neighbor]
|
||||
send_bufs[neighbor] = x_cpu[send_dofs]
|
||||
end
|
||||
|
||||
if haskey(partition.interface_recv, neighbor)
|
||||
recv_dofs = partition.interface_recv[neighbor]
|
||||
recv_bufs[neighbor] = zeros(Float32, length(recv_dofs))
|
||||
end
|
||||
end
|
||||
|
||||
# MPI communication
|
||||
requests = MPI.Request[]
|
||||
|
||||
# Post receives
|
||||
for neighbor in partition.interface_neighbors
|
||||
if haskey(recv_bufs, neighbor)
|
||||
req = MPI.Irecv!(recv_bufs[neighbor], comm; source=neighbor, tag=neighbor)
|
||||
push!(requests, req)
|
||||
end
|
||||
end
|
||||
|
||||
# Post sends
|
||||
for neighbor in partition.interface_neighbors
|
||||
if haskey(send_bufs, neighbor)
|
||||
req = MPI.Isend(send_bufs[neighbor], comm; dest=neighbor, tag=partition.rank)
|
||||
push!(requests, req)
|
||||
end
|
||||
end
|
||||
|
||||
# Wait for all communications
|
||||
MPI.Waitall(requests)
|
||||
|
||||
# Copy received data back to GPU
|
||||
for neighbor in partition.interface_neighbors
|
||||
if haskey(partition.interface_recv, neighbor)
|
||||
recv_dofs = partition.interface_recv[neighbor]
|
||||
x_cpu[recv_dofs] .= recv_bufs[neighbor]
|
||||
end
|
||||
end
|
||||
|
||||
# Update GPU array
|
||||
copyto!(x_local, x_cpu)
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmark
|
||||
# ============================================================================
|
||||
|
||||
function run_multigpu_benchmark(nx, ny, nz, n_warmup=5, n_runs=10)
|
||||
if rank == 0
|
||||
println("\n" * "="^70)
|
||||
println("Multi-GPU Benchmark: $nx × $ny × $nz mesh")
|
||||
println("="^70)
|
||||
end
|
||||
|
||||
# Create full mesh on all ranks
|
||||
nodes, elements = create_hex_mesh(nx, ny, nz)
|
||||
|
||||
if rank == 0
|
||||
println(" Total nodes: ", length(nodes))
|
||||
println(" Total elements: ", length(elements))
|
||||
println(" Total DOFs: ", 3 * length(nodes))
|
||||
end
|
||||
|
||||
# Partition for this rank
|
||||
partition = partition_mesh_for_rank(nodes, elements, rank, nranks)
|
||||
|
||||
n_owned = length(partition.owned_nodes)
|
||||
n_ghost = length(partition.ghost_nodes)
|
||||
n_local_dofs = 3 * (n_owned + n_ghost)
|
||||
|
||||
println("Rank $rank: $n_owned owned nodes, $n_ghost ghost nodes, " *
|
||||
"$(length(partition.local_elements)) elements")
|
||||
|
||||
# Prepare GPU data
|
||||
local_nodes = [nodes[i] for i in vcat(collect(partition.owned_nodes), partition.ghost_nodes)]
|
||||
|
||||
# Create mapping from global node ID to local index
|
||||
global_to_local_node = Dict{Int,Int32}()
|
||||
for (local_idx, global_nid) in enumerate(vcat(collect(partition.owned_nodes), partition.ghost_nodes))
|
||||
global_to_local_node[global_nid] = Int32(local_idx)
|
||||
end
|
||||
|
||||
# Remap element connectivity to local node indices
|
||||
local_elements = Element[]
|
||||
for global_eid in partition.local_elements
|
||||
element = elements[global_eid]
|
||||
# Convert global node IDs to local indices
|
||||
local_conn = ntuple(8) do i
|
||||
global_nid = Int(element.connectivity[i])
|
||||
global_to_local_node[global_nid]
|
||||
end
|
||||
push!(local_elements, Element(element.id, local_conn))
|
||||
end
|
||||
|
||||
# Create mapping from global element ID to local index (for CSR data)
|
||||
global_to_local_elem = Dict{Int,Int}()
|
||||
for (local_idx, global_id) in enumerate(partition.local_elements)
|
||||
global_to_local_elem[global_id] = local_idx
|
||||
end
|
||||
|
||||
# Convert node_to_elements to GPU-friendly flat format
|
||||
# Format: offsets array + flat data array (CSR-like)
|
||||
# IMPORTANT: Convert global element IDs to local indices
|
||||
node_to_elems_offsets = Int32[0]
|
||||
node_to_elems_data = Int32[]
|
||||
for arr in partition.node_to_elements
|
||||
# Map global element IDs to local indices
|
||||
local_indices = [global_to_local_elem[global_id] for global_id in arr]
|
||||
append!(node_to_elems_data, Int32.(local_indices))
|
||||
push!(node_to_elems_offsets, length(node_to_elems_data))
|
||||
end
|
||||
|
||||
# Debug: check element ID range
|
||||
if rank == 0 && length(node_to_elems_data) > 0
|
||||
min_elem_id = minimum(node_to_elems_data)
|
||||
max_elem_id = maximum(node_to_elems_data)
|
||||
println("\nCSR data element ID range: $min_elem_id to $max_elem_id")
|
||||
println("Local elements array size: $(length(local_elements))")
|
||||
if max_elem_id > length(local_elements)
|
||||
println("❌ WARNING: Element ID $max_elem_id > array size $(length(local_elements))")
|
||||
end
|
||||
end
|
||||
|
||||
# Transfer to GPU
|
||||
nodes_gpu = CuArray(local_nodes)
|
||||
elements_gpu = CuArray(local_elements)
|
||||
node_to_elems_offsets_gpu = CuArray(node_to_elems_offsets)
|
||||
node_to_elems_data_gpu = CuArray(node_to_elems_data)
|
||||
|
||||
# Debug: print array sizes
|
||||
if rank == 0
|
||||
println("\nArray sizes on GPU:")
|
||||
println(" nodes: $(length(nodes_gpu))")
|
||||
println(" elements: $(length(elements_gpu))")
|
||||
println(" node_to_elems_offsets: $(length(node_to_elems_offsets_gpu))")
|
||||
println(" node_to_elems_data: $(length(node_to_elems_data_gpu))")
|
||||
println(" Expected offsets length: $(n_owned + 1)")
|
||||
end
|
||||
|
||||
# Test vectors
|
||||
x_local = CUDA.rand(Float32, n_local_dofs)
|
||||
y_local = CUDA.zeros(Float32, n_local_dofs)
|
||||
|
||||
# Kernel launch parameters
|
||||
threads_per_block = 256
|
||||
n_blocks = cld(n_owned, threads_per_block)
|
||||
|
||||
if rank == 0
|
||||
println("\nGPU configuration:")
|
||||
println(" Threads per block: $threads_per_block")
|
||||
println(" Blocks per rank: $n_blocks")
|
||||
end
|
||||
|
||||
# Warmup
|
||||
for _ in 1:n_warmup
|
||||
exchange_ghost_values!(x_local, partition, comm)
|
||||
CUDA.@sync @cuda threads = threads_per_block blocks = n_blocks gpu_matvec_kernel!(
|
||||
y_local, x_local, nodes_gpu, elements_gpu,
|
||||
node_to_elems_offsets_gpu, node_to_elems_data_gpu, Int32(n_owned)
|
||||
)
|
||||
end
|
||||
|
||||
MPI.Barrier(comm)
|
||||
|
||||
# Benchmark
|
||||
times = Float64[]
|
||||
comm_times = Float64[]
|
||||
compute_times = Float64[]
|
||||
|
||||
for _ in 1:n_runs
|
||||
t_start = time_ns()
|
||||
|
||||
# Communication
|
||||
t_comm_start = time_ns()
|
||||
exchange_ghost_values!(x_local, partition, comm)
|
||||
MPI.Barrier(comm)
|
||||
t_comm_end = time_ns()
|
||||
|
||||
# Computation
|
||||
t_compute_start = time_ns()
|
||||
CUDA.@sync @cuda threads = threads_per_block blocks = n_blocks gpu_matvec_kernel!(
|
||||
y_local, x_local, nodes_gpu, elements_gpu,
|
||||
node_to_elems_offsets_gpu, node_to_elems_data_gpu, Int32(n_owned)
|
||||
)
|
||||
MPI.Barrier(comm)
|
||||
t_compute_end = time_ns()
|
||||
|
||||
t_end = time_ns()
|
||||
|
||||
push!(times, (t_end - t_start) / 1e9)
|
||||
push!(comm_times, (t_comm_end - t_comm_start) / 1e9)
|
||||
push!(compute_times, (t_compute_end - t_compute_start) / 1e9)
|
||||
end
|
||||
|
||||
# Gather results
|
||||
local_time = minimum(times)
|
||||
local_comm = minimum(comm_times)
|
||||
local_compute = minimum(compute_times)
|
||||
|
||||
all_times = MPI.Gather(local_time, 0, comm)
|
||||
all_comm = MPI.Gather(local_comm, 0, comm)
|
||||
all_compute = MPI.Gather(local_compute, 0, comm)
|
||||
|
||||
if rank == 0
|
||||
println("\nResults:")
|
||||
println(" Rank | Owned Nodes | Total Time | Comm Time | Compute Time | Comm %")
|
||||
println(" " * "-"^70)
|
||||
for r in 0:(nranks-1)
|
||||
nodes_str = lpad(string(length(partition.owned_nodes)), 11)
|
||||
total_str = @sprintf("%.3f ms", all_times[r+1] * 1000)
|
||||
comm_str = @sprintf("%.3f ms", all_comm[r+1] * 1000)
|
||||
compute_str = @sprintf("%.3f ms", all_compute[r+1] * 1000)
|
||||
comm_pct = @sprintf("%.1f%%", all_comm[r+1] / all_times[r+1] * 100)
|
||||
|
||||
println(" $r | $nodes_str | $(lpad(total_str, 10)) | " *
|
||||
"$(lpad(comm_str, 9)) | $(lpad(compute_str, 12)) | $(lpad(comm_pct, 6))")
|
||||
end
|
||||
|
||||
max_time = maximum(all_times)
|
||||
avg_compute = sum(all_compute) / length(all_compute)
|
||||
avg_comm = sum(all_comm) / length(all_comm)
|
||||
|
||||
println("\n Maximum time: ", @sprintf("%.3f ms", max_time * 1000))
|
||||
println(" Average compute: ", @sprintf("%.3f ms", avg_compute * 1000))
|
||||
println(" Average communication: ", @sprintf("%.3f ms", avg_comm * 1000))
|
||||
println(" Communication overhead: ", @sprintf("%.1f%%", avg_comm / max_time * 100))
|
||||
|
||||
throughput = length(nodes) / max_time / 1e6
|
||||
println(" Throughput: ", @sprintf("%.2f Mnodes/s", throughput))
|
||||
end
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Main
|
||||
# ============================================================================
|
||||
|
||||
if rank == 0
|
||||
println("Starting benchmarks...")
|
||||
println()
|
||||
end
|
||||
|
||||
# Run benchmarks with increasing mesh sizes
|
||||
run_multigpu_benchmark(30, 30, 30)
|
||||
run_multigpu_benchmark(50, 50, 50)
|
||||
run_multigpu_benchmark(70, 70, 70)
|
||||
|
||||
if rank == 0
|
||||
println("\n" * "="^70)
|
||||
println("✓ Multi-GPU Benchmark Complete")
|
||||
println("="^70)
|
||||
end
|
||||
|
||||
MPI.Finalize()
|
||||
@@ -1,329 +0,0 @@
|
||||
# Multi-GPU Nodal Assembly Results
|
||||
|
||||
**Date:** November 9, 2025
|
||||
**Hardware:** 2× MPI ranks, 1× NVIDIA RTX A2000 12GB (shared between ranks)
|
||||
**Software:** Julia 1.12.1, CUDA.jl, MPI.jl
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
**✅ Multi-GPU implementation WORKS!** Successfully ran multi-GPU nodal assembly with MPI + CUDA.
|
||||
|
||||
**Key Achievement:** Implemented nodal assembly on GPU with:
|
||||
- CSR-format node-to-elements connectivity (zero allocation)
|
||||
- Proper global-to-local index remapping for elements and nodes
|
||||
- MPI ghost value exchange for domain interfaces
|
||||
- Validated correctness (all ranks complete successfully)
|
||||
|
||||
**Performance:**
|
||||
- **Throughput:** 115-302 Mnodes/s (scales with mesh size)
|
||||
- **Communication overhead:** 29-61% (MPI transfers dominate for small/medium meshes)
|
||||
- **Compute performance:** GPU kernel is fast (0.16-0.56 ms), communication is bottleneck
|
||||
|
||||
---
|
||||
|
||||
## Detailed Results
|
||||
|
||||
### Benchmark Configuration
|
||||
- **MPI ranks:** 2
|
||||
- **GPU per rank:** 1 (shared device for both ranks in this test)
|
||||
- **Partitioning:** Slab decomposition (nodes split evenly)
|
||||
- **Element type:** Hex8 (8-node hexahedron)
|
||||
- **Kernel:** Mock stiffness (simplified matvec for validation)
|
||||
- **Warmup:** 10 iterations
|
||||
- **Measurement:** 100 iterations (timed)
|
||||
|
||||
### Performance Table
|
||||
|
||||
| Mesh Size | Total Nodes | Total DOFs | Owned/Rank | Ghost/Rank | Throughput | Comm % |
|
||||
|-----------|-------------|------------|------------|------------|------------|--------|
|
||||
| 30³ | 27,000 | 81,000 | 13,500 | 900 | 114.84 Mnodes/s | 29.1% |
|
||||
| 50³ | 125,000 | 375,000 | 62,500 | 2,500 | 130.64 Mnodes/s | 60.7% |
|
||||
| 70³ | 343,000 | 1,029,000 | 171,500 | 4,900 | 301.83 Mnodes/s | 50.9% |
|
||||
|
||||
### Detailed Timing Breakdown
|
||||
|
||||
**30×30×30 mesh:**
|
||||
```
|
||||
Rank 0: Total 0.235 ms = Compute 0.159 ms + Comm 0.069 ms (29.4%)
|
||||
Rank 1: Total 0.227 ms = Compute 0.159 ms + Comm 0.068 ms (29.9%)
|
||||
Throughput: 114.84 Mnodes/s
|
||||
```
|
||||
|
||||
**50×50×50 mesh:**
|
||||
```
|
||||
Rank 0: Total 0.957 ms = Compute 0.373 ms + Comm 0.581 ms (60.7%)
|
||||
Rank 1: Total 0.957 ms = Compute 0.375 ms + Comm 0.581 ms (60.7%)
|
||||
Throughput: 130.64 Mnodes/s
|
||||
```
|
||||
|
||||
**70×70×70 mesh:**
|
||||
```
|
||||
Rank 0: Total 1.136 ms = Compute 0.557 ms + Comm 0.579 ms (50.9%)
|
||||
Rank 1: Total 1.136 ms = Compute 0.557 ms + Comm 0.579 ms (50.9%)
|
||||
Throughput: 301.83 Mnodes/s
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Comparison with CPU Baseline
|
||||
|
||||
**From `nodal_assembly_scalability.jl` (validated Nov 9, 2025):**
|
||||
|
||||
### CPU Multi-Threading (8 threads, single node)
|
||||
|
||||
| Mesh Size | Nodes | Single-Thread | 8 Threads | Speedup | Efficiency |
|
||||
|-----------|---------|---------------|-----------|---------|------------|
|
||||
| 20³ | 8,000 | 3.4 Mnodes/s | 49.6 Mnodes/s | 14.6× | 182% |
|
||||
| 40³ | 64,000 | 3.6 Mnodes/s | 54.8 Mnodes/s | 15.2× | 189% |
|
||||
| 60³ | 216,000 | 3.6 Mnodes/s | 47.9 Mnodes/s | 13.3× | 166% |
|
||||
|
||||
### CPU Partitioned (4 partitions, sequential)
|
||||
|
||||
| Mesh Size | Nodes | Throughput | Speedup vs Single-Thread | Interface Overhead |
|
||||
|-----------|---------|------------|--------------------------|-------------------|
|
||||
| 20³ | 8,000 | 29.7 Mnodes/s | 8.5× | 40.1% |
|
||||
| 40³ | 64,000 | 28.8 Mnodes/s | 8.0× | 22.4% |
|
||||
| 60³ | 216,000 | 27.3 Mnodes/s | 7.6× | 10.3% |
|
||||
|
||||
### GPU vs CPU Comparison
|
||||
|
||||
**Throughput Comparison (approximate mesh sizes):**
|
||||
|
||||
| Mesh | CPU Single-Thread | CPU 8-Thread | CPU 4-Partition | GPU 2-Rank (MPI) | GPU Speedup vs 8-Thread |
|
||||
|------|-------------------|--------------|-----------------|------------------|-------------------------|
|
||||
| ~30³ | 3.5 Mnodes/s | ~50 Mnodes/s | ~29 Mnodes/s | 114.84 Mnodes/s | **2.3×** |
|
||||
| ~60³ | 3.6 Mnodes/s | 47.9 Mnodes/s | 27.3 Mnodes/s | ~200 Mnodes/s (interpolated) | **4.2×** |
|
||||
| 70³ | 3.6 Mnodes/s | ~48 Mnodes/s (est) | ~28 Mnodes/s (est) | 301.83 Mnodes/s | **6.3×** |
|
||||
|
||||
**Key Observations:**
|
||||
1. ✅ GPU is **2-6× faster** than CPU 8-thread for same mesh size
|
||||
2. ✅ GPU throughput scales better with mesh size (114 → 302 Mnodes/s)
|
||||
3. ⚠️ GPU communication overhead (29-61%) higher than CPU partitioned (10-40%)
|
||||
4. 🎯 GPU shines on larger meshes (70³: 6.3× faster than CPU)
|
||||
|
||||
---
|
||||
|
||||
## Analysis & Insights
|
||||
|
||||
### What Worked Well ✅
|
||||
|
||||
1. **Nodal assembly pattern on GPU:**
|
||||
- Each thread processes one node (no atomics!)
|
||||
- Gathers contributions from connected elements
|
||||
- Direct write to owned DOFs (no race conditions)
|
||||
|
||||
2. **CSR-format node-to-elements:**
|
||||
- `offsets[node_id]` → start of element list
|
||||
- `data[offsets[i]:offsets[i+1]]` → element IDs
|
||||
- Zero allocation, type-stable, GPU-friendly
|
||||
|
||||
3. **Global-to-local index remapping:**
|
||||
- Element IDs: Global mesh → Local partition indices
|
||||
- Node IDs in connectivity: Global mesh → Local partition indices
|
||||
- Critical for correctness with sliced arrays
|
||||
|
||||
4. **MPI ghost exchange:**
|
||||
- Interface nodes identified correctly
|
||||
- Ghost values exchanged between ranks
|
||||
- Enables domain decomposition
|
||||
|
||||
### Performance Bottlenecks ⚠️
|
||||
|
||||
1. **Communication overhead dominates small/medium meshes:**
|
||||
- 30³ mesh: 29% communication
|
||||
- 50³ mesh: **61% communication** (worst case!)
|
||||
- 70³ mesh: 51% communication
|
||||
- **Root cause:** MPI transfers CPU ↔ GPU for every iteration
|
||||
|
||||
2. **Single GPU shared between 2 MPI ranks:**
|
||||
- Both ranks compete for same GPU
|
||||
- No true parallelism in this test configuration
|
||||
- Need multiple GPUs for real multi-GPU scaling
|
||||
|
||||
3. **Mock kernel (simplified stiffness):**
|
||||
- Real FEM kernel would be more compute-intensive
|
||||
- Would reduce communication percentage
|
||||
- Current kernel is memory-bound
|
||||
|
||||
### Opportunities for Improvement 🎯
|
||||
|
||||
1. **CUDA-aware MPI:**
|
||||
- Direct GPU-to-GPU transfers (no CPU staging)
|
||||
- Can reduce communication time by 50-80%
|
||||
- Requires recompilation of MPI with CUDA support
|
||||
|
||||
2. **Multiple physical GPUs:**
|
||||
- Current test uses 1 GPU for 2 ranks (shared)
|
||||
- True multi-GPU: Each rank gets own GPU
|
||||
- Would enable concurrent execution
|
||||
|
||||
3. **Larger elements (higher-order):**
|
||||
- Tet10, Hex20, Hex27 have more work per element
|
||||
- More compute per node → reduces communication %
|
||||
- Better compute/communication ratio
|
||||
|
||||
4. **Full element stiffness:**
|
||||
- Real FEM: Integration loops, material models, plasticity
|
||||
- 10-100× more work per element
|
||||
- Communication becomes negligible (<5%)
|
||||
|
||||
5. **Batched assembly:**
|
||||
- Assemble multiple timesteps before MPI sync
|
||||
- Amortize communication cost
|
||||
- Useful for explicit dynamics
|
||||
|
||||
---
|
||||
|
||||
## Technical Details
|
||||
|
||||
### Data Structures
|
||||
|
||||
**Node (immutable, 32 bytes):**
|
||||
```julia
|
||||
struct Node
|
||||
id::Int32
|
||||
x::Float32
|
||||
y::Float32
|
||||
z::Float32
|
||||
end
|
||||
```
|
||||
|
||||
**Element (immutable, 36 bytes):**
|
||||
```julia
|
||||
struct Element
|
||||
id::Int32
|
||||
connectivity::NTuple{8, Int32} # Hex8
|
||||
end
|
||||
```
|
||||
|
||||
**Partition:**
|
||||
- `owned_nodes`: Nodes owned by this rank
|
||||
- `ghost_nodes`: Nodes owned by neighbors (interface)
|
||||
- `local_elements`: Elements touching owned nodes
|
||||
- `node_to_elements`: Inverse connectivity (node → elements)
|
||||
- `interface_send/recv`: MPI communication patterns
|
||||
|
||||
### GPU Kernel (Simplified)
|
||||
|
||||
```julia
|
||||
function gpu_matvec_kernel!(y, x, nodes, elements, offsets, data, n_owned)
|
||||
idx = (blockIdx().x - 1) * blockDim().x + threadIdx().x
|
||||
|
||||
if idx > n_owned
|
||||
return # Thread beyond owned nodes
|
||||
end
|
||||
|
||||
# Initialize accumulator
|
||||
fx = fy = fz = 0.0f0
|
||||
|
||||
# Loop over connected elements (CSR access)
|
||||
elem_start = offsets[idx] + 1
|
||||
elem_end = offsets[idx + 1]
|
||||
for i in elem_start:elem_end
|
||||
elem_id = data[i]
|
||||
element = elements[elem_id]
|
||||
|
||||
# Gather from element nodes
|
||||
for j in 1:8
|
||||
nid = element.connectivity[j]
|
||||
dof_base = (nid - 1) * 3 + 1
|
||||
fx += 0.1f0 * x[dof_base]
|
||||
fy += 0.1f0 * x[dof_base + 1]
|
||||
fz += 0.1f0 * x[dof_base + 2]
|
||||
end
|
||||
end
|
||||
|
||||
# Write result (owned DOFs only)
|
||||
dof_base = (idx - 1) * 3 + 1
|
||||
y[dof_base] = fx
|
||||
y[dof_base + 1] = fy
|
||||
y[dof_base + 2] = fz
|
||||
end
|
||||
```
|
||||
|
||||
### Key Implementation Challenges & Solutions
|
||||
|
||||
**Challenge 1:** CuArray{CuArray} not supported
|
||||
**Solution:** Flatten to CSR format (offsets + data arrays)
|
||||
|
||||
**Challenge 2:** Global element IDs in CSR data, but local array
|
||||
**Solution:** Create `global_to_local_elem` mapping, remap before GPU transfer
|
||||
|
||||
**Challenge 3:** Global node IDs in element connectivity
|
||||
**Solution:** Create `global_to_local_node` mapping, rebuild elements with local indices
|
||||
|
||||
**Challenge 4:** BoundsError during kernel execution
|
||||
**Solution:** All three index spaces must be consistent (nodes, elements, DOFs)
|
||||
|
||||
---
|
||||
|
||||
## Conclusions
|
||||
|
||||
### Claims We Can Now Make ✅
|
||||
|
||||
1. ✅ **Nodal assembly works on GPU** - Validated with working implementation
|
||||
2. ✅ **2-6× faster than CPU multi-threading** - Real measurements on same mesh
|
||||
3. ✅ **Scales to 343K nodes / 1M DOFs** - Successfully ran 70³ mesh
|
||||
4. ✅ **Communication overhead acceptable** - 29-61% (will improve with CUDA-aware MPI)
|
||||
5. ✅ **CSR format enables zero-allocation** - No dynamic memory in kernel
|
||||
|
||||
### Claims We CANNOT Yet Make ⚠️
|
||||
|
||||
1. ⚠️ **Multi-GPU strong scaling** - Only tested 1 GPU with 2 ranks (not true multi-GPU)
|
||||
2. ⚠️ **Production-ready performance** - Mock kernel, needs real FEM stiffness
|
||||
3. ⚠️ **Weak scaling to N GPUs** - Need cluster with multiple GPUs
|
||||
4. ⚠️ **Better than Gridap/Ferrite** - Haven't compared with other libraries
|
||||
5. ⚠️ **Contact mechanics on GPU** - Not yet implemented
|
||||
|
||||
### Next Steps 🎯
|
||||
|
||||
**Immediate (validate architecture):**
|
||||
1. Test with multiple physical GPUs (2-4 GPUs on cluster)
|
||||
2. Implement real element stiffness (not mock)
|
||||
3. Measure CUDA-aware MPI improvement
|
||||
4. Add higher-order elements (Tet10, Hex20)
|
||||
|
||||
**Short-term (production features):**
|
||||
1. Material state updates on GPU (plasticity, damage)
|
||||
2. Contact detection and assembly on GPU
|
||||
3. Preconditioned GMRES on GPU (full solver)
|
||||
4. Integration with JuliaFEM element library
|
||||
|
||||
**Long-term (scale-up):**
|
||||
1. Weak scaling study (1-64 GPUs)
|
||||
2. Strong scaling study (fixed problem, varying GPUs)
|
||||
3. Comparison with Gridap.jl + PETSc
|
||||
4. Real-world contact mechanics problem (1M+ DOFs)
|
||||
|
||||
---
|
||||
|
||||
## Files & Artifacts
|
||||
|
||||
**Benchmark code:**
|
||||
- `benchmarks/multigpu_mpi_benchmark.jl` (555 lines, working)
|
||||
|
||||
**CPU baseline (validated):**
|
||||
- `benchmarks/nodal_assembly_scalability.jl` (500 lines)
|
||||
|
||||
**Documentation:**
|
||||
- `docs/book/multigpu_nodal_assembly.md` (design, needs update with real data)
|
||||
- `docs/book/nodal_assembly_gpu_pattern.md` (architecture)
|
||||
- `demos/gpu_nodal_assembly_demo.jl` (educational demo)
|
||||
|
||||
**This report:**
|
||||
- `benchmarks/multigpu_results_2025-11-09.md`
|
||||
|
||||
---
|
||||
|
||||
## Acknowledgments
|
||||
|
||||
**User (Jukka):** Demanded real measurements, not designs. Caught AI making unvalidated claims. Insisted on "Just run" - forcing validation before documentation.
|
||||
|
||||
**Key Insight:** "Did you actually run that code?" - Best engineering feedback possible. No more design documents without validation!
|
||||
|
||||
---
|
||||
|
||||
**Status:** ✅ Multi-GPU architecture VALIDATED
|
||||
**Verdict:** Nodal assembly on GPU is **feasible and fast**. Communication overhead acceptable. Ready for production implementation.
|
||||
@@ -1,260 +0,0 @@
|
||||
"""
|
||||
Performance Analysis: NeoHookean Hyperelastic Material
|
||||
|
||||
Benchmarks automatic differentiation overhead in hyperelastic stress computation.
|
||||
|
||||
Key Questions:
|
||||
1. What is the cost of AD compared to LinearElastic?
|
||||
2. Is the implementation allocation-free?
|
||||
3. How does performance scale with problem size?
|
||||
4. What is the breakdown of strain energy vs stress vs tangent?
|
||||
|
||||
Results inform whether AD is suitable for production FEM assembly loops.
|
||||
"""
|
||||
|
||||
using BenchmarkTools
|
||||
using Tensors
|
||||
using LinearAlgebra
|
||||
using Printf
|
||||
|
||||
# Load implementations
|
||||
include("../src/materials/abstract_material.jl")
|
||||
include("../src/materials/linear_elastic.jl")
|
||||
include("../src/materials/neo_hookean.jl")
|
||||
|
||||
println("="^80)
|
||||
println("NeoHookean Hyperelastic Material - Performance Analysis")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# =============================================================================
|
||||
# Test 1: Single Stress Evaluation (typical FEM use case)
|
||||
# =============================================================================
|
||||
println("Test 1: Single Stress Evaluation")
|
||||
println("-"^80)
|
||||
|
||||
# Create materials
|
||||
neo = NeoHookean(E_mod=200e9, nu=0.3) # Steel-like properties
|
||||
linear = LinearElastic(E=200e9, ν=0.3)
|
||||
|
||||
# Test strain (moderate deformation)
|
||||
E_strain = SymmetricTensor{2,3}((0.01, 0.005, 0.003, -0.002, 0.004, 0.006))
|
||||
|
||||
# Compile first
|
||||
compute_stress(neo, E_strain, nothing, 0.0)
|
||||
compute_stress(linear, E_strain, nothing, 0.0)
|
||||
|
||||
# Benchmark
|
||||
println("\nLinearElastic (manual derivatives):")
|
||||
t_linear = @benchmark compute_stress($linear, $E_strain, nothing, 0.0)
|
||||
display(t_linear)
|
||||
println()
|
||||
|
||||
println("\nNeoHookean (automatic differentiation):")
|
||||
t_neo = @benchmark compute_stress($neo, $E_strain, nothing, 0.0)
|
||||
display(t_neo)
|
||||
println()
|
||||
|
||||
# Compute overhead
|
||||
overhead = median(t_neo).time / median(t_linear).time
|
||||
println("\nAD Overhead: $(round(overhead, digits=1))x")
|
||||
println("Absolute difference: $(round((median(t_neo).time - median(t_linear).time)/1e3, digits=1)) μs")
|
||||
|
||||
# =============================================================================
|
||||
# Test 2: Allocation Check
|
||||
# =============================================================================
|
||||
println("\n" * "="^80)
|
||||
println("Test 2: Allocation Analysis")
|
||||
println("-"^80)
|
||||
|
||||
allocs_neo = @allocated compute_stress(neo, E_strain, nothing, 0.0)
|
||||
allocs_linear = @allocated compute_stress(linear, E_strain, nothing, 0.0)
|
||||
|
||||
println("LinearElastic allocations: $allocs_linear bytes")
|
||||
println("NeoHookean allocations: $allocs_neo bytes")
|
||||
|
||||
if allocs_neo == 0
|
||||
println("✅ Zero-allocation achieved!")
|
||||
else
|
||||
println("⚠️ Allocations detected - investigate")
|
||||
end
|
||||
|
||||
# =============================================================================
|
||||
# Test 3: Component Breakdown (where does time go?)
|
||||
# =============================================================================
|
||||
println("\n" * "="^80)
|
||||
println("Test 3: Component Breakdown")
|
||||
println("-"^80)
|
||||
|
||||
# Test strain energy only
|
||||
C = 2E_strain + one(E_strain)
|
||||
println("\nStrain energy computation:")
|
||||
t_energy = @benchmark strain_energy($neo, $C)
|
||||
display(t_energy)
|
||||
println()
|
||||
|
||||
# Stress only (includes strain energy + gradient)
|
||||
println("\nComplete stress computation (energy + gradient + hessian):")
|
||||
display(t_neo)
|
||||
println()
|
||||
|
||||
energy_fraction = median(t_energy).time / median(t_neo).time
|
||||
println("Strain energy is $(round(energy_fraction*100, digits=1))% of total time")
|
||||
println("AD overhead (gradient + hessian) is $(round((1-energy_fraction)*100, digits=1))% of total time")
|
||||
|
||||
# =============================================================================
|
||||
# Test 4: Deformation Magnitude Sensitivity
|
||||
# =============================================================================
|
||||
println("\n" * "="^80)
|
||||
println("Test 4: Performance vs Deformation Magnitude")
|
||||
println("-"^80)
|
||||
|
||||
strain_levels = [1e-6, 1e-4, 1e-2, 0.1, 0.5]
|
||||
times = Float64[]
|
||||
|
||||
for ε_mag in strain_levels
|
||||
E_test = SymmetricTensor{2,3}((ε_mag, 0.0, 0.0, 0.0, 0.0, 0.0))
|
||||
compute_stress(neo, E_test, nothing, 0.0) # Compile
|
||||
t = @benchmark compute_stress($neo, $E_test, nothing, 0.0) samples = 1000
|
||||
push!(times, median(t).time)
|
||||
@printf("Strain magnitude: %.1e → Time: %.1f ns\n", ε_mag, median(t).time)
|
||||
end
|
||||
|
||||
time_variation = (maximum(times) - minimum(times)) / minimum(times) * 100
|
||||
println("\nTime variation across strain levels: $(round(time_variation, digits=1))%")
|
||||
if time_variation < 10
|
||||
println("✅ Performance is strain-independent (good for Newton solvers)")
|
||||
else
|
||||
println("⚠️ Performance varies with strain (may impact Newton convergence)")
|
||||
end
|
||||
|
||||
# =============================================================================
|
||||
# Test 5: Tangent Accuracy vs Finite Difference
|
||||
# =============================================================================
|
||||
println("\n" * "="^80)
|
||||
println("Test 5: Tangent Accuracy (AD vs Finite Difference)")
|
||||
println("-"^80)
|
||||
|
||||
E_base = SymmetricTensor{2,3}((0.01, 0.005, 0.003, -0.002, 0.004, 0.006))
|
||||
S_ad, 𝔻_ad, _ = compute_stress(neo, E_base, nothing, 0.0)
|
||||
|
||||
# Finite difference tangent
|
||||
ε_fd = 1e-8
|
||||
errors = Float64[]
|
||||
for i in 1:6
|
||||
E_pert_data = collect(E_base.data)
|
||||
E_pert_data[i] += ε_fd
|
||||
E_pert = SymmetricTensor{2,3}(tuple(E_pert_data...))
|
||||
|
||||
S_pert, _, _ = compute_stress(neo, E_pert, nothing, 0.0)
|
||||
|
||||
∂S∂E_fd = (S_pert - S_ad) / ε_fd
|
||||
error = norm(∂S∂E_fd) # Simplified error metric
|
||||
push!(errors, error)
|
||||
end
|
||||
|
||||
println("Tangent norm comparison:")
|
||||
println(" AD tangent: $(round(norm(𝔻_ad), sigdigits=6))")
|
||||
println(" FD derivative: $(round(mean(errors), sigdigits=6))")
|
||||
println(" Relative diff: $(round((norm(𝔻_ad) - mean(errors))/norm(𝔻_ad)*100, digits=2))%")
|
||||
println("\n✅ AD provides exact derivatives (limited only by machine precision)")
|
||||
|
||||
# =============================================================================
|
||||
# Test 6: Memory Footprint
|
||||
# =============================================================================
|
||||
println("\n" * "="^80)
|
||||
println("Test 6: Memory Footprint")
|
||||
println("-"^80)
|
||||
|
||||
println("Struct sizes:")
|
||||
println(" NeoHookean: $(sizeof(neo)) bytes")
|
||||
println(" LinearElastic: $(sizeof(linear)) bytes")
|
||||
println("\nReturn value sizes:")
|
||||
println(" SymmetricTensor{2,3}: $(sizeof(S_ad)) bytes")
|
||||
println(" SymmetricTensor{4,3}: $(sizeof(𝔻_ad)) bytes")
|
||||
println(" Total per evaluation: $(sizeof(S_ad) + sizeof(𝔻_ad)) bytes")
|
||||
|
||||
# =============================================================================
|
||||
# Test 7: Scaling with Multiple Evaluations
|
||||
# =============================================================================
|
||||
println("\n" * "="^80)
|
||||
println("Test 7: Assembly Loop Simulation (1000 evaluations)")
|
||||
println("-"^80)
|
||||
|
||||
n_evals = 1000
|
||||
strain_samples = [SymmetricTensor{2,3}((rand(), rand(), rand(), rand(), rand(), rand())) * 0.01
|
||||
for _ in 1:n_evals]
|
||||
|
||||
# Compile
|
||||
for E in strain_samples[1:10]
|
||||
compute_stress(neo, E, nothing, 0.0)
|
||||
compute_stress(linear, E, nothing, 0.0)
|
||||
end
|
||||
|
||||
println("\nLinearElastic ($n_evals evaluations):")
|
||||
t_linear_loop = @benchmark begin
|
||||
for E in $strain_samples
|
||||
compute_stress($linear, E, nothing, 0.0)
|
||||
end
|
||||
end
|
||||
display(t_linear_loop)
|
||||
println()
|
||||
|
||||
println("\nNeoHookean ($n_evals evaluations):")
|
||||
t_neo_loop = @benchmark begin
|
||||
for E in $strain_samples
|
||||
compute_stress($neo, E, nothing, 0.0)
|
||||
end
|
||||
end
|
||||
display(t_neo_loop)
|
||||
println()
|
||||
|
||||
overhead_loop = median(t_neo_loop).time / median(t_linear_loop).time
|
||||
per_eval_neo = median(t_neo_loop).time / n_evals
|
||||
per_eval_linear = median(t_linear_loop).time / n_evals
|
||||
|
||||
println("Per-evaluation time:")
|
||||
println(" LinearElastic: $(round(per_eval_linear, digits=1)) ns")
|
||||
println(" NeoHookean: $(round(per_eval_neo, digits=1)) ns")
|
||||
println(" Overhead: $(round(overhead_loop, digits=1))x")
|
||||
|
||||
# =============================================================================
|
||||
# Summary and Recommendations
|
||||
# =============================================================================
|
||||
println("\n" * "="^80)
|
||||
println("SUMMARY AND RECOMMENDATIONS")
|
||||
println("="^80)
|
||||
|
||||
total_overhead = median(t_neo).time / median(t_linear).time
|
||||
|
||||
println("\n📊 Performance Metrics:")
|
||||
println(" Single evaluation: $(round(median(t_neo).time, digits=1)) ns")
|
||||
println(" AD overhead: $(round(total_overhead, digits=1))x")
|
||||
println(" Allocations: $(allocs_neo) bytes")
|
||||
println(" Strain-independent: $(time_variation < 10 ? "Yes ✅" : "No ⚠️")")
|
||||
|
||||
println("\n🎯 Recommendations:")
|
||||
|
||||
if total_overhead < 5
|
||||
println(" ✅ EXCELLENT: AD overhead < 5x, suitable for production FEM")
|
||||
elseif total_overhead < 10
|
||||
println(" ✅ GOOD: AD overhead < 10x, acceptable for most applications")
|
||||
elseif total_overhead < 20
|
||||
println(" ⚠️ MODERATE: AD overhead < 20x, consider for prototyping only")
|
||||
else
|
||||
println(" ❌ HIGH: AD overhead > 20x, manual derivatives recommended for production")
|
||||
end
|
||||
|
||||
if allocs_neo == 0
|
||||
println(" ✅ Zero allocations achieved - suitable for tight loops")
|
||||
else
|
||||
println(" ⚠️ Allocations detected - profile and optimize")
|
||||
end
|
||||
|
||||
println("\n💡 Use Cases:")
|
||||
println(" • Research code: Strongly recommended (correctness > speed)")
|
||||
println(" • Prototyping: Excellent (rapid implementation)")
|
||||
println(" • Production: $(total_overhead < 10 ? "Acceptable" : "Profile first") ($(round(total_overhead, digits=1))x overhead)")
|
||||
println(" • Contact: Excellent (unsymmetric tangent, complex derivatives)")
|
||||
|
||||
println("\n" * "="^80)
|
||||
@@ -1,438 +0,0 @@
|
||||
#!/usr/bin/env julia
|
||||
#
|
||||
# Nodal Assembly Scalability Benchmark
|
||||
#
|
||||
# Tests:
|
||||
# 1. Single-threaded baseline
|
||||
# 2. Multi-threaded scaling (2, 4, 8 threads)
|
||||
# 3. Actual speedup measurements
|
||||
# 4. Interface communication overhead
|
||||
#
|
||||
# Run with: julia --project=. -t 8 benchmarks/nodal_assembly_scalability.jl
|
||||
|
||||
using LinearAlgebra
|
||||
using Printf
|
||||
using Base.Threads
|
||||
|
||||
println("="^70)
|
||||
println("Nodal Assembly Scalability Benchmark")
|
||||
println("="^70)
|
||||
println()
|
||||
println("Julia threads available: ", nthreads())
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# Data Structures
|
||||
# ============================================================================
|
||||
|
||||
struct Node
|
||||
id::Int
|
||||
x::Float64
|
||||
y::Float64
|
||||
z::Float64
|
||||
end
|
||||
|
||||
struct Element
|
||||
id::Int
|
||||
connectivity::NTuple{8,Int} # Hex8
|
||||
end
|
||||
|
||||
struct Partition
|
||||
rank::Int
|
||||
owned_nodes::UnitRange{Int}
|
||||
ghost_nodes::Vector{Int}
|
||||
local_elements::Vector{Int}
|
||||
node_to_elements::Vector{Vector{Int}}
|
||||
interface_nodes::Dict{Int,Vector{Int}}
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Mesh Generation
|
||||
# ============================================================================
|
||||
|
||||
function create_hex_mesh(nx, ny, nz)
|
||||
"""Create structured hexahedral mesh"""
|
||||
n_nodes = nx * ny * nz
|
||||
n_elements = (nx - 1) * (ny - 1) * (nz - 1)
|
||||
|
||||
# Create nodes
|
||||
nodes = Node[]
|
||||
for k in 1:nz, j in 1:ny, i in 1:nx
|
||||
node_id = (k - 1) * nx * ny + (j - 1) * nx + i
|
||||
push!(nodes, Node(node_id, Float64(i), Float64(j), Float64(k)))
|
||||
end
|
||||
|
||||
# Create elements (Hex8)
|
||||
elements = Element[]
|
||||
for k in 1:(nz-1), j in 1:(ny-1), i in 1:(nx-1)
|
||||
n1 = (k - 1) * nx * ny + (j - 1) * nx + i
|
||||
n2 = n1 + 1
|
||||
n3 = n2 + nx
|
||||
n4 = n1 + nx
|
||||
n5 = n1 + nx * ny
|
||||
n6 = n2 + nx * ny
|
||||
n7 = n3 + nx * ny
|
||||
n8 = n4 + nx * ny
|
||||
|
||||
elem_id = length(elements) + 1
|
||||
push!(elements, Element(elem_id, (n1, n2, n3, n4, n5, n6, n7, n8)))
|
||||
end
|
||||
|
||||
return nodes, elements
|
||||
end
|
||||
|
||||
function build_node_to_elements(nodes, elements)
|
||||
"""Build inverse connectivity"""
|
||||
node_to_elems = [Int[] for _ in 1:length(nodes)]
|
||||
|
||||
for (elem_id, element) in enumerate(elements)
|
||||
for node_id in element.connectivity
|
||||
push!(node_to_elems[node_id], elem_id)
|
||||
end
|
||||
end
|
||||
|
||||
return node_to_elems
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Partitioning
|
||||
# ============================================================================
|
||||
|
||||
function partition_mesh(nodes, elements, n_partitions)
|
||||
"""Partition mesh by nodes"""
|
||||
n_nodes = length(nodes)
|
||||
nodes_per_partition = ceil(Int, n_nodes / n_partitions)
|
||||
|
||||
node_to_elems = build_node_to_elements(nodes, elements)
|
||||
|
||||
partitions = Partition[]
|
||||
|
||||
for rank in 0:(n_partitions-1)
|
||||
# Owned nodes
|
||||
start_node = rank * nodes_per_partition + 1
|
||||
end_node = min((rank + 1) * nodes_per_partition, n_nodes)
|
||||
owned_nodes = start_node:end_node
|
||||
|
||||
# Find elements touching owned nodes
|
||||
local_elements = Int[]
|
||||
ghost_nodes = Set{Int}()
|
||||
|
||||
for (elem_id, element) in enumerate(elements)
|
||||
if any(nid in owned_nodes for nid in element.connectivity)
|
||||
push!(local_elements, elem_id)
|
||||
|
||||
# Mark ghost nodes
|
||||
for nid in element.connectivity
|
||||
if !(nid in owned_nodes)
|
||||
push!(ghost_nodes, nid)
|
||||
end
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
# Build local node_to_elements (owned nodes only)
|
||||
local_node_to_elems = [node_to_elems[nid] for nid in owned_nodes]
|
||||
|
||||
# Find interface nodes (owned nodes that couple to other partitions)
|
||||
interface = Dict{Int,Vector{Int}}()
|
||||
for neighbor_rank in 0:(n_partitions-1)
|
||||
if neighbor_rank == rank
|
||||
continue
|
||||
end
|
||||
|
||||
neighbor_start = neighbor_rank * nodes_per_partition + 1
|
||||
neighbor_end = min((neighbor_rank + 1) * nodes_per_partition, n_nodes)
|
||||
neighbor_owned = neighbor_start:neighbor_end
|
||||
|
||||
interface_with_neighbor = Int[]
|
||||
for elem_id in local_elements
|
||||
element = elements[elem_id]
|
||||
has_owned = any(nid in owned_nodes for nid in element.connectivity)
|
||||
has_neighbor = any(nid in neighbor_owned for nid in element.connectivity)
|
||||
|
||||
if has_owned && has_neighbor
|
||||
for nid in element.connectivity
|
||||
if nid in owned_nodes && !(nid in interface_with_neighbor)
|
||||
push!(interface_with_neighbor, nid)
|
||||
end
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
if !isempty(interface_with_neighbor)
|
||||
interface[neighbor_rank] = interface_with_neighbor
|
||||
end
|
||||
end
|
||||
|
||||
partition = Partition(
|
||||
rank,
|
||||
owned_nodes,
|
||||
collect(ghost_nodes),
|
||||
local_elements,
|
||||
local_node_to_elems,
|
||||
interface
|
||||
)
|
||||
|
||||
push!(partitions, partition)
|
||||
end
|
||||
|
||||
return partitions
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Nodal Assembly (Matrix-Vector Product)
|
||||
# ============================================================================
|
||||
|
||||
function matvec_single_threaded!(y, x, nodes, elements, node_to_elements)
|
||||
"""Single-threaded nodal assembly"""
|
||||
fill!(y, 0.0)
|
||||
|
||||
for node_id in 1:length(nodes)
|
||||
node = nodes[node_id]
|
||||
|
||||
# Get node DOFs
|
||||
dof_start = (node_id - 1) * 3 + 1
|
||||
y_nodal = zeros(3)
|
||||
|
||||
# Gather from connected elements
|
||||
for elem_id in node_to_elements[node_id]
|
||||
element = elements[elem_id]
|
||||
|
||||
# Mock element contribution (just for timing)
|
||||
for nid in element.connectivity
|
||||
x_dof_start = (nid - 1) * 3 + 1
|
||||
for d in 1:3
|
||||
y_nodal[d] += 0.1 * x[x_dof_start+d-1]
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
# Write to global
|
||||
for d in 1:3
|
||||
y[dof_start+d-1] = y_nodal[d]
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
function matvec_multi_threaded!(y, x, nodes, elements, node_to_elements)
|
||||
"""Multi-threaded nodal assembly (direct, no partitioning)"""
|
||||
fill!(y, 0.0)
|
||||
|
||||
@threads for node_id in 1:length(nodes)
|
||||
node = nodes[node_id]
|
||||
|
||||
dof_start = (node_id - 1) * 3 + 1
|
||||
y_nodal = zeros(3)
|
||||
|
||||
for elem_id in node_to_elements[node_id]
|
||||
element = elements[elem_id]
|
||||
|
||||
for nid in element.connectivity
|
||||
x_dof_start = (nid - 1) * 3 + 1
|
||||
for d in 1:3
|
||||
y_nodal[d] += 0.1 * x[x_dof_start+d-1]
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
for d in 1:3
|
||||
y[dof_start+d-1] = y_nodal[d]
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
function matvec_partitioned!(y, x, nodes, elements, partitions)
|
||||
"""Multi-threaded with explicit partitioning (simulates multi-GPU)"""
|
||||
fill!(y, 0.0)
|
||||
|
||||
# Each partition processed by one thread
|
||||
@threads for partition in partitions
|
||||
# Process owned nodes
|
||||
for (local_idx, node_id) in enumerate(partition.owned_nodes)
|
||||
node = nodes[node_id]
|
||||
|
||||
dof_start = (node_id - 1) * 3 + 1
|
||||
y_nodal = zeros(3)
|
||||
|
||||
for elem_id in partition.node_to_elements[local_idx]
|
||||
element = elements[elem_id]
|
||||
|
||||
for nid in element.connectivity
|
||||
x_dof_start = (nid - 1) * 3 + 1
|
||||
for d in 1:3
|
||||
y_nodal[d] += 0.1 * x[x_dof_start+d-1]
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
for d in 1:3
|
||||
y[dof_start+d-1] = y_nodal[d]
|
||||
end
|
||||
end
|
||||
end
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Benchmarks
|
||||
# ============================================================================
|
||||
|
||||
function run_benchmark(name, nx, ny, nz, n_warmup=2, n_runs=10)
|
||||
println("\n" * "="^70)
|
||||
println("Benchmark: $name")
|
||||
println(" Mesh: $nx × $ny × $nz = $(nx*ny*nz) nodes, $((nx-1)*(ny-1)*(nz-1)) elements")
|
||||
println("="^70)
|
||||
|
||||
# Create mesh
|
||||
print("Creating mesh... ")
|
||||
nodes, elements = create_hex_mesh(nx, ny, nz)
|
||||
node_to_elements = build_node_to_elements(nodes, elements)
|
||||
n_dofs = 3 * length(nodes)
|
||||
println("✓")
|
||||
|
||||
println(" Nodes: ", length(nodes))
|
||||
println(" Elements: ", length(elements))
|
||||
println(" DOFs: ", n_dofs)
|
||||
println(" Avg elements/node: ", sum(length.(node_to_elements)) / length(nodes))
|
||||
|
||||
# Test vectors
|
||||
x = randn(n_dofs)
|
||||
y_ref = zeros(n_dofs)
|
||||
y_test = zeros(n_dofs)
|
||||
|
||||
# ========================================================================
|
||||
# 1. Single-threaded baseline
|
||||
# ========================================================================
|
||||
println("\n1. Single-threaded baseline:")
|
||||
|
||||
# Warmup
|
||||
for _ in 1:n_warmup
|
||||
matvec_single_threaded!(y_ref, x, nodes, elements, node_to_elements)
|
||||
end
|
||||
|
||||
# Benchmark
|
||||
times = Float64[]
|
||||
for _ in 1:n_runs
|
||||
t_start = time_ns()
|
||||
matvec_single_threaded!(y_ref, x, nodes, elements, node_to_elements)
|
||||
t_end = time_ns()
|
||||
push!(times, (t_end - t_start) / 1e9)
|
||||
end
|
||||
|
||||
t_single = minimum(times)
|
||||
println(" Time: ", @sprintf("%.3f ms", t_single * 1000))
|
||||
println(" Throughput: ", @sprintf("%.2f Mnodes/s", length(nodes) / t_single / 1e6))
|
||||
|
||||
# ========================================================================
|
||||
# 2. Multi-threaded (all available threads)
|
||||
# ========================================================================
|
||||
if nthreads() > 1
|
||||
println("\n2. Multi-threaded ($(nthreads()) threads):")
|
||||
|
||||
# Warmup
|
||||
for _ in 1:n_warmup
|
||||
matvec_multi_threaded!(y_test, x, nodes, elements, node_to_elements)
|
||||
end
|
||||
|
||||
# Verify correctness
|
||||
error = norm(y_test - y_ref) / norm(y_ref)
|
||||
println(" Verification: ", error < 1e-10 ? "✓ PASS" : "✗ FAIL (error=$error)")
|
||||
|
||||
# Benchmark
|
||||
times = Float64[]
|
||||
for _ in 1:n_runs
|
||||
t_start = time_ns()
|
||||
matvec_multi_threaded!(y_test, x, nodes, elements, node_to_elements)
|
||||
t_end = time_ns()
|
||||
push!(times, (t_end - t_start) / 1e9)
|
||||
end
|
||||
|
||||
t_multi = minimum(times)
|
||||
speedup = t_single / t_multi
|
||||
efficiency = speedup / nthreads() * 100
|
||||
|
||||
println(" Time: ", @sprintf("%.3f ms", t_multi * 1000))
|
||||
println(" Speedup: ", @sprintf("%.2fx", speedup))
|
||||
println(" Efficiency: ", @sprintf("%.1f%%", efficiency))
|
||||
println(" Throughput: ", @sprintf("%.2f Mnodes/s", length(nodes) / t_multi / 1e6))
|
||||
end
|
||||
|
||||
# ========================================================================
|
||||
# 3. Partitioned (simulates multi-GPU)
|
||||
# ========================================================================
|
||||
if nthreads() >= 4
|
||||
n_partitions = 4
|
||||
println("\n3. Partitioned ($n_partitions partitions, 4 threads):")
|
||||
|
||||
# Create partitions
|
||||
print(" Creating partitions... ")
|
||||
partitions = partition_mesh(nodes, elements, n_partitions)
|
||||
println("✓")
|
||||
|
||||
# Print partition info
|
||||
for partition in partitions
|
||||
n_owned = length(partition.owned_nodes)
|
||||
n_ghost = length(partition.ghost_nodes)
|
||||
n_interface = sum(length.(values(partition.interface_nodes)))
|
||||
interface_pct = n_interface / n_owned * 100
|
||||
|
||||
println(" Partition $(partition.rank): $n_owned owned, $n_ghost ghost, " *
|
||||
"$n_interface interface ($(round(interface_pct, digits=1))%)")
|
||||
end
|
||||
|
||||
# Warmup
|
||||
for _ in 1:n_warmup
|
||||
matvec_partitioned!(y_test, x, nodes, elements, partitions)
|
||||
end
|
||||
|
||||
# Verify correctness
|
||||
error = norm(y_test - y_ref) / norm(y_ref)
|
||||
println(" Verification: ", error < 1e-10 ? "✓ PASS" : "✗ FAIL (error=$error)")
|
||||
|
||||
# Benchmark
|
||||
times = Float64[]
|
||||
for _ in 1:n_runs
|
||||
t_start = time_ns()
|
||||
matvec_partitioned!(y_test, x, nodes, elements, partitions)
|
||||
t_end = time_ns()
|
||||
push!(times, (t_end - t_start) / 1e9)
|
||||
end
|
||||
|
||||
t_partitioned = minimum(times)
|
||||
speedup = t_single / t_partitioned
|
||||
efficiency = speedup / n_partitions * 100
|
||||
|
||||
println(" Time: ", @sprintf("%.3f ms", t_partitioned * 1000))
|
||||
println(" Speedup: ", @sprintf("%.2fx", speedup))
|
||||
println(" Efficiency: ", @sprintf("%.1f%%", efficiency))
|
||||
println(" Throughput: ", @sprintf("%.2f Mnodes/s", length(nodes) / t_partitioned / 1e6))
|
||||
end
|
||||
|
||||
println()
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# Run Benchmarks
|
||||
# ============================================================================
|
||||
|
||||
# Small mesh
|
||||
run_benchmark("Small Mesh", 20, 20, 20)
|
||||
|
||||
# Medium mesh
|
||||
run_benchmark("Medium Mesh", 40, 40, 40)
|
||||
|
||||
# Large mesh (if enough threads)
|
||||
if nthreads() >= 4
|
||||
run_benchmark("Large Mesh", 60, 60, 60)
|
||||
end
|
||||
|
||||
println("="^70)
|
||||
println("✓ Benchmark Complete")
|
||||
println("="^70)
|
||||
println()
|
||||
println("Notes:")
|
||||
println(" - Single-threaded: Baseline performance")
|
||||
println(" - Multi-threaded: All available threads, direct parallelization")
|
||||
println(" - Partitioned: Simulates multi-GPU with explicit partitions")
|
||||
println(" - Efficiency = Speedup / N_threads × 100%")
|
||||
println(" - Near 100% efficiency = perfect scaling")
|
||||
println()
|
||||
@@ -1,320 +0,0 @@
|
||||
"""
|
||||
Performance Analysis: PerfectPlasticity Material Model
|
||||
|
||||
Comprehensive benchmarking of J2 plasticity implementation with radial return mapping.
|
||||
|
||||
Tests:
|
||||
1. Single evaluation performance (elastic vs plastic)
|
||||
2. Zero-allocation verification
|
||||
3. Type stability verification
|
||||
4. State overhead measurement
|
||||
5. Hardening parameter sensitivity
|
||||
6. Assembly loop simulation
|
||||
7. Comparison to LinearElastic and NeoHookean
|
||||
8. Strain level scalability
|
||||
|
||||
Run with:
|
||||
julia --project=. benchmarks/perfect_plasticity_analysis.jl
|
||||
"""
|
||||
|
||||
using BenchmarkTools
|
||||
using Tensors
|
||||
using Statistics
|
||||
using Printf
|
||||
using Dates
|
||||
|
||||
# Load implementations
|
||||
include("../src/materials/abstract_material.jl")
|
||||
include("../src/materials/linear_elastic.jl")
|
||||
include("../src/materials/neo_hookean.jl")
|
||||
include("../src/materials/perfect_plasticity.jl")
|
||||
|
||||
println("="^80)
|
||||
println("PERFECT PLASTICITY MATERIAL - PERFORMANCE ANALYSIS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 1: Single Evaluation - Elastic Path
|
||||
# ==============================================================================
|
||||
println("TEST 1: Single Evaluation - Elastic Path")
|
||||
println("-"^80)
|
||||
|
||||
steel = PerfectPlasticity(E=200e9, ν=0.3, σ_y=250e6, H=1e9)
|
||||
ε_elastic = SymmetricTensor{2,3}((1e-5, 0.0, 0.0, 0.0, 0.0, 0.0)) # Below yield
|
||||
state = PlasticityState()
|
||||
|
||||
# Benchmark elastic path
|
||||
bench_elastic = @benchmark compute_stress($steel, $ε_elastic, $state, 0.0)
|
||||
t_elastic = median(bench_elastic.times)
|
||||
allocs_elastic = bench_elastic.allocs
|
||||
|
||||
println("Elastic path (no yielding):")
|
||||
println(" Time: ", @sprintf("%.2f ns", t_elastic))
|
||||
println(" Allocations: ", allocs_elastic)
|
||||
println(" Memory: ", bench_elastic.memory, " bytes")
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 2: Single Evaluation - Plastic Path
|
||||
# ==============================================================================
|
||||
println("TEST 2: Single Evaluation - Plastic Path")
|
||||
println("-"^80)
|
||||
|
||||
ε_plastic = SymmetricTensor{2,3}((0.003, 0.0, 0.0, 0.0, 0.0, 0.0)) # Beyond yield
|
||||
|
||||
# Benchmark plastic path
|
||||
bench_plastic = @benchmark compute_stress($steel, $ε_plastic, $state, 0.0)
|
||||
t_plastic = median(bench_plastic.times)
|
||||
allocs_plastic = bench_plastic.allocs
|
||||
|
||||
println("Plastic path (radial return):")
|
||||
println(" Time: ", @sprintf("%.2f ns", t_plastic))
|
||||
println(" Allocations: ", allocs_plastic)
|
||||
println(" Memory: ", bench_plastic.memory, " bytes")
|
||||
println()
|
||||
|
||||
println("Plastic overhead:")
|
||||
println(" Ratio: ", @sprintf("%.2fx", t_plastic / t_elastic))
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 3: State Management Overhead
|
||||
# ==============================================================================
|
||||
println("TEST 3: State Management Overhead")
|
||||
println("-"^80)
|
||||
|
||||
# Compare with and without state history
|
||||
σ1, 𝔻1, state1 = compute_stress(steel, ε_plastic, nothing, 0.0) # Fresh state
|
||||
σ2, 𝔻2, state2 = compute_stress(steel, ε_plastic, state1, 0.0) # With history
|
||||
|
||||
bench_fresh = @benchmark compute_stress($steel, $ε_plastic, nothing, 0.0)
|
||||
bench_history = @benchmark compute_stress($steel, $ε_plastic, $state1, 0.0)
|
||||
|
||||
println("Fresh state (ε_p = 0, α = 0):")
|
||||
println(" Time: ", @sprintf("%.2f ns", median(bench_fresh.times)))
|
||||
println()
|
||||
println("With history (ε_p ≠ 0, α ≠ 0):")
|
||||
println(" Time: ", @sprintf("%.2f ns", median(bench_history.times)))
|
||||
println()
|
||||
println("State overhead: ", @sprintf("%.1f%%",
|
||||
(median(bench_history.times) - median(bench_fresh.times)) / median(bench_fresh.times) * 100))
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 4: Hardening Parameter Sensitivity
|
||||
# ==============================================================================
|
||||
println("TEST 4: Hardening Parameter Sensitivity")
|
||||
println("-"^80)
|
||||
|
||||
hardening_values = [0.0, 1e8, 1e9, 10e9, 100e9] # Perfect to strong hardening
|
||||
times_H = Float64[]
|
||||
|
||||
for H in hardening_values
|
||||
mat = PerfectPlasticity(E=200e9, ν=0.3, σ_y=250e6, H=H)
|
||||
bench = @benchmark compute_stress($mat, $ε_plastic, $state, 0.0)
|
||||
push!(times_H, median(bench.times))
|
||||
end
|
||||
|
||||
println("H (Pa) Time (ns) Overhead")
|
||||
println(repeat("-", 45))
|
||||
for (H, t) in zip(hardening_values, times_H)
|
||||
overhead = (t - times_H[1]) / times_H[1] * 100
|
||||
println(@sprintf("%-15.1e %8.2f %+6.1f%%", H, t, overhead))
|
||||
end
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 5: Comparison to Other Materials
|
||||
# ==============================================================================
|
||||
println("TEST 5: Comparison to Other Materials")
|
||||
println("-"^80)
|
||||
|
||||
# LinearElastic
|
||||
linear = LinearElastic(E=200e9, ν=0.3)
|
||||
bench_linear = @benchmark compute_stress($linear, $ε_plastic)
|
||||
t_linear = median(bench_linear.times)
|
||||
|
||||
# NeoHookean (uses Green-Lagrange strain for small deformation)
|
||||
μ_neo = 200e9 / (2 * (1 + 0.3))
|
||||
λ_neo = 200e9 * 0.3 / ((1 + 0.3) * (1 - 2 * 0.3))
|
||||
neo = NeoHookean(μ=μ_neo, λ=λ_neo)
|
||||
E_gl = ε_plastic # For small strains, E_GL ≈ ε
|
||||
bench_neo = @benchmark compute_stress($neo, $E_gl)
|
||||
t_neo = median(bench_neo.times)
|
||||
|
||||
println("Material Time (ns) Ratio vs Linear")
|
||||
println(repeat("-", 50))
|
||||
println(@sprintf("LinearElastic %8.2f 1.00x (baseline)", t_linear))
|
||||
println(@sprintf("PerfectPlasticity %8.2f %.2fx", t_plastic, t_plastic / t_linear))
|
||||
println(@sprintf("NeoHookean %8.2f %.2fx", t_neo, t_neo / t_linear))
|
||||
println()
|
||||
|
||||
println("Performance ranking:")
|
||||
println(" 1. LinearElastic (fastest, no state, manual derivatives)")
|
||||
println(" 2. PerfectPlasticity (", @sprintf("%.1fx", t_plastic / t_linear),
|
||||
" - state management + radial return)")
|
||||
println(" 3. NeoHookean (", @sprintf("%.1fx", t_neo / t_linear),
|
||||
" - automatic differentiation overhead)")
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 6: Assembly Loop Simulation
|
||||
# ==============================================================================
|
||||
println("TEST 6: Assembly Loop Simulation (1000 Gauss points)")
|
||||
println("-"^80)
|
||||
|
||||
n_gauss = 1000
|
||||
strains = [SymmetricTensor{2,3}((0.001 + 0.002 * rand(), 0.0, 0.0, 0.0, 0.0, 0.0))
|
||||
for _ in 1:n_gauss]
|
||||
|
||||
# Elastic assembly
|
||||
function assembly_elastic(material, strains)
|
||||
total = zero(SymmetricTensor{2,3})
|
||||
for ε in strains
|
||||
σ, _, _ = compute_stress(material, ε)
|
||||
total += σ
|
||||
end
|
||||
return total
|
||||
end
|
||||
|
||||
# Plastic assembly (stateful)
|
||||
function assembly_plastic(material, strains, state)
|
||||
total = zero(SymmetricTensor{2,3})
|
||||
for ε in strains
|
||||
σ, _, state = compute_stress(material, ε, state, 0.0)
|
||||
total += σ
|
||||
end
|
||||
return total, state
|
||||
end
|
||||
|
||||
bench_asm_linear = @benchmark assembly_elastic($linear, $strains)
|
||||
bench_asm_plastic = @benchmark assembly_plastic($steel, $strains, $state)
|
||||
|
||||
t_asm_linear = median(bench_asm_linear.times) / 1e6 # Convert to ms
|
||||
t_asm_plastic = median(bench_asm_plastic.times) / 1e6
|
||||
|
||||
println("LinearElastic assembly: ", @sprintf("%.3f ms", t_asm_linear))
|
||||
println("PerfectPlasticity assembly: ", @sprintf("%.3f ms", t_asm_plastic))
|
||||
println("Overhead: ", @sprintf("%.2fx", t_asm_plastic / t_asm_linear))
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 7: Strain Level Scalability
|
||||
# ==============================================================================
|
||||
println("TEST 7: Strain Level Scalability")
|
||||
println("-"^80)
|
||||
|
||||
strain_magnitudes = [0.0005, 0.001, 0.002, 0.005, 0.01, 0.02]
|
||||
times_strain = Float64[]
|
||||
yields = Bool[]
|
||||
|
||||
for ε_mag in strain_magnitudes
|
||||
ε_test = SymmetricTensor{2,3}((ε_mag, 0.0, 0.0, 0.0, 0.0, 0.0))
|
||||
σ_test, _, state_test = compute_stress(steel, ε_test, nothing, 0.0)
|
||||
|
||||
bench = @benchmark compute_stress($steel, $ε_test, nothing, 0.0)
|
||||
push!(times_strain, median(bench.times))
|
||||
push!(yields, state_test.κ > 0.0)
|
||||
end
|
||||
|
||||
println("ε_magnitude Time (ns) Yielded?")
|
||||
println(repeat("-", 40))
|
||||
for (ε_mag, t, y) in zip(strain_magnitudes, times_strain, yields)
|
||||
status = y ? "YES" : "no"
|
||||
println(@sprintf("%.4f %8.2f %s", ε_mag, t, status))
|
||||
end
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 8: Cyclic Loading Performance
|
||||
# ==============================================================================
|
||||
println("TEST 8: Cyclic Loading (Bauschinger Effect)")
|
||||
println("-"^80)
|
||||
|
||||
# Simulate cyclic loading path
|
||||
ε_cycle = [
|
||||
SymmetricTensor{2,3}((0.003, 0.0, 0.0, 0.0, 0.0, 0.0)), # Tension
|
||||
SymmetricTensor{2,3}((0.0, 0.0, 0.0, 0.0, 0.0, 0.0)), # Unload
|
||||
SymmetricTensor{2,3}((-0.002, 0.0, 0.0, 0.0, 0.0, 0.0)), # Compression
|
||||
SymmetricTensor{2,3}((0.0, 0.0, 0.0, 0.0, 0.0, 0.0)), # Unload
|
||||
]
|
||||
|
||||
function cyclic_loading(material, strains, state)
|
||||
for ε in strains
|
||||
σ, 𝔻, state = compute_stress(material, ε, state, 0.0)
|
||||
end
|
||||
return state
|
||||
end
|
||||
|
||||
bench_cyclic = @benchmark cyclic_loading($steel, $ε_cycle, $state)
|
||||
t_cyclic = median(bench_cyclic.times)
|
||||
|
||||
println("Cyclic loading (4 load steps):")
|
||||
println(" Total time: ", @sprintf("%.2f ns", t_cyclic))
|
||||
println(" Per load step: ", @sprintf("%.2f ns", t_cyclic / 4))
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# TEST 9: Type Stability Verification
|
||||
# ==============================================================================
|
||||
println("TEST 9: Type Stability")
|
||||
println("-"^80)
|
||||
|
||||
using InteractiveUtils
|
||||
|
||||
println("Return type inference:")
|
||||
result_type = @code_typed compute_stress(steel, ε_plastic, state, 0.0)
|
||||
println(" ✓ Type stable: ", result_type[2])
|
||||
println()
|
||||
|
||||
# ==============================================================================
|
||||
# SUMMARY
|
||||
# ==============================================================================
|
||||
println("="^80)
|
||||
println("SUMMARY")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
println("Performance Characteristics:")
|
||||
println(" • Elastic path: ", @sprintf("%.0f ns", t_elastic), " (no allocations)")
|
||||
println(" • Plastic path: ", @sprintf("%.0f ns", t_plastic), " (~128 bytes for state)")
|
||||
println(" • Plastic overhead:", @sprintf("%.2fx", t_plastic / t_elastic))
|
||||
println()
|
||||
|
||||
println("Comparison to other materials:")
|
||||
println(" • ", @sprintf("%.2fx", t_plastic / t_linear), " slower than LinearElastic (baseline)")
|
||||
println(" • ", @sprintf("%.2fx", t_neo / t_plastic), " faster than NeoHookean (AD)")
|
||||
println()
|
||||
|
||||
println("Key findings:")
|
||||
println(" ✓ Zero allocations on elastic path")
|
||||
println(" ✓ Minimal allocations on plastic path (state struct only)")
|
||||
println(" ✓ Type stable")
|
||||
println(" ✓ Hardening parameter has negligible performance impact")
|
||||
println(" ✓ Performance independent of strain level")
|
||||
println(" ✓ Suitable for production FEM with ~",
|
||||
@sprintf("%.0f", 1e9 / t_plastic), " evaluations/second")
|
||||
println()
|
||||
|
||||
println("Recommendations:")
|
||||
if t_plastic < 500
|
||||
println(" ✓ Excellent performance - suitable for all applications")
|
||||
elseif t_plastic < 1000
|
||||
println(" ✓ Good performance - suitable for most applications")
|
||||
println(" • Consider caching for problems with >10M DOF")
|
||||
else
|
||||
println(" ⚠ Acceptable performance - profile before using with >1M DOF")
|
||||
println(" • Consider precomputation for repeated analyses")
|
||||
end
|
||||
println()
|
||||
|
||||
println("Expected performance in FEM assembly:")
|
||||
println(" • Small problems (<10K DOF): Negligible overhead")
|
||||
println(" • Medium problems (10K-1M DOF): ", @sprintf("<%.1f seconds", 1e6 * t_plastic / 1e9))
|
||||
println(" • Large problems (>1M DOF): ", @sprintf("<%.1f seconds", 1e7 * t_plastic / 1e9))
|
||||
println()
|
||||
|
||||
println("="^80)
|
||||
println("Analysis complete: ", now())
|
||||
println("="^80)
|
||||
@@ -1,348 +0,0 @@
|
||||
Precompiling packages...
|
||||
669.3 ms ✓ EpollShim_jll
|
||||
725.5 ms ✓ libfdk_aac_jll
|
||||
760.9 ms ✓ Graphite2_jll
|
||||
752.3 ms ✓ fzf_jll
|
||||
756.4 ms ✓ LERC_jll
|
||||
822.1 ms ✓ Xorg_libICE_jll
|
||||
815.2 ms ✓ LAME_jll
|
||||
797.9 ms ✓ Ogg_jll
|
||||
788.9 ms ✓ x265_jll
|
||||
829.1 ms ✓ libaom_jll
|
||||
849.2 ms ✓ mtdev_jll
|
||||
843.9 ms ✓ MbedTLS_jll
|
||||
840.0 ms ✓ x264_jll
|
||||
860.8 ms ✓ XZ_jll
|
||||
674.8 ms ✓ libevdev_jll
|
||||
769.9 ms ✓ Opus_jll
|
||||
743.6 ms ✓ eudev_jll
|
||||
738.4 ms ✓ FriBidi_jll
|
||||
707.7 ms ✓ Dbus_jll
|
||||
730.8 ms ✓ Xorg_libxkbfile_jll
|
||||
779.8 ms ✓ Xorg_xcb_util_jll
|
||||
770.5 ms ✓ Xorg_libXi_jll
|
||||
772.1 ms ✓ Xorg_libXrandr_jll
|
||||
793.7 ms ✓ Xorg_libXcursor_jll
|
||||
892.3 ms ✓ Wayland_jll
|
||||
2016.1 ms ✓ ColorVectorSpace
|
||||
1008.1 ms ✓ HarfBuzz_jll
|
||||
711.6 ms ✓ JLFzf
|
||||
1304.3 ms ✓ Ghostscript_jll
|
||||
737.4 ms ✓ Xorg_libSM_jll
|
||||
1455.9 ms ✓ RecipesBase
|
||||
803.4 ms ✓ libvorbis_jll
|
||||
731.4 ms ✓ libinput_jll
|
||||
785.6 ms ✓ Libtiff_jll
|
||||
742.5 ms ✓ Xorg_xkbcomp_jll
|
||||
731.2 ms ✓ Xorg_xcb_util_image_jll
|
||||
733.0 ms ✓ Xorg_xcb_util_keysyms_jll
|
||||
741.3 ms ✓ Xorg_xcb_util_renderutil_jll
|
||||
2530.6 ms ✓ StatsBase
|
||||
764.4 ms ✓ Xorg_xcb_util_wm_jll
|
||||
1120.5 ms ✓ MbedTLS
|
||||
879.4 ms ✓ libass_jll
|
||||
653.5 ms ✓ Xorg_xkeyboard_config_jll
|
||||
962.5 ms ✓ Pango_jll
|
||||
711.2 ms ✓ Xorg_xcb_util_cursor_jll
|
||||
758.4 ms ✓ xkbcommon_jll
|
||||
1332.1 ms ✓ FFMPEG_jll
|
||||
999.6 ms ✓ Vulkan_Loader_jll
|
||||
1149.2 ms ✓ libdecor_jll
|
||||
839.5 ms ✓ FFMPEG
|
||||
2960.5 ms ✓ Latexify
|
||||
974.9 ms ✓ GLFW_jll
|
||||
4035.5 ms ✓ ColorSchemes
|
||||
852.9 ms ✓ Latexify → SparseArraysExt
|
||||
1490.1 ms ✓ Qt6Base_jll
|
||||
918.9 ms ✓ Qt6ShaderTools_jll
|
||||
1397.3 ms ✓ GR_jll
|
||||
2814.7 ms ✓ Qt6Declarative_jll
|
||||
1317.4 ms ✓ Qt6Wayland_jll
|
||||
7309.2 ms ✓ PlotUtils
|
||||
12361.6 ms ✓ HTTP
|
||||
2980.3 ms ✓ PlotThemes
|
||||
3588.3 ms ✓ RecipesPipeline
|
||||
5614.7 ms ✓ GR
|
||||
64840.3 ms ✓ Plots
|
||||
65 dependencies successfully precompiled in 87 seconds. 112 already precompiled.
|
||||
================================================================================
|
||||
SYSTEM INFORMATION
|
||||
================================================================================
|
||||
|
||||
CPU Model: Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz
|
||||
CPU Cores: 32 threads (32 physical cores)
|
||||
CPU Speed: 3300 MHz
|
||||
|
||||
Julia Version: 1.12.1
|
||||
OS: Linux x86_64-linux-gnu
|
||||
Word Size: 64 bits
|
||||
|
||||
Approximate CPU Cache Sizes:
|
||||
L1 Cache: ~32-64 KB per core (typical)
|
||||
L2 Cache: ~256-512 KB per core (typical)
|
||||
L3 Cache: ~8-32 MB shared (typical)
|
||||
|
||||
Note: Testing up to 8KB structs to exceed L1 cache
|
||||
|
||||
================================================================================
|
||||
SYSTEM INFORMATION
|
||||
================================================================================
|
||||
Julia Version: 1.12.1
|
||||
CPU Model: Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz
|
||||
CPU Cores: 32
|
||||
Total Memory: 503.35 GB
|
||||
L1 Cache: 48K
|
||||
L2 Cache: 1280K
|
||||
L3 Cache: 24576K
|
||||
|
||||
================================================================================
|
||||
STRUCT SIZE SCALING BENCHMARK
|
||||
================================================================================
|
||||
|
||||
Testing hypothesis: Immutable slows down with struct size, mutable stays constant
|
||||
|
||||
Testing struct with 1 Float64 fields (8 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.02ns Speedup=2.3x
|
||||
Update: Mut=7.2ns Imm=2.32ns Speedup=3.1x
|
||||
Iterate: Mut=11.77ns Imm=2.02ns Speedup=5.8x
|
||||
Copy: 2.02ns
|
||||
|
||||
Testing struct with 2 Float64 fields (16 bytes)...
|
||||
Access: Mut=4.9ns Imm=2.02ns Speedup=2.4x
|
||||
Update: Mut=7.21ns Imm=2.31ns Speedup=3.1x
|
||||
Iterate: Mut=14.49ns Imm=2.02ns Speedup=7.2x
|
||||
Copy: 2.03ns
|
||||
|
||||
Testing struct with 5 Float64 fields (40 bytes)...
|
||||
Access: Mut=4.9ns Imm=2.37ns Speedup=2.1x
|
||||
Update: Mut=7.2ns Imm=2.6ns Speedup=2.8x
|
||||
Iterate: Mut=16.71ns Imm=2.31ns Speedup=7.2x
|
||||
Copy: 2.38ns
|
||||
|
||||
Testing struct with 10 Float64 fields (80 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.03ns Speedup=2.3x
|
||||
Update: Mut=7.2ns Imm=2.6ns Speedup=2.8x
|
||||
Iterate: Mut=23.13ns Imm=2.6ns Speedup=8.9x
|
||||
Copy: 3.31ns
|
||||
|
||||
Testing struct with 20 Float64 fields (160 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.02ns Speedup=2.3x
|
||||
Update: Mut=7.2ns Imm=3.16ns Speedup=2.3x
|
||||
Iterate: Mut=63.13ns Imm=5.52ns Speedup=11.4x
|
||||
Copy: 3.18ns
|
||||
|
||||
Testing struct with 50 Float64 fields (400 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.03ns Speedup=2.3x
|
||||
Update: Mut=7.2ns Imm=6.02ns Speedup=1.2x
|
||||
Iterate: Mut=196.85ns Imm=24.78ns Speedup=7.9x
|
||||
Copy: 6.9ns
|
||||
|
||||
Testing struct with 100 Float64 fields (800 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.02ns Speedup=2.3x
|
||||
Update: Mut=7.2ns Imm=11.75ns Speedup=0.6x
|
||||
Iterate: Mut=260.56ns Imm=68.5ns Speedup=3.8x
|
||||
Copy: 9.48ns
|
||||
|
||||
Testing struct with 200 Float64 fields (1600 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.37ns Speedup=1.9x
|
||||
Update: Mut=7.41ns Imm=24.63ns Speedup=0.3x
|
||||
Iterate: Mut=777.91ns Imm=155.16ns Speedup=5.0x
|
||||
Copy: 20.95ns
|
||||
|
||||
Testing struct with 500 Float64 fields (4000 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.03ns Speedup=2.3x
|
||||
Update: Mut=7.2ns Imm=80.84ns Speedup=0.1x
|
||||
Iterate: Mut=1156.3ns Imm=499.36ns Speedup=2.3x
|
||||
Copy: 58.68ns
|
||||
|
||||
Testing struct with 1000 Float64 fields (8000 bytes)...
|
||||
Access: Mut=4.67ns Imm=2.08ns Speedup=2.2x
|
||||
Update: Mut=7.2ns Imm=193.26ns Speedup=0.0x
|
||||
Iterate: Mut=3465.75ns Imm=1100.2ns Speedup=3.2x
|
||||
Copy: 45.7ns
|
||||
|
||||
Testing struct with 2000 Float64 fields (16000 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.02ns Speedup=2.3x
|
||||
Update: Mut=7.2ns Imm=410.85ns Speedup=0.0x
|
||||
Iterate: Mut=4811.71ns Imm=2226.22ns Speedup=2.2x
|
||||
Copy: 82.95ns
|
||||
|
||||
Testing struct with 5000 Float64 fields (40000 bytes)...
|
||||
Access: Mut=4.62ns Imm=2.08ns Speedup=2.2x
|
||||
Update: Mut=7.2ns Imm=1936.1ns Speedup=0.0x
|
||||
Iterate: Mut=33969.0ns Imm=5667.17ns Speedup=6.0x
|
||||
Copy: 867.04ns
|
||||
|
||||
================================================================================
|
||||
RESULTS SUMMARY
|
||||
================================================================================
|
||||
|
||||
Field Access Performance:
|
||||
Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup
|
||||
----------------------------------------------------------------------
|
||||
1 | 8 | 4.62 | 2.02 | 2.3x
|
||||
2 | 16 | 4.90 | 2.02 | 2.4x
|
||||
5 | 40 | 4.90 | 2.37 | 2.1x
|
||||
10 | 80 | 4.62 | 2.03 | 2.3x
|
||||
20 | 160 | 4.62 | 2.02 | 2.3x
|
||||
50 | 400 | 4.62 | 2.03 | 2.3x
|
||||
100 | 800 | 4.62 | 2.02 | 2.3x
|
||||
200 | 1600 | 4.62 | 2.37 | 1.9x
|
||||
500 | 4000 | 4.62 | 2.03 | 2.3x
|
||||
1000 | 8000 | 4.67 | 2.08 | 2.2x
|
||||
2000 | 16000 | 4.62 | 2.02 | 2.3x
|
||||
5000 | 40000 | 4.62 | 2.08 | 2.2x
|
||||
|
||||
Field Update Performance:
|
||||
Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup
|
||||
----------------------------------------------------------------------
|
||||
1 | 8 | 7.20 | 2.31 | 3.1x
|
||||
2 | 16 | 7.21 | 2.31 | 3.1x
|
||||
5 | 40 | 7.20 | 2.60 | 2.8x
|
||||
10 | 80 | 7.20 | 2.60 | 2.8x
|
||||
20 | 160 | 7.20 | 3.16 | 2.3x
|
||||
50 | 400 | 7.20 | 6.02 | 1.2x
|
||||
100 | 800 | 7.20 | 11.75 | 0.6x
|
||||
200 | 1600 | 7.41 | 24.63 | 0.3x
|
||||
500 | 4000 | 7.20 | 80.84 | 0.1x
|
||||
1000 | 8000 | 7.20 | 193.26 | 0.0x
|
||||
2000 | 16000 | 7.20 | 410.85 | 0.0x
|
||||
5000 | 40000 | 7.20 | 1936.10 | 0.0x
|
||||
|
||||
Iteration Performance:
|
||||
Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup
|
||||
----------------------------------------------------------------------
|
||||
1 | 8 | 11.77 | 2.02 | 5.8x
|
||||
2 | 16 | 14.49 | 2.02 | 7.2x
|
||||
5 | 40 | 16.71 | 2.31 | 7.2x
|
||||
10 | 80 | 23.13 | 2.60 | 8.9x
|
||||
20 | 160 | 63.13 | 5.52 | 11.4x
|
||||
50 | 400 | 196.85 | 24.78 | 7.9x
|
||||
100 | 800 | 260.56 | 68.50 | 3.8x
|
||||
200 | 1600 | 777.91 | 155.16 | 5.0x
|
||||
500 | 4000 | 1156.30 | 499.36 | 2.3x
|
||||
1000 | 8000 | 3465.75 | 1100.20 | 3.2x
|
||||
2000 | 16000 | 4811.71 | 2226.22 | 2.2x
|
||||
5000 | 40000 | 33969.00 | 5667.17 | 6.0x
|
||||
|
||||
Immutable Copy Cost (ns):
|
||||
Size (fields) | Bytes | Copy Time (ns)
|
||||
----------------------------------------
|
||||
1 | 8 | 2.02
|
||||
2 | 16 | 2.03
|
||||
5 | 40 | 2.38
|
||||
10 | 80 | 3.31
|
||||
20 | 160 | 3.18
|
||||
50 | 400 | 6.90
|
||||
100 | 800 | 9.48
|
||||
200 | 1600 | 20.95
|
||||
500 | 4000 | 58.68
|
||||
1000 | 8000 | 45.70
|
||||
2000 | 16000 | 82.95
|
||||
5000 | 40000 | 867.04
|
||||
|
||||
================================================================================
|
||||
ANALYSIS
|
||||
================================================================================
|
||||
|
||||
✓ Immutable ALWAYS faster for field access (even at 1000 fields = 8KB)
|
||||
Minimum speedup: 1.9x at 5000 fields
|
||||
|
||||
⚠ Mutable wins for field update at 100 fields
|
||||
|
||||
✓ Immutable ALWAYS faster for iteration (even at 1000 fields = 8KB)
|
||||
Minimum speedup: 2.2x at 5000 fields
|
||||
|
||||
Scaling Analysis:
|
||||
|
||||
Copy time scaling:
|
||||
Linear fit: time(ns) = -25.47 + 0.1587 * nfields
|
||||
Per-field cost: 0.1587 ns/field
|
||||
Base overhead: -25.47 ns
|
||||
|
||||
Is copy time linear? (checking R²)
|
||||
R² = 0.9006
|
||||
⚠ Copy time not perfectly linear (compiler optimizations?)
|
||||
|
||||
KEY INSIGHT:
|
||||
--------------------------------------------------------------------------------
|
||||
|
||||
Even at 1000 fields (8KB struct), immutable is STILL faster because:
|
||||
1. Dict lookup cost (~40-50ns) >> copy cost per field (~0.1587ns)
|
||||
2. Type stability enables compiler optimizations (inlining, SIMD)
|
||||
3. Stack allocation has better cache locality than heap pointers
|
||||
|
||||
Theoretical crossover point (if it exists):
|
||||
Would occur at ~413 fields (3KB)
|
||||
|
||||
================================================================================
|
||||
CONCLUSION
|
||||
================================================================================
|
||||
|
||||
Your intuition about O(n) scaling is CORRECT, BUT:
|
||||
|
||||
• Dict lookup base cost is SO high (~40-50ns)
|
||||
• Copy cost per field is SO low (~0.1587ns)
|
||||
• Compiler optimizations are SO good (inlining, SIMD, escape analysis)
|
||||
|
||||
That immutable wins even for unrealistically large structs (8KB+)!
|
||||
|
||||
For typical FEM elements:
|
||||
• Material properties: 3-10 fields (24-80 bytes)
|
||||
• State variables: 10-50 fields (80-400 bytes)
|
||||
• Even with 100 fields (800 bytes), immutable is >10x faster
|
||||
|
||||
Type stability > Everything else.
|
||||
|
||||
================================================================================
|
||||
SAVING DATA
|
||||
================================================================================
|
||||
|
||||
✓ Data saved to: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/struct_size_scaling.json
|
||||
✓ CSV saved to: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/struct_size_scaling.csv
|
||||
|
||||
================================================================================
|
||||
GENERATING PLOTS
|
||||
================================================================================
|
||||
|
||||
]1337;ReportCellSizeP+q544e\GKS: cannot open display - headless operation mode active
|
||||
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/field_access_scaling.png
|
||||
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/field_update_scaling.png
|
||||
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/iteration_scaling.png
|
||||
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/copy_cost_linear.png
|
||||
✓ Plot saved: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results/speedup_ratios.png
|
||||
|
||||
All plots saved to: /home/juajukka/dev/JuliaFEM.jl/benchmarks/results
|
||||
|
||||
================================================================================
|
||||
SAVING DATA
|
||||
================================================================================
|
||||
|
||||
✓ Data saved to: benchmarks/results/struct_scaling_20251109_201256.json
|
||||
|
||||
✓ CSV saved to: benchmarks/results/struct_scaling_20251109_201256.csv
|
||||
|
||||
================================================================================
|
||||
GENERATING PLOTS
|
||||
================================================================================
|
||||
|
||||
┌ Warning: Assignment to `p1` in soft scope is ambiguous because a global variable by the same name exists: `p1` will be treated as a new local. Disambiguate by using `local p1` to suppress this warning or `global p1` to assign to the existing global variable.
|
||||
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:633
|
||||
┌ Warning: Assignment to `p2` in soft scope is ambiguous because a global variable by the same name exists: `p2` will be treated as a new local. Disambiguate by using `local p2` to suppress this warning or `global p2` to assign to the existing global variable.
|
||||
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:649
|
||||
┌ Warning: Assignment to `p3` in soft scope is ambiguous because a global variable by the same name exists: `p3` will be treated as a new local. Disambiguate by using `local p3` to suppress this warning or `global p3` to assign to the existing global variable.
|
||||
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:664
|
||||
┌ Warning: Assignment to `p4` in soft scope is ambiguous because a global variable by the same name exists: `p4` will be treated as a new local. Disambiguate by using `local p4` to suppress this warning or `global p4` to assign to the existing global variable.
|
||||
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:679
|
||||
┌ Warning: Assignment to `p5` in soft scope is ambiguous because a global variable by the same name exists: `p5` will be treated as a new local. Disambiguate by using `local p5` to suppress this warning or `global p5` to assign to the existing global variable.
|
||||
└ @ ~/dev/JuliaFEM.jl/benchmarks/struct_size_scaling.jl:697
|
||||
✓ Saved: field_access_20251109_201256.png
|
||||
✓ Saved: field_update_20251109_201256.png
|
||||
✓ Saved: iteration_20251109_201256.png
|
||||
✓ Saved: speedup_factors_20251109_201256.png
|
||||
✓ Saved: copy_cost_20251109_201256.png
|
||||
✓ Saved: combined_20251109_201256.png
|
||||
|
||||
All plots saved successfully!
|
||||
|
||||
================================================================================
|
||||
|
Before Width: | Height: | Size: 153 KiB |
|
Before Width: | Height: | Size: 36 KiB |
|
Before Width: | Height: | Size: 53 KiB |
|
Before Width: | Height: | Size: 34 KiB |
|
Before Width: | Height: | Size: 46 KiB |
|
Before Width: | Height: | Size: 31 KiB |
|
Before Width: | Height: | Size: 56 KiB |
|
Before Width: | Height: | Size: 36 KiB |
|
Before Width: | Height: | Size: 44 KiB |
|
Before Width: | Height: | Size: 54 KiB |
|
Before Width: | Height: | Size: 54 KiB |
@@ -1,13 +0,0 @@
|
||||
nfields,bytes,mut_access_ns,imm_access_ns,speedup_access,mut_update_ns,imm_update_ns,speedup_update,mut_iter_ns,imm_iter_ns,speedup_iter,imm_copy_ns
|
||||
1,8,4.615,2.023,2.2812654473554126,7.1991991991991995,2.315,3.1098052696324836,11.773773773773774,2.024,5.817081904038426,2.023
|
||||
2,16,4.904,2.023,2.4241225902125554,7.207207207207207,2.314,3.11460985618289,14.48997995991984,2.025,7.1555456592196744,2.029
|
||||
5,40,4.897,2.373,2.063632532659081,7.201201201201201,2.604,2.765438249309217,16.70941883767535,2.312,7.227257282731553,2.383
|
||||
10,80,4.621,2.029,2.2774765894529327,7.1991991991991995,2.602,2.7667944654877785,23.13152610441767,2.602,8.889902422912249,3.311
|
||||
20,160,4.618,2.023,2.282748393475037,7.201201201201201,3.1633266533066133,2.2764646179280326,63.131632653061224,5.523,11.43067764857165,3.177
|
||||
50,400,4.617,2.029,2.275505174963036,7.197197197197197,6.022,1.1951506471599465,196.84902597402598,24.783132530120483,7.942862982909166,6.895895895895896
|
||||
100,800,4.618,2.025,2.280493827160494,7.198198198198198,11.745745745745745,0.612834498039884,260.55786350148367,68.50307377049181,3.803593753682347,9.476476476476476
|
||||
200,1600,4.621,2.373,1.9473240623683101,7.41041041041041,24.633534136546185,0.30082611651798524,777.9142857142857,155.1606475716065,5.013605562294905,20.948897795591183
|
||||
500,4000,4.624,2.027,2.281203749383325,7.198198198198198,80.83854166666667,0.08904413723690831,1156.3,499.35567010309273,2.3155840000000003,58.68463886063072
|
||||
1000,8000,4.671,2.084,2.241362763915547,7.2002002002002,193.26257861635222,0.03725604952469045,3465.75,1100.2,3.1501090710779853,45.70171890798787
|
||||
2000,16000,4.625,2.024,2.2850790513833994,7.197197197197197,410.8542713567839,0.0175176399491468,4811.714285714285,2226.222222222222,2.1613809428742545,82.94813278008299
|
||||
5000,40000,4.623,2.083,2.2193951032165145,7.2042042042042045,1936.1,0.0037209876577677828,33969.0,5667.166666666667,5.994000529365056,867.0408163265306
|
||||
|
@@ -1,187 +0,0 @@
|
||||
{
|
||||
"julia_version": "1.12.1",
|
||||
"analysis": {
|
||||
"copy_intercept_ns": -25.467962218157183,
|
||||
"copy_slope_ns_per_field": 0.1586673181436861,
|
||||
"r_squared": 0.9005527188987935
|
||||
},
|
||||
"results": [
|
||||
{
|
||||
"mutable_update_ns": 7.1991991991991995,
|
||||
"immutable_copy_ns": 2.023,
|
||||
"mutable_access_ns": 4.615,
|
||||
"bytes": 8,
|
||||
"speedup_update": 3.1098052696324836,
|
||||
"speedup_iter": 5.817081904038426,
|
||||
"immutable_access_ns": 2.023,
|
||||
"speedup_access": 2.2812654473554126,
|
||||
"nfields": 1,
|
||||
"immutable_update_ns": 2.315,
|
||||
"mutable_iter_ns": 11.773773773773774,
|
||||
"immutable_iter_ns": 2.024
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.207207207207207,
|
||||
"immutable_copy_ns": 2.029,
|
||||
"mutable_access_ns": 4.904,
|
||||
"bytes": 16,
|
||||
"speedup_update": 3.11460985618289,
|
||||
"speedup_iter": 7.1555456592196744,
|
||||
"immutable_access_ns": 2.023,
|
||||
"speedup_access": 2.4241225902125554,
|
||||
"nfields": 2,
|
||||
"immutable_update_ns": 2.314,
|
||||
"mutable_iter_ns": 14.48997995991984,
|
||||
"immutable_iter_ns": 2.025
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.201201201201201,
|
||||
"immutable_copy_ns": 2.383,
|
||||
"mutable_access_ns": 4.897,
|
||||
"bytes": 40,
|
||||
"speedup_update": 2.765438249309217,
|
||||
"speedup_iter": 7.227257282731553,
|
||||
"immutable_access_ns": 2.373,
|
||||
"speedup_access": 2.063632532659081,
|
||||
"nfields": 5,
|
||||
"immutable_update_ns": 2.604,
|
||||
"mutable_iter_ns": 16.70941883767535,
|
||||
"immutable_iter_ns": 2.312
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.1991991991991995,
|
||||
"immutable_copy_ns": 3.311,
|
||||
"mutable_access_ns": 4.621,
|
||||
"bytes": 80,
|
||||
"speedup_update": 2.7667944654877785,
|
||||
"speedup_iter": 8.889902422912249,
|
||||
"immutable_access_ns": 2.029,
|
||||
"speedup_access": 2.2774765894529327,
|
||||
"nfields": 10,
|
||||
"immutable_update_ns": 2.602,
|
||||
"mutable_iter_ns": 23.13152610441767,
|
||||
"immutable_iter_ns": 2.602
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.201201201201201,
|
||||
"immutable_copy_ns": 3.177,
|
||||
"mutable_access_ns": 4.618,
|
||||
"bytes": 160,
|
||||
"speedup_update": 2.2764646179280326,
|
||||
"speedup_iter": 11.43067764857165,
|
||||
"immutable_access_ns": 2.023,
|
||||
"speedup_access": 2.282748393475037,
|
||||
"nfields": 20,
|
||||
"immutable_update_ns": 3.1633266533066133,
|
||||
"mutable_iter_ns": 63.131632653061224,
|
||||
"immutable_iter_ns": 5.523
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.197197197197197,
|
||||
"immutable_copy_ns": 6.895895895895896,
|
||||
"mutable_access_ns": 4.617,
|
||||
"bytes": 400,
|
||||
"speedup_update": 1.1951506471599465,
|
||||
"speedup_iter": 7.942862982909166,
|
||||
"immutable_access_ns": 2.029,
|
||||
"speedup_access": 2.275505174963036,
|
||||
"nfields": 50,
|
||||
"immutable_update_ns": 6.022,
|
||||
"mutable_iter_ns": 196.84902597402598,
|
||||
"immutable_iter_ns": 24.783132530120483
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.198198198198198,
|
||||
"immutable_copy_ns": 9.476476476476476,
|
||||
"mutable_access_ns": 4.618,
|
||||
"bytes": 800,
|
||||
"speedup_update": 0.612834498039884,
|
||||
"speedup_iter": 3.803593753682347,
|
||||
"immutable_access_ns": 2.025,
|
||||
"speedup_access": 2.280493827160494,
|
||||
"nfields": 100,
|
||||
"immutable_update_ns": 11.745745745745745,
|
||||
"mutable_iter_ns": 260.55786350148367,
|
||||
"immutable_iter_ns": 68.50307377049181
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.41041041041041,
|
||||
"immutable_copy_ns": 20.948897795591183,
|
||||
"mutable_access_ns": 4.621,
|
||||
"bytes": 1600,
|
||||
"speedup_update": 0.30082611651798524,
|
||||
"speedup_iter": 5.013605562294905,
|
||||
"immutable_access_ns": 2.373,
|
||||
"speedup_access": 1.9473240623683101,
|
||||
"nfields": 200,
|
||||
"immutable_update_ns": 24.633534136546185,
|
||||
"mutable_iter_ns": 777.9142857142857,
|
||||
"immutable_iter_ns": 155.1606475716065
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.198198198198198,
|
||||
"immutable_copy_ns": 58.68463886063072,
|
||||
"mutable_access_ns": 4.624,
|
||||
"bytes": 4000,
|
||||
"speedup_update": 0.08904413723690831,
|
||||
"speedup_iter": 2.3155840000000003,
|
||||
"immutable_access_ns": 2.027,
|
||||
"speedup_access": 2.281203749383325,
|
||||
"nfields": 500,
|
||||
"immutable_update_ns": 80.83854166666667,
|
||||
"mutable_iter_ns": 1156.3,
|
||||
"immutable_iter_ns": 499.35567010309273
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.2002002002002,
|
||||
"immutable_copy_ns": 45.70171890798787,
|
||||
"mutable_access_ns": 4.671,
|
||||
"bytes": 8000,
|
||||
"speedup_update": 0.03725604952469045,
|
||||
"speedup_iter": 3.1501090710779853,
|
||||
"immutable_access_ns": 2.084,
|
||||
"speedup_access": 2.241362763915547,
|
||||
"nfields": 1000,
|
||||
"immutable_update_ns": 193.26257861635222,
|
||||
"mutable_iter_ns": 3465.75,
|
||||
"immutable_iter_ns": 1100.2
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.197197197197197,
|
||||
"immutable_copy_ns": 82.94813278008299,
|
||||
"mutable_access_ns": 4.625,
|
||||
"bytes": 16000,
|
||||
"speedup_update": 0.0175176399491468,
|
||||
"speedup_iter": 2.1613809428742545,
|
||||
"immutable_access_ns": 2.024,
|
||||
"speedup_access": 2.2850790513833994,
|
||||
"nfields": 2000,
|
||||
"immutable_update_ns": 410.8542713567839,
|
||||
"mutable_iter_ns": 4811.714285714285,
|
||||
"immutable_iter_ns": 2226.222222222222
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.2042042042042045,
|
||||
"immutable_copy_ns": 867.0408163265306,
|
||||
"mutable_access_ns": 4.623,
|
||||
"bytes": 40000,
|
||||
"speedup_update": 0.0037209876577677828,
|
||||
"speedup_iter": 5.994000529365056,
|
||||
"immutable_access_ns": 2.083,
|
||||
"speedup_access": 2.2193951032165145,
|
||||
"nfields": 5000,
|
||||
"immutable_update_ns": 1936.1,
|
||||
"mutable_iter_ns": 33969.0,
|
||||
"immutable_iter_ns": 5667.166666666667
|
||||
}
|
||||
],
|
||||
"timestamp": "20251109_201256",
|
||||
"system": {
|
||||
"cpu_model": "Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz",
|
||||
"cpu_speed_mhz": 3300,
|
||||
"machine": "x86_64-linux-gnu",
|
||||
"word_size": 64,
|
||||
"cpu_cores": 32,
|
||||
"os": "Linux"
|
||||
}
|
||||
}
|
||||
@@ -1,13 +0,0 @@
|
||||
nfields,bytes,mut_access_ns,imm_access_ns,speedup_access,mut_update_ns,imm_update_ns,speedup_update,imm_copy_ns,mut_iter_ns,imm_iter_ns,speedup_iter
|
||||
1,8,4.615,2.023,2.2812654473554126,7.1991991991991995,2.315,3.1098052696324836,2.023,11.773773773773774,2.024,5.817081904038426
|
||||
2,16,4.904,2.023,2.4241225902125554,7.207207207207207,2.314,3.11460985618289,2.029,14.48997995991984,2.025,7.1555456592196744
|
||||
5,40,4.897,2.373,2.063632532659081,7.201201201201201,2.604,2.765438249309217,2.383,16.70941883767535,2.312,7.227257282731553
|
||||
10,80,4.621,2.029,2.2774765894529327,7.1991991991991995,2.602,2.7667944654877785,3.311,23.13152610441767,2.602,8.889902422912249
|
||||
20,160,4.618,2.023,2.282748393475037,7.201201201201201,3.1633266533066133,2.2764646179280326,3.177,63.131632653061224,5.523,11.43067764857165
|
||||
50,400,4.617,2.029,2.275505174963036,7.197197197197197,6.022,1.1951506471599465,6.895895895895896,196.84902597402598,24.783132530120483,7.942862982909166
|
||||
100,800,4.618,2.025,2.280493827160494,7.198198198198198,11.745745745745745,0.612834498039884,9.476476476476476,260.55786350148367,68.50307377049181,3.803593753682347
|
||||
200,1600,4.621,2.373,1.9473240623683101,7.41041041041041,24.633534136546185,0.30082611651798524,20.948897795591183,777.9142857142857,155.1606475716065,5.013605562294905
|
||||
500,4000,4.624,2.027,2.281203749383325,7.198198198198198,80.83854166666667,0.08904413723690831,58.68463886063072,1156.3,499.35567010309273,2.3155840000000003
|
||||
1000,8000,4.671,2.084,2.241362763915547,7.2002002002002,193.26257861635222,0.03725604952469045,45.70171890798787,3465.75,1100.2,3.1501090710779853
|
||||
2000,16000,4.625,2.024,2.2850790513833994,7.197197197197197,410.8542713567839,0.0175176399491468,82.94813278008299,4811.714285714285,2226.222222222222,2.1613809428742545
|
||||
5000,40000,4.623,2.083,2.2193951032165145,7.2042042042042045,1936.1,0.0037209876577677828,867.0408163265306,33969.0,5667.166666666667,5.994000529365056
|
||||
|
@@ -1,196 +0,0 @@
|
||||
{
|
||||
"struct_sizes": [
|
||||
1,
|
||||
2,
|
||||
5,
|
||||
10,
|
||||
20,
|
||||
50,
|
||||
100,
|
||||
200,
|
||||
500,
|
||||
1000,
|
||||
2000,
|
||||
5000
|
||||
],
|
||||
"system_info": {
|
||||
"cpu_model": "Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz",
|
||||
"julia_version": "1.12.1",
|
||||
"total_memory_gb": 503.35,
|
||||
"l1_cache": "48K",
|
||||
"l2_cache": "1280K",
|
||||
"l3_cache": "24576K",
|
||||
"cpu_cores": 32
|
||||
},
|
||||
"results": [
|
||||
{
|
||||
"mutable_update_ns": 7.1991991991991995,
|
||||
"immutable_copy_ns": 2.023,
|
||||
"mutable_access_ns": 4.615,
|
||||
"bytes": 8,
|
||||
"speedup_update": 3.1098052696324836,
|
||||
"speedup_iter": 5.817081904038426,
|
||||
"immutable_access_ns": 2.023,
|
||||
"speedup_access": 2.2812654473554126,
|
||||
"nfields": 1,
|
||||
"immutable_update_ns": 2.315,
|
||||
"mutable_iter_ns": 11.773773773773774,
|
||||
"immutable_iter_ns": 2.024
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.207207207207207,
|
||||
"immutable_copy_ns": 2.029,
|
||||
"mutable_access_ns": 4.904,
|
||||
"bytes": 16,
|
||||
"speedup_update": 3.11460985618289,
|
||||
"speedup_iter": 7.1555456592196744,
|
||||
"immutable_access_ns": 2.023,
|
||||
"speedup_access": 2.4241225902125554,
|
||||
"nfields": 2,
|
||||
"immutable_update_ns": 2.314,
|
||||
"mutable_iter_ns": 14.48997995991984,
|
||||
"immutable_iter_ns": 2.025
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.201201201201201,
|
||||
"immutable_copy_ns": 2.383,
|
||||
"mutable_access_ns": 4.897,
|
||||
"bytes": 40,
|
||||
"speedup_update": 2.765438249309217,
|
||||
"speedup_iter": 7.227257282731553,
|
||||
"immutable_access_ns": 2.373,
|
||||
"speedup_access": 2.063632532659081,
|
||||
"nfields": 5,
|
||||
"immutable_update_ns": 2.604,
|
||||
"mutable_iter_ns": 16.70941883767535,
|
||||
"immutable_iter_ns": 2.312
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.1991991991991995,
|
||||
"immutable_copy_ns": 3.311,
|
||||
"mutable_access_ns": 4.621,
|
||||
"bytes": 80,
|
||||
"speedup_update": 2.7667944654877785,
|
||||
"speedup_iter": 8.889902422912249,
|
||||
"immutable_access_ns": 2.029,
|
||||
"speedup_access": 2.2774765894529327,
|
||||
"nfields": 10,
|
||||
"immutable_update_ns": 2.602,
|
||||
"mutable_iter_ns": 23.13152610441767,
|
||||
"immutable_iter_ns": 2.602
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.201201201201201,
|
||||
"immutable_copy_ns": 3.177,
|
||||
"mutable_access_ns": 4.618,
|
||||
"bytes": 160,
|
||||
"speedup_update": 2.2764646179280326,
|
||||
"speedup_iter": 11.43067764857165,
|
||||
"immutable_access_ns": 2.023,
|
||||
"speedup_access": 2.282748393475037,
|
||||
"nfields": 20,
|
||||
"immutable_update_ns": 3.1633266533066133,
|
||||
"mutable_iter_ns": 63.131632653061224,
|
||||
"immutable_iter_ns": 5.523
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.197197197197197,
|
||||
"immutable_copy_ns": 6.895895895895896,
|
||||
"mutable_access_ns": 4.617,
|
||||
"bytes": 400,
|
||||
"speedup_update": 1.1951506471599465,
|
||||
"speedup_iter": 7.942862982909166,
|
||||
"immutable_access_ns": 2.029,
|
||||
"speedup_access": 2.275505174963036,
|
||||
"nfields": 50,
|
||||
"immutable_update_ns": 6.022,
|
||||
"mutable_iter_ns": 196.84902597402598,
|
||||
"immutable_iter_ns": 24.783132530120483
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.198198198198198,
|
||||
"immutable_copy_ns": 9.476476476476476,
|
||||
"mutable_access_ns": 4.618,
|
||||
"bytes": 800,
|
||||
"speedup_update": 0.612834498039884,
|
||||
"speedup_iter": 3.803593753682347,
|
||||
"immutable_access_ns": 2.025,
|
||||
"speedup_access": 2.280493827160494,
|
||||
"nfields": 100,
|
||||
"immutable_update_ns": 11.745745745745745,
|
||||
"mutable_iter_ns": 260.55786350148367,
|
||||
"immutable_iter_ns": 68.50307377049181
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.41041041041041,
|
||||
"immutable_copy_ns": 20.948897795591183,
|
||||
"mutable_access_ns": 4.621,
|
||||
"bytes": 1600,
|
||||
"speedup_update": 0.30082611651798524,
|
||||
"speedup_iter": 5.013605562294905,
|
||||
"immutable_access_ns": 2.373,
|
||||
"speedup_access": 1.9473240623683101,
|
||||
"nfields": 200,
|
||||
"immutable_update_ns": 24.633534136546185,
|
||||
"mutable_iter_ns": 777.9142857142857,
|
||||
"immutable_iter_ns": 155.1606475716065
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.198198198198198,
|
||||
"immutable_copy_ns": 58.68463886063072,
|
||||
"mutable_access_ns": 4.624,
|
||||
"bytes": 4000,
|
||||
"speedup_update": 0.08904413723690831,
|
||||
"speedup_iter": 2.3155840000000003,
|
||||
"immutable_access_ns": 2.027,
|
||||
"speedup_access": 2.281203749383325,
|
||||
"nfields": 500,
|
||||
"immutable_update_ns": 80.83854166666667,
|
||||
"mutable_iter_ns": 1156.3,
|
||||
"immutable_iter_ns": 499.35567010309273
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.2002002002002,
|
||||
"immutable_copy_ns": 45.70171890798787,
|
||||
"mutable_access_ns": 4.671,
|
||||
"bytes": 8000,
|
||||
"speedup_update": 0.03725604952469045,
|
||||
"speedup_iter": 3.1501090710779853,
|
||||
"immutable_access_ns": 2.084,
|
||||
"speedup_access": 2.241362763915547,
|
||||
"nfields": 1000,
|
||||
"immutable_update_ns": 193.26257861635222,
|
||||
"mutable_iter_ns": 3465.75,
|
||||
"immutable_iter_ns": 1100.2
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.197197197197197,
|
||||
"immutable_copy_ns": 82.94813278008299,
|
||||
"mutable_access_ns": 4.625,
|
||||
"bytes": 16000,
|
||||
"speedup_update": 0.0175176399491468,
|
||||
"speedup_iter": 2.1613809428742545,
|
||||
"immutable_access_ns": 2.024,
|
||||
"speedup_access": 2.2850790513833994,
|
||||
"nfields": 2000,
|
||||
"immutable_update_ns": 410.8542713567839,
|
||||
"mutable_iter_ns": 4811.714285714285,
|
||||
"immutable_iter_ns": 2226.222222222222
|
||||
},
|
||||
{
|
||||
"mutable_update_ns": 7.2042042042042045,
|
||||
"immutable_copy_ns": 867.0408163265306,
|
||||
"mutable_access_ns": 4.623,
|
||||
"bytes": 40000,
|
||||
"speedup_update": 0.0037209876577677828,
|
||||
"speedup_iter": 5.994000529365056,
|
||||
"immutable_access_ns": 2.083,
|
||||
"speedup_access": 2.2193951032165145,
|
||||
"nfields": 5000,
|
||||
"immutable_update_ns": 1936.1,
|
||||
"mutable_iter_ns": 33969.0,
|
||||
"immutable_iter_ns": 5667.166666666667
|
||||
}
|
||||
],
|
||||
"timestamp": "2025-11-09T20:12:52.503"
|
||||
}
|
||||
@@ -1,731 +0,0 @@
|
||||
# Benchmark: Does immutable performance degrade with struct size?
|
||||
# Theory: Stack copying is O(n), heap pointers are O(1)
|
||||
# Question: At what size does mutable win?
|
||||
|
||||
# NOTE: Using packages from global environment (not project)
|
||||
using BenchmarkTools
|
||||
using Printf
|
||||
using JSON
|
||||
using Plots
|
||||
using JSON
|
||||
using Dates
|
||||
|
||||
# Get system information
|
||||
println("="^80)
|
||||
println("SYSTEM INFORMATION")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# CPU info
|
||||
cpu_info = Sys.cpu_info()
|
||||
println("CPU Model: ", cpu_info[1].model)
|
||||
println("CPU Cores: ", Sys.CPU_THREADS, " threads (", length(cpu_info), " physical cores)")
|
||||
println("CPU Speed: ", cpu_info[1].speed, " MHz")
|
||||
println()
|
||||
|
||||
# Julia and system info
|
||||
println("Julia Version: ", VERSION)
|
||||
println("OS: ", Sys.KERNEL, " ", Sys.MACHINE)
|
||||
println("Word Size: ", Sys.WORD_SIZE, " bits")
|
||||
println()
|
||||
|
||||
# Memory and cache info (approximate)
|
||||
println("Approximate CPU Cache Sizes:")
|
||||
println(" L1 Cache: ~32-64 KB per core (typical)")
|
||||
println(" L2 Cache: ~256-512 KB per core (typical)")
|
||||
println(" L3 Cache: ~8-32 MB shared (typical)")
|
||||
println()
|
||||
println("Note: Testing up to 8KB structs to exceed L1 cache")
|
||||
println()
|
||||
|
||||
using BenchmarkTools
|
||||
using Printf
|
||||
using JSON
|
||||
using Plots
|
||||
|
||||
# Collect system information
|
||||
function get_system_info()
|
||||
info = Dict{String,Any}()
|
||||
info["julia_version"] = string(VERSION)
|
||||
info["cpu_model"] = Sys.cpu_info()[1].model
|
||||
info["cpu_cores"] = Sys.CPU_THREADS
|
||||
info["total_memory_gb"] = round(Sys.total_memory() / 1024^3, digits=2)
|
||||
|
||||
# Try to get CPU cache info (Linux)
|
||||
try
|
||||
if Sys.islinux()
|
||||
l1_cache = read("/sys/devices/system/cpu/cpu0/cache/index0/size", String) |> strip
|
||||
l2_cache = read("/sys/devices/system/cpu/cpu0/cache/index2/size", String) |> strip
|
||||
l3_cache = read("/sys/devices/system/cpu/cpu0/cache/index3/size", String) |> strip
|
||||
info["l1_cache"] = l1_cache
|
||||
info["l2_cache"] = l2_cache
|
||||
info["l3_cache"] = l3_cache
|
||||
end
|
||||
catch
|
||||
info["cache_info"] = "Not available"
|
||||
end
|
||||
|
||||
return info
|
||||
end
|
||||
|
||||
system_info = get_system_info()
|
||||
|
||||
println("="^80)
|
||||
println("SYSTEM INFORMATION")
|
||||
println("="^80)
|
||||
println("Julia Version: $(system_info["julia_version"])")
|
||||
println("CPU Model: $(system_info["cpu_model"])")
|
||||
println("CPU Cores: $(system_info["cpu_cores"])")
|
||||
println("Total Memory: $(system_info["total_memory_gb"]) GB")
|
||||
if haskey(system_info, "l1_cache")
|
||||
println("L1 Cache: $(system_info["l1_cache"])")
|
||||
println("L2 Cache: $(system_info["l2_cache"])")
|
||||
println("L3 Cache: $(system_info["l3_cache"])")
|
||||
end
|
||||
println()
|
||||
|
||||
println("="^80)
|
||||
println("STRUCT SIZE SCALING BENCHMARK")
|
||||
println("="^80)
|
||||
println()
|
||||
println("Testing hypothesis: Immutable slows down with struct size, mutable stays constant")
|
||||
println()
|
||||
|
||||
# Test different struct sizes (number of Float64 fields)
|
||||
# Extended range to go well beyond register file and L1 cache
|
||||
STRUCT_SIZES = [1, 2, 5, 10, 20, 50, 100, 200, 500, 1000, 2000, 5000]
|
||||
|
||||
results = []
|
||||
|
||||
for nfields in STRUCT_SIZES
|
||||
println("Testing struct with $nfields Float64 fields ($(nfields * 8) bytes)...")
|
||||
|
||||
# Generate mutable version
|
||||
field_names_mut = [Symbol("field$i") for i in 1:nfields]
|
||||
|
||||
# Mutable: Dict-based
|
||||
mutable_data = Dict{Symbol,Float64}()
|
||||
for fname in field_names_mut
|
||||
mutable_data[fname] = rand()
|
||||
end
|
||||
|
||||
# Immutable: NamedTuple-based
|
||||
immutable_data = NamedTuple{Tuple(field_names_mut)}(Tuple(rand() for _ in 1:nfields))
|
||||
|
||||
# Benchmark 1: Field Access (read first field)
|
||||
first_field = field_names_mut[1]
|
||||
|
||||
time_mut_access = @belapsed $mutable_data[$first_field]
|
||||
time_imm_access = @belapsed $immutable_data.$first_field
|
||||
|
||||
# Benchmark 2: Field Update (change first field)
|
||||
time_mut_update = @belapsed begin
|
||||
$mutable_data[$first_field] = 42.0
|
||||
end
|
||||
|
||||
time_imm_update = @belapsed begin
|
||||
$immutable_data = (; $immutable_data..., $first_field=42.0)
|
||||
end
|
||||
|
||||
# Benchmark 3: Struct Copy (merge with empty to force copy)
|
||||
time_imm_copy = @belapsed merge($immutable_data, NamedTuple())
|
||||
|
||||
# Benchmark 4: Iteration over all fields
|
||||
time_mut_iter = @belapsed begin
|
||||
sum = 0.0
|
||||
for (k, v) in $mutable_data
|
||||
sum += v
|
||||
end
|
||||
sum
|
||||
end
|
||||
|
||||
time_imm_iter = @belapsed begin
|
||||
sum = 0.0
|
||||
for v in $immutable_data
|
||||
sum += v
|
||||
end
|
||||
sum
|
||||
end
|
||||
|
||||
speedup_access = time_mut_access / time_imm_access
|
||||
speedup_update = time_mut_update / time_imm_update
|
||||
speedup_iter = time_mut_iter / time_imm_iter
|
||||
|
||||
push!(results, (
|
||||
nfields=nfields,
|
||||
bytes=nfields * 8,
|
||||
# Access times
|
||||
mut_access=time_mut_access,
|
||||
imm_access=time_imm_access,
|
||||
speedup_access=speedup_access,
|
||||
# Update times
|
||||
mut_update=time_mut_update,
|
||||
imm_update=time_imm_update,
|
||||
speedup_update=speedup_update,
|
||||
# Copy time
|
||||
imm_copy=time_imm_copy,
|
||||
# Iteration times
|
||||
mut_iter=time_mut_iter,
|
||||
imm_iter=time_imm_iter,
|
||||
speedup_iter=speedup_iter
|
||||
))
|
||||
|
||||
println(" Access: Mut=$(round(time_mut_access*1e9, digits=2))ns Imm=$(round(time_imm_access*1e9, digits=2))ns Speedup=$(round(speedup_access, digits=1))x")
|
||||
println(" Update: Mut=$(round(time_mut_update*1e9, digits=2))ns Imm=$(round(time_imm_update*1e9, digits=2))ns Speedup=$(round(speedup_update, digits=1))x")
|
||||
println(" Iterate: Mut=$(round(time_mut_iter*1e9, digits=2))ns Imm=$(round(time_imm_iter*1e9, digits=2))ns Speedup=$(round(speedup_iter, digits=1))x")
|
||||
println(" Copy: $(round(time_imm_copy*1e9, digits=2))ns")
|
||||
println()
|
||||
end
|
||||
|
||||
println("="^80)
|
||||
println("RESULTS SUMMARY")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
println("Field Access Performance:")
|
||||
println("Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup")
|
||||
println("-"^70)
|
||||
for r in results
|
||||
@printf("%13d | %5d | %12.2f | %14.2f | %6.1fx\n",
|
||||
r.nfields, r.bytes, r.mut_access * 1e9, r.imm_access * 1e9, r.speedup_access)
|
||||
end
|
||||
println()
|
||||
|
||||
println("Field Update Performance:")
|
||||
println("Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup")
|
||||
println("-"^70)
|
||||
for r in results
|
||||
@printf("%13d | %5d | %12.2f | %14.2f | %6.1fx\n",
|
||||
r.nfields, r.bytes, r.mut_update * 1e9, r.imm_update * 1e9, r.speedup_update)
|
||||
end
|
||||
println()
|
||||
|
||||
println("Iteration Performance:")
|
||||
println("Size (fields) | Bytes | Mutable (ns) | Immutable (ns) | Speedup")
|
||||
println("-"^70)
|
||||
for r in results
|
||||
@printf("%13d | %5d | %12.2f | %14.2f | %6.1fx\n",
|
||||
r.nfields, r.bytes, r.mut_iter * 1e9, r.imm_iter * 1e9, r.speedup_iter)
|
||||
end
|
||||
println()
|
||||
|
||||
println("Immutable Copy Cost (ns):")
|
||||
println("Size (fields) | Bytes | Copy Time (ns)")
|
||||
println("-"^40)
|
||||
for r in results
|
||||
@printf("%13d | %5d | %13.2f\n", r.nfields, r.bytes, r.imm_copy * 1e9)
|
||||
end
|
||||
println()
|
||||
|
||||
# Analysis
|
||||
println("="^80)
|
||||
println("ANALYSIS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Check if mutable ever wins
|
||||
access_wins = [r for r in results if r.speedup_access < 1.0]
|
||||
update_wins = [r for r in results if r.speedup_update < 1.0]
|
||||
iter_wins = [r for r in results if r.speedup_iter < 1.0]
|
||||
|
||||
if isempty(access_wins)
|
||||
println("✓ Immutable ALWAYS faster for field access (even at 1000 fields = 8KB)")
|
||||
min_speedup = minimum(r.speedup_access for r in results)
|
||||
println(" Minimum speedup: $(round(min_speedup, digits=1))x at $(results[end].nfields) fields")
|
||||
else
|
||||
println("⚠ Mutable wins for field access at $(access_wins[1].nfields) fields")
|
||||
end
|
||||
println()
|
||||
|
||||
if isempty(update_wins)
|
||||
println("✓ Immutable ALWAYS faster for field update (even at 1000 fields = 8KB)")
|
||||
min_speedup = minimum(r.speedup_update for r in results)
|
||||
println(" Minimum speedup: $(round(min_speedup, digits=1))x at $(results[end].nfields) fields")
|
||||
else
|
||||
println("⚠ Mutable wins for field update at $(update_wins[1].nfields) fields")
|
||||
end
|
||||
println()
|
||||
|
||||
if isempty(iter_wins)
|
||||
println("✓ Immutable ALWAYS faster for iteration (even at 1000 fields = 8KB)")
|
||||
min_speedup = minimum(r.speedup_iter for r in results)
|
||||
println(" Minimum speedup: $(round(min_speedup, digits=1))x at $(results[end].nfields) fields")
|
||||
else
|
||||
println("⚠ Mutable wins for iteration at $(iter_wins[1].nfields) fields")
|
||||
end
|
||||
println()
|
||||
|
||||
# Check scaling behavior
|
||||
println("Scaling Analysis:")
|
||||
println()
|
||||
|
||||
# Linear regression on copy time vs size
|
||||
sizes = [r.nfields for r in results]
|
||||
copy_times = [r.imm_copy * 1e9 for r in results] # Convert to ns
|
||||
|
||||
# Simple linear fit: time = a + b*size
|
||||
n = length(sizes)
|
||||
mean_size = sum(sizes) / n
|
||||
mean_time = sum(copy_times) / n
|
||||
cov = sum((sizes[i] - mean_size) * (copy_times[i] - mean_time) for i in 1:n) / n
|
||||
var_size = sum((s - mean_size)^2 for s in sizes) / n
|
||||
slope = cov / var_size
|
||||
intercept = mean_time - slope * mean_size
|
||||
|
||||
println("Copy time scaling:")
|
||||
println(" Linear fit: time(ns) = $(round(intercept, digits=2)) + $(round(slope, digits=4)) * nfields")
|
||||
println(" Per-field cost: $(round(slope, digits=4)) ns/field")
|
||||
println(" Base overhead: $(round(intercept, digits=2)) ns")
|
||||
println()
|
||||
|
||||
# Check if copy time grows linearly
|
||||
println("Is copy time linear? (checking R²)")
|
||||
ss_tot = sum((t - mean_time)^2 for t in copy_times)
|
||||
ss_res = sum((copy_times[i] - (intercept + slope * sizes[i]))^2 for i in 1:n)
|
||||
r_squared = 1 - ss_res / ss_tot
|
||||
println(" R² = $(round(r_squared, digits=4))")
|
||||
if r_squared > 0.95
|
||||
println(" ✓ Copy time is linear in struct size (as expected)")
|
||||
else
|
||||
println(" ⚠ Copy time not perfectly linear (compiler optimizations?)")
|
||||
end
|
||||
println()
|
||||
|
||||
# Key insight
|
||||
println("KEY INSIGHT:")
|
||||
println("-"^80)
|
||||
println()
|
||||
println("Even at 1000 fields (8KB struct), immutable is STILL faster because:")
|
||||
println(" 1. Dict lookup cost (~40-50ns) >> copy cost per field (~$(round(slope, digits=4))ns)")
|
||||
println(" 2. Type stability enables compiler optimizations (inlining, SIMD)")
|
||||
println(" 3. Stack allocation has better cache locality than heap pointers")
|
||||
println()
|
||||
println("Theoretical crossover point (if it exists):")
|
||||
crossover_fields = (40.0 - intercept) / slope # When copy cost = Dict lookup
|
||||
println(" Would occur at ~$(round(Int, crossover_fields)) fields ($(round(Int, crossover_fields*8/1024))KB)")
|
||||
if crossover_fields > 1000
|
||||
println(" But this is beyond any realistic FEM element!")
|
||||
end
|
||||
println()
|
||||
|
||||
println("="^80)
|
||||
println("CONCLUSION")
|
||||
println("="^80)
|
||||
println()
|
||||
println("Your intuition about O(n) scaling is CORRECT, BUT:")
|
||||
println()
|
||||
println(" • Dict lookup base cost is SO high (~40-50ns)")
|
||||
println(" • Copy cost per field is SO low (~$(round(slope, digits=4))ns)")
|
||||
println(" • Compiler optimizations are SO good (inlining, SIMD, escape analysis)")
|
||||
println()
|
||||
println("That immutable wins even for unrealistically large structs (8KB+)!")
|
||||
println()
|
||||
println("For typical FEM elements:")
|
||||
println(" • Material properties: 3-10 fields (24-80 bytes)")
|
||||
println(" • State variables: 10-50 fields (80-400 bytes)")
|
||||
println(" • Even with 100 fields (800 bytes), immutable is >10x faster")
|
||||
println()
|
||||
println("Type stability > Everything else.")
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# SAVE DATA TO DISK
|
||||
# ============================================================================
|
||||
|
||||
println("="^80)
|
||||
println("SAVING DATA")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Create results directory
|
||||
results_dir = joinpath(@__DIR__, "results")
|
||||
mkpath(results_dir)
|
||||
|
||||
# Prepare data for JSON
|
||||
data_to_save = Dict(
|
||||
"system_info" => system_info,
|
||||
"timestamp" => string(now()),
|
||||
"struct_sizes" => STRUCT_SIZES,
|
||||
"results" => [
|
||||
Dict(
|
||||
"nfields" => r.nfields,
|
||||
"bytes" => r.bytes,
|
||||
"mutable_access_ns" => r.mut_access * 1e9,
|
||||
"immutable_access_ns" => r.imm_access * 1e9,
|
||||
"speedup_access" => r.speedup_access,
|
||||
"mutable_update_ns" => r.mut_update * 1e9,
|
||||
"immutable_update_ns" => r.imm_update * 1e9,
|
||||
"speedup_update" => r.speedup_update,
|
||||
"immutable_copy_ns" => r.imm_copy * 1e9,
|
||||
"mutable_iter_ns" => r.mut_iter * 1e9,
|
||||
"immutable_iter_ns" => r.imm_iter * 1e9,
|
||||
"speedup_iter" => r.speedup_iter
|
||||
)
|
||||
for r in results
|
||||
]
|
||||
)
|
||||
|
||||
# Save as JSON
|
||||
json_file = joinpath(results_dir, "struct_size_scaling.json")
|
||||
open(json_file, "w") do f
|
||||
JSON.print(f, data_to_save, 2)
|
||||
end
|
||||
println("✓ Data saved to: $json_file")
|
||||
|
||||
# Save as CSV for easy plotting in other tools
|
||||
csv_file = joinpath(results_dir, "struct_size_scaling.csv")
|
||||
open(csv_file, "w") do f
|
||||
println(f, "nfields,bytes,mut_access_ns,imm_access_ns,speedup_access,mut_update_ns,imm_update_ns,speedup_update,imm_copy_ns,mut_iter_ns,imm_iter_ns,speedup_iter")
|
||||
for r in results
|
||||
println(f, "$(r.nfields),$(r.bytes),$(r.mut_access*1e9),$(r.imm_access*1e9),$(r.speedup_access),$(r.mut_update*1e9),$(r.imm_update*1e9),$(r.speedup_update),$(r.imm_copy*1e9),$(r.mut_iter*1e9),$(r.imm_iter*1e9),$(r.speedup_iter)")
|
||||
end
|
||||
end
|
||||
println("✓ CSV saved to: $csv_file")
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# GENERATE PLOTS
|
||||
# ============================================================================
|
||||
|
||||
println("="^80)
|
||||
println("GENERATING PLOTS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Extract data for plotting
|
||||
bytes_vals = [r.bytes for r in results]
|
||||
mut_access = [r.mut_access * 1e9 for r in results]
|
||||
imm_access = [r.imm_access * 1e9 for r in results]
|
||||
mut_update = [r.mut_update * 1e9 for r in results]
|
||||
imm_update = [r.imm_update * 1e9 for r in results]
|
||||
mut_iter = [r.mut_iter * 1e9 for r in results]
|
||||
imm_iter = [r.imm_iter * 1e9 for r in results]
|
||||
imm_copy = [r.imm_copy * 1e9 for r in results]
|
||||
|
||||
# Typical FEM element sizes
|
||||
fem_small = 40 # 5 fields (E, ν, ρ, etc.)
|
||||
fem_medium = 160 # 20 fields (material + state)
|
||||
fem_large = 400 # 50 fields (complex plasticity)
|
||||
|
||||
# Plot 1: Field Access Performance
|
||||
p1 = plot(bytes_vals, mut_access,
|
||||
label="Mutable (Dict)",
|
||||
xlabel="Struct Size (bytes)",
|
||||
ylabel="Time (nanoseconds)",
|
||||
title="Field Access Performance vs Struct Size",
|
||||
linewidth=2,
|
||||
marker=:circle,
|
||||
legend=:topleft,
|
||||
size=(800, 600))
|
||||
plot!(p1, bytes_vals, imm_access,
|
||||
label="Immutable (NamedTuple)",
|
||||
linewidth=2,
|
||||
marker=:square)
|
||||
vline!(p1, [fem_small, fem_medium, fem_large],
|
||||
label="Typical FEM sizes",
|
||||
linestyle=:dash,
|
||||
linecolor=:gray,
|
||||
linewidth=1)
|
||||
annotate!(p1, fem_small, maximum(mut_access) * 0.9, text("Small\n(5 fields)", 8, :left))
|
||||
annotate!(p1, fem_medium, maximum(mut_access) * 0.8, text("Medium\n(20 fields)", 8, :left))
|
||||
annotate!(p1, fem_large, maximum(mut_access) * 0.7, text("Large\n(50 fields)", 8, :left))
|
||||
|
||||
plot_file1 = joinpath(results_dir, "field_access_scaling.png")
|
||||
savefig(p1, plot_file1)
|
||||
println("✓ Plot saved: $plot_file1")
|
||||
|
||||
# Plot 2: Field Update Performance (showing crossover)
|
||||
p2 = plot(bytes_vals, mut_update,
|
||||
label="Mutable (Dict)",
|
||||
xlabel="Struct Size (bytes)",
|
||||
ylabel="Time (nanoseconds)",
|
||||
title="Field Update Performance vs Struct Size (Crossover at ~800 bytes)",
|
||||
linewidth=2,
|
||||
marker=:circle,
|
||||
legend=:topleft,
|
||||
size=(800, 600))
|
||||
plot!(p2, bytes_vals, imm_update,
|
||||
label="Immutable (NamedTuple)",
|
||||
linewidth=2,
|
||||
marker=:square)
|
||||
vline!(p2, [fem_small, fem_medium, fem_large, 800],
|
||||
label=["", "", "", "Crossover (~100 fields)"],
|
||||
linestyle=[:dash, :dash, :dash, :dot],
|
||||
linecolor=[:gray, :gray, :gray, :red],
|
||||
linewidth=[1, 1, 1, 2])
|
||||
annotate!(p2, fem_small, maximum(imm_update) * 0.2, text("Small", 8, :left))
|
||||
annotate!(p2, fem_medium, maximum(imm_update) * 0.3, text("Medium", 8, :left))
|
||||
annotate!(p2, fem_large, maximum(imm_update) * 0.4, text("Large", 8, :left))
|
||||
|
||||
plot_file2 = joinpath(results_dir, "field_update_scaling.png")
|
||||
savefig(p2, plot_file2)
|
||||
println("✓ Plot saved: $plot_file2")
|
||||
|
||||
# Plot 3: Iteration Performance
|
||||
p3 = plot(bytes_vals, mut_iter,
|
||||
label="Mutable (Dict)",
|
||||
xlabel="Struct Size (bytes)",
|
||||
ylabel="Time (nanoseconds)",
|
||||
title="Iteration Performance vs Struct Size",
|
||||
linewidth=2,
|
||||
marker=:circle,
|
||||
legend=:topleft,
|
||||
size=(800, 600),
|
||||
yscale=:log10)
|
||||
plot!(p3, bytes_vals, imm_iter,
|
||||
label="Immutable (NamedTuple)",
|
||||
linewidth=2,
|
||||
marker=:square)
|
||||
vline!(p3, [fem_small, fem_medium, fem_large],
|
||||
label="Typical FEM sizes",
|
||||
linestyle=:dash,
|
||||
linecolor=:gray,
|
||||
linewidth=1)
|
||||
|
||||
plot_file3 = joinpath(results_dir, "iteration_scaling.png")
|
||||
savefig(p3, plot_file3)
|
||||
println("✓ Plot saved: $plot_file3")
|
||||
|
||||
# Plot 4: Copy Cost (linear scaling)
|
||||
p4 = plot(bytes_vals, imm_copy,
|
||||
label="Measured",
|
||||
xlabel="Struct Size (bytes)",
|
||||
ylabel="Copy Time (nanoseconds)",
|
||||
title="Immutable Struct Copy Cost (Linear Scaling)",
|
||||
linewidth=2,
|
||||
marker=:circle,
|
||||
legend=:topright,
|
||||
size=(800, 600))
|
||||
# Add linear fit line
|
||||
plot!(p4, bytes_vals, [intercept + slope * (b / 8) for b in bytes_vals],
|
||||
label="Linear fit: $(round(intercept, digits=1)) + $(round(slope, digits=3)) × nfields",
|
||||
linestyle=:dash,
|
||||
linewidth=2)
|
||||
vline!(p4, [fem_small, fem_medium, fem_large],
|
||||
label="Typical FEM sizes",
|
||||
linestyle=:dash,
|
||||
linecolor=:gray,
|
||||
linewidth=1)
|
||||
|
||||
plot_file4 = joinpath(results_dir, "copy_cost_linear.png")
|
||||
savefig(p4, plot_file4)
|
||||
println("✓ Plot saved: $plot_file4")
|
||||
|
||||
# Plot 5: Speedup ratios (showing where immutable wins)
|
||||
p5 = plot(bytes_vals, [r.speedup_access for r in results],
|
||||
label="Field Access",
|
||||
xlabel="Struct Size (bytes)",
|
||||
ylabel="Speedup (Immutable / Mutable)",
|
||||
title="Performance Speedup: Immutable vs Mutable",
|
||||
linewidth=2,
|
||||
marker=:circle,
|
||||
legend=:right,
|
||||
size=(800, 600))
|
||||
plot!(p5, bytes_vals, [r.speedup_update for r in results],
|
||||
label="Field Update",
|
||||
linewidth=2,
|
||||
marker=:square)
|
||||
plot!(p5, bytes_vals, [r.speedup_iter for r in results],
|
||||
label="Iteration",
|
||||
linewidth=2,
|
||||
marker=:diamond)
|
||||
hline!(p5, [1.0],
|
||||
label="Break-even",
|
||||
linestyle=:dot,
|
||||
linecolor=:black,
|
||||
linewidth=2)
|
||||
vline!(p5, [fem_small, fem_medium, fem_large],
|
||||
label="",
|
||||
linestyle=:dash,
|
||||
linecolor=:gray,
|
||||
linewidth=1)
|
||||
annotate!(p5, fem_large, 0.5, text("Typical FEM range →", 8, :left))
|
||||
|
||||
plot_file5 = joinpath(results_dir, "speedup_ratios.png")
|
||||
savefig(p5, plot_file5)
|
||||
println("✓ Plot saved: $plot_file5")
|
||||
|
||||
println()
|
||||
println("All plots saved to: $results_dir")
|
||||
println()
|
||||
|
||||
# Save results to JSON
|
||||
println("="^80)
|
||||
println("SAVING DATA")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
timestamp = Dates.format(now(), "yyyymmdd_HHMMSS")
|
||||
output_dir = "benchmarks/results"
|
||||
mkpath(output_dir)
|
||||
|
||||
# Prepare data for saving
|
||||
benchmark_data = Dict(
|
||||
"timestamp" => timestamp,
|
||||
"julia_version" => string(VERSION),
|
||||
"system" => Dict(
|
||||
"cpu_model" => cpu_info[1].model,
|
||||
"cpu_cores" => Sys.CPU_THREADS,
|
||||
"cpu_speed_mhz" => cpu_info[1].speed,
|
||||
"os" => string(Sys.KERNEL),
|
||||
"machine" => string(Sys.MACHINE),
|
||||
"word_size" => Sys.WORD_SIZE
|
||||
),
|
||||
"results" => [
|
||||
Dict(
|
||||
"nfields" => r.nfields,
|
||||
"bytes" => r.bytes,
|
||||
"mutable_access_ns" => r.mut_access * 1e9,
|
||||
"immutable_access_ns" => r.imm_access * 1e9,
|
||||
"speedup_access" => r.speedup_access,
|
||||
"mutable_update_ns" => r.mut_update * 1e9,
|
||||
"immutable_update_ns" => r.imm_update * 1e9,
|
||||
"speedup_update" => r.speedup_update,
|
||||
"mutable_iter_ns" => r.mut_iter * 1e9,
|
||||
"immutable_iter_ns" => r.imm_iter * 1e9,
|
||||
"speedup_iter" => r.speedup_iter,
|
||||
"immutable_copy_ns" => r.imm_copy * 1e9
|
||||
) for r in results
|
||||
],
|
||||
"analysis" => Dict(
|
||||
"copy_slope_ns_per_field" => slope,
|
||||
"copy_intercept_ns" => intercept,
|
||||
"r_squared" => r_squared
|
||||
)
|
||||
)
|
||||
|
||||
json_file = joinpath(output_dir, "struct_scaling_$(timestamp).json")
|
||||
open(json_file, "w") do f
|
||||
JSON.print(f, benchmark_data, 2)
|
||||
end
|
||||
println("✓ Data saved to: $json_file")
|
||||
println()
|
||||
|
||||
# Also save as CSV for easy plotting
|
||||
csv_file = joinpath(output_dir, "struct_scaling_$(timestamp).csv")
|
||||
open(csv_file, "w") do f
|
||||
println(f, "nfields,bytes,mut_access_ns,imm_access_ns,speedup_access,mut_update_ns,imm_update_ns,speedup_update,mut_iter_ns,imm_iter_ns,speedup_iter,imm_copy_ns")
|
||||
for r in results
|
||||
println(f, "$(r.nfields),$(r.bytes),$(r.mut_access*1e9),$(r.imm_access*1e9),$(r.speedup_access),$(r.mut_update*1e9),$(r.imm_update*1e9),$(r.speedup_update),$(r.mut_iter*1e9),$(r.imm_iter*1e9),$(r.speedup_iter),$(r.imm_copy*1e9)")
|
||||
end
|
||||
end
|
||||
println("✓ CSV saved to: $csv_file")
|
||||
println()
|
||||
|
||||
println("="^80)
|
||||
println("GENERATING PLOTS")
|
||||
println("="^80)
|
||||
println()
|
||||
|
||||
# Note: Using Plots from global environment
|
||||
try
|
||||
# Import from global environment
|
||||
pushfirst!(LOAD_PATH, "@stdlib")
|
||||
import Plots
|
||||
|
||||
# Set backend
|
||||
Plots.gr()
|
||||
|
||||
# Extract data for plotting
|
||||
bytes_sizes = [r.bytes for r in results]
|
||||
|
||||
# Plot 1: Field Access Performance
|
||||
p1 = Plots.plot(bytes_sizes, [r.mut_access * 1e9 for r in results],
|
||||
label="Mutable (Dict)", linewidth=2, marker=:circle,
|
||||
xlabel="Struct Size (bytes)", ylabel="Time (nanoseconds)",
|
||||
title="Field Access Performance",
|
||||
legend=:topleft, xscale=:log10, grid=true)
|
||||
Plots.plot!(p1, bytes_sizes, [r.imm_access * 1e9 for r in results],
|
||||
label="Immutable (NamedTuple)", linewidth=2, marker=:square)
|
||||
|
||||
# Add typical FEM element size markers
|
||||
Plots.vline!(p1, [40, 400], label="Typical FEM (5-50 fields)",
|
||||
linestyle=:dash, linewidth=1, color=:gray)
|
||||
|
||||
Plots.savefig(p1, joinpath(output_dir, "field_access_$(timestamp).png"))
|
||||
println("✓ Saved: field_access_$(timestamp).png")
|
||||
|
||||
# Plot 2: Field Update Performance
|
||||
p2 = Plots.plot(bytes_sizes, [r.mut_update * 1e9 for r in results],
|
||||
label="Mutable (Dict)", linewidth=2, marker=:circle,
|
||||
xlabel="Struct Size (bytes)", ylabel="Time (nanoseconds)",
|
||||
title="Field Update Performance",
|
||||
legend=:topleft, xscale=:log10, grid=true)
|
||||
Plots.plot!(p2, bytes_sizes, [r.imm_update * 1e9 for r in results],
|
||||
label="Immutable (NamedTuple)", linewidth=2, marker=:square)
|
||||
|
||||
Plots.vline!(p2, [40, 400], label="Typical FEM (5-50 fields)",
|
||||
linestyle=:dash, linewidth=1, color=:gray)
|
||||
|
||||
Plots.savefig(p2, joinpath(output_dir, "field_update_$(timestamp).png"))
|
||||
println("✓ Saved: field_update_$(timestamp).png")
|
||||
|
||||
# Plot 3: Iteration Performance
|
||||
p3 = Plots.plot(bytes_sizes, [r.mut_iter * 1e9 for r in results],
|
||||
label="Mutable (Dict)", linewidth=2, marker=:circle,
|
||||
xlabel="Struct Size (bytes)", ylabel="Time (nanoseconds)",
|
||||
title="Field Iteration Performance",
|
||||
legend=:topleft, xscale=:log10, yscale=:log10, grid=true)
|
||||
Plots.plot!(p3, bytes_sizes, [r.imm_iter * 1e9 for r in results],
|
||||
label="Immutable (NamedTuple)", linewidth=2, marker=:square)
|
||||
|
||||
Plots.vline!(p3, [40, 400], label="Typical FEM (5-50 fields)",
|
||||
linestyle=:dash, linewidth=1, color=:gray)
|
||||
|
||||
Plots.savefig(p3, joinpath(output_dir, "iteration_$(timestamp).png"))
|
||||
println("✓ Saved: iteration_$(timestamp).png")
|
||||
|
||||
# Plot 4: Speedup Factors
|
||||
p4 = Plots.plot(bytes_sizes, [r.speedup_access for r in results],
|
||||
label="Access Speedup", linewidth=2, marker=:circle,
|
||||
xlabel="Struct Size (bytes)", ylabel="Speedup Factor (Immutable/Mutable)",
|
||||
title="Performance Advantage of Immutable Elements",
|
||||
legend=:right, xscale=:log10, grid=true)
|
||||
Plots.plot!(p4, bytes_sizes, [r.speedup_update for r in results],
|
||||
label="Update Speedup", linewidth=2, marker=:square)
|
||||
Plots.plot!(p4, bytes_sizes, [r.speedup_iter for r in results],
|
||||
label="Iteration Speedup", linewidth=2, marker=:diamond)
|
||||
|
||||
Plots.hline!(p4, [1.0], label="Break-even", linestyle=:dash, color=:black, linewidth=1)
|
||||
Plots.vline!(p4, [40, 400], label="Typical FEM",
|
||||
linestyle=:dash, linewidth=1, color=:gray)
|
||||
|
||||
Plots.savefig(p4, joinpath(output_dir, "speedup_factors_$(timestamp).png"))
|
||||
println("✓ Saved: speedup_factors_$(timestamp).png")
|
||||
|
||||
# Plot 5: Copy Cost Scaling
|
||||
p5 = Plots.plot(bytes_sizes, [r.imm_copy * 1e9 for r in results],
|
||||
label="Measured", linewidth=2, marker=:circle,
|
||||
xlabel="Struct Size (bytes)", ylabel="Copy Time (nanoseconds)",
|
||||
title="Immutable Struct Copy Cost",
|
||||
legend=:topleft, xscale=:log10, grid=true)
|
||||
|
||||
# Add linear fit
|
||||
fitted = [intercept + slope * r.nfields for r in results]
|
||||
Plots.plot!(p5, bytes_sizes, fitted,
|
||||
label="Linear Fit ($(round(slope, digits=4)) ns/field)",
|
||||
linewidth=2, linestyle=:dash)
|
||||
|
||||
Plots.vline!(p5, [40, 400], label="Typical FEM",
|
||||
linestyle=:dash, linewidth=1, color=:gray)
|
||||
|
||||
Plots.savefig(p5, joinpath(output_dir, "copy_cost_$(timestamp).png"))
|
||||
println("✓ Saved: copy_cost_$(timestamp).png")
|
||||
|
||||
# Combined plot
|
||||
layout = Plots.@layout [a b; c d]
|
||||
p_combined = Plots.plot(p1, p2, p3, p4, layout=layout, size=(1200, 900))
|
||||
Plots.savefig(p_combined, joinpath(output_dir, "combined_$(timestamp).png"))
|
||||
println("✓ Saved: combined_$(timestamp).png")
|
||||
|
||||
println()
|
||||
println("All plots saved successfully!")
|
||||
|
||||
catch e
|
||||
println("⚠ Could not generate plots (Plots.jl not available in global environment)")
|
||||
println(" Error: $e")
|
||||
println(" Install with: julia -e 'using Pkg; Pkg.add(\"Plots\")'")
|
||||
end
|
||||
|
||||
println()
|
||||
println("="^80)
|
||||
@@ -1,159 +0,0 @@
|
||||
#!/bin/bash
|
||||
# Quick test script for GPU benchmarks
|
||||
|
||||
echo "======================================================================"
|
||||
echo "GPU Benchmark Quick Test"
|
||||
echo "======================================================================"
|
||||
echo ""
|
||||
|
||||
# Check Julia
|
||||
echo "Checking Julia installation..."
|
||||
if ! command -v julia &> /dev/null; then
|
||||
echo "❌ Julia not found! Please install Julia 1.9+"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
julia_version=$(julia --version)
|
||||
echo "✅ Found: $julia_version"
|
||||
echo ""
|
||||
|
||||
# Check GPU
|
||||
echo "Checking GPU availability..."
|
||||
if command -v nvidia-smi &> /dev/null; then
|
||||
echo "✅ NVIDIA GPU detected:"
|
||||
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader
|
||||
echo ""
|
||||
else
|
||||
echo "⚠️ No NVIDIA GPU detected. Benchmarks will run CPU-only."
|
||||
echo ""
|
||||
fi
|
||||
|
||||
# Check CUDA.jl
|
||||
echo "Checking CUDA.jl..."
|
||||
julia --project=. -e '
|
||||
using Pkg
|
||||
try
|
||||
using CUDA
|
||||
if CUDA.functional()
|
||||
println("✅ CUDA.jl functional: ", CUDA.name(CUDA.device()))
|
||||
else
|
||||
println("⚠️ CUDA.jl installed but GPU not functional")
|
||||
end
|
||||
catch
|
||||
println("⚠️ CUDA.jl not installed. Run: Pkg.add(\"CUDA\")")
|
||||
end
|
||||
' 2>/dev/null
|
||||
echo ""
|
||||
|
||||
# Run quick state management test (small size)
|
||||
echo "======================================================================"
|
||||
echo "Test 1: State Management (10K elements)"
|
||||
echo "======================================================================"
|
||||
julia --project=. -e '
|
||||
n = 10_000
|
||||
println("Running state management benchmark with $n elements...")
|
||||
include("benchmarks/gpu_state_management_benchmark.jl")
|
||||
|
||||
# Override main() to run smaller test
|
||||
Δε_p, Δα, elements_s1, geometry_s2, state_s2 = setup_benchmark(n)
|
||||
|
||||
# CPU test
|
||||
println("\n📊 CPU Test:")
|
||||
state_s2_copy = deepcopy(state_s2)
|
||||
t = @elapsed update_state_strategy2_cpu!(state_s2_copy, Δε_p, Δα)
|
||||
println("Time: $(round(t * 1000, digits=2)) ms")
|
||||
println("✅ CPU benchmark works!")
|
||||
|
||||
# GPU test (if available)
|
||||
if USE_GPU
|
||||
println("\n📊 GPU Test:")
|
||||
try
|
||||
T = Float64
|
||||
state_gpu = AssemblyState{T}(
|
||||
CUDA.zeros(T, n * 6),
|
||||
CUDA.zeros(T, n),
|
||||
n
|
||||
)
|
||||
Δε_p_flat = zeros(T, n * 6)
|
||||
Δα_gpu = CuArray(Δα)
|
||||
for i in 1:n
|
||||
offset = (i - 1) * 6
|
||||
ε = Δε_p[i]
|
||||
Δε_p_flat[offset + 1] = ε[1, 1]
|
||||
Δε_p_flat[offset + 2] = ε[2, 2]
|
||||
Δε_p_flat[offset + 3] = ε[3, 3]
|
||||
Δε_p_flat[offset + 4] = ε[1, 2]
|
||||
Δε_p_flat[offset + 5] = ε[1, 3]
|
||||
Δε_p_flat[offset + 6] = ε[2, 3]
|
||||
end
|
||||
Δε_p_flat_gpu = CuArray(Δε_p_flat)
|
||||
|
||||
update_state_strategy2_gpu!(state_gpu, Δε_p_flat_gpu, Δα_gpu)
|
||||
println("✅ GPU benchmark works!")
|
||||
catch e
|
||||
println("⚠️ GPU test failed: $e")
|
||||
end
|
||||
end
|
||||
'
|
||||
echo ""
|
||||
|
||||
# Run quick matrix-free test (small size)
|
||||
echo "======================================================================"
|
||||
echo "Test 2: Matrix-Free Newton-Krylov (1K DOFs)"
|
||||
echo "======================================================================"
|
||||
julia --project=. -e '
|
||||
n = 1000
|
||||
println("Running matrix-free benchmark with $n DOFs...")
|
||||
include("benchmarks/matrix_free_gpu_benchmark.jl")
|
||||
|
||||
# Override to run small test
|
||||
T = Float64
|
||||
K = Matrix(Tridiagonal(-ones(T, n-1), 2ones(T, n), -ones(T, n-1)))
|
||||
f = ones(T, n) * 0.1
|
||||
β = T(1e-3)
|
||||
prob = NonlinearProblem(K, f, β, n)
|
||||
|
||||
println("\n📊 CPU Test:")
|
||||
u = zeros(T, n)
|
||||
r = zeros(T, n)
|
||||
du = zeros(T, n)
|
||||
temp = zeros(T, n)
|
||||
Jv = zeros(T, n)
|
||||
|
||||
t = @elapsed iters = newton_matrix_free!(u, prob, r, du, temp, Jv;
|
||||
max_iter=10, verbose=false)
|
||||
println("Time: $(round(t * 1000, digits=2)) ms")
|
||||
println("Iterations: $iters")
|
||||
println("✅ CPU benchmark works!")
|
||||
|
||||
if USE_GPU
|
||||
println("\n📊 GPU Test:")
|
||||
try
|
||||
K_gpu = CuArray(K)
|
||||
f_gpu = CuArray(f)
|
||||
u_gpu = CUDA.zeros(T, n)
|
||||
|
||||
t_gpu = CUDA.@elapsed begin
|
||||
iters_gpu = newton_matrix_free_gpu!(u_gpu, K_gpu, f_gpu, β;
|
||||
max_iter=10, verbose=false)
|
||||
CUDA.synchronize()
|
||||
end
|
||||
|
||||
println("Time: $(round(t_gpu * 1000, digits=2)) ms")
|
||||
println("Iterations: $iters_gpu")
|
||||
println("✅ GPU benchmark works!")
|
||||
catch e
|
||||
println("⚠️ GPU test failed: $e")
|
||||
end
|
||||
end
|
||||
'
|
||||
echo ""
|
||||
|
||||
echo "======================================================================"
|
||||
echo "Quick Test Complete!"
|
||||
echo "======================================================================"
|
||||
echo ""
|
||||
echo "To run full benchmarks:"
|
||||
echo " julia --project=. benchmarks/gpu_state_management_benchmark.jl"
|
||||
echo " julia --project=. benchmarks/matrix_free_gpu_benchmark.jl"
|
||||
echo ""
|
||||
@@ -1,284 +0,0 @@
|
||||
# Benchmark: Tet10 Shape Function Derivatives - Manual vs AD
|
||||
#
|
||||
# Compares two approaches:
|
||||
# 1. Manual: Hand-calculated derivatives (traditional FEM)
|
||||
# 2. AD: Automatic differentiation using Tensors.jl gradient()
|
||||
#
|
||||
# Run with: julia --project=. benchmarks/tet10_derivatives_benchmark.jl
|
||||
|
||||
using BenchmarkTools
|
||||
using Tensors
|
||||
using Printf
|
||||
|
||||
println("="^70)
|
||||
println("Tet10 Shape Function Derivatives: Manual vs AD Benchmark")
|
||||
println("="^70)
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# METHOD 1: MANUAL (Hand-Calculated Derivatives)
|
||||
# ============================================================================
|
||||
|
||||
"""
|
||||
Tet10 shape functions (manual implementation).
|
||||
Reference element: ξ ∈ [0,1], η ∈ [0,1], ζ ∈ [0,1], ξ+η+ζ ≤ 1
|
||||
"""
|
||||
module ManualTet10
|
||||
using Tensors
|
||||
|
||||
# Shape functions
|
||||
@inline N1(ξ, η, ζ) = (1 - ξ - η - ζ) * (2 * (1 - ξ - η - ζ) - 1)
|
||||
@inline N2(ξ, η, ζ) = ξ * (2 * ξ - 1)
|
||||
@inline N3(ξ, η, ζ) = η * (2 * η - 1)
|
||||
@inline N4(ξ, η, ζ) = ζ * (2 * ζ - 1)
|
||||
@inline N5(ξ, η, ζ) = 4 * ξ * (1 - ξ - η - ζ)
|
||||
@inline N6(ξ, η, ζ) = 4 * ξ * η
|
||||
@inline N7(ξ, η, ζ) = 4 * η * (1 - ξ - η - ζ)
|
||||
@inline N8(ξ, η, ζ) = 4 * ζ * (1 - ξ - η - ζ)
|
||||
@inline N9(ξ, η, ζ) = 4 * ξ * ζ
|
||||
@inline N10(ξ, η, ζ) = 4 * η * ζ
|
||||
|
||||
# Derivatives (calculated by hand - error-prone!)
|
||||
@inline dN1_dξ(ξ, η, ζ) = 4 * ξ + 4 * η + 4 * ζ - 3
|
||||
@inline dN1_dη(ξ, η, ζ) = 4 * ξ + 4 * η + 4 * ζ - 3
|
||||
@inline dN1_dζ(ξ, η, ζ) = 4 * ξ + 4 * η + 4 * ζ - 3
|
||||
|
||||
@inline dN2_dξ(ξ, η, ζ) = 4 * ξ - 1
|
||||
@inline dN2_dη(ξ, η, ζ) = 0.0
|
||||
@inline dN2_dζ(ξ, η, ζ) = 0.0
|
||||
|
||||
@inline dN3_dξ(ξ, η, ζ) = 0.0
|
||||
@inline dN3_dη(ξ, η, ζ) = 4 * η - 1
|
||||
@inline dN3_dζ(ξ, η, ζ) = 0.0
|
||||
|
||||
@inline dN4_dξ(ξ, η, ζ) = 0.0
|
||||
@inline dN4_dη(ξ, η, ζ) = 0.0
|
||||
@inline dN4_dζ(ξ, η, ζ) = 4 * ζ - 1
|
||||
|
||||
@inline dN5_dξ(ξ, η, ζ) = 4 * (1 - 2 * ξ - η - ζ)
|
||||
@inline dN5_dη(ξ, η, ζ) = -4 * ξ
|
||||
@inline dN5_dζ(ξ, η, ζ) = -4 * ξ
|
||||
|
||||
@inline dN6_dξ(ξ, η, ζ) = 4 * η
|
||||
@inline dN6_dη(ξ, η, ζ) = 4 * ξ
|
||||
@inline dN6_dζ(ξ, η, ζ) = 0.0
|
||||
|
||||
@inline dN7_dξ(ξ, η, ζ) = -4 * η
|
||||
@inline dN7_dη(ξ, η, ζ) = 4 * (1 - ξ - 2 * η - ζ)
|
||||
@inline dN7_dζ(ξ, η, ζ) = -4 * η
|
||||
|
||||
@inline dN8_dξ(ξ, η, ζ) = -4 * ζ
|
||||
@inline dN8_dη(ξ, η, ζ) = -4 * ζ
|
||||
@inline dN8_dζ(ξ, η, ζ) = 4 * (1 - ξ - η - 2 * ζ)
|
||||
|
||||
@inline dN9_dξ(ξ, η, ζ) = 4 * ζ
|
||||
@inline dN9_dη(ξ, η, ζ) = 0.0
|
||||
@inline dN9_dζ(ξ, η, ζ) = 4 * ξ
|
||||
|
||||
@inline dN10_dξ(ξ, η, ζ) = 0.0
|
||||
@inline dN10_dη(ξ, η, ζ) = 4 * ζ
|
||||
@inline dN10_dζ(ξ, η, ζ) = 4 * η
|
||||
|
||||
# Evaluation function (returns tuple - zero allocation)
|
||||
@inline function eval_basis_and_grad(xi::Vec{3})
|
||||
ξ, η, ζ = xi[1], xi[2], xi[3]
|
||||
|
||||
N = (N1(ξ, η, ζ), N2(ξ, η, ζ), N3(ξ, η, ζ), N4(ξ, η, ζ), N5(ξ, η, ζ),
|
||||
N6(ξ, η, ζ), N7(ξ, η, ζ), N8(ξ, η, ζ), N9(ξ, η, ζ), N10(ξ, η, ζ))
|
||||
|
||||
dN = (Vec(dN1_dξ(ξ, η, ζ), dN1_dη(ξ, η, ζ), dN1_dζ(ξ, η, ζ)),
|
||||
Vec(dN2_dξ(ξ, η, ζ), dN2_dη(ξ, η, ζ), dN2_dζ(ξ, η, ζ)),
|
||||
Vec(dN3_dξ(ξ, η, ζ), dN3_dη(ξ, η, ζ), dN3_dζ(ξ, η, ζ)),
|
||||
Vec(dN4_dξ(ξ, η, ζ), dN4_dη(ξ, η, ζ), dN4_dζ(ξ, η, ζ)),
|
||||
Vec(dN5_dξ(ξ, η, ζ), dN5_dη(ξ, η, ζ), dN5_dζ(ξ, η, ζ)),
|
||||
Vec(dN6_dξ(ξ, η, ζ), dN6_dη(ξ, η, ζ), dN6_dζ(ξ, η, ζ)),
|
||||
Vec(dN7_dξ(ξ, η, ζ), dN7_dη(ξ, η, ζ), dN7_dζ(ξ, η, ζ)),
|
||||
Vec(dN8_dξ(ξ, η, ζ), dN8_dη(ξ, η, ζ), dN8_dζ(ξ, η, ζ)),
|
||||
Vec(dN9_dξ(ξ, η, ζ), dN9_dη(ξ, η, ζ), dN9_dζ(ξ, η, ζ)),
|
||||
Vec(dN10_dξ(ξ, η, ζ), dN10_dη(ξ, η, ζ), dN10_dζ(ξ, η, ζ)))
|
||||
|
||||
return N, dN
|
||||
end
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# ============================================================================
|
||||
# METHOD 2: AD (Tensors.jl gradient)
|
||||
# ============================================================================
|
||||
|
||||
module ADTet10
|
||||
using Tensors
|
||||
|
||||
# Just shape functions (no manual derivatives!)
|
||||
@inline N1(xi) = (1 - xi[1] - xi[2] - xi[3]) * (2 * (1 - xi[1] - xi[2] - xi[3]) - 1)
|
||||
@inline N2(xi) = xi[1] * (2 * xi[1] - 1)
|
||||
@inline N3(xi) = xi[2] * (2 * xi[2] - 1)
|
||||
@inline N4(xi) = xi[3] * (2 * xi[3] - 1)
|
||||
@inline N5(xi) = 4 * xi[1] * (1 - xi[1] - xi[2] - xi[3])
|
||||
@inline N6(xi) = 4 * xi[1] * xi[2]
|
||||
@inline N7(xi) = 4 * xi[2] * (1 - xi[1] - xi[2] - xi[3])
|
||||
@inline N8(xi) = 4 * xi[3] * (1 - xi[1] - xi[2] - xi[3])
|
||||
@inline N9(xi) = 4 * xi[1] * xi[3]
|
||||
@inline N10(xi) = 4 * xi[2] * xi[3]
|
||||
|
||||
const shape_fns = (N1, N2, N3, N4, N5, N6, N7, N8, N9, N10)
|
||||
|
||||
@inline function eval_basis_and_grad(xi::Vec{3})
|
||||
# Evaluate basis functions
|
||||
N = ntuple(i -> shape_fns[i](xi), 10)
|
||||
|
||||
# Compute gradients with Tensors.jl gradient()
|
||||
dN = ntuple(i -> gradient(shape_fns[i], xi), 10)
|
||||
|
||||
return N, dN
|
||||
end
|
||||
end
|
||||
|
||||
# ============================================================================
|
||||
# BENCHMARKING
|
||||
# ============================================================================
|
||||
|
||||
println("Setting up benchmark...")
|
||||
println()
|
||||
|
||||
# Test point (typical integration point)
|
||||
const ξ_test = Vec(0.25, 0.25, 0.25)
|
||||
|
||||
# Verification: Both methods should give same results
|
||||
println("Verifying correctness...")
|
||||
N_manual, dN_manual = ManualTet10.eval_basis_and_grad(ξ_test)
|
||||
N_ad, dN_ad = ADTet10.eval_basis_and_grad(ξ_test)
|
||||
|
||||
println(" Manual basis: ", N_manual)
|
||||
println(" AD basis: ", N_ad)
|
||||
println()
|
||||
|
||||
# Check agreement
|
||||
rtol = 1e-10
|
||||
if !all(isapprox.(N_manual, N_ad, rtol=rtol))
|
||||
@warn "Manual and AD basis functions disagree!"
|
||||
end
|
||||
|
||||
# Check derivatives
|
||||
for i in 1:10
|
||||
if !isapprox(dN_manual[i], dN_ad[i], rtol=rtol)
|
||||
@warn "Manual and AD derivative $i disagree!" dN_manual[i] dN_ad[i]
|
||||
end
|
||||
end
|
||||
|
||||
println("✓ Both methods agree (within tolerance)")
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# RUN BENCHMARKS
|
||||
# ============================================================================
|
||||
|
||||
println("Running benchmarks (this may take a minute)...")
|
||||
println()
|
||||
|
||||
# Warm-up
|
||||
for _ in 1:1000
|
||||
ManualTet10.eval_basis_and_grad(ξ_test)
|
||||
ADTet10.eval_basis_and_grad(ξ_test)
|
||||
end
|
||||
|
||||
# Benchmark each method
|
||||
b_manual = @benchmark ManualTet10.eval_basis_and_grad($ξ_test)
|
||||
b_ad = @benchmark ADTet10.eval_basis_and_grad($ξ_test)
|
||||
|
||||
# ============================================================================
|
||||
# RESULTS
|
||||
# ============================================================================
|
||||
|
||||
# ============================================================================
|
||||
# RESULTS
|
||||
# ============================================================================
|
||||
|
||||
println("="^70)
|
||||
println("RESULTS")
|
||||
println("="^70)
|
||||
println()
|
||||
|
||||
# Extract median times
|
||||
t_manual = median(b_manual.times)
|
||||
t_ad = median(b_ad.times)
|
||||
|
||||
# Extract allocations
|
||||
alloc_manual = b_manual.allocs
|
||||
alloc_ad = b_ad.allocs
|
||||
|
||||
# Calculate relative speed
|
||||
rel_ad = t_ad / t_manual
|
||||
|
||||
println("Method | Time (ns) | Allocations | Relative Speed")
|
||||
println("----------------|-----------|-------------|----------------")
|
||||
@printf "Manual | %9.1f | %11d | %.2f× (baseline)\n" t_manual alloc_manual 1.0
|
||||
@printf "AD (Tensors.jl) | %9.1f | %11d | %.2f×\n" t_ad alloc_ad rel_ad
|
||||
println()
|
||||
|
||||
# Detailed stats
|
||||
println("Detailed Statistics:")
|
||||
println()
|
||||
println("Manual (Hand-Calculated):")
|
||||
display(b_manual)
|
||||
println()
|
||||
println()
|
||||
println("AD (Tensors.jl gradient):")
|
||||
display(b_ad)
|
||||
println()
|
||||
println()
|
||||
|
||||
# ============================================================================
|
||||
# ANALYSIS
|
||||
# ============================================================================
|
||||
|
||||
println("="^70)
|
||||
println("ANALYSIS")
|
||||
println("="^70)
|
||||
println()
|
||||
|
||||
if rel_ad < 2.0
|
||||
println("🎉 RECOMMENDATION: Use AD everywhere!")
|
||||
println()
|
||||
println("Tensors.jl AD is within 2× of manual, providing:")
|
||||
println(" ✓ Zero maintenance burden")
|
||||
println(" ✓ No manual derivative errors")
|
||||
println(" ✓ Easy to add new elements")
|
||||
println(" ✓ Supports any basis type")
|
||||
println()
|
||||
println("Small performance cost is acceptable for these benefits.")
|
||||
elseif rel_ad < 5.0
|
||||
println("⚠️ RECOMMENDATION: Hybrid approach")
|
||||
println()
|
||||
println("AD is 2-5× slower than manual. Consider:")
|
||||
println(" • Common elements (Tet10, Hex8, Quad4): Manual")
|
||||
println(" • Rare elements: AD-generated")
|
||||
println(" • Research/prototype elements: Always AD")
|
||||
println()
|
||||
println("This balances performance and maintainability.")
|
||||
else
|
||||
println("❌ RECOMMENDATION: Manual derivatives (with symbolic generation)")
|
||||
println()
|
||||
println("AD is >5× slower than manual. For performance-critical code:")
|
||||
println(" • Generate derivatives with SymPy/Symbolics.jl")
|
||||
println(" • Unit test against AD to verify correctness")
|
||||
println(" • Accept the maintenance burden")
|
||||
println()
|
||||
println("Consider AD only for prototyping.")
|
||||
end
|
||||
|
||||
println()
|
||||
println("Memory analysis:")
|
||||
if alloc_manual == 0 && alloc_ad == 0
|
||||
println(" ✓ Both methods achieve zero allocations (excellent!)")
|
||||
elseif alloc_manual == 0 && alloc_ad > 0
|
||||
println(" ⚠️ AD allocates (", alloc_ad, " allocs)")
|
||||
println(" This will hurt performance in tight loops.")
|
||||
else
|
||||
println(" ⚠️ Unexpected allocation pattern - investigate!")
|
||||
end
|
||||
|
||||
println()
|
||||
println("="^70)
|
||||
println("Benchmark complete! Results saved to console.")
|
||||
println("="^70)
|
||||