Commit Graph

1 Commits

Author SHA1 Message Date
Jukka Aho c65abfa5cc docs: Add 'Roadmap to HPC' - justifying hard performance choices
**Purpose:** Comprehensive justification for all technical decisions prioritizing
performance over convenience.

**Key Principles:**
- Efficiency > Educativeness (when forced to choose)
- Type stability over everything (100× performance difference)
- No free lunch - Julia doesn't make miracles
- HPC requires discipline and trade-offs

**Core Decisions Justified:**

1. **No Dynamic Field System**
   - field["foo"] = x is 100× slower (Dict{String,Any})
   - Type-stable structs only
   - Sacrifice: Runtime flexibility
   - Gain: Performance

2. **Immutable Data Structures**
   - struct over mutable struct
   - Sacrifice: Convenient mutation
   - Gain: 2-10× speedup, thread-safety, stack allocation

3. **NTuple Over Vector**
   - Compile-time size → SIMD optimization
   - Sacrifice: Dynamic sizing
   - Gain: Zero allocations, type stability

4. **Monolithic Over Multi-Package**
   - Learned from 2015-2019 mistake
   - Sacrifice: Small dependencies
   - Gain: It actually works

5. **Manual Derivatives (hot paths)**
   - 30× faster than AD for Tet10
   - Sacrifice: More code
   - Gain: Assembly loops stay fast

6. **Matrix-Free Methods**
   - Design for 1M+ DOF from day 1
   - Cannot retrofit later

7. **Explicit Over Implicit**
   - No magic, show the steps
   - Debuggable and teachable

**Hierarchy of Values:**
1. Correctness
2. Performance
3. Maintainability
4. Educativeness
5. Convenience

**What We're Giving Up:**
- Runtime flexibility (no element["custom_field"])
- Dynamic problem definition (no runtime topology changes)
- Duck typing convenience
- Small dependencies
- Beginner-friendly magic

**What We're Getting:**
- 10× single-thread speedup target
- 1M DOF contact problems
- Thread/GPU/distributed scalability
- Real HPC capability

**The Hard Truth:**
From Issue #266: "Do like Python, be slow like Python. Know what you do
before compiling, and be fast like C. There's no free lunch."

**Success Metrics:**
-  Zero allocations in assembly
-  Type-stable hot paths
- 🎯 10× faster than v0.5.1
- 🎯 1M DOF in < 1 hour
- 🎯 100+ thread scaling

**Use Cases:**
- "Why can't I use Dict?" → Point here
- "Why immutable?" → Point here
- "Why manual derivatives?" → Point here
- Any "why not convenience?" → Point here

**Status:** Living document, updated as we learn

See: Issue #266, TECHNICAL_VISION.md, benchmark results
2025-11-09 05:01:34 +02:00