Files
JuliaFEM.jl/demos/README_GPU_MPI.md
T
Jukka Aho 6d583eb30c docs(demos): Add GPU and MPI demonstration guide
New 77-line guide documenting:
- Prerequisites (MPI and CUDA globally installed)
- Running commands for MPI communication test
- Running commands for combined GPU+MPI test
- Single-process GPU test instructions
- What gets demonstrated (type stability requirement, MPI fast transfer, real hardware)
- Success indicators and result interpretation
- Key insight: same patterns enable CPU speedup, GPU execution, and MPI efficiency
2025-11-09 10:47:32 +02:00

78 lines
1.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Running GPU and MPI Demonstrations
## Prerequisites
These demonstrations require MPI and CUDA to be installed globally:
```bash
# MPI is already installed on your system
# CUDA is already installed on your system
```
The packages are loaded dynamically, so they don't need to be in `Project.toml`.
## Running the Demonstrations
### MPI Communication Test
Test data transfer between 2 MPI processes:
```bash
cd /home/juajukka/dev/JuliaFEM.jl
mpirun -np 2 julia benchmarks/gpu_mpi_demo.jl
```
Expected output:
- Rank 0 sends typed arrays to Rank 1
- Shows bytes transferred
- Validates data integrity
### GPU + MPI Combined Test
If CUDA GPU is available, runs full workflow:
```bash
mpirun -np 2 julia benchmarks/gpu_mpi_demo.jl
```
Expected output:
- Detects GPU (if available)
- Compiles type-stable kernel for GPU
- Executes on real hardware
- Transfers results via MPI
### Single-Process GPU Test
To test GPU without MPI:
```bash
julia benchmarks/gpu_only_demo.jl
```
## What Gets Demonstrated
1. **Type Stability Requirement**
- `Matrix{Float64}` transfers to GPU ✅
- `Dict{String,Any}` would FAIL GPU compilation ❌
2. **MPI Fast Transfer**
- Typed arrays: Fast buffer transfer (memcpy)
- Mixed types: Slow serialization (~100× slower)
3. **Real Hardware Execution**
- Actual CUDA kernel compilation and execution
- Actual MPI inter-process communication
- No mocks, no simulation
## Interpreting Results
Success indicators:
- ✅ "MPI communication successful" - Type-stable data transferred
- ✅ "GPU execution successful" - Kernel compiled and ran on GPU
- ✅ "Combined GPU+MPI workflow successful" - End-to-end validated
The key insight: The same type-stable patterns that give 9-92× CPU speedup also enable GPU execution and efficient MPI communication.