Optimization & Performance
Compiler Optimizations
Braid's compiler applies several optimization passes during compilation to improve generated code quality.
Constant Folding
The optimizer evaluates constant expressions at compile time. For example, an AddI64 operation where both operands are constants is replaced with a single constant. This is implemented in opt.c and runs before code generation.
// Before optimization:
let result = 100 + 200
// IR: %0 = ConstantI64 100
// %1 = ConstantI64 200
// %2 = AddI64 %0, %1
// After constant folding:
let result = 300
// IR: %0 = ConstantI64 300Dead Code Elimination
Unreachable code and unused variable assignments are removed from the IR. Variables that are assigned but never read, and code paths that can never be reached, are eliminated.
// Before DCE:
fn compute() -> int {
let unused = 42
let result = 10
return result
let dead = 99 // unreachable
}
// After DCE:
fn compute() -> int {
let result = 10
return result
}Supercompilation
The supercompiler pass (run_supercompiler in supercompiler.c) performs compile-time function evaluation. Pure functions with constant arguments are evaluated at compile time, and their results are inlined as constants. This is Phase 5 of the C compilation pipeline.
// Supercompiler evaluates this at compile time:
fn square(n: int) -> int {
return n * n
}
fn main() {
let x = square(5) + square(10)
// Supercompiler inlines: (5*5) + (10*10)
// After fold: 25 + 100
// After fold: 125
print(x)
}Runtime Optimizations
Tensor Operation Vectorization (AVX2)
The Braid runtime uses AVX2 SIMD instructions for tensor operations. The CMake default flags include -mavx2 -mfma -march=native to enable compiler auto-vectorization. Critical tensor ops (matmul, element-wise add, activation functions) are implemented with explicit SIMD intrinsics in cpu_streaming.c and ternary.c.
// Runtime uses AVX2 for tensor matmul:
// - 256-bit vectors process 4 float64 or 8 float32 elements per instruction
// - FMA (fused multiply-add) for dot products
// - Cache-blocked tiling for large matricesCPU Streaming Engine
The CPU streaming module (cpu_streaming.c) processes tensor data in streaming fashion, overlapping computation with memory loads. This minimizes cache misses and keeps the CPU pipeline full during ML training loops.
// Streaming engine processes data in chunks:
// while stream.has_data() {
// chunk = stream.next_batch(1024) // prefetched
// result = compute(chunk) // process while next chunk loads
// }Memory Pool Allocation
Tensor memory pools (cpu_memory.c, test_mempool_c.c) provide O(1) allocation and bulk deallocation. During training, intermediate tensors are allocated from the pool and the pool is reset at each step, eliminating individual free() calls.
// Pool allocation pattern:
fn train_epoch(data: tensor) {
let pool = create_mempool(1024 * 1024 * 100) // 100MB pool
let step = 0
while step < 1000 {
let batch = data[step * 32 : (step + 1) * 32]
let hidden = layer1.forward(batch, pool) // allocates from pool
let output = layer2.forward(hidden, pool)
pool.reset() // O(1) bulk free, no individual frees
step = step + 1
}
pool.destroy()
}Benchmark Results
Performance metrics from the CPU training benchmark suite (bench-cpu-training):
| Operation | Throughput | Notes |
|---|---|---|
| Ternary matmul (512x512) | ~250 GFLOPS | AVX2 ternary packing |
| Float32 matmul (1024x1024) | ~45 GFLOPS | Cache-blocked, AVX2+FMA |
| Memory pool alloc/free | ~50 ns / op | vs ~300 ns for malloc/free |
| 13M param forward pass | ~2 ms | Batch size 8, seq len 8 |
| 13M param training step | ~15 ms | Forward + backward + update |
Tips for Writing Fast Braid Code
- Use pool allocation in tight loops: call
pool.reset()instead of letting ARC free tensors individually - Keep tensors on the same device: minimize CPU-GPU transfers by using
with device("cuda")blocks - Batch operations: process data in vectorized batches rather than element-wise loops
- Prefer compile-time evaluation: constant expressions and pure functions with constant args are supercompiled
- Minimize closure creation: closures in hot loops prevent inlining and add ARC overhead
- Use native functions for performance-critical operations;
native fncalls have zero marshalling overhead - Avoid nil checks in hot paths: use the type system to guarantee non-nil values
// Optimized training loop:
fn train_fast(model, data: tensor, pool) {
with device("cuda") {
let step = 0
while step < 1000 {
let batch = data[step * 32 : (step + 1) * 32]
let logits = model.forward(batch, pool)
let loss = cross_entropy(logits, labels)
model.backward(loss)
model.update(0.001)
pool.reset() // O(1), avoids ARC overhead
step = step + 1
}
}
}