33. Documentation
Benchmarking
Bench-cpu-training
The bench_cpu_training.c benchmark measures LLM training performance on CPU, including three sub-benchmarks:
- bench_forward_small — small model (d_model=64, 2 layers, 8 tokens), measures forward pass latency
- bench_forward_13m — medium model (d_model=512, 6 layers, 13M param equivalent), measures throughput
- bench_train_step — full training step including forward + backward (d_model=128, 2 layers), measures step time
Key Metrics
| Metric | Description |
|---|---|
| Loss convergence | Cross-entropy loss decreasing over training steps, measured by hierarchical head |
| Steps/sec | Training steps per second, reported as tok/s (tokens per second) |
| Memory usage | CPU memory topology probe reports L1/L2/L3 cache sizes, NUMA nodes, and CXL memory |
| GFLOP/s | Estimated compute throughput from forward pass benchmark |
| Step time | Milliseconds per training step for the full update cycle |
Running Benchmarks
Build and run the benchmark from the braid-lang directory:
cd braid-lang/build cmake --build . --target bench-cpu-training ./tests/bench-cpu-training
Or build all test targets and run via ctest:
cmake --build . ctest -R bench
Sample Output
Running bench-cpu-training on a modern x86-64 system produces output similar to:
=== CPU Training Benchmarks === bench_forward_small ... OK dt=1.24ms tok/s=6451 GFLOP/s=12.3 bench_forward_13m ... OK dt=38.7ms tok/s=206.7 bench_train_step ... OK dt=152.3ms/step Result: ALL PASS
Metrics are measured after a warmup phase (typically 2-5 iterations) to stabilize cache behavior. The timer uses clock_gettime(CLOCK_MONOTONIC) on Linux or QueryPerformanceCounter on Windows.