BRAIDGROUP
RESEARCH & DEV
33. Documentation

Benchmarking

Bench-cpu-training

The bench_cpu_training.c benchmark measures LLM training performance on CPU, including three sub-benchmarks:

  • bench_forward_small — small model (d_model=64, 2 layers, 8 tokens), measures forward pass latency
  • bench_forward_13m — medium model (d_model=512, 6 layers, 13M param equivalent), measures throughput
  • bench_train_step — full training step including forward + backward (d_model=128, 2 layers), measures step time

Key Metrics

MetricDescription
Loss convergenceCross-entropy loss decreasing over training steps, measured by hierarchical head
Steps/secTraining steps per second, reported as tok/s (tokens per second)
Memory usageCPU memory topology probe reports L1/L2/L3 cache sizes, NUMA nodes, and CXL memory
GFLOP/sEstimated compute throughput from forward pass benchmark
Step timeMilliseconds per training step for the full update cycle

Running Benchmarks

Build and run the benchmark from the braid-lang directory:

cd braid-lang/build
cmake --build . --target bench-cpu-training
./tests/bench-cpu-training

Or build all test targets and run via ctest:

cmake --build .
ctest -R bench

Sample Output

Running bench-cpu-training on a modern x86-64 system produces output similar to:

=== CPU Training Benchmarks ===
bench_forward_small ... OK  dt=1.24ms  tok/s=6451  GFLOP/s=12.3
bench_forward_13m ...   OK  dt=38.7ms  tok/s=206.7
bench_train_step ...    OK  dt=152.3ms/step
Result: ALL PASS

Metrics are measured after a warmup phase (typically 2-5 iterations) to stabilize cache behavior. The timer uses clock_gettime(CLOCK_MONOTONIC) on Linux or QueryPerformanceCounter on Windows.