BRAIDGROUP
RESEARCH & DEV
25. Documentation

Compiler Internals

Compiler Architecture Overview

Braid has a dual-pipeline compiler architecture. The C pipeline uses a recursive descent parser and direct bytecode generation for the BraidVM. The C++ pipeline uses a ParserAdapter, AST lowering to Braid IR, and supports MLIR/LLVM backends. Both pipelines share the same lexer and runtime.

Pipeline (C):
Source (.br) -> Lexer -> Parser -> Type Checker -> Bytecode Codegen -> .bx

Pipeline (C++):
Source (.br) -> ParserAdapter -> AST -> ASTLowering -> IR -> IRVerifier -> CCodegen / MLIR / LLVM

File Formats:
  .br  - Braid source code
  .bd  - Braid script (interpreted)
  .bx  - Braid bytecode (binary)

Lexer and Parser (C Pipeline)

The lexer (lexer.c) tokenizes source code into a stream of Token structs with type, text, line, and column information. The parser (parser.c) is a recursive descent parser that produces an AST using AstNode structs with a tagged union for all node types.

Token Types:
  TOK_FN, TOK_LET, TOK_IF, TOK_ELSE, TOK_WHILE, TOK_RETURN
  TOK_STRUCT, TOK_ENUM, TOK_MATCH, TOK_IMPORT
  TOK_EXTERN, TOK_NATIVE, TOK_DIAMETER, TOK_POLE
  TOK_IDENTIFIER, TOK_NUMBER, TOK_FLOAT, TOK_STRING
  TOK_LBRACE, TOK_RBRACE, TOK_LBRACKET, TOK_RBRACKET
  TOK_AT (@), TOK_ELLIPSIS (...), TOK_MODEL

AST Node Types (braid.h):
  AST_PROGRAM, AST_FN_DECL, AST_VAR_DECL, AST_BLOCK
  AST_IF, AST_WHILE, AST_RETURN, AST_BINARY_OP
  AST_LITERAL, AST_FLOAT_LITERAL, AST_STRING_LITERAL
  AST_STRUCT_DECL, AST_STRUCT_LITERAL, AST_DIAMETER_DECL
  AST_AUTOGRAD_DECL, AST_LAYER_DECL, AST_MODEL_DEF
  AST_TENSOR_LITERAL, AST_SLICE_ACCESS, AST_TO_DEVICE
  AST_WITH_DEVICE

Type Checker

Braid uses a Hindley-Milner type inference system with unification. Type variables are resolved through unification, which merges types and creates parent links. Primitive types include int, float, bool, string. Functions have types like (arg1, arg2) -> return_type.

C++ Pipeline Stages

The CompilationPipeline class manages five stages:

  1. Parse (ParserAdapter): Parses source into a typed AST module
  2. IR Lowering (ASTLowering): Lowers AST to SSA-based Braid IR with opcodes like ConstantI64, AddI64, Call, Branch
  3. IR Verify (IRVerifier): Validates IR structure, operand types, and control flow
  4. MLIR (BraidIRToMLIRLowering): Converts Braid IR to MLIR dialect (stub)
  5. LLVM IR (BraidIRBackend): Emits LLVM IR (stub)

Bytecode Codegen

The C pipeline's bytecode codegen produces a stack-based instruction stream for the BraidVM. Opcodes include OP_LOAD_CONST, OP_ADD, OP_CALL, OP_INVOKE_FFI, OP_TENSOR_LITERAL, and OP_AUTOGRAD_FN. The bytecode is packaged into a .bx file with a header containing magic bytes BRDC, version, target kind, and payload size.

.bx Format:
Offset  Size  Field
0       4     Magic: 'BRDC'
4       4     Version
8       4     Target kind (1=LLVM IR, etc.)
12      4     Payload size (bytes)
16      n     Payload data

IR Opcodes:
  ConstantI64, ConstantBool, AddI64, SubI64, MulI64, DivI64
  LessThanI64, EqualI64, Branch, BranchIfFalse, Call, Return

Bytecode Opcodes:
  OP_HALT, OP_LOAD_CONST, OP_ADD, OP_SUB, OP_CALL, OP_RETURN
  OP_NEW_STRUCT, OP_GET_FIELD, OP_SET_FIELD
  OP_TENSOR_LITERAL, OP_SLICE_GET, OP_AUTOGRAD_FN
  OP_LAYER_STRUCT, OP_TO_DEVICE, OP_WITH_DEVICE

C Codegen

For the braidc run command, the C++ pipeline generates C source code via CCodegen. The emitted C code is compiled with gcc and executed as a native binary. The C codegen strips the LLM runtime include and patches void main() to int main() for standalone execution.

// braidc ccodegen hello.br output (simplified):
#include <stdio.h>
#include <stdlib.h>
#include <stdint.h>
#include <stdbool.h>

void print_value(int64_t v) { printf("%lld\n", (long long)v); }

int main() {
    int64_t result = 42;
    print_value(result);
    return 0;
}

Optimization Passes

The compiler includes several optimization passes in opt.c and supercompiler.c: constant folding (e.g., AddI64 with two constants simplifies to a constant), dead code elimination, polyhedral optimization for tensor loop nests, and supercompilation for compile-time evaluation of pure functions.