Compiler Internals
Compiler Architecture Overview
Braid has a dual-pipeline compiler architecture. The C pipeline uses a recursive descent parser and direct bytecode generation for the BraidVM. The C++ pipeline uses a ParserAdapter, AST lowering to Braid IR, and supports MLIR/LLVM backends. Both pipelines share the same lexer and runtime.
Pipeline (C):
Source (.br) -> Lexer -> Parser -> Type Checker -> Bytecode Codegen -> .bx
Pipeline (C++):
Source (.br) -> ParserAdapter -> AST -> ASTLowering -> IR -> IRVerifier -> CCodegen / MLIR / LLVM
File Formats:
.br - Braid source code
.bd - Braid script (interpreted)
.bx - Braid bytecode (binary)Lexer and Parser (C Pipeline)
The lexer (lexer.c) tokenizes source code into a stream of Token structs with type, text, line, and column information. The parser (parser.c) is a recursive descent parser that produces an AST using AstNode structs with a tagged union for all node types.
Token Types:
TOK_FN, TOK_LET, TOK_IF, TOK_ELSE, TOK_WHILE, TOK_RETURN
TOK_STRUCT, TOK_ENUM, TOK_MATCH, TOK_IMPORT
TOK_EXTERN, TOK_NATIVE, TOK_DIAMETER, TOK_POLE
TOK_IDENTIFIER, TOK_NUMBER, TOK_FLOAT, TOK_STRING
TOK_LBRACE, TOK_RBRACE, TOK_LBRACKET, TOK_RBRACKET
TOK_AT (@), TOK_ELLIPSIS (...), TOK_MODEL
AST Node Types (braid.h):
AST_PROGRAM, AST_FN_DECL, AST_VAR_DECL, AST_BLOCK
AST_IF, AST_WHILE, AST_RETURN, AST_BINARY_OP
AST_LITERAL, AST_FLOAT_LITERAL, AST_STRING_LITERAL
AST_STRUCT_DECL, AST_STRUCT_LITERAL, AST_DIAMETER_DECL
AST_AUTOGRAD_DECL, AST_LAYER_DECL, AST_MODEL_DEF
AST_TENSOR_LITERAL, AST_SLICE_ACCESS, AST_TO_DEVICE
AST_WITH_DEVICEType Checker
Braid uses a Hindley-Milner type inference system with unification. Type variables are resolved through unification, which merges types and creates parent links. Primitive types include int, float, bool, string. Functions have types like (arg1, arg2) -> return_type.
C++ Pipeline Stages
The CompilationPipeline class manages five stages:
- Parse (
ParserAdapter): Parses source into a typed AST module - IR Lowering (
ASTLowering): Lowers AST to SSA-based Braid IR with opcodes likeConstantI64,AddI64,Call,Branch - IR Verify (
IRVerifier): Validates IR structure, operand types, and control flow - MLIR (
BraidIRToMLIRLowering): Converts Braid IR to MLIR dialect (stub) - LLVM IR (
BraidIRBackend): Emits LLVM IR (stub)
Bytecode Codegen
The C pipeline's bytecode codegen produces a stack-based instruction stream for the BraidVM. Opcodes include OP_LOAD_CONST, OP_ADD, OP_CALL, OP_INVOKE_FFI, OP_TENSOR_LITERAL, and OP_AUTOGRAD_FN. The bytecode is packaged into a .bx file with a header containing magic bytes BRDC, version, target kind, and payload size.
.bx Format:
Offset Size Field
0 4 Magic: 'BRDC'
4 4 Version
8 4 Target kind (1=LLVM IR, etc.)
12 4 Payload size (bytes)
16 n Payload data
IR Opcodes:
ConstantI64, ConstantBool, AddI64, SubI64, MulI64, DivI64
LessThanI64, EqualI64, Branch, BranchIfFalse, Call, Return
Bytecode Opcodes:
OP_HALT, OP_LOAD_CONST, OP_ADD, OP_SUB, OP_CALL, OP_RETURN
OP_NEW_STRUCT, OP_GET_FIELD, OP_SET_FIELD
OP_TENSOR_LITERAL, OP_SLICE_GET, OP_AUTOGRAD_FN
OP_LAYER_STRUCT, OP_TO_DEVICE, OP_WITH_DEVICEC Codegen
For the braidc run command, the C++ pipeline generates C source code via CCodegen. The emitted C code is compiled with gcc and executed as a native binary. The C codegen strips the LLM runtime include and patches void main() to int main() for standalone execution.
// braidc ccodegen hello.br output (simplified):
#include <stdio.h>
#include <stdlib.h>
#include <stdint.h>
#include <stdbool.h>
void print_value(int64_t v) { printf("%lld\n", (long long)v); }
int main() {
int64_t result = 42;
print_value(result);
return 0;
}Optimization Passes
The compiler includes several optimization passes in opt.c and supercompiler.c: constant folding (e.g., AddI64 with two constants simplifies to a constant), dead code elimination, polyhedral optimization for tensor loop nests, and supercompilation for compile-time evaluation of pure functions.