BRAIDGROUP
RESEARCH & DEV
65. ML Library Docs

Tensor Operations

All tensor arithmetic and reduction operations support broadcasting, automatic type promotion, and both C and Braid APIs. SIMD acceleration is used where available for float32 and float64 paths.

Automatic Type Promotion

Binary operations promote both operands to a common dtype using the promote_dtype function. The promotion hierarchy is:

FLOAT64 > FLOAT32 > FLOAT16 > BFLOAT16 > INT64 > INT32 > INT16 > INT8

If either operand is FLOAT64, the result is FLOAT64. If either is FLOAT16 (and neither is FLOAT64 or FLOAT32), the result is FLOAT16, and so on.

Broadcasting

Broadcasting aligns shapes from the rightmost dimension. A dimension is compatible if both dimensions are equal, or one of them is 1. The result takes the larger dimension at each position. Scalar-like dimensions (size 1) are implemented with zero strides — no data copying occurs.

// Example broadcasting rules:
// [3, 1]  + [1, 4]  -> [3, 4]
// [5]     + [3, 5]  -> [3, 5]
// [2, 3]  + [1, 1]  -> [2, 3]
// [2, 3]  + [2, 3]  -> [2, 3]  (no broadcast needed)

Arithmetic Operations

tensor_add(a, b)

// Add two tensors with broadcasting
ObjTensor* a = /* [3, 1] */;
ObjTensor* b = /* [1, 4] */;
ObjTensor* sum = tensor_add(a, b);  // result: [3, 4]

tensor_sub(a, b)

ObjTensor* diff = tensor_sub(a, b);

tensor_mul(a, b)

ObjTensor* prod = tensor_mul(a, b);

tensor_div(a, b)

ObjTensor* quot = tensor_div(a, b);

All four are generated by the TENSOR_BINARY_OP macro and follow the same pattern: promote dtypes, compute broadcast shape, create output, iterate and apply the operator element-wise via tensor_get_flat / tensor_set_flat.

tensor_pow(a, exp)

ObjTensor* squared = tensor_pow(a, 2.0);
ObjTensor* sqrt_ = tensor_pow(a, 0.5);

tensor_neg(a)

ObjTensor* neg = tensor_neg(a);

tensor_scalar_mul(a, scalar)

ObjTensor* scaled = tensor_scalar_mul(a, 2.5);

Reduction Operations

All reductions operate along a single axis and preserve dimensions (the reduced axis becomes size 1).

tensor_sum(t, axis)

// Sum along axis 0: shape [2, 3] -> [1, 3]
ObjTensor* s = tensor_sum(t, 0);

tensor_mean(t, axis)

Computes sum divided by the dimension size along the given axis.

ObjTensor* m = tensor_mean(t, 1);  // shape [2, 3] -> [2, 1]

tensor_max(t, axis)

Initialized to -1e9 and picks the maximum along the axis.

ObjTensor* mx = tensor_max(t, -1);  // last axis

tensor_min(t, axis)

Initialized to 1e9 and picks the minimum along the axis.

ObjTensor* mn = tensor_min(t, axis);

Math Operations

tensor_exp(t)

Element-wise exponential. Falls back to SIMD-accelerated simd_exp when available.

ObjTensor* e = tensor_exp(t);

tensor_rsqrt(t)

Element-wise reciprocal square root (1 / sqrt(x)). Uses simd_rsqrt for SIMD acceleration.

ObjTensor* r = tensor_rsqrt(t);  // 1 / sqrt(x)

tensor_log(t)

Element-wise natural logarithm (defined in tensor_nn_ops.c).

ObjTensor* l = tensor_log(t);

tensor_sqrt(t)

Element-wise square root with SIMD fallback.

ObjTensor* s = tensor_sqrt(t);

Type Conversion

The tensor_to_dtype function (used internally by binary ops) converts a tensor to a target dtype when needed.

Complete Example

// Mixed types and broadcasting
int64_t shape_a[] = {3, 1};
int64_t shape_b[] = {1, 4};

ObjTensor* a = tensor_create(2, shape_a, TENSOR_FLOAT32);
ObjTensor* b = tensor_create(2, shape_b, TENSOR_INT32);

// Fill data
for (int i = 0; i < a->size; i++) tensor_set_flat(a, i, i + 1);
for (int i = 0; i < b->size; i++) tensor_set_flat(b, i, (i + 1) * 2);

// Operations with auto-promotion and broadcasting
ObjTensor* sum = tensor_add(a, b);   // FLOAT32: [3, 4]
ObjTensor* prod = tensor_mul(a, b);  // FLOAT32: [3, 4]

// Reduction
ObjTensor* means = tensor_mean(prod, 0);  // [1, 4]

// Cleanup
tensor_free(a); tensor_free(b);
tensor_free(sum); tensor_free(prod);
tensor_free(means);