docs / stdlib / nucleo

Stdlib index
  1. Overview
  1. Captured-borrow migration worklist (3.3.3)
  2. Owned-bind migration worklist (8.2.7)
  3. Return-side title audit — the ride-through enumeration
  4. stdlib ownership audit
  5. stdlib ownership dispositions (plan 1.3.1 / 4.3.1)

codec

  1. Base64

codec / csv

  1. Csv

codec / json

  1. Json

collection

  1. ArrayList
  2. BPlusTree
  3. Cache
  4. Collectors
  5. HashMap
  6. HashSet
  7. Heap
  8. ImmutableList
  9. ImmutableMap
  10. ImmutableSet
  11. LinkedList
  12. RedBlackTree
  13. Sort

collection / ltm

  1. LtmBPlusTree

concurrent

  1. AtomicInt32
  2. AtomicInt64
  3. Channel
  4. FiberLocal
  5. Lock
  6. Mutex
  7. RwLock
  8. Semaphore
  9. Tasks

error

  1. Exception
  2. NoOptionalValueException
  3. RecoverableException
  4. Throwable
  5. UnrecoverableException

gfx

  1. Sampler
  2. Texture2D

hash

  1. Blake3
  2. DefaultHasher
  3. Hash
  4. MD5
  5. Sha1
  6. Sha256
  7. SipHash
  8. XXHash3

ifx

  1. BackendRegistry
  2. Window

io

  1. Buffer

io / file

  1. File
  2. FileInfo
  3. FileReader
  4. FileWriter
  5. Path
  6. Watcher

io / net

  1. IpAddress
  2. Server
  3. ServerBuilder
  4. SocketAddress
  5. TcpListener
  6. TcpStream
  7. UdpSocket

io / net / dns

  1. Dns

io / net / tls

  1. TlsConnection
  2. TlsListener

io / net / uri

  1. Uri
  2. UriBuilder

lang

  1. Guid
  2. Math
  3. Optional
  4. Pair
  5. Slice
  6. String
  7. StringBuilder

lang / stream

  1. ArrayStream
  2. Stream

math

  1. Camera
  2. Color
  3. DType
  4. Ray
  5. Rotation
  6. Tensor
  7. Transform

math / fft

  1. Fft

math / linalg

  1. LinAlg

math / npio

  1. Npy

math / poly

  1. Poly

math / random

  1. Generator

math / stats

  1. Stats

nucleo

  1. Columns — the Arrow-laid-out substrate
  2. Fused tensor expressions — Fuse
  3. Table — the lazy, typed dataframe
  4. Tape — define-by-run autograd
  5. Transform intrinsics — Grad, Vmap, Jit

process

  1. Command
  2. Process

reflect

  1. Class

search / distance

  1. Distance

search / fuzzy

  1. Matcher

search / ngram

  1. Index

session

  1. PackageInstallException
  2. Packages

time

  1. Clock
  2. DateTimeFormatter
  3. Duration
  4. Instant
  5. LocalDate
  6. LocalDateTime
  7. LocalTime
  8. Period
  9. ZonedDateTime
  10. ZoneId
  11. ZoneOffset

wire

  1. Compressor
  2. Decompressor
  3. Encoder
  4. Schema
  5. SchemaEncoder

xpu

  1. Device
  2. KernelBuffer
  3. KernelStream

xpu / mesh

  1. MeshSimplifier

Transform intrinsics — Grad, Vmap, Jit

cajeta.nucleo.transform — compile-time function transforms. Grad, Vmap, and Jit are compiler intrinsics, not methods: applied to a statically-known function (a lambda literal, or a static method through the annotation sugar), they produce a new function at compile time. There is no runtime flag, no global gradient state, and no tape here — the compiled transforms specialize and synthesize code. (The runtime, define-by-run counterpart is Tape.)

The importable types live beside the intrinsics: GradResult<V, G> is the typed {value, grads} return bag, and Transforms documents the contract.

Grad — differentiation

Grad(f) for a scalar-valued f : (P...) -> S yields (P...) -> GradResult<S, G>: the forward value and the gradient with respect to argument 0. Grad<N>(f) selects argument N.

package snip.grad;

import cajeta.nucleo.transform.GradResult;

public final class Demo {
    public static float32 run() {
        (float32) -> GradResult<float32,float32> g = Grad((float32 x) -> x * x);
        GradResult<float32,float32> r = g(3.0f);
        return r.value * 100.0f + r.grads;   // 9*100 + 6 = 906
    }
}

Tensor arguments differentiate the same way — reduce to a scalar with Tensor.sum or Tensor.mean:

(Tensor<float32>) -> GradResult<float32, Tensor<float32>> g =
    Grad((Tensor<float32> t) -> Tensor.sum<float32,float32>(Tensor.mul<float32>(t, t)));
GradResult<float32, Tensor<float32>> r = g(x);   // r.grads == 2*x

The differentiable primitive set: + - * / and unary -, the scalar intrinsics Math.exp / Math.log / Math.sqrt, and the tensor statics Tensor.{add,sub,mul,div,matmul,sum,mean,exp,log,sqrt,relu}. Anything outside the set is a named, located compile error — never a silently wrong gradient. (relu’s gradient is the Tensor.reluMask step function, which is itself non-differentiable — second-order Grad through relu fails loud.)

GradAll — one backward, K parameters

GradAll(f) for a scalar-valued f : (P...) -> S yields (P...) -> GradResult<S, G[]>: the forward value and the gradients with respect to every parameter, as an array in argument order. GradAll<K>(f) grades only the leading K arguments — the functional-step convention: parameters first, data arguments after. One call returns all the grads; this is the form a training step uses (grads for each weight in one invocation).

(Tensor<float32>, Tensor<float32>, Tensor<float32>)
        -> GradResult<float32, Tensor<float32>[]> step =
    GradAll<2>((Tensor<float32> w, Tensor<float32> b, Tensor<float32> x) ->
        Tensor.sum<float32,float32>(
            Tensor.add<float32>(Tensor.matmul<float32>(x, w), b)));
GradResult<float32, Tensor<float32>[]> r = step(w, b, x);
// r.grads[0] = dL/dw, r.grads[1] = dL/db — x is data, not graded

The K differentiated parameters must share one type (they fill one array); mixing Tensor<float32> and float32 in the leading K is a named compile error, as are K < 1 and K > arity. Jit(GradAll<K>(f)) fuses; @NoGrad helpers are constants, exactly as under Grad.

  • Gradients are explicit return values. Calling g twice gives two independent results; there is no .grad accumulator and no zero_grad.
  • Grad differentiates through same-class static helpers (they inline into the gradient); a helper marked @NoGrad is a constant instead — its value flows, its gradient is statically zero.
  • Second order: Grad(Grad(f)) for scalar f — the backward is ordinary differentiable source, so it differentiates again.

Vmap — batching

Vmap(f) for f : (T) -> R yields (T[]) -> #R[]: f written for one example runs over a leading batch axis, no hand-written loop. Composed with Grad, it gives per-example gradients:

float32[] xs = {1.0f, 2.0f, 3.0f};
(float32[]) -> #GradResult<float32,float32>[] g =
    Vmap(Grad((float32 x) -> x * x));
GradResult<float32,float32>[] rs = g(xs);        // grads {2, 4, 6}

An op with no batching rule is a named compile error. Order matters and is not normalized: Grad(Vmap(f)) asks for the gradient of an array-valued function, which the scalar-valued Grad rejects.

Jit — fusion

Jit(f) has f’s exact signature and result; the body is fused (temporaries eliminated) by the standard optimization pipeline. It composes over the other transforms — Jit(Vmap(Grad(f))) fuses the batched, differentiated form into one body:

(float32[]) -> #GradResult<float32,float32>[] g =
    Jit(Vmap(Grad((float32 x) -> x * x)));       // same signature, fused

Annotation sugar

@Grad / @Vmap / @Jit on a static single-return method desugar to the nested combinator form. The annotation nearest the declaration applies first, so the stack reads like the nesting:

@Jit @Vmap @Grad
public static float32 sq(float32 x) { return x * x; }
// calls to sq are now Jit(Vmap(Grad(sq))):
GradResult<float32,float32>[] rs = T.sq(xs);     // per-example grads, fused

The sugar is only sugar — it runs the same driver as the explicit form, so compositions and error messages are identical.

Errors

Every misuse is named and located: a runtime-only function value is CAJETA_ERROR_TRANSFORM_NOT_SPECIALIZABLE; an op outside the differentiable or batchable set names the operator; a tensor rank mismatch names the ranks and dtype (Tensor.matmul<float32> rank mismatch: left operand is tensor, right operand is scalar). Pmap is recognized but reserved (CAJETA_ERROR_TRANSFORM_PMAP_UNIMPLEMENTED).

Source: docs/stdlib/nucleo/Transforms.md · 3 min read · 597 words