docs / stdlib / nucleo

Stdlib index
  1. Overview
  1. Captured-borrow migration worklist (3.3.3)
  2. Owned-bind migration worklist (8.2.7)
  3. Return-side title audit — the ride-through enumeration
  4. stdlib ownership audit
  5. stdlib ownership dispositions (plan 1.3.1 / 4.3.1)

codec

  1. Base64

codec / csv

  1. Csv

codec / json

  1. Json

collection

  1. ArrayList
  2. BPlusTree
  3. Cache
  4. Collectors
  5. HashMap
  6. HashSet
  7. Heap
  8. ImmutableList
  9. ImmutableMap
  10. ImmutableSet
  11. LinkedList
  12. RedBlackTree
  13. Sort

collection / ltm

  1. LtmBPlusTree

concurrent

  1. AtomicInt32
  2. AtomicInt64
  3. Channel
  4. FiberLocal
  5. Lock
  6. Mutex
  7. RwLock
  8. Semaphore
  9. Tasks

error

  1. Exception
  2. NoOptionalValueException
  3. RecoverableException
  4. Throwable
  5. UnrecoverableException

gfx

  1. Sampler
  2. Texture2D

hash

  1. Blake3
  2. DefaultHasher
  3. Hash
  4. MD5
  5. Sha1
  6. Sha256
  7. SipHash
  8. XXHash3

ifx

  1. BackendRegistry
  2. Window

io

  1. Buffer

io / file

  1. File
  2. FileInfo
  3. FileReader
  4. FileWriter
  5. Path
  6. Watcher

io / net

  1. IpAddress
  2. Server
  3. ServerBuilder
  4. SocketAddress
  5. TcpListener
  6. TcpStream
  7. UdpSocket

io / net / dns

  1. Dns

io / net / tls

  1. TlsConnection
  2. TlsListener

io / net / uri

  1. Uri
  2. UriBuilder

lang

  1. Guid
  2. Math
  3. Optional
  4. Pair
  5. Slice
  6. String
  7. StringBuilder

lang / stream

  1. ArrayStream
  2. Stream

math

  1. Camera
  2. Color
  3. DType
  4. Ray
  5. Rotation
  6. Tensor
  7. Transform

math / fft

  1. Fft

math / linalg

  1. LinAlg

math / npio

  1. Npy

math / poly

  1. Poly

math / random

  1. Generator

math / stats

  1. Stats

nucleo

  1. Columns — the Arrow-laid-out substrate
  2. Fused tensor expressions — Fuse
  3. Table — the lazy, typed dataframe
  4. Tape — define-by-run autograd
  5. Transform intrinsics — Grad, Vmap, Jit

process

  1. Command
  2. Process

reflect

  1. Class

search / distance

  1. Distance

search / fuzzy

  1. Matcher

search / ngram

  1. Index

session

  1. PackageInstallException
  2. Packages

time

  1. Clock
  2. DateTimeFormatter
  3. Duration
  4. Instant
  5. LocalDate
  6. LocalDateTime
  7. LocalTime
  8. Period
  9. ZonedDateTime
  10. ZoneId
  11. ZoneOffset

wire

  1. Compressor
  2. Decompressor
  3. Encoder
  4. Schema
  5. SchemaEncoder

xpu

  1. Device
  2. KernelBuffer
  3. KernelStream

xpu / mesh

  1. MeshSimplifier

Fused tensor expressions — Fuse

cajeta.nucleo.expr — compile-time fusion of elementwise tensor expressions. Fuse is a compiler intrinsic, not a method: applied to a statically-known tensor expression (a lambda literal, or a static method through @Fuse), it collapses N tensor operators into one loop at compile time. NumPy allocates a fresh tensor per operator; a fused expression reads each input element once, writes each output element once, and allocates only the result.

Fuse — one kernel, no temporaries

(Tensor<float32>) -> #Tensor<float32> g =
    Fuse((Tensor<float32> t) ->
        Tensor.sub<float32>(Tensor.add<float32>(Tensor.mul<float32>(t, t), t), t));
Tensor<float32> r = g(x);    // ONE pass; allocates only r

Building the fused function runs nothing and allocates nothing; the call is the force point — there is no .eval. The returned function is reusable: compile once, call per batch.

The fusible elementwise set: Tensor.{add,sub,mul,div}, the scalar-broadcast family Tensor.{add,sub,mul,div}Scalar, negate, and exp / log / sqrt. A shared sub-expression is computed once.

Reductions stage

A reduction (Tensor.sum / mean / std) bounds the fusion region: it runs as its own pass, and the elementwise tail fuses against its scalar result. The standardize shape is the headline:

(Tensor<float32>) -> #Tensor<float32> g =
    Fuse((Tensor<float32> t) ->
        Tensor.divScalar<float32>(
            Tensor.subScalar<float32>(t, Tensor.mean<float32,float32>(t)),
            Tensor.std<float32,float32>(t, 0)));
// mean and std stage once each; (t - mean) / std is one fused loop

A reduction as the whole expression fuses the elementwise body into the accumulation itself — Fuse(t -> Tensor.sum(Tensor.mul(t, t))) is scalar-valued and allocates no tensor at all.

@Fuse — the everyday shape

@Fuse
public static Tensor<float32> activate(Tensor<float32> t) {
    return Tensor.add<float32>(Tensor.mul<float32>(t, t), t);
}

Calls to the annotated method produce the fused result — same driver as the explicit form, identical values, identical allocation counts.

The autograd seam

Fuse and Grad consume the same forward DAG, so they compose without a translation layer:

(Tensor<float32>) -> GradResult<float32, Tensor<float32>> g =
    Grad(Fuse((Tensor<float32> t) ->
        Tensor.sum<float32,float32>(Tensor.mul<float32>(t, t))));
// identical gradient to the unfused form; @Grad @Fuse is the sugar spelling

The backward of a fused expression is ordinary IR — Jit fuses it like any other function — and an expression that is never differentiated carries no autograd machinery in its emitted code.

Errors

A body that cannot fuse is the named, located compile error CAJETA_ERROR_TRANSFORM_NOT_FUSIBLE — never a silent eager fallback: an op outside the elementwise set (matmul), or std as the whole fused expression (usable inside an elementwise body). A runtime-only function value is CAJETA_ERROR_TRANSFORM_NOT_SPECIALIZABLE, as for every transform intrinsic.

Source: docs/stdlib/nucleo/Fuse.md · 1 min read · 323 words