Serializers in T

T uses a first-class serializer system to manage data interchange between different runtimes (T, R, Python, Julia) and for materializing pipeline nodes as persistent artifacts.

1. Using Serializers

Serializers are identified by the ^ prefix. You can specify them when defining pipeline nodes:

p = pipeline {
  -- Use the built-in Arrow IPC serializer for a DataFrame
  data = node(command = read_csv("large.csv"), serializer = ^ipc)
  
  -- Use the PMML serializer for a model
  model = rn(command = <{ lm(y ~ x, data = data) }>, serializer = ^pmml)
  
  -- Use the JSON serializer for a simple dictionary
  config = node(command = { "debug": true, "retries": 5 }, serializer = ^json)
}

Symbols vs. Variables

T distinguishes between built-in symbols and custom serializer variables:

-- Built-in symbol (uses T's internal logic)
node(..., serializer = ^ipc)

-- Custom variable (passes the dictionary value)
import "src/my_ser.t" [my_ser]
node(..., serializer = my_ser)

[!IMPORTANT] String literals (e.g., serializer = "ipc") are strictly disallowed in strategy positions (rn(), pyn(), jln(), shn(), qn(), node(), mutate_node(), set_pipeline_global_options(), fetchurl()). You must use either a symbol with the ^ prefix for built-ins or a strategy dict for custom formats. Using a string literal will result in a TypeError.

Implicit Serialization

If you don’t specify a serializer, T uses the default serializer, which selects each runtime’s native binary format for in-language interchange (serialize for T, saveRDS for R, pickle for Python, and Julia’s Serialization package). For shell nodes, shn() defaults to ^text.

2. Built-in Serializers

Identifier Name Best For Write support Read support Notes
^tlang T-Native T-to-T interchange T T Internal binary format
^ipc Apache Arrow IPC Pipeline intermediates, cross-runtime exchange, fastest round trips T, R, Python, Julia T, R, Python, Julia Fully symmetric across all runtimes; uncompressed
^parquet Apache Parquet Long-term storage, archival, large datasets, external analytics tooling T, R, Python, Julia T, R, Python, Julia Fully symmetric across all runtimes; compressed
^csv CSV Tabular data T, R, Python, Julia T, R, Python, Julia Fully symmetric; R uses base write.csv/read.csv
^json JSON Config, lists, dicts T, R, Python, Julia T, R, Python, Julia Fully symmetric; Python uses stdlib
^pmml PMML Predictive Models T, R, Python, Julia T, R, Python, Julia Julia writer: GLM.jl → PMML 4.4; Julia reader: JPMML evaluator via JavaCall
^onnx ONNX ML Models T, R, Python T, R, Python, Julia Julia: inference only (ONNXRunTime.jl); export is experimental/limited
^text Plain Text Logs, shell output All All Raw text, no format constraints
^bin Binary Passthrough, fetchurl T T Opaque binary blob; default for fetchurl() nodes

Note: ^arrow was renamed to ^ipc in 0.55.0 with no alias. Unknown formats (anything outside this table, default, and custom strategies backed by the node’s functions) fail validation with a TypeError naming valid formats instead of failing at build time.

Choosing Between ^ipc and ^parquet

Both ^ipc and ^parquet are columnar, type-preserving Arrow formats that work symmetrically across every runtime. The distinction is live hand-off vs. durable artifact:

Rule of thumb: use ^ipc to pass data between nodes while a pipeline runs; use ^parquet for anything you persist, ship, or store. In one pipeline you can do both — move intermediates with ^ipc and materialize the final result to Parquet.

See Parquet vs Arrow IPC: How to Choose in the Data I/O guide for the full decision walkthrough.

3. The serializer Structure

A serializer is a first-class object in T. You can inspect its properties or even define your own.

type serializer = {
  format: string,
  writer: function(path: string, value: any) -> result[NA, string],
  reader: function(path: string) -> result[any, string]
}

Custom Serializers

You can create a custom serializer with a strategy dict. The dict has closed keys: format (always present, a ^-prefixed symbol) plus inline <{ ... }> code snippets per runtime (r_writer, r_reader, py_writer, py_reader, julia_writer, julia_reader). t check and pipeline validation enforce the shape: unknown keys, a missing format, non-code snippets, and custom formats without a snippet for the node’s runtime are all errors. Custom formats are not supported on T nodes (builtins only) or on sh and fetchurl nodes, and Quarto nodes take no serializer or deserializer at all.

my_log_serializer = [
  format: ^log,
  r_writer: <{ function(obj, path) writeLines(obj, path) }>,
  r_reader: <{ function(path) readLines(path) }>,
  py_writer: <{ lambda obj, path: open(path, 'w').write(str(obj)) }>,
  py_reader: <{ lambda path: open(path).read() }>
]

-- Usage: pass the variable name (no ^ hat on the variable itself!)
node(command = ..., serializer = my_log_serializer)

For a complete example of a cross-language custom serializer (YAML), see the Custom Polyglot Serializer Demo in the t_demos repository.

Per-Dependency Maps

A dict without a format key is a per-dependency map ([reader: ^csv, writer: ^json]): each key must name a real dependency of the node, otherwise validation fails naming the valid set. Two naming rules apply. A map with a format key is always a strategy dict, never a per-dependency map — even when format is also a dependency name, so a dependency literally named format cannot be keyed in map form (rename it). Map keys track pattern expansion: after expand_pipeline renames branch dependencies (mid → mid_branch_1), each branch entry keys the renamed dependency, so the chosen strategies keep applying instead of falling back to default.

4. Static Coherence Checks

One of the most powerful features of T’s serializer system is the static coherence check. When you build a pipeline, T verifies that the format produced by a source node matches the format expected by the consumer node.

node A {
  target: wn("data.csv", serializer = ^csv)
}

node B {
  source: rn("data.csv", serializer = ^ipc)
}

-- Result: Static Error
-- "Format mismatch: Node A produces ^csv, but Node B expects ^ipc."

This prevents runtime errors after long-running computations by catching interchange mismatches at the start of the build.

5. Serializer Runtime Dependencies

When you build a pipeline, T scans every node’s serializer and runtime to determine which packages are needed, then checks tproject.toml for those packages. If any are missing, T prompts you with the exact [r-dependencies], [py-dependencies], and [jl-dependencies] entries to add before proceeding. You must then run t update and re-enter nix develop for the packages to become available. (Set TLANG_AUTO_ADD_PIPELINE_DEPS=1 to skip the prompt in CI — T auto-appends the missing entries and exits with instructions to rerun the build. Per-command equivalents: t run --yes <file.t> answers yes, t run --no <file.t> declines; --no always wins, and --yes still requires tproject.toml to exist.)

The table below shows which packages each format pulls in per runtime:

Format R packages Python packages Julia packages
^csv (base R) pandas CSV, DataFrames
^ipc arrow pandas, pyarrow Arrow, DataFrames
^parquet arrow pandas, pyarrow Parquet2, DataFrames
^json jsonlite (stdlib) JSON
^pmml XML, jsonlite, r2pmml numpy, pandas, pyarrow, scikit-learn, scipy, sklearn2pmml, statsmodels GLM, JavaCall
^onnx onnx onnxruntime, skl2onnx ONNXRunTime, ONNX
^text (base R) (stdlib) (stdlib)
^bin (none) (none) (none)
default (none) (stdlib pickle) (stdlib Serialization)

The ^pmml format also requires the jre system tool for R, Python, and Julia nodes (for JPMML evaluator execution). Add "jre" to [additional-tools].packages in tproject.toml.

6. Polyglot Support

For cross-language nodes, serializers provide the necessary glue code for the target runtime. For example, when using ^ipc in an R node:

  1. T injects the arrow R library into the build environment.
  2. T generates the R code to call arrow::write_ipc_file().
  3. T ensures the resulting file is correctly tracked as a Nix artifact.

Custom Polyglot Serializers: R and Python Snippets

For a serializer to work across non-T runtimes, it can optionally provide code snippets for R and Python. These snippets are strings that T injects into the generated build scripts.

You can define these by adding r_writer, r_reader, py_writer, or py_reader keys to your serializer dictionary. You can use standard strings or foreign code blocks <{ ... }> for better readability:

my_custom_ser = [
  format: ^custom,
  
  -- T implementation
  writer: \(path, val) { Ok(NA) },
  reader: \(path) { Ok(42) },
  
  -- R snippets (using foreign code blocks)
  r_writer: <{ function(obj, path) { saveRDS(obj, path) } }>,
  r_reader: <{ function(path) { readRDS(path) } }>,
  
  -- Python snippets
  py_writer: <{ lambda obj, path: pickle.dump(obj, open(path, 'wb')) }>,
  py_reader: <{ lambda path: pickle.load(open(path, 'rb')) }>
]

Injected Code Patterns

When T processes a node with an R runtime and the above serializer: 1. It looks for the r_writer snippet. 2. It generates a call in the node’s R script: <r_writer>(node_result, "artifact_path").

Registering Custom Formats

If you use a custom format name (e.g., format: "myformat"), you should ensure that your R or Python scripts have the necessary libraries loaded to handle that format. You can do this by adding the libraries to your tproject.toml or using the functions / includes parameters in the node definition.

For ONNX specifically, Julia nodes read model artifacts through ONNXRunTime.jl via the built-in jl_read_onnx() helper. Julia ONNX export is not supported yet, so jl_write_onnx() fails explicitly instead of silently falling back to another format.

Python scikit-learn export stamps opset 21 (target_opset=21 in convert_sklearn). Newer skl2onnx defaults to opset 22, which ONNXRunTime rejects (official support ends at 21), so the pin keeps artifacts loadable in every supported consumer, including T-native predict.


Next Steps

  1. Pipeline Tutorial — Learn how pipelines use serializers for polyglot data interchange.
  2. Data I/O & Formats — Read and write CSV, Parquet, and Arrow IPC files; download data from URLs; see how to choose between Parquet and IPC.
  3. Project Development — Declare runtime dependencies so serializer packages are available at build time.