T uses a first-class serializer system to manage data interchange between different runtimes (T, R, Python, Julia) and for materializing pipeline nodes as persistent artifacts.
Serializers are identified by the ^ prefix. You can
specify them when defining pipeline nodes:
p = pipeline {
-- Use the built-in Arrow IPC serializer for a DataFrame
data = node(command = read_csv("large.csv"), serializer = ^ipc)
-- Use the PMML serializer for a model
model = rn(command = <{ lm(y ~ x, data = data) }>, serializer = ^pmml)
-- Use the JSON serializer for a simple dictionary
config = node(command = { "debug": true, "retries": 5 }, serializer = ^json)
}
T distinguishes between built-in symbols and custom serializer variables:
^ipc, ^json,
etc.): Use the ^ prefix for T’s built-in
serializers. These are registered symbols that the pipeline emitter
understands natively across all supported runtimes.my_serializer): If you have
defined a custom serializer in a variable (e.g., a dictionary imported
from another file), pass the variable name without the
^ prefix. This allows the evaluator to pass the actual
serializer definition to the node.-- Built-in symbol (uses T's internal logic)
node(..., serializer = ^ipc)
-- Custom variable (passes the dictionary value)
import "src/my_ser.t" [my_ser]
node(..., serializer = my_ser)
[!IMPORTANT] String literals (e.g.,
serializer = "ipc") are strictly disallowed in node constructors (rn(),pyn(),jln(),shn(),qn(),node()). You must use either a symbol with the^prefix for built-ins or a variable name for custom serializers. Using a string literal in a node constructor will result in aTypeError.
mutate_node()andset_pipeline_global_options()accept both strings and symbols:mutate_node($serializer = "pmml")andset_pipeline_global_options(p, serializer = ^pmml)are both valid.
If you don’t specify a serializer, T uses the default
serializer, which selects each runtime’s native binary format for
in-language interchange (serialize for T,
saveRDS for R, pickle for Python, and Julia’s
Serialization package). For shell nodes, shn()
defaults to ^text.
| Identifier | Name | Best For | Write support | Read support | Notes |
|---|---|---|---|---|---|
^tlang |
T-Native | T-to-T interchange | T | T | Internal binary format |
^ipc |
Apache Arrow IPC | Pipeline intermediates, cross-runtime exchange, fastest round trips | T, R, Python, Julia | T, R, Python, Julia | Fully symmetric across all runtimes; uncompressed |
^parquet |
Apache Parquet | Long-term storage, archival, large datasets, external analytics tooling | T, R, Python, Julia | T, R, Python, Julia | Fully symmetric across all runtimes; compressed |
^csv |
CSV | Tabular data | T, R, Python, Julia | T, R, Python, Julia | Fully symmetric; R uses base
write.csv/read.csv |
^json |
JSON | Config, lists, dicts | T, R, Python, Julia | T, R, Python, Julia | Fully symmetric; Python uses stdlib |
^pmml |
PMML | Predictive Models | T, R, Python, Julia | T, R, Python, Julia | Julia writer: GLM.jl → PMML 4.4; Julia reader: JPMML evaluator via
JavaCall |
^onnx |
ONNX | ML Models | T, R, Python | T, R, Python, Julia | Julia: inference only (ONNXRunTime.jl); export is
experimental/limited |
^text |
Plain Text | Logs, shell output | All | All | Raw text, no format constraints |
^bin |
Binary | Passthrough, fetchurl | T | T | Opaque binary blob; default for fetchurl() nodes |
^ipc and ^parquetBoth ^ipc and ^parquet are columnar,
type-preserving Arrow formats that work symmetrically across every
runtime. The distinction is live hand-off vs. durable
artifact:
^ipc is the Arrow in-memory format
written straight to disk — the fastest possible write/read, but
uncompressed and with a limited ecosystem outside
Arrow.^parquet is a compressed,
storage-optimized layout — smaller files (typically
4–10× for numeric data), column pruning on read, and first-class support
in Spark, DuckDB, pandas, dask, BigQuery, and Athena.Rule of thumb: use ^ipc to pass data
between nodes while a pipeline runs; use ^parquet
for anything you persist, ship, or store. In one pipeline you
can do both — move intermediates with ^ipc and materialize
the final result to Parquet.
See Parquet vs Arrow IPC: How to Choose in the Data I/O guide for the full decision walkthrough.
serializer
StructureA serializer is a first-class object in T. You can inspect its properties or even define your own.
type serializer = {
format: string,
writer: function(path: string, value: any) -> result[NA, string],
reader: function(path: string) -> result[any, string]
}
You can create a custom serializer by defining a record that matches
the required interface. Note that the format field should
use a Symbol (starting with ^) to remain
consistent with T’s symbol-based serialization mandate.
my_log_serializer = {
format: ^log,
writer: \(path, val) {
-- custom logic to write log
Ok(NA)
},
reader: \(path) {
-- custom logic to read log
Ok("log content")
}
}
-- Usage: Pass the variable name (no ^ hat on the variable itself!)
node(command = ..., serializer = my_log_serializer)
For a complete example of a cross-language custom serializer (YAML),
see the Custom
Polyglot Serializer Demo in the t_demos repository.
One of the most powerful features of T’s serializer system is the static coherence check. When you build a pipeline, T verifies that the format produced by a source node matches the format expected by the consumer node.
node A {
target: wn("data.csv", serializer = ^csv)
}
node B {
source: rn("data.csv", serializer = ^ipc)
}
-- Result: Static Error
-- "Format mismatch: Node A produces ^csv, but Node B expects ^ipc."
This prevents runtime errors after long-running computations by catching interchange mismatches at the start of the build.
When you build a pipeline, T scans every node’s serializer and
runtime to determine which packages are needed, then checks
tproject.toml for those packages. If any are missing, T
prompts you with the exact
[r-dependencies], [py-dependencies], and
[jl-dependencies] entries to add before proceeding. You
must then run t update and re-enter
nix develop for the packages to become available. (Set
TLANG_AUTO_ADD_PIPELINE_DEPS=1 to skip the prompt in CI — T
auto-appends the missing entries and exits with instructions to rerun
the build.)
The table below shows which packages each format pulls in per runtime:
| Format | R packages | Python packages | Julia packages |
|---|---|---|---|
^csv |
(base R) | pandas |
CSV, DataFrames |
^ipc |
arrow |
pandas, pyarrow |
Arrow, DataFrames |
^parquet |
arrow |
pandas, pyarrow |
Parquet2, DataFrames |
^json |
jsonlite |
(stdlib) | JSON |
^pmml |
XML, jsonlite, r2pmml |
numpy, pandas, pyarrow,
scikit-learn, scipy,
sklearn2pmml, statsmodels |
GLM, JavaCall |
^onnx |
onnx |
onnxruntime, skl2onnx |
ONNXRunTime, ONNX |
^text |
(base R) | (stdlib) | (stdlib) |
^bin |
(none) | (none) | (none) |
default |
(none) | (stdlib pickle) | (stdlib Serialization) |
The ^pmml format also requires the jre
system tool for R, Python, and Julia nodes (for JPMML evaluator
execution). Add "jre" to
[additional-tools].packages in
tproject.toml.
For cross-language nodes, serializers provide the necessary glue code
for the target runtime. For example, when using ^ipc in an
R node:
arrow R library into the build
environment.arrow::write_ipc_file().For a serializer to work across non-T runtimes, it can optionally provide code snippets for R and Python. These snippets are strings that T injects into the generated build scripts.
You can define these by adding r_writer,
r_reader, py_writer, or py_reader
keys to your serializer dictionary. You can use standard strings or
foreign code blocks <{ ... }> for
better readability:
my_custom_ser = [
format: ^custom,
-- T implementation
writer: \(path, val) { Ok(NA) },
reader: \(path) { Ok(42) },
-- R snippets (using foreign code blocks)
r_writer: <{ function(obj, path) { saveRDS(obj, path) } }>,
r_reader: <{ function(path) { readRDS(path) } }>,
-- Python snippets
py_writer: <{ lambda obj, path: pickle.dump(obj, open(path, 'wb')) }>,
py_reader: <{ lambda path: pickle.load(open(path, 'rb')) }>
]
When T processes a node with an R runtime and the above
serializer: 1. It looks for the r_writer snippet. 2. It
generates a call in the node’s R script:
<r_writer>(node_result, "artifact_path").
If you use a custom format name (e.g.,
format: "myformat"), you should ensure that your R or
Python scripts have the necessary libraries loaded to handle that
format. You can do this by adding the libraries to your
tproject.toml or using the functions /
includes parameters in the node definition.
For ONNX specifically, Julia nodes read model artifacts through
ONNXRunTime.jl via the built-in jl_read_onnx()
helper. Julia ONNX export is not supported yet, so
jl_write_onnx() fails explicitly instead of silently
falling back to another format.