What this is
Most questions about data are answered by the data model: a metric, a join, a calculated column, computed on the fly. Nothing is copied and no code runs.
Some jobs do not fit that shape. Tasks have to be pulled from a tracker's API. A supplier sends a spreadsheet in a layout no connector expects. A forecast needs pandas, not one SQL statement. Until now, jobs like these went to a script on somebody's laptop, which stopped running when that person left.
A pipeline is that script moved into the platform: Python code that reads declared sources, transforms them in an isolated sandbox, and writes the result as a dataset. It runs when you press run, or on a schedule. The result is an ordinary dataset, so dashboards, spreadsheets, portal pages and agents read it like any other.
Why it matters
Nobody has to write it by hand. Tell the assistant "I want to pull tasks from the tracker and these Excel files". It asks what it needs, writes the declaration and the code, runs a dry run, and shows you the result. You approve it and set a schedule.
It is reviewed before it runs on its own. A version runs on a schedule only after it has been approved, and approval is tied to a successful dry run. A scheduled run is then compared with that dry run's shape and row count, so a pipeline that silently starts producing half the rows gets noticed.
It has an owner that is not a person's laptop. Every pipeline runs as an AI employee with its own access: it reads only the sources it has been granted, and writes only where it is allowed to write. Change of staff does not stop it.
What it can read
- Datasets already in the platform.
- APIs from the organisation's API catalogue, with a trial read shown step by step before you commit to it. Access to each endpoint is granted to the pipeline separately.
- Uploaded files - a fixed set of files already loaded into the platform.
- Incremental windows - read only what is new since the last run, and backfill history in measured portions, with a clear status of whether the pipeline has caught up.
What it writes
One target dataset per pipeline, fixed when the pipeline is created: either replaced on each run, or updated by key, row by row. A pipeline never writes to an external system - it brings data in, it does not push it out.
How it looks
A pipelines screen lists each pipeline with its versions, its schedule and its runs. Every run has its log. A failed run can be restarted, only one run of a pipeline is active at a time, and a scheduled tick that was skipped records why.
When to use it, and when not
| Use | When |
|---|---|
| A view or a metric in the data model | Aggregations, joins, renames, calculated columns - anything one query expresses |
| A pipeline | Data has to be fetched from somewhere, parsed from an odd format, or transformed with real code |
The view is the default; a pipeline is the next step when a view cannot do the job.
How it is built
A pipeline's code runs in a sandbox container with no network, a read-only file system, no privileges and fixed limits on memory, CPU and processes; in the cloud it runs under an additional kernel isolation layer. Only one component of the platform can start containers, and it accepts nothing from a caller but the names of registered images and profiles. The platform, not the code, fetches the sources and commits the result as Parquet into the data engine.
The Python environment is fixed: pandas, polars, pyarrow, DuckDB, NumPy, openpyxl and the platform's own SDK, installed from pinned, hash-checked packages. Time inside the sandbox comes from the SDK, so a dry run and a scheduled run of the same version see the same "today".
Scheduled runs are ordinary orchestration processes with a pipeline step, and every change to a pipeline goes to the audit log. Everything is available to agents as tools, and those that change something need a person's approval in chat.