# Python pipelines

> When SQL is not enough - Python code that pulls data from an API, files or other datasets, transforms it in an isolated sandbox, and lands the result as a dataset, on demand or on a schedule

Source: https://aipril.dev/docs/engine/compute

## What this is

Most questions about data are answered by the [data model](https://aipril.dev/docs/engine/data-model.md): a metric, a
join, a calculated column, computed on the fly. Nothing is copied and no code runs.

Some jobs do not fit that shape. Tasks have to be pulled from a tracker's API. A supplier sends a
spreadsheet in a layout no connector expects. A forecast needs pandas, not one SQL statement.
Until now, jobs like these went to a script on somebody's laptop, which stopped running when that
person left.

A **pipeline** is that script moved into the platform: Python code that reads declared sources,
transforms them in an isolated sandbox, and writes the result as a dataset. It runs when you press
run, or on a schedule. The result is an ordinary dataset, so dashboards, spreadsheets, portal pages
and agents read it like any other.

## Why it matters

**Nobody has to write it by hand.** Tell the assistant "I want to pull tasks from the tracker and
these Excel files". It asks what it needs, writes the declaration and the code, runs a dry run, and
shows you the result. You approve it and set a schedule.

**It is reviewed before it runs on its own.** A version runs on a schedule only after it has been
approved, and approval is tied to a successful dry run. A scheduled run is then compared with that
dry run's shape and row count, so a pipeline that silently starts producing half the rows gets
noticed.

**It has an owner that is not a person's laptop.** Every pipeline runs as an
[AI employee](https://aipril.dev/docs/ai-organisation.md) with its own access: it reads only the sources it has been
granted, and writes only where it is allowed to write. Change of staff does not stop it.

## What it can read

- **Datasets** already in the platform.
- **APIs** from the organisation's API catalogue, with a trial read shown step by step before you
  commit to it. Access to each endpoint is granted to the pipeline separately.
- **Uploaded files** - a fixed set of files already loaded into the platform.
- **Incremental windows** - read only what is new since the last run, and backfill history in
  measured portions, with a clear status of whether the pipeline has caught up.

## What it writes

One target dataset per pipeline, fixed when the pipeline is created: either **replaced** on each
run, or **updated by key**, row by row. A pipeline never writes to an external system - it brings
data in, it does not push it out.

## How it looks

A pipelines screen lists each pipeline with its versions, its schedule and its runs. Every run has
its log. A failed run can be restarted, only one run of a pipeline is active at a time, and a
scheduled tick that was skipped records why.

## When to use it, and when not

| Use | When |
|---|---|
| A view or a metric in the data model | Aggregations, joins, renames, calculated columns - anything one query expresses |
| A pipeline | Data has to be fetched from somewhere, parsed from an odd format, or transformed with real code |

The view is the default; a pipeline is the next step when a view cannot do the job.

## How it is built

A pipeline's code runs in a **sandbox container with no network**, a read-only file system, no
privileges and fixed limits on memory, CPU and processes; in the cloud it runs under an additional
kernel isolation layer. Only one component of the platform can start containers, and it accepts
nothing from a caller but the names of registered images and profiles. The platform, not the code,
fetches the sources and commits the result as Parquet into the data engine.

The Python environment is fixed: pandas, polars, pyarrow, DuckDB, NumPy, openpyxl and the
platform's own SDK, installed from pinned, hash-checked packages. Time inside the sandbox comes from
the SDK, so a dry run and a scheduled run of the same version see the same "today".

Scheduled runs are ordinary [orchestration](https://aipril.dev/docs/processes/automation.md) processes with a pipeline
step, and every change to a pipeline goes to the [audit log](https://aipril.dev/docs/platform/audit.md). Everything is
available to agents as tools, and those that change something need a person's approval in chat.
