Use case · CI/CD for Data

Catch breaking changes before they hit production.

Software engineers don't ship code without tests, but most data teams ship transformations on faith. Limen brings CI to the pipeline: every contract is an executable test that runs on pull requests, scheduled jobs, and post-deploy checks.

  • Schema drift, null spikes, and broken joins get caught at the boundary instead of surfacing in a Monday morning dashboard.
  • The failure routes to the team that owns the data.
  • Every check produces a signed certificate: green to merge, red to block.
analytics-warehouse / pull / 482
analytics-warehouse · main ← feat/orders-v3

Add fulfilled_at to orders_fact

Open #482 · 3 commits · 14 files
Limen contract checks 6 of 7 passing · 1.4s
Pass orders_fact · schema 14 cols · drift 0
Pass orders_fact · row_count.delta +0.4% · within ±5%
Fail orders_fact · fulfilled_at.nullity 22.8% null · contract ≤ 1%
Pass orders_fact · joins(customers, items) no orphans
Pass revenue_daily · sla p95 fresh ≤ 09:00 UTC
Pass finance.ledger · invariants debit = credit ✓
Merge blocked fulfilled_at.nullity: 22.8% null exceeds contract bound (≤ 1%).
↳ routed to @team-orders · contract owner
PR
Runs on every pull request
cron
Runs on scheduled jobs
post-deploy
Runs after every deploy
↪ owner
Failures route to the right team
How it works

From "we'll catch it in review" to "the build is red."

Four mechanical steps. No new dashboard for anyone to ignore.

  1. Define the contract

    Schema, freshness, null bounds, distribution shape, referential keys, business invariants. Lives next to the table definition, in code.

    contract orders_fact {
      cols: 14
      fulfilled_at.nullity <= 1%
      row_count.delta   ±5%
      fresh_by                  09:00 UTC
    }
  2. Run on every change

    PRs trigger contract tests against a staging materialization. Scheduled jobs run the same tests against production. Same tests, same verdict logic, three trigger surfaces.

    pull_request schedule post_deploy
  3. Block at the boundary

    A contract failure blocks the merge or the deploy. The boundary is the contract, not the dashboard the next morning. No silent regressions reach downstream consumers.

    Merge blocked build #482 · 22.8% null
  4. Route to the owner

    Each contract has an owning team. Failures route there directly, not to the on-call who happened to notice. Owners see the failing assertion and the certificate of the last known-good run.

    team-orders → contract: orders_fact → last green: cert_3a4f7e
What gets caught

The four failure modes that ship by default.

Most pipeline regressions fall into the same handful of patterns. Contracts express each one as an assertion that the build can fail on, not a condition humans have to notice.

Schema drift

A column type changes upstream. The build catches the shape mismatch before any model that depends on the column gets refreshed.

Null spikes

A field that's been 99.7% populated for two years drops to 78% overnight. Limen flags the deviation against the historical contract bound.

Broken joins

A foreign key starts producing orphans. Referential contracts assert against both sides of the join, so the build goes red the first run that hits it.

SLA miss

The freshness deadline is part of the contract. A late table breaks the build for the consumers that depend on it, not for everyone in the company.

Get started

See how CI in your data pipeline can catch issues before they ship.

We'll connect to a single schema, generate baseline contracts from your historical data, and walk through how recent PRs would have been checked, so you can see where CI would flag breaking changes in your own pipeline.