Skip to content

Data pipelines that don’t break at 3am.

From a one-table MVP to a multi-petabyte lakehouse. We build the streaming and batch pipelines, the schemas, the orchestration, the tests, and the dashboards — so your team can ship product, not babysit data jobs.

Rate $100 – $120 USD / hour · fixed-fee and retainer available

What you get

Reference architecture
a written document with the data flow, ownership, SLAs, and a one-page recovery runbook for each pipeline.
Batch pipelines
Spark, dbt, Airflow, Dagster. Idempotent, retryable, with a real test suite.
Streaming pipelines
Kafka, Kinesis, Pub/Sub. Schema registry, exactly-once where it matters, dead-letter handling everywhere else.
Lakehouse & warehouse
BigQuery, Snowflake, ClickHouse, Iceberg, Delta. Cost dashboards, partition strategy, retention policy.
Reverse ETL
Hightouch / Census / Airbyte, with a written contract for each sync.
Observability
OpenLineage / DataHub / Monte Carlo for lineage, freshness, and volume anomalies.

How we engage

  1. 01

    Data audit (1–2 weeks, fixed fee)

    We instrument your existing pipelines, find the silent failures, and write the remediation plan.

  2. 02

    Reference build (4–6 weeks)

    One well-architected pipeline end-to-end — the pattern your team will copy.

  3. 03

    Embedded senior (retainer)

    A named data engineer in your Slack, in your standups, reviewing PRs.

Stack we work in

Batch

Spark, dbt, Airflow, Dagster, Beam

Streaming

Kafka, Kinesis, Pub/Sub, Pulsar, Redpanda

Warehouse

BigQuery, Snowflake, Redshift, Databricks, ClickHouse

Lakehouse

Iceberg, Delta, Hudi, Apache Paimon

Orchestration

Airflow, Dagster, Prefect, Kestra

Quality

Great Expectations, Soda, dbt tests, Monte Carlo

Lineage

OpenLineage, DataHub, Atlas, Marquez

Reference architectures

Anonymized patterns from real engagements. Client names omitted; details available under NDA.

Series A fintech — 8B rows/day

Rebuilt a Spark-on-Kafka pipeline that was silently dropping 0.4% of trades. Added schema registry, dead-letter queues, and a real test suite. Audit cleared in 6 weeks.

Healthtech — HIPAA lakehouse

Designed a Snowflake + dbt lakehouse with row-level security, PHI tagging, and BAA-compliant data sharing for hospital partners.

Logistics SaaS — real-time fleet analytics

Kafka → ClickHouse pipeline powering a live fleet dashboard. p99 query latency < 200ms on 12 months of historical data.

Questions we get asked

Do you do data modeling or just pipelines?

Both. We work at the schema layer (entities, contracts, lineage) and at the pipeline layer (ingestion, transformation, serving). A pipeline with a bad schema is a liability.

How do you handle schema changes?

Schema registry, versioned contracts, and a backfill plan written before the migration. We don’t ship breaking changes on a Friday.

What’s your take on lakehouse vs warehouse?

Depends on the workload. Most startups do not need a lakehouse; a well-modeled warehouse with a clean dbt project gets you 90% of the way for 10% of the complexity. We tell you when you actually need the extra weight.

Do you work with dbt?

Yes — dbt is the default for our transformation layer. We write dbt projects that a junior analyst can extend six months later.

Can you take over a half-built pipeline?

Yes. We start with a 1-week audit (no code written), and you get a written report of what’s working, what’s broken, and what to do next. No obligation to continue.

Let’s scope it properly.

A 30-minute call. No deck, no pitch — we read your repo or your architecture diagram and tell you what’s realistic.