How it works

From your vendor to your analysts, every morning

It all runs in your AWS account, to a deadline you set.

Your deadline sets the schedule

Before your deadline

We start a run every so often through the early morning. Each one picks up whatever has arrived. An extra run does no harm.

Once the feeds pass

Your dbt models build as soon as their feeds pass.

Cutoff

A late or held feed doesn't delay the rest. At a set time before your deadline, your models build with its last good data.

Your deadline

The morning report goes out.

What happens to each file

  1. We fetch it and keep a copy

    We store the file exactly as sent and never change it.

  2. We check it

    Each row gets the time the source published it. Then the data is checked against the feed's rules, a YAML file in your GitHub.

  3. It goes in, or it's held

    Data that passes goes into your tables. If it fails, the feed is held and you get an alert.

When a check fails

Some checks only warn. The data goes in with a note in the report. The rest put the feed on hold. Nothing from that file reaches your tables, and your team keeps reading the last data that passed. Later files wait behind it, so a table never skips a delivery.

Nobody can push held data through, us included. If a check turns out to be wrong, we fix the feed's rules in a pull request and the file goes in on the next run.

Your dbt models are checked too

dbt then builds your research tables from your own models and our package, which covers the security master, identifier mapping and point-in-time joins. If a dbt test fails, your models are held and the research tables stay as they were.

The morning report

One message to Slack or email at your deadline, with a line per feed and one for your models. It's sent from your account, so it arrives even if our systems are down. You also get an alert when a problem starts.

✅ Prices             published 06:02
⚠️ Estimates          published 47 min late
❌ Alt data           held since 05:40: 'score' changed from float to string. PR #12
⚠️ Models             built at the cutoff, Alt data stale
An example. Feed names and times are made up.

Where your team reads the tables

The tables are Apache Iceberg on S3 in your account, catalogued in AWS Glue. Snowflake, Athena and DuckDB read them in place. Postgres on RDS or Aurora can't read Iceberg, so it gets a copy.

If Snowflake or Postgres falls behind, you get an alert. The tables in S3 are still current.