How it works
From your vendor to your analysts, every morning
It all runs in your AWS account, to a deadline you set.
Your deadline sets the schedule
Before your deadline
We start a run every so often through the early morning. Each one picks up whatever has arrived. An extra run does no harm.
Once the feeds pass
Your dbt models build as soon as their feeds pass.
Cutoff
A late or held feed doesn't delay the rest. At a set time before your deadline, your models build with its last good data.
Your deadline
The morning report goes out.
What happens to each file
We fetch it and keep a copy
We store the file exactly as sent and never change it.
We check it
Each row gets the time the source published it. Then the data is checked against the feed's rules, a YAML file in your GitHub.
It goes in, or it's held
Data that passes goes into your tables. If it fails, the feed is held and you get an alert.
When a check fails
Some checks only warn. The data goes in with a note in the report. The rest put the feed on hold. Nothing from that file reaches your tables, and your team keeps reading the last data that passed. Later files wait behind it, so a table never skips a delivery.
Nobody can push held data through, us included. If a check turns out to be wrong, we fix the feed's rules in a pull request and the file goes in on the next run.
Your dbt models are checked too
dbt then builds your research tables from your own models and our package, which covers the security master, identifier mapping and point-in-time joins. If a dbt test fails, your models are held and the research tables stay as they were.
The morning report
One message to Slack or email at your deadline, with a line per feed and one for your models. It's sent from your account, so it arrives even if our systems are down. You also get an alert when a problem starts.
✅ Prices published 06:02
⚠️ Estimates published 47 min late
❌ Alt data held since 05:40: 'score' changed from float to string. PR #12
⚠️ Models built at the cutoff, Alt data staleWhere your team reads the tables
The tables are Apache Iceberg on S3 in your account, catalogued in AWS Glue. Snowflake, Athena and DuckDB read them in place. Postgres on RDS or Aurora can't read Iceberg, so it gets a copy.
If Snowflake or Postgres falls behind, you get an alert. The tables in S3 are still current.