What goes wrong

When data goes wrong

Vendor data breaks in a handful of familiar ways. None of them should reach your models, and none of them should be a surprise at 9 a.m.

A morning with a bad file

An example of how one broken delivery plays out.

2:10 a.m.

Your alt-data vendor ships a file where the score column has turned from numbers into text.

5:40 a.m.

The file fails its checks and the feed is held. Nothing from it reaches your tables.

6:00 a.m.

Your analysts start the day on yesterday's alt data. Everything else is current.

6:30 a.m.

Your morning report says which column changed and links the pull request that fixes it.

Later

The fix is tested against your recent files and merged. The next run brings the feed current.

✅ Prices        published 06:02
✅ Fundamentals  published 05:31
❌ Alt data      held since 05:40: 'score' changed from number to text. PR #12
⚠️ Models        built at the cutoff, Alt data stale
The morning report from that example.

Bad files

A renamed column, a changed type, a file that repeats the same records. Each feed has written rules for what its files must look like, and a file that breaks them is held before anyone reads it.

We fix the cause in a pull request: the feed's rules if the vendor changed its format, our parser if the mistake is ours.

Numbers that look wrong

A file with far fewer rows than usual, or values far from their history. Often it's a unit change, like thousands instead of dollars. Sometimes the market really did move.

Because real moves happen, these checks can let the data through with a flag in your report. Each feed's rules say which checks flag and which hold. Either way, we check with the vendor and fix the rule or the units.

Late or missing data

A vendor misses its schedule or its servers are down. One late feed doesn't hold up the rest. At a set time before your deadline, your models build with that feed's last good data, and the report marks it stale.

If the file turns up before your deadline, it still goes in. We retry, and chase the vendor if it keeps happening.

History that changes

Sources revise numbers they've already published. A revision goes in as a new row with its own timestamp, next to the original, so a backtest only sees it from when it was published.

A few sources quietly overwrite old numbers instead. That fails a check and the feed is held. We raise it with the source, and if it keeps happening, we treat each change as new from the moment it arrived. More on point-in-time data.

Tickers and splits

A ticker gets reassigned and two feeds disagree about which company is which. A split shows up in prices before it reaches per-share fundamentals.

Your models are held until the mapping or the split date is right, while your feed tables keep updating. We fix it in a pull request and rebuild.

On your side

A vendor key expires, or Snowflake or Postgres falls behind. You get an alert that names it. We renew tokens where the vendor allows it. We can't read your secrets, so for a fixed key we tell you which one to replace. We retry database updates, and your tables in S3 stay current either way.