Point-in-time data

A backtest should only see what was public at the time

Every row we store records when the source made it public.

A filing accepted at 4:05 p.m.

EDGAR dates a filing accepted after 5:30 p.m. ET to the next business day. So a filing accepted at 4:05 p.m. carries that day's date. A backtest that treats the date as known at the open trades on it hours early.

We use the filing's acceptance timestamp, acceptanceDateTime, instead.

Sources without a reliable timestamp

Some sources only give a date. For those we pick a cautious rule, such as treating each fact as known at the end of that day. If a source's timestamps can't be trusted, we use the time its file arrived.

Restated numbers are new rows

When a source revises a number, the revision becomes a new row with its own time. The original stays, and history goes back to the source's first record.

A few sources quietly overwrite old numbers instead. A check catches that and holds the feed. What we do about it.

Two ways to look at the past

We never delete an old version of a table.

What was public then

For each row, the latest version the source had published by that moment. It's what a backtest should read.

What your tables showed then

The table exactly as your team saw it then. Use it to reproduce an old report.

When the mistake is ours

Sometimes our parser reads a field wrong. We fix it and re-read the original files. The corrected rows go into a new version of the table, and the old version still shows what your team saw.