Skip to content

Archive

Postgres keeps the second-level rows of the top-100 for days, then only minute candles and rollups remain. A backtest wants months of trades, and it reads them end to end rather than at a point in time. The archive is that history as files: one Parquet file per listing, layer and UTC day, written the day after the day is over, checked against the storage it was written to, and listed and fetched through the API.

The principle is the platform’s own: stored as seen. A file holds exactly the rows the platform observed. Nothing is filled, nothing is interpolated, no row is repaired into a smoother series. Where the platform was not listening, the file has no rows, and the gaps of the day are in the list beside it.

group layers in the database in the archive
core trade, ticker, mark, depth-derived forever forever
top-100 trade, ticker, mark, depth-derived days: see Retention from the day after each day is over, with no end
tail none thirty days not archived

Funding, open interest, statistics, candles and liquidations are not archived: they are small and are read at a point in time, so they stay in the database for every group. A day’s row in the database is dropped only after its file has been read back from the storage and its row count and hash are what was written. There is no day that is neither in the database nor in the archive because a copy failed.

depth-derived is the book’s derived figures (mid, spread, liquidity in the bands around the mid and the value resting in them, the number of levels), not the book itself: levels are never stored.

Terminal window
curl -s -H "Authorization: Bearer $DEBYKO_KEY" \
"https://api.debyko.com/v2/archive?listing=BINANCE-USDM:BTCUSDT&layer=trade&from=2026-09-28&to=2026-09-30"
{
"listing": "BINANCE-USDM:BTCUSDT", "from": "2026-09-28", "to": "2026-09-30", "truncated": false,
"items": [
{ "layer": "trade", "day": "2026-09-30", "group": "core", "available": true, "rows": 4182113, "bytes": 187004211,
"sha256": "9f2c…e1", "gaps": 1, "anchored": false,
"download": "/v2/archive/BINANCE-USDM%3ABTCUSDT/trade/2026-09-30",
"coverage": [ { "cause": "ws_disconnected", "start": "2026-09-30T03:12:09.114000Z", "end": "2026-09-30T03:12:31.480000Z",
"end_unknown": false, "scope": "segment", "datasets": null, "repaired_at": null } ] }
]
}

(The figures are illustrative.) The list is free with any key and counted against nothing. It shows only files that were read back and checked. A day on which nothing was observed has available: false, no hash and no download, and its coverage says why.

Terminal window
curl -sL -H "Authorization: Bearer $DEBYKO_KEY" -o btcusdt-trade-2026-09-30.parquet \
"https://api.debyko.com/v2/archive/BINANCE-USDM%3ABTCUSDT/trade/2026-09-30"

The answer is a 302 to a link that reads that one file for fifteen minutes. The response carries X-Debyko-Sha256 (the file’s hash) and Link (the day’s entry in the list). The link needs no key and carries none: a client that drops the Authorization header when it follows a redirect to another host, as curl -L and requests do, never sends your key to the storage. The contents of a file are never paged through the API: the redirect is the whole of it.

  • With a key, a file counts against the plan’s files for the calendar month (UTC): Free and Mini none, Standard 100, Pro 1 000, Pro Plus 10 000. Past them the answer is 402 with the price of the file; X-Debyko-Archive-Remaining says what is left.
  • Without a key, when paying per call is on, a file costs 0.10 USDC (core) or 0.05 USDC (top-100): one payment, one file, one 302.
  • A day that is not archived yet is a 404: a day is written once it has been over for two days.

When as_of or a history request reaches a raw day the database no longer holds, the answer is not_retained as ever, and when a file exists for that day it carries archive: { "path": …, "sha256": … } beside the reason. That is a pointer and nothing more; the file is fetched as above.

  • Path in the storage: archive/{venue}/{listing}/{layer}/{yyyy}/{mm}/{listing}-{layer}-{yyyy-mm-dd}.parquet, the venue in lower case, the symbol with every character outside A-Za-z0-9._- written as %XX.
  • Rows: in time order (trades: then by the venue’s trade id), one row per observation, exactly as stored. A file has the rows of one listing and one UTC day.
  • Row groups: one per hour, a very busy hour in groups of at most 25 000 rows, each with minimum and maximum statistics on every column, so a reader can skip the hours it does not need. Compression: zstd.
  • Types: timestamps are INT64 with the logical type timestamp, UTC, nanosecond unit (the platform keeps microseconds; they are written as they are). Every number is DECIMAL(38, 18), never a float. A figure a venue published (prices, sizes, bids, asks, marks, mid, spread) is written as the venue gave it, never rounded: it fits the column exactly or the day is not written. A figure the platform computed (the derived columns, below) is the same type. generation, sequence are INT64, level counts INT32, words are strings.
  • Columns, in this order, per layer. ? marks a column that can be null (the venue did not publish it).

trade — event_time, price, size, side, venue_uid, received_at, generation, sequence, origin

ticker — bucket, bid, ask, last?, bid_size?, ask_size?, received_at, source_time?, generation, sequence, transport, origin, cadence?

mark — bucket, price, index_price?, premium?, received_at, source_time?, generation, sequence, transport, origin, cadence?

depth-derived — bucket, mid, spread, spread_bps, imbalance?, bid_liquidity_01?, bid_liquidity_025?, bid_liquidity_05?, bid_liquidity_1?, ask_liquidity_01?, ask_liquidity_025?, ask_liquidity_05?, ask_liquidity_1?, bid_value_01?, bid_value_025?, bid_value_05?, bid_value_1?, ask_value_01?, ask_value_025?, ask_value_05?, ask_value_1?, bid_levels, ask_levels, bid_band_levels_01?, bid_band_levels_025?, bid_band_levels_05?, bid_band_levels_1?, ask_band_levels_01?, ask_band_levels_025?, ask_band_levels_05?, ask_band_levels_1?, received_at, source_time?, generation, sequence, transport, origin, cadence?

Derived columns are the platform’s own calculation, not a venue’s figure, and are listed in the file’s footer (debyko.derived_columns). What the platform computes has the precision of its unit, half to even, fixed where it is computed (from 2026-10-02; before that a division carried up to 28 digits): a fraction or ratio has 12 digits after the decimal point, basis points have 6, and a price or quantity is exact arithmetic, as stored, with no rounding of ours. Where a stored historical value still has more digits than its unit, the file holds it rounded half to even to the digits of the unit (to 18, the column’s scale, for an amount): that is the only rounding there is, and only these columns have it. The mark layer’s premium is the one derived column that is a venue’s own figure on some rows (Hyperliquid, OKX and Kraken Futures state it; the platform divides it for Bybit, Binance, Coinbase and Weex, to 12 digits when it computes it), so the file holds it rounded half to even to 18 only: a venue’s figure is never cut.

Layer Columns Kind
trade, ticker every number venue figure, never rounded
mark price, index_price venue figure, never rounded
mark premium derived: fraction, 12 digits when the platform computes it (a stated premium is kept to 18)
depth-derived mid, spread venue figures (exact arithmetic on two venue prices), never rounded
depth-derived spread_bps derived: basis points, 6 digits
depth-derived imbalance derived: fraction, 12 digits
depth-derived bid_liquidity_01, bid_liquidity_025, bid_liquidity_05, bid_liquidity_1, ask_liquidity_01, ask_liquidity_025, ask_liquidity_05, ask_liquidity_1 derived: quantity, as stored (at most 18 digits)
depth-derived bid_value_01, bid_value_025, bid_value_05, bid_value_1, ask_value_01, ask_value_025, ask_value_05, ask_value_1 derived: amount in the quote currency, as stored (at most 18 digits)

bucket is the start of the sampling window the row stands for (cadence says how long it is: 1s, 10s, 15s or 1m); received_at is when the platform received the value and source_time when the venue says it happened. The same words mean the same as in the snapshots. Not in a file: the instrument (the file is of one listing) and the storage class a row was filed under, which is the platform’s own routing.

A file is not a claim that the market was quiet where it has no rows. It is the rows the platform saw. The gaps of the day (the platform’s connection to the venue broke, a poll was rate limited, the collector was down) are in coverage of the list entry with their causes, copied from the platform’s gap record when the file was written: scope says whether the whole venue connection was affected or this listing alone, datasets which layers, repaired_at whether a replay later filled it (such rows are in the file, with their own origin). A day with a gap is still written, rows as observed. A day with no rows at all has no file, and an entry that says why.

Each takes the downloaded file or, for the first two and DuckDB, the link itself (the Location of the 302; it is good for fifteen minutes).

# pandas
import pandas as pd
df = pd.read_parquet("btcusdt-trade-2026-09-30.parquet") # or the https link; prices are decimal.Decimal, times datetime64[ns, UTC]
# polars
import polars as pl
df = pl.read_parquet("btcusdt-trade-2026-09-30.parquet") # or pl.scan_parquet(link) to read only the row groups a filter needs
-- DuckDB
SELECT count(*), min(event_time), max(event_time) FROM read_parquet('btcusdt-trade-2026-09-30.parquet'); -- or the https link

Prices stay decimals so that nothing is rounded for you: cast to a float when you have decided it does not matter (price::DOUBLE in DuckDB, .cast(pl.Float64) in polars), not before.

Every file has a SHA-256 in the list and in the X-Debyko-Sha256 header of the download; what you download hashes to it (sha256sum file.parquet). The hash is what the manifest says the platform wrote and then read back from the storage before it let go of the rows. Each file’s hash will be a leaf of the day’s Merkle tree, so that a downloaded file can be verified against an on-chain root with the same proof endpoint as every other anchored value. Anchoring is not live yet: anchored is false in the list, and the field is there from the first day so that nothing about the list changes when it becomes true.