unrip/archive/implementation/20260612T150844Z-quote-lifecycle-retention-and-analytics-rollups-implementation.md
philipp 8d2ca61fa1 Plan verifier salt hot-path turn
Proof: Archived the quote lifecycle retention turn and opened detailed PROOF.md/IMPLEMENTATION.md for removing verifier salt refresh from quote-response execution hot path.

Assumptions: The current retention turn should be paused so the live salt-induced executor queue latency defect can take priority; no backlog item was selected because the turn was opened directly from live operator evidence.

Still fake: Planning only; salt cache refactor, executor rejection behavior, health/dashboard surfaces, deployment, and live latency evidence are not implemented yet.
2026-06-12 17:13:03 +02:00

290 lines
11 KiB
Markdown

# Implementation Turn: Quote lifecycle retention and analytics rollups
Status: open
Opened: 2026-06-06
## Goal
Keep durable successful trade evidence while preventing quote firehose and non-success lifecycle rows from filling Postgres by adding explicit retention policy, compact rollups, table-size guardrails, and operator-visible storage truth.
## Selected backlog items
- none selected; this turn was opened directly from the operator request on 2026-06-06.
## Design rules
- This turn is storage and analytics only.
- Do not change strategy actionability, quote response policy, edge thresholds, inventory checks, pair enablement, arming, signer checks, or relay submission behavior.
- Do not skip quotes to avoid relay errors.
- DB config and persisted runtime state are canonical; do not add active retention env vars.
- Completed successful trade evidence must survive every cleanup path.
- Non-success quote detail may be short-lived if aggregate rollups are persisted first.
- Old rows must remain readable.
- Deploy only through repo workflow.
## Problem statement
The previous raw quote retention fix did not prevent Postgres growth because normalized lifecycle tables became the dominant storage users:
- `trade_decisions`
- `swap_demand_events`
- `execute_trade_commands`
- `trade_execution_results`
- `quote_outcome_attributions`
Emergency cleanup proved the system can keep completed quote outcomes and linked rows while pruning non-success detail, but that cleanup was operational, not a complete product path. The repository now needs a durable retention and rollup layer so live quote analytics remain useful without unbounded row accumulation.
## Backend changes
### 1. Inspect current storage and retention paths
- Inspect:
- `ensureHistorySchema`
- quote lifecycle table creation and indexes
- history-writer batch ingestion
- current raw quote prune
- current non-success lifecycle prune
- dashboard bootstrap queries
- competitiveness aggregation helpers
- runtime health and alert summaries
- deploy/static config tests
- Identify all tables that grow with quote firehose or lifecycle activity.
- Record current row sizes and indexes in implementation notes or test fixtures where useful.
### 2. Define retention data classes
- Add a shared retention classifier for quote-linked rows:
- `successful_trade_evidence`
- `in_flight_detail`
- `non_success_detail`
- `raw_debug_firehose`
- `aggregate_rollup`
- Preserve successful evidence when:
- `quote_outcome_attributions.outcome_status = 'completed'`, or
- attribution status indicates successful linkage such as `heuristic_match`, `exact_match`, or a locally established equivalent.
- Treat submitted/not_filled/failed/rejected rows as non-success detail once outside the retention window.
- Keep classification pair-native and quote-id based.
- Add tests for classification using current and legacy payload shapes.
### 3. DB-backed retention policy
- Add DB-owned retention policy storage, either:
- a narrow `retention_policies` table, or
- a versioned config row if an established local config pattern fits better.
- Candidate fields:
- `policy_key`
- `enabled`
- `normal_detail_retention_ms`
- `pressure_detail_retention_ms`
- `raw_retention_ms`
- `rollup_granularity_ms`
- `rollup_retention_ms`
- `preserve_successful_evidence`
- `max_delete_rows_per_pass`
- `pressure_table_bytes`
- `pressure_database_bytes`
- `updated_at`
- `config_version`
- Seed conservative defaults in repo-owned DB seed/migration code.
- Missing or invalid policy must fail closed by skipping prune, not by deleting data.
- No env var fallback.
### 4. Rollup model
- Add compact quote lifecycle rollups before pruning detail.
- Store rollup rows keyed by:
- rollup window start/end
- pair id
- direction
- request kind
- edge bps bucket or exact configured edge
- notional asset
- notional bucket
- quote-age bucket
- result code
- failure category
- outcome status
- Store compact measures:
- quote count
- decision count
- command count
- executor result count
- relay accepted count
- relay failed count
- not_filled count
- completed count
- p50/p90-compatible timing bucket counts, or fixed timing histogram buckets
- min/max observed timestamps
- source row high-water mark or computed-at watermark
- Prefer bucket histograms over storing raw samples indefinitely.
- Add indexes for dashboard rollup queries.
### 5. Rollup writer
- Extend history-writer or add a narrow repo-owned maintenance module used by history-writer.
- On a schedule:
- find recent unrolled quote lifecycle rows
- aggregate them into rollup windows
- upsert rollup rows idempotently
- record rollup watermarks
- Rollup must run before detail prune.
- If rollup fails, detail prune for the affected window must not run.
- Add tests for idempotent upsert and no double-counting.
### 6. Detail retention prune
- Replace ad hoc non-success lifecycle prune with policy-driven retention.
- Prune these quote-linked detail tables only after successful rollup:
- `raw_near_intents_quotes`
- `swap_demand_events`
- `trade_decisions`
- `execute_trade_commands`
- `trade_execution_results`
- `quote_outcome_attributions`
- Preserve:
- all successful quote ids and linked lifecycle rows
- rows within the active retention window
- rows whose outcome attribution is still in flight and younger than the configured outcome grace window
- Delete in bounded batches and expose per-table counts.
- Prefer table/index bloat control using repo-owned `VACUUM`/`REINDEX` guidance or tested maintenance where appropriate.
- Do not use untracked manual SQL as the normal path.
### 7. Pressure guardrails
- Add DB/table size inspection helpers using Postgres catalog views.
- Add a retention mode decision:
- `normal`
- `pressure`
- `blocked_invalid_policy`
- `blocked_rollup_stale`
- In pressure mode:
- shorten non-success detail retention to the DB policy's pressure value
- keep successful evidence
- keep rollups
- log and expose the pressure reason
- Do not disarm, pause, or alter strategy behavior from storage pressure in this turn unless an existing safety invariant already does so.
- Add tests proving pressure mode does not delete successful evidence.
### 8. Operator dashboard
- Add a storage/retention section to the operator dashboard.
- Show:
- table sizes and approximate row counts
- retention mode
- configured retention windows
- last rollup time
- last prune time
- rows pruned by table
- successful quote evidence count
- rollup row count and freshness
- Update competitiveness views to use rollups for older windows and recent detail for live drilldown.
- Mark old pruned detail plainly as unavailable rather than empty.
- Keep pair-native labels.
- Add dashboard tests.
### 9. Runtime health and alerts
- Add runtime health fields for:
- retention mode
- storage pressure
- rollup freshness
- prune failures
- table size thresholds
- Add warning alerts for:
- rollup stale beyond threshold
- retention blocked by invalid policy
- quote lifecycle tables above pressure threshold
- Do not make storage pressure look like a trading result or settled PnL issue.
- Add tests for alert payloads and pair-independent storage warnings.
### 10. Repo-owned emergency command
- Add or extend an ops script for dry-run and execute retention maintenance.
- Required behavior:
- dry-run default
- show preserve count for successful quote ids
- show delete candidates by table
- refuse to run if preserve query fails
- refuse to run if rollups are stale unless explicitly marked as emergency and documented
- This script is for controlled recovery, not normal operation.
## Data and persistence
- Use additive schema changes where practical.
- Keep successful lifecycle rows in existing tables so old dashboard drilldowns continue to work.
- Keep rollups compact enough to retain for substantially longer than detail.
- Keep raw JSON payloads only in short-lived detail or successful evidence rows.
- Store original relay failure category/text in rollup dimensions where already normalized.
- Maintain old-row compatibility for missing timing/retention fields.
## Edge cases
- Missing quote id: do not classify as successful evidence; retain only inside short recent window unless rollup logic can safely count it.
- Missing outcome attribution table or failed preserve query: skip prune.
- Rollup watermark missing: skip prune for affected detail window.
- Clock skew or missing timestamps: bucket as `unknown`, do not invent age.
- Duplicate event ids: rollup idempotency must not double count.
- Completed quote has no raw row because raw was already pruned: preserve available linked lifecycle rows.
- Existing emergency-pruned history cannot be recreated.
- Old rows with legacy BTC/EURe fields remain readable and do not break rollups.
## Concrete implementation order
### Phase 1. Storage audit and contracts
- Inspect schema, indexes, history-writer retention, dashboard loaders, and runtime-health paths.
- Add retention classifier helpers and tests.
- Add storage metrics helper tests using SQL/static fixtures.
### Phase 2. Retention policy
- Add DB-backed retention policy schema/seed.
- Add policy loader and validation.
- Add fail-closed tests for missing/invalid policy.
### Phase 3. Rollup storage and writer
- Add rollup table/schema.
- Add rollup aggregation helper and SQL upsert path.
- Add idempotency and grouping tests for nBTC/EURe and nBTC/USDC.
### Phase 4. Prune integration
- Replace hardcoded prune constants with DB policy.
- Ensure rollup-before-prune ordering.
- Preserve completed successful quote ids across all quote lifecycle tables.
- Add bounded batch and pressure-mode tests.
### Phase 5. Dashboard and health
- Add storage/retention dashboard card or page.
- Wire rollup freshness and table sizes into dashboard bootstrap/WebSocket state.
- Add runtime-health warning fields and alert tests.
### Phase 6. Validation and deploy
- Run targeted retention/rollup/storage tests.
- Run full `npm test`.
- Build dashboard bundle.
- Deploy through repo workflow only.
- Collect live evidence:
- table sizes before/after at least one prune cycle
- rollup freshness
- retained successful evidence count
- recent quote flow still writing rows
- all deployments healthy
## Test plan
- Retention classifier tests.
- Successful quote preserve tests.
- Rollup grouping tests:
- pair
- direction
- request kind
- result code
- failure category
- quote-age bucket
- notional bucket
- outcome status
- Rollup idempotency tests.
- Policy validation tests.
- Prune blocked when rollup stale.
- Pressure-mode prune tests.
- Old-row compatibility tests.
- Dashboard storage/retention UI tests.
- Runtime-health/alert tests.
- Static config tests proving no new active retention env vars.
- Full `npm test`.
- Dashboard build.
## Validation checklist against the proof
- Successful completed quote evidence remains in detail tables after prune.
- Non-success detail tables stay bounded under live quote flow.
- Rollups continue to answer competitiveness questions after old detail is pruned.
- Dashboard shows retention mode, table sizes, rollup freshness, and prune results.
- No strategy behavior changes.
- Repo workflow deploys the result.
## Known fakes allowed at start of this turn
- Historical non-success detail already pruned cannot be recovered.
- Fee-complete realized PnL is still unavailable.
- Venue-native terminal fill ids remain unavailable.
- Rollups are analytics summaries, not proof of individual trades.