tv.dsh-market.com froze again: the 7-day view ended 2026-09-13 and the
30-day view 2026-09-16 while every raw day was intact. The summary still
ran COUNT(DISTINCT visitor) over the 4.96M-row event table, so the 30-day
pre-warm died with SQLITE_NOMEM, the per-item window scans failed at the
D1 API (internal error 7500) and the cache kept serving the last good
rollup.
Add per-UTC-day rollup tables (migration 0008) that the cron trigger
rewrites and backfills, and read every heartbeat window aggregate from
them: cost is days x catalog instead of days x events. A window whose
days are still being backfilled keeps its previous complete cache row
instead of publishing a short series, and interval item, channel and
version numbers become the sum of the per-day distinct counts, which the
dashboard now labels as 期间活跃.
Tests: scripts/market-worker.test.mjs pins the walk order, the atomic
per-day batch with its completion marker, the incompleteness gate and the
failing-day path; the rollup SQL is additionally verified against a real
SQLite engine on a two-day fixture.
The 30/90/365-day summary windows froze at their last successful rollup
from 2026-09-14 (SQLITE_NOMEM on the channel/version COUNT(DISTINCT)
scans at ~4.4M rows), so tv.dsh-market.com served a 09-13 snapshot that
read as lost days. Run the auxiliary breakdowns as their own batches
with fail-soft degradation (payload.degraded flags what is skipped),
stamp every payload with generated_at, and show the rollup age plus
degradation state on the dashboard so a lagging cache can no longer
masquerade as missing data.
The dashboard only showed the active-instance count as a KPI number, with
no series behind it. Render the heartbeat UV series as its own trend panel
above the site PV/UV chart, sharing one SVG line renderer whose tooltip,
crosshair and hit zone are now per chart box.
With the tv. forward in place the dashboard worker runs inside the store
worker's invocation chain; its public fetch back to dsh-market.com
re-entered that worker in the same context, tripped Cloudflare's loop
protection, and failed as a 522 from the custom-domain placeholder
origin. The summary fetch now rides a reverse MARKET binding that
bypasses zone routing entirely, falling back to public fetch only when
the binding is absent. Pin it with a signed-JWT test driving the
authenticated /data proxy.
The 2026-09-02 relay wildcard custom domain and zone route capture every
single-level subdomain for the store worker, so tv.dsh-market.com stopped
reaching dsh-market-telemetry-view (zero invocations since 09-03, store
SPA rendered behind the Access gate instead of the dashboard). The store
worker now forwards the whole tv. hostname through the TELEMETRY_VIEW
service binding before any other dispatch; /app.js joins run_worker_first
and falls back to ASSETS on every other hostname. Also deletes the
leftover root-only zone route *.dsh-market.com from the zone.
The total badge now sums every package name the family ever published
(25, retired names included) across the npm official registry, the
npmmirror registry, and this repository's GitHub release assets; ranges
are tiled into 365-day windows because both registries clamp or reject
wide windows, a failing channel is skipped instead of graying the badge,
and an optional GITHUB_TOKEN secret moves the GitHub channel to the
authenticated quota. README.md and README.en.md gain the badge next to
the npm version badge; the index.js route line passing env to the total
handler landed with the relay commits.
Cloudflare kills scheduled invocations after roughly two minutes, so the
five ~40s pre-warm windows died mid-list (only the first three landed).
Each tick now warms the tv first-paint key plus one rotation slot of the
range windows, and 90/365-day rollups tolerate 12 hours of staleness so
their longer recompute cadence never surfaces as a slow dashboard visit.
tv.dsh-market.com answered 502 (upstream 503) on every visit: telemetry_events
had no secondary index, the nine summary aggregates full-scanned 2M+ rows
(17-25s each), and the single nine-statement batch exceeded D1 limits.
- migration 0006: covering index (kind, day, ...), day index for retention,
pre-deduped telemetry_visitors table (backfilled), summary rollup cache
- summary serves a 30-minute rollup per exact query window; stale row on
live-aggregation failure; nine aggregates split into four batches so each
transaction stays under D1 limits
- cron pre-warms the five windows the tv dashboard requests
- badge counts telemetry_visitors (0.02s) instead of a full event scan;
heartbeat writes upsert the visitor hash in the same batch
Extend the badge resilience work to the remaining read paths that
surfaced raw D1 overload as worker exception pages:
- /api/stats gains a worker-level cache (one minute, one-hour stale
copy) and a 503 storage-unavailable fallback. Client responses stay
no-store and cached-hit headers are rewritten to no-store, so the
zone cache rules keep the previous freshness semantics instead of
pinning vote counts for hours.
- /api/telemetry/summary returns 503 storage-unavailable JSON instead
of an exception page. It is deliberately not edge-cached: the cache
key cannot carry the x-telemetry-key authorization.
- Retention pruning moves from the summary read path into the cron
trigger, so dashboard reads stop issuing opportunistic DELETEs.
- badge_cache migration renamed to 0005: 0004 was already taken by
install_events.
- Agent note, api-doc, OpenAPI and docs/telemetry.md updated.
Production tail showed shields' own fetcher (Shields.io/080e177)
aborting every badge fetch at ~3.45s wall time: the live
COUNT(DISTINCT) over ~1.1M telemetry rows exceeds shields' upstream
timeout, so the badge could never render through shields even with a
healthy database, and each aborted invocation died before caching.
Add a badge_cache D1 table (migration 0004) holding the count as one
row, refreshed every 30 minutes by a new cron trigger. The badge
handler reads that single indexed row (~26ms end to end for shields)
and keeps the 30 min edge cache plus a 24h stale fallback on top;
a missing row bootstraps through one full scan and re-seeds. Contract
texts and the agent note updated with the timeout evidence.
The README users badge rendered inaccessible: shields fetched
/api/telemetry/badge/users during D1 overload windows and the worker
threw D1_ERROR (full-table COUNT DISTINCT scan), so shields received
the Cloudflare 1101 page. The badge handler now caches its response in
the edge Cache API for 30 minutes, keeps a 24h stale copy, and serves
that or a valid 200 'unavailable' shields JSON instead of throwing.
Telemetry event writes fail soft with 503 storage-unavailable so
overloads stop surfacing as worker exception pages. Contract text in
docs/telemetry.md, api-doc.js and openapi.js updated.
The summary endpoint took TELEMETRY_READ_KEY as a ?key= query parameter,
persisting the credential in edge logs, browser history and referrers.
The key is now header-only with SHA-256 digest comparison; docs, OpenAPI
and tests updated.
- D1 migration 0004: install_events + install_counts for per-asset installs
- POST /api/install records one Turnstile-gated install event per success
- /api/stats now exposes installs alongside votes
- GET /api/npm-downloads batches manifest-derived plugin download counts
- Workshop card and public site render install and npm 30d metrics separately
- turnstile challenge iframe accepts market-install action
Display and source layers rename to dsh-web: GitHub slug, docs and prose,
aggregate package dir packages/dsh-web-all with npm name
@linxin666/dsh-web-all, settings package dir packages/dsh-web-settings,
private shared package dsh-web-shared, repo-local skill dirs, and the
docs banner asset.
Runtime, wire, and storage identifiers are frozen byte-identical so
installed profiles keep resolving with zero migration: web-ui-* bundle
ids, the dsh-web-ui-market settings section id, /api/dsh-web-ui-settings
and its proxy-token header, and the dsh-web-ui-telemetry-* storage keys.
Frozen history (docs/archive, docs/release-notes, archived notes), the
JAVA-LW fork reference, and local filesystem paths keep the old name.
npm migration: the next tag release dual-publishes @linxin666/dsh-web-all
alongside the final @linxin666/dsh-web-ui-all version, then the old name
is deprecated with a pointer; dual-publish lasts two releases.
Decision record: .agents/notes/implemented/architecture/2026-08-24-product-rename-dsh-web.md