Selection and provenance¶
Shared post-intake selection¶
The selection layer lives in tapestry.data.selection, between immutable raw
intake and consumers. The explorer uses it directly. Downstream analysis uses the
same table decisions, measure allowlists, and canonical signal descriptions.
The first selection step admits only native state and national observations.
Unsupported partitions are skipped before opening them. Mixed tables are filtered
row by row before measure selection and Hub conflict validation: HHS, HSA, county,
catchment, and wastewater-site observations are discarded from the selected stream.
A state field on a finer observation is context, not statewide support. National-only
sources may omit geographic columns. Raw snapshots are retained on disk; this filter
does not reclaim their disk space or select the model's final channel set.
Publisher → intake → immutable raw snapshots
↓
SelectedData (policy v6)
↙ ↘
ExplorerIndex native SelectedRecord stream
display variants downstream materialization
Source counts¶
| Family | Acquisition datasets | Logical source groups |
|---|---|---|
| NHSN | 4 | 1 |
| NSSP | 4 | 1 |
| NWSS (Delphi archive, auxiliary table, derived state indices) | 3 | 1 |
| Inpatient and outpatient claims | 2 | 2 |
| FluView (ILINet, clinical labs) | 2 | 1 |
| FluSurv-NET | 1 | 1 |
| Kinsa (via PopHIVE) | 1 | 1 |
| Current and historical forecast hubs | 6 | 3 additional (current hubs join NHSN/NSSP) |
| Total | 23 | 11 |
These are catalog definitions, not claims that every source has been downloaded or every geography is displayable. The selection summary reports missing acquisitions.
A signal is a meaningful measure within a logical group. A variant retains its acquisition source, cadence, release product, geography, units, smoothing, and demographic/fill facets. The explorer displays one checkbox per variant under each signal, so multiple releases or providers can be selected together. Grouping does not average providers or replace a missing variant with another provider's values. Finalized native-state CDC NHSN is the preferred display variant when available. Native-state support is the default UI filter; change it to see national context.
Choice counts are reported as indexed_signal_choices and indexed_variants
from /api/catalog after a rebuild. Counts vary with local acquisitions and
publisher schema changes.
NHSN: CDC measures and Delphi mirrors¶
The current research selection is all ages only. For each of COVID-19, influenza, and RSV, it retains total admissions, total hospitalized patients, total ICU patients, and hospitals reporting admissions. Two shared fields retain inpatient beds and occupied inpatient beds: 14 measures total. Adult, pediatric, individual age-band, and unknown-age fields are excluded from selection, but remain in immutable raw downloads. Rates, percentages, changes, and seasonal cumulative sums remain excluded from NHSN outcomes.
Each CDC pull saves the publisher's fieldName, name, description, and
dataTypeName from /api/views/<dataset-id>.json in columns.json, alongside
complete metadata.json. Raw rows keep API field identifiers. Selected signal
labels use the publisher's name, with descriptions available as explorer
hover text. Older downloads use their saved metadata directly.
Delphi and current Hub outcomes join their documented originating CDC column.
For example, confirmed_admissions_flu_ew and FluSight's wk inc flu hosp
join NHSN totalconfflunewadm (CDC label: “Total Influenza Admissions”).
Delphi's signal values are its measures; value is the storage field.
Hub target values likewise identify measures stored in observation.
Unmapped historical Hub sources retain their distinct origins.
The three CDC products remain distinguishable as Finalized, Preliminary, and First publication. Delphi contributes its configured NHSN signals as archive variants of their matching CDC columns; it does not supply every CDC measure (for example, adult-only hospitalized and ICU counts). The first-publication product is not a full revision archive. See Vintages and geography.
NSSP: one emergency department surveillance group¶
Signals are ED-visit percentages for COVID-19, influenza, RSV, ARI, and combined respiratory pathogens where published. Daily, weekly, and national demographic products remain variants under their corresponding CDC measure. Reported and smoothed measures use distinct CDC columns; Delphi joins the matching column. Current Hub ED targets join the reported column, with their 0–1 proportions explicitly distinguished from CDC/Delphi percentages. Values are not rescaled by grouping.
- Daily and demographic CDC tables retain
percent_visits. - CDC trajectories retain the four reported and four smoothed percentage fields.
Socrata's
percent_visits_smoothedmaps to combined andpercent_visits_smoothed_1maps to influenza, as defined by its metadata. - Delphi retains its nine configured weekly reported/smoothed signals.
- Trend classifications, location identifiers, and build metadata are not outcomes.
Both downstream selection and the explorer retain only state/national support.
CDC trajectories require county=All and no specific HSA; county copies of an HSA trajectory must
not be averaged into an apparent state observation. Demographic and smoothed
variants remain separate; percentages and hub ED decimal proportions are not
silently treated as equal units.
Wastewater and finer geography¶
The currently cataloged NWSS products contain site/sewershed observations, even when their
rows carry a state label. They remain in the raw acquisition catalog but are
excluded from both SelectedData and the explorer. No site-to-state averaging
or geographic crosswalk is performed. CDC also publishes a separate weekly
state/territory WVAL product
for influenza A, COVID-19, and RSV; it is not yet in this acquisition catalog. HHS and catchment observations are excluded
as well; a genuine national row from a mixed catchment/national source can remain.
FluView and FluSurv-NET¶
Policy v6 (2026-09-22) adds the Delphi V5 FluView and FluSurv-NET archives. Each
retains value for its configured catalog signals only.
- FluView is one group with two components. Signal keys carry a component
prefix (
fluview:ilinet_ili,fluview:clinical_pct_positive). ILINet signals are filed under Influenza-like illness, laboratory signals under Influenza.age_groupandfill_methodstay variant facets: ILINet visit counts are age-stratified, and statewide New York is Delphi'snyc_plus_ny_minus_nycpool of CDC's two New York jurisdictions (NYC and NY-minus-NYC are not states and are excluded). - FluSurv-NET signals are catchment hospitalization rates per 100,000,
labelled by stratum (
rate_age_0is ages 0–4, following Delphi's legacy V3 numbering). The state variants cover only FluSurv-NET catchment states; the national row is the network rate.miscnetwork rows (EIP, IHSP) and MSA rows are not state or national support and are excluded.
Delphi's public health laboratory source is not acquired (its state rows are season totals). See FluView and FluSurv-NET.
Hubs: canonical observations grouped under their origins¶
| Hub | Canonical file selection | Observed target choices |
|---|---|---|
| Current FluSight | target-data/time-series.csv |
2 |
| Current COVID | target-data/time-series.parquet |
2 |
| Current RSV | target-data/time-series.parquet |
2 |
| Legacy FluSight | data-truth/truth-Incident Hospitalizations.csv |
1 |
| Legacy COVID | Six top-level incident/cumulative case/death/hospitalization truth files | 6 |
| RSV-NET | Latest top-level dated rsvnet_hospitalization.csv |
2 |
| Expected definitions | 15 |
Current hubs use observation; legacy hubs and RSV-NET use value. The target
field, or the legacy truth filename, defines the signal. Age strata are variants.
Population, location codes, quantiles, oracle output, redundant exports,
auxiliary tables, and alternative truth providers do not create outcome series.
Derived peak/rate-change evaluation definitions are not independent observation
streams. The rate exclusion for NHSN does not remove the RSV-NET rate target.
Current hub files retain their as_of releases. Intake additionally materializes
configured target files without native release dates into git-history.ndjson.gz,
with a release/commit/path/blob index in git-history.json. These are canonical
Git-history variants of the same observations, not new outcome definitions.
Preferred filenames replace earlier aliases; dated backups and oracle outputs
remain excluded. The selected stream and explorer read these saved releases.
Complete-snapshot resolution respects omitted rows and whole-file deletions.
See the Git vintage policy and reproduction commands.
Canonical hub tables are checked across the entire file before yielding rows.
Identical duplicate keys are emitted once. Conflicting location/target/event/
release/age keys are quarantined as missing (_selection_conflict=True in native
metadata), preserving the release timestamp so an earlier value cannot silently
replace the conflict. Other observations remain available. Conflicts appear in
the source audit as quarantined.
Missing canonical files and Git LFS pointers are explicitly unavailable. There is no fallback to USAFacts, NYTimes, or another truth provider. In the local legacy COVID export, all six primary files are LFS pointers. Intake must retrieve their payloads before those choices can be plotted.
Downstream API¶
from tapestry.data.selection import SelectedData, describe
selected = SelectedData("data")
print(selected.summary())
for record in selected.iter_records(
group="nhsn",
dataset_key="cdc_nhsn_final",
):
# Only record.values contains selected outcomes.
# metadata carries native geography, dates, QC, and other raw context.
for column, value in record.values.items():
signal = describe(record.dataset_key, column, record.source_path,
record.metadata)
consume(signal["signal_key"], value, record.metadata,
record.available_at)
Specify dataset_key when a downstream task needs one particular provider or
release product. Omitting it exposes all selected variants with provenance; it
does not combine them. Raw suppression markers and nulls remain unchanged.
available_by="2026-01-09T12:00:00Z" filters out later releases using the publisher
vintage column, or conservatively the saved commit/retrieval time. Unknown
availability is excluded when a cutoff is supplied. It emits all eligible
revisions, not a resolved training panel. Downstream materialization must
choose the latest eligible observation or full release, handle retractions,
validate native support and units, and apply an event-date cutoff where required.
A recent raw download does not establish historical availability.
selected_tables() is the lower-level native table interface used by the
explorer. Apply measure_columns() to choose outcome columns. Consume each table
before advancing the iterator: temporary extraction of a tar member lasts only
for that iterator step. No training data should be built by reading the
explorer's aggregated state-point cache.
Commands, audit, and maintenance¶
## Policy counts, grouping, and missing downloads; no raw or index writes.
PYTHONPATH=src python -m tapestry.data.selection --data-root data
## Rebuild selected explorer data; reads raw snapshots without changing them.
python -m tapestry.explorer --data-root data index --force
The installed summary command is tapestry-select --data-root data.
/api/catalog includes the selection version, logical and indexed source counts,
canonical signal count, variant count, and unavailable source warnings.
SelectedData.audit records excluded hub files and missing canonical files after
iteration. The index's sources table also records excluded paths and read errors;
Index details emphasizes errors and unavailable inputs rather than listing
hundreds of deliberately excluded files.
To change selection, update the explicit mappings and POLICY_VERSION in
data/selection.py, update its regression tests, and rebuild. The policy version
participates in index invalidation. New NHSN/NSSP/NWSS numeric fields are not
silently admitted as outcomes. Other source families retain their existing
measure discovery and source grouping.
Vintages and geography¶
Revision semantics¶
versioned means the source can recover a historical information state; it does
not merely mean that local downloads are timestamped.
| Revision mode | Interpretation | Versioned |
|---|---|---|
snapshot_only |
Current mutable publisher view; local pulls remain immutable | no |
initial_release |
Frozen first-publication product | yes |
as_of_column |
Rows carry publisher release dates | yes |
report_time |
Delphi archive carries retained report-time revisions | yes |
git_history |
Files are recoverable at an exact Git commit | yes |
The raw repository preserves all available vintage fields. The explorer supports an explicit as-of cutoff and overlays of multiple versions.
Report-time archives select the latest eligible revision per event date; full
as_of tables select a complete release. Unversioned sources apply only an
event-date cutoff and cannot reconstruct past revisions.
No change does not mean no version. Delphi is a change history: for each
observation, the value advertised at report_time applies until another advertised
change replaces it. No new row on a Wednesday is required. A Hub commit gives the
actual repository state, including unchanged observations; that state remains in
force between target-file changes. Skipping duplicate blobs or unchanged rows is
storage compression, not a break in coverage. Explicit nulls, removals, and dates
before the first advertised version remain distinct from unchanged valid values.
This persistence is along publication time for the same observation; it never
fills a different event week or carries information backward before publication.
Hub acquisitions now include git-history.ndjson.gz and git-history.json for
configured canonical target CSVs without native release dates. Intake walks the
pinned main branch's first-parent history and saves complete target-file states,
including empty releases after deletion. The selected stream and explorer consume
that saved history; they do not run Git on each query. In the explorer these are
Git commit history variants beside the native as_of variants of the same
measure. The date cutoff selects a complete eligible Git snapshot, preserving
omissions and retractions rather than carrying removed values forward.
Repository ground truth: a commit fixes the complete target-file contents.
Timing assumption: its main-branch committer timestamp places that state on
the availability timeline; it does not establish the provider release time. Author time, observation dates,
filename dates, and retrieval time are not substituted. Backdated target changes
cannot predate the preceding target state; equal-time changes resolve to the last
first-parent state. Only history reachable from the pinned commit is exported.
Original commit hashes, paths and blob hashes are retained. CSV date and
target_end_date are explicit event-date aliases, never release dates. Current
Hub ED CSV values are already proportions; no percentage conversion is applied.
Once a preferred filename has appeared, an obsolete alias cannot replace it
after deletion, even if that older file remains in the Git tree.
hub-history augments the latest acquisition at its existing pinned commit;
ordinary Hub pulls also create these histories. No user's worktree is checked out.
.venv/bin/python -m tapestry.data --data-root data hub-history \
hub_flusight_current hub_covid_current hub_rsv_current hub_flusight_legacy
.venv/bin/python -m tapestry.explorer --data-root data index --force
.venv/bin/python -m tapestry.explorer --data-root data export
The path policy covers historical current-FluSight admissions and ED files,
current-COVID admissions, and legacy FluSight/COVID primary truth files. Current
RSV has native as_of history and no configured unversioned predecessor. Legacy
COVID LFS pointers still require their actual payloads; a pointer is not a usable
historical observation. RSV-NET catchment files are a separate source outside the
six scored NHSN/NSSP targets and are not folded into statewide Hub outcomes.
The September 17 backfill, pinned to the existing September 16 acquisitions:
| Source | Complete Git releases | Native state/DC/US rows | Commit-time range |
|---|---|---|---|
| Current FluSight | 132 | 1,577,922 | 2023-10-03–2026-07-09 |
| Current COVID | 93 | 241,072 | 2024-11-18–2026-09-09 |
| Current RSV | 0 | 0 | Native as_of history already present |
| Legacy FluSight | 527 | 3,736,470 | 2021-12-07–2023-11-22 |
Repeated event weeks appear in multiple complete releases; these counts are not
independent observations. Local acquisition IDs and exact pinned commit hashes
are in data/processed/hub-git-history-summary.json.
Geographic support¶
Native support is never silently relabeled:
- State observations are valid state-level inputs.
- National observations are stored once and shown with explicit national-context labels.
- HHS, HSA, and county observations are excluded at the shared post-intake filter.
- Catchment/network data are not treated as statewide values.
- Wastewater site and sewershed rows remain in raw storage but are excluded from the selected stream and explorer. No site-to-state average is computed.
This distinction is essential for the later mask: native state availability, parent context, and unavailable state data need not share the same mask policy.
Storage and provenance¶
On-disk layout¶
data/
├── catalog.json
├── mirrors/
│ └── hub_flusight_current.git/
├── raw/
│ └── cdc_nhsn_final/
│ ├── latest.json
│ └── snapshots/<retrieval-id>/
│ ├── data.ndjson.gz
│ ├── metadata.json
│ └── manifest.json
├── .staging/
└── .explorer/
├── series.sqlite3
└── revisions.parquet
raw/ and mirrors/ are durable source material. .explorer/ is disposable
and can always be reconstructed. series.sqlite3 contains searchable series
metadata plus the latest resolved point cache; revisions.parquet contains the
complete revision ledger used for as-of reads. .staging/ contains in-progress work and may
also contain validated Delphi partitions retained after an interrupted pull.
Atomic snapshots¶
A fetcher obtains a SnapshotWriter, streams its files into an isolated staging
directory, records row counts and metadata, and commits only after every file is
complete. Commit writes the manifest and atomically advances latest.json.
If a pull fails, an incomplete snapshot is not visible through latest(). Large
Delphi V5 pulls can deliberately preserve gzip CRC-valid partitions and reuse
them with --resume-from.
Manifest contents¶
Each snapshot manifest records:
- dataset key and immutable retrieval identifier;
- retrieval time, source URL, and revision mode;
- whether the source exposes historical versions;
- each payload path, byte count, row count, and SHA-256 digest;
- fetcher-specific provenance such as source metadata or a Hub Git commit.
Verify the latest snapshot with:
Data lifecycle¶
Raw snapshots are append-only. Cleaning the disposable explorer index does not remove source data. Before deleting staging or dated backup directories, inspect them for resumable partitions or recovery material.
Same-date conflicts¶
Different finite values for the same signal, location, observation date and report date have no release ordering. The model extractor leaves the ambiguous cell unavailable rather than choosing an arbitrary value. Availability evidence counts a finite published statement even when it conflicts; source publication and model usability are different measurements. Later unambiguous reports can restore usability. Missing values never become observed zeroes.