Files
LabDataStorageEvaluation/docs/specs_old/15_lineage_build_manifest.md
administrator b173ac82a9 fix(docs): fix formatting issues in README.md
- adapt meta-/code-/obs- rules to the local Python benchmark pipeline
- replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents
- add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks
- normalize specs and README typography to ASCII per code-data-formatting
- rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore
- specs: pipeline plan and task specs 00-11
- rules: 19 binding rule files adapted for this project
- docs: CLAUDE.md rule-trigger table
- config: .cursorrules commit convention, .gitignore excluding out/ and .venv/
- docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders
- config: requirements.txt for the closed dependency list
- datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
2026-07-11 13:39:13 -04:00

1.2 KiB

Task 15 — Lineage Consolidation, Build Manifest, Final Validation

Goal

Close the pipeline: consolidated lineage, reproducible build record, cross-format validation.

Input

  • All exports/*, lineage/, all config/schema versions.

Actions

  1. Write scripts/finalize.py:
    • consolidate lineage/ into lineage/lineage.sqlite (queryable: target row-range → source file + offsets + rule versions);
    • produce BUILD_MANIFEST.md: snapshot id, all config/mapping/schema versions, input inventory hash, per-artifact counts, run timestamps — the processing-flow documentation.
  2. Cross-format validation: row counts and key aggregates (e.g., sum of cycles per test, point counts per measurement) identical across canonical, JSON, SQLite, PostgreSQL staging.
  3. Lineage spot-check: resolve 20 random target rows back to source file + row offset; verify raw values match.

Final Result (pipeline done when)

  • BUILD_MANIFEST.md complete; all validations pass; lineage resolvable.
  • Final goal achieved: cleansed canonical dataset conforming to the ontology, all anomalies quarantined and explained, four export formats generated from one snapshot, full source→target lineage, entire pipeline reproducible from source + versioned config.