Files
LabDataStorageEvaluation/docs/rules/test-pipeline-validation.md
administrator b173ac82a9 fix(docs): fix formatting issues in README.md
- adapt meta-/code-/obs- rules to the local Python benchmark pipeline
- replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents
- add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks
- normalize specs and README typography to ASCII per code-data-formatting
- rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore
- specs: pipeline plan and task specs 00-11
- rules: 19 binding rule files adapted for this project
- docs: CLAUDE.md rule-trigger table
- config: .cursorrules commit convention, .gitignore excluding out/ and .venv/
- docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders
- config: requirements.txt for the closed dependency list
- datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
2026-07-11 13:39:13 -04:00

5.4 KiB

Test / Pipeline validation - markers, manifests, checksums

Purpose + scope. The validation contract between pipeline tasks: how a task proves it completed correctly and how downstream tasks verify their inputs before trusting them. Binding for every task implementation (01-11) and for any pytest suite added to this repository.

The pipeline has no human QA step between tasks - these gates are the only thing standing between a silent generation/conversion bug and a report full of wrong numbers.


1. Completion markers - written last, verified first

  • Every task writes ./out/.done/<NN>.ok only after all its outputs are complete and its own validation passed. The marker contains the task's headline counts (rows, bytes, triples) for downstream checks.
  • A task MUST verify the .done markers of its declared dependencies before starting, and abort with a clear message if one is missing.
  • A task that re-runs deletes its own marker first and rewrites it at the end. A marker is never written on failure and never from a finally block (see code-error-handling.md).

2. The CSV corpus manifest - the root of trust

  • Task 03 writes ./out/csv/MANIFEST.csv: one row per generated file with path, row count, byte size, sha256.
  • Every converter (04-07) validates what it read against the manifest: row counts per source file, and sha256 where the file is consumed whole. A mismatch is an abort, not a warning.
  • The manifest is also the reproducibility proof: regenerating with the same lab_config.yaml must reproduce identical sha256 values (see data-determinism.md).

3. Canonical expected volumes

The lab configuration fixes the corpus size exactly; validation asserts these numbers, not "roughly that many":

Entity Count
batches 4
wafers 12
coupons 588 (480 friction + 108 characterization-only)
tracks 1,440
friction cycle rows 1,440,000
friction loop points 2,880,000
XRF points 235,200 (588 x 400)
generated tree size 120-200 MB (hard abort bounds)

If a deliberate lab_config.yaml change moves these numbers, this table and the validation code change in the same commit.

4. Query-result checksums - all formats agree

  • Each benchmark query Q1-Q7 has ONE canonical expected result: rows sorted, floats compared with tolerance 1e-9, checksummed to ./out/bench/expected/q<N>.sha256.
  • The canonical checksums are computed from PostgreSQL and cross-validated against SQLite before any benchmarking starts; a disagreement between the two engines is a converter or query bug and blocks task 09.
  • Every measured benchmark run records a checksum_ok flag; a cell whose result does not match the expected checksum is an invalid measurement regardless of how fast it ran.

5. Format smoke tests

Each converter proves basic well-formedness of its output before writing its marker:

  • JSON-LD (04): one sample coupon file passes json.load AND an rdflib JSON-LD parse.
  • SQLite (05) / PostgreSQL (06): row counts per table match the manifest-derived expectations; track_summary row count equals 1,440.
  • RDF (07): the loaded store answers a SPARQL COUNT matching the serialized triple count recorded in the .done marker.

6. If a pytest suite is added

The specs define no pytest suite; validation is built into the tasks. If unit/integration tests are added later:

  • Skip, don't fail, when PostgreSQL is unreachable. A checkout without a running PostgreSQL still gets a green (skipped) run; connection failure in a fixture is pytest.skip(...), not an error.
  • Never mutate real ./out/ artifacts. Tests that need data build a small corpus into a temp directory (a reduced lab_config is fine) or copy the artifact; hours of generated/measured data are not test fixtures to be overwritten.
  • Ephemeral databases only. A test touching PostgreSQL creates and drops its own database (e.g. tribo_test) next to tribo; it never writes to tribo itself.
  • Independent and order-free. Every test passes run alone; no residue between tests, no execution-order dependencies.
  • Deterministic assertions. With the fixed seed, exact counts and exact values are assertable - prefer exact assertions over "greater than zero".

7. Anti-patterns

  • Writing a .done marker before validation, on failure, or in finally - the marker means "completed and verified".
  • Skipping manifest validation "because generation just ran" - the gate exists precisely for the runs where the assumption is wrong.
  • Warning-and-continuing on a count mismatch - a converter with a row deficit produces a database that LOOKS complete; abort.
  • Regenerating expected checksums to make a failing format pass - the canonical result comes from PostgreSQL cross-validated with SQLite; a mismatch in another format is a bug in that converter or query.
  • Tests writing into ./out/ or connecting to the real tribo database.
  • Failing (not skipping) when PostgreSQL is unavailable in a pytest run.