- adapt meta-/code-/obs- rules to the local Python benchmark pipeline - replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents - add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks - normalize specs and README typography to ASCII per code-data-formatting - rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore - specs: pipeline plan and task specs 00-11 - rules: 19 binding rule files adapted for this project - docs: CLAUDE.md rule-trigger table - config: .cursorrules commit convention, .gitignore excluding out/ and .venv/ - docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders - config: requirements.txt for the closed dependency list - datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
5.4 KiB
Test / Pipeline validation - markers, manifests, checksums
Purpose + scope. The validation contract between pipeline tasks: how a task proves it completed correctly and how downstream tasks verify their inputs before trusting them. Binding for every task implementation (01-11) and for any pytest suite added to this repository.
The pipeline has no human QA step between tasks - these gates are the only thing standing between a silent generation/conversion bug and a report full of wrong numbers.
1. Completion markers - written last, verified first
- Every task writes
./out/.done/<NN>.okonly after all its outputs are complete and its own validation passed. The marker contains the task's headline counts (rows, bytes, triples) for downstream checks. - A task MUST verify the
.donemarkers of its declared dependencies before starting, and abort with a clear message if one is missing. - A task that re-runs deletes its own marker first and rewrites it at the
end. A marker is never written on failure and never from a
finallyblock (see code-error-handling.md).
2. The CSV corpus manifest - the root of trust
- Task 03 writes
./out/csv/MANIFEST.csv: one row per generated file with path, row count, byte size, sha256. - Every converter (04-07) validates what it read against the manifest: row counts per source file, and sha256 where the file is consumed whole. A mismatch is an abort, not a warning.
- The manifest is also the reproducibility proof: regenerating with the
same
lab_config.yamlmust reproduce identical sha256 values (see data-determinism.md).
3. Canonical expected volumes
The lab configuration fixes the corpus size exactly; validation asserts these numbers, not "roughly that many":
| Entity | Count |
|---|---|
| batches | 4 |
| wafers | 12 |
| coupons | 588 (480 friction + 108 characterization-only) |
| tracks | 1,440 |
| friction cycle rows | 1,440,000 |
| friction loop points | 2,880,000 |
| XRF points | 235,200 (588 x 400) |
| generated tree size | 120-200 MB (hard abort bounds) |
If a deliberate lab_config.yaml change moves these numbers, this table
and the validation code change in the same commit.
4. Query-result checksums - all formats agree
- Each benchmark query Q1-Q7 has ONE canonical expected result: rows
sorted, floats compared with tolerance 1e-9, checksummed to
./out/bench/expected/q<N>.sha256. - The canonical checksums are computed from PostgreSQL and cross-validated against SQLite before any benchmarking starts; a disagreement between the two engines is a converter or query bug and blocks task 09.
- Every measured benchmark run records a
checksum_okflag; a cell whose result does not match the expected checksum is an invalid measurement regardless of how fast it ran.
5. Format smoke tests
Each converter proves basic well-formedness of its output before writing its marker:
- JSON-LD (04): one sample coupon file passes
json.loadAND an rdflib JSON-LD parse. - SQLite (05) / PostgreSQL (06): row counts per table match the
manifest-derived expectations;
track_summaryrow count equals 1,440. - RDF (07): the loaded store answers a SPARQL
COUNTmatching the serialized triple count recorded in the.donemarker.
6. If a pytest suite is added
The specs define no pytest suite; validation is built into the tasks. If unit/integration tests are added later:
- Skip, don't fail, when PostgreSQL is unreachable. A checkout without
a running PostgreSQL still gets a green (skipped) run;
connection failure in a fixture is
pytest.skip(...), not an error. - Never mutate real
./out/artifacts. Tests that need data build a small corpus into a temp directory (a reducedlab_configis fine) or copy the artifact; hours of generated/measured data are not test fixtures to be overwritten. - Ephemeral databases only. A test touching PostgreSQL creates and
drops its own database (e.g.
tribo_test) next totribo; it never writes totriboitself. - Independent and order-free. Every test passes run alone; no residue between tests, no execution-order dependencies.
- Deterministic assertions. With the fixed seed, exact counts and exact values are assertable - prefer exact assertions over "greater than zero".
7. Anti-patterns
- Writing a
.donemarker before validation, on failure, or infinally- the marker means "completed and verified". - Skipping manifest validation "because generation just ran" - the gate exists precisely for the runs where the assumption is wrong.
- Warning-and-continuing on a count mismatch - a converter with a row deficit produces a database that LOOKS complete; abort.
- Regenerating expected checksums to make a failing format pass - the canonical result comes from PostgreSQL cross-validated with SQLite; a mismatch in another format is a bug in that converter or query.
- Tests writing into
./out/or connecting to the realtribodatabase. - Failing (not skipping) when PostgreSQL is unavailable in a pytest run.
8. Related documents
- build-pipeline-tasks.md - marker layout and task order.
- data-determinism.md - why byte-identical regeneration is assertable.
- code-error-handling.md - fail-loud mechanics.
- bench-methodology.md - how
checksum_okenters the measurement protocol.