Files
LabDataStorageEvaluation/docs/rules/build-pipeline-tasks.md
administrator b173ac82a9 fix(docs): fix formatting issues in README.md
- adapt meta-/code-/obs- rules to the local Python benchmark pipeline
- replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents
- add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks
- normalize specs and README typography to ASCII per code-data-formatting
- rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore
- specs: pipeline plan and task specs 00-11
- rules: 19 binding rule files adapted for this project
- docs: CLAUDE.md rule-trigger table
- config: .cursorrules commit convention, .gitignore excluding out/ and .venv/
- docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders
- config: requirements.txt for the closed dependency list
- datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
2026-07-11 13:39:13 -04:00

4.3 KiB

Build / Pipeline tasks - the ./out tree, .done markers, task order

Purpose + scope. How the 11 pipeline tasks compose: where artifacts live, how completion is signaled, what order tasks run in, and what git does and does not track. Binding for every task implementation and for anyone running the pipeline.


1. The artifact tree - everything under ./out/

All generated artifacts live under ./out/ relative to the repo root (spec 00_PLAN):

out/
  config/    lab_config.yaml, hw_prices.yaml
  csv/       the canonical raw corpus + MANIFEST.csv
  json/      full/ (per-coupon JSON-LD), hybrid/dataset.jsonld
  sqlite/    tribo.db
  pg/        tribo.dump (pg_dump -Fc; the live DB is in PostgreSQL)
  rdf/       dataset.nt.gz, dataset.ttl, oxigraph_store/
  bench/     queries/, expected/, results_raw.csv, results_median.csv,
             storage_sizes.csv, extrapolation.csv, hardware_sizing.csv
  report/    REPORT.md, charts/, tables/, process_flow.md, diagrams/
  .done/     completion markers, one per task
  • A task writes ONLY under its own output paths; it treats its inputs as read-only.
  • Tasks communicate exclusively through these artifacts - never through shared in-process state (see code-python-style.md).
  • ./out/ is git-ignored (except nothing - the whole tree). Sources of truth are the specs, the rules, and the code; artifacts are reproducible from them (data-determinism.md). Never commit generated artifacts "for convenience".

2. Completion markers

  • Each task NN writes ./out/.done/<NN>.ok after - and only after - its outputs are complete and validated. The marker body carries the headline counts (rows, bytes, triples) downstream tasks check (see test-pipeline-validation.md).
  • A task verifies its dependencies' markers before starting and aborts loudly if one is missing.
  • Re-running a task: delete its own marker first, regenerate outputs, rewrite the marker. If its outputs changed, downstream markers are now stale - delete them too (the dependency graph below says which).

3. Task graph and execution order

01 -> 02
01 -> 03 -> (04 | 05 | 06 | 07) -> 08 -> 09 -> 10 -> 11
                                          09 ------> 11
  • Sequential order: 01, 02, 03, then 04-07 (independent, may run in parallel), 08, 09, 10, 11.
  • 02 (process diagrams) needs only 01 and can run any time after it.
  • Task 09 (benchmark execution) must run WITHOUT 04-07 style parallel work in the background - measurements need a quiet machine (bench-methodology.md section 6).

4. storage_sizes.csv - the shared append contract

  • Location: ./out/bench/storage_sizes.csv; columns: format, variant, bytes.
  • Tasks 04-07 APPEND their format's rows (e.g. json,full,..., json,hybrid,..., sqlite,db,..., pg,live,..., pg,dump,..., rdf,nt_gz,..., rdf,store,...); task 09 appends the gzip archival variants and consolidates.
  • Appends are idempotent per (format, variant): a re-run REPLACES its own prior rows, never duplicates them and never touches other formats' rows.
  • RDF load wall-time and similar per-format load metrics recorded per spec 07 go into the task's .done marker and the report, not into extra ad-hoc files.

5. Runs are resumable, not restartable-from-zero

  • The marker system exists so a failed pipeline resumes at the failed task; do not delete the whole ./out/ tree to retry one task.
  • Conversely, after a deliberate corpus regeneration (03), ALL downstream markers and artifacts (04-11) are stale and must be removed - a mixed tree of old and new artifacts is worse than an empty one (meta-no-assumptions.md applies before any such deletion).

6. Anti-patterns

  • Writing an artifact outside ./out/ or a task writing into another task's output directory.
  • Committing ./out/ content to git.
  • A marker written on partial success - see code-error-handling.md.
  • Running benchmarks while converters are still running.
  • Duplicate rows in storage_sizes.csv after re-runs - replace your own rows.
  • Hand-editing generated artifacts (a "quick fix" inside MANIFEST.csv or a result CSV) - regenerate them through the owning task instead.