- adapt meta-/code-/obs- rules to the local Python benchmark pipeline - replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents - add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks - normalize specs and README typography to ASCII per code-data-formatting - rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore - specs: pipeline plan and task specs 00-11 - rules: 19 binding rule files adapted for this project - docs: CLAUDE.md rule-trigger table - config: .cursorrules commit convention, .gitignore excluding out/ and .venv/ - docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders - config: requirements.txt for the closed dependency list - datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
104 lines
4.3 KiB
Markdown
104 lines
4.3 KiB
Markdown
# Build / Pipeline tasks - the ./out tree, .done markers, task order
|
|
|
|
**Purpose + scope.** How the 11 pipeline tasks compose: where artifacts
|
|
live, how completion is signaled, what order tasks run in, and what git
|
|
does and does not track. Binding for every task implementation and for
|
|
anyone running the pipeline.
|
|
|
|
---
|
|
|
|
## 1. The artifact tree - everything under `./out/`
|
|
|
|
All generated artifacts live under `./out/` relative to the repo root
|
|
(spec 00_PLAN):
|
|
|
|
```
|
|
out/
|
|
config/ lab_config.yaml, hw_prices.yaml
|
|
csv/ the canonical raw corpus + MANIFEST.csv
|
|
json/ full/ (per-coupon JSON-LD), hybrid/dataset.jsonld
|
|
sqlite/ tribo.db
|
|
pg/ tribo.dump (pg_dump -Fc; the live DB is in PostgreSQL)
|
|
rdf/ dataset.nt.gz, dataset.ttl, oxigraph_store/
|
|
bench/ queries/, expected/, results_raw.csv, results_median.csv,
|
|
storage_sizes.csv, extrapolation.csv, hardware_sizing.csv
|
|
report/ REPORT.md, charts/, tables/, process_flow.md, diagrams/
|
|
.done/ completion markers, one per task
|
|
```
|
|
|
|
- A task writes ONLY under its own output paths; it treats its inputs as
|
|
read-only.
|
|
- Tasks communicate exclusively through these artifacts - never through
|
|
shared in-process state
|
|
(see [code-python-style.md](code-python-style.md)).
|
|
- **`./out/` is git-ignored** (except nothing - the whole tree). Sources
|
|
of truth are the specs, the rules, and the code; artifacts are
|
|
reproducible from them ([data-determinism.md](data-determinism.md)).
|
|
Never commit generated artifacts "for convenience".
|
|
|
|
## 2. Completion markers
|
|
|
|
- Each task NN writes `./out/.done/<NN>.ok` after - and only after - its
|
|
outputs are complete and validated. The marker body carries the
|
|
headline counts (rows, bytes, triples) downstream tasks check
|
|
(see [test-pipeline-validation.md](test-pipeline-validation.md)).
|
|
- A task verifies its dependencies' markers before starting and aborts
|
|
loudly if one is missing.
|
|
- Re-running a task: delete its own marker first, regenerate outputs,
|
|
rewrite the marker. If its outputs changed, downstream markers are now
|
|
stale - delete them too (the dependency graph below says which).
|
|
|
|
## 3. Task graph and execution order
|
|
|
|
```
|
|
01 -> 02
|
|
01 -> 03 -> (04 | 05 | 06 | 07) -> 08 -> 09 -> 10 -> 11
|
|
09 ------> 11
|
|
```
|
|
|
|
- Sequential order: 01, 02, 03, then 04-07 (independent, may run in
|
|
parallel), 08, 09, 10, 11.
|
|
- 02 (process diagrams) needs only 01 and can run any time after it.
|
|
- Task 09 (benchmark execution) must run WITHOUT 04-07 style parallel
|
|
work in the background - measurements need a quiet machine
|
|
([bench-methodology.md](bench-methodology.md) section 6).
|
|
|
|
## 4. `storage_sizes.csv` - the shared append contract
|
|
|
|
- Location: `./out/bench/storage_sizes.csv`; columns:
|
|
`format, variant, bytes`.
|
|
- Tasks 04-07 APPEND their format's rows (e.g. `json,full,...`,
|
|
`json,hybrid,...`, `sqlite,db,...`, `pg,live,...`, `pg,dump,...`,
|
|
`rdf,nt_gz,...`, `rdf,store,...`); task 09 appends the gzip archival
|
|
variants and consolidates.
|
|
- Appends are idempotent per (format, variant): a re-run REPLACES its own
|
|
prior rows, never duplicates them and never touches other formats'
|
|
rows.
|
|
- RDF load wall-time and similar per-format load metrics recorded per
|
|
spec 07 go into the task's `.done` marker and the report, not into
|
|
extra ad-hoc files.
|
|
|
|
## 5. Runs are resumable, not restartable-from-zero
|
|
|
|
- The marker system exists so a failed pipeline resumes at the failed
|
|
task; do not delete the whole `./out/` tree to retry one task.
|
|
- Conversely, after a deliberate corpus regeneration (03), ALL downstream
|
|
markers and artifacts (04-11) are stale and must be removed - a mixed
|
|
tree of old and new artifacts is worse than an empty one
|
|
([meta-no-assumptions.md](meta-no-assumptions.md) applies before any
|
|
such deletion).
|
|
|
|
## 6. Anti-patterns
|
|
|
|
- **Writing an artifact outside `./out/`** or a task writing into another
|
|
task's output directory.
|
|
- **Committing `./out/` content to git.**
|
|
- **A marker written on partial success** - see
|
|
[code-error-handling.md](code-error-handling.md).
|
|
- **Running benchmarks while converters are still running.**
|
|
- **Duplicate rows in `storage_sizes.csv`** after re-runs - replace your
|
|
own rows.
|
|
- **Hand-editing generated artifacts** (a "quick fix" inside
|
|
`MANIFEST.csv` or a result CSV) - regenerate them through the owning
|
|
task instead.
|