fix(docs): fix formatting issues in README.md

- adapt meta-/code-/obs- rules to the local Python benchmark pipeline
- replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents
- add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks
- normalize specs and README typography to ASCII per code-data-formatting
- rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore
- specs: pipeline plan and task specs 00-11
- rules: 19 binding rule files adapted for this project
- docs: CLAUDE.md rule-trigger table
- config: .cursorrules commit convention, .gitignore excluding out/ and .venv/
- docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders
- config: requirements.txt for the closed dependency list
- datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
This commit is contained in:
administrator
2026-07-11 13:39:13 -04:00
parent d8353953b0
commit b173ac82a9
55 changed files with 3755 additions and 9 deletions

View File

@@ -0,0 +1,103 @@
# Build / Pipeline tasks - the ./out tree, .done markers, task order
**Purpose + scope.** How the 11 pipeline tasks compose: where artifacts
live, how completion is signaled, what order tasks run in, and what git
does and does not track. Binding for every task implementation and for
anyone running the pipeline.
---
## 1. The artifact tree - everything under `./out/`
All generated artifacts live under `./out/` relative to the repo root
(spec 00_PLAN):
```
out/
config/ lab_config.yaml, hw_prices.yaml
csv/ the canonical raw corpus + MANIFEST.csv
json/ full/ (per-coupon JSON-LD), hybrid/dataset.jsonld
sqlite/ tribo.db
pg/ tribo.dump (pg_dump -Fc; the live DB is in PostgreSQL)
rdf/ dataset.nt.gz, dataset.ttl, oxigraph_store/
bench/ queries/, expected/, results_raw.csv, results_median.csv,
storage_sizes.csv, extrapolation.csv, hardware_sizing.csv
report/ REPORT.md, charts/, tables/, process_flow.md, diagrams/
.done/ completion markers, one per task
```
- A task writes ONLY under its own output paths; it treats its inputs as
read-only.
- Tasks communicate exclusively through these artifacts - never through
shared in-process state
(see [code-python-style.md](code-python-style.md)).
- **`./out/` is git-ignored** (except nothing - the whole tree). Sources
of truth are the specs, the rules, and the code; artifacts are
reproducible from them ([data-determinism.md](data-determinism.md)).
Never commit generated artifacts "for convenience".
## 2. Completion markers
- Each task NN writes `./out/.done/<NN>.ok` after - and only after - its
outputs are complete and validated. The marker body carries the
headline counts (rows, bytes, triples) downstream tasks check
(see [test-pipeline-validation.md](test-pipeline-validation.md)).
- A task verifies its dependencies' markers before starting and aborts
loudly if one is missing.
- Re-running a task: delete its own marker first, regenerate outputs,
rewrite the marker. If its outputs changed, downstream markers are now
stale - delete them too (the dependency graph below says which).
## 3. Task graph and execution order
```
01 -> 02
01 -> 03 -> (04 | 05 | 06 | 07) -> 08 -> 09 -> 10 -> 11
09 ------> 11
```
- Sequential order: 01, 02, 03, then 04-07 (independent, may run in
parallel), 08, 09, 10, 11.
- 02 (process diagrams) needs only 01 and can run any time after it.
- Task 09 (benchmark execution) must run WITHOUT 04-07 style parallel
work in the background - measurements need a quiet machine
([bench-methodology.md](bench-methodology.md) section 6).
## 4. `storage_sizes.csv` - the shared append contract
- Location: `./out/bench/storage_sizes.csv`; columns:
`format, variant, bytes`.
- Tasks 04-07 APPEND their format's rows (e.g. `json,full,...`,
`json,hybrid,...`, `sqlite,db,...`, `pg,live,...`, `pg,dump,...`,
`rdf,nt_gz,...`, `rdf,store,...`); task 09 appends the gzip archival
variants and consolidates.
- Appends are idempotent per (format, variant): a re-run REPLACES its own
prior rows, never duplicates them and never touches other formats'
rows.
- RDF load wall-time and similar per-format load metrics recorded per
spec 07 go into the task's `.done` marker and the report, not into
extra ad-hoc files.
## 5. Runs are resumable, not restartable-from-zero
- The marker system exists so a failed pipeline resumes at the failed
task; do not delete the whole `./out/` tree to retry one task.
- Conversely, after a deliberate corpus regeneration (03), ALL downstream
markers and artifacts (04-11) are stale and must be removed - a mixed
tree of old and new artifacts is worse than an empty one
([meta-no-assumptions.md](meta-no-assumptions.md) applies before any
such deletion).
## 6. Anti-patterns
- **Writing an artifact outside `./out/`** or a task writing into another
task's output directory.
- **Committing `./out/` content to git.**
- **A marker written on partial success** - see
[code-error-handling.md](code-error-handling.md).
- **Running benchmarks while converters are still running.**
- **Duplicate rows in `storage_sizes.csv`** after re-runs - replace your
own rows.
- **Hand-editing generated artifacts** (a "quick fix" inside
`MANIFEST.csv` or a result CSV) - regenerate them through the owning
task instead.