- adapt meta-/code-/obs- rules to the local Python benchmark pipeline - replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents - add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks - normalize specs and README typography to ASCII per code-data-formatting - rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore - specs: pipeline plan and task specs 00-11 - rules: 19 binding rule files adapted for this project - docs: CLAUDE.md rule-trigger table - config: .cursorrules commit convention, .gitignore excluding out/ and .venv/ - docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders - config: requirements.txt for the closed dependency list - datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
7.3 KiB
LabDataStorageEvaluation - Project Rules
A reproducible evaluation pipeline for laboratory data storage formats, built around a simulated tribology laboratory (Sandia Pt-Au LDRD context). The pipeline (tasks 01-11, specs in docs/specs/):
- Defines the lab model (
lab_config.yaml) and process-flow diagrams. - Generates ~150 MB of physically plausible measurement data in CSV (the canonical raw format), deterministically from seed 20260711.
- Converts the corpus into 4 target formats: JSON-LD (full + hybrid), SQLite, PostgreSQL 16 (indexed + partitioned), RDF triplestore (oxigraph).
- Benchmarks 7 retrieval scenarios (Q1-Q7) x 5 formats, measuring wall time, peak RAM, CPU, disk I/O, and footprint.
- Extrapolates to 600 GB / 1.2 TB / 6 TB and produces a ranked report
(
REPORT.md+ charts).
Everything runs locally (Windows development host, local PostgreSQL 16,
no cloud services, no server deploy). The stack is Python 3.11+ (tabs for
indentation) with a closed dependency list; there is no UI, no HTTP
API, no authentication in this project. All artifacts land under
./out/ (git-ignored); tasks signal completion via ./out/.done/<NN>.ok
markers.
These rules are binding for every contributor (human or AI agent) touching this repository. They override any pattern found in imported specs, external tutorials, or older project docs that predate this file.
Mandatory rules - load before starting matching work
When the user's request matches a topic below, before writing any code or
producing a plan, run Read on the listed rule file(s). The rules are
binding; ignoring them produces broken work that the user will reject.
| If the task involves... | You MUST first Read |
|---|---|
| Writing or changing any Python (style, tabs, dependency list, streaming discipline) | docs/rules/code-python-style.md |
| Data generation, physical models, random values, seeds, reproducibility | docs/rules/data-determinism.md |
Naming entities/IDs, timestamps, measurement columns, units, au_wtpct_mean |
docs/rules/data-naming-units.md |
SQLite / PostgreSQL schema, DDL, keys, indexes, bulk loading, track_summary |
docs/rules/db-sql-schema.md |
| Benchmark queries, the runner, resource measurement, timeouts, extrapolation, scoring | docs/rules/bench-methodology.md |
The ./out/ tree, .done markers, task order/parallelism, storage_sizes.csv, re-runs |
docs/rules/build-pipeline-tasks.md |
| Validation: MANIFEST, row counts, result checksums, expected volumes, pytest | docs/rules/test-pipeline-validation.md |
Python error handling (try / except / raise, abort semantics) |
docs/rules/code-error-handling.md |
Logging, print() vs logger, log levels |
docs/rules/obs-python-logging.md |
| Configuration values: the YAML configs, constants, the PostgreSQL DSN, CLI flags | docs/rules/code-config-yaml.md |
| Writing, editing, or reviewing comments in source code (any language) | docs/rules/code-comments-policy.md |
| Prose strings in source/config/docs (typography / encoding / dashes) | docs/rules/code-data-formatting.md |
Adding or renaming a rule file in docs/rules/ |
docs/rules/meta-rules-naming.md |
Anything where misinterpretation could be destructive (deleting/regenerating ./out/, dropping the tribo DB, changing the seed or lab_config.yaml, editing specs, git history) |
docs/rules/meta-no-assumptions.md |
| Every response (length / formatting discipline) | docs/rules/meta-concise-answers.md |
| Asking the user any question | docs/rules/meta-concise-questions.md |
| Every response (communication language / register - Russian, no jargon or slang) | docs/rules/meta-language-style.md |
| Any git operation that would create a commit, tag, merge, revert, rebase, cherry-pick, or push | docs/rules/meta-git-commits.md |
Generating commit messages via Cursor, or maintaining the .cursorrules artifact |
docs/rules/meta-cursor-rules.md |
Multiple matches -> Read all matching rule files. Unsure -> Read the closest one; over-loading is cheap, missing a rule is expensive.
Other binding constraints
These are short, cross-cutting constraints. They are still binding; the linked rule file holds the detailed spec.
- Determinism is the foundation. Seed
20260711lives inlab_config.yamland is never hardcoded; only seeded RNGs; no wall-clock values inside the corpus; regeneration is byte-identical (proof:MANIFEST.csvsha256). See docs/rules/data-determinism.md. - A wrong artifact is worse than no artifact. Validation gates (row
counts vs MANIFEST, result checksums with 1e-9 tolerance, the
120-200 MB size gate) abort loudly; a
.donemarker is written only after full success. See docs/rules/code-error-handling.md and docs/rules/test-pipeline-validation.md. - Tasks communicate only through
./out/artifacts. The tree is git-ignored; sources of truth are specs + rules + code. Benchmarks run on a quiet machine. See docs/rules/build-pipeline-tasks.md. - Measurement honesty. Subprocess isolation per cell, 1 warm-up + 3
measured runs, randomized order, 30-min timeout recorded as
TIMEOUT, failed checksums invalidate the measurement, projected numbers are always labeled. See docs/rules/bench-methodology.md. - Python: tabs, closed dependency list. Stdlib first; only the packages listed in docs/rules/code-python-style.md; task 03 is stdlib + numpy only; stream everything bulk.
- Commit message format (mandatory). First line
type(scope): description, lowercase imperative English, under 72 chars. Types (closed):feat fix perf refactor security docs ci chore test. Scopes (closed):specs rules docs config datagen convert bench report tools all. The AI never creates commits - git is read-only for agents; the user runs all git write commands. See docs/rules/meta-git-commits.md.
Canonical source of truth
If this file and any docs/rules/<name>.md disagree, the rule file wins
within its topic - that is where the detailed binding spec lives. If a rule
and a task spec in docs/specs/ disagree on a convention, the rule wins
and the spec is updated to match; the specs remain authoritative for task
scope, inputs, outputs, and dependencies. If this file and any other
document (README, older docs, code comments) disagree, this file plus the
rule files win. Update the other document to match; do not work around
the rules.