fix(docs): fix formatting issues in README.md
- adapt meta-/code-/obs- rules to the local Python benchmark pipeline - replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents - add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks - normalize specs and README typography to ASCII per code-data-formatting - rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore - specs: pipeline plan and task specs 00-11 - rules: 19 binding rule files adapted for this project - docs: CLAUDE.md rule-trigger table - config: .cursorrules commit convention, .gitignore excluding out/ and .venv/ - docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders - config: requirements.txt for the closed dependency list - datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
This commit is contained in:
93
docs/rules/code-python-style.md
Normal file
93
docs/rules/code-python-style.md
Normal file
@@ -0,0 +1,93 @@
|
||||
# Code / Python style - tabs, 3.11+, closed dependency list, streaming
|
||||
|
||||
**Purpose + scope.** Baseline style and dependency policy for every Python
|
||||
file in this repository (pipeline tasks 01-11, query implementations,
|
||||
helper tools). Binding for all contributors.
|
||||
|
||||
---
|
||||
|
||||
## 1. Language and formatting
|
||||
|
||||
- **Python 3.11+.** Use modern stdlib freely (`pathlib`, `dataclasses`,
|
||||
`tomllib` if needed); no compatibility shims for older versions.
|
||||
- **Indentation: TABS.** Set by the project owner and repeated in the
|
||||
specs (00_PLAN, 03, 09). Every Python file in this repository indents
|
||||
with tab characters, never spaces. Configure the editor per file type;
|
||||
do not "fix" tab-indented files to spaces.
|
||||
- `from __future__ import annotations` at the top of every module.
|
||||
- Type hints on public function signatures; internal helpers at author's
|
||||
discretion.
|
||||
- English for all identifiers, docstrings, and comments (see
|
||||
[meta-language-style.md](meta-language-style.md) - Russian is the
|
||||
conversation language, not the code language).
|
||||
- Comments follow [code-comments-policy.md](code-comments-policy.md):
|
||||
default to none, keep only non-obvious WHY.
|
||||
|
||||
## 2. Dependencies - a closed list
|
||||
|
||||
The stdlib is the default. The ONLY permitted third-party packages:
|
||||
|
||||
| Package | Used by | For |
|
||||
|---|---|---|
|
||||
| `numpy` | 03 | vectorized generation, seeded RNG |
|
||||
| `PyYAML` | 01, 03, 10 | reading/writing the YAML configs |
|
||||
| `rdflib` | 04, 07 | JSON-LD validation, RDF serialization |
|
||||
| `oxigraph` | 07, 09 | embedded queryable triplestore |
|
||||
| `ijson` | 08, 09 | streaming JSON parsing |
|
||||
| `psycopg` (v3) | 06, 08, 09 | PostgreSQL COPY and queries |
|
||||
| `psutil` | 09 | RSS / CPU / IO measurement |
|
||||
| `matplotlib` | 11 | report charts |
|
||||
| `pandas` | 08 (optional) | ONLY as the separately-measured second CSV query variant; never in the pipeline itself |
|
||||
|
||||
Adding any other package requires updating this table first. In
|
||||
particular forbidden: ORMs, logging frameworks
|
||||
([obs-python-logging.md](obs-python-logging.md)), CLI frameworks
|
||||
(argparse suffices), `requests`/HTTP clients (nothing to call).
|
||||
|
||||
Task 03 (data generation) is stricter by spec: **stdlib + numpy only.**
|
||||
|
||||
## 3. Script shape - every task is a self-contained CLI
|
||||
|
||||
- One entry script per pipeline task, runnable as
|
||||
`python <script>.py [flags]`.
|
||||
- `argparse` with `--log-level` (see
|
||||
[obs-python-logging.md](obs-python-logging.md)) and task-appropriate
|
||||
per-run flags ([code-config-yaml.md](code-config-yaml.md) section 4).
|
||||
- A module docstring stating the task number, what it produces, and its
|
||||
spec file.
|
||||
- `main() -> int` returning the exit code; `sys.exit(main())` under
|
||||
`if __name__ == "__main__":`.
|
||||
- Exit code 0 only on full success (outputs written, validation passed,
|
||||
marker written).
|
||||
|
||||
## 4. Streaming discipline - the corpus never fits in memory by design
|
||||
|
||||
The generated corpus is ~150 MB and its scaled-up target is 600 GB; code
|
||||
must behave as if the data never fits in RAM:
|
||||
|
||||
- **Write generated rows as you produce them**; never accumulate the full
|
||||
cycle/loop dataset in a list (spec 03: "stream rows; never hold the
|
||||
full cycle dataset in memory").
|
||||
- **Read CSV/JSON with iterators** (`csv.reader` row-by-row, `ijson`
|
||||
events); never `read()` a bulk file whole or `json.load` the full
|
||||
corpus.
|
||||
- Batch database work (`executemany` / `COPY` in chunks), keeping only
|
||||
the current batch in memory.
|
||||
- Bounded aggregation state is fine (per-track summaries, counters);
|
||||
unbounded accumulation of raw rows is the thing being forbidden.
|
||||
|
||||
## 5. Anti-patterns
|
||||
|
||||
- **Spaces for indentation** in Python files - the project standard is
|
||||
tabs.
|
||||
- **An unlisted third-party import** - extend the section 2 table first
|
||||
or use the stdlib.
|
||||
- **`pd.read_csv` in pipeline code** - pandas exists only as the
|
||||
separately-measured Q-variant in spec 08.
|
||||
- **Loading a bulk artifact whole** (`json.load` on the full corpus,
|
||||
`readlines()` on a cycles CSV) - stream it.
|
||||
- **A task importable only from a specific CWD** - resolve paths from a
|
||||
repo-root anchor, not the current directory.
|
||||
- **Two tasks sharing mutable module state** - tasks communicate through
|
||||
`./out/` artifacts only (see
|
||||
[build-pipeline-tasks.md](build-pipeline-tasks.md)).
|
||||
Reference in New Issue
Block a user