- adapt meta-/code-/obs- rules to the local Python benchmark pipeline - replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents - add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks - normalize specs and README typography to ASCII per code-data-formatting - rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore - specs: pipeline plan and task specs 00-11 - rules: 19 binding rule files adapted for this project - docs: CLAUDE.md rule-trigger table - config: .cursorrules commit convention, .gitignore excluding out/ and .venv/ - docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders - config: requirements.txt for the closed dependency list - datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
1.8 KiB
1.8 KiB
05 - Convert to SQLite (optimized)
Context / Goal
Build the single-file relational variant optimized for size and single-user query speed.
Dependencies
Input
CSV tree.
Processing instructions
- One database
./out/sqlite/tribo.db. PRAGMAs during load:journal_mode=OFF, synchronous=OFF, cache_size=-262144; after load:journal_mode=WAL,ANALYZE,VACUUM. - Schema (16-table style, integer surrogate keys everywhere; no TEXT foreign keys in bulk tables):
batches, wafers, coupons, runs, instruments, tracks- dimension tables, TEXT natural keys kept as unique columns.xrf_points(coupon_id INT, gx INT, gy INT, pt_wtpct REAL, au_wtpct REAL).profilometry_points(coupon_id INT, gx INT, gy INT, thickness_um REAL).nanoindentation(coupon_id INT, indent_id INT, x_um REAL, y_um REAL, hardness_gpa REAL, reduced_modulus_gpa REAL).afm(coupon_id INT PRIMARY KEY, ra_nm REAL, rq_nm REAL).friction_cycles(track_id INT, cycle INT, cof REAL, PRIMARY KEY(track_id, cycle)) WITHOUT ROWID.friction_loop_points(track_id INT, cycle INT, pt INT, position_um REAL, force_mn REAL, PRIMARY KEY(track_id, cycle, pt)) WITHOUT ROWID.wear(track_id INT PRIMARY KEY, wear_volume_um3 REAL, k_archard REAL, sliding_distance_m REAL).track_summary(track_id INT PRIMARY KEY, cof_ss_mean REAL, cof_ss_std REAL, run_in_cycles INT)- precomputed during load.
- Indexes for Q1-Q7:
coupons(au_wtpct_mean),tracks(run_id),tracks(coupon_id),runs(environment), compositetracks(environment, load_mn)(denormalize environment/load onto tracks). - Batch inserts with
executemany, 50k rows per transaction. - Validate row counts against MANIFEST.
Output
./out/sqlite/tribo.db; size appended tostorage_sizes.csv.- Marker
./out/.done/05.ok.