diff --git a/README.md b/README.md
index 1debf8f..a418800 100644
--- a/README.md
+++ b/README.md
@@ -3,7 +3,8 @@
A reproducible pipeline that answers one practical question: **which storage
format should a materials-science laboratory use for its measurement data?**
-The pipeline simulates a tribology laboratory (Sandia Pt-Au LDRD context),
+The pipeline simulates a tribology laboratory (Pt-Au LDRD context, Sandia
+National Laboratories),
generates ~150 MB of physically plausible raw measurement data in CSV,
converts the corpus into four additional storage formats, benchmarks seven
laboratory-standard retrieval scenarios against all five formats, extrapolates
@@ -145,7 +146,8 @@ docs/
rules/ binding conventions (HOW work is done)
research/ measured design studies (e.g. FK key-schema study)
examples/ imported reference materials (not binding)
-out/ ALL generated artifacts (git-ignored, reproducible):
+out/ ALL generated artifacts (reproducible; bulk git-ignored,
+ report/ config/ .done/ committed - see .gitignore):
config/ csv/ json/ sqlite/ pg/ rdf/ bench/ report/ .done/
common/ shared helpers (storage_sizes.csv contract, ...)
queries/ Q1-Q7 implementations for csv / json / rdf formats
@@ -163,10 +165,13 @@ report.py task 11 entry script (REPORT.md, charts, tables, scoring)
requirements.txt closed dependency list (docs/rules/code-python-style.md)
```
-The `out/` tree - including the ~150 MB CSV corpus, the ~1 GB JSON-LD variant,
-databases, and benchmark results - is **excluded from git**. Sources of truth
-are the specs, the rules, and the code; every artifact regenerates
-deterministically from them. [docs/Initial_Prompt.md](docs/Initial_Prompt.md)
+The bulk of the `out/` tree - the ~150 MB CSV corpus, the JSON-LD variants,
+databases, stores, archives, and raw benchmark data - is **excluded from
+git**; every artifact regenerates deterministically from the specs, the
+rules, and the code. A curated results subset IS committed (decision of
+2026-07-11): `out/report/` (REPORT.md, charts, tables, diagrams),
+`out/config/` (lab_config.yaml, hw_prices.yaml), and the `out/.done/`
+markers - see `.gitignore`. [docs/Initial_Prompt.md](docs/Initial_Prompt.md)
is the original prompt the 12 specification files were generated from - kept
for provenance only, never authoritative: where it disagrees with the specs,
rules, or code (e.g. the SQLite key schema, the JSON-LD vocabulary), those win.
@@ -203,10 +208,14 @@ Benchmark execution is planned for a dedicated Linux host.
## Results
-This section is populated by tasks 02, 09, 10, and 11. All artifacts land
-under `out/` (git-ignored); the placeholders below name the exact files the
-pipeline produces, so the section can be filled in by copying the generated
-tables and linking the generated images.
+Results of the full pipeline run of 2026-07-11 (seed 20260711, 141.8 MiB
+corpus, all 35 benchmark cells checksum-valid). The generated sources of
+truth are committed under [out/report/](out/report/) - the full decision
+document is [out/report/REPORT.md](out/report/REPORT.md); the tables below
+are copied from it. An interactive single-file companion (sidebar
+navigation, per-query and per-scale views, cost explorer) is
+[docs/research/Tribology_Storage_Evaluation.html](docs/research/Tribology_Storage_Evaluation.html).
+Bulk measurement files (`out/bench/`) are git-ignored and regenerable.
### Process flow diagrams (task 02)
@@ -214,21 +223,58 @@ tables and linking the generated images.
D1 coupon assembly, D2 characterization and batch assembly, D3 testing tree,
D4 tribometer session sequence, D5 execution loop, D6 data hierarchy.
-### Measured results (task 09) - to be generated
+### Measured results (task 09)
-| Placeholder | Source file |
-|---|---|
-| Storage footprint per format (measured + compressed) | `out/bench/storage_sizes.csv` |
-| Q1-Q7 median wall time / peak RAM / CPU / read MB | `out/bench/results_median.csv` |
+Storage footprint per format (working + compressed archival):
-### Projections (task 10) - to be generated
+| Format | Measured | vs CSV | Archival (compressed) |
+|---|---|---|---|
+| CSV | 149 MB | 1.00x | 52.8 MB (tar.gz) |
+| JSON-LD (full) | 311 MB | 2.09x | 57.2 MB (tar.gz) |
+| SQLite | 123 MB | 0.83x | 55.1 MB (gzip) |
+| PostgreSQL | 397 MB | 2.67x | 53.1 MB (pg_dump -Fc) |
+| RDF triplestore | 1.74 GB | 11.70x | 126 MB (nt.gz) |
-| Placeholder | Source file |
-|---|---|
-| Projected wall time per format at 0.6 / 1.2 / 6 TB with IMPRACTICAL/FAIL flags | `out/bench/extrapolation.csv` |
-| Hardware sizing and cost per format at scale | `out/bench/hardware_sizing.csv` |
+Median wall time on the 150 MB corpus, warm cache (peak RAM / CPU / read
+bytes: [out/report/tables/t2_measured_medians.csv](out/report/tables/t2_measured_medians.csv)):
-### Charts (task 11) - to be generated
+| Format | Q1 | Q2 | Q3 | Q4 | Q5 | Q6 | Q7 |
+|---|---|---|---|---|---|---|---|
+| CSV | 0.06 s | 0.17 s | 1.15 s | 0.72 s | 1.09 s | 0.28 s | 1.27 s |
+| JSON-LD | 0.11 s | 1.48 s | 9.9 s | 9.9 s | 10 s | 9.8 s | 10 s |
+| SQLite | 0.06 s | 0.06 s | 0.11 s | 0.06 s | 0.77 s | 0.06 s | 0.71 s |
+| PostgreSQL | 0.22 s | 0.22 s | 0.22 s | 0.22 s | 0.44 s | 0.22 s | 0.33 s |
+| RDF triplestore | 0.11 s | 2.81 s | 46 s | 20 s | 60 s | 0.17 s | 41 s |
+
+### Projections (task 10)
+
+Projected wall times at 600 GB - `(!)` = IMPRACTICAL (over 1 hour),
+`FAIL` = over 24 hours; 1.2 TB and 6 TB scales:
+[out/report/tables/t3_projected_600gb.csv](out/report/tables/t3_projected_600gb.csv)
+and `out/bench/extrapolation.csv`:
+
+| Format | Q1 | Q2 | Q3 | Q4 | Q5 | Q6 | Q7 |
+|---|---|---|---|---|---|---|---|
+| CSV | 0.06 s | 7.1 min | 1.2 h (!) | 44 min | 1.2 h (!) | 15 min | 1.4 h (!) |
+| JSON-LD | 0.11 s | 1.5 h (!) | 11 h (!) | 11 h (!) | 11 h (!) | 10.9 h (!) | 11.2 h (!) |
+| SQLite | 0.06 s | 1.3 s | 5.7 min | 0.06 s | 47 min | 5.2 s | 44 min |
+| PostgreSQL | 0.22 s | 0.22 s | 3.4 s | 11 s | 3.7 min | 4.1 s | 1.9 min |
+| RDF triplestore | 0.11 s | 3.0 h (!) | 2.1 d FAIL | 22 h (!) | 2.8 d FAIL | 5.9 min | 1.9 d FAIL |
+
+Hardware sizing at 600 GB - bill of materials with street prices as of
+2026-07-11 (component sources: REPORT.md section 7 and
+[out/config/hw_prices.yaml](out/config/hw_prices.yaml)):
+
+| Format | Configuration | Total |
+|---|---|---|
+| CSV, JSON-LD, SQLite, PostgreSQL (each) | 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM ($1,887); 1 x 7.68 TB NVMe U.2 ($3,995) | $10,500 |
+| RDF triplestore | 2 nodes; 22 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2 | $58,740 |
+
+At 6 TB PostgreSQL still fits one 16-core node - 9 RDIMM modules, 3 NVMe
+drives, $33,586; the RDF triplestore needs 14 nodes and $527,732
+(projected; full table: `out/bench/hardware_sizing.csv`).
+
+### Charts (task 11)


@@ -237,8 +283,9 @@ D4 tribometer session sequence, D5 execution loop, D6 data hierarchy.


-(Images appear after running the full pipeline; task 11 writes them to
-`out/report/charts/` under exactly these names.)
+(Task 11 writes the charts to `out/report/charts/` under exactly these
+names; c3 covers all seven queries - `c3_q1..q7_wall_time.png` - and c4 has
+three query classes: `point_read`, `indexed`, `full_scan`.)
### Final scoreboard (task 11)
diff --git a/docs/research/Tribology_Storage_Evaluation.html b/docs/research/Tribology_Storage_Evaluation.html
new file mode 100644
index 0000000..83f8f46
--- /dev/null
+++ b/docs/research/Tribology_Storage_Evaluation.html
@@ -0,0 +1,812 @@
+
+
+
+
+
+Storage Format Evaluation - Tribology Lab Data
+
+
+
+
+
+
+
+
+
+
SE
+
+ Storage Format Evaluation
+ Tribology lab data · CSV vs JSON-LD vs SQLite vs PostgreSQL vs RDF
+
+
+ Git repository
+ 7 queries x 5 formats · 35 cells
+ seed 20260711
+ Measured 2026-07-11 · all checksums OK
+
+
+
+
+
+
+
+
+
+
+
+
Dashboard
+
A 141.8 MiB physically plausible CSV corpus (the canonical raw format) was converted into four other storage formats and benchmarked on seven laboratory retrieval scenarios with subprocess-isolated resource metering; results were extrapolated to 600 GB, 1.2 TB and 6 TB. Every measured run validated its rows against canonical results (PostgreSQL cross-validated with SQLite, tolerance 1e-9).
+
+
Benchmark cells
35 / 35
OK; 0 invalid, 0 timeouts
+
Corpus
141.8 MiB
8,718 CSV files, byte-reproducible
+
Winner @600 GB
58.6 / 80
PostgreSQL weighted score
+
Worst blow-up
11.7x
RDF store vs raw CSV size
+
Session
916 s
whole matrix, quiet machine
+
+
Recommendation (confirmed by data): PostgreSQL as the system of record; CSV retained as the raw archive (best compressor, ~2.8x); hybrid JSON-LD for inter-lab exchange; RDF only as a virtual layer (e.g. Ontop OBDA) over PostgreSQL - a materialized triplestore fails at scale on speed, RAM and cost simultaneously.
+
+
Typical query time per format (geometric mean Q1-Q7)
+
+
Left value: measured on the 150 MB corpus. Right value in parentheses: projected at 600 GB. Linear bars scaled within the measured values; every projected number in this document is labeled as projected.
+
+
+
Where each format breaks
+
+
Format
Strength
Breaking point
@600 GB flags
+
+
+
+
+
+
+
+
Lab Model & Corpus
+
A simulated tribology laboratory (Pt-Au LDRD study context, Sandia National Laboratories): Ti-6Al-4V coupons with a sputtered Cr adhesion layer and a Pt-Au composition-gradient coating, tested on a 6-probe tribometer in two environments. All data is generated deterministically from seed 20260711; regeneration is byte-identical (MANIFEST.csv sha256 proof).
SPARQL; run-in/steady-state derived from raw cycles
+
+
+
+
+
+
+
+
Methodology
+
The product of this study is its measurements; the protocol is designed so a sloppy number cannot survive it (docs/rules/bench-methodology.md).
+
+
+
Measurement protocol
+
+
One subprocess per run - RSS / CPU / read bytes attribute to exactly that cell; psutil samples the whole process tree every 50 ms (the venv launcher stub alone would report 4 MB).
+
1 warm-up + 3 measured runs per cell; medians reported, min/max kept in raw data.
+
Randomized cell order (deterministic shuffle from the corpus seed); order recorded.
+
30-minute hard timeout per run; a timeout is recorded as data, never retried.
+
Correctness gates speed: every run's rows are compared to the canonical expected results (sorted rows, float tolerance 1e-9); a fast run that fails the checksum is an invalid measurement.
+
+
+
+
Fairness and honesty
+
+
Identical result-set contract for all five implementations of each query; no format answers a simplified question.
+
Each format plays its idiomatic strength (CSV direct path read, SQL indexes, matviews) - but no LIMIT/sampling that changes the answer.
+
Q5 must scan raw cycles in every format; the PostgreSQL EXPLAIN plan is checked to prove track_summary was not used.
+
Cold-cache pass skipped (Windows host has no page-cache drop) - marked, not faked.
+
PostgreSQL is remote: client wall times include the LAN round-trip; server-side cost captured via pg_stat_statements deltas and Prometheus host windows.
+
+
+
+
+
The seven retrieval scenarios
+
+
ID
Scenario
Result rows
+
+
+
+
+
Environment (recorded per session)
+
+
+
+
+
+
+
Storage Footprint
+
Working footprint is what the query engine reads; archival is the compressed form for shelving. All archival sizes converge to 53-57 MB - the information content is the same, the working overhead is not.
+
+
Most compact
0.83x
SQLite: 123 MB, indexes included
+
PostgreSQL
2.67x
397 MB live; 34.1% of it is indexes
+
RDF store
11.70x
1.74 GB for the same data
+
Best archive
52.8 MB
csv tree tar.gz (2.8x compression)
+
+
+
Working footprint, measured on the 150 MB corpus
+
+
+
+
+
Archival (compressed) footprint
+
+
pg_dump -Fc is internally gzip-compressed; RDF also has a Turtle serialization (596 MB, not the archival form). Hybrid JSON-LD is 1.9 MB of metadata + sourceFile links.
+
+
+
Projected working footprint (linear scaling)
+
+
Format
@600 GB
@1.2 TB
@6 TB
+
+
+
All projected. Storage scales linearly in data volume; coefficients are the measured ratios above.
+
+
+
+
+
+
+
Measured Results 150 MB corpus · warm cache
+
Median of 3 isolated runs per cell. Pick a query and a metric; bars are linear with the actual value on each bar. Client peak RSS below ~50 ms of runtime under-reports (sampling cadence is spec-fixed at 50 ms) - wall times are exact.
+
+
+
+
+
+
+
+
Full median matrix - wall time
+
+
+
+
PostgreSQL server side (remote host, 4 cores / 3.8 GiB)
+
+
Query
Server exec per call
Shared blocks hit
Read from disk
+
+
+
pg_stat_statements deltas over the session (4 calls per query). Zero disk reads: the whole 397 MB database sits in the 1 GB shared_buffers. Host peak CPU during the session: 9.2%; RAM delta ~5 MB - the server is nowhere near its limits at this scale.
+
+
+
+
+
+
Projections every number here is projected, not measured
+
Per-cell scaling laws calibrated on the measured 150 MB point after subtracting each format's harness floor (interpreter start, imports, connection). Flags: IMPRACTICAL over 1 hour, FAIL over 24 hours. Pick a target scale.
+
+
+
+
+
+
+
+
+
Scaling law per access pattern
+
+
Law
Growth
Applied to
+
+
O(1)
flat
CSV / JSON-LD Q1 (one file, size fixed)
+
O(log n)
index depth
SQLite / PostgreSQL / RDF Q1
+
O(log n + k)
~linear (result set grows)
relational + RDF Q2, Q4
+
O(k log n)
linear x depth
relational Q3, Q6; RDF Q6
+
O(n)
linear, 1 core
all CSV/JSON scans; SQLite Q5/Q7; RDF Q3/Q5/Q7
+
O(n / cores)
linear / parallel workers
PostgreSQL Q5, Q7 (partitioned scan)
+
+
+
+
+
Full-scan degradation (Q5 + Q7 geometric mean)
+
+
First column measured, the rest projected. PostgreSQL rides its parallel partitioned scan; single-threaded engines grow linearly.
+
+
+
PostgreSQL is the only format with all seven queries below one hour at every scale - Q5 at 6 TB projects to 37 minutes on 16 cores. SQLite holds to 600 GB, then its single-threaded scans pass the hour mark. JSON-LD full scans reach 4.6 days at 6 TB; RDF Q5 reaches 27.8 days.
+
+
+
+
+
Hardware & Cost street prices as of 2026-07-11
+
Bill-of-materials costing: whole 64 GB RDIMM modules, whole 7.68 TB NVMe drives, a priced 1U chassis; disk includes 30% free-space headroom. RAM model: relational hot set = 10% of index size + 4 GB working set; RDF hot set = 20% of the store; streaming formats are flat. Single-node RAM ceiling: 1 TB.
Spot prices during a documented DRAM surge; re-cost by editing out/config/hw_prices.yaml and re-running tasks 10-11.
+
+
+
+
Cost vs speed at 600 GB (projected)
+
+
+
+
+
+
+
Scoring & Verdict
+
Subscores are 10 x best/value per metric, weighted: search speed x5 (0-50, inverse geometric mean of Q1-Q7), RAM economy x2 (0-20), disk economy x1 (0-10); maximum 80. A FAIL projection at 600 GB zeroes the search subscore.
+
+
+
Measured scale (150 MB)
+
search x5RAM x2disk x1
+
+
At laptop scale SQLite is the honest winner: near-instant indexed queries in one file with the smallest footprint.
+
+
+
Projected at 600 GB
+
search x5RAM x2disk x1
+
+
At production scale the ranking flips: parallel indexed retrieval dominates the weights, and PostgreSQL takes the lead.
+
+
+
+
Use-case mapping
+
+
Use case
Recommended format
+
+
Interactive analysis (joins, aggregations)
PostgreSQL; SQLite acceptable single-user up to ~600 GB
+
Report generation (repeated summaries)
PostgreSQL (track_summary materialized view)
+
Search / filtering
PostgreSQL or SQLite (indexed); flat formats need full scans
+
Archiving
CSV tree + tar.gz (canonical raw) plus pg_dump of the system of record
+
Inter-lab exchange
Hybrid JSON-LD (metadata + sourceFile links to CSV)
+
Semantic / ontology queries
Virtual RDF layer over PostgreSQL (e.g. Ontop OBDA), not a materialized triplestore
+
+
+
+
+
Limitations
+
+
Single-node measurements on one Windows host + one remote PostgreSQL host; no cluster variance.
+
Simulated corpus (physically plausible, fixed seed); real instrument data may distribute differently.
+
Extrapolation is analytic, calibrated on one 150 MB point; no intermediate-scale validation runs.
+
Warm-cache only (no page-cache drop on Windows); harness floor assumed constant across scales.
+
RDF numbers reflect compact modeling and client-side derivation of run-in/steady-state; a SPARQL-side aggregation engine could shift, not remove, the scan penalty.
+
Hardware prices are spot street prices (2026-07-11) during a DRAM price surge.
+
+
+
+ Generated from the LabDataStorageEvaluation pipeline artifacts (run of 2026-07-11, seed 20260711): out/report/REPORT.md, out/bench/results_median.csv, extrapolation.csv, hardware_sizing.csv, storage_sizes.csv, out/config/hw_prices.yaml. The full decision document with charts is out/report/REPORT.md; this page is the interactive companion. Related study: Foreign-Key Architecture Study.
+
+
+
+
+
+
About
+
Attribution and provenance for this document. It is distributed as a single standalone HTML file; everything needed to verify or reproduce the numbers is referenced below.
+
+
Author
+
This study was conducted by Mary Goncharenko, PhD student at the Tribology Laboratory, University of Florida, Gainesville, FL, during a research internship at the tribology laboratory of Sandia National Laboratories, with an assignment on laboratory ontology development.
+
The research was carried out with the technical assistance of her father, Vasiliy Goncharenko, Chief Technology Architect at SoftCreator, LLC (enterprise data architecture across relational and semantic stores; 35 years in technology).
This page is the interactive companion to the repository's out/report/REPORT.md; all numbers embedded here come from the pipeline run of 2026-07-11. The corpus regenerates byte-identically from seed 20260711 (MANIFEST.csv sha256 proof), so every figure is independently reproducible from the repository alone.
+
+
+
Data statement
+
The measurement corpus is simulated: physically plausible models of a Pt-Au LDRD tribology study context at Sandia National Laboratories, generated from a fixed random seed. It contains no experimental measurements and no export-controlled data. Hardware prices cited in the cost model are public street prices retrieved 2026-07-11 (sources in the Hardware & Cost section).
+
+
+
+
diff --git a/docs/rules/build-pipeline-tasks.md b/docs/rules/build-pipeline-tasks.md
index 2ac57c2..0e56837 100644
--- a/docs/rules/build-pipeline-tasks.md
+++ b/docs/rules/build-pipeline-tasks.md
@@ -31,10 +31,14 @@ out/
- Tasks communicate exclusively through these artifacts - never through
shared in-process state
(see [code-python-style.md](code-python-style.md)).
-- **`./out/` is git-ignored** (except nothing - the whole tree). Sources
- of truth are the specs, the rules, and the code; artifacts are
- reproducible from them ([data-determinism.md](data-determinism.md)).
- Never commit generated artifacts "for convenience".
+- **The bulk of `./out/` is git-ignored.** Sources of truth are the
+ specs, the rules, and the code; artifacts are reproducible from them
+ ([data-determinism.md](data-determinism.md)). A curated results subset
+ IS committed (owner decision, 2026-07-11): `out/report/` (REPORT.md,
+ charts, tables, diagrams), `out/config/` (lab_config.yaml,
+ hw_prices.yaml), and the `out/.done/` markers - exactly the exceptions
+ listed in `.gitignore`. Never commit the bulk artifacts (corpus,
+ databases, stores, archives, raw benchmark data) "for convenience".
## 2. Completion markers
@@ -92,7 +96,9 @@ out/
- **Writing an artifact outside `./out/`** or a task writing into another
task's output directory.
-- **Committing `./out/` content to git.**
+- **Committing bulk `./out/` content to git** (corpus, databases, stores,
+ archives, raw benchmark data) - only the curated subset in `.gitignore`
+ (`report/`, `config/`, `.done/`) is committed.
- **A marker written on partial success** - see
[code-error-handling.md](code-error-handling.md).
- **Running benchmarks while converters are still running.**