Compare commits

...

23 Commits

Author SHA1 Message Date
100e621705 Merge pull request 'feat(docs): add management summary link in evaluation documents' (#13) from dev_masha into master
Reviewed-on: #13
2026-07-13 00:19:41 -04:00
administrator
6d156f98c4 feat(docs): add management summary link in evaluation documents
- include link to one-page executive summary in README
- update dashboard section in evaluation HTML to reference summary
- revise report metadata to reflect additional CSV tables generated
2026-07-13 00:19:02 -04:00
ade412c92d Merge pull request 'fix(docs): update navigation links in Tribology Storage Evaluation' (#12) from dev_masha into master
Reviewed-on: #12
2026-07-12 12:52:51 -04:00
administrator
478d4f3201 fix(docs): update navigation links in Tribology Storage Evaluation 2026-07-12 12:52:30 -04:00
a491d8bedc Merge pull request 'dev_masha' (#11) from dev_masha into master
Reviewed-on: #11
2026-07-12 11:34:13 -04:00
administrator
2796b79170 chore(docs): remove obsolete example files from documentation
- delete Tribology_Ontology.html and TribologyData_MasterPlan_v1.md examples
- remove old planning documents from specs_old directory
2026-07-12 11:33:41 -04:00
administrator
ff68679000 feat(docs): enhance README with interactive result links and references
- add links to interactive HTML results for storage format evaluation and foreign-key architecture study
- update descriptions in README to reflect the new structure of research documents
- clarify the contents of the research directory in the README
2026-07-12 11:30:43 -04:00
6a04c030cf Merge pull request 'feat(docs): enhance README and build pipeline rules for clarity' (#10) from dev_masha into master
Reviewed-on: #10
2026-07-12 11:19:47 -04:00
administrator
b4270b0247 feat(docs): enhance README and build pipeline rules for clarity
- update README to clarify the exclusion of bulk artifacts from git
- specify committed curated results subset in README and rules
- improve descriptions of output directories and their contents
- refine details on the pipeline's data generation and reporting processes
2026-07-12 11:18:16 -04:00
0e8ef93f80 Merge pull request 'feat(report): update extrapolation parameters and scoring in README' (#9) from dev_masha into master
Reviewed-on: #9
2026-07-12 00:02:00 -04:00
administrator
35d2bdadbd feat(report): update extrapolation parameters and scoring in README
- revise hardware prices and parameters in extrapolate.py
- enhance scoring details in README for task 11
- clarify projected scoreboard values and sources
- improve chart readability and scoring methodology in report.py
- add new utility functions for formatting and reading hardware prices
2026-07-11 23:57:41 -04:00
841f5e1f26 Merge pull request 'feat(config): update .gitignore to exclude output artifacts and add config files' (#8) from dev_masha into master
Reviewed-on: #8
2026-07-11 23:52:50 -04:00
administrator
af91629f71 feat(config): update .gitignore to exclude output artifacts and add config files
- ignore all output files except for reports and config directories
- add new .done files for tasks 01-11 to track completion status
- include new configuration files for hardware prices and lab setup
- generate process flow and report documents for tasks 02 and 11
2026-07-11 23:52:28 -04:00
67206ddab3 Merge pull request 'feat(docs): update README and plan to reflect task implementations' (#7) from dev_masha into master
Reviewed-on: #7
2026-07-11 23:18:34 -04:00
administrator
a4c0d258eb feat(docs): update README and plan to reflect task implementations
- mark tasks 10 and 11 as implemented in the project overview
- add provenance note regarding the initial prompt in README.md
- clarify the relationship between the prompt and specifications in 00_PLAN.md
- update documentation structure to include Initial_Prompt.md for historical context
2026-07-11 23:18:07 -04:00
feb0d377b5 Merge pull request 'feat(bench): add SQLite and PostgreSQL wrappers for benchmark queries' (#6) from dev_masha into master
Reviewed-on: #6
2026-07-11 23:17:57 -04:00
administrator
eda11aa63e feat(bench): add SQLite and PostgreSQL wrappers for benchmark queries
- implement SQLite and PostgreSQL wrappers to execute benchmark queries
- update README to reflect task 09 as implemented
- enhance storage size updates to handle (format, variant) replacements
- clarify benchmarking methodology for remote PostgreSQL setup
2026-07-11 22:34:23 -04:00
8f8a871746 Merge pull request 'feat(convert): integrate process metrics into JSON and SQLite converters' (#5) from dev_masha into master
Reviewed-on: #5
2026-07-11 21:42:02 -04:00
administrator
48c102d2a4 feat(convert): integrate process metrics into JSON and SQLite converters
- add process_metrics function to record peak RSS and CPU time
- update JSON converter to include metrics in completion marker
- modify SQLite converter to utilize new metrics for performance tracking
- enhance storage size updates with locking mechanism for concurrent access
2026-07-11 21:39:08 -04:00
17981792fd Merge pull request 'feat(convert): add task 05 sqlite converter with variant B keys" -m "- convert: convert_sqlite.py builds tribo.db (composite PKs, enforced FKs, Q1-Q7 indexes, track_summary)' (#4) from dev_masha into master
Reviewed-on: #4
2026-07-11 16:08:50 -04:00
administrator
96ed7bc918 feat(convert): add task 05 sqlite converter with variant B keys" -m "- convert: convert_sqlite.py builds tribo.db (composite PKs, enforced FKs, Q1-Q7 indexes, track_summary)
- tools: common/track_summary.py fixes the cross-engine steady-state algorithm
- rules: db-sql-schema records the FK enforcement decision and track_summary definition
- specs: 05 aligned with data-naming-units, simtra_profiles added
- docs: FK key-schema study added under docs/research/, README updated
- replace manual CSV reading with CorpusReader for better data handling
- streamline argument parsing and dependency checks using common pipeline functions
- enhance marker writing for task completion tracking
- remove unused regex and validation error classes for cleaner code
2026-07-11 16:08:29 -04:00
4fcbe26ae6 Merge pull request 'feat(docs): update README to reflect JSON-LD conversion implementation' (#3) from dev_masha into master
Reviewed-on: #3
2026-07-11 16:08:23 -04:00
administrator
9364808b3d feat(docs): update README to reflect JSON-LD conversion implementation
- mark task 04 as implemented in the project overview
- add convert_json.py as the entry script for task 04
- clarify output paths for JSON-LD conversion artifacts
2026-07-11 14:21:13 -04:00
97 changed files with 7501 additions and 937 deletions

8
.gitignore vendored
View File

@@ -1,6 +1,12 @@
# Generated pipeline artifacts - reproducible from specs + rules + code # Generated pipeline artifacts - reproducible from specs + rules + code
# (see docs/rules/build-pipeline-tasks.md) # (see docs/rules/build-pipeline-tasks.md)
out/ out/*
!out/report/
!out/report/**
!out/config/
!out/config/**
!out/.done/
!out/.done/**
# Python # Python
__pycache__/ __pycache__/

171
README.md
View File

@@ -3,7 +3,8 @@
A reproducible pipeline that answers one practical question: **which storage A reproducible pipeline that answers one practical question: **which storage
format should a materials-science laboratory use for its measurement data?** format should a materials-science laboratory use for its measurement data?**
The pipeline simulates a tribology laboratory (Sandia Pt-Au LDRD context), The pipeline simulates a tribology laboratory (Pt-Au LDRD context, Sandia
National Laboratories),
generates ~150 MB of physically plausible raw measurement data in CSV, generates ~150 MB of physically plausible raw measurement data in CSV,
converts the corpus into four additional storage formats, benchmarks seven converts the corpus into four additional storage formats, benchmarks seven
laboratory-standard retrieval scenarios against all five formats, extrapolates laboratory-standard retrieval scenarios against all five formats, extrapolates
@@ -11,8 +12,26 @@ the measurements to 600 GB - 6 TB, and produces a ranked decision report for
five laboratory use cases: analysis, reporting, search/filtering, archiving, five laboratory use cases: analysis, reporting, search/filtering, archiving,
and inter-lab exchange. and inter-lab exchange.
**Results - interactive, single-file HTML (open in a browser or download and share):**
- **[Storage Decision](https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision.html)** -
one-page executive summary: the problem, the verdict (relational vs graph
database), and what it costs - written for management, every claim linked
to the full study
(source: [docs/research/Tribology_Storage_Decision.html](docs/research/Tribology_Storage_Decision.html)).
- **[Storage Format Evaluation](https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html)** -
the full study results: dashboard, measured matrix, projections to 6 TB,
hardware bill of materials, weighted scoring and verdict
(source: [docs/research/Tribology_Storage_Evaluation.html](docs/research/Tribology_Storage_Evaluation.html)).
- **[Foreign-Key Architecture Study](https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_FK_Architecture.html)** -
the measured key-schema study behind the relational design decision
(source: [docs/research/Tribology_FK_Architecture.html](docs/research/Tribology_FK_Architecture.html)).
Reference documents:
- Task specifications: [docs/specs/](docs/specs/) (plan + tasks 01-11). - Task specifications: [docs/specs/](docs/specs/) (plan + tasks 01-11).
- Binding conventions: [docs/rules/](docs/rules/) and [CLAUDE.md](CLAUDE.md). - Binding conventions: [docs/rules/](docs/rules/) and [CLAUDE.md](CLAUDE.md).
- Full decision report (markdown): [out/report/REPORT.md](out/report/REPORT.md).
## Goals ## Goals
@@ -60,14 +79,14 @@ machine (no parallel work). Every task writes a completion marker
| 01 Lab configuration | `make_lab_config.py` | `out/config/lab_config.yaml` | implemented | | 01 Lab configuration | `make_lab_config.py` | `out/config/lab_config.yaml` | implemented |
| 02 Process flow diagrams | `make_process_flow.py` | `out/report/process_flow.md`, `out/report/diagrams/` | implemented | | 02 Process flow diagrams | `make_process_flow.py` | `out/report/process_flow.md`, `out/report/diagrams/` | implemented |
| 03 CSV data generation | `generate_data.py` | `out/csv/` tree + `MANIFEST.csv` | implemented | | 03 CSV data generation | `generate_data.py` | `out/csv/` tree + `MANIFEST.csv` | implemented |
| 04 Convert to JSON-LD | `convert_json.py` | `out/json/full/`, `out/json/hybrid/dataset.jsonld` | planned | | 04 Convert to JSON-LD | `convert_json.py` | `out/json/full/`, `out/json/hybrid/dataset.jsonld` | implemented |
| 05 Convert to SQLite | `convert_sqlite.py` | `out/sqlite/tribo.db` | planned | | 05 Convert to SQLite | `convert_sqlite.py` | `out/sqlite/tribo.db` | implemented |
| 06 Convert to PostgreSQL | `convert_postgresql.py` | live `tribo` DB + `out/pg/tribo.dump` | planned | | 06 Convert to PostgreSQL | `convert_postgresql.py` | live PostgreSQL DB (+ dump when pg_dump is available) | implemented |
| 07 Convert to RDF | `convert_rdf.py` | `out/rdf/dataset.nt.gz`, oxigraph store | planned | | 07 Convert to RDF | `convert_rdf.py` | `out/rdf/dataset.nt.gz`, `dataset.ttl`, oxigraph store | implemented |
| 08 Benchmark queries | `make_queries.py` | `out/bench/queries/`, `out/bench/expected/` | planned | | 08 Benchmark queries | `make_queries.py` | `out/bench/queries/`, `out/bench/expected/` | implemented |
| 09 Benchmark execution | `bench_runner.py` | `out/bench/results_raw.csv`, `results_median.csv` | planned | | 09 Benchmark execution | `bench_runner.py` | `out/bench/results_raw.csv`, `results_median.csv` | implemented |
| 10 Extrapolation | `extrapolate.py` | `out/bench/extrapolation.csv`, `hardware_sizing.csv` | planned | | 10 Extrapolation | `extrapolate.py` | `out/bench/extrapolation.csv`, `hardware_sizing.csv` | implemented |
| 11 Reporting | `report.py` | `out/report/REPORT.md`, charts, tables | planned | | 11 Reporting | `report.py` | `out/report/REPORT.md`, charts, tables | implemented |
## The simulated laboratory ## The simulated laboratory
@@ -140,21 +159,39 @@ measurement counts. See
``` ```
docs/ docs/
Initial_Prompt.md the generative prompt behind specs 00-11 (provenance, non-binding)
specs/ task specifications 00-11 (WHAT to build) specs/ task specifications 00-11 (WHAT to build)
rules/ binding conventions (HOW work is done) rules/ binding conventions (HOW work is done)
examples/ imported reference materials (not binding) research/ measured design studies and interactive result pages
out/ ALL generated artifacts (git-ignored, reproducible): out/ ALL generated artifacts (reproducible; bulk git-ignored,
report/ config/ .done/ committed - see .gitignore):
config/ csv/ json/ sqlite/ pg/ rdf/ bench/ report/ .done/ config/ csv/ json/ sqlite/ pg/ rdf/ bench/ report/ .done/
common/ shared helpers (storage_sizes.csv contract, ...)
queries/ Q1-Q7 implementations for csv / json / rdf formats
make_lab_config.py task 01 entry script (one script per task, repo root) make_lab_config.py task 01 entry script (one script per task, repo root)
make_process_flow.py task 02 entry script make_process_flow.py task 02 entry script
generate_data.py task 03 entry script generate_data.py task 03 entry script
convert_json.py task 04 entry script
convert_sqlite.py task 05 entry script
convert_postgresql.py task 06 entry script (TRIBO_PG_DSN, TRIBO_PROM_URL)
convert_rdf.py task 07 entry script
make_queries.py task 08 entry script (canonical results need TRIBO_PG_DSN)
bench_runner.py task 09 entry script (TRIBO_PG_DSN, TRIBO_PROM_URL)
extrapolate.py task 10 entry script (seeds editable out/config/hw_prices.yaml)
report.py task 11 entry script (REPORT.md, charts, tables, scoring)
requirements.txt closed dependency list (docs/rules/code-python-style.md) requirements.txt closed dependency list (docs/rules/code-python-style.md)
``` ```
The `out/` tree - including the ~150 MB CSV corpus, the ~1 GB JSON-LD variant, The bulk of the `out/` tree - the ~150 MB CSV corpus, the JSON-LD variants,
databases, and benchmark results - is **excluded from git**. Sources of truth databases, stores, archives, and raw benchmark data - is **excluded from
are the specs, the rules, and the code; every artifact regenerates git**; every artifact regenerates deterministically from the specs, the
deterministically from them. rules, and the code. A curated results subset IS committed (decision of
2026-07-11): `out/report/` (REPORT.md, charts, tables, diagrams),
`out/config/` (lab_config.yaml, hw_prices.yaml), and the `out/.done/`
markers - see `.gitignore`. [docs/Initial_Prompt.md](docs/Initial_Prompt.md)
is the original prompt the 12 specification files were generated from - kept
for provenance only, never authoritative: where it disagrees with the specs,
rules, or code (e.g. the SQLite key schema, the JSON-LD vocabulary), those win.
## Getting started ## Getting started
@@ -188,10 +225,14 @@ Benchmark execution is planned for a dedicated Linux host.
## Results ## Results
This section is populated by tasks 02, 09, 10, and 11. All artifacts land Results of the full pipeline run of 2026-07-11 (seed 20260711, 141.8 MiB
under `out/` (git-ignored); the placeholders below name the exact files the corpus, all 35 benchmark cells checksum-valid). The generated sources of
pipeline produces, so the section can be filled in by copying the generated truth are committed under [out/report/](out/report/) - the full decision
tables and linking the generated images. document is [out/report/REPORT.md](out/report/REPORT.md); the tables below
are copied from it. An interactive single-file companion (sidebar
navigation, per-query and per-scale views, cost explorer) is
[docs/research/Tribology_Storage_Evaluation.html](docs/research/Tribology_Storage_Evaluation.html).
Bulk measurement files (`out/bench/`) are git-ignored and regenerable.
### Process flow diagrams (task 02) ### Process flow diagrams (task 02)
@@ -199,21 +240,58 @@ tables and linking the generated images.
D1 coupon assembly, D2 characterization and batch assembly, D3 testing tree, D1 coupon assembly, D2 characterization and batch assembly, D3 testing tree,
D4 tribometer session sequence, D5 execution loop, D6 data hierarchy. D4 tribometer session sequence, D5 execution loop, D6 data hierarchy.
### Measured results (task 09) - to be generated ### Measured results (task 09)
| Placeholder | Source file | Storage footprint per format (working + compressed archival):
|---|---|
| Storage footprint per format (measured + compressed) | `out/bench/storage_sizes.csv` |
| Q1-Q7 median wall time / peak RAM / CPU / read MB | `out/bench/results_median.csv` |
### Projections (task 10) - to be generated | Format | Measured | vs CSV | Archival (compressed) |
|---|---|---|---|
| CSV | 149 MB | 1.00x | 52.8 MB (tar.gz) |
| JSON-LD (full) | 311 MB | 2.09x | 57.2 MB (tar.gz) |
| SQLite | 123 MB | 0.83x | 55.1 MB (gzip) |
| PostgreSQL | 397 MB | 2.67x | 53.1 MB (pg_dump -Fc) |
| RDF triplestore | 1.74 GB | 11.70x | 126 MB (nt.gz) |
| Placeholder | Source file | Median wall time on the 150 MB corpus, warm cache (peak RAM / CPU / read
|---|---| bytes: [out/report/tables/t2_measured_medians.csv](out/report/tables/t2_measured_medians.csv)):
| Projected wall time per format at 0.6 / 1.2 / 6 TB with IMPRACTICAL/FAIL flags | `out/bench/extrapolation.csv` |
| Hardware sizing and cost per format at scale | `out/bench/hardware_sizing.csv` |
### Charts (task 11) - to be generated | Format | Q1 | Q2 | Q3 | Q4 | Q5 | Q6 | Q7 |
|---|---|---|---|---|---|---|---|
| CSV | 0.06 s | 0.17 s | 1.15 s | 0.72 s | 1.09 s | 0.28 s | 1.27 s |
| JSON-LD | 0.11 s | 1.48 s | 9.9 s | 9.9 s | 10 s | 9.8 s | 10 s |
| SQLite | 0.06 s | 0.06 s | 0.11 s | 0.06 s | 0.77 s | 0.06 s | 0.71 s |
| PostgreSQL | 0.22 s | 0.22 s | 0.22 s | 0.22 s | 0.44 s | 0.22 s | 0.33 s |
| RDF triplestore | 0.11 s | 2.81 s | 46 s | 20 s | 60 s | 0.17 s | 41 s |
### Projections (task 10)
Projected wall times at 600 GB - `(!)` = IMPRACTICAL (over 1 hour),
`FAIL` = over 24 hours; 1.2 TB and 6 TB scales:
[out/report/tables/t3_projected_600gb.csv](out/report/tables/t3_projected_600gb.csv)
and `out/bench/extrapolation.csv`:
| Format | Q1 | Q2 | Q3 | Q4 | Q5 | Q6 | Q7 |
|---|---|---|---|---|---|---|---|
| CSV | 0.06 s | 7.1 min | 1.2 h (!) | 44 min | 1.2 h (!) | 15 min | 1.4 h (!) |
| JSON-LD | 0.11 s | 1.5 h (!) | 11 h (!) | 11 h (!) | 11 h (!) | 10.9 h (!) | 11.2 h (!) |
| SQLite | 0.06 s | 1.3 s | 5.7 min | 0.06 s | 47 min | 5.2 s | 44 min |
| PostgreSQL | 0.22 s | 0.22 s | 3.4 s | 11 s | 3.7 min | 4.1 s | 1.9 min |
| RDF triplestore | 0.11 s | 3.0 h (!) | 2.1 d FAIL | 22 h (!) | 2.8 d FAIL | 5.9 min | 1.9 d FAIL |
Hardware sizing at 600 GB - bill of materials with street prices as of
2026-07-11 (component sources: REPORT.md section 7 and
[out/config/hw_prices.yaml](out/config/hw_prices.yaml)):
| Format | Configuration | Total |
|---|---|---|
| CSV, JSON-LD, SQLite, PostgreSQL (each) | 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM ($1,887); 1 x 7.68 TB NVMe U.2 ($3,995) | $10,500 |
| RDF triplestore | 2 nodes; 22 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2 | $58,740 |
At 6 TB PostgreSQL still fits one 16-core node - 9 RDIMM modules, 3 NVMe
drives, $33,586; the RDF triplestore needs 14 nodes and $527,732
(projected; full table: `out/bench/hardware_sizing.csv`).
### Charts (task 11)
![C1 - disk footprint per format, log scale](out/report/charts/c1_disk_footprint.png) ![C1 - disk footprint per format, log scale](out/report/charts/c1_disk_footprint.png)
![C2 - RAM requirement per format, log scale](out/report/charts/c2_ram_requirement.png) ![C2 - RAM requirement per format, log scale](out/report/charts/c2_ram_requirement.png)
@@ -222,29 +300,34 @@ D4 tribometer session sequence, D5 execution loop, D6 data hierarchy.
![C5 - weighted score, stacked bars](out/report/charts/c5_weighted_score.png) ![C5 - weighted score, stacked bars](out/report/charts/c5_weighted_score.png)
![C6 - cost vs performance at 600 GB, log-log](out/report/charts/c6_cost_vs_performance.png) ![C6 - cost vs performance at 600 GB, log-log](out/report/charts/c6_cost_vs_performance.png)
(Images appear after running the full pipeline; task 11 writes them to (Task 11 writes the charts to `out/report/charts/` under exactly these
`out/report/charts/` under exactly these names.) names; c3 covers all seven queries - `c3_q1..q7_wall_time.png` - and c4 has
three query classes: `point_read`, `indexed`, `full_scan`.)
### Final scoreboard (task 11) - to be generated ### Final scoreboard (task 11)
Weighted scoring, max 80 points: search speed x5 (0-50), RAM economy x2 Weighted scoring, max 80 points: search speed x5 (0-50), RAM economy x2
(0-20), disk economy x1 (0-10); TIMEOUT/FAIL projects to 0. (0-20), disk economy x1 (0-10); a FAIL projection at 600 GB zeroes the
search subscore. Values below are the projected-600 GB scoreboard from the
2026-07-11 pipeline run (seed 20260711); the generated source of truth is
`out/report/tables/t5_weighted_scores.csv`, and REPORT.md also carries the
measured-scale scoreboard.
| Format | Search (0-50) | RAM (0-20) | Disk (0-10) | Total (0-80) | | Format | Search (0-50) | RAM (0-20) | Disk (0-10) | Total (0-80) |
|---|---|---|---|---| |---|---|---|---|---|
| CSV | TBD | TBD | TBD | TBD | | CSV | 0.6 | 20.0 | 8.3 | 28.9 |
| JSON-LD | TBD | TBD | TBD | TBD | | JSON-LD | 0.1 | 20.0 | 4.0 | 24.0 |
| SQLite | TBD | TBD | TBD | TBD | | SQLite | 21.8 | 15.2 | 10.0 | 47.0 |
| PostgreSQL | TBD | TBD | TBD | TBD | | PostgreSQL | 50.0 | 5.5 | 3.1 | 58.6 |
| RDF | TBD | TBD | TBD | TBD | | RDF | 0.0 | 0.2 | 0.7 | 0.9 |
| Use case | Recommended format | | Use case | Recommended format |
|---|---| |---|---|
| Analysis | TBD | | Analysis | PostgreSQL (SQLite acceptable single-user up to ~600 GB) |
| Report generation | TBD | | Report generation | PostgreSQL (track_summary materialized view) |
| Search / filtering | TBD | | Search / filtering | PostgreSQL or SQLite (indexed) |
| Archiving | TBD | | Archiving | CSV tree + tar.gz, plus pg_dump of the system of record |
| Inter-lab exchange | TBD | | Inter-lab exchange | Hybrid JSON-LD (metadata + sourceFile links to CSV) |
## Conventions ## Conventions

527
bench_runner.py Normal file
View File

@@ -0,0 +1,527 @@
"""Task 09 - Benchmark execution and resource measurement.
Runs the 35-cell matrix (Q1-Q7 x 5 formats) in randomized order. Every
cell: 1 warm-up run (recorded as run 0, excluded from statistics) + 3
measured runs, each in its OWN subprocess sampled with psutil every 50 ms
for peak RSS / CPU time / read bytes. Result rows of every run are
validated against out/bench/expected/ (tolerance 1e-9) and recorded as
checksum_ok. Writes results_raw.csv, results_median.csv, PG EXPLAIN plans,
PG host/server metrics, the archival storage sizes, an environment record,
and the 09.ok marker. Spec: docs/specs/09_benchmark_execution.md; protocol:
docs/rules/bench-methodology.md.
Environment adaptations (bench-methodology section 6):
- the cold-cache matrix pass is SKIPPED (Windows host, no page-cache drop);
- PostgreSQL runs on a remote host: no backend-PID psutil sampling;
instead best-effort Prometheus host CPU/RAM windows per PG cell and per
session, plus best-effort pg_stat_statements per-query session deltas;
- on Windows the 50 ms sample reads the OS-maintained peak working set
(peak_wset), so sub-interval children still report a true peak RSS.
After a run times out, the cell's remaining runs are SKIPPED (never retry
a timeout - bench-methodology section 3); only executed runs appear in the
raw results and the cell reports TIMEOUT.
"""
from __future__ import annotations
import csv
import gzip
import logging
import random
import re
import shutil
import statistics
import subprocess
import sys
import tarfile
import time
from dataclasses import dataclass, field
from pathlib import Path
import psutil
import psycopg
from common.monitoring import get_prom_url, host_window_metrics
from common.pg import get_dsn
from common.pipeline import (
REPO_ROOT,
check_dependencies,
load_lab_config,
parse_task_args,
remove_stale_marker,
write_marker,
)
from common.results import compare, parse_rows
from common.storage_sizes import update_storage_sizes
logger = logging.getLogger(__name__)
TASK_ID = "09"
DEPENDS_ON = ["08"]
CELL_TIMEOUT_S = 30 * 60 # hard per-run timeout, spec 09
SAMPLE_INTERVAL_S = 0.05 # psutil sampling cadence, spec 09
MEASURED_RUNS = 3 # 1 warm-up + 3 measured per cell, spec 09
FORMATS = ("csv", "json", "sqlite", "pg", "rdf")
QUERY_NUMBERS = tuple(range(1, 8))
RAW_HEADER = [
"format", "query", "run", "wall_s", "peak_rss_mb", "cpu_s",
"cpu_util_pct", "read_mb", "rows", "checksum_ok", "status", "exec_order",
]
MEDIAN_HEADER = [
"format", "query", "wall_s_median", "wall_s_min", "wall_s_max",
"peak_rss_mb_median", "cpu_s_median", "cpu_util_pct_median",
"read_mb_median", "rows", "checksum_ok", "status",
]
# storage_sizes.csv completeness gate: every (format, variant) that must
# exist after consolidation (working + archival footprints, spec 09)
REQUIRED_SIZE_ROWS = {
("csv", "tree"), ("csv", "targz"),
("json", "full"), ("json", "hybrid"), ("json", "targz"),
("sqlite", "db"), ("sqlite", "db_gz"),
("pg", "live"), ("pg", "dump"),
("rdf", "nt_gz"), ("rdf", "ttl"), ("rdf", "store"),
}
class CellTimeoutError(Exception):
pass
@dataclass
class RunResult:
status: str
wall_s: float = 0.0
peak_rss_mb: float = 0.0
cpu_s: float = 0.0
cpu_util_pct: float = 0.0
read_mb: float = 0.0
rows: int | None = None
checksum_ok: bool | None = None
def raw_row(self, fmt: str, qn: int, run: int, order: int) -> list:
ok = "" if self.checksum_ok is None else str(self.checksum_ok).lower()
return [
fmt, f"q{qn}", run, round(self.wall_s, 4), round(self.peak_rss_mb, 1),
round(self.cpu_s, 3), round(self.cpu_util_pct, 1), round(self.read_mb, 2),
"" if self.rows is None else self.rows, ok, self.status, order,
]
def _kill_tree(proc: subprocess.Popen) -> None:
"""Kill the child and every descendant (the stub's real interpreter)."""
try:
for child in psutil.Process(proc.pid).children(recursive=True):
try:
child.kill()
except psutil.NoSuchProcess:
logger.debug("kill_tree: pid %d already exited", child.pid)
except psutil.Error:
logger.warning("kill_tree: descendants of pid %d not enumerable", proc.pid, exc_info=True)
proc.kill()
proc.wait()
def _run_child(cmd: list[str], stdout_path: Path, stderr_path: Path) -> tuple[int, RunResult]:
"""Spawn one query subprocess and sample its process TREE until exit.
The venv python.exe on Windows is a launcher stub: the real interpreter
runs as its child, so sampling only the spawned PID measures the stub
(~4 MB flat, ~0 CPU). Every sample therefore walks the descendant tree
and keeps per-PID last-known CPU / read bytes and the OS peak working
set; totals are sums over all PIDs ever seen. The flat stub adds ~4 MB
to the peak-RSS sum - negligible against real query footprints.
"""
cores = psutil.cpu_count(logical=True)
seen: dict[int, list] = {} # pid -> [cpu_s, read_bytes, peak_rss_bytes]
timed_out = False
start = time.perf_counter()
with open(stdout_path, "wb") as out_fh, open(stderr_path, "wb") as err_fh:
proc = subprocess.Popen(cmd, stdout=out_fh, stderr=err_fh, cwd=REPO_ROOT)
try:
try:
root = psutil.Process(proc.pid)
except psutil.Error:
root = None # child already gone: exited faster than one sample
while root is not None and proc.poll() is None:
try:
procs = [root, *root.children(recursive=True)]
except psutil.Error:
break
for p in procs:
try:
mem = p.memory_info()
cpu = p.cpu_times()
io = p.io_counters()
except psutil.Error:
continue # that process exited between listing and sampling
entry = seen.setdefault(p.pid, [0.0, 0, 0])
entry[0] = cpu.user + cpu.system
entry[1] = io.read_bytes
entry[2] = max(entry[2], getattr(mem, "peak_wset", 0) or mem.rss)
if time.perf_counter() - start > CELL_TIMEOUT_S:
timed_out = True
break
time.sleep(SAMPLE_INTERVAL_S)
if timed_out:
_kill_tree(proc)
raise CellTimeoutError(f"exceeded {CELL_TIMEOUT_S} s")
returncode = proc.wait()
finally:
if proc.poll() is None:
_kill_tree(proc)
wall = time.perf_counter() - start
cpu_s = sum(entry[0] for entry in seen.values())
result = RunResult(
status="", wall_s=wall,
peak_rss_mb=sum(entry[2] for entry in seen.values()) / 1048576,
cpu_s=cpu_s,
cpu_util_pct=100.0 * cpu_s / wall / cores if wall > 0 else 0.0,
read_mb=sum(entry[1] for entry in seen.values()) / 1048576,
)
return returncode, result
def _run_cell_once(fmt: str, qn: int, cmd: list[str], tmp_dir: Path, expected: list[tuple]) -> RunResult:
"""One benchmark run with pattern C isolation (code-error-handling)."""
stdout_path = tmp_dir / "stdout.txt"
stderr_path = tmp_dir / "stderr.txt"
try:
returncode, result = _run_child(cmd, stdout_path, stderr_path)
except CellTimeoutError:
logger.error("cell %s/q%d timed out after %d s", fmt, qn, CELL_TIMEOUT_S)
return RunResult(status="TIMEOUT", wall_s=CELL_TIMEOUT_S)
except Exception:
logger.error("cell %s/q%d failed to execute", fmt, qn, exc_info=True)
return RunResult(status="ERROR")
if returncode != 0:
tail = stderr_path.read_text(encoding="utf-8", errors="replace")[-2000:]
logger.error("cell %s/q%d exited %d; stderr tail: %s", fmt, qn, returncode, tail)
result.status = "ERROR"
return result
actual = parse_rows(stdout_path.read_text(encoding="ascii"))
diff = compare(expected, actual)
if diff:
logger.error("cell %s/q%d checksum mismatch: %s", fmt, qn, diff)
result.status = "OK"
result.rows = len(actual)
result.checksum_ok = diff is None
return result
@dataclass
class CellOutcome:
fmt: str
qn: int
measured: list[RunResult] = field(default_factory=list)
timed_out: bool = False # any run incl. warm-up: remaining runs skipped
window: tuple[float, float] | None = None
def status(self) -> str:
if self.timed_out or any(r.status == "TIMEOUT" for r in self.measured):
return "TIMEOUT"
if len(self.measured) < MEASURED_RUNS or any(r.status == "ERROR" for r in self.measured):
return "ERROR"
if not all(r.checksum_ok for r in self.measured):
return "INVALID"
return "OK"
def _median_row(outcome: CellOutcome) -> list:
status = outcome.status()
fmt, qn, measured = outcome.fmt, outcome.qn, outcome.measured
if status != "OK":
ok = "false" if status == "INVALID" else ""
return [fmt, f"q{qn}", "", "", "", "", "", "", "", "", ok, status]
walls = [r.wall_s for r in measured]
return [
fmt, f"q{qn}",
round(statistics.median(walls), 4), round(min(walls), 4), round(max(walls), 4),
round(statistics.median(r.peak_rss_mb for r in measured), 1),
round(statistics.median(r.cpu_s for r in measured), 3),
round(statistics.median(r.cpu_util_pct for r in measured), 1),
round(statistics.median(r.read_mb for r in measured), 2),
measured[0].rows, "true", status,
]
def _capture_explain(dsn: str, queries_dir: Path, explain_dir: Path) -> None:
"""EXPLAIN (ANALYZE, BUFFERS) per PG query; verify Q5 avoids track_summary."""
explain_dir.mkdir(parents=True, exist_ok=True)
with psycopg.connect(dsn) as conn:
for qn in QUERY_NUMBERS:
sql = (queries_dir / "pg" / f"q{qn}.sql").read_text(encoding="ascii")
plan = "\n".join(r[0] for r in conn.execute("EXPLAIN (ANALYZE, BUFFERS) " + sql).fetchall())
(explain_dir / f"q{qn}.txt").write_text(plan + "\n", encoding="utf-8", newline="\n")
if qn == 5 and "track_summary" in plan:
raise RuntimeError("q5 plan touches track_summary - the forced full scan is violated (bench-methodology section 4)")
logger.info("EXPLAIN plans captured to %s", explain_dir)
PG_STAT_QUERY = """
SELECT query, calls, total_exec_time, rows,
shared_blks_hit, shared_blks_read, temp_blks_read, temp_blks_written
FROM pg_stat_statements
"""
PG_STAT_COLUMNS = [
"calls", "total_exec_time_ms", "rows",
"shared_blks_hit", "shared_blks_read", "temp_blks_read", "temp_blks_written",
]
_QTAG = re.compile(r"^-- Q(\d):")
def _pg_statements_snapshot(dsn: str) -> dict[int, list[float]] | None:
"""Sum pg_stat_statements counters per benchmark query (matched by the
leading '-- Q<N>:' comment). Best-effort: None when the extension is
unavailable - server-side stats are auxiliary and never fail the task.
"""
try:
sums: dict[int, list[float]] = {}
with psycopg.connect(dsn) as conn:
for query_text, *values in conn.execute(PG_STAT_QUERY):
match = _QTAG.match(query_text)
if not match:
continue
acc = sums.setdefault(int(match.group(1)), [0.0] * len(values))
for i, v in enumerate(values):
acc[i] += float(v)
return sums
except Exception:
logger.warning("pg_stat_statements snapshot skipped (extension unavailable or not readable)", exc_info=True)
return None
def _write_pg_server_stats(path: Path, start: dict[int, list[float]], end: dict[int, list[float]]) -> None:
with open(path, "w", encoding="ascii", newline="") as fh:
writer = csv.writer(fh)
writer.writerow(["query"] + PG_STAT_COLUMNS)
for qn in QUERY_NUMBERS:
before = start.get(qn, [0.0] * len(PG_STAT_COLUMNS))
after = end.get(qn)
if after is None:
logger.warning("pg_stat_statements: q%d missing in the end snapshot (entry evicted?)", qn)
writer.writerow([f"q{qn}"] + [""] * len(PG_STAT_COLUMNS))
continue
delta = [after[i] - before[i] for i in range(len(PG_STAT_COLUMNS))]
writer.writerow([f"q{qn}"] + [round(v, 3) for v in delta])
logger.info("pg server-side statement deltas written to %s", path)
def _tree_bytes(root: Path) -> int:
return sum(p.stat().st_size for p in root.rglob("*") if p.is_file())
def _consolidate_storage_sizes(out_root: Path, bench_dir: Path) -> None:
"""Build archival variants, record all sizes, verify completeness."""
archives = bench_dir / "archives"
archives.mkdir(parents=True, exist_ok=True)
csv_tgz = archives / "csv_tree.tar.gz"
logger.info("archiving out/csv -> %s", csv_tgz)
with tarfile.open(csv_tgz, "w:gz") as tar:
tar.add(out_root / "csv", arcname="csv")
update_storage_sizes(bench_dir, "csv", [
("tree", _tree_bytes(out_root / "csv")), ("targz", csv_tgz.stat().st_size),
])
json_tgz = archives / "json.tar.gz"
logger.info("archiving out/json -> %s", json_tgz)
with tarfile.open(json_tgz, "w:gz") as tar:
tar.add(out_root / "json", arcname="json")
update_storage_sizes(bench_dir, "json", [("targz", json_tgz.stat().st_size)])
db_gz = archives / "tribo.db.gz"
logger.info("compressing out/sqlite/tribo.db -> %s", db_gz)
with open(out_root / "sqlite" / "tribo.db", "rb") as src, gzip.open(db_gz, "wb") as dst:
shutil.copyfileobj(src, dst, 1048576)
update_storage_sizes(bench_dir, "sqlite", [("db_gz", db_gz.stat().st_size)])
# pg_dump -Fc is internally gzip-compressed: it IS the pg archival variant
update_storage_sizes(bench_dir, "pg", [("dump", (out_root / "pg" / "tribo.dump").stat().st_size)])
present = set()
with open(bench_dir / "storage_sizes.csv", encoding="ascii") as fh:
for row in csv.reader(fh):
present.add((row[0], row[1]))
missing = REQUIRED_SIZE_ROWS - present
if missing:
raise ValueError(f"storage_sizes.csv incomplete after consolidation: missing {sorted(missing)}")
logger.info("storage_sizes.csv consolidated: all %d required rows present", len(REQUIRED_SIZE_ROWS))
def _environment_record(path: Path, dsn: str) -> None:
import platform
lines = [
f"benchmark_host_os: {platform.platform()}",
f"benchmark_host_cpu: {platform.processor()}",
f"benchmark_host_cores_logical: {psutil.cpu_count(logical=True)}",
f"benchmark_host_cores_physical: {psutil.cpu_count(logical=False)}",
f"benchmark_host_ram_gib: {round(psutil.virtual_memory().total / 1073741824, 1)}",
f"python: {platform.python_version()}",
]
try:
disks = subprocess.run(
["powershell", "-NoProfile", "-Command",
"Get-PhysicalDisk | ForEach-Object { $_.FriendlyName + ' / ' + $_.MediaType + ' / ' + [math]::Round($_.Size / 1GB) + ' GB' }"],
capture_output=True, text=True, timeout=60, check=True)
lines.append("benchmark_host_disks: " + "; ".join(ln.strip() for ln in disks.stdout.splitlines() if ln.strip()))
except Exception:
logger.warning("disk info skipped (Get-PhysicalDisk query failed - non-Windows host or no permission)", exc_info=True)
import sqlite3
import ijson
import pyoxigraph
lines += [
f"sqlite: {sqlite3.sqlite_version}",
f"psycopg: {psycopg.__version__}",
f"pyoxigraph: {pyoxigraph.__version__}",
f"ijson: {ijson.__version__}",
f"psutil: {psutil.__version__}",
]
with psycopg.connect(dsn) as conn:
pg_version = conn.execute("SELECT version()").fetchone()[0]
lines.append(f"postgresql_server: {pg_version}")
lines.append(f"postgresql_host: {conn.info.host} (remote; 4 cores / 3.8 GiB, see .done/06.ok)")
lines += [
"cold_cache_pass: skipped (Windows host, no page-cache drop - bench-methodology section 6)",
"pg_backend_sampling: not possible for a remote server; Prometheus windows + pg_stat_statements deltas instead",
]
path.write_text("\n".join(lines) + "\n", encoding="utf-8", newline="\n")
logger.info("environment record written to %s", path)
def _run_matrix(bench_dir: Path, expected: dict[int, list[tuple]], seed: int) -> list[CellOutcome]:
cells = [(fmt, qn) for fmt in FORMATS for qn in QUERY_NUMBERS]
random.Random(f"{seed}-bench-order").shuffle(cells)
tmp_dir = bench_dir / "tmp"
tmp_dir.mkdir(parents=True, exist_ok=True)
outcomes: list[CellOutcome] = []
with open(bench_dir / "results_raw.csv", "w", encoding="ascii", newline="") as raw_fh:
writer = csv.writer(raw_fh)
writer.writerow(RAW_HEADER)
for order, (fmt, qn) in enumerate(cells, start=1):
cmd = [sys.executable, str(bench_dir / "queries" / fmt / f"q{qn}.py")]
outcome = CellOutcome(fmt, qn)
window_start = time.time()
for run in range(MEASURED_RUNS + 1): # run 0 = warm-up
result = _run_cell_once(fmt, qn, cmd, tmp_dir, expected[qn])
writer.writerow(result.raw_row(fmt, qn, run, order))
raw_fh.flush()
if run > 0:
outcome.measured.append(result)
logger.info("cell %s/q%d run %d: %s wall=%.3fs rss=%.1fMB", fmt, qn, run, result.status, result.wall_s, result.peak_rss_mb)
if result.status == "TIMEOUT":
outcome.timed_out = True
logger.error("cell %s/q%d: skipping remaining runs after timeout", fmt, qn)
break
outcome.window = (window_start, time.time())
outcomes.append(outcome)
status = outcome.status()
walls = [r.wall_s for r in outcome.measured if r.status == "OK"]
detail = f"median {statistics.median(walls):.3f} s" if status == "OK" else status
print(f"[{order}/{len(cells)}] {fmt}/q{qn}: {detail}")
shutil.rmtree(tmp_dir)
return outcomes
def _write_pg_host_metrics(path: Path, prom_url: str, outcomes: list[CellOutcome]) -> None:
with open(path, "w", encoding="ascii", newline="") as fh:
writer = csv.writer(fh)
writer.writerow([
"query", "window_s", "host_cpu_avg_pct", "host_cpu_max_pct",
"host_ram_used_baseline_mb", "host_ram_used_peak_mb", "host_ram_used_delta_mb",
])
for outcome in sorted((o for o in outcomes if o.fmt == "pg"), key=lambda o: o.qn):
start, end = outcome.window
metrics = host_window_metrics(prom_url, start, end)
row = [f"q{outcome.qn}", round(end - start, 1)]
if metrics is None:
row += [""] * 5
else:
row += [
metrics["host_cpu_avg_pct"], metrics["host_cpu_max_pct"],
metrics["host_ram_used_baseline_mb"], metrics["host_ram_used_peak_mb"],
metrics["host_ram_used_delta_mb"],
]
writer.writerow(row)
logger.info("per-cell PG host metrics written to %s", path)
def main() -> int:
args = parse_task_args("Task 09: execute the Q1-Q7 x 5 formats benchmark matrix")
out_root: Path = args.out_root
bench_dir = out_root / "bench"
try:
check_dependencies(out_root, DEPENDS_ON)
remove_stale_marker(out_root, TASK_ID)
dump_path = out_root / "pg" / "tribo.dump"
if not dump_path.exists():
raise FileNotFoundError(f"{dump_path} missing - produce it with pg_dump -Fc before task 09 (storage consolidation needs it)")
dsn = get_dsn()
expected = {qn: parse_rows((bench_dir / "expected" / f"q{qn}.rows.csv").read_text(encoding="ascii")) for qn in QUERY_NUMBERS}
seed = load_lab_config(out_root)["seed"]
logger.warning("cold-cache pass skipped (Windows host - bench-methodology section 6)")
_environment_record(bench_dir / "environment.txt", dsn)
prom_url = get_prom_url()
stats_start = _pg_statements_snapshot(dsn)
session_start = time.time()
outcomes = _run_matrix(bench_dir, expected, seed)
session_end = time.time()
with open(bench_dir / "results_median.csv", "w", encoding="ascii", newline="") as fh:
writer = csv.writer(fh)
writer.writerow(MEDIAN_HEADER)
for outcome in sorted(outcomes, key=lambda o: (o.fmt, o.qn)):
writer.writerow(_median_row(outcome))
# snapshot BEFORE the EXPLAIN pass so its executions cannot pollute the deltas
if stats_start is not None:
stats_end = _pg_statements_snapshot(dsn)
if stats_end is not None:
_write_pg_server_stats(bench_dir / "pg_server_stats.csv", stats_start, stats_end)
_capture_explain(dsn, bench_dir / "queries", bench_dir / "explain")
session_host = None
if prom_url:
_write_pg_host_metrics(bench_dir / "pg_host_metrics.csv", prom_url, outcomes)
session_host = host_window_metrics(prom_url, session_start, session_end)
else:
logger.warning("TRIBO_PROM_URL unset - PG host metrics skipped")
_consolidate_storage_sizes(out_root, bench_dir)
by_status: dict[str, int] = {}
for outcome in outcomes:
by_status[outcome.status()] = by_status.get(outcome.status(), 0) + 1
entries = {
"cells_total": len(outcomes),
"cells_ok": by_status.get("OK", 0),
"cells_invalid": by_status.get("INVALID", 0),
"cells_timeout": by_status.get("TIMEOUT", 0),
"cells_error": by_status.get("ERROR", 0),
"session_wall_s": round(session_end - session_start, 1),
"cold_cache_pass": "skipped (Windows host)",
}
if session_host:
entries.update({f"session_{k}": v for k, v in session_host.items()})
marker_path = write_marker(out_root, TASK_ID, entries)
except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1
print(f"task 09 ok: {entries['cells_ok']}/{entries['cells_total']} cells OK "
f"(invalid={entries['cells_invalid']}, timeout={entries['cells_timeout']}, error={entries['cells_error']}), "
f"session {entries['session_wall_s']} s")
print(f"results: {bench_dir / 'results_raw.csv'}, {bench_dir / 'results_median.csv'}")
print(f"marker: {marker_path}")
return 0
if __name__ == "__main__":
sys.exit(main())

73
common/corpus.py Normal file
View File

@@ -0,0 +1,73 @@
"""Reading the CSV corpus with MANIFEST validation (converters, tasks 04-07).
Every file consumed from ./out/csv/ is verified against MANIFEST.csv (byte
size, sha256, row count) before its content is trusted - a mismatch aborts
the converter (docs/rules/test-pipeline-validation.md section 2).
"""
from __future__ import annotations
import hashlib
import logging
import re
from pathlib import Path
logger = logging.getLogger(__name__)
MANIFEST_HEADER = "path,rows,bytes,sha256"
INT_RE = re.compile(r"^-?[0-9]+$")
FLOAT_RE = re.compile(r"^-?([0-9]+\.[0-9]*|\.[0-9]+|[0-9]+)([eE][+-]?[0-9]+)?$")
class ValidationError(RuntimeError):
pass
def coerce(s: str) -> int | float | str:
if INT_RE.match(s):
return int(s)
if FLOAT_RE.match(s):
return float(s)
return s
class CorpusReader:
def __init__(self, csv_root: Path) -> None:
self.csv_root = csv_root
self.manifest = self._load_manifest()
def _load_manifest(self) -> dict[str, tuple[int, int, str]]:
manifest: dict[str, tuple[int, int, str]] = {}
with open(self.csv_root / "MANIFEST.csv", encoding="ascii") as fh:
header = fh.readline().strip()
if header != MANIFEST_HEADER:
raise ValidationError(f"MANIFEST.csv: unexpected header {header!r}")
for line in fh:
path, rows, nbytes, sha = line.strip().split(",")
manifest[path] = (int(rows), int(nbytes), sha)
logger.info("manifest loaded: %d entries", len(manifest))
return manifest
def read(self, relpath: str) -> tuple[list[str], list[list[str]]]:
"""Return (header, raw string rows), verified against MANIFEST."""
data = (self.csv_root / relpath).read_bytes()
exp_rows, exp_bytes, exp_sha = self.manifest[relpath]
if len(data) != exp_bytes or hashlib.sha256(data).hexdigest() != exp_sha:
raise ValidationError(f"{relpath}: bytes/sha256 mismatch vs MANIFEST")
lines = data.decode("ascii").split("\n")
header = lines[0].split(",")
rows = [ln.split(",") for ln in lines[1:] if ln]
if len(rows) != exp_rows:
raise ValidationError(f"{relpath}: {len(rows)} rows != manifest {exp_rows}")
return header, rows
def dicts(self, relpath: str) -> list[dict]:
"""Rows as dicts: lowercased CSV headers as keys, numerics coerced,
empty cells omitted (docs/rules/data-naming-units.md section 4)."""
header, rows = self.read(relpath)
keys = [h.lower() for h in header]
return [
{k: coerce(v) for k, v in zip(keys, row) if v != ""}
for row in rows
]

66
common/monitoring.py Normal file
View File

@@ -0,0 +1,66 @@
"""Remote host metrics via Prometheus / node_exporter (stdlib urllib only).
The PostgreSQL instance is remote, so psutil cannot attribute its CPU/RAM
(spec 09 assumed a local instance). The database host runs Prometheus;
its base URL comes from the TRIBO_PROM_URL environment variable
(e.g. http://192.168.10.73:9090). Unset = monitoring disabled.
"""
from __future__ import annotations
import json
import logging
import os
import urllib.parse
import urllib.request
logger = logging.getLogger(__name__)
QUERY_STEP_S = 5
HTTP_TIMEOUT_S = 15
# rate() window must span at least two scrapes (15 s default interval)
CPU_QUERY = '100 * (1 - avg(rate(node_cpu_seconds_total{mode="idle"}[30s])))'
RAM_QUERY = "node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes"
def get_prom_url() -> str:
return os.environ.get("TRIBO_PROM_URL", "")
def query_range(base_url: str, promql: str, start: float, end: float) -> list[float]:
params = urllib.parse.urlencode({"query": promql, "start": start, "end": end, "step": QUERY_STEP_S})
url = base_url.rstrip("/") + "/api/v1/query_range?" + params
with urllib.request.urlopen(url, timeout=HTTP_TIMEOUT_S) as resp:
data = json.load(resp)
if data.get("status") != "success":
raise RuntimeError(f"prometheus query failed: {data.get('errorType')}: {data.get('error')}")
result = data["data"]["result"]
if not result:
return []
return [float(v) for _, v in result[0]["values"]]
def host_window_metrics(base_url: str, start: float, end: float) -> dict | None:
"""CPU/RAM of the monitored host over [start, end] wall-clock seconds.
Returns None (with a warning) on any failure - monitoring is best-effort
and must never fail the owning task.
"""
try:
if end - start < 2 * QUERY_STEP_S:
end = start + 2 * QUERY_STEP_S
cpu = query_range(base_url, CPU_QUERY, start, end)
ram = query_range(base_url, RAM_QUERY, start, end)
if not cpu or not ram:
logger.warning("prometheus returned no samples for the window (expected condition: short window or scrape lag)")
return None
return {
"host_cpu_avg_pct": round(sum(cpu) / len(cpu), 1),
"host_cpu_max_pct": round(max(cpu), 1),
"host_ram_used_baseline_mb": round(ram[0] / 1048576),
"host_ram_used_peak_mb": round(max(ram) / 1048576),
"host_ram_used_delta_mb": round((max(ram) - ram[0]) / 1048576),
}
except Exception:
logger.warning("host metrics collection skipped (prometheus unreachable or query failed)", exc_info=True)
return None

16
common/pg.py Normal file
View File

@@ -0,0 +1,16 @@
"""PostgreSQL DSN access (docs/rules/code-config-yaml.md section 3).
The DSN comes from the TRIBO_PG_DSN environment variable and is read ONLY
here. It may embed a password: never log it, never hardcode it, never
commit it - log host/database names at most.
"""
from __future__ import annotations
import os
DEFAULT_DSN = "postgresql://postgres@localhost:5432/tribo"
def get_dsn() -> str:
return os.environ.get("TRIBO_PG_DSN", DEFAULT_DSN)

79
common/pipeline.py Normal file
View File

@@ -0,0 +1,79 @@
"""Task-orchestration helpers: CLI shape, dependency markers, completion
markers, lab config loading (docs/rules/build-pipeline-tasks.md,
docs/rules/code-python-style.md section 3).
"""
from __future__ import annotations
import argparse
import logging
import sys
from pathlib import Path
import psutil
import yaml
logger = logging.getLogger(__name__)
REPO_ROOT = Path(__file__).resolve().parent.parent
def parse_task_args(description: str) -> argparse.Namespace:
parser = argparse.ArgumentParser(description=description)
parser.add_argument("--log-level", default="INFO", choices=["DEBUG", "INFO", "WARNING", "ERROR"])
parser.add_argument("--out-root", type=Path, default=REPO_ROOT / "out", help="artifact tree root (default: ./out)")
args = parser.parse_args()
logging.basicConfig(
level=args.log_level,
stream=sys.stderr,
format="%(asctime)s %(levelname)-7s %(name)s: %(message)s",
)
return args
def check_dependencies(out_root: Path, deps: list[str]) -> None:
for dep in deps:
marker = out_root / ".done" / f"{dep}.ok"
if not marker.exists():
raise FileNotFoundError(f"dependency marker missing: {marker} - run task {dep} first")
def remove_stale_marker(out_root: Path, task_id: str) -> Path:
marker = out_root / ".done" / f"{task_id}.ok"
if marker.exists():
logger.info("re-run: removing stale marker %s", marker)
marker.unlink()
return marker
def write_marker(out_root: Path, task_id: str, entries: dict) -> Path:
marker = out_root / ".done" / f"{task_id}.ok"
lines = [f"task: '{task_id}'", "status: ok"] + [f"{k}: {v}" for k, v in entries.items()]
marker.parent.mkdir(parents=True, exist_ok=True)
with open(marker, "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(lines) + "\n")
return marker
def load_lab_config(out_root: Path) -> dict:
with open(out_root / "config" / "lab_config.yaml", encoding="utf-8") as fh:
return yaml.safe_load(fh)
def process_metrics() -> dict:
"""Peak RSS and CPU time of the current process.
Every converter records its own build cost in its marker so the task 10
hardware-sizing model can extrapolate per-format RAM/CPU needs, not just
disk. peak_wset exists on Windows; the rss fallback covers Linux, where
the value is the current (not peak) RSS - good enough for streaming
converters whose RSS is flat by design.
"""
proc = psutil.Process()
mem = proc.memory_info()
peak = getattr(mem, "peak_wset", 0) or mem.rss
cpu = proc.cpu_times()
return {
"proc_peak_rss_mb": round(peak / 1048576, 1),
"proc_cpu_seconds": round(cpu.user + cpu.system, 1),
}

View File

@@ -0,0 +1,15 @@
"""Benchmark query implementations for the file-based formats (spec 08).
One module per format (csv_queries, json_queries, rdf_queries), each with
functions q1..q7 returning canonical result rows (see common/results.py).
The runnable wrappers under out/bench/queries/ are generated by
make_queries.py and delegate here. SQLite / PostgreSQL implementations are
plain SQL files written by make_queries.py.
The relational formats answer Q2/Q3/Q4/Q7 from the precomputed
track_summary (their idiomatic strength); the file formats derive the same
values from raw cycles via common/track_summary.py - identical algorithm,
identical results (bench-methodology section 5).
"""
Q1_TRACK_CODE = "B722-W2-C13-T2" # fixed benchmark track, spec 08

View File

@@ -0,0 +1,139 @@
"""Q1-Q7 over the raw CSV tree: pure python + csv module, streaming.
Q1 is a direct path read (the honest CSV strength, spec 08); Q2-Q7 are
deterministic directory walks with programmatic joins. No manifest
validation here - queries measure retrieval, not integrity checking.
"""
from __future__ import annotations
import csv
from pathlib import Path
from common.queries import Q1_TRACK_CODE
from common.track_summary import summarize_track
def _rows(path: Path) -> tuple[list[str], list[list[str]]]:
with open(path, newline="", encoding="ascii") as fh:
reader = csv.reader(fh)
header = [h.lower() for h in next(reader)]
return header, list(reader)
def _dict_row(path: Path) -> dict[str, str]:
header, rows = _rows(path)
return dict(zip(header, rows[0]))
def _track_path(csv_root: Path, track_code: str) -> Path:
batch, wafer, coupon, track = track_code.split("-")
return csv_root / f"batch_{batch}" / f"wafer_{wafer}" / f"coupon_{coupon}" / f"track_{track}"
def _coupon_dirs(csv_root: Path):
for batch_dir in sorted(csv_root.glob("batch_*")):
batch_code = batch_dir.name[6:]
for wafer_dir in sorted(batch_dir.glob("wafer_*")):
for coupon_dir in sorted(wafer_dir.glob("coupon_*")):
yield batch_code, coupon_dir
def _track_dirs(coupon_dir: Path) -> list[Path]:
return sorted(coupon_dir.glob("track_T*"))
def _cofs(track_dir: Path) -> list[float]:
_, rows = _rows(track_dir / "cof_vs_cycle.csv")
return [float(r[1]) for r in rows]
def _mean(values: list[float]) -> float:
total = 0.0
for v in values:
total += v
return total / len(values)
def q1(csv_root: Path) -> list[tuple]:
_, rows = _rows(_track_path(csv_root, Q1_TRACK_CODE) / "cof_vs_cycle.csv")
return [(int(r[0]), float(r[1])) for r in rows]
def q2(csv_root: Path) -> list[tuple]:
out = []
for _, coupon_dir in _coupon_dirs(csv_root):
info = _dict_row(coupon_dir / "coupon_info.csv")
if info["run_code"] == "RESERVE" or not 9.5 <= float(info["au_wtpct_mean"]) <= 10.5:
continue
tracks = _track_dirs(coupon_dir)
env = _dict_row(tracks[0] / "track_info.csv")["environment"]
cof_ss = _mean([summarize_track(_cofs(t))[0] for t in tracks])
out.append((info["coupon_code"], env, cof_ss))
return out
def q3(csv_root: Path) -> list[tuple]:
out = []
for batch_code, coupon_dir in _coupon_dirs(csv_root):
info = _dict_row(coupon_dir / "coupon_info.csv")
if info["run_code"] == "RESERVE":
continue
header, rows = _rows(coupon_dir / "nanoindentation.csv")
h_col = header.index("hardness_gpa")
hardness = _mean([float(r[h_col]) for r in rows])
cof_ss = _mean([summarize_track(_cofs(t))[0] for t in _track_dirs(coupon_dir)])
out.append((info["coupon_code"], batch_code, hardness, cof_ss))
return out
def q4(csv_root: Path) -> list[tuple]:
out = []
for _, coupon_dir in _coupon_dirs(csv_root):
for track_dir in _track_dirs(coupon_dir):
info = _dict_row(track_dir / "track_info.csv")
if info["environment"] != "dry_n2" or float(info["load_mn"]) != 100.0:
continue
cof_ss = summarize_track(_cofs(track_dir))[0]
if cof_ss > 0.20:
out.append((info["track_code"], cof_ss))
return out
def q5(csv_root: Path) -> list[tuple]:
sums: dict[str, list[float]] = {}
for batch_code, coupon_dir in _coupon_dirs(csv_root):
for track_dir in _track_dirs(coupon_dir):
run_in = summarize_track(_cofs(track_dir))[2]
acc = sums.setdefault(batch_code, [0.0, 0.0])
acc[0] += run_in
acc[1] += 1
return [(batch, acc[0] / acc[1]) for batch, acc in sorted(sums.items())]
def q6(csv_root: Path) -> list[tuple]:
out = []
for _, coupon_dir in _coupon_dirs(csv_root):
info = _dict_row(coupon_dir / "coupon_info.csv")
if info["run_code"] == "RESERVE":
continue
tracks = _track_dirs(coupon_dir)
load = float(_dict_row(tracks[0] / "track_info.csv")["load_mn"])
wear = _mean([float(_dict_row(t / "wear.csv")["wear_volume_um3"]) for t in tracks])
out.append((info["coupon_code"], load, wear))
return out
def q7(csv_root: Path) -> list[tuple]:
groups: dict[tuple[str, int], list[float]] = {}
for _, coupon_dir in _coupon_dirs(csv_root):
for track_dir in _track_dirs(coupon_dir):
info = _dict_row(track_dir / "track_info.csv")
bucket = int(round(float(info["speed_mm_s"]) * float(info["load_mn"])))
cofs = _cofs(track_dir)
run_in = summarize_track(cofs)[2]
acc = groups.setdefault((info["environment"], bucket), [0.0, 0.0])
for cof in cofs[run_in:]: # cycles strictly past run-in
acc[0] += cof
acc[1] += 1
return [(env, bucket, acc[0] / acc[1]) for (env, bucket), acc in sorted(groups.items())]

View File

@@ -0,0 +1,169 @@
"""Q1-Q7 over the JSON-LD FULL variant: ijson streaming parser (spec 08).
Each coupon file is streamed event-by-event; scalar coupon fields appear
before the bulk arrays in the serialization, so filtering queries (Q2) can
abandon a non-matching file before parsing its megabytes of points.
"""
from __future__ import annotations
from pathlib import Path
import ijson
from common.queries import Q1_TRACK_CODE
from common.track_summary import summarize_track
class CouponDoc:
__slots__ = ("coupon_code", "batch_code", "run_code", "au", "hardness", "tracks")
def __init__(self) -> None:
self.coupon_code = ""
self.batch_code = ""
self.run_code = ""
self.au = 0.0
self.hardness: list[float] = []
self.tracks: list[dict] = []
def _scan(path: Path, au_range: tuple[float, float] | None = None) -> CouponDoc | None:
"""Stream one coupon file; with au_range set, bail out early on mismatch."""
doc = CouponDoc()
track: dict | None = None
with open(path, "rb") as fh:
for prefix, event, value in ijson.parse(fh, use_float=True):
if prefix == "coupon_code":
doc.coupon_code = value
elif prefix == "batch_code":
doc.batch_code = value
elif prefix == "run_code":
doc.run_code = value
if doc.run_code == "RESERVE":
return None
elif prefix == "au_wtpct_mean":
doc.au = value
if au_range and not au_range[0] <= value <= au_range[1]:
return None
elif prefix == "nanoindentation.indents.item.hardness_gpa":
doc.hardness.append(value)
elif prefix == "tracks.item" and event == "start_map":
track = {"cycles": [], "cofs": []}
elif track is not None:
if prefix == "tracks.item.track_code":
track["track_code"] = value
elif prefix == "tracks.item.environment":
track["environment"] = value
elif prefix == "tracks.item.load_mn":
track["load_mn"] = value
elif prefix == "tracks.item.speed_mm_s":
track["speed_mm_s"] = value
elif prefix == "tracks.item.cycles.item.cycle":
track["cycles"].append(value)
elif prefix == "tracks.item.cycles.item.cof":
track["cofs"].append(value)
elif prefix == "tracks.item.wear.wear_volume_um3":
track["wear"] = value
elif prefix == "tracks.item" and event == "end_map":
doc.tracks.append(track)
track = None
return doc
def _files(json_root: Path) -> list[Path]:
return sorted((json_root / "full").glob("*.jsonld"))
def _mean(values: list[float]) -> float:
total = 0.0
for v in values:
total += v
return total / len(values)
def _cof_ss(doc: CouponDoc) -> float:
return _mean([summarize_track(t["cofs"])[0] for t in doc.tracks])
def q1(json_root: Path) -> list[tuple]:
coupon_code = Q1_TRACK_CODE.rsplit("-", 1)[0]
doc = _scan(json_root / "full" / f"{coupon_code}.jsonld")
for track in doc.tracks:
if track["track_code"] == Q1_TRACK_CODE:
return [(int(c), cof) for c, cof in zip(track["cycles"], track["cofs"])]
raise LookupError(f"track {Q1_TRACK_CODE} not found")
def q2(json_root: Path) -> list[tuple]:
out = []
for path in _files(json_root):
doc = _scan(path, au_range=(9.5, 10.5))
if doc is None:
continue
out.append((doc.coupon_code, doc.tracks[0]["environment"], _cof_ss(doc)))
return out
def q3(json_root: Path) -> list[tuple]:
out = []
for path in _files(json_root):
doc = _scan(path)
if doc is None:
continue
out.append((doc.coupon_code, doc.batch_code, _mean(doc.hardness), _cof_ss(doc)))
return out
def q4(json_root: Path) -> list[tuple]:
out = []
for path in _files(json_root):
doc = _scan(path)
if doc is None:
continue
for track in doc.tracks:
if track["environment"] != "dry_n2" or track["load_mn"] != 100.0:
continue
cof_ss = summarize_track(track["cofs"])[0]
if cof_ss > 0.20:
out.append((track["track_code"], cof_ss))
return out
def q5(json_root: Path) -> list[tuple]:
sums: dict[str, list[float]] = {}
for path in _files(json_root):
doc = _scan(path)
if doc is None:
continue
acc = sums.setdefault(doc.batch_code, [0.0, 0.0])
for track in doc.tracks:
acc[0] += summarize_track(track["cofs"])[2]
acc[1] += 1
return [(batch, acc[0] / acc[1]) for batch, acc in sorted(sums.items())]
def q6(json_root: Path) -> list[tuple]:
out = []
for path in _files(json_root):
doc = _scan(path)
if doc is None:
continue
load = doc.tracks[0]["load_mn"]
out.append((doc.coupon_code, load, _mean([t["wear"] for t in doc.tracks])))
return out
def q7(json_root: Path) -> list[tuple]:
groups: dict[tuple[str, int], list[float]] = {}
for path in _files(json_root):
doc = _scan(path)
if doc is None:
continue
for track in doc.tracks:
bucket = int(round(track["speed_mm_s"] * track["load_mn"]))
run_in = summarize_track(track["cofs"])[2]
acc = groups.setdefault((track["environment"], bucket), [0.0, 0.0])
for cof in track["cofs"][run_in:]:
acc[0] += cof
acc[1] += 1
return [(env, bucket, acc[0] / acc[1]) for (env, bucket), acc in sorted(groups.items())]

View File

@@ -0,0 +1,173 @@
"""Q1-Q7 over the oxigraph triplestore: SPARQL retrieval (spec 08).
The store holds raw data only (no precomputed summaries), so queries that
need steady-state COF or run-in retrieve the raw cycles via SPARQL and
derive the values with the shared algorithm (common/track_summary.py) -
identical semantics to every other format. The SPARQL text of each query
is exported to out/bench/queries/rdf/q<N>.sparql by make_queries.py.
"""
from __future__ import annotations
from pathlib import Path
from pyoxigraph import Store
from common.queries import Q1_TRACK_CODE
from common.track_summary import summarize_track
PREFIX = "PREFIX tribo: <https://sandia.gov/ontology/tribology#>\n"
TRACK_IRI = f"<https://sandia.gov/ontology/tribology/id/track/{Q1_TRACK_CODE}>"
SPARQL = {
"q1": PREFIX + f"""SELECT ?cycle ?cof WHERE {{
{TRACK_IRI} tribo:hasCycle ?c .
?c tribo:cycle ?cycle ; tribo:cof ?cof .
}}""",
"q2": PREFIX + """SELECT ?ccode ?env ?track ?cycle ?cof WHERE {
?coupon tribo:au_wtpct_mean ?au ; tribo:coupon_code ?ccode .
FILTER(?au >= 9.5 && ?au <= 10.5)
?track tribo:partOf ?coupon ; tribo:environment ?env ; tribo:hasCycle ?c .
?c tribo:cycle ?cycle ; tribo:cof ?cof .
}""",
"q3_hardness": PREFIX + """SELECT ?ccode ?bcode (AVG(?h) AS ?hardness) WHERE {
?coupon tribo:coupon_code ?ccode ; tribo:cutFrom ?wafer ; tribo:nanoindentation ?m .
?wafer tribo:partOf ?batch .
?batch tribo:batch_code ?bcode .
?m tribo:hasIndent ?i .
?i tribo:hardness_gpa ?h .
} GROUP BY ?ccode ?bcode""",
"q3_cycles": PREFIX + """SELECT ?ccode ?track ?cycle ?cof WHERE {
?track a tribo:FrictionTest ; tribo:partOf ?coupon ; tribo:hasCycle ?c .
?coupon tribo:coupon_code ?ccode .
?c tribo:cycle ?cycle ; tribo:cof ?cof .
}""",
"q4": PREFIX + """SELECT ?tcode ?track ?cycle ?cof WHERE {
?track tribo:environment "dry_n2" ; tribo:load_mn ?load ; tribo:track_code ?tcode ; tribo:hasCycle ?c .
FILTER(?load = 100)
?c tribo:cycle ?cycle ; tribo:cof ?cof .
}""",
"q5": PREFIX + """SELECT ?bcode ?track ?cycle ?cof WHERE {
?track a tribo:FrictionTest ; tribo:partOf ?coupon ; tribo:hasCycle ?c .
?coupon tribo:cutFrom ?wafer .
?wafer tribo:partOf ?batch .
?batch tribo:batch_code ?bcode .
?c tribo:cycle ?cycle ; tribo:cof ?cof .
}""",
"q6": PREFIX + """SELECT ?ccode ?load (AVG(?wv) AS ?wear) WHERE {
?track tribo:partOf ?coupon ; tribo:load_mn ?load ; tribo:wear ?w .
?w tribo:wear_volume_um3 ?wv .
?coupon tribo:coupon_code ?ccode .
} GROUP BY ?ccode ?load""",
"q7": PREFIX + """SELECT ?env ?speed ?load ?track ?cycle ?cof WHERE {
?track tribo:environment ?env ; tribo:speed_mm_s ?speed ; tribo:load_mn ?load ; tribo:hasCycle ?c .
?c tribo:cycle ?cycle ; tribo:cof ?cof .
}""",
}
def _store(store_path: Path) -> Store:
try:
return Store.read_only(str(store_path))
except AttributeError:
return Store(str(store_path))
def _track_cofs(solutions, track_var: str, keep: tuple[str, ...]) -> dict[str, dict]:
"""Group cycle solutions by track: ordered cof list + kept literals."""
tracks: dict[str, dict] = {}
for sol in solutions:
key = str(sol[track_var])
entry = tracks.get(key)
if entry is None:
entry = {"pairs": []}
for name in keep:
entry[name] = sol[name].value
tracks[key] = entry
entry["pairs"].append((int(sol["cycle"].value), float(sol["cof"].value)))
for entry in tracks.values():
entry["pairs"].sort()
entry["cofs"] = [cof for _, cof in entry["pairs"]]
return tracks
def _mean(values: list[float]) -> float:
total = 0.0
for v in values:
total += v
return total / len(values)
def q1(store_path: Path) -> list[tuple]:
store = _store(store_path)
return [(int(s["cycle"].value), float(s["cof"].value)) for s in store.query(SPARQL["q1"])]
def q2(store_path: Path) -> list[tuple]:
store = _store(store_path)
tracks = _track_cofs(store.query(SPARQL["q2"]), "track", ("ccode", "env"))
coupons: dict[tuple[str, str], list[float]] = {}
for entry in sorted(tracks.values(), key=lambda e: e["ccode"]):
coupons.setdefault((entry["ccode"], entry["env"]), []).append(summarize_track(entry["cofs"])[0])
return [(code, env, _mean(values)) for (code, env), values in sorted(coupons.items())]
def q3(store_path: Path) -> list[tuple]:
store = _store(store_path)
hardness = {
s["ccode"].value: (s["bcode"].value, float(s["hardness"].value))
for s in store.query(SPARQL["q3_hardness"])
}
tracks = _track_cofs(store.query(SPARQL["q3_cycles"]), "track", ("ccode",))
coupons: dict[str, list[float]] = {}
for entry in sorted(tracks.values(), key=lambda e: e["ccode"]):
coupons.setdefault(entry["ccode"], []).append(summarize_track(entry["cofs"])[0])
return [
(code, hardness[code][0], hardness[code][1], _mean(values))
for code, values in sorted(coupons.items())
]
def q4(store_path: Path) -> list[tuple]:
store = _store(store_path)
tracks = _track_cofs(store.query(SPARQL["q4"]), "track", ("tcode",))
out = []
for entry in sorted(tracks.values(), key=lambda e: e["tcode"]):
cof_ss = summarize_track(entry["cofs"])[0]
if cof_ss > 0.20:
out.append((entry["tcode"], cof_ss))
return out
def q5(store_path: Path) -> list[tuple]:
store = _store(store_path)
tracks = _track_cofs(store.query(SPARQL["q5"]), "track", ("bcode",))
sums: dict[str, list[float]] = {}
for entry in sorted(tracks.values(), key=lambda e: e["bcode"]):
acc = sums.setdefault(entry["bcode"], [0.0, 0.0])
acc[0] += summarize_track(entry["cofs"])[2]
acc[1] += 1
return [(batch, acc[0] / acc[1]) for batch, acc in sorted(sums.items())]
def q6(store_path: Path) -> list[tuple]:
store = _store(store_path)
return [
(s["ccode"].value, float(s["load"].value), float(s["wear"].value))
for s in store.query(SPARQL["q6"])
]
def q7(store_path: Path) -> list[tuple]:
store = _store(store_path)
tracks = _track_cofs(store.query(SPARQL["q7"]), "track", ("env", "speed", "load"))
groups: dict[tuple[str, int], list[float]] = {}
for key in sorted(tracks):
entry = tracks[key]
bucket = int(round(float(entry["speed"]) * float(entry["load"])))
run_in = summarize_track(entry["cofs"])[2]
acc = groups.setdefault((entry["env"], bucket), [0.0, 0.0])
for cof in entry["cofs"][run_in:]:
acc[0] += cof
acc[1] += 1
return [(env, bucket, acc[0] / acc[1]) for (env, bucket), acc in sorted(groups.items())]

161
common/relational.py Normal file
View File

@@ -0,0 +1,161 @@
"""Shared relational row stream for the SQLite and PostgreSQL converters.
Walks the CSV corpus once (manifest-validated via CorpusReader) and yields
(table, row_tuple) pairs in FK-dependency-safe order: a parent row is always
yielded before any row that references it. Both engines load the SAME
logical schema from this stream (docs/rules/db-sql-schema.md section 6).
"""
from __future__ import annotations
import logging
from typing import Iterator
from common.corpus import CorpusReader
from common.track_summary import summarize_track
logger = logging.getLogger(__name__)
# Dependency-safe order; consumers that buffer rows must flush in this order.
TABLES = [
"instruments",
"batches",
"runs",
"wafers",
"coupons",
"tracks",
"simtra_profiles",
"xrf_points",
"profilometry_points",
"nanoindentation",
"afm",
"friction_cycles",
"friction_loop_points",
"wear",
"track_summary",
]
COLUMNS = {
"instruments": ("instrument_id", "instrument_code", "name", "role"),
"batches": ("batch_id", "batch_code", "pt_gun_tilt_deg", "au_gun_tilt_deg", "pt_power_w", "au_power_w", "pt_discharge_v", "au_discharge_v", "deposition_date"),
"runs": ("run_id", "run_code", "environment", "date", "plates", "operator"),
"wafers": ("wafer_id", "wafer_code", "batch_id", "wafer_index", "deposition_date", "coupons", "friction_coupons", "reserve_coupons"),
"coupons": ("coupon_id", "coupon_code", "wafer_id", "batch_id", "grid_row", "grid_col", "thickness_um", "ra_nm", "au_wtpct_mean", "run_id", "plate", "probe", "square"),
"tracks": ("track_id", "track_code", "coupon_id", "run_id", "environment", "load_mn", "stroke_mm", "speed_mm_s", "counterface_id", "started_at"),
"simtra_profiles": ("batch_id", "row_no", "angle_deg", "energy_ev", "pt_flux", "au_flux"),
"xrf_points": ("coupon_id", "grid_x", "grid_y", "pt_wtpct", "au_wtpct"),
"profilometry_points": ("coupon_id", "grid_x", "grid_y", "thickness_um"),
"nanoindentation": ("coupon_id", "indent_id", "x_um", "y_um", "hardness_gpa", "reduced_modulus_gpa", "max_load_mn"),
"afm": ("coupon_id", "ra_nm", "rq_nm", "image_file"),
"friction_cycles": ("track_id", "cycle", "cof"),
"friction_loop_points": ("track_id", "cycle", "pt", "position_um", "friction_force_mn"),
"wear": ("track_id", "wear_volume_um3", "k_archard", "sliding_distance_m"),
"track_summary": ("track_id", "cof_ss_mean", "cof_ss_std", "run_in_cycles"),
}
EXPECTED_COUNTS_KEYS = TABLES # every table is count-validated after load
def expected_counts(cfg: dict) -> dict[str, int]:
v = cfg["volumes"]
h = cfg["hierarchy"]
return {
"instruments": len(cfg["instruments"]),
"batches": h["batches"],
"runs": cfg["friction_assignment"]["runs"],
"wafers": h["batches"] * h["wafers_per_batch"],
"coupons": v["coupons_total"],
"tracks": v["tracks_total"],
"simtra_profiles": h["batches"] * v["simtra_rows_per_batch"],
"xrf_points": v["xrf_points_total"],
"profilometry_points": v["coupons_total"] * v["profilometry_points_per_coupon"],
"nanoindentation": v["coupons_total"] * v["nanoindentation_indents_per_coupon"],
"afm": v["coupons_total"],
"friction_cycles": v["cycle_rows_total"],
"friction_loop_points": v["loop_points_total"],
"wear": v["tracks_total"],
"track_summary": v["tracks_total"],
}
def stream_rows(cfg: dict, reader: CorpusReader) -> Iterator[tuple[str, tuple]]:
batch_ids: dict[str, int] = {}
run_ids: dict[str, int] = {}
for i, inst in enumerate(cfg["instruments"], 1):
yield "instruments", (i, inst["instrument_id"], inst["name"], inst["role"])
for i, row in enumerate(reader.dicts("batches.csv"), 1):
batch_ids[row["batch_code"]] = i
yield "batches", (
i, row["batch_code"], row["pt_gun_tilt_deg"], row["au_gun_tilt_deg"],
row["pt_power_w"], row["au_power_w"], row["pt_discharge_v"],
row["au_discharge_v"], row["deposition_date"],
)
for i, row in enumerate(reader.dicts("runs.csv"), 1):
run_ids[row["run_code"]] = i
yield "runs", (i, row["run_code"], row["environment"], row["date"], row["plates"], row["operator"])
for code, batch_id in batch_ids.items():
for row_no, row in enumerate(reader.dicts(f"simtra/simtra_profile_{code}.csv"), 1):
yield "simtra_profiles", (batch_id, row_no, row["angle_deg"], row["energy_ev"], row["pt_flux"], row["au_flux"])
h = cfg["hierarchy"]
fa = cfg["friction_assignment"]
wafer_id = coupon_id = track_id = 0
for batch in cfg["deposition_matrix"]:
b_code = batch["batch_code"]
for w in range(1, h["wafers_per_batch"] + 1):
wafer_id += 1
wafer_dir = f"batch_{b_code}/wafer_W{w}"
info = reader.dicts(f"{wafer_dir}/wafer_info.csv")[0]
yield "wafers", (
wafer_id, info["wafer_code"], batch_ids[b_code], info["wafer_index"],
info["deposition_date"], info["coupons"], info["friction_coupons"], info["reserve_coupons"],
)
for c in range(1, h["coupons_per_wafer"] + 1):
coupon_id += 1
rel_dir = f"{wafer_dir}/coupon_C{c:02d}"
info = reader.dicts(f"{rel_dir}/coupon_info.csv")[0]
is_friction = info["run_code"] != "RESERVE"
yield "coupons", (
coupon_id, info["coupon_code"], wafer_id, batch_ids[b_code],
info["grid_row"], info["grid_col"], info["thickness_um"], info["ra_nm"],
info["au_wtpct_mean"], run_ids[info["run_code"]] if is_friction else None,
info.get("plate"), info.get("probe"), info.get("square"),
)
for row in reader.dicts(f"{rel_dir}/xrf_map.csv"):
yield "xrf_points", (coupon_id, row["grid_x"], row["grid_y"], row["pt_wtpct"], row["au_wtpct"])
for row in reader.dicts(f"{rel_dir}/profilometry.csv"):
yield "profilometry_points", (coupon_id, row["grid_x"], row["grid_y"], row["thickness_um"])
for row in reader.dicts(f"{rel_dir}/nanoindentation.csv"):
yield "nanoindentation", (
coupon_id, row["indent_id"], row["x_um"], row["y_um"],
row["hardness_gpa"], row["reduced_modulus_gpa"], row["max_load_mn"],
)
afm = reader.dicts(f"{rel_dir}/afm.csv")[0]
yield "afm", (coupon_id, afm["ra_nm"], afm["rq_nm"], afm["image_file"])
if not is_friction:
continue
for t in range(1, fa["tracks_per_friction_coupon"] + 1):
track_id += 1
track_dir = f"{rel_dir}/track_T{t}"
tinfo = reader.dicts(f"{track_dir}/track_info.csv")[0]
yield "tracks", (
track_id, tinfo["track_code"], coupon_id, run_ids[tinfo["run_code"]],
tinfo["environment"], tinfo["load_mn"], tinfo["stroke_mm"], tinfo["speed_mm_s"],
tinfo["counterface_id"], tinfo["started_at"],
)
cofs: list[float] = []
for row in reader.dicts(f"{track_dir}/cof_vs_cycle.csv"):
cofs.append(row["cof"])
yield "friction_cycles", (track_id, row["cycle"], row["cof"])
ss_mean, ss_std, run_in = summarize_track(cofs)
yield "track_summary", (track_id, ss_mean, ss_std, run_in)
pt = 0
last_cycle = None
for row in reader.dicts(f"{track_dir}/friction_loops.csv"):
pt = pt + 1 if row["cycle"] == last_cycle else 1
last_cycle = row["cycle"]
yield "friction_loop_points", (track_id, row["cycle"], pt, row["position_um"], row["friction_force_mn"])
wear = reader.dicts(f"{track_dir}/wear.csv")[0]
yield "wear", (track_id, wear["wear_volume_um3"], wear["k_archard"], wear["sliding_distance_m"])
logger.info("batch %s streamed (through coupon %d, track %d)", b_code, coupon_id, track_id)

73
common/results.py Normal file
View File

@@ -0,0 +1,73 @@
"""Canonical benchmark result handling (spec 08, test-pipeline-validation).
Every implementation of a benchmark query returns the same canonical rows:
tuples of str / int / float columns. Rows are compared SORTED; floats agree
within 1e-9. The canonical expected artifacts are produced from PostgreSQL
and cross-validated against SQLite (task 08); every measured run (task 09)
re-validates its rows against them.
Two serializations:
- full: floats via repr() - lossless, used for stdout and expected/*.rows.csv;
- rounded (6 decimals): used ONLY for the sha256 fingerprint, so engines
whose aggregate arithmetic differs in the last bits hash identically.
The tolerant row comparison is the authoritative check; the sha256 is a
compact fingerprint for reports and markers.
"""
from __future__ import annotations
import hashlib
import sys
from common.corpus import coerce
FLOAT_TOL = 1e-9
Row = tuple
def sort_rows(rows: list[Row]) -> list[Row]:
return sorted(rows)
def _fmt(value, rounded: bool) -> str:
if isinstance(value, float):
return f"{round(value, 6):.6f}" if rounded else repr(value)
return str(value)
def serialize(rows: list[Row], rounded: bool = False) -> str:
return "\n".join(",".join(_fmt(v, rounded) for v in row) for row in rows) + "\n"
def sha256_of(rows: list[Row]) -> str:
return hashlib.sha256(serialize(sort_rows(rows), rounded=True).encode("ascii")).hexdigest()
def print_rows(rows: list[Row]) -> None:
sys.stdout.write(serialize(sort_rows(rows)))
def parse_rows(text: str) -> list[Row]:
rows = []
for line in text.splitlines():
if line:
rows.append(tuple(coerce(cell) for cell in line.split(",")))
return rows
def compare(expected: list[Row], actual: list[Row], tol: float = FLOAT_TOL) -> str | None:
"""Return None when equal within tolerance, else a first-difference message."""
exp, act = sort_rows(expected), sort_rows(actual)
if len(exp) != len(act):
return f"row count {len(act)} != expected {len(exp)}"
for i, (er, ar) in enumerate(zip(exp, act)):
if len(er) != len(ar):
return f"row {i}: arity {len(ar)} != {len(er)}"
for j, (ev, av) in enumerate(zip(er, ar)):
if isinstance(ev, float) or isinstance(av, float):
if abs(float(ev) - float(av)) > tol:
return f"row {i} col {j}: {av!r} != {ev!r} (tol {tol})"
elif ev != av:
return f"row {i} col {j}: {av!r} != {ev!r}"
return None

55
common/storage_sizes.py Normal file
View File

@@ -0,0 +1,55 @@
"""Idempotent per-(format, variant) updates to out/bench/storage_sizes.csv.
Contract (docs/rules/build-pipeline-tasks.md section 4): columns
format,variant,bytes; a task re-run REPLACES its own (format, variant) rows
and never touches other rows. Replacement is per variant, not per format:
converters 04-07 own the working-footprint variants while task 09 appends
archival variants of the same formats.
"""
from __future__ import annotations
import logging
import os
import time
from pathlib import Path
logger = logging.getLogger(__name__)
HEADER = "format,variant,bytes"
LOCK_TIMEOUT_S = 30.0 # converters 04-07 may run in parallel; the update is a
LOCK_POLL_S = 0.2 # read-modify-write, so it is serialized via a lock file
def update_storage_sizes(bench_dir: Path, fmt: str, variants: list[tuple[str, int]]) -> Path:
bench_dir.mkdir(parents=True, exist_ok=True)
path = bench_dir / "storage_sizes.csv"
lock = bench_dir / "storage_sizes.lock"
fd = _acquire(lock)
try:
replaced = {(fmt, variant) for variant, _ in variants}
kept: list[str] = []
if path.exists():
lines = path.read_text(encoding="ascii").splitlines()
if lines and lines[0] != HEADER:
raise ValueError(f"{path}: unexpected header {lines[0]!r}")
kept = [ln for ln in lines[1:] if ln and tuple(ln.split(",")[:2]) not in replaced]
rows = kept + [f"{fmt},{variant},{nbytes}" for variant, nbytes in variants]
with open(path, "w", encoding="ascii", newline="\n") as fh:
fh.write(HEADER + "\n" + "\n".join(rows) + "\n")
finally:
os.close(fd)
lock.unlink()
logger.info("storage_sizes.csv: %s -> %s", fmt, ", ".join(f"{v}={b}" for v, b in variants))
return path
def _acquire(lock: Path) -> int:
deadline = time.monotonic() + LOCK_TIMEOUT_S
while True:
try:
return os.open(lock, os.O_CREAT | os.O_EXCL | os.O_WRONLY)
except FileExistsError:
if time.monotonic() > deadline:
raise TimeoutError(f"storage_sizes lock held too long: {lock}")
time.sleep(LOCK_POLL_S)

48
common/track_summary.py Normal file
View File

@@ -0,0 +1,48 @@
"""Per-track steady-state summary (docs/rules/db-sql-schema.md section 5).
Both engines must produce identical values: SQLite calls this function
during load (task 05); the PostgreSQL materialized view (task 06)
implements the same three steps in SQL. Plain sequential float arithmetic
keeps results within the 1e-9 checksum tolerance across implementations.
"""
from __future__ import annotations
import math
TAIL_START_CYCLE = 501 # steady-state proxy window: cycles 501..N
def summarize_track(cofs: list[float]) -> tuple[float, float, int]:
"""Return (cof_ss_mean, cof_ss_std, run_in_cycles) for cofs[i] = cof at cycle i+1."""
n = len(cofs)
tail = cofs[TAIL_START_CYCLE - 1:]
m = _mean(tail)
s = _stddev_pop(tail, m)
run_in = TAIL_START_CYCLE - 1 # fallback when no cycle enters the 2-sigma band
band = 2.0 * s
for c in range(1, n + 1):
if abs(cofs[c - 1] - m) < band:
run_in = c
break
ss = cofs[run_in:] # cycles strictly greater than run_in
ss_mean = _mean(ss)
ss_std = _stddev_pop(ss, ss_mean)
return ss_mean, ss_std, run_in
def _mean(values: list[float]) -> float:
total = 0.0
for v in values:
total += v
return total / len(values)
def _stddev_pop(values: list[float], mean: float) -> float:
acc = 0.0
for v in values:
d = v - mean
acc += d * d
return math.sqrt(acc / len(values))

293
convert_json.py Normal file
View File

@@ -0,0 +1,293 @@
"""Task 04 - Convert the CSV corpus to JSON-LD.
Reads ./out/csv/ (validated file-by-file against MANIFEST.csv) and writes
two ontology-annotated variants:
- FULL: one compact JSON-LD file per coupon under ./out/json/full/, bulk
points embedded. Like spec 07's RDF modeling, bulk points (cycles, loop
points, map points) carry no @type - the containing predicate types them;
this keeps FULL lean, so its size lands below the 0.7-1.2 GB the spec
sketched for a fully typed serialization.
- HYBRID: single ./out/json/hybrid/dataset.jsonld holding every metadata
entity; every per-point array (cycles, loops, xrf, profilometry, indents)
is replaced by a {sourceFile, rows, sha256} reference into the CSV tree.
Appends both sizes to ./out/bench/storage_sizes.csv and writes
./out/.done/04.ok. Spec: docs/specs/04_convert_json.md.
"""
from __future__ import annotations
import json
import logging
import shutil
import sys
from pathlib import Path
from rdflib import Graph
from common.corpus import CorpusReader, ValidationError
from common.pipeline import (
check_dependencies,
load_lab_config,
parse_task_args,
process_metrics,
remove_stale_marker,
write_marker,
)
from common.storage_sizes import update_storage_sizes
logger = logging.getLogger(__name__)
TASK_ID = "04"
DEPENDS_ON = ["01", "03"]
VOCAB = "https://sandia.gov/ontology/tribology#"
ID_BASE = "https://sandia.gov/ontology/tribology/id/"
CONTEXT = {
"@vocab": VOCAB,
"tribo": VOCAB,
"qudt": "http://qudt.org/schema/qudt/",
"unit": "http://qudt.org/vocab/unit/",
"prov": "http://www.w3.org/ns/prov#",
"batch": ID_BASE + "batch/",
"wafer": ID_BASE + "wafer/",
"coupon": ID_BASE + "coupon/",
"track": ID_BASE + "track/",
"run": ID_BASE + "run/",
"instrument": ID_BASE + "instrument/",
"partOf": {"@id": "tribo:partOf", "@type": "@id"},
"cutFrom": {"@id": "tribo:cutFrom", "@type": "@id"},
"performedOn": {"@id": "tribo:performedOn", "@type": "@id"},
"performedBy": {"@id": "tribo:performedBy", "@type": "@id"},
"duringRun": {"@id": "tribo:duringRun", "@type": "@id"},
"started_at": {"@id": "prov:startedAtTime"},
}
MEASUREMENTS = [
# (csv file, json key, @type, instrument id, points key)
("xrf_map.csv", "xrf", "CompositionMeasurement", "m4_tornado", "points"),
("profilometry.csv", "profilometry", "ProfilometryMeasurement", "profilometer", "points"),
("nanoindentation.csv", "nanoindentation", "NanoindentationMeasurement", "ti980", "indents"),
]
class JsonConverter:
def __init__(self, cfg: dict, csv_root: Path, json_root: Path) -> None:
self.cfg = cfg
self.json_root = json_root
self.reader = CorpusReader(csv_root)
self.manifest = self.reader.manifest
self.graph: list[dict] = [] # hybrid metadata entities
self.full_bytes = 0
self.full_files = 0
self.counts = {"coupons": 0, "tracks": 0, "cycle_rows": 0, "loop_rows": 0, "xrf_rows": 0}
def rows_as_dicts(self, relpath: str) -> list[dict]:
return self.reader.dicts(relpath)
def source_ref(self, relpath: str) -> dict:
rows, _, sha = self.manifest[relpath]
return {"sourceFile": relpath, "rows": rows, "sha256": sha}
# --- flat entities (hybrid graph) ---
def convert_flat(self) -> None:
for inst in self.cfg["instruments"]:
self.graph.append({
"@id": f"instrument:{inst['instrument_id']}",
"@type": "Instrument",
"name": inst["name"],
"role": inst["role"],
})
for row in self.rows_as_dicts("batches.csv"):
self.graph.append({"@id": f"batch:{row['batch_code']}", "@type": "Batch", **row})
self.graph.append({
"@id": f"batch:{row['batch_code']}-simtra",
"@type": "SimtraProfile",
"performedOn": f"batch:{row['batch_code']}",
"performedBy": "instrument:simtra",
**self.source_ref(f"simtra/simtra_profile_{row['batch_code']}.csv"),
})
for row in self.rows_as_dicts("runs.csv"):
self.graph.append({"@id": f"run:{row['run_code']}", "@type": "Run", **row})
# --- per-coupon conversion ---
def convert_coupons(self) -> None:
h = self.cfg["hierarchy"]
full_dir = self.json_root / "full"
full_dir.mkdir(parents=True, exist_ok=True)
for batch in self.cfg["deposition_matrix"]:
b_code = batch["batch_code"]
for w in range(1, h["wafers_per_batch"] + 1):
wafer_code = f"{b_code}-W{w}"
wafer_dir = f"batch_{b_code}/wafer_W{w}"
info = self.rows_as_dicts(f"{wafer_dir}/wafer_info.csv")[0]
self.graph.append({
"@id": f"wafer:{wafer_code}", "@type": "Wafer",
"partOf": f"batch:{b_code}", **info,
})
for c in range(1, h["coupons_per_wafer"] + 1):
self.convert_coupon(f"{wafer_dir}/coupon_C{c:02d}", full_dir)
logger.debug("wafer %s converted", wafer_code)
logger.info("batch %s converted: %.1f MiB full so far", b_code, self.full_bytes / (1024 * 1024))
def convert_coupon(self, rel_dir: str, full_dir: Path) -> None:
info = self.rows_as_dicts(f"{rel_dir}/coupon_info.csv")[0]
code = info["coupon_code"]
coupon_id = f"coupon:{code}"
is_friction = info["run_code"] != "RESERVE"
full_node: dict = {
"@context": CONTEXT, "@id": coupon_id, "@type": "Coupon",
"cutFrom": f"wafer:{info['wafer_code']}", **info,
}
hybrid_node: dict = {
"@id": coupon_id, "@type": "Coupon",
"cutFrom": f"wafer:{info['wafer_code']}", **info,
}
for csv_name, key, cls, inst, points_key in MEASUREMENTS:
relpath = f"{rel_dir}/{csv_name}"
meta = {"@type": cls, "performedOn": coupon_id, "performedBy": f"instrument:{inst}"}
full_node[key] = {**meta, points_key: self.rows_as_dicts(relpath)}
hybrid_node[key] = {**meta, **self.source_ref(relpath)}
self.counts["xrf_rows"] += self.manifest[f"{rel_dir}/xrf_map.csv"][0]
afm = self.rows_as_dicts(f"{rel_dir}/afm.csv")[0]
afm_node = {"@type": "AFMMeasurement", "performedOn": coupon_id, "performedBy": "instrument:afm", **afm}
full_node["afm"] = afm_node
hybrid_node["afm"] = afm_node
if is_friction:
tracks_full = []
for t in range(1, self.cfg["friction_assignment"]["tracks_per_friction_coupon"] + 1):
full_track, hybrid_track = self.convert_track(f"{rel_dir}/track_T{t}", coupon_id, info["run_code"])
tracks_full.append(full_track)
self.graph.append(hybrid_track)
full_node["tracks"] = tracks_full
out_path = full_dir / f"{code}.jsonld"
text = json.dumps(full_node, separators=(",", ":"))
data = text.encode("ascii")
out_path.write_bytes(data)
self.full_bytes += len(data)
self.full_files += 1
self.graph.append(hybrid_node)
self.counts["coupons"] += 1
def convert_track(self, rel_dir: str, coupon_id: str, run_code: str) -> tuple[dict, dict]:
info = self.rows_as_dicts(f"{rel_dir}/track_info.csv")[0]
track_id = f"track:{info['track_code']}"
base = {
"@id": track_id, "@type": "FrictionTest",
"partOf": coupon_id, "duringRun": f"run:{run_code}",
"performedBy": "instrument:rapid", **info,
}
wear = self.rows_as_dicts(f"{rel_dir}/wear.csv")[0]
wear_node = {"@type": "WearMeasurement", **wear}
cycles_path = f"{rel_dir}/cof_vs_cycle.csv"
loops_path = f"{rel_dir}/friction_loops.csv"
full_track = {
**base,
"cycles": self.rows_as_dicts(cycles_path),
"loop_points": self.rows_as_dicts(loops_path),
"wear": wear_node,
}
hybrid_track = {
**base,
"cycles": self.source_ref(cycles_path),
"loop_points": self.source_ref(loops_path),
"wear": wear_node,
}
self.counts["cycle_rows"] += len(full_track["cycles"])
self.counts["loop_rows"] += len(full_track["loop_points"])
self.counts["tracks"] += 1
return full_track, hybrid_track
# --- outputs ---
def write_hybrid(self) -> int:
hybrid_dir = self.json_root / "hybrid"
hybrid_dir.mkdir(parents=True, exist_ok=True)
path = hybrid_dir / "dataset.jsonld"
with open(path, "w", encoding="ascii", newline="\n") as fh:
json.dump({"@context": CONTEXT, "@graph": self.graph}, fh, separators=(",", ":"))
return path.stat().st_size
def validate_counts(self) -> None:
v = self.cfg["volumes"]
expected = {
"coupons": v["coupons_total"],
"tracks": v["tracks_total"],
"cycle_rows": v["cycle_rows_total"],
"loop_rows": v["loop_points_total"],
"xrf_rows": v["xrf_points_total"],
}
for key, exp in expected.items():
if self.counts[key] != exp:
raise ValidationError(f"count mismatch: {key}: converted {self.counts[key]} != expected {exp}")
logger.info("all converted counts match lab_config volumes: %s", self.counts)
def validate_sample(paths: list[Path]) -> None:
"""Smoke test (spec 04 + test-pipeline-validation): json.load + rdflib parse."""
for path in paths:
with open(path, encoding="ascii") as fh:
json.load(fh)
g = Graph()
g.parse(path, format="json-ld")
if len(g) == 0:
raise ValidationError(f"{path}: rdflib parsed 0 triples")
logger.info("sample %s: valid JSON, %d triples via rdflib", path.name, len(g))
def main() -> int:
args = parse_task_args("Task 04: convert the CSV corpus to JSON-LD")
out_root: Path = args.out_root
json_root = out_root / "json"
try:
check_dependencies(out_root, DEPENDS_ON)
remove_stale_marker(out_root, TASK_ID)
if json_root.exists():
logger.info("re-run: removing previous output %s", json_root)
shutil.rmtree(json_root)
cfg = load_lab_config(out_root)
conv = JsonConverter(cfg, out_root / "csv", json_root)
conv.convert_flat()
conv.convert_coupons()
conv.validate_counts()
hybrid_bytes = conv.write_hybrid()
validate_sample([
json_root / "full" / "B722-W2-C13.jsonld",
json_root / "hybrid" / "dataset.jsonld",
])
update_storage_sizes(out_root / "bench", "json", [("full", conv.full_bytes), ("hybrid", hybrid_bytes)])
entries = {
"full_files": conv.full_files,
"full_bytes": conv.full_bytes,
"hybrid_bytes": hybrid_bytes,
"coupons": conv.counts["coupons"],
"tracks": conv.counts["tracks"],
"cycle_rows": conv.counts["cycle_rows"],
"loop_points": conv.counts["loop_rows"],
}
entries.update(process_metrics())
marker_path = write_marker(out_root, TASK_ID, entries)
except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1
print(f"task 04 ok: full {conv.full_files} files, {conv.full_bytes / (1024 * 1024):.1f} MiB; "
f"hybrid {hybrid_bytes / (1024 * 1024):.1f} MiB")
print(f"entities: coupons={conv.counts['coupons']} tracks={conv.counts['tracks']} "
f"cycles={conv.counts['cycle_rows']} loop_points={conv.counts['loop_rows']}")
print(f"marker: {marker_path}")
return 0
if __name__ == "__main__":
sys.exit(main())

334
convert_postgresql.py Normal file
View File

@@ -0,0 +1,334 @@
"""Task 06 - Convert the CSV corpus to PostgreSQL.
Loads the same logical schema as task 05 into the PostgreSQL 16 database
named by the TRIBO_PG_DSN environment variable: COPY FROM STDIN, RANGE
partitioning of the two bulk tables by track_id (12 partitions of 120
tracks), constraints and indexes created after the load (variant B:
enforced foreign keys everywhere - docs/research/Tribology_FK_Architecture.html),
track_summary as a materialized view (algorithm fixed in
docs/rules/db-sql-schema.md section 5, cross-checked against
common/track_summary.py). Host CPU/RAM during the load is sampled from the
database host's Prometheus (TRIBO_PROM_URL). Appends live (and, when
pg_dump is available, dump) sizes to storage_sizes.csv and writes
./out/.done/06.ok. Spec: docs/specs/06_convert_postgresql.md.
"""
from __future__ import annotations
import logging
import shutil
import subprocess
import sys
import time
from pathlib import Path
import psycopg
from common.corpus import CorpusReader, ValidationError
from common.monitoring import get_prom_url, host_window_metrics
from common.pg import get_dsn
from common.pipeline import (
check_dependencies,
load_lab_config,
parse_task_args,
process_metrics,
remove_stale_marker,
write_marker,
)
from common.relational import COLUMNS, TABLES, expected_counts, stream_rows
from common.storage_sizes import update_storage_sizes
from common.track_summary import summarize_track
logger = logging.getLogger(__name__)
TASK_ID = "06"
DEPENDS_ON = ["01", "03"]
BATCH_ROWS = 50000 # rows per COPY chunk (mirrors the 50k batching of spec 05)
PARTITIONS = 12 # RANGE partitions of 120 tracks each, spec 06
TRACKS_PER_PARTITION = 120
SUMMARY_PARITY_TRACKS = (1, 720, 1440) # matview cross-checked vs common/track_summary.py
PARITY_TOLERANCE = 1e-9 # cross-engine float tolerance, test-pipeline-validation
# Float columns are DOUBLE PRECISION, not REAL: the corpus is float64 and
# REAL (float4) would break the 1e-9 cross-engine result parity that the
# whole benchmark depends on (db-sql-schema section 6; spec updated).
CREATE_TABLES = [
"CREATE TABLE instruments (instrument_id INTEGER NOT NULL, instrument_code TEXT NOT NULL, name TEXT NOT NULL, role TEXT NOT NULL)",
"CREATE TABLE batches (batch_id INTEGER NOT NULL, batch_code TEXT NOT NULL, pt_gun_tilt_deg INTEGER NOT NULL, au_gun_tilt_deg INTEGER NOT NULL, pt_power_w INTEGER NOT NULL, au_power_w INTEGER NOT NULL, pt_discharge_v INTEGER NOT NULL, au_discharge_v INTEGER NOT NULL, deposition_date DATE NOT NULL)",
"CREATE TABLE runs (run_id INTEGER NOT NULL, run_code TEXT NOT NULL, environment TEXT NOT NULL, date DATE NOT NULL, plates INTEGER NOT NULL, operator TEXT NOT NULL)",
"CREATE TABLE wafers (wafer_id INTEGER NOT NULL, wafer_code TEXT NOT NULL, batch_id INTEGER NOT NULL, wafer_index INTEGER NOT NULL, deposition_date DATE NOT NULL, coupons INTEGER NOT NULL, friction_coupons INTEGER NOT NULL, reserve_coupons INTEGER NOT NULL)",
"CREATE TABLE coupons (coupon_id INTEGER NOT NULL, coupon_code TEXT NOT NULL, wafer_id INTEGER NOT NULL, batch_id INTEGER NOT NULL, grid_row INTEGER NOT NULL, grid_col INTEGER NOT NULL, thickness_um DOUBLE PRECISION NOT NULL, ra_nm DOUBLE PRECISION NOT NULL, au_wtpct_mean DOUBLE PRECISION NOT NULL, run_id INTEGER, plate INTEGER, probe INTEGER, square INTEGER)",
"CREATE TABLE tracks (track_id INTEGER NOT NULL, track_code TEXT NOT NULL, coupon_id INTEGER NOT NULL, run_id INTEGER NOT NULL, environment TEXT NOT NULL, load_mn DOUBLE PRECISION NOT NULL, stroke_mm DOUBLE PRECISION NOT NULL, speed_mm_s DOUBLE PRECISION NOT NULL, counterface_id TEXT NOT NULL, started_at TIMESTAMPTZ NOT NULL)",
"CREATE TABLE simtra_profiles (batch_id INTEGER NOT NULL, row_no INTEGER NOT NULL, angle_deg DOUBLE PRECISION NOT NULL, energy_ev DOUBLE PRECISION NOT NULL, pt_flux DOUBLE PRECISION NOT NULL, au_flux DOUBLE PRECISION NOT NULL)",
"CREATE TABLE xrf_points (coupon_id INTEGER NOT NULL, grid_x INTEGER NOT NULL, grid_y INTEGER NOT NULL, pt_wtpct DOUBLE PRECISION NOT NULL, au_wtpct DOUBLE PRECISION NOT NULL)",
"CREATE TABLE profilometry_points (coupon_id INTEGER NOT NULL, grid_x INTEGER NOT NULL, grid_y INTEGER NOT NULL, thickness_um DOUBLE PRECISION NOT NULL)",
"CREATE TABLE nanoindentation (coupon_id INTEGER NOT NULL, indent_id INTEGER NOT NULL, x_um DOUBLE PRECISION NOT NULL, y_um DOUBLE PRECISION NOT NULL, hardness_gpa DOUBLE PRECISION NOT NULL, reduced_modulus_gpa DOUBLE PRECISION NOT NULL, max_load_mn DOUBLE PRECISION NOT NULL)",
"CREATE TABLE afm (coupon_id INTEGER NOT NULL, ra_nm DOUBLE PRECISION NOT NULL, rq_nm DOUBLE PRECISION NOT NULL, image_file TEXT NOT NULL)",
"CREATE TABLE friction_cycles (track_id INTEGER NOT NULL, cycle INTEGER NOT NULL, cof DOUBLE PRECISION NOT NULL) PARTITION BY RANGE (track_id)",
"CREATE TABLE friction_loop_points (track_id INTEGER NOT NULL, cycle INTEGER NOT NULL, pt INTEGER NOT NULL, position_um DOUBLE PRECISION NOT NULL, friction_force_mn DOUBLE PRECISION NOT NULL) PARTITION BY RANGE (track_id)",
"CREATE TABLE wear (track_id INTEGER NOT NULL, wear_volume_um3 DOUBLE PRECISION NOT NULL, k_archard DOUBLE PRECISION NOT NULL, sliding_distance_m DOUBLE PRECISION NOT NULL)",
]
CONSTRAINTS = [
"ALTER TABLE instruments ADD PRIMARY KEY (instrument_id)",
"ALTER TABLE instruments ADD UNIQUE (instrument_code)",
"ALTER TABLE batches ADD PRIMARY KEY (batch_id)",
"ALTER TABLE batches ADD UNIQUE (batch_code)",
"ALTER TABLE runs ADD PRIMARY KEY (run_id)",
"ALTER TABLE runs ADD UNIQUE (run_code)",
"ALTER TABLE wafers ADD PRIMARY KEY (wafer_id)",
"ALTER TABLE wafers ADD UNIQUE (wafer_code)",
"ALTER TABLE coupons ADD PRIMARY KEY (coupon_id)",
"ALTER TABLE coupons ADD UNIQUE (coupon_code)",
"ALTER TABLE tracks ADD PRIMARY KEY (track_id)",
"ALTER TABLE tracks ADD UNIQUE (track_code)",
"ALTER TABLE simtra_profiles ADD PRIMARY KEY (batch_id, row_no)",
"ALTER TABLE xrf_points ADD PRIMARY KEY (coupon_id, grid_x, grid_y)",
"ALTER TABLE profilometry_points ADD PRIMARY KEY (coupon_id, grid_x, grid_y)",
"ALTER TABLE nanoindentation ADD PRIMARY KEY (coupon_id, indent_id)",
"ALTER TABLE afm ADD PRIMARY KEY (coupon_id)",
"ALTER TABLE friction_cycles ADD PRIMARY KEY (track_id, cycle)",
"ALTER TABLE friction_loop_points ADD PRIMARY KEY (track_id, cycle, pt)",
"ALTER TABLE wear ADD PRIMARY KEY (track_id)",
"ALTER TABLE wafers ADD FOREIGN KEY (batch_id) REFERENCES batches",
"ALTER TABLE coupons ADD FOREIGN KEY (wafer_id) REFERENCES wafers",
"ALTER TABLE coupons ADD FOREIGN KEY (batch_id) REFERENCES batches",
"ALTER TABLE coupons ADD FOREIGN KEY (run_id) REFERENCES runs",
"ALTER TABLE tracks ADD FOREIGN KEY (coupon_id) REFERENCES coupons",
"ALTER TABLE tracks ADD FOREIGN KEY (run_id) REFERENCES runs",
"ALTER TABLE simtra_profiles ADD FOREIGN KEY (batch_id) REFERENCES batches",
"ALTER TABLE xrf_points ADD FOREIGN KEY (coupon_id) REFERENCES coupons",
"ALTER TABLE profilometry_points ADD FOREIGN KEY (coupon_id) REFERENCES coupons",
"ALTER TABLE nanoindentation ADD FOREIGN KEY (coupon_id) REFERENCES coupons",
"ALTER TABLE afm ADD FOREIGN KEY (coupon_id) REFERENCES coupons",
"ALTER TABLE friction_cycles ADD FOREIGN KEY (track_id) REFERENCES tracks",
"ALTER TABLE friction_loop_points ADD FOREIGN KEY (track_id) REFERENCES tracks",
"ALTER TABLE wear ADD FOREIGN KEY (track_id) REFERENCES tracks",
]
INDEX_DDL = [
"CREATE INDEX ix_coupons_au_wtpct_mean ON coupons(au_wtpct_mean)", # Q2 selective filter
"CREATE INDEX ix_tracks_run_id ON tracks(run_id)", # Q5/Q7 joins to runs
"CREATE INDEX ix_tracks_coupon_id ON tracks(coupon_id)", # Q3/Q6 joins to coupons
"CREATE INDEX ix_runs_environment ON runs(environment)", # Q7 grouping
"CREATE INDEX ix_tracks_environment_load ON tracks(environment, load_mn)", # Q4 filter
# Q5 pruning within partitions; cascades to every partition (spec 06).
# nanoindentation(coupon_id) from the spec is omitted: the PK prefix serves it.
"CREATE INDEX ix_brin_friction_cycles_cycle ON friction_cycles USING BRIN (cycle)",
]
# Materialized view semantics fixed in docs/rules/db-sql-schema.md section 5.
MATVIEW_SQL = """
CREATE MATERIALIZED VIEW track_summary AS
WITH tail AS (
SELECT track_id, avg(cof) AS m, stddev_pop(cof) AS s
FROM friction_cycles
WHERE cycle >= 501
GROUP BY track_id
),
run_in AS (
SELECT fc.track_id,
COALESCE(MIN(fc.cycle) FILTER (WHERE abs(fc.cof - t.m) < 2 * t.s), 500) AS run_in_cycles
FROM friction_cycles fc
JOIN tail t USING (track_id)
GROUP BY fc.track_id
)
SELECT r.track_id,
avg(fc.cof) AS cof_ss_mean,
stddev_pop(fc.cof) AS cof_ss_std,
r.run_in_cycles
FROM run_in r
JOIN friction_cycles fc ON fc.track_id = r.track_id AND fc.cycle > r.run_in_cycles
GROUP BY r.track_id, r.run_in_cycles
"""
def partition_ddl() -> list[str]:
ddl = []
for parent in ("friction_cycles", "friction_loop_points"):
for i in range(PARTITIONS):
lo = i * TRACKS_PER_PARTITION + 1
hi = lo + TRACKS_PER_PARTITION
ddl.append(f"CREATE TABLE {parent}_p{i + 1:02d} PARTITION OF {parent} FOR VALUES FROM ({lo}) TO ({hi})")
return ddl
class PgLoader:
def __init__(self, conn: psycopg.Connection) -> None:
self.conn = conn
self.buffers: dict[str, list[tuple]] = {t: [] for t in TABLES}
self.summary_parity: dict[int, tuple] = {}
def create_schema(self) -> None:
with self.conn.cursor() as cur:
cur.execute("DROP MATERIALIZED VIEW IF EXISTS track_summary")
for table in reversed(TABLES):
if table != "track_summary":
cur.execute(f"DROP TABLE IF EXISTS {table} CASCADE")
for ddl in CREATE_TABLES + partition_ddl():
cur.execute(ddl)
self.conn.commit()
logger.info("schema created: %d tables, %d partitions", len(CREATE_TABLES), 2 * PARTITIONS)
def load(self, cfg: dict, reader: CorpusReader) -> None:
pending = 0
for table, row in stream_rows(cfg, reader):
if table == "track_summary":
# computed server-side as a materialized view; keep a sample
# of the Python values for the cross-engine parity check
if row[0] in SUMMARY_PARITY_TRACKS:
self.summary_parity[row[0]] = row
continue
self.buffers[table].append(row)
pending += 1
if pending >= BATCH_ROWS:
self.flush_all()
pending = 0
self.flush_all()
def flush_all(self) -> None:
for table in TABLES:
buf = self.buffers.get(table)
if not buf:
continue
cols = ", ".join(COLUMNS[table])
with self.conn.cursor() as cur:
with cur.copy(f"COPY {table} ({cols}) FROM STDIN") as copy:
for row in buf:
copy.write_row(row)
buf.clear()
self.conn.commit()
def finalize(self) -> None:
with self.conn.cursor() as cur:
for ddl in CONSTRAINTS:
cur.execute(ddl)
for ddl in INDEX_DDL:
cur.execute(ddl)
cur.execute(MATVIEW_SQL)
cur.execute("CREATE UNIQUE INDEX ix_track_summary_track_id ON track_summary(track_id)") # Q2/Q3 joins
self.conn.commit()
logger.info("constraints, indexes and track_summary materialized view created")
def validate(self, cfg: dict) -> dict[str, int]:
counts: dict[str, int] = {}
with self.conn.cursor() as cur:
for table, exp in expected_counts(cfg).items():
got = cur.execute(f"SELECT COUNT(*) FROM {table}").fetchone()[0]
counts[table] = got
if got != exp:
raise ValidationError(f"row count mismatch: {table}: loaded {got} != expected {exp}")
reserve = cur.execute("SELECT COUNT(*) FROM coupons WHERE run_id IS NULL").fetchone()[0]
if reserve != cfg["friction_assignment"]["reserve_coupons_total"]:
raise ValidationError(f"reserve coupons: {reserve} != expected")
for track_id, py_row in self.summary_parity.items():
db_row = cur.execute(
"SELECT cof_ss_mean, cof_ss_std, run_in_cycles FROM track_summary WHERE track_id = %s",
(track_id,),
).fetchone()
if db_row is None:
raise ValidationError(f"track_summary missing track {track_id}")
if (abs(db_row[0] - py_row[1]) > PARITY_TOLERANCE
or abs(db_row[1] - py_row[2]) > PARITY_TOLERANCE
or db_row[2] != py_row[3]):
raise ValidationError(
f"track_summary parity failed for track {track_id}: db={db_row} py={py_row[1:]}"
)
logger.info(
"all %d table counts match, %d reserve coupons, matview parity ok for tracks %s",
len(counts), reserve, list(self.summary_parity),
)
return counts
def table_sizes(self) -> dict[str, int]:
sizes: dict[str, int] = {}
with self.conn.cursor() as cur:
for table in TABLES:
total = cur.execute(
"""SELECT COALESCE(pg_total_relation_size(%s::regclass), 0)
+ COALESCE((SELECT sum(pg_total_relation_size(inhrelid))
FROM pg_inherits WHERE inhparent = %s::regclass), 0)""",
(table, table),
).fetchone()[0]
sizes[table] = int(total)
return sizes
def make_dump(dsn: str, dump_path: Path) -> int:
pg_dump = shutil.which("pg_dump")
if not pg_dump:
logger.warning("pg_dump skipped (client not installed on this host; produce the dump on the database host)")
return 0
dump_path.parent.mkdir(parents=True, exist_ok=True)
subprocess.run([pg_dump, "--format=custom", f"--file={dump_path}", dsn], check=True)
return dump_path.stat().st_size
def main() -> int:
args = parse_task_args("Task 06: convert the CSV corpus to PostgreSQL")
out_root: Path = args.out_root
dsn = get_dsn()
try:
check_dependencies(out_root, DEPENDS_ON)
remove_stale_marker(out_root, TASK_ID)
cfg = load_lab_config(out_root)
t_start = time.time()
with psycopg.connect(dsn) as conn:
database = conn.info.dbname
logger.info("connected to database %r on %s", database, conn.info.host)
loader = PgLoader(conn)
loader.create_schema()
loader.load(cfg, CorpusReader(out_root / "csv"))
load_done = time.time()
loader.finalize()
counts = loader.validate(cfg)
sizes = loader.table_sizes()
with psycopg.connect(dsn, autocommit=True) as conn:
conn.execute("VACUUM ANALYZE")
t_end = time.time()
live_bytes = sum(sizes.values())
variants = [("live", live_bytes)]
dump_bytes = make_dump(dsn, out_root / "pg" / "tribo.dump")
if dump_bytes:
variants.append(("dump", dump_bytes))
update_storage_sizes(out_root / "bench", "pg", variants)
entries = {
"database": database,
"tables": len(counts),
"total_rows": sum(counts.values()),
"live_bytes": live_bytes,
"friction_cycles_bytes": sizes["friction_cycles"],
"friction_loop_points_bytes": sizes["friction_loop_points"],
"copy_seconds": round(load_done - t_start, 1),
"total_seconds": round(t_end - t_start, 1),
"dump": dump_bytes if dump_bytes else "skipped (pg_dump unavailable on this host)",
}
entries.update(process_metrics()) # client-side converter cost
prom_url = get_prom_url()
if prom_url:
metrics = host_window_metrics(prom_url, t_start, t_end)
if metrics:
entries.update(metrics)
logger.info("db host during load: %s", metrics)
marker_path = write_marker(out_root, TASK_ID, entries)
except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1
print(f"task 06 ok: database {database} ({live_bytes / (1024 * 1024):.1f} MiB live, "
f"{len(counts)} tables, {sum(counts.values())} rows, {entries['total_seconds']} s)")
print(f"bulk rows: cycles={counts['friction_cycles']} loop_points={counts['friction_loop_points']} "
f"track_summary={counts['track_summary']}")
if prom_url and "host_cpu_avg_pct" in entries:
print(f"db host: cpu avg {entries['host_cpu_avg_pct']}% max {entries['host_cpu_max_pct']}%, "
f"ram peak {entries['host_ram_used_peak_mb']} MB (delta {entries['host_ram_used_delta_mb']} MB)")
print(f"marker: {marker_path}")
return 0
if __name__ == "__main__":
sys.exit(main())

395
convert_rdf.py Normal file
View File

@@ -0,0 +1,395 @@
"""Task 07 - Convert the CSV corpus to RDF.
Streams the corpus into ./out/rdf/dataset.nt.gz (N-Triples, gzip) and
./out/rdf/dataset.ttl (Turtle) with the task 04 vocabulary (tribo# + PROV-O
for started_at), then bulk-loads the N-Triples into an embedded on-disk
oxigraph store (./out/rdf/oxigraph_store/) and smoke-tests it with SPARQL
COUNT queries. Bulk points (cycles, loop points, map points) carry no
rdf:type - the containing predicate types them (documented decision per
spec 07), which is why the triple count lands below the spec's 25-45M
sketch. Appends serialization and store sizes to storage_sizes.csv and
writes ./out/.done/07.ok. Spec: docs/specs/07_convert_rdf.md.
"""
from __future__ import annotations
import gzip
import io
import logging
import shutil
import sys
import time
from pathlib import Path
from pyoxigraph import Store
try:
from pyoxigraph import RdfFormat
NT_FORMAT = RdfFormat.N_TRIPLES
except ImportError: # pyoxigraph < 0.5 uses MIME strings
NT_FORMAT = "application/n-triples"
from common.corpus import CorpusReader, ValidationError
from common.pipeline import (
check_dependencies,
load_lab_config,
parse_task_args,
process_metrics,
remove_stale_marker,
write_marker,
)
from common.storage_sizes import update_storage_sizes
logger = logging.getLogger(__name__)
TASK_ID = "07"
DEPENDS_ON = ["01", "03"]
VOCAB = "https://sandia.gov/ontology/tribology#"
ID_BASE = "https://sandia.gov/ontology/tribology/id/"
XSD = "http://www.w3.org/2001/XMLSchema#"
PREFIXES = {
"tribo": VOCAB,
"batch": ID_BASE + "batch/",
"wafer": ID_BASE + "wafer/",
"coupon": ID_BASE + "coupon/",
"track": ID_BASE + "track/",
"run": ID_BASE + "run/",
"instrument": ID_BASE + "instrument/",
"simtra": ID_BASE + "simtra/",
"prov": "http://www.w3.org/ns/prov#",
"xsd": XSD,
}
INT_COLS = {
"cycle", "grid_x", "grid_y", "indent_id", "plates", "plate", "probe", "square",
"row_no", "wafer_index", "coupons", "friction_coupons", "reserve_coupons",
"pt_gun_tilt_deg", "au_gun_tilt_deg", "pt_power_w", "au_power_w",
"pt_discharge_v", "au_discharge_v",
}
DOUBLE_COLS = {
"cof", "position_um", "friction_force_mn", "pt_wtpct", "au_wtpct", "thickness_um",
"x_um", "y_um", "hardness_gpa", "reduced_modulus_gpa", "max_load_mn", "ra_nm",
"rq_nm", "wear_volume_um3", "k_archard", "sliding_distance_m", "load_mn",
"stroke_mm", "speed_mm_s", "angle_deg", "energy_ev", "pt_flux", "au_flux",
"au_wtpct_mean",
}
DATE_COLS = {"date", "deposition_date"}
MEASUREMENTS = [
("xrf_map.csv", "xrf", "CompositionMeasurement", "m4_tornado"),
("profilometry.csv", "profilometry", "ProfilometryMeasurement", "profilometer"),
("nanoindentation.csv", "nanoindentation", "NanoindentationMeasurement", "ti980"),
]
def lit(col: str, token: str) -> tuple:
if col in INT_COLS:
return ("l", token, "integer")
if col in DOUBLE_COLS:
return ("l", token, "double")
if col in DATE_COLS:
return ("l", token, "date")
if col == "started_at":
return ("l", token, "dateTime")
return ("l", token, None)
def iri(curie: str) -> tuple:
return ("i", curie)
def node(pairs: list) -> tuple:
return ("n", pairs)
class RdfEmitter:
"""Writes the same subject blocks to N-Triples and Turtle streams."""
def __init__(self, nt_fh, ttl_fh) -> None:
self.nt = nt_fh
self.ttl = ttl_fh
self.triples = 0
self._blank = 0
for prefix, base in PREFIXES.items():
self.ttl.write(f"@prefix {prefix}: <{base}> .\n")
self.ttl.write("\n")
def _full(self, curie: str) -> str:
prefix, local = curie.split(":", 1)
return f"<{PREFIXES[prefix]}{local}>"
def _pred_nt(self, pred: str) -> str:
if pred == "a":
return "<http://www.w3.org/1999/02/22-rdf-syntax-ns#type>"
return self._full(pred)
@staticmethod
def _escape(text: str) -> str:
return text.replace("\\", "\\\\").replace('"', '\\"')
def _obj_nt(self, obj: tuple) -> str:
kind = obj[0]
if kind == "i":
return self._full(obj[1])
if kind == "l":
_, token, dtype = obj
if dtype is None:
return f'"{self._escape(token)}"'
return f'"{token}"^^<{XSD}{dtype}>'
self._blank += 1
label = f"_:b{self._blank}"
for pred, sub_obj in obj[1]:
self.nt.write(f"{label} {self._pred_nt(pred)} {self._obj_nt(sub_obj)} .\n")
self.triples += 1
return label
def _obj_ttl(self, obj: tuple) -> str:
kind = obj[0]
if kind == "i":
return obj[1]
if kind == "l":
_, token, dtype = obj
if dtype is None:
return f'"{self._escape(token)}"'
if dtype == "integer":
return token
return f'"{token}"^^xsd:{dtype}'
inner = " ; ".join(f"{p} {self._obj_ttl(o)}" for p, o in obj[1])
return f"[ {inner} ]"
def emit(self, subject: str, pairs: list) -> None:
subj_nt = self._full(subject)
ttl_parts = []
for pred, obj in pairs:
# N-Triples: nested nodes expand to labelled blanks first, so the
# referencing triple is written after the node's own triples.
obj_nt = self._obj_nt(obj)
self.nt.write(f"{subj_nt} {self._pred_nt(pred)} {obj_nt} .\n")
self.triples += 1
ttl_parts.append(f"{pred} {self._obj_ttl(obj)}")
self.ttl.write(subject + " " + " ;\n ".join(ttl_parts) + " .\n")
class RdfConverter:
def __init__(self, cfg: dict, csv_root: Path, emitter: RdfEmitter) -> None:
self.cfg = cfg
self.reader = CorpusReader(csv_root)
self.em = emitter
self.counts = {"coupons": 0, "tracks": 0, "cycle_nodes": 0, "loop_nodes": 0, "xrf_nodes": 0}
def raw_rows(self, relpath: str) -> list[dict[str, str]]:
header, rows = self.reader.read(relpath)
keys = [h.lower() for h in header]
return [dict(zip(keys, row)) for row in rows]
def prop_pairs(self, row: dict[str, str], skip: tuple = ()) -> list:
return [
("prov:startedAtTime" if col == "started_at" else f"tribo:{col}", lit(col, token))
for col, token in row.items()
if col not in skip and token != ""
]
def convert(self) -> None:
for inst in self.cfg["instruments"]:
self.em.emit(f"instrument:{inst['instrument_id']}", [
("a", iri("tribo:Instrument")),
("tribo:name", ("l", inst["name"], None)),
("tribo:role", ("l", inst["role"], None)),
])
for row in self.raw_rows("batches.csv"):
code = row["batch_code"]
self.em.emit(f"batch:{code}", [("a", iri("tribo:Batch"))] + self.prop_pairs(row))
points = [
("tribo:hasPoint", node([("tribo:row_no", ("l", str(no), "integer"))] + self.prop_pairs(p)))
for no, p in enumerate(self.raw_rows(f"simtra/simtra_profile_{code}.csv"), 1)
]
self.em.emit(f"simtra:{code}", [
("a", iri("tribo:SimtraProfile")),
("tribo:performedOn", iri(f"batch:{code}")),
("tribo:performedBy", iri("instrument:simtra")),
] + points)
for row in self.raw_rows("runs.csv"):
self.em.emit(f"run:{row['run_code']}", [("a", iri("tribo:Run"))] + self.prop_pairs(row))
h = self.cfg["hierarchy"]
for batch in self.cfg["deposition_matrix"]:
b_code = batch["batch_code"]
for w in range(1, h["wafers_per_batch"] + 1):
wafer_dir = f"batch_{b_code}/wafer_W{w}"
info = self.raw_rows(f"{wafer_dir}/wafer_info.csv")[0]
self.em.emit(f"wafer:{info['wafer_code']}", [
("a", iri("tribo:Wafer")),
("tribo:partOf", iri(f"batch:{b_code}")),
] + self.prop_pairs(info, skip=("batch_code",)))
for c in range(1, h["coupons_per_wafer"] + 1):
self.convert_coupon(f"{wafer_dir}/coupon_C{c:02d}")
logger.info("batch %s converted (%d triples so far)", b_code, self.em.triples)
def convert_coupon(self, rel_dir: str) -> None:
info = self.raw_rows(f"{rel_dir}/coupon_info.csv")[0]
code = info["coupon_code"]
subject = f"coupon:{code}"
is_friction = info["run_code"] != "RESERVE"
pairs = [("a", iri("tribo:Coupon")), ("tribo:cutFrom", iri(f"wafer:{info['wafer_code']}"))]
pairs += self.prop_pairs(info, skip=("batch_code", "wafer_code", "run_code"))
pairs.append(("tribo:run_code", ("l", info["run_code"], None)))
if is_friction:
pairs.append(("tribo:duringRun", iri(f"run:{info['run_code']}")))
for csv_name, key, cls, inst in MEASUREMENTS:
points = [
("tribo:hasPoint" if key != "nanoindentation" else "tribo:hasIndent", node(self.prop_pairs(p)))
for p in self.raw_rows(f"{rel_dir}/{csv_name}")
]
if key == "xrf":
self.counts["xrf_nodes"] += len(points)
pairs.append((f"tribo:{key}", node([
("a", iri(f"tribo:{cls}")),
("tribo:performedOn", iri(subject)),
("tribo:performedBy", iri(f"instrument:{inst}")),
] + points)))
afm = self.raw_rows(f"{rel_dir}/afm.csv")[0]
pairs.append(("tribo:afm", node([
("a", iri("tribo:AFMMeasurement")),
("tribo:performedOn", iri(subject)),
("tribo:performedBy", iri("instrument:afm")),
] + self.prop_pairs(afm))))
self.em.emit(subject, pairs)
self.counts["coupons"] += 1
if is_friction:
for t in range(1, self.cfg["friction_assignment"]["tracks_per_friction_coupon"] + 1):
self.convert_track(f"{rel_dir}/track_T{t}", subject)
def convert_track(self, rel_dir: str, coupon_subject: str) -> None:
info = self.raw_rows(f"{rel_dir}/track_info.csv")[0]
pairs = [
("a", iri("tribo:FrictionTest")),
("tribo:partOf", iri(coupon_subject)),
("tribo:duringRun", iri(f"run:{info['run_code']}")),
("tribo:performedBy", iri("instrument:rapid")),
] + self.prop_pairs(info, skip=("track_code", "run_code"))
pairs.insert(1, ("tribo:track_code", ("l", info["track_code"], None)))
cycles = self.raw_rows(f"{rel_dir}/cof_vs_cycle.csv")
pairs += [("tribo:hasCycle", node(self.prop_pairs(row))) for row in cycles]
self.counts["cycle_nodes"] += len(cycles)
loops = self.raw_rows(f"{rel_dir}/friction_loops.csv")
pairs += [("tribo:hasLoopPoint", node(self.prop_pairs(row))) for row in loops]
self.counts["loop_nodes"] += len(loops)
wear = self.raw_rows(f"{rel_dir}/wear.csv")[0]
pairs.append(("tribo:wear", node([("a", iri("tribo:WearMeasurement"))] + self.prop_pairs(wear))))
self.em.emit(f"track:{info['track_code']}", pairs)
self.counts["tracks"] += 1
def validate_counts(self) -> None:
v = self.cfg["volumes"]
expected = {
"coupons": v["coupons_total"],
"tracks": v["tracks_total"],
"cycle_nodes": v["cycle_rows_total"],
"loop_nodes": v["loop_points_total"],
"xrf_nodes": v["xrf_points_total"],
}
for key, exp in expected.items():
if self.counts[key] != exp:
raise ValidationError(f"count mismatch: {key}: converted {self.counts[key]} != expected {exp}")
logger.info("all converted counts match lab_config volumes: %s", self.counts)
def bulk_load(store: Store, nt_gz: Path) -> None:
with gzip.open(nt_gz, "rb") as fh:
try:
store.bulk_load(fh, NT_FORMAT)
except TypeError:
store.bulk_load(input=fh, format=NT_FORMAT)
def sparql_count(store: Store, query: str) -> int:
solutions = store.query(query)
return int(next(iter(solutions))[0].value)
def main() -> int:
args = parse_task_args("Task 07: convert the CSV corpus to RDF")
out_root: Path = args.out_root
rdf_root = out_root / "rdf"
nt_gz = rdf_root / "dataset.nt.gz"
ttl_path = rdf_root / "dataset.ttl"
store_dir = rdf_root / "oxigraph_store"
try:
check_dependencies(out_root, DEPENDS_ON)
remove_stale_marker(out_root, TASK_ID)
if rdf_root.exists():
logger.info("re-run: removing previous output %s", rdf_root)
shutil.rmtree(rdf_root)
rdf_root.mkdir(parents=True)
cfg = load_lab_config(out_root)
# mtime=0 keeps the gzip byte-identical across re-runs (data-determinism)
with open(nt_gz, "wb") as raw:
gz = gzip.GzipFile(filename="", mode="wb", fileobj=raw, compresslevel=6, mtime=0)
with io.TextIOWrapper(gz, encoding="ascii", newline="\n") as nt_fh, \
open(ttl_path, "w", encoding="ascii", newline="\n") as ttl_fh:
emitter = RdfEmitter(nt_fh, ttl_fh)
conv = RdfConverter(cfg, out_root / "csv", emitter)
conv.convert()
conv.validate_counts()
triples = emitter.triples
nt_bytes = nt_gz.stat().st_size
ttl_bytes = ttl_path.stat().st_size
logger.info("serialized %d triples: nt.gz %.1f MiB, ttl %.1f MiB",
triples, nt_bytes / 1048576, ttl_bytes / 1048576)
t0 = time.time()
store = Store(str(store_dir))
bulk_load(store, nt_gz)
store.flush()
load_seconds = round(time.time() - t0, 1)
total = sparql_count(store, "SELECT (COUNT(*) AS ?n) WHERE { ?s ?p ?o }")
if total != triples:
raise ValidationError(f"store holds {total} triples, serialized {triples}")
q1 = sparql_count(
store,
f"SELECT (COUNT(?c) AS ?n) WHERE {{ <{ID_BASE}track/B722-W2-C13-T2> <{VOCAB}hasCycle> ?c }}",
)
if q1 != cfg["volumes"]["cycles_per_track"]:
raise ValidationError(f"SPARQL Q1 smoke test returned {q1} cycles")
logger.info("store smoke tests ok: %d triples, Q1 track has %d cycles (load %.1f s)", total, q1, load_seconds)
del store
store_bytes = sum(f.stat().st_size for f in store_dir.rglob("*") if f.is_file())
update_storage_sizes(out_root / "bench", "rdf", [
("nt_gz", nt_bytes), ("ttl", ttl_bytes), ("store", store_bytes),
])
entries = {
"triples": triples,
"nt_gz_bytes": nt_bytes,
"ttl_bytes": ttl_bytes,
"store_bytes": store_bytes,
"store_load_seconds": load_seconds,
"coupons": conv.counts["coupons"],
"tracks": conv.counts["tracks"],
"cycle_nodes": conv.counts["cycle_nodes"],
"loop_nodes": conv.counts["loop_nodes"],
}
entries.update(process_metrics())
marker_path = write_marker(out_root, TASK_ID, entries)
except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1
print(f"task 07 ok: {triples} triples; nt.gz {nt_bytes / 1048576:.1f} MiB, "
f"ttl {ttl_bytes / 1048576:.1f} MiB, store {store_bytes / 1048576:.1f} MiB "
f"(bulk load {load_seconds} s)")
print(f"entities: coupons={conv.counts['coupons']} tracks={conv.counts['tracks']} "
f"cycles={conv.counts['cycle_nodes']} loop_points={conv.counts['loop_nodes']}")
print(f"marker: {marker_path}")
return 0
if __name__ == "__main__":
sys.exit(main())

286
convert_sqlite.py Normal file
View File

@@ -0,0 +1,286 @@
"""Task 05 - Convert the CSV corpus to SQLite.
Builds ./out/sqlite/tribo.db: variant B key schema (composite natural PKs
on bulk tables, WITHOUT ROWID, enforced foreign keys - see
docs/research/Tribology_FK_Architecture.html), Q1-Q7 indexes created after
the load, track_summary precomputed. Every consumed CSV is verified
against MANIFEST.csv. Appends the db size to storage_sizes.csv and writes
./out/.done/05.ok. Spec: docs/specs/05_convert_sqlite.md.
"""
from __future__ import annotations
import logging
import sqlite3
import sys
from pathlib import Path
from common.corpus import CorpusReader, ValidationError
from common.pipeline import (
check_dependencies,
load_lab_config,
parse_task_args,
process_metrics,
remove_stale_marker,
write_marker,
)
from common.relational import COLUMNS, TABLES, expected_counts, stream_rows
from common.storage_sizes import update_storage_sizes
logger = logging.getLogger(__name__)
TASK_ID = "05"
DEPENDS_ON = ["01", "03"]
BATCH_ROWS = 50000 # rows per executemany transaction, spec 05
LOAD_PRAGMAS = [
"PRAGMA journal_mode=OFF",
"PRAGMA synchronous=OFF",
"PRAGMA cache_size=-262144",
"PRAGMA foreign_keys=ON",
]
DDL = [
"""CREATE TABLE instruments (
instrument_id INTEGER PRIMARY KEY,
instrument_code TEXT NOT NULL UNIQUE,
name TEXT NOT NULL,
role TEXT NOT NULL
)""",
"""CREATE TABLE batches (
batch_id INTEGER PRIMARY KEY,
batch_code TEXT NOT NULL UNIQUE,
pt_gun_tilt_deg INTEGER NOT NULL,
au_gun_tilt_deg INTEGER NOT NULL,
pt_power_w INTEGER NOT NULL,
au_power_w INTEGER NOT NULL,
pt_discharge_v INTEGER NOT NULL,
au_discharge_v INTEGER NOT NULL,
deposition_date TEXT NOT NULL
)""",
"""CREATE TABLE runs (
run_id INTEGER PRIMARY KEY,
run_code TEXT NOT NULL UNIQUE,
environment TEXT NOT NULL,
date TEXT NOT NULL,
plates INTEGER NOT NULL,
operator TEXT NOT NULL
)""",
"""CREATE TABLE wafers (
wafer_id INTEGER PRIMARY KEY,
wafer_code TEXT NOT NULL UNIQUE,
batch_id INTEGER NOT NULL REFERENCES batches(batch_id),
wafer_index INTEGER NOT NULL,
deposition_date TEXT NOT NULL,
coupons INTEGER NOT NULL,
friction_coupons INTEGER NOT NULL,
reserve_coupons INTEGER NOT NULL
)""",
"""CREATE TABLE coupons (
coupon_id INTEGER PRIMARY KEY,
coupon_code TEXT NOT NULL UNIQUE,
wafer_id INTEGER NOT NULL REFERENCES wafers(wafer_id),
batch_id INTEGER NOT NULL REFERENCES batches(batch_id),
grid_row INTEGER NOT NULL,
grid_col INTEGER NOT NULL,
thickness_um REAL NOT NULL,
ra_nm REAL NOT NULL,
au_wtpct_mean REAL NOT NULL,
run_id INTEGER REFERENCES runs(run_id),
plate INTEGER,
probe INTEGER,
square INTEGER
)""",
"""CREATE TABLE tracks (
track_id INTEGER PRIMARY KEY,
track_code TEXT NOT NULL UNIQUE,
coupon_id INTEGER NOT NULL REFERENCES coupons(coupon_id),
run_id INTEGER NOT NULL REFERENCES runs(run_id),
environment TEXT NOT NULL,
load_mn REAL NOT NULL,
stroke_mm REAL NOT NULL,
speed_mm_s REAL NOT NULL,
counterface_id TEXT NOT NULL,
started_at TEXT NOT NULL
)""",
"""CREATE TABLE simtra_profiles (
batch_id INTEGER NOT NULL REFERENCES batches(batch_id),
row_no INTEGER NOT NULL,
angle_deg REAL NOT NULL,
energy_ev REAL NOT NULL,
pt_flux REAL NOT NULL,
au_flux REAL NOT NULL,
PRIMARY KEY (batch_id, row_no)
) WITHOUT ROWID""",
"""CREATE TABLE xrf_points (
coupon_id INTEGER NOT NULL REFERENCES coupons(coupon_id),
grid_x INTEGER NOT NULL,
grid_y INTEGER NOT NULL,
pt_wtpct REAL NOT NULL,
au_wtpct REAL NOT NULL,
PRIMARY KEY (coupon_id, grid_x, grid_y)
) WITHOUT ROWID""",
"""CREATE TABLE profilometry_points (
coupon_id INTEGER NOT NULL REFERENCES coupons(coupon_id),
grid_x INTEGER NOT NULL,
grid_y INTEGER NOT NULL,
thickness_um REAL NOT NULL,
PRIMARY KEY (coupon_id, grid_x, grid_y)
) WITHOUT ROWID""",
"""CREATE TABLE nanoindentation (
coupon_id INTEGER NOT NULL REFERENCES coupons(coupon_id),
indent_id INTEGER NOT NULL,
x_um REAL NOT NULL,
y_um REAL NOT NULL,
hardness_gpa REAL NOT NULL,
reduced_modulus_gpa REAL NOT NULL,
max_load_mn REAL NOT NULL,
PRIMARY KEY (coupon_id, indent_id)
) WITHOUT ROWID""",
"""CREATE TABLE afm (
coupon_id INTEGER PRIMARY KEY REFERENCES coupons(coupon_id),
ra_nm REAL NOT NULL,
rq_nm REAL NOT NULL,
image_file TEXT NOT NULL
)""",
"""CREATE TABLE friction_cycles (
track_id INTEGER NOT NULL REFERENCES tracks(track_id),
cycle INTEGER NOT NULL,
cof REAL NOT NULL,
PRIMARY KEY (track_id, cycle)
) WITHOUT ROWID""",
"""CREATE TABLE friction_loop_points (
track_id INTEGER NOT NULL REFERENCES tracks(track_id),
cycle INTEGER NOT NULL,
pt INTEGER NOT NULL,
position_um REAL NOT NULL,
friction_force_mn REAL NOT NULL,
PRIMARY KEY (track_id, cycle, pt)
) WITHOUT ROWID""",
"""CREATE TABLE wear (
track_id INTEGER PRIMARY KEY REFERENCES tracks(track_id),
wear_volume_um3 REAL NOT NULL,
k_archard REAL NOT NULL,
sliding_distance_m REAL NOT NULL
)""",
"""CREATE TABLE track_summary (
track_id INTEGER PRIMARY KEY REFERENCES tracks(track_id),
cof_ss_mean REAL NOT NULL,
cof_ss_std REAL NOT NULL,
run_in_cycles INTEGER NOT NULL
)""",
]
INDEX_DDL = [
"CREATE INDEX ix_coupons_au_wtpct_mean ON coupons(au_wtpct_mean)", # Q2 selective filter
"CREATE INDEX ix_tracks_run_id ON tracks(run_id)", # Q5/Q7 joins to runs
"CREATE INDEX ix_tracks_coupon_id ON tracks(coupon_id)", # Q3/Q6 joins to coupons
"CREATE INDEX ix_runs_environment ON runs(environment)", # Q7 grouping
"CREATE INDEX ix_tracks_environment_load ON tracks(environment, load_mn)", # Q4 filter
]
class SqliteLoader:
def __init__(self, db_path: Path) -> None:
db_path.parent.mkdir(parents=True, exist_ok=True)
self.conn = sqlite3.connect(db_path)
for pragma in LOAD_PRAGMAS:
self.conn.execute(pragma)
for ddl in DDL:
self.conn.execute(ddl)
self.inserts = {
t: f"INSERT INTO {t} VALUES ({','.join('?' * len(COLUMNS[t]))})" for t in TABLES
}
self.buffers: dict[str, list[tuple]] = {t: [] for t in TABLES}
def load(self, cfg: dict, reader: CorpusReader) -> None:
pending = 0
for table, row in stream_rows(cfg, reader):
self.buffers[table].append(row)
pending += 1
if pending >= BATCH_ROWS:
self.flush_all()
pending = 0
self.flush_all()
def flush_all(self) -> None:
# TABLES is FK-dependency ordered: parents flush before children.
for table in TABLES:
buf = self.buffers[table]
if buf:
self.conn.executemany(self.inserts[table], buf)
buf.clear()
self.conn.commit()
def finalize(self) -> None:
for ddl in INDEX_DDL:
self.conn.execute(ddl)
self.conn.commit()
self.conn.execute("PRAGMA journal_mode=WAL")
self.conn.execute("ANALYZE")
self.conn.execute("VACUUM")
logger.info("indexes created, WAL enabled, ANALYZE + VACUUM done")
def validate(self, cfg: dict) -> dict[str, int]:
counts: dict[str, int] = {}
for table, exp in expected_counts(cfg).items():
got = self.conn.execute(f"SELECT COUNT(*) FROM {table}").fetchone()[0]
counts[table] = got
if got != exp:
raise ValidationError(f"row count mismatch: {table}: loaded {got} != expected {exp}")
reserve = self.conn.execute("SELECT COUNT(*) FROM coupons WHERE run_id IS NULL").fetchone()[0]
if reserve != cfg["friction_assignment"]["reserve_coupons_total"]:
raise ValidationError(f"reserve coupons: {reserve} != expected")
fk_violations = self.conn.execute("PRAGMA foreign_key_check").fetchall()
if fk_violations:
raise ValidationError(f"foreign_key_check reported {len(fk_violations)} violations")
logger.info("all %d table counts match, %d reserve coupons, foreign_key_check clean", len(counts), reserve)
return counts
def main() -> int:
args = parse_task_args("Task 05: convert the CSV corpus to SQLite")
out_root: Path = args.out_root
db_path = out_root / "sqlite" / "tribo.db"
try:
check_dependencies(out_root, DEPENDS_ON)
remove_stale_marker(out_root, TASK_ID)
for stale in (db_path, db_path.with_suffix(".db-wal"), db_path.with_suffix(".db-shm")):
if stale.exists():
logger.info("re-run: removing %s", stale)
stale.unlink()
cfg = load_lab_config(out_root)
loader = SqliteLoader(db_path)
loader.load(cfg, CorpusReader(out_root / "csv"))
loader.finalize()
counts = loader.validate(cfg)
loader.conn.close()
db_bytes = db_path.stat().st_size
update_storage_sizes(out_root / "bench", "sqlite", [("db", db_bytes)])
entries = {
"db_file": db_path.as_posix(),
"db_bytes": db_bytes,
"tables": len(counts),
"total_rows": sum(counts.values()),
"friction_cycles": counts["friction_cycles"],
"friction_loop_points": counts["friction_loop_points"],
"tracks": counts["tracks"],
"track_summary": counts["track_summary"],
}
entries.update(process_metrics())
marker_path = write_marker(out_root, TASK_ID, entries)
except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1
print(f"task 05 ok: {db_path} ({db_bytes / (1024 * 1024):.1f} MiB, {len(counts)} tables, {sum(counts.values())} rows)")
print(f"bulk rows: cycles={counts['friction_cycles']} loop_points={counts['friction_loop_points']} "
f"xrf={counts['xrf_points']} track_summary={counts['track_summary']}")
print(f"marker: {marker_path}")
return 0
if __name__ == "__main__":
sys.exit(main())

90
docs/Initial_Prompt.md Normal file
View File

@@ -0,0 +1,90 @@
# Prompt: Generate a Multi-File Specification for a Tribology Data Storage Evaluation Pipeline
You are an AI assistant generating a specification for execution by Anthropic Claude Code
in a Cursor development environment (local filesystem, on-premise only, no cloud).
## Objective
Produce 12 Markdown files — `00_PLAN.md` plus task files `01``11` — that specify a complete,
reproducible pipeline: simulate a tribology laboratory, generate ~150 MB of realistic
measurement data in CSV, convert it into 4 target storage formats, benchmark 7 retrieval
scenarios across all 5 formats with resource metering, extrapolate results to 600 GB+,
and produce a decision report with tables, charts, and weighted scoring.
## File structure requirements
Each task file must contain: Context/Goal, Dependencies, Input, Processing instructions,
Output (artifacts + completion marker `./out/.done/<task>.ok`).
`00_PLAN.md` must contain: goals, execution environment, a Mermaid dependency flowchart,
parallelism notes, execution order, and global conventions (ISO-8601 UTC timestamps,
SI-explicit column names, fixed random seed 20260711, Python with tabs for indentation).
## Laboratory configuration (task 01)
- Material: Ti-6Al-4V coupons 10×10×3 mm, Cr adhesion layer, Pt-Au coating sputtered in a
composition gradient; film thickness 0.31.1 µm.
- Instruments: RAPID custom high-throughput parallelized 6-probe tribometer (friction);
Bruker TI980 TriboIndenter (nanoindentation); Bruker M4 Tornado micro-XRF (composition
mapping); Kurt J. Lesker PVD 200 sputter-down (deposition); AFM (roughness); optical
profilometry (thickness); SIMTRA (sputter-transport Monte-Carlo simulation).
- Deposition matrix, 4 batches (B721B724) with gun tilt (20°:0°, 20°:20°, 20°:20°, 0°:20°),
Pt/Au power (150/50, 100/100, 150/50, 50/150 W) and discharge voltages.
- Hierarchy and volumes: batch(×4) → wafer(×3) → coupon(×49, 7×7 grid) = 588 coupons;
480 friction coupons (4 RAPID runs × 5 plates × 6 probes × 4 coupons/square),
108 reserve; 2 environments (Lab Air / Dry N2, 2 runs each); 3 replicate tracks per
coupon with counterface rotation → 1,440 tracks × 1,000 cycles = 1.44 M cycle rows;
friction loops (200 points every 100th cycle) → 2.88 M loop points.
- Physical models for simulation: exponential COF run-in with steady-state depending on
Au% and environment; hardness/modulus linear in Au% with noise; log-normal roughness;
Archard wear; Au gradient across wafer position (batch-dependent center: 8/25/10/60 wt%).
## Process flow diagrams (task 02) — Mermaid
Coupon assembly (clean → sonicate → Cr adhesive → PVD sputter); characterization &
batch assembly (SIMTRA, profilometry, ×49/wafer, ×3/batch, batches 721724); testing tree
(friction µ_normal/µ_dry_nit, nanoindentation, AFM, XRF); tribometer session sequence
(request → samples → ball holders, 3 counterfaces/holder → Excel test plan → setup →
run → software stop → save avg files → teardown); execution loop (5 plates × 6 probes ×
4 coupons × 3 tracks with counterface rotation); data-hierarchy ERD.
## Data generation (task 03)
Hierarchical CSV file tree mirroring the lab structure (batches, simtra, per-coupon
xrf_map/profilometry/nanoindentation/afm, per-track track_info/cof_vs_cycle/
friction_loops/wear), streaming generation, MANIFEST.csv with row/byte/sha256,
size guard 120200 MB.
## Conversions (tasks 0407), each recording sizes into storage_sizes.csv
- JSON-LD: vocab + QUDT + PROV-O; two variants — FULL (bulk embedded, per-coupon files,
compact) and HYBRID (metadata only, sourceFile links to CSV).
- SQLite: single file, integer surrogate keys, WITHOUT ROWID bulk tables, precomputed
track_summary, indexes for the benchmark queries, load PRAGMAs, VACUUM/ANALYZE.
- PostgreSQL 16: same logical schema, RANGE partitioning of bulk tables, COPY loading,
b-tree + BRIN indexes, materialized track_summary, pg_dump archive.
- RDF: N-Triples.gz + Turtle serializations; queryable embedded store (oxigraph preferred,
Fuseki/TDB2 fallback); document triple-modeling decisions; expect 2545 M triples.
## Benchmarks (tasks 0809)
Seven scenarios: Q1 single-track COF curve; Q2 steady-state COF at Au=10±0.5 wt%;
Q3 hardness vs COF join; Q4 condition filter (Dry N2, 100 mN, cof>0.20); Q5 mean run-in
per batch aggregating ALL raw cycles (summary tables forbidden); Q6 wear volume vs load
join; Q7 Stribeck-style grouped aggregation. Implement per format (direct path read /
ijson streaming / SQL / SQL+EXPLAIN / SPARQL); identical result sets validated by checksum.
Harness: each cell in a separate subprocess; psutil metering (wall time, peak RSS, CPU
time/utilization, read bytes); 1 warm-up + 3 measured runs, medians; 30-min timeout;
outputs results_raw.csv / results_median.csv.
## Extrapolation (task 10)
Project storage linearly and query time by access-pattern complexity laws
(O(1)/O(log n)/O(n)/single-threaded vs parallel partitioned scan) from the 150 MB
calibration point to 600 GB, 1.2 TB, 6 TB; flag >1 h IMPRACTICAL, >24 h FAIL; RAM models
per format; hardware sizing and cost table driven by an editable hw_prices.yaml
(June 2026 defaults: DDR5 RDIMM ~$14/GB, enterprise NVMe ~$250/TB, 16-core 1U ~$6k).
## Reporting (task 11)
Tables (footprint + coefficients vs CSV, measured metrics, projected times with flags,
hardware costs, scoring matrix); matplotlib charts (log scales): footprint, RAM,
per-query bars, degradation curves 150 MB→6 TB, stacked weighted score, cost-vs-performance
scatter. Weighted scoring: search speed ×5 (050), RAM economy ×2 (020), disk economy
×1 (010), max 80; search score from inverse geometric mean of Q1Q7 projected at 600 GB.
Concise English narrative: executive summary and use-case mapping (analysis, reports,
search/filter, archiving, inter-lab exchange), limitations.
## Style
Specification prose in English, concise and directive. No implementation code in the
spec — only precise instructions an autonomous coding agent can execute.

View File

@@ -1,248 +0,0 @@
# Tribology Lab Data — Master Plan
**Discovery · Cleansing/Normalization · Ontology · Conversion**
Project: LDRD Mechanochemical Alloys FY25-27 — Sandia National Laboratories, Tribology Facility
Scope: ~700 GB raw CSV → normalized, self-documenting, multi-format store
Target of record: PostgreSQL (per ADR v1.0); deferred virtual RDF (Ontop OBDA) for FAIR federation
Status: Plan v1 — Draft
---
## 0. Objective
Transform ~700 GB of heterogeneous instrument-produced CSV into a formally modeled, cleansed, and normalized dataset with full source→target lineage, then generate reproducible exports (JSON, SQLite, PostgreSQL, RDF/triplestore) from a single canonical intermediate.
The existing artifacts (Business Description, ERD, ADR) already define the **target domain model**. This plan governs how the raw filesystem is reconciled *against* that model and materialized into the export formats.
---
## 1. Guiding principles
1. **Non-destructive.** Source CSV is read-only. All work products are derived; the raw tree is never mutated.
2. **Lineage-first.** Every target row carries provenance back to `(source_file, byte/row offset, transform rules applied)`. Lineage is a build-time output, not an afterthought.
3. **Phases 2 and 3 are iterative and co-dependent.** A draft ontology guides cleansing; findings in the data refine the ontology. Neither is frozen until both stabilize.
4. **Instruments Repository is built, not referenced.** No external etalon exists. Instruments and their characteristics accrete as an ontology sub-model as the data reveals them; instrument ranges later become anomaly criteria.
5. **Pure Python first.** Lightweight streaming pipeline, no DB engine during Phases 12. Migration to JSX Viewer / better-sqlite3 is deferred.
6. **Canonical intermediate.** All exports derive from one normalized representation, so formats never diverge.
7. **Reproducibility.** The pipeline is deterministic and re-runnable end-to-end from source + config; no manual editing of intermediates.
---
## 2. Phase overview
| Phase | Name | Nature | Primary output |
|-------|------|--------|----------------|
| 1 | Storage Discovery & Profiling | Non-destructive scan | Storage manifest + schema fingerprints |
| 2 ↔ 3 | Cleansing/Normalization ↔ Ontology | Iterative loop | Cleansed canonical data + relational schema + data dictionary + Instruments Repository |
| 4 | Conversion & Export | Deterministic build | JSON / SQLite / PostgreSQL / RDF outputs |
Cross-cutting throughout: **lineage**, **anomaly detection**, **configuration/rule registry**.
---
## 3. Phase 1 — Storage Discovery & Non-destructive Profiling
**Goal.** Reconstruct the actual (as-is) and intended (as-designed) structure of the filesystem and record structures, without loading full file bodies.
### 3.1 Inputs
- Root of the ~700 GB CSV tree (read-only).
### 3.2 Activities
**A. Filesystem inventory.** Walk the tree; for each file capture: full path, relative path components (folder semantics), filename, size, mtime, extension. Reconstruct the folder-level meaning (project → batch → coupon → sample, per the known hierarchy).
**B. Filename tokenization.** Parse filenames into structured tokens against the known conventions:
- coupon-level (no sample token) → XRF (`L_081423_1`), macro friction (`data1..data3`)
- sample-level (sample token `456/457`) → nanoindentation (`run1..run9`), friction test (`test_N`) / cycle (`cycleN` ambient, `_N` dry nitrogen)
- atmosphere tokens (`ambient`, `dryNitrogen`, …)
Produce a token-extraction report and a list of filenames that **fail** to match any known pattern (candidates for new rules or anomalies).
**C. Structural profiling (header-only + sampled body).** Per file, without full load:
- read header row → column names, count, order;
- infer per-column datatype from first/last N rows + a random sample;
- compute a **schema fingerprint** (normalized hash of the ordered column set) to group files of identical structure.
**D. Schema grouping.** Cluster files by (measurement type inferred from path/name) × (schema fingerprint). This surfaces the core Phase-2 problem: files of the *same purpose* with *different* column sets/naming.
### 3.3 Algorithm sketch (streaming, O(files), bounded memory)
```
for each file in walk(root):
meta = fs_metadata(file)
tokens = tokenize_filename(file) # -> measurement type, ids, atmosphere
header = read_first_line(file)
columns = parse_columns(header)
dtypes = infer_types(sample_rows(file, head=N, tail=N, random=M))
fp = fingerprint(columns) # ordered-set hash
emit_manifest_row(meta, tokens, columns, dtypes, fp)
```
### 3.4 Artifacts
- **`storage_manifest`** — machine-readable catalog (JSON lines + SQLite index) of every file with metadata, tokens, columns, dtypes, fingerprint.
- **`schema_fingerprint_report`** — fingerprint → {file count, member files, canonical column set} per measurement type.
- **`unmatched_report`** — files/filenames not matching known folder or naming rules.
- **`storage_map`** — human-readable reconstruction of folder/file purpose (as-is vs as-designed).
### 3.5 Exit criteria
- 100% of files inventoried and fingerprinted.
- Every file assigned a (measurement type, attachment level) or flagged as unmatched.
- Fingerprint clusters reviewed; the set of distinct schemas per measurement type is enumerated.
---
## 4. Phase 2 ↔ 3 — Cleansing/Normalization & Ontology (iterative)
**Goal.** Reconcile the enumerated raw schemas into the canonical relational model while building the ontology and Instruments Repository. Detect and quarantine anomalies.
The two phases run as a loop over each measurement type (XRF → nanoindentation → friction → macro friction), since each has its own schema and granularity.
### 4.1 The iteration loop (per measurement type)
1. **Draft canonical schema** from the target ERD (e.g., `friction_data_points`, `xrf_spectrum_points`).
2. **Map raw columns → canonical columns.** Build a column-mapping table per fingerprint: raw name/position → canonical name, unit, dtype. Resolve synonyms, unit differences, column reordering.
3. **Cleanse & normalize** on a streamed pass: apply mapping, coerce types, normalize units, standardize atmosphere/instrument codes, derive keys from filename tokens (FK resolution).
4. **Detect anomalies** (see 4.3); route offending rows/files to quarantine with reason codes.
5. **Feed findings back to the ontology** — new instrument, new column, new atmosphere value, unit variant, or an entity/attribute the ERD does not yet cover → update schema + data dictionary + Instruments Repository.
6. **Repeat** until the mapping covers all fingerprints of that type with no unresolved columns.
### 4.2 Instruments Repository (ontology sub-model)
Built incrementally as instruments are identified from path/name + column signatures. Each instrument entity carries:
- identity (type, model/label, measurement type it produces);
- measured quantities, units, and **valid physical ranges** (min/max, resolution);
- association to `measurement_types` / `column_definitions` in the data dictionary.
Instrument ranges become the source of "physically impossible value" criteria in anomaly detection — closing the loop between the repository and cleansing.
### 4.3 Anomaly detection
Criteria are **derived from the data**, not fixed a priori. Starting catalog:
| Class | Description | Source of criterion |
|-------|-------------|---------------------|
| Out-of-range | Value outside instrument's physical capability | Instruments Repository ranges |
| Sequence gap | Missing steps in ordered series (time, cycle, run) | Series continuity check |
| Structural mismatch | File of a purpose with unexpected column set | Fingerprint vs canonical |
| Type violation | Non-coercible value in a typed column | Mapping/coercion |
| Duplicate | Repeated rows / repeated identity keys | Key + row hashing |
| Missing required | Null in a non-nullable canonical field | Schema constraints |
| Encoding/format | Malformed rows, delimiter/decimal issues | Parser |
Anomalies are **quarantined, not dropped**: written to a rejects store with `(source_file, row_offset, reason_code, raw_payload)` for review and rule refinement.
### 4.4 Data dictionary / metadata catalog
Maintained in lockstep with the schema: `measurement_types` and `column_definitions` (column name, unit, data type, ordinal, description) — keeping heterogeneous raw series self-documenting, exactly as specified in the domain model.
### 4.5 Artifacts
- **`column_mapping`** — per (measurement type, fingerprint): raw→canonical mapping with unit/type rules.
- **`canonical_dataset`** — cleansed, normalized rows in the canonical intermediate (partitioned by measurement type; see §7).
- **`relational_schema`** — evolving DDL (PostgreSQL dialect) matching the ERD, versioned.
- **`data_dictionary`** — `measurement_types` + `column_definitions`.
- **`instruments_repository`** — instrument entities + characteristics + ranges.
- **`rejects_store`** — quarantined anomalies with reason codes.
- **`ontology_notes`** — running record of ERD changes and rationale.
### 4.6 Exit criteria
- Every fingerprint of every measurement type has a complete, unambiguous mapping.
- Canonical dataset produced for all in-scope files; rejects quarantined and categorized.
- Relational schema, data dictionary, and Instruments Repository mutually consistent and frozen for the release.
---
## 5. Phase 4 — Conversion & Export
**Goal.** Deterministically materialize the frozen canonical dataset into the required output formats. All formats derive from the same canonical intermediate, so they cannot diverge.
### 5.1 Targets
1. **JSON mega-file** — single hierarchical document (data note → … → measurements → points). Suitable for small/whole-dataset transport; streamed writer for size.
2. **Multiple JSON files** — split by natural boundary (per coupon / per sample / per measurement type). Directory layout mirrors the ontology hierarchy.
3. **SQLite** — schema generated from the relational model; bulk-loaded. Portable single-file store (aligns with future better-sqlite3 / JSX Viewer use).
4. **PostgreSQL** — system of record. DDL + `COPY`-based bulk load (parse tokens → FKs, insert headers, bulk-load numeric rows), per the ADR ingestion path.
5. **RDF / triplestore****metadata & provenance graph only** (never the bulk numeric series), per ADR. Preferred: virtual RDF via Ontop (SPARQL→SQL over PostgreSQL). Optional materialized export via R2RML/RML to Jena/GraphDB. Vocabularies: PROV-O, QUDT, Dublin Core, DCAT, materials/tribology terms; mint stable IRIs.
### 5.2 Design
- One **export driver** reads the canonical intermediate + schema and dispatches to per-format writers.
- Each writer is streaming and idempotent; re-running reproduces byte-comparable output (modulo timestamps).
- Lineage is emitted alongside every export (see §6).
### 5.3 Artifacts
- Per-format output sets + a **build manifest** (what was produced, from which canonical snapshot, with which config/rule versions).
### 5.4 Exit criteria
- All five target formats generated from one canonical snapshot.
- Round-trip validation: record counts and key aggregates match across formats.
- Lineage complete and resolvable for sampled records.
---
## 6. Cross-cutting concerns
### 6.1 Lineage & processing-flow documentation
- Per target row: `source_file`, source row offset, `source_modified_at`, ordered list of transform rules applied.
- Per pipeline run: config snapshot, rule/version set, input inventory hash, output manifest — a reproducible **processing-flow record**.
- Lineage is a first-class export, present in every target format.
### 6.2 Configuration & rule registry
- Filename-token patterns, column mappings, unit conversions, atmosphere/instrument code normalization, and anomaly thresholds live in **versioned config**, not code. This is what makes Phases 2↔3 iterative without rewrites.
### 6.3 Validation gates
- Between phases: schema conformance, referential completeness (all FKs resolvable), reject-rate thresholds, fingerprint coverage = 100%.
---
## 7. Target artifacts catalog (consolidated)
| Artifact | Phase | Form |
|----------|-------|------|
| storage_manifest | 1 | JSONL + SQLite |
| schema_fingerprint_report | 1 | JSON/CSV |
| unmatched_report | 1 | CSV |
| storage_map | 1 | Markdown |
| column_mapping | 2↔3 | CSV/JSON (versioned) |
| canonical_dataset | 2↔3 | partitioned intermediate |
| relational_schema (DDL) | 2↔3 | SQL |
| data_dictionary | 2↔3 | table/CSV |
| instruments_repository | 2↔3 | table/CSV |
| rejects_store | 2↔3 | JSONL |
| ontology_notes | 2↔3 | Markdown |
| exports (JSON/SQLite/PG/RDF) | 4 | per format |
| lineage + build manifest | all | JSONL |
---
## 8. Tooling & environment
- **Language:** Python (streaming, stdlib-first; `csv`, `pathlib`, `hashlib`, `json`, `sqlite3`). Add `pyarrow`/Parquet only if the canonical intermediate needs columnar efficiency at 700 GB.
- **No DB engine in Phases 12** beyond SQLite for the manifest index.
- **PostgreSQL** materialized in Phase 4 (system of record).
- **Ontop / R2RML** for the RDF surface (Phase 4, deferred activation).
- **Future:** migration of browse/query surface to JSX Viewer + better-sqlite3; canonical SQLite export is the natural bridge.
- Python indentation: **tabs**.
---
## 9. Performance strategy for 700 GB
- **Never full-load.** Phase 1 is header + sampled-row only; Phases 2/4 stream row-by-row.
- **Bounded memory.** Constant-memory transforms; no whole-file DataFrames on large series.
- **Partition the canonical intermediate** by measurement type and, for the high-volume series (`friction_data_points`), by a coarse key (e.g., test/cycle) — informed by volume estimates gathered in Phase 1.
- **Parallelize by file** (embarrassingly parallel scan/transform); aggregate manifests after.
- **Deterministic ordering** for reproducible, resumable runs (checkpoint by manifest offset).
- Volume estimates from Phase 1 drive the `friction_data_points` partitioning decision before Phase 4.
---
## 10. Open items & risks
1. **`macro_friction_points` schema is provisional** (columns `normal_load_mn`, `average_coefficient_of_friction` pending an actual macro-friction file). Resolve during Phase 2 mapping of that type.
2. **Unknown fingerprints.** True count of distinct raw schemas per type is unknown until Phase 1 completes; mapping effort in Phase 2 scales with it.
3. **Anomaly criteria maturity.** Instrument ranges (and thus out-of-range detection) are only as complete as the Instruments Repository, which is itself being built — expect multiple loop iterations.
4. **Atmosphere/instrument code drift.** Filename conventions vary (`cycleN` vs `_N`, atmosphere tokens); normalization rules must be discovered and versioned.
5. **Lineage granularity vs. size.** Row-level lineage on billion-row series must be compact (offsets/rule-ids, not copied payloads) to avoid inflating storage.
---
*End of Plan v1.*

File diff suppressed because one or more lines are too long

View File

@@ -0,0 +1,571 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Foreign-Key Architecture Study - Tribology Lab Data Storage</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=IBM+Plex+Mono:wght@400;500;600&family=IBM+Plex+Sans:wght@400;500;600;700&display=swap" rel="stylesheet">
<style>
*{box-sizing:border-box}
html,body{margin:0;padding:0}
body{font-family:"IBM Plex Sans",system-ui,sans-serif;color:#16181d;background:#eef0f2;-webkit-font-smoothing:antialiased}
::-webkit-scrollbar{width:10px;height:10px}
::-webkit-scrollbar-thumb{background:#cfd4da;border-radius:6px;border:2px solid #eef0f2}
::-webkit-scrollbar-track{background:transparent}
.mono{font-family:"IBM Plex Mono",monospace}
.app{display:flex;flex-direction:column;height:100vh;width:100%;overflow:hidden;background:#eef0f2}
header.top{display:flex;align-items:center;gap:14px;height:56px;flex:none;padding:0 22px;background:#fff;border-bottom:1px solid #e2e5ea}
.logo{display:flex;align-items:center;justify-content:center;width:30px;height:30px;border-radius:7px;background:#16181d;color:#fff;font-family:"IBM Plex Mono",monospace;font-weight:600;font-size:12px}
.htitle{display:flex;flex-direction:column;line-height:1.25;white-space:nowrap}
.htitle b{font-size:14.5px;font-weight:600;letter-spacing:-.01em}
.htitle span{font-size:11.5px;color:#838a95}
.hright{margin-left:auto;display:flex;align-items:center;gap:8px}
.chip{font-family:"IBM Plex Mono",monospace;font-size:11px;color:#838a95;background:#f3f5f7;border:1px solid #e6e8ec;padding:4px 10px;border-radius:6px}
.chip-ok{display:inline-flex;align-items:center;gap:6px;font-size:12px;color:#0f7a45;background:#e7f6ee;padding:4px 10px;border-radius:6px;font-weight:600}
.chip-ok i{width:7px;height:7px;border-radius:7px;background:#18a05a}
.frame{display:flex;flex:1;min-height:0}
nav.side{width:248px;flex:none;background:#fff;border-right:1px solid #e2e5ea;padding:14px 12px;display:flex;flex-direction:column;overflow-y:auto}
.navcap{font-family:"IBM Plex Mono",monospace;font-size:10px;letter-spacing:.1em;color:#a4abb4;padding:6px 11px 8px}
nav.side a{display:flex;align-items:center;gap:11px;width:100%;padding:8px 11px;border-radius:8px;cursor:pointer;text-decoration:none;font-size:13.5px;margin:1px 0;color:#475059;font-weight:500}
nav.side a .tag{width:24px;height:19px;border-radius:5px;display:flex;align-items:center;justify-content:center;font-family:"IBM Plex Mono",monospace;font-size:9px;font-weight:600;flex:none;background:#eceef1;color:#8a909b}
nav.side a.active{background:#eaf1fe;color:#16181d;font-weight:600}
nav.side a.active .tag{background:#1f5fdb;color:#fff}
.sidefoot{margin-top:auto;padding:14px 11px 6px;border-top:1px solid #eef0f2}
.sidefoot .cap{font-family:"IBM Plex Mono",monospace;font-size:10px;letter-spacing:.08em;color:#a4abb4;margin-bottom:8px}
.sidefoot .txt{font-size:12px;color:#6b727c;line-height:1.6}
main{flex:1;min-width:0;overflow-y:auto;padding:26px 30px 60px;scroll-behavior:smooth}
section{margin-bottom:44px;scroll-margin-top:10px}
h1{margin:0;font-size:21px;font-weight:600;letter-spacing:-.02em}
h2{margin:0 0 4px;font-size:18px;font-weight:600;letter-spacing:-.02em}
.sub{margin:4px 0 18px;font-size:13px;color:#838a95;max-width:860px;line-height:1.55}
.card{background:#fff;border:1px solid #e6e8ec;border-radius:11px;padding:18px 20px;margin-bottom:14px}
.card h3{margin:0 0 10px;font-size:14px;font-weight:600}
.card p{margin:0 0 10px;font-size:13px;color:#3d424b;line-height:1.6;max-width:880px}
.card p:last-child{margin-bottom:0}
.card ul{margin:6px 0 4px;padding-left:20px}
.card li{font-size:13px;color:#3d424b;line-height:1.65;margin-bottom:6px;max-width:840px}
.muted{color:#838a95}
.kpis{display:grid;grid-template-columns:repeat(4,1fr);gap:14px;margin-bottom:16px}
.kpi{background:#fff;border:1px solid #e6e8ec;border-radius:10px;padding:15px 16px}
.kpi .l{font-size:11.5px;color:#838a95;font-weight:500;text-transform:uppercase;letter-spacing:.04em}
.kpi .v{font-family:"IBM Plex Mono",monospace;font-size:26px;font-weight:600;margin-top:7px;line-height:1}
.kpi .s{font-size:11.5px;color:#a4abb4;margin-top:5px}
.v-bad{color:#b23b4e}.v-good{color:#0f7a45}.v-neutral{color:#16181d}
table.data{width:100%;border-collapse:collapse}
table.data th{font-size:10.5px;text-transform:uppercase;letter-spacing:.04em;color:#838a95;font-weight:600;text-align:left;padding:9px 12px;background:#f7f8fa;border-bottom:1px solid #eef0f2}
table.data td{font-size:12.5px;color:#3d424b;padding:11px 12px;border-bottom:1px solid #f4f5f7;vertical-align:top}
table.data td.mono, table.data th.num{font-family:"IBM Plex Mono",monospace}
table.data th.num, table.data td.num{text-align:right;font-family:"IBM Plex Mono",monospace}
.delta-bad{color:#b23b4e;font-weight:600}
.delta-ok{color:#0f7a45;font-weight:600}
.delta-zero{color:#838a95}
pre.code{background:#16181d;color:#dfe3e9;border-radius:9px;padding:14px 16px;font-family:"IBM Plex Mono",monospace;font-size:12px;line-height:1.65;overflow-x:auto;margin:12px 0 0}
pre.code .c{color:#7e8794}
pre.code .k{color:#7fb1ff}
pre.code .hl{color:#ff9db0}
.erd-wrap{overflow-x:auto;padding:8px 0}
.erd{position:relative;margin:0 auto}
.erd svg.wire{position:absolute;inset:0;pointer-events:none}
.tbl{position:absolute;background:#fff;border:1px solid #d4dae2;border-radius:9px;box-shadow:0 1px 4px rgba(20,30,50,.05);overflow:hidden}
.tbl.bulk{border:2px solid #1f5fdb;box-shadow:0 3px 12px rgba(31,95,219,.12)}
.tbl-h{background:#16181d;color:#fff;padding:9px 13px;font-family:"IBM Plex Mono",monospace;font-size:12px;font-weight:600;display:flex;justify-content:space-between;gap:8px}
.tbl.bulk .tbl-h{background:#1f5fdb}
.tbl-h .sub2{color:#7e8794;font-weight:400}
.tbl.bulk .tbl-h .sub2{color:#bcd2f7}
.tr{display:flex;justify-content:space-between;gap:8px;padding:6px 13px;font-size:11.5px;color:#3d424b;align-items:center}
.tr .f{font-family:"IBM Plex Mono",monospace}
.tr .t{font-family:"IBM Plex Mono",monospace;font-weight:600;font-size:10.5px;white-space:nowrap}
.tr.pk{background:#eaf1fe}.tr.pk .t{color:#1f5fdb}
.tr.fk{background:#fbf2e8}.tr.fk .t{color:#c08428}
.tr.ref .t{color:#9aa1ab}
.tr.idx{background:#fbe9ec}.tr.idx .t{color:#b23b4e}
.tr .ty{color:#a4abb4}
.notecard{position:absolute;background:#f7f8fa;border:1px dashed #c9cfd8;border-radius:9px;padding:12px 14px;font-size:11.5px;color:#5b6472;line-height:1.55}
.notecard b{font-family:"IBM Plex Mono",monospace;font-size:11px;color:#3d424b}
.hbars{display:flex;flex-direction:column;gap:12px;margin-top:6px}
.hbar{display:flex;align-items:center;gap:14px}
.hbar .lab{width:250px;flex:none;font-size:12.5px;color:#3d424b}
.hbar .lab small{display:block;color:#a4abb4;font-size:11px}
.hbar .trk{flex:1;background:#f1f3f6;border-radius:6px;overflow:hidden}
.hbar .fill{height:14px;border-radius:6px}
.hbar .val{font-family:"IBM Plex Mono",monospace;font-size:12.5px;color:#16181d;width:120px;flex:none;text-align:right}
.col-a{background:#9aa6b6}.col-b{background:#1f5fdb}.col-c{background:#c0476b}
.legend{display:flex;gap:16px;flex-wrap:wrap;margin:2px 0 10px}
.legend span{display:inline-flex;align-items:center;gap:6px;font-size:11.5px;color:#5b6472}
.legend i{width:10px;height:10px;border-radius:3px;display:inline-block}
.callout{border-left:3px solid #1f5fdb;background:#f5f8fe;border-radius:0 9px 9px 0;padding:12px 16px;font-size:13px;color:#2b3038;line-height:1.6;margin:12px 0;max-width:880px}
.callout.warn{border-left-color:#b23b4e;background:#fdf3f5}
.pill{display:inline-flex;align-items:center;gap:6px;padding:2px 9px;border-radius:20px;font-size:11px;font-weight:600}
.pill.rec{color:#0f7a45;background:#e7f6ee}
.pill.base{color:#5b6472;background:#eef0f2}
.pill.costly{color:#b23b4e;background:#fbe9ec}
.foot{font-size:11.5px;color:#a4abb4;border-top:1px solid #e6e8ec;padding-top:14px;margin-top:30px;line-height:1.7}
</style>
</head>
<body>
<div class="app">
<header class="top">
<div class="logo">FK</div>
<div class="htitle">
<b>Foreign-Key Architecture Study</b>
<span>Tribology lab data storage &middot; bulk-table key design</span>
</div>
<div class="hright">
<span class="chip">SQLite &middot; 4.32M rows measured</span>
<span class="chip-ok"><i></i>Measured 2026-07-11</span>
</div>
</header>
<div class="frame">
<nav class="side" id="sidenav">
<div class="navcap">NAVIGATE</div>
<a href="#overview" class="active"><span class="tag">OV</span><span>Overview</span></a>
<a href="#dataset"><span class="tag">DS</span><span>Sample Dataset</span></a>
<a href="#variant-a"><span class="tag">A</span><span>Variant A &middot; Composite PK</span></a>
<a href="#variant-b"><span class="tag">B</span><span>Variant B &middot; Enforced FK</span></a>
<a href="#variant-c"><span class="tag">C</span><span>Variant C &middot; Surrogate id</span></a>
<a href="#method"><span class="tag">BM</span><span>Benchmark Method</span></a>
<a href="#storage"><span class="tag">ST</span><span>Storage &amp; Load</span></a>
<a href="#reads"><span class="tag">RD</span><span>Read Performance</span></a>
<a href="#analysis"><span class="tag">AN</span><span>Why It Differs</span></a>
<a href="#verdict"><span class="tag">RC</span><span>Recommendation</span></a>
<div class="sidefoot">
<div class="cap">CONTEXT</div>
<div class="txt">Key-schema study for the bulk measurement tables of the LabDataStorageEvaluation pipeline (SQLite / PostgreSQL benchmark targets).</div>
</div>
</nav>
<main id="main">
<!-- ================= OVERVIEW ================= -->
<section id="overview">
<h1>Overview</h1>
<p class="sub">Three candidate key schemas for the two largest measurement tables (<span class="mono">friction_cycles</span>, 1.44M rows and <span class="mono">friction_loop_points</span>, 2.88M rows) were built from the same simulated corpus and measured for disk footprint, load time and six read patterns. In every variant the <b>relations between tables are single-column integer foreign keys</b> (<span class="mono">track_id</span>, <span class="mono">coupon_id</span>); the variants differ only in how a bulk row is identified inside its own table.</p>
<div class="kpis">
<div class="kpi"><div class="l">Disk, variant C vs A</div><div class="v v-bad">+74%</div><div class="s">190.2 vs 109.1 MiB</div></div>
<div class="kpi"><div class="l">GROUP BY scan, C vs A</div><div class="v v-bad">+59%</div><div class="s">135.9 vs 85.7 ms (Q5 pattern)</div></div>
<div class="kpi"><div class="l">Enforced FK cost (B)</div><div class="v v-good">+0.5 s</div><div class="s">load only; disk and reads = A</div></div>
<div class="kpi"><div class="l">Point lookups</div><div class="v v-neutral">&plusmn;2%</div><div class="s">identical across A / B / C</div></div>
</div>
<div class="card">
<h3>The three variants</h3>
<ul>
<li><b>Variant A - composite natural PK.</b> Row identity of a bulk row is its natural composite key <span class="mono">(track_id, cycle)</span>, stored as a clustered <span class="mono">WITHOUT ROWID</span> table. FK columns are plain integers, validated by the loader against the corpus manifest (no engine constraint).</li>
<li><b>Variant B - composite natural PK + enforced FK.</b> Same physical layout as A, plus <span class="mono">REFERENCES</span> constraints with <span class="mono">PRAGMA foreign_keys=ON</span>, so the engine itself guarantees referential integrity.</li>
<li><b>Variant C - surrogate key on every table.</b> Classic style: each table gets a single-column auto-increment key named <span class="mono">&lt;singular&gt;_id</span> (<span class="mono">friction_cycle_id</span>, never a bare <span class="mono">id</span>). Correctness then requires an additional <span class="mono">UNIQUE(track_id, cycle)</span> index - a second B-tree.</li>
</ul>
<div class="callout">Referential integrity is <b>not</b> what differs between the variants: every relation is a single integer FK in all three. What differs is the number of B-trees the engine must store and traverse per bulk table: one (A, B) or two (C).</div>
</div>
</section>
<!-- ================= DATASET ================= -->
<section id="dataset">
<h2>Sample Dataset</h2>
<p class="sub">All measurements ran against the real generated corpus of the evaluation pipeline - deterministic, physically plausible tribology data (seed 20260711). The two bulk tables and their parent dimension were loaded in full.</p>
<div class="card" style="padding:0;overflow:hidden">
<table class="data">
<thead><tr><th>Table</th><th>Grain</th><th class="num">Rows loaded</th><th>Columns (beyond keys)</th><th>Source files</th></tr></thead>
<tbody>
<tr><td class="mono">tracks</td><td>one friction track (3 per tested coupon)</td><td class="num">1,440</td><td>track_code, coupon_id, run_id, environment, load_mN, ...</td><td class="mono">track_info.csv &times; 1,440</td></tr>
<tr><td class="mono">friction_cycles</td><td>one reciprocating cycle of one track</td><td class="num">1,440,000</td><td>cof</td><td class="mono">cof_vs_cycle.csv &times; 1,440</td></tr>
<tr><td class="mono">friction_loop_points</td><td>one position-resolved loop point (10 loops &times; 200 pts per track)</td><td class="num">2,880,000</td><td>position_um, friction_force_mN</td><td class="mono">friction_loops.csv &times; 1,440</td></tr>
</tbody>
</table>
</div>
<div class="card">
<h3>Corpus context</h3>
<p>The corpus models a Pt-Au coating study: 4 deposition batches &rarr; 12 wafers &rarr; 588 coupons (480 friction-tested) &rarr; 1,440 tracks &rarr; 1.44M cycle records. Generated size: 141.8 MiB of CSV; the production target for extrapolation is 600 GB - 6 TB, which is why per-row bytes on the bulk tables matter.</p>
</div>
</section>
<!-- ================= VARIANT A ================= -->
<section id="variant-a">
<h2>Variant A &middot; Composite Natural Primary Key <span class="pill base" style="vertical-align:2px">baseline</span></h2>
<p class="sub">The bulk row is identified by what it physically is: cycle <span class="mono">n</span> of track <span class="mono">t</span>. The composite PK doubles as the clustering order, so the table itself is the only B-tree. FK columns are unenforced integers; integrity is guaranteed by row-count and checksum validation against the corpus manifest.</p>
<div class="card">
<div class="erd-wrap"><div class="erd" style="width:1000px;height:440px">
<svg class="wire" viewBox="0 0 1000 440" width="1000" height="440">
<path d="M270 200 H360 V77 H450" stroke="#c5cdd8" stroke-width="1.5" fill="none"/>
<path d="M270 240 H360 V269 H450" stroke="#c5cdd8" stroke-width="1.5" fill="none"/>
<circle cx="270" cy="200" r="3.5" fill="#1f5fdb"/><circle cx="270" cy="240" r="3.5" fill="#1f5fdb"/>
<text x="280" y="193" font-family="IBM Plex Mono, monospace" font-size="11" fill="#1f5fdb">1</text>
<text x="280" y="233" font-family="IBM Plex Mono, monospace" font-size="11" fill="#1f5fdb">1</text>
<text x="432" y="70" font-family="IBM Plex Mono, monospace" font-size="11" fill="#9aa1ab">&#8734;</text>
<text x="432" y="262" font-family="IBM Plex Mono, monospace" font-size="11" fill="#9aa1ab">&#8734;</text>
</svg>
<div class="tbl" style="left:20px;top:120px;width:250px">
<div class="tbl-h"><span>tracks</span><span class="sub2">dimension</span></div>
<div class="tr pk"><span class="f">track_id</span><span class="t">PK</span></div>
<div class="tr"><span class="f">track_code</span><span class="t" style="color:#1f5fdb">UK</span></div>
<div class="tr fk"><span class="f">coupon_id</span><span class="t">FK</span></div>
<div class="tr fk"><span class="f">run_id</span><span class="t">FK</span></div>
<div class="tr"><span class="f">environment</span><span class="ty">text</span></div>
<div class="tr"><span class="f">load_mn</span><span class="ty">real</span></div>
</div>
<div class="tbl bulk" style="left:450px;top:20px;width:280px">
<div class="tbl-h"><span>friction_cycles</span><span class="sub2">1.44M rows</span></div>
<div class="tr pk"><span class="f">track_id</span><span class="t">PK &middot; ref</span></div>
<div class="tr pk"><span class="f">cycle</span><span class="t">PK</span></div>
<div class="tr"><span class="f">cof</span><span class="ty">real</span></div>
</div>
<div class="tbl bulk" style="left:450px;top:185px;width:280px">
<div class="tbl-h"><span>friction_loop_points</span><span class="sub2">2.88M rows</span></div>
<div class="tr pk"><span class="f">track_id</span><span class="t">PK &middot; ref</span></div>
<div class="tr pk"><span class="f">cycle</span><span class="t">PK</span></div>
<div class="tr pk"><span class="f">pt</span><span class="t">PK</span></div>
<div class="tr"><span class="f">position_um</span><span class="ty">real</span></div>
<div class="tr"><span class="f">friction_force_mn</span><span class="ty">real</span></div>
</div>
<div class="notecard" style="left:770px;top:130px;width:210px">
<b>wear</b> and <b>track_summary</b> follow the same pattern: single-column PK = FK <b>track_id</b> (1:1 with tracks).<br><br>
<b>ref</b> = integer reference to tracks(track_id), validated by the loader, no engine constraint.
</div>
</div></div>
<pre class="code">CREATE TABLE friction_cycles (
track_id INTEGER NOT NULL, <span class="c">-- reference to tracks(track_id), loader-validated</span>
cycle INTEGER NOT NULL,
cof REAL NOT NULL,
<span class="k">PRIMARY KEY (track_id, cycle)</span>
) <span class="k">WITHOUT ROWID</span>; <span class="c">-- clustered: the PK is the table's only B-tree</span></pre>
</div>
</section>
<!-- ================= VARIANT B ================= -->
<section id="variant-b">
<h2>Variant B &middot; Composite PK + Enforced Foreign Keys <span class="pill rec" style="vertical-align:2px">recommended</span></h2>
<p class="sub">Physically identical to Variant A - same clustered layout, same single B-tree per bulk table. The only change: FK columns carry real <span class="mono">REFERENCES</span> constraints and every connection runs <span class="mono">PRAGMA foreign_keys=ON</span>, so the DBMS itself rejects an orphan row. Measured cost: +0.5 s of load time on 4.32M rows; zero cost on disk and reads.</p>
<div class="card">
<div class="erd-wrap"><div class="erd" style="width:1000px;height:440px">
<svg class="wire" viewBox="0 0 1000 440" width="1000" height="440">
<path d="M270 200 H360 V77 H450" stroke="#c5cdd8" stroke-width="1.5" fill="none"/>
<path d="M270 240 H360 V269 H450" stroke="#c5cdd8" stroke-width="1.5" fill="none"/>
<circle cx="270" cy="200" r="3.5" fill="#1f5fdb"/><circle cx="270" cy="240" r="3.5" fill="#1f5fdb"/>
<text x="280" y="193" font-family="IBM Plex Mono, monospace" font-size="11" fill="#1f5fdb">1</text>
<text x="280" y="233" font-family="IBM Plex Mono, monospace" font-size="11" fill="#1f5fdb">1</text>
<text x="432" y="70" font-family="IBM Plex Mono, monospace" font-size="11" fill="#9aa1ab">&#8734;</text>
<text x="432" y="262" font-family="IBM Plex Mono, monospace" font-size="11" fill="#9aa1ab">&#8734;</text>
</svg>
<div class="tbl" style="left:20px;top:120px;width:250px">
<div class="tbl-h"><span>tracks</span><span class="sub2">dimension</span></div>
<div class="tr pk"><span class="f">track_id</span><span class="t">PK</span></div>
<div class="tr"><span class="f">track_code</span><span class="t" style="color:#1f5fdb">UK</span></div>
<div class="tr fk"><span class="f">coupon_id</span><span class="t">FK</span></div>
<div class="tr fk"><span class="f">run_id</span><span class="t">FK</span></div>
<div class="tr"><span class="f">environment</span><span class="ty">text</span></div>
<div class="tr"><span class="f">load_mn</span><span class="ty">real</span></div>
</div>
<div class="tbl bulk" style="left:450px;top:20px;width:280px">
<div class="tbl-h"><span>friction_cycles</span><span class="sub2">1.44M rows</span></div>
<div class="tr pk"><span class="f">track_id</span><span class="t">PK</span></div>
<div class="tr fk"><span class="f">track_id &rarr; tracks</span><span class="t">FK &#10003;</span></div>
<div class="tr pk"><span class="f">cycle</span><span class="t">PK</span></div>
<div class="tr"><span class="f">cof</span><span class="ty">real</span></div>
</div>
<div class="tbl bulk" style="left:450px;top:190px;width:280px">
<div class="tbl-h"><span>friction_loop_points</span><span class="sub2">2.88M rows</span></div>
<div class="tr pk"><span class="f">track_id</span><span class="t">PK</span></div>
<div class="tr fk"><span class="f">track_id &rarr; tracks</span><span class="t">FK &#10003;</span></div>
<div class="tr pk"><span class="f">cycle</span><span class="t">PK</span></div>
<div class="tr pk"><span class="f">pt</span><span class="t">PK</span></div>
<div class="tr"><span class="f">position_um, friction_force_mn</span><span class="ty">real</span></div>
</div>
<div class="notecard" style="left:770px;top:130px;width:210px">
<b>PRAGMA foreign_keys = ON</b> on every connection - SQLite silently ignores FK clauses without it.<br><br>
FK checks are lookups into the already-indexed parent PK, which is why the measured overhead is only ~6% of load time.
</div>
</div></div>
<pre class="code">PRAGMA foreign_keys = ON;
CREATE TABLE friction_cycles (
track_id INTEGER NOT NULL <span class="k">REFERENCES tracks(track_id)</span>,
cycle INTEGER NOT NULL,
cof REAL NOT NULL,
<span class="k">PRIMARY KEY (track_id, cycle)</span>
) <span class="k">WITHOUT ROWID</span>;</pre>
</div>
</section>
<!-- ================= VARIANT C ================= -->
<section id="variant-c">
<h2>Variant C &middot; Surrogate <span class="mono" style="font-size:16px">&lt;singular&gt;_id</span> on Every Table <span class="pill costly" style="vertical-align:2px">+74% disk</span></h2>
<p class="sub">The uniform classic design: every table, including the bulk fact tables, gets a single-column auto-increment key (<span class="mono">friction_cycle_id</span>, <span class="mono">friction_loop_point_id</span> - by convention always <span class="mono">&lt;table name in singular&gt;_id</span>, never a bare <span class="mono">id</span>). In SQLite the id column itself is nearly free (it aliases the rowid), but correctness still demands <span class="mono">UNIQUE(track_id, cycle)</span> - and that unique index is a second B-tree holding a copy of almost every column.</p>
<div class="card">
<div class="erd-wrap"><div class="erd" style="width:1000px;height:520px">
<svg class="wire" viewBox="0 0 1000 520" width="1000" height="520">
<path d="M270 240 H360 V100 H450" stroke="#c5cdd8" stroke-width="1.5" fill="none"/>
<path d="M270 280 H360 V345 H450" stroke="#c5cdd8" stroke-width="1.5" fill="none"/>
<circle cx="270" cy="240" r="3.5" fill="#1f5fdb"/><circle cx="270" cy="280" r="3.5" fill="#1f5fdb"/>
<text x="280" y="233" font-family="IBM Plex Mono, monospace" font-size="11" fill="#1f5fdb">1</text>
<text x="280" y="273" font-family="IBM Plex Mono, monospace" font-size="11" fill="#1f5fdb">1</text>
<text x="432" y="93" font-family="IBM Plex Mono, monospace" font-size="11" fill="#9aa1ab">&#8734;</text>
<text x="432" y="338" font-family="IBM Plex Mono, monospace" font-size="11" fill="#9aa1ab">&#8734;</text>
</svg>
<div class="tbl" style="left:20px;top:160px;width:250px">
<div class="tbl-h"><span>tracks</span><span class="sub2">dimension</span></div>
<div class="tr pk"><span class="f">track_id</span><span class="t">PK</span></div>
<div class="tr"><span class="f">track_code</span><span class="t" style="color:#1f5fdb">UK</span></div>
<div class="tr fk"><span class="f">coupon_id</span><span class="t">FK</span></div>
<div class="tr fk"><span class="f">run_id</span><span class="t">FK</span></div>
<div class="tr"><span class="f">environment</span><span class="ty">text</span></div>
<div class="tr"><span class="f">load_mn</span><span class="ty">real</span></div>
</div>
<div class="tbl bulk" style="left:450px;top:20px;width:280px">
<div class="tbl-h"><span>friction_cycles</span><span class="sub2">1.44M rows</span></div>
<div class="tr pk"><span class="f">friction_cycle_id</span><span class="t">PK</span></div>
<div class="tr fk"><span class="f">track_id &rarr; tracks</span><span class="t">FK &#10003;</span></div>
<div class="tr"><span class="f">cycle</span><span class="ty">int</span></div>
<div class="tr"><span class="f">cof</span><span class="ty">real</span></div>
<div class="tr idx"><span class="f">UNIQUE (track_id, cycle)</span><span class="t">2nd B-tree</span></div>
</div>
<div class="tbl bulk" style="left:450px;top:235px;width:280px">
<div class="tbl-h"><span>friction_loop_points</span><span class="sub2">2.88M rows</span></div>
<div class="tr pk"><span class="f">friction_loop_point_id</span><span class="t">PK</span></div>
<div class="tr fk"><span class="f">track_id &rarr; tracks</span><span class="t">FK &#10003;</span></div>
<div class="tr"><span class="f">cycle</span><span class="ty">int</span></div>
<div class="tr"><span class="f">pt</span><span class="ty">int</span></div>
<div class="tr"><span class="f">position_um</span><span class="ty">real</span></div>
<div class="tr"><span class="f">friction_force_mn</span><span class="ty">real</span></div>
<div class="tr idx"><span class="f">UNIQUE (track_id, cycle, pt)</span><span class="t">2nd B-tree</span></div>
</div>
<div class="notecard" style="left:770px;top:180px;width:210px">
The surrogate key is <b>never referenced</b> by any other table - no row in the schema points at an individual cycle.<br><br>
The UNIQUE index is <b>mandatory for correctness</b> (it prevents duplicate cycles), so its cost cannot be avoided in this variant.
</div>
</div></div>
<pre class="code">CREATE TABLE friction_cycles (
<span class="hl">friction_cycle_id INTEGER PRIMARY KEY</span>, <span class="c">-- rowid alias; unused by any FK</span>
track_id INTEGER NOT NULL REFERENCES tracks(track_id),
cycle INTEGER NOT NULL,
cof REAL NOT NULL,
<span class="hl">UNIQUE (track_id, cycle)</span> <span class="c">-- required; duplicates the table</span>
);</pre>
</div>
</section>
<!-- ================= METHOD ================= -->
<section id="method">
<h2>Benchmark Method</h2>
<p class="sub">Identical protocol for all three variants; only the DDL differs.</p>
<div class="card">
<ul>
<li><b>Data:</b> the full generated corpus subset - <span class="mono">tracks</span> (1,440), <span class="mono">friction_cycles</span> (1,440,000), <span class="mono">friction_loop_points</span> (2,880,000) - read from the same CSV files in the same order for every variant.</li>
<li><b>Load:</b> SQLite, <span class="mono">PRAGMA journal_mode=OFF, synchronous=OFF, cache_size=-262144</span>; inserts via <span class="mono">executemany</span> in 50k-row batches; <span class="mono">VACUUM</span> before measuring the file size.</li>
<li><b>Reads:</b> warm OS page cache, best of 3 runs per pattern; fixed deterministic probe sets (39 tracks for range reads, 150 <span class="mono">(track_id, cycle)</span> pairs for point lookups); result-set sizes asserted, so every variant answers the identical question.</li>
<li><b>Caveat:</b> at 150 MB everything fits in RAM, which <i>flatters</i> variant C - at the 600 GB production target the working set exceeds RAM, and twice the pages means twice the cache pressure and disk reads. Warm-cache deltas are therefore a lower bound.</li>
</ul>
</div>
</section>
<!-- ================= STORAGE ================= -->
<section id="storage">
<h2>Results &middot; Storage &amp; Load</h2>
<p class="sub">File size after VACUUM; load time includes CSV parsing (identical work in all variants).</p>
<div class="card" style="padding:0;overflow:hidden">
<table class="data">
<thead><tr><th>Variant</th><th>Bulk-table keys</th><th class="num">Load time (s)</th><th class="num">DB size (MiB)</th><th class="num">Size vs A</th></tr></thead>
<tbody>
<tr><td><b>A</b> &middot; composite PK</td><td>PK (track_id, cycle), unenforced refs</td><td class="num">7.5</td><td class="num">109.1</td><td class="num delta-zero">baseline</td></tr>
<tr><td><b>B</b> &middot; composite PK + FK</td><td>PK (track_id, cycle), REFERENCES enforced</td><td class="num">8.0</td><td class="num">109.1</td><td class="num delta-ok">&plusmn;0%</td></tr>
<tr><td><b>C</b> &middot; surrogate id</td><td>friction_cycle_id PK + UNIQUE(track_id, cycle)</td><td class="num">8.8</td><td class="num">190.2</td><td class="num delta-bad">+74%</td></tr>
</tbody>
</table>
</div>
<div class="card">
<h3>Database size after VACUUM (MiB, scale 0-200)</h3>
<div class="hbars">
<div class="hbar"><span class="lab">Variant A<small>composite PK, unenforced refs</small></span><div class="trk"><div class="fill col-a" style="width:54.6%"></div></div><span class="val">109.1 MiB</span></div>
<div class="hbar"><span class="lab">Variant B<small>composite PK + enforced FK</small></span><div class="trk"><div class="fill col-b" style="width:54.6%"></div></div><span class="val">109.1 MiB</span></div>
<div class="hbar"><span class="lab">Variant C<small>surrogate id + UNIQUE index</small></span><div class="trk"><div class="fill col-c" style="width:95.1%"></div></div><span class="val">190.2 MiB</span></div>
</div>
<h3 style="margin-top:20px">Load time, 4.32M rows (s, scale 0-10)</h3>
<div class="hbars">
<div class="hbar"><span class="lab">Variant A</span><div class="trk"><div class="fill col-a" style="width:75%"></div></div><span class="val">7.5 s</span></div>
<div class="hbar"><span class="lab">Variant B</span><div class="trk"><div class="fill col-b" style="width:80%"></div></div><span class="val">8.0 s</span></div>
<div class="hbar"><span class="lab">Variant C</span><div class="trk"><div class="fill col-c" style="width:88%"></div></div><span class="val">8.8 s</span></div>
</div>
<div class="callout" style="margin-top:16px">Extrapolated to the 600 GB production target, the +74% of variant C on the two hottest tables translates into <b>hundreds of gigabytes of extra disk</b> and a doubled page count competing for the same RAM cache.</div>
</div>
</section>
<!-- ================= READS ================= -->
<section id="reads">
<h2>Results &middot; Read Performance</h2>
<p class="sub">Warm OS cache, best of 3 runs. Q1 / Q5 refer to the benchmark scenarios of the evaluation pipeline.</p>
<div class="card" style="padding:0;overflow:hidden">
<table class="data">
<thead><tr><th>Read pattern</th><th>Unit</th><th class="num">A</th><th class="num">B</th><th class="num">C</th><th class="num">C vs A</th></tr></thead>
<tbody>
<tr><td>Point lookup of one row by <span class="mono">(track_id, cycle)</span></td><td class="mono">&micro;s/op</td><td class="num">24.8</td><td class="num">23.6</td><td class="num">24.3</td><td class="num delta-zero">~0%</td></tr>
<tr><td>Q1: all 1,000 cycles of one track</td><td class="mono">ms/track</td><td class="num">0.473</td><td class="num">0.485</td><td class="num">0.533</td><td class="num delta-bad">+13%</td></tr>
<tr><td>All 2,000 loop points of one track</td><td class="mono">ms/track</td><td class="num">1.267</td><td class="num">1.299</td><td class="num">1.336</td><td class="num delta-bad">+5%</td></tr>
<tr><td>Full scan + AVG over 1.44M cycles</td><td class="mono">ms</td><td class="num">64.9</td><td class="num">64.9</td><td class="num">70.7</td><td class="num delta-bad">+9%</td></tr>
<tr><td>AVG with GROUP BY track_id over 1.44M cycles (Q5 pattern)</td><td class="mono">ms</td><td class="num">85.7</td><td class="num">89.1</td><td class="num">135.9</td><td class="num delta-bad">+59%</td></tr>
<tr><td>Full scan over 2.88M loop points</td><td class="mono">ms</td><td class="num">152.5</td><td class="num">150.8</td><td class="num">165.4</td><td class="num delta-bad">+8%</td></tr>
</tbody>
</table>
</div>
<div class="card">
<h3>Relative read time, variant A = 100% (lower is better)</h3>
<div class="legend">
<span><i class="col-a"></i>A &middot; composite PK</span>
<span><i class="col-b"></i>B &middot; + enforced FK</span>
<span><i class="col-c"></i>C &middot; surrogate id + UNIQUE</span>
</div>
<svg viewBox="0 0 1080 300" style="width:100%;height:auto;display:block">
<line x1="60" y1="246" x2="1080" y2="246" stroke="#d8dce1" stroke-width="1"/>
<line x1="60" y1="183.2" x2="1080" y2="183.2" stroke="#eef0f2" stroke-width="1"/>
<line x1="60" y1="120.4" x2="1080" y2="120.4" stroke="#c9d4e6" stroke-width="1" stroke-dasharray="4 3"/>
<line x1="60" y1="57.7" x2="1080" y2="57.7" stroke="#eef0f2" stroke-width="1"/>
<text x="54" y="246" text-anchor="end" dominant-baseline="middle" font-family="IBM Plex Mono, monospace" font-size="9" fill="#a4abb4">0%</text>
<text x="54" y="183.2" text-anchor="end" dominant-baseline="middle" font-family="IBM Plex Mono, monospace" font-size="9" fill="#a4abb4">50%</text>
<text x="54" y="120.4" text-anchor="end" dominant-baseline="middle" font-family="IBM Plex Mono, monospace" font-size="9" fill="#7e8ba1">100%</text>
<text x="54" y="57.7" text-anchor="end" dominant-baseline="middle" font-family="IBM Plex Mono, monospace" font-size="9" fill="#a4abb4">150%</text>
<!-- group 0: point lookup -->
<rect x="78" y="120.4" width="38" height="125.6" rx="2.5" fill="#9aa6b6"/>
<rect x="124" y="126.5" width="38" height="119.5" rx="2.5" fill="#1f5fdb"/>
<rect x="170" y="123.0" width="38" height="123.0" rx="2.5" fill="#c0476b"/>
<text x="143" y="262" text-anchor="middle" font-size="10" fill="#5b6472">Point lookup</text>
<!-- group 1: Q1 cycles -->
<rect x="245" y="120.4" width="38" height="125.6" rx="2.5" fill="#9aa6b6"/>
<rect x="291" y="117.3" width="38" height="128.7" rx="2.5" fill="#1f5fdb"/>
<rect x="337" y="104.5" width="38" height="141.5" rx="2.5" fill="#c0476b"/>
<text x="310" y="262" text-anchor="middle" font-size="10" fill="#5b6472">Q1: track cycles</text>
<!-- group 2: loop points range -->
<rect x="411" y="120.4" width="38" height="125.6" rx="2.5" fill="#9aa6b6"/>
<rect x="457" y="117.3" width="38" height="128.7" rx="2.5" fill="#1f5fdb"/>
<rect x="503" y="113.7" width="38" height="132.3" rx="2.5" fill="#c0476b"/>
<text x="476" y="262" text-anchor="middle" font-size="10" fill="#5b6472">Track loop points</text>
<!-- group 3: full scan avg -->
<rect x="578" y="120.4" width="38" height="125.6" rx="2.5" fill="#9aa6b6"/>
<rect x="624" y="120.4" width="38" height="125.6" rx="2.5" fill="#1f5fdb"/>
<rect x="670" y="109.3" width="38" height="136.7" rx="2.5" fill="#c0476b"/>
<text x="643" y="262" text-anchor="middle" font-size="10" fill="#5b6472">Full scan + AVG</text>
<!-- group 4: group by -->
<rect x="745" y="120.4" width="38" height="125.6" rx="2.5" fill="#9aa6b6"/>
<rect x="791" y="115.4" width="38" height="130.6" rx="2.5" fill="#1f5fdb"/>
<rect x="837" y="46.9" width="38" height="199.1" rx="2.5" fill="#c0476b"/>
<text x="856" y="40" text-anchor="middle" font-family="IBM Plex Mono, monospace" font-size="10" font-weight="600" fill="#b23b4e">+59%</text>
<text x="810" y="262" text-anchor="middle" font-size="10" fill="#5b6472">GROUP BY track (Q5)</text>
<!-- group 5: loops scan -->
<rect x="911" y="120.4" width="38" height="125.6" rx="2.5" fill="#9aa6b6"/>
<rect x="957" y="121.8" width="38" height="124.2" rx="2.5" fill="#1f5fdb"/>
<rect x="1003" y="109.8" width="38" height="136.2" rx="2.5" fill="#c0476b"/>
<text x="976" y="262" text-anchor="middle" font-size="10" fill="#5b6472">Loop points scan</text>
<text x="570" y="290" text-anchor="middle" font-size="10" fill="#838a95">dashed line = variant A baseline (100%)</text>
</svg>
</div>
</section>
<!-- ================= ANALYSIS ================= -->
<section id="analysis">
<h2>Why It Differs</h2>
<p class="sub">The whole difference reduces to one thing: how many B-trees per bulk table, and in what order the rows physically live.</p>
<div class="card">
<div class="erd-wrap"><div class="erd" style="width:1000px;height:230px">
<svg class="wire" viewBox="0 0 1000 230" width="1000" height="230">
<path d="M700 115 H760" stroke="#b23b4e" stroke-width="1.6" fill="none" stroke-dasharray="5 4"/>
<text x="730" y="105" text-anchor="middle" font-family="IBM Plex Mono, monospace" font-size="10" fill="#b23b4e">extra hop</text>
</svg>
<div class="tbl" style="left:20px;top:30px;width:380px;border-color:#0f9d8a">
<div class="tbl-h" style="background:#0f9d8a"><span>Variants A / B &middot; one B-tree</span></div>
<div class="tr"><span class="f">clustered table keyed by (track_id, cycle)</span></div>
<div class="tr"><span class="f">descent &rarr; leaf &rarr; sequential range scan</span></div>
<div class="tr"><span class="f muted">rows of one track are physically adjacent</span></div>
</div>
<div class="tbl" style="left:450px;top:10px;width:250px;border-color:#c0476b">
<div class="tbl-h" style="background:#c0476b"><span>Variant C &middot; B-tree 1</span></div>
<div class="tr"><span class="f">UNIQUE index (track_id, cycle)</span></div>
<div class="tr"><span class="f muted">stores key columns + rowid again</span></div>
</div>
<div class="tbl" style="left:760px;top:80px;width:220px;border-color:#c0476b">
<div class="tbl-h" style="background:#c0476b"><span>Variant C &middot; B-tree 2</span></div>
<div class="tr"><span class="f">table keyed by friction_cycle_id</span></div>
<div class="tr"><span class="f muted">per-row lookup to fetch cof</span></div>
</div>
</div></div>
<ul>
<li><b>+74% disk.</b> In SQLite the surrogate id column is a rowid alias and costs nothing in the table itself - but the mandatory <span class="mono">UNIQUE(track_id, cycle)</span> index is a full second B-tree whose entries contain almost every column of the row. Two trees instead of one on a 3-column table is close to a doubling.</li>
<li><b>+59% on grouped aggregation (the pipeline's forced-full-scan Q5).</b> In A/B the table is clustered by <span class="mono">track_id</span>, so <span class="mono">GROUP BY track_id</span> streams the tree in group order. In C the engine walks the UNIQUE index for ordering and hops into the table B-tree for every <span class="mono">cof</span> value - the classic double access path.</li>
<li><b>+5-13% on ranges and plain scans.</b> More pages (190 vs 109 MiB) means more page traversals and cell decoding, even fully cached.</li>
<li><b>~0% on point lookups.</b> Tree depth is the same either way; a single-row probe does not expose the second tree.</li>
<li><b>PostgreSQL outlook.</b> PG has no WITHOUT ROWID: the composite-PK table is a heap plus one PK index either way, but variant C still adds the id column to every heap tuple <i>and</i> the second index - expect +30-50% on the bulk tables, plus higher COPY-time FK validation cost.</li>
</ul>
<div class="callout warn"><b>Integrity is a constant across all variants.</b> Every relation is a single-column integer FK. Choosing C buys naming uniformity, not more integrity - and the surrogate key it adds is never referenced by anything.</div>
</div>
</section>
<!-- ================= RECOMMENDATION ================= -->
<section id="verdict">
<h2>Recommendation</h2>
<div class="card" style="border-left:4px solid #0f7a45">
<h3><span class="pill rec">Variant B</span> &nbsp;Composite natural PK on bulk tables + enforced foreign keys</h3>
<ul>
<li>Referential integrity is guaranteed by the DBMS itself at a measured cost of <b>+0.5 s of load time</b> and nothing else.</li>
<li>Disk footprint and read profile are identical to the leanest variant A: <b>109.1 MiB</b>, single clustered B-tree per bulk table.</li>
<li>Dimension tables (<span class="mono">batches</span>, <span class="mono">wafers</span>, <span class="mono">coupons</span>, <span class="mono">runs</span>, <span class="mono">tracks</span>, <span class="mono">instruments</span>) keep their classic single-column surrogate keys - <span class="mono">&lt;singular&gt;_id</span>, never a bare <span class="mono">id</span> - because other tables reference their rows.</li>
<li>Variant C remains a valid design where fact rows must be individually addressable from outside (annotations, reprocessing links). No table in this schema references an individual cycle or loop point, so its cost (+74% disk, up to +59% on aggregation, growing at 600 GB scale) buys nothing here.</li>
</ul>
</div>
<div class="foot">
Measured 2026-07-11 on the deterministic simulated corpus (seed 20260711, 141.8 MiB CSV) &middot; SQLite, warm OS cache, best of 3 runs &middot;
scripts: <span class="mono">measure_key_variants.py</span> (build + size + load + Q1) and <span class="mono">read_benchmarks.py</span> (six read patterns) &middot;
companion page: <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html" style="color:#1f5fdb;text-decoration:none">Storage Format Evaluation</a> (full study results).
</div>
</section>
</main>
</div>
</div>
<script>
(function(){
var links = Array.prototype.slice.call(document.querySelectorAll('#sidenav a[href^="#"]'));
var byId = {};
links.forEach(function(a){ byId[a.getAttribute('href').slice(1)] = a; });
var sections = Array.prototype.slice.call(document.querySelectorAll('main section'));
var main = document.getElementById('main');
function setActive(id){
links.forEach(function(a){ a.classList.remove('active'); });
if(byId[id]) byId[id].classList.add('active');
}
links.forEach(function(a){
a.addEventListener('click', function(){ setActive(a.getAttribute('href').slice(1)); });
});
var ticking = false;
main.addEventListener('scroll', function(){
if(ticking) return; ticking = true;
requestAnimationFrame(function(){
var top = main.scrollTop + 90, current = sections[0].id;
sections.forEach(function(s){ if(s.offsetTop <= top) current = s.id; });
setActive(current);
ticking = false;
});
});
})();
</script>
</body>
</html>

View File

@@ -0,0 +1,258 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Storage Decision - Executive Summary</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=IBM+Plex+Mono:wght@400;500;600&family=IBM+Plex+Sans:wght@400;500;600;700&display=swap" rel="stylesheet">
<style>
*{box-sizing:border-box}
html,body{margin:0;padding:0}
body{font-family:"IBM Plex Sans",system-ui,sans-serif;color:#16181d;background:#eef0f2;-webkit-font-smoothing:antialiased}
.mono{font-family:"IBM Plex Mono",monospace}
header.top{display:flex;align-items:center;gap:14px;height:56px;padding:0 22px;background:#fff;border-bottom:1px solid #e2e5ea;position:sticky;top:0;z-index:5}
.logo{display:flex;align-items:center;justify-content:center;width:30px;height:30px;border-radius:7px;background:#16181d;color:#fff;font-family:"IBM Plex Mono",monospace;font-weight:600;font-size:12px;flex:none}
.htitle{display:flex;flex-direction:column;line-height:1.25;white-space:nowrap}
.htitle b{font-size:14.5px;font-weight:600;letter-spacing:-.01em}
.htitle span{font-size:11.5px;color:#838a95}
.hright{margin-left:auto;display:flex;align-items:center;gap:8px}
.chip{font-family:"IBM Plex Mono",monospace;font-size:11px;color:#838a95;background:#f3f5f7;border:1px solid #e6e8ec;padding:4px 10px;border-radius:6px;text-decoration:none}
a.chip:hover{border-color:#1f5fdb;color:#1f5fdb}
.chip-go{display:inline-flex;align-items:center;gap:6px;font-size:12px;color:#fff;background:#1f5fdb;padding:5px 12px;border-radius:6px;font-weight:600;text-decoration:none}
.chip-go:hover{background:#1a4fb8}
.page{max-width:980px;margin:0 auto;padding:30px 24px 70px}
.card{background:#fff;border:1px solid #e6e8ec;border-radius:12px;padding:22px 26px;margin-bottom:16px}
h1{margin:0 0 6px;font-size:24px;font-weight:700;letter-spacing:-.02em}
h2{margin:0 0 12px;font-size:16px;font-weight:600;letter-spacing:-.01em}
p{margin:0 0 10px;font-size:13.5px;color:#3d424b;line-height:1.65}
p:last-child{margin-bottom:0}
.muted{color:#838a95}
a.ext{color:#1f5fdb;text-decoration:none}
a.ext:hover{text-decoration:underline}
a.fact{color:inherit;text-decoration:underline dashed;text-decoration-thickness:1px;text-underline-offset:3px}
a.fact:hover{text-decoration-style:solid}
.hero{display:grid;grid-template-columns:1fr 1fr;gap:16px;margin-bottom:16px}
.hero .card{margin:0}
.hero .cap{font-family:"IBM Plex Mono",monospace;font-size:10.5px;letter-spacing:.1em;color:#a4abb4;margin-bottom:10px}
.hero .big{font-size:15.5px;line-height:1.6;color:#16181d}
.verdict{border:1px solid #bfe6cf;background:#e9f8ef;border-radius:12px;padding:16px 22px;display:flex;align-items:flex-start;gap:16px;margin-bottom:20px}
.verdict .mark{width:34px;height:34px;border-radius:50%;background:#18a05a;color:#fff;display:flex;align-items:center;justify-content:center;font-size:18px;font-weight:700;flex:none;margin-top:2px}
.verdict b{color:#0f7a45}
.verdict p{margin:0 0 8px;font-size:14.5px}
table.vtab{border-collapse:collapse;width:100%}
table.vtab th{text-align:left;font-size:11px;font-weight:600;padding:5px 16px 5px 0;border-bottom:1px solid #cdeadb}
table.vtab th.win{color:#0f7a45}
table.vtab th.lose{color:#b23b4e}
table.vtab td{font-size:13px;color:#2b3038;padding:7px 16px 7px 0;border-bottom:1px solid #d9f0e3;line-height:1.5;vertical-align:top}
table.vtab tr:last-child td{border-bottom:0}
table.vtab td.crit{font-weight:600;color:#16181d;width:26%}
table.vtab td.lose{color:#b23b4e}
.kpis{display:grid;grid-template-columns:repeat(3,1fr);gap:16px;margin-bottom:16px}
.kpi{background:#fff;border:1px solid #e6e8ec;border-radius:12px;padding:18px 20px}
.kpi .v{font-family:"IBM Plex Mono",monospace;font-size:34px;font-weight:600;line-height:1;color:#16181d}
.kpi .l{font-size:12.5px;color:#5b6472;margin-top:9px;line-height:1.5}
table.duel{width:100%;border-collapse:collapse}
table.duel th{padding:6px 14px;text-align:left;font-size:12px;border-bottom:2px solid #eef0f2;vertical-align:bottom}
table.duel th .pill{display:inline-flex;padding:3px 10px;border-radius:20px;font-size:10.5px;font-weight:600;margin-top:4px}
table.duel th.win{color:#0f7a45}
table.duel th.win .pill{color:#0f7a45;background:#e7f6ee}
table.duel th.lose{color:#b23b4e}
table.duel th.lose .pill{color:#b23b4e;background:#fbe9ec}
table.duel td{padding:7px 14px;font-size:13px;color:#3d424b;border-bottom:1px solid #f4f5f7;line-height:1.45;vertical-align:top}
table.duel td.crit{font-weight:600;color:#16181d;width:26%}
table.duel td.crit small{display:block;font-weight:400;color:#a4abb4;margin-top:2px}
table.duel td .ok{color:#0f7a45;font-weight:700;margin-right:6px}
table.duel td .no{color:#b23b4e;font-weight:700;margin-right:6px}
table.duel td.more{width:12%;text-align:right;white-space:nowrap}
.panels{display:grid;grid-template-columns:1fr 1fr;gap:22px;margin-top:6px}
.panel h3{margin:0 0 12px;font-size:12.5px;font-weight:600;color:#5b6472;text-transform:uppercase;letter-spacing:.04em}
.brow{display:flex;align-items:center;gap:12px;margin-bottom:12px}
.brow .lab{width:118px;flex:none;font-size:12.5px;color:#3d424b}
.brow .trk{flex:1;background:#f1f3f6;border-radius:5px;height:20px;overflow:hidden;position:relative}
.brow .fill{height:20px;border-radius:0 4px 4px 0;min-width:3px}
.brow .inval{position:absolute;top:50%;transform:translateY(-50%);font-family:"IBM Plex Mono",monospace;font-size:12px;font-weight:600;white-space:nowrap}
.brow .inval.in-track{left:12px;color:#16181d}
.brow .inval.in-fill{right:12px;color:#fff}
.brow:hover .trk{background:#eaedf1}
.fill.pg{background:#1f5fdb}
.fill.rdf{background:#c0476b}
.steps{counter-reset:st;list-style:none;margin:6px 0 0;padding:0}
.steps li{counter-increment:st;display:flex;gap:14px;align-items:flex-start;padding:10px 0;font-size:13.5px;color:#3d424b;line-height:1.6;border-bottom:1px solid #f4f5f7}
.steps li:last-child{border-bottom:0}
.steps li::before{content:counter(st);width:26px;height:26px;border-radius:50%;background:#eaf1fe;color:#1f5fdb;font-weight:700;font-size:13px;display:flex;align-items:center;justify-content:center;flex:none;margin-top:1px}
.foot{font-size:11.5px;color:#a4abb4;border-top:1px solid #e6e8ec;padding-top:16px;margin-top:26px;line-height:1.8}
@media(max-width:840px){.hero,.kpis,.panels{grid-template-columns:1fr}}
</style>
</head>
<body>
<header class="top">
<div class="logo">SE</div>
<div class="htitle">
<b>Storage Decision</b>
<span>Executive summary &middot; laboratory measurement data</span>
</div>
<div class="hright">
<a class="chip" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation" target="_blank" rel="noopener">Git repository</a>
<a class="chip-go" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html" target="_blank" rel="noopener">Open the full study &rarr;</a>
</div>
</header>
<div class="page">
<h1>Which database should hold the laboratory's data?</h1>
<div class="hero">
<div class="card">
<div class="cap">THE PROBLEM</div>
<p class="big"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#premises" target="_blank" rel="noopener"><b>600 GB</b></a> of laboratory data today.<br/><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#premises" target="_blank" rel="noopener"><b>6 TB</b></a> within a few years.<br/>It needs one working database.</p>
</div>
<div class="card">
<div class="cap">THE CANDIDATES</div>
<p class="big"><b>Relational</b> (PostgreSQL)<br/>vs<br/><b>Graph</b> (RDF triplestore).</p>
</div>
</div>
<div class="verdict">
<div class="mark">&#10003;</div>
<div style="flex:1">
<p><b>Decision: Relational vs Graph (RDF):</b></p>
<table class="vtab">
<tr>
<th></th>
<th class="win">Relational (PostgreSQL)</th>
<th class="lose">Graph (RDF)</th>
</tr>
<tr>
<td class="crit">Average report, today</td>
<td><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">~6 seconds</a></td>
<td class="lose"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">~1.6 hours</a></td>
</tr>
<tr>
<td class="crit">Average report at 6 TB</td>
<td><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">~30 seconds</a></td>
<td class="lose"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">~12 hours</a></td>
</tr>
<tr>
<td class="crit">Hardware at 6 TB</td>
<td><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#hardware" target="_blank" rel="noopener">one ordinary server</a></td>
<td class="lose"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#hardware" target="_blank" rel="noopener">a 14-server cluster</a></td>
</tr>
<tr>
<td class="crit">Ontology</td>
<td colspan="2">remains accessible with Relational (PostgreSQL)</td>
</tr>
</table>
</div>
</div>
<div class="kpis">
<div class="kpi"><div class="v"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#kpis" target="_blank" rel="noopener">1,000x</a></div><div class="l">faster on the average everyday report</div></div>
<div class="kpi"><div class="v"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#kpis" target="_blank" rel="noopener">$494k saved</a></div><div class="l">hardware at 6 TB: $33.6k for PGSQL server instead of a $527.7k RDF cluster</div></div>
<div class="kpi"><div class="v"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#kpis" target="_blank" rel="noopener">1 vs 14</a></div><div class="l">servers needed at 6 TB</div></div>
</div>
<div class="card">
<h2>Side by side</h2>
<table class="duel">
<tr>
<th></th>
<th class="win">Relational (PostgreSQL)<br><span class="pill">&#10003; recommended</span></th>
<th class="lose">Graph (RDF triplestore)<br><span class="pill">&#10007; rejected at scale</span></th>
<th></th>
</tr>
<tr>
<td class="crit">Average report, today<small>at 600 GB</small></td>
<td><span class="ok">&#10003;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">about 6 seconds</a></td>
<td><span class="no">&#10007;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">about 1.6 hours; 3 of 7 tasks need more than a day</a></td>
<td class="more"><a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html#proj" target="_blank" rel="noopener">details &rarr;</a></td>
</tr>
<tr>
<td class="crit">Average report at 6 TB<small>the 3-year horizon</small></td>
<td><span class="ok">&#10003;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">about 30 seconds</a></td>
<td><span class="no">&#10007;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">about 12 hours; 5 of 7 tasks need a day or more</a></td>
<td class="more"><a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html#proj" target="_blank" rel="noopener">details &rarr;</a></td>
</tr>
<tr>
<td class="crit">Servers needed today<small>at 600 GB</small></td>
<td><span class="ok">&#10003;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#hardware" target="_blank" rel="noopener">one standard server, $10.5k</a></td>
<td><span class="no">&#10007;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#hardware" target="_blank" rel="noopener">already two servers with 1.4 TB of memory, $58.7k</a></td>
<td class="more"><a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html#hw" target="_blank" rel="noopener">details &rarr;</a></td>
</tr>
<tr>
<td class="crit">Servers at 6 TB</td>
<td><span class="ok">&#10003;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#hardware" target="_blank" rel="noopener">still one server, $33.6k</a></td>
<td><span class="no">&#10007;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#hardware" target="_blank" rel="noopener">a 14-server cluster, $527.7k</a></td>
<td class="more"><a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html#hw" target="_blank" rel="noopener">details &rarr;</a></td>
</tr>
<tr>
<td class="crit">Disk for the same data<small>at 6 TB</small></td>
<td><span class="ok">&#10003;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#storage" target="_blank" rel="noopener">21 TB</a></td>
<td><span class="no">&#10007;</span><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#storage" target="_blank" rel="noopener">91 TB - the graph form stores the same data 4.4x bigger</a></td>
<td class="more"><a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html#storage" target="_blank" rel="noopener">details &rarr;</a></td>
</tr>
</table>
</div>
<div class="card">
<h2>Average everyday report</h2>
<div class="panels">
<div class="panel">
<h3>Today &middot; 600 GB</h3>
<div class="brow" title="PostgreSQL, average of the 7 everyday tasks at 600 GB: about 6 seconds (projected)">
<div class="lab">PostgreSQL</div>
<div class="trk"><div class="fill pg" style="width:0.4%"></div><span class="inval in-track"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">6 s</a></span></div>
</div>
<div class="brow" title="RDF triplestore, average of the 7 everyday tasks at 600 GB: about 1.6 hours (projected)">
<div class="lab">Graph (RDF)</div>
<div class="trk"><div class="fill rdf" style="width:100%"></div><span class="inval in-fill"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">1.6 h</a></span></div>
</div>
</div>
<div class="panel">
<h3>In a few years &middot; 6 TB</h3>
<div class="brow" title="PostgreSQL, average of the 7 everyday tasks at 6 TB: about 30 seconds (projected)">
<div class="lab">PostgreSQL</div>
<div class="trk"><div class="fill pg" style="width:0.4%"></div><span class="inval in-track"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">30 s</a></span></div>
</div>
<div class="brow" title="RDF triplestore, average of the 7 everyday tasks at 6 TB: about 12 hours (projected)">
<div class="lab">Graph (RDF)</div>
<div class="trk"><div class="fill rdf" style="width:100%"></div><span class="inval in-fill"><a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#performance" target="_blank" rel="noopener">12 h</a></span></div>
</div>
</div>
</div>
<p class="muted" style="margin-top:14px">Bars are to scale within each panel - the blue bar is real, it is just that small. An everyday report that takes hours is a report nobody runs.</p>
</div>
<div class="card">
<h2>What about the other formats?</h2>
<p><b>CSV</b> stays exactly where it belongs - the raw archive coming off the instruments (best compression, byte-reproducible). <b>JSON-LD</b> stays the exchange format for sending data between laboratories. Neither is a working database, so neither was a candidate. <b>SQLite</b> is excellent for a single user <a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#notes" target="_blank" rel="noopener">up to ~600 GB</a> but does not scale to the shared 6 TB horizon.</p>
<p><b>And the ontology?</b> It survives intact: semantic, ontology-style access is provided as a <b>virtual layer on top of PostgreSQL</b> (e.g. Ontop OBDA) - the graph view without the graph database's cost.</p>
</div>
<div class="card">
<h2>Recommended next steps</h2>
<ol class="steps">
<li>Adopt <b>PostgreSQL as the system of record</b> for laboratory measurement data; keep the raw CSV archive as the source of truth for reprocessing.</li>
<li>Provide ontology and semantic queries as a <b>virtual RDF layer</b> over PostgreSQL - no data is copied into a triplestore.</li>
<li>Revisit hardware once volume approaches 6 TB: the measured plan is <a class="fact" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html#hardware" target="_blank" rel="noopener">one 16-core server with 551 GB RAM (~$33.6k at July 2026 prices)</a>.</li>
</ol>
</div>
<div class="foot">
Figures at 600 GB and 6 TB are model projections calibrated on a measured benchmark (7 everyday retrieval tasks x 5 storage formats, every result checksum-validated); the "average report" is the geometric mean of the seven tasks, the same averaging the study's scoring uses. Method, raw data and hardware price sources: <a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision_Facts.html" target="_blank" rel="noopener">facts ledger (every number traced to its artifact)</a> &middot; <a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html" target="_blank" rel="noopener">full interactive study</a> &middot; <a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation" target="_blank" rel="noopener">git.boskadoff.com/Public/LabDataStorageEvaluation</a>.<br>
Study: Mary Goncharenko (Tribology Laboratory, University of Florida; internship at Sandia National Laboratories), with the technical assistance of Vasiliy Goncharenko (SoftCreator, LLC).
</div>
</div>
</body>
</html>

View File

@@ -0,0 +1,131 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Decision Facts Ledger - Storage Decision</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=IBM+Plex+Mono:wght@400;500;600&family=IBM+Plex+Sans:wght@400;500;600;700&display=swap" rel="stylesheet">
<style>
*{box-sizing:border-box}
html,body{margin:0;padding:0}
body{font-family:"IBM Plex Sans",system-ui,sans-serif;color:#16181d;background:#eef0f2;-webkit-font-smoothing:antialiased}
.mono{font-family:"IBM Plex Mono",monospace}
header.top{display:flex;align-items:center;gap:14px;height:56px;padding:0 22px;background:#fff;border-bottom:1px solid #e2e5ea;position:sticky;top:0;z-index:5}
.logo{display:flex;align-items:center;justify-content:center;width:30px;height:30px;border-radius:7px;background:#16181d;color:#fff;font-family:"IBM Plex Mono",monospace;font-weight:600;font-size:12px;flex:none}
.htitle{display:flex;flex-direction:column;line-height:1.25;white-space:nowrap}
.htitle b{font-size:14.5px;font-weight:600;letter-spacing:-.01em}
.htitle span{font-size:11.5px;color:#838a95}
.hright{margin-left:auto;display:flex;align-items:center;gap:8px}
.chip{font-family:"IBM Plex Mono",monospace;font-size:11px;color:#838a95;background:#f3f5f7;border:1px solid #e6e8ec;padding:4px 10px;border-radius:6px;text-decoration:none}
a.chip:hover{border-color:#1f5fdb;color:#1f5fdb}
.page{max-width:1080px;margin:0 auto;padding:28px 24px 70px}
h1{margin:0 0 6px;font-size:22px;font-weight:700;letter-spacing:-.02em}
h2{margin:26px 0 10px;font-size:15px;font-weight:600}
p{margin:0 0 10px;font-size:13px;color:#3d424b;line-height:1.6;max-width:940px}
.muted{color:#838a95}
a{color:#1f5fdb;text-decoration:none}
a:hover{text-decoration:underline}
.card{background:#fff;border:1px solid #e6e8ec;border-radius:11px;padding:6px 18px 10px;margin-bottom:8px;overflow-x:auto}
table{width:100%;border-collapse:collapse;min-width:760px}
th{font-size:10.5px;text-transform:uppercase;letter-spacing:.04em;color:#838a95;font-weight:600;text-align:left;padding:9px 12px;border-bottom:1px solid #eef0f2}
td{font-size:12.5px;color:#3d424b;padding:9px 12px;border-bottom:1px solid #f4f5f7;vertical-align:top;line-height:1.55}
tr:last-child td{border-bottom:0}
td.id{font-family:"IBM Plex Mono",monospace;font-weight:600;color:#16181d;white-space:nowrap}
td.val{font-family:"IBM Plex Mono",monospace;white-space:nowrap}
.foot{font-size:11.5px;color:#a4abb4;border-top:1px solid #e6e8ec;padding-top:14px;margin-top:26px;line-height:1.7}
</style>
</head>
<body>
<header class="top">
<div class="logo">SE</div>
<div class="htitle">
<b>Decision Facts Ledger</b>
<span>Every fact on the Storage Decision page &middot; value, formula, source</span>
</div>
<div class="hright">
<a class="chip" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision.html" target="_blank" rel="noopener">&larr; Storage Decision</a>
<a class="chip" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html" target="_blank" rel="noopener">Full study</a>
<a class="chip" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation" target="_blank" rel="noopener">Git repository</a>
</div>
</header>
<div class="page">
<h1>Decision Facts Ledger</h1>
<p>Every number shown on the <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision.html" target="_blank" rel="noopener">Storage Decision</a> executive summary, traced to the artifact that proves it. All source files are committed in the public repository and regenerate deterministically from seed 20260711. Projection tables carry one row per format x task x scale; "geometric mean" is the same averaging the study's scoring uses.</p>
<p class="muted">Sources: <span class="mono">t6</span> = <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6_projected_all_scales.csv</a> &middot; <span class="mono">t7</span> = <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7_hardware_all_scales.csv</a> &middot; <span class="mono">t1</span> = <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t1_storage_footprint.csv" target="_blank" rel="noopener">t1_storage_footprint.csv</a> &middot; <span class="mono">prices</span> = <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/config/hw_prices.yaml" target="_blank" rel="noopener">hw_prices.yaml</a> &middot; <span class="mono">spec</span> = <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/docs/specs/10_extrapolation_model.md" target="_blank" rel="noopener">10_extrapolation_model.md</a></p>
<h2 id="premises">Premises</h2>
<div class="card">
<table>
<tr><th>ID</th><th>Fact on the page</th><th>Exact value</th><th>How obtained</th><th>Source</th></tr>
<tr><td class="id">P1</td><td>"600 GB of laboratory data today"</td><td class="val">600 GB</td><td>The laboratory's current operational volume; also the first extrapolation scale of the study (scale_gb = 600).</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/docs/specs/10_extrapolation_model.md" target="_blank" rel="noopener">spec</a>, <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
<tr><td class="id">P2</td><td>"6 TB within a few years"</td><td class="val">6,000 GB</td><td>The largest extrapolation scale of the study (scale_gb = 6000), the growth horizon.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/docs/specs/10_extrapolation_model.md" target="_blank" rel="noopener">spec</a>, <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
</table>
</div>
<h2 id="performance">Performance (average everyday report)</h2>
<div class="card">
<table>
<tr><th>ID</th><th>Fact on the page</th><th>Exact value</th><th>How obtained</th><th>Source</th></tr>
<tr><td class="id">F1</td><td>PostgreSQL, average report today: "~6 seconds"</td><td class="val">5.66 s</td><td>Geometric mean of the 7 projected_wall_s values where format = PostgreSQL, scale_gb = 600.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
<tr><td class="id">F2</td><td>PostgreSQL, average report at 6 TB: "~30 seconds"</td><td class="val">29.6 s</td><td>Geometric mean, format = PostgreSQL, scale_gb = 6000.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
<tr><td class="id">F3</td><td>Graph (RDF), average report today: "~1.6 hours"</td><td class="val">5,902 s = 1.64 h</td><td>Geometric mean, format = RDF triplestore, scale_gb = 600.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
<tr><td class="id">F4</td><td>Graph (RDF), average report at 6 TB: "~12 hours"</td><td class="val">43,157 s = 11.99 h</td><td>Geometric mean, format = RDF triplestore, scale_gb = 6000.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
<tr><td class="id">F5</td><td>"3 of 7 tasks need more than a day" (RDF, today)</td><td class="val">q3 2.1 d &middot; q5 2.8 d &middot; q7 1.9 d</td><td>Count of RDF rows at scale_gb = 600 with projected_wall_s &gt; 86,400.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
<tr><td class="id">F6</td><td>"5 of 7 tasks need a day or more" (RDF, 6 TB)</td><td class="val">q2, q3, q4, q5, q7</td><td>Count of RDF rows at scale_gb = 6000 with projected_wall_s &gt;= 86,400.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
</table>
</div>
<h2 id="hardware">Hardware and cost</h2>
<div class="card">
<table>
<tr><th>ID</th><th>Fact on the page</th><th>Exact value</th><th>How obtained</th><th>Source</th></tr>
<tr><td class="id">F7</td><td>"one standard server, $10.5k" (PostgreSQL, today)</td><td class="val">$10,500 &middot; 1 node</td><td>Row format = PostgreSQL, scale_gb = 600: est_cost_usd, nodes; bill of materials in the configuration column.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7</a></td></tr>
<tr><td class="id">F8</td><td>"already two servers with 1.4 TB of memory, $58.7k" (RDF, today)</td><td class="val">$58,740 &middot; 2 nodes &middot; 1,404.1 GB RAM</td><td>Row format = RDF triplestore, scale_gb = 600: est_cost_usd, nodes, ram_gb.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7</a></td></tr>
<tr><td class="id">F9</td><td>"still one server, $33.6k" / "one 16-core server with 551 GB RAM" (PostgreSQL, 6 TB)</td><td class="val">$33,586 &middot; 1 node &middot; 16 cores &middot; 551.3 GB</td><td>Row format = PostgreSQL, scale_gb = 6000: est_cost_usd, nodes, cores, ram_gb.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7</a></td></tr>
<tr><td class="id">F10</td><td>"a 14-server cluster, $527.7k" (RDF, 6 TB)</td><td class="val">$527,732 &middot; 14 nodes</td><td>Row format = RDF triplestore, scale_gb = 6000: est_cost_usd, nodes.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7</a></td></tr>
<tr><td class="id">F11</td><td>"July 2026 prices"</td><td class="val">as_of 2026-07-11</td><td>Street prices with per-component source URLs recorded in the operator-editable price file.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/config/hw_prices.yaml" target="_blank" rel="noopener">prices</a></td></tr>
</table>
</div>
<h2 id="storage">Storage</h2>
<div class="card">
<table>
<tr><th>ID</th><th>Fact on the page</th><th>Exact value</th><th>How obtained</th><th>Source</th></tr>
<tr><td class="id">F12</td><td>"21 TB" disk (PostgreSQL, 6 TB)</td><td class="val">20.855 TB</td><td>Row format = PostgreSQL, scale_gb = 6000: disk_tb (includes 30% free-space headroom).</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7</a></td></tr>
<tr><td class="id">F13</td><td>"91 TB" disk (RDF, 6 TB)</td><td class="val">91.268 TB</td><td>Row format = RDF triplestore, scale_gb = 6000: disk_tb.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7</a></td></tr>
<tr><td class="id">F14</td><td>"stores the same data 4.4x bigger"</td><td class="val">11.70 / 2.67 = 4.38</td><td>Ratio of the measured working-footprint coefficients vs raw CSV (RDF 11.70x over PostgreSQL 2.67x).</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t1_storage_footprint.csv" target="_blank" rel="noopener">t1</a></td></tr>
</table>
</div>
<h2 id="kpis">Headline figures</h2>
<div class="card">
<table>
<tr><th>ID</th><th>Fact on the page</th><th>Exact value</th><th>How obtained</th><th>Source</th></tr>
<tr><td class="id">K1</td><td>"1,000x faster on the average everyday report"</td><td class="val">1,043x today &middot; 1,459x at 6 TB</td><td>F3 / F1 = 5,902 / 5.66; F4 / F2 = 43,157 / 29.6. Stated conservatively as 1,000x.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
<tr><td class="id">K2</td><td>"$494k saved" (hardware at 6 TB)</td><td class="val">$494,146</td><td>F10 - F9 = 527,732 - 33,586.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7</a></td></tr>
<tr><td class="id">K3</td><td>"1 vs 14 servers needed at 6 TB"</td><td class="val">nodes: 1 vs 14</td><td>nodes column, rows PostgreSQL/6000 and RDF triplestore/6000.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t7_hardware_all_scales.csv" target="_blank" rel="noopener">t7</a></td></tr>
</table>
</div>
<h2 id="notes">Notes</h2>
<div class="card">
<table>
<tr><th>ID</th><th>Fact on the page</th><th>Exact value</th><th>How obtained</th><th>Source</th></tr>
<tr><td class="id">N1</td><td>"SQLite is excellent for a single user up to ~600 GB"</td><td class="val">all 7 tasks OK at 600 GB; q5 = 1.6 h, q7 = 1.5 h at 1.2 TB</td><td>SQLite rows in t6: every flag OK at scale_gb = 600; its single-threaded full scans exceed the 1-hour bound from scale_gb = 1200.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/src/branch/master/out/report/tables/t6_projected_all_scales.csv" target="_blank" rel="noopener">t6</a></td></tr>
<tr><td class="id">N2</td><td>Method behind every projected number</td><td class="val">measured 35-cell benchmark</td><td>Scaling laws calibrated on the measured corpus; every measured run checksum-validated. Full methodology and measured medians in the interactive study.</td><td><a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Evaluation.html" target="_blank" rel="noopener">full study</a></td></tr>
</table>
</div>
<div class="foot">
Part of the LabDataStorageEvaluation study &middot; run of 2026-07-11, seed 20260711 &middot; canonical source: <a href="https://git.boskadoff.com/Public/LabDataStorageEvaluation" target="_blank" rel="noopener">git.boskadoff.com/Public/LabDataStorageEvaluation</a>.<br>
Study: Mary Goncharenko (Tribology Laboratory, University of Florida; internship at Sandia National Laboratories), with the technical assistance of Vasiliy Goncharenko (SoftCreator, LLC).
</div>
</div>
</body>
</html>

View File

@@ -0,0 +1,814 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Storage Format Evaluation - Tribology Lab Data</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=IBM+Plex+Mono:wght@400;500;600&family=IBM+Plex+Sans:wght@400;500;600;700&display=swap" rel="stylesheet">
<style>
*{box-sizing:border-box}
html,body{margin:0;padding:0}
body{font-family:"IBM Plex Sans",system-ui,sans-serif;color:#16181d;background:#eef0f2;-webkit-font-smoothing:antialiased}
::-webkit-scrollbar{width:10px;height:10px}
::-webkit-scrollbar-thumb{background:#cfd4da;border-radius:6px;border:2px solid #eef0f2}
::-webkit-scrollbar-track{background:transparent}
.mono{font-family:"IBM Plex Mono",monospace}
.app{display:flex;flex-direction:column;height:100vh;width:100%;overflow:hidden;background:#eef0f2}
header.top{display:flex;align-items:center;gap:14px;height:56px;flex:none;padding:0 22px;background:#fff;border-bottom:1px solid #e2e5ea}
.logo{display:flex;align-items:center;justify-content:center;width:30px;height:30px;border-radius:7px;background:#16181d;color:#fff;font-family:"IBM Plex Mono",monospace;font-weight:600;font-size:12px}
.htitle{display:flex;flex-direction:column;line-height:1.25;white-space:nowrap}
.htitle b{font-size:14.5px;font-weight:600;letter-spacing:-.01em}
.htitle span{font-size:11.5px;color:#838a95}
.hright{margin-left:auto;display:flex;align-items:center;gap:8px}
.chip{font-family:"IBM Plex Mono",monospace;font-size:11px;color:#838a95;background:#f3f5f7;border:1px solid #e6e8ec;padding:4px 10px;border-radius:6px}
.chip-ok{display:inline-flex;align-items:center;gap:6px;font-size:12px;color:#0f7a45;background:#e7f6ee;padding:4px 10px;border-radius:6px;font-weight:600}
.chip-ok i{width:7px;height:7px;border-radius:7px;background:#18a05a}
.frame{display:flex;flex:1;min-height:0}
nav.side{width:252px;flex:none;background:#fff;border-right:1px solid #e2e5ea;padding:14px 12px;display:flex;flex-direction:column;overflow-y:auto}
.navcap{font-family:"IBM Plex Mono",monospace;font-size:10px;letter-spacing:.1em;color:#a4abb4;padding:6px 11px 8px}
nav.side a{display:flex;align-items:center;gap:11px;width:100%;padding:8px 11px;border-radius:8px;cursor:pointer;text-decoration:none;font-size:13.5px;margin:1px 0;color:#475059;font-weight:500}
nav.side a .tag{width:24px;height:19px;border-radius:5px;display:flex;align-items:center;justify-content:center;font-family:"IBM Plex Mono",monospace;font-size:9px;font-weight:600;flex:none;background:#eceef1;color:#8a909b}
nav.side a.active{background:#eaf1fe;color:#16181d;font-weight:600}
nav.side a.active .tag{background:#1f5fdb;color:#fff}
.sidefoot{margin-top:auto;padding:14px 11px 6px;border-top:1px solid #eef0f2}
.sidefoot .cap{font-family:"IBM Plex Mono",monospace;font-size:10px;letter-spacing:.08em;color:#a4abb4;margin-bottom:8px}
.sidefoot .txt{font-size:12px;color:#6b727c;line-height:1.6}
main{flex:1;min-width:0;overflow-y:auto;padding:26px 30px 60px}
main section{display:none}
main section.on{display:block}
h1{margin:0;font-size:21px;font-weight:600;letter-spacing:-.02em}
.sub{margin:4px 0 18px;font-size:13px;color:#838a95;max-width:900px;line-height:1.55}
.card{background:#fff;border:1px solid #e6e8ec;border-radius:11px;padding:18px 20px;margin-bottom:14px}
.card h3{margin:0 0 10px;font-size:14px;font-weight:600}
.card p{margin:0 0 10px;font-size:13px;color:#3d424b;line-height:1.6;max-width:900px}
.card p:last-child{margin-bottom:0}
.card ul{margin:6px 0 4px;padding-left:20px}
.card li{font-size:13px;color:#3d424b;line-height:1.65;margin-bottom:6px;max-width:860px}
.muted{color:#838a95}
.kpis{display:grid;grid-template-columns:repeat(5,1fr);gap:14px;margin-bottom:16px}
.kpis.four{grid-template-columns:repeat(4,1fr)}
.kpi{background:#fff;border:1px solid #e6e8ec;border-radius:10px;padding:15px 16px}
.kpi .l{font-size:11px;color:#838a95;font-weight:500;text-transform:uppercase;letter-spacing:.04em}
.kpi .v{font-family:"IBM Plex Mono",monospace;font-size:24px;font-weight:600;margin-top:7px;line-height:1}
.kpi .s{font-size:11.5px;color:#a4abb4;margin-top:5px}
.v-bad{color:#b23b4e}.v-good{color:#0f7a45}.v-neutral{color:#16181d}
table.data{width:100%;border-collapse:collapse}
table.data th{font-size:10.5px;text-transform:uppercase;letter-spacing:.04em;color:#838a95;font-weight:600;text-align:left;padding:9px 12px;background:#f7f8fa;border-bottom:1px solid #eef0f2}
table.data td{font-size:12.5px;color:#3d424b;padding:10px 12px;border-bottom:1px solid #f4f5f7;vertical-align:top}
table.data th.num, table.data td.num{text-align:right;font-family:"IBM Plex Mono",monospace}
table.data td.mono{font-family:"IBM Plex Mono",monospace}
.hbars{display:flex;flex-direction:column;gap:11px;margin-top:6px}
.hbar{display:flex;align-items:center;gap:14px}
.hbar .lab{width:190px;flex:none;font-size:12.5px;color:#3d424b}
.hbar .lab small{display:block;color:#a4abb4;font-size:11px}
.hbar .trk{flex:1;background:#f1f3f6;border-radius:6px;overflow:hidden}
.hbar .fill{height:14px;border-radius:6px;min-width:2px}
.hbar .val{font-family:"IBM Plex Mono",monospace;font-size:12.5px;color:#16181d;width:130px;flex:none;text-align:right}
.legend{display:flex;gap:16px;flex-wrap:wrap;margin:2px 0 10px}
.legend span{display:inline-flex;align-items:center;gap:6px;font-size:11.5px;color:#5b6472}
.legend i{width:10px;height:10px;border-radius:3px;display:inline-block}
.stack{display:flex;align-items:center;gap:14px;margin-bottom:11px}
.stack .lab{width:190px;flex:none;font-size:12.5px;color:#3d424b}
.stack .trk{flex:1;display:flex;background:#f1f3f6;border-radius:6px;overflow:hidden;height:16px}
.stack .seg{height:16px}
.stack .val{font-family:"IBM Plex Mono",monospace;font-size:12.5px;width:70px;flex:none;text-align:right;font-weight:600}
.tabs{display:inline-flex;gap:4px;background:#eceef1;border-radius:8px;padding:3px;margin:0 0 14px}
.tabs button{border:0;background:transparent;font-family:"IBM Plex Sans",system-ui,sans-serif;font-size:12px;font-weight:600;color:#5b6472;padding:5px 12px;border-radius:6px;cursor:pointer}
.tabs button.on{background:#fff;color:#16181d;box-shadow:0 1px 3px rgba(20,30,50,.08)}
.callout{border-left:3px solid #1f5fdb;background:#f5f8fe;border-radius:0 9px 9px 0;padding:12px 16px;font-size:13px;color:#2b3038;line-height:1.6;margin:12px 0;max-width:900px}
.callout.warn{border-left-color:#b23b4e;background:#fdf3f5}
.callout.good{border-left-color:#18a05a;background:#f0faf4}
.pill{display:inline-flex;align-items:center;padding:2px 9px;border-radius:20px;font-size:10.5px;font-weight:600;white-space:nowrap}
.pill.ok{color:#0f7a45;background:#e7f6ee}
.pill.imp{color:#a06a12;background:#fdf3e0}
.pill.fail{color:#b23b4e;background:#fbe9ec}
.pill.rec{color:#0f7a45;background:#e7f6ee}
.pill.base{color:#5b6472;background:#eef0f2}
pre.env{background:#16181d;color:#dfe3e9;border-radius:9px;padding:14px 16px;font-family:"IBM Plex Mono",monospace;font-size:11.5px;line-height:1.65;overflow-x:auto;margin:10px 0 0}
.tree{font-family:"IBM Plex Mono",monospace;font-size:12px;line-height:1.8;color:#3d424b;background:#f7f8fa;border:1px solid #eef0f2;border-radius:9px;padding:14px 18px;margin-top:8px}
.grid2{display:grid;grid-template-columns:1fr 1fr;gap:14px}
@media(max-width:1100px){.grid2{grid-template-columns:1fr}.kpis{grid-template-columns:repeat(2,1fr)}}
.foot{font-size:11.5px;color:#a4abb4;border-top:1px solid #e6e8ec;padding-top:14px;margin-top:30px;line-height:1.7}
.src a, a.ext{color:#1f5fdb;text-decoration:none;word-break:break-all}
.src a:hover, a.ext:hover{text-decoration:underline}
a.chip{text-decoration:none}
a.chip:hover{border-color:#1f5fdb;color:#1f5fdb}
svg text{font-family:"IBM Plex Sans",system-ui,sans-serif}
</style>
</head>
<body>
<div class="app">
<header class="top">
<div class="logo">SE</div>
<div class="htitle">
<b>Storage Format Evaluation</b>
<span>Tribology lab data &middot; CSV vs JSON-LD vs SQLite vs PostgreSQL vs RDF</span>
</div>
<div class="hright">
<a class="chip" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation" target="_blank" rel="noopener">Git repository</a>
<span class="chip">7 queries x 5 formats &middot; 35 cells</span>
<span class="chip">seed 20260711</span>
<span class="chip-ok"><i></i>Measured 2026-07-11 &middot; all checksums OK</span>
</div>
</header>
<div class="frame">
<nav class="side" id="sidenav">
<div class="navcap">NAVIGATE</div>
<a href="#dashboard" class="active"><span class="tag">DB</span><span>Dashboard</span></a>
<a href="#lab"><span class="tag">LM</span><span>Lab Model &amp; Corpus</span></a>
<a href="#method"><span class="tag">ME</span><span>Methodology</span></a>
<a href="#storage"><span class="tag">SF</span><span>Storage Footprint</span></a>
<a href="#measured"><span class="tag">MR</span><span>Measured Results</span></a>
<a href="#proj"><span class="tag">PR</span><span>Projections 0.6-6 TB</span></a>
<a href="#hw"><span class="tag">HW</span><span>Hardware &amp; Cost</span></a>
<a href="#score"><span class="tag">SC</span><span>Scoring &amp; Verdict</span></a>
<a href="#about"><span class="tag">AB</span><span>About</span></a>
<div class="sidefoot">
<div class="cap">CONTEXT</div>
<div class="txt">Final results of the LabDataStorageEvaluation pipeline (tasks 01-11): a simulated Pt-Au tribology laboratory corpus benchmarked across five storage formats. Sources of truth: out/report/ and out/bench/.</div>
<div class="cap" style="margin-top:12px">AUTHOR &middot; SOURCE</div>
<div class="txt">Mary Goncharenko - University of Florida / Sandia National Laboratories.<br>
<a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation" target="_blank" rel="noopener">git.boskadoff.com/Public/LabDataStorageEvaluation</a></div>
</div>
</nav>
<main id="main">
<!-- ================= DASHBOARD ================= -->
<section id="dashboard" class="on">
<h1>Dashboard</h1>
<p class="sub">A 141.8 MiB physically plausible CSV corpus (the canonical raw format) was converted into four other storage formats and benchmarked on seven laboratory retrieval scenarios with subprocess-isolated resource metering; results were extrapolated to 600 GB, 1.2 TB and 6 TB. Every measured run validated its rows against canonical results (PostgreSQL cross-validated with SQLite, tolerance 1e-9). Short on time? The one-page management summary is <a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision.html" target="_blank" rel="noopener">Storage Decision</a>.</p>
<div class="kpis">
<div class="kpi"><div class="l">Benchmark cells</div><div class="v v-good">35 / 35</div><div class="s">OK; 0 invalid, 0 timeouts</div></div>
<div class="kpi"><div class="l">Corpus</div><div class="v v-neutral">141.8 MiB</div><div class="s">8,718 CSV files, byte-reproducible</div></div>
<div class="kpi"><div class="l">Winner @600 GB</div><div class="v v-good">58.6 / 80</div><div class="s">PostgreSQL weighted score</div></div>
<div class="kpi"><div class="l">Worst blow-up</div><div class="v v-bad">11.7x</div><div class="s">RDF store vs raw CSV size</div></div>
<div class="kpi"><div class="l">Session</div><div class="v v-neutral">916 s</div><div class="s">whole matrix, quiet machine</div></div>
</div>
<div class="callout good"><b>Recommendation (confirmed by data):</b> PostgreSQL as the system of record; CSV retained as the raw archive (best compressor, ~2.8x); hybrid JSON-LD for inter-lab exchange; RDF only as a virtual layer (e.g. Ontop OBDA) over PostgreSQL - a materialized triplestore fails at scale on speed, RAM and cost simultaneously.</div>
<div class="card">
<h3>Typical query time per format (geometric mean Q1-Q7)</h3>
<div id="dash-geo" class="hbars"></div>
<p class="muted" style="margin-top:12px">Left value: measured on the 150 MB corpus. Right value in parentheses: projected at 600 GB. Linear bars scaled within the measured values; every projected number in this document is labeled as projected.</p>
</div>
<div class="card">
<h3>Where each format breaks</h3>
<table class="data">
<thead><tr><th>Format</th><th>Strength</th><th>Breaking point</th><th>@600 GB flags</th></tr></thead>
<tbody id="dash-breaks"></tbody>
</table>
</div>
</section>
<!-- ================= LAB MODEL ================= -->
<section id="lab">
<h1>Lab Model &amp; Corpus</h1>
<p class="sub">A simulated tribology laboratory (Pt-Au LDRD study context, Sandia National Laboratories): Ti-6Al-4V coupons with a sputtered Cr adhesion layer and a Pt-Au composition-gradient coating, tested on a 6-probe tribometer in two environments. All data is generated deterministically from seed 20260711; regeneration is byte-identical (MANIFEST.csv sha256 proof).</p>
<div class="kpis four">
<div class="kpi"><div class="l">Deposition</div><div class="v v-neutral">4 &middot; 12 &middot; 588</div><div class="s">batches &middot; wafers &middot; coupons (7x7 grid)</div></div>
<div class="kpi"><div class="l">Friction testing</div><div class="v v-neutral">1,440</div><div class="s">tracks: 480 coupons x 3 replicates</div></div>
<div class="kpi"><div class="l">Bulk rows</div><div class="v v-neutral">4.32 M</div><div class="s">1.44M COF cycles + 2.88M loop points</div></div>
<div class="kpi"><div class="l">Characterization</div><div class="v v-neutral">235,200</div><div class="s">XRF points + nanoindentation, AFM, profilometry</div></div>
</div>
<div class="grid2">
<div class="card">
<h3>Data hierarchy</h3>
<div class="tree">batch (B721-B724, deposition matrix)
&nbsp;&nbsp;-&gt; wafer (x3 per batch)
&nbsp;&nbsp;&nbsp;&nbsp;-&gt; coupon (x49 per wafer; 480 friction + 108 reserve)
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;-&gt; xrf_map (400 pts) &middot; nanoindentation (25) &middot; afm &middot; profilometry (100)
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;-&gt; track (x3, counterface rotation)
&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;-&gt; cof_vs_cycle (1,000 rows) &middot; friction_loops (10x200) &middot; wear</div>
<p style="margin-top:10px">Runs R1-R4 map one-to-one onto batches; two environments (lab air / dry N2). The Q1 benchmark target track is <span class="mono">B722-W2-C13-T2</span>.</p>
</div>
<div class="card">
<h3>Physical models behind the numbers</h3>
<ul>
<li>COF: exponential run-in to a steady state depending on Au wt% and environment; noise floor 0.12.</li>
<li>Hardness / modulus: linear in Au wt% with measurement noise (25 indents per coupon).</li>
<li>Wear: Archard model, k0 = 1.5e-7 (lab air) / 6e-8 (dry N2) mm3/(N*m), log-normal noise.</li>
<li>Au gradient across wafer position; batch-dependent center 8 / 25 / 10 / 60 wt%.</li>
</ul>
<p>Benchmark queries therefore return physically meaningful answers (e.g. Q2 selects the Au 10 +/- 0.5 wt% composition window), not random noise.</p>
</div>
</div>
<div class="card">
<h3>The five storage formats under test</h3>
<table class="data">
<thead><tr><th>Format</th><th>Variant benchmarked</th><th>How Q1-Q7 run</th></tr></thead>
<tbody>
<tr><td><b>CSV</b> (canonical raw)</td><td>hierarchical file tree + MANIFEST</td><td>python csv module, streaming; Q1 = direct path read</td></tr>
<tr><td><b>JSON-LD</b></td><td>FULL (per-coupon files, 588); HYBRID kept for exchange</td><td>ijson event streaming with early exit</td></tr>
<tr><td><b>SQLite</b></td><td>single file, composite natural PKs WITHOUT ROWID, enforced FKs</td><td>SQL; precomputed track_summary for Q2-Q4/Q7</td></tr>
<tr><td><b>PostgreSQL 16</b></td><td>remote host, range-partitioned bulk tables (12+12), BRIN + b-tree</td><td>SQL; parallel scans (4 workers), track_summary matview</td></tr>
<tr><td><b>RDF</b></td><td>pyoxigraph on-disk store, 17,419,713 triples (compact modeling)</td><td>SPARQL; run-in/steady-state derived from raw cycles</td></tr>
</tbody>
</table>
</div>
</section>
<!-- ================= METHODOLOGY ================= -->
<section id="method">
<h1>Methodology</h1>
<p class="sub">The product of this study is its measurements; the protocol is designed so a sloppy number cannot survive it (docs/rules/bench-methodology.md).</p>
<div class="grid2">
<div class="card">
<h3>Measurement protocol</h3>
<ul>
<li><b>One subprocess per run</b> - RSS / CPU / read bytes attribute to exactly that cell; psutil samples the whole process tree every 50 ms (the venv launcher stub alone would report 4 MB).</li>
<li><b>1 warm-up + 3 measured runs</b> per cell; medians reported, min/max kept in raw data.</li>
<li><b>Randomized cell order</b> (deterministic shuffle from the corpus seed); order recorded.</li>
<li><b>30-minute hard timeout</b> per run; a timeout is recorded as data, never retried.</li>
<li><b>Correctness gates speed:</b> every run's rows are compared to the canonical expected results (sorted rows, float tolerance 1e-9); a fast run that fails the checksum is an invalid measurement.</li>
</ul>
</div>
<div class="card">
<h3>Fairness and honesty</h3>
<ul>
<li>Identical result-set contract for all five implementations of each query; no format answers a simplified question.</li>
<li>Each format plays its idiomatic strength (CSV direct path read, SQL indexes, matviews) - but no LIMIT/sampling that changes the answer.</li>
<li><b>Q5 must scan raw cycles in every format</b>; the PostgreSQL EXPLAIN plan is checked to prove track_summary was not used.</li>
<li>Cold-cache pass skipped (Windows host has no page-cache drop) - marked, not faked.</li>
<li>PostgreSQL is remote: client wall times include the LAN round-trip; server-side cost captured via pg_stat_statements deltas and Prometheus host windows.</li>
</ul>
</div>
</div>
<div class="card">
<h3>The seven retrieval scenarios</h3>
<table class="data">
<thead><tr><th>ID</th><th>Scenario</th><th class="num">Result rows</th></tr></thead>
<tbody id="method-queries"></tbody>
</table>
</div>
<div class="card">
<h3>Environment (recorded per session)</h3>
<pre class="env" id="method-env"></pre>
</div>
</section>
<!-- ================= STORAGE ================= -->
<section id="storage">
<h1>Storage Footprint</h1>
<p class="sub">Working footprint is what the query engine reads; archival is the compressed form for shelving. All archival sizes converge to 53-57 MB - the information content is the same, the working overhead is not.</p>
<div class="kpis four">
<div class="kpi"><div class="l">Most compact</div><div class="v v-good">0.83x</div><div class="s">SQLite: 123 MB, indexes included</div></div>
<div class="kpi"><div class="l">PostgreSQL</div><div class="v v-neutral">2.67x</div><div class="s">397 MB live; 34.1% of it is indexes</div></div>
<div class="kpi"><div class="l">RDF store</div><div class="v v-bad">11.70x</div><div class="s">1.74 GB for the same data</div></div>
<div class="kpi"><div class="l">Best archive</div><div class="v v-good">52.8 MB</div><div class="s">csv tree tar.gz (2.8x compression)</div></div>
</div>
<div class="card">
<h3>Working footprint, measured on the 150 MB corpus</h3>
<div id="st-work" class="hbars"></div>
</div>
<div class="grid2">
<div class="card">
<h3>Archival (compressed) footprint</h3>
<div id="st-arch" class="hbars"></div>
<p class="muted" style="margin-top:10px">pg_dump -Fc is internally gzip-compressed; RDF also has a Turtle serialization (596 MB, not the archival form). Hybrid JSON-LD is 1.9 MB of metadata + sourceFile links.</p>
</div>
<div class="card">
<h3>Projected working footprint (linear scaling)</h3>
<table class="data">
<thead><tr><th>Format</th><th class="num">@600 GB</th><th class="num">@1.2 TB</th><th class="num">@6 TB</th></tr></thead>
<tbody id="st-proj"></tbody>
</table>
<p class="muted" style="margin-top:10px">All projected. Storage scales linearly in data volume; coefficients are the measured ratios above.</p>
</div>
</div>
</section>
<!-- ================= MEASURED ================= -->
<section id="measured">
<h1>Measured Results <span class="pill base" style="vertical-align:3px">150 MB corpus &middot; warm cache</span></h1>
<p class="sub">Median of 3 isolated runs per cell. Pick a query and a metric; bars are linear with the actual value on each bar. Client peak RSS below ~50 ms of runtime under-reports (sampling cadence is spec-fixed at 50 ms) - wall times are exact.</p>
<div class="tabs" id="mq-tabs"></div>
<div class="tabs" id="mm-tabs" style="margin-left:10px"></div>
<div class="card">
<h3 id="m-title"></h3>
<div id="m-bars" class="hbars"></div>
</div>
<div class="card">
<h3>Full median matrix - wall time</h3>
<table class="data"><thead id="m-mat-h"></thead><tbody id="m-mat"></tbody></table>
</div>
<div class="card">
<h3>PostgreSQL server side (remote host, 4 cores / 3.8 GiB)</h3>
<table class="data">
<thead><tr><th>Query</th><th class="num">Server exec per call</th><th class="num">Shared blocks hit</th><th class="num">Read from disk</th></tr></thead>
<tbody id="m-pgsrv"></tbody>
</table>
<p class="muted" style="margin-top:10px">pg_stat_statements deltas over the session (4 calls per query). Zero disk reads: the whole 397 MB database sits in the 1 GB shared_buffers. Host peak CPU during the session: 9.2%; RAM delta ~5 MB - the server is nowhere near its limits at this scale.</p>
</div>
</section>
<!-- ================= PROJECTIONS ================= -->
<section id="proj">
<h1>Projections <span class="pill imp" style="vertical-align:3px">every number here is projected, not measured</span></h1>
<p class="sub">Per-cell scaling laws calibrated on the measured 150 MB point after subtracting each format's harness floor (interpreter start, imports, connection). Flags: IMPRACTICAL over 1 hour, FAIL over 24 hours. Pick a target scale.</p>
<div class="tabs" id="p-tabs"></div>
<div class="card">
<h3 id="p-title"></h3>
<table class="data"><thead id="p-mat-h"></thead><tbody id="p-mat"></tbody></table>
<p class="muted" style="margin-top:10px" id="p-note"></p>
</div>
<div class="grid2">
<div class="card">
<h3>Scaling law per access pattern</h3>
<table class="data">
<thead><tr><th>Law</th><th>Growth</th><th>Applied to</th></tr></thead>
<tbody>
<tr><td class="mono">O(1)</td><td>flat</td><td>CSV / JSON-LD Q1 (one file, size fixed)</td></tr>
<tr><td class="mono">O(log n)</td><td>index depth</td><td>SQLite / PostgreSQL / RDF Q1</td></tr>
<tr><td class="mono">O(log n + k)</td><td>~linear (result set grows)</td><td>relational + RDF Q2, Q4</td></tr>
<tr><td class="mono">O(k log n)</td><td>linear x depth</td><td>relational Q3, Q6; RDF Q6</td></tr>
<tr><td class="mono">O(n)</td><td>linear, 1 core</td><td>all CSV/JSON scans; SQLite Q5/Q7; RDF Q3/Q5/Q7</td></tr>
<tr><td class="mono">O(n / cores)</td><td>linear / parallel workers</td><td>PostgreSQL Q5, Q7 (partitioned scan)</td></tr>
</tbody>
</table>
</div>
<div class="card">
<h3>Full-scan degradation (Q5 + Q7 geometric mean)</h3>
<div id="p-degr"></div>
<p class="muted" style="margin-top:8px">First column measured, the rest projected. PostgreSQL rides its parallel partitioned scan; single-threaded engines grow linearly.</p>
</div>
</div>
<div class="callout">PostgreSQL is the only format with all seven queries below one hour at every scale - Q5 at 6 TB projects to 37 minutes on 16 cores. SQLite holds to 600 GB, then its single-threaded scans pass the hour mark. JSON-LD full scans reach 4.6 days at 6 TB; RDF Q5 reaches 27.8 days.</div>
</section>
<!-- ================= HARDWARE ================= -->
<section id="hw">
<h1>Hardware &amp; Cost <span class="pill base" style="vertical-align:3px">street prices as of 2026-07-11</span></h1>
<p class="sub">Bill-of-materials costing: whole 64 GB RDIMM modules, whole 7.68 TB NVMe drives, a priced 1U chassis; disk includes 30% free-space headroom. RAM model: relational hot set = 10% of index size + 4 GB working set; RDF hot set = 20% of the store; streaming formats are flat. Single-node RAM ceiling: 1 TB.</p>
<div class="tabs" id="h-tabs"></div>
<div class="card">
<h3 id="h-title"></h3>
<table class="data">
<thead><tr><th>Format</th><th>Exact configuration</th><th class="num">Chassis</th><th class="num">RAM</th><th class="num">Disk</th><th class="num">Extra nodes</th><th class="num">Total</th></tr></thead>
<tbody id="h-bom"></tbody>
</table>
</div>
<div class="grid2">
<div class="card">
<h3>Total cost comparison</h3>
<div id="h-bars" class="hbars"></div>
</div>
<div class="card">
<h3>Component prices &amp; sources <span class="pill base">retrieved 2026-07-11</span></h3>
<ul class="src">
<li>64 GB DDR5-5600 ECC RDIMM (A-Tech, 2x32GB): <b>$1,887</b><br><a href="https://atechmemory.com/collections/ddr5-memory-ram">atechmemory.com/collections/ddr5-memory-ram</a></li>
<li>7.68 TB NVMe U.2 1DWPD enterprise SSD (Cloud Ninjas): <b>$3,995</b><br><a href="https://cloudninjas.com/products/new-7-68tb-nvme-u-2-1dwpd-sie-2-5-enterprise-solid-state-drive-for-14th-15th-16th-gen-dell">cloudninjas.com/products/new-7-68tb-nvme-u-2-1dwpd...</a></li>
<li>1U AMD EPYC chassis (Broadberry starting configs): AS-1015CS-TNR <b>$4,618</b>, AS-1115CS-TNR <b>$6,272</b><br><a href="https://www.broadberry.com/amd-epyc-9004-supermicro-servers">broadberry.com/amd-epyc-9004-supermicro-servers</a></li>
<li>Market context - server DRAM 2025-2026 price surge:<br><a href="https://www.techpowerup.com/342331/server-dram-pricing-jumps-50-only-70-of-orders-getting-filled">techpowerup.com/342331/server-dram-pricing-jumps-50...</a></li>
</ul>
<p class="muted" style="margin-top:10px">Spot prices during a documented DRAM surge; re-cost by editing out/config/hw_prices.yaml and re-running tasks 10-11.</p>
</div>
</div>
<div class="card">
<h3>Cost vs speed at 600 GB (projected)</h3>
<div id="h-scatter"></div>
</div>
</section>
<!-- ================= SCORING ================= -->
<section id="score">
<h1>Scoring &amp; Verdict</h1>
<p class="sub">Subscores are 10 x best/value per metric, weighted: search speed x5 (0-50, inverse geometric mean of Q1-Q7), RAM economy x2 (0-20), disk economy x1 (0-10); maximum 80. A FAIL projection at 600 GB zeroes the search subscore.</p>
<div class="grid2">
<div class="card">
<h3>Measured scale (150 MB)</h3>
<div class="legend"><span><i style="background:#1f5fdb"></i>search x5</span><span><i style="background:#e0892e"></i>RAM x2</span><span><i style="background:#0f9d8a"></i>disk x1</span></div>
<div id="sc-measured"></div>
<p class="muted" style="margin-top:8px">At laptop scale SQLite is the honest winner: near-instant indexed queries in one file with the smallest footprint.</p>
</div>
<div class="card">
<h3>Projected at 600 GB</h3>
<div class="legend"><span><i style="background:#1f5fdb"></i>search x5</span><span><i style="background:#e0892e"></i>RAM x2</span><span><i style="background:#0f9d8a"></i>disk x1</span></div>
<div id="sc-proj"></div>
<p class="muted" style="margin-top:8px">At production scale the ranking flips: parallel indexed retrieval dominates the weights, and PostgreSQL takes the lead.</p>
</div>
</div>
<div class="card">
<h3>Use-case mapping</h3>
<table class="data">
<thead><tr><th>Use case</th><th>Recommended format</th></tr></thead>
<tbody>
<tr><td>Interactive analysis (joins, aggregations)</td><td><b>PostgreSQL</b>; SQLite acceptable single-user up to ~600 GB</td></tr>
<tr><td>Report generation (repeated summaries)</td><td><b>PostgreSQL</b> (track_summary materialized view)</td></tr>
<tr><td>Search / filtering</td><td><b>PostgreSQL or SQLite</b> (indexed); flat formats need full scans</td></tr>
<tr><td>Archiving</td><td><b>CSV tree + tar.gz</b> (canonical raw) plus pg_dump of the system of record</td></tr>
<tr><td>Inter-lab exchange</td><td><b>Hybrid JSON-LD</b> (metadata + sourceFile links to CSV)</td></tr>
<tr><td>Semantic / ontology queries</td><td><b>Virtual RDF layer</b> over PostgreSQL (e.g. Ontop OBDA), not a materialized triplestore</td></tr>
</tbody>
</table>
</div>
<div class="card">
<h3>Limitations</h3>
<ul>
<li>Single-node measurements on one Windows host + one remote PostgreSQL host; no cluster variance.</li>
<li>Simulated corpus (physically plausible, fixed seed); real instrument data may distribute differently.</li>
<li>Extrapolation is analytic, calibrated on one 150 MB point; no intermediate-scale validation runs.</li>
<li>Warm-cache only (no page-cache drop on Windows); harness floor assumed constant across scales.</li>
<li>RDF numbers reflect compact modeling and client-side derivation of run-in/steady-state; a SPARQL-side aggregation engine could shift, not remove, the scan penalty.</li>
<li>Hardware prices are spot street prices (2026-07-11) during a DRAM price surge.</li>
</ul>
</div>
<div class="foot">
Generated from the LabDataStorageEvaluation pipeline artifacts (run of 2026-07-11, seed 20260711): out/report/REPORT.md, out/bench/results_median.csv, extrapolation.csv, hardware_sizing.csv, storage_sizes.csv, out/config/hw_prices.yaml. The full decision document with charts is out/report/REPORT.md; this page is the interactive companion. Related study: <a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_FK_Architecture.html" target="_blank" rel="noopener">Foreign-Key Architecture Study</a>.
</div>
</section>
<!-- ================= ABOUT ================= -->
<section id="about">
<h1>About</h1>
<p class="sub">Attribution and provenance for this document. It is distributed as a single standalone HTML file; everything needed to verify or reproduce the numbers is referenced below.</p>
<div class="card">
<h3>Author</h3>
<p>This study was conducted by <b>Mary Goncharenko</b>, PhD student at the Tribology Laboratory, University of Florida, Gainesville, FL, during a research internship at the tribology laboratory of Sandia National Laboratories, with an assignment on laboratory ontology development.</p>
<p>The research was carried out with the technical assistance of her father, <b>Vasiliy Goncharenko</b>, Chief Technology Architect at SoftCreator, LLC (enterprise data architecture across relational and semantic stores; 35 years in technology).</p>
<p class="muted">Profiles: <a class="ext" href="https://www.linkedin.com/in/marygoncharenko/" target="_blank" rel="noopener">linkedin.com/in/marygoncharenko</a> &middot; <a class="ext" href="https://linkedin.com/in/vasiliy-goncharenko" target="_blank" rel="noopener">linkedin.com/in/vasiliy-goncharenko</a></p>
</div>
<div class="card">
<h3>Source repository</h3>
<p>The complete evaluation pipeline - task specifications, binding conventions, source code, and the generated report - is publicly available at:</p>
<p class="mono"><a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation" target="_blank" rel="noopener">https://git.boskadoff.com/Public/LabDataStorageEvaluation</a></p>
<p>This page is the interactive companion to the repository's <span class="mono">out/report/REPORT.md</span>; all numbers embedded here come from the pipeline run of 2026-07-11. The corpus regenerates byte-identically from seed 20260711 (MANIFEST.csv sha256 proof), so every figure is independently reproducible from the repository alone.</p>
</div>
<div class="card">
<h3>Data statement</h3>
<p>The measurement corpus is simulated: physically plausible models of a Pt-Au LDRD tribology study context at Sandia National Laboratories, generated from a fixed random seed. It contains no experimental measurements and no export-controlled data. Hardware prices cited in the cost model are public street prices retrieved 2026-07-11 (sources in the Hardware &amp; Cost section).</p>
<p>Companion pages in the same repository: <a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_Storage_Decision.html" target="_blank" rel="noopener">Storage Decision (one-page executive summary)</a> &middot; <a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation/raw/branch/master/docs/research/Tribology_FK_Architecture.html" target="_blank" rel="noopener">Foreign-Key Architecture Study</a>.</p>
</div>
<div class="foot">Standalone distributable file. If you received this document outside the repository, the canonical source is <a class="ext" href="https://git.boskadoff.com/Public/LabDataStorageEvaluation" target="_blank" rel="noopener">git.boskadoff.com/Public/LabDataStorageEvaluation</a>.</div>
</section>
</main>
</div>
</div>
<script>
"use strict";
/* ---------- embedded data (from the 2026-07-11 pipeline run) ---------- */
const FMT = ["csv","json","sqlite","pg","rdf"];
const NAME = {csv:"CSV", json:"JSON-LD", sqlite:"SQLite", pg:"PostgreSQL", rdf:"RDF triplestore"};
const COLOR = {csv:"#9aa6b6", json:"#e0892e", sqlite:"#0f9d8a", pg:"#1f5fdb", rdf:"#c0476b"};
const QN = [1,2,3,4,5,6,7];
const QDESC = {
1:["COF vs cycle curve for one track (B722-W2-C13-T2)",1000],
2:["Steady-state COF for coupons with mean Au = 10 +/- 0.5 wt%",33],
3:["Hardness vs steady-state COF across all batches",480],
4:["Anomaly filter: dry N2, 100 mN, cof_ss > 0.20",469],
5:["Mean run-in cycles per batch over ALL raw cycles (no summaries)",4],
6:["Wear volume vs load per coupon",480],
7:["Stribeck-style mean COF by (environment, speed x load)",2]
};
/* measured medians, per format: [q1..q7] */
const WALL = {
csv:[0.06,0.1661,1.1521,0.715,1.0948,0.2756,1.2735],
json:[0.1113,1.4788,9.9408,9.9218,9.9623,9.7927,10.0781],
sqlite:[0.0617,0.0607,0.1134,0.0604,0.7659,0.0612,0.7094],
pg:[0.2214,0.2206,0.2211,0.2232,0.4405,0.2212,0.3342],
rdf:[0.1131,2.8054,46.1273,19.8245,59.5884,0.1682,40.5534]
};
const RSS = {
csv:[3.9,20.5,21.1,20.9,20.7,20.2,21.0],
json:[22.2,23.9,24.4,24.5,24.4,24.1,24.1],
sqlite:[3.9,3.9,4.1,3.9,25.0,3.9,25.0],
pg:[42.2,42.1,42.0,41.9,42.0,41.6,42.1],
rdf:[31.6,85.4,263.9,167.9,264.0,45.9,263.6]
};
const READ = {
csv:[0,2.45,31.76,16.82,31.78,1.23,32.47],
json:[1.4,53.08,297.57,297.75,297.57,297.82,297.94],
sqlite:[0,0,0,0,117.91,0,92.97],
pg:[4.08,4.08,4.08,4.08,4.08,4.08,4.08],
rdf:[3.62,375.56,5445.1,2694.56,5390.46,15.92,5380.51]
};
/* projections: scale -> format -> [[seconds,flag] x7]; flag "" | "I" | "F" */
const PROJ = {
600: {
csv:[[0.06,""],[428.301,""],[4408.0,"I"],[2643.775,""],[4176.725,"I"],[870.266,""],[4897.995,"I"]],
json:[[0.111,""],[5519.622,"I"],[39673.993,"I"],[39597.306,"I"],[39760.772,"I"],[39076.231,"I"],[40228.164,"I"]],
sqlite:[[0.062,""],[1.271,""],[339.238,""],[0.06,""],[2847.603,""],[5.18,""],[2619.558,""]],
pg:[[0.222,""],[0.221,""],[3.42,""],[10.715,""],[222.111,""],[4.06,""],[114.849,""]],
rdf:[[0.113,""],[10866.789,"I"],[185722.881,"F"],[79559.375,"I"],[240054.647,"F"],[352.729,""],[163225.475,"F"]]
},
1200: {
csv:[[0.06,""],[856.543,""],[8815.94,"I"],[5287.489,"I"],[8353.391,"I"],[1740.472,""],[9795.931,"I"]],
json:[[0.111,""],[11039.133,"I"],[79347.876,"I"],[79194.5,"I"],[79521.432,"I"],[78152.351,"I"],[80456.218,"I"]],
sqlite:[[0.063,""],[2.482,""],[699.328,""],[0.06,""],[5695.146,"I"],[10.615,""],[5239.055,"I"]],
pg:[[0.222,""],[0.221,""],[6.817,""],[21.209,""],[444.001,""],[8.137,""],[229.477,""]],
rdf:[[0.113,""],[21733.465,"I"],[371445.649,"F"],[159118.636,"F"],[480109.181,"F"],[727.088,""],[326450.837,"F"]]
},
6000: {
csv:[[0.06,""],[4282.474,"I"],[44079.46,"I"],[26437.206,"I"],[41766.715,"I"],[8702.119,"I"],[48979.414,"I"]],
json:[[0.111,""],[55195.222,"I"],[396738.933,"F"],[395972.054,"F"],[397606.717,"F"],[390761.312,"F"],[402280.643,"F"]],
sqlite:[[0.063,""],[12.169,""],[3739.195,"I"],[0.06,""],[28475.49,"I"],[56.5,""],[26195.034,"I"]],
pg:[[0.222,""],[0.221,""],[35.495,""],[105.162,""],[2219.125,""],[42.55,""],[1146.503,""]],
rdf:[[0.113,""],[108666.874,"F"],[1857227.792,"F"],[795592.727,"F"],[2400545.453,"F"],[3887.402,"I"],[1632253.733,"F"]]
}
};
/* storage bytes */
const WORK = {csv:148654472, json:311300222, sqlite:123400192, pg:397467648, rdf:1739414085};
const ARCH = {csv:[52827471,"tar.gz of the tree"], json:[57162674,"tar.gz (full + hybrid)"], sqlite:[55149977,"gzip of tribo.db"], pg:[53054380,"pg_dump -Fc"], rdf:[126487650,"N-Triples gz"]};
const COEFF = {csv:1.0, json:2.09, sqlite:0.83, pg:2.67, rdf:11.70};
/* hardware sizing: scale -> format -> {modules,drives,nodes,cores,base,ram,disk,extra,total} */
const HW = {
600: {
csv:{m:1,d:1,n:1,c:16,base:4618,ram:1887,disk:3995,extra:0,total:10500},
json:{m:1,d:1,n:1,c:16,base:4618,ram:1887,disk:3995,extra:0,total:10500},
sqlite:{m:1,d:1,n:1,c:16,base:4618,ram:1887,disk:3995,extra:0,total:10500},
pg:{m:1,d:1,n:1,c:16,base:4618,ram:1887,disk:3995,extra:0,total:10500},
rdf:{m:22,d:2,n:2,c:16,base:4618,ram:41514,disk:7990,extra:4618,total:58740}
},
1200: {
csv:{m:1,d:1,n:1,c:16,base:4618,ram:1887,disk:3995,extra:0,total:10500},
json:{m:1,d:1,n:1,c:16,base:4618,ram:1887,disk:3995,extra:0,total:10500},
sqlite:{m:1,d:1,n:1,c:16,base:4618,ram:1887,disk:3995,extra:0,total:10500},
pg:{m:2,d:1,n:1,c:16,base:4618,ram:3774,disk:3995,extra:0,total:12387},
rdf:{m:44,d:3,n:3,c:16,base:4618,ram:83028,disk:11985,extra:9236,total:108867}
},
6000: {
csv:{m:1,d:2,n:1,c:16,base:4618,ram:1887,disk:7990,extra:0,total:14495},
json:{m:1,d:3,n:1,c:16,base:4618,ram:1887,disk:11985,extra:0,total:18490},
sqlite:{m:3,d:1,n:1,c:16,base:4618,ram:5661,disk:3995,extra:0,total:14274},
pg:{m:9,d:3,n:1,c:16,base:4618,ram:16983,disk:11985,extra:0,total:33586},
rdf:{m:220,d:12,n:14,c:16,base:4618,ram:415140,disk:47940,extra:60034,total:527732}
}
};
/* weighted scores: [search, ram, disk, total] */
const SCORES = {
measured:{csv:[15.5,20.0,8.3,43.8], json:[1.7,17.2,4.0,22.9], sqlite:[50.0,16.9,10.0,76.9], pg:[26.2,10.0,3.1,39.3], rdf:[1.3,1.6,0.7,3.6]},
proj:{csv:[0.6,20.0,8.3,28.9], json:[0.1,20.0,4.0,24.0], sqlite:[21.8,15.2,10.0,47.0], pg:[50.0,5.5,3.1,58.6], rdf:[0.0,0.2,0.7,0.9]}
};
/* pg server per-call exec ms + shared blocks (session deltas / 4 calls) */
const PGSRV = {1:[0.42,72],2:[0.43,1312],3:[4.92,6500],4:[1.09,140],5:[237.26,189692],6:[1.68,172],7:[107.14,32456]};
const ENVIRONMENT = [
"benchmark_host_os: Windows-2025Server-10.0.26100-SP0",
"benchmark_host_cpu: Intel64 Family 6 Model 141 (4 cores physical/logical), 8 GiB RAM",
"benchmark_host_disks: VMware Virtual NVMe Disk / SSD / 100 GB",
"python: 3.12.10 &middot; sqlite 3.49.1 &middot; psycopg 3.3.4 &middot; pyoxigraph 0.5.9 &middot; ijson 3.5.1 &middot; psutil 7.2.2",
"postgresql_server: PostgreSQL 16.14 (Ubuntu 24.04), remote host, 4 cores / 3.8 GiB",
"pg tuning: shared_buffers=1GB &middot; work_mem=32MB &middot; effective_cache_size=2500MB &middot; max_parallel_workers_per_gather=4",
"cold_cache_pass: skipped (Windows host, no page-cache drop)",
"pg_backend_sampling: remote server - Prometheus windows + pg_stat_statements deltas instead"
].join("\n");
/* ---------- formatting helpers ---------- */
function fmtTime(s){
if(s < 10) return s.toFixed(2) + " s";
if(s < 120) return s.toFixed(0) + " s";
if(s < 7200) return (s/60).toFixed(1) + " min";
if(s < 172800) return (s/3600).toFixed(1) + " h";
return (s/86400).toFixed(1) + " d";
}
function fmtSize(b){
if(b >= 1e12) return (b/1e12).toFixed(2) + " TB";
if(b >= 1e9) return (b/1e9).toFixed(2) + " GB";
return (b/1e6).toFixed(0) + " MB";
}
function fmtMoney(v){ return "$" + v.toLocaleString("en-US"); }
function esc(t){ return String(t).replace(/&/g,"&amp;").replace(/</g,"&lt;"); }
function el(id){ return document.getElementById(id); }
function flagPill(f){
if(f === "F") return ' <span class="pill fail">FAIL</span>';
if(f === "I") return ' <span class="pill imp">&gt; 1 h</span>';
return "";
}
function geomean(a){ return Math.exp(a.reduce((s,v)=>s+Math.log(v),0)/a.length); }
/* one horizontal bar row; linear scale against max */
function barRow(label, sub, frac, color, val){
return '<div class="hbar"><div class="lab">' + label + (sub ? '<small>' + sub + '</small>' : '') + '</div>' +
'<div class="trk"><div class="fill" style="width:' + Math.max(0.4, frac*100).toFixed(2) + '%;background:' + color + '"></div></div>' +
'<div class="val">' + val + '</div></div>';
}
function tabs(containerId, items, onPick){
const box = el(containerId);
box.innerHTML = items.map((t,i) => '<button data-i="' + i + '"' + (i===0 ? ' class="on"' : '') + '>' + t + '</button>').join("");
box.querySelectorAll("button").forEach(b => b.addEventListener("click", () => {
box.querySelectorAll("button").forEach(x => x.classList.remove("on"));
b.classList.add("on");
onPick(parseInt(b.dataset.i,10));
}));
}
/* ---------- navigation ---------- */
/* menu items only: the sidebar footer also holds a regular external link */
const navLinks = document.querySelectorAll('nav.side a[href^="#"]');
function show(id){
if(!document.getElementById(id)) return;
document.querySelectorAll("main section").forEach(s => s.classList.toggle("on", s.id === id));
navLinks.forEach(a => a.classList.toggle("active", a.getAttribute("href") === "#" + id));
el("main").scrollTop = 0;
}
navLinks.forEach(a => a.addEventListener("click", e => {
e.preventDefault();
const id = a.getAttribute("href").slice(1);
show(id);
history.replaceState(null, "", "#" + id);
}));
if(location.hash && document.getElementById(location.hash.slice(1))) show(location.hash.slice(1));
/* ---------- dashboard ---------- */
(function(){
const rows = FMT.map(f => {
const g150 = geomean(WALL[f]);
const g600 = geomean(PROJ[600][f].map(x => x[0]));
return [NAME[f], "geometric mean of Q1-Q7", g150, fmtTime(g150) + ' <span class="muted">(' + fmtTime(g600) + ' @600 GB)</span>', COLOR[f]];
});
const max = Math.max(...rows.map(r => r[2]));
el("dash-geo").innerHTML = rows.map(r => barRow(r[0], r[1], r[2]/max, r[4], r[3])).join("");
const breaks = {
csv:["Direct path reads; canonical archive; best compression","Full scans pass 1 h just above 600 GB"],
json:["Self-describing exchange files; hybrid variant is 1.9 MB","Every scan reads the whole corpus: hours at 600 GB"],
sqlite:["Fastest at 150 MB; one file; smallest footprint","Single-threaded scans pass 1 h at ~1.2 TB"],
pg:["Parallel partitioned scans + matviews; flat 0.2-0.4 s today","None found up to 6 TB (Q5 = 37 min)"],
rdf:["Native semantics, SPARQL","11.7x disk, 20% hot set in RAM, scans measured in days"]
};
el("dash-breaks").innerHTML = FMT.map(f => {
const fl = PROJ[600][f];
const nI = fl.filter(x => x[1]==="I").length, nF = fl.filter(x => x[1]==="F").length;
let flags = (nI+nF)===0 ? '<span class="pill ok">all 7 OK</span>' :
(nI ? '<span class="pill imp">' + nI + ' x &gt;1 h</span> ' : '') + (nF ? '<span class="pill fail">' + nF + ' x FAIL</span>' : '');
return "<tr><td><b>" + NAME[f] + "</b></td><td>" + breaks[f][0] + "</td><td>" + breaks[f][1] + "</td><td>" + flags + "</td></tr>";
}).join("");
})();
/* ---------- methodology ---------- */
el("method-queries").innerHTML = QN.map(q =>
'<tr><td class="mono">Q' + q + '</td><td>' + QDESC[q][0] + '</td><td class="num">' + QDESC[q][1].toLocaleString("en-US") + '</td></tr>').join("");
el("method-env").innerHTML = ENVIRONMENT;
/* ---------- storage ---------- */
(function(){
const workMax = WORK.rdf;
el("st-work").innerHTML = FMT.map(f =>
barRow(NAME[f], COEFF[f].toFixed(2) + "x vs raw CSV", WORK[f]/workMax, COLOR[f], fmtSize(WORK[f]))).join("");
const archMax = ARCH.rdf[0];
el("st-arch").innerHTML = FMT.map(f =>
barRow(NAME[f], ARCH[f][1], ARCH[f][0]/archMax, COLOR[f], fmtSize(ARCH[f][0]))).join("");
el("st-proj").innerHTML = FMT.map(f =>
"<tr><td><b>" + NAME[f] + "</b></td>" + [600,1200,6000].map(s =>
'<td class="num">' + fmtSize(COEFF[f]*s*1e9) + "</td>").join("") + "</tr>").join("");
})();
/* ---------- measured results ---------- */
(function(){
const METRICS = [
["Wall time", WALL, fmtTime],
["Client peak RSS", RSS, v => v.toFixed(0) + " MB"],
["Read bytes", READ, v => v >= 1000 ? (v/1000).toFixed(2) + " GB" : v.toFixed(1) + " MB"]
];
let q = 0, m = 0;
function render(){
const [mname, data, fmt] = METRICS[m];
el("m-title").textContent = "Q" + QN[q] + " - " + QDESC[QN[q]][0] + " - " + mname.toLowerCase() + " (median of 3 runs)";
const max = Math.max(...FMT.map(f => data[f][q])) || 1;
el("m-bars").innerHTML = FMT.map(f =>
barRow(NAME[f], "", (data[f][q])/max, COLOR[f], fmt(data[f][q]))).join("");
}
tabs("mq-tabs", QN.map(n => "Q" + n), i => { q = i; render(); });
tabs("mm-tabs", METRICS.map(x => x[0]), i => { m = i; render(); });
render();
el("m-mat-h").innerHTML = "<tr><th>Format</th>" + QN.map(n => '<th class="num">Q' + n + "</th>").join("") + "</tr>";
el("m-mat").innerHTML = FMT.map(f =>
"<tr><td><b>" + NAME[f] + "</b></td>" + WALL[f].map(v => '<td class="num">' + fmtTime(v) + "</td>").join("") + "</tr>").join("");
el("m-pgsrv").innerHTML = QN.map(n =>
'<tr><td class="mono">Q' + n + "</td><td class=\"num\">" + PGSRV[n][0].toFixed(2) + " ms</td><td class=\"num\">" +
PGSRV[n][1].toLocaleString("en-US") + '</td><td class="num">0</td></tr>').join("");
})();
/* ---------- projections ---------- */
(function(){
const SCALES = [600, 1200, 6000];
const SLBL = {600:"600 GB", 1200:"1.2 TB", 6000:"6 TB"};
function render(scale){
el("p-title").textContent = "Projected wall times at " + SLBL[scale] + " of raw data";
el("p-mat-h").innerHTML = "<tr><th>Format</th>" + QN.map(n => '<th class="num">Q' + n + "</th>").join("") + "</tr>";
el("p-mat").innerHTML = FMT.map(f =>
"<tr><td><b>" + NAME[f] + "</b></td>" + PROJ[scale][f].map(x =>
'<td class="num">' + fmtTime(x[0]) + flagPill(x[1]) + "</td>").join("") + "</tr>").join("");
const total = FMT.reduce((acc,f) => {
PROJ[scale][f].forEach(x => { if(x[1]==="I") acc.i++; if(x[1]==="F") acc.f++; });
return acc;
}, {i:0, f:0});
el("p-note").textContent = "Flags at this scale: " + total.i + " x IMPRACTICAL (over 1 hour), " + total.f +
" x FAIL (over 24 hours), " + (35 - total.i - total.f) + " x OK. Scaling ratio vs the measured corpus: x" +
Math.round(scale*1e9/WORK.csv).toLocaleString("en-US") + ".";
}
tabs("p-tabs", SCALES.map(s => SLBL[s]), i => render(SCALES[i]));
render(600);
/* degradation mini-table: full scan class geometric mean */
const cols = [["measured 150 MB", f => geomean([WALL[f][4], WALL[f][6]])]]
.concat(SCALES.map(s => ["projected " + SLBL[s], f => geomean([PROJ[s][f][4][0], PROJ[s][f][6][0]])]));
let html = '<table class="data"><thead><tr><th>Format</th>' +
cols.map(c => '<th class="num">' + c[0] + "</th>").join("") + "</tr></thead><tbody>";
html += FMT.map(f => "<tr><td><b>" + NAME[f] + "</b></td>" +
cols.map(c => '<td class="num">' + fmtTime(c[1](f)) + "</td>").join("") + "</tr>").join("");
html += "</tbody></table>";
el("p-degr").innerHTML = html;
})();
/* ---------- hardware ---------- */
(function(){
const SCALES = [600, 1200, 6000];
const SLBL = {600:"600 GB", 1200:"1.2 TB", 6000:"6 TB"};
function config(h){
return h.n + " x 1U Supermicro AS-1015CS-TNR (" + h.c + " cores) &middot; " +
h.m + " x 64 GB DDR5 ECC RDIMM &middot; " + h.d + " x 7.68 TB NVMe U.2";
}
function render(scale){
el("h-title").textContent = "Exact configuration and cost breakdown at " + SLBL[scale] + " (projected sizing)";
el("h-bom").innerHTML = FMT.map(f => {
const h = HW[scale][f];
return "<tr><td><b>" + NAME[f] + "</b></td><td>" + config(h) + "</td>" +
[h.base, h.ram, h.disk, h.extra, h.total].map((v,i) =>
'<td class="num"' + (i===4 ? ' style="font-weight:600"' : '') + ">" + fmtMoney(v) + "</td>").join("") + "</tr>";
}).join("");
const max = Math.max(...FMT.map(f => HW[scale][f].total));
el("h-bars").innerHTML = FMT.map(f => {
const h = HW[scale][f];
return barRow(NAME[f], h.n > 1 ? h.n + " nodes" : "single node", h.total/max, COLOR[f], fmtMoney(h.total));
}).join("");
scatter(scale);
}
function scatter(scale){
const W = 860, H = 340, L = 70, R = 30, T = 20, B = 46;
const pts = FMT.map(f => ({f, cost: HW[scale][f].total, geo: geomean(PROJ[scale][f].map(x => x[0]))}));
const maxC = Math.max(...pts.map(p => p.cost)) * 1.15;
const maxG = Math.max(...pts.map(p => p.geo)) * 1.15;
const X = c => L + (W-L-R) * c / maxC;
const Y = g => T + (H-T-B) * (1 - g / maxG);
let s = '<svg viewBox="0 0 ' + W + ' ' + H + '" style="width:100%;max-width:900px">';
/* axes + linear ticks in real units */
for(let i = 1; i <= 4; i++){
const c = maxC * i / 4, x = X(c);
s += '<line x1="' + x + '" y1="' + T + '" x2="' + x + '" y2="' + (H-B) + '" stroke="#eef0f2"/>';
s += '<text x="' + x + '" y="' + (H-B+18) + '" font-size="11" fill="#838a95" text-anchor="middle">$' + Math.round(c/1000) + 'k</text>';
const g = maxG * i / 4, y = Y(g);
s += '<line x1="' + L + '" y1="' + y + '" x2="' + (W-R) + '" y2="' + y + '" stroke="#eef0f2"/>';
s += '<text x="' + (L-8) + '" y="' + (y+4) + '" font-size="11" fill="#838a95" text-anchor="end">' + fmtTime(g) + '</text>';
}
s += '<line x1="' + L + '" y1="' + (H-B) + '" x2="' + (W-R) + '" y2="' + (H-B) + '" stroke="#cfd4da"/>';
s += '<line x1="' + L + '" y1="' + T + '" x2="' + L + '" y2="' + (H-B) + '" stroke="#cfd4da"/>';
s += '<text x="' + ((L+W-R)/2) + '" y="' + (H-6) + '" font-size="11.5" fill="#5b6472" text-anchor="middle">estimated hardware cost (bill of materials, linear)</text>';
s += '<text x="14" y="' + ((T+H-B)/2) + '" font-size="11.5" fill="#5b6472" text-anchor="middle" transform="rotate(-90 14 ' + ((T+H-B)/2) + ')">typical query time, projected (linear)</text>';
/* better = lower-left cue */
s += '<text x="' + (L+16) + '" y="' + (H-B-10) + '" font-size="11.5" fill="#0f7a45" font-weight="600">better: lower-left</text>';
const OFF = {csv:[12,-8], json:[12,4], sqlite:[12,14], pg:[12,-6], rdf:[-8,-12]};
const ANCH = {csv:"start", json:"start", sqlite:"start", pg:"start", rdf:"end"};
pts.forEach(p => {
s += '<circle cx="' + X(p.cost) + '" cy="' + Y(p.geo) + '" r="7" fill="' + COLOR[p.f] + '"/>';
s += '<text x="' + (X(p.cost)+OFF[p.f][0]) + '" y="' + (Y(p.geo)+OFF[p.f][1]) + '" font-size="11.5" fill="#3d424b" text-anchor="' + ANCH[p.f] + '">' +
NAME[p.f] + ': ' + fmtMoney(p.cost) + ', ' + fmtTime(p.geo) + '</text>';
});
s += "</svg>";
el("h-scatter").innerHTML = s;
}
tabs("h-tabs", SCALES.map(s => SLBL[s]), i => render(SCALES[i]));
render(600);
})();
/* ---------- scoring ---------- */
function renderScores(containerId, board){
const order = FMT.slice().sort((a,b) => board[b][3] - board[a][3]);
el(containerId).innerHTML = order.map(f => {
const [se, ra, di, total] = board[f];
return '<div class="stack"><div class="lab">' + NAME[f] + '</div><div class="trk">' +
'<div class="seg" style="width:' + (se/80*100) + '%;background:#1f5fdb"></div>' +
'<div class="seg" style="width:' + (ra/80*100) + '%;background:#e0892e"></div>' +
'<div class="seg" style="width:' + (di/80*100) + '%;background:#0f9d8a"></div>' +
'</div><div class="val">' + total.toFixed(1) + '</div></div>';
}).join("");
}
renderScores("sc-measured", SCORES.measured);
renderScores("sc-proj", SCORES.proj);
</script>
</body>
</html>

View File

@@ -82,6 +82,15 @@ invalidates everything downstream.
pass is therefore SKIPPED and marked as such in results and report - pass is therefore SKIPPED and marked as such in results and report -
never approximated by a fake (e.g. file re-copy tricks) without a never approximated by a fake (e.g. file re-copy tricks) without a
rules update describing the method. rules update describing the method.
- **PostgreSQL is remote on this setup** (dedicated DB host on the LAN),
so spec 09's backend-PID sampling via `pg_stat_activity` + psutil is
impossible - psutil cannot attribute a remote process. Instead the
runner records (a) client-side child metrics exactly like every other
format (wall times therefore include LAN round-trips - stated in the
report), (b) best-effort Prometheus / node_exporter host CPU/RAM
windows per PG cell and for the whole session (`TRIBO_PROM_URL`), and
(c) best-effort `pg_stat_statements` per-query session deltas
(server-side execution time and buffer traffic).
- Quiesce the machine during measured runs (no parallel pipeline tasks, - Quiesce the machine during measured runs (no parallel pipeline tasks,
no background loads); the runner runs cells sequentially by design. no background loads); the runner runs cells sequentially by design.

View File

@@ -31,10 +31,14 @@ out/
- Tasks communicate exclusively through these artifacts - never through - Tasks communicate exclusively through these artifacts - never through
shared in-process state shared in-process state
(see [code-python-style.md](code-python-style.md)). (see [code-python-style.md](code-python-style.md)).
- **`./out/` is git-ignored** (except nothing - the whole tree). Sources - **The bulk of `./out/` is git-ignored.** Sources of truth are the
of truth are the specs, the rules, and the code; artifacts are specs, the rules, and the code; artifacts are reproducible from them
reproducible from them ([data-determinism.md](data-determinism.md)). ([data-determinism.md](data-determinism.md)). A curated results subset
Never commit generated artifacts "for convenience". IS committed (owner decision, 2026-07-11): `out/report/` (REPORT.md,
charts, tables, diagrams), `out/config/` (lab_config.yaml,
hw_prices.yaml), and the `out/.done/` markers - exactly the exceptions
listed in `.gitignore`. Never commit the bulk artifacts (corpus,
databases, stores, archives, raw benchmark data) "for convenience".
## 2. Completion markers ## 2. Completion markers
@@ -92,7 +96,9 @@ out/
- **Writing an artifact outside `./out/`** or a task writing into another - **Writing an artifact outside `./out/`** or a task writing into another
task's output directory. task's output directory.
- **Committing `./out/` content to git.** - **Committing bulk `./out/` content to git** (corpus, databases, stores,
archives, raw benchmark data) - only the curated subset in `.gitignore`
(`report/`, `config/`, `.done/`) is committed.
- **A marker written on partial success** - see - **A marker written on partial success** - see
[code-error-handling.md](code-error-handling.md). [code-error-handling.md](code-error-handling.md).
- **Running benchmarks while converters are still running.** - **Running benchmarks while converters are still running.**

View File

@@ -78,8 +78,6 @@ an en-dash.
content, not typographic substitutions. UTF-8 throughout is the right content, not typographic substitutions. UTF-8 throughout is the right
answer. (Repository artifacts are English; this matters only for quoted answer. (Repository artifacts are English; this matters only for quoted
material.) material.)
- **Imported reference materials** in `docs/examples/` are kept verbatim as
received; they are not binding documents.
## 4. Detection / enforcement ## 4. Detection / enforcement
@@ -89,8 +87,7 @@ future detector is documented here so it can be built consistently:
- **What it flags.** Any occurrence of the Unicode codepoints in the section 1 - **What it flags.** Any occurrence of the Unicode codepoints in the section 1
substitution table inside in-scope files. The section 3 exemptions substitution table inside in-scope files. The section 3 exemptions
(scientific notation, international content, `docs/examples/`) are the only (scientific notation, international content) are the only carve-outs.
carve-outs.
- **Exit behavior.** Non-zero exit when any flagged codepoint is found under - **Exit behavior.** Non-zero exit when any flagged codepoint is found under
the scan roots, so it can later gate a commit hook or a build step. the scan roots, so it can later gate a commit hook or a build step.
- **Tooling.** When this detector is created it MUST be a PowerShell script - **Tooling.** When this detector is created it MUST be a PowerShell script

View File

@@ -30,12 +30,12 @@ The stdlib is the default. The ONLY permitted third-party packages:
| Package | Used by | For | | Package | Used by | For |
|---|---|---| |---|---|---|
| `numpy` | 03 | vectorized generation, seeded RNG | | `numpy` | 03 | vectorized generation, seeded RNG |
| `PyYAML` | 01, 03, 10 | reading/writing the YAML configs | | `PyYAML` | 01, 03, 10, 11 | reading/writing the YAML configs |
| `rdflib` | 04, 07 | JSON-LD validation, RDF serialization | | `rdflib` | 04, 07 | JSON-LD validation, RDF serialization |
| `oxigraph` | 07, 09 | embedded queryable triplestore | | `oxigraph` | 07, 09 | embedded queryable triplestore |
| `ijson` | 08, 09 | streaming JSON parsing | | `ijson` | 08, 09 | streaming JSON parsing |
| `psycopg` (v3) | 06, 08, 09 | PostgreSQL COPY and queries | | `psycopg` (v3) | 06, 08, 09 | PostgreSQL COPY and queries |
| `psutil` | 09 | RSS / CPU / IO measurement | | `psutil` | 04-07, 09 | RSS / CPU / IO measurement; converters record their own peak RSS / CPU time into their `.done` markers for the task 10 sizing model |
| `matplotlib` | 11 | report charts | | `matplotlib` | 11 | report charts |
| `pandas` | 08 (optional) | ONLY as the separately-measured second CSV query variant; never in the pipeline itself | | `pandas` | 08 (optional) | ONLY as the separately-measured second CSV query variant; never in the pipeline itself |

View File

@@ -74,11 +74,14 @@ and it wins where a spec detail is ambiguous.
- Loads are all-or-nothing per table and validated against `MANIFEST.csv` - Loads are all-or-nothing per table and validated against `MANIFEST.csv`
row counts - see [code-error-handling.md](code-error-handling.md) row counts - see [code-error-handling.md](code-error-handling.md)
section 3 and [test-pipeline-validation.md](test-pipeline-validation.md). section 3 and [test-pipeline-validation.md](test-pipeline-validation.md).
- On every SQLite connection that relies on FK behaviour: - **Foreign keys are ENFORCED in both engines on every table, including
`PRAGMA foreign_keys = ON` (SQLite silently ignores FK clauses without bulk tables** (owner decision 2026-07-11, after measurement on the real
it). Bulk tables MAY declare FK columns without enforced constraints for corpus: +0.5 s load on 4.3M rows in SQLite, zero disk and read cost;
load speed - row-count validation is the integrity gate; whichever choice full study with ER diagrams and measurements:
is made must be the same in both engines. [docs/research/Tribology_FK_Architecture.html](../research/Tribology_FK_Architecture.html)).
Every SQLite connection sets `PRAGMA foreign_keys = ON` (SQLite silently
ignores FK clauses without it). Row-count validation against
`MANIFEST.csv` remains a second, independent integrity gate.
## 5. `track_summary` - precomputed, and off-limits to Q5 ## 5. `track_summary` - precomputed, and off-limits to Q5
@@ -89,6 +92,20 @@ never read it** - Q5 is the forced full-scan benchmark over raw
`friction_cycles` in every format (spec `friction_cycles` in every format (spec
[08](../specs/08_benchmark_queries.md)). [08](../specs/08_benchmark_queries.md)).
Binding computation (identical semantics in both engines; float results
agree within the 1e-9 checksum tolerance):
1. Steady-state proxy window = cycles 501-1000 (the second half);
`m` = avg(cof), `s` = stddev_pop(cof) over that window.
2. `run_in_cycles` = the first cycle `c` with `|cof(c) - m| < 2*s`;
fallback 500 if no cycle qualifies.
3. `cof_ss_mean` / `cof_ss_std` = avg / stddev_pop of cof over cycles
`> run_in_cycles`.
SQLite computes this in Python during load
([common/track_summary.py](../../common/track_summary.py)); PostgreSQL as
a materialized view implementing the same three steps.
## 6. Engine parity ## 6. Engine parity
- Same logical schema, same table names, same column names, same row - Same logical schema, same table names, same column names, same row

View File

@@ -1,5 +1,10 @@
# 00_PLAN - Tribology Lab Data Storage Evaluation Pipeline # 00_PLAN - Tribology Lab Data Storage Evaluation Pipeline
> Provenance: this plan and the 11 task specifications were generated from
> [../Initial_Prompt.md](../Initial_Prompt.md). The prompt is a historical
> artifact, not a contract; the specs and [docs/rules/](../rules/) supersede
> it wherever they differ.
## Goal ## Goal
Build a complete, reproducible pipeline that: Build a complete, reproducible pipeline that:
1. Defines a simulated tribology laboratory (Sandia Pt-Au LDRD context) with real instruments and a realistic test workflow. 1. Defines a simulated tribology laboratory (Sandia Pt-Au LDRD context) with real instruments and a realistic test workflow.

View File

@@ -10,17 +10,19 @@ Build the single-file relational variant optimized for size and single-user quer
CSV tree. CSV tree.
## Processing instructions ## Processing instructions
- One database `./out/sqlite/tribo.db`. PRAGMAs during load: `journal_mode=OFF, synchronous=OFF, cache_size=-262144`; after load: `journal_mode=WAL`, `ANALYZE`, `VACUUM`. - One database `./out/sqlite/tribo.db`. PRAGMAs during load: `journal_mode=OFF, synchronous=OFF, cache_size=-262144, foreign_keys=ON`; after load: `journal_mode=WAL`, `ANALYZE`, `VACUUM`.
- Schema (16-table style, integer surrogate keys everywhere; no TEXT foreign keys in bulk tables): - Foreign keys are enforced on every table, including bulk tables (variant B of the measured key-schema study; owner decision 2026-07-11, see db-sql-schema rule). Full study with ER diagrams and measurements: [../research/Tribology_FK_Architecture.html](../research/Tribology_FK_Architecture.html).
- Schema (16-table style, integer surrogate keys everywhere; no TEXT foreign keys in bulk tables; SQL column names are the lowercase of the CSV headers per data-naming-units):
- `batches, wafers, coupons, runs, instruments, tracks` - dimension tables, TEXT natural keys kept as unique columns. - `batches, wafers, coupons, runs, instruments, tracks` - dimension tables, TEXT natural keys kept as unique columns.
- `xrf_points(coupon_id INT, gx INT, gy INT, pt_wtpct REAL, au_wtpct REAL)`. - `simtra_profiles(batch_id INT, row_no INT, angle_deg REAL, energy_ev REAL, pt_flux REAL, au_flux REAL, PRIMARY KEY(batch_id, row_no)) WITHOUT ROWID`.
- `profilometry_points(coupon_id INT, gx INT, gy INT, thickness_um REAL)`. - `xrf_points(coupon_id INT, grid_x INT, grid_y INT, pt_wtpct REAL, au_wtpct REAL, PRIMARY KEY(coupon_id, grid_x, grid_y)) WITHOUT ROWID`.
- `nanoindentation(coupon_id INT, indent_id INT, x_um REAL, y_um REAL, hardness_gpa REAL, reduced_modulus_gpa REAL)`. - `profilometry_points(coupon_id INT, grid_x INT, grid_y INT, thickness_um REAL, PRIMARY KEY(coupon_id, grid_x, grid_y)) WITHOUT ROWID`.
- `afm(coupon_id INT PRIMARY KEY, ra_nm REAL, rq_nm REAL)`. - `nanoindentation(coupon_id INT, indent_id INT, x_um REAL, y_um REAL, hardness_gpa REAL, reduced_modulus_gpa REAL, max_load_mn REAL, PRIMARY KEY(coupon_id, indent_id)) WITHOUT ROWID`.
- `afm(coupon_id INT PRIMARY KEY, ra_nm REAL, rq_nm REAL, image_file TEXT)`.
- `friction_cycles(track_id INT, cycle INT, cof REAL, PRIMARY KEY(track_id, cycle)) WITHOUT ROWID`. - `friction_cycles(track_id INT, cycle INT, cof REAL, PRIMARY KEY(track_id, cycle)) WITHOUT ROWID`.
- `friction_loop_points(track_id INT, cycle INT, pt INT, position_um REAL, force_mn REAL, PRIMARY KEY(track_id, cycle, pt)) WITHOUT ROWID`. - `friction_loop_points(track_id INT, cycle INT, pt INT, position_um REAL, friction_force_mn REAL, PRIMARY KEY(track_id, cycle, pt)) WITHOUT ROWID`.
- `wear(track_id INT PRIMARY KEY, wear_volume_um3 REAL, k_archard REAL, sliding_distance_m REAL)`. - `wear(track_id INT PRIMARY KEY, wear_volume_um3 REAL, k_archard REAL, sliding_distance_m REAL)`.
- `track_summary(track_id INT PRIMARY KEY, cof_ss_mean REAL, cof_ss_std REAL, run_in_cycles INT)` - precomputed during load. - `track_summary(track_id INT PRIMARY KEY, cof_ss_mean REAL, cof_ss_std REAL, run_in_cycles INT)` - precomputed during load (algorithm fixed in db-sql-schema rule, section 5).
- Indexes for Q1-Q7: `coupons(au_wtpct_mean)`, `tracks(run_id)`, `tracks(coupon_id)`, `runs(environment)`, composite `tracks(environment, load_mn)` (denormalize environment/load onto tracks). - Indexes for Q1-Q7: `coupons(au_wtpct_mean)`, `tracks(run_id)`, `tracks(coupon_id)`, `runs(environment)`, composite `tracks(environment, load_mn)` (denormalize environment/load onto tracks).
- Batch inserts with `executemany`, 50k rows per transaction. - Batch inserts with `executemany`, 50k rows per transaction.
- Validate row counts against MANIFEST. - Validate row counts against MANIFEST.

View File

@@ -1,27 +0,0 @@
# Data Processing Plan — Tribology Lab Data (LDRD Mechanochemical Alloys FY25-27)
Executor: Cursor AI agent with filesystem access to the raw CSV tree (read-only) and a writable `work/` directory.
Source tree: `SOURCE_ROOT` (set in `config/pipeline.json`, Task 01). All outputs under `work/`.
Rules: source is never mutated; every task is re-runnable; Python, tabs for indentation.
## Tasks
1. **01_setup_workspace.md** — create `work/` layout + pipeline config.
2. **02_fs_inventory.md** — walk source tree; produce file inventory.
3. **03_filename_tokenization.md** — parse filenames into tokens; flag unmatched.
4. **04_schema_profiling.md** — read headers + sampled rows; infer dtypes; compute schema fingerprints.
5. **05_schema_grouping.md** — cluster files by (measurement type × fingerprint); report structural inconsistencies.
6. **06_storage_map.md** — human-readable as-is / as-designed storage map.
7. **07_column_mapping.md** — per fingerprint: raw→canonical column mapping (ERD-aligned).
8. **08_ontology_schema.md** — relational DDL + data dictionary + Instruments Repository (iterates with 07/09).
9. **09_cleanse_normalize.md** — streamed cleanse → canonical dataset; quarantine anomalies with reason codes.
10. **10_anomaly_report.md** — consolidate rejects; refine rules; loop 0709 until coverage = 100%.
11. **11_export_json.md** — JSON mega-file + per-entity JSON set from canonical.
12. **12_export_sqlite.md** — SQLite database per DDL, bulk-loaded.
13. **13_export_postgresql.md** — PostgreSQL DDL + COPY load scripts.
14. **14_export_rdf.md** — RDF metadata/provenance graph (R2RML mappings; no numeric series).
15. **15_lineage_build_manifest.md** — lineage consolidation + final build manifest; end-to-end validation.
## Final Goal
From ~700 GB of raw heterogeneous CSV: a single cleansed **canonical dataset** conforming to the ontology (ERD + data dictionary + Instruments Repository), with all anomalies quarantined and explained, materialized into **JSON, SQLite, PostgreSQL, and RDF** exports, each carrying complete **source→target lineage**, reproducible from source + versioned config in one pipeline run.

View File

@@ -1,26 +0,0 @@
# Task 01 — Setup Workspace
## Goal
Create the working directory layout and versioned pipeline configuration.
## Input
- Path to the raw CSV root (read-only). Ask the user if unknown.
## Actions
1. Create:
```
work/
config/ # pipeline.json, token_patterns.json, column_mappings/, anomaly_rules.json
manifest/ # inventory, fingerprints, reports
canonical/ # cleansed partitioned dataset
rejects/ # quarantined anomalies
schema/ # DDL, data dictionary, instruments repository
exports/{json,sqlite,postgresql,rdf}/
lineage/
scripts/ # all Python scripts produced by tasks
```
2. Write `work/config/pipeline.json`: `SOURCE_ROOT`, sample sizes (head/tail/random rows), version stamp.
3. Write empty `work/config/token_patterns.json` and `work/config/anomaly_rules.json` skeletons.
## Intermediate Result (done when)
- Directory tree exists; `pipeline.json` valid and points to an existing `SOURCE_ROOT`.

View File

@@ -1,17 +0,0 @@
# Task 02 — Filesystem Inventory
## Goal
Non-destructive inventory of every file under `SOURCE_ROOT`.
## Input
- `work/config/pipeline.json` (Task 01).
## Actions
1. Write `scripts/fs_inventory.py` (streaming walk, constant memory).
2. For each file capture: relative path, path components (folder tokens), filename, extension, size bytes, mtime (ISO), sha-free quick id (path hash).
3. Emit `manifest/inventory.jsonl` (one JSON object per file) and load an index copy into `manifest/manifest.sqlite` table `inventory`.
4. Emit `manifest/inventory_summary.md`: file count, total size, count by extension, count and size by top-level folder, depth histogram.
## Intermediate Result (done when)
- `inventory.jsonl` row count == filesystem file count.
- `inventory_summary.md` produced; non-CSV extensions listed explicitly.

View File

@@ -1,17 +0,0 @@
# Task 03 — Filename & Path Tokenization
## Goal
Parse every path/filename into structured tokens; classify each file by measurement type and attachment level.
## Input
- `manifest/inventory.jsonl` (Task 02).
- Known conventions: hierarchy `data_note → batch → coupon(die) → sample`; XRF `L_<MMDDYY>_<seq>`; macro friction `data1..data3` (+ atmosphere); nanoindentation `run1..run9`; friction `test_N` / `cycleN` (ambient) or `_N` (dryNitrogen); sample tokens like `456/457`.
## Actions
1. Encode patterns as regex rules in `config/token_patterns.json` (versioned; extend as new patterns are found).
2. Write `scripts/tokenize.py`: apply rules to each inventory row; extract `{batch, coupon, die, sample, measurement_type, attachment_level(coupon|sample), atmosphere, run/test/cycle/seq, lot, date}`.
3. Emit `manifest/tokens.jsonl` and `manifest/unmatched.csv` (files matching no rule, with best-guess notes).
## Intermediate Result (done when)
- Every inventory row is in `tokens.jsonl` with either a full classification or an `unmatched` flag.
- `unmatched.csv` reviewed; genuinely new conventions added to `token_patterns.json` and the script re-run (repeat until unmatched are only true anomalies).

View File

@@ -1,18 +0,0 @@
# Task 04 — Schema Profiling & Fingerprints
## Goal
Capture the record structure of every CSV without full load; compute schema fingerprints.
## Input
- `manifest/inventory.jsonl`, `manifest/tokens.jsonl`; sample sizes from `pipeline.json`.
## Actions
1. Write `scripts/profile_schema.py` (streaming; head N + tail N + random M rows only).
2. Per file capture: delimiter, header presence, column names (raw, ordered), column count, inferred dtype per column, null-rate in sample, min/max of numeric columns in sample, total row estimate (size/avg-row-len).
3. Fingerprint = hash of normalized ordered column-name list (lowercased, trimmed). Also compute a name-insensitive positional fingerprint (dtype sequence) for headerless files.
4. Emit `manifest/schema_profiles.jsonl`; add table `schema_profiles` to `manifest.sqlite`.
5. Log unreadable/malformed files to `manifest/profile_errors.csv`.
## Intermediate Result (done when)
- 100% of CSV files have a profile row or an entry in `profile_errors.csv`.
- Every profile carries a fingerprint.

View File

@@ -1,16 +0,0 @@
# Task 05 — Schema Grouping & Inconsistency Report
## Goal
Enumerate distinct record structures per measurement type; expose structural inconsistencies.
## Input
- `manifest/tokens.jsonl`, `manifest/schema_profiles.jsonl`.
## Actions
1. Write `scripts/group_schemas.py`: group files by `measurement_type × fingerprint`.
2. Emit `manifest/fingerprint_report.md` per measurement type: list of distinct fingerprints, file count each, canonical column set each, column-name diffs between fingerprints (added/removed/renamed candidates), dtype conflicts.
3. Emit `manifest/inconsistencies.csv`: one row per detected inconsistency `{measurement_type, kind(column_set|naming|dtype|placement), fingerprints involved, example files}`.
## Intermediate Result (done when)
- The number of distinct schemas per measurement type is known and documented.
- `inconsistencies.csv` complete — this is the work queue for Task 07.

View File

@@ -1,19 +0,0 @@
# Task 06 — Storage Map (As-Is / As-Designed)
## Goal
Human-readable reconstruction of the storage logic: purpose of each folder level and file kind.
## Input
- All Task 0205 artifacts.
## Actions
1. Write `manifest/storage_map.md`:
- annotated folder-tree template (one exemplar branch per pattern, not all 700 GB);
- purpose of each nesting level mapped to entities (`data_note → batch → coupon/die → sample`);
- purpose of each file kind mapped to measurement types and attachment levels;
- deviations: folders/files violating the intended structure (from `unmatched.csv`, `inconsistencies.csv`).
2. Mark reconstructed intent vs observed fact explicitly (`AS-DESIGNED` / `AS-IS` labels).
## Intermediate Result (done when)
- `storage_map.md` explains every folder level and file kind; deviations enumerated with examples.
- Phase 1 complete. Gate: fingerprint coverage 100%, unmatched files are catalogued anomalies only.

View File

@@ -1,18 +0,0 @@
# Task 07 — Column Mapping (raw → canonical)
## Goal
For every `measurement_type × fingerprint`, a complete mapping of raw columns to canonical columns.
## Input
- `manifest/fingerprint_report.md`, `manifest/inconsistencies.csv`.
- Target canonical schemas (ERD): `xrf_spectrum_points`, `nanoindentation_points`, `friction_data_points`, `macro_friction_points`.
## Actions
1. For each fingerprint create `config/column_mappings/<type>__<fingerprint>.json`: `{raw_name|position → canonical_name, unit_in, unit_out, dtype, transform(optional)}`.
2. Resolve synonyms, reordering, unit variants; record every decision inline as `"note"`.
3. Columns with no canonical target → escalate to Task 08 (ontology change) instead of dropping silently.
4. `macro_friction_points` schema is PROVISIONAL — finalize from an actual macro-friction file profile here; propagate to Task 08.
## Intermediate Result (done when)
- Every fingerprint has a mapping file; no raw column is unmapped or undocumented.
- Escalation list for Task 08 written to `schema/ontology_escalations.md`.

View File

@@ -1,18 +0,0 @@
# Task 08 — Ontology: Relational Schema, Data Dictionary, Instruments Repository
## Goal
Formal, versioned data model driving cleansing and all exports. Iterates with Tasks 07/09/10.
## Input
- ERD (project artifact `Mary_G__LabResults__ERD.txt`), `schema/ontology_escalations.md`.
## Actions
1. Generate `schema/ddl_postgresql.sql` from the ERD: full hierarchy `ptau_data_notes → batches → coupons → samples`, four measurement branches, `atmospheres` lookup, `measurement_types` + `column_definitions`.
2. Generate `schema/data_dictionary.csv`: per measurement type — column name, unit, dtype, ordinal, description (source: the mappings of Task 07).
3. Create `schema/instruments_repository.csv`: instrument id, type, measurement type produced, measured quantities, units, valid min/max ranges, resolution. Populate incrementally from tokens + column signatures; ranges feed Task 09 anomaly rules.
4. Maintain `schema/ontology_notes.md`: every schema change with rationale (append-only).
5. Version the DDL (`-- schema_version: N` header); bump on each iteration.
## Intermediate Result (done when)
- DDL, data dictionary, and Instruments Repository are mutually consistent and cover all mapped fingerprints.
- No open items in `ontology_escalations.md`.

View File

@@ -1,18 +0,0 @@
# Task 09 — Cleanse & Normalize → Canonical Dataset
## Goal
Streamed transformation of all raw CSV into the canonical dataset; anomalies quarantined, lineage recorded.
## Input
- Mappings (Task 07), schema + instrument ranges (Task 08), `config/anomaly_rules.json`, tokens (Task 03).
## Actions
1. Write `scripts/cleanse.py` — streaming, constant memory, parallel by file, resumable (checkpoint by inventory offset).
2. Per file: apply column mapping → coerce dtypes → normalize units → standardize codes (atmosphere, instrument) → resolve hierarchy keys from tokens (natural keys: batch/coupon/sample/run/test/cycle numbers).
3. Anomaly checks per row: out-of-instrument-range, non-coercible type, missing required, duplicate key, sequence gap (per ordered series), malformed row. Reasons coded per `anomaly_rules.json`.
4. Clean rows → `canonical/<measurement_type>/…` partitioned files (JSONL or Parquet; partition friction points by test/cycle per Phase-1 volume estimates). Rejected rows → `rejects/<type>.jsonl` with `{source_file, row_offset, reason_code, raw_payload}`.
5. Lineage per output row: `{source_file_id, row_offset, mapping_version, schema_version}` → compact companion files in `lineage/`.
## Intermediate Result (done when)
- Every in-scope file processed; row counts reconcile: `source ≈ canonical + rejects` per file.
- Canonical partitions written; lineage present for every canonical row.

View File

@@ -1,17 +0,0 @@
# Task 10 — Anomaly Report & Iteration Gate
## Goal
Consolidate anomalies, refine rules/mappings/ontology, and gate the Phase 2↔3 loop.
## Input
- `rejects/*.jsonl`, `manifest/unmatched.csv`, current rule/mapping/schema versions.
## Actions
1. Write `scripts/anomaly_report.py``manifest/anomaly_report.md`: counts by reason code × measurement type, top offending files, examples per class, reject-rate per fingerprint.
2. Triage each anomaly class: (a) true data anomaly — keep quarantined, document; (b) rule defect — fix `anomaly_rules.json` / mapping / token pattern; (c) ontology gap — escalate to Task 08.
3. If (b) or (c) occurred: bump versions and re-run Tasks 07→09 for affected types.
4. Repeat until: fingerprint mapping coverage 100%, reject classes are all category (a), reject rate stable between runs.
## Intermediate Result (done when)
- `anomaly_report.md` final; every reject class documented as a true anomaly.
- Canonical dataset frozen: `canonical/SNAPSHOT.md` written with versions + row counts. Phases 2↔3 complete.

View File

@@ -1,17 +0,0 @@
# Task 11 — Export: JSON
## Goal
Two JSON products from the frozen canonical snapshot.
## Input
- `canonical/` snapshot, `schema/` (Tasks 0910).
## Actions
1. Write `scripts/export_json.py` (streaming writer; never build the whole document in memory).
2. **Mega-file** `exports/json/dataset.json`: full hierarchy `data_note → batches → coupons(die) → samples → measurements → points`, plus embedded `data_dictionary` and `instruments`.
3. **Split set** `exports/json/split/`: directory mirrors the hierarchy; one file per measurement (points inline), plus `index.json` per level.
4. Each measurement object carries its lineage block `{source_file, mapping_version, schema_version}`.
## Intermediate Result (done when)
- Both products written; spot-validation: point counts per measurement match canonical.
- `exports/json/EXPORT_MANIFEST.json` written (snapshot id, versions, counts).

View File

@@ -1,16 +0,0 @@
# Task 12 — Export: SQLite
## Goal
Portable single-file SQLite database conforming to the ontology.
## Input
- `canonical/` snapshot, `schema/ddl_postgresql.sql`.
## Actions
1. Write `scripts/export_sqlite.py`: translate DDL to SQLite dialect (bigint→INTEGER, timestamptz→TEXT ISO, keep FKs), create `exports/sqlite/labdata.sqlite`.
2. Bulk-load in dependency order (hierarchy → lookups → headers → points); `executemany` in transactions; create indexes after load; `PRAGMA journal_mode=OFF, synchronous=OFF` during load, `ANALYZE` after.
3. Load `measurement_types`, `column_definitions`, `instruments` tables; add `lineage` table (row-range level, not per-row, to bound size).
## Intermediate Result (done when)
- `labdata.sqlite` opens; per-table row counts equal canonical counts; FK integrity check passes (`PRAGMA foreign_key_check` empty).
- `exports/sqlite/EXPORT_MANIFEST.json` written.

View File

@@ -1,22 +0,0 @@
# Task 13 — Export: PostgreSQL
## Goal
System-of-record load package: DDL + COPY-based bulk load (per ADR ingestion path).
## Input
- `canonical/` snapshot, `schema/ddl_postgresql.sql`.
## Actions
1. Generate `exports/postgresql/`:
- `01_schema.sql` (DDL, versioned);
- `02_lookups.sql` (atmospheres, measurement_types, column_definitions, instruments);
- per-table CSV staging files from canonical (streamed);
- `03_copy_load.sql` (`\copy` statements in dependency order);
- `04_post_load.sql` (indexes, constraints validation, ANALYZE);
- `load.sh` / `load.ps1` orchestration.
2. Point tables loaded via COPY; headers via INSERT; partitioning DDL for `friction_data_points` per Phase-1 volume estimates.
3. Include `lineage` table load (row-range granularity).
## Intermediate Result (done when)
- Package executes cleanly on an empty database (test locally if a PG instance is available; otherwise dry-validate SQL syntax).
- Post-load counts script confirms parity with canonical. `EXPORT_MANIFEST.json` written.

View File

@@ -1,17 +0,0 @@
# Task 14 — Export: RDF (Metadata & Provenance Graph)
## Goal
RDF surface for the metadata/provenance layer only — never the numeric series (per ADR).
## Input
- `canonical/` snapshot headers + hierarchy, `schema/`, lineage.
## Actions
1. Mint stable IRIs for every data note, batch, coupon, sample, measurement (pattern `https://tribology.ufl.example/id/<entity>/<key>`; base IRI in config).
2. Write `scripts/export_rdf.py``exports/rdf/metadata.ttl` (Turtle): hierarchy, measurement headers, atmosphere, instrument, source-file provenance. Vocabularies: PROV-O (provenance), QUDT (units), Dublin Core (citation), DCAT (dataset catalog), local `lab:` terms.
3. Emit `exports/rdf/mappings.r2rml.ttl` — R2RML mappings from the PostgreSQL schema, ready for Ontop virtual deployment.
4. Numeric series: reference-only — each measurement IRI carries a `lab:dataService` pointer, no point triples.
## Intermediate Result (done when)
- `metadata.ttl` parses (validate with `rdflib`); triple counts reported per class.
- R2RML mappings cover all metadata tables. `EXPORT_MANIFEST.json` written.

View File

@@ -1,18 +0,0 @@
# Task 15 — Lineage Consolidation, Build Manifest, Final Validation
## Goal
Close the pipeline: consolidated lineage, reproducible build record, cross-format validation.
## Input
- All `exports/*`, `lineage/`, all config/schema versions.
## Actions
1. Write `scripts/finalize.py`:
- consolidate `lineage/` into `lineage/lineage.sqlite` (queryable: target row-range → source file + offsets + rule versions);
- produce `BUILD_MANIFEST.md`: snapshot id, all config/mapping/schema versions, input inventory hash, per-artifact counts, run timestamps — the processing-flow documentation.
2. Cross-format validation: row counts and key aggregates (e.g., sum of cycles per test, point counts per measurement) identical across canonical, JSON, SQLite, PostgreSQL staging.
3. Lineage spot-check: resolve 20 random target rows back to source file + row offset; verify raw values match.
## Final Result (pipeline done when)
- `BUILD_MANIFEST.md` complete; all validations pass; lineage resolvable.
- **Final goal achieved:** cleansed canonical dataset conforming to the ontology, all anomalies quarantined and explained, four export formats generated from one snapshot, full source→target lineage, entire pipeline reproducible from source + versioned config.

295
extrapolate.py Normal file
View File

@@ -0,0 +1,295 @@
"""Task 10 - Extrapolation to 600 GB / 1.2 TB / 6 TB.
Calibrates per-cell scaling laws on the measured 150 MB point (task 09
medians) and projects wall times and hardware needs per format x scale.
Model (spec 10, docs/rules/bench-methodology.md section 7):
- per-format fixed overhead = the format's cheapest measured median
(interpreter start, imports, connection); the remainder is the
data-dependent calibration constant;
- growth laws per access pattern: o1 (flat), log (index depth,
1 + ln r / ln n0 with n0 = measured friction-cycle rows), logk (result
set grows linearly with data), k_log_n (linear x index depth), scan
(linear), scan_parallel (linear x base_cores / cores_at_scale; range
partitions outnumber cores at every scale);
- RDF join penalties are already embedded in the measured constants;
- RDF q3/q5/q7 are scans because the implementation derives run-in /
steady-state from raw cycles (no summary triples exist by design).
Every number written here is a PROJECTION, never a measurement; rows are
flagged IMPRACTICAL (> 1 h) and FAIL (> 24 h).
Spec: docs/specs/10_extrapolation_model.md.
"""
from __future__ import annotations
import csv
import logging
import math
import sys
from pathlib import Path
import psycopg
import yaml
from common.pg import get_dsn
from common.pipeline import (
check_dependencies,
load_lab_config,
parse_task_args,
remove_stale_marker,
write_marker,
)
from common.relational import expected_counts
logger = logging.getLogger(__name__)
TASK_ID = "10"
DEPENDS_ON = ["09"]
SCALES_GB = (600, 1200, 6000) # raw-CSV target scales (decimal GB), spec 10
IMPRACTICAL_S = 3600.0 # projection > 1 h, spec 10
FAIL_S = 24 * 3600.0 # projection > 24 h, spec 10
DISK_HEADROOM = 1.3 # 30% free-space headroom, spec 10
CACHE_INDEX_FRACTION = 0.10 # relational hot-set: 10% of index size, spec 10
WORKING_SET_GB = 4.0 # relational session working set, spec 10 default
RDF_HOT_FRACTION = 0.20 # dictionary+index working set, 15-25% midpoint, spec 10
STREAM_RAM_GB = 16.0 # flat streaming baseline (csv/json servers)
MIN_RAM_GB = 16.0 # smallest sensible server RAM
NODE_RAM_CEILING_GB = 1024.0 # single-node RAM threshold, spec 10 RDF note
BASE_CORES = 4 # cores behind the measured PG parallel scans (task 09 host)
CORE_OPTIONS = (16, 32) # priced server tiers, hw_prices.yaml
FORMATS = ("csv", "json", "sqlite", "pg", "rdf")
QUERY_NUMBERS = tuple(range(1, 8))
# working (queryable) footprint variant per format for storage scaling
SIZE_VARIANT = {
"csv": ("csv", "tree"), "json": ("json", "full"), "sqlite": ("sqlite", "db"),
"pg": ("pg", "live"), "rdf": ("rdf", "store"),
}
# access-pattern law per (format, query) - spec 10 table, adjusted to the
# ACTUAL implementations (rdf q3/q5/q7 read raw cycles; see module docstring)
LAWS = {
("csv", 1): "o1", ("json", 1): "o1", ("sqlite", 1): "log", ("pg", 1): "log", ("rdf", 1): "log",
("csv", 2): "scan", ("json", 2): "scan", ("sqlite", 2): "logk", ("pg", 2): "logk", ("rdf", 2): "logk",
("csv", 3): "scan", ("json", 3): "scan", ("sqlite", 3): "k_log_n", ("pg", 3): "k_log_n", ("rdf", 3): "scan",
("csv", 4): "scan", ("json", 4): "scan", ("sqlite", 4): "logk", ("pg", 4): "logk", ("rdf", 4): "logk",
("csv", 5): "scan", ("json", 5): "scan", ("sqlite", 5): "scan", ("pg", 5): "scan_parallel", ("rdf", 5): "scan",
("csv", 6): "scan", ("json", 6): "scan", ("sqlite", 6): "k_log_n", ("pg", 6): "k_log_n", ("rdf", 6): "k_log_n",
("csv", 7): "scan", ("json", 7): "scan", ("sqlite", 7): "scan", ("pg", 7): "scan_parallel", ("rdf", 7): "scan",
}
HW_PRICES_DEFAULT = """# Hardware cost parameters for extrapolation (task 10).
# Street prices retrieved 2026-07-11 from the sources below (server DRAM is
# in a documented 2025-2026 price surge - see market_context). The cost
# model buys whole components: 64 GB RDIMM modules and 7.68 TB NVMe drives.
# Operator-editable; re-run extrapolate.py after changes (code-config-yaml).
as_of: "2026-07-11"
ram_module_gb: 64
ram_module_usd: 1887.0 # A-Tech 64GB (2x32GB) DDR5-5600 ECC RDIMM
nvme_drive_tb: 7.68
nvme_drive_usd: 3995.0 # Cloud Ninjas NEW 7.68TB NVMe U.2 1DWPD (Dell 14-16G)
server_16_core_usd: 4618.0 # Supermicro AS-1015CS-TNR 1U (EPYC 9004/9005), Broadberry starting config
server_32_core_usd: 6272.0 # Supermicro AS-1115CS-TNR 1U, Broadberry starting config
extra_node_usd: 4618.0 # one additional entry 1U node (AS-1015CS-TNR chassis)
sources:
ram: https://atechmemory.com/collections/ddr5-memory-ram
nvme: https://cloudninjas.com/products/new-7-68tb-nvme-u-2-1dwpd-sie-2-5-enterprise-solid-state-drive-for-14th-15th-16th-gen-dell
servers: https://www.broadberry.com/amd-epyc-9004-supermicro-servers
market_context: https://www.techpowerup.com/342331/server-dram-pricing-jumps-50-only-70-of-orders-getting-filled
"""
HW_PRICE_KEYS = (
"as_of", "ram_module_gb", "ram_module_usd", "nvme_drive_tb", "nvme_drive_usd",
"server_16_core_usd", "server_32_core_usd", "extra_node_usd", "sources",
)
def read_medians(path: Path) -> dict[tuple[str, int], float]:
"""Median walls of the 35 cells; abort unless every cell is OK + checksummed
(projections built on invalid measurements are forbidden, bench-methodology).
"""
medians: dict[tuple[str, int], float] = {}
with open(path, encoding="ascii", newline="") as fh:
for row in csv.DictReader(fh):
if row["status"] != "OK" or row["checksum_ok"] != "true":
raise ValueError(f"{path}: cell {row['format']}/{row['query']} is {row['status']}/checksum={row['checksum_ok']} - cannot calibrate on it")
medians[(row["format"], int(row["query"][1:]))] = float(row["wall_s_median"])
expected_cells = {(fmt, qn) for fmt in FORMATS for qn in QUERY_NUMBERS}
if set(medians) != expected_cells:
raise ValueError(f"{path}: {len(medians)} cells, expected the full 35-cell matrix")
return medians
def read_sizes(path: Path) -> dict[tuple[str, str], int]:
sizes: dict[tuple[str, str], int] = {}
with open(path, encoding="ascii", newline="") as fh:
reader = csv.reader(fh)
next(reader)
for fmt, variant, nbytes in reader:
sizes[(fmt, variant)] = int(nbytes)
missing = set(SIZE_VARIANT.values()) - set(sizes)
if missing:
raise ValueError(f"{path}: missing rows {sorted(missing)}")
return sizes
def pg_index_share(dsn: str) -> tuple[float, int]:
"""Index share of the live PG footprint (read-only catalog query)."""
with psycopg.connect(dsn) as conn:
index_bytes, total_bytes = conn.execute(
"SELECT COALESCE(sum(pg_indexes_size(c.oid)), 0)::bigint,"
" COALESCE(sum(pg_total_relation_size(c.oid)), 0)::bigint"
" FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace"
" WHERE n.nspname = 'public' AND c.relkind IN ('r', 'p', 'm')"
).fetchone()
if total_bytes == 0:
raise ValueError("pg catalog reports zero relation bytes - is the database loaded?")
return index_bytes / total_bytes, index_bytes
def load_hw_prices(config_dir: Path) -> dict:
path = config_dir / "hw_prices.yaml"
if not path.exists():
path.write_text(HW_PRICES_DEFAULT, encoding="ascii", newline="\n")
logger.info("seeded editable defaults: %s", path)
with open(path, encoding="ascii") as fh:
prices = yaml.safe_load(fh)
missing = [k for k in HW_PRICE_KEYS if k not in prices]
if missing:
raise ValueError(f"{path}: missing keys {missing}")
return prices
def growth(law: str, r: float, n0: int, cores: int) -> float:
depth = 1.0 + math.log(r) / math.log(n0)
if law == "o1":
return 1.0
if law == "log":
return depth
if law == "logk":
return r
if law == "k_log_n":
return r * depth
if law == "scan":
return r
if law == "scan_parallel":
return r * BASE_CORES / cores
logger.error("growth: unknown law=%r", law)
raise ValueError(f"unknown scaling law: {law!r}")
def flag_of(seconds: float) -> str:
if seconds > FAIL_S:
return "FAIL"
if seconds > IMPRACTICAL_S:
return "IMPRACTICAL"
return "OK"
def project(fmt: str, qn: int, medians: dict, overhead: dict, r: float, n0: int, cores: int) -> float:
variable = max(0.0, medians[(fmt, qn)] - overhead[fmt])
return overhead[fmt] + variable * growth(LAWS[(fmt, qn)], r, n0, cores)
def pick_pg_cores(medians: dict, overhead: dict, r: float, n0: int) -> int:
"""Smallest priced tier; escalate only when it fixes an IMPRACTICAL flag
on a parallel-scan cell (the only law that benefits from cores)."""
for cores in CORE_OPTIONS:
worst = max(project("pg", qn, medians, overhead, r, n0, cores) for qn in QUERY_NUMBERS if LAWS[("pg", qn)] == "scan_parallel")
if worst <= IMPRACTICAL_S:
return cores
return CORE_OPTIONS[-1]
def main() -> int:
args = parse_task_args("Task 10: extrapolate measurements to 600 GB / 1.2 TB / 6 TB")
out_root: Path = args.out_root
bench_dir = out_root / "bench"
try:
check_dependencies(out_root, DEPENDS_ON)
remove_stale_marker(out_root, TASK_ID)
medians = read_medians(bench_dir / "results_median.csv")
sizes = read_sizes(bench_dir / "storage_sizes.csv")
prices = load_hw_prices(out_root / "config")
n0 = expected_counts(load_lab_config(out_root))["friction_cycles"]
csv_bytes = sizes[SIZE_VARIANT["csv"]]
coeff = {fmt: sizes[SIZE_VARIANT[fmt]] / csv_bytes for fmt in FORMATS}
index_share, pg_index_bytes = pg_index_share(get_dsn())
logger.info("storage coefficients vs csv: %s; pg index share %.3f",
", ".join(f"{fmt}={coeff[fmt]:.2f}" for fmt in FORMATS), index_share)
overhead = {fmt: min(medians[(fmt, qn)] for qn in QUERY_NUMBERS) for fmt in FORMATS}
flags = {"OK": 0, "IMPRACTICAL": 0, "FAIL": 0}
with open(bench_dir / "extrapolation.csv", "w", encoding="ascii", newline="") as fh:
writer = csv.writer(fh)
writer.writerow(["format", "query", "scale_gb", "projected_wall_s", "flag"])
for scale_gb in SCALES_GB:
r = scale_gb * 1e9 / csv_bytes
cores = {fmt: CORE_OPTIONS[0] for fmt in FORMATS}
cores["pg"] = pick_pg_cores(medians, overhead, r, n0)
for fmt in FORMATS:
for qn in QUERY_NUMBERS:
seconds = project(fmt, qn, medians, overhead, r, n0, cores[fmt])
flag = flag_of(seconds)
flags[flag] += 1
writer.writerow([fmt, f"q{qn}", scale_gb, round(seconds, 3), flag])
with open(bench_dir / "hardware_sizing.csv", "w", encoding="ascii", newline="") as fh:
writer = csv.writer(fh)
writer.writerow(["format", "scale_gb", "disk_tb", "ram_gb", "cores", "nodes",
"est_cost_usd", "ram_modules", "nvme_drives", "base_usd", "ram_usd", "disk_usd", "extra_nodes_usd"])
for scale_gb in SCALES_GB:
r = scale_gb * 1e9 / csv_bytes
for fmt in FORMATS:
fmt_bytes = coeff[fmt] * scale_gb * 1e9
disk_tb = fmt_bytes * DISK_HEADROOM / 1e12
if fmt in ("csv", "json"):
ram_gb = STREAM_RAM_GB
elif fmt in ("sqlite", "pg"):
# same index share assumed for sqlite: identical rows,
# keys, and indexes as the PG schema (variant B)
ram_gb = max(MIN_RAM_GB, CACHE_INDEX_FRACTION * index_share * fmt_bytes / 1e9 + WORKING_SET_GB)
else:
ram_gb = max(MIN_RAM_GB, RDF_HOT_FRACTION * fmt_bytes / 1e9)
cores = pick_pg_cores(medians, overhead, r, n0) if fmt == "pg" else CORE_OPTIONS[0]
nodes = max(1, math.ceil(ram_gb / NODE_RAM_CEILING_GB))
# bill of materials: whole RDIMM modules and whole NVMe drives
ram_modules = math.ceil(ram_gb / prices["ram_module_gb"])
nvme_drives = math.ceil(disk_tb / prices["nvme_drive_tb"])
base_usd = prices["server_16_core_usd"] if cores == 16 else prices["server_32_core_usd"]
ram_usd = ram_modules * prices["ram_module_usd"]
disk_usd = nvme_drives * prices["nvme_drive_usd"]
extra_nodes_usd = (nodes - 1) * prices["extra_node_usd"]
cost = base_usd + ram_usd + disk_usd + extra_nodes_usd
writer.writerow([fmt, scale_gb, round(disk_tb, 3), round(ram_gb, 1), cores, nodes,
round(cost), ram_modules, nvme_drives,
round(base_usd), round(ram_usd), round(disk_usd), round(extra_nodes_usd)])
entries = {
"extrapolation_rows": len(SCALES_GB) * len(FORMATS) * len(QUERY_NUMBERS),
"sizing_rows": len(SCALES_GB) * len(FORMATS),
"flags_ok": flags["OK"],
"flags_impractical": flags["IMPRACTICAL"],
"flags_fail": flags["FAIL"],
"pg_index_share": round(index_share, 3),
"pg_index_mb": round(pg_index_bytes / 1048576, 1),
"prices_as_of": prices["as_of"],
"note_pandas": "pandas CSV variant not measured; its RAM is O(n) - infeasible above ~10 GB (spec 10)",
}
for fmt in FORMATS:
entries[f"coeff_{fmt}"] = round(coeff[fmt], 3)
marker_path = write_marker(out_root, TASK_ID, entries)
except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1
print(f"task 10 ok: {entries['extrapolation_rows']} projections "
f"(ok={flags['OK']}, impractical={flags['IMPRACTICAL']}, fail={flags['FAIL']}), "
f"{entries['sizing_rows']} sizing rows")
print(f"outputs: {bench_dir / 'extrapolation.csv'}, {bench_dir / 'hardware_sizing.csv'}")
print(f"marker: {marker_path}")
return 0
if __name__ == "__main__":
sys.exit(main())

View File

@@ -9,21 +9,26 @@ instruments would write it, driven entirely by out/config/lab_config.yaml
from __future__ import annotations from __future__ import annotations
import argparse
import hashlib import hashlib
import logging import logging
import math import math
import shutil import shutil
import sys import sys
from datetime import datetime, timedelta, timezone from datetime import datetime, timedelta
from pathlib import Path from pathlib import Path
import numpy as np import numpy as np
import yaml
from common.pipeline import (
check_dependencies,
load_lab_config,
parse_task_args,
remove_stale_marker,
write_marker,
)
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
REPO_ROOT = Path(__file__).resolve().parent
TASK_ID = "03" TASK_ID = "03"
DEPENDS_ON = ["01"] DEPENDS_ON = ["01"]
@@ -361,56 +366,17 @@ class CorpusGenerator:
return len(data) return len(data)
def check_dependencies(out_root: Path) -> None:
for dep in DEPENDS_ON:
marker = out_root / ".done" / f"{dep}.ok"
if not marker.exists():
raise FileNotFoundError(f"dependency marker missing: {marker} - run task {dep} first")
def write_marker(marker_path: Path, gen: CorpusGenerator, tree_bytes: int) -> None:
lines = [
f"task: '{TASK_ID}'",
"status: ok",
f"csv_root: {gen.csv_root.as_posix()}",
f"files: {len(gen.manifest)}",
f"total_rows: {gen.total_rows}",
f"total_bytes: {tree_bytes}",
f"coupons: {gen.counts['coupons']}",
f"tracks: {gen.counts['tracks']}",
f"cycle_rows: {gen.counts['cycle_rows']}",
f"loop_points: {gen.counts['loop_rows']}",
f"xrf_points: {gen.counts['xrf_rows']}",
]
marker_path.parent.mkdir(parents=True, exist_ok=True)
with open(marker_path, "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(lines) + "\n")
def main() -> int: def main() -> int:
parser = argparse.ArgumentParser(description="Task 03: generate the CSV corpus") args = parse_task_args("Task 03: generate the CSV corpus")
parser.add_argument("--log-level", default="INFO", choices=["DEBUG", "INFO", "WARNING", "ERROR"])
parser.add_argument("--out-root", type=Path, default=REPO_ROOT / "out", help="artifact tree root (default: ./out)")
args = parser.parse_args()
logging.basicConfig(
level=args.log_level,
stream=sys.stderr,
format="%(asctime)s %(levelname)-7s %(name)s: %(message)s",
)
out_root: Path = args.out_root out_root: Path = args.out_root
csv_root = out_root / "csv" csv_root = out_root / "csv"
marker_path = out_root / ".done" / f"{TASK_ID}.ok"
try: try:
check_dependencies(out_root) check_dependencies(out_root, DEPENDS_ON)
if marker_path.exists(): remove_stale_marker(out_root, TASK_ID)
logger.info("re-run: removing stale marker %s", marker_path)
marker_path.unlink()
if csv_root.exists(): if csv_root.exists():
logger.info("re-run: removing previous corpus %s", csv_root) logger.info("re-run: removing previous corpus %s", csv_root)
shutil.rmtree(csv_root) shutil.rmtree(csv_root)
with open(out_root / "config" / "lab_config.yaml", encoding="utf-8") as fh: cfg = load_lab_config(out_root)
cfg = yaml.safe_load(fh)
gen = CorpusGenerator(cfg, csv_root) gen = CorpusGenerator(cfg, csv_root)
gen.generate() gen.generate()
@@ -424,7 +390,17 @@ def main() -> int:
raise ValidationError( raise ValidationError(
f"tree size {tree_mib:.1f} MiB outside gate {SIZE_GATE_MIB[0]}-{SIZE_GATE_MIB[1]} MiB" f"tree size {tree_mib:.1f} MiB outside gate {SIZE_GATE_MIB[0]}-{SIZE_GATE_MIB[1]} MiB"
) )
write_marker(marker_path, gen, tree_bytes) marker_path = write_marker(out_root, TASK_ID, {
"csv_root": gen.csv_root.as_posix(),
"files": len(gen.manifest),
"total_rows": gen.total_rows,
"total_bytes": tree_bytes,
"coupons": gen.counts["coupons"],
"tracks": gen.counts["tracks"],
"cycle_rows": gen.counts["cycle_rows"],
"loop_points": gen.counts["loop_rows"],
"xrf_points": gen.counts["xrf_rows"],
})
except Exception: except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True) logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1 return 1

View File

@@ -8,16 +8,16 @@ completion marker. Spec: docs/specs/01_lab_configuration.md.
from __future__ import annotations from __future__ import annotations
import argparse
import logging import logging
import sys import sys
from pathlib import Path from pathlib import Path
import yaml import yaml
from common.pipeline import parse_task_args, remove_stale_marker, write_marker
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
REPO_ROOT = Path(__file__).resolve().parent
TASK_ID = "01" TASK_ID = "01"
# Transcription of docs/specs/01_lab_configuration.md. This literal is the # Transcription of docs/specs/01_lab_configuration.md. This literal is the
@@ -301,51 +301,30 @@ def write_yaml(cfg: dict, config_path: Path) -> int:
return len(text.encode("utf-8")) return len(text.encode("utf-8"))
def write_marker(marker_path: Path, config_path: Path, config_bytes: int, cfg: dict) -> None:
v = cfg["volumes"]
f = cfg["friction_assignment"]
h = cfg["hierarchy"]
lines = [
f"task: '{TASK_ID}'",
"status: ok",
f"config_file: {config_path.as_posix()}",
f"config_bytes: {config_bytes}",
f"seed: {cfg['seed']}",
f"batches: {h['batches']}",
f"wafers: {h['batches'] * h['wafers_per_batch']}",
f"coupons: {v['coupons_total']}",
f"friction_coupons: {f['friction_coupons_total']}",
f"tracks: {v['tracks_total']}",
f"cycle_rows: {v['cycle_rows_total']}",
f"loop_points: {v['loop_points_total']}",
f"xrf_points: {v['xrf_points_total']}",
]
marker_path.parent.mkdir(parents=True, exist_ok=True)
with open(marker_path, "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(lines) + "\n")
def main() -> int: def main() -> int:
parser = argparse.ArgumentParser(description="Task 01: write out/config/lab_config.yaml") args = parse_task_args("Task 01: write out/config/lab_config.yaml")
parser.add_argument("--log-level", default="INFO", choices=["DEBUG", "INFO", "WARNING", "ERROR"])
parser.add_argument("--out-root", type=Path, default=REPO_ROOT / "out", help="artifact tree root (default: ./out)")
args = parser.parse_args()
logging.basicConfig(
level=args.log_level,
stream=sys.stderr,
format="%(asctime)s %(levelname)-7s %(name)s: %(message)s",
)
out_root: Path = args.out_root out_root: Path = args.out_root
config_path = out_root / "config" / "lab_config.yaml" config_path = out_root / "config" / "lab_config.yaml"
marker_path = out_root / ".done" / f"{TASK_ID}.ok"
try: try:
if marker_path.exists(): remove_stale_marker(out_root, TASK_ID)
logger.info("re-run: removing stale marker %s", marker_path)
marker_path.unlink()
validate_config(LAB_CONFIG) validate_config(LAB_CONFIG)
config_bytes = write_yaml(LAB_CONFIG, config_path) config_bytes = write_yaml(LAB_CONFIG, config_path)
write_marker(marker_path, config_path, config_bytes, LAB_CONFIG) v = LAB_CONFIG["volumes"]
f = LAB_CONFIG["friction_assignment"]
h = LAB_CONFIG["hierarchy"]
marker_path = write_marker(out_root, TASK_ID, {
"config_file": config_path.as_posix(),
"config_bytes": config_bytes,
"seed": LAB_CONFIG["seed"],
"batches": h["batches"],
"wafers": h["batches"] * h["wafers_per_batch"],
"coupons": v["coupons_total"],
"friction_coupons": f["friction_coupons_total"],
"tracks": v["tracks_total"],
"cycle_rows": v["cycle_rows_total"],
"loop_points": v["loop_points_total"],
"xrf_points": v["xrf_points_total"],
})
except Exception: except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True) logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1 return 1

View File

@@ -8,16 +8,20 @@ the laboratory workflow: ./out/report/process_flow.md plus one standalone
from __future__ import annotations from __future__ import annotations
import argparse
import logging import logging
import sys import sys
from pathlib import Path from pathlib import Path
import yaml from common.pipeline import (
check_dependencies,
load_lab_config,
parse_task_args,
remove_stale_marker,
write_marker,
)
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
REPO_ROOT = Path(__file__).resolve().parent
TASK_ID = "02" TASK_ID = "02"
DEPENDS_ON = ["01"] DEPENDS_ON = ["01"]
@@ -190,13 +194,6 @@ def build_diagrams(cfg: dict) -> list[tuple[str, str, str, str]]:
] ]
def check_dependencies(out_root: Path) -> None:
for dep in DEPENDS_ON:
marker = out_root / ".done" / f"{dep}.ok"
if not marker.exists():
raise FileNotFoundError(f"dependency marker missing: {marker} - run task {dep} first")
def validate_diagram(stem: str, text: str) -> None: def validate_diagram(stem: str, text: str) -> None:
first = text.splitlines()[0].strip() if text.strip() else "" first = text.splitlines()[0].strip() if text.strip() else ""
if not first.startswith(DIAGRAM_TYPES): if not first.startswith(DIAGRAM_TYPES):
@@ -235,42 +232,20 @@ def write_outputs(cfg: dict, out_root: Path) -> tuple[Path, int, int]:
return md_path, len(md_text.encode("utf-8")), total_diagram_bytes return md_path, len(md_text.encode("utf-8")), total_diagram_bytes
def write_marker(marker_path: Path, md_path: Path, md_bytes: int, diagram_bytes: int) -> None:
lines = [
f"task: '{TASK_ID}'",
"status: ok",
f"process_flow_file: {md_path.as_posix()}",
f"process_flow_bytes: {md_bytes}",
"diagrams: 6",
f"diagram_bytes: {diagram_bytes}",
]
marker_path.parent.mkdir(parents=True, exist_ok=True)
with open(marker_path, "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(lines) + "\n")
def main() -> int: def main() -> int:
parser = argparse.ArgumentParser(description="Task 02: write process flow diagrams") args = parse_task_args("Task 02: write process flow diagrams")
parser.add_argument("--log-level", default="INFO", choices=["DEBUG", "INFO", "WARNING", "ERROR"])
parser.add_argument("--out-root", type=Path, default=REPO_ROOT / "out", help="artifact tree root (default: ./out)")
args = parser.parse_args()
logging.basicConfig(
level=args.log_level,
stream=sys.stderr,
format="%(asctime)s %(levelname)-7s %(name)s: %(message)s",
)
out_root: Path = args.out_root out_root: Path = args.out_root
marker_path = out_root / ".done" / f"{TASK_ID}.ok"
try: try:
check_dependencies(out_root) check_dependencies(out_root, DEPENDS_ON)
if marker_path.exists(): remove_stale_marker(out_root, TASK_ID)
logger.info("re-run: removing stale marker %s", marker_path) cfg = load_lab_config(out_root)
marker_path.unlink()
with open(out_root / "config" / "lab_config.yaml", encoding="utf-8") as fh:
cfg = yaml.safe_load(fh)
md_path, md_bytes, diagram_bytes = write_outputs(cfg, out_root) md_path, md_bytes, diagram_bytes = write_outputs(cfg, out_root)
write_marker(marker_path, md_path, md_bytes, diagram_bytes) marker_path = write_marker(out_root, TASK_ID, {
"process_flow_file": md_path.as_posix(),
"process_flow_bytes": md_bytes,
"diagrams": 6,
"diagram_bytes": diagram_bytes,
})
except Exception: except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True) logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1 return 1

339
make_queries.py Normal file
View File

@@ -0,0 +1,339 @@
"""Task 08 - Benchmark query definitions (Q1-Q7 x 5 formats).
Writes the runnable query files under ./out/bench/queries/ (SQL for
SQLite/PostgreSQL, SPARQL text + python drivers for RDF, python drivers
for CSV/JSON-LD), then produces the canonical expected results: executed
on PostgreSQL, cross-validated against SQLite row-by-row within 1e-9, and
stored as ./out/bench/expected/q<N>.rows.csv + q<N>.sha256. CSV, JSON and
RDF implementations are spot-checked on Q1. Writes ./out/.done/08.ok.
Spec: docs/specs/08_benchmark_queries.md.
"""
from __future__ import annotations
import logging
import sqlite3
import sys
from pathlib import Path
import psycopg
from common.pg import get_dsn
from common.pipeline import (
check_dependencies,
parse_task_args,
process_metrics,
remove_stale_marker,
write_marker,
)
from common.queries import Q1_TRACK_CODE, csv_queries, json_queries, rdf_queries
from common.results import compare, serialize, sha256_of, sort_rows
logger = logging.getLogger(__name__)
TASK_ID = "08"
DEPENDS_ON = ["03", "04", "05", "06", "07"]
EXACT_ROW_COUNTS = {1: 1000, 3: 480, 5: 4, 6: 480, 7: 2} # fixed by corpus volumes
# --- SQL implementations -------------------------------------------------
# Q2/Q3/Q4/Q7 use the precomputed track_summary (the relational strength);
# Q5 must scan raw friction_cycles in every format (spec 08).
Q1_SQL = f"""-- Q1: COF vs cycle curve for track {Q1_TRACK_CODE} (point read)
SELECT fc.cycle, fc.cof
FROM friction_cycles fc
JOIN tracks t ON t.track_id = fc.track_id
WHERE t.track_code = '{Q1_TRACK_CODE}'
ORDER BY fc.cycle;
"""
Q2_SQL = """-- Q2: steady-state COF per coupon with mean Au = 10 +/- 0.5 wt%
SELECT c.coupon_code, t.environment, AVG(s.cof_ss_mean) AS cof_ss
FROM coupons c
JOIN tracks t ON t.coupon_id = c.coupon_id
JOIN track_summary s ON s.track_id = t.track_id
WHERE c.au_wtpct_mean BETWEEN 9.5 AND 10.5
GROUP BY c.coupon_code, t.environment
ORDER BY c.coupon_code;
"""
Q3_SQL = """-- Q3: hardness vs steady-state COF per friction coupon, across all batches
SELECT c.coupon_code,
b.batch_code,
(SELECT AVG(n.hardness_gpa) FROM nanoindentation n WHERE n.coupon_id = c.coupon_id) AS hardness_gpa,
AVG(s.cof_ss_mean) AS cof_ss
FROM coupons c
JOIN batches b ON b.batch_id = c.batch_id
JOIN tracks t ON t.coupon_id = c.coupon_id
JOIN track_summary s ON s.track_id = t.track_id
GROUP BY c.coupon_id, c.coupon_code, b.batch_code
ORDER BY c.coupon_code;
"""
Q4_SQL = """-- Q4: anomaly filter - Dry N2 tracks at 100 mN with cof_ss > 0.20
SELECT t.track_code, s.cof_ss_mean
FROM tracks t
JOIN track_summary s ON s.track_id = t.track_id
WHERE t.environment = 'dry_n2' AND t.load_mn = 100 AND s.cof_ss_mean > 0.20
ORDER BY t.track_code;
"""
# Q5: forced full scan over raw cycles; the run-in algorithm is inlined
# (docs/rules/db-sql-schema.md section 5). SQLite has no stddev built-in,
# so the 2-sigma band compares squares: (cof-m)^2 < 4*var <=> |cof-m| < 2s.
Q5_SQLITE = """-- Q5: mean run-in cycles per batch over ALL raw cycle rows (track_summary forbidden)
WITH tail AS (
SELECT track_id, AVG(cof) AS m, AVG(cof * cof) - AVG(cof) * AVG(cof) AS var
FROM friction_cycles
WHERE cycle >= 501
GROUP BY track_id
),
run_in AS (
SELECT fc.track_id,
COALESCE(MIN(CASE WHEN (fc.cof - tl.m) * (fc.cof - tl.m) < 4 * tl.var THEN fc.cycle END), 500) AS run_in
FROM friction_cycles fc
JOIN tail tl ON tl.track_id = fc.track_id
GROUP BY fc.track_id
)
SELECT b.batch_code, AVG(r.run_in) AS avg_run_in_cycles
FROM run_in r
JOIN tracks t ON t.track_id = r.track_id
JOIN coupons c ON c.coupon_id = t.coupon_id
JOIN batches b ON b.batch_id = c.batch_id
GROUP BY b.batch_code
ORDER BY b.batch_code;
"""
Q5_PG = """-- Q5: mean run-in cycles per batch over ALL raw cycle rows (track_summary forbidden)
WITH tail AS (
SELECT track_id, avg(cof) AS m, stddev_pop(cof) AS s
FROM friction_cycles
WHERE cycle >= 501
GROUP BY track_id
),
run_in AS (
SELECT fc.track_id,
COALESCE(MIN(fc.cycle) FILTER (WHERE abs(fc.cof - tl.m) < 2 * tl.s), 500) AS run_in
FROM friction_cycles fc
JOIN tail tl ON tl.track_id = fc.track_id
GROUP BY fc.track_id
)
SELECT b.batch_code, AVG(r.run_in)::double precision AS avg_run_in_cycles
FROM run_in r
JOIN tracks t ON t.track_id = r.track_id
JOIN coupons c ON c.coupon_id = t.coupon_id
JOIN batches b ON b.batch_id = c.batch_id
GROUP BY b.batch_code
ORDER BY b.batch_code;
"""
Q6_SQL = """-- Q6: mean wear volume vs load per friction coupon
SELECT c.coupon_code, t.load_mn, AVG(w.wear_volume_um3) AS wear_volume_um3
FROM wear w
JOIN tracks t ON t.track_id = w.track_id
JOIN coupons c ON c.coupon_id = t.coupon_id
GROUP BY c.coupon_id, c.coupon_code, t.load_mn
ORDER BY c.coupon_code;
"""
Q7_SQL = """-- Q7: Stribeck-style mean COF by (environment, speed*load bucket) past run-in
SELECT t.environment,
CAST(ROUND(t.speed_mm_s * t.load_mn) AS INTEGER) AS speed_load_bucket,
AVG(fc.cof) AS mean_cof
FROM friction_cycles fc
JOIN tracks t ON t.track_id = fc.track_id
JOIN track_summary s ON s.track_id = fc.track_id
WHERE fc.cycle > s.run_in_cycles
GROUP BY t.environment, speed_load_bucket
ORDER BY t.environment, speed_load_bucket;
"""
SQLITE_SQL = {1: Q1_SQL, 2: Q2_SQL, 3: Q3_SQL, 4: Q4_SQL, 5: Q5_SQLITE, 6: Q6_SQL, 7: Q7_SQL}
PG_SQL = {1: Q1_SQL, 2: Q2_SQL, 3: Q3_SQL, 4: Q4_SQL, 5: Q5_PG, 6: Q6_SQL, 7: Q7_SQL}
RDF_SPARQL_FILES = {
1: ("q1",), 2: ("q2",), 3: ("q3_hardness", "q3_cycles"),
4: ("q4",), 5: ("q5",), 6: ("q6",), 7: ("q7",),
}
WRAPPER = '''"""Benchmark query q{n} for the {fmt} format (generated by task 08).
Prints the canonical result rows to stdout at full float precision; the
benchmark runner (task 09) captures and validates them against
out/bench/expected/q{n}.rows.csv.
"""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[4]
sys.path.insert(0, str(REPO_ROOT))
from common.queries.{module} import q{n}
from common.results import print_rows
if __name__ == "__main__":
print_rows(q{n}(REPO_ROOT / "out" / {subpath}))
'''
PY_FORMATS = [
("csv", "csv_queries", '"csv"'),
("json", "json_queries", '"json"'),
("rdf", "rdf_queries", '"rdf" / "oxigraph_store"'),
]
SQLITE_WRAPPER = '''"""Benchmark query q{n} for the sqlite format (generated by task 08).
Executes the adjacent q{n}.sql against out/sqlite/tribo.db (read-only) and
prints the canonical result rows to stdout at full float precision; the
benchmark runner (task 09) captures and validates them against
out/bench/expected/q{n}.rows.csv.
"""
import sqlite3
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[4]
sys.path.insert(0, str(REPO_ROOT))
from common.results import print_rows
if __name__ == "__main__":
sql = Path(__file__).with_suffix(".sql").read_text(encoding="ascii")
db = REPO_ROOT / "out" / "sqlite" / "tribo.db"
conn = sqlite3.connect(f"file:{{db.as_posix()}}?mode=ro", uri=True)
print_rows([tuple(r) for r in conn.execute(sql).fetchall()])
'''
PG_WRAPPER = '''"""Benchmark query q{n} for the pg format (generated by task 08).
Executes the adjacent q{n}.sql on the PostgreSQL instance addressed by
TRIBO_PG_DSN and prints the canonical result rows to stdout at full float
precision; the benchmark runner (task 09) captures and validates them
against out/bench/expected/q{n}.rows.csv.
"""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[4]
sys.path.insert(0, str(REPO_ROOT))
import psycopg
from common.pg import get_dsn
from common.results import print_rows
if __name__ == "__main__":
sql = Path(__file__).with_suffix(".sql").read_text(encoding="ascii")
with psycopg.connect(get_dsn()) as conn:
print_rows([tuple(r) for r in conn.execute(sql).fetchall()])
'''
SQL_WRAPPERS = {"sqlite": SQLITE_WRAPPER, "pg": PG_WRAPPER}
def write_query_files(bench_dir: Path) -> int:
queries_dir = bench_dir / "queries"
written = 0
for fmt, module, subpath in PY_FORMATS:
fmt_dir = queries_dir / fmt
fmt_dir.mkdir(parents=True, exist_ok=True)
for n in range(1, 8):
text = WRAPPER.format(n=n, fmt=fmt, module=module, subpath=subpath)
(fmt_dir / f"q{n}.py").write_text(text, encoding="ascii", newline="\n")
written += 1
for fmt, sql in (("sqlite", SQLITE_SQL), ("pg", PG_SQL)):
fmt_dir = queries_dir / fmt
fmt_dir.mkdir(parents=True, exist_ok=True)
for n, text in sql.items():
(fmt_dir / f"q{n}.sql").write_text(text, encoding="ascii", newline="\n")
(fmt_dir / f"q{n}.py").write_text(SQL_WRAPPERS[fmt].format(n=n), encoding="ascii", newline="\n")
written += 2
rdf_dir = queries_dir / "rdf"
for n, keys in RDF_SPARQL_FILES.items():
text = "\n\n".join(f"# retrieval query: {key}\n{rdf_queries.SPARQL[key]}" for key in keys)
(rdf_dir / f"q{n}.sparql").write_text(text + "\n", encoding="ascii", newline="\n")
written += 1
logger.info("wrote %d query files under %s", written, queries_dir)
return written
def run_sqlite(db_path: Path) -> dict[int, list[tuple]]:
conn = sqlite3.connect(db_path)
try:
return {n: [tuple(r) for r in conn.execute(sql).fetchall()] for n, sql in SQLITE_SQL.items()}
finally:
conn.close()
def run_pg(dsn: str) -> dict[int, list[tuple]]:
results: dict[int, list[tuple]] = {}
with psycopg.connect(dsn) as conn:
for n, sql in PG_SQL.items():
results[n] = [tuple(r) for r in conn.execute(sql).fetchall()]
logger.info("pg q%d: %d rows", n, len(results[n]))
return results
def main() -> int:
args = parse_task_args("Task 08: define benchmark queries and canonical expected results")
out_root: Path = args.out_root
bench_dir = out_root / "bench"
try:
check_dependencies(out_root, DEPENDS_ON)
remove_stale_marker(out_root, TASK_ID)
files_written = write_query_files(bench_dir)
logger.info("computing canonical results on PostgreSQL")
pg_rows = run_pg(get_dsn())
logger.info("cross-validating against SQLite")
sqlite_rows = run_sqlite(out_root / "sqlite" / "tribo.db")
for n in range(1, 8):
diff = compare(pg_rows[n], sqlite_rows[n])
if diff:
raise ValueError(f"q{n}: SQLite disagrees with PostgreSQL: {diff}")
exact = EXACT_ROW_COUNTS.get(n)
if exact is not None and len(pg_rows[n]) != exact:
raise ValueError(f"q{n}: {len(pg_rows[n])} rows, expected {exact}")
if not pg_rows[n]:
raise ValueError(f"q{n}: empty result set")
logger.info("spot-checking csv/json/rdf on Q1")
for label, rows in (
("csv", csv_queries.q1(out_root / "csv")),
("json", json_queries.q1(out_root / "json")),
("rdf", rdf_queries.q1(out_root / "rdf" / "oxigraph_store")),
):
diff = compare(pg_rows[1], rows)
if diff:
raise ValueError(f"q1 via {label} disagrees with canonical: {diff}")
expected_dir = bench_dir / "expected"
expected_dir.mkdir(parents=True, exist_ok=True)
checksums = {}
for n in range(1, 8):
rows = sort_rows(pg_rows[n])
(expected_dir / f"q{n}.rows.csv").write_text(serialize(rows), encoding="ascii", newline="\n")
checksums[n] = sha256_of(rows)
(expected_dir / f"q{n}.sha256").write_text(checksums[n] + "\n", encoding="ascii", newline="\n")
entries = {"query_files": files_written}
for n in range(1, 8):
entries[f"q{n}_rows"] = len(pg_rows[n])
entries[f"q{n}_sha256"] = checksums[n][:12]
entries.update(process_metrics())
marker_path = write_marker(out_root, TASK_ID, entries)
except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1
counts = ", ".join(f"q{n}={len(pg_rows[n])}" for n in range(1, 8))
print(f"task 08 ok: {files_written} query files, canonical results validated (pg == sqlite, 1e-9)")
print(f"rows: {counts}")
print(f"expected: {bench_dir / 'expected'} (q<N>.rows.csv + q<N>.sha256)")
print(f"marker: {marker_path}")
return 0
if __name__ == "__main__":
sys.exit(main())

13
out/.done/01.ok Normal file
View File

@@ -0,0 +1,13 @@
task: '01'
status: ok
config_file: C:/Projects/Public/LabDataStorageEvaluation/out/config/lab_config.yaml
config_bytes: 5827
seed: 20260711
batches: 4
wafers: 12
coupons: 588
friction_coupons: 480
tracks: 1440
cycle_rows: 1440000
loop_points: 2880000
xrf_points: 235200

6
out/.done/02.ok Normal file
View File

@@ -0,0 +1,6 @@
task: '02'
status: ok
process_flow_file: C:/Projects/Public/LabDataStorageEvaluation/out/report/process_flow.md
process_flow_bytes: 4205
diagrams: 6
diagram_bytes: 2924

11
out/.done/03.ok Normal file
View File

@@ -0,0 +1,11 @@
task: '03'
status: ok
csv_root: C:/Projects/Public/LabDataStorageEvaluation/out/csv
files: 8718
total_rows: 4636776
total_bytes: 148654472
coupons: 588
tracks: 1440
cycle_rows: 1440000
loop_points: 2880000
xrf_points: 235200

11
out/.done/04.ok Normal file
View File

@@ -0,0 +1,11 @@
task: '04'
status: ok
full_files: 588
full_bytes: 311300222
hybrid_bytes: 1887322
coupons: 588
tracks: 1440
cycle_rows: 1440000
loop_points: 2880000
proc_peak_rss_mb: 119.1
proc_cpu_seconds: 19.1

12
out/.done/05.ok Normal file
View File

@@ -0,0 +1,12 @@
task: '05'
status: ok
db_file: C:/Projects/Public/LabDataStorageEvaluation/out/sqlite/tribo.db
db_bytes: 123400192
tables: 15
total_rows: 4638223
friction_cycles: 1440000
friction_loop_points: 2880000
tracks: 1440
track_summary: 1440
proc_peak_rss_mb: 323.0
proc_cpu_seconds: 19.9

18
out/.done/06.ok Normal file
View File

@@ -0,0 +1,18 @@
task: '06'
status: ok
database: lab_data
tables: 15
total_rows: 4638223
live_bytes: 397467648
friction_cycles_bytes: 99753984
friction_loop_points_bytes: 267796480
copy_seconds: 19.8
total_seconds: 23.0
dump: skipped (pg_dump unavailable on this host)
proc_peak_rss_mb: 58.6
proc_cpu_seconds: 18.7
host_cpu_avg_pct: 2.4
host_cpu_max_pct: 3.7
host_ram_used_baseline_mb: 617
host_ram_used_peak_mb: 681
host_ram_used_delta_mb: 64

13
out/.done/07.ok Normal file
View File

@@ -0,0 +1,13 @@
task: '07'
status: ok
triples: 17419713
nt_gz_bytes: 126487650
ttl_bytes: 596401229
store_bytes: 1739414085
store_load_seconds: 33.6
coupons: 588
tracks: 1440
cycle_nodes: 1440000
loop_nodes: 2880000
proc_peak_rss_mb: 2185.2
proc_cpu_seconds: 151.8

19
out/.done/08.ok Normal file
View File

@@ -0,0 +1,19 @@
task: '08'
status: ok
query_files: 56
q1_rows: 1000
q1_sha256: bb6d871275d0
q2_rows: 33
q2_sha256: 62ca30e729a8
q3_rows: 480
q3_sha256: a3a73414201f
q4_rows: 469
q4_sha256: 38281dcf7896
q5_rows: 4
q5_sha256: f3ee32dbc981
q6_rows: 480
q6_sha256: 9cc28d51ee35
q7_rows: 2
q7_sha256: 1991f8bf2217
proc_peak_rss_mb: 61.1
proc_cpu_seconds: 1.7

14
out/.done/09.ok Normal file
View File

@@ -0,0 +1,14 @@
task: '09'
status: ok
cells_total: 35
cells_ok: 35
cells_invalid: 0
cells_timeout: 0
cells_error: 0
session_wall_s: 915.5
cold_cache_pass: skipped (Windows host)
session_host_cpu_avg_pct: 0.4
session_host_cpu_max_pct: 9.2
session_host_ram_used_baseline_mb: 909
session_host_ram_used_peak_mb: 914
session_host_ram_used_delta_mb: 5

16
out/.done/10.ok Normal file
View File

@@ -0,0 +1,16 @@
task: '10'
status: ok
extrapolation_rows: 105
sizing_rows: 15
flags_ok: 53
flags_impractical: 35
flags_fail: 17
pg_index_share: 0.341
pg_index_mb: 128.5
prices_as_of: 2026-07-11
note_pandas: pandas CSV variant not measured; its RAM is O(n) - infeasible above ~10 GB (spec 10)
coeff_csv: 1.0
coeff_json: 2.094
coeff_sqlite: 0.83
coeff_pg: 2.674
coeff_rdf: 11.701

12
out/.done/11.ok Normal file
View File

@@ -0,0 +1,12 @@
task: '11'
status: ok
charts: 14
tables: 7
report_bytes: 16091
winner_600gb: pg
prices_as_of: 2026-07-11
score600_csv: 28.9
score600_json: 24.0
score600_sqlite: 47.0
score600_pg: 58.6
score600_rdf: 0.9

18
out/config/hw_prices.yaml Normal file
View File

@@ -0,0 +1,18 @@
# Hardware cost parameters for extrapolation (task 10).
# Street prices retrieved 2026-07-11 from the sources below (server DRAM is
# in a documented 2025-2026 price surge - see market_context). The cost
# model buys whole components: 64 GB RDIMM modules and 7.68 TB NVMe drives.
# Operator-editable; re-run extrapolate.py after changes (code-config-yaml).
as_of: "2026-07-11"
ram_module_gb: 64
ram_module_usd: 1887.0 # A-Tech 64GB (2x32GB) DDR5-5600 ECC RDIMM
nvme_drive_tb: 7.68
nvme_drive_usd: 3995.0 # Cloud Ninjas NEW 7.68TB NVMe U.2 1DWPD (Dell 14-16G)
server_16_core_usd: 4618.0 # Supermicro AS-1015CS-TNR 1U (EPYC 9004/9005), Broadberry starting config
server_32_core_usd: 6272.0 # Supermicro AS-1115CS-TNR 1U, Broadberry starting config
extra_node_usd: 4618.0 # one additional entry 1U node (AS-1015CS-TNR chassis)
sources:
ram: https://atechmemory.com/collections/ddr5-memory-ram
nvme: https://cloudninjas.com/products/new-7-68tb-nvme-u-2-1dwpd-sie-2-5-enterprise-solid-state-drive-for-14th-15th-16th-gen-dell
servers: https://www.broadberry.com/amd-epyc-9004-supermicro-servers
market_context: https://www.techpowerup.com/342331/server-dram-pricing-jumps-50-only-70-of-orders-getting-filled

246
out/config/lab_config.yaml Normal file
View File

@@ -0,0 +1,246 @@
# Generated by make_lab_config.py (task 01) from docs/specs/01_lab_configuration.md.
# Do not hand-edit: regenerate via `python make_lab_config.py`.
project: LabDataStorageEvaluation
description: Simulated tribology laboratory (Sandia Pt-Au LDRD context)
spec: docs/specs/01_lab_configuration.md
seed: 20260711
material_system:
substrate:
material: Ti-6Al-4V
dimensions_mm:
- 10
- 10
- 3
adhesion_layer:
material: Cr
process: sputtered
coating:
material: Pt-Au
deposition: composition gradient across each wafer
thickness_um_range:
- 0.3
- 1.1
instruments:
- instrument_id: rapid
role: friction
name: RAPID custom high-throughput parallelized 6-probe tribometer
data_produced: COF vs cycle per track; test conditions
- instrument_id: ti980
role: nanoindentation
name: Bruker TI980 TriboIndenter
data_produced: hardness, reduced modulus (25 indents per coupon)
- instrument_id: m4_tornado
role: composition
name: Bruker M4 Tornado micro-XRF
data_produced: Pt/Au wt% map per coupon (20x20 grid)
- instrument_id: pvd200
role: deposition
name: Kurt J. Lesker PVD 200 sputter-down
data_produced: batch deposition parameters
- instrument_id: afm
role: surface_roughness
name: AFM
data_produced: topography image metadata + Ra/Rq per coupon
- instrument_id: profilometer
role: film_thickness
name: Optical profilometry
data_produced: thickness map per coupon (10x10 grid)
- instrument_id: simtra
role: simulation
name: SIMTRA sputter transport Monte-Carlo
data_produced: deposition atom-energy / composition profiles per deposition run
deposition_matrix:
- batch_code: B721
pt_gun_tilt_deg: 20
au_gun_tilt_deg: 0
pt_power_W: 150
au_power_W: 50
pt_discharge_V: 432
au_discharge_V: 311
- batch_code: B722
pt_gun_tilt_deg: 20
au_gun_tilt_deg: 20
pt_power_W: 100
au_power_W: 100
pt_discharge_V: 399
au_discharge_V: 352
- batch_code: B723
pt_gun_tilt_deg: 20
au_gun_tilt_deg: 20
pt_power_W: 150
au_power_W: 50
pt_discharge_V: 430
au_discharge_V: 311
- batch_code: B724
pt_gun_tilt_deg: 0
au_gun_tilt_deg: 20
pt_power_W: 50
au_power_W: 150
pt_discharge_V: 376
au_discharge_V: 340
hierarchy:
batches: 4
wafers_per_batch: 3
coupon_grid_per_wafer:
- 7
- 7
coupons_per_wafer: 49
friction_assignment:
friction_coupons_total: 480
reserve_coupons_total: 108
runs: 4
coupons_per_run: 120
plates_per_run: 5
probes_per_plate: 6
coupons_per_probe_square: 4
tracks_per_friction_coupon: 3
counterfaces_per_holder: 3
fresh_counterface_per_track: true
reserve_grid_positions:
- 1
- 4
- 7
- 22
- 25
- 28
- 43
- 46
- 49
volumes:
coupons_total: 588
tracks_total: 1440
cycles_per_track: 1000
cycle_rows_total: 1440000
loop_every_n_cycles: 100
loops_per_track: 10
loop_points_per_loop: 200
loop_points_total: 2880000
xrf_grid:
- 20
- 20
xrf_points_per_coupon: 400
xrf_points_total: 235200
profilometry_grid:
- 10
- 10
profilometry_points_per_coupon: 100
nanoindentation_indents_per_coupon: 25
simtra_rows_per_batch: 1000
runs:
- run_code: R1
batch_code: B721
environment: lab_air
date: '2026-04-06'
operator: A. Reyes
- run_code: R2
batch_code: B722
environment: lab_air
date: '2026-04-13'
operator: K. Patel
- run_code: R3
batch_code: B723
environment: dry_n2
date: '2026-04-20'
operator: A. Reyes
- run_code: R4
batch_code: B724
environment: dry_n2
date: '2026-04-27'
operator: M. Novak
schedule:
deposition_start_date: '2026-03-02'
batch_interval_days: 7
characterization_lag_days: 3
run_start_time: 09:00:00
track_interval_s: 2100
test_conditions:
normal_load_mN: 100
stroke_mm: 1.0
speed_mm_s: 1.0
counterface:
material: ruby (Al2O3)
diameter_mm: 3.175
rh_pct:
lab_air: 45
dry_n2: 2
temperature_C: 23
temperature_tolerance_C: 1
physical_models:
au_gradient:
description: linear Au wt% gradient along wafer x, batch-dependent center
batch_center_wtpct:
B721: 8
B722: 25
B723: 10
B724: 60
wafer_span_wtpct: 5
xrf_noise_sd_wtpct: 0.3
cof_run_in:
formula: cof(c) = cof_ss + (cof_0 - cof_ss) * exp(-c / tau) + N(0, sigma)
cof_0_range:
- 0.35
- 0.5
tau_cycles_range:
- 100
- 400
noise_sd: 0.01
cof_steady_state:
formula: cof_ss = intercept + slope_per_au_wtpct * au_wtpct_mean, clamped to >= floor
lab_air:
intercept: 0.32
slope_per_au_wtpct: -0.0015
dry_n2:
intercept: 0.24
slope_per_au_wtpct: -0.0012
floor: 0.12
hardness_GPa:
formula: H = intercept + slope_per_au_wtpct * au_wtpct_mean + N(0, noise_sd)
intercept: 8.5
slope_per_au_wtpct: -0.05
noise_sd: 0.25
reduced_modulus_GPa:
formula: Er = intercept + slope_per_au_wtpct * au_wtpct_mean + N(0, noise_sd)
intercept: 190
slope_per_au_wtpct: -0.6
noise_sd: 4
roughness_ra_nm:
distribution: lognormal
median_nm: 5
sigma_log: 0.3
rq_over_ra: 1.25
thickness_um:
profile: radial parabolic across wafer plus noise
center_um: 1.05
edge_um: 0.35
noise_sd_um: 0.02
range_um:
- 0.3
- 1.1
wear_archard:
formula: V = k * F * s; k = k0 * (1 + slope_per_au_wtpct * au_wtpct_mean), lognormal noise, clamped
to >= k_floor
k0_mm3_per_N_m:
lab_air: 1.5e-07
dry_n2: 6.0e-08
slope_per_au_wtpct:
lab_air: -0.008
dry_n2: -0.01
k_floor_mm3_per_N_m: 1.0e-09
noise_sigma_log: 0.15
friction_loop:
force_noise_sd_mN: 0.3
nanoindentation:
max_load_mN: 10
indent_grid:
- 5
- 5
grid_pitch_um: 20
simtra:
angle_deg_range:
- 0
- 90
energy_eV_max: 50
surface_binding_energy_eV:
pt: 5.84
au: 3.81
flux_per_W: 0.02

318
out/report/REPORT.md Normal file
View File

@@ -0,0 +1,318 @@
# Laboratory Data Storage Format Evaluation - Final Report
Generated by task 11 (report.py) from the measured benchmark matrix (task 09) and the extrapolation model (task 10). Every number at scales beyond the 150 MB corpus is a PROJECTION and is labeled as such; scaling assumptions are listed in section 6.
## 1. Executive summary
**Recommendation: PostgreSQL as the system of record.** The measured matrix and the projections confirm the expected outcome:
- **PostgreSQL** is the only format whose seven benchmark queries stay below the 1-hour practicality bound at every projected scale (600 GB, 1.2 TB, 6 TB); at 6 TB it needs one 16-core node with 551 GB RAM at an estimated $33,586 (bill of materials, prices as of 2026-07-11). Partitioned parallel scans and the track_summary materialized view are the scaling levers.
- **CSV** remains the canonical raw archive: byte-reproducible, instrument-native, and the best compressor (tar.gz ~2.8x); its full scans exceed 1 h beyond ~600 GB, so it is an archive, not a query layer.
- **SQLite** is an excellent single-user store up to ~600 GB (all queries OK), then its single-threaded scans degrade (Q5/Q7 IMPRACTICAL at 1.2 TB).
- **Hybrid JSON-LD** (metadata + sourceFile links, 1.8 MiB at 150 MB corpus) is the exchange format; the FULL variant re-reads the whole corpus per scan and becomes impractical past 600 GB.
- **RDF (materialized triplestore)** fails at scale: 11.7x storage blow-up, 5 of 7 queries FAIL at 6 TB, and the hot set needs 13.7 TB RAM across 14 nodes (~$527,732 projected). Semantic access should be a virtual layer (e.g. Ontop OBDA) over PostgreSQL instead.
## 2. Methodology snapshot
35 cells (Q1-Q7 x 5 formats); per cell 1 warm-up + 3 measured runs in isolated subprocesses sampled with psutil at 50 ms; randomized cell order; 30-min timeout; every run validated against canonical results (PostgreSQL cross-validated with SQLite, tolerance 1e-9). Medians are reported; raw runs and min/max live in out/bench/results_raw.csv. The cold-cache pass is skipped on this Windows host and marked as such. PostgreSQL runs on a remote host; its wall times include the LAN round-trip, and server-side metrics come from pg_stat_statements and Prometheus (docs/rules/bench-methodology.md section 6).
## 3. Environment
```
benchmark_host_os: Windows-2025Server-10.0.26100-SP0
benchmark_host_cpu: Intel64 Family 6 Model 141 Stepping 1, GenuineIntel
benchmark_host_cores_logical: 4
benchmark_host_cores_physical: 4
benchmark_host_ram_gib: 8.0
python: 3.12.10
benchmark_host_disks: VMware Virtual NVMe Disk / SSD / 100 GB
sqlite: 3.49.1
psycopg: 3.3.4
pyoxigraph: 0.5.9
ijson: 3.5.1
psutil: 7.2.2
postgresql_server: PostgreSQL 16.14 (Ubuntu 16.14-0ubuntu0.24.04.1) on x86_64-pc-linux-gnu, compiled by gcc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0, 64-bit
postgresql_host: 192.168.10.73 (remote; 4 cores / 3.8 GiB, see .done/06.ok)
cold_cache_pass: skipped (Windows host, no page-cache drop - bench-methodology section 6)
pg_backend_sampling: not possible for a remote server; Prometheus windows + pg_stat_statements deltas instead
```
## 4. Storage footprint
| format | measured | vs CSV | archival (compressed) | projected @600 GB | projected @1.2 TB | projected @6 TB |
|---|---|---|---|---|---|---|
| CSV | 149 MB | 1.00x | 53 MB | 600 GB | 1.20 TB | 6.00 TB |
| JSON-LD | 311 MB | 2.09x | 57 MB | 1.26 TB | 2.51 TB | 13 TB |
| SQLite | 123 MB | 0.83x | 55 MB | 498 GB | 996 GB | 4.98 TB |
| PostgreSQL | 397 MB | 2.67x | 53 MB | 1.60 TB | 3.21 TB | 16 TB |
| RDF triplestore | 1.74 GB | 11.70x | 126 MB | 7.02 TB | 14 TB | 70 TB |
![C1 - disk footprint](charts/c1_disk_footprint.png)
## 5. Measured performance (150 MB corpus)
| format | query | wall time | peak RSS | CPU util | read |
|---|---|---|---|---|---|
| CSV | q1 | 0.06 s | 4 MB | 0% | 0.0 MB |
| CSV | q2 | 0.17 s | 20 MB | 16% | 2.5 MB |
| CSV | q3 | 1.15 s | 21 MB | 24% | 31.8 MB |
| CSV | q4 | 0.71 s | 21 MB | 22% | 16.8 MB |
| CSV | q5 | 1.09 s | 21 MB | 23% | 31.8 MB |
| CSV | q6 | 0.28 s | 20 MB | 17% | 1.2 MB |
| CSV | q7 | 1.27 s | 21 MB | 23% | 32.5 MB |
| JSON-LD | q1 | 0.11 s | 22 MB | 10% | 1.4 MB |
| JSON-LD | q2 | 1.48 s | 24 MB | 24% | 53.1 MB |
| JSON-LD | q3 | 9.94 s | 24 MB | 25% | 297.6 MB |
| JSON-LD | q4 | 9.92 s | 24 MB | 25% | 297.8 MB |
| JSON-LD | q5 | 9.96 s | 24 MB | 25% | 297.6 MB |
| JSON-LD | q6 | 9.79 s | 24 MB | 24% | 297.8 MB |
| JSON-LD | q7 | 10 s | 24 MB | 24% | 297.9 MB |
| SQLite | q1 | 0.06 s | 4 MB | 0% | 0.0 MB |
| SQLite | q2 | 0.06 s | 4 MB | 0% | 0.0 MB |
| SQLite | q3 | 0.11 s | 4 MB | 0% | 0.0 MB |
| SQLite | q4 | 0.06 s | 4 MB | 0% | 0.0 MB |
| SQLite | q5 | 0.77 s | 25 MB | 23% | 117.9 MB |
| SQLite | q6 | 0.06 s | 4 MB | 0% | 0.0 MB |
| SQLite | q7 | 0.71 s | 25 MB | 23% | 93.0 MB |
| PostgreSQL | q1 | 0.22 s | 42 MB | 18% | 4.1 MB |
| PostgreSQL | q2 | 0.22 s | 42 MB | 20% | 4.1 MB |
| PostgreSQL | q3 | 0.22 s | 42 MB | 18% | 4.1 MB |
| PostgreSQL | q4 | 0.22 s | 42 MB | 18% | 4.1 MB |
| PostgreSQL | q5 | 0.44 s | 42 MB | 10% | 4.1 MB |
| PostgreSQL | q6 | 0.22 s | 42 MB | 19% | 4.1 MB |
| PostgreSQL | q7 | 0.33 s | 42 MB | 12% | 4.1 MB |
| RDF triplestore | q1 | 0.11 s | 32 MB | 17% | 3.6 MB |
| RDF triplestore | q2 | 2.81 s | 85 MB | 24% | 375.6 MB |
| RDF triplestore | q3 | 46 s | 264 MB | 25% | 5445.1 MB |
| RDF triplestore | q4 | 20 s | 168 MB | 25% | 2694.6 MB |
| RDF triplestore | q5 | 60 s | 264 MB | 25% | 5390.5 MB |
| RDF triplestore | q6 | 0.17 s | 46 MB | 16% | 15.9 MB |
| RDF triplestore | q7 | 41 s | 264 MB | 25% | 5380.5 MB |
Sub-interval note: cells faster than the 50 ms sampling cadence (fast SQLite/CSV point reads) under-report client RSS/CPU; wall times are exact.
![C2 - RAM requirement](charts/c2_ram_requirement.png)
Per-query charts: [Q1](charts/c3_q1_wall_time.png), [Q2](charts/c3_q2_wall_time.png), [Q3](charts/c3_q3_wall_time.png), [Q4](charts/c3_q4_wall_time.png), [Q5](charts/c3_q5_wall_time.png), [Q6](charts/c3_q6_wall_time.png), [Q7](charts/c3_q7_wall_time.png)
## 6. Projections (600 GB / 1.2 TB / 6 TB)
All values in this section are projected, never measured. Laws per access pattern (spec 10): flat / index-depth log / linear result set / linear x depth / full scan / parallel partitioned scan; constants calibrated on the measured point after subtracting the per-format harness floor. Flags: IMPRACTICAL > 1 h, FAIL > 24 h.
Projected wall times at 600 GB:
| format | query | projected wall time | flag |
|---|---|---|---|
| CSV | q1 | 0.06 s | OK |
| CSV | q2 | 7.1 min | OK |
| CSV | q3 | 73.5 min | IMPRACTICAL |
| CSV | q4 | 44.1 min | OK |
| CSV | q5 | 69.6 min | IMPRACTICAL |
| CSV | q6 | 14.5 min | OK |
| CSV | q7 | 81.6 min | IMPRACTICAL |
| JSON-LD | q1 | 0.11 s | OK |
| JSON-LD | q2 | 92.0 min | IMPRACTICAL |
| JSON-LD | q3 | 11.0 h | IMPRACTICAL |
| JSON-LD | q4 | 11.0 h | IMPRACTICAL |
| JSON-LD | q5 | 11.0 h | IMPRACTICAL |
| JSON-LD | q6 | 10.9 h | IMPRACTICAL |
| JSON-LD | q7 | 11.2 h | IMPRACTICAL |
| SQLite | q1 | 0.06 s | OK |
| SQLite | q2 | 1.27 s | OK |
| SQLite | q3 | 5.7 min | OK |
| SQLite | q4 | 0.06 s | OK |
| SQLite | q5 | 47.5 min | OK |
| SQLite | q6 | 5.18 s | OK |
| SQLite | q7 | 43.7 min | OK |
| PostgreSQL | q1 | 0.22 s | OK |
| PostgreSQL | q2 | 0.22 s | OK |
| PostgreSQL | q3 | 3.42 s | OK |
| PostgreSQL | q4 | 11 s | OK |
| PostgreSQL | q5 | 3.7 min | OK |
| PostgreSQL | q6 | 4.06 s | OK |
| PostgreSQL | q7 | 115 s | OK |
| RDF triplestore | q1 | 0.11 s | OK |
| RDF triplestore | q2 | 3.0 h | IMPRACTICAL |
| RDF triplestore | q3 | 2.1 d | FAIL |
| RDF triplestore | q4 | 22.1 h | IMPRACTICAL |
| RDF triplestore | q5 | 2.8 d | FAIL |
| RDF triplestore | q6 | 5.9 min | OK |
| RDF triplestore | q7 | 45.3 h | FAIL |
Degradation by data size: [point read](charts/c4_degradation_point_read.png), [indexed / join](charts/c4_degradation_indexed.png), [full scan](charts/c4_degradation_full_scan.png)
## 7. Hardware sizing and cost at 600 GB (projected)
Bill-of-materials costing with street prices as of **2026-07-11** (sources below). Every configuration buys whole components: 64 GB DDR5 ECC RDIMM modules, 7.68 TB enterprise NVMe U.2 drives, and a priced 1U chassis; disk capacity includes 30% free-space headroom.
| format | configuration | chassis | RAM | disk | extra nodes | total |
|---|---|---|---|---|---|---|
| CSV | 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 | $4,618 | $1,887 | $3,995 | $0 | $10,500 |
| JSON-LD | 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 | $4,618 | $1,887 | $3,995 | $0 | $10,500 |
| SQLite | 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 | $4,618 | $1,887 | $3,995 | $0 | $10,500 |
| PostgreSQL | 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 | $4,618 | $1,887 | $3,995 | $0 | $10,500 |
| RDF triplestore | 2 x 1U Supermicro AS-1015CS-TNR (16 cores); 22 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2 | $4,618 | $41,514 | $7,990 | $4,618 | $58,740 |
Component prices (retrieved 2026-07-11):
- 64 GB DDR5-5600 ECC RDIMM (A-Tech, 2x32GB kit): $1,887 - <https://atechmemory.com/collections/ddr5-memory-ram>
- 7.68 TB NVMe U.2 1DWPD enterprise SSD (Dell 14-16G compatible, Cloud Ninjas): $3,995 - <https://cloudninjas.com/products/new-7-68tb-nvme-u-2-1dwpd-sie-2-5-enterprise-solid-state-drive-for-14th-15th-16th-gen-dell>
- 1U AMD EPYC 9004/9005 chassis (Broadberry starting configurations): Supermicro AS-1015CS-TNR $4,618, Supermicro AS-1115CS-TNR $6,272 - <https://www.broadberry.com/amd-epyc-9004-supermicro-servers>
- Market context - server DRAM is in a documented 2025-2026 price surge: <https://www.techpowerup.com/342331/server-dram-pricing-jumps-50-only-70-of-orders-getting-filled>
Full sizing at 1.2 TB and 6 TB: out/bench/hardware_sizing.csv. RDF exceeds the 1 TB single-node RAM ceiling already at 600 GB and reaches 14 nodes at 6 TB.
![C6 - cost vs performance](charts/c6_cost_vs_performance.png)
## 8. Weighted scoring
Subscores are 10 x best/value per metric; weights: search x5, RAM economy x2, disk economy x1; max 80. Search uses the inverse geometric mean of Q1-Q7 (FAIL at 600 GB zeroes the search subscore). Left: measured scale; right: projected 600 GB.
Measured scale (150 MB):
| format | search (x5) | RAM (x2) | disk (x1) | total (max 80) |
|---|---|---|---|---|
| SQLite | 50.0 | 16.9 | 10.0 | 76.9 |
| CSV | 15.5 | 20.0 | 8.3 | 43.8 |
| PostgreSQL | 26.2 | 10.0 | 3.1 | 39.3 |
| JSON-LD | 1.7 | 17.2 | 4.0 | 22.9 |
| RDF triplestore | 1.3 | 1.6 | 0.7 | 3.6 |
Projected at 600 GB:
| format | search (x5) | RAM (x2) | disk (x1) | total (max 80) |
|---|---|---|---|---|
| PostgreSQL | 50.0 | 5.5 | 3.1 | 58.6 |
| SQLite | 21.8 | 15.2 | 10.0 | 47.0 |
| CSV | 0.6 | 20.0 | 8.3 | 28.9 |
| JSON-LD | 0.1 | 20.0 | 4.0 | 24.0 |
| RDF triplestore | 0.0 | 0.2 | 0.7 | 0.9 |
![C5 - weighted score](charts/c5_weighted_score.png)
## 9. Use-case mapping
| Use case | Recommended format(s) |
|---|---|
| Interactive analysis (joins, aggregations) | PostgreSQL; SQLite acceptable single-user up to ~600 GB |
| Report generation (repeated summaries) | PostgreSQL (track_summary materialized view) |
| Search / filtering | PostgreSQL or SQLite (indexed); flat formats need full scans |
| Archiving | CSV tree + tar.gz (canonical raw) plus pg_dump of the system of record |
| Inter-lab exchange | Hybrid JSON-LD (metadata + sourceFile links to CSV) |
| Semantic / ontology queries | Virtual RDF layer over PostgreSQL (e.g. Ontop OBDA), not a materialized triplestore |
## 10. Limitations
- Single-node measurements on one Windows host + one remote PostgreSQL host; no cluster or cloud variance.
- The corpus is simulated (physically plausible models, fixed seed 20260711); real instrument data may have different value distributions, though volumes and shapes match the lab.
- Extrapolation uses analytic laws calibrated on a single 150 MB point; no intermediate-scale validation runs were performed.
- Client wall times include a per-format harness floor (interpreter start, imports, connection); the floor is assumed constant across scales. The cheapest cell of each format therefore projects flat.
- The cold-cache matrix pass is skipped (no page-cache drop on Windows); all measured runs are warm-cache.
- The pandas CSV variant was not measured; its RAM is O(n) and it is infeasible above roughly 10 GB (spec 10).
- RDF numbers reflect the compact modeling and a client that derives run-in / steady-state from raw cycles; SPARQL-side aggregation engines could shift (not remove) the scan penalty.
- Hardware prices are spot street prices retrieved 2026-07-11 during a documented server-DRAM price surge; re-cost by editing out/config/hw_prices.yaml and re-running tasks 10-11.
## Appendix A - process flow diagrams (task 02)
### D1 - Coupon assembly
```mermaid
flowchart LR
A([Ti-6Al-4V Base 10x10x3 mm]) --> B[Cleaning]
B --> B1([Rinsed in deionized water])
B1 --> B2([Sonicated in cleaning solution])
B --> C[Smearing Adhesive]
C --> C1([Cr adhesive coating])
C --> D[Vapor Deposition System<br/>Kurt J. Lesker PVD 200]
D --> D1([PtAu sputter coating, gradient])
D --> E([Test Coupon])
```
### D2 - Characterization and batch assembly
```mermaid
flowchart LR
TC([Test Coupon]) --> S[SIMTRA simulation<br/>composition / atom energies]
S --> OP[Optical profilometry<br/>film thickness 0.3-1.1 um]
OP --> TW[Assemble Test Wafer<br/>x49 coupons]
TW --> TB[Assemble Test Batch<br/>x3 wafers]
TB --> B721[B721<br/>Pt 150 W / Au 50 W]
TB --> B722[B722<br/>Pt 100 W / Au 100 W]
TB --> B723[B723<br/>Pt 150 W / Au 50 W]
TB --> B724[B724<br/>Pt 50 W / Au 150 W]
```
### D3 - Testing tree
```mermaid
flowchart TD
TC([Test Coupon]) --> F[Friction<br/>RAPID 6-probe tribometer]
F --> FA([Lab Air -> COF dataset mu_normal])
F --> FN([Dry N2 -> COF dataset mu_dry_nit])
TC --> NI[Nanoindentation<br/>Bruker TI980]
NI --> NIR([Hardness, reduced modulus<br/>25 indents per coupon])
TC --> AF[AFM]
AF --> AFR([Topography -> Ra / Rq])
TC --> XR[micro-XRF<br/>Bruker M4 Tornado]
XR --> XRR([Pt/Au wt% map, 20x20 grid])
```
### D4 - Tribometer session sequence
```mermaid
sequenceDiagram
actor Op as Operator
participant Prep as Sample Prep
participant Plan as Excel Test Plan
participant SW as Control Software
participant TR as RAPID Tribometer
Op->>Prep: Request test
Op->>Prep: Get samples
Op->>Prep: Load samples onto plates
Op->>Prep: Load ball holders (3 counterfaces per holder)
Op->>Plan: Create test plan<br/>(Sample Plate, Plate Location, Sample ID, Save Location,<br/>Folder Name, X/Y/Z Offset, Load, Iterations)
Op->>TR: Transfer plates to tribometer
Op->>TR: Set up platter
Op->>SW: Activate software
Op->>SW: Load plan
SW->>TR: Activate tribometer
TR-->>SW: Run reciprocating tests (execution loop, see D5)
Op->>SW: Software stop
SW-->>Op: Save avg files
Op->>TR: Remove plates and platters
Op->>SW: Close software
Op->>TR: Pull equipment out
```
### D5 - Execution loop
```mermaid
flowchart TD
L[Load all counterfaces] --> P{5 plates}
P -->|for each plate| PR{6 probes in parallel}
PR -->|for each probe| C{4 coupons on coupon square}
C -->|for each coupon| T[Draw track]
T --> RC[Rotate counterface - fresh ball per track]
RC -->|3 tracks per coupon| T
T --> DONE([120 coupons x 3 tracks = 360 tracks per run])
```
### D6 - Data hierarchy
```mermaid
erDiagram
DEPOSITION_RUN ||--|| SIMTRA_PROFILE : produces
DEPOSITION_RUN ||--|| BATCH : creates
BATCH ||--|{ WAFER : contains
WAFER ||--|{ COUPON : contains
COUPON ||--|| XRF_MAP : has
COUPON ||--|| NANOINDENTATION : has
COUPON ||--|| AFM : has
COUPON ||--|| PROFILOMETRY : has
COUPON ||--o{ TRACK : "friction coupons only"
TRACK ||--|{ CYCLE : has
TRACK ||--|{ LOOP_POINT : has
TRACK ||--|| WEAR : has
```

Binary file not shown.

After

Width:  |  Height:  |  Size: 71 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 74 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 38 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 41 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 36 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 53 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 51 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 73 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 76 KiB

View File

@@ -0,0 +1,9 @@
flowchart LR
A([Ti-6Al-4V Base 10x10x3 mm]) --> B[Cleaning]
B --> B1([Rinsed in deionized water])
B1 --> B2([Sonicated in cleaning solution])
B --> C[Smearing Adhesive]
C --> C1([Cr adhesive coating])
C --> D[Vapor Deposition System<br/>Kurt J. Lesker PVD 200]
D --> D1([PtAu sputter coating, gradient])
D --> E([Test Coupon])

View File

@@ -0,0 +1,9 @@
flowchart LR
TC([Test Coupon]) --> S[SIMTRA simulation<br/>composition / atom energies]
S --> OP[Optical profilometry<br/>film thickness 0.3-1.1 um]
OP --> TW[Assemble Test Wafer<br/>x49 coupons]
TW --> TB[Assemble Test Batch<br/>x3 wafers]
TB --> B721[B721<br/>Pt 150 W / Au 50 W]
TB --> B722[B722<br/>Pt 100 W / Au 100 W]
TB --> B723[B723<br/>Pt 150 W / Au 50 W]
TB --> B724[B724<br/>Pt 50 W / Au 150 W]

View File

@@ -0,0 +1,10 @@
flowchart TD
TC([Test Coupon]) --> F[Friction<br/>RAPID 6-probe tribometer]
F --> FA([Lab Air -> COF dataset mu_normal])
F --> FN([Dry N2 -> COF dataset mu_dry_nit])
TC --> NI[Nanoindentation<br/>Bruker TI980]
NI --> NIR([Hardness, reduced modulus<br/>25 indents per coupon])
TC --> AF[AFM]
AF --> AFR([Topography -> Ra / Rq])
TC --> XR[micro-XRF<br/>Bruker M4 Tornado]
XR --> XRR([Pt/Au wt% map, 20x20 grid])

View File

@@ -0,0 +1,23 @@
sequenceDiagram
actor Op as Operator
participant Prep as Sample Prep
participant Plan as Excel Test Plan
participant SW as Control Software
participant TR as RAPID Tribometer
Op->>Prep: Request test
Op->>Prep: Get samples
Op->>Prep: Load samples onto plates
Op->>Prep: Load ball holders (3 counterfaces per holder)
Op->>Plan: Create test plan<br/>(Sample Plate, Plate Location, Sample ID, Save Location,<br/>Folder Name, X/Y/Z Offset, Load, Iterations)
Op->>TR: Transfer plates to tribometer
Op->>TR: Set up platter
Op->>SW: Activate software
Op->>SW: Load plan
SW->>TR: Activate tribometer
TR-->>SW: Run reciprocating tests (execution loop, see D5)
Op->>SW: Software stop
SW-->>Op: Save avg files
Op->>TR: Remove plates and platters
Op->>SW: Close software
Op->>TR: Pull equipment out

View File

@@ -0,0 +1,8 @@
flowchart TD
L[Load all counterfaces] --> P{5 plates}
P -->|for each plate| PR{6 probes in parallel}
PR -->|for each probe| C{4 coupons on coupon square}
C -->|for each coupon| T[Draw track]
T --> RC[Rotate counterface - fresh ball per track]
RC -->|3 tracks per coupon| T
T --> DONE([120 coupons x 3 tracks = 360 tracks per run])

View File

@@ -0,0 +1,13 @@
erDiagram
DEPOSITION_RUN ||--|| SIMTRA_PROFILE : produces
DEPOSITION_RUN ||--|| BATCH : creates
BATCH ||--|{ WAFER : contains
WAFER ||--|{ COUPON : contains
COUPON ||--|| XRF_MAP : has
COUPON ||--|| NANOINDENTATION : has
COUPON ||--|| AFM : has
COUPON ||--|| PROFILOMETRY : has
COUPON ||--o{ TRACK : "friction coupons only"
TRACK ||--|{ CYCLE : has
TRACK ||--|{ LOOP_POINT : has
TRACK ||--|| WEAR : has

127
out/report/process_flow.md Normal file
View File

@@ -0,0 +1,127 @@
# Process Flow - Simulated Tribology Laboratory
Generated by make_process_flow.py (task 02) from out/config/lab_config.yaml.
Standalone diagram sources: ./diagrams/D1..D6.mermaid.
## D1 - Coupon Assembly
Substrate preparation, adhesion layer, and gradient Pt-Au sputter deposition.
```mermaid
flowchart LR
A([Ti-6Al-4V Base 10x10x3 mm]) --> B[Cleaning]
B --> B1([Rinsed in deionized water])
B1 --> B2([Sonicated in cleaning solution])
B --> C[Smearing Adhesive]
C --> C1([Cr adhesive coating])
C --> D[Vapor Deposition System<br/>Kurt J. Lesker PVD 200]
D --> D1([PtAu sputter coating, gradient])
D --> E([Test Coupon])
```
## D2 - Characterization and Batch Assembly
Per-coupon simulation and thickness characterization, then assembly into wafers (x49 coupons) and batches (x3 wafers), deposited per the parameter table below.
```mermaid
flowchart LR
TC([Test Coupon]) --> S[SIMTRA simulation<br/>composition / atom energies]
S --> OP[Optical profilometry<br/>film thickness 0.3-1.1 um]
OP --> TW[Assemble Test Wafer<br/>x49 coupons]
TW --> TB[Assemble Test Batch<br/>x3 wafers]
TB --> B721[B721<br/>Pt 150 W / Au 50 W]
TB --> B722[B722<br/>Pt 100 W / Au 100 W]
TB --> B723[B723<br/>Pt 150 W / Au 50 W]
TB --> B724[B724<br/>Pt 50 W / Au 150 W]
```
Deposition matrix:
| Batch | Pt:Au gun tilt | Pt power W | Au power W | Pt discharge V | Au discharge V |
|---|---|---|---|---|---|
| B721 | 20:0 deg | 150 | 50 | 432 | 311 |
| B722 | 20:20 deg | 100 | 100 | 399 | 352 |
| B723 | 20:20 deg | 150 | 50 | 430 | 311 |
| B724 | 0:20 deg | 50 | 150 | 376 | 340 |
## D3 - Testing Tree
The four characterization/testing paths every coupon can take.
```mermaid
flowchart TD
TC([Test Coupon]) --> F[Friction<br/>RAPID 6-probe tribometer]
F --> FA([Lab Air -> COF dataset mu_normal])
F --> FN([Dry N2 -> COF dataset mu_dry_nit])
TC --> NI[Nanoindentation<br/>Bruker TI980]
NI --> NIR([Hardness, reduced modulus<br/>25 indents per coupon])
TC --> AF[AFM]
AF --> AFR([Topography -> Ra / Rq])
TC --> XR[micro-XRF<br/>Bruker M4 Tornado]
XR --> XRR([Pt/Au wt% map, 20x20 grid])
```
## D4 - Tribometer Session Sequence
Operator workflow for one RAPID tribometer session, from test request to teardown.
```mermaid
sequenceDiagram
actor Op as Operator
participant Prep as Sample Prep
participant Plan as Excel Test Plan
participant SW as Control Software
participant TR as RAPID Tribometer
Op->>Prep: Request test
Op->>Prep: Get samples
Op->>Prep: Load samples onto plates
Op->>Prep: Load ball holders (3 counterfaces per holder)
Op->>Plan: Create test plan<br/>(Sample Plate, Plate Location, Sample ID, Save Location,<br/>Folder Name, X/Y/Z Offset, Load, Iterations)
Op->>TR: Transfer plates to tribometer
Op->>TR: Set up platter
Op->>SW: Activate software
Op->>SW: Load plan
SW->>TR: Activate tribometer
TR-->>SW: Run reciprocating tests (execution loop, see D5)
Op->>SW: Software stop
SW-->>Op: Save avg files
Op->>TR: Remove plates and platters
Op->>SW: Close software
Op->>TR: Pull equipment out
```
## D5 - Execution Loop
Nested iteration executed by the tribometer within one run.
```mermaid
flowchart TD
L[Load all counterfaces] --> P{5 plates}
P -->|for each plate| PR{6 probes in parallel}
PR -->|for each probe| C{4 coupons on coupon square}
C -->|for each coupon| T[Draw track]
T --> RC[Rotate counterface - fresh ball per track]
RC -->|3 tracks per coupon| T
T --> DONE([120 coupons x 3 tracks = 360 tracks per run])
```
## D6 - Data Hierarchy
Entity containment and per-entity measurement datasets (ERD-style).
```mermaid
erDiagram
DEPOSITION_RUN ||--|| SIMTRA_PROFILE : produces
DEPOSITION_RUN ||--|| BATCH : creates
BATCH ||--|{ WAFER : contains
WAFER ||--|{ COUPON : contains
COUPON ||--|| XRF_MAP : has
COUPON ||--|| NANOINDENTATION : has
COUPON ||--|| AFM : has
COUPON ||--|| PROFILOMETRY : has
COUPON ||--o{ TRACK : "friction coupons only"
TRACK ||--|{ CYCLE : has
TRACK ||--|{ LOOP_POINT : has
TRACK ||--|| WEAR : has
```

View File

@@ -0,0 +1,6 @@
format,measured,vs CSV,archival (compressed),projected @600 GB,projected @1.2 TB,projected @6 TB
CSV,149 MB,1.00x,53 MB,600 GB,1.20 TB,6.00 TB
JSON-LD,311 MB,2.09x,57 MB,1.26 TB,2.51 TB,13 TB
SQLite,123 MB,0.83x,55 MB,498 GB,996 GB,4.98 TB
PostgreSQL,397 MB,2.67x,53 MB,1.60 TB,3.21 TB,16 TB
RDF triplestore,1.74 GB,11.70x,126 MB,7.02 TB,14 TB,70 TB
1 format measured vs CSV archival (compressed) projected @600 GB projected @1.2 TB projected @6 TB
2 CSV 149 MB 1.00x 53 MB 600 GB 1.20 TB 6.00 TB
3 JSON-LD 311 MB 2.09x 57 MB 1.26 TB 2.51 TB 13 TB
4 SQLite 123 MB 0.83x 55 MB 498 GB 996 GB 4.98 TB
5 PostgreSQL 397 MB 2.67x 53 MB 1.60 TB 3.21 TB 16 TB
6 RDF triplestore 1.74 GB 11.70x 126 MB 7.02 TB 14 TB 70 TB

View File

@@ -0,0 +1,36 @@
format,query,wall time,peak RSS,CPU util,read
CSV,q1,0.06 s,4 MB,0%,0.0 MB
CSV,q2,0.17 s,20 MB,16%,2.5 MB
CSV,q3,1.15 s,21 MB,24%,31.8 MB
CSV,q4,0.71 s,21 MB,22%,16.8 MB
CSV,q5,1.09 s,21 MB,23%,31.8 MB
CSV,q6,0.28 s,20 MB,17%,1.2 MB
CSV,q7,1.27 s,21 MB,23%,32.5 MB
JSON-LD,q1,0.11 s,22 MB,10%,1.4 MB
JSON-LD,q2,1.48 s,24 MB,24%,53.1 MB
JSON-LD,q3,9.94 s,24 MB,25%,297.6 MB
JSON-LD,q4,9.92 s,24 MB,25%,297.8 MB
JSON-LD,q5,9.96 s,24 MB,25%,297.6 MB
JSON-LD,q6,9.79 s,24 MB,24%,297.8 MB
JSON-LD,q7,10 s,24 MB,24%,297.9 MB
SQLite,q1,0.06 s,4 MB,0%,0.0 MB
SQLite,q2,0.06 s,4 MB,0%,0.0 MB
SQLite,q3,0.11 s,4 MB,0%,0.0 MB
SQLite,q4,0.06 s,4 MB,0%,0.0 MB
SQLite,q5,0.77 s,25 MB,23%,117.9 MB
SQLite,q6,0.06 s,4 MB,0%,0.0 MB
SQLite,q7,0.71 s,25 MB,23%,93.0 MB
PostgreSQL,q1,0.22 s,42 MB,18%,4.1 MB
PostgreSQL,q2,0.22 s,42 MB,20%,4.1 MB
PostgreSQL,q3,0.22 s,42 MB,18%,4.1 MB
PostgreSQL,q4,0.22 s,42 MB,18%,4.1 MB
PostgreSQL,q5,0.44 s,42 MB,10%,4.1 MB
PostgreSQL,q6,0.22 s,42 MB,19%,4.1 MB
PostgreSQL,q7,0.33 s,42 MB,12%,4.1 MB
RDF triplestore,q1,0.11 s,32 MB,17%,3.6 MB
RDF triplestore,q2,2.81 s,85 MB,24%,375.6 MB
RDF triplestore,q3,46 s,264 MB,25%,5445.1 MB
RDF triplestore,q4,20 s,168 MB,25%,2694.6 MB
RDF triplestore,q5,60 s,264 MB,25%,5390.5 MB
RDF triplestore,q6,0.17 s,46 MB,16%,15.9 MB
RDF triplestore,q7,41 s,264 MB,25%,5380.5 MB
1 format query wall time peak RSS CPU util read
2 CSV q1 0.06 s 4 MB 0% 0.0 MB
3 CSV q2 0.17 s 20 MB 16% 2.5 MB
4 CSV q3 1.15 s 21 MB 24% 31.8 MB
5 CSV q4 0.71 s 21 MB 22% 16.8 MB
6 CSV q5 1.09 s 21 MB 23% 31.8 MB
7 CSV q6 0.28 s 20 MB 17% 1.2 MB
8 CSV q7 1.27 s 21 MB 23% 32.5 MB
9 JSON-LD q1 0.11 s 22 MB 10% 1.4 MB
10 JSON-LD q2 1.48 s 24 MB 24% 53.1 MB
11 JSON-LD q3 9.94 s 24 MB 25% 297.6 MB
12 JSON-LD q4 9.92 s 24 MB 25% 297.8 MB
13 JSON-LD q5 9.96 s 24 MB 25% 297.6 MB
14 JSON-LD q6 9.79 s 24 MB 24% 297.8 MB
15 JSON-LD q7 10 s 24 MB 24% 297.9 MB
16 SQLite q1 0.06 s 4 MB 0% 0.0 MB
17 SQLite q2 0.06 s 4 MB 0% 0.0 MB
18 SQLite q3 0.11 s 4 MB 0% 0.0 MB
19 SQLite q4 0.06 s 4 MB 0% 0.0 MB
20 SQLite q5 0.77 s 25 MB 23% 117.9 MB
21 SQLite q6 0.06 s 4 MB 0% 0.0 MB
22 SQLite q7 0.71 s 25 MB 23% 93.0 MB
23 PostgreSQL q1 0.22 s 42 MB 18% 4.1 MB
24 PostgreSQL q2 0.22 s 42 MB 20% 4.1 MB
25 PostgreSQL q3 0.22 s 42 MB 18% 4.1 MB
26 PostgreSQL q4 0.22 s 42 MB 18% 4.1 MB
27 PostgreSQL q5 0.44 s 42 MB 10% 4.1 MB
28 PostgreSQL q6 0.22 s 42 MB 19% 4.1 MB
29 PostgreSQL q7 0.33 s 42 MB 12% 4.1 MB
30 RDF triplestore q1 0.11 s 32 MB 17% 3.6 MB
31 RDF triplestore q2 2.81 s 85 MB 24% 375.6 MB
32 RDF triplestore q3 46 s 264 MB 25% 5445.1 MB
33 RDF triplestore q4 20 s 168 MB 25% 2694.6 MB
34 RDF triplestore q5 60 s 264 MB 25% 5390.5 MB
35 RDF triplestore q6 0.17 s 46 MB 16% 15.9 MB
36 RDF triplestore q7 41 s 264 MB 25% 5380.5 MB

View File

@@ -0,0 +1,36 @@
format,query,projected wall time,flag
CSV,q1,0.06 s,OK
CSV,q2,7.1 min,OK
CSV,q3,73.5 min,IMPRACTICAL
CSV,q4,44.1 min,OK
CSV,q5,69.6 min,IMPRACTICAL
CSV,q6,14.5 min,OK
CSV,q7,81.6 min,IMPRACTICAL
JSON-LD,q1,0.11 s,OK
JSON-LD,q2,92.0 min,IMPRACTICAL
JSON-LD,q3,11.0 h,IMPRACTICAL
JSON-LD,q4,11.0 h,IMPRACTICAL
JSON-LD,q5,11.0 h,IMPRACTICAL
JSON-LD,q6,10.9 h,IMPRACTICAL
JSON-LD,q7,11.2 h,IMPRACTICAL
SQLite,q1,0.06 s,OK
SQLite,q2,1.27 s,OK
SQLite,q3,5.7 min,OK
SQLite,q4,0.06 s,OK
SQLite,q5,47.5 min,OK
SQLite,q6,5.18 s,OK
SQLite,q7,43.7 min,OK
PostgreSQL,q1,0.22 s,OK
PostgreSQL,q2,0.22 s,OK
PostgreSQL,q3,3.42 s,OK
PostgreSQL,q4,11 s,OK
PostgreSQL,q5,3.7 min,OK
PostgreSQL,q6,4.06 s,OK
PostgreSQL,q7,115 s,OK
RDF triplestore,q1,0.11 s,OK
RDF triplestore,q2,3.0 h,IMPRACTICAL
RDF triplestore,q3,2.1 d,FAIL
RDF triplestore,q4,22.1 h,IMPRACTICAL
RDF triplestore,q5,2.8 d,FAIL
RDF triplestore,q6,5.9 min,OK
RDF triplestore,q7,45.3 h,FAIL
1 format query projected wall time flag
2 CSV q1 0.06 s OK
3 CSV q2 7.1 min OK
4 CSV q3 73.5 min IMPRACTICAL
5 CSV q4 44.1 min OK
6 CSV q5 69.6 min IMPRACTICAL
7 CSV q6 14.5 min OK
8 CSV q7 81.6 min IMPRACTICAL
9 JSON-LD q1 0.11 s OK
10 JSON-LD q2 92.0 min IMPRACTICAL
11 JSON-LD q3 11.0 h IMPRACTICAL
12 JSON-LD q4 11.0 h IMPRACTICAL
13 JSON-LD q5 11.0 h IMPRACTICAL
14 JSON-LD q6 10.9 h IMPRACTICAL
15 JSON-LD q7 11.2 h IMPRACTICAL
16 SQLite q1 0.06 s OK
17 SQLite q2 1.27 s OK
18 SQLite q3 5.7 min OK
19 SQLite q4 0.06 s OK
20 SQLite q5 47.5 min OK
21 SQLite q6 5.18 s OK
22 SQLite q7 43.7 min OK
23 PostgreSQL q1 0.22 s OK
24 PostgreSQL q2 0.22 s OK
25 PostgreSQL q3 3.42 s OK
26 PostgreSQL q4 11 s OK
27 PostgreSQL q5 3.7 min OK
28 PostgreSQL q6 4.06 s OK
29 PostgreSQL q7 115 s OK
30 RDF triplestore q1 0.11 s OK
31 RDF triplestore q2 3.0 h IMPRACTICAL
32 RDF triplestore q3 2.1 d FAIL
33 RDF triplestore q4 22.1 h IMPRACTICAL
34 RDF triplestore q5 2.8 d FAIL
35 RDF triplestore q6 5.9 min OK
36 RDF triplestore q7 45.3 h FAIL

View File

@@ -0,0 +1,6 @@
format,configuration,chassis,RAM,disk,extra nodes,total
CSV,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,"$4,618","$1,887","$3,995",$0,"$10,500"
JSON-LD,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,"$4,618","$1,887","$3,995",$0,"$10,500"
SQLite,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,"$4,618","$1,887","$3,995",$0,"$10,500"
PostgreSQL,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,"$4,618","$1,887","$3,995",$0,"$10,500"
RDF triplestore,2 x 1U Supermicro AS-1015CS-TNR (16 cores); 22 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2,"$4,618","$41,514","$7,990","$4,618","$58,740"
1 format configuration chassis RAM disk extra nodes total
2 CSV 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 $4,618 $1,887 $3,995 $0 $10,500
3 JSON-LD 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 $4,618 $1,887 $3,995 $0 $10,500
4 SQLite 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 $4,618 $1,887 $3,995 $0 $10,500
5 PostgreSQL 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 $4,618 $1,887 $3,995 $0 $10,500
6 RDF triplestore 2 x 1U Supermicro AS-1015CS-TNR (16 cores); 22 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2 $4,618 $41,514 $7,990 $4,618 $58,740

View File

@@ -0,0 +1,11 @@
scale,format,search_pts,ram_pts,disk_pts,total
measured,CSV,15.5,20.0,8.3,43.8
measured,JSON-LD,1.7,17.2,4.0,22.9
measured,SQLite,50.0,16.9,10.0,76.9
measured,PostgreSQL,26.2,10.0,3.1,39.3
measured,RDF triplestore,1.3,1.6,0.7,3.6
projected_600gb,CSV,0.6,20.0,8.3,28.9
projected_600gb,JSON-LD,0.1,20.0,4.0,24.0
projected_600gb,SQLite,21.8,15.2,10.0,47.0
projected_600gb,PostgreSQL,50.0,5.5,3.1,58.6
projected_600gb,RDF triplestore,0.0,0.2,0.7,0.9
1 scale format search_pts ram_pts disk_pts total
2 measured CSV 15.5 20.0 8.3 43.8
3 measured JSON-LD 1.7 17.2 4.0 22.9
4 measured SQLite 50.0 16.9 10.0 76.9
5 measured PostgreSQL 26.2 10.0 3.1 39.3
6 measured RDF triplestore 1.3 1.6 0.7 3.6
7 projected_600gb CSV 0.6 20.0 8.3 28.9
8 projected_600gb JSON-LD 0.1 20.0 4.0 24.0
9 projected_600gb SQLite 21.8 15.2 10.0 47.0
10 projected_600gb PostgreSQL 50.0 5.5 3.1 58.6
11 projected_600gb RDF triplestore 0.0 0.2 0.7 0.9

View File

@@ -0,0 +1,106 @@
format,query,scale_gb,projected_wall_s,projected_wall,flag
CSV,q1,600,0.06,0.06 s,OK
CSV,q1,1200,0.06,0.06 s,OK
CSV,q1,6000,0.06,0.06 s,OK
CSV,q2,600,428.301,7.1 min,OK
CSV,q2,1200,856.543,14.3 min,OK
CSV,q2,6000,4282.474,71.4 min,IMPRACTICAL
CSV,q3,600,4408.0,73.5 min,IMPRACTICAL
CSV,q3,1200,8815.94,2.4 h,IMPRACTICAL
CSV,q3,6000,44079.46,12.2 h,IMPRACTICAL
CSV,q4,600,2643.775,44.1 min,OK
CSV,q4,1200,5287.489,88.1 min,IMPRACTICAL
CSV,q4,6000,26437.206,7.3 h,IMPRACTICAL
CSV,q5,600,4176.725,69.6 min,IMPRACTICAL
CSV,q5,1200,8353.391,2.3 h,IMPRACTICAL
CSV,q5,6000,41766.715,11.6 h,IMPRACTICAL
CSV,q6,600,870.266,14.5 min,OK
CSV,q6,1200,1740.472,29.0 min,OK
CSV,q6,6000,8702.119,2.4 h,IMPRACTICAL
CSV,q7,600,4897.995,81.6 min,IMPRACTICAL
CSV,q7,1200,9795.931,2.7 h,IMPRACTICAL
CSV,q7,6000,48979.414,13.6 h,IMPRACTICAL
JSON-LD,q1,600,0.111,0.11 s,OK
JSON-LD,q1,1200,0.111,0.11 s,OK
JSON-LD,q1,6000,0.111,0.11 s,OK
JSON-LD,q2,600,5519.622,92.0 min,IMPRACTICAL
JSON-LD,q2,1200,11039.133,3.1 h,IMPRACTICAL
JSON-LD,q2,6000,55195.222,15.3 h,IMPRACTICAL
JSON-LD,q3,600,39673.993,11.0 h,IMPRACTICAL
JSON-LD,q3,1200,79347.876,22.0 h,IMPRACTICAL
JSON-LD,q3,6000,396738.933,4.6 d,FAIL
JSON-LD,q4,600,39597.306,11.0 h,IMPRACTICAL
JSON-LD,q4,1200,79194.5,22.0 h,IMPRACTICAL
JSON-LD,q4,6000,395972.054,4.6 d,FAIL
JSON-LD,q5,600,39760.772,11.0 h,IMPRACTICAL
JSON-LD,q5,1200,79521.432,22.1 h,IMPRACTICAL
JSON-LD,q5,6000,397606.717,4.6 d,FAIL
JSON-LD,q6,600,39076.231,10.9 h,IMPRACTICAL
JSON-LD,q6,1200,78152.351,21.7 h,IMPRACTICAL
JSON-LD,q6,6000,390761.312,4.5 d,FAIL
JSON-LD,q7,600,40228.164,11.2 h,IMPRACTICAL
JSON-LD,q7,1200,80456.218,22.3 h,IMPRACTICAL
JSON-LD,q7,6000,402280.643,4.7 d,FAIL
SQLite,q1,600,0.062,0.06 s,OK
SQLite,q1,1200,0.063,0.06 s,OK
SQLite,q1,6000,0.063,0.06 s,OK
SQLite,q2,600,1.271,1.27 s,OK
SQLite,q2,1200,2.482,2.48 s,OK
SQLite,q2,6000,12.169,12 s,OK
SQLite,q3,600,339.238,5.7 min,OK
SQLite,q3,1200,699.328,11.7 min,OK
SQLite,q3,6000,3739.195,62.3 min,IMPRACTICAL
SQLite,q4,600,0.06,0.06 s,OK
SQLite,q4,1200,0.06,0.06 s,OK
SQLite,q4,6000,0.06,0.06 s,OK
SQLite,q5,600,2847.603,47.5 min,OK
SQLite,q5,1200,5695.146,94.9 min,IMPRACTICAL
SQLite,q5,6000,28475.49,7.9 h,IMPRACTICAL
SQLite,q6,600,5.18,5.18 s,OK
SQLite,q6,1200,10.615,11 s,OK
SQLite,q6,6000,56.5,56 s,OK
SQLite,q7,600,2619.558,43.7 min,OK
SQLite,q7,1200,5239.055,87.3 min,IMPRACTICAL
SQLite,q7,6000,26195.034,7.3 h,IMPRACTICAL
PostgreSQL,q1,600,0.222,0.22 s,OK
PostgreSQL,q1,1200,0.222,0.22 s,OK
PostgreSQL,q1,6000,0.222,0.22 s,OK
PostgreSQL,q2,600,0.221,0.22 s,OK
PostgreSQL,q2,1200,0.221,0.22 s,OK
PostgreSQL,q2,6000,0.221,0.22 s,OK
PostgreSQL,q3,600,3.42,3.42 s,OK
PostgreSQL,q3,1200,6.817,6.82 s,OK
PostgreSQL,q3,6000,35.495,35 s,OK
PostgreSQL,q4,600,10.715,11 s,OK
PostgreSQL,q4,1200,21.209,21 s,OK
PostgreSQL,q4,6000,105.162,105 s,OK
PostgreSQL,q5,600,222.111,3.7 min,OK
PostgreSQL,q5,1200,444.001,7.4 min,OK
PostgreSQL,q5,6000,2219.125,37.0 min,OK
PostgreSQL,q6,600,4.06,4.06 s,OK
PostgreSQL,q6,1200,8.137,8.14 s,OK
PostgreSQL,q6,6000,42.55,43 s,OK
PostgreSQL,q7,600,114.849,115 s,OK
PostgreSQL,q7,1200,229.477,3.8 min,OK
PostgreSQL,q7,6000,1146.503,19.1 min,OK
RDF triplestore,q1,600,0.113,0.11 s,OK
RDF triplestore,q1,1200,0.113,0.11 s,OK
RDF triplestore,q1,6000,0.113,0.11 s,OK
RDF triplestore,q2,600,10866.789,3.0 h,IMPRACTICAL
RDF triplestore,q2,1200,21733.465,6.0 h,IMPRACTICAL
RDF triplestore,q2,6000,108666.874,30.2 h,FAIL
RDF triplestore,q3,600,185722.881,2.1 d,FAIL
RDF triplestore,q3,1200,371445.649,4.3 d,FAIL
RDF triplestore,q3,6000,1857227.792,21.5 d,FAIL
RDF triplestore,q4,600,79559.375,22.1 h,IMPRACTICAL
RDF triplestore,q4,1200,159118.636,44.2 h,FAIL
RDF triplestore,q4,6000,795592.727,9.2 d,FAIL
RDF triplestore,q5,600,240054.647,2.8 d,FAIL
RDF triplestore,q5,1200,480109.181,5.6 d,FAIL
RDF triplestore,q5,6000,2400545.453,27.8 d,FAIL
RDF triplestore,q6,600,352.729,5.9 min,OK
RDF triplestore,q6,1200,727.088,12.1 min,OK
RDF triplestore,q6,6000,3887.402,64.8 min,IMPRACTICAL
RDF triplestore,q7,600,163225.475,45.3 h,FAIL
RDF triplestore,q7,1200,326450.837,3.8 d,FAIL
RDF triplestore,q7,6000,1632253.733,18.9 d,FAIL
1 format query scale_gb projected_wall_s projected_wall flag
2 CSV q1 600 0.06 0.06 s OK
3 CSV q1 1200 0.06 0.06 s OK
4 CSV q1 6000 0.06 0.06 s OK
5 CSV q2 600 428.301 7.1 min OK
6 CSV q2 1200 856.543 14.3 min OK
7 CSV q2 6000 4282.474 71.4 min IMPRACTICAL
8 CSV q3 600 4408.0 73.5 min IMPRACTICAL
9 CSV q3 1200 8815.94 2.4 h IMPRACTICAL
10 CSV q3 6000 44079.46 12.2 h IMPRACTICAL
11 CSV q4 600 2643.775 44.1 min OK
12 CSV q4 1200 5287.489 88.1 min IMPRACTICAL
13 CSV q4 6000 26437.206 7.3 h IMPRACTICAL
14 CSV q5 600 4176.725 69.6 min IMPRACTICAL
15 CSV q5 1200 8353.391 2.3 h IMPRACTICAL
16 CSV q5 6000 41766.715 11.6 h IMPRACTICAL
17 CSV q6 600 870.266 14.5 min OK
18 CSV q6 1200 1740.472 29.0 min OK
19 CSV q6 6000 8702.119 2.4 h IMPRACTICAL
20 CSV q7 600 4897.995 81.6 min IMPRACTICAL
21 CSV q7 1200 9795.931 2.7 h IMPRACTICAL
22 CSV q7 6000 48979.414 13.6 h IMPRACTICAL
23 JSON-LD q1 600 0.111 0.11 s OK
24 JSON-LD q1 1200 0.111 0.11 s OK
25 JSON-LD q1 6000 0.111 0.11 s OK
26 JSON-LD q2 600 5519.622 92.0 min IMPRACTICAL
27 JSON-LD q2 1200 11039.133 3.1 h IMPRACTICAL
28 JSON-LD q2 6000 55195.222 15.3 h IMPRACTICAL
29 JSON-LD q3 600 39673.993 11.0 h IMPRACTICAL
30 JSON-LD q3 1200 79347.876 22.0 h IMPRACTICAL
31 JSON-LD q3 6000 396738.933 4.6 d FAIL
32 JSON-LD q4 600 39597.306 11.0 h IMPRACTICAL
33 JSON-LD q4 1200 79194.5 22.0 h IMPRACTICAL
34 JSON-LD q4 6000 395972.054 4.6 d FAIL
35 JSON-LD q5 600 39760.772 11.0 h IMPRACTICAL
36 JSON-LD q5 1200 79521.432 22.1 h IMPRACTICAL
37 JSON-LD q5 6000 397606.717 4.6 d FAIL
38 JSON-LD q6 600 39076.231 10.9 h IMPRACTICAL
39 JSON-LD q6 1200 78152.351 21.7 h IMPRACTICAL
40 JSON-LD q6 6000 390761.312 4.5 d FAIL
41 JSON-LD q7 600 40228.164 11.2 h IMPRACTICAL
42 JSON-LD q7 1200 80456.218 22.3 h IMPRACTICAL
43 JSON-LD q7 6000 402280.643 4.7 d FAIL
44 SQLite q1 600 0.062 0.06 s OK
45 SQLite q1 1200 0.063 0.06 s OK
46 SQLite q1 6000 0.063 0.06 s OK
47 SQLite q2 600 1.271 1.27 s OK
48 SQLite q2 1200 2.482 2.48 s OK
49 SQLite q2 6000 12.169 12 s OK
50 SQLite q3 600 339.238 5.7 min OK
51 SQLite q3 1200 699.328 11.7 min OK
52 SQLite q3 6000 3739.195 62.3 min IMPRACTICAL
53 SQLite q4 600 0.06 0.06 s OK
54 SQLite q4 1200 0.06 0.06 s OK
55 SQLite q4 6000 0.06 0.06 s OK
56 SQLite q5 600 2847.603 47.5 min OK
57 SQLite q5 1200 5695.146 94.9 min IMPRACTICAL
58 SQLite q5 6000 28475.49 7.9 h IMPRACTICAL
59 SQLite q6 600 5.18 5.18 s OK
60 SQLite q6 1200 10.615 11 s OK
61 SQLite q6 6000 56.5 56 s OK
62 SQLite q7 600 2619.558 43.7 min OK
63 SQLite q7 1200 5239.055 87.3 min IMPRACTICAL
64 SQLite q7 6000 26195.034 7.3 h IMPRACTICAL
65 PostgreSQL q1 600 0.222 0.22 s OK
66 PostgreSQL q1 1200 0.222 0.22 s OK
67 PostgreSQL q1 6000 0.222 0.22 s OK
68 PostgreSQL q2 600 0.221 0.22 s OK
69 PostgreSQL q2 1200 0.221 0.22 s OK
70 PostgreSQL q2 6000 0.221 0.22 s OK
71 PostgreSQL q3 600 3.42 3.42 s OK
72 PostgreSQL q3 1200 6.817 6.82 s OK
73 PostgreSQL q3 6000 35.495 35 s OK
74 PostgreSQL q4 600 10.715 11 s OK
75 PostgreSQL q4 1200 21.209 21 s OK
76 PostgreSQL q4 6000 105.162 105 s OK
77 PostgreSQL q5 600 222.111 3.7 min OK
78 PostgreSQL q5 1200 444.001 7.4 min OK
79 PostgreSQL q5 6000 2219.125 37.0 min OK
80 PostgreSQL q6 600 4.06 4.06 s OK
81 PostgreSQL q6 1200 8.137 8.14 s OK
82 PostgreSQL q6 6000 42.55 43 s OK
83 PostgreSQL q7 600 114.849 115 s OK
84 PostgreSQL q7 1200 229.477 3.8 min OK
85 PostgreSQL q7 6000 1146.503 19.1 min OK
86 RDF triplestore q1 600 0.113 0.11 s OK
87 RDF triplestore q1 1200 0.113 0.11 s OK
88 RDF triplestore q1 6000 0.113 0.11 s OK
89 RDF triplestore q2 600 10866.789 3.0 h IMPRACTICAL
90 RDF triplestore q2 1200 21733.465 6.0 h IMPRACTICAL
91 RDF triplestore q2 6000 108666.874 30.2 h FAIL
92 RDF triplestore q3 600 185722.881 2.1 d FAIL
93 RDF triplestore q3 1200 371445.649 4.3 d FAIL
94 RDF triplestore q3 6000 1857227.792 21.5 d FAIL
95 RDF triplestore q4 600 79559.375 22.1 h IMPRACTICAL
96 RDF triplestore q4 1200 159118.636 44.2 h FAIL
97 RDF triplestore q4 6000 795592.727 9.2 d FAIL
98 RDF triplestore q5 600 240054.647 2.8 d FAIL
99 RDF triplestore q5 1200 480109.181 5.6 d FAIL
100 RDF triplestore q5 6000 2400545.453 27.8 d FAIL
101 RDF triplestore q6 600 352.729 5.9 min OK
102 RDF triplestore q6 1200 727.088 12.1 min OK
103 RDF triplestore q6 6000 3887.402 64.8 min IMPRACTICAL
104 RDF triplestore q7 600 163225.475 45.3 h FAIL
105 RDF triplestore q7 1200 326450.837 3.8 d FAIL
106 RDF triplestore q7 6000 1632253.733 18.9 d FAIL

View File

@@ -0,0 +1,16 @@
format,scale_gb,configuration,disk_tb,ram_gb,cores,nodes,ram_modules,nvme_drives,base_usd,ram_usd,disk_usd,extra_nodes_usd,est_cost_usd
CSV,600,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,0.78,16.0,16,1,1,1,4618,1887,3995,0,10500
CSV,1200,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,1.56,16.0,16,1,1,1,4618,1887,3995,0,10500
CSV,6000,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2,7.8,16.0,16,1,1,2,4618,1887,7990,0,14495
JSON-LD,600,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,1.633,16.0,16,1,1,1,4618,1887,3995,0,10500
JSON-LD,1200,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,3.267,16.0,16,1,1,1,4618,1887,3995,0,10500
JSON-LD,6000,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 3 x 7.68 TB NVMe U.2,16.334,16.0,16,1,1,3,4618,1887,11985,0,18490
SQLite,600,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,0.647,21.0,16,1,1,1,4618,1887,3995,0,10500
SQLite,1200,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,1.295,38.0,16,1,1,1,4618,1887,3995,0,10500
SQLite,6000,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 3 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,6.475,173.9,16,1,3,1,4618,5661,3995,0,14274
PostgreSQL,600,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,2.086,58.7,16,1,1,1,4618,1887,3995,0,10500
PostgreSQL,1200,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 2 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2,4.171,113.5,16,1,2,1,4618,3774,3995,0,12387
PostgreSQL,6000,1 x 1U Supermicro AS-1015CS-TNR (16 cores); 9 x 64 GB DDR5 ECC RDIMM; 3 x 7.68 TB NVMe U.2,20.855,551.3,16,1,9,3,4618,16983,11985,0,33586
RDF triplestore,600,2 x 1U Supermicro AS-1015CS-TNR (16 cores); 22 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2,9.127,1404.1,16,2,22,2,4618,41514,7990,4618,58740
RDF triplestore,1200,3 x 1U Supermicro AS-1015CS-TNR (16 cores); 44 x 64 GB DDR5 ECC RDIMM; 3 x 7.68 TB NVMe U.2,18.254,2808.3,16,3,44,3,4618,83028,11985,9236,108867
RDF triplestore,6000,14 x 1U Supermicro AS-1015CS-TNR (16 cores); 220 x 64 GB DDR5 ECC RDIMM; 12 x 7.68 TB NVMe U.2,91.268,14041.3,16,14,220,12,4618,415140,47940,60034,527732
1 format scale_gb configuration disk_tb ram_gb cores nodes ram_modules nvme_drives base_usd ram_usd disk_usd extra_nodes_usd est_cost_usd
2 CSV 600 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 0.78 16.0 16 1 1 1 4618 1887 3995 0 10500
3 CSV 1200 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 1.56 16.0 16 1 1 1 4618 1887 3995 0 10500
4 CSV 6000 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2 7.8 16.0 16 1 1 2 4618 1887 7990 0 14495
5 JSON-LD 600 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 1.633 16.0 16 1 1 1 4618 1887 3995 0 10500
6 JSON-LD 1200 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 3.267 16.0 16 1 1 1 4618 1887 3995 0 10500
7 JSON-LD 6000 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 3 x 7.68 TB NVMe U.2 16.334 16.0 16 1 1 3 4618 1887 11985 0 18490
8 SQLite 600 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 0.647 21.0 16 1 1 1 4618 1887 3995 0 10500
9 SQLite 1200 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 1.295 38.0 16 1 1 1 4618 1887 3995 0 10500
10 SQLite 6000 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 3 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 6.475 173.9 16 1 3 1 4618 5661 3995 0 14274
11 PostgreSQL 600 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 1 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 2.086 58.7 16 1 1 1 4618 1887 3995 0 10500
12 PostgreSQL 1200 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 2 x 64 GB DDR5 ECC RDIMM; 1 x 7.68 TB NVMe U.2 4.171 113.5 16 1 2 1 4618 3774 3995 0 12387
13 PostgreSQL 6000 1 x 1U Supermicro AS-1015CS-TNR (16 cores); 9 x 64 GB DDR5 ECC RDIMM; 3 x 7.68 TB NVMe U.2 20.855 551.3 16 1 9 3 4618 16983 11985 0 33586
14 RDF triplestore 600 2 x 1U Supermicro AS-1015CS-TNR (16 cores); 22 x 64 GB DDR5 ECC RDIMM; 2 x 7.68 TB NVMe U.2 9.127 1404.1 16 2 22 2 4618 41514 7990 4618 58740
15 RDF triplestore 1200 3 x 1U Supermicro AS-1015CS-TNR (16 cores); 44 x 64 GB DDR5 ECC RDIMM; 3 x 7.68 TB NVMe U.2 18.254 2808.3 16 3 44 3 4618 83028 11985 9236 108867
16 RDF triplestore 6000 14 x 1U Supermicro AS-1015CS-TNR (16 cores); 220 x 64 GB DDR5 ECC RDIMM; 12 x 7.68 TB NVMe U.2 91.268 14041.3 16 14 220 12 4618 415140 47940 60034 527732

703
report.py Normal file
View File

@@ -0,0 +1,703 @@
"""Task 11 - Reporting: tables, charts, weighted scoring, REPORT.md.
Aggregates the measured matrix (task 09), the projections (task 10) and the
process diagrams (task 02) into the final decision document. Writes 14 PNG
charts (150 dpi) under ./out/report/charts/ with the exact names fixed in
README.md, seven CSV tables under ./out/report/tables/ (t6/t7 carry every
projection and hardware row for all three scales - they are the committed
proof artifacts the executive summary links to), and ./out/report/REPORT.md.
Chart readability contract (owner feedback 2026-07-11): linear scales with
actual units everywhere (panels instead of log axes when magnitudes clash),
no scientific tick notation, full format names, a value label on every bar
and point, titles that state the takeaway. Hardware costs are a discrete
bill of materials (whole RDIMM modules, whole NVMe drives, priced chassis)
with street-price sources and an as-of date from out/config/hw_prices.yaml.
Weighted scoring: search speed x5 (0-50), RAM economy x2 (0-20), disk
economy x1 (0-10), max 80; each subscore is 10 x best/value; a format with
a FAIL projection at 600 GB scores 0 on search there (spec 11). Projected
numbers are always labeled as projected (bench-methodology section 7).
Spec: docs/specs/11_reporting.md.
"""
from __future__ import annotations
import csv
import logging
import math
import sys
from pathlib import Path
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
import yaml
from matplotlib.ticker import FuncFormatter
from common.pipeline import (
check_dependencies,
parse_task_args,
remove_stale_marker,
write_marker,
)
logger = logging.getLogger(__name__)
TASK_ID = "11"
DEPENDS_ON = ["02", "09", "10"]
DPI = 150 # chart resolution, spec 11
FORMATS = ("csv", "json", "sqlite", "pg", "rdf")
QUERY_NUMBERS = tuple(range(1, 8))
SCALES_GB = (600, 1200, 6000)
REPORT_SCALE_GB = 600 # the headline projection scale, spec 11
SCALE_LABELS = {600: "600 GB", 1200: "1.2 TB", 6000: "6 TB"}
WEIGHT_SEARCH = 5 # x5 -> 0-50, spec 11
WEIGHT_RAM = 2 # x2 -> 0-20
WEIGHT_DISK = 1 # x1 -> 0-10
ARCHIVAL_VARIANT = {
"csv": ("csv", "targz"), "json": ("json", "targz"), "sqlite": ("sqlite", "db_gz"),
"pg": ("pg", "dump"), "rdf": ("rdf", "nt_gz"),
}
WORKING_VARIANT = {
"csv": ("csv", "tree"), "json": ("json", "full"), "sqlite": ("sqlite", "db"),
"pg": ("pg", "live"), "rdf": ("rdf", "store"),
}
# query classes for the degradation charts (C4), spec 11
QUERY_CLASSES = {
"point_read": (1,),
"indexed": (2, 3, 4, 6),
"full_scan": (5, 7),
}
CLASS_LABELS = {"point_read": "point read", "indexed": "indexed / join", "full_scan": "full scan"}
QUERY_LABELS = {
1: "single-track COF curve", 2: "coupons at Au 10+/-0.5 wt%",
3: "hardness vs steady-state COF", 4: "anomaly filter",
5: "run-in over raw cycles", 6: "wear volume vs load", 7: "Stribeck aggregation",
}
DIAGRAM_TITLES = {
"D1": "Coupon assembly", "D2": "Characterization and batch assembly",
"D3": "Testing tree", "D4": "Tribometer session sequence",
"D5": "Execution loop", "D6": "Data hierarchy",
}
FMT_COLORS = {"csv": "tab:gray", "json": "tab:orange", "sqlite": "tab:green", "pg": "tab:blue", "rdf": "tab:red"}
FMT_NAMES = {"csv": "CSV", "json": "JSON-LD", "sqlite": "SQLite", "pg": "PostgreSQL", "rdf": "RDF triplestore"}
CHASSIS = {16: "Supermicro AS-1015CS-TNR", 32: "Supermicro AS-1115CS-TNR"}
def read_medians(path: Path) -> dict[tuple[str, int], dict]:
cells: dict[tuple[str, int], dict] = {}
with open(path, encoding="ascii", newline="") as fh:
for row in csv.DictReader(fh):
if row["status"] != "OK" or row["checksum_ok"] != "true":
raise ValueError(f"{path}: cell {row['format']}/{row['query']} is not a valid measurement")
cells[(row["format"], int(row["query"][1:]))] = {
"wall_s": float(row["wall_s_median"]),
"peak_rss_mb": float(row["peak_rss_mb_median"]),
"cpu_util_pct": float(row["cpu_util_pct_median"]),
"read_mb": float(row["read_mb_median"]),
}
if len(cells) != len(FORMATS) * len(QUERY_NUMBERS):
raise ValueError(f"{path}: {len(cells)} cells, expected the full matrix")
return cells
def read_projections(path: Path) -> dict[tuple[str, int, int], tuple[float, str]]:
proj: dict[tuple[str, int, int], tuple[float, str]] = {}
with open(path, encoding="ascii", newline="") as fh:
for row in csv.DictReader(fh):
proj[(row["format"], int(row["query"][1:]), int(row["scale_gb"]))] = (
float(row["projected_wall_s"]), row["flag"])
if len(proj) != len(FORMATS) * len(QUERY_NUMBERS) * len(SCALES_GB):
raise ValueError(f"{path}: incomplete projection table")
return proj
def read_sizes(path: Path) -> dict[tuple[str, str], int]:
with open(path, encoding="ascii", newline="") as fh:
reader = csv.reader(fh)
next(reader)
return {(fmt, variant): int(nbytes) for fmt, variant, nbytes in reader}
def read_sizing(path: Path) -> dict[tuple[str, int], dict]:
sizing: dict[tuple[str, int], dict] = {}
with open(path, encoding="ascii", newline="") as fh:
for row in csv.DictReader(fh):
sizing[(row["format"], int(row["scale_gb"]))] = {
"disk_tb": float(row["disk_tb"]), "ram_gb": float(row["ram_gb"]),
"cores": int(row["cores"]), "nodes": int(row["nodes"]),
"est_cost_usd": float(row["est_cost_usd"]),
"ram_modules": int(row["ram_modules"]), "nvme_drives": int(row["nvme_drives"]),
"base_usd": float(row["base_usd"]), "ram_usd": float(row["ram_usd"]),
"disk_usd": float(row["disk_usd"]), "extra_nodes_usd": float(row["extra_nodes_usd"]),
}
return sizing
def read_hw_prices(path: Path) -> dict:
with open(path, encoding="ascii") as fh:
return yaml.safe_load(fh)
def geometric_mean(values: list[float]) -> float:
return math.exp(sum(math.log(v) for v in values) / len(values))
def md_table(headers: list[str], rows: list[list]) -> str:
lines = ["| " + " | ".join(headers) + " |", "|" + "---|" * len(headers)]
lines += ["| " + " | ".join(str(c) for c in row) + " |" for row in rows]
return "\n".join(lines)
def human_time(seconds: float) -> str:
if seconds < 120:
return f"{seconds:.2f} s" if seconds < 10 else f"{seconds:.0f} s"
if seconds < 2 * 3600:
return f"{seconds / 60:.1f} min"
if seconds < 2 * 86400:
return f"{seconds / 3600:.1f} h"
return f"{seconds / 86400:.1f} d"
def human_size(nbytes: float) -> str:
for unit, div in (("TB", 1e12), ("GB", 1e9), ("MB", 1e6)):
if nbytes >= div:
return f"{nbytes / div:.2f} {unit}" if nbytes < 10 * div else f"{nbytes / div:.0f} {unit}"
return f"{nbytes / 1e3:.0f} KB"
def money(value: float) -> str:
return f"${value:,.0f}"
def write_csv(path: Path, headers: list[str], rows: list[list]) -> None:
with open(path, "w", encoding="ascii", newline="") as fh:
writer = csv.writer(fh)
writer.writerow(headers)
writer.writerows(rows)
# --- weighted scoring -----------------------------------------------------
def scores_for(metrics: dict[str, dict[str, float]], failed_search: set[str]) -> dict[str, dict]:
"""10 x best/value per metric; search zeroed for formats in failed_search."""
best = {key: min(m[key] for m in metrics.values()) for key in ("search", "ram", "disk")}
out: dict[str, dict] = {}
for fmt, m in metrics.items():
search = 0.0 if fmt in failed_search else 10.0 * best["search"] / m["search"]
ram = 10.0 * best["ram"] / m["ram"]
disk = 10.0 * best["disk"] / m["disk"]
out[fmt] = {
"search_pts": round(search * WEIGHT_SEARCH, 1),
"ram_pts": round(ram * WEIGHT_RAM, 1),
"disk_pts": round(disk * WEIGHT_DISK, 1),
"total": round(search * WEIGHT_SEARCH + ram * WEIGHT_RAM + disk * WEIGHT_DISK, 1),
}
return out
def build_scoreboards(medians, proj, sizes, sizing) -> dict[str, dict[str, dict]]:
measured_metrics = {
fmt: {
"search": geometric_mean([medians[(fmt, qn)]["wall_s"] for qn in QUERY_NUMBERS]),
"ram": max(medians[(fmt, qn)]["peak_rss_mb"] for qn in QUERY_NUMBERS),
"disk": float(sizes[WORKING_VARIANT[fmt]]),
}
for fmt in FORMATS
}
projected_metrics = {
fmt: {
"search": geometric_mean([proj[(fmt, qn, REPORT_SCALE_GB)][0] for qn in QUERY_NUMBERS]),
"ram": sizing[(fmt, REPORT_SCALE_GB)]["ram_gb"],
"disk": sizing[(fmt, REPORT_SCALE_GB)]["disk_tb"],
}
for fmt in FORMATS
}
failed = {fmt for fmt in FORMATS if any(proj[(fmt, qn, REPORT_SCALE_GB)][1] == "FAIL" for qn in QUERY_NUMBERS)}
return {
"measured": scores_for(measured_metrics, set()),
"projected_600gb": scores_for(projected_metrics, failed),
}
# --- charts ---------------------------------------------------------------
# Shared conventions: linear axes, real units, a value label on every bar
# and point, full format names, no scientific notation.
def _vbar_panel(ax, values: dict[str, float], labels: dict[str, str], unit: str, title: str) -> None:
bars = ax.bar(range(len(FORMATS)), [values[f] for f in FORMATS], color=[FMT_COLORS[f] for f in FORMATS])
for bar, fmt in zip(bars, FORMATS):
ax.annotate(labels[fmt], (bar.get_x() + bar.get_width() / 2, bar.get_height()),
ha="center", va="bottom", fontsize=8)
ax.set_xticks(range(len(FORMATS)))
ax.set_xticklabels([FMT_NAMES[f] for f in FORMATS], rotation=12, fontsize=8)
ax.set_ylabel(unit)
ax.set_title(title, fontsize=10)
ax.margins(y=0.18)
ax.grid(axis="y", alpha=0.25)
def _hbar_panel(ax, values: dict[str, float], labels: dict[str, str], title: str, names: bool) -> None:
y = range(len(FORMATS))
bars = ax.barh(list(y), [values[f] for f in FORMATS], color=[FMT_COLORS[f] for f in FORMATS])
for bar, fmt in zip(bars, FORMATS):
ax.annotate(" " + labels[fmt], (bar.get_width(), bar.get_y() + bar.get_height() / 2),
va="center", fontsize=8)
ax.set_yticks(list(y))
ax.set_yticklabels([FMT_NAMES[f] for f in FORMATS] if names else [""] * len(FORMATS), fontsize=8)
ax.invert_yaxis()
ax.set_xticks([]) # linear axis; the value labels on the bars carry the numbers
ax.margins(x=0.30)
ax.set_title(title, fontsize=10)
def chart_c1(charts_dir, sizes, proj_bytes, coeff) -> None:
fig, axes = plt.subplots(1, 2, figsize=(11, 4.5))
measured = {f: sizes[WORKING_VARIANT[f]] for f in FORMATS}
_vbar_panel(axes[0], {f: measured[f] / 1e9 for f in FORMATS},
{f: human_size(measured[f]) for f in FORMATS}, "GB (linear)", "measured on the 150 MB corpus")
_vbar_panel(axes[1], {f: proj_bytes[f] / 1e9 for f in FORMATS},
{f: human_size(proj_bytes[f]) for f in FORMATS}, "GB (linear)", "projected at 600 GB of raw data")
fig.suptitle(f"C1 - disk footprint: RDF stores the same data in {coeff['rdf']:.0f}x the raw CSV size, SQLite in {coeff['sqlite']:.2f}x")
fig.tight_layout()
fig.savefig(charts_dir / "c1_disk_footprint.png", dpi=DPI)
plt.close(fig)
def chart_c2(charts_dir, medians, sizing) -> None:
fig, axes = plt.subplots(1, 2, figsize=(11, 4.5))
rss = {f: max(medians[(f, qn)]["peak_rss_mb"] for qn in QUERY_NUMBERS) for f in FORMATS}
ram = {f: sizing[(f, REPORT_SCALE_GB)]["ram_gb"] for f in FORMATS}
_vbar_panel(axes[0], rss, {f: f"{rss[f]:.0f} MB" for f in FORMATS},
"MB (linear)", "measured client peak RSS on 150 MB")
_vbar_panel(axes[1], ram, {f: f"{ram[f]:.0f} GB" for f in FORMATS},
"GB (linear)", "recommended server RAM at 600 GB (projected)")
axes[1].axhline(1024, color="gray", linestyle="--", linewidth=1)
axes[1].annotate("single-node ceiling: 1 TB RAM", (0, 1024), fontsize=8, va="bottom")
fig.suptitle("C2 - RAM: only the RDF hot set outgrows a single 1 TB node at 600 GB")
fig.tight_layout()
fig.savefig(charts_dir / "c2_ram_requirement.png", dpi=DPI)
plt.close(fig)
def chart_c3(charts_dir, medians, proj) -> None:
for qn in QUERY_NUMBERS:
fig, axes = plt.subplots(1, 2, figsize=(11, 3.6))
measured = {f: medians[(f, qn)]["wall_s"] for f in FORMATS}
projected = {f: proj[(f, qn, REPORT_SCALE_GB)][0] for f in FORMATS}
_hbar_panel(axes[0], measured, {f: human_time(measured[f]) for f in FORMATS},
"measured on 150 MB", names=True)
_hbar_panel(axes[1], projected, {f: human_time(projected[f]) for f in FORMATS},
"projected at 600 GB", names=False)
fastest = min(FORMATS, key=lambda f: projected[f])
slowest = max(FORMATS, key=lambda f: projected[f])
fig.suptitle(f"C3 - Q{qn} ({QUERY_LABELS[qn]}) at 600 GB: {FMT_NAMES[fastest]} {human_time(projected[fastest])}"
f" vs {FMT_NAMES[slowest]} {human_time(projected[slowest])}")
fig.tight_layout()
fig.savefig(charts_dir / f"c3_q{qn}_wall_time.png", dpi=DPI)
plt.close(fig)
def chart_c4(charts_dir, medians, proj, measured_label: str) -> None:
for cls, queries in QUERY_CLASSES.items():
fig, axes = plt.subplots(1, 1 + len(SCALES_GB), figsize=(13, 3.6))
measured = {f: geometric_mean([medians[(f, qn)]["wall_s"] for qn in queries]) for f in FORMATS}
_hbar_panel(axes[0], measured, {f: human_time(measured[f]) for f in FORMATS}, measured_label, names=True)
for ax, scale in zip(axes[1:], SCALES_GB):
values = {f: geometric_mean([proj[(f, qn, scale)][0] for qn in queries]) for f in FORMATS}
_hbar_panel(ax, values, {f: human_time(values[f]) for f in FORMATS},
f"projected {SCALE_LABELS[scale]}", names=False)
queries_label = ", ".join(f"Q{qn}" for qn in queries)
fig.suptitle(f"C4 - {CLASS_LABELS[cls]} queries ({queries_label}): wall time growth by data size"
" (every panel after the first is projected; linear scale within each panel)")
fig.tight_layout()
fig.savefig(charts_dir / f"c4_degradation_{cls}.png", dpi=DPI)
plt.close(fig)
def chart_c5(charts_dir, scoreboards, winner: str) -> None:
fig, axes = plt.subplots(1, 2, figsize=(11, 4.5), sharey=True)
for ax, (title, board) in zip(axes, [
("measured scale (150 MB)", scoreboards["measured"]),
("projected at 600 GB", scoreboards["projected_600gb"]),
]):
x = range(len(FORMATS))
search = [board[fmt]["search_pts"] for fmt in FORMATS]
ram = [board[fmt]["ram_pts"] for fmt in FORMATS]
disk = [board[fmt]["disk_pts"] for fmt in FORMATS]
ax.bar(x, search, label=f"search speed x{WEIGHT_SEARCH} (0-50)")
ax.bar(x, ram, bottom=search, label=f"RAM economy x{WEIGHT_RAM} (0-20)")
ax.bar(x, disk, bottom=[s + r for s, r in zip(search, ram)], label=f"disk economy x{WEIGHT_DISK} (0-10)")
for i, fmt in enumerate(FORMATS):
ax.text(i, board[fmt]["total"] + 1.5, f"{board[fmt]['total']:.0f}", ha="center", fontsize=9)
ax.set_xticks(list(x))
ax.set_xticklabels([FMT_NAMES[f] for f in FORMATS], rotation=12, fontsize=8)
ax.set_ylim(0, 85)
ax.set_title(title, fontsize=10)
ax.grid(axis="y", alpha=0.25)
axes[0].set_ylabel("weighted score (linear, max 80)")
axes[0].legend(loc="upper left", fontsize=8)
board600 = scoreboards["projected_600gb"]
fig.suptitle(f"C5 - weighted score (max 80): {FMT_NAMES[winner]} leads the 600 GB scoreboard with {board600[winner]['total']:.0f} points")
fig.tight_layout()
fig.savefig(charts_dir / "c5_weighted_score.png", dpi=DPI)
plt.close(fig)
def chart_c6(charts_dir, proj, sizing) -> None:
fig, ax = plt.subplots(figsize=(9, 5.5))
offsets = {"csv": (12, 0), "json": (12, -4), "sqlite": (12, 10), "pg": (12, -16), "rdf": (-40, -20)}
geo = {f: geometric_mean([proj[(f, qn, REPORT_SCALE_GB)][0] for qn in QUERY_NUMBERS]) for f in FORMATS}
cost = {f: sizing[(f, REPORT_SCALE_GB)]["est_cost_usd"] for f in FORMATS}
for fmt in FORMATS:
ax.scatter(cost[fmt], geo[fmt] / 60, s=70, color=FMT_COLORS[fmt])
ax.annotate(f"{FMT_NAMES[fmt]}: {money(cost[fmt])}, {human_time(geo[fmt])}",
(cost[fmt], geo[fmt] / 60), textcoords="offset points", xytext=offsets[fmt], fontsize=8)
ax.xaxis.set_major_formatter(FuncFormatter(lambda v, _: f"${v / 1000:.0f}k"))
ax.yaxis.set_major_formatter(FuncFormatter(lambda v, _: f"{v:.0f} min"))
ax.set_xlabel("estimated hardware cost at 600 GB (bill of materials, linear)")
ax.set_ylabel("typical query time at 600 GB, projected (geometric mean Q1-Q7, linear)")
ax.annotate("better: lower-left", xy=(0.05, 0.04), xytext=(0.42, 0.10),
xycoords="axes fraction", textcoords="axes fraction",
fontsize=10, color="darkgreen", fontweight="bold",
arrowprops={"arrowstyle": "->", "color": "darkgreen"})
ax.margins(x=0.18, y=0.12)
ax.grid(alpha=0.25)
fastest = min(FORMATS, key=lambda f: geo[f])
ax.set_title(f"C6 - cost vs speed at 600 GB (projected): {FMT_NAMES[fastest]} is the fastest at {money(cost[fastest])}")
fig.tight_layout()
fig.savefig(charts_dir / "c6_cost_vs_performance.png", dpi=DPI)
plt.close(fig)
# --- report assembly ------------------------------------------------------
def scoreboard_rows(board: dict[str, dict]) -> list[list]:
ranked = sorted(FORMATS, key=lambda f: -board[f]["total"])
return [[FMT_NAMES[fmt], board[fmt]["search_pts"], board[fmt]["ram_pts"], board[fmt]["disk_pts"], board[fmt]["total"]] for fmt in ranked]
def config_string(row: dict, prices: dict) -> str:
return (f"{row['nodes']} x 1U {CHASSIS[row['cores']]} ({row['cores']} cores);"
f" {row['ram_modules']} x {prices['ram_module_gb']} GB DDR5 ECC RDIMM;"
f" {row['nvme_drives']} x {prices['nvme_drive_tb']} TB NVMe U.2")
def assemble_report(tables: dict[str, tuple[list[str], list[list]]], scoreboards,
environment: str, diagrams: dict[str, str], winner: str, sizing, prices: dict) -> str:
t1_h, t1_r = tables["t1_storage_footprint"]
t2_h, t2_r = tables["t2_measured_medians"]
t3_h, t3_r = tables["t3_projected_600gb"]
t4_h, t4_r = tables["t4_hardware_600gb"]
pg6 = sizing[("pg", 6000)]
rdf6 = sizing[("rdf", 6000)]
as_of = prices["as_of"]
sources = prices["sources"]
parts = [
"# Laboratory Data Storage Format Evaluation - Final Report",
"",
"Generated by task 11 (report.py) from the measured benchmark matrix"
" (task 09) and the extrapolation model (task 10). Every number at"
" scales beyond the 150 MB corpus is a PROJECTION and is labeled as"
" such; scaling assumptions are listed in section 6.",
"",
"## 1. Executive summary",
"",
f"**Recommendation: {FMT_NAMES[winner]} as the system of record.** The measured"
" matrix and the projections confirm the expected outcome:",
"",
"- **PostgreSQL** is the only format whose seven benchmark queries stay"
" below the 1-hour practicality bound at every projected scale (600 GB,"
f" 1.2 TB, 6 TB); at 6 TB it needs one {pg6['cores']}-core node with"
f" {pg6['ram_gb']:.0f} GB RAM at an estimated {money(pg6['est_cost_usd'])}"
f" (bill of materials, prices as of {as_of}). Partitioned parallel scans"
" and the track_summary materialized view are the scaling levers.",
"- **CSV** remains the canonical raw archive: byte-reproducible,"
" instrument-native, and the best compressor (tar.gz ~2.8x); its full"
" scans exceed 1 h beyond ~600 GB, so it is an archive, not a query"
" layer.",
"- **SQLite** is an excellent single-user store up to ~600 GB (all"
" queries OK), then its single-threaded scans degrade (Q5/Q7"
" IMPRACTICAL at 1.2 TB).",
"- **Hybrid JSON-LD** (metadata + sourceFile links, 1.8 MiB at 150 MB"
" corpus) is the exchange format; the FULL variant re-reads the whole"
" corpus per scan and becomes impractical past 600 GB.",
"- **RDF (materialized triplestore)** fails at scale: 11.7x storage"
" blow-up, 5 of 7 queries FAIL at 6 TB, and the hot set needs"
f" {rdf6['ram_gb'] / 1024:.1f} TB RAM across {rdf6['nodes']} nodes"
f" (~{money(rdf6['est_cost_usd'])} projected). Semantic access should be"
" a virtual layer (e.g. Ontop OBDA) over PostgreSQL instead.",
"",
"## 2. Methodology snapshot",
"",
"35 cells (Q1-Q7 x 5 formats); per cell 1 warm-up + 3 measured runs in"
" isolated subprocesses sampled with psutil at 50 ms; randomized cell"
" order; 30-min timeout; every run validated against canonical results"
" (PostgreSQL cross-validated with SQLite, tolerance 1e-9). Medians are"
" reported; raw runs and min/max live in out/bench/results_raw.csv."
" The cold-cache pass is skipped on this Windows host and marked as"
" such. PostgreSQL runs on a remote host; its wall times include the"
" LAN round-trip, and server-side metrics come from pg_stat_statements"
" and Prometheus (docs/rules/bench-methodology.md section 6).",
"",
"## 3. Environment",
"",
"```",
environment.rstrip(),
"```",
"",
"## 4. Storage footprint",
"",
md_table(t1_h, t1_r),
"",
"![C1 - disk footprint](charts/c1_disk_footprint.png)",
"",
"## 5. Measured performance (150 MB corpus)",
"",
md_table(t2_h, t2_r),
"",
"Sub-interval note: cells faster than the 50 ms sampling cadence"
" (fast SQLite/CSV point reads) under-report client RSS/CPU; wall"
" times are exact.",
"",
"![C2 - RAM requirement](charts/c2_ram_requirement.png)",
"",
"Per-query charts: " + ", ".join(f"[Q{qn}](charts/c3_q{qn}_wall_time.png)" for qn in QUERY_NUMBERS),
"",
"## 6. Projections (600 GB / 1.2 TB / 6 TB)",
"",
"All values in this section are projected, never measured. Laws per"
" access pattern (spec 10): flat / index-depth log / linear result set"
" / linear x depth / full scan / parallel partitioned scan; constants"
" calibrated on the measured point after subtracting the per-format"
" harness floor. Flags: IMPRACTICAL > 1 h, FAIL > 24 h.",
"",
"Projected wall times at 600 GB:",
"",
md_table(t3_h, t3_r),
"",
"Degradation by data size: [point read](charts/c4_degradation_point_read.png),"
" [indexed / join](charts/c4_degradation_indexed.png),"
" [full scan](charts/c4_degradation_full_scan.png)",
"",
"## 7. Hardware sizing and cost at 600 GB (projected)",
"",
f"Bill-of-materials costing with street prices as of **{as_of}**"
" (sources below). Every configuration buys whole components: 64 GB"
" DDR5 ECC RDIMM modules, 7.68 TB enterprise NVMe U.2 drives, and a"
" priced 1U chassis; disk capacity includes 30% free-space headroom.",
"",
md_table(t4_h, t4_r),
"",
f"Component prices (retrieved {as_of}):",
"",
f"- {prices['ram_module_gb']} GB DDR5-5600 ECC RDIMM (A-Tech, 2x32GB kit):"
f" {money(prices['ram_module_usd'])} - <{sources['ram']}>",
f"- {prices['nvme_drive_tb']} TB NVMe U.2 1DWPD enterprise SSD (Dell 14-16G"
f" compatible, Cloud Ninjas): {money(prices['nvme_drive_usd'])} - <{sources['nvme']}>",
f"- 1U AMD EPYC 9004/9005 chassis (Broadberry starting configurations):"
f" {CHASSIS[16]} {money(prices['server_16_core_usd'])},"
f" {CHASSIS[32]} {money(prices['server_32_core_usd'])} - <{sources['servers']}>",
f"- Market context - server DRAM is in a documented 2025-2026 price"
f" surge: <{sources['market_context']}>",
"",
"Full sizing at 1.2 TB and 6 TB: out/bench/hardware_sizing.csv."
f" RDF exceeds the 1 TB single-node RAM ceiling already at 600 GB and"
f" reaches {rdf6['nodes']} nodes at 6 TB.",
"",
"![C6 - cost vs performance](charts/c6_cost_vs_performance.png)",
"",
"## 8. Weighted scoring",
"",
f"Subscores are 10 x best/value per metric; weights: search x{WEIGHT_SEARCH},"
f" RAM economy x{WEIGHT_RAM}, disk economy x{WEIGHT_DISK}; max"
f" {10 * (WEIGHT_SEARCH + WEIGHT_RAM + WEIGHT_DISK)}. Search uses the"
" inverse geometric mean of Q1-Q7 (FAIL at 600 GB zeroes the search"
" subscore). Left: measured scale; right: projected 600 GB.",
"",
"Measured scale (150 MB):",
"",
md_table(["format", f"search (x{WEIGHT_SEARCH})", f"RAM (x{WEIGHT_RAM})", f"disk (x{WEIGHT_DISK})", "total (max 80)"],
scoreboard_rows(scoreboards["measured"])),
"",
"Projected at 600 GB:",
"",
md_table(["format", f"search (x{WEIGHT_SEARCH})", f"RAM (x{WEIGHT_RAM})", f"disk (x{WEIGHT_DISK})", "total (max 80)"],
scoreboard_rows(scoreboards["projected_600gb"])),
"",
"![C5 - weighted score](charts/c5_weighted_score.png)",
"",
"## 9. Use-case mapping",
"",
md_table(["Use case", "Recommended format(s)"], [
["Interactive analysis (joins, aggregations)", "PostgreSQL; SQLite acceptable single-user up to ~600 GB"],
["Report generation (repeated summaries)", "PostgreSQL (track_summary materialized view)"],
["Search / filtering", "PostgreSQL or SQLite (indexed); flat formats need full scans"],
["Archiving", "CSV tree + tar.gz (canonical raw) plus pg_dump of the system of record"],
["Inter-lab exchange", "Hybrid JSON-LD (metadata + sourceFile links to CSV)"],
["Semantic / ontology queries", "Virtual RDF layer over PostgreSQL (e.g. Ontop OBDA), not a materialized triplestore"],
]),
"",
"## 10. Limitations",
"",
"- Single-node measurements on one Windows host + one remote"
" PostgreSQL host; no cluster or cloud variance.",
"- The corpus is simulated (physically plausible models, fixed seed"
" 20260711); real instrument data may have different value"
" distributions, though volumes and shapes match the lab.",
"- Extrapolation uses analytic laws calibrated on a single 150 MB"
" point; no intermediate-scale validation runs were performed.",
"- Client wall times include a per-format harness floor (interpreter"
" start, imports, connection); the floor is assumed constant across"
" scales. The cheapest cell of each format therefore projects flat.",
"- The cold-cache matrix pass is skipped (no page-cache drop on"
" Windows); all measured runs are warm-cache.",
"- The pandas CSV variant was not measured; its RAM is O(n) and it is"
" infeasible above roughly 10 GB (spec 10).",
"- RDF numbers reflect the compact modeling and a client that derives"
" run-in / steady-state from raw cycles; SPARQL-side aggregation"
" engines could shift (not remove) the scan penalty.",
f"- Hardware prices are spot street prices retrieved {as_of} during a"
" documented server-DRAM price surge; re-cost by editing"
" out/config/hw_prices.yaml and re-running tasks 10-11.",
"",
"## Appendix A - process flow diagrams (task 02)",
"",
]
for code in sorted(diagrams):
parts += [f"### {code} - {DIAGRAM_TITLES[code]}", "", "```mermaid", diagrams[code].rstrip(), "```", ""]
return "\n".join(parts) + "\n"
def main() -> int:
args = parse_task_args("Task 11: final report - tables, charts, weighted scoring")
out_root: Path = args.out_root
bench_dir = out_root / "bench"
report_dir = out_root / "report"
try:
check_dependencies(out_root, DEPENDS_ON)
remove_stale_marker(out_root, TASK_ID)
medians = read_medians(bench_dir / "results_median.csv")
proj = read_projections(bench_dir / "extrapolation.csv")
sizes = read_sizes(bench_dir / "storage_sizes.csv")
sizing = read_sizing(bench_dir / "hardware_sizing.csv")
prices = read_hw_prices(out_root / "config" / "hw_prices.yaml")
environment = (bench_dir / "environment.txt").read_text(encoding="utf-8")
diagrams = {p.stem: p.read_text(encoding="ascii") for p in sorted((report_dir / "diagrams").glob("D*.mermaid"))}
if len(diagrams) != 6:
raise ValueError(f"expected 6 process diagrams, found {sorted(diagrams)}")
charts_dir = report_dir / "charts"
tables_dir = report_dir / "tables"
charts_dir.mkdir(parents=True, exist_ok=True)
tables_dir.mkdir(parents=True, exist_ok=True)
csv_bytes = sizes[WORKING_VARIANT["csv"]]
coeff = {fmt: sizes[WORKING_VARIANT[fmt]] / csv_bytes for fmt in FORMATS}
proj_bytes_600 = {fmt: coeff[fmt] * REPORT_SCALE_GB * 1e9 for fmt in FORMATS}
tables: dict[str, tuple[list[str], list[list]]] = {}
tables["t1_storage_footprint"] = (
["format", "measured", "vs CSV", "archival (compressed)",
"projected @600 GB", "projected @1.2 TB", "projected @6 TB"],
[[FMT_NAMES[fmt], human_size(sizes[WORKING_VARIANT[fmt]]), f"{coeff[fmt]:.2f}x",
human_size(sizes[ARCHIVAL_VARIANT[fmt]]),
*[human_size(coeff[fmt] * s * 1e9) for s in SCALES_GB]] for fmt in FORMATS],
)
tables["t2_measured_medians"] = (
["format", "query", "wall time", "peak RSS", "CPU util", "read"],
[[FMT_NAMES[fmt], f"q{qn}", human_time(medians[(fmt, qn)]["wall_s"]),
f"{medians[(fmt, qn)]['peak_rss_mb']:.0f} MB",
f"{medians[(fmt, qn)]['cpu_util_pct']:.0f}%",
f"{medians[(fmt, qn)]['read_mb']:.1f} MB"]
for fmt in FORMATS for qn in QUERY_NUMBERS],
)
tables["t3_projected_600gb"] = (
["format", "query", "projected wall time", "flag"],
[[FMT_NAMES[fmt], f"q{qn}", human_time(proj[(fmt, qn, REPORT_SCALE_GB)][0]), proj[(fmt, qn, REPORT_SCALE_GB)][1]]
for fmt in FORMATS for qn in QUERY_NUMBERS],
)
tables["t4_hardware_600gb"] = (
["format", "configuration", "chassis", "RAM", "disk", "extra nodes", "total"],
[[FMT_NAMES[fmt], config_string(sizing[(fmt, REPORT_SCALE_GB)], prices),
money(sizing[(fmt, REPORT_SCALE_GB)]["base_usd"]),
money(sizing[(fmt, REPORT_SCALE_GB)]["ram_usd"]),
money(sizing[(fmt, REPORT_SCALE_GB)]["disk_usd"]),
money(sizing[(fmt, REPORT_SCALE_GB)]["extra_nodes_usd"]),
money(sizing[(fmt, REPORT_SCALE_GB)]["est_cost_usd"])] for fmt in FORMATS],
)
tables["t6_projected_all_scales"] = (
["format", "query", "scale_gb", "projected_wall_s", "projected_wall", "flag"],
[[FMT_NAMES[fmt], f"q{qn}", s, proj[(fmt, qn, s)][0], human_time(proj[(fmt, qn, s)][0]), proj[(fmt, qn, s)][1]]
for fmt in FORMATS for qn in QUERY_NUMBERS for s in SCALES_GB],
)
tables["t7_hardware_all_scales"] = (
["format", "scale_gb", "configuration", "disk_tb", "ram_gb", "cores", "nodes",
"ram_modules", "nvme_drives", "base_usd", "ram_usd", "disk_usd", "extra_nodes_usd", "est_cost_usd"],
[[FMT_NAMES[fmt], s, config_string(sizing[(fmt, s)], prices),
sizing[(fmt, s)]["disk_tb"], sizing[(fmt, s)]["ram_gb"],
sizing[(fmt, s)]["cores"], sizing[(fmt, s)]["nodes"],
sizing[(fmt, s)]["ram_modules"], sizing[(fmt, s)]["nvme_drives"],
round(sizing[(fmt, s)]["base_usd"]), round(sizing[(fmt, s)]["ram_usd"]),
round(sizing[(fmt, s)]["disk_usd"]), round(sizing[(fmt, s)]["extra_nodes_usd"]),
round(sizing[(fmt, s)]["est_cost_usd"])] for fmt in FORMATS for s in SCALES_GB],
)
scoreboards = build_scoreboards(medians, proj, sizes, sizing)
tables["t5_weighted_scores"] = (
["scale", "format", "search_pts", "ram_pts", "disk_pts", "total"],
[[scale, FMT_NAMES[fmt], board[fmt]["search_pts"], board[fmt]["ram_pts"], board[fmt]["disk_pts"], board[fmt]["total"]]
for scale, board in scoreboards.items() for fmt in FORMATS],
)
for name, (headers, rows) in tables.items():
write_csv(tables_dir / f"{name}.csv", headers, rows)
winner = max(FORMATS, key=lambda f: scoreboards["projected_600gb"][f]["total"])
chart_c1(charts_dir, sizes, proj_bytes_600, coeff)
chart_c2(charts_dir, medians, sizing)
chart_c3(charts_dir, medians, proj)
chart_c4(charts_dir, medians, proj, f"measured {human_size(csv_bytes)}")
chart_c5(charts_dir, scoreboards, winner)
chart_c6(charts_dir, proj, sizing)
charts = sorted(p.name for p in charts_dir.glob("*.png"))
if len(charts) != 14:
raise ValueError(f"expected 14 charts, wrote {len(charts)}: {charts}")
report_text = assemble_report(tables, scoreboards, environment, diagrams, winner, sizing, prices)
report_path = report_dir / "REPORT.md"
report_path.write_text(report_text, encoding="ascii", newline="\n")
entries = {
"charts": len(charts),
"tables": len(tables),
"report_bytes": report_path.stat().st_size,
"winner_600gb": winner,
"prices_as_of": prices["as_of"],
}
for fmt in FORMATS:
entries[f"score600_{fmt}"] = scoreboards["projected_600gb"][fmt]["total"]
marker_path = write_marker(out_root, TASK_ID, entries)
except Exception:
logger.critical("task %s failed", TASK_ID, exc_info=True)
return 1
print(f"task 11 ok: REPORT.md ({entries['report_bytes']} bytes), {entries['charts']} charts, {entries['tables']} tables")
print(f"winner @600 GB (projected): {winner} - " + ", ".join(
f"{fmt}={scoreboards['projected_600gb'][fmt]['total']}" for fmt in FORMATS))
print(f"report: {report_path}")
print(f"marker: {marker_path}")
return 0
if __name__ == "__main__":
sys.exit(main())