- adapt meta-/code-/obs- rules to the local Python benchmark pipeline - replace db-sql-ddl, code-config-env-scope, test-e2e-pytest with pipeline equivalents - add code-python-style, data-determinism, data-naming-units, bench-methodology, build-pipeline-tasks - normalize specs and README typography to ASCII per code-data-formatting - rewrite root CLAUDE.md trigger table; add .cursorrules and .gitignore - specs: pipeline plan and task specs 00-11 - rules: 19 binding rule files adapted for this project - docs: CLAUDE.md rule-trigger table - config: .cursorrules commit convention, .gitignore excluding out/ and .venv/ - docs: rewrite README.md with pipeline diagrams, setup guide, result placeholders - config: requirements.txt for the closed dependency list - datagen: make_lab_config.py writes out/config/lab_config.yaml and .done marker
1.4 KiB
1.4 KiB
04 - Convert to JSON-LD
Context / Goal
Produce the ontology-annotated JSON-LD representation of the corpus: the format evaluated for inter-lab data exchange (FAIR). Two variants: (a) full - bulk points embedded; (b) hybrid - metadata JSON-LD with sourceFile links to CSV.
Dependencies
03 (./out/csv/, MANIFEST.csv).
Input
CSV tree.
Processing instructions
@context: project vocabhttps://sandia.gov/ontology/tribology#+ QUDT units + PROV-O provenance; predicatespartOf,cutFrom,performedOn,performedByas@idlinks.- Entity classes:
Batch, Wafer, Coupon, Instrument, CompositionMeasurement, NanoindentationMeasurement, AFMMeasurement, ProfilometryMeasurement, FrictionTest (track), FrictionCycle, FrictionLoopPoint, WearMeasurement, SimtraProfile, Run. - Variant FULL: one JSON-LD file per coupon (embedding its tracks/cycles/loops), compact serialization (no pretty-print). Directory
./out/json/full/. - Variant HYBRID: single
./out/json/hybrid/dataset.jsonldwith all metadata entities; cycle/loop arrays replaced by{"sourceFile": "<relative csv path>", "rows": N, "sha256": "..."}. - Stream-write; validate one sample file with
json.load+ rdflib JSON-LD parse.
Output
./out/json/full/(expect ~0.7-1.2 GB),./out/json/hybrid/dataset.jsonld(expect < 20 MB).- Record byte sizes of both variants into
./out/bench/storage_sizes.csv(append: format, variant, bytes). - Marker
./out/.done/04.ok.