# Code / Data / Formatting - ASCII-only typographic characters in source files **Binding scope.** Strings that end up in any of the following MUST use ASCII typographic characters; never their Unicode "smart" equivalents: - Python source, SQL and SPARQL query files - Generated data and artifacts written to disk (CSV, JSON-LD, N-Triples, Turtle, MANIFEST and benchmark result CSVs, log messages) - Configuration files (`.yaml`, `.json`, `.toml`) - PowerShell / shell scripts (any file processed by Windows tooling) - Markdown documentation (`docs/`, `README.md`, `CLAUDE.md`) - the specs were normalized to ASCII typography on 2026-07-11 and stay that way **Long dashes are an absolute exception, no carve-outs.** Em-dash (U+2014) and en-dash (U+2013) MUST be replaced with ASCII hyphen-minus (`-`) in **every** file in this repository, including markdown documentation, code comments, and report prose. Reason: long dashes are visually indistinguishable from a hyphen in many monospace fonts, break diffs / grep / IDE search, and round-trip badly through Windows tooling chains (section 2). --- ## 1. The substitution table | Unicode char | Codepoint | Name | ASCII replacement | |---|---|---|---| | em-dash | U+2014 | em-dash | `-` (single hyphen-minus) | | en-dash | U+2013 | en-dash | `-` (single hyphen-minus) | | horizontal bar | U+2015 | horizontal bar | `-` | | minus sign | U+2212 | minus sign | `-` | | ellipsis | U+2026 | horizontal ellipsis | `...` (three ASCII dots) | | left double quote | U+201C | left double quotation mark | `"` | | right double quote | U+201D | right double quotation mark | `"` | | left single quote | U+2018 | left single quotation mark | `'` | | right single quote | U+2019 | right single quotation mark | `'` | | left guillemet | U+00AB | left guillemet | `"` | | right guillemet | U+00BB | right guillemet | `"` | | nbsp | U+00A0 | non-breaking space | regular space | **Single hyphen, not double.** When replacing an em-dash that separates two clauses, use a single `-` surrounded by spaces (`text - clause`), NOT `--`. **Ranges use a single hyphen:** `Q1-Q7`, `120-200 MB`, `0.3-1.1 um` - never an en-dash. ## 2. Why - documented failure modes 1. **PowerShell 5.1 reads no-BOM files as Windows-1252.** Non-ASCII UTF-8 bytes become mojibake and break string-literal parsing. This project is developed on Windows; any script or config touched by PowerShell is exposed. 2. **Windows -> git-bash -> tar / packaging round-trip.** At any step an unset locale or wrong encoding assumption produces mojibake. A UTF-8 byte sequence for U+2014 interpreted under a non-UTF-8 locale renders as a `?` replacement character in the packaged output. 3. **Copy-paste from Word / Google Docs / Notion / web articles.** These tools auto-substitute typography (`"` -> smart quotes, `'` -> smart apostrophe, `--` -> em-dash, `...` -> ellipsis) silently. The substitution survives copy into any editor and lands in source files unnoticed. The original specs arrived with en/em dashes this way. 4. **Data files must diff and grep cleanly.** The CSV corpus, MANIFEST, and benchmark results are compared by checksum and processed by streaming parsers; a stray typographic character breaks byte-level reproducibility. ## 3. Where the rule does NOT apply - **Scientific and unit notation is content, not typography.** Characters such as `µ` (micro), `±`, `°C`, `×`, `Ø`, and Greek letters (`τ`, `σ`) in formulas inside the specs are meaningful scientific notation and are kept. Note the boundary: in **column names and code identifiers** the units are always plain ASCII (`_um`, `_nm`, `_GPa` - see [data-naming-units.md](data-naming-units.md)); the scientific characters live only in explanatory prose. - **International language strings** (Cyrillic, CJK, etc.) are meaningful content, not typographic substitutions. UTF-8 throughout is the right answer. (Repository artifacts are English; this matters only for quoted material.) ## 4. Detection / enforcement There is no automated check for this rule yet - **not in CI yet**. The rule is enforced by review and by the editing discipline in this file. The intent of a future detector is documented here so it can be built consistently: - **What it flags.** Any occurrence of the Unicode codepoints in the section 1 substitution table inside in-scope files. The section 3 exemptions (scientific notation, international content) are the only carve-outs. - **Exit behavior.** Non-zero exit when any flagged codepoint is found under the scan roots, so it can later gate a commit hook or a build step. - **Tooling.** When this detector is created it MUST be a PowerShell script (this is a Windows project) and live under `tools/`. ## 5. Examples ### Good ```python # safe through every Windows tooling step print("Corpus size 152.3 MB - within the 120-200 MB gate.") label = "steady-state COF, cycles 500-1000" ``` ### Bad ```python # em-dash literal (U+2014) - mojibake risk on PowerShell read/write print("Corpus size 152.3 MB [U+2014] within the gate.") # en-dash range (U+2013) - breaks grep for "Q1-Q7" label = "queries Q1[U+2013]Q7" ``` ## 6. Anti-patterns - **Copy-pasting strings from a word processor or web page** into source / specs / config without normalization. - **Trusting tar / packaging to "preserve as-is".** They preserve bytes; the *encoding interpretation* is what gets lost when locale or encoding declaration is wrong on either end. - **Assuming the source file encoding alone is enough.** The source file may be fine - the failure is at downstream consumers (PowerShell read, build write, OS locale). - **Using `--` instead of `-` "for emphasis".** Single ASCII hyphen surrounded by spaces is the project's chosen separator. - **Putting `µ` or `°` into a column name or identifier** - units in names are ASCII (`_um`, `_C`); see [data-naming-units.md](data-naming-units.md).