4 The sample format§
PROVISIONAL The grammar is normative and closed. The machine-readable form of every fixed fact in this clause is the sampleFormat member of bus-timeseries.schema.json.
4.1 What the format must survive§
The samples are the part of a package most likely to be read in 2056 by someone with none of today’s software: an energy baseline is evidence, and evidence outlives tools. So the format is judged against four requirements. Readable with no library — a person with a text editor can verify what a tool asserts. Streamable — a receiver processes a decade row by row without holding it in memory. Appendable — an exporter writes a large series in one pass without seeking backward. And outside the schema-validation bottleneck that 1.3 measured.
bus-csv-1 is CSV per [RFC4180], profiled hard enough to have exactly one form. Two candidate classes were considered and rejected, on the record:
- Parquet and HDF5 — rejected — both are binary formats whose readability depends on a maintained library, and both fail the 2056 test outright: a format that needs this decade’s toolchain to open is a vendor dependency with the vendor generalized. Their compression advantage is real and mostly recovered by DEFLATE over repetitive timestamped text; their query advantage is irrelevant to a package. A future binary
formattoken remains a MINOR change if scale evidence ever demands one. - Inline JSON arrays — rejected — samples as entity fields or document arrays would put hundreds of millions of rows through the JSON parser, the schema validator, and the canonicalizer. 1.3 is the measurement; this is the decision it forces.
- One multi-series file with a series column — rejected — a
seriesRefcolumn makes the manifest digest, the byte length, and the row count meaningless per series, obliges a receiver to partition the file before using it, and turns appending by independent writers into a merge. One series per file keeps every declared quantity checkable against exactly one artifact.
4.2 Container and encoding§
- A sample file is plain text, UTF-8, no byte order mark. Media type
text/csvwith thecharset=utf-8parameter, stated in the manifest resource entry. - Line endings are LF. This is a deliberate deviation from [RFC4180]’s CRLF, and the reason is canonical form: a format with two permitted line endings has two encodings of the same series, and clause 7 needs there to be one. LF is what current tools write; the carriage return buys nothing.
- The final row ends with a newline, so files append and concatenate cleanly and the row count is the line count minus one.
4.3 The header§
The first line is REQUIRED and is exactly timestamp,value,quality — these names, this order, lowercase, no spaces, no extra columns. A fixed header on a fixed schema is redundant, and deliberately so: a sample file separated from its package is still self-describing to the person who finds it.
4.4 Timestamps§
Storage is UTC; display is the building’s. The building already declares its IANA time zone (BUS-1 8.4), so any reader can render wall-clock time correctly across every daylight-saving rule that applied over the series’ life — which per-row offsets cannot do, and which per-row zone suffixes would pay for in bytes on every one of a hundred million rows. A source system that stored local time converts on export and records what it converted from (6.5).
Rows are in non-decreasing timestamp order. Equal timestamps SHOULD NOT occur; where a source system genuinely holds two records at one instant — change-of-value data at coarse clock resolution does this — they are preserved in source order rather than invented apart.
4.5 Values§
- For a
numberpoint: an optional minus sign, decimal digits, an optional.fraction. No exponent, no leading+, no thousands separator, and.as the only decimal mark regardless of locale.55.4,-3.25,0.5— never5.54e1, never55,4. - For a
booleanpoint:0or1. - For an
enumpoint: the member code verbatim, as declared by the point’s enumeration. - For a
stringpoint: [RFC4180] quoting where the value contains a comma, a double quote, or a newline; otherwise unquoted. - The value field is empty exactly where quality is
missing(4.6). An empty value with any other quality, or amissingrow carrying a value, is a grammar violation (test TS-1).
A gap is the absence of rows, visible against the declared sampling interval, and an exporter MUST NOT invent rows to fill one. A missing row and no row are different statements — the historian recorded that it failed to record versus the historian recorded nothing — and both age better than a plausible interpolation.
4.6 Quality§
The quality column carries one of good, uncertain, bad, missing, or is empty.
| Token | Meaning |
|---|---|
good | The source system recorded the sample as good. |
uncertain | Recorded, flagged doubtful: out of range, sensor fault suspected, manually entered. |
bad | Recorded and known bad. Carried because deleting evidence is worse than flagging it. |
missing | The source recorded that no value exists at this timestamp. The value field is empty. |
An empty quality field means the source system recorded no quality, which is unstated — NOT good. This is the same rule as provenance confidence (BUS-1 14.1): writing good for a source that never assessed quality asserts something nobody knows, which is the failure Requirement 9.4-3 of BUS-1 names for ports and this clause refuses for samples.
4.7 Worked example§
The supply air temperature of AHU-1 from the Cedar Street worked example (BUS-1 Annex D), one hour of it, with one failed scan and one doubtful reading. First the file:
resources/series/ahu1-sat-2026h1.csv
timestamp,value,quality
2026-06-18T05:00:00Z,13.3,good
2026-06-18T05:15:00Z,13.0,good
2026-06-18T05:30:00Z,,missing
2026-06-18T05:45:00Z,13.9,uncertain
2026-06-18T06:00:00Z,13.4,goodThen the series that declares it, in a timeSeries model document. Note that rowCount is 5 — the missing row is a row — and that the unit is the point’s:
{ "id": "urn:uuid:...0a1",
"type": "bus:Series",
"localId": "AHU1-SAT-TREND",
"pointRef": "urn:uuid:...061",
"format": "bus-csv-1",
"resourceRef": "resources/series/ahu1-sat-2026h1.csv",
"coverage": { "start": "2026-06-18T05:00:00Z", "end": "2026-06-18T06:00:00Z" },
"sampling": { "kind": "sampled",
"interval": { "value": 15, "unitCode": "min" } },
"unitCode": "degC",
"rowCount": 5,
"externalIds": [ { "scheme": "com.example.historian.trendId",
"value": "TL-201-14" } ],
"provenance": { "assertedBy": { "kind": "system", "name": "busref 0.3.0" },
"assertedAt": "2026-06-18T06:10:00Z",
"method": "measured",
"source": { "kind": "system", "label": "Historian PRD" } } }And the manifest entry that makes the file verifiable before a single row is read (5.2):
"resources": [
{ "path": "resources/series/ahu1-sat-2026h1.csv",
"mediaType": "text/csv; charset=utf-8",
"digest": "sha-256-k7r7wZyhDvFH70aMBcbl3cfiRWXXQnPC4SDvSiLhTqU=",
"byteLength": 183 } ]