Files
shiro-neko/src/skills-md/data.md
T

41 lines
1.8 KiB
Markdown

---
name: data
description: Process, validate, or transform data. Use when parsing files, cleaning datasets, designing a data pipeline, or debugging a transform that produces wrong output.
---
# Data
Bad data fails silently and far away from where it entered. Validate at the boundary, keep the
raw, and make every transform checkable.
## Validate at the boundary
Parse and validate when data enters the system, not when it is used. A schema check at the edge
turns a corrupt record into a clear rejection; skipping it turns the same record into a wrong
answer three layers later. Reject loudly, with the record and the reason — never coerce and
carry on.
## Keep the raw
Store the untransformed input alongside the derived. When a transform turns out to be wrong,
the raw lets you recompute; without it, the information is gone. Derived data is rebuildable;
source data is not.
## Transformations are pure and tested
A transform takes input and returns output with no hidden state, so it can be tested on a
fixture and re-run safely. Test the edge cases that actually occur in data: the empty field,
the wrong type, the unexpected null, the duplicate, the encoding that is not UTF-8.
## Duplicates, nulls, and ranges are the usual corruption
Check for: unexpected duplicates on a key, nulls where a value is required, values outside a
sane range (a negative age, a date in the future), and referential breaks (an id pointing at
nothing). These four catch most real-world data problems before they reach a report.
## Idempotent pipelines
A step that can be re-run without duplicating or corrupting its output is a step you can retry
after a failure. Key on a stable id and upsert rather than blind-insert. A pipeline you cannot
safely re-run is a pipeline you will one day have to fix by hand at 2am.