Guide · Data documentation · 9 minute read

Create a data dictionary that stays useful.

A useful data dictionary is more than a table of column names and inferred types. It records what each field means, who owns that meaning, how sensitive the field is, which values are valid, and how the next delivery will be checked. Build those decisions beside observed evidence, then keep the definition separate from the rows.

Start here

Document meaning, not just shape

A schema describes machine-facing shape: field names, types, required values, and sometimes enumerations. A data dictionary adds the human contract: the business name, definition, owner, role, sensitivity, unit, format, allowed values, and stewardship notes. Both matter, but they answer different questions.

For example, mrr being a number does not tell a finance analyst whether it is gross or net, which currency it uses, whether tax is included, or which date determines the reporting month. The dictionary should make those decisions explicit enough that a new teammate can use the field without asking the person who created the export.

Start with one dataset-level purpose statement and owner. If nobody owns the definition, conflicting interpretations will accumulate even when the CSV remains structurally valid.

→ Open the CSV Data Dictionary workspace

Evidence

Profile the file before you define it

Observed evidence is a useful starting point, not a substitute for meaning. Record row and column counts, missingness, distinctness, inferred value shapes, candidate keys, short category candidates, and length bounds. These facts reveal where the written contract needs care: a supposedly required ID with blanks, a plan field with unexpected categories, or a timestamp column mixing shapes.

Keep observed and authored facts visibly separate. A column containing only free and pro in today's sample does not prove those are the only legitimate plans. An apparently unique customer ID in 500 rows is a candidate key, not proof that the upstream system enforces uniqueness.

Avoid copying raw sample values into broadly shared documentation. Aggregate counts and explicit allowed lists are usually enough. Raw rows belong in a separately controlled issue ledger when remediation actually needs them.

→ Run a deeper data-quality profile

Field definitions

Use a repeatable definition template

For every column, write a readable business name and one-sentence definition that states what the value represents. Then record the documented type, role, required state, unit, format, allowed values, and stewardship notes. A strong definition is specific enough to rule out a nearby interpretation.

Use roles consistently. An identifier joins records; a dimension groups them; a measure is aggregated; a timestamp places an event in time; an attribute describes it. Roles make the dictionary easier to scan and help reviewers catch contradictions such as summing a numeric-looking account code.

Write formats as conventions rather than examples where possible: ISO 8601 UTC, ISO 3166-1 alpha-2, or decimal USD before tax. Examples age quickly; a named convention remains testable.

Governance

Classify personal data and sensitivity explicitly

Mark whether a field contains personal data and give it a sensitivity level such as public, internal, confidential, or restricted. Header names like email, phone, address, passport, or customer ID can prompt review, but a header-only match is not a legal conclusion. Context, jurisdiction, combinations of fields, and organizational policy still matter.

Do not hide classification behind an icon or color. Write the state in full, include a note when handling depends on context, and keep ownership nearby so a reviewer knows who can resolve ambiguity.

A local-first tool reduces one risk because source rows need not leave the device, but it does not decide whether an exported dictionary or issue ledger is safe to share. Treat explicit allowed values and raw failing values according to the same policy as the data they describe.

→ Create a safer sample before sharing rows

Contract

Turn documentation into recurring validation

Once definitions are reviewed, use the dictionary as a contract for the next delivery. Resolve configured field names exactly after documented header normalization. Report missing reviewed columns and unexpected new columns instead of silently changing the dictionary to match the new file.

Run required, documented-type, and explicit allowed-value checks across every row, not just a preview. Keep schema drift separate from row failures: a missing field is a delivery-shape problem, while one invalid email or plan is a record-level problem with a source row.

A pass means the checks represented in the dictionary passed. It does not certify that every business definition is correct, that personal-data classification meets every law, or that an upstream system will accept the file. Those claims require human ownership and destination-specific evidence.

Maintenance

Version the definition, not the dataset

Save the reusable dictionary as versioned metadata: dataset title, purpose, owner, cadence, and ordered field definitions. Exclude filenames, source rows, samples, observed counts, signals, issues, and results. That boundary lets a teammate reuse the contract without inheriting yesterday's data.

Change the version when meaning or validation behavior changes: a field is renamed, a unit changes, a category is added, a required rule tightens, or sensitivity is reclassified. Record delivery-to-delivery observations in an audit report rather than rewriting the stable definition around every file.

Export formats should serve different readers. CSV works for spreadsheet review, Markdown for repositories and wikis, hardened HTML for a portable human document, JSON Schema for machine structure, and dictionary JSON for another run in the same workspace. Keep the raw-value issue CSV separate and disclosed.

Common questions
  • ·

    What fields should a data dictionary include?

    At minimum: source field name, business name, definition, documented type, role, required state, personal-data flag, sensitivity, unit, format, allowed values, and notes. At dataset level include title, purpose, owner, version, and refresh cadence.

  • ·

    What is the difference between a schema and a data dictionary?

    A schema is primarily machine-facing structure. A data dictionary adds business meaning, ownership, governance, and observed evidence. A mature workflow often exports both from the same reviewed definition.

  • ·

    Should sample values appear in a data dictionary?

    Only when they are safe and genuinely clarify the field. Aggregate evidence and authored allowed values are usually safer. Keep raw failing values in a separately controlled issue ledger.

  • ·

    Can types and descriptions be inferred automatically?

    Value shapes and aggregate evidence can be inferred deterministically. Business meaning, ownership, and sensitivity still require review; a current sample should not silently redefine the business contract.

  • ·

    How often should a data dictionary be updated?

    Review it whenever a field, meaning, unit, format, category, required rule, owner, or sensitivity changes. Validate every recurring delivery against the current reviewed version.

  • ·

    How do I validate a future CSV against the dictionary?

    Arm the saved definition before loading the next file. Check missing and unexpected columns, then run required, type, and allowed-value rules across every row and retain aggregate evidence for the run.

  • ·

    Is a personal-data header signal a compliance classification?

    No. It is a review prompt based on the header text. Legal and policy classification depends on context, jurisdiction, combinations of fields, and your organization's rules.

  • ·

    Does the data dictionary tool upload my CSV?

    No. Profiling, authoring, validation, persistence, and export run locally in your browser.

Keep going