14 KiB
Type Inference & Data Profiling
How Astrolabe looks at a tabular dataset and figures out, for each column, what
kind of data it holds — number, string, date, or boolean — and how it
rolls those facts up into the profile stored on a dataset record.
This is pure, portable logic. It lives in src/core/, touches no browser APIs
and no React, takes plain values in and returns plain data out, and is covered
by Vitest unit tests. Anything that needs a profile (the create form, the edit
flow, the detail panel) calls into this module; nothing here reaches back out.
1. Why infer types at all
A dataset is just rows of values. The UI wants to describe it without re-parsing the payload every time: how many rows and columns, what the columns are called, and roughly what each column contains. The inferred type drives the small type indicator next to each column name in the dataset detail panel and the meta line in the list. It is a display hint, not a contract — nothing downstream coerces values based on it, and Vega-Lite does its own type handling at render time. Because it is only a hint, a wrong guess is cheap, and the rules below favour being simple and predictable over being clever.
We support exactly four inferred types:
| Type | Meaning |
|---|---|
number |
Every non-empty value is numeric. |
boolean |
Every non-empty value is true/false (any case). |
date |
Every non-empty value parses as a date. |
string |
The fallback — anything that isn't one of the above. |
There is deliberately no integer/float split, no datetime-vs-date distinction, and no JSON type. Those distinctions add branches and edge cases without changing what the user sees. Keep it at four.
2. Inferring one column
Given the values of a single column, decide its type.
The shape of the algorithm
- Drop the empties. Filter out
null,undefined, and empty/whitespace-only strings before doing anything. Empty cells carry no type signal — a column of numbers with a few blanks is still a number column. - All-empty →
string. If nothing survives the filter (the column is entirely empty, or there are zero rows), default tostring. There is no evidence for any other type. - Run the type checks in precedence order. For each candidate type, ask:
does every surviving value match this type? The first candidate for
which the answer is yes wins. This is the "all values match → that type,
else fall back" rule: one stray value that doesn't fit knocks the column
down to the next candidate, and ultimately to
string.
Precedence order matters
The order of the checks is not arbitrary — it exists because the value-sets overlap, and we want the most specific interpretation that fits.
- boolean first. The strings
"true"/"false"are not numbers and not dates, so booleans never collide with the other checks — but putting them first keeps a0/1-free true/false column out ofstring. (We do not treat0/1as boolean; that's a number column.) - number second.
Number("2024")is a perfectly good number, so a column of bare years would read asnumber— which is the honest answer. Numbers are checked before dates so that plain numeric columns never get mis-classified as dates by an over-eager date parser. - date third. Date parsing is the loosest, most permissive check, so it goes last among the positive checks. By the time we reach it we already know the column isn't all-boolean and isn't all-numeric.
- string is the fallback when no positive check matches every value.
Mnemonic: boolean → number → date → string, narrowest evidence to widest.
What counts as each type
- numeric: trim the string form; reject empty;
Number(trimmed)must be finite and notNaN. (Nativenumbervalues pass directly.) Reject blank and whitespace soNumber("") === 0doesn't sneak through. - boolean: native
booleanvalues pass; otherwise the trimmed, lower-cased string must be exactly"true"or"false". - date: guard before parsing. Require the trimmed value to look
date-shaped (a leading
YYYY-MM-DDorYYYY/MM/DD, orM/D/YYYY) and then confirmDate.parsereturns a finite timestamp. The shape guard is essential:Date.parsewill happily accept"42"or"March"on some engines, which would swallow number and string columns. Never rely onDate.parsealone.
Sketch
// src/core/type-inference.ts
export type ColumnType = 'number' | 'string' | 'date' | 'boolean';
const isEmpty = (v: unknown): boolean =>
v === null || v === undefined || (typeof v === 'string' && v.trim() === '');
const isNumeric = (v: unknown): boolean => {
if (typeof v === 'number') return Number.isFinite(v);
if (typeof v !== 'string') return false;
const t = v.trim();
if (t === '') return false;
const n = Number(t);
return !Number.isNaN(n) && Number.isFinite(n);
};
const isBoolean = (v: unknown): boolean => {
if (typeof v === 'boolean') return true;
if (typeof v !== 'string') return false;
const t = v.trim().toLowerCase();
return t === 'true' || t === 'false';
};
// Shape guard first, then confirm it actually parses.
const DATE_SHAPE = /^\d{4}[-/]\d{2}[-/]\d{2}|^\d{1,2}\/\d{1,2}\/\d{4}/;
const isDate = (v: unknown): boolean => {
if (typeof v !== 'string') return false;
const t = v.trim();
return DATE_SHAPE.test(t) && !Number.isNaN(Date.parse(t));
};
/**
* Infer one of four column types from a sample of column values.
* Empty cells are ignored; an all-empty column is `string`.
* Precedence: boolean → number → date → string.
*/
export function inferColumnType(values: readonly unknown[]): ColumnType {
const present = values.filter((v) => !isEmpty(v));
if (present.length === 0) return 'string';
if (present.every(isBoolean)) return 'boolean';
if (present.every(isNumeric)) return 'number';
if (present.every(isDate)) return 'date';
return 'string';
}
Robustness notes
- Mixed columns fall through to
stringautomatically — theeverycheck fails on the first non-conforming value, so a column of mostly-numbers with one label isstring, which is the safe, honest answer. - Whitespace is trimmed in every check, so
" 42 "reads as numeric and" "is treated as empty. - Empty columns (all cells blank, or a zero-row dataset) return
stringby the all-empty rule — never throw, never guess. - Large columns: see §4.
inferColumnTypeitself just consumes whatever array it's handed; the caller decides whether to sample.
Do / Don't
- Do ignore empty cells before classifying.
- Do keep the precedence boolean → number → date → string.
- Do guard date detection with a shape regex before trusting
Date.parse. - Don't classify a column unless every present value matches — one
outlier means
string. - Don't add more types (integer, float, datetime, json). Four, no more.
- Don't let
Number(""),Date.parse("42"), or0/1leak into the wrong bucket.
3. Profiling a dataset
A profile is the set of derived summary fields stored on a dataset record so the UI can describe it without re-parsing the payload. Per the data model, a profiled dataset carries:
| Field | Type | Meaning |
|---|---|---|
rowCount |
number | null |
Data rows, or null when N/A. |
columnCount |
number | null |
Columns, or null when N/A. |
columns |
string[] |
Column names, in order. |
columnTypes |
Array<{ name; type }> |
Per-column inferred type (see §2). |
size |
number |
Approximate payload size in bytes. |
null row/column counts and an empty columns/columnTypes are how the UI
shows "N/A" — see §3.2.
3.1 What gets profiled
Profiling applies only to tabular inline data:
- JSON that is an array of objects.
- CSV (comma-separated, header row).
- TSV (tab-separated, header row).
Everything else is not profiled:
- URL datasets — the library holds only the link, not the data, so there is
nothing to scan. Counts are
null/ N/A. - Non-tabular data — a single JSON object, TopoJSON, or anything we can't
read as rows-of-columns. Counts are
null/ N/A.
For the not-profiled cases, size is still computed (it's just the byte length
of the stored payload), but rowCount and columnCount are null, and
columns/columnTypes are empty.
3.2 The algorithm
- Compute
sizefrom the raw payload regardless of whether it's tabular — byte length of the text (CSV/TSV) or of the JSON-serialized value. - Decide if it's tabular. Map
(format, parsed shape)to a row set:csv/tsv→ parse into rows-of-objects using the matching delimiter.jsonthat is a non-empty array of objects → use it directly.- anything else (
topojson, a lone JSON object, an empty array) → not tabular; return the N/A profile (rowCount: null,columnCount: null,columns: [],columnTypes: [], plussize).
- Derive columns from the union of keys across the rows (or the CSV/TSV header), preserving first-seen order.
- Infer each column's type by collecting that column's values across the
rows and calling
inferColumnType(§2), sampling per §4. - Assemble
rowCount,columnCount,columns,columnTypes,size.
Sketch
// src/core/profile.ts
import { inferColumnType, type ColumnType } from './type-inference';
export interface DatasetProfile {
rowCount: number | null;
columnCount: number | null;
columns: string[];
columnTypes: Array<{ name: string; type: ColumnType }>;
size: number;
}
const NA = (size: number): DatasetProfile => ({
rowCount: null,
columnCount: null,
columns: [],
columnTypes: [],
size,
});
/** Profile a dataset payload. `rows` is the tabular form (CSV/TSV/JSON-array)
* already parsed to rows-of-objects, or null for non-tabular / URL data. */
export function profileData(
rows: ReadonlyArray<Record<string, unknown>> | null,
size: number,
): DatasetProfile {
if (!rows || rows.length === 0) return NA(size);
// Column order = first-seen order across all rows.
const columns: string[] = [];
const seen = new Set<string>();
for (const row of rows) {
for (const key of Object.keys(row)) {
if (!seen.has(key)) {
seen.add(key);
columns.push(key);
}
}
}
if (columns.length === 0) return NA(size);
const sample = sampleRows(rows);
const columnTypes = columns.map((name) => ({
name,
type: inferColumnType(sample.map((r) => r[name])),
}));
return {
rowCount: rows.length,
columnCount: columns.length,
columns,
columnTypes,
size,
};
}
Parsing CSV/TSV text and detecting the payload shape happen upstream of
profileData; this function takes already-parsed rows so it stays pure and
trivially testable. The caller passes null for URL and non-tabular datasets.
4. Sampling vs. full scan
rowCount/columnCount/size always reflect the whole dataset — they're
cheap (a length and a byte count). Only type inference has a per-value cost,
and it's the one place a huge dataset could hurt.
So: infer types from a bounded sample of rows, not the full column. A fixed cap (e.g. the first ~200 rows) keeps profiling fast and predictable on large pasted datasets while still being more than enough signal to classify a column.
const SAMPLE_SIZE = 200;
const sampleRows = <T>(rows: ReadonlyArray<T>): ReadonlyArray<T> =>
rows.length <= SAMPLE_SIZE ? rows : rows.slice(0, SAMPLE_SIZE);
Trade-off to be aware of: a column that is numeric for its first 200 rows but
turns to text later will be mis-typed as number. That's an accepted cost — the
type is a display hint, the mistake is cheap, and the speed win on large
datasets is worth it. Sampling the head (rather than randomly) keeps results
deterministic, which matters for tests and for not surprising the user when
the same paste profiles the same way twice.
Do / Don't
- Do count rows/columns and size over the full payload.
- Do cap type-inference sampling at a fixed head slice for determinism.
- Don't randomly sample — non-deterministic profiles break tests and confuse users.
- Don't scan every value of a million-row paste to guess a type.
5. Testing
Both functions are pure, so tests are plain input/output assertions in Vitest — no mocks, no DOM, no fixtures beyond literal arrays.
Cover at least:
inferColumnType: each type detected from a clean column; mixed columns fall tostring; empty/whitespace cells ignored; all-empty and zero-length →string; precedence (a["true","false"]column isbooleannotstring; a["2024","2025"]column isnumbernotdate); date shape guard rejects"42"and"hello"even though one engine'sDate.parsemight accept them;0/1arenumber, notboolean.profileData: a JSON-array dataset profiles fully;nullrows (URL) and an empty array (non-tabular) return the N/A profile but still carrysize; column order follows first-seen key order across ragged rows; sampling cap is respected (a dataset longer than the cap still profiles, using only the head).
Summary
- Four types only: boolean → number → date → string, checked in that order.
- All present values must match a type or the column falls through; empty
cells are ignored; an all-empty column is
string. - Guard date detection with a shape regex before trusting
Date.parse. - A profile carries
rowCount,columnCount,columns,columnTypes,size; URL and non-tabular datasets get a null/N-A profile (still sized). - Counts and size scan the whole payload; type inference samples the head for speed and determinism.
- All of it is pure
src/core/logic, unit-tested with Vitest.