Files
astrolabe/docs/spec/06-chart-builder.md
T
oleh af9ee1e4c0 Add aggregation, binning, granularity, sort, and stacking to the Chart Builder
- Per-channel transforms: aggregate (sum/mean/median/min/max), quantitative
  bin, and temporal timeUnit granularity; bin and aggregate are mutually
  exclusive. A field-less "Count of records" measure (Voyager's count(*)).
- Chart-level sort (rank a categorical axis by its measure) and stacking
  (zero / 100% normalize), each shown only when it applies.
- Field type is a fixed N|O|Q|T segmented control with the column's invalid
  types disabled; SegmentedControl gains APG-correct disabled options.
- A crowded-category-axis warning (a raw measure drawing one mark per row over
  a large dataset) and a disabled-Create hint (says why it's disabled).
- Drop the Create success toast — the new snippet is immediately visible.
- Docs: spec §06, research-doc §8 backlog (incl. the cardinality/extent
  profiling TODO), architecture 01 (stable-selector rule) and 05 (builder-local
  preview), and a profiling breadcrumb.
2026-06-06 18:04:24 +03:00

10 KiB

06 · Chart Builder

The Chart Builder is a visual, no-JSON way to compose a Vega-Lite chart from a selected dataset. The user picks a mark type and maps the dataset's columns to encoding channels; the builder produces a complete Vega-Lite spec and saves it as a new snippet that references the dataset. It is intended for users who want to start a chart quickly without hand-writing JSON in the Spec Editor & Draft/Published Workflow.

Design level — "smart + guarded" (Tier B). The builder is mark-first and stays within the inputs below, but it is not a dumb composer: it picks a sensible default mark for the data shape, offers only field types valid for each column, keeps unsuitable channel mappings out of reach, and surfaces non-blocking guidance for encodings that render poorly. These behaviors are derived from cross-source chart-choice research recorded in docs/chart-builder-research.md (the convergence of Draco, Voyager, the FT Visual Vocabulary, and Datawrapper). The richer "intent-first" front door (ask what do you want to show? and recommend a chart) is explicitly out of scope for now and noted there as a future tier.

Opening

  • Launched from a selected dataset in the Datasets manager via that dataset's "build chart" action.
  • Opens as a modal dialog over the application; the URL reflects the dataset's "build" action so the open builder is shareable/restorable (see Application Shell & Navigation).
  • On open, the builder loads the selected dataset, displays its name, and pre-populates sensible defaults (see below). If no dataset is available, it shows a "No dataset loaded" message and offers no controls.

Layout

A two-pane modal:

  • Left — configuration: dataset name, mark type selector, one row per encoding channel, optional width/height inputs, and a "Create Snippet" action.
  • Right — live preview: a rendered chart that updates as the configuration changes, with a placeholder/error area.

Inputs and Controls

Mark type

  • Single selection from an exact set of five mark types: Bar, Line, Point, Area, Circle.
  • On open, the mark defaults to the type that best fits the pre-populated X/Y field-type shape (Tier B smart default): a temporal axis against a measure → Line; two measures → Point; a category against a measure → Bar; two categories → Point; and Bar as the fallback when only one axis (or none) is mapped. The user can switch to any of the five afterward.
  • Exactly one mark type is active at any time; selecting one updates the preview.

Encoding channels

  • Exactly four channels are offered, in this order: X, Y, Color, Size.
  • For each channel the user:
    • Picks a dataset column from a dropdown of the dataset's detected columns (see Datasets for column detection). A "None" option leaves the channel unmapped. Each column option shows a small type indicator alongside the column name.
    • Optionally overrides the channel's field type. The override appears only once a column is selected, and offers only the types valid for that column (Tier B valid-type locking) — a string/boolean column never offers Quantitative, and only a date column offers Temporal. Concretely: number → {Quantitative (default), Ordinal, Nominal}; date → {Temporal}; text → {Nominal (default), Ordinal}; boolean → {Nominal}. When a column admits only one valid type, no override control is shown.
  • When a column is chosen, its field type defaults from the dataset's inferred column type (numeric → Quantitative, date → Temporal, otherwise Nominal); the user may change it within the valid set above.
  • Size discipline: the Size channel accepts only Quantitative or Ordinal columns — size implies an ordered magnitude, so categorical (Nominal) and Temporal columns are not offered for Size (they remain available on X/Y/Color). A column that can't go on Size is shown disabled there with a brief reason.
  • The column dropdown also offers a field-less "Count of records" measure (Vega-Lite count) — a quantitative count of the rows, with no column.
  • Clearing a channel back to "None" leaves it out of the produced spec.
  • A Swap X/Y control exchanges the X and Y mappings (field and type) in one click, for quickly flipping the axes of the pre-populated default without re-selecting both columns.

The field-type override is presented as a fixed N | O | Q | T segmented control (abbreviations with full-name tooltips, after Datasets' Nominal/Ordinal/Quantitative/Temporal), always showing all four with the column's invalid types disabled rather than hidden — so the control keeps one shape on every channel.

Transforms (per channel)

Once a column is mapped, the channel offers the transforms that apply to its field type — and only those:

  • Aggregate (a measure / Quantitative field): one of Sum, Mean, Median, Min, Max, or None. (The field-less Count measure is chosen via the "Count of records" column option above.)
  • Bin (a Quantitative field): bins the values into ranges — e.g. a Quantitative X binned with a Count Y is a histogram. Binning and aggregating the same field are mutually exclusive (setting one clears the other).
  • Granularity (a Temporal field): a Vega-Lite timeUnit — Year, Year-Quarter, Year-Month, Year-Month-Day, Quarter, Month, Week, Day of month, Day of week, Hour — or None (raw timestamps). Defaults to None (no silent change to what the raw data shows).

Sort and stacking (chart-level)

These controls appear only when they apply:

  • Sort (when X and Y form a category-vs-measure pair): sorts the categorical axis by the measure — Ascending, Descending, or None — the standard way to rank a bar chart.
  • Stacking (a Bar or Area mark with a Color series): Stacked (absolute) or 100% (normalized, part-to-whole). Bars/areas without a Color series, or other marks, show no stacking control.

Default pre-population

  • On open, the first detected column is assigned to X and the second (if any) to Y, each with its derived field type and no transforms. Remaining channels start unmapped. The mark starts at the smart default for that X/Y shape (see Mark type), not unconditionally Bar.

Guidance (non-blocking)

The builder surfaces short, plain-language hints for configurations that render but read poorly — advisory only, never blocking the Create Snippet action (validation below is the sole gate). These follow the chart-choice research (docs/chart-builder-research.md) and include, for example:

  • A Line or Area mark with only one axis mapped (both axes are needed to draw it).
  • A Bar/Line/Area whose X and Y are both categories (nothing to measure).
  • Two measures on a non-scatter mark (a scatter — Point/Circle — usually reads better).
  • An Area chart split into multiple colour series (per-series change is hard to see).
  • A Bar/Line/Area that pairs a category axis with a raw (un-aggregated) measure over a many-row dataset — it draws one mark, and one axis label, per row, so the category axis becomes an unreadable picket fence. The hint suggests aggregating the measure (one mark per category) or, for a bar, flipping to a horizontal bar (Swap X/Y) where long labels stay readable (FT Visual Vocabulary / Datawrapper). Only the un-aggregated case (mark-count = row-count) is detected; flagging an aggregated axis that still has many distinct categories needs per-column distinct counts the profiler does not yet compute (a known gap).

A clean configuration shows no hints.

Dimensions (optional)

  • Optional numeric Width and Height inputs in pixels.
  • When left empty, the chart uses default/responsive sizing (consistent with Live Preview); when provided, the values are written into the spec.

Live Preview

  • The right pane renders the chart described by the current mark, encodings, and dimensions, resolving the dataset reference to its actual data (same rendering behavior as Live Preview).
  • Updates are debounced: changes to mark, encodings, or dimensions trigger a re-render after a short pause rather than on every keystroke.
  • While no encoding is mapped, the pane shows a placeholder instructing the user to configure at least one encoding.
  • If the spec fails to render, the pane shows an inline error message describing the problem instead of a chart.

Validation

  • A chart requires at least one channel mapped to a column.
  • While no channel is mapped, the "Create Snippet" action is disabled and the preview shows the configuration prompt.

Output / Create

Selecting "Create Snippet" produces the final artifact:

  • Builds a complete Vega-Lite spec containing: the schema reference, a named data reference to the dataset, the chosen mark (with tooltips enabled), the mapped encodings (each with its field and field type, plus any aggregate / bin / timeUnit transform), chart-level sort and stacking where set, and any explicit width/height.
  • Channels left unmapped are omitted; if no encodings exist the spec omits the encoding block entirely (prevented by validation here).
  • Creates a new snippet from that spec with an auto-generated descriptive name, adds it to the snippet library, and records that it was built from the dataset.
  • Links the snippet to the dataset by recording the dataset reference, so the bidirectional snippet↔dataset relationship is established (see Datasets).
  • Closes the builder; the newly created snippet becomes the active snippet in the library/editor. No success toast — the result is immediately visible (the new snippet opens in the editor), so a toast would be noise (architecture 10 §1, "toast only what the user can't already see"). This refines the earlier blanket "every action toasts" rule, consistent with the Extract-to-dataset / publish reconciliation.

Closing

  • The builder can be dismissed without creating anything (close control / modal dismissal).
  • Closing resets all builder state (dataset, mark type, encodings, dimensions, preview) so a later open starts fresh, and any pending preview render is cancelled.