diff --git a/docs/architecture/11-learning-section.md b/docs/architecture/11-learning-section.md index e7f8cba..bb01ef8 100644 --- a/docs/architecture/11-learning-section.md +++ b/docs/architecture/11-learning-section.md @@ -42,7 +42,11 @@ Inside a `:::progression`, each `##` heading is a stage: heading → tab label, the following fenced `vega-lite` block → spec. A `:::data` block names datasets once for the whole lesson; specs reference them with `{ "data": { "name": … } }` and `injectDatasets` merges the rows in at render time — so a shared dataset isn't repeated per stage, and the -source pane keeps a stage's grammar legible instead of burying it under data. +source pane keeps a stage's grammar legible instead of burying it under data. Every lesson +uses the `:::data` form (never per-stage inline rows), targets a misconception rather than +a chart type, and closes with a "take it further" prose beat that leans on the per-stage +"Open in Astrolabe" links — the roster and per-lesson briefs live in +`docs/exploration/lessons-roadmap.md`. ## The pipeline diff --git a/src/learn/lessons/binning.md b/src/learn/lessons/binning.md index 691af18..899c99d 100644 --- a/src/learn/lessons/binning.md +++ b/src/learn/lessons/binning.md @@ -5,10 +5,18 @@ tagline: How one word turns a noisy bar chart into a distribution — and where --- Counting a continuous field is the most common chart that's _almost_ right. The naive -version compiles, renders, and quietly lies about your data. Walk the four stages below — -each tab is one small edit — and watch a picket fence become a distribution, then a finished +version compiles, renders, and quietly lies about your data. Walk the stages below — each +tab is one small edit — and watch a picket fence become a distribution, then a finished chart. +:::data +{ +"orders": [ +{"minutes":12},{"minutes":14},{"minutes":15},{"minutes":16},{"minutes":18},{"minutes":19},{"minutes":20},{"minutes":21},{"minutes":22},{"minutes":23},{"minutes":24},{"minutes":25},{"minutes":26},{"minutes":28},{"minutes":30},{"minutes":31},{"minutes":33},{"minutes":35},{"minutes":37},{"minutes":40},{"minutes":43},{"minutes":46},{"minutes":50},{"minutes":55},{"minutes":61},{"minutes":68},{"minutes":77},{"minutes":86} +] +} +::: + :::progression ## raw counts @@ -20,12 +28,12 @@ fence, not a distribution. It looks almost right, which is exactly what makes it ```vega-lite { "$schema": "https://vega.github.io/schema/vega-lite/v6.json", + "data": { "name": "orders" }, "mark": "bar", "encoding": { "x": { "field": "minutes", "type": "quantitative" }, "y": { "aggregate": "count" } - }, - "data": { "values": [{"minutes":12},{"minutes":14},{"minutes":15},{"minutes":16},{"minutes":18},{"minutes":19},{"minutes":20},{"minutes":21},{"minutes":22},{"minutes":23},{"minutes":24},{"minutes":25},{"minutes":26},{"minutes":28},{"minutes":30},{"minutes":31},{"minutes":33},{"minutes":35},{"minutes":37},{"minutes":40},{"minutes":43},{"minutes":46},{"minutes":50},{"minutes":55},{"minutes":61},{"minutes":68},{"minutes":77},{"minutes":86}] } + } } ``` @@ -37,12 +45,12 @@ counts each bucket — a real histogram from a one-word edit. ```vega-lite { "$schema": "https://vega.github.io/schema/vega-lite/v6.json", + "data": { "name": "orders" }, "mark": "bar", "encoding": { "x": { "field": "minutes", "type": "quantitative", "bin": true }, "y": { "aggregate": "count" } - }, - "data": { "values": [{"minutes":12},{"minutes":14},{"minutes":15},{"minutes":16},{"minutes":18},{"minutes":19},{"minutes":20},{"minutes":21},{"minutes":22},{"minutes":23},{"minutes":24},{"minutes":25},{"minutes":26},{"minutes":28},{"minutes":30},{"minutes":31},{"minutes":33},{"minutes":35},{"minutes":37},{"minutes":40},{"minutes":43},{"minutes":46},{"minutes":50},{"minutes":55},{"minutes":61},{"minutes":68},{"minutes":77},{"minutes":86}] } + } } ``` @@ -54,24 +62,51 @@ Take control of the resolution: `bin: { step: 10 }` forces clean 10-minute bucke ```vega-lite { "$schema": "https://vega.github.io/schema/vega-lite/v6.json", + "data": { "name": "orders" }, "mark": "bar", "encoding": { "x": { "field": "minutes", "type": "quantitative", "bin": { "step": 10 }, "title": "Delivery time (min)" }, "y": { "aggregate": "count", "title": "Orders" } - }, - "data": { "values": [{"minutes":12},{"minutes":14},{"minutes":15},{"minutes":16},{"minutes":18},{"minutes":19},{"minutes":20},{"minutes":21},{"minutes":22},{"minutes":23},{"minutes":24},{"minutes":25},{"minutes":26},{"minutes":28},{"minutes":30},{"minutes":31},{"minutes":33},{"minutes":35},{"minutes":37},{"minutes":40},{"minutes":43},{"minutes":46},{"minutes":50},{"minutes":55},{"minutes":61},{"minutes":68},{"minutes":77},{"minutes":86}] } + } +} +``` + +## + a mean line + +The first act of _composition_: the spec becomes a `layer` of two marks sharing the same +data and x-axis — the bars, and a `rule` at the mean. Each layer keeps its own mark and +encoding; the shared scale is what makes them one chart. This one jump is most of what +"layering" means. + +```vega-lite +{ + "$schema": "https://vega.github.io/schema/vega-lite/v6.json", + "data": { "name": "orders" }, + "layer": [ + { + "mark": "bar", + "encoding": { + "x": { "field": "minutes", "type": "quantitative", "bin": { "step": 10 }, "title": "Delivery time (min)" }, + "y": { "aggregate": "count", "title": "Orders" } + } + }, + { + "mark": { "type": "rule", "color": "#d6336c", "size": 2 }, + "encoding": { "x": { "field": "minutes", "aggregate": "mean" } } + } + ] } ``` ## polished -The last 20%: round the bar tops, add a tooltip that reports each bucket's range and count, -and _layer_ a mean line over the bars. Composition is where polish lives — one spec, two -marks sharing an axis. +The last 20% is legibility, not structure: round the bar tops, and add a tooltip that +reports each bucket's range and count. ```vega-lite { "$schema": "https://vega.github.io/schema/vega-lite/v6.json", + "data": { "name": "orders" }, "layer": [ { "mark": { "type": "bar", "cornerRadiusEnd": 2 }, @@ -88,8 +123,7 @@ marks sharing an axis. "mark": { "type": "rule", "color": "#d6336c", "size": 2 }, "encoding": { "x": { "field": "minutes", "aggregate": "mean" } } } - ], - "data": { "values": [{"minutes":12},{"minutes":14},{"minutes":15},{"minutes":16},{"minutes":18},{"minutes":19},{"minutes":20},{"minutes":21},{"minutes":22},{"minutes":23},{"minutes":24},{"minutes":25},{"minutes":26},{"minutes":28},{"minutes":30},{"minutes":31},{"minutes":33},{"minutes":35},{"minutes":37},{"minutes":40},{"minutes":43},{"minutes":46},{"minutes":50},{"minutes":55},{"minutes":61},{"minutes":68},{"minutes":77},{"minutes":86}] } + ] } ``` @@ -99,6 +133,9 @@ marks sharing an axis. **The sharp edge.** Tempted to wire a slider to the bin width? It won't work. In Vega-Lite, `bin.step` and `maxbins` are fixed numbers — no parameter or expression drives them, so a bound slider compiles fine and then does nothing. The one binning property you _can_ make -interactive is `extent`: brush an interval selection to re-bin a chosen range. That's a -lesson of its own. +interactive is `extent`: brush an interval selection to re-bin a chosen range. ::: + +Take it further: open the last stage in Astrolabe (the link under the chart) and try a +`step` of 5 against 20 — watch the story the histogram tells change with the resolution. +Then swap the rule's `mean` for `median` and see which one your outliers drag. diff --git a/src/learn/lessons/linked-views.md b/src/learn/lessons/linked-views.md index 68e61f5..9c0b660 100644 --- a/src/learn/lessons/linked-views.md +++ b/src/learn/lessons/linked-views.md @@ -12,8 +12,8 @@ re-aggregate by hand for every window you got curious about. So we build three views over one interval selection. Brush a span of the run on the top chart and the other two recompute over exactly that window — the cumulative gap between the -arms, and their share split. Five edits, and all three stay in sync through a single -`param`. +arms, and their share split. A handful of small edits, and everything stays in sync +through a single `param`. :::data { @@ -522,14 +522,165 @@ you'll drag across, because a selection lives on the view whose marks define it. } ``` -## + a linked gap +## + filter by the brush -Append a second view that reads the selection. `filter: {param: brush}` keeps only the -brushed minutes; then it pivots A and B into columns, re-accumulates each, and plots -A − B. Now dragging the band on the top chart rebases this one — brush the first half-hour -and the early gap shows; brush the tail and it flattens. "Is the imbalance localised in -time?" gets answered by dragging instead of re-querying. With nothing brushed the selection -is empty, which by default means _every_ row, so it opens on the full run. +Append a second view that reads the selection — `filter: {param: brush}` is the entire +link. Everything downstream of that filter re-runs over only the brushed rows, so this +copy of the cumulative lines rebases to whatever window you drag: brush the first +half-hour and the early race shows; brush the tail and both lines restart from zero +there. With nothing brushed the selection is empty, which by default matches _every_ +row — that's why it opens on the full run. + +```vega-lite +{ + "$schema": "https://vega.github.io/schema/vega-lite/v6.json", + "vconcat": [ + { + "data": { + "name": "ev" + }, + "transform": [ + { + "sort": [ + { + "field": "minute" + } + ], + "window": [ + { + "op": "sum", + "field": "n", + "as": "cum" + } + ], + "groupby": [ + "arm" + ], + "frame": [ + null, + 0 + ] + } + ], + "params": [ + { + "name": "brush", + "select": { + "type": "interval", + "encodings": [ + "x" + ] + } + } + ], + "mark": "line", + "encoding": { + "x": { + "field": "minute", + "type": "temporal", + "title": "Time (UTC)", + "axis": { + "format": "%H:%M" + } + }, + "y": { + "field": "cum", + "type": "quantitative", + "title": "Cumulative assignments" + }, + "color": { + "field": "arm", + "type": "nominal", + "scale": { + "domain": [ + "A", + "B" + ], + "range": [ + "#3a7ca5", + "#e0913a" + ] + }, + "title": "Branch" + } + } + }, + { + "data": { + "name": "ev" + }, + "transform": [ + { + "filter": { + "param": "brush" + } + }, + { + "sort": [ + { + "field": "minute" + } + ], + "window": [ + { + "op": "sum", + "field": "n", + "as": "cum" + } + ], + "groupby": [ + "arm" + ], + "frame": [ + null, + 0 + ] + } + ], + "mark": "line", + "encoding": { + "x": { + "field": "minute", + "type": "temporal", + "title": "Time (UTC)", + "axis": { + "format": "%H:%M" + } + }, + "y": { + "field": "cum", + "type": "quantitative", + "title": "Rebased in the window" + }, + "color": { + "field": "arm", + "type": "nominal", + "scale": { + "domain": [ + "A", + "B" + ], + "range": [ + "#3a7ca5", + "#e0913a" + ] + }, + "legend": null + } + } + } + ] +} +``` + +## + the gap math + +Two rebased lines still make you subtract by eye. To plot A − B directly, the rows have +to change _shape_: `pivot` turns the A and B rows into columns of one row per minute, +two `calculate`s patch the minutes where an arm is missing, and a two-op `window` +re-accumulates each column before the final subtraction. Four transforms in service of +one line — reshaping long rows into wide columns is a craft of its own, but the link to +the brush is unchanged: the same `filter` still heads the pipeline. ```vega-lite { @@ -1183,3 +1334,7 @@ brush}` inherits Vega-Lite's default `empty: "all"`, so until you drag something matches every row — which is why the dashboard opens on the full run. To make it open blank instead, set `empty: "none"` on the selection. ::: + +Take it further: open the last stage in Astrolabe (the link under the chart) and try both +traps for yourself — move the `params` block onto the gap view and feel the dashboard go +dead, then put it back and set `empty: "none"` to see the blank-until-brushed variant.