Practice — Hot · Solution
1 · Q1, answered by the grid
The heat grid runs 2.44% to 3.82% — a 1.4-point band with no outlier. The honest answer: no region has a returns problem, and the quality team should not be dispatched on this data. Furniture–Midwest tops the grid at 3.82%, which is a fact worth one sentence and zero plane tickets: it is 0.6 points above the company average, not a fire. (The stretched scale from Medium could make it LOOK like a fire — which is exactly why you built and deleted that version.)
2–3 · Q2, split into its answerable and unanswerable halves
What the data shows: among current states, revenue concentration — CA ($2.91M), NY, TX lead; IN trails ($0.93M); 23 states have stores, 27 have none and therefore no retail data at all. A map of this answers "where are we strong?" — not "where should we go?"
The schema investigation (task 3): customers carry a State column — and online orders carry a CustomerID. So the model can show online revenue by the customer's state, including states with no stores: relate through customers, and suddenly the 27 dark states light up with whatever online demand exists there. This is a genuine, buildable analysis — online revenue by customer state as a demand proxy — and finding it is the exercise's payoff. Its limits still deserve a sentence: customer state ≈ shipping demand, a proxy, not a market study.
What "underserved" still requires that no table contains: population and income by state, competitor presence, retail cost structures. The dataset can rank our presence and approximate our online demand; it cannot see the market.
4 · A briefing that would hold
"On returns: no region has a problem — rates run a narrow 2.4% to 3.8% band with no outlier, and I'd hold the quality team until a cell actually moves. On store siting: the retail map ranks where we're strong today — California, New York, Texas lead — but it's silent on 27 states where we have no stores. The better instrument is one we can build from this model: online revenue by the customer's home state, which shows demand arriving from states we don't serve in person. I'd use that as the shortlist-maker, then buy the two things this data can't see — population-adjusted market size and competitor presence — before committing either store."
Reflection answers
- The tip-off was the word "underserved" — a claim about demand we are NOT capturing, which by definition is thin in our own sales data. Questions containing a counterfactual ("where would we succeed?") always deserve a data-inventory pass before a chart.
- A defensible ranking: the missing-data list first (it prevents a two-store mistake), the online-demand-by-state analysis second (a genuinely new instrument this briefing invented), the no-problem grid third (valuable, but it mostly prevents a smaller mistake). Reasonable people can swap the first two; the grid last is firm — and note that all three deliverables are sentences about scope as much as charts.