Bucketing and diagnosis: what actually happens to her answers

You asked whether we bucket people into the four problems correctly, whether the diagnosis follows from the questions, and whether the reporting screens match what she told us. I ran two opposite people through the whole quiz to find out.

The answer

Nothing is bucketed. There is no scoring. Every diagnosis screen is a fixed mockup showing one invented persona, identical no matter what anyone answers.

I walked the quiz twice — once answering as badly as the options allow, once as well as they allow — and compared every screen that reports back. They are byte-for-byte identical.

Reporting screenWeakest possible answersStrongest possible answersDifference
2.11 Your starting pointthe diagnostic reveal Mobility 28% · Balance 50% · Strength 62% · Energy 58% Mobility 28% · Balance 50% · Strength 62% · Energy 58%NONE
5.2 Your planpriorities + safety line Mobility first, then strengthMobility first, then strengthNONE
5.3 What changes firstweek 2/4/8/12 milestones identical 468 charactersidentical 468 charactersNONE
5.5 Paywallbefore/after on four measures identical 4,816 charactersidentical 4,816 charactersNONE

So the honest status is: the quiz collects the right things and routes correctly — that part is verified — but it does not yet compute anything. The persona on screen belongs to a woman whose mobility is poor, balance middling, strength and energy fine. Everyone currently gets her.

What the screens claim to know

This matters more than a missing feature, because six screens make first-person assertions about her. Until scoring exists, each one is a guess that is right for some people and wrong for the rest.

Two of them are safety claims

5.2 — “Working around your back: no floor work in weeks 1–2, and no loaded bending at any point. No jumping — we'll load your bones another way.”

Answer “None of these” on 2.10 and it still says it is working around your back. Answer “Never” on 2.8 — the answer that correctly skips the “we won't ask you to jump” interstitial — and the plan still says it is removing jumping. The routing already respects those answers; the plan text does not. A woman who told us she has no back problem is being told her plan is built around one.

ScreenWhat it assertsWrong when she…
1.6“You said you want to be stronger for everyday things picked any of the other four goals on 1.4 — 4 of 5 people
2.11“You said you want to be stronger… your answers say strength is holding up” picked another goal, or strength is in fact her weakest axis
2.11Verdict: “Mobility first, then strength is anyone whose weakest axis is not mobility
3.6“what your answers pointed to” + a profile-matched statistic has a different profile — the stat is fixed
5.2Three priorities: stiffness, stairs, single-leg balance reported none of those things
5.5“Because you mentioned stiffness” + before/after on four measures answered “Nowhere much” on 2.6

The mapping the spec already intends

The design is not vague about this — the spec annotates what each answer is supposed to drive, with weights. Collected in one place, it is a coherent scoring model that simply has not been built.

QuestionMobilityBalanceStrengthEnergyAlso drives
2.2 Up from the floor without hands●● ●●starting difficulty
2.3 Shopping up a flight of stairs ●●
2.4 Stand on one leg for 10 seconds●● supported vs free-standing work
2.5 By mid-afternoon ●●
2.6 Where stiff or achy●● region targeting · fires 2.7
2.1 How much you move in a weekmodifier across all four axes starting difficulty
1.3 Age bandshifts the benchmark each axis is graded against
Mobility2.2 ●● + 2.6 ●●4 points of signal
Energy2.3 ●● + 2.5 ●●4 points of signal
Balance2.4 ●● + 2.2 ●3 points of signal
Strength2.2 ●● + 2.3 ●3 points of signal

Four problems with that model, before anyone builds it

1. The most popular goal maps to no axis at all

1.4 offers five goals. Four map cleanly — steadier → Balance, stronger → Strength, without stiffness → Mobility, energy → Energy. The fifth, “Lose weight, especially around my middle”, maps to nothing we measure. You deliberately kept that option because it is what this audience actually searches for — so the funnel's most commercially important answer is the one the diagnosis cannot speak to. Decide now whether it routes to Strength (the honest mechanism for body composition) or gets its own reconciliation line.

2. “I've never tried” is not a low score

On 2.4, options are Yes either leg / A few seconds / Not really / I've never tried. The first three are a performance scale; the fourth is missing data. Scoring it as the worst answer would mark a woman who simply never attempted a one-leg stand as having the worst balance in the funnel, and send her down the balance track. It should be treated as unknown and imputed from 2.2, or prompt her to try it.

3. One question carries three of the four axes

2.2 — getting off the floor — feeds Mobility, Strength and Balance. It is a good clinical marker, but it means a single misread answer moves three quarters of the diagnosis at once, and it makes the axes correlate with each other rather than measure independently. Worth either adding a second strength item or lowering its balance weight.

4. Nothing distinguishes “can't” from “won't”

Both 2.2 and 2.3 end in avoidance answers — “I avoid getting down there”, “I avoid stairs with anything heavy”. Avoidance usually does indicate inability, so scoring it low is defensible, but it is a judgement rather than a measurement, and it is the answer most likely to come from fear rather than capacity. It is also the answer where getting the diagnosis wrong costs the most trust.

A scoring model you could approve

This is the smallest thing that would make the screens honest. Every axis scores 0–100.

  1. Score each answer 0–1 by position: best option 1.0, worst 0.0, evenly spaced between. 2.6 scores 1 − (regions ÷ 6), with “Nowhere much” = 1.0.
  2. Weight and combine per the table above, so each axis is a weighted mean of its sources × 100.
  3. Apply the activity modifier from 2.1: −8 / −3 / +3 / +8 points across all four axes.
  4. Apply the age allowance from 1.3: +0 / +3 / +6 / +9 / +12 as the band rises, so a 67-year-old is not graded against a 47-year-old's benchmark.
  5. Band into the four levels already used on screen — Needs attention 0–34, Room to build 35–54, Holding up 55–74, Going well 75–100.
  6. The verdict is the lowest axis, tie-broken Mobility → Balance → Strength → Energy. The headline becomes “{lowest} first, then {second lowest}”.

And keep the mechanic that makes this worth doing

2.11's copy already does something no competitor does: it lets the measurement disagree with the stated goal — “you said you want to be stronger; your answers say strength is holding up, it's stiffness making everything feel harder.” That sentence only lands when it is true. Wire it as three cases: goal matches the weakest axis (confirm it), goal is a different axis (reconcile, as the copy does now), or the goal is weight (no axis — say what you can honestly say instead).

What each screen needs before launch

ScreenMust become
1.6Goal echoed from 1.4; which of bone/muscle/tendon/cartilage to lead on
2.11Four scores + labels, the verdict headline, and the goal-vs-measurement reconciliation
3.6The statistic matched to the winning axis — falls, walking speed or body composition
5.2Three priorities from the two weakest axes + her stiffness regions; safety line assembled only from what she actually selected on 2.10 and 2.8
5.3Milestones ordered by axis rather than a fixed four
5.5Before/after rows built from her own 2.2–2.5 answers; “because you mentioned…” from 2.6

The one I would fix first

Not the scoring — the safety line on 5.2. Everything else is a personalisation that is merely generic until it is built. That line is the only place the quiz tells a woman something factually untrue about her own body, and it does it in the two areas where this audience is most sensitive: continence and back pain. It can be fixed today by assembling the sentence from her 2.8 and 2.10 answers, without any scoring model at all.