Bet
Chapter 01 — Premise
Product teams remember what they shipped.
Not what they believed.
Bet treats a decision as a claim: stated with confidence, then checked when reality answers.
Ticket
Move checkout to a single page
chosen over improving the existing flow
Spec doc
Checkout v2
multi-step assumed to be the problem
Message
shipping this sprint
confidence — high? never written down
Retro
Q3 review
looked back after the outcome
confidence & resolution
— nothing here at all
Chapter 02 — What the stack loses
Reasoning is the thing the stack throws away.
Trackers keep the what. Docs keep it once. Chat forgets by Friday. Retros remember it after the outcome.
Tracker
Keeps
the what
Loses
the why
Document
Keeps
it once
Loses
whether it held
Chat
Keeps
the moment
Loses
it by Friday
Retro
Keeps
it eventually
Loses
the time to act on it
What Bet keeps
Chapter 03 — Capture & draft
Start from what you already wrote.
Paste the note. Bet drafts the claim, assumption, and alternative. A draft, not a verdict — the structure is suggested, the commitment is yours.
We’re moving checkout to a single page. Multi-step is where we lose people — a single page should lift completion. Shipping this sprint.
quoted from you
Claim
Single-page checkout will raise completion rate.
Rests on
multi-step is the main driver of drop-off
InferredInstead of
improving the existing multi-step flow
InferredConfidence
& resolution
left blank — yours to set
Chapter 04 — Commit
Now it becomes a claim.
Your confidence. Your criterion. Set before the result is in.
Confidence · four fixed bands
Resolves when · set now
Metric
overall checkout completion rate
Threshold
rises ≥ 3 pts absolute vs. prior 4 weeks
Window
4 weeks after launch
Chapter 05 — Reveal
The insight is not generated. It is revealed.
One assumption is carrying the weight. And a related bet from March points toward a different explanation.
One assumption is carrying the weight
Single-page checkout raises completion
multi-step drives the drop-off
most of the payoffimproving the multi-step flow
A related bet from March
Payment-method friction drives checkout loss
vs.same
drop-off
Multi-step friction drives checkout loss
These bets place weight on different explanations. They may apply to different segments — or one may be overstated.
Chapter 06 — Settlement
Measured against what you committed to.
+1.4, not +3. Not met. The small-cart result is a new finding — not a rescued bet.
completion ≥ +3 pts / 4 wks
+1.4 pts overall
+1.4 fell short of the +3 threshold
Evidence quality · separate axis
Additional finding does not change the result
+3.1 pts on small carts, flat on large.
Hindsight may add to the record. It does not rewrite it.
Chapter 07 — Record, restraint & the hinge
Not enough yet to say. So it does not.
Probable2 met1 not
Confident3 met1 not
Not enough comparable bets yet to say whether your confidence is well-calibrated — this is a record, not an assessment.
The screens show what Bet does. The full case explains why it is shaped this way.
The written case
The Reasoning Behind Bet
The screens above show what the product does. This explains what they cannot: why it is shaped this way, and what remains unproven.
From journal to reasoning layer
Bet began as a “decision journal.” The deeper I took the concept, the less convincing that container became. The effort is front-loaded, the payoff distant and probabilistic, and the core promise is difficult to validate, since decision quality and outcome quality are only loosely coupled. It’s the profile of a tool people admire and quietly stop using.
The durable idea survived the container. Product organisations keep excellent records of what they did, and almost no memory of why they believed it was right. Most of the product stack records actions, artifacts, and outcomes. Very little preserves the reasoning behind a decision in a form that can be revisited when reality answers. Bet is an attempt at that instrument.
Choosing the object
Everything turns on the unit. I tested several — hypothesis, assumption, forecast, decision — and each missed something: too formal to speak aloud, too passive, too binary. “Bet” was the only one that combined belief, commitment, and accountability in language product teams already use. That combination matters practically: you can’t measure whether someone’s judgment is well-calibrated using objects that were never actually risked on.
A bet carries a claim, the confidence behind it, the assumption it rests on, the alternative it passed over, and the part most decisions never get — a resolution condition, set in advance.
Three product decisions shaped the concept more than any visual choice
It never sets your confidence.
The system can read a claim from your text and infer the assumption underneath it — but confidence is left blank for you to own. This adds friction, and I kept it: a calibration score is meaningless if the confidence being scored was the model’s guess rather than yours.
Confidence is structured, not free-form.
I used a small fixed scale rather than free-entry percentages, choosing consistency over precision — a “73%” you can’t actually feel isn’t more honest than a band that means the same thing every time. This is one of the assumptions I’d test first with real users; the specific scale is a reasoned bet, not a proven one.
Where the criterion is explicit, settlement is anchored to it.
The bet resolves against the threshold set in advance, rather than against a retrospective story. In the checkout example, the redesign lifted completion by 1.4 points against a 3-point criterion: not met, cleanly, even though it lifted small-cart completion sharply. That small-cart result is real, but it’s a new finding — the seed of a new bet — not grounds to relabel a failed one as a partial win. Letting the flattering subset soften the verdict would reintroduce exactly the bias the product exists to remove.
Where AI helps — and where it stops
The AI reads structure out of prose written for other purposes, and it can propose possible links between a current bet and earlier ones, with its reasoning visible so you can confirm or reject the connection. That kind of cross-time recall becomes increasingly difficult to maintain by hand as decisions accumulate and language changes.
Its role is bounded on purpose. It doesn’t set your confidence. It may calculate whether an explicit threshold was met, but it doesn’t interpret ambiguous evidence or redefine success after the fact. And it proposes tensions for you to judge — it doesn’t assert them. A product about calibration has no business being miscalibrated about its own conclusions, which is why the record view, with only a handful of settled bets, says plainly that there isn’t yet enough to say.
What remains unproven
Four things I haven’t validated. Whether the fixed confidence scale matches how people actually hold uncertainty. Whether the tension-detection stays trustworthy as history grows, or starts surfacing coincidences that read as insight. The cold-start problem — the structural value works from the first bet, but the cross-bet value only arrives with volume. And adoption: whether extraction genuinely lowers the capture cost enough to beat the friction that killed decision journals in the first place.