Open Question

How do we know when to trust our own judgement?

What I suspect

We usually know how certain we feel long before we know whether we’re correct. The two get treated as one signal. They aren’t.

What I can support

That confidence and accuracy diverge more than people expect, and not randomly. It tracks emotional state and expected reward more reliably than it tracks the evidence actually in front of someone.

What would change my mind

Someone who’s just reliably well-calibrated on their own – no structure, no check, no external prompt. So far the structure keeps outperforming the person, mine included.

The deepest investigation so far

Dissertation

2025 – present

Does seeing an AI recommendation change how well a professional’s confidence tracks their actual accuracy?

Why it matters. Every team shipping an AI feature makes one quiet decision without noticing it’s a decision: when the recommendation appears. Before the user has thought, or after. Most teams inherit the interaction before they question it.

What I’m testing. Whether the order changes the calibration, not just the answer. Show the recommendation before someone commits to their own confidence and you may be anchoring them. Show it after and you may be measuring something else entirely. The assumption baked into most tooling is that early is free. I don’t think early is free.

Where it stands. The effect looks real but also situational. It shows up where people are on unfamiliar ground and nearly vanishes where they already know the terrain. That’s either an interesting boundary condition or a sign I’m measuring “familiar” badly. I’m no longer sure which explanation is correct. I want to rerun it with a sharper definition of “unfamiliar” before I trust either reading.

  1. I assumed this was mostly a product-design problem. The more I read, the less convincing that explanation became.

  2. Kept circling back to calibration and eventually stopped treating it as a side detail. It might be the whole thing.

  3. Started looking for evidence that AI reliably harms calibration. The literature wouldn’t cooperate.

Other investigations

Draft

2026

What keeps a person moving toward a goal when no one is watching and nothing is due?

The neurocomputational models are good on the first few months and quiet on the fifth year. That’s the stretch I actually care about, and it’s the stretch nobody seems to have.

Reading

2026

Why does confidence rise and fall on its own, without new information?

Reward-prediction models currently seem to explain the swings better than the decision-making literature itself. I’m still not sure whether that’s because they’re genuinely more explanatory, or because they’re quietly answering a different question.

Complete

2026

Is temperament something you’re handed, or something that gets negotiated?

I went in expecting handed, then shaped a bit. The differential-susceptibility and transactional work has me somewhere less tidy – closer to a thing that keeps getting re-negotiated between a child and an environment that are both reacting to each other. I still don’t know whether that negotiation ever really settles, or whether it’s supposed to keep going.

Complete

2026

Does early life stress leave a mark you can still read in the adult wiring?

The behavioural-neuroscience side points to functional connectivity, and adult reactivity downstream of it. I find the evidence persuasive and I don’t fully trust how persuasive I find it. It’s the kind of story that’s easy to want to be true.

Partial

2025

Does making reasoning legible improve the decision, or just the record of it?

This is the one that surfaced in Bet, and I still can’t cleanly separate the two. A legible decision is easier to defend later. Whether it was a better decision is a different claim, and I keep catching myself sliding from one to the other.

Ideas that stayed

Gigerenzer

Made me much less sure that more information reliably makes a decision better. Some of the best-calibrated judgement comes from the crudest possible model, which is uncomfortable to sit with if you build tools for a living.

Kahneman

The one everyone reaches for. What actually stuck wasn’t that people are biased. It’s that I’m least reliable in the exact moment I feel most clear, and nothing on the inside flags it.

The confidence-calibration literature

Plainly the reason the dissertation changed shape. I went in with a design question and found a psychology one underneath it. It had been the real question the whole time.

Mandelbrot

Taught me that outcomes cluster far more in the extremes than a bell curve admits – markets, storms, failures that arrive all at once rather than gradually. I trust a smooth-looking chart a little less now.

Edges

Two people get the same recommendation and both walk away more certain, for opposite reasons

A good explanation can raise confidence without touching accuracy. It’s supposed to feel like understanding. Sometimes it just feels like understanding.

I don’t know whether confidence should be calibrated to being correct or to being useful, and I’ve stopped assuming those point the same way.