The finder’s methodology

How we weigh risk

The setup finder scores six concerns. Four of them you walk through — each starts at a research-based default and moves with facts about your situation. This page is the homework behind it — where every default comes from, what shifts it, which numbers are estimates rather than research, what we refuse to ask, and where we’re honestly uncertain. No other part of the recommendation is hand-waved either, but this is the part worth publishing.

Our starting bias, stated up front

For money whose loss would genuinely hurt, we will not leave you on a single point of failure — one hardware wallet, one seed, no passphrase, nothing else that has to also be true. Above the smallest stakes, the finder’s recommendation always clears that bar. That is a stance, not a calculation, and the Coldcard seed-generation flaw is why it is stated this plainly: every wallet drained in that incident was a lone key with no second factor, and the people who came through it untouched were the ones with a passphrase, their own dice throws, or a key from a second maker. One thing failed, and for some holders that one thing was everything.

And the reason it is a stance rather than a number is the premise this whole guide rests on: security is something you learn and practise, and there is no product that does it for you. The failures that actually reach people are the ones nobody modelled — the Coldcard flaw above was invisible to every buyer and to us — and the only general defence against a threat you did not imagine is not owning a component whose failure takes everything. So the question we ask of a recommendation is not is this safe now but is this still standing in ten years, operated by a real person having a bad week. Everything else on this page is arithmetic; this is what the arithmetic is in service of.

If losing your Bitcoin genuinely wouldn’t change your life, one hardware wallet kept properly is still an honest answer, and the finder will still say so. Everywhere above that, levelling up and removing the single point of failure is our recommendation — and it is checked on every path the engine can produce, not just asserted here.

The standard picture — the typical holder’s four researched risks

Each bar is a concern’s share of expected loss for a typical holder — someone with meaningful but not life-changing money at stake, mixed custody, no public profile. Not a probability, a share: “of the ways this is likely to go wrong, how much of the danger lives here.” Beside each bar is the research band we publish with it — the honest range, not just the point.

Scams and remote theft band 30–50

Phishing, fake support and fake apps, SIM swaps, malware, address poisoning, long-con investment scams — theft that never touches your door.

Locking yourself out band 25–45

Lost or untested backups, a forgotten passphrase, no way for anyone to recover if something happens to you, a botched migration.

A company failing you band 10–30

A meaningful share of your Bitcoin sitting with an exchange or custodian that goes under, freezes your account, or loses your coins — the Mt. Gox to FTX class of loss. Small change on an app is not what this is about.

Targeted physical theft band 1–9

Someone coming after you specifically — coercion, burglary for devices and backups, insider theft by people who know you.

Where each default comes from

Four anchors, four very different kinds of evidence — a reported-loss flow, a lost-coin stock, a corporate failure history, and an incident registry. That mismatch is exactly why the bands above are wide (more on that below).

Scams and remote theft — the largest measured flow

The biggest bucket, because it is the biggest measured number in the field. The FBI’s Internet Crime Complaint Center logged $11.4 billion in reported US cryptocurrency-fraud losses across 181,565 complaints in its latest annual report, and Chainalysis’s blockchain-side measurement puts global crypto scam revenue around $17 billion a year. Both are counts of money that actually moved — not projections — and both are understatements, since most victims never report.

Sources: FBI IC3 annual report ↗ · Chainalysis crypto crime reporting ↗

Locking yourself out — the largest all-time stock

Estimates of Bitcoin lost forever cluster between 11% and 23% of everything ever mined — Chainalysis’s classic dormancy study put it at 17–23% of then-supply; Ledger’s 2025 restatement lands at 11–18%. One honest footnote: River’s deliberately conservative 2025 estimate (~1.57 million BTC, about 7.5%) sits below that range — and 98% of the losses River counts happened before 2020, which is real evidence that modern wallets and steel backups have cut the flow. On the people side, surveys put a lockout — a forgotten password, a lost backup — in the history of roughly 35–40% of holders, and the oldest recovery firm succeeds only about a third of the time.

Sources: River — what happens to lost Bitcoin (incl. the conservative estimate) ↗

A company failing you — the longest track record

This is the best-documented failure class in Bitcoin’s history. The foundational academic study of exchange failure, from 2013, already found a 45% failure rate with a median exchange lifetime of 381 days — and the pattern held: of all the exchanges ever launched, roughly six in ten have closed, from Mt. Gox through the 2022 bankruptcies to FTX. Cumulative customer losses run to tens of billions of dollars. It sits third, not first, for one reason: it is the only risk you can take to zero in an afternoon, by withdrawing — which is why the finder’s answer to it is the checklist, not a bigger bar.

Sources: Moore & Christin, Financial Cryptography 2013 — exchange-failure study ↗

Targeted physical theft — rare, documented, and deliberately over-weighted

The public registry of physical bitcoin attacks — maintained by Jameson Lopp — documents on the order of 344 incidents worldwide across twelve years, with a sharp rise recently. Against tens of millions of holders that is a baseline on the order of one in a hundred thousand per year, and our default sits above that raw rate on purpose: it is a severity premium, stated openly, because an attack at your door is catastrophic in a way a frozen account is not. The academic study of these attacks (AFT 2024) adds the finding the whole section is built on: victim selection is essentially never random — it starts with someone knowing.

Sources: Lopp’s physical-attack registry ↗ · Ordekian et al., “Investigating Wrench Attacks,” AFT 2024 ↗

The other two concerns — scored, but not from base rates

The four above are the ones with published research behind them, and they are the four you walk through. The engine scores six things in total. The other two come from questions you answer in plain words, and they are scored exactly like the rest rather than applied as a thumb on the scale afterwards:

Neither has a research band, because neither is a base rate — they are facts about you that only you can supply, and the finder asks for them directly.

The typical holder’s bundle — and this is the estimated part

An untouched bar has to mean something. The obvious choice — start everyone at zero and let the questions add up — would quietly assert that the typical holder has none of these problems, which is false: it would make a section you walked and cleared look identical to one you skipped, and those are opposite signals. So the baseline is a model of what a typical holder actually carries. Score above it and your bar rises; below it and it falls; land on it exactly and you sit on the published default, which is what “typical” is supposed to mean.

Read these as estimates, because that is what they are

Of the 29 statements the assessment can put to you, three carry a published figure for how common that situation is among holders. The other 26 are our own estimates. They are the softest numbers in this engine — softer than the bands above, which at least have papers behind them — and we would rather publish them and be argued with than keep them in the code where nobody can check them.

A company failing you 22 of 48 available
Locking yourself out 42 of 88 available
Scams and remote theft 36 of 108 available
Targeted physical theft 22 of 84 available

Points, not a count of statements — each one carries a small, medium or large weight, so “three of eight” describes different people depending on which three. Locking yourself out is deliberately the largest share, because it is the thing holders are genuinely worst at: most have never test-restored a backup, and nearly nine in ten have told nobody. Targeted physical theft is the smallest, matching its small share above — most holders are known to somebody, but very few are targets.

What moves the bars

Your stakes shift the starting point. The finder asks how much losing this Bitcoin would hurt — never how much it is. If losing it wouldn’t change your life, the company-failure share starts larger (small holders are the ones most likely to still be entirely on an exchange) and targeted physical theft starts near nothing. If losing it would be devastating, targeted theft starts meaningfully higher and company failure lower — big holders are mostly off exchanges already, and perceived worth is what attracts targeting.

Then your situation moves each bar individually. The assessment is 29 plain yes/no statements — “true of me?” — and every single one carries its evidence in the same breath, so you can see why checking it moves the bar. Checking everything in a section pushes that bar high, but never to the top of the scale: total certainty about the future is not something this page gets to claim, and neither does the bar.

Low, typical, elevated, high — and why there are no numbers

Your result is always one of four words per concern. Typical means you sit inside the research band above — you are the base rate. Low means below it, elevated meaningfully above it, high well above it. That’s the whole scale.

We don’t show you a score, and that is a decision, not a limitation. The sources behind these defaults measure different things, over different periods, with different denominators — a “your remote-theft risk is 43.7” would be false precision, dressing a judgment call up as a measurement. The words carry exactly as much confidence as the evidence does. No more.

The questions we deliberately don’t ask

Most risk quizzes ask some version of these. We looked for the evidence and didn’t find it — so the assessment doesn’t ask, and this list stays published as part of the method.

The concern we tested and did not add

One risk on this list nearly became a seventh row: the device you bought is quietly broken and you cannot tell. It is not hypothetical — a hardware wallet we rate shipped a seed generator that produced predictable seeds, and the wallets drained because of it are the reason for our standing advisory. It is a real way to lose Bitcoin and it belongs somewhere on this site.

So before writing a word of it we ran it through the engine as if it were live, across every combination of answers the finder can produce, to see what it would actually change. The answer was the opposite of what we expected:

What we did instead is the part that matters. The defence against a broken device is not a different rung — it is not having a component whose failure is total, which is already the floor we hold you to. What it really argues for is a property of how you build your setup: keys from more than one maker, and randomness that did not come only from the device. Those are decisions you make on your plan and your checklist, so that is where they now bite — including a plan that is fully filled in and still has one brand holding enough keys to spend on its own. A risk score would have been a worse version of a thing we can simply check.

We are publishing this because a finding that says “don’t build it” is worth as much as one that says build it, and because a method you cannot see us change our minds inside is not a method.

Where we’re uncertain — on purpose, out loud

Honesty about the seams: dormant coins aren’t proof of loss — the lost-coin estimates rest on heuristics about coins that haven’t moved, and some of those owners are just patient. Recovery firms publish no case-type breakdowns, so “forgotten passphrases dominate their caseload” is a qualitative claim from the firms, not a percentage. The physical-attack registry counts only media-reported cases — in one academic sample, two of eleven interviewed victims had reported at all, so the true count is plausibly several times higher. And no academic study splits expected loss across these four buckets — we composed it from flow data, stock estimates, and an incident registry with different denominators. The typical holder’s bundle above is our estimate, not a measurement, and it is the number here we would most expect to be wrong. That is exactly why the bands are wide, and why the finder speaks in words instead of decimals.

The headline statistics behind these defaults — the FBI’s IC3 annual report, Chainalysis’s crime reporting, and the physical-attack registry — are on our freshness watch and re-checked on a schedule, like the prices and claims everywhere else on this guide. The verified date at the bottom of this page is real, and it moves when the evidence does.

Now put it to work

The defaults are everyone’s picture; the finder’s job is to find out where yours differs — walk your own risk assessment. And if you’d rather read the base rates as a lesson first, they’re the same ones taught in how people actually lose Bitcoin.

Last verified: August 4, 2026