The Source Ledger: The Content System That Stops You Publishing Numbers You Can't Defend

Yesterday we wrote that the cheapest pre-publish check is tracing every number in a post rather than remembering it. That check takes ten seconds to write down and it is the one founders fail most often, because tracing a number at 11pm means opening a tab, finding the article you half-remember, discovering it cites another article, and then deciding whether to publish anyway.

Most founders publish anyway.

The fix is not more discipline at the end. It's a source ledger — a single file where every external number you might cite lives with its origin attached, built while you read instead of while you write. It is the least glamorous content system we run and it prevents the only content failure that costs you a reader permanently.

Three numbers that became facts by repetition

These are not hypotheticals. They're claims we've watched circulate through the Amazon and LinkedIn operator world, and in each case the number was checkable and wrong.

The invented fee schedule. In July, a wave of seller content described a new $0.38 surcharge on sub-$15 items and a 12-21% fee increase, complete with effective dates. Checking Amazon's own published fee page turned up a January change averaging eight cents and explicit language that no new fee types were being introduced. Somebody had invented a fee schedule. Operators were repricing products off a blog post.

The zoom stat. "Image zoom lifts conversion up to 20%" and "one in five shoppers zooms" both circulate constantly in Amazon creative content. Neither has a study behind it that anyone can produce. They're plausible, they're directionally reasonable, and they have no origin.

The misattributed table. A widely shared table of LinkedIn edit penalties — edit within ten minutes, lose 3%; edit after ninety, lose 40-50% — cites Richard van der Blom's Algorithm Insights report as its foundation. That report is real, large, and genuinely useful. It contains no finding about edited posts at all. A real researcher's name got attached to a claim he never made, and downstream repetition gave it the texture of a fact.

The pattern in all three: nobody lied. A number got written down without a source, repeated by someone who assumed the first writer had checked, and by the fourth repetition it read like consensus.

Why founders are the most exposed

The operator content that works is specific. Numbers are the specificity. So you cite more numbers than a generalist would, which means you carry more exposure per post than a generalist does.

Three things make it worse right now:

The fastest source is a search result, and search results in ecommerce are increasingly written by people summarizing other people. The top ten results for most seller statistics are ten articles citing each other.

AI drafting compounds it. A model asked for supporting statistics will confidently return numbers that were common in its training data, which is exactly the population of repeated-until-true claims. It's not hallucinating. It's reflecting the same broken chain back to you.

Your audience contains people who can check. That's the whole point of your positioning. The reader you want most — the operator who could hire you — is also the one most likely to know that fee page. When they catch it, they don't comment. They quietly reclassify you as someone who repeats things.

What goes in the ledger

One row per external number. Six fields, and one of them does most of the work.

  • The claim, written the way you'd say it in a post. Not the raw stat — the sentence. You'll paste this straight into drafts.
  • The number.
  • The source, and whether it's primary or secondary. This is the column that matters.
  • The date of the data, not the date of the article.
  • The link.
  • A re-check date.

Primary versus secondary is the field that changes what you publish. A source is primary when the entity that would actually know published it. Amazon's fee page is primary. A blog about Amazon's fee page is secondary. A platform's own announcement is primary. A tool vendor's report is primary for their own data and secondary for everything else in it. A study is primary when you can see the sample size and the method; a study that gives you only a headline percentage is a press release wearing a lab coat.

Most numbers in circulation are secondary and most of them will hold up. The point isn't to ban them. The point is that you know which is which before you're on the hook for it, and you write about them differently.

Date of the data, not the article catches the second most common failure. Articles get updated, republished, and re-dated. The study underneath is from 2023 and describes a platform that has since changed twice. If you can't find the data date, that's itself a finding — write "date unknown" and treat the entry as fragile.

If you want a seventh field, add who benefits if this is true. A tool vendor's study of the feature their tool optimizes is not disqualified — it's frequently the only data anyone has collected. But it changes how you frame it, and framing it honestly is free.

What to do when a number won't verify

Don't kill the post. Demote the claim. There are three moves and all of them are stronger than they feel.

Cite the mechanism instead of the magnitude. You don't need a percentage to say that shoppers who zoom have already decided they're interested, or that a shorter title forces a choice about which differentiator survives. The mechanism is the part you actually know. The number was decoration you were using as proof.

Attribute the uncertainty out loud. "A figure that circulates constantly in this space, which I've never been able to source" costs you nothing and buys you the reputation of someone who checks. We've watched posts built on that exact sentence outperform the posts that quoted the number confidently, because a reader who has also seen that stat now knows something about you.

Replace it with your own observation. You have proprietary numbers. Not as clean, not as quotable, and considerably harder to argue with. The founder who says "across the accounts I've worked in, it usually looks like X" is making a weaker claim and a more credible one.

There's an inversion here worth naming: the number is what makes a post feel authoritative, and removing it is often what makes the post authoritative. Anyone can repeat a statistic. Saying "this one doesn't hold up and here's what I checked" is a thing only someone who looked can say.

The ledger pays twice

The first return is speed. A well-sourced number is reusable — you'll cite the same three or four across six months of content without re-verifying each time, and you'll do it with the specificity intact because the ledger holds the sentence, not just the digit.

The second return is bigger and most people miss it. Every entry you fail to verify is a post. The claim you couldn't source, the table with the misattributed researcher, the fee schedule nobody could find on the fee page — those are the highest-performing pieces of content available to an operator in a category full of confident repeaters. You're not being contrarian for sport. You went and looked, which almost nobody does, and the finding is the content.

The ledger is where those posts come from. Without it, you notice a shaky number, feel briefly annoyed, and forget it by Thursday.

How to actually run it

Ten minutes a week, and entries go in while you read, not while you write. That's the whole trick. A claim gets into the ledger before it gets into a draft, which means the verification happens when you're curious rather than when you're on deadline with a hook you already like.

Set the re-check cadence by claim type: platform mechanics every 6-12 months, fee and policy numbers whenever the platform announces a change, industry studies annually. When something structural changes on a channel you write about, open the ledger the same week — that's when half your entries expire at once, and it's also when the demand for content about it peaks.

If you're already running a proof bank, this sits next to it and does the opposite job. The proof bank holds your receipts. The source ledger holds everyone else's.

FAQ

Isn't this overkill for LinkedIn posts? The volume is low — most founders cite fewer than twenty external numbers a year. The cost of one bad one is a reader who never tells you they stopped believing you.

What if the only available source is a vendor's blog? Use it and say so. "According to a study by a tool vendor in the space" is an accurate sentence and readers are fine with it. What breaks trust is presenting vendor data as though it were the platform's own.

Do I have to cite sources in the post itself? No. The ledger is for you. Cite when the source strengthens the claim, attribute vaguely when it doesn't, and never state something flatly that you'd have to hedge if asked.

What about my own numbers? Different file. Your results, client outcomes, and observed patterns belong in the proof bank, with their own decisions about how much of the number to publish. The source ledger is strictly for numbers you didn't generate.


If you're publishing operator content and you've never traced the statistics you repeat, start the ledger with the three you use most. We build these for the founders we write for, because the fastest way to lose a technical audience is a number they can check and you can't.

Ready to turn your LinkedIn into a revenue channel?

We write operator-level content for e-commerce founders. No fluff. No generic posts. Just content that drives pipeline.

Book a Strategy Call