How Much Data Do You Need Before Trusting a Result?
The Count That Matters Is Rarely the One You Have
Short answer
The range runs from one observation to several hundred, depending on the question. Confirming that a competitor changed their pricing page needs one. Publishing a per-account benchmark needs at least twenty contributing accounts. Comparing two conversion rates needs hundreds of conversions in each group.
1. How much data do I need before trusting a result?
It depends entirely on the question type. Presence or absence of an event needs one observation. A benchmark needs at least twenty contributing accounts. A rate comparison needs hundreds of events per group.
- Sample size
- The number of independent observations a figure is computed from, which sets how much the figure would move if you collected a different set of observations of the same kind.
| Question | What it takes | Why |
|---|---|---|
| Did a competitor change their pricing page? | One observation each side of the change | Presence or absence is not a statistical question. It either happened or it did not |
| Is a competitor posting more than last quarter? | Two complete quarters, every post captured | If you captured every post, this is a count of the whole population rather than a sample |
| Which of our posts performed best? | Around thirty posts in the period | Ranking the top few is reliable long before any average is, because order survives noise better than level |
| What is a typical value for accounts like these? | At least twenty contributing accounts | A mean over a handful of accounts can move tens of percent when one more account is added |
| Did version B convert better than version A? | Hundreds of conversions per variant | A gap of one or two percentage points sits well inside the range chance produces at small counts |
| Is engagement declining, or is this seasonal? | Two complete annual cycles | One cycle produces a hypothesis. Separating a season from a trend requires seeing the season repeat |
2. Why is the number of accounts in my sample not the number that matters?
Because every metric is computed only from the accounts that actually report it. We hold data on 113 tracked competitors and still could not publish an engagement rate, because far too few of them had a usable follower denominator.
This is the most common way a sample size gets overstated, and it happens inside the reporting rather than in the collection. The headline number is the size of the tracked set. The number that governs whether a figure is publishable is how many accounts contributed to that specific figure, and the two can differ by a factor of several.
| Figure | Accounts contributing | Published? |
|---|---|---|
| Facebook posts per week | 77 | Yes |
| Instagram posts per week | 71 | Yes |
| LinkedIn posts per week | 56 | Yes |
| Facebook likes per post | 54 | Yes |
| Instagram likes per post | 51 | Yes |
| LinkedIn likes per post | 33 | Yes |
| Engagement rate, any platform | Fewer than 20 on every platform | No, suppressed |
| Any X figure | One account for the whole month | No, suppressed |
Set the floor before you look at the numbers, and write it down. A threshold chosen after seeing the result is a threshold chosen to include the result, and that is a decision nobody can audit afterwards, including you.
3. How many observations does a best-time-to-post finding need?
Far more than a single account produces. Seven days crossed with four time bands is 28 cells, and each cell needs several posts before the comparison between cells means anything at all.
| Grid | Cells | Posts needed | Time at two posts per week |
|---|---|---|---|
| Day of week only | 7 | 35 | About 18 weeks |
| Day of week by morning or afternoon | 14 | 70 | About 35 weeks |
| Day of week by four time bands | 28 | 140 | About 70 weeks |
This is why single-account posting-time analysis produces a different answer every quarter. With one or two posts per cell, the best cell is whichever one happened to contain a post that did well, and it moves as soon as new data arrives. Nothing is wrong with the arithmetic; there is simply nothing in it.
Three ways out, in order of how much they cost:
- Collapse the grid. Three bands and weekday versus weekend is six cells, which a single account can fill in a quarter.
- Pool across your tracked set. Twenty accounts on one platform fill a 28-cell grid in weeks, and the finding is about the platform rather than about you.
- Ask a different question. "Does posting on weekends do worse for us" needs two cells and answers most of what the grid was for.
4. How do I know whether an A/B test result is real?
Ask how large the gap is relative to the counts behind it, not which number is bigger. Five conversions out of 100 against eight out of 100 is a difference that random assignment produces constantly.
| Visitors per variant | A converts | B converts | What it supports |
|---|---|---|---|
| 100 | 5 | 8 | Nothing. A three-conversion gap at this size turns up constantly by chance alone |
| 1,000 | 50 | 80 | A real difference, if the test ran to a sample size fixed in advance |
| 1,000 | 50 | 56 | Nothing yet. The gap is small relative to the variation normal at this size |
Two practical rules follow. Fix the sample size and the end date before starting, because a test stopped the moment it looks good will look good roughly whenever noise is in your favour. And detecting a smaller difference costs disproportionately more data: roughly four times the sample to halve the difference you can resolve.
For a team with modest traffic, the honest conclusion is often that a proper test is out of reach for small effects, and that a large change tested once is better use of the traffic than three small changes tested badly.
5. Which findings need far less data than people assume?
Anything factual rather than statistical. Whether an event happened, the order of the top few items, counts over a complete population, and a single counter-example to a universal claim all need very little data.
Five results a small team can act on immediately:
- Did it happen. A competitor changed their pricing, shipped a page, entered a market. One before and one after is the whole evidence base, and no amount of extra observation improves it.
- The order of the top few. Which three posts did best is stable well before any average is, because ranking only needs the gaps to be large, not the levels to be precise.
- Counts over a complete population. If you captured every post a competitor published, the post count is not an estimate at all, so no sample-size question arises.
- The direction of a large change. A halving or a doubling is outside the range small-sample noise produces, so it is readable long before a ten percent move is.
- A single disqualifying case. One counter-example is enough to retire a universal claim, which is why one lost deal to a feature gap deserves attention that one won deal does not.
Recognising these saves a surprising amount of time, because they are exactly the questions competitive monitoring is good at. Public competitor data is thin for averages and excellent for events, which is a reason to build the programme around events.
6. What should I do with a result that falls under the floor?
Keep it, label it an observation with the count attached, and let it prioritise work rather than decide it. Suppressing a thin result from reports is different from deleting the underlying data.
An observation under the floor still has uses. It tells you where to look next, it justifies collecting more, and it is often enough to rule something out even when it cannot rule anything in. What it cannot do is appear as a bare number in a report, because a number without its count travels further than the caveat attached to it.
Four rules for a thin result:
- Write the count next to it every single time, in the same sentence rather than in a footnote.
- Call it an observation rather than a benchmark, in exactly those words.
- Use it to choose what to investigate, never as the sole input to a spend or positioning decision.
- Keep collecting on the same definition, so that in two quarters it either crosses the floor or clearly will not.
7. How do I report a number so a reader can judge it?
Print the count, the basis and the period next to the figure. A reader who can see what a number was computed from can decide how much weight to give it without asking you.
The format that works is short: the value, then how it was measured, then how many observations, then the date. "1.04 posts per week, unweighted mean across 56 tracked LinkedIn accounts, June 2026" carries everything a reader needs to argue with it, which is the point.
Also state what was excluded and why. Suppressed cells are more informative than they look, because they tell the reader which questions the data cannot answer yet, and that stops the same question being asked every month.
Oppira records the contributing account count behind every aggregate it computes, which is what makes the difference between a benchmark and an observation something the tool can enforce rather than something a person has to remember.
Key Takeaways
The requirement varies by question, not by team
One observation for whether an event happened, twenty contributing accounts for a benchmark, hundreds of conversions per variant for a rate comparison.
Count the contributors, not the sample
We track 113 competitors and could not publish an engagement rate, because the accounts with a usable follower denominator were far under our floor.
Grids eat data
A seven-day by four-band posting grid is 28 cells, needing about 140 posts at five per cell, which is well over a year for one account.
Percentages hide sample size
Five out of 100 versus eight out of 100 supports nothing. The same split at 1,000 each supports a decision. Always print the raw counts.
Events need almost no data
Whether a competitor changed their pricing, the order of the top three posts, and any halving or doubling are all readable immediately.
Set the floor before you look
A threshold chosen after seeing the result is a threshold chosen to include the result. Write the number down first, then compute.
Frequently Asked Questions
Explore More
Related analyses, benchmarks, and industry insights
Related Guides
Glossary Terms
Turn reading into a reaction
Oppira watches your competitors, keeps your playbook current, and drafts the response. Start free and see your market clearly by tomorrow.