Updated July 29, 2026
10 min read
Analytics

How Much Data Do You Need Before Trusting a Result?

The Count That Matters Is Rarely the One You Have

Short answer

The range runs from one observation to several hundred, depending on the question. Confirming that a competitor changed their pricing page needs one. Publishing a per-account benchmark needs at least twenty contributing accounts. Comparing two conversion rates needs hundreds of conversions in each group.

GB
Written byGabor BartaCo-founder, Oppira

Gabor leads product and content at Oppira. He has spent over a decade building tools and writing about competitive intelligence, social media analytics, and growth marketing for B2B SaaS companies.

Published July 29, 2026

1. How much data do I need before trusting a result?

It depends entirely on the question type. Presence or absence of an event needs one observation. A benchmark needs at least twenty contributing accounts. A rate comparison needs hundreds of events per group.

Sample size
The number of independent observations a figure is computed from, which sets how much the figure would move if you collected a different set of observations of the same kind.
What each kind of marketing question actually requires
QuestionWhat it takesWhy
Did a competitor change their pricing page?One observation each side of the changePresence or absence is not a statistical question. It either happened or it did not
Is a competitor posting more than last quarter?Two complete quarters, every post capturedIf you captured every post, this is a count of the whole population rather than a sample
Which of our posts performed best?Around thirty posts in the periodRanking the top few is reliable long before any average is, because order survives noise better than level
What is a typical value for accounts like these?At least twenty contributing accountsA mean over a handful of accounts can move tens of percent when one more account is added
Did version B convert better than version A?Hundreds of conversions per variantA gap of one or two percentage points sits well inside the range chance produces at small counts
Is engagement declining, or is this seasonal?Two complete annual cyclesOne cycle produces a hypothesis. Separating a season from a trend requires seeing the season repeat
What each kind of marketing question actually requiresThe two extremes matter more than the middle. Teams routinely demand significance for questions that are simply factual, and act confidently on rate comparisons built from a few dozen events.

2. Why is the number of accounts in my sample not the number that matters?

Because every metric is computed only from the accounts that actually report it. We hold data on 113 tracked competitors and still could not publish an engagement rate, because far too few of them had a usable follower denominator.

This is the most common way a sample size gets overstated, and it happens inside the reporting rather than in the collection. The headline number is the size of the tracked set. The number that governs whether a figure is publishable is how many accounts contributed to that specific figure, and the two can differ by a factor of several.

Contributing accounts behind each figure in our June 2026 aggregate
FigureAccounts contributingPublished?
Facebook posts per week77Yes
Instagram posts per week71Yes
LinkedIn posts per week56Yes
Facebook likes per post54Yes
Instagram likes per post51Yes
LinkedIn likes per post33Yes
Engagement rate, any platformFewer than 20 on every platformNo, suppressed
Any X figureOne account for the whole monthNo, suppressed
Contributing accounts behind each figure in our June 2026 aggregate Oppira Benchmark, accounts contributing to each measured figure, as of June 30, 2026.One tracked set, one month, and contributing counts running from a single account to 77. Cadence is available for nearly every account because a post either exists or does not. Engagement rate needs a follower count sampled at the right time, which far fewer accounts had.

Set the floor before you look at the numbers, and write it down. A threshold chosen after seeing the result is a threshold chosen to include the result, and that is a decision nobody can audit afterwards, including you.

3. How many observations does a best-time-to-post finding need?

Far more than a single account produces. Seven days crossed with four time bands is 28 cells, and each cell needs several posts before the comparison between cells means anything at all.

Illustrative example: posts required to fill a posting-time grid at five posts per cell
GridCellsPosts neededTime at two posts per week
Day of week only735About 18 weeks
Day of week by morning or afternoon1470About 35 weeks
Day of week by four time bands28140About 70 weeks
Illustrative example: posts required to fill a posting-time grid at five posts per cellCell and post counts are arithmetic on an illustrative five posts per cell. In the Oppira Benchmark for June 2026, tracked accounts averaged 1.98 posts per week on Instagram and 1.04 on LinkedIn, so two per week is a generous assumption rather than a conservative one.

This is why single-account posting-time analysis produces a different answer every quarter. With one or two posts per cell, the best cell is whichever one happened to contain a post that did well, and it moves as soon as new data arrives. Nothing is wrong with the arithmetic; there is simply nothing in it.

Three ways out, in order of how much they cost:

  1. Collapse the grid. Three bands and weekday versus weekend is six cells, which a single account can fill in a quarter.
  2. Pool across your tracked set. Twenty accounts on one platform fill a 28-cell grid in weeks, and the finding is about the platform rather than about you.
  3. Ask a different question. "Does posting on weekends do worse for us" needs two cells and answers most of what the grid was for.

4. How do I know whether an A/B test result is real?

Ask how large the gap is relative to the counts behind it, not which number is bigger. Five conversions out of 100 against eight out of 100 is a difference that random assignment produces constantly.

Illustrative example: the same percentage gap at three different sample sizes
Visitors per variantA convertsB convertsWhat it supports
10058Nothing. A three-conversion gap at this size turns up constantly by chance alone
1,0005080A real difference, if the test ran to a sample size fixed in advance
1,0005056Nothing yet. The gap is small relative to the variation normal at this size
Illustrative example: the same percentage gap at three different sample sizesAll figures are illustrative. The first and second rows describe the identical five versus eight percent split, and only one of them supports a decision. Percentages hide sample size, which is why a test report should always print the raw counts.

Two practical rules follow. Fix the sample size and the end date before starting, because a test stopped the moment it looks good will look good roughly whenever noise is in your favour. And detecting a smaller difference costs disproportionately more data: roughly four times the sample to halve the difference you can resolve.

For a team with modest traffic, the honest conclusion is often that a proper test is out of reach for small effects, and that a large change tested once is better use of the traffic than three small changes tested badly.

5. Which findings need far less data than people assume?

Anything factual rather than statistical. Whether an event happened, the order of the top few items, counts over a complete population, and a single counter-example to a universal claim all need very little data.

Five results a small team can act on immediately:

  • Did it happen. A competitor changed their pricing, shipped a page, entered a market. One before and one after is the whole evidence base, and no amount of extra observation improves it.
  • The order of the top few. Which three posts did best is stable well before any average is, because ranking only needs the gaps to be large, not the levels to be precise.
  • Counts over a complete population. If you captured every post a competitor published, the post count is not an estimate at all, so no sample-size question arises.
  • The direction of a large change. A halving or a doubling is outside the range small-sample noise produces, so it is readable long before a ten percent move is.
  • A single disqualifying case. One counter-example is enough to retire a universal claim, which is why one lost deal to a feature gap deserves attention that one won deal does not.

Recognising these saves a surprising amount of time, because they are exactly the questions competitive monitoring is good at. Public competitor data is thin for averages and excellent for events, which is a reason to build the programme around events.

6. What should I do with a result that falls under the floor?

Keep it, label it an observation with the count attached, and let it prioritise work rather than decide it. Suppressing a thin result from reports is different from deleting the underlying data.

An observation under the floor still has uses. It tells you where to look next, it justifies collecting more, and it is often enough to rule something out even when it cannot rule anything in. What it cannot do is appear as a bare number in a report, because a number without its count travels further than the caveat attached to it.

Four rules for a thin result:

  1. Write the count next to it every single time, in the same sentence rather than in a footnote.
  2. Call it an observation rather than a benchmark, in exactly those words.
  3. Use it to choose what to investigate, never as the sole input to a spend or positioning decision.
  4. Keep collecting on the same definition, so that in two quarters it either crosses the floor or clearly will not.

7. How do I report a number so a reader can judge it?

Print the count, the basis and the period next to the figure. A reader who can see what a number was computed from can decide how much weight to give it without asking you.

The format that works is short: the value, then how it was measured, then how many observations, then the date. "1.04 posts per week, unweighted mean across 56 tracked LinkedIn accounts, June 2026" carries everything a reader needs to argue with it, which is the point.

Also state what was excluded and why. Suppressed cells are more informative than they look, because they tell the reader which questions the data cannot answer yet, and that stops the same question being asked every month.

Oppira records the contributing account count behind every aggregate it computes, which is what makes the difference between a benchmark and an observation something the tool can enforce rather than something a person has to remember.

Key Takeaways

The requirement varies by question, not by team

One observation for whether an event happened, twenty contributing accounts for a benchmark, hundreds of conversions per variant for a rate comparison.

Count the contributors, not the sample

We track 113 competitors and could not publish an engagement rate, because the accounts with a usable follower denominator were far under our floor.

Grids eat data

A seven-day by four-band posting grid is 28 cells, needing about 140 posts at five per cell, which is well over a year for one account.

Percentages hide sample size

Five out of 100 versus eight out of 100 supports nothing. The same split at 1,000 each supports a decision. Always print the raw counts.

Events need almost no data

Whether a competitor changed their pricing, the order of the top three posts, and any halving or doubling are all readable immediately.

Set the floor before you look

A threshold chosen after seeing the result is a threshold chosen to include the result. Write the number down first, then compute.

Frequently Asked Questions

Oppira

Turn reading into a reaction

Oppira watches your competitors, keeps your playbook current, and drafts the response. Start free and see your market clearly by tomorrow.