Research Tips

Statistical Power, Explained: Why 80% Is the Default and How to Reach It

By Arie Lindenburg Published on August 16, 2026 6 min read

Somewhere in your methods chapter, a supervisor expects a sentence like "a power analysis indicated a required sample of 128 participants". If you are not sure what that sentence means, or you have been copying it from other theses and hoping nobody asks, this guide is for you.

Statistical power is a simple idea wearing intimidating clothes. By the end of this post you will understand what it is, why everyone uses 80%, and how to produce your own defensible number with our free statistical power calculator, a G*Power alternative that runs in your browser.

What statistical power actually is

Suppose the effect you are studying is real: the intervention works, the two groups genuinely differ, the correlation exists. Statistical power is the probability that your study will detect it.

Power of 80% means: if the effect is real, your test has an 80% chance of coming back significant, and a 20% chance of missing it entirely. That miss is called a Type II error, and it is the quiet tragedy of underpowered research. The effect was there. Your study just could not see it.

📝

The two ways a study can be wrong

A false positive (Type I error) is calling an effect real when it is not; its probability is your alpha, usually 5%. A false negative (Type II error) is missing a real effect; its probability is 1 minus power. At 80% power you accept a 20% miss risk.

The four dials, and how they trade off

Power analysis has four quantities. Fix any three and the fourth follows. That is all a power calculator does.

  • Effect size: how big the difference or relationship is. Small effects need big samples to see.
  • Alpha: your significance threshold, almost always 0.05.
  • Power: your target detection probability, conventionally 0.80.
  • Sample size: the thing you usually solve for.

Effect size is the dial that dominates everything. For a two-group comparison (Cohen's d):

Required sample per group, two-tailed t-test, alpha 0.05, power 80%

Effect size Cohen's d Per group Total
Small 0.2 394 788
Medium 0.5 64 128
Large 0.8 26 52

Read that table twice. Chasing a small effect costs six times the sample of a medium one. This is why honest effect size expectations matter more than any other planning decision, and why "we'll just see what we find" studies so often find nothing.

Why 80% power became the convention

The 80% default traces back to Jacob Cohen, who reasoned that a false positive is roughly four times as bad as a false negative for science as a whole. With alpha at 5%, setting the miss risk at 20% keeps that 4:1 ratio. It is a judgment call that hardened into convention, the same way 95% confidence did.

Should you ever deviate? Sometimes. Clinical and high-stakes research often uses 90% power. Exploratory student projects sometimes accept less, openly. What you should not do is run the analysis backwards: collecting 40 responses, then choosing whatever power target makes 40 look adequate.

⚠️

Power analysis comes before data collection

A power analysis after fieldwork is a description, not a justification. Run it while planning, report the inputs you chose, and let it set your recruitment target. Reviewers can tell the difference.

How to pick an effect size honestly

This is the part that stalls people, because the honest answer is "you have to make a judgment". Three usable strategies, in order of preference:

  1. Prior literature. Find studies measuring something similar and use their reported effect sizes, discounted a little, since published effects skew optimistic.
  2. Smallest effect of interest. Ask what the smallest difference is that would still matter in practice, and power for that. This is the most defensible framing in a thesis.
  3. Cohen's benchmarks. Small (d = 0.2), medium (d = 0.5), large (d = 0.8) as a last resort when the literature is thin. Say explicitly that you are using conventional benchmarks.

A worked example you can copy

Say your thesis compares stress scores between students who use a mindfulness app and students who do not. Prior studies suggest a medium effect, around d = 0.5.

  1. Test: independent two-group comparison, two-tailed.
  2. Effect size: d = 0.5, justified from literature.
  3. Alpha: 0.05. Power: 0.80.
  4. Result: 64 per group, 128 total completed responses.

Then plan backwards from 128. With a realistic response rate, how many invitations is that? Our response rate calculator turns the target into an invitation plan, and the sample size calculator covers the simpler case where you are describing one population instead of comparing groups. If you are unsure which of those two situations you are in, the difference is explained in the sample size formula, explained.

128 responses sounds like a lot?

The SurveySwap exchange delivers real, engaged respondents for free, and you can top up with paid, targeted responses when the deadline is close.

Get Free Responses

Underpowered anyway? Say so

Sometimes the honest power analysis returns a number you cannot reach. You have three respectable options: measure a bigger effect (stronger manipulation, more reliable instruments), simplify the design (fewer groups, a within-subjects design), or proceed and report the limitation transparently, framing your study as preliminary. What is not respectable is quietly ignoring the analysis you ran. And whatever your sample size ends up being, protect its quality: a powered study full of bots and speeders is worse than a smaller clean one. Start with attention checks.

Frequently asked questions

What does 80% statistical power mean?

If the effect you are studying is real, a study with 80% power has an 80% chance of detecting it as statistically significant, and a 20% chance of missing it (a Type II error).

What is the difference between power analysis and a sample size calculation?

Margin-of-error sample size calculations describe one population ("what percentage of students feel X"). Power analysis plans for detecting effects or differences ("do group A and B differ"). Comparisons and hypothesis tests need power analysis.

Is G*Power the only way to run a power analysis?

No. G*Power is free desktop software and a fine choice, but for the common cases (t-tests, correlations, comparing two proportions) a browser-based power calculator gives you the same numbers faster.

Can I run a power analysis after collecting data?

Post-hoc power analysis is widely discouraged. Run the analysis before fieldwork to set your target; afterwards, report confidence intervals instead.

Find every calculator mentioned here in the free research tools collection.

Tags

statistical-power sample-size quantitative-research

Related articles

Get free respondents

Join 5,000+ researchers who collect survey responses for free.

Get started for free