Statistical Power Analysis Calculator
Paste your group values or upload a CSV file. Get effect size, achieved power, required sample size per group, and the minimum detectable effect — with four colourful charts you can read on any screen.
Quick answer: Statistical power is the probability your study detects a real effect of a given size. Power = 1 − β, and the standard target is 0.80. The three inputs that set it are effect size, alpha and sample size. Key landmarks for an independent-samples t-test at α = .05, two-tailed, 80% power: d = 0.20 needs 394 per group, d = 0.50 needs 64 per group, d = 0.80 needs 26 per group. Halve the effect size and you need roughly four times the data.
🔑 Key Takeaways
- 1Power = the chance of detecting an effect that is really there. Power of .80 means that if the effect exists at the size you assumed, you would find it in 8 of every 10 repeats of the study.
- 2Required sample size scales with 1/d². Halving the effect you want to detect roughly quadruples the n you need. This one relationship drives every power decision you will ever make.
- 3Effect size is the input you must justify, not guess. Use a meta-analysis, a pilot's lower confidence bound, or a defensible smallest effect of interest. Cohen's 0.2 / 0.5 / 0.8 anchors are a last resort.
- 4Post-hoc power adds nothing to your p-value. Observed power is a one-to-one function of p, so it cannot defend a null result. Report the effect size with a confidence interval instead.
- 5Report the minimum detectable effect when your result is null. "We had 80% power to detect d ≥ 0.81" is informative; "our power was low" is not.
- 6Under-powered studies exaggerate the effects they do find. Only unusually large sample effects clear significance at small n, so published estimates come out inflated — the Type M error behind much replication failure.
- 7Reducing noise beats recruiting more subjects. Better instruments, tighter protocols, blocking and strong covariates all shrink the denominator of your effect size and are usually far cheaper than doubling n.
1. Enter Your Data
Comma-separated values (default). One box per group. Group names are editable and appear in every chart and table. Add as many groups as you need, and use the × button on any group to remove it.
Upload CSV file
Supported: .csv only. After upload, click the columns that should each become a cluster — selected columns will be loaded as separate clusters (groups). If your data is in Excel, use File → Save As → CSV first.
Test settings
3. Interpretation of Results — In Detail
3.1 What each number is actually telling you
Effect size is the headline. Everything else in power analysis is downstream of it. Cohen's d (two groups) is the distance between two means expressed in pooled standard deviations. A d of 0.50 means the two group means sit half a standard deviation apart. Cohen's f (three or more groups) is the standard deviation of the group means divided by the within-group standard deviation. The two are linked: for exactly two equal-sized groups, f = d/2.
The conventional anchors are Cohen's, and they are rough by design: d ≈ 0.20 small, 0.50 medium, 0.80 large; f ≈ 0.10 small, 0.25 medium, 0.40 large. Treat them as a fallback when you have no field-specific benchmark. In ecology and field biology, a d of 0.30 in a hard-to-sample population can matter enormously; in a tightly controlled lab assay, a d of 0.80 can be trivial. Always ask "is this difference large enough to change a decision?" before you ask "is this difference statistically significant?"
Achieved power (also called observed or post-hoc power) is the probability that a study with your exact sample sizes, your alpha, and an effect of exactly the size you observed would return a significant result. Read it as a design diagnostic, not as evidence. It is mathematically a one-to-one function of your p-value, so it adds no new information about whether your specific result is real. What it does tell you is whether your design was ever capable of catching effects of that magnitude.
Required sample size per group is the forward-looking number and the one most people actually need. It answers: to have a target power chance of detecting an effect of this size at this alpha, how many observations do I need in each group? Note the steepness — halving the effect size roughly quadruples the required n. This is the single most important intuition in experimental design.
Minimum detectable effect (MDE) flips the question. Given the sample size you already have or can afford, what is the smallest standardised effect you have a fair chance of detecting? If your MDE is 0.90 and the literature says the real effect is around 0.30, your study cannot succeed no matter how carefully you run it. Reporting the MDE alongside a non-significant result is far more informative than reporting post-hoc power.
3.2 Reading the four charts
Chart 1 (group means ± SD) shows you whether the effect is driven by a genuine shift in central tendency or by one group having a much wider spread. Wide, heavily overlapping bars with similar means mean a small standardised effect and a large required sample size, regardless of how far apart the raw means look on the original scale.
Chart 2 (power curve) is the workhorse. Find your current n on the x-axis, read power off the y-axis, then find where the curve crosses your target power line — that crossing is your required n. Notice the curve is concave: early increases in n buy a lot of power, later increases buy very little. Beyond about 0.95 power the curve is nearly flat, which is why chasing 0.99 power is almost always a waste of resources.
Chart 3 (power vs effect size) tells you the robustness of your design. If your planned n gives 0.80 power at d = 0.60 but only 0.35 power at d = 0.40, and you are not confident the true effect exceeds 0.50, your design is fragile. The three curves show what happens if you can recruit 1.5× or 2× more.
Chart 4 (required n vs target power) makes the cost of ambition explicit. Moving from 0.80 to 0.90 power typically costs about 30% more subjects; moving from 0.90 to 0.95 costs another 25% on top. Use this chart in grant budgeting conversations.
3.3 Decision guide
3.4 Common misreadings to avoid
- "My post-hoc power was low, so the effect might still be real." Circular. Low observed power and a high p-value are the same fact stated twice. Report the confidence interval on the effect instead.
- "I'll keep collecting until p < 0.05." Optional stopping inflates the false positive rate well above alpha. Fix n in advance, or use a formal sequential design with alpha spending.
- "The effect was significant so it is large." With a large enough n, trivially small effects become significant. Judge magnitude from the effect size, never from p.
- "Unequal group sizes are fine." They are allowed, but power is governed by the harmonic mean of the group sizes. A 10 vs 200 split has roughly the power of two groups of 19.
- "I'll use the effect size from my pilot." Pilot effect sizes are noisy and biased upward when selected for being interesting. Use the lower bound of the pilot's confidence interval, or the smallest effect you would care about.
4. How to Write Your Results in Research
4.1 A priori power analysis (Methods section)
Write this before collecting data. It belongs in Methods, not Results.
Every a priori statement needs four elements: the test, the assumed effect size with its justification, alpha, and target power. The justification is the part reviewers scrutinise — cite a meta-analysis, a pilot study, or state a smallest effect size of interest and defend it substantively.
4.2 Reporting an ANOVA design
4.3 Reporting results with a significant effect
Order matters: descriptive statistics first, then the test statistic with degrees of freedom, then the exact p-value, then the effect size with a confidence interval, then a plain-language magnitude statement.
4.4 Reporting a non-significant result honestly
That MDE sentence is what separates a defensible null report from an uninformative one.
4.5 Sensitivity power analysis (when n was fixed for you)
4.6 APA-style checklist before submission
- Test name, degrees of freedom in parentheses, test statistic to 2 decimals
- Exact p-value to 3 decimals (write p < .001 below that threshold); no leading zero on p, d, or r
- Effect size with a 95% confidence interval — required by APA 7th edition
- Means and SDs for every group, in text or in a table (never both)
- Statistical symbols italicised: t, F, p, d, f, M, SD, n
- Software and version named, e.g. "power analyses were conducted using G*Power 3.1 / the StatsUnlock Power Analysis Calculator"
- Assumption checks reported: normality, homogeneity of variance, independence
- Any deviation from the pre-registered plan disclosed explicitly
4.7 Table template for a manuscript
| Group | n | M | SD | Test | p | Effect size [95% CI] | Power |
|---|---|---|---|---|---|---|---|
| Control | 30 | 19.80 | 3.90 | t(58) = 4.63 | < .001 | d = 1.20 [0.64, 1.75] | .99 |
| Treatment | 30 | 24.60 | 4.10 |
5. Formulas Used
6. How to Use This Calculator
- Enter your data. Type or paste comma-separated values into each group box, or click Load sample data to see a worked example immediately.
- Or upload a CSV file. Choose a .csv file (export from Excel with File > Save As > CSV). The tool reads the header row, shows every column as a clickable chip, and loads each column you select as its own cluster.
- Rename or remove groups. Click a group name field and type to rename it — names flow through to the summary table, the charts and both downloads. Use the × button on any group to delete that specific group; the minimum is two.
- Set alpha, tails and target power. Defaults are α = .05, two-tailed, power = .80, which match the great majority of published designs.
- Click Run Power Analysis. Two groups triggers an independent-samples t-test; three or more triggers a one-way ANOVA automatically.
- Read the KPI cards first — effect size, achieved power, required n per group and MDE are the four numbers you will report.
- Check the group summary table for n, mean, SD, SEM and variance per group, and confirm none of your groups has an unexpectedly large spread.
- Study the four charts to understand how sensitive your conclusions are to the assumed effect size and to your recruitment target.
- Download the results as .txt for your lab notebook or .csv for a manuscript table.
- Write it up using the templates in Section 4, quoting the effect size with its confidence interval alongside the p-value.
7. Sample Size Reference Table
The numbers researchers look up most often. Use these to sanity-check the calculator, or to plan a study before you have any data.
Independent-samples t-test — n required per group (α = .05, two-tailed)
| Cohen's d | Magnitude | Power .80 | Power .90 | Power .95 |
|---|---|---|---|---|
| 0.10 | very small | 1571 | 2103 | 2600 |
| 0.20 | small | 394 | 527 | 651 |
| 0.30 | small–medium | 176 | 235 | 290 |
| 0.40 | small–medium | 100 | 133 | 164 |
| 0.50 | medium | 64 | 86 | 105 |
| 0.60 | medium–large | 45 | 60 | 74 |
| 0.70 | medium–large | 34 | 44 | 55 |
| 0.80 | large | 26 | 34 | 42 |
| 1.00 | large | 17 | 23 | 27 |
| 1.20 | very large | 12 | 16 | 20 |
Multiply by 2 for the total N. For a one-tailed test, required n falls by roughly 20%. Moving to α = .01 raises it by roughly 50% (d = 0.50 needs 96 per group instead of 64).
One-way ANOVA — n required per group (α = .05, power .80)
| Groups (k) | f = 0.10 (small) | f = 0.25 (medium) | f = 0.40 (large) |
|---|---|---|---|
| 2 | 382 | 62 | 25 |
| 3 | 317 | 52 | 21 |
| 4 | 271 | 45 | 18 |
| 5 | 238 | 39 | 16 |
Effect size conversions and benchmarks
| Measure | Small | Medium | Large | Conversion |
|---|---|---|---|---|
| Cohen's d | 0.20 | 0.50 | 0.80 | d = 2f (two equal groups) |
| Cohen's f | 0.10 | 0.25 | 0.40 | f = √(η² / (1 − η²)) |
| Eta squared (η²) | 0.01 | 0.06 | 0.14 | η² = f² / (1 + f²) |
| Correlation r | 0.10 | 0.30 | 0.50 | r = d / √(d² + 4) |
8. Example Results
Six worked examples you can reproduce in the calculator above. Each one shows a decision a real researcher has to make.
Example 1 — The textbook case: planning for a medium effect
You are planning a two-group experiment and the literature suggests a medium effect. You set α = .05, two-tailed, and target 80% power.
Achieved power at n = 64: 0.802
This is the single most quoted number in applied statistics. Note what happens if you were over-optimistic: if the true effect is only d = 0.30, those same 128 participants give you just 0.35 power — you would miss the effect two times out of three.
Example 2 — A large effect, but too few subjects
You ran 15 animals per group and observed a large standardised difference of d = 0.80.
Required n for .80 power: 26 per group
Minimum detectable effect at n = 15: d = 1.06
Even a large effect can be missed at this sample size — you had a coin-flip chance. The MDE line is the honest summary: with 15 per group, only effects above d = 1.06 were reliably within reach. Eleven more animals per group would have fixed it.
Example 3 — A pilot study that cannot succeed
A pilot with 10 participants per group looking for a small-to-medium effect of d = 0.40.
Required n for .80 power: 100 per group
Minimum detectable effect at n = 10: d = 1.33
This design detects the target effect roughly one time in seven. It is not a failed study — it is a legitimate pilot whose job is to estimate the effect size and the variance, not to test the hypothesis. Say exactly that in the write-up and use the numbers to justify the funded follow-up.
Example 4 — One-way ANOVA with three groups
Three treatment levels, planning for a medium ANOVA effect of f = 0.25 at α = .05 and 80% power.
Achieved power at n = 52: 0.804
Equivalent η² = 0.059
Adding a fourth group actually lowers the per-group requirement to 45 (total N = 180) because the omnibus test gains from the extra between-group information — but the total sample still rises. If your real question is one specific contrast rather than the omnibus test, power that contrast as a two-group comparison instead.
Example 5 — Why unbalanced groups waste data
Two designs, both with 50 participants in total, both looking for d = 0.60 at α = .05.
Balanced 25 vs 25: power 0.547
The same 50 people give you 43% more power when split evenly. Power is driven by the harmonic mean of the group sizes, which is dominated by the smaller group — a 10 vs 40 design has roughly the power of two groups of 16. Whenever allocation is under your control, split evenly.
Example 6 — What alpha and tails cost you
The same medium effect (d = 0.50) at 80% power, under three different testing choices.
α = .05, two-tailed: 64 per group
α = .01, two-tailed: 96 per group
A one-tailed test saves 20% of your sample — but only if the direction was pre-registered and you are genuinely willing to treat an effect the other way as a non-finding. Tightening alpha to .01, as you must after a Bonferroni correction for five tests, costs 50% more subjects. Decide both before you collect data.
9. When to Use This Tool
Use this power analysis calculator whenever a sample size question is on the table: planning an experiment and needing a defensible recruitment target, writing the sample size justification for a grant, thesis or ethics application, checking whether a study you are reviewing was ever capable of detecting the effect it claims, turning pilot data into a properly powered follow-up, or working out the minimum detectable effect when the sample size was fixed by budget, permits or the number of animals available.
It suits postgraduate students designing a first study, ecologists and wildlife biologists planning field seasons, psychologists pre-registering an experiment, clinical researchers sizing a trial arm, peer reviewers assessing methodological adequacy, and anyone who needs a power calculation on a phone or a shared computer without installing G*Power or opening R.
Use something else when: your outcome is binary, a count or a time-to-event (those need proportion, Poisson or log-rank formulas); your design has repeated measures, nesting or random effects (use a simulation approach such as the R package simr); or you need power for regression coefficients, correlations or chi-square tests, which G*Power covers and this tool does not.
10. Assumptions and Limitations
- Independence. Observations must be independent. Repeated measures, litters, plots nested in sites or students nested in classrooms all violate this and need a mixed-model power approach instead.
- Approximate normality. The t-test and ANOVA tolerate mild non-normality at moderate n, but heavy skew or strong outliers will distort both the effect size and the power estimate.
- Homogeneity of variance. The pooled-SD effect size assumes similar spread across groups. If variances differ by more than roughly 4:1, use Welch's test and treat the power figures here as indicative only.
- Continuous outcome. Binary, count and time-to-event outcomes need their own power formulas (proportions, Poisson, log-rank).
- Equal allocation for planning. Required n is reported per group under equal allocation, which is the most efficient design for a fixed total N.
- Noncentral distributions are approximated. Power uses the Johnson–Welch approximation for noncentral t and the Patnaik approximation for noncentral F. Both are accurate to about ±0.005 across the useful range, which is well within the precision any power analysis deserves.
- Post-hoc power is descriptive. Achieved power computed from the observed effect adds no inferential information beyond the p-value. Use it to describe the design, not to defend a null result.
- No correction for multiplicity. If you are running many tests, divide alpha accordingly (or use FDR) before entering it here — power falls as alpha falls.
- Attrition is not modelled. Inflate the required n by your expected dropout rate: nrecruit = n* / (1 − dropout).
11. Conclusion
Statistical power is the quiet variable that decides whether a study was ever capable of answering its own question. A beautifully designed protocol, a carefully calibrated instrument and a well-chosen model all lose their value if the sample size was never large enough to separate signal from noise. That is why power analysis belongs at the start of a project, in the Methods section and in the grant budget, rather than being reconstructed afterwards to explain a disappointing p-value.
This calculator gives you the four numbers that carry almost all of the practical information. The effect size puts your difference on a scale that transfers across studies and instruments. The achieved power tells you what your realised design could see. The required sample size per group converts an intended effect into a concrete recruitment target. And the minimum detectable effect tells you, given a sample size you may not control, exactly which effects were within reach and which were never going to appear. Between them they turn a vague worry — "maybe we needed more subjects" — into a defensible, quotable statement.
Two habits will improve your work more than any statistical refinement. The first is to decide the smallest effect size that would actually matter before you look at data, and to justify it substantively rather than by reaching for Cohen's conventions. Conventions are a fallback for when you know nothing; you usually know something. The second is to report effect sizes with confidence intervals as your primary result. A confidence interval communicates both the estimate and its uncertainty in one object, and it answers the question a reader really has — how big is the effect, and how sure are we? — in a way that a bare p-value never can.
Remember the arithmetic that drives everything: required sample size scales with the inverse square of the effect size. Halve the effect you want to detect and you need roughly four times the data. This is why small effects are expensive, why under-powered studies produce exaggerated estimates when they do reach significance, and why reducing measurement noise is often a far cheaper route to power than recruiting more subjects. Tighter protocols, better instruments, blocking on known sources of variation and adding strong covariates all shrink the denominator of your effect size and can be worth more than doubling n.
Finally, treat power analysis as a planning conversation rather than a compliance box. Run the numbers for the effect you expect, for half that effect, and for the smallest effect you would care about. If the design only works under the most optimistic assumption, you have learned something important while it is still cheap to act on. Use the charts above to have that conversation with your supervisor, your collaborators or your reviewers, and you will design studies that answer their questions the first time.
12. Frequently Asked Questions
What is statistical power in simple terms?
It is the probability your study finds an effect that is genuinely there. Power of .80 means that if the effect really exists at the size you assumed, you would detect it in 8 of every 10 repetitions of the study.
Why is 0.80 the standard target power?
It is a convention proposed by Jacob Cohen, reflecting a judgement that a Type II error is about four times less costly than a Type I error at α = .05. It is not a law. Confirmatory or expensive studies often target .90 or .95.
What is the difference between a priori, post-hoc and sensitivity power analysis?
A priori solves for sample size before data collection. Post-hoc computes achieved power from observed data. Sensitivity solves for the minimum detectable effect given a fixed n. A priori and sensitivity are useful; post-hoc is descriptive only.
Is post-hoc power meaningful?
Only as a description of the design. Observed power is a deterministic function of the p-value, so it cannot provide independent evidence about whether an effect exists. Report confidence intervals instead.
How do I choose an effect size before I have data?
In order of preference: a meta-analysis in your field, a published study using the same measure, your own pilot data (use the lower confidence bound), or a smallest effect size of interest justified on practical grounds. Cohen's conventions are the last resort.
What is Cohen's d and how is it interpreted?
It is the difference between two means in pooled standard deviation units. Rough anchors: 0.20 small, 0.50 medium, 0.80 large. Field context should override these anchors whenever you have it.
When should I use Hedges' g instead of Cohen's d?
Whenever group sizes are small, roughly under 20 per group. Cohen's d is biased upward in small samples and Hedges' g applies the standard correction factor.
What is Cohen's f and how does it relate to d?
Cohen's f is the ANOVA effect size — the SD of group means over the within-group SD. For two equal groups, f = d/2. Anchors: 0.10 small, 0.25 medium, 0.40 large.
What is eta squared and partial eta squared?
Eta squared is the proportion of total variance explained by the grouping factor. Partial eta squared removes other factors' variance from the denominator and is standard in factorial designs; the two are identical in a one-way design.
How does alpha affect the required sample size?
Lowering alpha makes the test more conservative and demands more data. Moving from α = .05 to α = .01 typically increases required n by roughly 45–50% at the same power and effect size.
Does a one-tailed test really need fewer subjects?
Yes, about 20% fewer, but only if you can justify the direction in advance and are willing to treat an effect in the opposite direction as a non-finding. Most reviewers expect two-tailed tests unless the design is explicitly directional and pre-registered.
What sample size do I need to detect a medium effect?
For a two-tailed independent t-test at α = .05, power .80, and d = 0.50, you need about 64 per group (128 total). At d = 0.80 you need about 26 per group; at d = 0.20 you need about 394 per group.
Can I use unequal group sizes?
Yes, and this calculator handles them. Be aware that power is driven by the harmonic mean, so a badly unbalanced design wastes data. Aim for equal allocation when you have a choice.
What if my data are not normally distributed?
At moderate n the t-test is robust to mild departures. For strong skew, transform the outcome, use a Welch or robust test, or run a simulation-based power analysis for the non-parametric alternative. Rank tests typically need about 5% more subjects than the t-test when normality does hold.
How do I handle unequal variances?
Use Welch's t-test for the inference, and treat the pooled-SD effect size here as an approximation. If the variance ratio exceeds about 4:1, simulate power under your actual variance structure.
Should I add extra subjects for dropout?
Yes. Inflate the required n by dividing by (1 − expected dropout rate). With a 15% expected loss and a required 64 per group, recruit 64 / 0.85 ≈ 76 per group.
What is the minimum detectable effect and when should I report it?
It is the smallest standardised effect your sample size can detect at your target power. Report it whenever your result is non-significant, and whenever n was fixed by resource constraints rather than by design.
Can I stop collecting data once the result is significant?
Not with a fixed-sample design — optional stopping badly inflates false positives. If you need interim looks, use a group-sequential design with an alpha-spending function, or a Bayesian sequential approach with a pre-specified stopping rule.
How do I power a study with more than two groups?
Enter all groups here for the omnibus one-way ANOVA power. If your real question is a specific contrast, power that contrast as a two-group comparison instead — omnibus power and contrast power can differ substantially.
Do I need to correct alpha for multiple comparisons?
If you are testing many hypotheses, yes. Enter the corrected alpha (for example .05/6 = .0083 for six Bonferroni-corrected tests) before computing required sample size, since power falls as alpha falls.
What does it mean if my study is under-powered but significant?
Be cautious. In low-power designs only unusually large sample effects clear the significance threshold, so the published estimate is systematically inflated. This is the Type M error, and it is a major driver of replication failure.
Is this calculator equivalent to G*Power?
For independent-samples t-tests and one-way ANOVA it reproduces G*Power to about three decimal places. G*Power covers many more designs — regression, correlations, chi-square, repeated measures — so use it for those.
























