01The number is the last thing decided
A defensible sample size is calculated rather than chosen, and a calculation needs inputs. Which analysis you will run, since comparing two means and fitting a regression with six predictors do not ask for the same thing. The smallest difference that would genuinely matter in practice. How much risk you will accept of missing a real effect, and how much of announcing one that is not there. Supply those and a number comes out. Withhold any of them and no number can be justified to anybody.
That order also explains why nobody can answer this for you over a coffee. The input carrying the most weight is not a statistical quantity at all: it is what size of difference is worth detecting in your setting, which is your professional judgment, argued from comparable studies rather than chosen quietly to keep the total manageable.
02Four inputs and one thing added afterwards
Write these into the proposal in this order, each with its source beside it. Reviewers are not checking your arithmetic so much as checking that every input came from somewhere other than convenience.
- The analysis you will actually run, named precisely, including how many groups or predictors it carries
- The smallest effect worth finding, taken from comparable published work or from what would be enough to change practice
- Your tolerance for the two kinds of error, which most fields settle by convention and your committee will expect stated anyway
- The design overhead: repeated measures, clustering, stratification and unequal groups all move the requirement, usually upward
- Attrition, which is not part of the calculation but is added on top, estimated from studies asking a similar amount of their participants
03When the number is bigger than your world
The usual result of a first calculation is a figure you cannot reach, because your site does not contain that many people, or not that many who meet your criteria and would agree. That is information rather than failure, and there are only a few honest ways to respond to it.
Widen the population or add a site, accepting that this means another permission and probably another review. Simplify the design, since fewer groups and fewer predictors need fewer people. Choose a more efficient design, because measuring the same participants twice often buys more than recruiting strangers does. Or change the claim: a study presented as feasibility or as a pilot, said so from the beginning and analysed on those terms, is respectable work. What is not respectable is running the underpowered version and writing it up as though the shortfall never happened.
04Interview studies count differently
If your work is qualitative, none of the above applies and borrowing its vocabulary will actively hurt you. Numbers there are defended through the design and through saturation, meaning the point at which further interviews stop producing anything new. You estimate a range in the proposal, explain what that range rests on, and then report what actually happened during collection.
Whichever kind of study you are running, write the justification into the methods chapter as a short paragraph naming every input, where each came from, and the method or software used to produce the figure. It is the first thing a reviewer looks for and the first thing a defence question lands on. Where the calculation itself is what is holding up the proposal, the statistics side can be worked through with somebody who does this daily, and the chapter it sits in belongs to the wider build.
FAQQuestions to control
Can I use a rule of thumb instead of a calculation?
Rules of thumb are useful for a first sketch and weak under questioning, which makes them a poor thing to put in a proposal. A reviewer asking where your number came from wants inputs, not a convention someone repeated in a workshop. Run the calculation properly, state the inputs, and mention the rule of thumb only as a sanity check that your result is in a plausible region.
What if I genuinely cannot recruit enough people?
Say so early, in writing, to your chair. The options are widening the population, adding a site, simplifying the design, switching to a design that reuses participants, or reframing the study as a pilot. All five are survivable and all five are far easier before collection than after it. The one outcome to avoid is discovering the gap when you are already analysing.
Do I need special software to do this?
You need something that documents its assumptions, and free tools built for power analysis do that perfectly well. What matters to a committee is not which package produced the figure but whether the inputs are named, sourced and defensible. Save the output, record the version and the settings you used, and keep them with your methods notes so the number can be reproduced a year later.
How many interviews are enough for a qualitative study?
Enough that new interviews stop yielding new content, which is a judgment you make during collection rather than a number you can promise in advance. Proposals normally give a range with reasoning from similar studies, then explain how saturation will be assessed and reported. Committees accept that readily when the reasoning is written out; they push back when a bare number appears with nothing behind it.