The sample-size line reviewers actually check
Most proposals do not fail because the idea is weak. They fail because n was guessed.
Three patterns I see every cycle:
1. Defaulting p to 0.5 when you already know better. 0.5 maximises n. That is conservative, not precise. If your pilot, DHS table, or last round put prevalence at 18%, use 18% (and a sensitivity at 25% if you must). Reviewers can tell when the number was copied from a template.
2. Ignoring clustering. A school, clinic, or village is not 400 independent people. If you sample clusters, apply a design effect (DEFF = 1 + (m − 1)ρ). Skip it and your confidence interval is fiction.
3. Forgetting non-response. You calculated 312. You will not get 312 complete records. Inflate for the response rate you can defend, not the one you hope for.
Write the sentence before you write the rest of the protocol: “A sample of N detects a difference of Δ at 80% power and 5% two-sided alpha, allowing DEFF and 10% non-response.” If you cannot fill those tokens, you are not ready to collect.
That is the whole job of a sample-size sheet: make that sentence true, then stop.
