You've been running trials for weeks. The data is piling up, patterns are emerging, and part of you wants to declare victory. Another part whispers: just a few more measurements. So when do you actually stop?
This question haunts every experimentalist, from the undergraduate titrating solutions to the physicist tuning a particle detector. Stopping too early means unreliable conclusions. Stopping too late wastes resources and, more dangerously, invites the temptation to keep collecting until the data tells you what you want to hear. The craft lies in deciding before you begin.
Stopping Rules: Deciding Before You Start
A stopping rule is a commitment you make to your future self. Before the first data point is collected, you write down the exact conditions under which you will stop gathering more. This might sound bureaucratic, but it is one of the most powerful tools for honest science.
The reason is subtle. Once you start looking at data, your brain begins to negotiate. A trend appears, and suddenly you want to collect just enough to confirm it. This is called optional stopping, and it silently inflates false positive rates. What looked like a real effect at N=30 may vanish at N=60, but by then you've already stopped and written the abstract.
Good stopping rules are specific: collect exactly 100 samples, or stop when the confidence interval width falls below 0.05, or stop after 20 trials if no effect is detected. The rule should be blind to whether the results please you. Write it in your lab notebook. Date it. Then follow it.
TakeawayA stopping rule made before you see data protects you from yourself. Decide the finish line while you can still be objective about where it should be.
Convergence Testing: When More Data Stops Mattering
Sometimes you don't know in advance how much data you'll need. In these cases, convergence testing helps you recognize the moment when additional measurements no longer meaningfully change your conclusions.
The technique is straightforward. As you collect data, periodically recalculate your key statistic, whether that's a mean, a slope, or a fitted parameter. Plot it against sample size. Early on, the value will bounce around wildly. As N grows, the plot flattens. When adding new measurements produces changes smaller than your required precision, you have converged.
This method is especially valuable in fields like computational chemistry, spectroscopy, and Monte Carlo simulation, where each measurement is expensive. But beware: convergence in appearance is not always convergence in truth. A stable-looking plateau can hide systematic error. Always ask whether your measurement is converging toward the right answer, not just an answer. Cross-check with independent methods when you can.
TakeawayConvergence tells you when your measurement has stabilized, but not whether it has stabilized on the truth. Precision and accuracy are different questions.
Power Analysis: Sizing the Experiment Before You Run It
Power analysis is the practice of calculating, in advance, how large a sample you need to reliably detect an effect of a given size. It answers the question: if the effect I'm looking for really exists, how many measurements will I need to see it?
The math depends on three ingredients. First, the smallest effect size you consider meaningful. Second, your acceptable false positive rate, typically 5%. Third, your desired statistical power, usually 80%, meaning you want an 80% chance of detecting the effect if it's really there. Plug these into a power calculator, and out comes a required sample size.
Skipping this step is common and costly. Underpowered experiments waste resources and fail to detect real effects, contributing to the replication crisis across many fields. Overpowered ones burn time and money detecting effects too small to matter. Power analysis forces you to specify what you're looking for before you look, which is a healthy discipline even when the numbers themselves are approximate.
TakeawayAn experiment without power analysis is a question asked without knowing whether you have the tools to hear the answer.
Knowing when to stop is not a technicality. It is a form of intellectual honesty. The experimenter who decides in advance protects their conclusions from the subtle pull of wishful thinking.
Stopping rules, convergence tests, and power analysis are three complementary tools for making that decision well. Learn to use them, and your experiments will yield conclusions you can actually stand behind, however the data falls.