What sample size means in herbal research is simpler than most journal methods sections make it sound: it is the number of people, plants, specimens or observations a study actually analyses, and it decides how much weight the result can carry. In herbal research the unit changes with the design, so a field survey of thirty informants and a placebo-controlled trial of sixty patients are both “small” but not comparable things.
Sample size alone never settles whether a study is good. What matters is whether the researchers chose the number before collecting data, whether it was large enough to detect the difference they were hunting for, and whether the people analysed match the people recruited. The rest of this guide walks through those questions in plain language.
Table of Contents
- 1What Sample Size Means in Herbal Research
- 2Why Sample Size Matters
- 3How Researchers Choose a Sample Size
- 4What an effect size is
- 5What statistical power is
- 6What the alpha level is
- 7What Sample Size Means in Herbal Research: From Goal to Number
- 8Different Herbal Research Designs Need Different Samples
- 9How to Tell Whether a Herbal Study Has a Strong Sample
- 10Common Sample-Size Mistakes and How to Avoid Them
- 11Frequently Asked Questions
- 12What is the minimum sample size for herbal research?
- 13Is a sample size of 30 enough for an herbal study?
- 14Why do the treatment and placebo groups have different numbers?
- 15How should researchers account for participants who drop out?
- 16Can a small ethnobotanical study still provide useful evidence?
- 17Conclusion: What to Check First in an Herbal Study
What Sample Size Means in Herbal Research
Sample size is the number of participants, plants, specimens or observations included in a study. In herbal research it decides whether a result can be trusted or whether it is simply too small to reveal anything, and the unit being counted changes with the study design.

That last part trips people up. A cross-sectional survey of how many hospital patients use herbs counts people. An ethnobotanical survey counts informants, and often records several plant species per informant. A laboratory study of one plant might count chemical assays, or batches of extract, rather than human beings at all.
Think about a specific case. A team wants to know whether a standardised extract of a leaf used for cough across two districts changes symptom scores over four weeks. They run a randomised, placebo-controlled trial with 40 people in the extract arm and 40 in the placebo arm. The sample size there is 80 people, analysed as two groups of 40.
The same team, a year earlier, interviewed traditional healers about plants used for fever. Thirty informants mention eleven species. Here the sample size is 30 informants, and the eleven species are the outcomes being described, not the sample.
Either way the number alone tells you very little. A study can have a large sample and still be weak if the participants were recruited from one clinic in one city, or if the researchers picked the number after seeing the data. A study with a small sample can still be honest and informative, especially in ethnobotanical fieldwork, where no one can recruit a representative sample of a continent’s healers.
Why Sample Size Matters
Sample size mainly controls one thing: how reliably a study separates a real effect from random noise. Bigger samples narrow the range of results that chance alone can produce, so a difference that shows up is more likely to be a genuine difference.
Four practical consequences follow from that.
Precision. Larger samples give tighter confidence intervals. If a trial estimates a symptom score improvement, a wide interval stretching from a small benefit to a large one tells you very little about which is closer to the truth.
Power. Statistical power is the chance of detecting a real effect if one exists. Small studies have low power, so they miss real effects and publish non-significant results that say nothing about the herb either way.
Subgroup work. If a study wants to compare response by age group, dose, or preparation type, each subgroup needs its own adequate number. Small samples quietly rule that kind of analysis out.
Attrition. The planned sample and the analysed sample are rarely the same. People withdraw, lose interest, or stop reporting, and the difference between planned and analysed numbers is often the most informative line in a paper.
Larger is not automatically better, though. A thousand participants recruited from a single online forum can be less useful than eighty recruited from a defined community with careful randomisation. Sample size buys precision, not representativeness.
It also has an ethical side. Recruiter pressure is real: over-recruiting exposes people to risk, cost and time for an answer nobody needs, while under-recruiting wastes everyone’s effort on a study that cannot conclude.
How Researchers Choose a Sample Size
A formal power calculation is the usual route. Researchers set the effect they expect, the confidence level they want, the power they want, and how much the measurements naturally vary, and the calculation returns the number they need per group. Know any three of these and you can solve for the fourth, which is why power analysis is described as a relationship between quantities rather than a formula with a fixed answer.
What an effect size is
An effect size expresses how big a difference is in units of natural variation, rather than as a raw number. Cohen’s benchmarks put a small effect around 0.2, a medium effect around 0.5, and a large effect around 0.8 on that scale. Researchers either take a value from prior work, or settle on the smallest effect they would still consider worth detecting clinically. That last choice matters: asking to detect a tiny difference demands a very large group.
What statistical power is
Power is the probability of getting a statistically significant result when a real effect of the assumed size exists. Trials usually aim for 80% or 90%. That target is also where Type II error comes in, which is the chance of missing a real effect entirely and reporting a non-significant result by mistake.
What the alpha level is
Alpha is the tolerated false positive rate, conventionally 0.05, meaning a 5% chance of calling a real non-effect a genuine finding. This is Type I error. Alpha and power pull against each other, so tightening one usually enlarges the required sample.
Two more inputs sit behind those three. Variance, expressed as a standard deviation, comes from how much scores differ between people or samples. And the design matters: comparing two independent groups needs more people than measuring the same people twice, and a planned comparison of four groups needs more than a comparison of two.
Ethnographic and ethnobotanical studies usually skip power calculations entirely, because the goal is to describe what a community knows and uses rather than to test an intervention. Those projects plan on saturation, meaning interviews continue until new informants stop introducing new species or uses.
What Sample Size Means in Herbal Research: From Goal to Number
Here is the calculation chain, worked through on a fictional extract trial for a cough preparation used traditionally in two districts.

Step one: state the question. Does the standardised extract reduce daytime cough score after four weeks compared with matched placebo?
Step two: pick the outcome and timepoint. One primary outcome measured at four weeks. Adding a second co-primary outcome changes the calculation, because the chance of at least one false positive rises across outcomes.
Step three: set the assumptions. Assume a medium expected effect of 0.5, a standard deviation of 0.8, alpha of 0.05 two-sided, and 80% power.
Step four: read off the group size. At those settings a two-arm comparison needs roughly 64 people per arm, 128 in total.
Step five: add attrition. Herbal trials built on people already using a traditional remedy often see high dropout, since participants can simply stop taking the study product and carry on as before. Assuming 15% attrition means inflating each arm to about 76, so 152 recruited in total.
Step six: add screening failures. In the published gastrodia-uncaria trial reported in Circulation, 779 people were screened and 251 were randomised, a conversion of about 32%. Screening exclusions of that size have to be planned for, and they are a big reason herbal trials miss their recruitment target.
Step seven: write the number down twice. The protocol records 152 planned. The published paper reports the number actually analysed, and those two figures rarely match.
Different Herbal Research Designs Need Different Samples
No single sample size fits every herbal research design, because the unit of analysis and the question change completely between designs.
| Design | Unit counted | What the sample represents | Strength | Limitation |
|---|---|---|---|---|
| Case report | One patient | A single described outcome | Documents unusual responses | Says nothing about how common anything is |
| Cross-sectional survey | Respondents | Prevalence of herb use in a defined population | Estimates how common a practice is | Association only, no cause and effect |
| Ethnobotanical field survey | Informants | Reported uses of plant species in a community | Captures knowledge that is not in any clinic record | Cannot reach a representative sample; index values shift as informants increase |
| Laboratory study | Assays, extracts or batches | Composition, potency or activity of a material | Precise and repeatable | Results may not translate to people |
| Randomised controlled trial | Randomised patients | Cause and effect of a preparation on an outcome | Supports causal claims | Expensive, slow, and hard to blind with plant products |
| Longitudinal cohort | Followed participants | Change over time within the same people | Shows trajectories and durability | Loses people over time, inflating attrition |
| Qualitative interviews | Interviewees | Reasons, practices and meanings | Explains why practices happen | No numeric estimate of effect |
The two numbers most often confused are the survey figure and the trial figure. A cross-sectional herb-use survey at 95% confidence and a 5% margin of error, assuming half the population uses herbs, lands on 384 participants, and that is the pattern repeated across several published national surveys. A therapeutic trial is sized completely differently, because it estimates an effect and a difference rather than a proportion.
How to Tell Whether a Herbal Study Has a Strong Sample
Run these checks in order when you read a paper or a news summary of one.
- Look for a stated calculation. A methods section should say what effect size, power, alpha and variance were assumed. If it only says “a sample of 100 was selected”, there was no planning.
- Compare the planned and the analysed number. The gap between recruitment target and final analysis is where underpowering hides.
- Check group balance. Unequal groups are not automatically wrong, since random allocation produces imbalance by chance, but large gaps deserve a look at why they happened.
- Ask about attrition and where it went. Compare the intention-to-treat analysis with any per-protocol analysis, and check whether the two tell the same story.
- Check missing data handling. Dropping participants who did not finish can quietly favour whichever group did well.
- See whether subgroups were planned. Subgroup splits after the fact rarely support firm claims, especially in small samples.
- Read the baseline table. Similar groups at the start make a comparison meaningful; wildly unequal groups weaken it.
- Find the effect estimate and its confidence interval. A p-value on its own hides the size of the effect and the range of plausible values.
- Judge significance against clinical importance. A difference that is statistically detectable but too small to matter to a patient is not a useful result.
Judging a paper only by whether it clears 30 or 100 participants will mislead you. Thirty participants in a randomised comparison of two treatments and thirty interview subjects describing plant uses are not on the same scale at all.
Common Sample-Size Mistakes and How to Avoid Them
Treating a pilot as proof. Pilot and feasibility studies are built to test procedures, estimate variance and check recruitment, not to settle effectiveness. Their numbers should feed a later calculation rather than be read as findings.
Assuming a convenient sample is representative. Patients attending one hospital, or herb buyers posting in one online group, are not the population. Multi-herb formula trials also face a hard attribution problem, since with eleven herbs in one preparation you cannot tell which ingredient produced any change, and the expected effect gets diluted across all of them.
Calculating power after seeing the results. Running a power calculation once the data are in is circular, since post hoc power calculated from observed results simply restates the confidence interval. Plan before, not after.
Reading too much into small subgroups. Twenty people split across four groups gives five per group, which cannot support a comparison no matter how clean the data look.
Counting repeated measurements as extra participants. Twelve weekly symptom scores from ten people is still ten people. Repeated measures help some designs, but they do not add independent participants.
Treating a non-significant p-value as proof of no effect. In a study with 20 per arm, a p-value above 0.05 usually means the trial could not detect anything. The honest reading is “not large enough to tell”, not “the herb does not work”.
Ignoring variance from the plant material itself. Herbal preparations vary by species, growing site, harvest time and batch, and an unstandardised product adds spread to the results. That inflates the standard deviation and pushes the required sample size up.
The scale of the problem is documented. A systematic review of 167 clinical herbal medicine studies found that only 0.6% carried out an a priori sample size calculation, and an analysis of 31,873 clinical trials across all fields found only 12% to 13% exceeded 80% power. Herbal research sits inside a wider pattern of underpowered trials, not a special case of it.
Frequently Asked Questions
What is the minimum sample size for herbal research?
There is no universal minimum, because the required number depends on the design, the expected effect, the chosen confidence level and the power. A two-arm randomised trial often needs 60 to 80 people per group for a medium effect, while a prevalence survey sized for a 5% margin of error at 95% confidence lands near 384. Ethnobotanical field studies use saturation rather than a fixed number. Ask what the study was powered to detect, not whether it passed an arbitrary threshold.
Is a sample size of 30 enough for an herbal study?
Thirty per group is a sometimes-cited convention, not a standard. Thirty per group can detect large differences reasonably well but is poorly suited to modest effects, and it leaves nothing for attrition or subgroup analysis. One published review of Chinese herbal formula literature noted that no prior trial reported a sample size estimation and fewer than 130 patients were included per trial. Thirty also works fine in many ethnobotanical interviews, where the goal is describing use, not estimating an effect.
Why do the treatment and placebo groups have different numbers?
Random allocation alone produces unequal group sizes by chance, especially in small trials, and it is rarely a sign of manipulation. A large imbalance usually reflects recruitment problems, dropouts concentrated in one arm, or exclusions applied after randomisation. Read the participant flow diagram to see where participants went between randomisation and analysis. If attrition was concentrated in the treatment arm and the analysis only counts people who completed the study, the comparison is much weaker than the headline numbers suggest.
How should researchers account for participants who drop out?
Recruit more than the required number using a documented attrition assumption, then state clearly in the paper both the planned sample and the number actually analysed. Herbal trials need a generous allowance because participants already using a traditional remedy can stop taking the study product without stopping their usual care. Keep an intention-to-treat analysis of everyone randomised, report any per-protocol analysis separately, and describe missing outcomes rather than quietly removing the participants who left.
Can a small ethnobotanical study still provide useful evidence?
Yes, as long as it is read for what it can actually support. Field studies document knowledge, preparation methods and local names held by a community, and a carefully conducted study of a defined group is valuable even at a modest informant count. The key checks are whether saturation was reached before stopping, whether informants were selected transparently, and whether plant identification was verified by herbarium specimens. These studies estimate prevalence and acceptability, not therapeutic effect.
Conclusion: What to Check First in an Herbal Study
Sample size is one component of credibility, not a quality score you can read off the abstract. A number on its own tells you almost nothing about whether an herbal study deserves weight.
Start with the design, because it tells you what the sample actually was: people, informants, or laboratory batches. Then check whether the researchers planned the number before collecting data, whether they stated the effect they were trying to detect, how many people were analysed against how many were planned, and where the dropouts went. Read the effect estimate and its confidence interval rather than the p-value alone, and decide whether the change would matter to an actual person.
Two findings are worth keeping in mind as you read. Only 0.6% of 167 clinical herbal medicine studies in one systematic review used an a priori sample size calculation, and the gastrodia-uncaria trial in Circulation stopped at 63% of its planned enrolment, leaving post hoc power at 72.3%. Between them they explain much of the field’s reputation for small, inconclusive trials.
None of this settles whether a traditional plant helps anyone. That question needs well-designed, adequately powered trials, and for your own health decisions a doctor, pharmacist or qualified practitioner is the right person to ask.


