Every clinical trial eventually produces a table of numbers that someone has to defend — to a regulator, an investor, or a peer reviewer. The statistical analysis plan (SAP) is the document that determines, in advance, exactly how those numbers will be derived. Written and finalized before anyone on the study team sees unblinded data, it converts the protocol’s high-level objectives into a fully specified, executable statistical program: which populations are analyzed, which model is fitted to which endpoint, how missing data are handled, and what the output tables will look like.
Trial complexity has made this discipline harder to skip. Protocols benchmarked by the Tufts Center for the Study of Drug Development now specify roughly 60% more procedures and enroll subjects at over 60% more investigative sites than protocols a decade earlier1, a trend that pushes more data, more endpoints, and more analytical decisions into every study — decisions that belong in the SAP, not in ad hoc discussions after the database is locked.
The SAP Is Not the Protocol — And That Distinction Matters
The protocol and the SAP answer different questions. The protocol describes what the trial will do: the design, treatment groups, eligibility criteria, and endpoints to be measured. The SAP describes how the resulting data will be analyzed once collected — the statistical model for each endpoint, the exact populations included in each analysis, and the precise rules for handling the messy realities of real-world data collection, such as missed visits, protocol deviations, and dropouts.
Regulators expect this separation. ICH E9, the foundational guideline for statistical principles in clinical trials, calls for the protocol to describe the principal features of the planned analysis, while the SAP supplies the detailed, technical elaboration needed to actually execute it2. The protocol describes only the planned confirmatory analyses of the trial, expediting protocol creation and shifting alignment toward the SAP. In practice, sponsors increasingly draft the SAP in parallel with the protocol rather than waiting until after protocol finalization, since many analytical decisions — how a composite endpoint is derived, which covariates adjust for baseline imbalance — are easier to get right while the study design is still open to discussion.
The Regulatory Foundation
A defensible SAP is built on a small set of internationally harmonized guidelines, and AICROS biostatisticians reference each of these by name in every SAP they produce:
- ICH E9 (Statistical Principles for Clinical Trials) — the core guideline defining the SAP’s purpose: minimizing bias and ensuring that the analysis reflects decisions made prospectively, not after seeing the data.
- ICH E9(R1) Addendum — introduced the estimand framework in 2019, requiring sponsors to explicitly define the target population, the treatment condition, how intercurrent events (such as treatment discontinuation) are handled, and the population-level summary measure for each key objective.
- ICH E3 (Structure and Content of Clinical Study Reports) — governs how SAP-driven results are ultimately reported in the clinical study report (CSR), tying the SAP directly to the submission package.
- ICH E6 (Good Clinical Practice) — provides the quality and documentation framework the SAP operates within, including version control and sign-off requirements.
Because SAPs are increasingly posted publicly on ClinicalTrials.gov alongside the protocol, the estimand framework has also become a tool for external scrutiny — reviewers, journal editors, and competitors can now check whether a trial’s published conclusions match what was pre-specified.
The Core Components of a Statistical Analysis Plan
While formats vary by sponsor and therapeutic area, a well-constructed SAP consistently addresses the following elements. This is the structure AICROS biostatisticians build from on every engagement, adapted to the specific design and phase of the study:
| Component | What It Specifies |
|---|---|
| Study objectives & estimands | Primary, secondary, and exploratory objectives translated into ICH E9(R1) estimands — the target population, endpoint, treatment condition, intercurrent-event strategy, and population-level summary measure. |
| Trial design summary | Design type (parallel, crossover, factorial, adaptive), randomization method, stratification factors, and blinding — restated from the protocol so the SAP stands on its own. |
| Analysis populations | Precise definitions of the intention-to-treat (ITT) / modified ITT, per-protocol (PP), and safety populations, including rules for excluding or reclassifying subjects. |
| Endpoint definitions | Operational definitions of primary, secondary, and exploratory endpoints, including derivation rules, visit windows, and handling of composite or time-to-event outcomes. |
| Statistical methods & models | The exact model for each endpoint (e.g., ANCOVA, mixed-effects model for repeated measures, Cox proportional hazards), covariates, hypothesis tests, and significance levels. |
| Missing data & sensitivity analyses | Pre-specified imputation strategy (e.g., multiple imputation, tipping-point analysis) and sensitivity analyses that test whether conclusions hold under alternative assumptions. |
| Multiplicity control | Methods for controlling type I error across multiple endpoints, doses, or analysis time points—gatekeeping procedures, hierarchical testing, or alpha-splitting. |
| Interim analyses & stopping rules | Timing and statistical boundaries for interim looks, the role of an independent Data Monitoring Committee (DMC), and early-stopping criteria. |
| Subgroup analyses | Pre-specified subgroups (e.g., by age, biomarker status, region), acknowledging that these are typically descriptive rather than confirmatory. |
| Mock shells for TLFs | Table, listing, and figure (TLF) templates that show the exact layout of the outputs before a single data point is analyzed. |
| Sample size & power | The statistical justification for the planned sample size, restated with any refinements made after protocol finalization. |
Why Timing Is as Important as Content
An SAP that is technically complete but finalized late offers little protection. The prevailing standard — consistent with ICH E9 — is that the SAP must be finalized and signed off before database lock or study unblinding, whichever comes first. For open-label studies, sign-off must occur before any substantial data accumulation that could bias analytical decisions.
Sponsors increasingly push this timeline even earlier, for reasons beyond regulatory compliance. Finalizing the SAP alongside the protocol lets statistical programmers begin building analysis datasets and mock output tables as data arrives, rather than after the last patient’s last visit, compressing the path to database lock and CSR delivery. It also surfaces design ambiguities — an underspecified composite endpoint, an untested assumption about dropout patterns — while the protocol can still be revised, rather than after enrollment has begun.
That second point has real financial weight.
Tufts CSDD research found that 76% of Phase I–IV trials now require at least one substantial protocol amendment, up from 57% a decade earlier, with each amendment costing between $141,000 and $535,000 in direct expense alone4 — a separate Tufts analysis found nearly half of such amendments were avoidable5. An SAP finalized early, in coordination with protocol development, is one of the more cost-effective safeguards against analysis-driven amendments discovered too late to avoid.
Why the SAP Matters Beyond Compliance
It minimizes bias and protects scientific credibility
The core purpose of the SAP, as stated in ICH E9, is to ensure that analytical choices — which covariates to adjust for, how to define a responder, which subgroup comparisons to report — are locked in before anyone can see how those choices affect the result. Without this discipline, post hoc analytical flexibility (sometimes called ‘p-hacking’ when taken to extremes) can turn a null result into an apparently positive one, undermining the trial’s scientific and regulatory credibility.
It reduces costly rework and accelerates submission
A complete SAP with mock TLF shells lets statistical programmers build SDTM and ADaM datasets and validate output programs well before database lock, rather than starting from a blank page once the data are final. This directly shortens the interval between the last patient’s last visit and CSR delivery — the metric sponsors watch most closely when planning a submission timeline. It also provides opportunities to identify issues in the recorded data earlier, supporting data-cleaning efforts.
It supports reproducibility and regulatory scrutiny
Regulatory statisticians reviewing an NDA or MAA submission compare the analyses reported in the CSR against the pre-specified SAP. Deviations must be justified explicitly. An SAP that anticipates likely data issues — informative censoring, competing risks, non-proportional hazards — and pre-specifies sensitivity analyses to address them gives reviewers confidence that the primary result is robust rather than fragile.
It reflects a broader shift toward specialized biostatistics capability
Sponsors are increasingly separating data management and biostatistics from general clinical operations when selecting outsourcing partners, partly driven by the growing technical demands of estimand-based analysis, adaptive designs, and Bayesian methods that require more specialized skills. Across several CRO market segments tracked by Mordor Intelligence, the data management and biostatistics service line is now growing faster than the broader market it sits within:
| Market Segment | Data Management & Biostatistics Growth | Segment Leader / Context |
|---|---|---|
| Hybrid & tech-enabled CRO models | 10.50% CAGR (fastest-growing service line in this segment) | Adaptive designs requiring continuous Bayesian re-estimation and dropout prediction |
| Rare disease CRO services | 9.20% CAGR through 2031 | Outpaces the segment’s 55.10%-share clinical trial management line |
| Early-phase CRO services | 7.00% CAGR through 2031 | Rising biomarker-driven, genetically stratified Phase IIa designs |
Frequently Asked Questions
How is an SAP different from the statistical section of the protocol?
The protocol states the trial’s design and objectives and typically summarizes the intended analysis in broad terms, with only the primary analysis and accompanying analysis population being fully specified . The SAP is the detailed, technical elaboration of that summary — the specific model, population definitions, and handling rules needed to actually execute the complete analysis, as described in ICH E9.
When exactly must the SAP be finalized?
Before database lock or study unblinding, whichever occurs first. For open-label trials, sign-off should occur before a substantial volume of data accumulates. Many sponsors now target finalization well ahead of that deadline — some as early as protocol approval — to reduce downstream risk.
Who is responsible for writing and approving the SAP?
A lead biostatistician typically drafts the SAP in collaboration with the medical monitor, data management, and statistical programming teams. Final sign-off usually requires the sponsor’s responsible statistician.
What happens if the actual analysis deviates from the SAP?
Deviations are permitted but must be disclosed and justified in the clinical study report. Undisclosed deviations invite regulatory scrutiny and can undermine the credibility of the reported results, which is precisely why the SAP is finalized and version-controlled before unblinding.
Does every study need a full-length SAP, including early-phase and exploratory trials?
Scope should match risk and regulatory intent. Confirmatory Phase II/III trials warrant a comprehensive SAP with full estimand specification and mock shells. Early-phase or exploratory studies may use a lighter-weight version, but ICH E9’s core principle — pre-specifying analytical decisions before unblinding — still applies.
About AICROS
The Association of International CROs (AICROS) is a global alliance of specialist-led, small-to-midsize contract research organizations, offering multinational reach across six core service lines: biostatistics, clinical operations, data management, medical writing, pharmacovigilance, and regulatory affairs. AICROS member organizations combine the depth of dedicated statistical and scientific specialists with the coordination of a global network, giving sponsors direct access to experienced biostatisticians who build SAPs aligned with ICH E9(R1), CDISC standards, and current regulatory expectations from first draft through CSR delivery.
To discuss a statistical analysis plan, a biostatistics engagement, or any of AICROS’s six service lines, contact the AICROS team at info@aicros.org.
References
- Tufts Center for the Study of Drug Development (Tufts CSDD), benchmarking data on Phase III protocol procedures and site counts, 2015–2025, as summarized by CRIO. https://clinicalresearch.io/blog/the-rising-complexity-of-study-design-what-it-means-for-clinical-research-sites/
- ICH E9: Statistical Principles for Clinical Trials, European Medicines Agency. https://www.ema.europa.eu/en/ich-e9-statistical-principles-clinical-trials
- Incorporating estimands into clinical trial statistical analysis plans, PubMed. https://pubmed.ncbi.nlm.nih.gov/35257600/
- The Amendment Trap: Why 76% of Clinical Trials Face Six-Figure Protocol Changes, Precision for Medicine, citing Tufts CSDD. https://www.precisionformedicine.com/blog/the-amendment-trap-why-76-of-clinical-trials-face-six-figure-protocol-changes
- Getz K, et al. The Impact of Protocol Amendments on Clinical Trial Performance and Cost. Therapeutic Innovation & Regulatory Science, Springer Nature. https://link.springer.com/article/10.1177/2168479016632271
- Guidelines for the Content of Statistical Analysis Plans in Clinical Trials, American University of Beirut (SPIRIT/CONSORT guidance). https://www.aub.edu.lb/SHARP/PublishingImages/Pages/publications/Guidelines-for-the-Content-of-Statistical-Analysis-Plans-in-Clinical-Trials.pdf
- Hybrid CRO Models And Tech-Enabled CRO Market Size, Share & 2031 Growth Trends Report, Mordor Intelligence. https://www.mordorintelligence.com/industry-reports/hybrid-cro-models-and-tech-enabled-cro-market
- Rare Disease Contract Research Organization (CRO) Market Size, Share & 2031 Growth Trends Report, Mordor Intelligence. https://www.mordorintelligence.com/industry-reports/rare-disease-contract-research-organization-cro-market
- Early Phase Contract Research Organization (CRO) Services Market Size, Share & 2031 Growth Trends Report, Mordor Intelligence. https://www.mordorintelligence.com/industry-reports/early-phase-contract-research-organization-cro-services-market

