|

Sampling and Statistical Strategy for Process Validation

Sampling and statistical methods should operate as a cross-lifecycle governance framework, not as isolated calculations performed for individual validation protocols. The objective is to ensure that data collected during process development, PPQ, continued process verification, investigations, and post-change verification are capable of supporting the decisions for which they are intended.

This article therefore does not repeat the detailed PPQ sampling design covered in PPQ Sampling Plan and Data Collection Strategy, the PPQ-specific acceptance and statistical evaluation addressed in PPQ Acceptance Criteria and Statistical Evaluation, experimental-design methodology covered in Design of Experiments for Process Characterization and Validation, or routine statistical monitoring addressed in Continued Process Verification (CPV) Program and Monitoring Strategy. Its purpose is to establish the principles by which sampling and statistical approaches are selected, justified, reviewed, and maintained consistently across those lifecycle applications.

FDA’s Process Validation: General Principles and Practices connects sampling directly with representative data, statistical confidence, process variability, and continued monitoring. FDA also recommends that statistical methods and data-collection plans used for process stability and capability be developed by a statistician or an individual with adequate statistical-process-control training.


Statistical Governance Across the Validation Lifecycle

A statistical strategy should begin with the decision that the data must support, not with selection of a preferred statistical test. The same numerical method can be appropriate for one validation question and inappropriate for another because population structure, variability, sampling design, independence, distribution, or required confidence differ.

The governing sequence should therefore be: Decision question → population and variability → sampling design → statistical method → assumptions → analysis → interpretation → lifecycle decision

That sequence should remain recognizable from Stage 1 development through PPQ and CPV. The amount of data and the analytical methods may change, but the scientific basis for the decision should remain traceable.

During Stage 1, the emphasis is primarily on learning: identifying sources of variability, characterizing process relationships, estimating variation, and determining where additional knowledge is needed. PPQ shifts the emphasis toward commercial-scale confirmation and reproducibility. CPV then applies accumulated commercial data to detect change, assess process stability, and maintain the state of control. FDA specifically describes a transition from heightened PPQ sampling toward a statistically appropriate and representative routine monitoring level as sufficient commercial variability data become available.

Sampling and statistical strategy across the process validation lifecycle from Process Design through PPQ, CPV, lifecycle review, and improvement.
A common statistical governance framework supports different validation objectives across the lifecycle. Development emphasizes knowledge generation, PPQ emphasizes commercial confirmation, and CPV uses accumulated data to maintain and improve process control.

Start With the Decision Question

Sampling and statistical plans should state what the analysis is intended to determine. Examples include estimating variability, comparing process conditions, confirming uniformity, demonstrating a performance characteristic, evaluating a trend, estimating process capability, or determining whether a post-change process behaves consistently with the established state.

Without a defined decision question, sample-size calculations and statistical tests can appear mathematically rigorous while providing little useful evidence. A study intended to estimate a process mean, for example, has a different design objective from one intended to detect a shift, demonstrate equivalence, estimate a tail proportion, or characterize within-batch variability.

The statistical plan should therefore identify the population of interest, the unit being sampled, the parameter or attribute being evaluated, and the decision that will be made from the resulting evidence.


Representative Sampling

21 CFR 211.160 requires scientifically sound sampling plans and states that samples used for components, in-process materials, and drug products must be representative and properly identified. FDA’s process-validation guidance similarly emphasizes that validation samples should represent the batch under evaluation.

Representativeness is broader than randomness. A purely random set of samples can still provide poor coverage of meaningful process variability if the population contains important structural differences such as beginning versus end of operation, multiple equipment positions, different filling heads, material lots, shifts, campaigns, or process locations.

A representative strategy should therefore reflect the sources of variability that matter to the process and product. The objective is not to sample everywhere; it is to obtain evidence that appropriately represents the manufacturing population relevant to the validation decision.


Risk-Based Sampling

Risk should influence where sampling effort is concentrated, particularly when process knowledge indicates that some locations, times, materials, or conditions are more likely to expose meaningful variability.

Risk-based sampling should not mean sampling only the condition believed most likely to fail. The strategy still needs adequate representation of routine manufacturing. A useful design can combine routine representative coverage with additional samples at locations or conditions where process knowledge indicates higher uncertainty or potential quality impact.

Quality Risk Management in Process Validation provides the broader lifecycle framework for determining how process knowledge, uncertainty, importance, and risk influence validation effort.


Stratification

Stratification divides a process population into meaningful subgroups so that variability that might otherwise be hidden can be evaluated explicitly. Relevant strata may include time within a batch, process location, manufacturing line, equipment position, filling head, material lot, equipment train, shift, campaign, or other scientifically justified factors.

The purpose of stratification is not to create more categories than the dataset can reasonably support. It should be used when there is a plausible process reason for expecting differences among subgroups or when demonstrating uniformity across those subgroups is important to the validation conclusion.

For example, samples collected from beginning, middle, and end positions may be useful when process knowledge indicates that time-dependent variation is credible. The same pattern should not be imposed automatically on every process merely because it is a familiar validation convention.


Sampling Unit and Statistical Independence

A sound statistical plan must distinguish between the number of measurements and the number of independent units represented by those measurements. Ten samples from one batch may provide valuable information about within-batch variability, but they do not provide the same evidence about batch-to-batch reproducibility as ten independent commercial batches.

Repeated measurements from the same physical sample, multiple aliquots from one location, or many observations generated from one production event can improve measurement precision while adding limited information about broader process variability. Treating dependent observations as if they were independent can exaggerate the apparent amount of evidence.

The sampling plan should therefore identify the experimental or observational unit relevant to the conclusion and recognize nested structures such as measurements within samples, samples within locations, and locations within batches.

This principle is particularly important when interpreting PPQ evidence and is treated more specifically in PPQ Acceptance Criteria and Statistical Evaluation.


Sample-Size Rationale

There is no single statistically correct sample size for process validation. Sample size should be driven by the objective of the study and the amount and structure of variability expected.

A scientifically useful rationale can consider the population being represented, expected variability, desired precision, confidence level, magnitude of a difference that should be detectable, risk associated with an incorrect conclusion, number of strata, available prior knowledge, and practical limits of destructive or resource-intensive testing.

FDA specifically expects PPQ sampling to include an adequate number of samples to provide sufficient statistical confidence regarding quality within and between batches, with the selected confidence level potentially based on risk associated with the attribute. This principle is important, but the detailed PPQ sampling calculation and allocation belong in PPQ Sampling Plan and Data Collection Strategy.

A sample-size statement such as “n = 10 based on previous validation practice” provides little scientific justification. A stronger rationale explains what n = 10 is expected to demonstrate and why that amount of information is adequate for the decision being made.


Precision and Statistical Power

Two common statistical concepts can influence sample-size planning: precision and power.

Precision describes how narrowly a population parameter can be estimated. Increasing independent sample size generally reduces uncertainty around estimates such as the mean, although the relationship also depends on process variability. NIST describes confidence intervals as a way of expressing the uncertainty associated with estimating an underlying population parameter from a sample.

Power addresses a different question: whether a study is sufficiently sensitive to detect a meaningful difference when such a difference actually exists. Power is particularly relevant for comparison, equivalence, or change-assessment questions but does not need to be calculated mechanically for every validation dataset.

The statistical objective should determine which concept matters. A study designed primarily to characterize variation may require a precision-based rationale, while a comparison study may be more appropriately justified through detectable difference or power.


Sampling Design and Execution

A statistically justified plan can still fail if sampling is executed inconsistently. Sampling procedures should define where and when samples are collected, how sampling units are selected, how stratification or randomization is performed where applicable, how samples are identified, and how exceptions are handled.

21 CFR 211.165 requires written procedures describing the sampling method and number of units per batch to be tested and requires appropriate statistical quality-control criteria for batch approval.

The statistical strategy should also distinguish planned missing observations from unplanned missing data. A planned design may intentionally omit some combinations because they add little information, while missing results caused by sample loss, laboratory failure, or execution error may create bias and should be evaluated rather than silently ignored.

Sampling design and execution governance showing objective definition, representative and risk-based sampling, sample-size rationale, execution controls, and data management.
Sampling begins with the decision objective and population to be represented. The selected approach, sample size, execution method, and data handling should remain scientifically connected and documented.

Within-Batch and Between-Batch Variability

Process validation frequently requires understanding more than one level of variability. Within-batch variability can arise from location, time, unit operation, equipment position, or process dynamics, while between-batch variability can reflect materials, environment, equipment state, operators, or other manufacturing differences.

FDA specifically expects PPQ statistical methods to consider both intra-batch and inter-batch variability and expects CPV to scrutinize both forms of variation.

These sources should not automatically be pooled into one standard deviation when the distinction matters. The sampling structure should preserve enough information to determine whether variability is predominantly within batches, between batches, associated with a particular stratum, or related to another identifiable source.


Descriptive Statistics Before Inferential Statistics

Statistical analysis should begin with understanding the data. Graphs, distributions, ranges, medians, means, standard deviations, percentiles, subgroup summaries, and time-ordered plots often reveal features that a formal hypothesis test can obscure.

Descriptive statistics are not merely preliminary calculations. They are often the clearest representation of process behavior and should be reviewed before selecting more complex methods.

A statistically significant result can be scientifically unimportant, while a practically important change may fail to reach a conventional significance threshold when the dataset is small. Validation decisions therefore should not be reduced to a p-value.


Distributional Assumptions

Many statistical procedures assume that the data follow a particular probability model, commonly a normal distribution. That assumption should be evaluated rather than imposed automatically.

NIST notes that statistical intervals and hypothesis tests often depend on distributional assumptions and that the selected distribution should be sufficiently appropriate for the statistical technique to produce valid conclusions.

Pharmaceutical process data may be skewed, bounded, discrete, censored, multimodal, zero-inflated, or affected by batch structure. Examples can include microbial counts, impurity results near a reporting limit, hold times, failure proportions, particle measurements, or process durations.

When an assumption is not reasonable, alternatives can include transformation, robust estimation, nonparametric methods, distribution-specific models, or other justified techniques. The statistical plan should explain the choice rather than simply report that “normality passed” or “normality failed.”


Independence and Data Structure

Statistical assumptions also include independence. Time-series observations, repeated measurements, multiple results from the same batch, and samples taken from neighboring locations can be correlated.

Ignoring correlation can underestimate uncertainty and produce overly optimistic confidence intervals or significance tests. The method should therefore reflect the structure of the data where that structure can materially affect the conclusion.

For complex hierarchical datasets, the appropriate analysis may involve separate evaluation of within- and between-group variation, mixed models, variance components, or another method consistent with the decision objective.

The goal is not statistical complexity for its own sake. The model should be only as complicated as required to represent the process adequately.


Confidence Intervals

A confidence interval expresses uncertainty associated with estimating a population parameter from sample data. It can often provide more useful information than a single point estimate because it shows the range of values compatible with the evidence at a stated confidence level.

Confidence intervals can be useful for means, differences, variability measures, proportions, regression parameters, or other quantities depending on the application. The selected confidence level should be justified in relation to the decision and risk rather than assumed automatically to be 95 percent in every validation application.

Detailed use of confidence intervals in PPQ acceptance decisions remains within PPQ Acceptance Criteria and Statistical Evaluation.


Confidence Intervals Versus Tolerance Intervals

Confidence intervals and tolerance intervals answer different questions. A confidence interval estimates uncertainty around a population parameter such as a mean. A statistical tolerance interval is intended to contain a specified proportion of the underlying population at a stated confidence level.

That distinction is important when the validation question concerns the expected distribution of individual process results rather than the precision of the estimated mean.

Tolerance intervals can be useful in selected validation applications, but their assumptions and potentially substantial sample-size requirements should be understood. NIST notes that nonparametric tolerance coverage can require much larger sample sizes when strong population coverage is required.


Process Capability

Capability indices can provide useful information when a stable process distribution can meaningfully be compared with established specification limits. They should not be used automatically for every process parameter or CQA.

Capability analysis depends on assumptions about process stability, distribution, independence, and the relevance of the specification limits being used. A capability number calculated from unstable or highly structured data can give a misleading impression of process performance.

Capability calculations are therefore one statistical tool within the broader strategy, not a universal process-validation acceptance requirement. PPQ-specific use is addressed in PPQ Acceptance Criteria and Statistical Evaluation, while routine capability assessment belongs primarily within Continued Process Verification (CPV) Program and Monitoring Strategy.


Outliers

Unexpected observations should be investigated rather than mechanically deleted. An extreme result can represent data error, analytical error, an assignable process event, or genuine process variability.

The statistical treatment should therefore follow the scientific investigation. Excluding an observation because its inclusion makes a statistical conclusion less favorable is not acceptable justification.

FDA specifically states that PPQ data should not be excluded from consideration without documented, science-based justification. The same principle is useful across validation lifecycle analyses.


Missing Data

Missing data should also be evaluated for potential bias. The impact of a missing result depends on why it is missing and whether the absence is related to process performance.

For example, a randomly broken sample vial may have limited statistical significance, while a missing result caused by an instrument overload associated with unusually high process concentration could be directly related to the characteristic being evaluated.

The statistical plan should define how missing data will be identified, investigated, and handled where foreseeable. Post-hoc replacement or imputation should not be performed without a justified methodology.


Statistical Method Selection

The appropriate method should follow from the data and the question. Descriptive statistics may be adequate when the objective is to summarize variability; comparative methods may be needed when assessing process conditions; regression can evaluate relationships; confidence or tolerance intervals can quantify uncertainty; capability methods can assess performance relative to specifications; and control-chart methods can evaluate time-ordered stability.

No method should be selected solely because it is familiar or available in the statistical software package.

The selected method should also be interpretable by the technical and Quality personnel responsible for the validation decision. A highly complex model offers little value if its assumptions cannot be explained or its result cannot be connected to process understanding.

Statistical method selection framework showing evaluation of data and assumptions, selection of descriptive and inferential methods, interpretation, and lifecycle application.
Statistical methods should be selected from the validation question and structure of the data. Assumptions, limitations, interpretation, and lifecycle application should be documented together with the calculation.

Stage 1 Application

During Process Design, statistical methods primarily support knowledge generation. Sampling can help identify variability, characterize process behavior, estimate distributions, evaluate measurement performance, and provide inputs to process models or designed experiments.

The detailed application of experimental design belongs in Design of Experiments for Process Characterization and Validation and Process Characterization and Development Studies for Process Validation.

The important governance output from Stage 1 is not a final commercial sample size. It is an increasingly quantitative understanding of expected variability, important strata, relevant process relationships, uncertainty, and which questions must still be addressed during PPQ.


PPQ Application

PPQ applies the statistical framework to commercial-scale confirmation. Sampling should be more extensive than routine production where appropriate and should provide adequate information about both within-batch and between-batch performance. FDA specifically expects the PPQ protocol to define sampling points, number of samples, frequency, and statistical methods used to evaluate the collected data.

The detailed design belongs in PPQ Sampling Plan and Data Collection Strategy, while evaluation of acceptance criteria, confidence, tolerance concepts, limited datasets, and capability belongs in PPQ Acceptance Criteria and Statistical Evaluation.

This article establishes the governance expectation that those methods should be predefined, scientifically justified, appropriate to the data structure, and consistent with the lifecycle statistical strategy.


CPV Application

CPV expands the dataset substantially and changes the statistical question from initial confirmation toward ongoing detection of change. Time order, process stability, trends, shifts, variability, capability, and recurrence therefore become more important than they were in a limited PPQ dataset.

FDA recommends statistically trending relevant product and process data and states that monitoring can transition to a statistically appropriate and representative level once sufficient data are available to estimate variability.

The detailed monitoring architecture belongs in Continued Process Verification (CPV) Program and Monitoring Strategy, while interpretation of emerging shifts and trends is addressed in Process Drift, Statistical Signals, and CPV Investigation.

The cross-lifecycle statistical strategy should ensure that PPQ and CPV data remain comparable where meaningful and that changes in methods, limits, or sampling frequency are scientifically justified.


Changes and Post-Change Verification

Sampling and statistical principles also apply when manufacturing changes are introduced. A post-change verification plan should define what evidence is needed to demonstrate that the revised process or control continues to perform acceptably.

The amount of evidence should reflect the nature of the change, affected process knowledge, uncertainty, existing validation evidence, and the ability of the selected measurements to detect an adverse effect.

Process Change Control, Revalidation, and Lifecycle Management addresses the broader validation-impact and revalidation decision. This article provides the statistical governance for determining whether the supporting sampling and analysis are capable of answering that decision question.


Statistical Plans Should Be Predefined

Statistical approaches should be defined prospectively whenever practical. The plan should identify the data to be collected, sampling structure, statistical method, assumptions, acceptance or decision criteria where applicable, and planned handling of exceptional data.

This does not prohibit scientifically justified adaptation. Development studies may generate new knowledge, and CPV may identify previously unknown variation. When a statistical approach is changed after data have been reviewed, the reason should be documented so that the analysis is not perceived as selecting a method merely because it produces a preferred conclusion.

FDA requires PPQ protocols to describe how collected data will be evaluated and the statistical methods used, and expects departures from approved protocols to be justified and controlled.


Documentation of Statistical Rationale

The validation record should allow another technically qualified reviewer to understand why the approach was selected and reproduce the logic of the decision.

A statistical rationale should normally identify the objective, population, sampling unit, sample-selection method, relevant strata, sample-size basis, expected variability, assumptions, statistical method, software or calculation approach where relevant, confidence or significance level where applicable, treatment of missing or unusual data, results, limitations, and final interpretation.

The level of detail should be proportionate to complexity and risk. A simple descriptive analysis may require only a concise justification, while a multilevel or model-based analysis may require considerably more explanation.


Statistical Review and Competence

Statistical analysis should be reviewed by personnel who understand both the statistical method and the manufacturing process. Pure mathematical review without process context can miss invalid assumptions, while process expertise without adequate statistical understanding can produce unsupported conclusions.

FDA specifically recommends that a statistician or individual with adequate training in statistical process-control techniques develop the Stage 3 data-collection plan and statistical methods used to evaluate process stability and capability.

The same principle is useful more broadly: the complexity and importance of the validation decision should determine the level of statistical expertise involved.


Maintaining the Statistical Strategy

The statistical strategy should evolve as process knowledge grows. Stage 1 estimates may be replaced by commercial observations, PPQ assumptions may be refined by CPV, and monitoring frequency may change when sufficient representative data demonstrate stable performance.

Changes should not erase the historical basis of the strategy. The organization should be able to trace why the sampling plan or method changed, what evidence supported the change, and whether the revised approach continues to detect meaningful process variability.

This makes statistical governance part of Process Control Strategy Lifecycle Management rather than a series of disconnected validation calculations.


Key Principles

  • Sampling and statistical methods should be selected from the decision question, population, variability, and risk, not from convention.
  • Representative sampling is not synonymous with random sampling; meaningful process strata should be addressed when they can affect the validation conclusion.
  • The number of measurements should not be confused with the number of independent manufacturing units represented by those measurements.
  • Sample size should be justified using the study objective, expected variability, precision or power where relevant, confidence, population structure, and risk.
  • Statistical assumptions—including distribution and independence—should be evaluated and documented rather than applied automatically.
  • Confidence intervals, tolerance intervals, capability indices, hypothesis tests, and control charts answer different questions and should be used only when appropriate.
  • Stage 1, PPQ, and CPV use the same statistical governance principles but for different lifecycle objectives.
  • Unexpected or missing data should be scientifically evaluated rather than automatically excluded or replaced.
  • Statistical methods, assumptions, limitations, changes, results, and conclusions should remain traceable within the validation record.
  • Sampling and monitoring strategies should evolve as commercial data increase, while remaining scientifically justified and capable of detecting meaningful variability.