Risk Assessment and FMEA-Based Analysis

Risk Assessment is a structured evaluation of potential failure modes associated with GMP-relevant systems, equipment, utilities, and processes that may impact product quality, patient safety, data integrity, or regulatory compliance. Within a validation lifecycle, risk assessment is used to define validation scope, testing depth, and control strategy, ensuring that validation effort is proportionate to system criticality and intended use.

Risk assessment may use FMEA, Fault Tree Analysis, hazard analysis, process mapping, checklists, preliminary hazard analysis, or another suitable qualitative or quantitative method. The selected method should reflect the system’s complexity, available knowledge, type of decision, and potential GMP impact. FMEA is particularly useful when individual failure modes, their causes, and their effects must be evaluated systematically, but it is not required for every validation assessment.

The level of formality and documentation should be proportionate to the risk, complexity, and uncertainty of the decision. A simple documented assessment may be sufficient for a well-understood, lower-risk system, while a complex or GMP-critical system may require a multidisciplinary FMEA or another structured analysis. Previous assessments may be applied to functionally equivalent systems only when equivalency is evaluated, justified, and documented.


Purpose of Risk Assessment in Validation

The objective of risk assessment is to understand and manage risk, not to eliminate it entirely. By identifying credible failure modes and evaluating their potential impact, organizations can focus validation activities on what truly matters to product quality and patient safety.

Risk assessment supports:


When Risk Analysis Is Required

A formal risk analysis shall be performed when:

  • A lean validation approach is planned
  • The quality or business impact of a new system or utility is unknown
  • Changes are introduced under change control and the impact to quality, compliance, or business continuity is unclear

In these cases, documented risk analysis provides the justification for validation strategy and downstream controls.


Selecting an Appropriate Risk Analysis Method

Failure Mode and Effects Analysis (FMEA) is one structured method used in validation risk assessment. It evaluates potential failure modes, their causes, their effects, and the controls available to prevent or detect them. FMEA is most useful when the system can be broken down into defined functions, components, process steps, or requirements.

FMEA may not be the most effective method for every assessment. Fault Tree Analysis may be more appropriate when the analysis begins with a specific undesirable event and works backward to identify contributing causes. Process mapping, hazard analysis, checklists, or simpler qualitative assessments may be sufficient for less complex or well-understood decisions.

The selected method should fit the question being evaluated. Regardless of the method used, the assessment should identify credible risks, document their potential consequences, define necessary controls, and evaluate the residual risk after those controls are implemented.

Once risks of failure modes are identified, risk reduction measures may be applied to eliminate, reduce, contain, or control risk through engineering controls, procedural controls, monitoring, maintenance, training, or validation activities.

Risk assessment method selection leading to an FMEA sequence of failure mode, effect, severity, controls, and residual risk
FMEA is one of several risk-analysis methods. When selected, it evaluates failure modes and effects, gives high severity independent attention, and verifies controls before residual risk is accepted.

Risk Assessment Methodology

The Risk Assessment identifies, analyzes, and evaluates parameters that are critical to GMP compliance and system performance.

The assessment typically begins with system requirements derived from the User Requirements Specification (URS), when available. Each GMP-critical requirement is evaluated, and one or more risk scenarios are defined.

The Risk Assessment team is responsible for:

  • Identifying potential failure modes
  • Assigning Severity (S), Probability (P), and Detectability (D) scores
  • Applying the defined scoring model consistently
  • Documenting rationale and assumptions

Risk Scoring Model

Severity (S)

ScoreClassificationDescription
1NegligibleFailure causes no impact on product quality, no interruption to manufacturing, no compliance risk. E.g., non-critical display light failure.
2MinorFailure affects a non-GMP utility or secondary function. No direct product or data impact, minimal downtime (e.g., minor HVAC fluctuation outside manufacturing area).
3ModerateFailure impacts a non-critical step but could lead to deviation or rework. Possible short production delay (e.g., buffer tank temperature drift detected and corrected before batch impact).
4MajorFailure affects a GMP-critical utility (WFI loop, clean steam) or key equipment control, likely to lead to batch rejection or process deviation requiring investigation.
5CriticalFailure results in confirmed product contamination, sterility breach, or major data integrity issue. Could trigger regulatory action or product recall.

The potential impact of the failure on product quality, patient safety, or regulatory compliance.


Probability of Occurrence (P)

ScoreClassificationDescription
1RemoteFailure has never been observed; robust preventive maintenance (PM) and monitoring in place (e.g., validated UPS for control systems).
2UnlikelyFailure could occur due to unusual conditions; historical data shows rare events (e.g., filter housing gasket failure once in 5 years).
3PossibleOccasional failure modes documented; some known wear points (e.g., autoclave door seal replacement required once or twice a year).
4LikelyFrequent operational issues; dependent on manual intervention or aging components (e.g., recurring chiller trips, known PLC faults).
5FrequentHigh likelihood of recurring failure unless mitigated (e.g., known history of valve sticking or flow meter sensor drift impacting batches).

The likelihood that a specific failure mode will occur under normal operating conditions.


Detectability (D)

ScoreClassificationDescription
1High DetectabilityAutomatic alarms, interlocks, or monitoring systems reliably catch failures (e.g., System temperature deviation alarms).
2GoodSingle automated or manual control that typically detects failure (e.g., in-process pH checks).
3ModerateFailure might go unnoticed until later QA checks or trending (e.g., pressure drop across filters reviewed post-run).
4LowDetection is only possible via operator observation or delayed test results (e.g., microbial excursions in WFI).
5UndetectableNo reliable detection until after product release or significant impact (e.g., hidden PLC logic error with no alarms).

The likelihood that a failure or its effect will be detected before adverse impact.


Risk Priority Number (PRN)

When FMEA is used, individual scores may be assigned for Severity (S), Probability of Occurrence (P), and Detectability (D). The Risk Priority Number (RPN) is then calculated as: RPN = S × P × D

RPN is a prioritization tool, not an absolute measure of risk or an automatic acceptance criterion. Different combinations of severity, probability, and detectability can produce the same RPN while representing substantially different risk conditions.

Severity must be evaluated independently of the total RPN. A failure mode with a potentially critical effect on product quality, patient safety, sterility assurance, or data integrity must not be accepted solely because low probability or favorable detectability produces a low RPN. High-severity risks require documented evaluation and appropriate controls regardless of the numerical category.

RPN Priority Categories

PriorityPRN RangeActions
Low Risk1-19May be acceptable with documented justification and routine controls, provided severity is not independently unacceptable
Medium Risk20-39Evaluate additional controls or verification before acceptance
High Risk40-125Risk reduction required through design, procedural, monitoring, maintenance, or validation controls

These ranges are examples and should not be treated as universally applicable acceptance limits. Each organization should establish and approve its own scoring definitions, escalation rules, severity overrides, and risk-acceptance criteria. Numerical thresholds support consistent prioritization but do not replace scientific judgment or Quality oversight.


Risk Classification and Mitigation

Failure modes with potential adverse impact on product quality or patient safety are classified as Critical or Direct Impact. Failure modes with no such impact are classified as Non-Critical or No Impact.

For all Critical or Direct Impact risks, mitigation measures shall be defined to reduce probability, severity, or improve detectability. These measures must be necessary, appropriate, and integrated into qualification, validation, and operational controls.


Documentation

All identified failure modes, risk scores, mitigation measures, and supporting justifications shall be documented as part of the formal Risk Assessment record. Risk Assessments are subject to review and approval by the System Owner, appropriate Subject Matter Experts, and Quality Assurance to ensure accuracy, consistency, and regulatory compliance.

Risk assessments previously executed for existing equipment or systems may be applied to functionally equivalent systems, provided that equivalency is appropriately evaluated, justified, and documented.

The table below is provided as an example to demonstrate how Failure Mode and Effects Analysis (FMEA) outputs may be documented and evaluated within a validation risk assessment. The example demonstrates the application of Severity, Probability, and Detectability scoring, the calculation of the Risk Priority Number (RPN), and the reassessment of residual risk following the implementation of mitigation measures.

Failure Mode and Effects Analysis (FMEA) outputs documentation example

This example does not represent a complete or prescriptive risk assessment and does not establish acceptance criteria for any specific system or process. Actual risk assessments shall be performed based on system-specific requirements, intended use, operating conditions, historical performance, and Quality oversight.