Knowledge is Power

Sitewide Search

Search Bare Metal Cyber

Search courses, individual lessons, wiki entries, books, podcasts, magazine articles, Daily Cyber News, and Darwin.

NIST SP 800-53 Learning Center

SI-13 — Predictable Failure Prevention

Read the official control and assessment content, then use the separately labeled Bare Metal Cyber perspective to connect the requirement to implementation, evidence, and sustained operation.

5Enhancements
2Parameters
0Baseline memberships
3Assessment methods

SI — System and Information Integrity · NIST SP 800-53 Release 5.2.0

Official NIST control content

Control statement

  1. a.Determine mean time to failure (MTTF) for the following system components in specific environments of operation: [Organization-defined: system components] ; and
  2. b.Provide substitute system components and a means to exchange active and standby components in accordance with the following criteria: [Organization-defined: mean time to failure (MTTF) substitution criteria].
Official NIST discussion

Discussion

While MTTF is primarily a reliability issue, predictable failure prevention is intended to address potential failures of system components that provide security capabilities. Failure rates reflect installation-specific consideration rather than the industry-average. Organizations define the criteria for the substitution of system components based on the MTTF value with consideration for the potential harm from component failures. The transfer of responsibilities between active and standby components does not compromise safety, operational readiness, or security capabilities. The preservation of system state variables is also critical to help ensure a successful transfer process. Standby components remain available at all times except for maintenance issues or recovery failures in progress.

Official OSCAL parameters

Organization-defined parameters

These values must be resolved through the organization’s tailoring and governance process. Bracketed parameter references in the control text identify where a decision is required.

system componentssystem components for which mean time to failure (MTTF) should be determined are defined;
mean time to failure (MTTF) substitution criteriamean time to failure (MTTF) substitution criteria to be used as a means to exchange active and standby components are defined;
Original Bare Metal Cyber perspective

From control text to operational evidence

Use Predictable Failure Prevention as a testable risk decision. Translate the official statement into accountable people, repeatable processes, configured technology, and evidence that demonstrates the outcome over time. In this family, pay particular attention to flaw remediation, malicious-code protection, monitoring, integrity, and trustworthy information handling.

Implementation workflow

  • Define the control boundary, responsible owner, inherited portions, and systems or processes in scope.
  • Resolve each organization-defined parameter before declaring the control implemented.
  • Document how the implementation satisfies every clause of the official control statement.
  • Collect evidence as a normal byproduct of operation rather than only before an assessment.
  • Review exceptions, changes, and monitoring results on a risk-based cadence.

Evidence examples

  • patch and remediation records
  • malware protection configuration
  • monitoring alerts and response records
  • integrity validation and exception reports

Common failure patterns

  • patch compliance hides unsupported assets
  • alerts generated without response ownership
  • exceptions never expire
  • integrity monitoring excludes critical configurations

Questions practitioners should ask

  • What risk decision is this control intended to support in this system?
  • Which parts are implemented locally, inherited, shared, or not applicable—and what evidence supports that decision?
  • Do the documented narrative, deployed configuration, operating process, and collected evidence agree?
  • What event or threshold requires the implementation to be reviewed or changed?
Official NIST SP 800-53A content

Assessment objectives and methods

Show the assessment objective
  1. SI-13a.mean time to failure (MTTF) is determined for [Organization-defined: system components] in specific environments of operation;
  2. SI-13b.substitute system components and a means to exchange active and standby components are provided in accordance with [Organization-defined: mean time to failure (MTTF) substitution criteria].

Examine

  • System and information integrity policy
  • system and information integrity procedures
  • procedures addressing predictable failure prevention
  • system design documentation
  • system configuration settings and associated documentation
  • list of MTTF substitution criteria
  • system audit records
  • system security plan
  • other relevant documents or records

Interview

  • Organizational personnel responsible for MTTF determinations and activities
  • organizational personnel with information security responsibilities
  • system/network administrators
  • organizational personnel with contingency planning responsibilities

Test

  • Organizational processes for managing MTTF
Official relationships

Related controls

These relationships come from the official OSCAL catalog. They indicate useful dependencies or context, not automatic inheritance or equivalence.

Official NIST enhancements

Control enhancements

Enhancements add specificity, strength, or scope to the base control. Baseline badges show explicit selections in the official SP 800-53B OSCAL profiles.

Official NIST control enhancement

SI-13(1) — Transferring Component Responsibilities

Take system components out of service by transferring component responsibilities to substitute components no later than [Organization-defined: fraction or percentage] of mean time to failure.

Official discussion

Transferring primary system component responsibilities to other substitute components prior to primary component failure is important to reduce the risk of degraded or debilitated mission or business functions. Making such transfers based on a percentage of mean time to failure allows organizations to be proactive based on their risk tolerance. However, the premature replacement of system components can result in the increased cost of system operations.

Organization-defined parameters (1)
fraction or percentagethe fraction or percentage of mean time to failure within which to transfer the responsibilities of a system component to a substitute component is defined;
Assessment objectives and methods

system components are taken out of service by transferring component responsibilities to substitute components no later than [Organization-defined: fraction or percentage] of mean time to failure.

Examine

  • System and information integrity policy
  • system and information integrity procedures
  • procedures addressing predictable failure prevention
  • system design documentation
  • system configuration settings and associated documentation
  • system audit records
  • system security plan
  • other relevant documents or records

Interview

  • Organizational personnel responsible for MTTF activities
  • organizational personnel with information security responsibilities
  • system/network administrators
  • organizational personnel with contingency planning responsibilities

Test

  • Organizational processes for managing MTTF
  • automated mechanisms supporting and/or implementing the transfer of component responsibilities to substitute components
Official NIST control enhancement

SI-13(2) — Time Limit on Process Execution Without Supervision

Withdrawn

This enhancement is marked withdrawn in the official OSCAL catalog. Related-control metadata below may identify where its intent was incorporated.

Official NIST control enhancement

SI-13(3) — Manual Transfer Between Components

Manually initiate transfers between active and standby system components when the use of the active component reaches [Organization-defined: percentage] of the mean time to failure.

Official discussion

For example, if the MTTF for a system component is 100 days and the MTTF percentage defined by the organization is 90 percent, the manual transfer would occur after 90 days.

Organization-defined parameters (1)
percentagethe percentage of the mean time to failure for transfers to be manually initiated is defined;
Assessment objectives and methods

transfers are initiated manually between active and standby system components when the use of the active component reaches [Organization-defined: percentage] of the mean time to failure.

Examine

  • System and information integrity policy
  • system and information integrity procedures
  • procedures addressing predictable failure prevention
  • system design documentation
  • system configuration settings and associated documentation
  • system audit records
  • system security plan
  • other relevant documents or records

Interview

  • Organizational personnel responsible for MTTF activities
  • organizational personnel with information security responsibilities
  • system/network administrators
  • organizational personnel with contingency planning responsibilities

Test

  • Organizational processes for managing MTTF and conducting the manual transfer between active and standby components
Official NIST control enhancement

SI-13(4) — Standby Component Installation and Notification

If system component failures are detected:

  1. (a)Ensure that the standby components are successfully and transparently installed within [Organization-defined: time period] ; and
  2. (b)[Organization-defined: si-13.04_odp.02].
Official discussion

Automatic or manual transfer of components from standby to active mode can occur upon the detection of component failures.

Organization-defined parameters (4)
time periodtime period for standby components to be installed is defined;
si-13.04_odp.02
alarmalarm to be activated when system component failures are detected is defined (if selected);
actionaction to be taken when system component failures are detected is defined (if selected);
Assessment objectives and methods
  1. SI-13(04)(a)the standby components are successfully and transparently installed within [Organization-defined: time period] if system component failures are detected;
  2. SI-13(04)(b)[Organization-defined: si-13.04_odp.02] are performed if system component failures are detected.

Examine

  • System and information integrity policy
  • system and information integrity procedures
  • procedures addressing predictable failure prevention
  • system design documentation
  • system configuration settings and associated documentation
  • list of actions to be taken once system component failure is detected
  • system audit records
  • system security plan
  • other relevant documents or records

Interview

  • Organizational personnel responsible for MTTF activities
  • organizational personnel with information security responsibilities
  • system/network administrators
  • organizational personnel with contingency planning responsibilities

Test

  • Organizational processes for managing MTTF
  • automated mechanisms supporting and/or implementing the transparent installation of standby components
  • automated mechanisms supporting and/or implementing alarms or system shutdown if component failures are detected
Official NIST control enhancement

SI-13(5) — Failover Capability

Provide [Organization-defined: si-13.05_odp.01] [Organization-defined: failover capability] for the system.

Official discussion

Failover refers to the automatic switchover to an alternate system upon the failure of the primary system. Failover capability includes incorporating mirrored system operations at alternate processing sites or periodic data mirroring at regular intervals defined by the recovery time periods of organizations.

Organization-defined parameters (2)
si-13.05_odp.01
failover capabilitya failover capability for the system has been defined;
Assessment objectives and methods

[Organization-defined: si-13.05_odp.01] [Organization-defined: failover capability] is provided for the system.

Examine

  • System and information integrity policy
  • system and information integrity procedures
  • procedures addressing predictable failure prevention
  • system design documentation
  • system configuration settings and associated documentation
  • documentation describing the failover capability provided for the system
  • system audit records
  • system security plan
  • other relevant documents or records

Interview

  • Organizational personnel responsible for the failover capability
  • organizational personnel with information security responsibilities
  • system/network administrators
  • organizational personnel with contingency planning responsibilities

Test

  • Organizational processes for managing the failover capability
  • automated mechanisms supporting and/or implementing the failover capability
Related controls
Source record

Authoritative sources