P

Probability

System Availability Risk Calculator

Screen the probability that planned downtime plus Poisson-modeled corrective failures will reach a downtime threshold, with expected downtime and probability-weighted exposure.

DOWNTIME-THRESHOLD TAIL RISK

Measure the chance of crossing a consequential downtime boundary

The model converts horizon and MTBF into an expected failure count, maps each count to downtime through a fixed restoration duration, and sums every count that reaches the entered threshold.

Threshold breach probability -
Probability-weighted exposure -
Failures needed to trigger -
Expected failure count -
Expected total downtime -
Expected availability screen -

LIVE DECISION RECORD

Failure-count threshold ladder

Each row links a possible failure count to total downtime, exact Poisson probability, cumulative mass, and threshold status.

Reliability manager watching failure tokens approach a red downtime threshold on an operations board
Risk is defined by a boundary: planned hours consume part of it, and each additional failure moves the system toward the consequence zone.
Failure-count threshold ladderCurrent inputs; unrounded model values
Each row links a possible failure count to total downtime, exact Poisson probability, cumulative mass, and threshold status.
FailuresTotal downtimeExact probabilityCumulative probabilityThreshold status

CURRENT CALCULATION PROCESS

Formula, current substitution, intermediate values, and reconciliation

lambda = H/MTBF; m = max(0, ceil((T-D_p)/MTTR)); Risk = P(N >= m), N ~ Poisson(lambda)

Current symbol, unit, and entered-value register
SymbolMeaning and unitCurrent value
horizonHoursRisk horizon (hours) - Calendar exposure period for the failure-count screen.2160
mtbfHoursMean time between failures (hours) - Constant-rate failure spacing used to set the Poisson mean.600
mttrHoursRestoration time per failure (hours) - Fixed downtime increment assigned to each modeled failure.7
plannedDowntimeHoursPlanned downtime (hours) - Known downtime already consuming part of the threshold.10
downtimeThresholdHoursDowntime risk threshold (hours) - Total downtime level that defines the adverse event.32
thresholdExposureConsequence if threshold is reached - Contract, margin, or recovery exposure in one currency.180000

    Waiting for valid inputs.

    WHO THIS MODEL SERVES

    A scoped decision aid, not a universal forecast

    Primary audience: Reliability, operations-risk, insurance, and SLA teams screening a count-driven downtime threshold.

    Decision boundary: Use for a constant-rate Poisson failure count and fixed repair duration; do not treat it as a safety-integrity calculation or repair-duration distribution.

    HOW TO RUN THE RISK SCREEN

    Five steps from horizon to escalation probability

    1. Choose the exposure horizon and confirm MTBF uses operating hours compatible with it.
    2. Enter a restoration duration that represents the scenario, not an optimistic best case.
    3. Separate committed planned downtime from random corrective events.
    4. Define the downtime threshold from an SLA, production, cash, or recovery decision.
    5. Inspect the integer trigger count and retain the table with the consequence source.

    RISK FUNDAMENTALS

    Five concepts required by the tail model

    Poisson mean
    lambda=H/MTBF is the expected count under a constant event rate.
    Trigger count
    The smallest whole failure count that reaches the remaining downtime threshold.
    Upper tail
    The combined probability of the trigger count and every more severe count.
    Expected downtime
    Planned hours plus lambda times repair hours; it is distinct from threshold probability.
    Consequence exposure
    A conditional amount attached to the threshold event, not an event probability.

    FORMULA AND DEFAULT SUBSTITUTION

    Convert a time boundary into a count boundary

    lambda=H/MTBF; m=ceil((T-D_p)/MTTR); P(breach)=1-sum from k=0 to m-1 e^(-lambda)lambda^k/k!

    Defaults give lambda=2160/600=3.6 failures. The unplanned hours needed are 32-10=22, so m=ceil(22/7)=4 failures. The live risk is P(N>=4), and expected downtime is 10+3.6 x 7=35.2 hours.

    DEEPER RISK ANALYSIS

    Three controls that can move the tail

    Reliability versus restoration

    Increasing MTBF lowers lambda and shifts probability toward fewer failures. Reducing MTTR can change the integer count needed to breach even when failure frequency is unchanged.

    Threshold cliff

    Because failures are whole events, a small threshold or repair-time change can move m by one and create a discrete jump in risk.

    Severity calibration

    The weighted exposure is meaningful only when the entered consequence corresponds to the same breach definition, horizon, and accounting basis.

    WORKED RISK CASES

    Two boundaries with different escalation logic

    Quarterly production SLA

    With lambda=3.6 and four failures needed, the risk team can compare a spares project that cuts repair time with a reliability project that increases MTBF; each acts on a different side of the trigger equation.

    Planned-stop saturation

    If the threshold is 8 hours while planned downtime is 10 hours, the event is already true. A 100% result is correct and signals that the threshold definition or maintenance commitment needs review.

    RISK TERMINOLOGY

    Six terms for the decision record

    Failure intensity
    Constant event rate represented by 1/MTBF.
    Poisson process
    Independent event-count model with stationary rate.
    Downtime threshold
    Total planned-plus-corrective hours defining the adverse event.
    Trigger failures
    Minimum corrective count needed to cross that boundary.
    Tail risk
    Probability assigned to all counts at or beyond the trigger.
    Weighted exposure
    Tail probability multiplied by one conditional consequence amount.

    EVIDENCE RETENTION

    Keep event and consequence definitions aligned

    Retain failure timestamps, operating-hour denominator, repair-close criteria, planned-stop calendar, threshold authority, and consequence calculation. Record whether repeated alarms were deduplicated into incidents.

    LIMITS AND EXCLUSIONS

    Where the screen stops

    • Failure arrivals use one constant Poisson rate with independent increments.
    • Every failure receives the same fixed repair duration.
    • Overlapping failures, shared causes, queueing, and resource contention are excluded.
    • The entered threshold exposure is applied once, not by downtime severity.
    • Safety, regulatory, and contractual risk require domain-specific review.

    RELIABLE SOURCES

    Primary references for count and availability assumptions

    DOWNTIME-RISK FAQ

    Questions about the breach tail

    Why use a Poisson count instead of MTBF/(MTBF+MTTR)?

    Availability gives an average time share; risk asks for a tail event. A Poisson count model can sum the chance of enough failures to cross a specified downtime threshold.

    What if planned downtime already exceeds the threshold?

    The adverse event is already triggered, so the calculator returns 100% threshold probability with zero additional failures required.

    Why is repair time fixed per failure?

    It keeps the threshold mapping auditable. If repair durations vary materially, use a compound count-severity model or simulation.

    Does probability-weighted exposure equal expected loss?

    Only if the entered consequence applies once whenever the threshold is reached and no other state-dependent costs matter.

    Can I enter a wear-out asset?

    Not without caution. The model assumes a constant failure rate; age-dependent hazard violates the Poisson-process premise.

    Why cap expected failures at 200?

    The page is a screening tool with explicit exact summation; very high counts need a more suitable numerical or continuous approximation and a reviewed model.

    IMPORTANT RISK NOTE

    A stable rate is a hypothesis, not a label

    Check for trends, clustering, maintenance effects, and common causes before using the Poisson tail. Escalation decisions should retain uncertainty around both event frequency and consequence.