How likely is a series system to fail within a year?
In this example, 79.0% of new units fail within their first year of continuous operation (8,760 hours): a unit of four components in series stops at its first component failure, so it runs a full year only if all four parts do. A Monte Carlo simulation of 20,000 units gives 78.97%, against an exact 79.05%.
The four components are a pump, a motor, a seal and a bearing, each with a Weibull lifetime. The exact chance of a failure within a year is 1 − Π exp(−(t/scale)shape) = 79.05%, the product running over the four parts. The mean life is 5,712 hours (exact: 5,696), and 10% of units fail within 1,091 hours. The seal, with the shortest scale and a shape close to 1, is the part that fails first in 71.0% of units.
The model
The workbook has one sheet, Fleet, with a row for each component: Pump, Motor, Seal and Bearing in rows 2 to 5. Each row holds the component’s Weibull shape (column B), its scale in hours (column C) and a Life (h) cell (column D). Two formulas do the rest. System life (h) in D7 is =MIN(D2:D5): the components are in series, so the unit stops at the first failure. Fails within a year in D8 is =IF(D7<8760,1,0), 1 when the unit fails within 8,760 hours, a year of running around the clock, and 0 otherwise. There are no add-in functions.
For the simulation, each Life cell is an input drawn from a Weibull distribution with the shape and scale of its row, and D7 and D8 are the outputs. These distributions belong to the setup that opens with the example in xellstorm, where the shapes and scales are typed in with the same numbers as columns B and C; they are not linked to those cells, so editing column B or C in Excel does not change the simulation. In the downloadable file each Life cell holds the component’s scale as a placeholder, which every trial replaces with a drawn lifetime. The example also sets one target: a system life under 8,760 hours.
| Component | Shape | Scale | Mean life | Fails within a year |
|---|---|---|---|---|
| Pump | 1.5 | 20,000 | 18,055 | 25.2% |
| Motor | 2.0 | 30,000 | 26,587 | 8.2% |
| Seal | 1.2 | 8,000 | 7,525 | 67.2% |
| Bearing | 2.5 | 25,000 | 22,182 | 7.0% |
Each part’s numbers follow from its shape and scale: its chance of failing within a year if it ran alone with the survival formula in the next section, and its mean life as the area under its survival curve.
The exact answer: multiply the survival chances
A Weibull lifetime has two parameters. The scale is the age by which 63.2% of parts have failed, whatever the shape. The shape says how the failure rate changes with age: below 1 it falls (early failures), at 1 it stays constant (random failures, the exponential distribution), and above 1 it rises as the part wears out. The chance that a part is still running at age t is exp(−(t/scale)shape).
In a series system every part must survive, so with independent parts the survival chances multiply. After a year, the pump is still running with a chance of 74.8%, the motor 91.8%, the seal 32.8% and the bearing 93.0%. Their product, 20.95%, is the chance that the unit runs the whole year, so it fails within the year with a chance of 79.05%. In one formula: P(failure by t) = 1 − exp(−Σ (t/scale)shape).
The mean time to failure (MTTF) is the area under the system’s survival curve, the integral of exp(−Σ (t/scale)shape) from zero to infinity. It has no simple closed form for mixed shapes, but a numerical integral gives 5,696 hours, shorter than the mean life of any single part, even the seal’s 7,525 hours: the unit always stops at the earliest of four failures.
Results
The simulation draws a lifetime for each of the four parts in every trial, lets the workbook take the smallest, and counts. With 20,000 trials it reproduces the exact answers: 78.97% of units fail within a year against 79.05%, a gap smaller than one standard error of the simulation (0.29 percentage points), and the simulated mean of 5,712 hours is within one standard error (30 hours) of the exact 5,696.
| Measure | Simulated | Exact |
|---|---|---|
| Chance of failing within a year (8,760 h) | 78.97% | 79.05% |
| Mean life (MTTF) | 5,712 | 5,696 |
| P10, the B10 life (10% have failed) | 1,091 | 1,086 |
| P50, the median life | 4,769 | 4,759 |
| P90 (90% have failed) | 11,624 | 11,610 |
The P10 of system life is also called the B10 life, the age by which 10% of units have failed: 1,091 hours here, about 45 days of continuous running, far less than the mean. Lifetimes are skewed to the right, so half the units fail within 4,769 hours, before the mean of 5,712. See P50, P80 and P90 for how to read percentiles.
Which component stops the system
In each trial the part with the shortest life is the one that stopped the unit. Counting them gives each part’s share of the failures; the exact share is the integral of that part’s failure rate times the system’s survival over all ages.
| Component | Fails first, simulated | Exact | Contribution to variance |
|---|---|---|---|
| Seal | 71.0% | 70.8% | 92.3% |
| Pump | 18.0% | 18.2% | 7.1% |
| Motor | 5.7% | 5.6% | 0.4% |
| Bearing | 5.3% | 5.4% | 0.2% |
The seal fails first in 71.0% of units (exact: 70.8%), and in 74.5% of the units that fail within the first year (exact: 74.2%). The pump comes second with 18.0%. xellstorm’s contribution to variance, estimated from the ranks of the trials as the app shows it, tells the same story: of the variation in system life that the inputs explain, the seal accounts for 92.3%, far ahead of the pump’s 7.1%. In the tornado, which sets one input at a time to its P10 and P90 with the others at their medians, only the seal and the pump move system life at all; the motor and the bearing outlast a median seal even at their own P10. See tornado charts and sensitivity analysis.
Why the seal dominates: close to random failures
Two things make the seal the weak link. Its scale of 8,000 hours is the shortest of the four and less than a year, so on its own it fails within a year with a chance of 67.2%. And its shape of 1.2, the closest to 1, means its failure rate rises only slowly with age, close to the constant rate of random failures: over a year, a seal that has already run a year is only slightly more likely to fail than a new one (see the table below). The bearing, with a shape of 2.5, rarely fails young: its failures come as it wears out, and in most units another part has failed by then, so the bearing stops only 5.3% of them.
The shape decides what maintenance can do. Compare a new part with one that has already run a year without failing:
| Component | Shape | New part | Survived a year | Ratio |
|---|---|---|---|---|
| Seal | 1.2 | 67.2% | 76.5% | 1.1× |
| Pump | 1.5 | 25.2% | 41.1% | 1.6× |
| Motor | 2.0 | 8.2% | 22.6% | 2.8× |
| Bearing | 2.5 | 7.0% | 28.7% | 4.1× |
A seal that has survived a year fails in the next year with a chance of 76.5%, against 67.2% for a new seal: age adds little. A bearing that has survived a year is 4.1 times as likely to fail in the next year as a new one (28.7% against 7.0%). Replacing parts at a fixed age pays off when they wear out, as the bearing does; for a part whose failures are close to random, a new part is hardly safer than the old one.
What a maintenance planner does with it
- Put the effort on the seal. If seals never failed, the chance of a failure within a year would drop from 79.0% to 36.0% (exact: 36.1%), the other three parts alone. That is the most any improvement to the seal can buy, and far more than any other part offers.
- Choose the remedy by the shape. Replacing working seals on a schedule gains little. The levers for the seal are a longer-lived seal (a larger scale), condition monitoring that catches a leak before it stops the unit, and spares and a quick replacement procedure that keep each stop short. Age-based replacement suits the parts that wear out.
- Test the change before buying it. In xellstorm, add a scenario that changes the seal’s distribution, for example to the shape and scale of a supplier’s longer-lived seal, and compare the chance of a failure within a year. Scenarios use the same random draws as the base case, so the difference comes from the change rather than from sampling noise.
- Read other questions off the S-curve. It gives the chance that a new unit has failed by any age: the longest warranty in which at most 10% of new units fail is the B10 life, 1,091 hours.
What the model assumes
The four lives are independent, every part starts new, and the unit runs around the clock. The chance on this page is that of a first stop within the first year of a new unit: after a repair, the other parts are no longer new, so the number of stops per year needs a model of repairs, which a workbook can hold in more cells. A shared cause, such as contamination that shortens the lives of the seal and the bearing together, can be modeled with a rank correlation between their inputs. Shapes and scales come from failure data and are themselves uncertain; if you have observed lifetimes of failed parts, the From data tab next to each input on the Distributions step fits a Weibull, among other families, by maximum likelihood. It treats every value as a completed life and cannot use units that are still running (censored data). Fitting only the failures then understates the lives, badly when many units are still running, so in that case estimate the shape and scale with a method that handles suspensions and type them in.
Try it yourself
- Open the model in xellstorm. The four Life cells are inputs, D7 and D8 are the outputs, and a system life under 8,760 hours is the target; there is nothing to install and no sign-up, and the workbook is calculated in your browser.
- Run it: with the same seed and 20,000 trials you get the numbers on this page. For system life, the results show the mean, the chance of the target (79.0%), the P10 (1,091) and the average of the shortest 5% of lives; the percentile table lists P1 to P99.
- Hover over the S-curve to read the chance that a unit has failed by any number of hours, or type a period into the probability box, for example a warranty or a planned overhaul interval, to see the chance of a failure before it.
- On the Distributions step, change the seal’s scale or add a scenario for a better seal, and run again to see how the chance of a failure within a year moves.
- To have the simulation read the shapes and scales from the sheet, type
=B2in the pump’s shape field and=C2in its scale field, and so on for each row. Then open your own reliability model and click the cells you are unsure of on the model map.
Can I do this in plain Excel?
The exact answer, yes: for a series system with Weibull lifetimes, =1-EXP(-SUMPRODUCT((8760/C2:C5)^B2:B5)) in this workbook gives the chance of a failure within a year, and =WEIBULL.DIST(8760,B4,C4,TRUE) gives the seal’s chance on its own. A simulation needs a random lifetime in each Life cell, =C2*(-LN(1-RAND()))^(1/B2) (the inverse of the Weibull distribution function), then a data table to repeat the calculation and formulas to summarize it. Monte Carlo simulation in Excel shows the method. The exact formula stops working as soon as the model does more than take the first failure, such as a redundant pump or a repair time and cost per stop; a simulation of the workbook handles whatever its formulas compute.
Questions
What is a series system in reliability?
A series system is one that stops when any of its components fails, like a chain that breaks at its weakest link. Its life is the shortest of the component lives, and, with independent components, its chance of surviving a period is the product of the components’ survival chances, so it is never more reliable than its least reliable part.
What do the Weibull shape and scale mean?
The Weibull scale is the age by which 63.2% of parts have failed. The shape describes how the failure rate changes with age: below 1 it falls (early failures), at 1 it is constant (random failures), and above 1 it rises (wear-out). In scipy and in xellstorm’s setup files the distribution is weibull_min, with the shape as c, and the app labels the fields shape and scale; Excel’s WEIBULL.DIST takes the shape as alpha and the scale as beta.
What is the B10 life?
The B10 life is the age by which 10% of a population of units has failed, the 10th percentile of the lifetime. The measure is common in bearing ratings, where it is called L10, and in warranty planning. For the four-part unit in this example it is 1,091 hours simulated and 1,086 hours exact, against a mean life of 5,696 hours.
Why simulate when there is an exact formula?
The exact formula for a series system holds only while the model is a series system of independent parts with known lifetime distributions. Real reliability models add redundancy, spares, repair times, costs of downtime or correlated failures, and then the formula no longer applies. This example is kept simple so that the simulation can be checked against the formula; the same simulation works unchanged on a workbook that computes much more.
How many trials does a reliability simulation need?
The noise in a simulated probability shrinks with the square root of the number of trials: with 20,000 trials, the standard error of the chance of a failure within a year (79.0%) is 0.29 percentage points. That is the standard error of plain random sampling; Latin Hypercube sampling, the default, usually does a little better, and never worse when the result moves only one way with each input, as here. Rare events, such as a failure within the first few days, need many more trials than common ones. xellstorm can also keep running until the mean and percentiles settle within a tolerance you choose.
Related
- Mechanical design
Tolerance stack-up analysis: worst case, RSS or Monte Carlo?
A housing and five parts: worst case lets the gap shrink to 0.120 mm; RSS predicts 3.29% of gaps below 0.4 mm and 20,000 Monte Carlo trials find 3.21%. - Guide
Tornado charts and sensitivity analysis explained
A tornado chart moves each input alone, from P10 to P90 by default in xellstorm. In a building estimate, one risk moves the cost by 150 (USD thousands). - Guide
Monte Carlo simulation in Excel, with or without add-ins
Three ways to run a Monte Carlo simulation on an Excel model: RAND() and a data table, an add-in, or a browser tool with no add-in. PERT formula included.
xellstorm is a browser-based Monte Carlo simulation tool for Excel models: no add-in, and the workbook never leaves your computer.