Viser opslag med etiketten RAMS. Vis alle opslag
Viser opslag med etiketten RAMS. Vis alle opslag

fredag den 12. februar 2010

Quantitative Risk Analysis


In some situations the qualitative risk analysis or the ALARP principle is insufficient: The safety people are torn and disagrees internally. Consequently, it is time to use the heavier "quantitative risk analysis"-tool.
The fault tree is integrated into Excel and models a scenario, where a passenger is trapped between closing doors. (All numbers and technical barriers are hypothetical).



Interpretation

The quantitative risk analysis is the right way to estimate the frequency of a hazard.
It removes personal obsessions from a safety problem and ensures that the discussions are conducted on an objective basis.
The fault tree above concerns a commuter fleet operating 365 days pr. year with 80 trains with 100 departures pr train pr. day. This result in c = 2.9E-06 departures every year pr. fleet.
In order to have an accident, there have to be squeezed a passenger arm, leg or items like e.g. a baby carriage, umbrella etc. between the closing doors. This is judged to happen continuously when passengers passes the doors, meaning d = 1.
There are three barriers that prevent the hazard:
A human based departure procedure, where the train driver looks out of the window and checks the doors before departure (e). It is estimated that the driver miss a check every 4'Th day due to distraction or lacking of concentration, meaning e = 1/(4*b).
There are also two technical functions:
- A traction blocking that prevents the train from driving if the door controllers indicate the doors are open (f). This function is part of the train computer and is expected to be reliable with a failure rate of 1 failure pr. 1,000,000 departures.
- A trap detection system in the door controller that prevents the passengers from being squeezed in a closing door (g). This function is sensitive to door mechanics; the FRACAS system indicates a failure rate of 1 failure pr. 10,000 departures.
As it can be seen we will end up having an accident where the train departs with a passenger trapped between doors every year. The Safety department has recorded an incident the recent year, indicating the fault tree is trustable.
Can we accept this? What are our quantitative acceptance criterion? It should be written and stated in the safety management system of the Operator.
The safety management now decides that the above result is unacceptable. We can only allow the hazard to occur every 10,000'End year.
A deeper analysis shows that the failures on the detection function only occurs for thin objects like a small child's arm.
It is judged that the detection system in the daily life is activated by large objects like a person; thin objects only occurs 3 times pr. day, meaning d = 3/b.
The sensors are adjusted and a maintenance program introduced; the following test result shows an improved reliability in the area of 1 failure pr. 100,000 departures (g).
An additional departure procedure is introduced stating the train conductor has to supervise the train doors before departure in front of a dedicated door. A new technical feature makes it possible to firstly close the other doors and finally, the train conductor enters the last door before departure. The improved procedure is expected to be more reliable with an estimated human failure rate of 1 failure pr. 10,000 departures (e).
These mitigating actions result in a dramatically lowering of the frequency of the hazard to app. 10,000 years between accidents, hereby fulfilling the acceptance criterion.
  

As a side effect, the analysis proves the importance of the departure procedure and the detection function.

The old rule, KISS, (Keep It Simple Stupid) is recommended for quantitative analysis. The fault trees easily swell up into large trees with several undocumented values based on engineering judgement. This only starts new discussions instead.

Next chapter >> 4.5 Common Cause Failures (CCF)

Focus on the Source

The "Guide to the application of EN 50126-1 for safety", TR 50126-2: Feb. 2007, concerns risk modelling and quantitative risk models.

Chapter 5.2, "Generic Risk Model" says:

Modelling predominantly represents a simplification and generalisation of reality but, enhances our understanding of causal relationships, highlights important factors and provides a useful tool for anticipation and potentially prediction of future.
A risk model may be created for a specific task (e.g., occurrence of a hazard, a combination of hazards, an operation, a sub-system, etc.) for a particular application or for a whole railway system by applying the risk assessment process to the relevant task or to the railway system.
[.....]
Developing a risk model for a whole railway system is a demanding task [....] the report does not recommend a single generic risk model for a whole railway system. [....]
Annex D lists essential steps for building such a model [....]


Read more...

søndag den 29. november 2009

Common Cause Failures (CCF)


Special precautions have to be taken against common cause failures.

It is a single failure that causes a safety function to collaps e.g. a mechanical or logic error in a product as shown below in Figure A.7 from EN 50129.

It can be handled by using redundant systems, inherited charactheristics of components, safety analysis, independent reviews, FRACAS system, etc.



Interpretation

Common causes failures can e.g. be a sleeping tricky error in Function A that cause a dramatic failure in Function B.
If we have installed hundreds of systems we have a possibly accident.

Train fleet example:

Let’s say the developer of a diesel traction system in a train uses the exhaust gas to power a turbo. The turbo powers an air inlet compressor. The compressed air enters the combustion chamber.
A hose clamp on the air tubes are under dimensioned, nevertheless the design passes design reviews and burn-in tests.
The hose clamp is slowly loosened during operation and this causes a decrease of air in the combustion chamber that again causes an overheated exhaust gas that again causes the turbo to overheat and crack and finally cause an oil leakage in the turbo driven power transmission to the compressor located near the exhaust pipe.
The operational staff reports of occasional small fires in the turbo driven power transmission, the maintenance staff discovers the cracked turbo, and it is concluded that cracked turbo's must be changed.
In this case we have an undisclosed common cause failure in the train fleet (the loose hose clamp).
One day, under the right circumstances, the oil leakage will cause a larger fire. If the daily train route furthermore passes a tunnel we might end up with a "fire in train in tunnel" scenario.

Interlocking logic example:

See "Quick guide to safety management based on EN50126"

Next chapter >> 4.6 Safety Integrity Levels (SIL)

Focus on the Source

See "Quick guide to safety management based on EN50126"



Read more...

fredag den 6. november 2009

Failure Reporting And Corrective Action System (FRACAS)


Once the operation starts, the product enters phase 12, “Performance monitoring” in the V-model. In this phase it is time to implement a monitoring system.
If e.g. the maintenance staff discovers that a certain type of points tend to have loose bolts then there have to be an office, guard, database or other report system where the incident can be reported.
There also have to be somebody in the organization that reads this report and takes appropriate corrective action.


Potters Bar accident in UK, 2002, due to loose bolts in a point

Interpretation

It sounds easy; but investigation reports from accidents and “near miss” incidents often shows that the implemented FRACAS did not work properly: The points were poorly maintained, the failure report was shelved, the engineers misjudged data, the supplier never fixed it, the appropriate procedure was not updated or the purchasers lacked time to buy new bolts.

Regularly audits are a suitable tool for examining the implemented FRACAS system.

Statistics is a helpful tool, because it removes emotions from a problem and forces the safety management to take action.

Next chapter >> 4.3 Using the ALARP principle

Focus on the Source

Chapter 6.12 , “Phase 12: Performance monitoring” in EN 50126:1999 describes the objectives and requirements to this phase:

6.12.3 Requirements
6.12.3.1 Requirement 1 of this phase shall be to establish, implement and regularly review a process for:
- the collection of operational performance and RAMS statistics;
-the acquisition, analysis and evaluation of performance and RAMS data; checking that the assumptions made in the safety case remain valid.
6.12.3.2 Requirement 2 of this phase shall be to analyse performance and RAMS data and statistics to influence:
- new operating and maintenance procedures;
- changes in logistic support for the system.



Read more...

lørdag den 16. august 2008

RAMS and how to control it


EN 50126 is all about controlling the RAMS parameters of a Railway system (e.g. a complete train, an LED lamp etc.).

It appears directly from the title: “Railway applications - The specification and demonstration of Reliability, Availability, Maintainability and Safety (RAMS)”.
The RAMS parameters are linked as shown at Figure 2:



Interpretation
The RAMS parameters are useful when categorizing different items e.g.:
  • the requirements to and specifications of the system
  • faults and findings during design and service.
This way, all parties (Operator, Supplier, Safety Authority, Assessor) know what we are talking about if e.g. an error is disclosed during testing: Is the fault a Reliability issue, an Availability issue, a Maintainability problem or a Safety problem.

Lets say we have a new train ready and approved for operation, but some errors exists. The errors have been categorized as Reliability issues, which are not directly safety-related. In this case Figure 2 above would look as shown on the left:

The yellow "Reliability" in the bottom will cause a Yellow "Availability" in the middle, which again will cause a yellow top level "Railway RAMS".


Since we have a green "Maintainability" in the bottom, it might be possible to increase the "Maintenance" work and hereby compensate for the yellow "Reliability", so we obtain a green "Availability", which again will cause a green top level "Railway RAMS". See the Figure below.

This is controlling RAMS!






Next chapter >> 2.2 The V-model






From the Source (EN 50126)

The links are described more detailed in chapter 4.3.2 and 4.3.3 in EN 50126:1999:

"Safety and availability are inter-linked in the sense that a weakness in either or mismanagement of conflicts between safety and availability requirements may prevent achievement of a dependable system. The inter-linking of railway RAMS elements, reliability, availability, maintainability and safety is shown in figure 2."
"Attainment of in-service safety and availability targets can only be achieved by meeting all reliability and maintainability requirements and controlling the ongoing, long-term, maintenance and operational activities and the system environment."

A more elaborated version of Figure 2 is given in Figure 5 (not shown here).



Read more...