هندسة الوثوقية
تعد وثوقية النظم system reliability من الاهتمامات الرئيسة لهندسة الوثوقية reliability engineering، وهي أحد فروع العلوم الهندسية الحديثة. ابتدأ الاهتمام بهندسة الوثوقية خلال الحرب العالمية الثانية، وتطور هذا العلم في العقود الأخيرة خصوصاً مع تطور المنظومات الكهربائية والإلكترونية والميكانيكية. إن خبرة استثمار تلك المنظومات الهندسية واستخدامها بيّنت أنه مهما كانت إجراءات تحسين جودة المنتج، فإن الحقيقة التي يجب التعايش معها هي أن هذه المنظومات معرضة للتعطل وفقدان القدرة على القيام بالمهام والوظائف المخصصة لأجلها خلال فترة ما في أثناء الاستثمار، كما أن ظهور بعض الأعطال الحرجة خلال العمل يمكن أن يؤدي إلى خسائر مادية وبشرية كبيرة يمكن أن تؤدي في بعض الأحيان إلى حوادث مؤسفة أو كوارث.
نظرة عامة

تهتم هندسة الوثوقية بدراسة أعطال النظم واحتمال حدوثها وحديتها وتأثيرها في وثوقية النظم والسلامة والبيئة، وتواجه الشركات تحديات كبيرة لتحسين وثوقية منتجاتها ورفعها؛ لكي تفي بمتطلبات الزبون ورضاه، وتقلل كلفة الضمان والصيانة في فترة الاستثمار.
تستخدم مجموعة من التعابير والمصطلحات في هندسة الوثوقية وهندسة النظم، حيث يطلق مصطلح المكوّن item على أي مستوى من مستويات النظم وخصوصاً المنظومات الإلكترونية (عنصر component ـ دارة circuit ـ كتلة module ـ نظام جزئي subsystem ـ نظام system ـ منظومة super system).
نظرية الوثوقية
وتعابير هندسة الوثوقية فأهمها ما يأتي:
الوثوقية reliability: هي قابلية مكوّن ما تنفيذ وظائفه المطلوبة تحت شروط محددة فترة زمنية معينة. أو هي احتمال نجاح مكوّن ما بتنفيذ وظائفه المطلوبة من دون عطل تحت شروط محددة فترة زمنية معينة.
العطل (الفشل) failure: ويعرف بأنه توقف مكوّن ما عن تنفيذ أحد وظائفه المطلوبة.
معدل الأعطال :failure rate هو عدد الأعطال الكلي لمكون ما مقسوماً على عدد الوحدات الزمنية (غالباً الساعة) خلال فترة القياس تحت شروط محددة.
نمط العطل failure mode: هو الشكل أو الطريقة التي يتعطل بها مكوّن ما عن تنفيذ وظائفه.
قابلية الإصلاح maintainability: هي خاصية المكوّن في التنبيه على الأعطال والكشف عن أسبابها وإزالتها عن طريق الإصلاح أو الخدمات الفنية وإرجاع المكوّن إلى تنفيذ وظائفه.
الجاهزية availability: هي قابلية مكوّن (تحت مجموعة عوامل من وثوقية وإصلاح) ليقوم بتنفيذ وظائفه المطلوبة تحت شروط محددة فترة زمنية معينة.
الزمن الوسطي بين الأعطال mean time between failure (MTBF): هو مقياس لوثوقية مكوّن ما قابل للإصلاح، وهو القيمة الوسطى لفترة كافية من الزمن لحدوث عدة أعطال، وتمثلها النسبة بين ذلك الزمن على عدد الأعطال خلال تلك الفترة من زمن الحياة.
الوثوقية خلال دورة حياة المنتج
خطة برنامج الوثوقية
تقوم فعاليات هندسة الوثوقية بدور أساسي في دورة حياة المنتج ابتداءً من دراسة الجدوى، مروراً بالتصميم ثم الإنتاج وانتهاءً بالاستثمار. ويتم عادة دراسة الوثوقية في أربع مراحل كما يأتي:
1 ـ الوثوقية في مرحلة دراسة الجدوى: يجري خلال هذه المرحلة تحديد برنامج الوثوقية اللازم للمنتج وفقاً لمتطلبات الزبون والبيئة التي سيعمل بها المنتج أو النظام.
2 ـ الوثوقية في مرحلة التطوير والتصميم: تهتم فعاليات الوثوقية في أثناء مرحلة التطوير والتصميم بالتنبؤ والتحليل والقياس لوثوقية المنتج للوصول إلى الوثوقية المطلوبة له.

أ ـ التنبؤ بالوثوقية reliability prediction:
يعتمد التنبؤ بالوثوقية لمكوّن ما على معدل الأعطال الذي بينت الحياة العملية أنه تابع للزمن ، ويتألف من ثلاث مراحل:
¬ مرحلة الطفولة التي يتناقص فيها معدل الأعطال.
¬ مرحلة الحياة المفيدة يثبت بها معدل الأعطال (يرمز له بـ λ).
¬ مرحلة التقادم ويزداد فيها معدل الأعطال.
وتعطى الوثوقية في فترة الحياة المفيدة لمكوّن ما R(t) كتابع أسي لزمن المهمة (t) ولمعدل الأعطال (λ) على الشكل الآتي:
R (t) = e-λt… (1)
كما يحسب الزمن الوسطي بين الأعطال (MTBF) بدلالة معدل الأعطال على النحو الآتي:
MTBF = 1/ λ … (2)
يتعلق معدل الأعطال λ بعدة عوامل أهمها الحرارة والبيئة وجودة المكوّن والإجهاد المتعلق بالتصميم. يقاس معدل الأعطال «بعدد الأعطال في وحدة زمنية»، ويؤخذ معدل الأعطال والعوامل المؤثرة فيه من مصادر معلومات إحصائية عديدة أهمها المعيار الأمريكي MIL-HDBK-217 وBellcore TR-332.
يستخدم المخطط الصندوقي للوثوقية reliability block diagram للتنبؤ بوثوقية نظام مؤلف من عدة مكونات، ويعد المخطط الصندوقي للوثوقية من أهم الطرق لنمذجة الوثوقية reliability modeling ويتكون من صناديق مرتبطة على التسلسل أو التفرع، ويمثل كل صندوق وثوقية كل جزء من النظام. ويأخذ المخطط الصندوقي أحد شكلين:
Rsys =R1 Rn (3)x*... * R3 * R2 *
λsys =λ1 + λ2 + λ3 + ... + Rn (4)

ـ الربط التسلسلي: ترتبط فيها الصناديق بشكل متتالٍ ولكي يعمل النظام بنجاح يجب أن تعمل جميع مكوناته بنجاح، ويخفق النظام إذا أخفق واحد على الأقل من مكوناته. وتكون وثوقية النظام Rsys أصغر من وثوقية أي من مكوناته وتساوي جداء وثوقية مكوناته، ويكون معدل الأعطال للنظام λsys مساوياً مجموع معدل أعطال مكوناته.
ـ الربط التفرعي: ترتبط فيها الصناديق بشكل تفرعي ويعمل النظام بنجاح إذا عمل بنجاح واحد على الأقل من مكوناته، ويخفق النظام إذا أخفقت جميع مكوناته. ويمثل هذا الربط استخدام المكونات الرديفة أو الاحتياط redundancy وتكون وثوقية النظام أكبر من وثوقية مكوناته وتحسب وثوقية النظام Rsys على النحو المبين في الشكل (3).
Rsys = 1- (1-R1)* (1-R2) * (1-R3)*… *(1-Rn) (5)

ـ الربط المختلط: ترتبط فيها الصناديق بشكل تسلسلي وتفرعي، ويتم الحساب والتنبؤ بوثوقية النظام اعتماداً على المعادلات (1), (3), (5). يجري وضع المخطط الصندوقي للنظام ومكوناته هرمياً من الأعلى إلى الأسفل أو بالعكس .

ب ـ تحليل الوثوقية reliability analysis
تستخدم مجموعة من أدوات تحليل الوثوقية خلال مرحلة التطوير والتصميم لتجنب الحوادث والمشكلات، وذلك من خلال تحليل الأعطال ومسبباتها وعلاقتها مع بعضها واحتمال حدوثها وتأثيراتها في وثوقية النظام والسلامة والبيئة. من أهم أدوات تحليل الوثوقية: تحليل آثار وحراجة أنماط العطل failure mode effects and criticality analysis (FMECA)، و تحليل شجرة الأعطال fault tree analysis (FTA).
تحليل آثار وحراجة أنماط العطل FMECA
تحليل FMECA هو منهجية تحليل الأعطال من الأسفل إلى الأعلى bottom-up أي من مستوى العنصر إلى مستوى النظام. يهدف التحليل إلى تحديد وتحليل الأعطال الخاصة المحتملة الكامنة potential failure للمكونات وتأثيرها في الوثوقية ودرجة حرجيتها على السلامة، وكيفية تجنب تأثير هذه الأعطال في إخفاق النظام، وتحديد الإجراءات التصحيحية للتخلص من تلك الأعطال الناجمة عن التصميم أو عن عمليات التصنيع والإنتاج قبل تسليم المنتج إلى الزبون.
تطبق FMECA مبكراً من مرحلة التطوير والتصميم وتحدّث مع تطور التصميم، ويجري تنفيذ إجراءات FMECA وتوثيقها بجداول وفقاً لمعايير عالمية عديدة من بينها المعيار العسكري الأمريكي MIL-STD-1629 والمعيار البريطاني BS 5760-part 5.[1]
تحليل شجرة الأعطال FTA
تحليل شجرة الأعطال FTA هو منهجية تحليل من الأعلى إلى الأسفل top-down أي من مستوى إخفاق حرج للنظام إلى مستوى إخفاق العنصر. يقوم هذا التحليل بتبيان العلاقة بين الأعطال ومسبباتها التي تؤدي إلى إخفاق حرج للنظام وحساب احتمال حدوث ذلك الإخفاق الذي يمكن أن يؤدي إلى حادث أو كارثة.
يتألف تحليل شجرة الأعطال من ثلاث مراحل:
¬ بناء شجرة الأعطال، وهي رسم يمثل العلاقة بين الأعطال ببوابات منطقية AND-OR من مستوى إخفاق النظام الحرج ويسمى الحدث الأعلى top event إلى مستوى إخفاق العنصر الذي يسمى الحدث الأساسي basic event.
¬ التحليل النوعي qualitative analysis الذي يتضمن تحديد مجموعات القص الصغرى minimal cut sets للأعطال التي إذا حدثت يحدث الحدث الأعلى.
¬ التحليل الكمي quantitative analysis الذي يتضمن حساب احتمال حدوث الحدث الأعلى.
ج ـ قياس الوثوقية reliability measuring
الغاية من قياس الوثوقية هو تحديد فيما إذا كان المنتج في الشكل النهائي قد حقق متطلبات الوثوقية وتحديد نقاط الضعف التي لا تزال غير مغطاة وتحتاج إلى تصحيح، والوصول إلى قناعة عالية المستوى بأن متطلبات الزبون قد تم تحقيقها. إذا كانت نتائج القياس غير مقبولة فإن قبول المنتج أو تحريره للإنتاج يتأجل إلى حين اتخاذ إجراء تصحيحي. من أهم وسائل قياس الوثوقية مجموعة من الاختبارات المحيطية environmental Tests مثل:
¬ اختبار الحياة المسرّعة accelerated life test يضمن قياس الوثوقية في زمن اختبار قصير.
¬ اختبار رفع الوثوقية reliability growth test (RGT).
¬ اختبر وحلل وصحح test analyze and fix (TAAF) لكشف الأعطال وتصحيحها.
د ـ مراجعة التصميم design review
تقوم لجنة خاصة بمراجعة التصميم للتحقق من أن التصميم المنفذ يحقق متطلبات الزبون من ناحية الأداء والوثوقية. تتألف لجنة مراجعة التصميم من عدة اختصاصات منها: هندسة نظام، هندسة التصميم، هندسة الوثوقية، هندسة الإنتاج، التسويق وشؤون الاستثمار.
تجرى المراجعة التصميمية عدة مرات خلال مراحل تطور النظام، وتسمح بتحرير النظام المصمم إلى مرحلة الإنتاج الكمي في حال تحقيقه لمتطلبات الزبون، وتحرير المنتجات المصنعة إلى الزبون أو الأسواق بنهاية مرحلة الإنتاج والتصنيع.
3ـ الوثوقية في مرحلة الإنتاج والتصنيع
تتعرض وثوقية المنتج للانخفاض في أثناء مرحلة الإنتاج والتصنيع لعدة أسباب من العيوب أهمها جودة المواد الأولية وتباينها، التباين الناتج من عمليات التصنيع والتجميع. تهدف فعاليات الوثوقية والجودة quality في المصانع إلى المحافظة على وثوقية المنتج التي تم التوصل إليها في أثناء التصميم.
من أهم فعاليات الوثوقية والجودة في المصنع على مستوى المواد الأولية والمواد المصنعة والتجميع هي الآتية:
¬ التفتيش inspection من خلال التفتيش عن المكونات المعيبة الأولية والإنتاجية.
¬ ضبط الوثوقية والجودة reliability/quality control من خلال ضبط العمليات الإنتاجية.
¬ اختبارات الغربلة screen test and burn-in وهي اختبارات محيطية للتخلص من الأعطال الطفولية في المنتجات، وتسليم المنتج إلى الزبون بالوثوقية ومعدل الأعطال المطلوبة.
¬ اختبارات القبول acceptance tests وهي اختبارات الأداء والوثوقية لقبول الزبون المنتج المتفق عليه مسبقاً.
4ـ الوثوقية في مرحلة الاستثمار
تهتم فعاليات الوثوقية والجودة بمرحلة الاستثمار أيضاً من خلال ضبط عمليات التغليف والتخزين والنقل، وفترة الضمان وخدمات الاستثمار من تدريب وصيانة وإصلاح. كما تقوم تلك الفعاليات بجمع المعلومات والملاحظات وتقوّمها للتحقق من أن وثوقية النظام قد طابقت المواصفات المطلوبة، كما تتابع الإجراءات التصحيحية للعيوب التي قد تظهر وتقدم تقاريرها إلى إدارة هندسة الوثوقية.
تتعلق مشاركة بعض فعاليات هندسة الوثوقية أو غالبيتها أو جميعها ببرنامج الوثوقية المطلوب بسبب الكلفة ووفقاً لأهمية النظام أو المنتج، مثلاً: المركبات الفضائية ، الطائرات المدنية والعسكرية، منظومات الطائرات من دون طيار ، محطات الطاقة , القطارات ، أجهزة طبية للتشخيص والعمل الجراحي وجميعها يتطلب وثوقية عالية، في حين أن الأجهزة المنزلية مثل المذياع «الراديو» والتلفاز والحاسوب الشخصي والأجهزة الهاتفية عمرها الاستثماري محدود ومن ثم تحتاج إلى مشاركة أقل من فعاليات هندسة الوثوقية.
الوثوقية الهندسية ضد السلامة الهندسية
توقع الوثوقية وتحسينها
Reliability prediction combines:
- creation of a proper reliability model (see further on this page)
- estimation (and justification) of input parameters for this model (e.g. failure rates for a particular failure mode or event and the mean time to repair the system for a particular failure)
- estimation of output reliability parameters at system or part level (i.e. system availability or frequency of a particular functional failure) The emphasis on quantification and target setting (e.g. MTBF) might imply there is a limit to achievable reliability, however, there is no inherent limit and development of higher reliability does not need to be more costly. In addition, they[من؟] argue that prediction of reliability from historic data can be very misleading, with comparisons only valid for identical designs, products, manufacturing processes, and maintenance with identical operating loads and usage environments. Even minor changes in any of these could have major effects on reliability. Furthermore, the most unreliable and important items (i.e. the most interesting candidates for a reliability investigation) are most likely to be modified and re-engineered since historical data was gathered, making the standard (re-active or pro-active) statistical methods and processes used in e.g. medical or insurance industries less effective. Another argument is that to be able to accurately predict reliability by testing, the exact mechanisms of failure must be known and therefore – in most cases – could be prevented. Following the incorrect route of trying to quantify and solve a complex reliability engineering problem in terms of MTBF or probability using an-incorrect – for example, the re-active – approach is referred to by Barnard as "Playing the Numbers Game" and is regarded as bad practice.[2]
For existing systems, it is arguable that any attempt by a responsible program to correct the root cause of discovered failures may render the initial MTBF estimate invalid, as new assumptions (themselves subject to high error levels) of the effect of this correction must be made. Another practical issue is the general unavailability of detailed failure data, with those available often featuring inconsistent filtering of failure (feedback) data, and ignoring statistical errors (which are very high for rare events like reliability related failures). Very clear guidelines must be present to count and compare failures related to different type of root-causes (e.g. manufacturing-, maintenance-, transport-, system-induced or inherent design failures). Comparing different types of causes may lead to incorrect estimations and incorrect business decisions about the focus of improvement.
To perform a proper quantitative reliability prediction for systems may be difficult and very expensive if done by testing. At the individual part-level, reliability results can often be obtained with comparatively high confidence, as testing of many sample parts might be possible using the available testing budget. However, these tests may lack validity at a system-level due to assumptions made at part-level testing. These authors emphasized the importance of initial part- or system-level testing until failure, and to learn from such failures to improve the system or part. The general conclusion is drawn that an accurate and absolute prediction – by either field-data comparison or testing – of reliability is in most cases not possible. An exception might be failures due to wear-out problems such as fatigue failures. In the introduction of MIL-STD-785 it is written that reliability prediction should be used with great caution, if not used solely for comparison in trade-off studies.
Design for reliability
Design for Reliability (DfR) is a process that encompasses tools and procedures to ensure that a product meets its reliability requirements, under its use environment, for the duration of its lifetime. DfR is implemented in the design stage of a product to proactively improve product reliability.[3] DfR is often used as part of an overall Design for Excellence (DfX) strategy.
Statistics-based approach (i.e. MTBF)
Reliability design begins with the development of a (system) model. Reliability and availability models use block diagrams and Fault Tree Analysis to provide a graphical means of evaluating the relationships between different parts of the system. These models may incorporate predictions based on failure rates taken from historical data. While the (input data) predictions are often not accurate in an absolute sense, they are valuable to assess relative differences in design alternatives. Maintainability parameters, for example Mean time to repair (MTTR), can also be used as inputs for such models.
The most important fundamental initiating causes and failure mechanisms are to be identified and analyzed with engineering tools. A diverse set of practical guidance as to performance and reliability should be provided to designers so that they can generate low-stressed designs and products that protect, or are protected against, damage and excessive wear. Proper validation of input loads (requirements) may be needed, in addition to verification for reliability "performance" by testing.

One of the most important design techniques is redundancy. This means that if one part of the system fails, there is an alternate success path, such as a backup system. The reason why this is the ultimate design choice is related to the fact that high-confidence reliability evidence for new parts or systems is often not available, or is extremely expensive to obtain. By combining redundancy, together with a high level of failure monitoring, and the avoidance of common cause failures; even a system with relatively poor single-channel (part) reliability, can be made highly reliable at a system level (up to mission critical reliability). No testing of reliability has to be required for this. In conjunction with redundancy, the use of dissimilar designs or manufacturing processes (e.g. via different suppliers of similar parts) for single independent channels, can provide less sensitivity to quality issues (e.g. early childhood failures at a single supplier), allowing very-high levels of reliability to be achieved at all moments of the development cycle (from early life to long-term). Redundancy can also be applied in systems engineering by double checking requirements, data, designs, calculations, software, and tests to overcome systematic failures.
Another effective way to deal with reliability issues is to perform analysis that predicts degradation, enabling the prevention of unscheduled downtime events / failures. RCM (Reliability Centered Maintenance) programs can be used for this.
Physics-of-failure-based approach
For electronic assemblies, there has been an increasing shift towards a different approach called physics of failure. This technique relies on understanding the physical static and dynamic failure mechanisms. It accounts for variation in load, strength, and stress that lead to failure with a high level of detail, made possible with the use of modern finite element method (FEM) software programs that can handle complex geometries and mechanisms such as creep, stress relaxation, fatigue, and probabilistic design (Monte Carlo Methods/DOE). The material or component can be re-designed to reduce the probability of failure and to make it more robust against such variations. Another common design technique is component derating: i.e. selecting components whose specifications significantly exceed the expected stress levels, such as using heavier gauge electrical wire than might normally be specified for the expected electric current.
Common tools and techniques
Many of the tasks, techniques, and analyses used in Reliability Engineering are specific to particular industries and applications, but can commonly include:
- Physics of failure (PoF)
- Built-in self-test (BIT or BIST) (testability analysis)
- Failure mode and effects analysis (FMEA)
- Reliability hazard analysis
- Reliability block-diagram analysis
- Dynamic reliability block-diagram analysis[4]
- Fault tree analysis
- Root cause analysis
- Statistical engineering, design of experiments – e.g. on simulations / FEM models or with testing
- Sneak circuit analysis
- Accelerated testing
- Reliability growth analysis (re-active reliability)
- Weibull analysis (for testing or mainly "re-active" reliability)
- Hypertabastic survival models
- Thermal analysis by finite element analysis (FEA) and / or measurement
- Thermal induced, shock and vibration fatigue analysis by FEA and / or measurement
- Electromagnetic analysis
- Avoidance of single point of failure (SPOF)
- Functional analysis and functional failure analysis (e.g., function FMEA, FHA or FFA)
- Predictive and preventive maintenance: reliability centered maintenance (RCM) analysis
- Testability analysis
- Failure diagnostics analysis (normally also incorporated in FMEA)
- Human error analysis
- Operational hazard analysis
- Preventative/Planned Maintenance Optimization (PMO)
- Manual screening
- Integrated logistics support
Results from these methods are presented during reviews of part or system design, and logistics. Reliability is just one requirement among many for a complex part or system. Engineering trade-off studies are used to determine the optimum balance between reliability requirements and other constraints.
The importance of language
Reliability engineers, whether using quantitative or qualitative methods to describe a failure or hazard, rely on language to pinpoint the risks and enable issues to be solved. The language used must help create an orderly description of the function/item/system and its complex surrounding as it relates to the failure of these functions/items/systems. Systems engineering is very much about finding the correct words to describe the problem (and related risks), so that they can be readily solved via engineering solutions. Jack Ring said that a systems engineer's job is to "language the project." (Ring et al. 2000)[5] For part/system failures, reliability engineers should concentrate more on the "why and how", rather that predicting "when". Understanding "why" a failure has occurred (e.g. due to over-stressed components or manufacturing issues) is far more likely to lead to improvement in the designs and processes used[6] than quantifying "when" a failure is likely to occur (e.g. via determining MTBF). To do this, first the reliability hazards relating to the part/system need to be classified and ordered (based on some form of qualitative and quantitative logic if possible) to allow for more efficient assessment and eventual improvement. This is partly done in pure language and proposition logic, but also based on experience with similar items. This can for example be seen in descriptions of events in fault tree analysis, FMEA analysis, and hazard (tracking) logs. In this sense language and proper grammar (part of qualitative analysis) plays an important role in reliability engineering, just like it does in safety engineering or in-general within systems engineering.
Correct use of language can also be key to identifying or reducing the risks of human error, which are often the root cause of many failures. This can include proper instructions in maintenance manuals, operation manuals, emergency procedures, and others to prevent systematic human errors that may result in system failures. These should be written by trained or experienced technical authors using so-called simplified English or Simplified Technical English, where words and structure are specifically chosen and created so as to reduce ambiguity or risk of confusion (e.g. an "replace the old part" could ambiguously refer to a swapping a worn-out part with a non-worn-out part, or replacing a part with one using a more recent and hopefully improved design).
اختبار الموثوقية

The purpose of reliability testing or reliability verification is to discover potential problems with the design as early as possible and, ultimately, provide confidence that the system meets its reliability requirements. The reliability of the product in all environments such as expected use, transportation, or storage during the specified lifespan should be considered.[7] It is to expose the product to natural or artificial environmental conditions to undergo its action to evaluate the performance of the product under the environmental conditions of actual use, transportation, and storage, and to analyze and study the degree of influence of environmental factors and their mechanism of action.[8] Through the use of various environmental test equipment to simulate the high temperature, low temperature, and high humidity, and temperature changes in the climate environment, to accelerate the reaction of the product in the use environment, to verify whether it reaches the expected quality in R&D, design, and manufacturing.[9]
Reliability verification is also called reliability testing, which refers to the use of modeling, statistics, and other methods to evaluate the reliability of the product based on the product's life span and expected performance.[10] Most product on the market requires reliability testing, such as automotive, integrated circuit, heavy machinery used to mine nature resources, Aircraft auto software.[11][12]
Reliability testing may be performed at several levels and there are different types of testing. Complex systems may be tested at component, circuit board, unit, assembly, subsystem and system levels.[13] (The test level nomenclature varies among applications.) For example, performing environmental stress screening tests at lower levels, such as piece parts or small assemblies, catches problems before they cause failures at higher levels. Testing proceeds during each level of integration through full-up system testing, developmental testing, and operational testing, thereby reducing program risk. However, testing does not mitigate unreliability risk.
With each test both statistical type I and type II errors could be made, depending on sample size, test time, assumptions and the needed discrimination ratio. There is risk of incorrectly rejecting a good design (type I error) and the risk of incorrectly accepting a bad design (type II error).
It is not always feasible to test all system requirements. Some systems are prohibitively expensive to test; some failure modes may take years to observe; some complex interactions result in a huge number of possible test cases; and some tests require the use of limited test ranges or other resources. In such cases, different approaches to testing can be used, such as (highly) accelerated life testing, design of experiments, and simulations.
The desired level of statistical confidence also plays a role in reliability testing. Statistical confidence is increased by increasing either the test time or the number of items tested. Reliability test plans are designed to achieve the specified reliability at the specified confidence level with the minimum number of test units and test time. Different test plans result in different levels of risk to the producer and consumer. The desired reliability, statistical confidence, and risk levels for each side influence the ultimate test plan. The customer and developer should agree in advance on how reliability requirements will be tested.
A key aspect of reliability testing is to define "failure". Although this may seem obvious, there are many situations where it is not clear whether a failure is really the fault of the system. Variations in test conditions, operator differences, weather and unexpected situations create differences between the customer and the system developer. One strategy to address this issue is to use a scoring conference process. A scoring conference includes representatives from the customer, the developer, the test organization, the reliability organization, and sometimes independent observers. The scoring conference process is defined in the statement of work. Each test case is considered by the group and "scored" as a success or failure. This scoring is the official result used by the reliability engineer.
As part of the requirements phase, the reliability engineer develops a test strategy with the customer. The test strategy makes trade-offs between the needs of the reliability organization, which wants as much data as possible, and constraints such as cost, schedule and available resources. Test plans and procedures are developed for each reliability test, and results are documented.
Reliability testing is common in the Photonics industry. Examples of reliability tests of lasers are life test and burn-in. These tests consist of the highly accelerated aging, under controlled conditions, of a group of lasers. The data collected from these life tests are used to predict laser life expectancy under the intended operating characteristics.[14]
Reliability test requirements
There are many criteria to test depends on the product or process that are testing on, and mainly, there are five components that are most common:[15][16]
- Product life span
- Intended function
- Operating Condition
- Probability of Performance
- User exceptions[17]
The product life span can be split into four different for analysis. Useful life is the estimated economic life of the product, which is defined as the time can be used before the cost of repair do not justify the continue use to the product. Warranty life is the product should perform the function within the specified time period. Design life is where during the design of the product, designer take into consideration on the life time of competitive product and customer desire and ensure that the product do not result in customer dissatisfaction.[18][19]
Reliability test requirements can follow from any analysis for which the first estimate of failure probability, failure mode or effect needs to be justified. Evidence can be generated with some level of confidence by testing. With software-based systems, the probability is a mix of software and hardware-based failures. Testing reliability requirements is problematic for several reasons. A single test is in most cases insufficient to generate enough statistical data. Multiple tests or long-duration tests are usually very expensive. Some tests are simply impractical, and environmental conditions can be hard to predict over a systems life-cycle.
Reliability engineering is used to design a realistic and affordable test program that provides empirical evidence that the system meets its reliability requirements. Statistical confidence levels are used to address some of these concerns. A certain parameter is expressed along with a corresponding confidence level: for example, an MTBF of 1000 hours at 90% confidence level. From this specification, the reliability engineer can, for example, design a test with explicit criteria for the number of hours and number of failures until the requirement is met or failed. Different sorts of tests are possible.
The combination of required reliability level and required confidence level greatly affects the development cost and the risk to both the customer and producer. Care is needed to select the best combination of requirements—e.g. cost-effectiveness. Reliability testing may be performed at various levels, such as component, subsystem and system. Also, many factors must be addressed during testing and operation, such as extreme temperature and humidity, shock, vibration, or other environmental factors (like loss of signal, cooling or power; or other catastrophes such as fire, floods, excessive heat, physical or security violations or other myriad forms of damage or degradation). For systems that must last many years, accelerated life tests may be needed.
Testing method
A systematic approach to reliability testing is to, first, determine reliability goal, then do tests that are linked to performance and determine the reliability of the product.[20] A reliability verification test in modern industries should clearly determine how they relate to the product's overall reliability performance and how individual tests impact the warranty cost and customer satisfaction.[21]
Accelerated testing
The purpose of accelerated life testing (ALT test) is to induce field failure in the laboratory at a much faster rate by providing a harsher, but nonetheless representative, environment. In such a test, the product is expected to fail in the lab just as it would have failed in the field—but in much less time. The main objective of an accelerated test is either of the following:
- To discover failure modes
- To predict the normal field life from the high stress lab life
An accelerated testing program can be broken down into the following steps:
- Define objective and scope of the test
- Collect required information about the product
- Identify the stress(es)
- Determine level of stress(es)
- Conduct the accelerated test and analyze the collected data.
Common ways to determine a life stress relationship are:
- Arrhenius model
- Eyring model
- Inverse power law model
- Temperature–humidity model
- Temperature non-thermal model
الاختبار السريع
برمجيات الموثوقية
تقييم الموثوقية التشغيلية
المنظمات الموثوقية
شهادة
تعليم الموثوقية الهندسية
انظر أيضاً
- Brittle Systems
- Bayesian inference
- Burn-in
- Failing badly
- Failure rate
- Human reliability
- Integrated Logistics Support
- Highly accelerated stress test
- Highly Accelerated Life Test
- Logistic engineering
- Performance engineering
- Professional engineer
- Product qualification
- Quality engineering
- Reliability
- Reliable system design
- Reliability theory
- Reliability theory of aging and longevity
- Risk assessment
- Redundancy (total quality management)
- Security engineering
- Single point of failure (SPOF)
- Software engineering
- Systems engineering
- Safety engineering
- Statistics
- Temperature cycling
- Spurious trip level
- Safety integrity level
المصادر
- ^ وثوقية النظم, الموسوعة العربية
- ^ Barnard, R.W.A. (2008). "What is wrong with Reliability Engineering?" (PDF). Lambda Consulting. Retrieved 30 October 2014.
- ^ "Best Practices in Design for Reliability" (PDF). Archived from the original (PDF) on 2017-11-17.
- ^ Salvatore Distefano, Antonio Puliafito: Dependability Evaluation with Dynamic Reliability Block Diagrams and Dynamic Fault Trees. IEEE Trans. Dependable Sec. Comput. 6(1): 4–17 (2009)
- ^ The Seven Samurais of Systems Engineering, James Martin (2008) Archived 1 ديسمبر 2023 at the Wayback Machine
- ^ خطأ استشهاد: وسم
<ref>غير صحيح؛ لا نص تم توفيره للمراجع المسماةO'Connor_2002 - ^ خطأ استشهاد: وسم
<ref>غير صحيح؛ لا نص تم توفيره للمراجع المسماةsciencedirect.com - ^ Zhang, J.; Geiger, C.; Sun, F. (January 2016). "A system approach to reliability verification test design". 2016 Annual Reliability and Maintainability Symposium (RAMS). pp. 1–6. doi:10.1109/RAMS.2016.7448014. ISBN 978-1-5090-0249-8. S2CID 24770411.
- ^ Dai, Wei; Maropoulos, Paul G.; Zhao, Yu (2015-01-02). "Reliability modelling and verification of manufacturing processes based on process knowledge management". International Journal of Computer Integrated Manufacturing. 28 (1): 98–111. doi:10.1080/0951192X.2013.834462. ISSN 0951-192X. S2CID 32995968.
- ^ "Reliability Verification for AI and ML Processors - White Paper". www.allaboutcircuits.com (in الإنجليزية). Retrieved 2020-12-11.
- ^ Weber, Wolfgang; Tondok, Heidemarie; Bachmayer, Michael (2005-07-01). "Enhancing software safety by fault trees: experiences from an application to flight critical software". Reliability Engineering & System Safety. Safety, Reliability and Security of Industrial Computer Systems (in الإنجليزية). 89 (1): 57–70. doi:10.1016/j.ress.2004.08.007. ISSN 0951-8320.
- ^ Ren, Yuanqiang; Tao, Jingya; Xue, Zhaopeng (January 2020). "Design of a Large-Scale Piezoelectric Transducer Network Layer and Its Reliability Verification for Space Structures". Sensors (in الإنجليزية). 20 (15): 4344. Bibcode:2020Senso..20.4344R. doi:10.3390/s20154344. PMC 7435873. PMID 32759794.
- ^ Ben-Gal I., Herer Y. and Raz T. (2003). "Self-correcting inspection procedure under inspection errors" (PDF). IIE Transactions on Quality and Reliability, 34(6), pp. 529–540. Archived from the original (PDF) on 13 October 2013. Retrieved 10 January 2014.
{{cite journal}}: Cite journal requires|journal=(help) - ^ "Yelo Reliability Testing". Archived from the original on 5 March 2016. Retrieved 6 November 2014.
- ^ Matheson, Granville J. (2019-05-24). "We need to talk about reliability: making better use of test-retest studies for study design and interpretation". PeerJ. 7 e6918. doi:10.7717/peerj.6918. ISSN 2167-8359. PMC 6536112. PMID 31179173.
- ^ Pronskikh, Vitaly (2019-03-01). "Computer Modeling and Simulation: Increasing Reliability by Disentangling Verification and Validation". Minds and Machines (in الإنجليزية). 29 (1): 169–186. doi:10.1007/s11023-019-09494-7. ISSN 1572-8641. OSTI 1556973. S2CID 84187280.
- ^ Halamay, D. A.; Starrett, M.; Brekken, T. K. A. (2019). "Hardware Testing of Electric Hot Water Heaters Providing Energy Storage and Demand Response Through Model Predictive Control". IEEE Access. 7: 139047–139057. Bibcode:2019IEEEA...7m9047H. doi:10.1109/ACCESS.2019.2932978. ISSN 2169-3536.
- ^ Chen, Jing; Wang, Yinglong; Guo, Ying; Jiang, Mingyue (2019-02-19). "A metamorphic testing approach for event sequences". PLOS ONE (in الإنجليزية). 14 (2) e0212476. Bibcode:2019PLoSO..1412476C. doi:10.1371/journal.pone.0212476. ISSN 1932-6203. PMC 6380623. PMID 30779769.
- ^ Bieńkowska, Agnieszka; Tworek, Katarzyna; Zabłocka-Kluczka, Anna (January 2020). "Organizational Reliability Model Verification in the Crisis Escalation Phase Caused by the COVID-19 Pandemic". Sustainability (in الإنجليزية). 12 (10): 4318. Bibcode:2020Sust...12.4318B. doi:10.3390/su12104318.
- ^ Jenihhin, M.; Lai, X.; Ghasempouri, T.; Raik, J. (October 2018). "Towards Multidimensional Verification: Where Functional Meets Non-Functional". 2018 IEEE Nordic Circuits and Systems Conference (NORCAS): NORCHIP and International Symposium of System-on-Chip (SoC). pp. 1–7. arXiv:1908.00314. doi:10.1109/NORCHIP.2018.8573495. ISBN 978-1-5386-7656-1. S2CID 56170277.
- ^ Rackwitz, R. (2000-02-21). "Optimization — the basis of code-making and reliability verification". Structural Safety (in الإنجليزية). 22 (1): 27–60. doi:10.1016/S0167-4730(99)00037-5. ISSN 0167-4730.
هذه المقالة بحاجة لمصادر إضافية لتحسين وثوقيتها. (October 2008) |
قراءات أخرى
- Blanchard, Benjamin S. (1992), Logistics Engineering and Management (Fourth Ed.), Prentice-Hall, Inc., Englewood Cliffs, New Jersey.
- Breitler, Alan L. and Sloan, C. (2005), Proceedings of the American Institute of Aeronautics and Astronautics (AIAA) Air Force T&E Days Conference, Nashville, TN, December, 2005: System Reliability Prediction: towards a General Approach Using a Neural Network.
- Ebeling, Charles E., (1997), An Introduction to Reliability and Maintainability Engineering, McGraw-Hill Companies, Inc., Boston.
- Denney, Richard (2005) Succeeding with Use Cases: Working Smart to Deliver Quality. Addison-Wesley Professional Publishing. ISBN . Discusses the use of software reliability engineering in use case driven software development.
- Gano, Dean L. (2007), "Apollo Root Cause Analysis" (Third Edition), Apollonian Publications, LLC., Richland, Washington
- Holmes, Oliver Wendell, Sr. The Deacon's Masterpiece
- Kapur, K.C., and Lamberson, L.R., (1977), Reliability in Engineering Design, John Wiley & Sons, New York.
- Kececioglu, Dimitri, (1991) "Reliability Engineering Handbook", Prentice-Hall, Englewood Cliffs, New Jersey
- Trevor Kletz (1998) Process Plants: A Handbook for Inherently Safer Design CRC ISBN 1-56032-619-0
- Leemis, Lawrence, (1995) Reliability: Probabilistic Models and Statistical Methods, 1995, Prentice-Hall. ISBN 0-13-720517-1
- Frank Lees (2005). Loss Prevention in the Process Industries (3rdEdition ed.). Elsevier. ISBN 978-0-7506-7555-0.
- MacDiarmid, Preston; Morris, Seymour; et al., (1995), Reliability Toolkit: Commercial Practices Edition, Reliability Analysis Center and Rome Laboratory, Rome, New York.
- Modarres, Mohammad; Kaminskiy, Mark; Krivtsov, Vasiliy (1999), "Reliability Engineering and Risk Analysis: A Practical Guide, CRC Press, ISBN 0-8247-2000-8.
- Musa, John (2005) Software Reliability Engineering: More Reliable Software Faster and Cheaper, 2nd. Edition, AuthorHouse. ISBN
- Neubeck, Ken (2004) "Practical Reliability Analysis", Prentice Hall, New Jersey
- Neufelder, Ann Marie, (1993), Ensuring Software Reliability, Marcel Dekker, Inc., New York.
- O'Connor, Patrick D. T. (2002), Practical Reliability Engineering (Fourth Ed.), John Wiley & Sons, New York.
- Shooman, Martin, (1987), Software Engineering: Design, Reliability, and Management, McGraw-Hill, New York.
- Tobias, Trindade, (1995), Applied Reliability, Chapman & Hall/CRC, ISBN 0-442-00469-9
- Springer Series in Reliability Engineering
- Nelson, Wayne B., (2004), Accelerated Testing - Statistical Models, Test Plans, and Data Analysis, John Wiley & Sons, New York, ISBN 0-471-69736-2
- Bagdonavicius, V., Nikulin, M., (2002), "Accelerated Life Models. Modeling and Statistical analysis", CHAPMAN&HALL/CRC, Boca Raton, ISBN 1-58488-186-0
معايير الولايات المتحدة ، والمواصفات، والكتيبات
- Aerospace Report Number: TOR-2007(8583)-6889 Reliability Program Requirements for Space Systems, The Aerospace Corporation (10 Jul 2007)
- DoD 3235.1-H (3rd Ed) Test and Evaluation of System Reliability, Availability, and Maintainability (A Primer), U.S. Department of Defense (March 1982) .
- NASA GSFC 431-REF-000370 Flight Assurance Procedure: Performing a Failure Mode and Effects Analysis, National Aeronautics and Space Administration Goddard Space Flight Center (10 Aug 1996).
- IEEE 1332-1998 IEEE Standard Reliability Program for the Development and Production of Electronic Systems and Equipment, Institute of Electrical and Electronics Engineers (1998).
- JPL D-5703 Reliability Analysis Handbook, National Aeronautics and Space Administration Jet Propulsion Laboratory (July 1990).
- MIL-STD-785B Reliability Program for Systems and Equipment Development and Production, U.S. Department of Defense (15 Sep 1980). (*Obsolete, superseded by ANSI/GEIA-STD-0009-2008 titled Reliability Program Standard for Systems Design, Development, and Manufacturing, 13 Nov 2008)
- MIL-HDBK-217F Reliability Prediction of Electronic Equipment, U.S. Department of Defense (2 Dec 1991).
- MIL-HDBK-217F (Notice 1) Reliability Prediction of Electronic Equipment, U.S. Department of Defense (10 Jul 1992).
- MIL-HDBK-217F (Notice 2) Reliability Prediction of Electronic Equipment, U.S. Department of Defense (28 Feb 1995).
- MIL-STD-690D Failure Rate Sampling Plans and Procedures, U.S. Department of Defense (10 Jun 2005).
- MIL-HDBK-338B Electronic Reliability Design Handbook, U.S. Department of Defense (1 Oct 1998).
- MIL-HDBK-2173 Reliability-Centered Maintenance (RCM) Requirements for Naval Aircraft, Weapon Systems, and Support Equipment, U.S. Department of Defense (30 JAN 1998); (superseded by NAVAIR 00-25-403).
- MIL-STD-1543B Reliability Program Requirements for Space and Launch Vehicles, U.S. Department of Defense (25 Oct 1988).
- MIL-STD-1629A Procedures for Performing a Failure Mode Effects and Criticality Analysis, U.S. Department of Defense (24 Nov 1980).
- MIL-HDBK-781A Reliability Test Methods, Plans, and Environments for Engineering Development, Qualification, and Production, U.S. Department of Defense (1 Apr 1996).
- NSWC-06 (Part A) Handbook of Reliability Prediction Procedures for Mechanical Equipment, Naval Surface Warfare Center (10 Jan 2006).
- NSWC-06 (Part B) Handbook of Reliability Prediction Procedures for Mechanical Equipment, Naval Surface Warfare Center (10 Jan 2006).
معايير المملكة المتحدة
In the UK, there are more up to date standards maintained under the sponsorship of UK MOD as Defence Standards. The relevant Standards include:
DEF STAN 00-40 Reliability and Maintainability (R&M)
- PART 1: Issue 5: Management Responsibilities and Requirements for Programmes and Plans
- PART 4: (ARMP-4)Issue 2: Guidance for Writing NATO R&M Requirements Documents
- PART 6: Issue 1: IN-SERVICE R & M
- PART 7 (ARMP-7) Issue 1: NATO R&M Terminology Applicable to ARMP’s
DEF STAN 00-42 RELIABILITY AND MAINTAINABILITY ASSURANCE GUIDES
- PART 1: Issue 1: ONE-SHOT DEVICES/SYSTEMS
- PART 2: Issue 1: SOFTWARE
- PART 3: Issue 2: R&M CASE
- PART 4: Issue 1: Testability
- PART 5: Issue 1: IN-SERVICE RELIABILITY DEMONSTRATIONS
DEF STAN 00-43 RELIABILITY AND MAINTAINABILITY ASSURANCE ACTIVITY
- PART 2: Issue 1: IN-SERVICE MAINTAINABILITY DEMONSTRATIONS
DEF STAN 00-44 RELIABILITY AND MAINTAINABILITY DATA COLLECTION AND CLASSIFICATION
- PART 1: Issue 2: MAINTENANCE DATA & DEFECT REPORTING IN THE ROYAL NAVY, THE ARMY AND THE ROYAL AIR FORCE
- PART 2: Issue 1: DATA CLASSIFICATION AND INCIDENT SENTENCING - GENERAL
- PART 3: Issue 1: INCIDENT SENTENCING - SEA
- PART 4: Issue 1: INCIDENT SENTENCING - LAND
DEF STAN 00-45 Issue 1: RELIABILITY CENTERED MAINTENANCE
DEF STAN 00-49 Issue 1: RELIABILITY AND MAINTAINABILITY MOD GUIDE TO TERMINOLOGY DEFINITIONS
These can be obtained from DSTAN. There are also many commercial standards, produced by many organisations including the SAE, MSG, ARP, and IEE.
المعايير الفرنسية
- FIDES [2]. The FIDES methodology (UTE-C 80-811) is based on the physics of failures and supported by the analysis of test data, field returns and existing modelling.
- UTE-C 80-810 or RDF2000 [3]. The RDF2000 methodology is based on the French telecom experience.
المعايير الدولية
وصلات خارجية
الوصلات الخارجية في هذه المقالة قد لا تتبع سياسات المحتوى أو الإرشادات. من فضلك حسـِّن هذه المقالة بإزالة الوصلات الخارجية الزائدة أو غير المناسبة. |
- American Society for Quality
- Carnegie Mellon Software Engineering Institute
- IEEE Reliability Society
- NASA Hardware and Software Reliability report
- NIST/SEMATECH, "Engineering Statistics Handbook", [4]
- Society of Reliability Engineers
- University of Maryland Reliability Engineering Program
- Reliability Information Analysis Center
- Models and methods regarding reliability analysis
- UK Defence Standardization Organisation's Home on the Web
- Center for Risk and Reliability at University of Maryland, College Park
- Reliability Engineering services and software
- On-line Reliability Engineering Resources for the Reliability Professional
- EURELNET European Reliability Network - Failure Mechanisms and Materials Database
- Articles with hatnote templates targeting a nonexistent page
- جميع المقالات الحاوية على عبارات مبهمة
- جميع المقالات الحاوية على عبارات مبهمة from December 2023
- Pages with empty portal template
- Articles needing additional references from October 2008
- All articles needing additional references
- تنظيف الوصلات الخارجية بالمعرفة
- Design for X
- أعطال
- هندسة الوثوقية
- Sهندسة النظم
- جودة البرمجيات
- إحصاءات هندسية
- Survival analysis
- علوم المواد