Skip to content
Reliability Engineering Training – Classroom & Virtual cover image

Reliability Engineering Training – Classroom & Virtual
Web Infra Academy

High Availability, Resilience, SRE & Observability for solution architects, software architects and senior engineers

Summary

Price
£1,795 inc VAT
Study method
Classroom
Duration
2 days · Full-time
Qualification
No formal qualification

Add to basket or enquire

Location & dates

Location
Address
44 Hallam St
Hallam Street
West London
London
W1W6JJ
United Kingdom

Overview

Learn how to design, operate and improve reliable modern IT systems.

This two-day Reliability Engineering training combines high availability, software resilience, SRE, observability, data reliability and safe delivery. You will learn how systems behave under real-world failures and how architecture, software, infrastructure and operations work together to create resilient services.

Through explanation, practical examples and discussion, you will learn to evaluate architecture decisions, identify reliability risks and apply reliability principles to real-world systems.

Course media

Description

Modern IT systems are distributed, interconnected and constantly changing. Reliability therefore cannot be added afterwards: it must be designed into applications, architectures, delivery processes and operations.

This two-day course provides a practical, architecture-focused approach to Reliability Engineering. You will learn how to build systems that remain available, recover predictably and continue to provide useful service when components, dependencies or infrastructure fail.

The course covers the complete reliability lifecycle.

You will explore:

• Reliability Engineering, SRE principles, SLIs, SLOs and error budgets
• Designing software and microservices for failure
• Resilience patterns including retries, backoff, circuit breakers, bulkheads and backpressure
• High availability architectures, redundancy, failover and failure isolation
• Multi-zone and multi-region architectures and controlling blast radius
• Safe change using CI/CD, Infrastructure as Code, GitOps and progressive delivery
• Canary and blue/green deployments, feature flags and automated guardrails
• Data reliability, replication, recovery and state management
• CAP, PACELC, consistency and availability trade-offs in distributed systems
• Resilience testing and chaos engineering
• Observability using metrics, logs, traces and the Four Golden Signals
• Moving from monitoring towards system intelligence
• Platform engineering, policy-as-code and reliability governance
• Continuous improvement and reliability maturity

Rather than treating high availability, resilience and observability as separate disciplines, the course shows how they interact across the complete system lifecycle.

You will learn to recognise architectural weaknesses, understand how failures propagate through distributed systems and make informed trade-offs between availability, consistency, performance, complexity and cost.

During the training, concepts are connected to practical examples and real-world architecture decisions. There is room for discussion and for relating reliability principles to the systems and challenges participants encounter in practice.

After completing the course, you will be better equipped to design and evaluate reliable architectures, make informed resilience decisions, detect and understand production problems faster and systematically improve reliability over time.

Who is this course for?

This course is designed for experienced IT professionals involved in designing, building or improving modern production systems.

It is particularly suitable for:

• Solution architects
• Software and application architects
• Senior software engineers
• DevOps engineers
• Platform engineers
• Site Reliability Engineers (SREs)
• Cloud architects and engineers
• Technical leads and engineering leads

The course is especially relevant for professionals who make architectural and engineering decisions about availability, resilience, recoverability and operational behaviour.

Requirements

No specific certification or programming experience is required.

Participants should have a general understanding of modern IT systems, applications, infrastructure or cloud environments. Experience in software engineering, architecture, DevOps, operations or platform engineering will help you get the most from the course.

The course is aimed at experienced IT professionals rather than complete IT beginners.

Career path

This course supports progression towards senior technical roles such as Solution Architect, Software or Application Architect, Reliability Architect, Site Reliability Engineer (SRE), Senior Software Engineer, Platform Engineer, Cloud Architect and Technical or Engineering Lead.

Questions and answers

Reviews

Currently there are no reviews for this course. Be the first to leave a review.

FAQs