Course Outline
Introduction
- Integrating traditional IT infrastructure with software development methodologies.
- Defining the operational distinctives between software engineers and system administrators.
- Differentiating Site Reliability Engineers from DevOps engineers in a government context.
Overview of an IT System
- System architecture components across on-premise and cloud environments.
Overview of SRE Principles and Practices
- Implementation of Infrastructure as Code.
- Application of containerization and orchestration tools (e.g., Docker, Kubernetes).
- Adoption of Continuous Integration, Continuous Deployment, and Continuous Delivery pipelines.
- Establishing comprehensive observability frameworks.
Evaluating an IT System
- Assessing team composition and available organizational resources.
- Documenting existing systems and operational processes.
- Analyzing the projected impact of SRE adoption on public sector operations.
- Defining the responsibilities of the software engineering team.
- Defining the responsibilities of the operational team.
- Determining the strategic role of management oversight.
Maintaining the Reliability of a System
- Defining and quantifying desired service reliability metrics.
- Analyzing Service Level Objectives (SLOs) for government compliance.
- Interpreting Service Level Indicators (SLIs) and Service Level Agreements (SLAs).
- Managing error budgets to balance innovation with stability.
- Formulating robust SLOs for critical public services.
Optimizing System Administration
- Configuring standardized development environments.
- Selecting and evaluating SRE tooling for government use.
- Prioritizing administrative tasks for automated execution.
- Developing software solutions for operational efficiency.
Deploying "Infrastructure as Code"
- Executing rigorous code testing and iterative refinement.
- Engineering systems for anti-fragility and resilience.
- Conducting post-incident reviews to derive lessons from failure.
Monitoring a System
- Tracking system performance against established baselines.
- Utilizing SRE-specific tools and analytical techniques.
The Future of SRE
Summary and Conclusion
Requirements
- Foundational understanding of IT infrastructure components.
- Basic knowledge of the software development lifecycle.
- Experience with programming or scripting in any language.
Audience
- Software Developers
- System Administrators
- Software Architects
- DevOps Engineers
- IT Managers
Testimonials (7)
How detailed subjects are explained with real world examples
Brian Hlabane - African Bank
Course - Site Reliability Engineering (SRE) Fundamentals
She is expert in area and provide really nice training. Material, training was really mix of examples , discussion and
Peter Tutka - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
View on the SRE/ DevOps from more business/ theoretical point of view. Most helpful for people who already have the practical view.
Michael Varhol - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
Approach of the training to send questionnaire before the training, so the training was planned accordingly to expectations. Brings the participants more active.
Stefan Girman - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
Sticking to the initial survey from attendees about what should be the focus of training.
Denis Majorsky - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
discussions , SRE definition
Daniel Horvath - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
Concept of the training, keeping the people focused by asking them a questions and triggering discussions. Also group breakout sessions were great to think about things in groups and see different outcomes from other group.