Course Outline
Module 1: Microservices Architecture Design
• Establishing effective microservice boundaries
• Application of Domain-Driven Design (DDD) principles
• Alternative approaches to defining business domain boundaries, including considerations for volatility, data integrity, technology stacks, and organizational structure
• Strategies for decomposing monolithic systems
• Mitigating risks associated with premature decomposition
• Decomposition by architectural layer
• Utilization of established decomposition patterns, such as the Strangler, Parallel Run, and Feature Toggle models
• Addressing data decomposition challenges related to performance, integrity, and transactional consistency
Module 2: Docker and Runtime Optimization
• Selection of appropriate base images for government workloads
• Strategies to minimize the number of image layers
• Implementation of multi-stage builds to enhance efficiency
• Image optimization techniques, including the sorting of multi-line arguments
• Leveraging build caches to reduce deployment times
• Version pinning for image stability and reproducibility
• Precise tuning of resource allocation
• Adherence to secure container practices
• Runtime configuration adjustments to maximize performance
Module 3: Kubernetes and Release Strategies
Overview of Kubernetes Deployment Mechanisms
• Creation and execution of initial deployments
• Evaluation of various Kubernetes deployment options
Execution of Rolling Update Deployments
• Mechanisms and principles behind rolling updates
• Step-by-step creation and execution of rolling updates
• Procedures for rolling back deployments to maintain service stability
Execution of Canary Deployments
• Conceptual understanding of canary deployment strategies
• Implementation and execution of canary deployments for controlled release
Execution of Blue-Green Deployments
• Principles of blue-green deployment for zero-downtime releases
• Creation and execution of blue-green deployments
Management of Jobs and CronJobs
• Configuration and creation of batch jobs and scheduled cron jobs
Monitoring and Troubleshooting Protocols
• Application of troubleshooting techniques using kubectl for system diagnostics
Module 4: Automation and Operational Efficiency
Utilizing Python for Kubernetes Automation
• Execution of administrative operations in Kubernetes using Python
• Definition of configuration objects via Python scripting
• Creation of deployment objects using Python
• Monitoring of Kubernetes events through Python-based scripts
• Scaling of deployments using automated Python processes
Challenges in Deployment Automation
• Implementation of declarative configuration standards in Kubernetes
• Strategies for maintaining configuration integrity and consistency
Application of GitOps Methodologies
• Core principles of the GitOps framework
• Introduction to Flux as a GitOps engine
• Installation and integration of Flux into a Kubernetes cluster
Configuration of Flux for Automated Releases
• Integration of notification systems for operational awareness
• Structure and management of source repositories for configuration data
Managing Application Updates via Image Automation
• Updating application deployments utilizing Flux
• Scanning container image repositories for version tags
• Definition of policies for selecting the latest available images
• Configuration of Flux to facilitate automatic image updates
Module 5: Observability and Root Cause Analysis
Kubernetes Logging and Tracing Capabilities
• Importance of comprehensive logging and tracing in cloud-native environments
• Methods for accessing Kubernetes log data
• Analysis of Pod and Container logs
• Review of Control Plane logs
• Monitoring resource usage across Nodes and Pods
Log Collection and Analytical Methods
• Implementation of log aggregation strategies
• Visualization of log data for enhanced insight
Distributed Tracing Implementation
• Definition and utility of distributed tracing
• Application of OpenTelemetry standards
• Utilization of distributed tracing tools for service observability
• Instrumentation of applications for detailed tracing
• Identification of performance issues through tracing analysis
Monitoring with Prometheus and Grafana
• Core concepts of system observability
• Overview of standard monitoring tools
• Implementation of Prometheus instrumentation for metrics collection
Advanced Logging Use Cases
• Processing of log data for analytical purposes
• Filtering and enrichment of log entries to improve signal-to-noise ratio
• Implementation of event sourcing patterns
Module 6: Cluster Crisis Simulation and Incident Response
• Identification of various failure modes in cluster environments
• Simulation of Node failures to test resilience
• Analysis of Pod Eviction and Resource Exhaustion scenarios
• Diagnosis of network-related issues
• Handling of DNS failures and application timeout management
• Simulation of API Server outages to assess impact
• Stress testing for system stability under high traffic conditions
• Evaluation of Storage Failure impacts
• Identification and correction of Configuration Errors
• Understanding of standardized incident reporting and response procedures
Module 7: AI-Assisted Troubleshooting
• Benefits of Generative AI in Kubernetes operations
• Architectural overview of the K8sGPT CLI
• Installation procedures for the K8sGPT CLI
• Comprehensive usage of K8sGPT commands
• Application of K8sGPT Analyzers, including podAnalyzer, pvcAnalyzer, and rsAnalyzer
• Cluster-wide analysis utilizing K8sGPT
• Analysis of real-time operational issues using K8sGPT
• Deployment of the In-Cluster Operator for K8sGPT
Requirements
- Fundamental proficiency with the Linux command line interface
- Practical experience in application development or system administration
- Familiarity with containerization concepts, specifically Docker
- Basic understanding of core Kubernetes concepts, including pods, deployments, and services
- General understanding of software architecture principles, such as APIs and microservices
Target Audience:
- DevOps Engineers
- Site Reliability Engineers (SREs)
- Backend and Software Developers specializing in microservices
- Cloud Engineers and Platform Engineers
-
System Administrators transitioning to Kubernetes-based environments
Testimonials (2)
Craig was extremely involved in the training, always making sure we are paying attention, adapted the examples to our day-to-day activities and always provided an answer when asked, even if the information was not added in the presentation.
Ecaterina Ioana Nicoale - BOOKING HOLDINGS ROMANIA SRL
Course - DevOps Foundation®
High level of commitment and knowledge of the trainer