Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
System Monitoring Capabilities for Government Entities
- Rationale for evaluating cloud-based monitoring solutions regarding infrastructure topology visibility and performance metrics.
- Technical architecture of Uptime Kuma, utilizing Node.js, SQLite, and a Vue.js interface for operational clarity.
- Evaluation against established monitoring platforms such as Nagios, Zabbix, and Grafana OnCall.
Rapid Deployment Procedures
- Implementation via Docker containerization with persistent volume storage.
- Configuration of reverse proxy services and Transport Layer Security (TLS) encryption.
- Initialization of system parameters and administrative account provisioning.
- Utilization of environment variables to secure authentication credentials and define base URLs.
Monitoring Modalities for Government Operations
- HTTP/HTTPS service verification, including keyword validation and status code assessment.
- Transmission Control Protocol (TCP) port accessibility and Internet Control Message Protocol (ICMP) ping tests.
- DNS resolution accuracy and query-type verification.
- Push-based monitoring for cron job execution and backup system heartbeats.
- Support for MQTT, gRPC protocols, and dedicated game server health checks.
Alerting and Notification Channels
- Delivery via SMTP email and Microsoft Teams webhooks.
- Integration with Slack, Discord, Telegram, and Signal messaging platforms.
- Compatibility with PagerDuty, Opsgenie, and custom webhook payloads.
- Implementation of notification throttling mechanisms and escalation policies for critical incidents.
Public Status Page Management
- Creation of branded public-facing status dashboards.
- Documentation of incident timelines and maintenance schedules.
- Customization through Cascading Style Sheets (CSS) and domain mapping.
- Provision of RSS and JSON feeds to facilitate automated status aggregation.
System Integration and Maintenance Protocols
- Exposure of Prometheus metrics endpoints for external data scraping.
- Application Programming Interface (API) support for bulk monitor creation and lifecycle management.
- Database backup strategies and migration procedures.
- Procedures for system updates and security hardening to ensure compliance for government use cases.
Requirements
- Demonstrated proficiency in Linux operating system management and Docker container orchestration.
- Comprehensive knowledge of network protocols, including HTTP and TCP, alongside established system monitoring principles.
- Operational experience with diverse alerting mechanisms, such as email distribution lists, Slack integrations, and Discord webhooks.
Target Audience
- Site Reliability Engineering (SRE) and DevOps personnel seeking to transition away from commercial cloud-based monitoring dashboards toward self-hosted solutions.
- Agile operational units requiring lightweight, sovereign-compliant uptime verification tools.
- Institutions adhering to strict regulatory frameworks that prohibit the use of Software-as-a-Service (SaaS) monitoring platforms in favor of on-premises or government-approved infrastructure for government applications.
7 Hours