Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
EXO Infrastructure as Code Automation
- Overview of EXO deployment architectures, including single-node, multi-node, and Remote Direct Memory Access (RDMA) cluster configurations.
- Implementation of configuration management tools to automate the installation of required dependencies, such as Xcode, uv, Node.js, and Rust.
- Leveraging Nix flakes to ensure reproducible EXO builds and consistent developer environments for government IT staff.
- Development of Ansible playbooks or shell scripts to facilitate unattended cluster provisioning and lifecycle management.
Reproducible Builds and Continuous Integration Integration
- Pinning dependencies and executing dashboard builds within continuous integration (CI) pipelines to ensure consistency for government systems.
- Execution of EXO smoke tests within GitHub Actions or GitLab CI runner environments.
- Establishment of golden images and snapshot-based rollback workflows for macOS and Linux virtual machines to support rapid recovery.
- Management of custom model card versioning in parallel with application code repositories.
Cluster Discovery and Network Automation
- Configuration of multicast DNS (mDNS) and static DNS records to ensure reliable libp2p node discovery within federal networks.
- Automation of network profile creation and Thunderbolt bridge management on macOS endpoints for government users.
- Utilization of custom namespaces, specifically EXO_LIBP2P_NAMESPACE, to isolate development, staging, and production clusters.
- Implementation of firewall rules and network segmentation strategies to support multi-tenant environments and secure data separation.
Storage Management and Model Lifecycle Control
- Design and implementation of strategies for EXO_MODELS_DIRS and EXO_MODELS_READ_ONLY_DIRS to manage model storage effectively.
- Mouting of Network File System (NFS) or Storage Area Network (SAN) shares as read-only repositories to enable rapid model provisioning.
- Execution of garbage collection protocols for stale caches and enforcement of retention policies for versioned model weights.
- Automation of pre-download processes and health checks prior to rolling updates to maintain service continuity.
Monitoring, Logging, and Alerting
- Ingestion of EXO logs into centralized logging platforms, such as Elasticsearch, Logstash, Kibana (ELK), Loki, or Splunk.
- Construction of Grafana dashboards utilizing output from EXO_TRACING_ENABLED for real-time operational visibility.
- Configuration of alerts for critical events, including cluster membership changes, out-of-memory (OOM) conditions, and inference latency spikes.
- Correlation of hardware telemetry data from macmon with model performance metrics to identify regressions.
System Updates, Rollback, and Disaster Recovery
- Execution of canary deployments for EXO binary updates on individual nodes prior to fleet-wide rollout to government systems.
- Implementation of model-level rollback capabilities, allowing switching between quantized versions without requiring full re-downloads.
- Procedures for backing up and restoring cluster state, custom namespaces, and cached weights to ensure data integrity.
- Development of documented recovery runbooks for scenarios involving total cluster rebuilds.
Security Hardening and Regulatory Compliance
- Application of Transport Layer Security (TLS) at the reverse proxy layer, utilizing tools such as nginx or Traefik, for the dashboard and API interfaces.
- Implementation of API rate limiting and IP whitelisting controls to secure EXO endpoints against unauthorized access.
- Isolation of clusters through Virtual Local Area Networks (VLANs) and enforcement of zero-trust network policies.
- Maintenance of audit trails for user access and an inventory of deployed models and their respective versions for government accountability.
Requirements
**Technical Competencies**
* Demonstrated proficiency in DevOps methodologies, including continuous integration and deployment (CI/CD), infrastructure as code (IaC), and container orchestration platforms.
* Working knowledge of system administration and package management within macOS or Linux environments.
* Fundamental understanding of networking protocols, Domain Name System (DNS) resolution, and storage architecture principles.
**Intended Audience**
This content is designed for government agencies and entities requiring specialized training for:
* DevOps engineers
* Infrastructure architects
* Site Reliability Engineers (SREs) managing on-premise artificial intelligence workloads
21 Hours
Testimonials (2)
Craig was extremely involved in the training, always making sure we are paying attention, adapted the examples to our day-to-day activities and always provided an answer when asked, even if the information was not added in the presentation.
Ecaterina Ioana Nicoale - BOOKING HOLDINGS ROMANIA SRL
Course - DevOps Foundation®
High level of commitment and knowledge of the trainer