Get in Touch

Course Outline

EXO Infrastructure as Code Automation

  • Overview of EXO deployment architectures, including single-node, multi-node, and Remote Direct Memory Access (RDMA) cluster configurations.
  • Implementation of configuration management tools to automate the installation of required dependencies, such as Xcode, uv, Node.js, and Rust.
  • Leveraging Nix flakes to ensure reproducible EXO builds and consistent developer environments for government IT staff.
  • Development of Ansible playbooks or shell scripts to facilitate unattended cluster provisioning and lifecycle management.

Reproducible Builds and Continuous Integration Integration

  • Pinning dependencies and executing dashboard builds within continuous integration (CI) pipelines to ensure consistency for government systems.
  • Execution of EXO smoke tests within GitHub Actions or GitLab CI runner environments.
  • Establishment of golden images and snapshot-based rollback workflows for macOS and Linux virtual machines to support rapid recovery.
  • Management of custom model card versioning in parallel with application code repositories.

Cluster Discovery and Network Automation

  • Configuration of multicast DNS (mDNS) and static DNS records to ensure reliable libp2p node discovery within federal networks.
  • Automation of network profile creation and Thunderbolt bridge management on macOS endpoints for government users.
  • Utilization of custom namespaces, specifically EXO_LIBP2P_NAMESPACE, to isolate development, staging, and production clusters.
  • Implementation of firewall rules and network segmentation strategies to support multi-tenant environments and secure data separation.

Storage Management and Model Lifecycle Control

  • Design and implementation of strategies for EXO_MODELS_DIRS and EXO_MODELS_READ_ONLY_DIRS to manage model storage effectively.
  • Mouting of Network File System (NFS) or Storage Area Network (SAN) shares as read-only repositories to enable rapid model provisioning.
  • Execution of garbage collection protocols for stale caches and enforcement of retention policies for versioned model weights.
  • Automation of pre-download processes and health checks prior to rolling updates to maintain service continuity.

Monitoring, Logging, and Alerting

  • Ingestion of EXO logs into centralized logging platforms, such as Elasticsearch, Logstash, Kibana (ELK), Loki, or Splunk.
  • Construction of Grafana dashboards utilizing output from EXO_TRACING_ENABLED for real-time operational visibility.
  • Configuration of alerts for critical events, including cluster membership changes, out-of-memory (OOM) conditions, and inference latency spikes.
  • Correlation of hardware telemetry data from macmon with model performance metrics to identify regressions.

System Updates, Rollback, and Disaster Recovery

  • Execution of canary deployments for EXO binary updates on individual nodes prior to fleet-wide rollout to government systems.
  • Implementation of model-level rollback capabilities, allowing switching between quantized versions without requiring full re-downloads.
  • Procedures for backing up and restoring cluster state, custom namespaces, and cached weights to ensure data integrity.
  • Development of documented recovery runbooks for scenarios involving total cluster rebuilds.

Security Hardening and Regulatory Compliance

  • Application of Transport Layer Security (TLS) at the reverse proxy layer, utilizing tools such as nginx or Traefik, for the dashboard and API interfaces.
  • Implementation of API rate limiting and IP whitelisting controls to secure EXO endpoints against unauthorized access.
  • Isolation of clusters through Virtual Local Area Networks (VLANs) and enforcement of zero-trust network policies.
  • Maintenance of audit trails for user access and an inventory of deployed models and their respective versions for government accountability.

Requirements

**Technical Competencies** * Demonstrated proficiency in DevOps methodologies, including continuous integration and deployment (CI/CD), infrastructure as code (IaC), and container orchestration platforms. * Working knowledge of system administration and package management within macOS or Linux environments. * Fundamental understanding of networking protocols, Domain Name System (DNS) resolution, and storage architecture principles. **Intended Audience** This content is designed for government agencies and entities requiring specialized training for: * DevOps engineers * Infrastructure architects * Site Reliability Engineers (SREs) managing on-premise artificial intelligence workloads
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories