Get in Touch

Course Outline

EXO Infrastructure as Code

  • Review of EXO deployment architectures: single-node, multi-node, and RDMA-based clusters
  • Automation of dependency deployment (Xcode, uv, Node.js, Rust) using configuration management tools
  • Leveraging Nix flakes to ensure consistent EXO builds and standardized developer environments for government operations
  • Developing Ansible playbooks or shell scripts to execute unattended cluster provisioning tasks

Reproducible Builds and CI Integration

  • Pin dependencies and compile dashboard components within continuous integration pipelines
  • Execute EXO smoke tests utilizing GitHub Actions or GitLab CI runners
  • Generate golden images and implement snapshot-based rollback procedures for macOS and Linux virtual machines
  • Version control custom model cards in parallel with application code

Cluster Discovery and Networking Automation

  • Configure mDNS and static DNS entries to ensure reliable libp2p node discovery
  • Automate the creation of network profiles and management of Thunderbolt bridges on macOS systems
  • Utilize custom namespaces (EXO_LIBP2P_NAMESPACE) to segregate development, staging, and production clusters
  • Establish firewall rules and network segmentation protocols for multi-tenant environments

Storage and Model Lifecycle Management

  • Define strategies for EXO_MODELS_DIRS and EXO_MODELS_READ_ONLY_DIRS configurations
  • Mount NFS or SAN shares as read-only model repositories to accelerate provisioning
  • Implement garbage collection for obsolete caches and establish retention policies for versioned weights
  • Automate pre-downloading of models and execute health checks prior to rolling updates

Monitoring and Alerting

  • Transmit EXO logs to centralized logging systems (ELK, Loki, or Splunk)
  • Construct Grafana dashboards based on EXO_TRACING_ENABLED outputs
  • Trigger alerts for cluster membership changes, out-of-memory events, and inference latency anomalies
  • Correlate macmon hardware telemetry data with model performance regressions

Update, Rollback, and Disaster Recovery

  • Stage EXO binary updates on canary nodes prior to full fleet deployment
  • Execute model-level rollbacks by switching between quantized versions without re-downloading assets
  • Perform backups and restores of cluster state, custom namespaces, and cached weights
  • Document recovery runbooks for total cluster reconstruction scenarios

Security Hardening and Compliance

  • Apply TLS encryption at the reverse proxy layer (nginx, traefik) for dashboard and API access
  • Enforce API rate limiting and IP whitelisting for EXO endpoints
  • Isolate clusters using VLANs and zero-trust network policies
  • Conduct access audits and maintain a detailed inventory of deployed models and their versions

Requirements

  • Proficiency in DevOps practices (CI/CD, IaC, container orchestration)
  • Familiarity with macOS or Linux system administration and package management
  • Understanding of networking, DNS, and storage concepts

Target Audience

  • DevOps engineers
  • Infrastructure architects
  • SREs responsible for on-premise AI workloads
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories