Get in Touch

Course Outline

Introduction to the EXO Framework and Local AI Clustering

  • Overview of the EXO architecture and the exo-explore ecosystem
  • Comparison of centralized cloud-based inference versus distributed local processing
  • Technical stack: libp2p device discovery, MLX backend, dashboard, and API interfaces
  • Hardware prerequisites: Apple Silicon processors (M3 Ultra, M4 Pro/Max), Thunderbolt 5 connectivity, and shared storage

Installation of EXO on macOS

  • Configuration of Xcode, Metal ToolChain, and macOS system requirements
  • Installation of uv, Node.js, and the Rust nightly toolchain
  • Deployment of the designated macmon fork for Apple Silicon performance monitoring
  • Repository cloning and dashboard compilation via npm
  • Execution of EXO from source and validation of the localhost:52415 dashboard

Installation of EXO on Linux

  • Dependency management using apt or Homebrew on Linux distributions
  • Setup of uv, Node.js version 18 or higher, and the Rust nightly toolchain
  • Dashboard compilation and EXO execution in CPU-only operational mode
  • Directory structure: XDG Base Directory specifications for configuration, data, cache, and logs

Automated Device Discovery and Cluster Formation

  • Mechanisms of libp2p-based automatic discovery within local network segments
  • Configuration of isolated clusters using custom namespaces via EXO_LIBP2P_NAMESPACE
  • Verification of node participation in the dashboard cluster view
  • Management of discovery failures and network segmentation challenges

Activation of RDMA over Thunderbolt 5

  • RDMA architectural design and the reported 99 percent reduction in latency
  • Enabling RDMA functionality via macOS Recovery mode using rdma_ctl
  • Physical cable requirements and port topology limitations on Mac Studio units
  • Ensuring macOS version consistency across all cluster nodes
  • Diagnostics for RDMA discovery issues and DHCP configuration problems

Deployment of Frontier Models

  • Utilization of the dashboard to load and shard DeepSeek v3.1, Qwen3-235B, and Llama family models
  • Previewing instance distribution via the /instance/previews API endpoint
  • Establishment of model instances using pipeline or tensor-parallel sharding strategies
  • Configuration of custom model cards sourced from the HuggingFace hub

Monitoring and Troubleshooting Procedures

  • Interpretation of EXO logs and analysis of distributed tracing data
  • Assessment of cluster health status within the dashboard cluster view
  • Diagnosis of worker node failures and evaluation of reconnection behaviors
  • Application of EXO_TRACING_ENABLED for performance bottleneck identification

Cluster Maintenance and Update Protocols

  • Procedures for updating EXO binaries and rebuilding the dashboard
  • Migration of model caches and management of pre-downloaded assets over NFS
  • Graceful removal of nodes and reallocation of computational workloads

Requirements

  • A solid understanding of networking fundamentals, including IP addressing, subnetting, and firewall configurations
  • Proficiency in macOS or Linux command-line administration
  • Familiarity with Python package management (pip/uv) and Node.js tooling

Target Audience

  • System administrators
  • DevOps engineers
  • AI infrastructure architects responsible for on-premise LLM deployment
 21 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories