EXO: End-to-End Local AI Cluster Deployment Training Course
EXO serves as an open-source framework that integrates Apple Silicon devices into a distributed AI cluster, facilitating the local inference of frontier models that exceed the memory capacity of any single device.
This instructor-led, live training (available online or on-site) is designed for system administrators and DevOps engineers tasked with deploying, configuring, and managing EXO clusters for private LLM inference across multiple Apple Silicon or Linux nodes for government applications.
Upon completion of this training, participants will be capable of:
- Installing and configuring EXO on macOS and Linux nodes.
- Enabling automatic device discovery and establishing multi-node clusters.
- Activating and verifying RDMA over Thunderbolt 5 for ultra-low-latency inter-device communication.
- Deploying frontier models (DeepSeek, Qwen, Llama) across clustered devices.
- Monitoring cluster health and resolving common deployment challenges.
Course Format
- Interactive lectures and facilitated discussions.
- Extensive exercises and practical application tasks.
- Hands-on implementation within a live-lab environment.
Course Customization Options
- To request customized training for government, please contact the provider to arrange specific requirements.
Course Outline
Introduction to the EXO Framework and Local AI Clustering
- Overview of the EXO architecture and the exo-explore ecosystem
- Comparison of centralized cloud-based inference versus distributed local processing
- Technical stack: libp2p device discovery, MLX backend, dashboard, and API interfaces
- Hardware prerequisites: Apple Silicon processors (M3 Ultra, M4 Pro/Max), Thunderbolt 5 connectivity, and shared storage
Installation of EXO on macOS
- Configuration of Xcode, Metal ToolChain, and macOS system requirements
- Installation of uv, Node.js, and the Rust nightly toolchain
- Deployment of the designated macmon fork for Apple Silicon performance monitoring
- Repository cloning and dashboard compilation via npm
- Execution of EXO from source and validation of the localhost:52415 dashboard
Installation of EXO on Linux
- Dependency management using apt or Homebrew on Linux distributions
- Setup of uv, Node.js version 18 or higher, and the Rust nightly toolchain
- Dashboard compilation and EXO execution in CPU-only operational mode
- Directory structure: XDG Base Directory specifications for configuration, data, cache, and logs
Automated Device Discovery and Cluster Formation
- Mechanisms of libp2p-based automatic discovery within local network segments
- Configuration of isolated clusters using custom namespaces via EXO_LIBP2P_NAMESPACE
- Verification of node participation in the dashboard cluster view
- Management of discovery failures and network segmentation challenges
Activation of RDMA over Thunderbolt 5
- RDMA architectural design and the reported 99 percent reduction in latency
- Enabling RDMA functionality via macOS Recovery mode using rdma_ctl
- Physical cable requirements and port topology limitations on Mac Studio units
- Ensuring macOS version consistency across all cluster nodes
- Diagnostics for RDMA discovery issues and DHCP configuration problems
Deployment of Frontier Models
- Utilization of the dashboard to load and shard DeepSeek v3.1, Qwen3-235B, and Llama family models
- Previewing instance distribution via the /instance/previews API endpoint
- Establishment of model instances using pipeline or tensor-parallel sharding strategies
- Configuration of custom model cards sourced from the HuggingFace hub
Monitoring and Troubleshooting Procedures
- Interpretation of EXO logs and analysis of distributed tracing data
- Assessment of cluster health status within the dashboard cluster view
- Diagnosis of worker node failures and evaluation of reconnection behaviors
- Application of EXO_TRACING_ENABLED for performance bottleneck identification
Cluster Maintenance and Update Protocols
- Procedures for updating EXO binaries and rebuilding the dashboard
- Migration of model caches and management of pre-downloaded assets over NFS
- Graceful removal of nodes and reallocation of computational workloads
Requirements
- A solid understanding of networking fundamentals, including IP addressing, subnetting, and firewall configurations
- Proficiency in macOS or Linux command-line administration
- Familiarity with Python package management (pip/uv) and Node.js tooling
Target Audience
- System administrators
- DevOps engineers
- AI infrastructure architects responsible for on-premise LLM deployment
Runs with a minimum of 4 + people. For 1-to-1 or private group training, request a quote.
EXO: End-to-End Local AI Cluster Deployment Training Course - Booking
EXO: End-to-End Local AI Cluster Deployment Training Course - Enquiry
EXO: End-to-End Local AI Cluster Deployment - Consultancy Enquiry
Upcoming Courses
Related Courses
Advanced LangGraph: Optimization, Debugging, and Monitoring Complex Graphs
35 HoursLangGraph serves as a framework for developing stateful, multi-actor LLM applications structured as composable graphs, featuring persistent state management and precise control over execution flows.
This instructor-led, live training session (available online or on-site) is designed for advanced AI platform engineers, DevOps professionals specializing in AI, and ML architects seeking to optimize, troubleshoot, monitor, and manage production-grade LangGraph systems for government environments.
Upon completion of this training, participants will be equipped to:
- Design and refine complex LangGraph structures to enhance velocity, cost-efficiency, and scalability.
- Implement robust reliability standards through retry mechanisms, timeout controls, idempotency, and checkpoint-based recovery.
- Diagnose and trace graph executions, inspect internal states, and methodically replicate production anomalies.
- Instrument systems with logging, metrics, and tracing capabilities, deploy to production, and monitor SLAs and associated costs.
Training Delivery Model
- Interactive instructional sessions and facilitated discussions.
- Extensive exercises and practical application tasks.
- Real-time implementation within a live laboratory environment.
Customization Options
- Interested parties may contact the provider to coordinate a tailored training program aligned with specific organizational needs.
Building Coding Agents with Devstral: From Agent Design to Tooling
14 HoursDevstral is an open-source framework engineered for the development and deployment of coding agents. These agents interact with codebases, developer tools, and application programming interfaces to improve engineering efficiency and operational standards for government systems.
This instructor-led training, available online or on-site, targets intermediate to advanced machine learning engineers, developer tooling specialists, and site reliability engineers. The objective is to enable participants to design, implement, and optimize coding agents using Devstral within a formal operational context.
Upon completion of this training, participants will be capable of:
- Installing and configuring Devstral for the development of coding agents.
- Designing agentic workflows for codebase analysis and modification.
- Integrating coding agents with developer tools and external APIs.
- Applying best practices for the secure and efficient deployment of agents.
Instructional Methodology
- Interactive lectures and facilitated discussions.
- Extensive practical exercises and skill-building activities.
- Hands-on implementation within a live laboratory environment.
Customization Options for Government
- Contact our team to arrange a customized training session tailored to specific government requirements.
Open-Source Model Ops: Self-Hosting, Fine-Tuning and Governance with Devstral & Mistral Models
14 HoursDevstral and Mistral models represent open-source artificial intelligence technologies engineered for flexible deployment, precise fine-tuning, and scalable integration within organizational infrastructure.
This instructor-led training, available in online or onsite formats, targets intermediate to advanced machine learning engineers, platform teams, and research professionals. The curriculum focuses on self-hosting, fine-tuning, and governing Mistral and Devstral models for production environments, with specific relevance to secure and accountable operations for government.
Upon completion, participants will be equipped to:
- Configure and manage self-hosted environments for Mistral and Devstral models.
- Execute fine-tuning processes to enhance domain-specific performance.
- Deploy versioning, monitoring, and lifecycle governance frameworks.
- Maintain security, compliance, and responsible usage protocols for open-source models for government.
Course Format
- Interactive lectures facilitated by expert discussion.
- Practical exercises focused on self-hosting and fine-tuning techniques.
- Live-lab implementation of governance and monitoring pipelines.
Customization Availability
- Prospective learners may contact our office to arrange customized training sessions for this course.
Fiji: Image Processing for Biotechnology and Toxicology
14 HoursLangGraph Applications in Finance
35 HoursLangGraph serves as a framework for constructing stateful, multi-agent large language model (LLM) applications through composable graphs, featuring persistent state management and precise execution control.
This instructor-led, live training session (available online or on-site) is tailored for intermediate to advanced professionals seeking to design, implement, and manage LangGraph-based financial solutions with robust governance, observability, and compliance for government and regulated sectors.
Upon completion of this training, participants will be proficient in the following areas:
- Designing financial workflows within LangGraph that adhere to regulatory and audit mandates.
- Integrating financial data standards and ontologies into graph state management and operational tooling.
- Implementing reliability, safety, and human-in-the-loop controls for critical financial processes.
- Deploying, monitoring, and optimizing LangGraph systems to meet performance targets, cost efficiency, and service level agreements (SLAs).
Course Format
- Interactive lectures and structured discussions.
- Extensive practical exercises and skill-building activities.
- Hands-on implementation within a live laboratory environment.
Course Customization Options
- To request a customized training curriculum for this course, please contact the provider to coordinate arrangements.
LangGraph Foundations: Graph-Based LLM Prompting and Chaining
14 HoursLangGraph serves as a specialized framework for constructing graph-based LLM applications, enabling advanced features such as strategic planning, conditional branching, tool utilization, persistent memory, and controlled execution environments.
This instructor-led, live training program, available online or onsite, is designed for junior-level developers, prompt engineers, and data professionals seeking to engineer robust, multi-step LLM workflows utilizing LangGraph.
Upon completion of this program, participants will possess the capability to:
- Articulate fundamental LangGraph components (nodes, edges, state) and their applicable contexts.
- Construct prompt chains that support branching, tool invocation, and memory retention.
- Integrate retrieval mechanisms and external APIs into graph-based workflows for government use cases.
- Test, debug, and assess LangGraph applications to ensure reliability and safety standards.
Instructional Methodology
- Interactive lectures combined with facilitated group discussions.
- Guided laboratory sessions and code analysis within a secure sandbox environment.
- Scenario-driven exercises focused on architectural design, testing, and evaluation.
Customization Capabilities
- To arrange a customized training program for government agencies, please contact our team for specific requirements.
LangGraph in Healthcare: Workflow Orchestration for Regulated Environments
35 HoursLangGraph facilitates stateful, multi-actor workflows leveraging Large Language Models (LLMs) with precise control over execution paths and state persistence. Within the public health sector, these capabilities are essential for ensuring regulatory compliance, achieving interoperability, and developing decision-support systems aligned with government medical operations.
This instructor-led, live training (available online or onsite) is designed for intermediate to advanced professionals seeking to design, implement, and manage LangGraph-based healthcare solutions for government, while addressing regulatory, ethical, and operational complexities.
Upon completion of this training, participants will be able to:
- Design healthcare-specific LangGraph workflows with a focus on regulatory compliance and auditability.
- Integrate LangGraph applications with medical ontologies and standards (FHIR, SNOMED CT, ICD) required for government data exchange.
- Apply best practices for reliability, traceability, and explainability in sensitive public sector environments.
- Deploy, monitor, and validate LangGraph applications in healthcare production settings for government use.
Format of the Course
- Interactive lectures and facilitated discussions.
- Hands-on exercises utilizing real-world public sector case studies.
- Implementation practice in a live laboratory environment.
Course Customization Options
- To request a customized training program tailored to specific government agency needs, please contact us to arrange details.
LangGraph for Legal Applications
35 HoursBuilding Dynamic Workflows with LangGraph and LLM Agents
14 HoursLangGraph for Marketing Automation
14 HoursLangGraph serves as a graph-based orchestration framework that facilitates conditional, multi-step workflows involving large language models (LLMs) and external tools. This infrastructure is well-suited for the automation and customization of content pipelines for government operations.
This instructor-led, live training session, available in online or onsite formats, targets intermediate-level professionals, including content strategists and automation developers. The objective is to implement dynamic, branching email campaigns and automated content generation systems using LangGraph.
Upon completion of this training, participants will possess the ability to:
- Design graph-structured content and email workflows incorporating conditional logic.
- Integrate LLMs, application programming interfaces (APIs), and external data sources to drive automated personalization.
- Manage state, memory, and contextual data across multi-step campaign executions.
- Evaluate, monitor, and optimize workflow performance and final delivery outcomes.
Course Structure and Delivery Method
- Interactive instructional sessions and facilitated group discussions.
- Practical laboratory exercises focused on implementing email workflows and content pipelines.
- Scenario-based assignments emphasizing personalization, audience segmentation, and branching logic.
Course Customization and Tailored Solutions
- For inquiries regarding customized training solutions tailored to specific operational needs for government, please contact the appropriate administrative office.
Le Chat Enterprise: Private ChatOps, Integrations & Admin Controls
14 HoursLe Chat Enterprise serves as a private ChatOps solution that delivers secure, customizable, and governed conversational AI capabilities for organizations, specifically designed for government operations, with robust support for RBAC, SSO, connectors, and enterprise application integrations.
This instructor-led, live training program (available online or onsite) is designed for intermediate-level product managers, IT leads, solution engineers, and security/compliance teams seeking to deploy, configure, and govern Le Chat Enterprise within enterprise and government environments.
Upon completion of this training, participants will be equipped to:
- Establish and configure Le Chat Enterprise for secure government deployments.
- Activate RBAC, SSO, and compliance-driven control mechanisms.
- Integrate Le Chat with enterprise applications and secure data stores for government use.
- Develop and implement governance and administrative playbooks for ChatOps operations.
Course Delivery Format
- Interactive lectures and structured discussions.
- Extensive exercises and practical application tasks.
- Hands-on implementation within a live laboratory environment.
Course Customization Options for Government
- To request a tailored training program for this course, please contact us to arrange specific requirements.
Cost-Effective LLM Architectures: Mistral at Scale (Performance / Cost Engineering)
14 HoursProductizing Conversational Assistants with Mistral Connectors & Integrations
14 HoursMistral AI serves as an open platform that empowers organizations to develop and embed conversational assistants within both internal operations and external service delivery channels for government and enterprise environments.
This instructor-led, live training session—available in online or on-site formats—is designed for product managers, full-stack developers, and integration engineers at beginner to intermediate levels. The curriculum focuses on the strategic design, integration, and operationalization of conversational assistants utilizing Mistral’s connector and integration frameworks.
Upon completion of this training, participants will possess the capability to:
- Integrate Mistral conversational models with enterprise-grade and SaaS infrastructure components.
- Implement retrieval-augmented generation (RAG) to ensure grounded, fact-based responses.
- Architect user experience patterns for both internal operational tools and external service points.
- Deploy assistants into functional workflows to address specific operational requirements.
Training Delivery Format
- Interactive instructional sessions and facilitated discussions.
- Practical exercises focused on system integration.
- Live laboratory sessions for the development of conversational prototypes.
Customization Availability
- Organizations may request tailored training modules for government or specific operational needs by contacting the provider for scheduling.
Enterprise-Grade Deployments with Mistral Medium 3
14 HoursMistral Medium 3 is a high-performance, multimodal large language model engineered for production-grade implementation within enterprise and government sectors.
This instructor-led, live training session, available online or on-site, is designed for intermediate to advanced AI/ML engineers, platform architects, and MLOps teams seeking to deploy, optimize, and secure Mistral Medium 3 for critical enterprise and public sector applications.
Upon completion of this training, participants will be equipped to:
- Deploy Mistral Medium 3 through API integrations and self-hosted infrastructure options.
- Enhance inference performance while managing operational costs effectively.
- Execute multimodal use cases leveraging Mistral Medium 3 capabilities.
- Enforce security and compliance standards aligned with enterprise and governmental requirements.
Course Delivery Format
- Interactive instructional sessions and facilitated discussions.
- Extensive practical exercises and proficiency drills.
- Hands-on implementation within a controlled live-lab environment.
Customization Options for Government Needs
- To request a customized training curriculum tailored for government requirements, please contact the program office to coordinate scheduling and scope.
Mistral for Responsible AI: Privacy, Data Residency & Enterprise Controls
14 HoursMistral AI operates as an open, enterprise-oriented platform designed to facilitate the secure, compliant, and responsible deployment of artificial intelligence solutions.
This instructor-led, live instructional session, available in either virtual or on-site formats, is tailored for mid-level compliance officers, security engineers, and legal/operational stakeholders. The objective is to implement responsible AI practices using Mistral by applying privacy safeguards, data residency controls, and enterprise-level governance mechanisms relevant for government and public sector operations.
Upon completion of this training, participants will demonstrate the ability to:
- Deploy privacy-preserving methodologies within Mistral environments.
- Execute data residency strategies to satisfy statutory and regulatory obligations.
- Configure enterprise-grade security controls, including RBAC, SSO, and audit logging.
- Assess vendor capabilities and deployment models for regulatory alignment.
Instructional Methodology
- Facilitated instruction through lectures and collaborative discussion.
- Analysis of compliance-centric scenarios and practical exercises.
- Practical application of enterprise AI control frameworks.
Program Customization
- Interested parties may contact the organization to discuss tailored training options for this course.