Get in Touch

Course Outline

Preparation of Machine Learning Models for Operational Deployment

  • Encapsulation of models using Docker
  • Model export from TensorFlow and PyTorch frameworks
  • Protocols for versioning and storage management

Serving Models on Kubernetes Infrastructure

  • Overview of inference server capabilities
  • Implementation of TensorFlow Serving and TorchServe
  • Configuration of model endpoints

Techniques for Inference Optimization

  • Strategies for batch processing
  • Management of concurrent request handling
  • Tuning for latency and throughput performance

Autoscaling Mechanisms for ML Workloads

  • Horizontal Pod Autoscaler (HPA) configuration
  • Vertical Pod Autoscaler (VPA) usage
  • Kubernetes Event-Driven Autoscaling (KEDA) integration

GPU Provisioning and Resource Governance

  • Configuration of GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Model Rollout and Release Methodologies

  • Blue/green deployment strategies
  • Canary rollout patterns
  • A/B testing for model performance evaluation

Monitoring and Observability for Production ML Systems

  • Key metrics for inference workloads
  • Best practices for logging and tracing
  • Dashboard creation and alerting configuration

Security and Reliability Frameworks

  • Securing model endpoints for government use
  • Implementation of network policies and access controls
  • Ensuring high availability and resilience

Summary and Future Initiatives

Requirements

  • Knowledge of containerized application workflows
  • Practical experience with Python-based machine learning models
  • Familiarity with core Kubernetes concepts

Target Audience

  • ML engineers
  • DevOps engineers
  • Platform engineering teams
 14 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories